今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio  Source Separation

标题:基于多任务音频源分离的语音和歌声联合识别
链接:https://arxiv.org/abs/2404.11275
作者:Ye Bai,Chenxing Li,Hao Li,Yuanyuan Zhao,Xiaorui Wang
备注:Accepted by ICME 2024
摘要:在短视频和直播中,语音、歌声和背景音乐往往相互重叠、相互模糊。这种复杂性在构造和识别音频内容方面造成困难,这可能会损害后续的ASR和音乐理解应用。本文提出了一种基于多任务音频源分离(MTASS)的语音识别模型(JRSV),该模型能够联合识别语音和歌声。具体地,MTASS模块将混合音频分离成不同的语音和歌声音轨,同时去除背景音乐。CTC/注意力混合识别模块识别两个轨道。为了进一步提高识别的鲁棒性,提出了在线蒸馏的方法。为了评估所提出的方法,基准数据集的构建和发布。实验结果表明,JRSV可以显着提高识别精度的混合音频的每个轨道。
摘要:In short video and live broadcasts, speech, singing voice, and background music often overlap and obscure each other. This complexity creates difficulties in structuring and recognizing the audio content, which may impair subsequent ASR and music understanding applications. This paper proposes a multi-task audio source separation (MTASS) based ASR model called JRSV, which Jointly Recognizes Speech and singing Voices. Specifically, the MTASS module separates the mixed audio into distinct speech and singing voice tracks while removing background music. The CTC/attention hybrid recognition module recognizes both tracks. Online distillation is proposed to improve the robustness of recognition further. To evaluate the proposed methods, a benchmark dataset is constructed and released. Experimental results demonstrate that JRSV can significantly improve recognition accuracy on each track of the mixed audio.


【2】 Music Enhancement with Deep Filters: A Technical Report for The ICASSP  2024 Cadenza Challenge
标题:利用深度滤镜增强音乐:ICASP 2024年华彩挑战赛技术报告
链接:https://arxiv.org/abs/2404.11116
作者:Keren Shao,Ke Chen,Shlomo Dubnov
备注:2 pages, 2 figures, 1 tables, Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2024
摘要:在这个挑战中,我们将深度过滤器从原始的DeepfilterNet中分离出来,并将它们合并到我们基于Spec-UNet的网络中,以进一步改进基于混合Demucs(hdemucs)的混音管道。使用深度滤波器组件背后的动机在于其更好地处理时间精细结构的潜力。我们展示了一个渐进的改善信号失真比(SDR)和助听器音频质量指数(HAAQI)指标时,比较hdemucs的性能对我们的模型的不同版本。
摘要:In this challenge, we disentangle the deep filters from the original DeepfilterNet and incorporate them into our Spec-UNet-based network to further improve a hybrid Demucs (hdemucs) based remixing pipeline. The motivation behind the use of the deep filter component lies at its potential in better handling temporal fine structures. We demonstrate an incremental improvement in both the Signal-to-Distortion Ratio (SDR) and the Hearing Aid Audio Quality Index (HAAQI) metrics when comparing the performance of hdemucs against different versions of our model.

【3】 FairSSD: Understanding Bias in Synthetic Speech Detectors
标题:FairSSD:了解合成语音检测器中的偏差
链接:https://arxiv.org/abs/2404.10989
作者:Amit Kumar Singh Yadav,Kratika Bhagtani,Davide Salvi,Paolo Bestagini,Edward J. Delp
备注:Accepted at CVPR 2024 (WMF)
摘要:可以生成与人类说话者记录的语音在感知上不可区分的合成语音的方法是容易获得的。有几起事件报告了滥用这些方法生成的合成语音进行欺诈的情况。为了对抗这种误用,已经提出了许多方法来检测合成语音。这些检测器中的一些更具可解释性,可以概括为检测野生合成语音,并且对噪声具有鲁棒性。然而,有限的工作已经做了了解这些探测器的偏见。在这项工作中,我们研究了现有的合成语音检测器的偏见,以确定他们是否会不公平地针对特定的性别,年龄和口音组。我们还检查这些检测器是否会有一个更高的错误分类率为真正的语音从语音障碍的扬声器w.r.t流利的扬声器。使用超过90万个语音信号的6个现有的合成语音检测器的广泛实验表明,大多数检测器的性别,年龄和口音偏见,和未来的工作是必要的,以确保公平。为了支持未来的研究,我们在https://gitlab.com/viper-purdue/fairssd上发布了我们的评估数据集、研究中使用的模型和源代码。
摘要:Methods that can generate synthetic speech which is perceptually indistinguishable from speech recorded by a human speaker, are easily available. Several incidents report misuse of synthetic speech generated from these methods to commit fraud. To counter such misuse, many methods have been proposed to detect synthetic speech. Some of these detectors are more interpretable, can generalize to detect synthetic speech in the wild and are robust to noise. However, limited work has been done on understanding bias in these detectors. In this work, we examine bias in existing synthetic speech detectors to determine if they will unfairly target a particular gender, age and accent group. We also inspect whether these detectors will have a higher misclassification rate for bona fide speech from speech-impaired speakers w.r.t fluent speakers. Extensive experiments on 6 existing synthetic speech detectors using more than 0.9 million speech signals demonstrate that most detectors are gender, age and accent biased, and future work is needed to ensure fairness. To support future research, we release our evaluation dataset, models used in our study and source code at https://gitlab.com/viper-purdue/fairssd.

【4】 Teaching a Multilingual Large Language Model to Understand Multilingual  Speech via Multi-Instructional Training
标题:通过多教学训练教授多语言大语言模型以理解多语言语音
链接:https://arxiv.org/abs/2404.10922
作者:Pavel Denisov,Ngoc Thang Vu
备注:NAACL Findings 2024
摘要:语言建模的最新进展导致了能够执行各种自然语言处理任务的大型语言模型(LLM)的出现。尽管他们在基于文本的任务中取得了成功,但将LLM应用于语音领域仍然有限且具有挑战性。本文介绍了BLOOMZMMS,一种新的模型,集成了多语言LLM与多语言语音编码器,旨在利用LLM的语音识别和超越的能力。利用多教学训练方法,我们证明了语言知识从文本到语音模态的可转移性。我们对来自139种语言的1900小时转录数据进行的实验表明,多语言语音表示可以有效地学习并与多语言LLM保持一致。虽然这种学习表示最初显示出任务泛化的局限性,我们通过在多教学风格中生成合成目标来解决这个问题。我们的zero-shot评估结果证实了我们的方法在多个任务中的鲁棒性,包括语音翻译和多语言口语理解,从而为在语音领域应用LLM开辟了新的途径。
摘要:Recent advancements in language modeling have led to the emergence of Large Language Models (LLMs) capable of various natural language processing tasks. Despite their success in text-based tasks, applying LLMs to the speech domain remains limited and challenging. This paper presents BLOOMZMMS, a novel model that integrates a multilingual LLM with a multilingual speech encoder, aiming to harness the capabilities of LLMs for speech recognition and beyond. Utilizing a multi-instructional training approach, we demonstrate the transferability of linguistic knowledge from the text to the speech modality. Our experiments, conducted on 1900 hours of transcribed data from 139 languages, establish that a multilingual speech representation can be effectively learned and aligned with a multilingual LLM. While this learned representation initially shows limitations in task generalization, we address this issue by generating synthetic targets in a multi-instructional style. Our zero-shot evaluation results confirm the robustness of our approach across multiple tasks, including speech translation and multilingual spoken language understanding, thereby opening new avenues for applying LLMs in the speech domain.

【5】 Unsupervised Speaker Diarization in Distributed IoT Networks Using  Federated Learning
标题:使用联邦学习在分布式物联网网络中进行无监督发言人拨号
链接:https://arxiv.org/abs/2404.10842
作者:Amit Kumar Bhuyan,Hrishikesh Dutta,Subir Biswas
备注:11 pages, 7 figures, 1 table
摘要:本文提出了一种用于联网IoT风格音频设备的计算高效的分布式扬声器日志化框架。这项工作提出了一个联邦学习模型,它可以识别在一个会话中的参与者,而不需要一个大型的音频数据库进行训练。针对基于余弦相似度的联邦学习模型,提出了一种无监督的在线更新机制。此外,所提出的日记化系统解决了说话人变化检测的问题,通过。无监督分割技术使用霍特林的t平方统计和贝叶斯信息准则。在这种新的方法中,说话人变化检测偏置周围检测到的准沉默,这降低了漏检率和误检率之间的权衡的严重性。此外,由于逐帧识别扬声器而导致的计算开销通过。语音片段的无监督聚类。实验结果表明,在非IID语音数据的存在下,所提出的训练方法是有效的。它还显示了相当大的改进,在分割阶段的错误和遗漏检测的减少,同时减少了计算开销。提高的准确性和降低的计算成本使该机制适用于分布式物联网音频网络中的实时扬声器日志。
摘要:This paper presents a computationally efficient and distributed speaker diarization framework for networked IoT-style audio devices. The work proposes a Federated Learning model which can identify the participants in a conversation without the requirement of a large audio database for training. An unsupervised online update mechanism is proposed for the Federated Learning model which depends on cosine similarity of speaker embeddings. Moreover, the proposed diarization system solves the problem of speaker change detection via. unsupervised segmentation techniques using Hotelling's t-squared Statistic and Bayesian Information Criterion. In this new approach, speaker change detection is biased around detected quasi-silences, which reduces the severity of the trade-off between the missed detection and false detection rates. Additionally, the computational overhead due to frame-by-frame identification of speakers is reduced via. unsupervised clustering of speech segments. The results demonstrate the effectiveness of the proposed training method in the presence of non-IID speech data. It also shows a considerable improvement in the reduction of false and missed detection at the segmentation stage, while reducing the computational overhead. Improved accuracy and reduced computational cost makes the mechanism suitable for real-time speaker diarization across a distributed IoT audio network.

【6】 In situ sound absorption estimation with the discrete complex image  source method
标题:离散复图像源法现场声吸收估计
链接:https://arxiv.org/abs/2404.11399
作者:Eric Brandao,William Fonseca,Paulo Mareze,Carlos Resende,Gabriel Azzuz,Joao Pontalti,Efren Fernandez-Grande
备注:37 pages, 12 figures, original manuscript to be submitted to the Journal of Sound and Vibration
摘要:估计现场的声音吸收依赖于准确地描述测量的声场。有证据表明,模拟撞击球面波的反射是很重要的,特别是对于紧凑的测量系统。本文提出了一种通过将麦克风阵列测量的声压映射到复平面中沿直线的单极子分布来估计材料样品的吸声系数的方法。所提出的方法相比,作为两个源(一个图像源和一个图像源)的叠加建模的声场。用Tikhonov正则化方法求解反问题,并根据L曲线准则自动选择正则化参数。通过模拟无限和有限多孔吸收体上方的声场来测试吸声测量。比较了平面波吸收系数和球面波入射时的吸收系数。两个多孔样品和一个共振吸收体的实验分析也进行了现场。四个阵列进行了测试,增加孔径和传感器的数量。结果表明,即使在只有几个麦克风的阵列中,测量也是可行的。积分方程的离散化导致在样品的表面处的声压和粒子速度的更准确的重建。所得的吸收系数与球面波入射时获得的吸收系数一致,表明沿复线包括更多的单极子是声场的基本特征。
摘要:Estimating the sound absorption in situ relies on accurately describing the measured sound field. Evidence suggests that modeling the reflection of impinging spherical waves is important, especially for compact measurement systems. This article proposes a method for estimating the sound absorption coefficient of a material sample by mapping the sound pressure, measured by a microphone array, to a distribution of monopoles along a line in the complex plane. The proposed method is compared to modeling the sound field as a superposition of two sources (a monopole and an image source). The obtained inverse problems are solved with Tikhonov regularization, with automatic choice of the regularization parameter by the L-curve criterion. The sound absorption measurement is tested with simulations of the sound field above infinite and finite porous absorbers. The approaches are compared to the plane-wave absorption coefficient and the one obtained by spherical wave incidence. Experimental analysis of two porous samples and one resonant absorber is also carried out in situ. Four arrays were tested with an increasing aperture and number of sensors. It was demonstrated that measurements are feasible even with an array with only a few microphones. The discretization of the integral equation led to a more accurate reconstruction of the sound pressure and particle velocity at the sample's surface. The resulting absorption coefficient agrees with the one obtained for spherical wave incidence, indicating that including more monopoles along the complex line is an essential feature of the sound field.

eess.AS音频处理
【1】 In situ sound absorption estimation with the discrete complex image  source method
标题:离散复图像源法现场声吸收估计
链接:https://arxiv.org/abs/2404.11399
作者:Eric Brandao,William Fonseca,Paulo Mareze,Carlos Resende,Gabriel Azzuz,Joao Pontalti,Efren Fernandez-Grande
备注:37 pages, 12 figures, original manuscript to be submitted to the Journal of Sound and Vibration
摘要:估计现场的声音吸收依赖于准确地描述测量的声场。有证据表明,模拟撞击球面波的反射是很重要的,特别是对于紧凑的测量系统。本文提出了一种通过将麦克风阵列测量的声压映射到复平面中沿直线的单极子分布来估计材料样品的吸声系数的方法。所提出的方法相比,作为两个源(一个图像源和一个图像源)的叠加建模的声场。用Tikhonov正则化方法求解反问题,并根据L曲线准则自动选择正则化参数。通过模拟无限和有限多孔吸收体上方的声场来测试吸声测量。比较了平面波吸收系数和球面波入射时的吸收系数。两个多孔样品和一个共振吸收体的实验分析也进行了现场。四个阵列进行了测试,增加孔径和传感器的数量。结果表明,即使在只有几个麦克风的阵列中,测量也是可行的。积分方程的离散化导致在样品的表面处的声压和粒子速度的更准确的重建。所得的吸收系数与球面波入射时获得的吸收系数一致,表明沿复线包括更多的单极子是声场的基本特征。
摘要:Estimating the sound absorption in situ relies on accurately describing the measured sound field. Evidence suggests that modeling the reflection of impinging spherical waves is important, especially for compact measurement systems. This article proposes a method for estimating the sound absorption coefficient of a material sample by mapping the sound pressure, measured by a microphone array, to a distribution of monopoles along a line in the complex plane. The proposed method is compared to modeling the sound field as a superposition of two sources (a monopole and an image source). The obtained inverse problems are solved with Tikhonov regularization, with automatic choice of the regularization parameter by the L-curve criterion. The sound absorption measurement is tested with simulations of the sound field above infinite and finite porous absorbers. The approaches are compared to the plane-wave absorption coefficient and the one obtained by spherical wave incidence. Experimental analysis of two porous samples and one resonant absorber is also carried out in situ. Four arrays were tested with an increasing aperture and number of sensors. It was demonstrated that measurements are feasible even with an array with only a few microphones. The discretization of the integral equation led to a more accurate reconstruction of the sound pressure and particle velocity at the sample's surface. The resulting absorption coefficient agrees with the one obtained for spherical wave incidence, indicating that including more monopoles along the complex line is an essential feature of the sound field.

【2】 Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio  Source Separation
标题:基于多任务音频源分离的语音和歌声联合识别
链接:https://arxiv.org/abs/2404.11275
作者:Ye Bai,Chenxing Li,Hao Li,Yuanyuan Zhao,Xiaorui Wang
备注:Accepted by ICME 2024
摘要:在短视频和直播中,语音、歌声和背景音乐往往相互重叠、相互模糊。这种复杂性在构造和识别音频内容方面造成困难,这可能会损害后续的ASR和音乐理解应用。本文提出了一种基于多任务音频源分离(MTASS)的语音识别模型(JRSV),该模型能够联合识别语音和歌声。具体地,MTASS模块将混合音频分离成不同的语音和歌声音轨,同时去除背景音乐。CTC/注意力混合识别模块识别两个轨道。为了进一步提高识别的鲁棒性,提出了在线蒸馏的方法。为了评估所提出的方法,基准数据集的构建和发布。实验结果表明,JRSV可以显着提高识别精度的混合音频的每个轨道。
摘要:In short video and live broadcasts, speech, singing voice, and background music often overlap and obscure each other. This complexity creates difficulties in structuring and recognizing the audio content, which may impair subsequent ASR and music understanding applications. This paper proposes a multi-task audio source separation (MTASS) based ASR model called JRSV, which Jointly Recognizes Speech and singing Voices. Specifically, the MTASS module separates the mixed audio into distinct speech and singing voice tracks while removing background music. The CTC/attention hybrid recognition module recognizes both tracks. Online distillation is proposed to improve the robustness of recognition further. To evaluate the proposed methods, a benchmark dataset is constructed and released. Experimental results demonstrate that JRSV can significantly improve recognition accuracy on each track of the mixed audio.

【3】 Music Enhancement with Deep Filters: A Technical Report for The ICASSP  2024 Cadenza Challenge
标题:利用深度滤镜增强音乐:ICASP 2024年华彩挑战赛技术报告
链接:https://arxiv.org/abs/2404.11116
作者:Keren Shao,Ke Chen,Shlomo Dubnov
备注:2 pages, 2 figures, 1 tables, Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2024
摘要:在这个挑战中,我们将深度过滤器从原始的DeepfilterNet中分离出来,并将它们合并到我们基于Spec-UNet的网络中,以进一步改进基于混合Demucs(hdemucs)的混音管道。使用深度滤波器组件背后的动机在于其更好地处理时间精细结构的潜力。我们展示了一个渐进的改善信号失真比(SDR)和助听器音频质量指数(HAAQI)指标时,比较hdemucs的性能对我们的模型的不同版本。
摘要:In this challenge, we disentangle the deep filters from the original DeepfilterNet and incorporate them into our Spec-UNet-based network to further improve a hybrid Demucs (hdemucs) based remixing pipeline. The motivation behind the use of the deep filter component lies at its potential in better handling temporal fine structures. We demonstrate an incremental improvement in both the Signal-to-Distortion Ratio (SDR) and the Hearing Aid Audio Quality Index (HAAQI) metrics when comparing the performance of hdemucs against different versions of our model.


【4】 FairSSD: Understanding Bias in Synthetic Speech Detectors
标题:FairSSD:了解合成语音检测器中的偏差
链接:https://arxiv.org/abs/2404.10989
作者:Amit Kumar Singh Yadav,Kratika Bhagtani,Davide Salvi,Paolo Bestagini,Edward J. Delp
备注:Accepted at CVPR 2024 (WMF)
摘要:可以生成与人类说话者记录的语音在感知上不可区分的合成语音的方法是容易获得的。有几起事件报告了滥用这些方法生成的合成语音进行欺诈的情况。为了对抗这种误用,已经提出了许多方法来检测合成语音。这些检测器中的一些更具可解释性,可以概括为检测野生合成语音,并且对噪声具有鲁棒性。然而,有限的工作已经做了了解这些探测器的偏见。在这项工作中,我们研究了现有的合成语音检测器的偏见,以确定他们是否会不公平地针对特定的性别,年龄和口音组。我们还检查这些检测器是否会有一个更高的错误分类率为真正的语音从语音障碍的扬声器w.r.t流利的扬声器。使用超过90万个语音信号的6个现有的合成语音检测器的广泛实验表明,大多数检测器的性别,年龄和口音偏见,和未来的工作是必要的,以确保公平。为了支持未来的研究,我们在https://gitlab.com/viper-purdue/fairssd上发布了我们的评估数据集、研究中使用的模型和源代码。
摘要:Methods that can generate synthetic speech which is perceptually indistinguishable from speech recorded by a human speaker, are easily available. Several incidents report misuse of synthetic speech generated from these methods to commit fraud. To counter such misuse, many methods have been proposed to detect synthetic speech. Some of these detectors are more interpretable, can generalize to detect synthetic speech in the wild and are robust to noise. However, limited work has been done on understanding bias in these detectors. In this work, we examine bias in existing synthetic speech detectors to determine if they will unfairly target a particular gender, age and accent group. We also inspect whether these detectors will have a higher misclassification rate for bona fide speech from speech-impaired speakers w.r.t fluent speakers. Extensive experiments on 6 existing synthetic speech detectors using more than 0.9 million speech signals demonstrate that most detectors are gender, age and accent biased, and future work is needed to ensure fairness. To support future research, we release our evaluation dataset, models used in our study and source code at https://gitlab.com/viper-purdue/fairssd.


【5】 Teaching a Multilingual Large Language Model to Understand Multilingual  Speech via Multi-Instructional Training
标题:通过多教学训练教授多语言大语言模型以理解多语言语音
链接:https://arxiv.org/abs/2404.10922
作者:Pavel Denisov,Ngoc Thang Vu
备注:NAACL Findings 2024
摘要:语言建模的最新进展导致了能够执行各种自然语言处理任务的大型语言模型(LLM)的出现。尽管他们在基于文本的任务中取得了成功,但将LLM应用于语音领域仍然有限且具有挑战性。本文介绍了BLOOMZMMS,一种新的模型,集成了多语言LLM与多语言语音编码器,旨在利用LLM的语音识别和超越的能力。利用多教学训练方法,我们证明了语言知识从文本到语音模态的可转移性。我们对来自139种语言的1900小时转录数据进行的实验表明,多语言语音表示可以有效地学习并与多语言LLM保持一致。虽然这种学习表示最初显示出任务泛化的局限性,我们通过在多教学风格中生成合成目标来解决这个问题。我们的zero-shot评估结果证实了我们的方法在多个任务中的鲁棒性,包括语音翻译和多语言口语理解,从而为在语音领域应用LLM开辟了新的途径。
摘要:Recent advancements in language modeling have led to the emergence of Large Language Models (LLMs) capable of various natural language processing tasks. Despite their success in text-based tasks, applying LLMs to the speech domain remains limited and challenging. This paper presents BLOOMZMMS, a novel model that integrates a multilingual LLM with a multilingual speech encoder, aiming to harness the capabilities of LLMs for speech recognition and beyond. Utilizing a multi-instructional training approach, we demonstrate the transferability of linguistic knowledge from the text to the speech modality. Our experiments, conducted on 1900 hours of transcribed data from 139 languages, establish that a multilingual speech representation can be effectively learned and aligned with a multilingual LLM. While this learned representation initially shows limitations in task generalization, we address this issue by generating synthetic targets in a multi-instructional style. Our zero-shot evaluation results confirm the robustness of our approach across multiple tasks, including speech translation and multilingual spoken language understanding, thereby opening new avenues for applying LLMs in the speech domain.


【6】 Unsupervised Speaker Diarization in Distributed IoT Networks Using  Federated Learning
标题:使用联邦学习在分布式物联网网络中进行无监督发言人拨号
链接:https://arxiv.org/abs/2404.10842
作者:Amit Kumar Bhuyan,Hrishikesh Dutta,Subir Biswas
备注:11 pages, 7 figures, 1 table
摘要:本文提出了一种用于联网IoT风格音频设备的计算高效的分布式扬声器日志化框架。这项工作提出了一个联邦学习模型,它可以识别在一个会话中的参与者,而不需要一个大型的音频数据库进行训练。针对基于余弦相似度的联邦学习模型,提出了一种无监督的在线更新机制。此外,所提出的日记化系统解决了说话人变化检测的问题,通过。无监督分割技术使用霍特林的t平方统计和贝叶斯信息准则。在这种新的方法中,说话人变化检测偏置周围检测到的准沉默,这降低了漏检率和误检率之间的权衡的严重性。此外,由于逐帧识别扬声器而导致的计算开销通过。语音片段的无监督聚类。实验结果表明,在非IID语音数据的存在下,所提出的训练方法是有效的。它还显示了相当大的改进,在分割阶段的错误和遗漏检测的减少,同时减少了计算开销。提高的准确性和降低的计算成本使该机制适用于分布式物联网音频网络中的实时扬声器日志。
摘要:This paper presents a computationally efficient and distributed speaker diarization framework for networked IoT-style audio devices. The work proposes a Federated Learning model which can identify the participants in a conversation without the requirement of a large audio database for training. An unsupervised online update mechanism is proposed for the Federated Learning model which depends on cosine similarity of speaker embeddings. Moreover, the proposed diarization system solves the problem of speaker change detection via. unsupervised segmentation techniques using Hotelling's t-squared Statistic and Bayesian Information Criterion. In this new approach, speaker change detection is biased around detected quasi-silences, which reduces the severity of the trade-off between the missed detection and false detection rates. Additionally, the computational overhead due to frame-by-frame identification of speakers is reduced via. unsupervised clustering of speech segments. The results demonstrate the effectiveness of the proposed training method in the presence of non-IID speech data. It also shows a considerable improvement in the reduction of false and missed detection at the segmentation stage, while reducing the computational overhead. Improved accuracy and reduced computational cost makes the mechanism suitable for real-time speaker diarization across a distributed IoT audio network.

机器翻译由腾讯交互翻译提供,仅供参考