微信公众号:arXiv_Daily
cs.SD语音
标题: 重新想象舞蹈:舞者与人工智能之间的实时音乐共创
链接:https://arxiv.org/abs/2506.12008
备注:Accepted for publication at ICCC 2025 (International Conference on Computational Creativity)
摘要:舞蹈表演传统上遵循一种单向的关系,即运动对音乐的反应。虽然人工智能在各个创意领域都取得了进展,但它在舞蹈中的应用主要集中在从音乐输入中生成编舞。我们提出了一个系统,使舞者通过他们的动作动态地塑造音乐环境。我们的多模态架构通过智能地结合预先录制的音乐片段来响应舞蹈动作,从而创建连贯的音乐作品,建立了一种双向的创造性合作伙伴关系,舞者既是表演者又是作曲家。通过相关性分析的性能数据,我们展示了新兴的通信模式之间的运动质量和音频功能。这种方法将人工智能在表演艺术中的作用重新定义为一个响应式的合作者,为更广泛的人群提供专业舞蹈表演和即兴艺术表达的可能性。
摘要:Dance performance traditionally follows a unidirectional relationship where movement responds to music. While AI has advanced in various creative domains, its application in dance has primarily focused on generating choreography from musical input. We present a system that enables dancers to dynamically shape musical environments through their movements. Our multi-modal architecture creates a coherent musical composition by intelligently combining pre-recorded musical clips in response to dance movements, establishing a bidirectional creative partnership where dancers function as both performers and composers. Through correlation analysis of performance data, we demonstrate emergent communication patterns between movement qualities and audio features. This approach reconceptualizes the role of AI in performing arts as a responsive collaborator that expands possibilities for both professional dance performance and improvisational artistic expression across broader populations.
标题: 基于信心的肌电信号到语音自我训练:利用合成肌电信号进行稳健建模
链接:https://arxiv.org/abs/2506.11862
摘要:语音肌电图(EMG)到语音(V-ETS)模型从肌肉活动信号重建语音,促进神经喉诊断等应用。尽管V-ETS具有潜力,但由于缺乏成对的EMG-语音数据,V-ETS的发展受到阻碍。为了解决这个问题,我们提出了一种新的基于置信度的多说话者自训练(CoM 2S)方法,以及一个新策划的Libri-EMG数据集。该方法利用由预训练模型生成的合成EMG数据,然后基于音素级置信度提出过滤机制,通过提出的自我训练技术来增强ETS模型。实验结果表明,该方法提高了音素准确率,减少了语音混乱,降低了单词错误率,证实了我们的CoM 2S方法的V-ETS的有效性。为了支持未来的研究,我们将发布代码和建议Libri-EMG数据库-一个开放访问的,时间对齐的,多扬声器语音EMG和语音记录。
摘要:Voiced Electromyography (EMG)-to-Speech (V-ETS) models reconstruct speech from muscle activity signals, facilitating applications such as neurolaryngologic diagnostics. Despite its potential, the advancement of V-ETS is hindered by a scarcity of paired EMG-speech data. To address this, we propose a novel Confidence-based Multi-Speaker Self-training (CoM2S) approach, along with a newly curated Libri-EMG dataset. This approach leverages synthetic EMG data generated by a pre-trained model, followed by a proposed filtering mechanism based on phoneme-level confidence to enhance the ETS model through the proposed self-training techniques. Experiments demonstrate our method improves phoneme accuracy, reduces phonological confusion, and lowers word error rate, confirming the effectiveness of our CoM2S approach for V-ETS. In support of future research, we will release the codes and the proposed Libri-EMG dataset-an open-access, time-aligned, multi-speaker voiced EMG and speech recordings.
标题: 无条件倒置模型的抽象声音融合
链接:https://arxiv.org/abs/2506.11811
摘要:抽象声音被定义为不向收听者公开可识别的真实世界声音事件的声音。声音融合的目的是合成一个原始声音和一个参考声音,以产生一个新的声音,表现出的听觉特征,而不仅仅是添加剂叠加的声音成分。为了实现这种融合,我们采用反演技术,保留原始样品的基本特征,同时使可控合成。我们提出了基于DPMSolver++采样器的新的非线性和常微分方程反演模型,该采样器通过将模型输出配置为常数来反转采样过程,消除了噪声预测项引起的循环依赖性。我们的反演方法不需要及时调节,同时在采样过程中保持灵活的指导。
摘要:An abstract sound is defined as a sound that does not disclose identifiable real-world sound events to a listener. Sound fusion aims to synthesize an original sound and a reference sound to generate a novel sound that exhibits auditory features beyond mere additive superposition of the sound constituents. To achieve this fusion, we employ inversion techniques that preserve essential features of the original sample while enabling controllable synthesis. We propose novel SDE and ODE inversion models based on DPMSolver++ samplers that reverse the sampling process by configuring model outputs as constants, eliminating circular dependencies incurred by noise prediction terms. Our inversion approach requires no prompt conditioning while maintaining flexible guidance during sampling.
标题: 实现从现实世界环境中自动转录以儿童为中心的录音
链接:https://arxiv.org/abs/2506.11747
备注:pre-print
摘要:使用儿童佩戴的麦克风获得的长格式音频录音(也称为以儿童为中心的全天录音)已成为研究儿童语言体验及其对后续语言发展影响的标准方法。长篇语音音频的文字记录可以在各个语言层面进行丰富的分析,但典型长篇语料库的大规模却无法进行全面的手动注释。与此同时,由于现实世界音频的嘈杂、不受约束的性质,基于自动语音识别(ASR)的转录面临着重大挑战,并且没有现有的研究成功地应用ASR来转录这些数据。然而,以前的尝试都假设ASR必须完整地处理每个长格式记录。在这项工作中,我们提出了一种方法来自动检测那些可以可靠地转录与现代ASR系统的长格式音频的话语,允许自动和相对准确的转录的一个显着比例的所有语音在典型的长格式数据。我们验证的方法在四个英语长篇音频语料库,显示它实现了一个中位数字错误率(WER)为0%,平均WER为18%时,转录13%的总语音数据集。相比之下,转录所有语音而不进行任何过滤产生52%的中值WER和51%的平均WER。我们还比较了来自自动成绩单与手动注释的单词对数频率,并表明所有转录单词的频率相关性为r = 0.92(Pearson),自动成绩单中出现至少5次的单词的频率相关性为r = 0.98。总的来说,这项工作为以儿童为中心的长篇音频的日益详细的自动语言分析迈出了具体的一步。
摘要:Longform audio recordings obtained with microphones worn by children-also known as child-centered daylong recordings-have become a standard method for studying children's language experiences and their impact on subsequent language development. Transcripts of longform speech audio would enable rich analyses at various linguistic levels, yet the massive scale of typical longform corpora prohibits comprehensive manual annotation. At the same time, automatic speech recognition (ASR)-based transcription faces significant challenges due to the noisy, unconstrained nature of real-world audio, and no existing study has successfully applied ASR to transcribe such data. However, previous attempts have assumed that ASR must process each longform recording in its entirety. In this work, we present an approach to automatically detect those utterances in longform audio that can be reliably transcribed with modern ASR systems, allowing automatic and relatively accurate transcription of a notable proportion of all speech in typical longform data. We validate the approach on four English longform audio corpora, showing that it achieves a median word error rate (WER) of 0% and a mean WER of 18% when transcribing 13% of the total speech in the dataset. In contrast, transcribing all speech without any filtering yields a median WER of 52% and a mean WER of 51%. We also compare word log-frequencies derived from the automatic transcripts with those from manual annotations and show that the frequencies correlate at r = 0.92 (Pearson) for all transcribed words and r = 0.98 for words that appear at least five times in the automatic transcripts. Overall, the work provides a concrete step toward increasingly detailed automated linguistic analyses of child-centered longform audio.
标题: (SimPhon语音测试):语音平衡语音测试的数字驱动方法
链接:https://arxiv.org/abs/2506.11620
摘要:传统的测听通常不能完全表征听力损失对言语理解的功能影响,特别是对于老年性耳聋常见的阈上缺陷。这促使了更具体的诊断言语感知测试的发展。我们介绍了模拟音素语音测试(SimPhon语音测试)的方法,一种新颖的,多阶段的计算流水线在硅片上的设计和验证的语音平衡的最小对语音测试。该方法利用现代自动语音识别(ASR)系统作为人类听众的代理来模拟感音神经性听力损失的感知效果。通过在受控的声学退化下处理语音刺激,我们首先识别最常见的音素混淆模式。然后,这些模式指导从综合语言语料库中导出的大量候选词对的数据驱动的策展。随后的阶段包括模拟诊断测试,专家人工策展和最终的有针对性的敏感性分析,系统地将候选人减少到最终的优化的25对(SimPhon Speech Test-25)。一个关键的发现是,SimPhon Speech Test-25测试项目的诊断性能与标准语音可懂度指数(SII)的预测没有显着相关性,这表明SimPhon Speech Test捕获了简单可听度之外的感知缺陷。这种经过计算优化的测试集显着提高了听力测试开发的效率,为初步人体试验做好了准备。
摘要:Traditional audiometry often provides an incomplete characterization of the functional impact of hearing loss on speech understanding, particularly for supra-threshold deficits common in presbycusis. This motivates the development of more diagnostically specific speech perception tests. We introduce the Simulated Phoneme Speech Test (SimPhon Speech Test) methodology, a novel, multi-stage computational pipeline for the in silico design and validation of a phonetically balanced minimal-pair speech test. This methodology leverages a modern Automatic Speech Recognition (ASR) system as a proxy for a human listener to simulate the perceptual effects of sensorineural hearing loss. By processing speech stimuli under controlled acoustic degradation, we first identify the most common phoneme confusion patterns. These patterns then guide the data-driven curation of a large set of candidate word pairs derived from a comprehensive linguistic corpus. Subsequent phases involving simulated diagnostic testing, expert human curation, and a final, targeted sensitivity analysis systematically reduce the candidates to a final, optimized set of 25 pairs (the SimPhon Speech Test-25). A key finding is that the diagnostic performance of the SimPhon Speech Test-25 test items shows no significant correlation with predictions from the standard Speech Intelligibility Index (SII), suggesting the SimPhon Speech Test captures perceptual deficits beyond simple audibility. This computationally optimized test set offers a significant increase in efficiency for audiological test development, ready for initial human trials.
标题: 用矩阵分解端到端扩张的分割模型
链接:https://arxiv.org/abs/2506.11605
备注:37 pages, 18 figures. Submitted to Computer Speech & Language
摘要:带向量聚类的端到端神经日志化是一种强大而实用的说话人日志化方法。针对这些管道的分割模型提出了多项增强措施,但尚未对其协同作用进行彻底评估。在这项工作中,我们提供了一个深入的分析主要的架构选择对管道的性能的影响。我们研究了不同的编码器(SincNet,预训练和微调的WavLM),不同的解码器(LSTM,Mamba和Conformer),不同的损失(多标签和多类幂集)以及不同的块大小。通过覆盖九个数据集的深入实验,我们发现基于WavLM的微调编码器总是产生最好的系统。LSTM解码器被基于Mamba和Conformer的解码器所超越,虽然我们发现Mamba对其他架构选择更强大,但它略逊于我们使用Conformer编码器的最佳架构。我们发现,多标签和多类功率集损失不具有相同的错误分布。我们证实,多类损失有助于几乎所有模型获得卓越的性能,除了微调WavLM时,在这种情况下,多标签是更好的选择。我们还评估了块大小对所有上述架构选择的影响,发现较新的架构往往能更好地处理长块大小,这可以大大提高管道性能。我们最好的系统在五个广泛使用的说话人日记数据集上取得了最先进的结果。
摘要:End-to-End Neural Diarization with Vector Clustering is a powerful and practical approach to perform Speaker Diarization. Multiple enhancements have been proposed for the segmentation model of these pipelines, but their synergy had not been thoroughly evaluated. In this work, we provide an in-depth analysis on the impact of major architecture choices on the performance of the pipeline. We investigate different encoders (SincNet, pretrained and finetuned WavLM), different decoders (LSTM, Mamba, and Conformer), different losses (multilabel and multiclass powerset), and different chunk sizes. Through in-depth experiments covering nine datasets, we found that the finetuned WavLM-based encoder always results in the best systems by a wide margin. The LSTM decoder is outclassed by Mamba- and Conformer-based decoders, and while we found Mamba more robust to other architecture choices, it is slightly inferior to our best architecture, which uses a Conformer encoder. We found that multilabel and multiclass powerset losses do not have the same distribution of errors. We confirmed that the multiclass loss helps almost all models attain superior performance, except when finetuning WavLM, in which case, multilabel is the superior choice. We also evaluated the impact of the chunk size on all aforementioned architecture choices and found that newer architectures tend to better handle long chunk sizes, which can greatly improve pipeline performance. Our best system achieved state-of-the-art results on five widely used speaker diarization datasets.
标题: 语音反欺骗中的语音增强放大伪影
链接:https://arxiv.org/abs/2506.11542
备注:Accepted to Interspeech2025
摘要:欺骗话语总是包含生成模型引入的伪像。虽然已经提出了几种对策来检测欺骗性话语,但大多数主要集中在架构改进上。在这项工作中,我们研究如何文物仍然隐藏在欺骗性的讲话,以及如何提高他们的存在。我们提出了一个模型不可知的管道,使用语音增强和各种类型的噪声放大文物。我们的方法包括三个关键步骤:噪声添加、噪声提取和噪声放大。首先,我们在原始语音中引入噪声。然后,我们应用语音增强来提取纠缠噪声和伪影。最后,我们放大这些提取的特征。此外,我们的管道是兼容不同的语音增强模型和对策架构。我们的方法在ASVspoof2019上将欺骗检测性能提高了44.44\%,在ASVspoof2021上提高了26.34\%。
摘要:Spoofed utterances always contain artifacts introduced by generative models. While several countermeasures have been proposed to detect spoofed utterances, most primarily focus on architectural improvements. In this work, we investigate how artifacts remain hidden in spoofed speech and how to enhance their presence. We propose a model-agnostic pipeline that amplifies artifacts using speech enhancement and various types of noise. Our approach consists of three key steps: noise addition, noise extraction, and noise amplification. First, we introduce noise into the raw speech. Then, we apply speech enhancement to extract the entangled noise and artifacts. Finally, we amplify these extracted features. Moreover, our pipeline is compatible with different speech enhancement models and countermeasure architectures. Our method improves spoof detection performance by up to 44.44\% on ASVspoof2019 and 26.34\% on ASVspoof2021.
标题: LiLAC:用于音乐音频生成的轻量级潜在控制网络
链接:https://arxiv.org/abs/2506.11476
备注:Accepted at ISMIR 2025
摘要:文本到音频扩散模型产生高质量和多样化的音乐,但许多(如果不是大多数的话)SOTA模型缺乏音乐制作所必需的细粒度、时变控制。ControlNet通过克隆和微调新条件下的编码器,将外部控件附加到预训练的生成模型。但是,这种方法会占用大量内存,并将用户限制在一组固定的控件中。我们提出了一个轻量级的,模块化的架构,大大减少了参数计数,同时匹配ControlNet的音频质量和条件的遵守。我们的方法提供了更大的灵活性和显着降低内存使用,使独立控件的更有效的培训和部署。我们进行了广泛的客观和主观评估,并在附带的网站https://lightlatentcontrol.github.io上提供了大量的音频示例
摘要:Text-to-audio diffusion models produce high-quality and diverse music but many, if not most, of the SOTA models lack the fine-grained, time-varying controls essential for music production. ControlNet enables attaching external controls to a pre-trained generative model by cloning and fine-tuning its encoder on new conditionings. However, this approach incurs a large memory footprint and restricts users to a fixed set of controls. We propose a lightweight, modular architecture that considerably reduces parameter count while matching ControlNet in audio quality and condition adherence. Our method offers greater flexibility and significantly lower memory usage, enabling more efficient training and deployment of independent controls. We conduct extensive objective and subjective evaluations and provide numerous audio examples on the accompanying website at https://lightlatentcontrol.github.io
标题: 语音-音乐编码器模型合并的相关排列方法
链接:https://arxiv.org/abs/2506.11403
备注:Under review
摘要:创建统一的语音和音乐模型需要昂贵的预训练。模型合并可以替代地创建具有最小计算开销的统一音频模型。然而,当模型在权重空间中不对齐时,直接合并具有挑战性。受Git Re-Basin的启发,我们引入了一种相关置换方法,该方法将音乐编码器的内部层与语音编码器对齐。我们将以前的工作扩展到合并Transformer层的情况。该方法计算一个置换矩阵,使模型的逐层特征互相关性最大化,从而实现这些不相交模型的有效融合。通过这种方法,合并后的模型保留了语音功能,同时显著提高了音乐性能,与线性插值模型合并相比,平均得分提高了14.83分。这项工作允许从独立训练的编码器创建统一的音频模型。
摘要:Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, direct merging is challenging when the models are not aligned in the weight space. Motivated by Git Re-Basin, we introduce a correlation-permutation approach that aligns a music encoder's internal layers with a speech encoder. We extend previous work to the case of merging transformer layers. The method computes a permutation matrix that maximizes the model's features-wise cross-correlations layer by layer, enabling effective fusion of these otherwise disjoint models. The merged model retains speech capabilities through this method while significantly enhancing music performance, achieving an improvement of 14.83 points in average score compared to linear interpolation model merging. This work allows the creation of unified audio models from independently trained encoders.
标题: GSYS:跨领域和语言的一般对比音频文本预训练
链接:https://arxiv.org/abs/2506.11350
摘要:对比语言音频预训练(CLAP)是一种广泛使用的方法,以弥合音频和文本域之间的差距。目前的CLAP方法支持英语的声音和音乐检索,忽略了多语言的口语内容。为了解决这个问题,我们引入了通用语言音频预训练(Gestival),它扩展了CLAP的多语言和多领域能力。通过在Clotho和AudioCaps等标准音频文本检索基准上实现具有竞争力的性能,Gestro展示了其多功能性,同时在语音检索和分类任务中显着超越了现有方法。此外,Gestro在广泛使用的声音事件zero-shot基准上取得了很好的结果,同时在语音内容基准上优于以前的方法。对50种语言的进一步关键字识别评估强调了GLAP的高级多语言功能。最后,在四种语言中评估多语言声音和音乐理解。检查点和资料来源:https://github.com/xiaomi-research/dasheng-glap。
摘要:Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in English, ignoring multilingual spoken content. To address this, we introduce general language audio pretraining (GLAP), which expands CLAP with multilingual and multi-domain abilities. GLAP demonstrates its versatility by achieving competitive performance on standard audio-text retrieval benchmarks like Clotho and AudioCaps, while significantly surpassing existing methods in speech retrieval and classification tasks. Additionally, GLAP achieves strong results on widely used sound-event zero-shot benchmarks, while simultaneously outperforming previous methods on speech content benchmarks. Further keyword spotting evaluations across 50 languages emphasize GLAP's advanced multilingual capabilities. Finally, multilingual sound and music understanding is evaluated across four languages. Checkpoints and Source: https://github.com/xiaomi-research/dasheng-glap.
标题: MUPAS:多标签声音分类中的点尺度无监督域自适应
链接:https://arxiv.org/abs/2506.11331
摘要:无监督域自适应(UDA)对于使机器学习模型适应新的、无标签的环境至关重要,在这些环境中,数据分布的变化会降低性能。现有的UDA算法是针对单标签任务设计的,并且依赖于大量的计算资源,限制了它们在多标签场景和资源受限的物联网设备中的使用。在城市声音分类等环境中,克服这些限制尤其具有挑战性,其中重叠的声音和变化的声学需要低功耗设备上系统的鲁棒的自适应多标签功能。为了解决这些限制,我们引入了声音的无监督域自适应(MUDAS),这是一个为资源受限的物联网环境中的多标签声音分类而开发的UDA框架。MUDAS通过使用高置信度数据选择性地原位重新训练分类器来有效地适应模型,最大限度地减少计算和内存需求,以适应设备上的部署。此外,MUDAS采用了类特定的自适应阈值来生成可靠的伪标签,并应用多样性正则化来提高多标签分类的准确性。在对纽约市多个地点记录的SONYC Urban Sound Tagging(SONYC-UST)数据集进行评估时,MUDAS在现有UDA算法的分类准确性方面取得了显着改进,在资源受限的物联网环境中实现了良好的性能。
摘要:Unsupervised Domain Adaptation (UDA) is essential for adapting machine learning models to new, unlabeled environments where data distribution shifts can degrade performance. Existing UDA algorithms are designed for single-label tasks and rely on significant computational resources, limiting their use in multi-label scenarios and in resource-constrained IoT devices. Overcoming these limitations is particularly challenging in contexts such as urban sound classification, where overlapping sounds and varying acoustics require robust, adaptive multi-label capabilities on low-power, on-device systems. To address these limitations, we introduce Mote-scale Unsupervised Domain Adaptation for Sounds (MUDAS), a UDA framework developed for multi-label sound classification in resource-constrained IoT settings. MUDAS efficiently adapts models by selectively retraining the classifier in situ using high-confidence data, minimizing computational and memory requirements to suit on-device deployment. Additionally, MUDAS incorporates class-specific adaptive thresholds to generate reliable pseudo-labels and applies diversity regularization to improve multi-label classification accuracy. In evaluations on the SONYC Urban Sound Tagging (SONYC-UST) dataset recorded at various New York City locations, MUDAS demonstrates notable improvements in classification accuracy over existing UDA algorithms, achieving good performance in a resource-constrained IoT setting.
标题: 使用TTS合成数据增强ASB的自细化框架
链接:https://arxiv.org/abs/2506.11130
摘要:我们提出了一个自我完善的框架,提高ASR的性能,只有未标记的数据集。该过程从现有的ASR模型开始,在未注释的语音上生成伪标签,然后将其用于训练高保真文本到语音(TTS)系统。然后,合成的语音文本对被引导到原始ASR系统中,完成闭环自我改进循环。我们证明了该框架对台湾普通话语音的有效性。利用6,000小时的未标记语音,适量的文本数据和人工智能模型的合成内容,我们将Whisper-large-v2改编为一个专门的模型Twister。与Whisper相比,Twister在普通话上减少了20%的错误率,在普通话-英语代码转换基准测试中减少了50%。结果突出的框架作为一个令人信服的替代伪标签自蒸馏方法,并提供了一个实用的途径,以提高ASR性能在低资源或特定领域的设置。
摘要:We propose a self-refining framework that enhances ASR performance with only unlabeled datasets. The process starts with an existing ASR model generating pseudo-labels on unannotated speech, which are then used to train a high-fidelity text-to-speech (TTS) system. Then, synthesized speech text pairs are bootstrapped into the original ASR system, completing the closed-loop self-improvement cycle. We demonstrated the effectiveness of the framework on Taiwanese Mandarin speech. Leveraging 6,000 hours of unlabeled speech, a moderate amount of text data, and synthetic content from the AI models, we adapt Whisper-large-v2 into a specialized model, Twister. Twister reduces error rates by up to 20% on Mandarin and 50% on Mandarin-English code-switching benchmarks compared to Whisper. Results highlight the framework as a compelling alternative to pseudo-labeling self-distillation approaches and provides a practical pathway for improving ASR performance in low-resource or domain-specific settings.
标题: SUTA-LM:连接测试时自适应和语言模型重新评分以实现稳健的ASB
链接:https://arxiv.org/abs/2506.11121
摘要:尽管在端到端ASR方面取得了进展,但现实世界中的领域不匹配仍然会导致性能下降,测试时自适应(TTA)旨在通过在推理期间调整模型来缓解这种情况。最近的工作探索将TTA与外部语言模型相结合,使用波束搜索重新评分或生成错误校正等技术。在这项工作中,我们确定了一个以前被忽视的挑战:TTA可以干扰语言模型重新评分,揭示了有效结合这两种方法的重要性。基于这种见解,我们提出了SUTA-LM,一个简单而有效的扩展SUTA,一个基于熵最小化的TTA方法,语言模型重新评分。SUTA-LM首先应用由自动步骤选择机制引导的受控自适应过程,该机制利用声学和语言信息,然后通过语言模型重新评分来细化输出。在18个不同的ASR数据集上的实验表明,SUTA-LM在广泛的领域中实现了稳健的结果。
摘要:Despite progress in end-to-end ASR, real-world domain mismatches still cause performance drops, which Test-Time Adaptation (TTA) aims to mitigate by adjusting models during inference. Recent work explores combining TTA with external language models, using techniques like beam search rescoring or generative error correction. In this work, we identify a previously overlooked challenge: TTA can interfere with language model rescoring, revealing the nontrivial nature of effectively combining the two methods. Based on this insight, we propose SUTA-LM, a simple yet effective extension of SUTA, an entropy-minimization-based TTA approach, with language model rescoring. SUTA-LM first applies a controlled adaptation process guided by an auto-step selection mechanism leveraging both acoustic and linguistic information, followed by language model rescoring to refine the outputs. Experiments on 18 diverse ASR datasets show that SUTA-LM achieves robust results across a wide range of domains.
标题: 基准基础语音和语言模型用于阿尔茨海默病和相关痴呆症的自发语音检测
链接:https://arxiv.org/abs/2506.11119
摘要:背景资料:阿尔茨海默病和相关痴呆症(ADRD)是一种进行性神经退行性疾病,早期发现对于及时干预和护理至关重要。自发言语包含丰富的声学和语言标记,可以作为认知能力下降的非侵入性生物标志物。在大规模音频或文本数据上预训练的基础模型产生编码上下文和声学特征的高维嵌入。 研究方法:我们使用了CHTE挑战数据集,其中包括来自1,600多名参与者的录音,这些参与者具有三种认知状态:健康对照(HC),轻度认知障碍(MCI)和阿尔茨海默病(AD)。我们排除了非英语、非自发或质量差的录音。最终数据集包括703例(59.13%)HC、81例(6.81%)MCI和405例(34.06%)AD病例。我们对一系列开源基础语音和语言模型进行了基准测试,将认知状态分为三类。 结果:Whisper-medium模型在语音模型中获得了最高的性能(准确率= 0.731,AUC = 0.802)。在语言模型中,带有停顿注释的BERT表现最好(准确度= 0.662,AUC = 0.744)。使用最先进的自动语音识别(ASR)模型生成的音频嵌入的ADRD检测优于其他方法。包括非语义特征,如暂停模式,始终改进了基于文本的分类。 结论:本研究介绍了一个基准框架,使用基础模型和临床相关数据集。基于声学的方法-特别是ASR衍生的嵌入-表现出可扩展的,非侵入性的和具有成本效益的ADRD早期检测的强大潜力。
摘要:Background: Alzheimer's disease and related dementias (ADRD) are progressive neurodegenerative conditions where early detection is vital for timely intervention and care. Spontaneous speech contains rich acoustic and linguistic markers that may serve as non-invasive biomarkers for cognitive decline. Foundation models, pre-trained on large-scale audio or text data, produce high-dimensional embeddings encoding contextual and acoustic features. Methods: We used the PREPARE Challenge dataset, which includes audio recordings from over 1,600 participants with three cognitive statuses: healthy control (HC), mild cognitive impairment (MCI), and Alzheimer's Disease (AD). We excluded non-English, non-spontaneous, or poor-quality recordings. The final dataset included 703 (59.13%) HC, 81 (6.81%) MCI, and 405 (34.06%) AD cases. We benchmarked a range of open-source foundation speech and language models to classify cognitive status into the three categories. Results: The Whisper-medium model achieved the highest performance among speech models (accuracy = 0.731, AUC = 0.802). Among language models, BERT with pause annotation performed best (accuracy = 0.662, AUC = 0.744). ADRD detection using state-of-the-art automatic speech recognition (ASR) model-generated audio embeddings outperformed others. Including non-semantic features like pause patterns consistently improved text-based classification. Conclusion: This study introduces a benchmarking framework using foundation models and a clinically relevant dataset. Acoustic-based approaches -- particularly ASR-derived embeddings -- demonstrate strong potential for scalable, non-invasive, and cost-effective early detection of ADRD.
标题: 评估语音神经表示中各向异性的影响:关键词发现的案例研究
链接:https://arxiv.org/abs/2506.11096
摘要:像wav2vec2和HuBERT这样的预训练语音表示表现出很强的各向异性,导致随机嵌入之间的高度相似性。虽然已被广泛观察到,但这一特性对下游任务的影响仍不清楚。这项工作评估各向异性的关键字发现计算文献语言学。使用动态时间规整,我们表明,尽管各向异性,wav2vec2相似性措施有效地识别单词没有转录。我们的研究结果突出了这些表征的鲁棒性,这些表征捕获了语音结构并在说话者之间进行概括。我们的研究结果强调了预训练在学习丰富和不变的语音表示中的重要性。
摘要:Pretrained speech representations like wav2vec2 and HuBERT exhibit strong anisotropy, leading to high similarity between random embeddings. While widely observed, the impact of this property on downstream tasks remains unclear. This work evaluates anisotropy in keyword spotting for computational documentary linguistics. Using Dynamic Time Warping, we show that despite anisotropy, wav2vec2 similarity measures effectively identify words without transcription. Our results highlight the robustness of these representations, which capture phonetic structures and generalize across speakers. Our results underscore the importance of pretraining in learning rich and invariant speech representations.
标题: 利用大语言模型反馈定制语音识别模型
链接:https://arxiv.org/abs/2506.11091
摘要:自动语音识别(ASR)系统在一般的转录任务上取得了很好的性能。然而,他们仍然在努力识别罕见的命名实体并适应域不匹配。相比之下,在大量互联网规模的数据集上训练的大型语言模型(LLM)在广泛的领域中通常更有效。在这项工作中,我们提出了一种基于强化学习的无监督域自适应方法,利用未标记的数据来提高转录质量,特别是通过LLM的反馈来提高受域不匹配影响的命名实体。鉴于上下文信息,我们的框架采用LLM作为奖励模型,从ASR模型中对假设进行评分。这些分数作为奖励信号,通过强化学习来微调ASR模型。与传统的自训练方法相比,该方法在实体词错误率上提高了21%.
摘要:Automatic speech recognition (ASR) systems have achieved strong performance on general transcription tasks. However, they continue to struggle with recognizing rare named entities and adapting to domain mismatches. In contrast, large language models (LLMs), trained on massive internet-scale datasets, are often more effective across a wide range of domains. In this work, we propose a reinforcement learning based approach for unsupervised domain adaptation, leveraging unlabeled data to enhance transcription quality, particularly the named entities affected by domain mismatch, through feedback from a LLM. Given contextual information, our framework employs a LLM as the reward model to score the hypotheses from the ASR model. These scores serve as reward signals to fine-tune the ASR model via reinforcement learning. Our method achieves a 21\% improvement on entity word error rate over conventional self-training methods.
标题: 利用吸引器深度集群的端到端扩展
链接:https://arxiv.org/abs/2506.11090
备注:To appear at INTERSPEECH 2025
摘要:说话人日记化仍然具有挑战性,由于需要结构化的说话人表示,有效的建模和鲁棒性变化的条件。我们提出了一个高性能的,紧凑的日记化框架,集成了构象解码器,变换器更新的吸引子,和一个深聚类风格的角度损失。我们的方法改进了扬声器表示与增强的构象结构,将交叉注意力吸引子和一个额外的卷积模块。为了加强结构化嵌入,我们通过构建标签吸引子向量来扩展深度聚类,将其方向结构与音频嵌入对齐。我们还施加正交约束的积极吸引更好的扬声器分离,同时抑制非积极吸引,以防止假激活。最后,一个排列不变的训练二进制交叉熵损失细化说话人检测。实验结果表明,该方法在保持参数个数不变的情况下,实现了较低的日志化误差。
摘要:Speaker diarization remains challenging due to the need for structured speaker representations, efficient modeling, and robustness to varying conditions. We propose a performant, compact diarization framework that integrates conformer decoders, transformer-updated attractors, and a deep clustering style angle loss. Our approach refines speaker representations with an enhanced conformer structure, incorporating cross-attention to attractors and an additional convolution module. To enforce structured embeddings, we extend deep clustering by constructing label-attractor vectors, aligning their directional structure with audio embeddings. We also impose orthogonality constraints on active attractors for better speaker separation while suppressing non-active attractors to prevent false activations. Finally, a permutation invariant training binary cross-entropy loss refines speaker detection. Experiments show that our method achieves low diarization error while maintaining parameter count.
标题: 从清晰度到更好的泛化,用于语音Deepfake检测
链接:https://arxiv.org/abs/2506.11532
备注:Accepted to Interspeech 2025
摘要:泛化仍然是语音深度伪造检测(SDD)的关键挑战。虽然各种方法的目的是提高鲁棒性,泛化通常是通过性能指标,如相等的错误率,没有一个理论框架来解释模型的性能进行评估。这项工作调查的清晰度作为一个理论代理泛化SDD。我们分析了锐度如何响应域的变化,并发现它在看不见的条件下增加,表明更高的模型灵敏度。在此基础上,我们应用锐度感知最小化(SAM)来显式降低锐度,从而在各种不可见的测试集上获得更好、更稳定的性能。此外,相关性分析证实了在大多数测试设置的锐度和泛化之间的统计上显着的关系。这些研究结果表明,清晰度可以作为SDD泛化的理论指标,清晰度感知训练提供了一个有前途的策略,以提高鲁棒性。
摘要:Generalization remains a critical challenge in speech deepfake detection (SDD). While various approaches aim to improve robustness, generalization is typically assessed through performance metrics like equal error rate without a theoretical framework to explain model performance. This work investigates sharpness as a theoretical proxy for generalization in SDD. We analyze how sharpness responds to domain shifts and find it increases in unseen conditions, indicating higher model sensitivity. Based on this, we apply Sharpness-Aware Minimization (SAM) to reduce sharpness explicitly, leading to better and more stable performance across diverse unseen test sets. Furthermore, correlation analysis confirms a statistically significant relationship between sharpness and generalization in most test settings. These findings suggest that sharpness can serve as a theoretical indicator for generalization in SDD and that sharpness-aware training offers a promising strategy for improving robustness.
标题: 通过预训练生成音频编码器的嵌入实现高效语音增强
链接:https://arxiv.org/abs/2506.11514
备注:Accepted by Interspeech 2025
摘要:最近的研究深入研究了利用来自预训练模型的音频嵌入的语音增强(SE)方法,与时频掩蔽或信号预测技术不同。本文介绍了一种有效的和可扩展的SE方法。我们的方法涉及最初提取音频嵌入从嘈杂的语音使用预训练的音频编码器,然后由一个紧凑的编码器网络去噪。随后,声码器从去噪嵌入合成干净的语音。消融研究证实了预训练的音频编码器和声码器的降噪编码器的参数效率。语音增强和扬声器保真度的实验结果表明,我们的生成audiocoder为基础的SE系统优于模型利用歧视性audiocoder。此外,主观听力测试验证,我们提出的系统超越了现有的国家的最先进的SE模型的感知质量。
摘要:Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal prediction techniques. This paper introduces an efficient and extensible SE method. Our approach involves initially extracting audio embeddings from noisy speech using a pre-trained audioencoder, which are then denoised by a compact encoder network. Subsequently, a vocoder synthesizes the clean speech from denoised embeddings. An ablation study substantiates the parameter efficiency of the denoise encoder with a pre-trained audioencoder and vocoder. Experimental results on both speech enhancement and speaker fidelity demonstrate that our generative audioencoder-based SE system outperforms models utilizing discriminative audioencoders. Furthermore, subjective listening tests validate that our proposed system surpasses an existing state-of-the-art SE model in terms of perceptual quality.
标题: 小足迹关键词发现的进展:高效模型和算法的全面回顾
链接:https://arxiv.org/abs/2506.11169
备注:61 pages, 21 figures
摘要:小占用空间关键字定位(SF-KWS)在当今智能语音激活设备、智能手机和物联网(IoT)应用领域中越来越受欢迎。这一激增归功于深度学习的进步,它能够从连续的单词流中识别预定义的单词或关键字。为了在实际场景中的低功耗和有限内存的边缘设备上实现SF-KWS模型,高效的Tiny Machine Learning(TinyML)框架至关重要。在这项研究中,我们探讨了七个不同类别的技术,即模型架构,学习技术,模型压缩,注意力感知架构,特征优化,神经网络搜索,和混合方法,这是适合开发一个SF-KWS系统。这一全面的概述将作为那些希望了解,利用或有助于SF-KWS领域的宝贵资源。在这项工作中进行的分析,使许多潜在的研究方向的识别,包括自动语音识别研究的见解和那些具体有关的领域的口语SF-KWS。
摘要:Small-Footprint Keyword Spotting (SF-KWS) has gained popularity in today's landscape of smart voice-activated devices, smartphones, and Internet of Things (IoT) applications. This surge is attributed to the advancements in Deep Learning, enabling the identification of predefined words or keywords from a continuous stream of words. To implement the SF-KWS model on edge devices with low power and limited memory in real-world scenarios, a efficient Tiny Machine Learning (TinyML) framework is essential. In this study, we explore seven distinct categories of techniques namely, Model Architecture, Learning Techniques, Model Compression, Attention Awareness Architecture, Feature Optimization, Neural Network Search, and Hybrid Approaches, which are suitable for developing an SF-KWS system. This comprehensive overview will serve as a valuable resource for those looking to understand, utilize, or contribute to the field of SF-KWS. The analysis conducted in this work enables the identification of numerous potential research directions, encompassing insights from automatic speech recognition research and those specifically pertinent to the realm of spoken SF-KWS.
标题: S2 ST-Omni:通过完美的语音文本对齐和流语音解码器实现高效且可扩展的多语言语音到语音翻译框架
链接:https://arxiv.org/abs/2506.11160
备注:Working in progress
摘要:多语言语音到语音翻译(S2 ST)的目的是直接将多个源语言的口语转换为目标语言的自然和可理解的语音。尽管最近取得了进展,但仍存在重大挑战:(1)实现高质量和低延迟的S2 ST仍然是一个关键的障碍;(2)现有的S2 ST方法严重依赖于大规模的并行语音语料库,这是非常困难的收集。为了解决这些问题,我们提出了S2 ST-Omni,一个高效和可扩展的多语种语音到语音翻译框架。具体来说,我们将S2 ST任务分解为语音到文本翻译(S2 TT)和文本到语音合成(TTS),将它们统一在一个端到端的语音语言模型中。为了实现高质量的S2 TT,同时减少对并行语料库的依赖,我们利用大规模的预训练模型- Whisper用于音频理解,Qwen 3.0用于文本理解。引入轻量级语音适配器来对齐语音和文本表示,从而能够有效使用预训练的多模式知识。为了保证翻译质量和实时性,我们在TTS阶段采用预训练的流式语音解码器以自回归方式生成目标语音。在CVSS基准测试上进行的大量实验表明,S2 ST-Omni的性能优于最先进的S2 ST基线,同时保持相当的延迟,突出了其在实际部署中的有效性和实用潜力。
摘要:Multilingual speech-to-speech translation (S2ST) aims to directly convert spoken utterances from multiple source languages into natural and intelligible speech in a target language. Despite recent progress, significant challenges remain: (1) achieving high-quality and low-latency S2ST remains a critical hurdle; (2) existing S2ST approaches heavily rely on large-scale parallel speech corpora, which are extremely difficult to collect. To address these issues, we propose S2ST-Omni, an efficient and scalable framework for multilingual speech-to-speech translation. Specifically, we decompose the S2ST task into speech-to-text translation (S2TT) and text-to-speech synthesis (TTS), unifying them within a single end-to-end speech-language model. To achieve high-quality S2TT while reducing dependence on parallel corpora, we leverage large-scale pretrained models -- Whisper for audio understanding and Qwen 3.0 for text understanding. A lightweight speech adapter is introduced to align speech and text representations, enabling effective use of pretrained multimodal knowledge. To ensure both translation quality and real-time performance, we adopt a pretrained streaming speech decoder in the TTS stage to generate target speech in an autoregressive manner. Extensive experiments on the CVSS benchmark demonstrate that S2ST-Omni outperforms state-of-the-art S2ST baselines while maintaining comparable latency, highlighting its effectiveness and practical potential for real-world deployment.
标题: 间歇性和移动的说话人的跟踪:数据集和网络
链接:https://arxiv.org/abs/2506.11145
备注:None
摘要:本文提出的问题跟踪间歇和移动源,即可能改变位置时,他们是不活跃的来源。这个问题很少被探讨,目前大多数跟踪方法依赖于空间观测的跟踪身份管理。它们或者基于先前的定位步骤,或者被设计为通过预测有序的位置估计来执行联合定位和跟踪。这引起了对这样的方法是否可以在处理不连续的空间轨道时保持可靠的轨道标识分配性能的关注,所述不连续的空间轨道可能由静默期间的方向改变引起。我们介绍LibriJump,一个新的数据集的声学场景的一阶高保真度立体声格式,专注于扬声器跟踪。该数据集包含在不活动期间具有变化位置的扬声器,从而模拟不连续的音轨。为了测量身份分配性能,我们建议使用从计算机视觉社区改编的跟踪关联度量。我们提供的实验显示的互补性的关联度量与以前使用的跟踪度量,连续和不连续的空间轨道。
摘要:This paper presents the problem of tracking intermittent and moving sources, i.e, sources that may change position when they are inactive. This issue is seldom explored, and most current tracking methods rely on spatial observations for track identity management. They are either based on a previous localization step, or designed to perform joint localization and tracking by predicting ordered position estimates. This raises concerns about whether such methods can maintain reliable track identity assignment performance when dealing with discontinuous spatial tracks, which may be caused by a change of direction during silence. We introduce LibriJump, a novel dataset of acoustic scenes in the First Order Ambisonics format focusing on speaker tracking. The dataset contains speakers with changing positions during inactivity periods, thus simulating discontinuous tracks. To measure the identity assignment performance, we propose to use tracking association metrics adapted from the computer vision community. We provide experiments showing the complementarity of association metrics with previously used tracking metrics, given continuous and discontinuous spatial tracks.
标题: 利用预设提高儿童语音识别和阅读错误检测
链接:https://arxiv.org/abs/2506.11079
备注:This paper is accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:自动朗读评估可以为教师提供有价值的支持,使阅读练习更有效的评分。然而,阅读评价系统和应用的研究仍然有限。我们提出了一种新的多模态方法,利用音频和知识的文本资源。特别是,我们探索了使用Whisper和提示调整的大型语言模型(LLM)的潜力,以改善儿童语音识别的transmittance,以及它们在下游阅读错误检测中的有效性。我们的研究结果表明,与没有提示的基线Whisper模型相比,提示Whisper和提示LLM的有效性。表现最好的系统在荷兰儿童阅读语音中实现了最先进的识别性能,单词错误率(WER)为5.1%,提高了9.4%的基线WER。此外,它显着提高了阅读错误检测,将F1分数从0.39提高到0.73。
摘要:Automatic reading aloud evaluation can provide valuable support to teachers by enabling more efficient scoring of reading exercises. However, research on reading evaluation systems and applications remains limited. We present a novel multimodal approach that leverages audio and knowledge from text resources. In particular, we explored the potential of using Whisper and instruction-tuned large language models (LLMs) with prompts to improve transcriptions for child speech recognition, as well as their effectiveness in downstream reading mistake detection. Our results demonstrate the effectiveness of prompting Whisper and prompting LLM, compared to the baseline Whisper model without prompting. The best performing system achieved state-of-the-art recognition performance in Dutch child read speech, with a word error rate (WER) of 5.1%, improving the baseline WER of 9.4%. Furthermore, it significantly improved reading mistake detection, increasing the F1 score from 0.39 to 0.73.
标题: 十五年以儿童为中心的长篇录音:承诺、资源和有效性的剩余挑战
链接:https://arxiv.org/abs/2506.11075
备注:5 pages, 3 figures
摘要:使用儿童佩戴的设备收集的音频记录是儿童语言研究的基本工具。全天收集的长形式录音承诺以最小的观察者偏差捕获儿童的输入和生产,因此具有较高的有效性。所产生的数据量之大,需要进行自动化分析,为研究人员和临床医生提取相关指标。本文总结了关于此技术的集体知识,为现有资源提供了切入点。我们还强调了威胁自动化注释的准确性和对结果指标的解释的各种错误来源。为了解决这个问题,我们提出了潜在的故障排除指标,以帮助用户评估数据质量。虽然完全自动化的质量控制系统是不可行的,但我们概述了研究人员改进数据收集和分析的实用策略。
摘要:Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and therefore high validity. The sheer volume of resulting data necessitates automated analysis to extract relevant metrics for researchers and clinicians. This paper summarizes collective knowledge on this technique, providing entry points to existing resources. We also highlight various sources of error that threaten the accuracy of automated annotations and the interpretation of resulting metrics. To address this, we propose potential troubleshooting metrics to help users assess data quality. While a fully automated quality control system is not feasible, we outline practical strategies for researchers to improve data collection and contextualize their analyses.
标题: 儿童可穿戴设备语音自动处理面临的挑战:语音类型分类器的案例
链接:https://arxiv.org/abs/2506.11074
备注:5 pages, 3 figures
摘要:用儿童佩戴的设备收集的录音有望通过毫不费力地捕捉儿童的自然语言环境和语言产生来彻底改变基础和应用语言科学。这一前景取决于语音技术,它可以将如此收集到的大量数据转化为可用的信息。本文通过总结三年来旨在改进一项基本任务:语音类型分类的实验,展示了阻碍进展的几个障碍。我们的实验表明,改进表示功能,架构和参数搜索有助于只有边际收益的性能。通过关注数据的相关性和数量,取得了更多进展,这突出了收集具有适当权限的数据以允许共享的重要性。
摘要:Recordings gathered with child-worn devices promised to revolutionize both fundamental and applied speech sciences by allowing the effortless capture of children's naturalistic speech environment and language production. This promise hinges on speech technologies that can transform the sheer mounds of data thus collected into usable information. This paper demonstrates several obstacles blocking progress by summarizing three years' worth of experiments aimed at improving one fundamental task: Voice Type Classification. Our experiments suggest that improvements in representation features, architecture, and parameter search contribute to only marginal gains in performance. More progress is made by focusing on data relevance and quantity, which highlights the importance of collecting data with appropriate permissions to allow sharing.
标题: 我们可以信任机器学习吗?用于语音建模的开源语音分析工具特征的可靠性
链接:https://arxiv.org/abs/2506.11072
备注:5 pages, 1 figure, 3 tables
摘要:基于机器学习的行为模型依赖于从视听记录中提取的特征。使用开源工具处理录音,以提取分类模型的语音特征。这些工具往往缺乏验证,以确保在捕获行为相关信息的可靠性。这一差距引起了对不同人群和背景下的可重复性和公平性的关注。语音处理工具在其设计环境之外使用时,可能无法公平地捕捉行为变化,从而导致偏见。我们评估了从两个广泛使用的语音分析工具,OpenSMILE和Praat提取的语音特征,以评估其可靠性时,考虑青少年自闭症。我们观察到不同工具的功能存在相当大的差异,这影响了不同背景和人口统计学群体的模型性能。我们鼓励领域相关验证,以提高机器学习模型在临床应用中的可靠性。
摘要:Machine learning-based behavioral models rely on features extracted from audio-visual recordings. The recordings are processed using open-source tools to extract speech features for classification models. These tools often lack validation to ensure reliability in capturing behaviorally relevant information. This gap raises concerns about reproducibility and fairness across diverse populations and contexts. Speech processing tools, when used outside of their design context, can fail to capture behavioral variations equitably and can then contribute to bias. We evaluate speech features extracted from two widely used speech analysis tools, OpenSMILE and Praat, to assess their reliability when considering adolescents with autism. We observed considerable variation in features across tools, which influenced model performance across context and demographic groups. We encourage domain-relevant verification to enhance the reliability of machine learning models in clinical applications.
标题: 用于隐私保护性发音障碍和老年人语音识别的正规联邦学习
链接:https://arxiv.org/abs/2506.11069
摘要:迄今为止,构音障碍和老年言语的准确识别仍然具有挑战性。虽然隐私问题已经推动了从集中式方法向联邦学习(FL)的转变,以确保数据的机密性,但这进一步加剧了数据稀缺,数据分布不平衡和说话人异质性的挑战。为此,本文进行了一个系统的调查,隐私保护构音障碍和老年人语音识别的正则化FL技术,解决不同层次的FL过程1)参数为基础,2)嵌入为基础,3)新的损失为基础的正则化。在基准UASpeech构音障碍和DementiaBank Pitt老年人语音语料库上的实验表明,正则化FL系统的性能始终优于基线FedAvg系统,其统计学显著WER降低高达0.55%绝对值(2.13%相对值)。进一步将通信频率增加到每批一次交换接近集中训练性能。
摘要:Accurate recognition of dysarthric and elderly speech remains challenging to date. While privacy concerns have driven a shift from centralized approaches to federated learning (FL) to ensure data confidentiality, this further exacerbates the challenges of data scarcity, imbalanced data distribution and speaker heterogeneity. To this end, this paper conducts a systematic investigation of regularized FL techniques for privacy-preserving dysarthric and elderly speech recognition, addressing different levels of the FL process by 1) parameter-based, 2) embedding-based and 3) novel loss-based regularization. Experiments on the benchmark UASpeech dysarthric and DementiaBank Pitt elderly speech corpora suggest that regularized FL systems consistently outperform the baseline FedAvg system by statistically significant WER reductions of up to 0.55\% absolute (2.13\% relative). Further increasing communication frequency to one exchange per batch approaches centralized training performance.
标题: PMF-DEC:音素增强多模式融合,用于上下文感知的ASB错误纠正,并具有特定错误的选择性解码
链接:https://arxiv.org/abs/2506.11064
备注:Accepted by IEEE TASLP 2025
摘要:端到端自动语音识别(ASR)模型通常难以准确识别罕见单词。在此之前,我们介绍了一种称为错误检测和上下文感知纠错(ED-CEC)的ASR后处理方法,该方法利用命名实体和技术术语等上下文信息来提高ASR转录的准确性。虽然ED-CEC在纠正生僻字方面取得了显著的成功,但在处理发音相似但拼写不同的生僻字时,其准确率仍然很低。针对这一问题,本文提出了一种基于音素增强的多模态融合上下文感知纠错方法(PMF-CEC),该方法能够更好地区分目标稀有词和同音异义词。此外,我们观察到,以前的ASR错误检测模块遭受过度检测。为了缓解这一问题,我们引入了保留概率机制来过滤掉置信度低于设定阈值的编辑操作,保留原始操作以提高错误检测准确性。在5个数据集上进行的实验表明,与ED-CEC相比,PMF-CEC在保持合理的推理速度的同时,进一步降低了偏误率,在同音异义词的纠正方面表现出更强的优势。此外,我们的方法优于其他上下文偏置方法,并且与基于LLM的方法相比,在大偏置列表下,在更快的推理和更好的鲁棒性方面仍然很有价值。
摘要:End-to-end automatic speech recognition (ASR) models often struggle to accurately recognize rare words. Previously, we introduced an ASR postprocessing method called error detection and context-aware error correction (ED-CEC), which leverages contextual information such as named entities and technical terms to improve the accuracy of ASR transcripts. Although ED-CEC achieves a notable success in correcting rare words, its accuracy remains low when dealing with rare words that have similar pronunciations but different spellings. To address this issue, we proposed a phoneme-augmented multimodal fusion method for context-aware error correction (PMF-CEC) method on the basis of ED-CEC, which allowed for better differentiation between target rare words and homophones. Additionally, we observed that the previous ASR error detection module suffers from overdetection. To mitigate this, we introduced a retention probability mechanism to filter out editing operations with confidence scores below a set threshold, preserving the original operation to improve error detection accuracy. Experiments conducted on five datasets demonstrated that our proposed PMF-CEC maintains reasonable inference speed while further reducing the biased word error rate compared with ED-CEC, showing a stronger advantage in correcting homophones. Moreover, our method outperforms other contextual biasing methods, and remains valuable compared with LLM-based methods in terms of faster inference and better robustness under large biasing lists.
【1】 Tracking of Spatially Dynamic Room Impulse Responses Along Locally Linearized Trajectories
链接:https://arxiv.org/abs/2506.11703
备注:8 pages, 6 figures. Accepted paper for conference: Forum Acousticum Euronoise 2025 (fa-euronoise2025)
摘要:在多个空间点测量房间脉冲响应(RIR)是一项耗时的任务,而模拟需要详细了解房间的声学环境。在以前的工作中,我们提出了一种方法,用于估计早期部分的RIR沿线性轨迹在一个时变的声学场景,涉及一个静态声源和一个麦克风以恒定的速度移动。该方法依赖于在轨迹的起点和终点处测量的RIR,并且假设由直达声和沿着轨迹的各个反射占据的时间间隔是不重叠的。因此,该方法的适用性仅限于房间内相对较小的区域,其性能尚未通过真实数据进行验证。在本文中,我们提出了一个实际的扩展的方法,更现实的情况下,通过分割较长的轨迹到较小的线性区间的假设近似举行。沿着这些段分段应用该方法将其适用性扩展到更复杂的房间环境。我们证明了它的有效性,使用的mantoRIR数据库,其中包括移动麦克风录音和RIR测量在离散点沿一个控制的L形轨迹在一个真实的房间。
摘要:Measuring room impulse responses (RIRs) at multiple spatial points is a time-consuming task, while simulations require detailed knowledge of the room's acoustic environment. In prior work, we proposed a method for estimating the early part of RIRs along a linear trajectory in a time-varying acoustic scenario involving a static sound source and a microphone moving at constant velocity. This approach relies on measured RIRs at the start and end points of the trajectory and assumes that the time intervals occupied by the direct sound and individual reflections along the trajectory are non-overlapping. The method's applicability is therefore restricted to relatively small areas within a room, and its performance has yet to be validated with real-world data. In this paper, we propose a practical extension of the method to more realistic scenarios by segmenting longer trajectories into smaller linear intervals where the assumptions approximately hold. Applying the method piecewise along these segments extends its applicability to more complex room environments. We demonstrate its effectiveness using the trajectoRIR database, which includes moving microphone recordings and RIR measurements at discrete points along a controlled L-shaped trajectory in a real room.
【2】 Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
链接:https://arxiv.org/abs/2506.11630
备注:Interspeech 2025
摘要:本文介绍了SHTNet,一个轻量级的球谐变换(SHT)为基础的框架,旨在通过三个关键的创新,以解决跨阵列泛化的挑战,在多通道自动语音识别(ASR)。首先,基于SHT的空间声场分解将麦克风信号转换为几何不变的球谐系数,将信号处理与阵列几何隔离。其次,空间-频谱注意力融合网络(SSAFN)结合了坐标感知空间建模,精细的自注意力信道组合器和频谱噪声抑制,而无需传统的波束形成。第三,Rand-SHT训练通过随机信道选择和阵列几何重构来增强鲁棒性。该系统在异构阵列(例如,圆形、方形和双耳),计算量比传统的神经波束形成器少97.1\%。
摘要:This paper presents SHTNet, a lightweight spherical harmonic transform (SHT) based framework, which is designed to address cross-array generalization challenges in multi-channel automatic speech recognition (ASR) through three key innovations. First, SHT based spatial sound field decomposition converts microphone signals into geometry-invariant spherical harmonic coefficients, isolating signal processing from array geometry. Second, the Spatio-Spectral Attention Fusion Network (SSAFN) combines coordinate-aware spatial modeling, refined self-attention channel combinator, and spectral noise suppression without conventional beamforming. Third, Rand-SHT training enhances robustness through random channel selection and array geometry reconstruction. The system achieves 39.26\% average CER across heterogeneous arrays (e.g., circular, square, and binaural) on datasets including Aishell-4, Alimeeting, and XMOS, with 97.1\% fewer computations than conventional neural beamformers.
【3】 From Sharpness to Better Generalization for Speech Deepfake Detection
链接:https://arxiv.org/abs/2506.11532
备注:Accepted to Interspeech 2025
摘要:泛化仍然是语音深度伪造检测(SDD)的关键挑战。虽然各种方法的目的是提高鲁棒性,泛化通常是通过性能指标,如相等的错误率,没有一个理论框架来解释模型的性能进行评估。这项工作调查的清晰度作为一个理论代理泛化SDD。我们分析了锐度如何响应域的变化,并发现它在看不见的条件下增加,表明更高的模型灵敏度。在此基础上,我们应用锐度感知最小化(SAM)来显式降低锐度,从而在各种不可见的测试集上获得更好、更稳定的性能。此外,相关性分析证实了在大多数测试设置的锐度和泛化之间的统计上显着的关系。这些研究结果表明,清晰度可以作为SDD泛化的理论指标,清晰度感知训练提供了一个有前途的策略,以提高鲁棒性。
摘要:Generalization remains a critical challenge in speech deepfake detection (SDD). While various approaches aim to improve robustness, generalization is typically assessed through performance metrics like equal error rate without a theoretical framework to explain model performance. This work investigates sharpness as a theoretical proxy for generalization in SDD. We analyze how sharpness responds to domain shifts and find it increases in unseen conditions, indicating higher model sensitivity. Based on this, we apply Sharpness-Aware Minimization (SAM) to reduce sharpness explicitly, leading to better and more stable performance across diverse unseen test sets. Furthermore, correlation analysis confirms a statistically significant relationship between sharpness and generalization in most test settings. These findings suggest that sharpness can serve as a theoretical indicator for generalization in SDD and that sharpness-aware training offers a promising strategy for improving robustness.
【4】 Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
链接:https://arxiv.org/abs/2506.11514
备注:Accepted by Interspeech 2025
摘要:最近的研究深入研究了利用来自预训练模型的音频嵌入的语音增强(SE)方法,与时频掩蔽或信号预测技术不同。本文介绍了一种有效的和可扩展的SE方法。我们的方法涉及最初提取音频嵌入从嘈杂的语音使用预训练的音频编码器,然后由一个紧凑的编码器网络去噪。随后,声码器从去噪嵌入合成干净的语音。消融研究证实了预训练的音频编码器和声码器的降噪编码器的参数效率。语音增强和扬声器保真度的实验结果表明,我们的生成audiocoder为基础的SE系统优于模型利用歧视性audiocoder。此外,主观听力测试验证,我们提出的系统超越了现有的国家的最先进的SE模型的感知质量。
摘要:Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal prediction techniques. This paper introduces an efficient and extensible SE method. Our approach involves initially extracting audio embeddings from noisy speech using a pre-trained audioencoder, which are then denoised by a compact encoder network. Subsequently, a vocoder synthesizes the clean speech from denoised embeddings. An ablation study substantiates the parameter efficiency of the denoise encoder with a pre-trained audioencoder and vocoder. Experimental results on both speech enhancement and speaker fidelity demonstrate that our generative audioencoder-based SE system outperforms models utilizing discriminative audioencoders. Furthermore, subjective listening tests validate that our proposed system surpasses an existing state-of-the-art SE model in terms of perceptual quality.
【5】 Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms
链接:https://arxiv.org/abs/2506.11169
备注:61 pages, 21 figures
摘要:小占用空间关键字定位(SF-KWS)在当今智能语音激活设备、智能手机和物联网(IoT)应用领域中越来越受欢迎。这一激增归功于深度学习的进步,它能够从连续的单词流中识别预定义的单词或关键字。为了在实际场景中的低功耗和有限内存的边缘设备上实现SF-KWS模型,高效的Tiny Machine Learning(TinyML)框架至关重要。在这项研究中,我们探讨了七个不同类别的技术,即模型架构,学习技术,模型压缩,注意力感知架构,特征优化,神经网络搜索,和混合方法,这是适合开发一个SF-KWS系统。这一全面的概述将作为那些希望了解,利用或有助于SF-KWS领域的宝贵资源。在这项工作中进行的分析,使许多潜在的研究方向的识别,包括自动语音识别研究的见解和那些具体有关的领域的口语SF-KWS。
摘要:Small-Footprint Keyword Spotting (SF-KWS) has gained popularity in today's landscape of smart voice-activated devices, smartphones, and Internet of Things (IoT) applications. This surge is attributed to the advancements in Deep Learning, enabling the identification of predefined words or keywords from a continuous stream of words. To implement the SF-KWS model on edge devices with low power and limited memory in real-world scenarios, a efficient Tiny Machine Learning (TinyML) framework is essential. In this study, we explore seven distinct categories of techniques namely, Model Architecture, Learning Techniques, Model Compression, Attention Awareness Architecture, Feature Optimization, Neural Network Search, and Hybrid Approaches, which are suitable for developing an SF-KWS system. This comprehensive overview will serve as a valuable resource for those looking to understand, utilize, or contribute to the field of SF-KWS. The analysis conducted in this work enables the identification of numerous potential research directions, encompassing insights from automatic speech recognition research and those specifically pertinent to the realm of spoken SF-KWS.
【6】 S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamlessly Speech-Text Alignment and Streaming Speech Decoder
链接:https://arxiv.org/abs/2506.11160
备注:Working in progress
摘要:多语言语音到语音翻译(S2 ST)的目的是直接将多个源语言的口语转换为目标语言的自然和可理解的语音。尽管最近取得了进展,但仍存在重大挑战:(1)实现高质量和低延迟的S2 ST仍然是一个关键的障碍;(2)现有的S2 ST方法严重依赖于大规模的并行语音语料库,这是非常困难的收集。为了解决这些问题,我们提出了S2 ST-Omni,一个高效和可扩展的多语种语音到语音翻译框架。具体来说,我们将S2 ST任务分解为语音到文本翻译(S2 TT)和文本到语音合成(TTS),将它们统一在一个端到端的语音语言模型中。为了实现高质量的S2 TT,同时减少对并行语料库的依赖,我们利用大规模的预训练模型- Whisper用于音频理解,Qwen 3.0用于文本理解。一个轻量级的语音适配器被引入对齐语音和文本表示,使预训练的多模态知识的有效使用。为了保证翻译质量和实时性,我们在TTS阶段采用预训练的流式语音解码器以自回归方式生成目标语音。在CVSS基准测试上进行的大量实验表明,S2 ST-Omni的性能优于最先进的S2 ST基线,同时保持相当的延迟,突出了其在实际部署中的有效性和实用潜力。
摘要:Multilingual speech-to-speech translation (S2ST) aims to directly convert spoken utterances from multiple source languages into natural and intelligible speech in a target language. Despite recent progress, significant challenges remain: (1) achieving high-quality and low-latency S2ST remains a critical hurdle; (2) existing S2ST approaches heavily rely on large-scale parallel speech corpora, which are extremely difficult to collect. To address these issues, we propose S2ST-Omni, an efficient and scalable framework for multilingual speech-to-speech translation. Specifically, we decompose the S2ST task into speech-to-text translation (S2TT) and text-to-speech synthesis (TTS), unifying them within a single end-to-end speech-language model. To achieve high-quality S2TT while reducing dependence on parallel corpora, we leverage large-scale pretrained models -- Whisper for audio understanding and Qwen 3.0 for text understanding. A lightweight speech adapter is introduced to align speech and text representations, enabling effective use of pretrained multimodal knowledge. To ensure both translation quality and real-time performance, we adopt a pretrained streaming speech decoder in the TTS stage to generate target speech in an autoregressive manner. Extensive experiments on the CVSS benchmark demonstrate that S2ST-Omni outperforms state-of-the-art S2ST baselines while maintaining comparable latency, highlighting its effectiveness and practical potential for real-world deployment.
【7】 Improved in-car sound pick-up using multichannel Wiener filter
链接:https://arxiv.org/abs/2506.11157
备注:6 pages
摘要:随着汽车电子和传感器的进步,使用多个麦克风的声音拾取对于免提电话和语音命令车载应用已经变得可行。然而,由于带宽或处理限制,在有效地处理多个麦克风信号方面仍然存在挑战。这项工作探讨了使用多通道维纳滤波器算法与双麦克风车载系统,以提高语音质量的司机和乘客的声音,即,以减轻由回波引起的陷波滤波效应并改善背景噪声降低。我们使用现代客观指标(如深度噪声抑制平均意见得分)评估其在各种噪声条件下的性能。还研究了驾驶员/乘客头部运动的影响。所提出的方法示出提供显着的改进,在一个简单的混合麦克风信号。
摘要:With advancements in automotive electronics and sensors, the sound pick-up using multiple microphones has become feasible for hands-free telephony and voice command in-car applications. However, challenges remain in effectively processing multiple microphone signals due to bandwidth or processing limitations. This work explores the use of the Multichannel Wiener Filter algorithm with a two-microphone in-car system, to enhance speech quality for driver and passenger voice, i.e., to mitigate notch-filtering effects caused by echoes and improve background noise reduction. We evaluate its performance under various noise conditions using modern objective metrics like Deep Noise Suppression Mean Opinion Score. The effect of head movements of driver/passenger is also investigated. The proposed method is shown to provide significant improvements over a simple mixing of microphone signals.
【8】 Tracking of Intermittent and Moving Speakers : Dataset and Metrics
链接:https://arxiv.org/abs/2506.11145
备注:None
摘要:本文提出的问题跟踪间歇和移动源,即可能改变位置时,他们是不活跃的来源。这个问题很少被探讨,目前大多数跟踪方法依赖于空间观测的跟踪身份管理。它们或者基于先前的定位步骤,或者被设计为通过预测有序的位置估计来执行联合定位和跟踪。这引起了对这样的方法是否可以在处理不连续的空间轨道时保持可靠的轨道标识分配性能的关注,所述不连续的空间轨道可能由静默期间的方向改变引起。我们介绍LibriJump,一个新的数据集的声学场景的一阶高保真度立体声格式,专注于扬声器跟踪。该数据集包含在不活动期间具有变化位置的扬声器,从而模拟不连续的音轨。为了测量身份分配性能,我们建议使用从计算机视觉社区改编的跟踪关联度量。我们提供的实验显示的互补性的关联度量与以前使用的跟踪度量,连续和不连续的空间轨道。
摘要:This paper presents the problem of tracking intermittent and moving sources, i.e, sources that may change position when they are inactive. This issue is seldom explored, and most current tracking methods rely on spatial observations for track identity management. They are either based on a previous localization step, or designed to perform joint localization and tracking by predicting ordered position estimates. This raises concerns about whether such methods can maintain reliable track identity assignment performance when dealing with discontinuous spatial tracks, which may be caused by a change of direction during silence. We introduce LibriJump, a novel dataset of acoustic scenes in the First Order Ambisonics format focusing on speaker tracking. The dataset contains speakers with changing positions during inactivity periods, thus simulating discontinuous tracks. To measure the identity assignment performance, we propose to use tracking association metrics adapted from the computer vision community. We provide experiments showing the complementarity of association metrics with previously used tracking metrics, given continuous and discontinuous spatial tracks.
【9】 Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
链接:https://arxiv.org/abs/2506.11089
摘要:自动语音识别(ASR)模型依赖于高质量的转录数据进行有效的训练。为大型未标记音频数据集生成伪标签通常依赖于复杂的流水线,这些流水线通过多级处理组合多个ASR输出,导致错误传播、信息丢失和不相交优化。我们提出了一个统一的多ASR语音驱动的框架,使用基于文本或语音的大型语言模型(LLM)进行后处理,取代投票或其他仲裁逻辑,以协调集成输出。我们进行了比较研究的多个架构,有和没有LLM,转录准确性显着改善相比,传统的方法。此外,我们使用各种方法生成的伪标签来训练不同数据集的半监督ASR模型,与基线相比,文本和语音LLM转录的性能再次得到提高。
摘要:Automatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on complex pipelines that combine multiple ASR outputs through multi-stage processing, leading to error propagation, information loss and disjoint optimization. We propose a unified multi-ASR prompt-driven framework using postprocessing by either textual or speech-based large language models (LLMs), replacing voting or other arbitration logic for reconciling the ensemble outputs. We perform a comparative study of multiple architectures with and without LLMs, showing significant improvements in transcription accuracy compared to traditional methods. Furthermore, we use the pseudo-labels generated by the various approaches to train semi-supervised ASR models for different datasets, again showing improved performance with textual and speechLLM transcriptions compared to baselines.
【10】 Intelligibility of Text-to-Speech Systems for Mathematical Expressions
链接:https://arxiv.org/abs/2506.11086
备注:Accepted at Interspeech 2025
摘要:已经有有限的评估先进的文本到语音(TTS)模型与数学example(MX)作为输入。在这项工作中,我们设计的实验,以评估质量和可懂度的五个TTS模型,通过听力和转录测试的各种类别的MX。由于TTS模型不能直接处理LaTeX,我们使用两个大语言模型(LLM)从LaTeX MX生成英语发音。我们使用来自用户评级的平均意见得分,并使用三个指标通过转录正确性来量化可理解性。我们还比较了听众偏好的TTS输出与人类专家再现相同的MX。结果表明,TTS模型的输出MX不一定是可理解的,在可理解性的差距不同的TTS模型和MX类别。对于大多数类别,TTS模型的性能显着不如专家再现。选择LLM的效果有限。这就需要为MX改进TTS模型。
摘要:There has been limited evaluation of advanced Text-to-Speech (TTS) models with Mathematical eXpressions (MX) as inputs. In this work, we design experiments to evaluate quality and intelligibility of five TTS models through listening and transcribing tests for various categories of MX. We use two Large Language Models (LLMs) to generate English pronunciation from LaTeX MX as TTS models cannot process LaTeX directly. We use Mean Opinion Score from user ratings and quantify intelligibility through transcription correctness using three metrics. We also compare listener preference of TTS outputs with respect to human expert rendition of same MX. Results establish that output of TTS models for MX is not necessarily intelligible, the gap in intelligibility varies across TTS models and MX category. For most categories, performance of TTS models is significantly worse than that of expert rendition. The effect of choice of LLM is limited. This establishes the need to improve TTS models for MX.
【11】 Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts
链接:https://arxiv.org/abs/2506.11079
备注:This paper is accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:自动朗读评估可以为教师提供有价值的支持,使阅读练习更有效的评分。然而,阅读评价系统和应用的研究仍然有限。我们提出了一种新的多模态方法,利用音频和知识的文本资源。特别是,我们探索了使用Whisper和提示调整的大型语言模型(LLM)的潜力,以改善儿童语音识别的transmittance,以及它们在下游阅读错误检测中的有效性。我们的研究结果表明,与没有提示的基线Whisper模型相比,提示Whisper和提示LLM的有效性。表现最好的系统在荷兰儿童阅读语音中实现了最先进的识别性能,单词错误率(WER)为5.1%,提高了9.4%的基线WER。此外,它显着提高了阅读错误检测,将F1分数从0.39提高到0.73。
摘要:Automatic reading aloud evaluation can provide valuable support to teachers by enabling more efficient scoring of reading exercises. However, research on reading evaluation systems and applications remains limited. We present a novel multimodal approach that leverages audio and knowledge from text resources. In particular, we explored the potential of using Whisper and instruction-tuned large language models (LLMs) with prompts to improve transcriptions for child speech recognition, as well as their effectiveness in downstream reading mistake detection. Our results demonstrate the effectiveness of prompting Whisper and prompting LLM, compared to the baseline Whisper model without prompting. The best performing system achieved state-of-the-art recognition performance in Dutch child read speech, with a word error rate (WER) of 5.1%, improving the baseline WER of 9.4%. Furthermore, it significantly improved reading mistake detection, increasing the F1 score from 0.39 to 0.73.
【12】 Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
链接:https://arxiv.org/abs/2506.11075
备注:5 pages, 3 figures
摘要:使用儿童佩戴的设备收集的音频记录是儿童语言研究的基本工具。全天收集的长形式录音承诺以最小的观察者偏差捕获儿童的输入和生产,因此具有较高的有效性。所产生的数据量之大,需要进行自动化分析,为研究人员和临床医生提取相关指标。本文总结了关于此技术的集体知识,为现有资源提供了切入点。我们还强调了威胁自动化注释的准确性和对结果指标的解释的各种错误来源。为了解决这个问题,我们提出了潜在的故障排除指标,以帮助用户评估数据质量。虽然完全自动化的质量控制系统是不可行的,但我们概述了研究人员改进数据收集和分析的实用策略。
摘要:Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and therefore high validity. The sheer volume of resulting data necessitates automated analysis to extract relevant metrics for researchers and clinicians. This paper summarizes collective knowledge on this technique, providing entry points to existing resources. We also highlight various sources of error that threaten the accuracy of automated annotations and the interpretation of resulting metrics. To address this, we propose potential troubleshooting metrics to help users assess data quality. While a fully automated quality control system is not feasible, we outline practical strategies for researchers to improve data collection and contextualize their analyses.
【13】 Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
链接:https://arxiv.org/abs/2506.11074
备注:5 pages, 3 figures
摘要:用儿童佩戴的设备收集的录音有望通过毫不费力地捕捉儿童的自然语言环境和语言产生来彻底改变基础和应用语言科学。这一前景取决于语音技术,它可以将如此收集到的大量数据转化为可用的信息。本文通过总结三年来旨在改进一项基本任务:语音类型分类的实验,展示了阻碍进展的几个障碍。我们的实验表明,改进表示功能,架构和参数搜索有助于只有边际收益的性能。通过关注数据的相关性和数量,取得了更多进展,这突出了收集具有适当权限的数据以允许共享的重要性。
摘要:Recordings gathered with child-worn devices promised to revolutionize both fundamental and applied speech sciences by allowing the effortless capture of children's naturalistic speech environment and language production. This promise hinges on speech technologies that can transform the sheer mounds of data thus collected into usable information. This paper demonstrates several obstacles blocking progress by summarizing three years' worth of experiments aimed at improving one fundamental task: Voice Type Classification. Our experiments suggest that improvements in representation features, architecture, and parameter search contribute to only marginal gains in performance. More progress is made by focusing on data relevance and quantity, which highlights the importance of collecting data with appropriate permissions to allow sharing.
【14】 Can We Trust Machine Learning? The Reliability of Features from Open-Source Speech Analysis Tools for Speech Modeling
链接:https://arxiv.org/abs/2506.11072
备注:5 pages, 1 figure, 3 tables
摘要:基于机器学习的行为模型依赖于从视听记录中提取的特征。使用开源工具处理录音,以提取分类模型的语音特征。这些工具往往缺乏验证,以确保在捕获行为相关信息的可靠性。这一差距引起了对不同人群和背景下的可重复性和公平性的关注。语音处理工具在其设计环境之外使用时,可能无法公平地捕捉行为变化,从而导致偏见。我们评估了从两个广泛使用的语音分析工具,OpenSMILE和Praat提取的语音特征,以评估其可靠性时,考虑青少年自闭症。我们观察到不同工具的功能存在相当大的差异,这影响了不同背景和人口统计学群体的模型性能。我们鼓励领域相关验证,以提高机器学习模型在临床应用中的可靠性。
摘要:Machine learning-based behavioral models rely on features extracted from audio-visual recordings. The recordings are processed using open-source tools to extract speech features for classification models. These tools often lack validation to ensure reliability in capturing behaviorally relevant information. This gap raises concerns about reproducibility and fairness across diverse populations and contexts. Speech processing tools, when used outside of their design context, can fail to capture behavioral variations equitably and can then contribute to bias. We evaluate speech features extracted from two widely used speech analysis tools, OpenSMILE and Praat, to assess their reliability when considering adolescents with autism. We observed considerable variation in features across tools, which influenced model performance across context and demographic groups. We encourage domain-relevant verification to enhance the reliability of machine learning models in clinical applications.
【15】 Embedded Acoustic Intelligence for Automotive Systems
链接:https://arxiv.org/abs/2506.11071
摘要:将声音洞察转化为可操作的数据流,这篇摘要利用学位论文研究的发现来增强汽车系统智能,使我们能够解决道路类型[1]。通过提取和解释安装在汽车轴距内的麦克风的声学特征,我们专注于对道路类型进行分类。利用深度神经网络和由开放AI生态系统的预训练模型提供支持的特征提取,(通过拥抱脸[2]),我们的方法使自动驾驶和高级驾驶员辅助系统(AD/ADAS)能够预测路面,支持主动道路噪音消除的自适应学习,并为城市规划提供有价值的见解。本研究的结果专门用于支持下一代汽车系统的引人注目的商业案例。这种前瞻性的方法不仅有望重新定义乘客的舒适性和提高车辆的安全性,而且还为智能化、数据驱动的城市道路管理铺平了道路,使未来的移动性既可实现又可持续。
摘要:Transforming sound insights into actionable streams of data, this abstract leverages findings from degree thesis research to enhance automotive system intelligence, enabling us to address road type [1].By extracting and interpreting acoustic signatures from microphones installed within the wheelbase of a car, we focus on classifying road type.Utilizing deep neural networks and feature extraction powered by pre-trained models from the Open AI ecosystem (via Hugging Face [2]), our approach enables Autonomous Driving and Advanced Driver- Assistance Systems (AD/ADAS) to anticipate road surfaces, support adaptive learning for active road noise cancellation, and generate valuable insights for urban planning. The results of this study were specifically captured to support a compelling business case for next-generation automotive systems. This forward-looking approach not only promises to redefine passenger comfort and improve vehicle safety, but also paves the way for intelligent, data-driven urban road management, making the future of mobility both achievable and sustainable.
【16】 Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
链接:https://arxiv.org/abs/2506.11069
摘要:迄今为止,构音障碍和老年言语的准确识别仍然具有挑战性。虽然隐私问题已经推动了从集中式方法向联邦学习(FL)的转变,以确保数据的机密性,但这进一步加剧了数据稀缺,数据分布不平衡和说话人异质性的挑战。为此,本文进行了一个系统的调查,隐私保护构音障碍和老年人语音识别的正则化FL技术,解决不同层次的FL过程1)参数为基础,2)嵌入为基础,3)新的损失为基础的正则化。在基准UASpeech构音障碍和DementiaBank Pitt老年人语音语料库上的实验表明,正则化FL系统的性能始终优于基线FedAvg系统,其统计学显著WER降低高达0.55%绝对值(2.13%相对值)。进一步将通信频率增加到每批一次交换接近集中训练性能。
摘要:Accurate recognition of dysarthric and elderly speech remains challenging to date. While privacy concerns have driven a shift from centralized approaches to federated learning (FL) to ensure data confidentiality, this further exacerbates the challenges of data scarcity, imbalanced data distribution and speaker heterogeneity. To this end, this paper conducts a systematic investigation of regularized FL techniques for privacy-preserving dysarthric and elderly speech recognition, addressing different levels of the FL process by 1) parameter-based, 2) embedding-based and 3) novel loss-based regularization. Experiments on the benchmark UASpeech dysarthric and DementiaBank Pitt elderly speech corpora suggest that regularized FL systems consistently outperform the baseline FedAvg system by statistically significant WER reductions of up to 0.55\% absolute (2.13\% relative). Further increasing communication frequency to one exchange per batch approaches centralized training performance.
【17】 PMF-CEC: Phoneme-augmented Multimodal Fusion for Context-aware ASR Error Correction with Error-specific Selective Decoding
链接:https://arxiv.org/abs/2506.11064
备注:Accepted by IEEE TASLP 2025
摘要:端到端自动语音识别(ASR)模型通常难以准确识别罕见单词。在此之前,我们介绍了一种称为错误检测和上下文感知纠错(ED-CEC)的ASR后处理方法,该方法利用命名实体和技术术语等上下文信息来提高ASR转录的准确性。虽然ED-CEC在纠正生僻字方面取得了显著的成功,但在处理发音相似但拼写不同的生僻字时,其准确率仍然很低。针对这一问题,提出了一种音素增强的多模态融合上下文感知纠错方法(PMF-CEC),该方法基于ED-CEC,能够更好地区分目标稀有词和同音异义词。此外,我们观察到,以前的ASR错误检测模块遭受过度检测。为了缓解这一问题,我们引入了保留概率机制来过滤掉置信度低于设定阈值的编辑操作,保留原始操作以提高错误检测准确性。在5个数据集上进行的实验表明,与ED-CEC相比,PMF-CEC在保持合理的推理速度的同时,进一步降低了偏误率,在同音异义词的纠正方面表现出更强的优势。此外,我们的方法优于其他上下文偏置方法,并且与基于LLM的方法相比,在大偏置列表下,在更快的推理和更好的鲁棒性方面仍然很有价值。
摘要:End-to-end automatic speech recognition (ASR) models often struggle to accurately recognize rare words. Previously, we introduced an ASR postprocessing method called error detection and context-aware error correction (ED-CEC), which leverages contextual information such as named entities and technical terms to improve the accuracy of ASR transcripts. Although ED-CEC achieves a notable success in correcting rare words, its accuracy remains low when dealing with rare words that have similar pronunciations but different spellings. To address this issue, we proposed a phoneme-augmented multimodal fusion method for context-aware error correction (PMF-CEC) method on the basis of ED-CEC, which allowed for better differentiation between target rare words and homophones. Additionally, we observed that the previous ASR error detection module suffers from overdetection. To mitigate this, we introduced a retention probability mechanism to filter out editing operations with confidence scores below a set threshold, preserving the original operation to improve error detection accuracy. Experiments conducted on five datasets demonstrated that our proposed PMF-CEC maintains reasonable inference speed while further reducing the biased word error rate compared with ED-CEC, showing a stronger advantage in correcting homophones. Moreover, our method outperforms other contextual biasing methods, and remains valuable compared with LLM-based methods in terms of faster inference and better robustness under large biasing lists.
【18】 Reimagining Dance: Real-time Music Co-creation between Dancers and AI
链接:https://arxiv.org/abs/2506.12008
备注:Accepted for publication at ICCC 2025 (International Conference on Computational Creativity)
摘要:舞蹈表演传统上遵循一种单向的关系,即运动对音乐的反应。虽然人工智能在各个创意领域都取得了进展,但它在舞蹈中的应用主要集中在从音乐输入中生成编舞。我们提出了一个系统,使舞者通过他们的动作动态地塑造音乐环境。我们的多模态架构通过智能地结合预先录制的音乐片段来响应舞蹈动作,从而创建连贯的音乐作品,建立了一种双向的创造性合作伙伴关系,舞者既是表演者又是作曲家。通过相关性分析的性能数据,我们展示了新兴的通信模式之间的运动质量和音频功能。这种方法将人工智能在表演艺术中的作用重新定义为一个响应式的合作者,为更广泛的人群提供专业舞蹈表演和即兴艺术表达的可能性。
摘要:Dance performance traditionally follows a unidirectional relationship where movement responds to music. While AI has advanced in various creative domains, its application in dance has primarily focused on generating choreography from musical input. We present a system that enables dancers to dynamically shape musical environments through their movements. Our multi-modal architecture creates a coherent musical composition by intelligently combining pre-recorded musical clips in response to dance movements, establishing a bidirectional creative partnership where dancers function as both performers and composers. Through correlation analysis of performance data, we demonstrate emergent communication patterns between movement qualities and audio features. This approach reconceptualizes the role of AI in performing arts as a responsive collaborator that expands possibilities for both professional dance performance and improvisational artistic expression across broader populations.
【19】 Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
链接:https://arxiv.org/abs/2506.11862
摘要:语音肌电图(EMG)到语音(V-ETS)模型从肌肉活动信号重建语音,促进神经喉诊断等应用。尽管V-ETS具有潜力,但由于缺乏成对的EMG-语音数据,V-ETS的发展受到阻碍。为了解决这个问题,我们提出了一种新的基于置信度的多说话者自训练(CoM 2S)方法,以及一个新策划的Libri-EMG数据集。该方法利用由预训练模型生成的合成EMG数据,然后基于音素级置信度提出过滤机制,通过提出的自我训练技术来增强ETS模型。实验结果表明,该方法提高了音素准确率,减少了语音混乱,降低了单词错误率,证实了我们的CoM 2S方法的V-ETS的有效性。为了支持未来的研究,我们将发布代码和建议Libri-EMG数据库-一个开放访问的,时间对齐的,多扬声器语音EMG和语音记录。
摘要:Voiced Electromyography (EMG)-to-Speech (V-ETS) models reconstruct speech from muscle activity signals, facilitating applications such as neurolaryngologic diagnostics. Despite its potential, the advancement of V-ETS is hindered by a scarcity of paired EMG-speech data. To address this, we propose a novel Confidence-based Multi-Speaker Self-training (CoM2S) approach, along with a newly curated Libri-EMG dataset. This approach leverages synthetic EMG data generated by a pre-trained model, followed by a proposed filtering mechanism based on phoneme-level confidence to enhance the ETS model through the proposed self-training techniques. Experiments demonstrate our method improves phoneme accuracy, reduces phonological confusion, and lowers word error rate, confirming the effectiveness of our CoM2S approach for V-ETS. In support of future research, we will release the codes and the proposed Libri-EMG dataset-an open-access, time-aligned, multi-speaker voiced EMG and speech recordings.
【20】 Abstract Sound Fusion with Unconditioned Inversion Model
链接:https://arxiv.org/abs/2506.11811
摘要:抽象声音被定义为不向收听者公开可识别的真实世界声音事件的声音。声音融合的目的是合成一个原始声音和一个参考声音,以产生一个新的声音,表现出的听觉特征,而不仅仅是添加剂叠加的声音成分。为了实现这种融合,我们采用反演技术,保留原始样品的基本特征,同时使可控合成。我们提出了基于DPMSolver++采样器的新的非线性和常微分方程反演模型,该采样器通过将模型输出配置为常数来反转采样过程,消除了噪声预测项引起的循环依赖性。我们的反演方法不需要及时调节,同时在采样过程中保持灵活的指导。
摘要:An abstract sound is defined as a sound that does not disclose identifiable real-world sound events to a listener. Sound fusion aims to synthesize an original sound and a reference sound to generate a novel sound that exhibits auditory features beyond mere additive superposition of the sound constituents. To achieve this fusion, we employ inversion techniques that preserve essential features of the original sample while enabling controllable synthesis. We propose novel SDE and ODE inversion models based on DPMSolver++ samplers that reverse the sampling process by configuring model outputs as constants, eliminating circular dependencies incurred by noise prediction terms. Our inversion approach requires no prompt conditioning while maintaining flexible guidance during sampling.
【21】 (SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
链接:https://arxiv.org/abs/2506.11620
摘要:传统的测听通常不能完全表征听力损失对言语理解的功能影响,特别是对于老年性耳聋常见的阈上缺陷。这促使了更具体的诊断言语感知测试的发展。我们介绍了模拟音素语音测试(SimPhon语音测试)的方法,一种新颖的,多阶段的计算流水线在硅片上的设计和验证的语音平衡的最小对语音测试。该方法利用现代自动语音识别(ASR)系统作为人类听众的代理来模拟感音神经性听力损失的感知效果。通过在受控的声学退化下处理语音刺激,我们首先识别最常见的音素混淆模式。然后,这些模式指导从综合语言语料库中导出的大量候选词对的数据驱动的策展。随后的阶段包括模拟诊断测试,专家人工策展和最终的有针对性的敏感性分析,系统地将候选人减少到最终的优化的25对(SimPhon Speech Test-25)。一个关键的发现是,SimPhon Speech Test-25测试项目的诊断性能与标准语音可懂度指数(SII)的预测没有显着相关性,这表明SimPhon Speech Test捕获了简单可听度之外的感知缺陷。这种经过计算优化的测试集显着提高了听力测试开发的效率,为初步人体试验做好了准备。
摘要:Traditional audiometry often provides an incomplete characterization of the functional impact of hearing loss on speech understanding, particularly for supra-threshold deficits common in presbycusis. This motivates the development of more diagnostically specific speech perception tests. We introduce the Simulated Phoneme Speech Test (SimPhon Speech Test) methodology, a novel, multi-stage computational pipeline for the in silico design and validation of a phonetically balanced minimal-pair speech test. This methodology leverages a modern Automatic Speech Recognition (ASR) system as a proxy for a human listener to simulate the perceptual effects of sensorineural hearing loss. By processing speech stimuli under controlled acoustic degradation, we first identify the most common phoneme confusion patterns. These patterns then guide the data-driven curation of a large set of candidate word pairs derived from a comprehensive linguistic corpus. Subsequent phases involving simulated diagnostic testing, expert human curation, and a final, targeted sensitivity analysis systematically reduce the candidates to a final, optimized set of 25 pairs (the SimPhon Speech Test-25). A key finding is that the diagnostic performance of the SimPhon Speech Test-25 test items shows no significant correlation with predictions from the standard Speech Intelligibility Index (SII), suggesting the SimPhon Speech Test captures perceptual deficits beyond simple audibility. This computationally optimized test set offers a significant increase in efficiency for audiological test development, ready for initial human trials.
【22】 Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
链接:https://arxiv.org/abs/2506.11605
备注:37 pages, 18 figures. Submitted to Computer Speech & Language
摘要:带向量聚类的端到端神经日志化是一种强大而实用的说话人日志化方法。针对这些管道的分割模型提出了多项增强措施,但尚未对其协同作用进行彻底评估。在这项工作中,我们提供了一个深入的分析主要的架构选择对管道的性能的影响。我们研究了不同的编码器(SincNet,预训练和微调的WavLM),不同的解码器(LSTM,Mamba和Conformer),不同的损失(多标签和多类幂集)以及不同的块大小。通过覆盖九个数据集的深入实验,我们发现基于WavLM的微调编码器总是产生最好的系统。LSTM解码器被基于Mamba和Conformer的解码器所超越,虽然我们发现Mamba对其他架构选择更强大,但它略逊于我们使用Conformer编码器的最佳架构。我们发现,多标签和多类功率集损失不具有相同的错误分布。我们证实,多类损失有助于几乎所有模型获得卓越的性能,除了微调WavLM时,在这种情况下,多标签是更好的选择。我们还评估了块大小对所有上述架构选择的影响,发现较新的架构往往能更好地处理长块大小,这可以大大提高管道性能。我们最好的系统在五个广泛使用的说话人日记数据集上取得了最先进的结果。
摘要:End-to-End Neural Diarization with Vector Clustering is a powerful and practical approach to perform Speaker Diarization. Multiple enhancements have been proposed for the segmentation model of these pipelines, but their synergy had not been thoroughly evaluated. In this work, we provide an in-depth analysis on the impact of major architecture choices on the performance of the pipeline. We investigate different encoders (SincNet, pretrained and finetuned WavLM), different decoders (LSTM, Mamba, and Conformer), different losses (multilabel and multiclass powerset), and different chunk sizes. Through in-depth experiments covering nine datasets, we found that the finetuned WavLM-based encoder always results in the best systems by a wide margin. The LSTM decoder is outclassed by Mamba- and Conformer-based decoders, and while we found Mamba more robust to other architecture choices, it is slightly inferior to our best architecture, which uses a Conformer encoder. We found that multilabel and multiclass powerset losses do not have the same distribution of errors. We confirmed that the multiclass loss helps almost all models attain superior performance, except when finetuning WavLM, in which case, multilabel is the superior choice. We also evaluated the impact of the chunk size on all aforementioned architecture choices and found that newer architectures tend to better handle long chunk sizes, which can greatly improve pipeline performance. Our best system achieved state-of-the-art results on five widely used speaker diarization datasets.
【23】 Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
链接:https://arxiv.org/abs/2506.11542
备注:Accepted to Interspeech2025
摘要:欺骗话语总是包含生成模型引入的伪像。虽然已经提出了几种对策来检测欺骗性话语,但大多数主要集中在架构改进上。在这项工作中,我们研究如何文物仍然隐藏在欺骗性的讲话,以及如何提高他们的存在。我们提出了一个模型不可知的管道,使用语音增强和各种类型的噪声放大文物。我们的方法包括三个关键步骤:噪声添加、噪声提取和噪声放大。首先,我们在原始语音中引入噪声。然后,我们应用语音增强来提取纠缠噪声和伪影。最后,我们放大这些提取的特征。此外,我们的管道是兼容不同的语音增强模型和对策架构。我们的方法在ASVspoof2019上将欺骗检测性能提高了44.44\%,在ASVspoof2021上提高了26.34\%。
摘要:Spoofed utterances always contain artifacts introduced by generative models. While several countermeasures have been proposed to detect spoofed utterances, most primarily focus on architectural improvements. In this work, we investigate how artifacts remain hidden in spoofed speech and how to enhance their presence. We propose a model-agnostic pipeline that amplifies artifacts using speech enhancement and various types of noise. Our approach consists of three key steps: noise addition, noise extraction, and noise amplification. First, we introduce noise into the raw speech. Then, we apply speech enhancement to extract the entangled noise and artifacts. Finally, we amplify these extracted features. Moreover, our pipeline is compatible with different speech enhancement models and countermeasure architectures. Our method improves spoof detection performance by up to 44.44\% on ASVspoof2019 and 26.34\% on ASVspoof2021.
【24】 LiLAC: A Lightweight Latent ControlNet for Musical Audio Generation
链接:https://arxiv.org/abs/2506.11476
备注:Accepted at ISMIR 2025
摘要:文本到音频扩散模型产生高质量和多样化的音乐,但许多(如果不是大多数的话)SOTA模型缺乏音乐制作所必需的细粒度、时变控制。ControlNet通过克隆和微调新条件下的编码器,将外部控件附加到预训练的生成模型。但是,这种方法会占用大量内存,并将用户限制在一组固定的控件中。我们提出了一个轻量级的,模块化的架构,大大减少了参数计数,同时匹配ControlNet的音频质量和条件的遵守。我们的方法提供了更大的灵活性和显着降低内存使用,使独立控件的更有效的培训和部署。我们进行了广泛的客观和主观评估,并在附带的网站https://lightlatentcontrol.github.io上提供了大量的音频示例
摘要:Text-to-audio diffusion models produce high-quality and diverse music but many, if not most, of the SOTA models lack the fine-grained, time-varying controls essential for music production. ControlNet enables attaching external controls to a pre-trained generative model by cloning and fine-tuning its encoder on new conditionings. However, this approach incurs a large memory footprint and restricts users to a fixed set of controls. We propose a lightweight, modular architecture that considerably reduces parameter count while matching ControlNet in audio quality and condition adherence. Our method offers greater flexibility and significantly lower memory usage, enabling more efficient training and deployment of independent controls. We conduct extensive objective and subjective evaluations and provide numerous audio examples on the accompanying website at https://lightlatentcontrol.github.io
【25】 A correlation-permutation approach for speech-music encoders model merging
链接:https://arxiv.org/abs/2506.11403
备注:Under review
摘要:创建统一的语音和音乐模型需要昂贵的预训练。模型合并可以替代地创建具有最小计算开销的统一音频模型。然而,当模型在权重空间中不对齐时,直接合并具有挑战性。受Git Re-Basin的启发,我们引入了一种相关置换方法,该方法将音乐编码器的内部层与语音编码器对齐。我们将以前的工作扩展到合并Transformer层的情况。该方法计算一个置换矩阵,使模型的逐层特征互相关性最大化,从而实现这些不相交模型的有效融合。通过这种方法,合并后的模型保留了语音功能,同时显著提高了音乐性能,与线性插值模型合并相比,平均得分提高了14.83分。这项工作允许从独立训练的编码器创建统一的音频模型。
摘要:Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, direct merging is challenging when the models are not aligned in the weight space. Motivated by Git Re-Basin, we introduce a correlation-permutation approach that aligns a music encoder's internal layers with a speech encoder. We extend previous work to the case of merging transformer layers. The method computes a permutation matrix that maximizes the model's features-wise cross-correlations layer by layer, enabling effective fusion of these otherwise disjoint models. The merged model retains speech capabilities through this method while significantly enhancing music performance, achieving an improvement of 14.83 points in average score compared to linear interpolation model merging. This work allows the creation of unified audio models from independently trained encoders.
【26】 GLAP: General contrastive audio-text pretraining across domains and languages
链接:https://arxiv.org/abs/2506.11350
摘要:对比语言音频预训练(CLAP)是一种广泛使用的方法,以弥合音频和文本域之间的差距。目前的CLAP方法支持英语的声音和音乐检索,忽略了多语言的口语内容。为了解决这个问题,我们引入了通用语言音频预训练(Gestival),它扩展了CLAP的多语言和多领域能力。通过在Clotho和AudioCaps等标准音频文本检索基准上实现具有竞争力的性能,Gestro展示了其多功能性,同时在语音检索和分类任务中显着超越了现有方法。此外,Gestro在广泛使用的声音事件zero-shot基准上取得了很好的结果,同时在语音内容基准上优于以前的方法。对50种语言的进一步关键字识别评估强调了Gandroid先进的多语言功能。最后,在四种语言中评估多语言声音和音乐理解。检查点和资料来源:https://github.com/xiaomi-research/dasheng-glap。
摘要:Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in English, ignoring multilingual spoken content. To address this, we introduce general language audio pretraining (GLAP), which expands CLAP with multilingual and multi-domain abilities. GLAP demonstrates its versatility by achieving competitive performance on standard audio-text retrieval benchmarks like Clotho and AudioCaps, while significantly surpassing existing methods in speech retrieval and classification tasks. Additionally, GLAP achieves strong results on widely used sound-event zero-shot benchmarks, while simultaneously outperforming previous methods on speech content benchmarks. Further keyword spotting evaluations across 50 languages emphasize GLAP's advanced multilingual capabilities. Finally, multilingual sound and music understanding is evaluated across four languages. Checkpoints and Source: https://github.com/xiaomi-research/dasheng-glap.
【27】 MUDAS: Mote-scale Unsupervised Domain Adaptation in Multi-label Sound Classification
链接:https://arxiv.org/abs/2506.11331
摘要:无监督域自适应(UDA)对于使机器学习模型适应新的、无标签的环境至关重要,在这些环境中,数据分布的变化会降低性能。现有的UDA算法是为单标签任务设计的,依赖于大量的计算资源,限制了它们在多标签场景和资源受限的物联网设备中的使用。在城市声音分类等环境中,克服这些限制尤其具有挑战性,其中重叠的声音和变化的声学需要低功耗设备上系统的鲁棒的自适应多标签功能。为了解决这些限制,我们引入了声音的无监督域自适应(MUDAS),这是一个为资源受限的物联网环境中的多标签声音分类而开发的UDA框架。MUDAS通过使用高置信度数据选择性地原位重新训练分类器来有效地适应模型,最大限度地减少计算和内存需求,以适应设备上的部署。此外,MUDAS采用了类特定的自适应阈值来生成可靠的伪标签,并应用多样性正则化来提高多标签分类的准确性。在对纽约市多个地点记录的SONYC Urban Sound Tagging(SONYC-UST)数据集进行评估时,MUDAS在现有UDA算法的分类准确性方面取得了显着改进,在资源受限的物联网环境中实现了良好的性能。
摘要:Unsupervised Domain Adaptation (UDA) is essential for adapting machine learning models to new, unlabeled environments where data distribution shifts can degrade performance. Existing UDA algorithms are designed for single-label tasks and rely on significant computational resources, limiting their use in multi-label scenarios and in resource-constrained IoT devices. Overcoming these limitations is particularly challenging in contexts such as urban sound classification, where overlapping sounds and varying acoustics require robust, adaptive multi-label capabilities on low-power, on-device systems. To address these limitations, we introduce Mote-scale Unsupervised Domain Adaptation for Sounds (MUDAS), a UDA framework developed for multi-label sound classification in resource-constrained IoT settings. MUDAS efficiently adapts models by selectively retraining the classifier in situ using high-confidence data, minimizing computational and memory requirements to suit on-device deployment. Additionally, MUDAS incorporates class-specific adaptive thresholds to generate reliable pseudo-labels and applies diversity regularization to improve multi-label classification accuracy. In evaluations on the SONYC Urban Sound Tagging (SONYC-UST) dataset recorded at various New York City locations, MUDAS demonstrates notable improvements in classification accuracy over existing UDA algorithms, achieving good performance in a resource-constrained IoT setting.
【28】 A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
链接:https://arxiv.org/abs/2506.11130
摘要:我们提出了一个自我完善的框架,提高ASR的性能,只有未标记的数据集。该过程从现有的ASR模型开始,在未注释的语音上生成伪标签,然后将其用于训练高保真文本到语音(TTS)系统。然后,合成的语音文本对被引导到原始ASR系统中,完成闭环自我改进循环。我们证明了该框架对台湾普通话语音的有效性。利用6,000小时的未标记语音,适量的文本数据和人工智能模型的合成内容,我们将Whisper-large-v2改编为一个专门的模型Twister。与Whisper相比,Twister在普通话上减少了20%的错误率,在普通话-英语代码转换基准测试中减少了50%。结果突出的框架作为一个令人信服的替代伪标签自蒸馏方法,并提供了一个实用的途径,以提高ASR性能在低资源或特定领域的设置。
摘要:We propose a self-refining framework that enhances ASR performance with only unlabeled datasets. The process starts with an existing ASR model generating pseudo-labels on unannotated speech, which are then used to train a high-fidelity text-to-speech (TTS) system. Then, synthesized speech text pairs are bootstrapped into the original ASR system, completing the closed-loop self-improvement cycle. We demonstrated the effectiveness of the framework on Taiwanese Mandarin speech. Leveraging 6,000 hours of unlabeled speech, a moderate amount of text data, and synthetic content from the AI models, we adapt Whisper-large-v2 into a specialized model, Twister. Twister reduces error rates by up to 20% on Mandarin and 50% on Mandarin-English code-switching benchmarks compared to Whisper. Results highlight the framework as a compelling alternative to pseudo-labeling self-distillation approaches and provides a practical pathway for improving ASR performance in low-resource or domain-specific settings.
【29】 SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR
链接:https://arxiv.org/abs/2506.11121
摘要:尽管在端到端ASR方面取得了进展,但现实世界中的领域不匹配仍然会导致性能下降,测试时自适应(TTA)旨在通过在推理期间调整模型来缓解这种情况。最近的工作探索将TTA与外部语言模型相结合,使用波束搜索重新评分或生成错误校正等技术。在这项工作中,我们确定了一个以前被忽视的挑战:TTA可以干扰语言模型重新评分,揭示了有效结合这两种方法的重要性。基于这种见解,我们提出了SUTA-LM,一个简单而有效的扩展SUTA,一个基于熵最小化的TTA方法,语言模型重新评分。SUTA-LM首先应用由自动步骤选择机制引导的受控自适应过程,该机制利用声学和语言信息,然后通过语言模型重新评分来细化输出。在18个不同的ASR数据集上的实验表明,SUTA-LM在广泛的领域中实现了稳健的结果。
摘要:Despite progress in end-to-end ASR, real-world domain mismatches still cause performance drops, which Test-Time Adaptation (TTA) aims to mitigate by adjusting models during inference. Recent work explores combining TTA with external language models, using techniques like beam search rescoring or generative error correction. In this work, we identify a previously overlooked challenge: TTA can interfere with language model rescoring, revealing the nontrivial nature of effectively combining the two methods. Based on this insight, we propose SUTA-LM, a simple yet effective extension of SUTA, an entropy-minimization-based TTA approach, with language model rescoring. SUTA-LM first applies a controlled adaptation process guided by an auto-step selection mechanism leveraging both acoustic and linguistic information, followed by language model rescoring to refine the outputs. Experiments on 18 diverse ASR datasets show that SUTA-LM achieves robust results across a wide range of domains.
【30】 Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech
链接:https://arxiv.org/abs/2506.11119
摘要:背景资料:阿尔茨海默病和相关痴呆症(ADRD)是一种进行性神经退行性疾病,早期发现对于及时干预和护理至关重要。自发言语包含丰富的声学和语言标记,可以作为认知能力下降的非侵入性生物标志物。在大规模音频或文本数据上预训练的基础模型产生编码上下文和声学特征的高维嵌入。 研究方法:我们使用了CHTE挑战数据集,其中包括来自1,600多名参与者的录音,这些参与者具有三种认知状态:健康对照(HC),轻度认知障碍(MCI)和阿尔茨海默病(AD)。我们排除了非英语、非自发或质量差的录音。最终数据集包括703例(59.13%)HC、81例(6.81%)MCI和405例(34.06%)AD病例。我们对一系列开源基础语音和语言模型进行了基准测试,将认知状态分为三类。 结果:Whisper-medium模型在语音模型中获得了最高的性能(准确率= 0.731,AUC = 0.802)。在语言模型中,带有停顿注释的BERT表现最好(准确度= 0.662,AUC = 0.744)。使用最先进的自动语音识别(ASR)模型生成的音频嵌入的ADRD检测优于其他方法。包括非语义特征,如暂停模式,始终改进了基于文本的分类。 结论:本研究介绍了一个基准框架,使用基础模型和临床相关数据集。基于声学的方法-特别是ASR衍生的嵌入-表现出可扩展的,非侵入性的和具有成本效益的ADRD早期检测的强大潜力。
摘要:Background: Alzheimer's disease and related dementias (ADRD) are progressive neurodegenerative conditions where early detection is vital for timely intervention and care. Spontaneous speech contains rich acoustic and linguistic markers that may serve as non-invasive biomarkers for cognitive decline. Foundation models, pre-trained on large-scale audio or text data, produce high-dimensional embeddings encoding contextual and acoustic features. Methods: We used the PREPARE Challenge dataset, which includes audio recordings from over 1,600 participants with three cognitive statuses: healthy control (HC), mild cognitive impairment (MCI), and Alzheimer's Disease (AD). We excluded non-English, non-spontaneous, or poor-quality recordings. The final dataset included 703 (59.13%) HC, 81 (6.81%) MCI, and 405 (34.06%) AD cases. We benchmarked a range of open-source foundation speech and language models to classify cognitive status into the three categories. Results: The Whisper-medium model achieved the highest performance among speech models (accuracy = 0.731, AUC = 0.802). Among language models, BERT with pause annotation performed best (accuracy = 0.662, AUC = 0.744). ADRD detection using state-of-the-art automatic speech recognition (ASR) model-generated audio embeddings outperformed others. Including non-semantic features like pause patterns consistently improved text-based classification. Conclusion: This study introduces a benchmarking framework using foundation models and a clinically relevant dataset. Acoustic-based approaches -- particularly ASR-derived embeddings -- demonstrate strong potential for scalable, non-invasive, and cost-effective early detection of ADRD.
【31】 Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
链接:https://arxiv.org/abs/2506.11096
摘要:像wav2vec2和HuBERT这样的预训练语音表示表现出很强的各向异性,导致随机嵌入之间的高度相似性。虽然已被广泛观察到,但这一特性对下游任务的影响仍不清楚。这项工作评估各向异性的关键字发现计算文献语言学。使用动态时间规整,我们表明,尽管各向异性,wav2vec2相似性措施有效地识别单词没有转录。我们的研究结果突出了这些表征的鲁棒性,这些表征捕获了语音结构并在说话者之间进行概括。我们的研究结果强调了预训练在学习丰富和不变的语音表示中的重要性。
摘要:Pretrained speech representations like wav2vec2 and HuBERT exhibit strong anisotropy, leading to high similarity between random embeddings. While widely observed, the impact of this property on downstream tasks remains unclear. This work evaluates anisotropy in keyword spotting for computational documentary linguistics. Using Dynamic Time Warping, we show that despite anisotropy, wav2vec2 similarity measures effectively identify words without transcription. Our results highlight the robustness of these representations, which capture phonetic structures and generalize across speakers. Our results underscore the importance of pretraining in learning rich and invariant speech representations.
【32】 Customizing Speech Recognition Model with Large Language Model Feedback
链接:https://arxiv.org/abs/2506.11091
摘要:自动语音识别(ASR)系统在一般的转录任务上取得了很好的性能。然而,他们仍然在努力识别罕见的命名实体并适应域不匹配。相比之下,在大量互联网规模数据集上训练的大型语言模型(LLM)在广泛的领域中通常更有效。在这项工作中,我们提出了一种基于强化学习的无监督域自适应方法,利用未标记的数据来提高转录质量,特别是通过LLM的反馈来提高受域不匹配影响的命名实体。鉴于上下文信息,我们的框架采用LLM作为奖励模型,从ASR模型中对假设进行评分。这些分数作为奖励信号,通过强化学习来微调ASR模型。与传统的自训练方法相比,该方法在实体词错误率上提高了21%.
摘要:Automatic speech recognition (ASR) systems have achieved strong performance on general transcription tasks. However, they continue to struggle with recognizing rare named entities and adapting to domain mismatches. In contrast, large language models (LLMs), trained on massive internet-scale datasets, are often more effective across a wide range of domains. In this work, we propose a reinforcement learning based approach for unsupervised domain adaptation, leveraging unlabeled data to enhance transcription quality, particularly the named entities affected by domain mismatch, through feedback from a LLM. Given contextual information, our framework employs a LLM as the reward model to score the hypotheses from the ASR model. These scores serve as reward signals to fine-tune the ASR model via reinforcement learning. Our method achieves a 21\% improvement on entity word error rate over conventional self-training methods.
【33】 End-to-End Diarization utilizing Attractor Deep Clustering
链接:https://arxiv.org/abs/2506.11090
备注:To appear at INTERSPEECH 2025
摘要:说话人日记化仍然具有挑战性,由于需要结构化的说话人表示,有效的建模和鲁棒性变化的条件。我们提出了一个高性能的,紧凑的日记化框架,集成了构象解码器,变换器更新的吸引子,和一个深聚类风格的角度损失。我们的方法改进了扬声器表示与增强的构象结构,将交叉注意力吸引子和一个额外的卷积模块。为了加强结构化嵌入,我们通过构建标签吸引子向量来扩展深度聚类,将其方向结构与音频嵌入对齐。我们还施加正交约束的积极吸引更好的扬声器分离,同时抑制非积极吸引,以防止假激活。最后,一个排列不变的训练二进制交叉熵损失细化说话人检测。实验结果表明,该方法在保持参数个数不变的情况下,实现了较低的日志化误差。
摘要:Speaker diarization remains challenging due to the need for structured speaker representations, efficient modeling, and robustness to varying conditions. We propose a performant, compact diarization framework that integrates conformer decoders, transformer-updated attractors, and a deep clustering style angle loss. Our approach refines speaker representations with an enhanced conformer structure, incorporating cross-attention to attractors and an additional convolution module. To enforce structured embeddings, we extend deep clustering by constructing label-attractor vectors, aligning their directional structure with audio embeddings. We also impose orthogonality constraints on active attractors for better speaker separation while suppressing non-active attractors to prevent false activations. Finally, a permutation invariant training binary cross-entropy loss refines speaker detection. Experiments show that our method achieves low diarization error while maintaining parameter count.
机器翻译由腾讯交互翻译提供,仅供参考
