今日论文合集:cs.SD语音28篇,eess.AS音频处理30篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Effective and Efficient One-pass Compression of Speech Foundation Models  Using Sparsity-aware Self-pinching Gates
标题: 使用稀疏感知自挤压门对语音基础模型进行有效且高效的单程压缩
链接:https://arxiv.org/abs/2505.22608
作者: Haoning Xu,  Zhaoqing Li,  Youjun Chen,  Huimeng Wang,  Guinan Li,  Mengzhe Geng,  Chengxi Deng,  Xunying Liu 
备注:Submitted to Interspeech 2025
摘要:本文提出了一种新的语音基础模型压缩方法,紧密结合模型修剪和参数更新到一个单一的阶段。高度紧凑的层级绑定自箍缩门,每个门只包含一个可学习的阈值,与未压缩的模型联合训练,并用于细粒度的神经元级修剪。在LibriSpeech-100 hr语料库上进行的实验表明,我们的方法将wav2vec2.0-base和HuBERT-large模型的参数数量分别减少了65%和60%,同时在测试干净的数据集上没有产生统计上显著的单词错误率(WER)增加。与以前发表的相同任务的方法相比,我们的方法不仅在4.26倍的可比模型压缩比下在测试干净数据集上实现了7.05%的最低WER,而且模型压缩时间至少减少了25%。
摘要:This paper presents a novel approach for speech foundation models compression that tightly integrates model pruning and parameter update into a single stage. Highly compact layer-level tied self-pinching gates each containing only a single learnable threshold are jointly trained with uncompressed models and used in fine-grained neuron level pruning. Experiments conducted on the LibriSpeech-100hr corpus suggest that our approach reduces the number of parameters of wav2vec2.0-base and HuBERT-large models by 65% and 60% respectively, while incurring no statistically significant word error rate (WER) increase on the test-clean dataset. Compared to previously published methods on the same task, our approach not only achieves the lowest WER of 7.05% on the test-clean dataset under a comparable model compression ratio of 4.26x, but also operates with at least 25% less model compression time.


【2】 Towards General Discrete Speech Codec for Complex Acoustic Environments:  A Study of Reconstruction and Downstream Task Consistency

标题: 面向复杂声学环境的通用离散语音编解码器:重建和下游任务一致性研究
链接:https://arxiv.org/abs/2505.22515
作者: Haoran Wang,  Guanyu Chen,  Bohan Li,  Hankun Wang,  Yiwei Guo,  Zhihan Li,  Xie Chen,  Kai Yu 
备注:Initial Upload
摘要:神经语音编解码器在重建干净的语音信号方面表现出色;然而,它们在复杂声学环境和下游信号处理任务中的功效仍然有待探索。在这项研究中,我们引入了一个新的基准称为环境弹性语音编解码器基准(ERSB),系统地评估神经语音编解码器是否具有环境弹性。具体来说,我们评估了两个关键能力:(1)鲁棒重建,它测量语音和非语音声学细节的保留,以及(2)下游任务一致性,它确保在使用重建语音而不是原始语音时下游信号处理任务的偏差最小。我们的综合实验表明,复杂的声学环境显着降低信号重建和下游任务的一致性。这项工作突出了当前语音编解码器的局限性,并提出了一个未来的方向,以改善它们,提高环境适应力。
摘要:Neural speech codecs excel in reconstructing clean speech signals; however, their efficacy in complex acoustic environments and downstream signal processing tasks remains underexplored. In this study, we introduce a novel benchmark named Environment-Resilient Speech Codec Benchmark (ERSB) to systematically evaluate whether neural speech codecs are environment-resilient. Specifically, we assess two key capabilities: (1) robust reconstruction, which measures the preservation of both speech and non-speech acoustic details, and (2) downstream task consistency, which ensures minimal deviation in downstream signal processing tasks when using reconstructed speech instead of the original. Our comprehensive experiments reveal that complex acoustic environments significantly degrade signal reconstruction and downstream task consistency. This work highlights the limitations of current speech codecs and raises a future direction that improves them for greater environmental resilience.


【3】 Effective Context in Neural Speech Models

标题: 神经语音模型中的有效上下文
链接:https://arxiv.org/abs/2505.22487
作者: Yen Meng,  Sharon Goldwater,  Hao Tang 
备注:Accepted to Interspeech 2025
摘要:现代神经语音模型受益于更长的上下文,并且已经提出了许多方法来增加模型可以使用的最大上下文。然而,很少有人试图衡量这些模型实际使用了多少上下文,即,有效的上下文。在这里,我们提出了两种测量有效语境的方法,并用它们来分析不同的言语Transformers。对于监督模型,我们发现有效上下文与任务的性质相关,基频跟踪,电话分类和单词分类需要越来越多的有效上下文。对于自监督模型,我们发现有效上下文主要在早期层中增加,并且保持相对较短-类似于监督电话模型。鉴于这些模型在预测过程中不使用长上下文,我们证明了HuBERT可以在流模式下运行,而无需修改架构,也无需进一步微调。
摘要:Modern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e., the effective context. Here, we propose two approaches to measuring the effective context, and use them to analyze different speech Transformers. For supervised models, we find that the effective context correlates well with the nature of the task, with fundamental frequency tracking, phone classification, and word classification requiring increasing amounts of effective context. For self-supervised models, we find that effective context increases mainly in the early layers, and remains relatively short -- similar to the supervised phone model. Given that these models do not use a long context during prediction, we show that HuBERT can be run in streaming mode without modification to the architecture and without further fine-tuning.


【4】 FGS-Audio: Fixed-Decoder Framework for Audio Steganography with  Adversarial Perturbation Generation

标题: FGS-音频:具有对抗性微扰生成的音频隐写术的固定解码器框架
链接:https://arxiv.org/abs/2505.22266
作者: Jialin Yan,  Yu Cheng,  Zhaoxia Yin,  Xinpeng Zhang,  Shilin Wang,  Tanfeng Sun,  Xinghao Jiang 
摘要:人工智能生成内容(AIGC)的快速发展使得高保真生成的音频在互联网上广泛可用,为隐蔽通信提供了丰富而通用的掩护信号源。在深度学习的推动下,当前的音频隐写框架主要基于编解码网络架构。虽然这些方法大大提高了音频隐写的安全性,但它们通常采用精心设计的训练工作流程,并依赖于广泛的预训练模型。为了解决上述问题,本文开创性地提出了一种基于对抗扰动生成的音频隐写固定解码器框架(FGS-Audio)。携带秘密信息的对抗性扰动被嵌入到覆盖音频中以生成隐写音频。接收方只需共享固定解码网络的结构和权值,就能准确地从隐写音频中提取秘密信息,从而消除了对大型预训练模型的依赖。在FGS-Audio中,我们提出了一种音频对抗扰动生成(APG)策略,并设计了一个轻量级的固定解码器。固定解码器保证可靠地提取隐藏消息,而对抗性扰动经过优化,以保持隐写音频在感知和统计上接近覆盖音频,从而提高对隐写分析的抵抗力。实验结果表明,该方法在不同的相对负载下具有良好的抗隐写分析性能,优于现有的SOTA方法。在隐写音频质量方面,与SOTA方法相比,FGS-Audio实现了超过10 dB的平均PSNR改善。
摘要:The rapid development of Artificial Intelligence Generated Content (AIGC) has made high-fidelity generated audio widely available across the Internet, offering an abundant and versatile source of cover signals for covert communication. Driven by advances in deep learning, current audio steganography frameworks are mainly based on encoding-decoding network architectures. While these methods greatly improve the security of audio steganography, they typically employ elaborate training workflows and rely on extensive pre-trained models. To address the aforementioned issues, this paper pioneers a Fixed-Decoder Framework for Audio Steganography with Adversarial Perturbation Generation (FGS-Audio). The adversarial perturbations that carry secret information are embedded into cover audio to generate stego audio. The receiver only needs to share the structure and weights of the fixed decoding network to accurately extract the secret information from the stego audio, thus eliminating the reliance on large pre-trained models. In FGS-Audio, we propose an audio Adversarial Perturbation Generation (APG) strategy and design a lightweight fixed decoder. The fixed decoder guarantees reliable extraction of the hidden message, while the adversarial perturbations are optimized to keep the stego audio perceptually and statistically close to the cover audio, thereby improving resistance to steganalysis. The experimental results show that the method exhibits excellent anti-steganalysis performance under different relative payloads, outperforming existing SOTA approaches. In terms of stego audio quality, FGS-Audio achieves an average PSNR improvement of over 10 dB compared to SOTA method.


【5】 Advancing Hearing Assessment: An ASR-Based Frequency-Specific Speech  Test for Diagnosing Presbycusis

标题: 高级听力评估:用于诊断老年性耳聋的基于SVR的特定频率言语测试
链接:https://arxiv.org/abs/2505.22231
作者: Stefan Bleeck 
摘要:传统的测听通常无法完全表征听力损失对言语理解的功能影响,特别是在老年性耳聋等情况下的阈上缺陷和频率特异性感知挑战。本文介绍了一种新的自动语音识别(ASR)为基础的频率特定的语音测试,旨在提供粒度诊断的见解的发展和模拟评估。我们的方法利用ASR来模拟中度倾斜听力损失的感知效果,通过在受控的声学退化下处理语音刺激,随后分析音素级混淆模式。主要研究结果表明,模拟听力损失会引入特定的音素混淆,主要影响高频辅音(例如,齿槽/腭到唇齿替换)并导致显著的音素缺失,与老年性聋中的声学线索退化一致。从这些ASR衍生的混淆中策划的测试电池证明了诊断价值,在综合模拟中有效区分了模拟的正常听力和听力受损的听众。这种ASR驱动的方法为开发客观,粒度和频率特定的听力评估工具提供了一个有前途的途径,这些工具补充了传统的听力测试。未来的工作将集中在与人类参与者一起验证这些发现,并探索先进人工智能模型的整合,以提高诊断精度。
摘要:Traditional audiometry often fails to fully characterize the functional impact of hearing loss on speech understanding, particularly supra-threshold deficits and frequency-specific perception challenges in conditions like presbycusis. This paper presents the development and simulated evaluation of a novel Automatic Speech Recognition (ASR)-based frequency-specific speech test designed to provide granular diagnostic insights. Our approach leverages ASR to simulate the perceptual effects of moderate sloping hearing loss by processing speech stimuli under controlled acoustic degradation and subsequently analyzing phoneme-level confusion patterns. Key findings indicate that simulated hearing loss introduces specific phoneme confusions, predominantly affecting high-frequency consonants (e.g., alveolar/palatal to labiodental substitutions) and leading to significant phoneme deletions, consistent with the acoustic cues degraded in presbycusis. A test battery curated from these ASR-derived confusions demonstrated diagnostic value, effectively differentiating between simulated normal-hearing and hearing-impaired listeners in a comprehensive simulation. This ASR-driven methodology offers a promising avenue for developing objective, granular, and frequency-specific hearing assessment tools that complement traditional audiometry. Future work will focus on validating these findings with human participants and exploring the integration of advanced AI models for enhanced diagnostic precision.


【6】 Two-stage Audio-Visual Target Speaker Extraction System for Real-Time  Processing On Edge Device

标题: 边缘设备实时处理的两级视听目标说话人提取系统
链接:https://arxiv.org/abs/2505.22229
作者: Zixuan Li,  Xueliang Zhang,  Lei Miao,  Zhipeng Yan 
摘要:视听目标说话人提取(AVTSE)的目的是在多说话人环境中,以视觉线索为辅助,分离出目标说话人的声音。大多数现有的AVTSE方法同时对视觉和音频特征进行编码,导致极高的计算复杂度,并使其无法在边缘设备上进行实时处理。为了解决这个问题,我们提出了一个两级超紧凑型AVTSE系统。具体地,在第一阶段中,采用紧凑的网络用于使用视觉信息的语音活动检测(VAD)。在第二阶段,VAD结果与音频输入相结合,以隔离目标说话者的语音。实验表明,该系统有效地抑制了背景噪声和干扰语音,而花费很少的计算资源。
摘要:Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously, resulting in extremely high computational complexity and making it impractical for real-time processing on edge devices. To tackle this issue, we proposed a two-stage ultra-compact AVTSE system. Specifically, in the first stage, a compact network is employed for voice activity detection (VAD) using visual information. In the second stage, the VAD results are combined with audio inputs to isolate the target speaker's voice. Experiments show that the proposed system effectively suppresses background noise and interfering voices while spending little computational resources.


【7】 Developing a Top-tier Framework in Naturalistic Conditions Challenge for  Categorized Emotion Prediction: From Speech Foundation Models and Learning  Objective to Data Augmentation and Engineering Choices

标题: 在自然条件挑战中开发顶级框架以实现分类情绪预测:从言语基础模型和学习目标到数据增强和工程选择
链接:https://arxiv.org/abs/2505.22133
作者: Tiantian Feng,  Thanathai Lertpetchpun,  Dani Byrd,  Shrikanth Narayanan 
备注:Accepted to INTERSPEECH 2025
摘要:语音情感识别(SER),特别是自然表达的情感,仍然是一个具有挑战性的计算任务。主要挑战包括情感标注中固有的主观性和数据集中情感标签的不平衡分布。本文介绍了为参加INTERSPEECH 2025情感识别挑战赛(任务1)而开发的\texttt{SAILER}系统。挑战数据集包含来自播客的自然情感语音,是研究不平衡和主观情感注释的宝贵资源。我们的系统设计简单,可重复,有效,突出了建模,学习目标,数据增强和工程选择的关键选择。结果表明,即使是一个单一的系统(没有集成)可以超过95%的提交,宏观F1得分超过0.4。此外,三个系统的组合进一步提高了性能,获得了竞争性排名分数(表现前三的团队)。我们的模型在:https://github.com/tiantiaf0627/vox-profile-release。
摘要:Speech emotion recognition (SER), particularly for naturally expressed emotions, remains a challenging computational task. Key challenges include the inherent subjectivity in emotion annotation and the imbalanced distribution of emotion labels in datasets. This paper introduces the \texttt{SAILER} system developed for participation in the INTERSPEECH 2025 Emotion Recognition Challenge (Task 1). The challenge dataset, which contains natural emotional speech from podcasts, serves as a valuable resource for studying imbalanced and subjective emotion annotations. Our system is designed to be simple, reproducible, and effective, highlighting critical choices in modeling, learning objectives, data augmentation, and engineering choices. Results show that even a single system (without ensembling) can outperform more than 95\% of the submissions, with a Macro-F1 score exceeding 0.4. Moreover, an ensemble of three systems further improves performance, achieving a competitively ranked score (top-3 performing team). Our model is at: https://github.com/tiantiaf0627/vox-profile-release.


【8】 AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

标题: AudioTurbo:具有矫正扩散的快速文本到音频生成
链接:https://arxiv.org/abs/2505.22106
作者: Junqi Zhao,  Jinzheng Zhao,  Haohe Liu,  Yun Chen,  Lu Han,  Xubo Liu,  Mark Plumbley,  Wenwu Wang 
摘要:扩散模型显着提高了音频生成的质量和多样性,但受到推理速度慢的阻碍。修正流通过学习直线常微分方程(ODE)路径来提高推理速度。然而,这种方法需要从头开始训练流匹配模型,并且在低步数下往往表现不佳,甚至表现不佳。为了解决整流流的局限性,同时利用先进的预训练扩散模型的优势,本研究将预训练模型与整流扩散方法相结合,以提高文本到音频(TTA)生成的效率。具体来说,我们提出了AudioTurbo,它从预先训练的TTA模型生成的确定性噪声样本对中学习一阶ODE路径。AudioCaps数据集上的实验表明,我们的模型,只有10个采样步骤,优于以前的模型,并减少推理到3个步骤相比,基于流匹配的加速模型。
摘要:Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by learning straight-line ordinary differential equation (ODE) paths. However, this approach requires training a flow-matching model from scratch and tends to perform suboptimally, or even poorly, at low step counts. To address the limitations of rectified flow while leveraging the advantages of advanced pre-trained diffusion models, this study integrates pre-trained models with the rectified diffusion method to improve the efficiency of text-to-audio (TTA) generation. Specifically, we propose AudioTurbo, which learns first-order ODE paths from deterministic noise sample pairs generated by a pre-trained TTA model. Experiments on the AudioCaps dataset demonstrate that our model, with only 10 sampling steps, outperforms prior models and reduces inference to 3 steps compared to a flow-matching-based acceleration model.


【9】 Visual Cues Support Robust Turn-taking Prediction in Noise

标题: 视觉线索支持噪音中稳健的轮流预测
链接:https://arxiv.org/abs/2505.22088
作者: Sam O'Connor Russell,  Naomi Harte 
备注:5 pages
摘要:准确的预测话轮转换模型(PTTMs)是自然的人机交互的基础。然而,人们对它们在噪音方面的表现知之甚少。因此,本研究探讨PTTM性能的类型的噪音可能会遇到一次部署。我们的分析表明PTTM对噪声高度敏感。保持/移位精度从清晰语音中的84%下降到10 dB音乐噪声中的52%。使用噪声数据进行训练可以实现多模态PTTM,其中包括视觉特征,以更好地利用视觉线索,在10 dB音乐噪声中具有72%的准确率。多模态PTTM在所有噪声类型和SNR上都优于仅音频PTTM,突出了其利用视觉线索的能力;然而,这并不总是适用于新类型的噪声。分析还表明,成功的训练依赖于准确的转录,限制了ASR衍生的translantine在清洁条件下的使用。我们将代码公开以供将来的研究。
摘要:Accurate predictive turn-taking models (PTTMs) are essential for naturalistic human-robot interaction. However, little is known about their performance in noise. This study therefore explores PTTM performance in types of noise likely to be encountered once deployed. Our analyses reveal PTTMs are highly sensitive to noise. Hold/shift accuracy drops from 84% in clean speech to just 52% in 10 dB music noise. Training with noisy data enables a multimodal PTTM, which includes visual features to better exploit visual cues, with 72% accuracy in 10 dB music noise. The multimodal PTTM outperforms the audio-only PTTM across all noise types and SNRs, highlighting its ability to exploit visual cues; however, this does not always generalise to new types of noise. Analysis also reveals that successful training relies on accurate transcription, limiting the use of ASR-derived transcriptions to clean conditions. We make code publicly available for future research.


【10】 On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech  Foundation Models for Dysarthric Speech Recognition

标题: Zero-ShotMoE说话人的实时路由用于合成障碍语音识别的语音基础模型的适应
链接:https://arxiv.org/abs/2505.22072
作者: Shujie HU,  Xurong Xie,  Mengzhe Geng,  Jiajun Deng,  Huimeng Wang,  Guinan Li,  Chengxi Deng,  Tianzi Wang,  Mingyu Cui,  Helen Meng,  Xunying Liu 
备注:Accepted by Interspeech 2025
摘要:针对构音障碍语音识别的基础模型,提出了一种基于MoE的说话人自适应框架。该方法在结合领域知识的同时实现了zero-shot自适应和实时处理。语音障碍的严重程度和性别条件的适配器专家动态结合使用的飞行预测扬声器相关的路由参数。KL-分歧用于进一步加强专家之间的多样性和他们的泛化看不见的扬声器。UASpeech语料库的实验结果表明,在飞行MoE为基础的适应产生统计上显着的WER减少高达1.34%的绝对值(6.36%相对)超过未适应的基线HuBERT/WavLM模型。一致的WER减少高达2.55%的绝对(11.44%相对)和RTF加速高达7倍,获得了在不同的扬声器级数据量的批处理模式自适应。最低公布的WER为16.35%(46.77%的非常低的可懂度)。
摘要:This paper proposes a novel MoE-based speaker adaptation framework for foundation models based dysarthric speech recognition. This approach enables zero-shot adaptation and real-time processing while incorporating domain knowledge. Speech impairment severity and gender conditioned adapter experts are dynamically combined using on-the-fly predicted speaker-dependent routing parameters. KL-divergence is used to further enforce diversity among experts and their generalization to unseen speakers. Experimental results on the UASpeech corpus suggest that on-the-fly MoE-based adaptation produces statistically significant WER reductions of up to 1.34% absolute (6.36% relative) over the unadapted baseline HuBERT/WavLM models. Consistent WER reductions of up to 2.55% absolute (11.44% relative) and RTF speedups of up to 7 times are obtained over batch-mode adaptation across varying speaker-level data quantities. The lowest published WER of 16.35% (46.77% on very low intelligibility) is obtained.


【11】 Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency  Streaming ASR

标题: Delayed-KD:基于延迟知识蒸馏的CIC,用于低延迟流媒体ASB
链接:https://arxiv.org/abs/2505.22069
作者: Longhao Li,  Yangze Li,  Hongfei Xue,  Jie Liu,  Shuai Fang,  Kai Wang,  Lei Xie 
备注:Accepted by Interspeech2025
摘要:基于CTC的流式ASR在现实世界的应用中获得了极大的关注,但面临两个主要挑战:小块的准确性下降和令牌发射延迟。为了缓解这些挑战,我们提出了延迟KD,它将延迟知识蒸馏应用于从非流模型到流模型的CTC后验概率。具体来说,在块大小很小的情况下,我们引入了一个时间对齐缓冲区(TAB),该缓冲区定义了与非流教师模型相比的相对延迟范围,以对齐CTC输出并减轻非空白令牌不匹配。此外,TAB支持对令牌发送延迟进行细粒度控制。在178小时的AISHELL-1和10,000小时的WenetSpeech汉语数据集上的实验表明,Delayed-KD具有一致的优越性。令人印象深刻的是,延迟KD在40 ms延迟时在AISHELL-1上实现了5.42%的较低字符错误率(CER),与在320 ms延迟时运行的竞争对手U2++模型相当。
摘要:CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency. To mitigate these challenges, we propose Delayed-KD, which applies delayed knowledge distillation on CTC posterior probabilities from a non-streaming to a streaming model. Specifically, with a tiny chunk size, we introduce a Temporal Alignment Buffer (TAB) that defines a relative delay range compared to the non-streaming teacher model to align CTC outputs and mitigate non-blank token mismatches. Additionally, TAB enables fine-grained control over token emission delay. Experiments on 178-hour AISHELL-1 and 10,000-hour WenetSpeech Mandarin datasets show consistent superiority of Delayed-KD. Impressively, Delayed-KD at 40 ms latency achieves a lower character error rate (CER) of 5.42% on AISHELL-1, comparable to the competitive U2++ model running at 320 ms latency.


【12】 Weakly Supervised Data Refinement and Flexible Sequence Compression for  Efficient Thai LLM-based ASR

标题: 弱监督数据细化和灵活的序列压缩,以实现高效的泰国基于LLM的ASB
链接:https://arxiv.org/abs/2505.22063
作者: Mingchen Shao,  Xinfa Zhu,  Chengyou Wang,  Bingshen Mu,  Hai Li,  Ying Yan,  Junhui Liu,  Danming Xie,  Lei Xie 
备注:Accepted by INTERSPEECH 2025
摘要:尽管取得了显著的成就,但低资源场景下的自动语音识别(ASR)仍然面临两个挑战:高质量数据稀缺和高计算需求。本文提出了EThai-ASR,第一个将大型语言模型(LLM)应用于泰国ASR,并创建了一个高效的基于LLM的ASR系统。EThai-ASR包括语音编码器、连接模块和泰语LLM解码器。为了解决数据稀缺问题并获得强大的语音编码器,EThai-ASR引入了一种自进化的数据细化策略来细化弱标签,从而产生增强的语音编码器。此外,我们提出了一个可插拔的序列压缩模块中使用的连接模块,设计了三种模式,以减少序列长度,从而减少计算需求,同时保持体面的性能。大量的实验表明,EThai-ASR在多个数据集上达到了最先进的精度。我们发布我们的精炼文本记录,以促进进一步的研究。
摘要:Despite remarkable achievements, automatic speech recognition (ASR) in low-resource scenarios still faces two challenges: high-quality data scarcity and high computational demands. This paper proposes EThai-ASR, the first to apply large language models (LLMs) to Thai ASR and create an efficient LLM-based ASR system. EThai-ASR comprises a speech encoder, a connection module and a Thai LLM decoder. To address the data scarcity and obtain a powerful speech encoder, EThai-ASR introduces a self-evolving data refinement strategy to refine weak labels, yielding an enhanced speech encoder. Moreover, we propose a pluggable sequence compression module used in the connection module with three modes designed to reduce the sequence length, thus decreasing computational demands while maintaining decent performance. Extensive experiments demonstrate that EThai-ASR has achieved state-of-the-art accuracy in multiple datasets. We release our refined text transcripts to promote further research.


【13】 Voice Adaptation for Swiss German

标题: 瑞士德语语音改编
链接:https://arxiv.org/abs/2505.22054
作者: Samuel Stucki,  Jan Deriu,  Mark Cieliebak 
备注:Submitted to Interspeech
摘要:这项工作研究了瑞士德语方言的语音适应模型的性能,即,将标准德语文本翻译成瑞士德语方言语音。为此,我们预处理了一个大型的瑞士播客数据集,我们自动转录并用方言类进行注释,产生了大约5000小时的弱标记训练材料。我们在这个数据集上对XTTSv 2模型进行了微调,并表明它在人工和自动评估中取得了良好的成绩,并且可以正确地呈现所需的方言。我们的工作表明了一个步骤,使语音克隆技术适应代表性不足的语言。由此产生的模型实现了高达-0.28的CMOS分数和3.8的SMOS分数。
摘要:This work investigates the performance of Voice Adaptation models for Swiss German dialects, i.e., translating Standard German text to Swiss German dialect speech. For this, we preprocess a large dataset of Swiss podcasts, which we automatically transcribe and annotate with dialect classes, yielding approximately 5000 hours of weakly labeled training material. We fine-tune the XTTSv2 model on this dataset and show that it achieves good scores in human and automated evaluations and can correctly render the desired dialect. Our work shows a step towards adapting Voice Cloning technology to underrepresented languages. The resulting model achieves CMOS scores of up to -0.28 and SMOS scores of 3.8.


【14】 AudioGenie: A Training-Free Multi-Agent Framework for Diverse  Multimodality-to-Multiaudio Generation

标题: AudioGenie:一个免训练的多代理框架,用于多样化的多模式到多音频生成
链接:https://arxiv.org/abs/2505.22053
作者: Yan Rong,  Jinting Wang,  Shan Yang,  Guangzhi Lei,  Li Liu 
摘要:多模态到多音频(MM 2 MA)生成在合成不同的和上下文对齐的音频类型(例如,音效、语音、音乐和歌曲)从多模态输入(例如,视频,文本,图像),由于缺乏高质量的配对数据集和缺乏强大的多任务学习框架。近年来,多智能体系统在解决上述问题方面显示出巨大的潜力。然而,直接将其应用于MM 2 MA任务带来了三个关键挑战:(1)对多模式输入(尤其是视频)的细粒度理解不足,(2)单个模型无法处理不同的音频事件,以及(3)缺乏可靠输出的自我纠正机制。为此,我们提出了AudioGenie,一种新的训练免费多代理系统具有双层架构的生成团队和监督团队。对于生成团队,设计了细粒度的任务分解和自适应的混合专家(MoE)协作实体进行动态模型选择,设计了试错迭代精化模块进行自校正.主管团队确保时空一致性,并通过反馈回路核实产出。此外,我们建立了MA-Bench,MM 2 MA任务的第一个基准测试,包括198个带注释的视频和多类型的音频。实验表明,我们的AudioGenie优于国家的最先进的(SOTA)方法在9个指标在8个任务。用户研究进一步验证了所提出的方法在质量,准确性,对齐和美观方面的有效性。可以在https://audiogenie.github.io/上找到带有示例的匿名项目网站。
摘要:Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images), owing to the scarcity of high-quality paired datasets and the lack of robust multi-task learning frameworks. Recently, multi-agent system shows great potential in tackling the above issues. However, directly applying it to MM2MA task presents three critical challenges: (1) inadequate fine-grained understanding of multimodal inputs (especially for video), (2) the inability of single models to handle diverse audio events, and (3) the absence of self-correction mechanisms for reliable outputs. To this end, we propose AudioGenie, a novel training-free multi-agent system featuring a dual-layer architecture with a generation team and a supervisor team. For the generation team, a fine-grained task decomposition and an adaptive Mixture-of-Experts (MoE) collaborative entity are designed for dynamic model selection, and a trial-and-error iterative refinement module is designed for self-correction. The supervisor team ensures temporal-spatial consistency and verifies outputs through feedback loops. Moreover, we build MA-Bench, the first benchmark for MM2MA tasks, comprising 198 annotated videos with multi-type audios. Experiments demonstrate that our AudioGenie outperforms state-of-the-art (SOTA) methods across 9 metrics in 8 tasks. User study further validate the effectiveness of the proposed method in terms of quality, accuracy, alignment, and aesthetic. The anonymous project website with samples can be found at https://audiogenie.github.io/.


【15】 Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

标题: 缓解视觉指南音频字幕中的视听不匹配
链接:https://arxiv.org/abs/2505.22045
作者: Le Xu,  Chenxing Li,  Yong Ren,  Yujie Chen,  Yu Gu,  Ruibo Fu,  Shan Yang,  Dong Yu 
备注:Accepted by INTERSPEECH 2025
摘要:当前的视觉引导音频字幕系统经常无法解决现实世界场景中的视听错位,例如配音内容或屏幕外声音。为了弥合这一关键差距,我们提出了一个熵感知的门控融合框架,通过跨模态不确定性量化动态调制视觉信息流。我们的新方法采用注意熵分析的交叉注意层自动识别和抑制误导性的视觉线索在模态融合。作为对这种架构的补充,我们开发了一种批量视听洗牌技术,该技术生成合成的不匹配训练对,大大增强了模型对对齐噪声的弹性。AudioCaps基准的评估表明,我们的系统的性能优于现有的基线,特别是在不匹配的模态场景。此外,与基线相比,我们的解决方案在推理速度上提高了约6倍。
摘要:Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge this critical gap, we present an entropy-aware gated fusion framework that dynamically modulates visual information flow through cross-modal uncertainty quantification. Our novel approach employs attention entropy analysis in cross-attention layers to automatically identify and suppress misleading visual cues during modal fusion. Complementing this architecture, we develop a batch-wise audiovisual shuffling technique that generates synthetic mismatched training pairs, greatly enhancing model resilience against alignment noise. Evaluations on the AudioCaps benchmark demonstrate our system's superior performance over existing baselines, especially in mismatched modality scenarios. Furthermore, our solution demonstrates an approximately 6x improvement in inference speed compared to the baseline.


【16】 Improving Respiratory Sound Classification with Architecture-Agnostic  Knowledge Distillation from Ensembles

标题: 通过从合奏中提取建筑不可知知识来改进呼吸声分类
链接:https://arxiv.org/abs/2505.22027
作者: Miika Toikkanen,  June-Woo Kim 
备注:Accepted to Interspeech 2025
摘要:呼吸声数据集在大小和质量上受到限制,使得难以实现高性能。集成模型有所帮助,但不可避免地会增加推理时的计算成本。软标签培训有效地提取知识,仅在培训时增加额外成本。在这项研究中,我们探索呼吸声分类的软标签作为一个架构不可知的方法,蒸馏成一个学生模型的教师模型的合奏。我们研究了我们的方法的不同变化,发现即使是一个单一的老师,相同的学生,大大提高了性能超出了自己的能力,实现了最佳收益,只使用几个老师。我们在ICHBI上获得了64.39的最新得分,超过了之前的最好成绩0.85,并将各个架构的平均得分提高了1.16以上。我们的研究结果突出了知识蒸馏的有效性与呼吸声分类的软标签,无论大小或架构。
摘要:Respiratory sound datasets are limited in size and quality, making high performance difficult to achieve. Ensemble models help but inevitably increase compute cost at inference time. Soft label training distills knowledge efficiently with extra cost only at training. In this study, we explore soft labels for respiratory sound classification as an architecture-agnostic approach to distill an ensemble of teacher models into a student model. We examine different variations of our approach and find that even a single teacher, identical to the student, considerably improves performance beyond its own capability, with optimal gains achieved using only a few teachers. We achieve the new state-of-the-art Score of 64.39 on ICHBI, surpassing the previous best by 0.85 and improving average Scores across architectures by more than 1.16. Our results highlight the effectiveness of knowledge distillation with soft labels for respiratory sound classification, regardless of size or architecture.


【17】 RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic  Decomposed Modeling

标题: RESOUND:通过声学-语义分解建模从无声视频中重建语音
链接:https://arxiv.org/abs/2505.22024
作者: Long-Khanh Pham,  Thanh V. T. Tran,  Minh-Tan Pham,  Van Nguyen 
备注:accepted in Interspeech 2025
摘要:唇语合成(L2S),从视觉线索重建语音,面临着准确性和自然性的挑战,由于在捕捉语言内容,口音和韵律的监督有限。在本文中,我们提出了RESOUND,一种新的L2S系统,从无声的说话人脸视频生成可理解和表达的语音。利用源过滤器理论,我们的方法包括两个组成部分:声学路径预测韵律和语义路径提取语言特征。这种分离简化了学习,允许对每个表示进行独立优化。此外,我们通过将语音单元(一种经过验证的无监督语音表示技术)集成到梅尔频谱图旁边的波形生成中来提高性能。这允许RESOUND合成韵律语音,同时保留内容和说话者身份。两个标准的L2S基准进行的实验证实了所提出的方法在各种指标的有效性。
摘要:Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a novel L2S system that generates intelligible and expressive speech from silent talking face videos. Leveraging source-filter theory, our method involves two components: an acoustic path to predict prosody and a semantic path to extract linguistic features. This separation simplifies learning, allowing independent optimization of each representation. Additionally, we enhance performance by integrating speech units, a proven unsupervised speech representation technique, into waveform generation alongside mel-spectrograms. This allows RESOUND to synthesize prosodic speech while preserving content and speaker identity. Experiments conducted on two standard L2S benchmarks confirm the effectiveness of the proposed method across various metrics.


【18】 Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation  Addition for MISP 2025 Challenge

标题: MISP 2025挑战赛的重叠自适应混合扬声器拨号和ASB感知观察添加
链接:https://arxiv.org/abs/2505.22013
作者: Shangkun Huang,  Yuxuan Du,  Jingwen Yang,  Dejun Zhang,  Xupeng Jia,  Jing Deng,  Jintao Kang,  Rong Zheng 
备注:Accepted to Interspeech 2025
摘要:本文介绍了为应对MISP 2025挑战而开发的系统。对于日志化系统,我们提出了一种混合方法相结合的WavLM端到端的分割方法与传统的多模块聚类技术,以自适应地选择适当的模型来处理不同程度的重叠语音。针对自动语音识别(ASR)系统,提出了一种ASR感知的观测值添加方法,该方法可以弥补低信噪比条件下引导源分离(GSS)的性能限制。最后,我们在级联架构中集成了扬声器日志和ASR系统,以解决Track 3问题。我们的系统在轨道2上实现了9.48%的字符错误率(CER),在轨道3上实现了11.56%的级联最小排列字符错误率(cpCER),最终在两个轨道上都获得了第一名,从而证明了所提出的方法在现实世界的会议场景中的有效性。
摘要:This paper presents the system developed to address the MISP 2025 Challenge. For the diarization system, we proposed a hybrid approach combining a WavLM end-to-end segmentation method with a traditional multi-module clustering technique to adaptively select the appropriate model for handling varying degrees of overlapping speech. For the automatic speech recognition (ASR) system, we proposed an ASR-aware observation addition method that compensates for the performance limitations of Guided Source Separation (GSS) under low signal-to-noise ratio conditions. Finally, we integrated the speaker diarization and ASR systems in a cascaded architecture to address Track 3. Our system achieved character error rates (CER) of 9.48% on Track 2 and concatenated minimum permutation character error rate (cpCER) of 11.56% on Track 3, ultimately securing first place in both tracks and thereby demonstrating the effectiveness of the proposed methods in real-world meeting scenarios.


【19】 Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging  Recognition and Event Detection

标题: 利用LLM进行口吃语音:一种桥接识别和事件检测的统一架构
链接:https://arxiv.org/abs/2505.22005
作者: Shangkun Huang,  Jing Deng,  Jintao Kang,  Rong Zheng 
备注:Accepted to Interspeech 2025
摘要:自动语音识别(ASR)在口吃语音场景下的性能瓶颈限制了其在语音康复等领域的应用。提出了一种LLM驱动的ASR-SED多任务学习框架,该框架联合优化了ASR和口吃事件检测(SED)任务。我们提出了一种动态交互机制,其中ASR分支利用CTC生成的软提示来辅助LLM上下文建模,而SED分支输出口吃嵌入来增强LLM对口吃语音的理解。我们采用对比学习来加强口吃声学特征的辨别能力,并应用焦点损失来减轻口吃事件类别中的长尾分布。在AS-70汉语口吃数据集上的评估表明,我们的框架将ASR字符错误率(CER)降低到5.45%(相对减少-37.71%),并实现了平均73.63%的SED F1分数(相对改善+46.58%)。
摘要:The performance bottleneck of Automatic Speech Recognition (ASR) in stuttering speech scenarios has limited its applicability in domains such as speech rehabilitation. This paper proposed an LLM-driven ASR-SED multi-task learning framework that jointly optimized the ASR and Stuttering Event Detection (SED) tasks. We proposed a dynamic interaction mechanism where the ASR branch leveraged CTC-generated soft prompts to assist LLM context modeling, while the SED branch output stutter embeddings to enhance LLM comprehension of stuttered speech. We incorporated contrastive learning to strengthen the discriminative power of stuttering acoustic features and applied Focal Loss to mitigate the long-tailed distribution in stuttering event categories. Evaluations on the AS-70 Mandarin stuttering dataset demonstrated that our framework reduced the ASR character error rate (CER) to 5.45% (-37.71% relative reduction) and achieved an average SED F1-score of 73.63% (+46.58% relative improvement).


【20】 Music Source Restoration

标题: 音乐资源恢复
链接:https://arxiv.org/abs/2505.21827
作者: Yongyi Zang,  Zheqi Dai,  Mark D. Plumbley,  Qiuqiang Kong 
备注:A modified version of this paper is in review
摘要:我们介绍了音乐源恢复(MSR),一个新的任务,解决理想化的源分离和现实世界的音乐制作之间的差距。当前的音乐源分离(MSS)方法假设混合是源的简单和,忽略了在音乐制作期间使用的信号降级,如均衡、压缩和混响。MSR将混合物建模为单独退化源的退化总和,目标是恢复原始的、未退化的信号。由于缺乏MSR的数据,我们提出了RawStems,一个数据集注释的578首歌曲与未经处理的源信号组织成8个主要和17个次要乐器组,共计354.13小时。据我们所知,RawStems是第一个包含具有分层类别的未处理音乐词干的数据集。我们考虑频谱滤波,动态范围压缩,谐波失真,混响和有损编解码器作为可能的退化,并建立U形作为基线方法,证明了我们的数据集上的MSR的可行性。我们发布了RawStems数据集注释、退化模拟管道、训练代码和预训练模型,供公众使用。
摘要:We introduce Music Source Restoration (MSR), a novel task addressing the gap between idealized source separation and real-world music production. Current Music Source Separation (MSS) approaches assume mixtures are simple sums of sources, ignoring signal degradations employed during music production like equalization, compression, and reverb. MSR models mixtures as degraded sums of individually degraded sources, with the goal of recovering original, undegraded signals. Due to the lack of data for MSR, we present RawStems, a dataset annotation of 578 songs with unprocessed source signals organized into 8 primary and 17 secondary instrument groups, totaling 354.13 hours. To the best of our knowledge, RawStems is the first dataset that contains unprocessed music stems with hierarchical categories. We consider spectral filtering, dynamic range compression, harmonic distortion, reverb and lossy codec as possible degradations, and establish U-Former as a baseline method, demonstrating the feasibility of MSR on our dataset. We release the RawStems dataset annotations, degradation simulation pipeline, training code and pre-trained models to be publicly available.


【21】 Voice Quality Dimensions as Interpretable Primitives for Speaking Style  for Atypical Speech and Affect

标题: 语音质量维度作为非典型言语和情感说话风格的可解释基本要素
链接:https://arxiv.org/abs/2505.21809
作者: Jaya Narain,  Vasudha Kowtha,  Colin Lea,  Lauren Tooley,  Dianna Yee,  Vikramjit Mitra,  Zifang Huang,  Miquel Espi Marques,  Jon Huang,  Carlos Avendano,  Shirley Ren 
备注:accepted for Interspeech 2025
摘要:感知语音质量维度描述了非典型语音和其他语音调制的关键特征。在这里,我们开发和评估七个语音和语音维度(可懂度,不精确的辅音,刺耳的声音,自然度,单响度,呼吸困难和呼吸)的语音质量模型。探针在公共语音可访问性(SAP)项目数据集上进行了训练,该数据集包含来自434名扬声器的11,184个样本,使用来自冻结预训练模型的嵌入作为特征。我们发现,我们的探测器在SAP数据集中的语音启发类别中具有很强的性能和很强的泛化能力。我们进一步验证了其他数据集上的zero-shot性能,包括看不见的语言和任务:意大利语非典型语音,英语非典型语音和情感语音。强大的zero-shot性能和跨一系列评估结果的可解释性表明在说话风格相关任务中使用语音质量维度的实用性。
摘要:Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, imprecise consonants, harsh voice, naturalness, monoloudness, monopitch, and breathiness). Probes were trained on the public Speech Accessibility (SAP) project dataset with 11,184 samples from 434 speakers, using embeddings from frozen pre-trained models as features. We found that our probes had both strong performance and strong generalization across speech elicitation categories in the SAP dataset. We further validated zero-shot performance on additional datasets, encompassing unseen languages and tasks: Italian atypical speech, English atypical speech, and affective speech. The strong zero-shot performance and the interpretability of results across an array of evaluations suggests the utility of using voice quality dimensions in speaking style-related tasks.


【22】 An Investigation on Speaker Augmentation for End-to-End Speaker  Extraction

标题: 端到端说话人提取的说话人增强研究
链接:https://arxiv.org/abs/2505.21805
作者: Zhenghai You,  Zhenyu Zhou,  Lantian Li,  Dong Wang 
摘要:目标混淆是指偶尔切换到非目标说话人,是端到端说话人提取(E2 E-SE)系统面临的一个关键挑战。我们认为,这个问题在很大程度上是由于缺乏概括性和歧视性的说话人嵌入,并介绍了一个简单而有效的说话人增强策略来解决这个问题。具体来说,我们提出了一个时域的rescaling和重新缩放管道,改变扬声器的特点,同时保留其他语音属性。这会生成各种各样的伪说话者,以帮助建立一个可推广的说话者嵌入空间,而特定于说话者特征的增强会创建硬样本,迫使模型专注于真正的说话者特征。在WSJ 0 - 2 Mix和LibriMix上的实验表明,该方法有效地缓解了目标混淆,提高了提取性能。此外,它可以与度量学习相结合,这是解决目标混淆的另一种有效方法,从而带来进一步的收益。
摘要:Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and discrimination of the speaker embeddings, and introduce a simple yet effective speaker augmentation strategy to tackle the problem. Specifically, we propose a time-domain resampling and rescaling pipeline that alters speaker traits while preserving other speech properties. This generates a variety of pseudo-speakers to help establish a generalizable speaker embedding space, while the speaker-trait-specific augmentation creates hard samples that force the model to focus on genuine speaker characteristics. Experiments on WSJ0-2Mix and LibriMix show that our method mitigates the target confusion and improves extraction performance. Moreover, it can be combined with metric learning, another effective approach to address target confusion, leading to further gains.


【23】 Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech  Recognition Data for Research and Commercial Use

标题: Loquacious Set:25,000小时的转录和多样化的英语语音识别数据,用于研究和商业用途
链接:https://arxiv.org/abs/2505.21578
作者: Titouan Parcollet,  Yuan Tseng,  Shucong Zhang,  Rogier van Dalen 
备注:Accepted at Interspeech 2025
摘要:自动语音识别(ASR)研究是由工业研究人员和学术界之间的通用数据集驱动的,鼓励进行比较和评估。尽管LibriSpeech作为ASR基准取得了长期的成功,但现在受到其规模的限制,并专注于干净,阅读语音,导致单词错误率接近零。最近的数据集,包括MOSEL,YODAS,Gigaspeech,OWSM,Libriheavy或People's Speech,都受到重大限制,包括行业研究人员无法使用的许可证,不可靠的传输,不正确的音频数据或缺乏评估集。这部作品展示了一个长达25,000小时的商业英语演讲集。Loquacious Set拥有数十万具有不同口音和各种语音类型(阅读,自发,谈话,干净,嘈杂)的扬声器,旨在为业内学者和研究人员在现实世界中构建ASR系统。
摘要:Automatic speech recognition (ASR) research is driven by the availability of common datasets between industrial researchers and academics, encouraging comparisons and evaluations. LibriSpeech, despite its long success as an ASR benchmark, is now limited by its size and focus on clean, read speech, leading to near-zero word error rates. More recent datasets, including MOSEL, YODAS, Gigaspeech, OWSM, Libriheavy or People's Speech suffer from major limitations including licenses that researchers in the industry cannot use, unreliable transcriptions, incorrect audio data, or the lack of evaluation sets. This work presents the Loquacious Set, a 25,000-hour curated collection of commercially usable English speech. Featuring hundreds of thousands of speakers with diverse accents and a wide range of speech types (read, spontaneous, talks, clean, noisy), the Loquacious Set is designed to work for academics and researchers in the industry to build ASR systems in real-world scenarios.


【24】 VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach  Leveraging Speaker-Specific Latents

标题: VoiceMark:一种利用特定说话人特征的抗零采样语音克隆水印方法
链接:https://arxiv.org/abs/2505.21568
作者: Haiyun Li,  Zhiyong Wu,  Xiaofeng Xie,  Jingran Xie,  Yaoxun Xu,  Hanyang Peng 
备注:Accepted by Interspeech 2025
摘要:抗语音克隆水印是一种新兴的跟踪和防止未授权克隆的技术。现有的方法通过在带水印的音频上训练传统的VC模型来有效地跟踪传统的VC模型,但是在zero-shot VC场景中失败,其中模型从音频提示合成音频而不进行训练。为了解决这个问题,我们提出了VoiceMark,第一个zero-shot抗VC水印方法,利用特定于说话人的潜伏期作为水印载体,允许水印通过zero-shot VC过程转移到合成音频中。此外,我们引入VC模拟的增强和基于VAD的损失,以提高对失真的鲁棒性。在多个zero-shot VC模型上的实验表明,经过zero-shot VC合成后,VoiceMark的水印检测准确率达到了95%以上,明显优于现有的只能达到50%左右的方法。查看我们的代码和演示:https://huggingface.co/spaces/haiyunli/VoiceMark
摘要:Voice cloning (VC)-resistant watermarking is an emerging technique for tracing and preventing unauthorized cloning. Existing methods effectively trace traditional VC models by training them on watermarked audio but fail in zero-shot VC scenarios, where models synthesize audio from an audio prompt without training. To address this, we propose VoiceMark, the first zero-shot VC-resistant watermarking method that leverages speaker-specific latents as the watermark carrier, allowing the watermark to transfer through the zero-shot VC process into the synthesized audio. Additionally, we introduce VC-simulated augmentations and VAD-based loss to enhance robustness against distortions. Experiments on multiple zero-shot VC models demonstrate that VoiceMark achieves over 95% accuracy in watermark detection after zero-shot VC synthesis, significantly outperforming existing methods, which only reach around 50%. See our code and demos at: https://huggingface.co/spaces/haiyunli/VoiceMark


【25】 Articulatory modeling of the S-shaped F2 trajectories observed in  Öhman's spectrographic analysis of VCV syllables

标题: Öhman对VCV音节的光谱分析中观察到的S形F2轨迹的关节建模
链接:https://arxiv.org/abs/2505.22455
作者: Frédéric Berthommier 
备注:5 pages, 4 figures, submitted to Interspeech 2025
摘要:Ohman的VCV序列与元音间爆破辅音的合成在30年前首次使用DRM模型实现。然而,这种方法仍然主要是声学的,缺乏发音的限制。在这项研究中,相同的75 VCV进行了分析,但产生的前田模型,使用轨迹规划,区分元音到元音的过渡辅音的影响。合成数据表现出与Ohman序列相似的特征,包括S形F2轨迹的存在。此外,轨迹方程(LE)为F2和F3计算合成CV数据调查其潜在的决定性,导致重新评估传统的解释。研究结果表明,虽然发音规划的结构分别为元音和辅音组,S形F2轨迹出现从一个复合机制所管辖的协调协同作用的所有发音。
摘要:The synthesis of Ohman's VCV sequences with intervocalic plosive consonants was first achieved 30 years ago using the DRM model. However, this approach remains primarily acoustic and lacks articulatory constraints. In this study, the same 75 VCVs are analyzed, but generated with the Maeda model, using trajectory planning that differentiates vowel-to-vowel transitions from consonantal influences. Synthetic data exhibit similar characteristics to Ohman's sequences, including the presence of S-shaped F2 trajectories. Furthermore, locus equations (LEs) for F2 and F3 are computed from synthetic CV data to investigate their underlying determinism, leading to a reassessment of conventional interpretations. The findings indicate that, although articulatory planning is structured separately for vowel and consonant groups, S-shaped F2 trajectories emerge from a composite mechanism governed by the coordinated synergy of all articulators.


【26】 Analysis and Evaluation of Synthetic Data Generation in Speech  Dysfluency Detection

标题: 语音不流畅检测中合成数据生成的分析与评估
链接:https://arxiv.org/abs/2505.22029
作者: Jinming Zhang,  Xuanru Zhou,  Jiachen Lian,  Shuhe Li,  William Li,  Zoe Ezzes,  Rian Bogley,  Lisa Wauters,  Zachary Miller,  Jet Vonk,  Brittany Morin,  Maria Gorno-Tempini,  Gopala Anumanchipalli 
备注:Submitted to Interspeech 2025
摘要:言语不流利检测对于临床诊断和语言评估至关重要,但现有方法受到缺乏高质量注释数据的限制。虽然TTS模型的最新进展已经实现了合成不流利生成,但现有的合成数据集存在不自然的韵律和有限的上下文多样性。为了解决这些局限性,我们提出了LLM-Dys -最全面的不流利语音语料库与LLM增强的不流利模拟。该数据集捕获了11个跨单词和音素水平的不流利类别。在此资源的基础上,我们改进了端到端的不流利检测框架。实验验证证明了最先进的性能。所有数据、模型和代码都在https://github.com/Berkeley-Speech-Group/LLM-Dys上开源。
摘要:Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency generation, existing synthetic datasets suffer from unnatural prosody and limited contextual diversity. To address these limitations, we propose LLM-Dys -- the most comprehensive dysfluent speech corpus with LLM-enhanced dysfluency simulation. This dataset captures 11 dysfluency categories spanning both word and phoneme levels. Building upon this resource, we improve an end-to-end dysfluency detection framework. Experimental validation demonstrates state-of-the-art performance. All data, models, and code are open-sourced at https://github.com/Berkeley-Speech-Group/LLM-Dys.


【27】 WhisperD: Dementia Speech Recognition and Filler Word Detection with  Whisper

标题: WhisperD:使用Whisper的痴呆症语音识别和填充词检测
链接:https://arxiv.org/abs/2505.21551
作者: Emmanuel Akinrintoyo,  Nadine Abdelhalim,  Nicole Salomons 
备注:Submitted to Interspeech 2025 (Accepted)
摘要:Whisper无法正确转录痴呆症语音,因为痴呆症患者(PwDs)经常表现出不规则的语音模式和不流利,如停顿,重复和支离破碎的句子。它接受了标准言语训练,可能很少或根本没有接触过受痴呆症影响的言语。然而,正确的转录是至关重要的痴呆症的语音成本效益的诊断和辅助技术的发展。在这项工作中,我们使用开源痴呆症语音数据集(DementiaBank)和我们的内部数据集微调Whisper,以提高其单词错误率(WER)。微调还包括填充词,以确定填充词包含率(FIR)和F1分数。微调后的模型明显优于现成的模型。中型模型的WER为0.24,优于以前的工作。同样,对于看不见的数据和语音模式也有显著的普遍性。
摘要:Whisper fails to correctly transcribe dementia speech because persons with dementia (PwDs) often exhibit irregular speech patterns and disfluencies such as pauses, repetitions, and fragmented sentences. It was trained on standard speech and may have had little or no exposure to dementia-affected speech. However, correct transcription is vital for dementia speech for cost-effective diagnosis and the development of assistive technology. In this work, we fine-tune Whisper with the open-source dementia speech dataset (DementiaBank) and our in-house dataset to improve its word error rate (WER). The fine-tuning also includes filler words to ascertain the filler inclusion rate (FIR) and F1 score. The fine-tuned models significantly outperformed the off-the-shelf models. The medium-sized model achieved a WER of 0.24, outperforming previous work. Similarly, there was a notable generalisability to unseen data and speech patterns.


【28】 VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled  data and Large-Scale Speech Pretraining

标题: VietASB:通过50小时标记数据和大规模语音预训练实现行业级越南语ASB
链接:https://arxiv.org/abs/2505.21527
作者: Jianheng Zhuo,  Yifan Yang,  Yiwen Shao,  Yong Xu,  Dong Yu,  Kai Yu,  Xie Chen 
摘要:自动语音识别(ASR)已经取得了显着的进展,但严重依赖于大规模的标记数据,这对于像越南语这样的低资源语言来说是稀缺的。虽然Whisper、USM和MMS等现有系统实现了令人鼓舞的性能,但在培训成本、延迟和可访问性方面,它们的功效仍然不足。为了解决这些问题,我们提出了VietASR,这是一种新型的ASR训练管道,它利用了大量的未标记数据和一小部分标记数据。通过在大规模未标记数据集上进行多迭代ASR偏置自监督学习,VietASR为增强ASR性能提供了一种具有成本效益和实用性的解决方案。实验表明,对70,000小时的未标记数据进行预训练并对仅50小时的标记数据进行微调,可以产生轻量级但功能强大的ASR模型。它在实际数据上优于Whisper Large-v3和商业ASR系统。我们的代码和模型将开源,以促进低资源ASR的研究。
摘要:Automatic speech recognition (ASR) has made remarkable progress but heavily relies on large-scale labeled data, which is scarce for low-resource languages like Vietnamese. While existing systems such as Whisper, USM, and MMS achieve promising performance, their efficacy remains inadequate in terms of training costs, latency, and accessibility. To address these issues, we propose VietASR, a novel ASR training pipeline that leverages vast amounts of unlabeled data and a small set of labeled data. Through multi-iteration ASR-biased self-supervised learning on a large-scale unlabeled dataset, VietASR offers a cost-effective and practical solution for enhancing ASR performance. Experiments demonstrate that pre-training on 70,000-hour unlabeled data and fine-tuning on merely 50-hour labeled data yield a lightweight but powerful ASR model. It outperforms Whisper Large-v3 and commercial ASR systems on real-world data. Our code and models will be open-sourced to facilitate research in low-resource ASR.


eess.AS音频处理


【1】 Articulatory modeling of the S-shaped F2 trajectories observed in  Öhman's spectrographic analysis of VCV syllables

标题: Öhman对VCV音节的光谱分析中观察到的S形F2轨迹的关节建模
链接:https://arxiv.org/abs/2505.22455
作者: Frédéric Berthommier 
备注:5 pages, 4 figures, submitted to Interspeech 2025
摘要:Ohman的VCV序列与元音间爆破辅音的合成在30年前首次使用DRM模型实现。然而,这种方法仍然主要是声学的,缺乏发音的限制。在这项研究中,相同的75 VCV进行了分析,但产生的前田模型,使用轨迹规划,区分元音到元音的过渡辅音的影响。合成数据表现出与Ohman序列相似的特征,包括S形F2轨迹的存在。此外,轨迹方程(LE)为F2和F3计算合成CV数据调查其潜在的决定性,导致重新评估传统的解释。研究结果表明,虽然发音规划的结构分别为元音和辅音组,S形F2轨迹出现从一个复合机制所管辖的协调协同作用的所有发音。
摘要:The synthesis of Ohman's VCV sequences with intervocalic plosive consonants was first achieved 30 years ago using the DRM model. However, this approach remains primarily acoustic and lacks articulatory constraints. In this study, the same 75 VCVs are analyzed, but generated with the Maeda model, using trajectory planning that differentiates vowel-to-vowel transitions from consonantal influences. Synthetic data exhibit similar characteristics to Ohman's sequences, including the presence of S-shaped F2 trajectories. Furthermore, locus equations (LEs) for F2 and F3 are computed from synthetic CV data to investigate their underlying determinism, leading to a reassessment of conventional interpretations. The findings indicate that, although articulatory planning is structured separately for vowel and consonant groups, S-shaped F2 trajectories emerge from a composite mechanism governed by the coordinated synergy of all articulators.


【2】 Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in  Large Language Models for Speech Recognition

标题: 语音中LLM的评估经常存在缺陷:语音识别大型语言模型中的测试集污染
链接:https://arxiv.org/abs/2505.22251
作者: Yuan Tseng,  Titouan Parcollet,  Rogier van Dalen,  Shucong Zhang,  Sourav Bhattacharya 
摘要:最近的工作表明,与现有系统相比,大型语言模型(LLM)可以提高语音任务的性能。为了支持他们的主张,LibriSpeech和Common Voice的结果经常被引用。然而,这项工作发现,大量的LibriSpeech和Common Voice评估集出现在公共LLM预训练语料库中。这就对这两个数据集得出的结论的可靠性提出了质疑。为了测量污染的影响,比较了有或没有污染的LLM训练,表明受污染的LLM更有可能生成它在训练期间看到的测试句子。使用受污染的LLM的语音识别器在错误率上仅显示出细微的差异,但在训练期间看到的transmittance分配了显着更高的概率。结果表明,LLM的输出可能会受到少量数据污染的影响,这突出了使用保留数据评估基于LLM的语音系统的重要性。
摘要:Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Common Voice are often quoted. However, this work finds that a substantial amount of the LibriSpeech and Common Voice evaluation sets appear in public LLM pretraining corpora. This calls into question the reliability of findings drawn from these two datasets. To measure the impact of contamination, LLMs trained with or without contamination are compared, showing that a contaminated LLM is more likely to generate test sentences it has seen during training. Speech recognisers using contaminated LLMs shows only subtle differences in error rates, but assigns significantly higher probabilities to transcriptions seen during training. Results show that LLM outputs can be biased by tiny amounts of data contamination, highlighting the importance of evaluating LLM-based speech systems with held-out data.


【3】 ARiSE: Auto-Regressive Multi-Channel Speech Enhancement

标题: ARiSE:自回归多通道语音增强
链接:https://arxiv.org/abs/2505.22051
作者: Pengjie Shen,  Xueliang Zhang,  Zhong-Qiu Wang 
摘要:我们提出了ARiSE,一个自回归算法的多通道语音增强。ARiSE通过引入自回归连接来改进现有的基于深度神经网络(DNN)的帧在线多通道语音增强模型,其中利用先前帧处的估计目标语音作为额外的输入特征来帮助DNN估计当前帧处的目标语音。额外输入特征可以从(a)先前帧中的估计目标语音;以及(b)基于先前估计目标语音计算的具有波束形成器的波束形成混合物中导出。另一方面,以自回归的方式天真地训练DNN是非常缓慢的。为了解决这个问题,我们提出了一个平行的培训机制,以加快培训。噪声混响条件下的评估结果表明,所提出的算法的有效性和潜力。
摘要:We propose ARiSE, an auto-regressive algorithm for multi-channel speech enhancement. ARiSE improves existing deep neural network (DNN) based frame-online multi-channel speech enhancement models by introducing auto-regressive connections, where the estimated target speech at previous frames is leveraged as extra input features to help the DNN estimate the target speech at the current frame. The extra input features can be derived from (a) the estimated target speech in previous frames; and (b) a beamformed mixture with the beamformer computed based on the previous estimated target speech. On the other hand, naively training the DNN in an auto-regressive manner is very slow. To deal with this, we propose a parallel training mechanism to speed up the training. Evaluation results in noisy-reverberant conditions show the effectiveness and potential of the proposed algorithms.


【4】 Analysis and Evaluation of Synthetic Data Generation in Speech  Dysfluency Detection

标题: 语音不流畅检测中合成数据生成的分析与评估
链接:https://arxiv.org/abs/2505.22029
作者: Jinming Zhang,  Xuanru Zhou,  Jiachen Lian,  Shuhe Li,  William Li,  Zoe Ezzes,  Rian Bogley,  Lisa Wauters,  Zachary Miller,  Jet Vonk,  Brittany Morin,  Maria Gorno-Tempini,  Gopala Anumanchipalli 
备注:Submitted to Interspeech 2025
摘要:言语不流利检测对于临床诊断和语言评估至关重要,但现有方法受到缺乏高质量注释数据的限制。尽管TTS模型的最新进展已经实现了合成不流利生成,但现有的合成数据集存在不自然的韵律和有限的上下文多样性。为了解决这些局限性,我们提出了LLM-Dys -最全面的不流利语音语料库与LLM增强的不流利模拟。该数据集捕获了11个跨单词和音素水平的不流利类别。在此资源的基础上,我们改进了端到端的不流利检测框架。实验验证证明了最先进的性能。所有数据、模型和代码都在https://github.com/Berkeley-Speech-Group/LLM-Dys上开源。
摘要:Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency generation, existing synthetic datasets suffer from unnatural prosody and limited contextual diversity. To address these limitations, we propose LLM-Dys -- the most comprehensive dysfluent speech corpus with LLM-enhanced dysfluency simulation. This dataset captures 11 dysfluency categories spanning both word and phoneme levels. Building upon this resource, we improve an end-to-end dysfluency detection framework. Experimental validation demonstrates state-of-the-art performance. All data, models, and code are open-sourced at https://github.com/Berkeley-Speech-Group/LLM-Dys.


【5】 WhisperD: Dementia Speech Recognition and Filler Word Detection with  Whisper

标题: WhisperD:使用Whisper的痴呆症语音识别和填充词检测
链接:https://arxiv.org/abs/2505.21551
作者: Emmanuel Akinrintoyo,  Nadine Abdelhalim,  Nicole Salomons 
备注:Submitted to Interspeech 2025 (Accepted)
摘要:Whisper无法正确转录痴呆症语音,因为痴呆症患者(PwDs)经常表现出不规则的语音模式和不流利,如停顿,重复和支离破碎的句子。它接受过标准语言训练,可能很少或根本没有接触过受痴呆症影响的语言。然而,正确的转录是至关重要的痴呆症的语音成本效益的诊断和辅助技术的发展。在这项工作中,我们使用开源痴呆症语音数据集(DementiaBank)和我们的内部数据集微调Whisper,以提高其单词错误率(WER)。微调还包括填充词,以确定填充词包含率(FIR)和F1分数。微调后的模型明显优于现成的模型。中型模型的WER为0.24,优于以前的工作。同样,对于看不见的数据和语音模式也有显著的普遍性。
摘要:Whisper fails to correctly transcribe dementia speech because persons with dementia (PwDs) often exhibit irregular speech patterns and disfluencies such as pauses, repetitions, and fragmented sentences. It was trained on standard speech and may have had little or no exposure to dementia-affected speech. However, correct transcription is vital for dementia speech for cost-effective diagnosis and the development of assistive technology. In this work, we fine-tune Whisper with the open-source dementia speech dataset (DementiaBank) and our in-house dataset to improve its word error rate (WER). The fine-tuning also includes filler words to ascertain the filler inclusion rate (FIR) and F1 score. The fine-tuned models significantly outperformed the off-the-shelf models. The medium-sized model achieved a WER of 0.24, outperforming previous work. Similarly, there was a notable generalisability to unseen data and speech patterns.


【6】 VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled  data and Large-Scale Speech Pretraining

标题: VietASB:通过50小时标记数据和大规模语音预训练实现行业级越南语ASB
链接:https://arxiv.org/abs/2505.21527
作者: Jianheng Zhuo,  Yifan Yang,  Yiwen Shao,  Yong Xu,  Dong Yu,  Kai Yu,  Xie Chen 
摘要:自动语音识别(ASR)已经取得了显着的进展,但严重依赖于大规模的标记数据,这对于像越南语这样的低资源语言来说是稀缺的。虽然Whisper、USM和MMS等现有系统实现了令人鼓舞的性能,但在培训成本、延迟和可访问性方面,它们的功效仍然不足。为了解决这些问题,我们提出了VietASR,这是一种新型的ASR训练管道,它利用了大量的未标记数据和一小部分标记数据。通过在大规模未标记数据集上进行多迭代ASR偏置自监督学习,VietASR为增强ASR性能提供了一种具有成本效益和实用性的解决方案。实验表明,对70,000小时的未标记数据进行预训练,并对仅50小时的标记数据进行微调,可以产生轻量级但功能强大的ASR模型。它在实际数据上优于Whisper Large-v3和商业ASR系统。我们的代码和模型将开源,以促进低资源ASR的研究。
摘要:Automatic speech recognition (ASR) has made remarkable progress but heavily relies on large-scale labeled data, which is scarce for low-resource languages like Vietnamese. While existing systems such as Whisper, USM, and MMS achieve promising performance, their efficacy remains inadequate in terms of training costs, latency, and accessibility. To address these issues, we propose VietASR, a novel ASR training pipeline that leverages vast amounts of unlabeled data and a small set of labeled data. Through multi-iteration ASR-biased self-supervised learning on a large-scale unlabeled dataset, VietASR offers a cost-effective and practical solution for enhancing ASR performance. Experiments demonstrate that pre-training on 70,000-hour unlabeled data and fine-tuning on merely 50-hour labeled data yield a lightweight but powerful ASR model. It outperforms Whisper Large-v3 and commercial ASR systems on real-world data. Our code and models will be open-sourced to facilitate research in low-resource ASR.


【7】 Effective and Efficient One-pass Compression of Speech Foundation Models  Using Sparsity-aware Self-pinching Gates

标题: 使用稀疏感知自挤压门对语音基础模型进行有效且高效的单程压缩
链接:https://arxiv.org/abs/2505.22608
作者: Haoning Xu,  Zhaoqing Li,  Youjun Chen,  Huimeng Wang,  Guinan Li,  Mengzhe Geng,  Chengxi Deng,  Xunying Liu 
备注:Submitted to Interspeech 2025
摘要:本文提出了一种新的语音基础模型压缩方法,紧密结合模型修剪和参数更新到一个单一的阶段。高度紧凑的层级绑定自箍缩门,每个门只包含一个可学习的阈值,与未压缩的模型联合训练,并用于细粒度的神经元级修剪。在LibriSpeech-100 hr语料库上进行的实验表明,我们的方法将wav2vec2.0-base和HuBERT-large模型的参数数量分别减少了65%和60%,同时在测试干净的数据集上没有产生统计上显著的单词错误率(WER)增加。与以前发表的相同任务的方法相比,我们的方法不仅在4.26倍的可比模型压缩比下在测试干净数据集上实现了7.05%的最低WER,而且模型压缩时间至少减少了25%。
摘要:This paper presents a novel approach for speech foundation models compression that tightly integrates model pruning and parameter update into a single stage. Highly compact layer-level tied self-pinching gates each containing only a single learnable threshold are jointly trained with uncompressed models and used in fine-grained neuron level pruning. Experiments conducted on the LibriSpeech-100hr corpus suggest that our approach reduces the number of parameters of wav2vec2.0-base and HuBERT-large models by 65% and 60% respectively, while incurring no statistically significant word error rate (WER) increase on the test-clean dataset. Compared to previously published methods on the same task, our approach not only achieves the lowest WER of 7.05% on the test-clean dataset under a comparable model compression ratio of 4.26x, but also operates with at least 25% less model compression time.


【8】 Towards General Discrete Speech Codec for Complex Acoustic Environments:  A Study of Reconstruction and Downstream Task Consistency

标题: 面向复杂声学环境的通用离散语音编解码器:重建和下游任务一致性研究
链接:https://arxiv.org/abs/2505.22515
作者: Haoran Wang,  Guanyu Chen,  Bohan Li,  Hankun Wang,  Yiwei Guo,  Zhihan Li,  Xie Chen,  Kai Yu 
备注:Initial Upload
摘要:神经语音编解码器在重建干净的语音信号方面表现出色;然而,它们在复杂声学环境和下游信号处理任务中的功效仍然有待探索。在这项研究中,我们引入了一个新的基准称为环境弹性语音编解码器基准(ERSB),系统地评估神经语音编解码器是否具有环境弹性。具体来说,我们评估了两个关键能力:(1)鲁棒重建,它测量语音和非语音声学细节的保留,以及(2)下游任务一致性,它确保在使用重建语音而不是原始语音时下游信号处理任务的偏差最小。我们的综合实验表明,复杂的声学环境显着降低信号重建和下游任务的一致性。这项工作突出了当前语音编解码器的局限性,并提出了一个未来的方向,以改善它们,提高环境适应力。
摘要:Neural speech codecs excel in reconstructing clean speech signals; however, their efficacy in complex acoustic environments and downstream signal processing tasks remains underexplored. In this study, we introduce a novel benchmark named Environment-Resilient Speech Codec Benchmark (ERSB) to systematically evaluate whether neural speech codecs are environment-resilient. Specifically, we assess two key capabilities: (1) robust reconstruction, which measures the preservation of both speech and non-speech acoustic details, and (2) downstream task consistency, which ensures minimal deviation in downstream signal processing tasks when using reconstructed speech instead of the original. Our comprehensive experiments reveal that complex acoustic environments significantly degrade signal reconstruction and downstream task consistency. This work highlights the limitations of current speech codecs and raises a future direction that improves them for greater environmental resilience.


【9】 Effective Context in Neural Speech Models

标题: 神经语音模型中的有效上下文
链接:https://arxiv.org/abs/2505.22487
作者: Yen Meng,  Sharon Goldwater,  Hao Tang 
备注:Accepted to Interspeech 2025
摘要:现代神经语音模型受益于更长的上下文,并且已经提出了许多方法来增加模型可以使用的最大上下文。然而,很少有人试图衡量这些模型实际使用了多少上下文,即,有效的上下文。在这里,我们提出了两种测量有效语境的方法,并用它们来分析不同的言语Transformers。对于监督模型,我们发现有效上下文与任务的性质相关,基频跟踪,电话分类和单词分类需要越来越多的有效上下文。对于自监督模型,我们发现有效上下文主要在早期层中增加,并且保持相对较短-类似于监督电话模型。鉴于这些模型在预测过程中不使用长上下文,我们证明了HuBERT可以在流模式下运行,而无需修改架构,也无需进一步微调。
摘要:Modern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e., the effective context. Here, we propose two approaches to measuring the effective context, and use them to analyze different speech Transformers. For supervised models, we find that the effective context correlates well with the nature of the task, with fundamental frequency tracking, phone classification, and word classification requiring increasing amounts of effective context. For self-supervised models, we find that effective context increases mainly in the early layers, and remains relatively short -- similar to the supervised phone model. Given that these models do not use a long context during prediction, we show that HuBERT can be run in streaming mode without modification to the architecture and without further fine-tuning.


【10】 FGS-Audio: Fixed-Decoder Framework for Audio Steganography with  Adversarial Perturbation Generation

标题: FGS-音频:具有对抗性微扰生成的音频隐写术的固定解码器框架
链接:https://arxiv.org/abs/2505.22266
作者: Jialin Yan,  Yu Cheng,  Zhaoxia Yin,  Xinpeng Zhang,  Shilin Wang,  Tanfeng Sun,  Xinghao Jiang 
摘要:人工智能生成内容(AIGC)的快速发展使得高保真生成的音频在互联网上广泛可用,为隐蔽通信提供了丰富而通用的掩护信号源。在深度学习的推动下,当前的音频隐写框架主要基于编解码网络架构。虽然这些方法大大提高了音频隐写的安全性,但它们通常采用精心设计的训练工作流程,并依赖于广泛的预训练模型。为了解决上述问题,本文开创性地提出了一种基于对抗扰动生成的音频隐写固定解码器框架(FGS-Audio)。携带秘密信息的对抗性扰动被嵌入到覆盖音频中以生成隐写音频。接收方只需共享固定解码网络的结构和权值,就能准确地从隐写音频中提取秘密信息,从而消除了对大型预训练模型的依赖。在FGS-Audio中,我们提出了一种音频对抗扰动生成(APG)策略,并设计了一个轻量级的固定解码器。固定的解码器保证可靠地提取隐藏的消息,而对抗性扰动进行了优化,以保持隐写音频感知和统计接近的封面音频,从而提高抗隐写分析。实验结果表明,该方法在不同的相对负载下具有良好的抗隐写分析性能,优于现有的SOTA方法。在隐写音频质量方面,与SOTA方法相比,FGS-Audio实现了超过10 dB的平均PSNR改善。
摘要:The rapid development of Artificial Intelligence Generated Content (AIGC) has made high-fidelity generated audio widely available across the Internet, offering an abundant and versatile source of cover signals for covert communication. Driven by advances in deep learning, current audio steganography frameworks are mainly based on encoding-decoding network architectures. While these methods greatly improve the security of audio steganography, they typically employ elaborate training workflows and rely on extensive pre-trained models. To address the aforementioned issues, this paper pioneers a Fixed-Decoder Framework for Audio Steganography with Adversarial Perturbation Generation (FGS-Audio). The adversarial perturbations that carry secret information are embedded into cover audio to generate stego audio. The receiver only needs to share the structure and weights of the fixed decoding network to accurately extract the secret information from the stego audio, thus eliminating the reliance on large pre-trained models. In FGS-Audio, we propose an audio Adversarial Perturbation Generation (APG) strategy and design a lightweight fixed decoder. The fixed decoder guarantees reliable extraction of the hidden message, while the adversarial perturbations are optimized to keep the stego audio perceptually and statistically close to the cover audio, thereby improving resistance to steganalysis. The experimental results show that the method exhibits excellent anti-steganalysis performance under different relative payloads, outperforming existing SOTA approaches. In terms of stego audio quality, FGS-Audio achieves an average PSNR improvement of over 10 dB compared to SOTA method.


【11】 Advancing Hearing Assessment: An ASR-Based Frequency-Specific Speech  Test for Diagnosing Presbycusis

标题: 高级听力评估:用于诊断老年性耳聋的基于SVR的特定频率言语测试
链接:https://arxiv.org/abs/2505.22231
作者: Stefan Bleeck 
摘要:传统的测听通常无法完全表征听力损失对言语理解的功能影响,特别是在老年性耳聋等情况下的阈上缺陷和频率特异性感知挑战。本文介绍了一种新的自动语音识别(ASR)为基础的频率特定的语音测试,旨在提供粒度诊断的见解的发展和模拟评估。我们的方法利用ASR来模拟中度倾斜听力损失的感知效果,通过在受控的声学退化下处理语音刺激,随后分析音素级混淆模式。主要研究结果表明,模拟听力损失会引入特定的音素混淆,主要影响高频辅音(例如,齿槽/腭到唇齿替换)并导致显著的音素缺失,与老年性聋中的声学线索退化一致。从这些ASR衍生的混淆中策划的测试电池显示了诊断价值,在综合模拟中有效区分了模拟的正常听力和听力受损的听众。这种ASR驱动的方法为开发客观,粒度和频率特定的听力评估工具提供了一个有前途的途径,这些工具补充了传统的听力测试。未来的工作将集中在与人类参与者一起验证这些发现,并探索先进人工智能模型的整合,以提高诊断精度。
摘要:Traditional audiometry often fails to fully characterize the functional impact of hearing loss on speech understanding, particularly supra-threshold deficits and frequency-specific perception challenges in conditions like presbycusis. This paper presents the development and simulated evaluation of a novel Automatic Speech Recognition (ASR)-based frequency-specific speech test designed to provide granular diagnostic insights. Our approach leverages ASR to simulate the perceptual effects of moderate sloping hearing loss by processing speech stimuli under controlled acoustic degradation and subsequently analyzing phoneme-level confusion patterns. Key findings indicate that simulated hearing loss introduces specific phoneme confusions, predominantly affecting high-frequency consonants (e.g., alveolar/palatal to labiodental substitutions) and leading to significant phoneme deletions, consistent with the acoustic cues degraded in presbycusis. A test battery curated from these ASR-derived confusions demonstrated diagnostic value, effectively differentiating between simulated normal-hearing and hearing-impaired listeners in a comprehensive simulation. This ASR-driven methodology offers a promising avenue for developing objective, granular, and frequency-specific hearing assessment tools that complement traditional audiometry. Future work will focus on validating these findings with human participants and exploring the integration of advanced AI models for enhanced diagnostic precision.


【12】 Two-stage Audio-Visual Target Speaker Extraction System for Real-Time  Processing On Edge Device

标题: 边缘设备实时处理的两级视听目标说话人提取系统
链接:https://arxiv.org/abs/2505.22229
作者: Zixuan Li,  Xueliang Zhang,  Lei Miao,  Zhipeng Yan 
摘要:视听目标说话人提取(AVTSE)的目的是在多说话人环境中,以视觉线索为辅助,分离出目标说话人的声音。大多数现有的AVTSE方法同时对视觉和音频特征进行编码,导致极高的计算复杂度,并使其无法在边缘设备上进行实时处理。为了解决这个问题,我们提出了一个两级超紧凑型AVTSE系统。具体地,在第一阶段中,采用紧凑的网络用于使用视觉信息的语音活动检测(VAD)。在第二阶段,VAD结果与音频输入相结合,以隔离目标说话者的语音。实验表明,该系统有效地抑制了背景噪声和干扰语音,而花费很少的计算资源。
摘要:Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously, resulting in extremely high computational complexity and making it impractical for real-time processing on edge devices. To tackle this issue, we proposed a two-stage ultra-compact AVTSE system. Specifically, in the first stage, a compact network is employed for voice activity detection (VAD) using visual information. In the second stage, the VAD results are combined with audio inputs to isolate the target speaker's voice. Experiments show that the proposed system effectively suppresses background noise and interfering voices while spending little computational resources.


【13】 Developing a Top-tier Framework in Naturalistic Conditions Challenge for  Categorized Emotion Prediction: From Speech Foundation Models and Learning  Objective to Data Augmentation and Engineering Choices

标题: 在自然条件挑战中开发顶级框架以实现分类情绪预测:从言语基础模型和学习目标到数据增强和工程选择
链接:https://arxiv.org/abs/2505.22133
作者: Tiantian Feng,  Thanathai Lertpetchpun,  Dani Byrd,  Shrikanth Narayanan 
备注:Accepted to INTERSPEECH 2025
摘要:语音情感识别(SER),特别是自然表达的情感,仍然是一个具有挑战性的计算任务。主要挑战包括情感标注中固有的主观性和数据集中情感标签的不平衡分布。本文介绍了为参加INTERSPEECH 2025情感识别挑战赛(任务1)而开发的\texttt{SAILER}系统。挑战数据集包含来自播客的自然情感语音,是研究不平衡和主观情感注释的宝贵资源。我们的系统设计简单,可重复,有效,突出了建模,学习目标,数据增强和工程选择的关键选择。结果表明,即使是一个单一的系统(没有集成)可以超过95%的提交,宏观F1得分超过0.4。此外,三个系统的组合进一步提高了性能,获得了竞争性排名分数(表现前三的团队)。我们的模型在:https://github.com/tiantiaf0627/vox-profile-release。
摘要:Speech emotion recognition (SER), particularly for naturally expressed emotions, remains a challenging computational task. Key challenges include the inherent subjectivity in emotion annotation and the imbalanced distribution of emotion labels in datasets. This paper introduces the \texttt{SAILER} system developed for participation in the INTERSPEECH 2025 Emotion Recognition Challenge (Task 1). The challenge dataset, which contains natural emotional speech from podcasts, serves as a valuable resource for studying imbalanced and subjective emotion annotations. Our system is designed to be simple, reproducible, and effective, highlighting critical choices in modeling, learning objectives, data augmentation, and engineering choices. Results show that even a single system (without ensembling) can outperform more than 95\% of the submissions, with a Macro-F1 score exceeding 0.4. Moreover, an ensemble of three systems further improves performance, achieving a competitively ranked score (top-3 performing team). Our model is at: https://github.com/tiantiaf0627/vox-profile-release.


【14】 AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

标题: AudioTurbo:具有矫正扩散的快速文本到音频生成
链接:https://arxiv.org/abs/2505.22106
作者: Junqi Zhao,  Jinzheng Zhao,  Haohe Liu,  Yun Chen,  Lu Han,  Xubo Liu,  Mark Plumbley,  Wenwu Wang 
摘要:扩散模型显着提高了音频生成的质量和多样性,但受到推理速度慢的阻碍。修正流通过学习直线常微分方程(ODE)路径来提高推理速度。然而,这种方法需要从头开始训练流匹配模型,并且在低步数下往往表现不佳,甚至表现不佳。为了解决整流流的局限性,同时利用先进的预训练扩散模型的优势,本研究将预训练模型与整流扩散方法相结合,以提高文本到音频(TTA)生成的效率。具体来说,我们提出了AudioTurbo,它从预先训练的TTA模型生成的确定性噪声样本对中学习一阶ODE路径。AudioCaps数据集上的实验表明,我们的模型,只有10个采样步骤,优于以前的模型,并减少推理到3个步骤相比,基于流匹配的加速模型。
摘要:Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by learning straight-line ordinary differential equation (ODE) paths. However, this approach requires training a flow-matching model from scratch and tends to perform suboptimally, or even poorly, at low step counts. To address the limitations of rectified flow while leveraging the advantages of advanced pre-trained diffusion models, this study integrates pre-trained models with the rectified diffusion method to improve the efficiency of text-to-audio (TTA) generation. Specifically, we propose AudioTurbo, which learns first-order ODE paths from deterministic noise sample pairs generated by a pre-trained TTA model. Experiments on the AudioCaps dataset demonstrate that our model, with only 10 sampling steps, outperforms prior models and reduces inference to 3 steps compared to a flow-matching-based acceleration model.


【15】 Visual Cues Support Robust Turn-taking Prediction in Noise

标题: 视觉线索支持噪音中稳健的轮流预测
链接:https://arxiv.org/abs/2505.22088
作者: Sam O'Connor Russell,  Naomi Harte 
备注:5 pages
摘要:准确的预测话轮转换模型(PTTMs)是自然的人机交互的基础。然而,人们对它们在噪音方面的表现知之甚少。因此,本研究探讨PTTM性能的类型的噪音可能会遇到一次部署。我们的分析表明PTTM对噪声高度敏感。保持/移位精度从清晰语音中的84%下降到10 dB音乐噪声中的52%。使用噪声数据进行训练可以实现多模态PTTM,其中包括视觉特征,以更好地利用视觉线索,在10 dB音乐噪声中具有72%的准确率。多模态PTTM在所有噪声类型和SNR上都优于仅音频PTTM,突出了其利用视觉线索的能力;然而,这并不总是适用于新类型的噪声。分析还表明,成功的训练依赖于准确的转录,限制了ASR衍生的translantine在清洁条件下的使用。我们将代码公开以供将来的研究。
摘要:Accurate predictive turn-taking models (PTTMs) are essential for naturalistic human-robot interaction. However, little is known about their performance in noise. This study therefore explores PTTM performance in types of noise likely to be encountered once deployed. Our analyses reveal PTTMs are highly sensitive to noise. Hold/shift accuracy drops from 84% in clean speech to just 52% in 10 dB music noise. Training with noisy data enables a multimodal PTTM, which includes visual features to better exploit visual cues, with 72% accuracy in 10 dB music noise. The multimodal PTTM outperforms the audio-only PTTM across all noise types and SNRs, highlighting its ability to exploit visual cues; however, this does not always generalise to new types of noise. Analysis also reveals that successful training relies on accurate transcription, limiting the use of ASR-derived transcriptions to clean conditions. We make code publicly available for future research.


【16】 On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech  Foundation Models for Dysarthric Speech Recognition

标题: Zero-ShotMoE说话人的实时路由用于合成障碍语音识别的语音基础模型的适应
链接:https://arxiv.org/abs/2505.22072
作者: Shujie HU,  Xurong Xie,  Mengzhe Geng,  Jiajun Deng,  Huimeng Wang,  Guinan Li,  Chengxi Deng,  Tianzi Wang,  Mingyu Cui,  Helen Meng,  Xunying Liu 
备注:Accepted by Interspeech 2025
摘要:针对构音障碍语音识别的基础模型,提出了一种基于MoE的说话人自适应框架。该方法在结合领域知识的同时实现了zero-shot自适应和实时处理。语音障碍的严重程度和性别条件的适配器专家动态结合使用的飞行预测扬声器相关的路由参数。KL-分歧用于进一步加强专家之间的多样性和他们的泛化看不见的扬声器。UASpeech语料库的实验结果表明,在飞行MoE为基础的适应产生统计上显着的WER减少高达1.34%的绝对值(6.36%相对)超过未适应的基线HuBERT/WavLM模型。通过批量模式自适应,在不同的扬声器级别数据量下,WER可持续降低高达2.55%(相对11.44%),RTF加速可达7倍。最低公布的WER为16.35%(46.77%的非常低的可懂度)。
摘要:This paper proposes a novel MoE-based speaker adaptation framework for foundation models based dysarthric speech recognition. This approach enables zero-shot adaptation and real-time processing while incorporating domain knowledge. Speech impairment severity and gender conditioned adapter experts are dynamically combined using on-the-fly predicted speaker-dependent routing parameters. KL-divergence is used to further enforce diversity among experts and their generalization to unseen speakers. Experimental results on the UASpeech corpus suggest that on-the-fly MoE-based adaptation produces statistically significant WER reductions of up to 1.34% absolute (6.36% relative) over the unadapted baseline HuBERT/WavLM models. Consistent WER reductions of up to 2.55% absolute (11.44% relative) and RTF speedups of up to 7 times are obtained over batch-mode adaptation across varying speaker-level data quantities. The lowest published WER of 16.35% (46.77% on very low intelligibility) is obtained.


【17】 Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency  Streaming ASR

标题: Delayed-KD:基于延迟知识蒸馏的CIC,用于低延迟流媒体ASB
链接:https://arxiv.org/abs/2505.22069
作者: Longhao Li,  Yangze Li,  Hongfei Xue,  Jie Liu,  Shuai Fang,  Kai Wang,  Lei Xie 
备注:Accepted by Interspeech2025
摘要:基于CTC的流式ASR在现实世界的应用中获得了极大的关注,但面临两个主要挑战:小块的准确性下降和令牌发射延迟。为了缓解这些挑战,我们提出了延迟KD,它将延迟知识蒸馏应用于从非流模型到流模型的CTC后验概率。具体来说,在块大小很小的情况下,我们引入了一个时间对齐缓冲区(TAB),该缓冲区定义了与非流教师模型相比的相对延迟范围,以对齐CTC输出并减轻非空白令牌不匹配。此外,TAB支持对令牌发送延迟进行细粒度控制。在178小时的AISHELL-1和10,000小时的WenetSpeech汉语数据集上的实验表明,Delayed-KD具有一致的优越性。令人印象深刻的是,延迟KD在40 ms延迟时在AISHELL-1上实现了5.42%的较低字符错误率(CER),与在320 ms延迟时运行的竞争对手U2++模型相当。
摘要:CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency. To mitigate these challenges, we propose Delayed-KD, which applies delayed knowledge distillation on CTC posterior probabilities from a non-streaming to a streaming model. Specifically, with a tiny chunk size, we introduce a Temporal Alignment Buffer (TAB) that defines a relative delay range compared to the non-streaming teacher model to align CTC outputs and mitigate non-blank token mismatches. Additionally, TAB enables fine-grained control over token emission delay. Experiments on 178-hour AISHELL-1 and 10,000-hour WenetSpeech Mandarin datasets show consistent superiority of Delayed-KD. Impressively, Delayed-KD at 40 ms latency achieves a lower character error rate (CER) of 5.42% on AISHELL-1, comparable to the competitive U2++ model running at 320 ms latency.


【18】 Weakly Supervised Data Refinement and Flexible Sequence Compression for  Efficient Thai LLM-based ASR

标题: 弱监督数据细化和灵活的序列压缩,以实现高效的泰国基于LLM的ASB
链接:https://arxiv.org/abs/2505.22063
作者: Mingchen Shao,  Xinfa Zhu,  Chengyou Wang,  Bingshen Mu,  Hai Li,  Ying Yan,  Junhui Liu,  Danming Xie,  Lei Xie 
备注:Accepted by INTERSPEECH 2025
摘要:尽管取得了显著的成就,但低资源场景下的自动语音识别(ASR)仍然面临两个挑战:高质量数据稀缺和高计算需求。本文提出了EThai-ASR,第一个将大型语言模型(LLM)应用于泰国ASR,并创建了一个高效的基于LLM的ASR系统。EThai-ASR包括语音编码器、连接模块和泰语LLM解码器。为了解决数据稀缺问题并获得强大的语音编码器,EThai-ASR引入了一种自进化的数据细化策略来细化弱标签,从而产生增强的语音编码器。此外,我们提出了一个可插拔的序列压缩模块中使用的连接模块,设计了三种模式,以减少序列长度,从而减少计算需求,同时保持体面的性能。大量的实验表明,EThai-ASR在多个数据集上达到了最先进的精度。我们发布我们的精炼文本记录,以促进进一步的研究。
摘要:Despite remarkable achievements, automatic speech recognition (ASR) in low-resource scenarios still faces two challenges: high-quality data scarcity and high computational demands. This paper proposes EThai-ASR, the first to apply large language models (LLMs) to Thai ASR and create an efficient LLM-based ASR system. EThai-ASR comprises a speech encoder, a connection module and a Thai LLM decoder. To address the data scarcity and obtain a powerful speech encoder, EThai-ASR introduces a self-evolving data refinement strategy to refine weak labels, yielding an enhanced speech encoder. Moreover, we propose a pluggable sequence compression module used in the connection module with three modes designed to reduce the sequence length, thus decreasing computational demands while maintaining decent performance. Extensive experiments demonstrate that EThai-ASR has achieved state-of-the-art accuracy in multiple datasets. We release our refined text transcripts to promote further research.


【19】 Voice Adaptation for Swiss German

标题: 瑞士德语语音改编
链接:https://arxiv.org/abs/2505.22054
作者: Samuel Stucki,  Jan Deriu,  Mark Cieliebak 
备注:Submitted to Interspeech
摘要:这项工作研究了瑞士德语方言的语音适应模型的性能,即,将标准德语文本翻译成瑞士德语方言语音。为此,我们预处理了一个大型的瑞士播客数据集,我们自动转录并用方言类进行注释,产生了大约5000小时的弱标记训练材料。我们在这个数据集上对XTTSv 2模型进行了微调,并表明它在人工和自动评估中取得了良好的成绩,并且可以正确地呈现所需的方言。我们的工作表明了一个步骤,使语音克隆技术适应代表性不足的语言。由此产生的模型实现了高达-0.28的CMOS分数和3.8的SMOS分数。
摘要:This work investigates the performance of Voice Adaptation models for Swiss German dialects, i.e., translating Standard German text to Swiss German dialect speech. For this, we preprocess a large dataset of Swiss podcasts, which we automatically transcribe and annotate with dialect classes, yielding approximately 5000 hours of weakly labeled training material. We fine-tune the XTTSv2 model on this dataset and show that it achieves good scores in human and automated evaluations and can correctly render the desired dialect. Our work shows a step towards adapting Voice Cloning technology to underrepresented languages. The resulting model achieves CMOS scores of up to -0.28 and SMOS scores of 3.8.


【20】 AudioGenie: A Training-Free Multi-Agent Framework for Diverse  Multimodality-to-Multiaudio Generation

标题: AudioGenie:一个免训练的多代理框架,用于多样化的多模式到多音频生成
链接:https://arxiv.org/abs/2505.22053
作者: Yan Rong,  Jinting Wang,  Shan Yang,  Guangzhi Lei,  Li Liu 
摘要:多模态到多音频(MM 2 MA)生成在合成不同的和上下文对齐的音频类型(例如,音效、语音、音乐和歌曲)从多模态输入(例如,视频,文本,图像),这是由于缺乏高质量的配对数据集和缺乏强大的多任务学习框架。近年来,多智能体系统在解决上述问题方面显示出巨大的潜力。然而,直接将其应用于MM 2 MA任务提出了三个关键挑战:(1)对多模态输入(特别是视频)的细粒度理解不足,(2)单个模型无法处理不同的音频事件,以及(3)缺乏可靠输出的自校正机制。为此,我们提出了AudioGenie,一种新的训练免费多代理系统具有双层架构的生成团队和监督团队。对于生成团队,设计了细粒度的任务分解和自适应的混合专家(MoE)协作实体进行动态模型选择,设计了试错迭代精化模块进行自校正.主管团队确保时空一致性,并通过反馈回路核实产出。此外,我们建立了MA-Bench,MM 2 MA任务的第一个基准测试,包括198个带注释的视频和多类型的音频。实验表明,我们的AudioGenie优于国家的最先进的(SOTA)方法在9个指标在8个任务。用户研究进一步验证了所提出的方法在质量,准确性,对齐和美观方面的有效性。可以在https://audiogenie.github.io/上找到带有示例的匿名项目网站。
摘要:Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images), owing to the scarcity of high-quality paired datasets and the lack of robust multi-task learning frameworks. Recently, multi-agent system shows great potential in tackling the above issues. However, directly applying it to MM2MA task presents three critical challenges: (1) inadequate fine-grained understanding of multimodal inputs (especially for video), (2) the inability of single models to handle diverse audio events, and (3) the absence of self-correction mechanisms for reliable outputs. To this end, we propose AudioGenie, a novel training-free multi-agent system featuring a dual-layer architecture with a generation team and a supervisor team. For the generation team, a fine-grained task decomposition and an adaptive Mixture-of-Experts (MoE) collaborative entity are designed for dynamic model selection, and a trial-and-error iterative refinement module is designed for self-correction. The supervisor team ensures temporal-spatial consistency and verifies outputs through feedback loops. Moreover, we build MA-Bench, the first benchmark for MM2MA tasks, comprising 198 annotated videos with multi-type audios. Experiments demonstrate that our AudioGenie outperforms state-of-the-art (SOTA) methods across 9 metrics in 8 tasks. User study further validate the effectiveness of the proposed method in terms of quality, accuracy, alignment, and aesthetic. The anonymous project website with samples can be found at https://audiogenie.github.io/.


【21】 Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

标题: 缓解视觉指南音频字幕中的视听不匹配
链接:https://arxiv.org/abs/2505.22045
作者: Le Xu,  Chenxing Li,  Yong Ren,  Yujie Chen,  Yu Gu,  Ruibo Fu,  Shan Yang,  Dong Yu 
备注:Accepted by INTERSPEECH 2025
摘要:当前的视觉引导音频字幕系统经常无法解决现实世界场景中的视听错位,例如配音内容或屏幕外声音。为了弥合这一关键差距,我们提出了一个熵感知的门控融合框架,通过跨模态不确定性量化动态调制视觉信息流。我们的新方法采用注意熵分析的交叉注意层自动识别和抑制误导性的视觉线索在模态融合。作为对这种架构的补充,我们开发了一种批量视听洗牌技术,该技术生成合成的不匹配训练对,大大增强了模型对对齐噪声的弹性。AudioCaps基准的评估表明,我们的系统的性能优于现有的基线,特别是在不匹配的模态场景。此外,与基线相比,我们的解决方案在推理速度上提高了约6倍。
摘要:Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge this critical gap, we present an entropy-aware gated fusion framework that dynamically modulates visual information flow through cross-modal uncertainty quantification. Our novel approach employs attention entropy analysis in cross-attention layers to automatically identify and suppress misleading visual cues during modal fusion. Complementing this architecture, we develop a batch-wise audiovisual shuffling technique that generates synthetic mismatched training pairs, greatly enhancing model resilience against alignment noise. Evaluations on the AudioCaps benchmark demonstrate our system's superior performance over existing baselines, especially in mismatched modality scenarios. Furthermore, our solution demonstrates an approximately 6x improvement in inference speed compared to the baseline.


【22】 Improving Respiratory Sound Classification with Architecture-Agnostic  Knowledge Distillation from Ensembles

标题: 通过从合奏中提取建筑不可知知识来改进呼吸声分类
链接:https://arxiv.org/abs/2505.22027
作者: Miika Toikkanen,  June-Woo Kim 
备注:Accepted to Interspeech 2025
摘要:呼吸声数据集在大小和质量上受到限制,使得难以实现高性能。包围模型有助于但不可避免地增加了推理时的计算成本。软标签培训有效地提取知识,仅在培训时增加额外成本。在这项研究中,我们探索呼吸声分类的软标签作为一个架构不可知的方法,蒸馏成一个学生模型的教师模型的合奏。我们研究了我们的方法的不同变化,发现即使是一个单一的老师,相同的学生,大大提高了性能超出了自己的能力,实现了最佳收益,只使用几个老师。我们在ICHBI上获得了64.39的最新得分,超过了之前的最好成绩0.85,并将各个架构的平均得分提高了1.16以上。我们的研究结果突出了知识蒸馏的有效性与呼吸声分类的软标签,无论大小或架构。
摘要:Respiratory sound datasets are limited in size and quality, making high performance difficult to achieve. Ensemble models help but inevitably increase compute cost at inference time. Soft label training distills knowledge efficiently with extra cost only at training. In this study, we explore soft labels for respiratory sound classification as an architecture-agnostic approach to distill an ensemble of teacher models into a student model. We examine different variations of our approach and find that even a single teacher, identical to the student, considerably improves performance beyond its own capability, with optimal gains achieved using only a few teachers. We achieve the new state-of-the-art Score of 64.39 on ICHBI, surpassing the previous best by 0.85 and improving average Scores across architectures by more than 1.16. Our results highlight the effectiveness of knowledge distillation with soft labels for respiratory sound classification, regardless of size or architecture.


【23】 RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic  Decomposed Modeling

标题: RESOUND:通过声学-语义分解建模从无声视频中重建语音
链接:https://arxiv.org/abs/2505.22024
作者: Long-Khanh Pham,  Thanh V. T. Tran,  Minh-Tan Pham,  Van Nguyen 
备注:accepted in Interspeech 2025
摘要:唇语合成(L2S),从视觉线索重建语音,面临着准确性和自然性的挑战,由于在捕捉语言内容,口音和韵律的监督有限。在本文中,我们提出了RESOUND,一种新的L2S系统,从无声的说话人脸视频生成可理解和表达的语音。利用源过滤器理论,我们的方法包括两个组成部分:声学路径预测韵律和语义路径提取语言特征。这种分离简化了学习,允许对每个表示进行独立优化。此外,我们通过将语音单元(一种经过验证的无监督语音表示技术)集成到梅尔频谱图旁边的波形生成中来提高性能。这允许RESOUND合成韵律语音,同时保留内容和说话者身份。两个标准的L2S基准进行的实验证实了所提出的方法在各种指标的有效性。
摘要:Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a novel L2S system that generates intelligible and expressive speech from silent talking face videos. Leveraging source-filter theory, our method involves two components: an acoustic path to predict prosody and a semantic path to extract linguistic features. This separation simplifies learning, allowing independent optimization of each representation. Additionally, we enhance performance by integrating speech units, a proven unsupervised speech representation technique, into waveform generation alongside mel-spectrograms. This allows RESOUND to synthesize prosodic speech while preserving content and speaker identity. Experiments conducted on two standard L2S benchmarks confirm the effectiveness of the proposed method across various metrics.


【24】 Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation  Addition for MISP 2025 Challenge

标题: MISP 2025挑战赛的重叠自适应混合扬声器拨号和ASB感知观察添加
链接:https://arxiv.org/abs/2505.22013
作者: Shangkun Huang,  Yuxuan Du,  Jingwen Yang,  Dejun Zhang,  Xupeng Jia,  Jing Deng,  Jintao Kang,  Rong Zheng 
备注:Accepted to Interspeech 2025
摘要:本文介绍了为应对MISP 2025挑战而开发的系统。对于日志化系统,我们提出了一种混合方法相结合的WavLM端到端的分割方法与传统的多模块聚类技术,以自适应地选择适当的模型来处理不同程度的重叠语音。针对自动语音识别(ASR)系统,提出了一种ASR感知的观测值添加方法,该方法可以弥补低信噪比条件下引导源分离(GSS)的性能限制。最后,我们在级联架构中集成了扬声器日志和ASR系统,以解决Track 3问题。我们的系统在轨道2上实现了9.48%的字符错误率(CER),在轨道3上实现了11.56%的级联最小排列字符错误率(cpCER),最终在两个轨道上都获得了第一名,从而证明了所提出的方法在现实世界的会议场景中的有效性。
摘要:This paper presents the system developed to address the MISP 2025 Challenge. For the diarization system, we proposed a hybrid approach combining a WavLM end-to-end segmentation method with a traditional multi-module clustering technique to adaptively select the appropriate model for handling varying degrees of overlapping speech. For the automatic speech recognition (ASR) system, we proposed an ASR-aware observation addition method that compensates for the performance limitations of Guided Source Separation (GSS) under low signal-to-noise ratio conditions. Finally, we integrated the speaker diarization and ASR systems in a cascaded architecture to address Track 3. Our system achieved character error rates (CER) of 9.48% on Track 2 and concatenated minimum permutation character error rate (cpCER) of 11.56% on Track 3, ultimately securing first place in both tracks and thereby demonstrating the effectiveness of the proposed methods in real-world meeting scenarios.


【25】 Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging  Recognition and Event Detection

标题: 利用LLM进行口吃语音:一种桥接识别和事件检测的统一架构
链接:https://arxiv.org/abs/2505.22005
作者: Shangkun Huang,  Jing Deng,  Jintao Kang,  Rong Zheng 
备注:Accepted to Interspeech 2025
摘要:自动语音识别(ASR)在口吃语音场景下的性能瓶颈限制了其在语音康复等领域的应用。提出了一种LLM驱动的ASR-SED多任务学习框架,该框架联合优化了ASR和口吃事件检测(SED)任务。我们提出了一种动态交互机制,其中ASR分支利用CTC生成的软提示来辅助LLM上下文建模,而SED分支输出口吃嵌入来增强LLM对口吃语音的理解。我们采用对比学习来加强口吃声学特征的辨别能力,并应用焦点损失来减轻口吃事件类别中的长尾分布。在AS-70汉语口吃数据集上的评估表明,我们的框架将ASR字符错误率(CER)降低到5.45%(相对减少-37.71%),并实现了平均73.63%的SED F1分数(相对改善+46.58%)。
摘要:The performance bottleneck of Automatic Speech Recognition (ASR) in stuttering speech scenarios has limited its applicability in domains such as speech rehabilitation. This paper proposed an LLM-driven ASR-SED multi-task learning framework that jointly optimized the ASR and Stuttering Event Detection (SED) tasks. We proposed a dynamic interaction mechanism where the ASR branch leveraged CTC-generated soft prompts to assist LLM context modeling, while the SED branch output stutter embeddings to enhance LLM comprehension of stuttered speech. We incorporated contrastive learning to strengthen the discriminative power of stuttering acoustic features and applied Focal Loss to mitigate the long-tailed distribution in stuttering event categories. Evaluations on the AS-70 Mandarin stuttering dataset demonstrated that our framework reduced the ASR character error rate (CER) to 5.45% (-37.71% relative reduction) and achieved an average SED F1-score of 73.63% (+46.58% relative improvement).


【26】 Music Source Restoration

标题: 音乐资源恢复
链接:https://arxiv.org/abs/2505.21827
作者: Yongyi Zang,  Zheqi Dai,  Mark D. Plumbley,  Qiuqiang Kong 
备注:A modified version of this paper is in review
摘要:我们介绍了音乐源恢复(MSR),一个新的任务,解决理想化的源分离和现实世界的音乐制作之间的差距。当前的音乐源分离(MSS)方法假设混合是源的简单和,忽略了在音乐制作期间使用的信号降级,如均衡、压缩和混响。MSR将混合物建模为单独退化源的退化总和,目标是恢复原始的、未退化的信号。由于缺乏MSR的数据,我们提出了RawStems,一个数据集注释的578首歌曲与未经处理的源信号组织成8个主要和17个次要乐器组,共计354.13小时。据我们所知,RawStems是第一个包含未处理的音乐词干的数据集。我们考虑频谱滤波,动态范围压缩,谐波失真,混响和有损编解码器作为可能的退化,并建立U形作为基线方法,证明了我们的数据集上的MSR的可行性。我们发布了RawStems数据集注释、退化模拟管道、训练代码和预训练模型,供公众使用。
摘要:We introduce Music Source Restoration (MSR), a novel task addressing the gap between idealized source separation and real-world music production. Current Music Source Separation (MSS) approaches assume mixtures are simple sums of sources, ignoring signal degradations employed during music production like equalization, compression, and reverb. MSR models mixtures as degraded sums of individually degraded sources, with the goal of recovering original, undegraded signals. Due to the lack of data for MSR, we present RawStems, a dataset annotation of 578 songs with unprocessed source signals organized into 8 primary and 17 secondary instrument groups, totaling 354.13 hours. To the best of our knowledge, RawStems is the first dataset that contains unprocessed music stems with hierarchical categories. We consider spectral filtering, dynamic range compression, harmonic distortion, reverb and lossy codec as possible degradations, and establish U-Former as a baseline method, demonstrating the feasibility of MSR on our dataset. We release the RawStems dataset annotations, degradation simulation pipeline, training code and pre-trained models to be publicly available.


【27】 Voice Quality Dimensions as Interpretable Primitives for Speaking Style  for Atypical Speech and Affect

标题: 语音质量维度作为非典型言语和情感说话风格的可解释基本要素
链接:https://arxiv.org/abs/2505.21809
作者: Jaya Narain,  Vasudha Kowtha,  Colin Lea,  Lauren Tooley,  Dianna Yee,  Vikramjit Mitra,  Zifang Huang,  Miquel Espi Marques,  Jon Huang,  Carlos Avendano,  Shirley Ren 
备注:accepted for Interspeech 2025
摘要:感知语音质量维度描述了非典型语音和其他语音调制的关键特征。在这里,我们开发和评估七个语音和语音维度(可懂度,不精确的辅音,刺耳的声音,自然度,单响度,呼吸困难和呼吸)的语音质量模型。探针在公共语音可访问性(SAP)项目数据集上进行了训练,该数据集包含来自434名扬声器的11,184个样本,使用来自冻结预训练模型的嵌入作为特征。我们发现,我们的探测器在SAP数据集中的语音启发类别中具有很强的性能和很强的泛化能力。我们进一步验证了其他数据集上的zero-shot性能,包括看不见的语言和任务:意大利语非典型语音,英语非典型语音和情感语音。强大的zero-shot性能和跨一系列评估结果的可解释性表明在说话风格相关任务中使用语音质量维度的实用性。
摘要:Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, imprecise consonants, harsh voice, naturalness, monoloudness, monopitch, and breathiness). Probes were trained on the public Speech Accessibility (SAP) project dataset with 11,184 samples from 434 speakers, using embeddings from frozen pre-trained models as features. We found that our probes had both strong performance and strong generalization across speech elicitation categories in the SAP dataset. We further validated zero-shot performance on additional datasets, encompassing unseen languages and tasks: Italian atypical speech, English atypical speech, and affective speech. The strong zero-shot performance and the interpretability of results across an array of evaluations suggests the utility of using voice quality dimensions in speaking style-related tasks.


【28】 An Investigation on Speaker Augmentation for End-to-End Speaker  Extraction

标题: 端到端说话人提取的说话人增强研究
链接:https://arxiv.org/abs/2505.21805
作者: Zhenghai You,  Zhenyu Zhou,  Lantian Li,  Dong Wang 
摘要:目标混淆是指偶尔切换到非目标说话人,是端到端说话人提取(E2 E-SE)系统面临的一个关键挑战。我们认为,这个问题在很大程度上是由于缺乏概括性和歧视性的说话人嵌入,并介绍了一个简单而有效的说话人增强策略来解决这个问题。具体来说,我们提出了一个时域的rescaling和重新缩放管道,改变扬声器的特点,同时保留其他语音属性。这会生成各种各样的伪说话者,以帮助建立一个可推广的说话者嵌入空间,而特定于说话者特征的增强会创建硬样本,迫使模型专注于真正的说话者特征。在WSJ 0 - 2 Mix和LibriMix上的实验表明,该方法有效地缓解了目标混淆,提高了提取性能。此外,它可以与指标学习相结合,这是解决目标混乱的另一种有效方法,从而带来进一步的收益。
摘要:Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and discrimination of the speaker embeddings, and introduce a simple yet effective speaker augmentation strategy to tackle the problem. Specifically, we propose a time-domain resampling and rescaling pipeline that alters speaker traits while preserving other speech properties. This generates a variety of pseudo-speakers to help establish a generalizable speaker embedding space, while the speaker-trait-specific augmentation creates hard samples that force the model to focus on genuine speaker characteristics. Experiments on WSJ0-2Mix and LibriMix show that our method mitigates the target confusion and improves extraction performance. Moreover, it can be combined with metric learning, another effective approach to address target confusion, leading to further gains.


【29】 Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech  Recognition Data for Research and Commercial Use

标题: Loquacious Set:25,000小时的转录和多样化的英语语音识别数据,用于研究和商业用途
链接:https://arxiv.org/abs/2505.21578
作者: Titouan Parcollet,  Yuan Tseng,  Shucong Zhang,  Rogier van Dalen 
备注:Accepted at Interspeech 2025
摘要:自动语音识别(ASR)研究是由工业研究人员和学术界之间的通用数据集驱动的,鼓励进行比较和评估。尽管LibriSpeech作为ASR基准取得了长期的成功,但现在受到其规模的限制,并专注于干净,阅读语音,导致单词错误率接近零。最近的数据集,包括MOSEL,YODAS,Gigaspeech,OWSM,Libriheavy或People's Speech,都受到重大限制,包括行业研究人员无法使用的许可证,不可靠的传输,不正确的音频数据或缺乏评估集。这部作品展示了一个长达25,000小时的商业英语演讲集。Loquacious Set拥有数十万具有不同口音和各种语音类型(阅读,自发,谈话,干净,嘈杂)的扬声器,旨在为业内学者和研究人员在现实世界中构建ASR系统。
摘要:Automatic speech recognition (ASR) research is driven by the availability of common datasets between industrial researchers and academics, encouraging comparisons and evaluations. LibriSpeech, despite its long success as an ASR benchmark, is now limited by its size and focus on clean, read speech, leading to near-zero word error rates. More recent datasets, including MOSEL, YODAS, Gigaspeech, OWSM, Libriheavy or People's Speech suffer from major limitations including licenses that researchers in the industry cannot use, unreliable transcriptions, incorrect audio data, or the lack of evaluation sets. This work presents the Loquacious Set, a 25,000-hour curated collection of commercially usable English speech. Featuring hundreds of thousands of speakers with diverse accents and a wide range of speech types (read, spontaneous, talks, clean, noisy), the Loquacious Set is designed to work for academics and researchers in the industry to build ASR systems in real-world scenarios.


【30】 VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach  Leveraging Speaker-Specific Latents

标题: VoiceMark:一种利用特定说话人特征的抗零采样语音克隆水印方法
链接:https://arxiv.org/abs/2505.21568
作者: Haiyun Li,  Zhiyong Wu,  Xiaofeng Xie,  Jingran Xie,  Yaoxun Xu,  Hanyang Peng 
备注:Accepted by Interspeech 2025
摘要:抗语音克隆水印是一种新兴的跟踪和防止未授权克隆的技术。现有的方法通过在带水印的音频上训练传统的VC模型来有效地跟踪传统的VC模型,但是在zero-shot VC场景中失败,其中模型从音频提示合成音频而不进行训练。为了解决这个问题,我们提出了VoiceMark,这是第一个抗零次VC(zero-shot VC)水印方法,它利用特定于说话者的潜伏期作为水印载体,允许水印通过zero-shot VC过程传输到合成音频中。此外,我们引入VC模拟的增强和基于VAD的损失,以提高对失真的鲁棒性。在多个zero-shot VC模型上的实验表明,经过zero-shot VC合成后,VoiceMark的水印检测准确率达到了95%以上,明显优于现有的只能达到50%左右的方法。查看我们的代码和演示:https://huggingface.co/spaces/haiyunli/VoiceMark
摘要:Voice cloning (VC)-resistant watermarking is an emerging technique for tracing and preventing unauthorized cloning. Existing methods effectively trace traditional VC models by training them on watermarked audio but fail in zero-shot VC scenarios, where models synthesize audio from an audio prompt without training. To address this, we propose VoiceMark, the first zero-shot VC-resistant watermarking method that leverages speaker-specific latents as the watermark carrier, allowing the watermark to transfer through the zero-shot VC process into the synthesized audio. Additionally, we introduce VC-simulated augmentations and VAD-based loss to enhance robustness against distortions. Experiments on multiple zero-shot VC models demonstrate that VoiceMark achieves over 95% accuracy in watermark detection after zero-shot VC synthesis, significantly outperforming existing methods, which only reach around 50%. See our code and demos at: https://huggingface.co/spaces/haiyunli/VoiceMark


机器翻译由腾讯交互翻译提供,仅供参考