今日论文合集:cs.SD语音21篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 ZeroSep: Separate Anything in Audio with Zero Training
标题: ZeroSep:通过零训练分离音频中的任何内容
链接:https://arxiv.org/abs/2505.23625
作者: Chao Huang,  Yuesheng Ma,  Junxuan Huang,  Susan Liang,  Yunlong Tang,  Jing Bi,  Wenqiang Liu,  Nima Mesgarani,  Chenliang Xu 
备注:Project page: this https URL
摘要:音频源分离是机器理解复杂声学环境的基础,也是众多音频应用的基础。目前的监督式深度学习方法虽然功能强大,但受限于对大量特定于任务的标记数据的需求,并且难以推广到现实世界声学场景的巨大可变性和开放集性质。受生成基础模型成功的启发,我们研究了预训练的文本引导音频扩散模型是否可以克服这些限制。我们有一个惊人的发现:zero-shot源分离可以完全通过正确配置下的预训练文本引导音频扩散模型来实现。我们的方法名为ZeroSep,其工作原理是将混合音频反转到扩散模型的潜在空间中,然后使用文本条件来指导去噪过程以恢复单个源。在没有任何特定于任务的训练或微调的情况下,ZeroSep将生成扩散模型重新用于区分分离任务,并通过其丰富的文本先验知识本质上支持开放集场景。ZeroSep与各种预训练的文本引导音频扩散骨干兼容,并在多个分离基准上提供强大的分离性能,甚至超过了监督方法。
摘要:Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the immense variability and open-set nature of real-world acoustic scenes. Inspired by the success of generative foundation models, we investigate whether pre-trained text-guided audio diffusion models can overcome these limitations. We make a surprising discovery: zero-shot source separation can be achieved purely through a pre-trained text-guided audio diffusion model under the right configuration. Our method, named ZeroSep, works by inverting the mixed audio into the diffusion model's latent space and then using text conditioning to guide the denoising process to recover individual sources. Without any task-specific training or fine-tuning, ZeroSep repurposes the generative diffusion model for a discriminative separation task and inherently supports open-set scenarios through its rich textual priors. ZeroSep is compatible with a variety of pre-trained text-guided audio diffusion backbones and delivers strong separation performance on multiple separation benchmarks, surpassing even supervised methods.


【2】 Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

标题: 采用高斯过程的Few-Shot语音深度伪造检测自适应
链接:https://arxiv.org/abs/2505.23619
作者: Neta Glazer,  David Chernin,  Idan Achituve,  Sharon Gannot,  Ethan Fetaya 
摘要:文本到语音(TTS)模型的最新进展,特别是在语音克隆方面,加剧了对适应性强且有效的深度伪造检测方法的需求。随着TTS系统的不断发展,检测模型必须能够以最少的数据有效地适应以前看不见的生成模型。本文介绍了ADD-GP,一种基于高斯过程(GP)分类器的用于音频深度伪造检测(ADD)的Few-Shot自适应框架。我们展示了强大的深度嵌入模型与高斯过程灵活性的结合如何实现强大的性能和适应性。此外,我们还证明了这种方法也可以用于个性化检测,对新的TTS模型和一次性适应性具有更强的鲁棒性。为了支持我们的评估,使用新的最先进的语音克隆模型为这项任务构建了一个基准数据集。
摘要:Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS systems continue to evolve, detection models must be able to efficiently adapt to previously unseen generation models with minimal data. This paper introduces ADD-GP, a few-shot adaptive framework based on a Gaussian Process (GP) classifier for Audio Deepfake Detection (ADD). We show how the combination of a powerful deep embedding model with the Gaussian processes flexibility can achieve strong performance and adaptability. Additionally, we show this approach can also be used for personalized detection, with greater robustness to new TTS models and one-shot adaptability. To support our evaluation, a benchmark dataset is constructed for this task using new state-of-the-art voice cloning models.


【3】 Spectrotemporal Modulation: Efficient and Interpretable Feature  Representation for Classifying Speech, Music, and Environmental Sounds

标题: 光谱时间调制:用于分类语音、音乐和环境声音的高效且可解释的特征表示
链接:https://arxiv.org/abs/2505.23509
作者: Andrew Chang,  Yike Li,  Iran R. Roman,  David Poeppel 
备注:Interspeech 2025
摘要:音频DNN在各种机器监听任务中表现出令人印象深刻的性能;然而,它们的大多数表示都是计算成本高且不可解释的,这为优化留下了空间。在这里,我们提出了一种新的方法集中在频谱时间调制(STM)功能,信号处理方法,模仿人类听觉皮层的神经生理表示。我们基于STM的模型的分类性能,在没有任何预训练的情况下,与预训练的音频DNN在不同的自然语音,音乐和环境声音中的分类性能相当,这些都是人类认知和机器感知的基本类别。这些结果表明,STM是音频分类的有效和可解释的特征表示,推进了机器听力的发展,为语音和听觉科学的基本理解解锁了令人兴奋的新可能性,以及开发音频BCI和认知计算。
摘要:Audio DNNs have demonstrated impressive performance on various machine listening tasks; however, most of their representations are computationally costly and uninterpretable, leaving room for optimization. Here, we propose a novel approach centered on spectrotemporal modulation (STM) features, a signal processing method that mimics the neurophysiological representation in the human auditory cortex. The classification performance of our STM-based model, without any pretraining, is comparable to that of pretrained audio DNNs across diverse naturalistic speech, music, and environmental sounds, which are essential categories for both human cognition and machine perception. These results show that STM is an efficient and interpretable feature representation for audio classification, advancing the development of machine listening and unlocking exciting new possibilities for basic understanding of speech and auditory sciences, as well as developing audio BCI and cognitive computing.


【4】 Semantics-Aware Human Motion Generation from Audio Instructions

标题: 从音频指令生成语义感知的人体运动
链接:https://arxiv.org/abs/2505.23465
作者: Zi-An Wang,  Shihao Zou,  Shiyao Yu,  Mingyuan Zhang,  Chao Dong 
备注:None
摘要:交互技术的最新进展已经突出了用于语义编码的音频信号的重要性。本文探讨了一个新的任务,其中音频信号被用作条件输入,以生成符合音频语义的运动。与基于文本的交互不同,音频提供了一种更自然、更直观的交流方式。然而,现有的方法通常专注于将运动与音乐或语音节奏相匹配,这通常导致音频的语义与所生成的运动之间的弱连接。我们提出了一个端到端的框架,使用掩蔽生成Transformer,增强了记忆检索注意力模块来处理稀疏和冗长的音频输入。此外,我们丰富了现有的数据集,将描述转换为会话风格,并生成相应的音频与不同的扬声器身份。实验证明了该框架的有效性和效率,表明音频指令可以传达类似于文本的语义,同时提供更实用和用户友好的交互。
摘要:Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions that align with the semantics of the audio. Unlike text-based interactions, audio provides a more natural and intuitive communication method. However, existing methods typically focus on matching motions with music or speech rhythms, which often results in a weak connection between the semantics of the audio and generated motions. We propose an end-to-end framework using a masked generative transformer, enhanced by a memory-retrieval attention module to handle sparse and lengthy audio inputs. Additionally, we enrich existing datasets by converting descriptions into conversational style and generating corresponding audio with varied speaker identities. Experiments demonstrate the effectiveness and efficiency of the proposed framework, demonstrating that audio instructions can convey semantics similar to text while providing more practical and user-friendly interactions.


【5】 Nosey: Open-source hardware for acoustic nasalance

标题: Nosey:用于声学鼻喉科的开源硬件
链接:https://arxiv.org/abs/2505.23339
作者: Maya Dewhurst,  Jack Collins,  Justin J. H. Lo,  Roy Alderton,  Sam Kirkham 
备注:Accepted to Interspeech 2025
摘要:我们介绍了Nosey(Nasalance开源估计系统),这是一个低成本,可定制的3D打印系统,用于记录我们作为开源硬件提供的声学鼻测数据(http://github.com/phoneticslab/nosey)。我们首先概述了我们的硬件nasalance系统背后的动机和设计原则,然后提出了一个比较Nosey和商业nasalance设备。爱管闲事的表现一贯较高的nasalance分数比商业设备,但语音环境之间的对比度的大小是系统之间的可比性。我们还审查了定制硬件以方便测试的方法,例如麦克风和不同建筑材料的比较。我们的结论是,Nosey是一个灵活的和具有成本效益的替代商业nasometry设备,并提出了一些方法上的考虑,其用于数据收集。
摘要:We introduce Nosey (Nasalance Open Source Estimation sYstem), a low-cost, customizable, 3D-printed system for recording acoustic nasalance data that we have made available as open-source hardware (http://github.com/phoneticslab/nosey). We first outline the motivations and design principles behind our hardware nasalance system, and then present a comparison between Nosey and a commercial nasalance device. Nosey shows consistently higher nasalance scores than the commercial device, but the magnitude of contrast between phonological environments is comparable between systems. We also review ways of customizing the hardware to facilitate testing, such as comparison of microphones and different construction materials. We conclude that Nosey is a flexible and cost-effective alternative to commercial nasometry devices and propose some methodological considerations for its use in data collection.


【6】 MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and  Source Extraction

标题: MGE-LDM:同时音乐生成和源提取的联合潜在扩散
链接:https://arxiv.org/abs/2505.23305
作者: Yunkee Chae,  Kyogu Lee 
备注:27 pages, 4 figures
摘要:我们提出了MGE-LDM,一个统一的潜在的扩散框架,同时音乐生成,源插补和查询驱动的源分离。与以前的方法局限于固定的仪器类,MGE-LDM学习一个联合分布在一个单一的紧凑的潜在扩散模型的完整的混合物,子混合物,和个别的茎。在推断时,MGE-LDM使得能够(1)完全混合生成,(2)部分生成(即,源插补),以及(3)任意源的文本条件提取。通过将分离和归责制定为潜在空间中的条件修复任务,我们的方法支持对任意仪器源的灵活的、类不可知的操作。值得注意的是,MGE-LDM可以跨异构多轨道数据集(例如,Slakh 2100,MUSDB 18,MoisesDB),而不依赖于预定义的仪器类别。音频样本可以在我们的项目页面上找到:https://yoongi43.github.io/MGELDM_Samples/。
摘要:We present MGE-LDM, a unified latent diffusion framework for simultaneous music generation, source imputation, and query-driven source separation. Unlike prior approaches constrained to fixed instrument classes, MGE-LDM learns a joint distribution over full mixtures, submixtures, and individual stems within a single compact latent diffusion model. At inference, MGE-LDM enables (1) complete mixture generation, (2) partial generation (i.e., source imputation), and (3) text-conditioned extraction of arbitrary sources. By formulating both separation and imputation as conditional inpainting tasks in the latent space, our approach supports flexible, class-agnostic manipulation of arbitrary instrument sources. Notably, MGE-LDM can be trained jointly across heterogeneous multi-track datasets (e.g., Slakh2100, MUSDB18, MoisesDB) without relying on predefined instrument categories. Audio samples are available at our project page: https://yoongi43.github.io/MGELDM_Samples/.


【7】 Bridging the Gap Between Semantic and User Preference Spaces for  Multi-modal Music Representation Learning

标题: 多模态音乐表征学习中语义空间与用户偏好空间之间的桥梁
链接:https://arxiv.org/abs/2505.23298
作者: Xiaofeng Pan,  Jing Chen,  Haitong Zhang,  Menglin Xing,  Jiayi Wei,  Xuefeng Mu,  Zhongqian Xie 
备注:ICMR 2025
摘要:近年来的音乐表征学习研究主要集中在未标记音频的声学音乐表征学习,或者进一步尝试在缺乏注释的音频-文本对的情况下获得多模态音乐表征。它们要么忽略语言语义,要么依赖于难以创建且昂贵的标记音频数据集。此外,由于忽略了用户偏好空间,仅仅对语义空间进行建模通常无法在音乐推荐任务上取得令人满意的性能。在本文中,我们提出了一种新的分层两阶段对比学习(HTCL)方法,从语义的角度到用户的角度分层模型的相似性,学习一个全面的音乐表示弥合语义和用户偏好空间之间的差距。我们设计了一个可扩展的音频编码器,并利用一个预先训练好的BERT模型作为文本编码器,通过大规模的对比预训练来学习音频-文本语义。此外,我们探索了一种简单而有效的方法来利用我们的在线音乐平台的交互数据,通过对比微调,这不同于以前的作品,遵循协同过滤的想法,以适应语义空间的用户偏好空间。因此,我们得到一个功能强大的音频编码器,不仅从文本编码器提取语言语义,而且在用户偏好空间中建模相似性,保留语义空间的完整性。音乐语义和推荐任务的实验结果证实了我们的方法的有效性。
摘要:Recent works of music representation learning mainly focus on learning acoustic music representations with unlabeled audios or further attempt to acquire multi-modal music representations with scarce annotated audio-text pairs. They either ignore the language semantics or rely on labeled audio datasets that are difficult and expensive to create. Moreover, merely modeling semantic space usually fails to achieve satisfactory performance on music recommendation tasks since the user preference space is ignored. In this paper, we propose a novel Hierarchical Two-stage Contrastive Learning (HTCL) method that models similarity from the semantic perspective to the user perspective hierarchically to learn a comprehensive music representation bridging the gap between semantic and user preference spaces. We devise a scalable audio encoder and leverage a pre-trained BERT model as the text encoder to learn audio-text semantics via large-scale contrastive pre-training. Further, we explore a simple yet effective way to exploit interaction data from our online music platform to adapt the semantic space to user preference space via contrastive fine-tuning, which differs from previous works that follow the idea of collaborative filtering. As a result, we obtain a powerful audio encoder that not only distills language semantics from the text encoder but also models similarity in user preference space with the integrity of semantic space preserved. Experimental results on both music semantic and recommendation tasks confirm the effectiveness of our method.


【8】 Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven  Facial Animation

标题: Wav2Sem:即插即用音频语义去耦合3D语音驱动面部动画
链接:https://arxiv.org/abs/2505.23290
作者: Hao Li,  Ju Dai,  Xin Zhao,  Feng Zhou,  Junjun Pan,  Lei Li 
备注:Accepted to CVPR 2025
摘要:在3D语音驱动的面部动画生成中,现有方法通常采用预训练的自监督音频模型作为编码器。然而,由于语音相似的音节与不同的唇形在语言中的流行,这些近同音字音节往往表现出显着的耦合在自我监督的音频特征空间,导致在随后的嘴唇运动生成的平均效应。针对这一问题,提出了一种即插即用的语义去相关模块Wav2Sem。该模块提取对应于整个音频序列的语义特征,利用添加的语义信息在特征空间内对音频编码进行去相关,从而实现更具表现力的音频特征。在多个语音驱动模型上进行的大量实验表明,Wav2Sem模块有效地融合了音频特征,显著减轻了唇形生成中语音相似音节的平均效应,从而提高了面部动画的精度和自然度。我们的源代码可在https://github.com/wslh852/Wav2Sem.git上获得。
摘要:In 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant coupling in self-supervised audio feature spaces, leading to the averaging effect in subsequent lip motion generation. To address this issue, this paper proposes a plug-and-play semantic decorrelation module-Wav2Sem. This module extracts semantic features corresponding to the entire audio sequence, leveraging the added semantic information to decorrelate audio encodings within the feature space, thereby achieving more expressive audio features. Extensive experiments across multiple Speech-driven models indicate that the Wav2Sem module effectively decouples audio features, significantly alleviating the averaging effect of phonetically similar syllables in lip shape generation, thereby enhancing the precision and naturalness of facial animations. Our source code is available at https://github.com/wslh852/Wav2Sem.git.


【9】 Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable  Emotion Recognition

标题: 面向可解释情感识别的LLM授权细粒度语音描述符
链接:https://arxiv.org/abs/2505.23236
作者: Youjun Chen,  Xurong Xie,  Haoning Xu,  Mengzhe Geng,  Guinan Li,  Chengxi Deng,  Huimeng Wang,  Shujie Hu,  Xunying Liu 
备注:Accepted by INTERSPEECH2025
摘要:本文提出了一种新的端到端LLM授权的可解释语音情感识别(SER)方法。细粒度语音情感描述符(SED)特征,例如,音高、音调和强调,通过交替的LLM微调从HuBERT SSL表示中解脱出来,以联合SER-SED预测和ASR任务。通过信息瓶颈(IB)得到的VAE压缩HuBERT特征用于调整特征粒度。IEMOCAP和MELD基准测试的实验表明,我们的方法始终优于可比的基于LLaMA的SER基线,包括那些使用(a)单独交替多任务微调或(b)功能解纠缠。统计学显着增加SER未加权的准确性高达4.0%和3.7%的绝对(5.4%和6.6%的相对)。更重要的是,情感描述符提供了进一步的解释SER。
摘要:This paper presents a novel end-to-end LLM-empowered explainable speech emotion recognition (SER) approach. Fine-grained speech emotion descriptor (SED) features, e.g., pitch, tone and emphasis, are disentangled from HuBERT SSL representations via alternating LLM fine-tuning to joint SER-SED prediction and ASR tasks. VAE compressed HuBERT features derived via Information Bottleneck (IB) are used to adjust feature granularity. Experiments on the IEMOCAP and MELD benchmarks demonstrate that our approach consistently outperforms comparable LLaMA-based SER baselines, including those using either (a) alternating multi-task fine-tuning alone or (b) feature disentanglement only. Statistically significant increase of SER unweighted accuracy by up to 4.0% and 3.7% absolute (5.4% and 6.6% relative) are obtained. More importantly, emotion descriptors offer further explainability for SER.


【10】 Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive  Approach Using WavLM

标题: 迈向稳健的重叠语音检测:使用WavLM的扬声器感知渐进方法
链接:https://arxiv.org/abs/2505.23207
作者: Zhaokai Sun,  Li Zhang,  Qing Wang,  Pan Zhou,  Lei Xie 
摘要:重叠语音检测(OSD)的目的是识别多个说话人在对话中重叠的区域,这是多方语音处理中的一个关键挑战。这项工作提出了一个扬声器感知渐进OSD模型,利用渐进的训练策略,以提高子任务,如语音活动检测(VAD)和重叠检测之间的相关性。为了改善声学表示,我们探索了最先进的自监督学习(SSL)模型的有效性,包括WavLM和wav 2 vec 2.0,同时结合了扬声器注意力模块,以丰富帧级扬声器信息的功能。实验结果表明,该方法在AMI测试集上的F1值为82.76%,具有较好的性能,证明了该方法在OSD中的鲁棒性和有效性.
摘要:Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a progressive training strategy to enhance the correlation between subtasks such as voice activity detection (VAD) and overlap detection. To improve acoustic representation, we explore the effectiveness of state-of-the-art self-supervised learning (SSL) models, including WavLM and wav2vec 2.0, while incorporating a speaker attention module to enrich features with frame-level speaker information. Experimental results show that the proposed method achieves state-of-the-art performance, with an F1 score of 82.76\% on the AMI test set, demonstrating its robustness and effectiveness in OSD.


【11】 ZIPA: A family of efficient models for multilingual phone recognition

标题: ZIPA:一系列高效的多语言电话识别模型
链接:https://arxiv.org/abs/2505.23170
作者: Jian Zhu,  Farhan Samir,  Eleanor Chodroff,  David R. Mortensen 
备注:ACL 2025 Main
摘要:我们提出了ZIPA,一个高效的语音模型,提高了跨语言电话识别的最先进的性能。我们首先策划了IPAPack++,这是一个大规模的多语言语音语料库,拥有17,132小时的标准化电话传输和一个新颖的评估集,捕捉了看不见的语言和社会语音变化。通过大规模的训练数据,ZIPA,包括换能器(ZIPA-T)和基于CTC的变体(ZIPA-CR),利用高效的Zipformer主干,并以更少的参数胜过现有的电话识别系统。通过对11,000小时的伪标记多语言数据进行噪声学生训练,进一步扩展会产生进一步的改进。虽然ZIPA在基准测试中表现出色,但错误分析揭示了在建模社会语音多样性方面的持续局限性,强调了未来研究的挑战。
摘要:We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPAPack++, a large-scale multilingual speech corpus with 17,132 hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation. With the large-scale training data, ZIPA, including transducer (ZIPA-T) and CTC-based (ZIPA-CR) variants, leverage the efficient Zipformer backbones and outperform existing phone recognition systems with much fewer parameters. Further scaling via noisy student training on 11,000 hours of pseudo-labeled multilingual data yields further improvement. While ZIPA achieves strong performance on benchmarks, error analysis reveals persistent limitations in modeling sociophonetic diversity, underscoring challenges for future research.


【12】 Patient Domain Supervised Contrastive Learning for Lung Sound  Classification Using Mobile Phone

标题: 使用手机进行肺音分类的患者领域监督对比学习
链接:https://arxiv.org/abs/2505.23132
作者: Seung Gyu Jeong,  Seong Eun Kim 
备注:ITS-CSCC 2024
摘要:听诊对于诊断肺部疾病至关重要。COVID-19大流行揭示了传统的面对面肺音评估的局限性。为了克服这些问题,数字听诊器和人工智能(AI)的进步导致了新诊断方法的发展。在这种情况下,我们的研究旨在使用智能手机麦克风来记录和分析肺音。我们面临两大挑战:电子听诊器和智能手机麦克风之间的音频风格差异,以及患者之间的差异。为了应对这些挑战,我们开发了一种称为患者域监督对比学习(PD-SCL)的方法。通过将该方法与音频频谱图Transformer(AST)模型相结合,与原始AST模型相比,其性能显著提高了2.4%.这一进展表明,智能手机可以有效地诊断肺音,解决患者数据的不一致性,并显示出在传统临床环境之外广泛使用的潜力。我们的研究有助于在COVID-19后的世界中更容易检测肺部疾病。
摘要:Auscultation is crucial for diagnosing lung diseases. The COVID-19 pandemic has revealed the limitations of traditional, in-person lung sound assessments. To overcome these issues, advancements in digital stethoscopes and artificial intelligence (AI) have led to the development of new diagnostic methods. In this context, our study aims to use smartphone microphones to record and analyze lung sounds. We faced two major challenges: the difference in audio style between electronic stethoscopes and smartphone microphones, and the variability among patients. To address these challenges, we developed a method called Patient Domain Supervised Contrastive Learning (PD-SCL). By integrating this method with the Audio Spectrogram Transformer (AST) model, we significantly improved its performance by 2.4\% compared to the original AST model. This progress demonstrates that smartphones can effectively diagnose lung sounds, addressing inconsistencies in patient data and showing potential for broad use beyond traditional clinical settings. Our research contributes to making lung disease detection more accessible in the post-COVID-19 world.


【13】 Contextualized Automatic Speech Recognition with Dynamic Vocabulary  Prediction and Activation

标题: 具有动态词汇预测和激活的上下文自动语音识别
链接:https://arxiv.org/abs/2505.23077
作者: Zhennan Lin,  Kaixun Huang,  Wei Ren,  Linju Yang,  Lei Xie 
备注:Accepted by interspeech 2025
摘要:深度偏置通过合并上下文短语来提高自动语音识别(ASR)性能。然而,大多数现有的方法增强子词在上下文短语作为独立的单元,潜在地损害上下文短语的完整性,导致准确性降低。在本文中,我们提出了一个基于编码器的短语级上下文ASR方法,利用动态词汇预测和激活。我们引入了架构优化,并集成了偏置损失,以扩展基于帧级输出的短语级预测。我们还介绍了一个信心激活的解码方法,确保上下文短语的完整输出,同时抑制不正确的偏见。在Librispeech和Wenetspeech数据集上的实验表明,与基线相比,我们的方法实现了28.31%和23.49%的相对WER降低,其中上下文短语的WER相对降低了72.04%和75.69%。
摘要:Deep biasing improves automatic speech recognition (ASR) performance by incorporating contextual phrases. However, most existing methods enhance subwords in a contextual phrase as independent units, potentially compromising contextual phrase integrity, leading to accuracy reduction. In this paper, we propose an encoder-based phrase-level contextualized ASR method that leverages dynamic vocabulary prediction and activation. We introduce architectural optimizations and integrate a bias loss to extend phrase-level predictions based on frame-level outputs. We also introduce a confidence-activated decoding method that ensures the complete output of contextual phrases while suppressing incorrect bias. Experiments on Librispeech and Wenetspeech datasets demonstrate that our approach achieves relative WER reductions of 28.31% and 23.49% compared to baseline, with the WER on contextual phrases decreasing relatively by 72.04% and 75.69%.


【14】 AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker  Speech Dataset for Automatic Speech Diarization and Recognition

标题: AISHELL-5:第一个用于自动语音规模化和识别的开源车载多通道多说话人语音数据集
链接:https://arxiv.org/abs/2505.23036
作者: Yuhang Dai,  He Wang,  Xingchen Li,  Zihan Zhang,  Shuiyuan Wang,  Lei Xie,  Xin Xu,  Hongxiao Guo,  Shaoji Zhang,  Hui Bu,  Wei Chen 
备注:5 pages, 1 figures, 3 tables, accepted by InterSpeech 2025
摘要:None
摘要:This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an electric vehicle across more than 60 real driving scenarios. This audio data consists of four far-field speech signals captured by microphones located on each car door, as well as near-field signals obtained from high-fidelity headset microphones worn by each speaker. (2) a collection of 40 hours of real-world environmental noise recordings, which supports the in-car speech data simulation. Moreover, we also provide an open-access, reproducible baseline system based on this dataset. This system features a speech frontend model that employs speech source separation to extract each speaker's clean speech from the far-field signals, along with a speech recognition module that accurately transcribes the content of each individual speaker. Experimental results demonstrate the challenges faced by various mainstream ASR models when evaluated on the AISHELL-5. We firmly believe the AISHELL-5 dataset will significantly advance the research on ASR systems under complex driving scenarios by establishing the first publicly available in-car ASR benchmark.


【15】 Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of  Pre-trained Multimodal Representation via Text Updates

标题: LLM可以欺骗CLIP吗?通过文本更新对预训练多模式表示的对抗性组合进行基准测试
链接:https://arxiv.org/abs/2505.22943
作者: Jaewoo Ahn,  Heeseung Yun,  Dayoon Ko,  Gunhee Kim 
备注:ACL 2025 Main. Code is released at this https URL
摘要:虽然预先训练的多模态表示(例如,CLIP)已经显示出令人印象深刻的能力,它们表现出显著的组成漏洞,导致违反直觉的判断。我们引入了多模态对抗组合性(MAC),这是一个基准测试,它利用大型语言模型(LLM)生成欺骗性文本样本,以利用不同模态的漏洞,并通过样本攻击成功率和基于熵的多样性来评估它们。为了改进zero-shot方法,我们提出了一种自训练方法,该方法利用拒绝采样微调与多样性促进过滤,从而提高攻击成功率和样本多样性。使用Llama-3.1-8B等较小的语言模型,我们的方法在揭示各种多模态表示(包括图像,视频和音频)的组成漏洞方面表现出卓越的性能。
摘要:While pre-trained multimodal representations (e.g., CLIP) have shown impressive capabilities, they exhibit significant compositional vulnerabilities leading to counterintuitive judgments. We introduce Multimodal Adversarial Compositionality (MAC), a benchmark that leverages large language models (LLMs) to generate deceptive text samples to exploit these vulnerabilities across different modalities and evaluates them through both sample-wise attack success rate and group-wise entropy-based diversity. To improve zero-shot methods, we propose a self-training approach that leverages rejection-sampling fine-tuning with diversity-promoting filtering, which enhances both attack success rate and sample diversity. Using smaller language models like Llama-3.1-8B, our approach demonstrates superior performance in revealing compositional vulnerabilities across various multimodal representations, including images, videos, and audios.


【16】 BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural  Speech Synthesis with Flow Matching Models

标题: BinauralFlow:一种具有流匹配模型的高质量双耳语音合成的因果和可流化方法
链接:https://arxiv.org/abs/2505.22865
作者: Susan Liang,  Dejan Markovic,  Israel D. Gebru,  Steven Krenn,  Todd Keebler,  Jacob Sandakly,  Frank Yu,  Samuel Hassel,  Chenliang Xu,  Alexander Richard 
备注:ICML 2025, 18 pages
摘要:双耳渲染旨在基于单声道音频以及说话者和收听者的位置来合成模仿自然听觉的双耳音频。尽管已经提出了许多方法来解决这个问题,但它们在渲染质量和流式推理方面存在困难。合成与真实世界录音难以区分的高质量双耳音频需要对双耳线索、房间混响和环境声音进行精确建模。此外,现实世界的应用程序需要流式推理。为了解决这些挑战,我们提出了一个流匹配的流双耳语音合成框架称为BinauralFlow。我们认为双耳渲染是一个生成问题,而不是一个回归问题,并设计了一个条件流匹配模型来渲染高质量的音频。此外,我们设计了一个因果U-Net架构,仅根据过去的信息来估计当前的音频帧,以定制流推理的生成模型。最后,我们介绍了一个连续的推理流水线,结合流STFT/ISTFT操作,缓冲区银行,中点求解器,和早期跳过时间表,以提高渲染的连续性和速度。定量和定性评价表明,我们的方法比SOTA方法的优越性。感知研究进一步表明,我们的模型几乎无法与真实世界的录音区分开来,混淆率为42%。
摘要:Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-quality binaural audio that is indistinguishable from real-world recordings requires precise modeling of binaural cues, room reverb, and ambient sounds. Additionally, real-world applications demand streaming inference. To address these challenges, we propose a flow matching based streaming binaural speech synthesis framework called BinauralFlow. We consider binaural rendering to be a generation problem rather than a regression problem and design a conditional flow matching model to render high-quality audio. Moreover, we design a causal U-Net architecture that estimates the current audio frame solely based on past information to tailor generative models for streaming inference. Finally, we introduce a continuous inference pipeline incorporating streaming STFT/ISTFT operations, a buffer bank, a midpoint solver, and an early skip schedule to improve rendering continuity and speed. Quantitative and qualitative evaluations demonstrate the superiority of our method over SOTA approaches. A perceptual study further reveals that our model is nearly indistinguishable from real-world recordings, with a $42\%$ confusion rate.


【17】 StressTest: Can YOUR Speech LM Handle the Stress?

标题: 压力测试:您的言语LM能承受压力吗?
链接:https://arxiv.org/abs/2505.22765
作者: Iddo Yosha,  Gallil Maimon,  Yossi Adi 
摘要:句子重音指的是强调,放在口语中的特定单词上,以突出或对比一个想法,或者引入新的信息。它经常被用来暗示一个没有明确说明的潜在意图。语音感知语言模型(SLM)的最新进展已经实现了对音频的直接处理,允许模型绕过转录并访问语音信号的全部内容,并执行音频推理任务,例如口语问题回答。尽管句子重音在形成意义和说话者意图方面起着至关重要的作用,但在评估和开发此类模型时,它仍然在很大程度上被忽视。在这项工作中,我们通过引入StressTest来解决这一差距,StressTest是一个专门设计用于评估模型根据压力模式区分口语句子解释的能力的基准。我们评估了几个领先的SLM的性能,并发现,尽管他们的整体能力,他们在这些任务上表现不佳。为了克服这一限制,我们提出了一种新的合成数据生成管道,并创建了Stress 17 k,这是一个模拟重音变化所暗示的含义变化的训练集。然后,我们根据经验表明,使用该合成数据集优化模型与真实世界的记录非常一致,并能够有效地微调空间光调制器。结果表明,我们的微调模型,StresSLM,显着优于现有的模型在句子重音推理和检测任务。代码、模型、数据和音频示例-pages.cs.huji.ac.il/adiyoss-lab/stresstest。
摘要:Sentence stress refers to emphasis, placed on specific words within a spoken utterance to highlight or contrast an idea, or to introduce new information. It is often used to imply an underlying intention that is not explicitly stated. Recent advances in speech-aware language models (SLMs) have enabled direct processing of audio, allowing models to bypass transcription and access the full richness of the speech signal and perform audio reasoning tasks such as spoken question answering. Despite the crucial role of sentence stress in shaping meaning and speaker intent, it remains largely overlooked in evaluation and development of such models. In this work, we address this gap by introducing StressTest, a benchmark specifically designed to evaluate a model's ability to distinguish between interpretations of spoken sentences based on the stress pattern. We assess the performance of several leading SLMs and find that, despite their overall capabilities, they perform poorly on such tasks. To overcome this limitation, we propose a novel synthetic data generation pipeline, and create Stress17k, a training set that simulates change of meaning implied by stress variation. Then, we empirically show that optimizing models with this synthetic dataset aligns well with real-world recordings and enables effective finetuning of SLMs. Results suggest, that our finetuned model, StresSLM, significantly outperforms existing models on both sentence stress reasoning and detection tasks. Code, models, data, and audio samples - pages.cs.huji.ac.il/adiyoss-lab/stresstest.


【18】 FAMA: The First Large-Scale Open-Science Speech Foundation Model for  English and Italian

标题: FAMA:第一个英语和意大利语大规模开放科学演讲基金会模型
链接:https://arxiv.org/abs/2505.22759
作者: Sara Papi,  Marco Gaido,  Luisa Bentivogli,  Alessio Brutti,  Mauro Cettolo,  Roberto Gretter,  Marco Matassoni,  Mohamed Nabih,  Matteo Negri 
摘要:语音基础模型(SFM)的发展,如Whisper和M4 T,大大推进了语音处理领域。然而,它们的封闭性--无法访问训练数据和代码--带来了重大的可重复性和公平评估挑战。虽然其他领域通过开发基于开源(OS)代码和数据的完全透明的模型,在开放科学方面取得了实质性进展,但在语音方面的类似努力仍然有限。为了填补这一空白,我们引入了FAMA,这是第一个用于英语和意大利语的开放科学SFM系列,它在150 k+小时的OS语音数据上进行了训练。此外,我们还提出了一个新的数据集,其中包含两种语言的16 k小时的清洁和伪标记语音。结果表明,与现有的SFM相比,FAMA实现了具有竞争力的性能,同时速度提高了8倍。所有工件,包括代码、数据集和模型,都是在符合操作系统的许可证下发布的,这促进了语音技术研究的开放性。
摘要:The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inaccessible training data and code--poses major reproducibility and fair evaluation challenges. While other domains have made substantial progress toward open science by developing fully transparent models trained on open-source (OS) code and data, similar efforts in speech remain limited. To fill this gap, we introduce FAMA, the first family of open science SFMs for English and Italian, trained on 150k+ hours of OS speech data. Moreover, we present a new dataset containing 16k hours of cleaned and pseudo-labeled speech for both languages. Results show that FAMA achieves competitive performance compared to existing SFMs while being up to 8 times faster. All artifacts, including code, datasets, and models, are released under OS-compliant licenses, promoting openness in speech technology research.


【19】 Vision-Integrated High-Quality Neural Speech Coding

标题: 视觉集成高质量神经语音编码
链接:https://arxiv.org/abs/2505.23379
作者: Yao Guo,  Yang Ai,  Rui-Chen Zheng,  Hui-Peng Du,  Xiao-Hang Jiang,  Zhen-Hua Ling 
备注:Accepted by interspeech2025
摘要:本文提出了一种新的视觉集成神经语音编解码器(VNSC),其目的是提高语音编码质量,利用视觉模态信息。在VNSC中,图像分析-合成模块从嘴唇图像中提取视觉特征,而特征融合模块促进图像分析-合成模块和语音编码模块之间的交互,传输视觉信息以辅助语音编码过程。根据视觉信息是否是在推理阶段,功能融合模块集成到语音编码模块使用显式集成或隐式蒸馏策略的视觉特征。实验结果表明,在不增加码率的情况下,融合视觉信息有效地提高了解码语音的质量,增强了神经语音编解码器的噪声鲁棒性。
摘要:This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality information. In VNSC, the image analysis-synthesis module extracts visual features from lip images, while the feature fusion module facilitates interaction between the image analysis-synthesis module and the speech coding module, transmitting visual information to assist the speech coding process. Depending on whether visual information is available during the inference stage, the feature fusion module integrates visual features into the speech coding module using either explicit integration or implicit distillation strategies. Experimental results confirm that integrating visual information effectively improves the quality of the decoded speech and enhances the noise robustness of the neural speech codec, without increasing the bitrate.


【20】 LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable  Data for Custom Keyword Spotting

标题: LLM-Synth 4KWS:可扩展的自动生成和合成混淆数据以用于自定义关键字定位
链接:https://arxiv.org/abs/2505.22995
作者: Pai Zhu,  Quan Wang,  Dhruuv Agarwal,  Kurt Partridge 
摘要:自定义关键字定位(KWS)允许从流式音频中检测用户定义的口语关键字。这是通过比较来自语音注册和输入音频的嵌入来实现的。最先进的自定义KWS模型通常使用从训练数据集中随机抽样的关键字的话语进行对比训练。这些KWS模型经常遇到令人困惑的关键字,比如“blue”和“glue”。本文介绍了一种有效的方法来增强训练与易混淆的话语,其中关键字生成和分组从大型语言模型(LLM),和语音信号合成不同的说话风格从文本到语音(TTS)引擎。为了更好地衡量易混淆KWS的用户体验,我们定义了一个新的北极星指标,使用来自易混淆组的平均DET曲线下面积(c-AUC)。该方法具有高可扩展性和零人工成本,在语音命令测试集上,AUC提高了3.7%,c-AUC提高了11.3%。
摘要:Custom keyword spotting (KWS) allows detecting user-defined spoken keywords from streaming audio. This is achieved by comparing the embeddings from voice enrollments and input audio. State-of-the-art custom KWS models are typically trained contrastively using utterances whose keywords are randomly sampled from training dataset. These KWS models often struggle with confusing keywords, such as "blue" versus "glue". This paper introduces an effective way to augment the training with confusable utterances where keywords are generated and grouped from large language models (LLMs), and speech signals are synthesized with diverse speaking styles from text-to-speech (TTS) engines. To better measure user experience on confusable KWS, we define a new northstar metric using the average area under DET curve from confusable groups (c-AUC). Featuring high scalability and zero labor cost, the proposed method improves AUC by 3.7% and c-AUC by 11.3% on the Speech Commands testing set.


【21】 NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in  Greedy ASR Decoding

标题: NGPU-LM:贪婪ASB解码中用于上下文偏置的GOP加速N-Gram语言模型
链接:https://arxiv.org/abs/2505.22857
作者: Vladimir Bataev,  Andrei Andrusenko,  Lilit Grigoryan,  Aleksandr Laptev,  Vitaly Lavrukhin,  Boris Ginsburg 
备注:Accepted to Interspeech 2025
摘要:统计n-gram语言模型被广泛用于自动语音识别(ASR)中的上下文偏置任务。然而,现有的实现缺乏计算效率,由于不良的并行化,使上下文偏置不太吸引人的工业用途。这项工作重新考虑了统计n-gram语言模型的数据结构,以实现GPU优化推理的快速和并行操作。我们的方法名为NGPU-LM,为所有主要的ASR模型类型(包括传感器,注意力编码器-解码器模型和CTC)引入了可定制的贪婪解码,计算开销不到7%。所提出的方法可以消除超过50%的精度差距贪婪和波束搜索域外的情况下,同时避免显着放缓所造成的波束搜索。拟议的NGPU-LM的实现是开源的。
摘要:Statistical n-gram language models are widely used for context-biasing tasks in Automatic Speech Recognition (ASR). However, existing implementations lack computational efficiency due to poor parallelization, making context-biasing less appealing for industrial use. This work rethinks data structures for statistical n-gram language models to enable fast and parallel operations for GPU-optimized inference. Our approach, named NGPU-LM, introduces customizable greedy decoding for all major ASR model types - including transducers, attention encoder-decoder models, and CTC - with less than 7% computational overhead. The proposed approach can eliminate more than 50% of the accuracy gap between greedy and beam search for out-of-domain scenarios while avoiding significant slowdown caused by beam search. The implementation of the proposed NGPU-LM is open-sourced.


eess.AS音频处理


【1】 DeepFilterGAN: A Full-band Real-time Speech Enhancement System with  GAN-based Stochastic Regeneration

标题: DeepGridGAN:一个具有基于GAN随机再生的全频段实时语音增强系统
链接:https://arxiv.org/abs/2505.23515
作者: Sanberk Serbest,  Tijana Stojkovic,  Milos Cernak,  Andrew Harper 
备注:Accepted to Interspeech 2025
摘要:在这项工作中,我们提出了一个基于GAN的随机再生的全频带实时语音增强系统。预测模型侧重于估计目标分布的均值,而生成模型旨在学习完整的分布。预测模型的这种行为可能导致过度抑制,即语音内容的移除。在文献中,它表明,在随机再生框架内结合预测模型与生成模型可以减少输出中的失真。我们使用这个框架,以获得一个实时语音增强系统。具有3.58M参数和低延迟,我们的系统是专为实时流与轻量级架构。实验表明,我们的系统提高了第一阶段的NISQA-MOS度量。最后,通过消融研究,我们显示了在我们的系统中的噪声条件反射的重要性。我们用我们的模型参加了2025年紧急挑战赛,后来又做了进一步的改进。
摘要:In this work, we propose a full-band real-time speech enhancement system with GAN-based stochastic regeneration. Predictive models focus on estimating the mean of the target distribution, whereas generative models aim to learn the full distribution. This behavior of predictive models may lead to over-suppression, i.e. the removal of speech content. In the literature, it was shown that combining a predictive model with a generative one within the stochastic regeneration framework can reduce the distortion in the output. We use this framework to obtain a real-time speech enhancement system. With 3.58M parameters and a low latency, our system is designed for real-time streaming with a lightweight architecture. Experiments show that our system improves over the first stage in terms of NISQA-MOS metric. Finally, through an ablation study, we show the importance of noisy conditioning in our system. We participated in 2025 Urgent Challenge with our model and later made further improvements.


【2】 Vision-Integrated High-Quality Neural Speech Coding

标题: 视觉集成高质量神经语音编码
链接:https://arxiv.org/abs/2505.23379
作者: Yao Guo,  Yang Ai,  Rui-Chen Zheng,  Hui-Peng Du,  Xiao-Hang Jiang,  Zhen-Hua Ling 
备注:Accepted by interspeech2025
摘要:本文提出了一种新的视觉集成神经语音编解码器(VNSC),其目的是提高语音编码质量,利用视觉模态信息。在VNSC中,图像分析-合成模块从嘴唇图像中提取视觉特征,而特征融合模块促进图像分析-合成模块和语音编码模块之间的交互,传输视觉信息以辅助语音编码过程。根据视觉信息是否是在推理阶段,功能融合模块集成到语音编码模块使用显式集成或隐式蒸馏策略的视觉特征。实验结果表明,在不增加码率的情况下,融合视觉信息有效地提高了解码语音的质量,增强了神经语音编解码器的噪声鲁棒性。
摘要:This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality information. In VNSC, the image analysis-synthesis module extracts visual features from lip images, while the feature fusion module facilitates interaction between the image analysis-synthesis module and the speech coding module, transmitting visual information to assist the speech coding process. Depending on whether visual information is available during the inference stage, the feature fusion module integrates visual features into the speech coding module using either explicit integration or implicit distillation strategies. Experimental results confirm that integrating visual information effectively improves the quality of the decoded speech and enhances the noise robustness of the neural speech codec, without increasing the bitrate.


【3】 Spoken question answering for visual queries

标题: 针对可视查询的口语问答
链接:https://arxiv.org/abs/2505.23308
作者: Nimrod Shabtay,  Zvi Kons,  Avihu Dekel,  Hagai Aronowitz,  Ron Hoory,  Assaf Arbelle 
备注:Accepted for Interspeech 2025 (with additional results)
摘要:问答系统(QA)是用来回答自然语言问题的。视觉QA(VQA)和口语QA(SQA)系统扩展了文本QA系统,分别接受视觉和口语输入。   这项工作旨在创建一个系统,使用户能够通过语音和图像进行交互。这是通过融合文本,语音和图像模态来解决口语VQA(SVQA)的任务来实现的。由此产生的多模态模型具有文本、视觉和语音输入,并且可以回答图像上的语音问题。   训练和评估SVQA模型需要所有三种模式的数据集,但目前不存在这样的数据集。我们通过使用两个zero-shot TTS模型合成VQA数据集来解决这个问题。我们的初步研究结果表明,仅用合成语音训练的模型几乎达到了在文本QA上训练的上界模型的性能。此外,我们表明,TTS模型的选择对准确性有轻微的影响。
摘要:Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spoken input respectively.   This work aims to create a system that enables user interaction through both speech and images. That is achieved through the fusion of text, speech, and image modalities to tackle the task of spoken VQA (SVQA). The resulting multi-modal model has textual, visual, and spoken inputs and can answer spoken questions on images.   Training and evaluating SVQA models requires a dataset for all three modalities, but no such dataset currently exists. We address this problem by synthesizing VQA datasets using two zero-shot TTS models. Our initial findings indicate that a model trained only with synthesized speech nearly reaches the performance of the upper-bounding model trained on textual QAs. In addition, we show that the choice of the TTS model has a minor impact on accuracy.


【4】 Interspeech 2025 URGENT Speech Enhancement Challenge

标题: Interspeech 2025紧急语音增强挑战赛
链接:https://arxiv.org/abs/2505.23212
作者: Kohei Saijo,  Wangyou Zhang,  Samuele Cornell,  Robin Scheibler,  Chenda Li,  Zhaoheng Ni,  Anurag Kumar,  Marvin Sach,  Yihui Fu,  Wei Wang,  Tim Fingscheidt,  Shinji Watanabe 
备注:Accepted to Interspeech 2025
摘要:已经有越来越多的努力来开发通用语音增强(SE)以处理具有各种语音失真和记录条件的输入。紧急挑战系列旨在通过包含广泛的失真类型、增加数据多样性和纳入广泛的评估指标来促进这种普遍的SE。本文介绍了Interspeech 2025 URGENT Challenge,该系列的第二版,以探索迄今为止受到有限关注的几个方面:语言依赖性,更多失真类型的通用性,数据可扩展性以及使用噪声训练数据的有效性。我们收到了32份提交,其中最好的系统使用判别模型,而其他大多数竞争对手都是混合方法。分析揭示了一些关键发现:(i)一些生成或混合方法在主观评价中优于顶级判别模型,以及(ii)纯生成SE模型可以表现出语言依赖性。
摘要:There has been a growing effort to develop universal speech enhancement (SE) to handle inputs with various speech distortions and recording conditions. The URGENT Challenge series aims to foster such universal SE by embracing a broad range of distortion types, increasing data diversity, and incorporating extensive evaluation metrics. This work introduces the Interspeech 2025 URGENT Challenge, the second edition of the series, to explore several aspects that have received limited attention so far: language dependency, universality for more distortion types, data scalability, and the effectiveness of using noisy training data. We received 32 submissions, where the best system uses a discriminative model, while most other competitive ones are hybrid methods. Analysis reveals some key findings: (i) some generative or hybrid approaches are preferred in subjective evaluations over the top discriminative model, and (ii) purely generative SE models can exhibit language dependency.


【5】 LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable  Data for Custom Keyword Spotting

标题: LLM-Synth 4KWS:可扩展的自动生成和合成混淆数据以用于自定义关键字定位
链接:https://arxiv.org/abs/2505.22995
作者: Pai Zhu,  Quan Wang,  Dhruuv Agarwal,  Kurt Partridge 
摘要:自定义关键字定位(KWS)允许从流式音频中检测用户定义的口语关键字。这是通过比较来自语音注册和输入音频的嵌入来实现的。最先进的自定义KWS模型通常使用从训练数据集中随机抽样的关键字的话语进行对比训练。这些KWS模型经常遇到令人困惑的关键字,比如“blue”和“glue”。本文介绍了一种有效的方法来增强训练与易混淆的话语,其中关键字生成和分组从大型语言模型(LLM),和语音信号合成不同的说话风格从文本到语音(TTS)引擎。为了更好地衡量易混淆KWS的用户体验,我们定义了一个新的北极星指标,使用来自易混淆组的平均DET曲线下面积(c-AUC)。该方法具有高可扩展性和零人工成本,在语音命令测试集上,AUC提高了3.7%,c-AUC提高了11.3%。
摘要:Custom keyword spotting (KWS) allows detecting user-defined spoken keywords from streaming audio. This is achieved by comparing the embeddings from voice enrollments and input audio. State-of-the-art custom KWS models are typically trained contrastively using utterances whose keywords are randomly sampled from training dataset. These KWS models often struggle with confusing keywords, such as "blue" versus "glue". This paper introduces an effective way to augment the training with confusable utterances where keywords are generated and grouped from large language models (LLMs), and speech signals are synthesized with diverse speaking styles from text-to-speech (TTS) engines. To better measure user experience on confusable KWS, we define a new northstar metric using the average area under DET curve from confusable groups (c-AUC). Featuring high scalability and zero labor cost, the proposed method improves AUC by 3.7% and c-AUC by 11.3% on the Speech Commands testing set.


【6】 NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in  Greedy ASR Decoding

标题: NGPU-LM:贪婪ASB解码中用于上下文偏置的GOP加速N-Gram语言模型
链接:https://arxiv.org/abs/2505.22857
作者: Vladimir Bataev,  Andrei Andrusenko,  Lilit Grigoryan,  Aleksandr Laptev,  Vitaly Lavrukhin,  Boris Ginsburg 
备注:Accepted to Interspeech 2025
摘要:统计n-gram语言模型被广泛用于自动语音识别(ASR)中的上下文偏置任务。然而,现有的实现缺乏计算效率,由于不良的并行化,使上下文偏置不太吸引人的工业用途。这项工作重新考虑了统计n-gram语言模型的数据结构,以实现GPU优化推理的快速和并行操作。我们的方法名为NGPU-LM,为所有主要的ASR模型类型(包括传感器,注意力编码器-解码器模型和CTC)引入了可定制的贪婪解码,计算开销不到7%。所提出的方法可以消除超过50%的精度差距贪婪和波束搜索域外的情况下,同时避免显着放缓所造成的波束搜索。拟议的NGPU-LM的实现是开源的。
摘要:Statistical n-gram language models are widely used for context-biasing tasks in Automatic Speech Recognition (ASR). However, existing implementations lack computational efficiency due to poor parallelization, making context-biasing less appealing for industrial use. This work rethinks data structures for statistical n-gram language models to enable fast and parallel operations for GPU-optimized inference. Our approach, named NGPU-LM, introduces customizable greedy decoding for all major ASR model types - including transducers, attention encoder-decoder models, and CTC - with less than 7% computational overhead. The proposed approach can eliminate more than 50% of the accuracy gap between greedy and beam search for out-of-domain scenarios while avoiding significant slowdown caused by beam search. The implementation of the proposed NGPU-LM is open-sourced.


【7】 Spectrotemporal Modulation: Efficient and Interpretable Feature  Representation for Classifying Speech, Music, and Environmental Sounds

标题: 光谱时间调制:用于分类语音、音乐和环境声音的高效且可解释的特征表示
链接:https://arxiv.org/abs/2505.23509
作者: Andrew Chang,  Yike Li,  Iran R. Roman,  David Poeppel 
备注:Interspeech 2025
摘要:音频DNN在各种机器监听任务中表现出令人印象深刻的性能;然而,它们的大多数表示都是计算成本高且不可解释的,这为优化留下了空间。在这里,我们提出了一种新的方法集中在频谱时间调制(STM)功能,信号处理方法,模仿人类听觉皮层的神经生理表示。我们基于STM的模型的分类性能,在没有任何预训练的情况下,与预训练的音频DNN在不同的自然语音,音乐和环境声音中的分类性能相当,这些都是人类认知和机器感知的基本类别。这些结果表明,STM是音频分类的有效和可解释的特征表示,推进了机器听力的发展,为语音和听觉科学的基本理解解锁了令人兴奋的新可能性,以及开发音频BCI和认知计算。
摘要:Audio DNNs have demonstrated impressive performance on various machine listening tasks; however, most of their representations are computationally costly and uninterpretable, leaving room for optimization. Here, we propose a novel approach centered on spectrotemporal modulation (STM) features, a signal processing method that mimics the neurophysiological representation in the human auditory cortex. The classification performance of our STM-based model, without any pretraining, is comparable to that of pretrained audio DNNs across diverse naturalistic speech, music, and environmental sounds, which are essential categories for both human cognition and machine perception. These results show that STM is an efficient and interpretable feature representation for audio classification, advancing the development of machine listening and unlocking exciting new possibilities for basic understanding of speech and auditory sciences, as well as developing audio BCI and cognitive computing.


【8】 Spoken Language Modeling with Duration-Penalized Self-Supervised Units

标题: 具有持续时间惩罚自我监督单元的口语建模
链接:https://arxiv.org/abs/2505.23494
作者: Nicol Visser,  Herman Kamper 
备注:Accepted to Interspeech 2025
摘要:口语模型(SLM)对通过离散化自监督语音表示获得的声学单元进行操作。虽然这些单元的特性直接影响性能,但是码本大小和单元粗糙度之间的相互作用(即,时间)仍然未被探索。我们调查SLM的性能,因为我们不同的码本大小和单位粗糙度使用简单的持续时间惩罚动态规划(DPDP)的方法。新的分析是在不同的语言水平。在音素和单词级别,只要码本大小选择得当,粗糙度几乎没有什么好处。然而,当在再合成任务中产生整个句子时,SLM在较粗糙的单元中表现得更好。在词汇和句法语言建模任务中,较粗糙的单元在较低的比特率下也具有较高的准确性。因此,我们表明,粗糙的单位并不总是更好的,但DPDP是一个简单而有效的方法来获得粗糙的单位的任务,他们是有益的。
摘要:Spoken language models (SLMs) operate on acoustic units obtained by discretizing self-supervised speech representations. Although the characteristics of these units directly affect performance, the interaction between codebook size and unit coarseness (i.e., duration) remains unexplored. We investigate SLM performance as we vary codebook size and unit coarseness using the simple duration-penalized dynamic programming (DPDP) method. New analyses are performed across different linguistic levels. At the phone and word levels, coarseness provides little benefit, as long as the codebook size is chosen appropriately. However, when producing whole sentences in a resynthesis task, SLMs perform better with coarser units. In lexical and syntactic language modeling tasks, coarser units also give higher accuracies at lower bitrates. We therefore show that coarser units aren't always better, but that DPDP is a simple and efficient way to obtain coarser units for the tasks where they are beneficial.


【9】 EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic,  Expressiveness, and Linguistic Challenges Using Model-as-a-Judge

标题: EmergentTTS-Eval:使用模型作为评判者评估DTS模型的复杂韵律、表达性和语言挑战
链接:https://arxiv.org/abs/2505.23009
作者: Ruskin Raj Manku,  Yuzhi Tang,  Xingjian Shi,  Mu Li,  Alex Smola 
摘要:文本到语音(TTS)基准测试通常无法捕捉模型如何处理细微差别和语义复杂的文本。在$\textit{EmergentTTS}$的基础上,我们引入了$\textit{EmergentTTS-Eval}$,这是一个全面的基准测试,涵盖了六个具有挑战性的TTS场景:情感,非语言学,外国词,句法复杂性,复杂发音(例如URL,公式)和问题。至关重要的是,我们的框架自动化测试用例生成和评估,使基准易于扩展。从一小组人类书写的种子提示开始,我们使用LLM迭代地扩展它们,以针对特定的结构,语音和韵律挑战,从而产生1,645个不同的测试用例。此外,我们采用了一个模型作为一个判断的方法,使用大型音频语言模型(LALM)来评估跨多个维度的语音,如表达的情感,韵律,语调和发音准确性。我们在EmergentTTS-Eval上评估了最先进的开源和专有TTS系统,如11 Labs、Deepgram和OpenAI的4 o-mini-TTS,展示了其揭示细粒度性能差异的能力。结果表明,模型作为一个判断的方法提供了强大的TTS评估和高相关性与人类的喜好。我们开源了评估$\href{https://github.com/boson-ai/EmergentTTS-Eval-public}{code}$和$\href{https://huggingface.co/bosonets/bosonai/EmergentTTS-Eval}{dataset}$。
摘要:Text-to-Speech (TTS) benchmarks often fail to capture how well models handle nuanced and semantically complex text. Building on $\textit{EmergentTTS}$, we introduce $\textit{EmergentTTS-Eval}$, a comprehensive benchmark covering six challenging TTS scenarios: emotions, paralinguistics, foreign words, syntactic complexity, complex pronunciation (e.g. URLs, formulas), and questions. Crucially, our framework automates both test-case generation and evaluation, making the benchmark easily extensible. Starting from a small set of human-written seed prompts, we iteratively extend them using LLMs to target specific structural, phonetic and prosodic challenges, resulting in 1,645 diverse test cases. Moreover, we employ a model-as-a-judge approach, using a Large Audio Language Model (LALM) to assess the speech across multiple dimensions such as expressed emotion, prosodic, intonational, and pronunciation accuracy. We evaluate state-of-the-art open-source and proprietary TTS systems, such as 11Labs, Deepgram, and OpenAI's 4o-mini-TTS, on EmergentTTS-Eval, demonstrating its ability to reveal fine-grained performance differences. Results show that the model-as-a-judge approach offers robust TTS assessment and a high correlation with human preferences. We open source the evaluation $\href{https://github.com/boson-ai/EmergentTTS-Eval-public}{code}$ and the $\href{https://huggingface.co/datasets/bosonai/EmergentTTS-Eval}{dataset}$.


【10】 StressTest: Can YOUR Speech LM Handle the Stress?

标题: 压力测试:您的言语LM能承受压力吗?
链接:https://arxiv.org/abs/2505.22765
作者: Iddo Yosha,  Gallil Maimon,  Yossi Adi 
摘要:句子重音指的是强调,放在口语中的特定单词上,以突出或对比一个想法,或者引入新的信息。它经常被用来暗示一个没有明确说明的潜在意图。语音感知语言模型(SLM)的最新进展已经实现了对音频的直接处理,允许模型绕过转录并访问语音信号的全部内容,并执行音频推理任务,例如口语问题回答。尽管句子重音在形成意义和说话者意图方面起着至关重要的作用,但在评估和开发此类模型时,它仍然在很大程度上被忽视。在这项工作中,我们通过引入StressTest来解决这一差距,StressTest是一个专门设计用于评估模型根据压力模式区分口语句子解释的能力的基准。我们评估了几个领先的SLM的性能,并发现,尽管他们的整体能力,他们在这些任务上表现不佳。为了克服这一限制,我们提出了一种新的合成数据生成管道,并创建了Stress 17 k,这是一个模拟重音变化所暗示的含义变化的训练集。然后,我们根据经验表明,使用此合成数据集优化模型与真实世界的记录非常一致,并能够有效地微调SLM。结果表明,我们的微调模型,StresSLM,显着优于现有的模型在句子重音推理和检测任务。代码、模型、数据和音频示例-pages.cs.huji.ac.il/adiyoss-lab/stresstest。
摘要:Sentence stress refers to emphasis, placed on specific words within a spoken utterance to highlight or contrast an idea, or to introduce new information. It is often used to imply an underlying intention that is not explicitly stated. Recent advances in speech-aware language models (SLMs) have enabled direct processing of audio, allowing models to bypass transcription and access the full richness of the speech signal and perform audio reasoning tasks such as spoken question answering. Despite the crucial role of sentence stress in shaping meaning and speaker intent, it remains largely overlooked in evaluation and development of such models. In this work, we address this gap by introducing StressTest, a benchmark specifically designed to evaluate a model's ability to distinguish between interpretations of spoken sentences based on the stress pattern. We assess the performance of several leading SLMs and find that, despite their overall capabilities, they perform poorly on such tasks. To overcome this limitation, we propose a novel synthetic data generation pipeline, and create Stress17k, a training set that simulates change of meaning implied by stress variation. Then, we empirically show that optimizing models with this synthetic dataset aligns well with real-world recordings and enables effective finetuning of SLMs. Results suggest, that our finetuned model, StresSLM, significantly outperforms existing models on both sentence stress reasoning and detection tasks. Code, models, data, and audio samples - pages.cs.huji.ac.il/adiyoss-lab/stresstest.


机器翻译由腾讯交互翻译提供,仅供参考