今日论文合集:cs.SD语音28篇,eess.AS音频处理35篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音



【1】WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
标题:WildFX:一个基于AW的In-the-Wild音频FX图形建模管道
链接:http://arxiv.org/pdf/2507.10534v1

作者:g, Taylor Berg-Kirkpatrick, Julian McAuley, Zachary Novack
摘要:尽管端到端人工智能音乐生成取得了快速进展,但专业数字信号处理(DSP)工作流程的人工智能驱动建模仍然具有挑战性。特别是,虽然人们对音频效果图(例如混响,压缩,均衡)的神经黑箱建模越来越感兴趣,但基于AI的方法很难复制专业工作流程中使用的细微差别的信号流和参数交互。现有的可微插件方法通常与现实世界的工具不同,在等效的计算约束下,相对于简化的神经控制器表现出较差的性能。我们介绍了WildFX,这是一个与Docker容器化的管道,用于生成具有丰富效果图的多轨音频混合数据集,由专业的数字音频工作站(Digital Audio Workstation,简称DTS)后端提供支持。WildFX支持跨平台商业插件或任何插件的无缝集成,以VST VST 3 LV 2 CLAP格式,实现结构复杂性(例如,侧链、交叉)并实现有效的并行化处理。一个极简的元数据接口简化了项目 插件配置。实验表明,通过混合图,插件 增益参数的盲估计,以及它的桥梁AI研究与实际的DSP需求的能力,流水线的有效性。该代码可从以下网址获得:https: github.com IsaacYQH WildFX。
摘要:Despite rapid progress in end-to-end AI music generation, AI-driven modeling of professional Digital Signal Processing (DSP) workflows remains challenging. In particular, while there is growing interest in neural black-box modeling of audio effect graphs (e.g. reverb, compression, equalization), AI-based approaches struggle to replicate the nuanced signal flow and parameter interactions used in professional workflows. Existing differentiable plugin approaches often diverge from real-world tools, exhibiting inferior performance relative to simplified neural controllers under equivalent computational constraints. We introduce WildFX, a pipeline containerized with Docker for generating multi-track audio mixing datasets with rich effect graphs, powered by a professional Digital Audio Workstation (DAW) backend. WildFX supports seamless integration of cross-platform commercial plugins or any plugins in the wild, in VST VST3 LV2 CLAP formats, enabling structural complexity (e.g., sidechains, crossovers) and achieving efficient parallelized processing. A minimalist metadata interface simplifies project plugin configuration. Experiments demonstrate the pipeline's validity through blind estimation of mixing graphs, plugin gain parameters, and its ability to bridge AI research with practical DSP demands. The code is available on: https: github.com IsaacYQH WildFX.


【2】AudioMAE++: learning better masked audio representations with SwiGLU FFNs
标题:AudioMAE++:使用SwiGLU FFN学习更好的掩蔽音频表示
链接:https://arxiv.org/abs/2507.10464

作者:adav, Sergios Theodoridis, Zheng-Hua Tan
备注:TO APPEAR AT IEEE MLSP 2025
摘要:在音频频谱图补丁上训练的掩蔽自动编码器(MAE)已经成为学习自监督音频表示的一种重要方法。虽然最近的几篇论文已经评估了在音频数据上训练MAE的关键方面,但这些方法中的大多数仍然利用普通的Transformer构建块,而Transformer社区已经看到了更新的架构进步的稳定集成。在这项工作中,我们提出了AudioMAE++,一个改进的音频掩蔽自动编码器与两个这样的增强,即马卡龙风格的Transformer块与门控线性单元。当在AudioSet数据集上进行预训练时,所提出的AudioMAE++模型在10个不同的下游任务上优于现有的基于MAE的方法,在音频分类和基于语音的基准测试上表现出出色的性能。所提出的AudioMAE++模型还表现出出色的缩放特性,性能优于直接可比的标准MAE基线,参数多达4倍。
摘要:Masked Autoencoders (MAEs) trained on audio spectrogram patches have emerged as a prominent approach for learning self-supervised audio representations. While several recent papers have evaluated key aspects of training MAEs on audio data, the majority of these approaches still leverage vanilla transformer building blocks, whereas the transformer community has seen steady integration of newer architectural advancements. In this work, we propose AudioMAE++, a revamped audio masked autoencoder with two such enhancements, namely macaron-style transformer blocks with gated linear units. When pretrained on the AudioSet dataset, the proposed AudioMAE++ models outperform existing MAE based approaches on 10 diverse downstream tasks, demonstrating excellent performance on audio classification and speech-based benchmarks. The proposed AudioMAE++ models also demonstrate excellent scaling characteristics, outperforming directly comparable standard MAE baselines with up to 4x more parameters.


【3】Radif corpus: a symbolic dataset for non-metric iranian classical music
标题:Radif文集:非度量伊朗古典音乐的符号数据集
链接:https://arxiv.org/abs/2507.10456

作者:nani, Sean O Leary, James McDermott
摘要:非格律音乐是伊朗古典音乐的核心。Dastgahi音乐是伊朗艺术音乐和某些民间传统的基础理论体系。在伊朗古典音乐的核心在于radif,一个基本的曲目,组织旋律材料的核心性能和教学。   在这项研究中,我们介绍了第一个数字语料库代表完整的非韵律radif剧目,涵盖了所有13个现有的组成部分,这一剧目。我们为228首音乐提供了描述音符、音符持续时间、音程和层次结构的文件(总共约281分钟)和数据电子表格。我们忠实地代表音调,包括四分之一色调,和非公制方面。此外,我们还提供了支持的基本统计数据,以及对语料库的复杂性和相似性的度量。   我们的语料库为伊朗古典音乐的计算研究提供了一个平台。研究人员可以将其用于研究旋律模式,调查即兴风格,或用于音乐信息检索,音乐理论和计算(民族)音乐学的其他任务。
摘要:Non-metric music forms the core of the repertoire in Iranian classical music. Dastgahi music serves as the underlying theoretical system for both Iranian art music and certain folk traditions. At the heart of Iranian classical music lies the radif, a foundational repertoire that organizes melodic material central to performance and pedagogy.   In this study, we introduce the first digital corpus representing the complete non-metrical radif repertoire, covering all 13 existing components of this repertoire. We provide MIDI files (about 281 minutes in total) and data spreadsheets describing notes, note durations, intervals, and hierarchical structures for 228 pieces of music. We faithfully represent the tonality including quarter-tones, and the non-metric aspect. Furthermore, we provide supporting basic statistics, and measures of complexity and similarity over the corpus.   Our corpus provides a platform for computational studies of Iranian classical music. Researchers might employ it in studying melodic patterns, investigating improvisational styles, or for other tasks in music information retrieval, music theory, and computational (ethno)musicology.


【4】Evaluating Fake Music Detection Performance Under Audio Augmentations
标题:评估音频增强下的假音乐检测性能
链接:https://arxiv.org/abs/2507.10447

作者:oka, Tomasz Wężowicz, Dominik Sidorczuk, Mateusz Modrzejewski
备注:ISMIR 2025 LBD, 2 pages + bibliography, 1 figure
摘要:随着生成音频模型的快速发展,区分人类创作和生成的音乐变得越来越具有挑战性。作为响应,已经提出了用于检测假音乐的模型。在这项工作中,我们探讨了这种系统的鲁棒性下的音频增强。为了评估模型的泛化能力,我们构建了一个由使用多个系统生成的真实音乐和合成音乐组成的数据集。然后,我们应用一系列音频转换,并分析它们如何影响分类准确性。我们测试了最近最先进的音乐deepfake检测模型在存在音频增强的情况下的性能。该模型的性能显着下降,即使与光增强的引入。
摘要:With the rapid advancement of generative audio models, distinguishing between human-composed and generated music is becoming increasingly challenging. As a response, models for detecting fake music have been proposed. In this work, we explore the robustness of such systems under audio augmentations. To evaluate model generalization, we constructed a dataset consisting of both real and synthetic music generated using several systems. We then apply a range of audio transformations and analyze how they affect classification accuracy. We test the performance of a recent state-of-the-art musical deepfake detection model in the presence of audio augmentations. The performance of the model decreases significantly even with the introduction of light augmentations.


【5】DQLoRA: A Lightweight Domain-Aware Denoising ASR via Adapter-guided Distillation
标题:DQLoRA:一种轻量级领域感知,通过适配器引导蒸馏降噪ASB
链接:https://arxiv.org/abs/2507.10313

作者:
摘要:我们展示了一个DQLoRA的演示,这是一个适配器引导的蒸馏框架,用于在低资源和噪声条件下进行鲁棒的语音识别。我们的方法采用冻结的Whisper模型作为教师来提供语义监督,以及配备基于QLoRA的适配器的轻量级Wav2Vec2学生。训练是在使用DNS风格噪声增强的FLEURS数据集上进行的。通过联合最小化CTC损失和基于KL的蒸馏损失来优化学生,从而在保持识别准确性的同时实现有效的自适应。
摘要:We present a demo of DQLoRA, an Adapter-Guided Distillation framework for robust speech recognition under low-resource and noisy conditions. Our method employs a frozen Whisper model as the teacher to provide semantic supervision, and a lightweight Wav2Vec2 student equipped with QLoRA-based Adapters. Training is conducted on the FLEURS dataset augmented with DNS-style noise. The student is optimized by jointly minimizing CTC loss and KL-based distillation loss, enabling efficient adaptation while preserving recognition accuracy.


【6】DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
标题:DualDub:通过语音和背景音频联合合成的视频到音轨生成
链接:https://arxiv.org/abs/2507.10109

作者:an, Xinfa Zhu, Haohe Liu, Zhixian Zhao, Zihao Chen, Chaofan Ding, Xinhan Di, Junjie Zheng, Lei Xie
摘要:虽然最近的视频到音频(V2A)模型可以从视觉输入生成逼真的背景音频,但它们在很大程度上忽略了语音,这是许多视频配乐的重要组成部分。本文提出了一个新的任务,视频到配乐(V2ST)的生成,其目的是在一个统一的框架内联合产生同步的背景音频和语音。为了解决V2ST问题,我们引入了DualDub,这是一个基于多模态语言模型的统一框架,它集成了多模态编码器、跨模态对齐器和双解码头,用于同时生成背景音频和语音。具体来说,我们提出的跨模态对准器采用因果和非因果的注意机制,以提高同步和声学和谐。此外,为因应资料缺乏的问题,本研究设计了一套课程学习策略,逐步建立多模态能力。最后,我们介绍了DualBench,这是V2ST评估的第一个基准,具有精心策划的测试集和全面的指标。实验结果表明,DualDub实现了最先进的性能,生成高质量和良好同步的音轨与语音和背景音频。
摘要:While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST) generation, which aims to jointly produce synchronized background audio and speech within a unified framework. To tackle V2ST, we introduce DualDub, a unified framework built on a multimodal language model that integrates a multimodal encoder, a cross-modal aligner, and dual decoding heads for simultaneous background audio and speech generation. Specifically, our proposed cross-modal aligner employs causal and non-causal attention mechanisms to improve synchronization and acoustic harmony. Besides, to handle data scarcity, we design a curriculum learning strategy that progressively builds the multimodal capability. Finally, we introduce DualBench, the first benchmark for V2ST evaluation with a carefully curated test set and comprehensive metrics. Experimental results demonstrate that DualDub achieves state-of-the-art performance, generating high-quality and well-synchronized soundtracks with both speech and background audio.


【7】The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents
标题:声音背后的人:通过多模式大型语言模型代理揭开音频私人属性分析的神秘面纱
链接:https://arxiv.org/abs/2507.10016

作者:, Kaixiang Yao, Xinfeng Li, Dong Yang, Haoyang Li, Xiaofeng Wang, Wei Dong
备注:22 pages, 4 figures
摘要:我们的研究揭示了与多模态大型语言模型(MLLM)相关的一种新的隐私风险:从音频数据中推断敏感个人属性的能力-我们称之为音频私有属性分析的技术。这种能力构成了重大威胁,因为音频可以在没有直接交互或可见性的情况下被秘密捕获。此外,与图像和文本相比,音频具有独特的特征,例如音调和音高,可以用于更详细的分析。然而,在理解MLLM采用的来自音频的私有属性剖析中存在两个关键挑战:(1)缺乏具有敏感属性注释的音频基准数据集,以及(2)当前MLLM直接从音频推断这些属性的能力有限。为了解决这些挑战,我们引入了AP^2,这是一个音频基准数据集,由从真实世界数据中收集和组成的两个子集组成,并且都用敏感属性标签进行了注释。此外,我们还提出了Gifts,这是一个混合多代理框架,它利用了音频语言模型(ALMs)和大型语言模型(LLM)的互补优势来增强推理能力。Gifts采用LLM来指导ALM推断敏感属性,然后从法医学角度分析和巩固ALM的推断,克服现有ALM在生成长上下文响应时的严重幻觉。我们的评估表明,礼物显着优于基线方法在推断敏感属性。最后,我们研究了模型级和数据级防御策略,以减轻音频私有属性分析的风险。我们的工作验证了使用MLLM进行基于音频的隐私攻击的可行性,强调了强大防御的必要性,并提供了一个数据集和框架,以促进未来的研究。
摘要:Our research uncovers a novel privacy risk associated with multimodal large language models (MLLMs): the ability to infer sensitive personal attributes from audio data -- a technique we term audio private attribute profiling. This capability poses a significant threat, as audio can be covertly captured without direct interaction or visibility. Moreover, compared to images and text, audio carries unique characteristics, such as tone and pitch, which can be exploited for more detailed profiling. However, two key challenges exist in understanding MLLM-employed private attribute profiling from audio: (1) the lack of audio benchmark datasets with sensitive attribute annotations and (2) the limited ability of current MLLMs to infer such attributes directly from audio. To address these challenges, we introduce AP^2, an audio benchmark dataset that consists of two subsets collected and composed from real-world data, and both are annotated with sensitive attribute labels. Additionally, we propose Gifts, a hybrid multi-agent framework that leverages the complementary strengths of audio-language models (ALMs) and large language models (LLMs) to enhance inference capabilities. Gifts employs an LLM to guide the ALM in inferring sensitive attributes, then forensically analyzes and consolidates the ALM's inferences, overcoming severe hallucinations of existing ALMs in generating long-context responses. Our evaluations demonstrate that Gifts significantly outperforms baseline approaches in inferring sensitive attributes. Finally, we investigate model-level and data-level defense strategies to mitigate the risks of audio private attribute profiling. Our work validates the feasibility of audio-based privacy attacks using MLLMs, highlighting the need for robust defenses, and provides a dataset and framework to facilitate future research.


【8】ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
标题:ASTAR-NTU解决方案AudioMOS挑战赛2025 Track 1
链接:https://arxiv.org/abs/2507.09904

作者:tter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei, Jeremy H.M. Wong, Nancy F. Chen, Hung-yi Lee
备注:Under Review - Submitted to AudioMOS Challenge 2025 - ASRU 2025
摘要:文本到音乐系统的评估受到成本和收集专家进行评估的可用性的限制。AudioMOS 2025 Challenge音轨1被创建为自动预测音乐印象(MI)以及提示和生成的音乐作品之间的文本对齐(TA)。本文报告了我们的获奖系统,该系统使用双分支架构,并将预训练的MuQ和RoBERTa模型作为音频和文本编码器。交叉注意机制融合了音频和文本表示。对于训练,我们将MI和TA预测重新构建为分类任务。为了结合MOS分数的顺序性质,使用高斯内核将独热标签转换为软分布。在官方测试集上,用该方法训练的单个模型实现了MI的系统级斯皮尔曼等级相关系数(SRCC)为0.991,TA的SRCC为0.952,对应于MI SRCC中的21.21%和TA SRCC中的31.47%的挑战基线的相对改善。
摘要:Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically predict music impression (MI) as well as text alignment (TA) between the prompt and the generated musical piece. This paper reports our winning system, which uses a dual-branch architecture with pre-trained MuQ and RoBERTa models as audio and text encoders. A cross-attention mechanism fuses the audio and text representations. For training, we reframe the MI and TA prediction as a classification task. To incorporate the ordinal nature of MOS scores, one-hot labels are converted to a soft distribution using a Gaussian kernel. On the official test set, a single model trained with this method achieves a system-level Spearman's Rank Correlation Coefficient (SRCC) of 0.991 for MI and 0.952 for TA, corresponding to a relative improvement of 21.21\% in MI SRCC and 31.47\% in TA SRCC over the challenge baseline.


【9】Knowing When to Quit: Probabilistic Early Exits for Speech Separation
标题:知道何时退出:可能因言语分离而提前退出
链接:https://arxiv.org/abs/2507.09768

作者:kær Olsen. Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen, Rasmus Malik Høegh Lindrup, Bjørn Sand Jensen, Morten Mørup
摘要:近年来,基于深度学习的单通道语音分离有了很大的改进,这在很大程度上是由越来越多的计算和参数高效的神经网络架构驱动的。然而,大多数这样的架构被设计为具有固定的计算和参数预算,并且因此不能扩展到变化的计算需求或资源,这限制了它们在嵌入式和异构设备(诸如移动电话和耳机)中的使用。为了实现这样的用例,我们设计了一个能够提前退出的语音分离神经网络架构,我们提出了一个不确定性感知的概率框架来联合建模干净的语音信号和误差方差,我们使用它来推导概率提前退出条件所需的信噪比。我们在语音分离和增强任务上评估了我们的方法,并且我们表明,单个早期退出模型可以与在许多计算和参数预算下训练的最先进的模型竞争。我们的框架能够实现语音分离网络的细粒度动态计算缩放,同时实现最先进的性能和可解释的退出条件。
摘要:In recent years, deep learning-based single-channel speech separation has improved considerably, in large part driven by increasingly compute- and parameter-efficient neural network architectures. Most such architectures are, however, designed with a fixed compute and parameter budget, and consequently cannot scale to varying compute demands or resources, which limits their use in embedded and heterogeneous devices such as mobile phones and hearables. To enable such use-cases we design a neural network architecture for speech separation capable of early-exit, and we propose an uncertainty-aware probabilistic framework to jointly model the clean speech signal and error variance which we use to derive probabilistic early-exit conditions in terms of desired signal-to-noise ratios. We evaluate our methods on both speech separation and enhancement tasks, and we show that a single early-exit model can be competitive with state-of-the-art models trained at many compute and parameter budgets. Our framework enables fine-grained dynamic compute-scaling of speech separation networks while achieving state-of-the-art performance and interpretable exit conditions.


【10】MB-RIRs: a Synthetic Room Impulse Response Dataset with Frequency-Dependent Absorption Coefficients
标题:MB-RIR:具有频率相关吸收系数的合成房间脉冲响应数据集
链接:https://arxiv.org/abs/2507.09750

作者:ó, Joanna Luberadzka, Umut Sayin, Xavier Serra
备注:Accepted to WASPAA25
摘要:我们研究了四种策略对改善单声道语音增强(SE)的合成房间脉冲响应(RIR)数据集的生态有效性的影响。在传统的基于图像源方法(ISM)鞋盒RIR的基础上,我们实现了三个特征:多波段吸收系数、源方向性和接收器方向性。我们还考虑了SoundSpaces数据集中基于网格的RIR。然后,我们为每个RIR数据集训练DeepFilternet 3模型,并客观和主观地评估真实RIR测试集的性能。我们发现使用频率相关吸声系数(MB-RIR)的RIR在真实RIR上评估时可以获得+0.51dB的SDR和+8.9的MUSHRA评分。MB-RIR数据集公开提供免费下载。
摘要:We investigate the effects of four strategies for improving the ecological validity of synthetic room impulse response (RIR) datasets for monoaural Speech Enhancement (SE). We implement three features on top of the traditional image source method-based (ISM) shoebox RIRs: multiband absorption coefficients, source directivity and receiver directivity. We additionally consider mesh-based RIRs from the SoundSpaces dataset. We then train a DeepFilternet3 model for each RIR dataset and evaluate the performance on a test set of real RIRs both objectively and subjectively. We find that RIRs which use frequency-dependent acoustic absorption coefficients (MB-RIRs) can obtain +0.51dB of SDR and a +8.9 MUSHRA score when evaluated on real RIRs. The MB-RIRs dataset is publicly available for free download.


【11】THAI Speech Emotion Recognition (THAI-SER) corpus
标题:泰语语音情感识别(THAI-SER)语料库
链接:https://arxiv.org/abs/2507.09618

作者:Wongpithayadisai, Chompakorn Chaksangchaichot, Soravitt Sangnark, Patawee Prakrankamanant, Krit Gangwanpongpun, Siwa Boonpunmongkol, Premmarin Milindasuta, Dangkamon Na-Pombejra, Sarana Nutanong, Ekapol Chuangsuwanich
摘要:我们提出了第一个相当大的语料库泰国语音情感识别,THAI-SER,包含41小时36分钟(27,854话语),从100个录音在不同的录音环境:变焦和两个工作室设置。这些录音包括脚本和即兴表演,由200名专业演员(112名女性和88名男性,年龄在18至55岁之间)表演,并由专业导演执导。有五种主要的情绪:中性,愤怒,快乐,悲伤和沮丧,在记录话语时分配给演员。使用众包用情感类别注释话语。为了控制注释过程的质量,我们还设计了一个广泛的过滤和质量控制方案,以确保大多数协议得分保持在0.71以上。我们使用两个指标来评估我们的注释语料库:注释者间的可靠性和人类识别的准确性。使用Krippendorff的alpha计算注释者间的可靠性得分,其中我们的语料库在过滤后达到了0.692的alpha得分,高于推荐的0.667。对于人类识别准确性,我们的语料库在过滤后得分高达0.772。我们还提供了在语料库内和跨语料库设置上评估的语料库上训练的模型的结果。语料库是公开的知识共享BY-SA 4.0下,以及我们的实验代码。
摘要:We present the first sizeable corpus of Thai speech emotion recognition, THAI-SER, containing 41 hours and 36 minutes (27,854 utterances) from 100 recordings made in different recording environments: Zoom and two studio setups. The recordings contain both scripted and improvised sessions, acted by 200 professional actors (112 females and 88 males, aged 18 to 55) and were directed by professional directors. There are five primary emotions: neutral, angry, happy, sad, and frustrated, assigned to the actors when recording utterances. The utterances are annotated with an emotional category using crowdsourcing. To control the annotation process's quality, we also design an extensive filtering and quality control scheme to ensure that the majority agreement score remains above 0.71. We evaluate our annotated corpus using two metrics: inter-annotator reliability and human recognition accuracy. Inter-annotator reliability score was calculated using Krippendorff's alpha, where our corpus, after filtering, achieved an alpha score of 0.692, higher than a recommendation of 0.667. For human recognition accuracy, our corpus scored up to 0.772 post-filtering. We also provide the results of the model trained on the corpus evaluated on both in-corpus and cross-corpus setups. The corpus is publicly available under a Creative Commons BY-SA 4.0, as well as our codes for the experiments.


【12】Ensemble Confidence Calibration for Sound Event Detection in Open-environment
标题:开放环境下声事件检测的包围置信度标定
链接:https://arxiv.org/abs/2507.09606

作者:Chen, Han Yin
摘要:声音事件检测(SED)在具有明确事件类别的受控环境中取得了很大进展。然而,现实世界的应用程序往往发生在开放的环境中。在这种情况下,目前的方法往往会产生过于自信的预测,缺乏适当的方法来衡量不确定性。这限制了他们在新情况下的适应能力和良好表现。为了解决这个问题,我们是第一个使用集成方法在SED,以提高对域外(OOD)输入的鲁棒性。我们提出了一种称为基于能量的开放世界Softmax(EOW-Softmax)的置信度校准方法,该方法可以帮助系统更好地处理未知场景中的不确定性。我们进一步将EOW-Softmax应用于声音发生和重叠检测(SOD),通过调整预测。通过这种方式,模型变得更具适应性,同时保持其检测重叠事件的能力。实验表明,我们的方法提高了在开放环境中的性能。它减少过度自信,提高处理OOD情况的能力。
摘要:Sound event detection (SED) has made strong progress in controlled environments with clear event categories. However, real-world applications often take place in open environments. In such cases, current methods often produce predictions with too much confidence and lack proper ways to measure uncertainty. This limits their ability to adapt and perform well in new situations. To solve this problem, we are the first to use ensemble methods in SED to improve robustness against out-of-domain (OOD) inputs. We propose a confidence calibration method called Energy-based Open-World Softmax (EOW-Softmax), which helps the system better handle uncertainty in unknown scenes. We further apply EOW-Softmax to sound occurrence and overlap detection (SOD) by adjusting the prediction. In this way, the model becomes more adaptable while keeping its ability to detect overlapping events. Experiments show that our method improves performance in open environments. It reduces overconfidence and increases the ability to handle OOD situations.


【13】SC-TSE: Speaker Consistency-Aware Target Speaker Extraction
标题:SC-PSE:说话者一致性感知目标说话者提取
链接:https://arxiv.org/abs/2507.09510

作者:nbin Qi, Yanzhang Xie, Xiang Xie
备注:Accept to Interspeech2025
摘要:目标说话人提取(TSE)使用参考线索从混合语音中提取目标语音。在依赖于音频线索的TSE系统中,从注册语音中嵌入说话人对性能至关重要。然而,这些嵌入可能遭受说话人身份混淆。与以往的研究侧重于提高说话人嵌入提取不同,本文从说话人一致性的角度提高了TSE的性能。在本文中,我们提出了一个说话人一致性意识的目标说话人提取方法,结合了基于质心的说话人一致性损失。该方法通过确保登记的语音和提取的语音之间的说话者一致性来增强TSE性能。此外,我们将条件损失抑制集成到训练过程中。实验结果验证了我们提出的方法在提高TSE性能的有效性。在线提供演讲演示。\脚注{https:sc-tse.netlify.app/
摘要:Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled speech is crucial to performance. However, these embeddings may suffer from speaker identity confusion. Unlike previous studies that focus on improving speaker embedding extraction, we improve TSE performance from the perspective of speaker consistency. In this paper, we propose a speaker consistency-aware target speaker extraction method that incorporates a centroid-based speaker consistency loss. This approach enhances TSE performance by ensuring speaker consistency between the enrolled and extracted speech. In addition, we integrate conditional loss suppression into the training process. The experimental results validate the effectiveness of our proposed methods in advancing the TSE performance. A speech demo is available online.\footnote{https://sc-tse.netlify.app/


【14】Acoustic Wave Modeling Using 2D FDTD: Applications in Unreal Engine For Dynamic Sound Rendering
标题:使用2D MTD的声波建模:在动态声音渲染的虚幻引擎中的应用
链接:https://arxiv.org/abs/2507.09376

作者:amsurya
备注:Accepted to the 50th International Computer Music Conference (ICMC), 2025
摘要:准确的声音传播仿真对于在虚拟应用中提供沉浸式体验至关重要,但声学建模的行业方法通常无法解释声波现象的全部范围。本文提出了一种新的二维时域有限差分(FDTD)框架,在虚幻引擎中将声音传播模拟为基于波的模型,重点是捕获低频波现象,在生成的脉冲响应中嵌入遮挡、衍射、反射和干涉。该过程首先通过自上而下的投影将场景几何形状离散化为2D网格,从中导出障碍物遮罩和边界条件。基于Python的FDTD求解器在源位置注入正弦扫描,虚拟四声道麦克风阵列在预定义的收听者位置记录压力场响应。压力响应的去卷积产生多通道脉冲响应,这些脉冲响应保持空间方向性,然后将其集成到虚幻引擎的音频管道中进行动态播放。基准测试证实了与分析预期的一致性,论文概述了旨在实现商业可行性的混合扩展。
摘要:Accurate sound propagation simulation is essential for delivering immersive experiences in virtual applications, yet industry methods for acoustic modeling often do not account for the full breadth of acoustic wave phenomena. This paper proposes a novel two-dimensional (2D) finite-difference time-domain (FDTD) framework that simulates sound propagation as a wave-based model in Unreal Engine, with an emphasis on capturing lower frequency wave phenomena, embedding occlusion, diffraction, reflection and interference in generated impulse responses. The process begins by discretizing the scene geometry into a 2D grid via a top-down projection from which obstacle masks and boundary conditions are derived. A Python-based FDTD solver injects a sine sweep at a source position, and virtual quadraphonic microphone arrays record pressure field responses at pre-defined listener positions. De-convolution of the pressure responses yields multi-channel impulse responses that retain spatial directionality which are then integrated into Unreal Engine's audio pipeline for dynamic playback. Benchmark tests confirm agreement with analytical expectations, and the paper outlines hybrid extensions aimed at commercial viability.


【15】BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus
标题:BENYO-S2 ST-Corpus-1:英语到约鲁巴语的双语直接Speech-to-Speech翻译库
链接:https://arxiv.org/abs/2507.09342

作者:Adetiba, Abdultaofeek Abayomi, Raymond J. Kala, Ayodele H. Ifijeh, Oluwatobi E. Dare, Olabode Idowu-Bismark, Gabriel O. Sobola, Joy N. Adetiba, Monsurat Adepeju Lateef, Heather Cole-Lewis
摘要:语音到语音翻译(S2 ST)数据集严重短缺,用于高资源到低资源的语言对,如英语到约鲁巴语。因此,在这项研究中,我们策划了双语英语到约鲁巴语语音到语音翻译语料库版本1(BENYO-S2 ST-Corpus-1)。语料库是基于一个混合架构,我们开发的大规模直接S2 ST语料库创建以降低成本。为了实现这一目标,我们利用了非语音到语音的标准约鲁巴语(SY)实时音频和成绩单在YORULECT语料库以及相应的标准英语(SE)成绩单。YORULECT语料库是小规模的(1,504)样本,它没有配对的英语音频。因此,我们使用预训练的AI模型(即Facebook MMS)生成SE音频。我们还开发了一种名为AcoustAug的音频增强算法,基于三个潜在的声学特征,从两种语言的原始音频生成增强音频。BENYO-S2 ST-Corpus-1每种语言有12,032个音频样本,总共有24,064个样本大小。两种语言的音频总时长为41.20小时。这个尺寸是相当重要的。除了构建S2 ST模型之外,BENYO-S2 ST-Corpus-1还可以用于构建预训练模型或改进现有模型。使用创建的语料库和Coqui框架构建预训练的约鲁巴语TTS模型(命名为YoruTTS-0.5)作为概念证明。YoruTTS-0.5在1,000个epoch之后给出了63.54的F0 RMSE值,这表明与参考实时音频具有中等的基本音高相似性。最终,研究人员和开发人员可以利用本研究中的语料库架构来管理多语言高资源到低资源非洲语言的数据集。这将弥合高资源和低资源语言对之间翻译的巨大数字鸿沟。BENYO-S2 ST-Corpus-1和YoruTTS-0.5可在(https://bit.ly/40bGMwi)公开获得。
摘要:There is a major shortage of Speech-to-Speech Translation (S2ST) datasets for high resource-to-low resource language pairs such as English-to-Yoruba. Thus, in this study, we curated the Bilingual English-to-Yoruba Speech-to-Speech Translation Corpus Version 1 (BENYO-S2ST-Corpus-1). The corpus is based on a hybrid architecture we developed for large-scale direct S2ST corpus creation at reduced cost. To achieve this, we leveraged non speech-to-speech Standard Yoruba (SY) real-time audios and transcripts in the YORULECT Corpus as well as the corresponding Standard English (SE) transcripts. YORULECT Corpus is small scale(1,504) samples, and it does not have paired English audios. Therefore, we generated the SE audios using pre-trained AI models (i.e. Facebook MMS). We also developed an audio augmentation algorithm named AcoustAug based on three latent acoustic features to generate augmented audios from the raw audios of the two languages. BENYO-S2ST-Corpus-1 has 12,032 audio samples per language, which gives a total of 24,064 sample size. The total audio duration for the two languages is 41.20 hours. This size is quite significant. Beyond building S2ST models, BENYO-S2ST-Corpus-1 can be used to build pretrained models or improve existing ones. The created corpus and Coqui framework were used to build a pretrained Yoruba TTS model (named YoruTTS-0.5) as a proof of concept. The YoruTTS-0.5 gave a F0 RMSE value of 63.54 after 1,000 epochs, which indicates moderate fundamental pitch similarity with the reference real-time audio. Ultimately, the corpus architecture in this study can be leveraged by researchers and developers to curate datasets for multilingual high-resource-to-low-resource African languages. This will bridge the huge digital divides in translations among high and low-resource language pairs. BENYO-S2ST-Corpus-1 and YoruTTS-0.5 are publicly available at (https://bit.ly/40bGMwi).


【16】Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning
标题:具有隐性和显式声学特征条件反射的伦巴第说话风格的声音转换
链接:https://arxiv.org/abs/2507.09310

作者:Woszczyk, Manuel Sam Ribeiro, Thomas Merritt, Daniel Korzekwa
备注:Presented at Clarity Challenge 2023
摘要:朗伯说话风格的文本到语音(TTS)系统可以提高语音的整体清晰度,适用于听力损失和嘈杂环境。然而,训练这些模型需要大量的数据,并且由于扬声器和噪声的可变性以及令人疲劳的记录条件,Lombard效应的记录具有挑战性。语音转换(VC)已被证明是一种有用的增强技术,训练TTS系统的记录数据的情况下,从目标说话人在目标说话风格。在本文中,我们关注的是伦巴第语风格的迁移。我们的目标是转换扬声器身份,同时保留定义伦巴第语风格的声学属性。我们比较语音转换模型与隐式和显式的声学特征条件。我们观察到,我们提出的隐式条件反射策略实现了可懂度增益相比,模型条件显式的声学特征,同时还保持扬声器的相似性。
摘要:Text-to-Speech (TTS) systems in Lombard speaking style can improve the overall intelligibility of speech, useful for hearing loss and noisy conditions. However, training those models requires a large amount of data and the Lombard effect is challenging to record due to speaker and noise variability and tiring recording conditions. Voice conversion (VC) has been shown to be a useful augmentation technique to train TTS systems in the absence of recorded data from the target speaker in the target speaking style. In this paper, we are concerned with Lombard speaking style transfer. Our goal is to convert speaker identity while preserving the acoustic attributes that define the Lombard speaking style. We compare voice conversion models with implicit and explicit acoustic feature conditioning. We observe that our proposed implicit conditioning strategy achieves an intelligibility gain comparable to the model conditioned on explicit acoustic features, while also preserving speaker similarity.


【17】ClaritySpeech: Dementia Obfuscation in Speech
标题:ClaritySpeech:言语中的痴呆症混淆
链接:https://arxiv.org/abs/2507.09282

作者:Woszczyk, Ranya Aloufi, Soteris Demetriou
备注:Accepted at Interspeech 2025
摘要:痴呆症是一种神经退行性疾病,会改变言语模式,造成沟通障碍,并引发隐私问题。目前的语音技术,如自动语音转录(ASR),与痴呆症和非典型语音斗争,进一步挑战可访问性。本文提出了一种新的痴呆模糊语音框架ClaritySpeech,集成ASR,文本模糊,和zero-shot文本到语音(TTS),以纠正受痴呆影响的语音,同时在低数据环境中保持说话人身份,无需微调。结果显示,在ADReSS和ADReSSo的各种对抗性设置和模式(音频,文本,融合)中,平均F1得分分别下降了16%和10%,保持了50%的说话人相似性。我们还发现,我们的系统提高了WER(从0.73到0.08 ADReSS和0.15 ADReSSo)和语音质量从1.65到~2.15,提高了隐私和可访问性。
摘要:Dementia, a neurodegenerative disease, alters speech patterns, creating communication barriers and raising privacy concerns. Current speech technologies, such as automatic speech transcription (ASR), struggle with dementia and atypical speech, further challenging accessibility. This paper presents a novel dementia obfuscation in speech framework, ClaritySpeech, integrating ASR, text obfuscation, and zero-shot text-to-speech (TTS) to correct dementia-affected speech while preserving speaker identity in low-data environments without fine-tuning. Results show a 16% and 10% drop in mean F1 score across various adversarial settings and modalities (audio, text, fusion) for ADReSS and ADReSSo, respectively, maintaining 50% speaker similarity. We also find that our system improves WER (from 0.73 to 0.08 for ADReSS and 0.15 for ADReSSo) and speech quality from 1.65 to ~2.15, enhancing privacy and accessibility.


【18】Towards Spatial Audio Understanding via Question Answering
标题:基于问题生成的空间音频理解
链接:https://arxiv.org/abs/2507.09195

作者:rathy Sudarsanam, Archontis Politis
摘要:在本文中,我们介绍了一种新的框架,通过问答(QA)范式的一阶立体混响(FOA)信号的空间音频理解,旨在扩展的范围内的声音事件定位和检测(SELD)对空间场景的理解和推理。首先,我们使用基于规则的方法为STARSS 23数据集策划和发布细粒度的时空文本描述,并使用基于大型语言模型(LLM)的改写进一步增强语言多样性。我们还介绍了与STARSS23场景对齐的QA数据集,涵盖了事件存在,本地化,空间和时间关系等各个方面。为了增加语言的多样性,我们再次利用LLM为每个问题生成多个改写。最后,我们开发了一个基线空间音频QA模型,该模型以FOA信号和自然语言问题为输入,并提供关于场景中声音事件的各种发生,时间和空间关系的答案,该模型被制定为分类任务。尽管仅使用场景级问答监督进行训练,但我们的模型实现了与使用帧级时空注释训练的完全监督的声音事件定位和检测模型相当的性能。结果突出了空间音频理解的语言指导方法的潜力,并为将语言监督整合到空间场景分析中开辟了新的方向。
摘要:In this paper, we introduce a novel framework for spatial audio understanding of first-order ambisonic (FOA) signals through a question answering (QA) paradigm, aiming to extend the scope of sound event localization and detection (SELD) towards spatial scene understanding and reasoning. First, we curate and release fine-grained spatio-temporal textual descriptions for the STARSS23 dataset using a rule-based approach, and further enhance linguistic diversity using large language model (LLM)-based rephrasing. We also introduce a QA dataset aligned with the STARSS23 scenes, covering various aspects such as event presence, localization, spatial, and temporal relationships. To increase language variety, we again leverage LLMs to generate multiple rephrasings per question. Finally, we develop a baseline spatial audio QA model that takes FOA signals and natural language questions as input and provides answers regarding various occurrences, temporal, and spatial relationships of sound events in the scene formulated as a classification task. Despite being trained solely with scene-level question answering supervision, our model achieves performance that is comparable to a fully supervised sound event localization and detection model trained with frame-level spatiotemporal annotations. The results highlight the potential of language-guided approaches for spatial audio understanding and open new directions for integrating linguistic supervision into spatial scene analysis.


【19】Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
标题:LoRA专家的混合具有多模式和多粒度LLM生成错误纠正,用于强调语音识别
链接:https://arxiv.org/abs/2507.09116

作者:Mu, Kun Wei, Pengcheng Guo, Lei Xie
备注:IEEE Transactions on Audio, Speech and Language Processing
摘要:尽管ASR有了很大的改进,但在面对不利条件(如说话者口音)时,性能往往会下降。生成式纠错(GER)利用LLM丰富的语言知识和卓越的推理能力,显著优于典型的LM方法。然而,它在口音语音场景中缺乏特异性。在这项研究中,我们利用GER,以提高准确性的转录预测,解决口音语音识别的两个主要特征。为了充分利用发音信息,我们提出了多模态GER,它集成了语音模态的发音信息,以及多粒度GER,它包含了与发音相关的细粒度音素级信息。这两种方法使LLM能够利用口音语音的发音信息和来自单词级别假设的语义信息,通过LoRA微调进行准确的转录预测。一方面,我们采用三阶段训练策略来为每个口音训练单独的多模态GER模型,以获得单口音LoRA专家。通过采用我们提出的HDMoLE方法,该方法在LoRA专家的混合物中结合了分层路由和动态阈值,我们有效地将多个单口音LoRA专家合并在单个多模态GER中,以克服口音多样性带来的挑战。另一方面,多粒度GER利用由HDMoLE模型生成的N-最佳词级和音素级假设来预测最终的带口音语音传输。多口音英语数据集上的实验结果证明了我们提出的方法的有效性。与Whisper-large-v3基线相比,我们的方法实现了67.35%的显著相对WER降低。
摘要:Despite substantial improvements in ASR, performance tends to degrade when faced with adverse conditions such as speaker accents. Generative error correction (GER) leverages the rich linguistic knowledge and exceptional reasoning ability of LLMs, significantly outperforming typical LM methods. However, it lacks specificity in accented speech scenarios. In this study, we leverage GER to improve the accuracy of transcription predictions by addressing the two primary features of accented speech recognition. To fully leverage pronunciation information, we propose the multi-modal GER, which integrates pronunciation information from the speech modality, and the multi-granularity GER, which incorporates fine-grained phoneme-level information related to pronunciation. These two methods enable the LLM to utilize the pronunciation information of accented speech and the semantic information from word-level hypotheses for accurate transcription predictions through LoRA fine-tuning. On the one hand, we employ a three-stage training strategy to train separate multi-modal GER models for each accent to obtain mono-accent LoRA experts. By adopting our proposed HDMoLE method, which incorporates hierarchical routing and dynamic thresholds within the mixture of LoRA experts, we effectively merge multiple mono-accent LoRA experts within a single multi-modal GER to overcome the challenges posed by accent diversity. On the other hand, multi-granularity GER leverages the N-best word-level and phoneme-level hypotheses generated by the HDMoLE model to predict the final accented speech transcriptions. Experimental results on the multi-accent English dataset demonstrate the efficacy of our proposed methods. Our methods achieve a remarkable relative WER reduction of 67.35% compared to the Whisper-large-v3 baseline.


【20】Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
标题:更少的压力,更多的隐私:空中交通管制员神经化语音的压力检测
链接:https://arxiv.org/abs/2507.08882

作者:swanathan, Alexander Blatt, Konrad Hagemann, Dietrich Klakow
备注:8 pages, 2 figures, 4 tables, publication identification number   (URN)- urn:nbn:de:101:1-2022122008393409239462, see archived online   publication- https://d-nb.info/127614606X/34 & Katalogeintrag:   https://d-nb.info/127614606X/
摘要:空中交通管制(ATC)要求在时间压力下进行多任务处理,错误的后果很高。这可能会引起压力。压力检测是保持ATC高安全标准的关键。然而,处理ATC语音数据需要隐私限制,例如通用数据保护条例(GDPR)法。分析ATC语音数据是遵守这些限制的一种方法。在本文中,不同的架构,匿名ATCO语音的压力检测进行评估。我们最好的网络在模拟和实际压力下的语音(SUSAS)数据集的匿名版本上达到了93.6%的压力检测准确率,在我们的匿名ATC模拟数据集上达到了80.1%的准确率。这表明,隐私不一定是构建性能良好的基于深度学习的模型的障碍。
摘要:Air traffic control (ATC) demands multi-tasking under time pressure with high consequences of an error. This can induce stress. Detecting stress is a key point in maintaining the high safety standards of ATC. However, processing ATC voice data entails privacy restrictions, e.g. the General Data Protection Regulation (GDPR) law. Anonymizing the ATC voice data is one way to comply with these restrictions. In this paper, different architectures for stress detection for anonymized ATCO speech are evaluated. Our best networks reach a stress detection accuracy of 93.6% on an anonymized version of the Speech Under Simulated and Actual Stress (SUSAS) dataset and an accuracy of 80.1% on our anonymized ATC simulation dataset. This shows that privacy does not have to be an impediment in building well-performing deep-learning-based models.


【21】Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
标题:使用连续值令牌和掩蔽下一令牌预测的生成性音频语言建模
链接:https://arxiv.org/abs/2507.09834

作者:ang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang, Huy Phan, Bo-Ru Lu, Harsha Sundar, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao, Chao Wang
备注:Accepted by ICML 2025. Project website: this https URL
摘要:使用Transformer解码器的自回归下一个标记预测已经成为大型语言模型(LLM)的事实标准,在大规模自然语言处理(NLP)中取得了显着的成功。由于其固有的连续性,将这种范式扩展到音频带来了独特的挑战。我们研究了无离散标记的因果语言模型(LM)的音频生成。我们利用令牌扩散来模拟下一个连续值令牌的连续分布。我们的方法比以前的离散解决方案AudioGen有了显著的改进,在Frechet音频距离(FAD)和Kullback-Leibler(KL)发散度方面分别实现了AudioCaps的20%和40%的相对增益。此外,我们提出了一种新的掩蔽下一个令牌预测任务,将掩蔽预测的因果LM框架。在AudioCaps上,该创新分别比AudioGen Base(285 M)和AudioGen Large(1B)模型产生了41%和33%的相对FAD改进,与最先进的(SOTA)扩散模型相当。此外,我们使用更少的参数实现了这些结果--Base模型为193 M,Large模型为462 M。
摘要:Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio poses unique challenges due to its inherently continuous nature. We research audio generation with a causal language model (LM) without discrete tokens. We leverage token-wise diffusion to model the continuous distribution of the next continuous-valued token. Our approach delivers significant improvements over previous discrete solution, AudioGen, achieving 20% and 40% relative gains on AudioCaps in Frechet Audio Distance (FAD) and Kullback-Leibler (KL) divergence, respectively. Additionally, we propose a novel masked next-token prediction task that incorporates masked prediction into the causal LM framework. On AudioCaps, the innovation yields 41% and 33% relative FAD improvements over AudioGen Base (285M) and AudioGen Large (1B) models, respectively, and is on par with the state-of-the-art (SOTA) diffusion models. Furthermore, we achieve these results with significantly fewer parameters -- 193M for our Base and 462M for our Large models.


【22】Low-Rank Adaptation of Deep Prior Neural Networks For Room Impulse Response Reconstruction
标题:深度先验神经网络的低等级自适应用于房间脉冲响应重建
链接:https://arxiv.org/abs/2507.09806

作者:zoli, Federico Miotello, Shoichi Koyama, Fabio Antonacci
备注:to appear in IEEE WASPAA
摘要:Deep Prior框架已经成为一种强大的生成工具,可以用于从很少的稀疏压力测量重建环境中的声场。它采用了一个神经网络,该网络仅在有限的可用数据集上进行训练,并作为一个隐含的先验知识,指导底层优化问题的解决方案。然而,深度先验方法的一个重要限制是它不能推广到新的声学配置,例如声源位置的变化。因此,网络必须从头开始重新训练每个新的设置,这是计算密集型和耗时的。为了解决这个问题,我们通过低秩自适应(LoRA)研究了深度先验中的迁移学习,它通过引入可训练参数的低秩分解来实现预训练神经网络的有效微调,从而使网络能够以最小的计算开销适应新的测量集。我们将LoRA嵌入到基于MultiResUNet的Deep Prior模型中,并将其自适应性能与所有参数的完全微调以及经典再训练进行比较,特别是在仅使用有限数量麦克风的情况下。结果表明,无论是完全还是通过LoRA进行微调,当源位置是唯一的变化参数时,都特别有利,可以保持较高的物理保真度,并突出迁移学习在声学应用中的价值。
摘要:The Deep Prior framework has emerged as a powerful generative tool which can be used for reconstructing sound fields in an environment from few sparse pressure measurements. It employs a neural network that is trained solely on a limited set of available data and acts as an implicit prior which guides the solution of the underlying optimization problem. However, a significant limitation of the Deep Prior approach is its inability to generalize to new acoustic configurations, such as changes in the position of a sound source. As a consequence, the network must be retrained from scratch for every new setup, which is both computationally intensive and time-consuming. To address this, we investigate transfer learning in Deep Prior via Low-Rank Adaptation (LoRA), which enables efficient fine-tuning of a pre-trained neural network by introducing a low-rank decomposition of trainable parameters, thus allowing the network to adapt to new measurement sets with minimal computational overhead. We embed LoRA into a MultiResUNet-based Deep Prior model and compare its adaptation performance against full fine-tuning of all parameters as well as classical retraining, particularly in scenarios where only a limited number of microphones are used. The results indicate that fine-tuning, whether done completely or via LoRA, is especially advantageous when the source location is the sole changing parameter, preserving high physical fidelity, and highlighting the value of transfer learning for acoustics applications.


【23】Enhancing Stereo Sound Event Detection with BiMamba and Pretrained PSELDnet
标题:使用BiMamba和预训练的PSELDnet增强立体声事件检测
链接:https://arxiv.org/abs/2507.09570

作者:ao, Han Yin
摘要:预训练方法大大提高了声音事件定位和检测(SELD)的性能。然而,现有的基于transformer的模型仍然面临着很高的计算成本。为了解决这个问题,我们提出了一个立体SELD系统,使用预训练的PSELDnet和双向Mamba序列模型。具体来说,我们将Conformer模块替换为BiMamba模块。我们还使用非对称卷积来更好地捕捉音频信号中的时间和频率关系。在DCASE 2025任务3开发数据集上的测试结果表明,我们的方法比基线和带有Conformer解码器的原始PSELDnet性能更好。此外,拟议的模式比基准花费更少的计算资源。这些结果表明,BiMamba架构可以有效解决SELD任务中的关键挑战。源代码可在https://github.com/ alexandergwm/DCASE 2025 TASK 3 Stereo PSELD Mamba上公开访问。
摘要:Pre-training methods have greatly improved the performance of sound event localization and detection (SELD). However, existing Transformer-based models still face high computational cost. To solve this problem, we present a stereo SELD system using a pre-trained PSELDnet and a bidirectional Mamba sequence model. Specifically, we replace the Conformer module with a BiMamba module. We also use asymmetric convolutions to better capture the time and frequency relationships in the audio signal. Test results on the DCASE2025 Task 3 development dataset show that our method performs better than both the baseline and the original PSELDnet with a Conformer decoder. In addition, the proposed model costs fewer computing resources than the baselines. These results show that the BiMamba architecture is effective for solving key challenges in SELD tasks. The source code is publicly accessible at https://github.com/ alexandergwm/DCASE2025 TASK3 Stereo PSELD Mamba.


【24】The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
标题:MLC-LM挑战中用于多说话人自动语音识别的DKU系统
链接:https://arxiv.org/abs/2507.09499

作者: Ming Cheng, Ze Li, Ming Li
备注:Technical Report for MLC-SLM Challenge in Interspeech2025
摘要:我们为MLC-SLM挑战赛的任务2提供了DKU系统,该系统旨在直接从原始音频中执行多说话人自动语音识别,而无需Oracle说话人标签或时间边界。我们的方法建立在日记感知框架的基础上,将说话者嵌入和时间话语边界集成到基于Qwen2.5的大型语言模型(LLM)中。然后,我们通过微调LLM解码器中的语言特定适配器和LoRA模块来增强系统的多语言性能。最后,我们的系统在MLC-SLM数据集的开发和测试集上实现了23.56%和18.08%的tcpWER,大大超过了官方基线。
摘要:We present the DKU system for Task 2 of the MLC-SLM Challenge, which aims to perform multi-speaker automatic speech recognition directly from raw audio without Oracle speaker labels or time boundaries. Our approach builds upon a diarization-aware framework integrating speaker embeddings and temporal utterance boundaries into a Qwen2.5-based large language model (LLM). Then, we enhance the system's multilingual performance by fine-tuning language-specific adapters and LoRA modules within the LLM decoder. Finally, our system achieves the tcpWER of 23.56\% and 18.08\% on the development and test sets of the MLC-SLM dataset, substantially outperforming the official baseline.


【25】Controllable joint noise reduction and hearing loss compensation using a differentiable auditory model
标题:使用可微听觉模型实现可控联合降噪和听力损失补偿
链接:https://arxiv.org/abs/2507.09372

作者:Gonzalez, Torsten Dau, Tobias May
备注:Accepted to Clarity 2025 Workshop
摘要:基于深度学习的听力损失补偿(HLC)旨在使用神经网络提高听力受损听众的语音清晰度和质量。人道主义法律中心的一个主要挑战是缺乏一个地面实况目标。最近的工作已经使用神经网络在闭环框架中模拟不可微的听觉外围模型,但这种方法缺乏灵活性。或者,可微分的听觉模型允许直接优化,但以前的研究集中在个人的听众配置文件,或联合降噪(NR)和HLC没有平衡每个任务。这项工作制定NR和HLC作为一个多任务学习问题,训练系统,同时预测去噪和补偿信号从嘈杂的语音和听力图使用可微的听觉模型。结果表明,该系统实现了类似的目标度量性能分别为每个任务训练的系统,同时能够在推理过程中调整NR和HLC之间的平衡。
摘要:Deep learning-based hearing loss compensation (HLC) seeks to enhance speech intelligibility and quality for hearing impaired listeners using neural networks. One major challenge of HLC is the lack of a ground-truth target. Recent works have used neural networks to emulate non-differentiable auditory peripheral models in closed-loop frameworks, but this approach lacks flexibility. Alternatively, differentiable auditory models allow direct optimization, yet previous studies focused on individual listener profiles, or joint noise reduction (NR) and HLC without balancing each task. This work formulates NR and HLC as a multi-task learning problem, training a system to simultaneously predict denoised and compensated signals from noisy speech and audiograms using a differentiable auditory model. Results show the system achieves similar objective metric performance to systems trained for each task separately, while being able to adjust the balance between NR and HLC during inference.


【26】Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
标题:我们真的可以重新利用多说话人ASB数据库来进行说话人分区化吗?
链接:https://arxiv.org/abs/2507.09226

作者:iguchi, Naohiro Tawara, Takanori Ashihara, Atsushi Ando, Marc Delcroix
摘要:神经说话人日记化被广泛用于语音感知说话人日记化,但它需要大量的多说话人数据集进行训练。为了满足这种数据需求,通常通过组合多个语料库来构建大型数据集,包括最初为多说话人自动语音识别(ASR)设计的语料库。然而,ASR数据集通常具有松散定义的段边界,这些边界与日志化基准的更严格约定不一致。在这项工作中,我们表明,这种边界松散显着影响的diarization错误率,降低评估的可靠性。我们还发现,在具有不同边界精度的数据上训练的模型往往会学习特定于数据集的松散性,从而导致域外数据集的泛化能力较差。通过强制对齐使用标准化紧密边界进行训练不仅可以提高日志化性能,特别是在流媒体场景中,而且还可以在与简单的后处理相结合时提高ASR性能。
摘要:Neural speaker diarization is widely used for overlap-aware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large datasets are often constructed by combining multiple corpora, including those originally designed for multi-speaker automatic speech recognition (ASR). However, ASR datasets often feature loosely defined segment boundaries that do not align with the stricter conventions of diarization benchmarks. In this work, we show that such boundary looseness significantly impacts the diarization error rate, reducing evaluation reliability. We also reveal that models trained on data with varying boundary precision tend to learn dataset-specific looseness, leading to poor generalization across out-of-domain datasets. Training with standardized tight boundaries via forced alignment improves not only diarization performance, especially in streaming scenarios, but also ASR performance when combined with simple post-processing.


【27】Large Language Models and Non-Negative Matrix Factorization for Bioacoustic Signal Decomposition
标题:生物声学信号分解的大语言模型和非负矩阵分解
链接:https://arxiv.org/abs/2507.09161

作者:orabi, Shahram Shirani, James P. Reilly
备注:Presented at Queen's University Biological Station Seminars of Graduate Research in Ontario, Lake Shift Dissertation Camp (QUBS '25)
摘要:大型语言模型已经显示出从非结构化数据中提取意义的非凡能力,提供了超越传统数值方法的解释生物医学信号的新方法。在这项研究中,我们提出了一个矩阵分解框架的生物声学信号分析,这是加强了大型语言模型。重点是分离临床记录中通常重叠的生物声信号,使用矩阵分解将混合物分解为可解释的成分。然后,将大型语言模型应用于分离的信号,以将不同的声学模式与潜在的医学状况(例如心律紊乱或呼吸异常)相关联。从应用于临床人体模型的数字听诊器获得记录,以确保受控和高保真的采集环境。这种混合方法不需要标记的数据或源类型的先验知识,并且它为临床决策支持提供了一个更可解释和可访问的框架。该方法表明集成到未来的智能诊断工具的承诺。
摘要:Large language models have shown a remarkable ability to extract meaning from unstructured data, offering new ways to interpret biomedical signals beyond traditional numerical methods. In this study, we present a matrix factorization framework for bioacoustic signal analysis which is enhanced by large language models. The focus is on separating bioacoustic signals that commonly overlap in clinical recordings, using matrix factorization to decompose the mixture into interpretable components. A large language model is then applied to the separated signals to associate distinct acoustic patterns with potential medical conditions such as cardiac rhythm disturbances or respiratory abnormalities. Recordings were obtained from a digital stethoscope applied to a clinical manikin to ensure a controlled and high-fidelity acquisition environment. This hybrid approach does not require labeled data or prior knowledge of source types, and it provides a more interpretable and accessible framework for clinical decision support. The method demonstrates promise for integration into future intelligent diagnostic tools.


【28】SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
标题:SemAlignVC:使用语义对齐增强Zero-Shot音色转换
链接:https://arxiv.org/abs/2507.09070

作者:hta, Yingru Liu, Zhenyu Tang, Kainan Peng, Vimal Manohar, Shun Zhang, Mike Seltzer, Qing He, Mingbo Ma
备注:6 pages, 2 figures, Accepted at the ISCA Speech Synthesis Workshop (SSW) 2025
摘要:Zero-shot语音转换(VC)在保留语言和非语言内容的同时,以目标说话人的声音合成语音。然而,音色泄漏源扬声器的特质持续存在仍然是一个挑战,特别是在神经编解码器和基于LLM的VC,量化表示纠缠扬声器的身份与内容。我们介绍SemAlignVC,一个架构,旨在防止音色泄漏使用SemAlign,一种新的方法,对齐文本和音频表示,以确保扬声器独立的语义编码。这种解纠缠表示条件的自回归Transformer高保真转换没有显式的扬声器嵌入。实验表明,SemAlignVC显着降低了音色泄漏,在扬声器音色相似性,可懂度和自然度方面优于基线,使其成为一个强大的,保护隐私的,可推广的VC解决方案。音频样本可访问https://shivammehta25.github.io/SemAlignVC/
摘要:Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/



eess.AS音频处理



【1】ASDKit: A Toolkit for Comprehensive Evaluation of Anomalous Sound Detection Methods
标题:ASDKit:异常声音检测方法综合评估的工具包
链接:https://arxiv.org/abs/2507.10264

作者:jimura, Kevin Wilkinghoff, Keisuke Imoto, Tomoki Toda
摘要:在本文中,我们介绍了ASDKit,一个工具包异常声音检测(ASD)的任务。我们的目标是通过提供一个开源框架来收集和仔细评估各种ASD方法,以促进ASD研究。首先,ASDKit为广泛的ASD方法提供培训和评估脚本,所有这些都在统一的框架内处理。例如,它包括基于自动编码器的官方DCASE基线,代表性判别方法和基于自我监督学习的方法。其次,它支持对DCASE 2020- 2024数据集进行全面评估,从而能够仔细评估ASD性能,这对数据集和随机种子等因素高度敏感。在我们的实验中,我们使用ASDKit重新评估各种ASD方法,并在多个数据集和试验中识别一致有效的技术。我们还证明了ASDKit在所考虑的数据集上再现了最先进的性能。
摘要:In this paper, we introduce ASDKit, a toolkit for anomalous sound detection (ASD) task. Our aim is to facilitate ASD research by providing an open-source framework that collects and carefully evaluates various ASD methods. First, ASDKit provides training and evaluation scripts for a wide range of ASD methods, all handled within a unified framework. For instance, it includes the autoencoder-based official DCASE baseline, representative discriminative methods, and self-supervised learning-based methods. Second, it supports comprehensive evaluation on the DCASE 2020--2024 datasets, enabling careful assessment of ASD performance, which is highly sensitive to factors such as datasets and random seeds. In our experiments, we re-evaluate various ASD methods using ASDKit and identify consistently effective techniques across multiple datasets and trials. We also demonstrate that ASDKit reproduces the state-of-the-art-level performance on the considered datasets.


【2】Natural Language-based Assessment of L2 Oral Proficiency using LLMs
标题:使用LLM基于自然百分比的L2口语能力评估
链接:https://arxiv.org/abs/2507.10200

作者:annò, Rao Ma, Mengjie Qian, Siyuan Tang, Kate Knill, Mark Gales
备注:Accepted for the 10th Workshop on Speech and Language Technology in Education (SLaTE 2025)
摘要:基于自然语言的评估(NLA)是一种第二语言评估方法,它使用指令-以can-do描述符的形式表示-最初用于人类考官,旨在确定大型语言模型(LLM)是否可以以与人类评估相当的方式解释和应用它们。在这项工作中,我们探索使用这样的描述符与开源LLM,Qwen 2.5 72 B,以评估响应从公开的S&I语料库中的zero-shot设置。我们的研究结果表明,这种方法-仅依赖于文本信息-实现了有竞争力的性能:虽然它没有超过为任务微调的最先进的语音LLM,但它超过了专门为此目的训练的基于BERT的模型。NLA被证明在不匹配的任务设置中特别有效,可推广到其他数据类型和语言,并提供更大的可解释性,因为它基于清晰可解释的,广泛适用的语言描述符。
摘要:Natural language-based assessment (NLA) is an approach to second language assessment that uses instructions - expressed in the form of can-do descriptors - originally intended for human examiners, aiming to determine whether large language models (LLMs) can interpret and apply them in ways comparable to human assessment. In this work, we explore the use of such descriptors with an open-source LLM, Qwen 2.5 72B, to assess responses from the publicly available S&I Corpus in a zero-shot setting. Our results show that this approach - relying solely on textual information - achieves competitive performance: while it does not outperform state-of-the-art speech LLMs fine-tuned for the task, it surpasses a BERT-based model trained specifically for this purpose. NLA proves particularly effective in mismatched task settings, is generalisable to other data types and languages, and offers greater interpretability, as it is grounded in clearly explainable, widely applicable language descriptors.


【3】Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process
标题:救援的和声:为什么语音不是Wss过程
链接:https://arxiv.org/abs/2507.10176

作者:Bologni, Richard Heusdens, Richard C. Hendriks
备注:Comments: Accepted at the 2024 International Workshop on Acoustic   Signal Enhancement (IWAENC 2024)
摘要:语音处理算法通常依赖于底层过程的统计知识。然而,尽管经过多年的研究,关于最合适的语音统计模型的争论仍在继续。语音通常被建模为广义平稳(WSS)过程。然而,将WSS模型用于光谱相关过程是根本错误的,因为WSS意味着光谱不相关。在本文中,我们证明,浊音语音可以更准确地表示为一个循环平稳(CS)过程。通过采用CS而不是WSS模型的过程中,固有地跨频率相关的,它是可能的,以改善互功率谱密度(PSD),源分离,和波束形成的估计。我们说明了CS过程的谐波频率之间的相关性可以提高系统识别,并验证我们的研究结果使用模拟和真实的语音数据。
摘要:Speech processing algorithms often rely on statistical knowledge of the underlying process. Despite many years of research, however, the debate on the most appropriate statistical model for speech still continues. Speech is commonly modeled as a wide-sense stationary (WSS) process. However, the use of the WSS model for spectrally correlated processes is fundamentally wrong, as WSS implies spectral uncorrelation. In this paper, we demonstrate that voiced speech can be more accurately represented as a cyclostationary (CS) process. By employing the CS rather than the WSS model for processes that are inherently correlated across frequency, it is possible to improve the estimation of cross-power spectral densities (PSDs), source separation, and beamforming. We illustrate how the correlation between harmonic frequencies of CS processes can enhance system identification, and validate our findings using both simulated and real speech data.


【4】Cyclic Multichannel Wiener Filter for Acoustic Beamforming
标题:用于声束形成的循环多通道维纳过滤器
链接:https://arxiv.org/abs/2507.10159

作者:Bologni, Richard Heusdens, Richard C. Hendriks
备注:Comments: Accepted for publication at the 2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA 2025). IEEE retains copyright
摘要:声学波束形成模型通常假设语音信号在短时间帧内的广义平稳性。然而,有声语音更好地建模为循环平稳(CS)过程,这是一个随机过程,其平均值和自相关是T_1 $周期性的,其中$\alpha_1=1/T_1$对应于元音的基频。高次谐波频率是基波的整数倍。本文介绍了一种基于循环平稳模型的循环多通道维纳滤波器(cMWF)。该波束形成器利用信号的谐波频率之间的频谱相关性,以进一步降低目标和处理后的输入之间的均方误差(MSE)。所提出的cMWF在MSE意义上是最优的,并且当目标是广义静止时,其降低到MWF。模拟数据上的实验表明,合成数据上的尺度不变信号失真比(SI-SDR)有相当大的改善,但也表明对估计的基频的准确性的高敏感性,这限制了对真实数据的有效性。
摘要:Acoustic beamforming models typically assume wide-sense stationarity of speech signals within short time frames. However, voiced speech is better modeled as a cyclostationary (CS) process, a random process whose mean and autocorrelation are $T_1$-periodic, where $\alpha_1=1/T_1$ corresponds to the fundamental frequency of vowels. Higher harmonic frequencies are found at integer multiples of the fundamental. This work introduces a cyclic multichannel Wiener filter (cMWF) for speech enhancement derived from a cyclostationary model. This beamformer exploits spectral correlation across the harmonic frequencies of the signal to further reduce the mean-squared error (MSE) between the target and the processed input. The proposed cMWF is optimal in the MSE sense and reduces to the MWF when the target is wide-sense stationary. Experiments on simulated data demonstrate considerable improvements in scale-invariant signal-to-distortion ratio (SI-SDR) on synthetic data but also indicate high sensitivity to the accuracy of the estimated fundamental frequency $\alpha_1$, which limits effectiveness on real data.


【5】Aligning Generative Speech Enhancement with Human Preferences via Direct Preference Optimization
标题:通过直接偏好优化将生成语音增强与人类偏好保持一致
链接:https://arxiv.org/abs/2507.09929

作者:i, Nana Hou, Yuchen Hu, Jixun Yao, Sabato Marco Siniscalchi, Eng Siong Chng
摘要:本文从语言模型的角度研究了语音增强技术。我们提出了一种新的方法,利用直接偏好优化(DPO),以提高增强语音的感知质量。使用UTMOS,神经MOS预测模型,作为人类评级的代理,我们的方法引导优化感知首选输出。这不同于现有的基于LM的SE方法,其专注于最大化干净语音令牌的可能性,这可能与人类感知不一致并降低质量,尽管预测误差很低。在2020年深度噪声抑制挑战测试集上的实验表明,将DPO应用于预训练的基于LM的SE模型,可以在各种语音质量指标上获得一致的改善,相对增益高达56%。据我们所知,这是DPO在SE中的第一个应用,也是第一个将代理感知反馈纳入基于LM的SE训练的应用,为感知对齐SE指明了一个有希望的方向。
摘要:This work investigates speech enhancement (SE) from the perspective of language models (LMs). We propose a novel method that leverages Direct Preference Optimization (DPO) to improve the perceptual quality of enhanced speech. Using UTMOS, a neural MOS prediction model, as a proxy for human ratings, our approach guides optimization toward perceptually preferred outputs. This differs from existing LM-based SE methods that focus on maximizing the likelihood of clean speech tokens, which may misalign with human perception and degrade quality despite low prediction error. Experiments on the 2020 Deep Noise Suppression Challenge test sets demonstrate that applying DPO to a pretrained LM-based SE model yields consistent improvements across various speech quality metrics, with relative gains of up to 56%. To our knowledge, this is the first application of DPO to SE and the first to incorporate proxy perceptual feedback into LM-based SE training, pointing to a promising direction for perceptually aligned SE.


【6】Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
标题:使用连续值令牌和掩蔽下一令牌预测的生成性音频语言建模
链接:https://arxiv.org/abs/2507.09834

作者:ang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang, Huy Phan, Bo-Ru Lu, Harsha Sundar, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao, Chao Wang
备注:Accepted by ICML 2025. Project website: this https URL
摘要:使用Transformer解码器的自回归下一个标记预测已经成为大型语言模型(LLM)的事实标准,在大规模自然语言处理(NLP)中取得了显着的成功。由于其固有的连续性,将这种范式扩展到音频带来了独特的挑战。我们研究了无离散标记的因果语言模型(LM)的音频生成。我们利用令牌扩散来模拟下一个连续值令牌的连续分布。我们的方法比以前的离散解决方案AudioGen有了显著的改进,在Frechet音频距离(FAD)和Kullback-Leibler(KL)发散度方面分别实现了AudioCaps的20%和40%的相对增益。此外,我们提出了一种新的掩蔽下一个令牌预测任务,将掩蔽预测的因果LM框架。在AudioCaps上,该创新分别比AudioGen Base(285 M)和AudioGen Large(1B)模型产生了41%和33%的相对FAD改进,与最先进的(SOTA)扩散模型相当。此外,我们使用更少的参数实现了这些结果--Base模型为193 M,Large模型为462 M。
摘要:Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio poses unique challenges due to its inherently continuous nature. We research audio generation with a causal language model (LM) without discrete tokens. We leverage token-wise diffusion to model the continuous distribution of the next continuous-valued token. Our approach delivers significant improvements over previous discrete solution, AudioGen, achieving 20% and 40% relative gains on AudioCaps in Frechet Audio Distance (FAD) and Kullback-Leibler (KL) divergence, respectively. Additionally, we propose a novel masked next-token prediction task that incorporates masked prediction into the causal LM framework. On AudioCaps, the innovation yields 41% and 33% relative FAD improvements over AudioGen Base (285M) and AudioGen Large (1B) models, respectively, and is on par with the state-of-the-art (SOTA) diffusion models. Furthermore, we achieve these results with significantly fewer parameters -- 193M for our Base and 462M for our Large models.


【7】Low-Rank Adaptation of Deep Prior Neural Networks For Room Impulse Response Reconstruction
标题:深度先验神经网络的低等级自适应用于房间脉冲响应重建
链接:https://arxiv.org/abs/2507.09806

作者:zoli, Federico Miotello, Shoichi Koyama, Fabio Antonacci
备注:to appear in IEEE WASPAA
摘要:Deep Prior框架已经成为一种强大的生成工具,可以用于从很少的稀疏压力测量重建环境中的声场。它采用了一个神经网络,该网络仅在有限的可用数据集上进行训练,并作为一个隐含的先验知识,指导底层优化问题的解决方案。然而,深度先验方法的一个重要限制是它不能推广到新的声学配置,例如声源位置的变化。因此,网络必须从头开始重新训练每个新的设置,这是计算密集型和耗时的。为了解决这个问题,我们通过低秩自适应(LoRA)研究了深度先验中的迁移学习,它通过引入可训练参数的低秩分解来实现预训练神经网络的有效微调,从而使网络能够以最小的计算开销适应新的测量集。我们将LoRA嵌入到基于MultiResUNet的Deep Prior模型中,并将其自适应性能与所有参数的完全微调以及经典再训练进行比较,特别是在仅使用有限数量麦克风的情况下。结果表明,无论是完全还是通过LoRA进行微调,当源位置是唯一的变化参数时,都特别有利,可以保持较高的物理保真度,并突出迁移学习在声学应用中的价值。
摘要:The Deep Prior framework has emerged as a powerful generative tool which can be used for reconstructing sound fields in an environment from few sparse pressure measurements. It employs a neural network that is trained solely on a limited set of available data and acts as an implicit prior which guides the solution of the underlying optimization problem. However, a significant limitation of the Deep Prior approach is its inability to generalize to new acoustic configurations, such as changes in the position of a sound source. As a consequence, the network must be retrained from scratch for every new setup, which is both computationally intensive and time-consuming. To address this, we investigate transfer learning in Deep Prior via Low-Rank Adaptation (LoRA), which enables efficient fine-tuning of a pre-trained neural network by introducing a low-rank decomposition of trainable parameters, thus allowing the network to adapt to new measurement sets with minimal computational overhead. We embed LoRA into a MultiResUNet-based Deep Prior model and compare its adaptation performance against full fine-tuning of all parameters as well as classical retraining, particularly in scenarios where only a limited number of microphones are used. The results indicate that fine-tuning, whether done completely or via LoRA, is especially advantageous when the source location is the sole changing parameter, preserving high physical fidelity, and highlighting the value of transfer learning for acoustics applications.


【8】Enhancing Stereo Sound Event Detection with BiMamba and Pretrained PSELDnet
标题:使用BiMamba和预训练的PSELDnet增强立体声事件检测
链接:https://arxiv.org/abs/2507.09570

作者:ao, Han Yin
摘要:预训练方法大大提高了声音事件定位和检测(SELD)的性能。然而,现有的基于transformer的模型仍然面临着很高的计算成本。为了解决这个问题,我们提出了一个立体SELD系统,使用预训练的PSELDnet和双向Mamba序列模型。具体来说,我们将Conformer模块替换为BiMamba模块。我们还使用非对称卷积来更好地捕捉音频信号中的时间和频率关系。在DCASE 2025任务3开发数据集上的测试结果表明,我们的方法比基线和带有Conformer解码器的原始PSELDnet性能更好。此外,拟议的模式比基准花费更少的计算资源。这些结果表明,BiMamba架构可以有效解决SELD任务中的关键挑战。源代码可在https://github.com/ alexandergwm/DCASE 2025 TASK 3 Stereo PSELD Mamba上公开访问。
摘要:Pre-training methods have greatly improved the performance of sound event localization and detection (SELD). However, existing Transformer-based models still face high computational cost. To solve this problem, we present a stereo SELD system using a pre-trained PSELDnet and a bidirectional Mamba sequence model. Specifically, we replace the Conformer module with a BiMamba module. We also use asymmetric convolutions to better capture the time and frequency relationships in the audio signal. Test results on the DCASE2025 Task 3 development dataset show that our method performs better than both the baseline and the original PSELDnet with a Conformer decoder. In addition, the proposed model costs fewer computing resources than the baselines. These results show that the BiMamba architecture is effective for solving key challenges in SELD tasks. The source code is publicly accessible at https://github.com/ alexandergwm/DCASE2025 TASK3 Stereo PSELD Mamba.


【9】The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
标题:MLC-LM挑战中用于多说话人自动语音识别的DKU系统
链接:https://arxiv.org/abs/2507.09499

作者: Ming Cheng, Ze Li, Ming Li
备注:Technical Report for MLC-SLM Challenge in Interspeech2025
摘要:我们为MLC-SLM挑战赛的任务2提供了DKU系统,该系统旨在直接从原始音频中执行多说话人自动语音识别,而无需Oracle说话人标签或时间边界。我们的方法建立在日记感知框架的基础上,将说话者嵌入和时间话语边界集成到基于Qwen2.5的大型语言模型(LLM)中。然后,我们通过微调LLM解码器中的语言特定适配器和LoRA模块来增强系统的多语言性能。最后,我们的系统在MLC-SLM数据集的开发和测试集上实现了23.56%和18.08%的tcpWER,大大超过了官方基线。
摘要:We present the DKU system for Task 2 of the MLC-SLM Challenge, which aims to perform multi-speaker automatic speech recognition directly from raw audio without Oracle speaker labels or time boundaries. Our approach builds upon a diarization-aware framework integrating speaker embeddings and temporal utterance boundaries into a Qwen2.5-based large language model (LLM). Then, we enhance the system's multilingual performance by fine-tuning language-specific adapters and LoRA modules within the LLM decoder. Finally, our system achieves the tcpWER of 23.56\% and 18.08\% on the development and test sets of the MLC-SLM dataset, substantially outperforming the official baseline.


【10】Controllable joint noise reduction and hearing loss compensation using a differentiable auditory model
标题:使用可微听觉模型实现可控联合降噪和听力损失补偿
链接:https://arxiv.org/abs/2507.09372

作者:Gonzalez, Torsten Dau, Tobias May
备注:Accepted to Clarity 2025 Workshop
摘要:基于深度学习的听力损失补偿(HLC)旨在使用神经网络提高听力受损听众的语音清晰度和质量。人道主义法律中心的一个主要挑战是缺乏一个地面实况目标。最近的工作已经使用神经网络在闭环框架中模拟不可微的听觉外围模型,但这种方法缺乏灵活性。或者,可微分的听觉模型允许直接优化,但以前的研究集中在个人的听众配置文件,或联合降噪(NR)和HLC没有平衡每个任务。这项工作制定NR和HLC作为一个多任务学习问题,训练系统,同时预测去噪和补偿信号从嘈杂的语音和听力图使用可微的听觉模型。结果表明,该系统实现了类似的目标度量性能分别为每个任务训练的系统,同时能够在推理过程中调整NR和HLC之间的平衡。
摘要:Deep learning-based hearing loss compensation (HLC) seeks to enhance speech intelligibility and quality for hearing impaired listeners using neural networks. One major challenge of HLC is the lack of a ground-truth target. Recent works have used neural networks to emulate non-differentiable auditory peripheral models in closed-loop frameworks, but this approach lacks flexibility. Alternatively, differentiable auditory models allow direct optimization, yet previous studies focused on individual listener profiles, or joint noise reduction (NR) and HLC without balancing each task. This work formulates NR and HLC as a multi-task learning problem, training a system to simultaneously predict denoised and compensated signals from noisy speech and audiograms using a differentiable auditory model. Results show the system achieves similar objective metric performance to systems trained for each task separately, while being able to adjust the balance between NR and HLC during inference.


【11】Microphone Occlusion Mitigation for Own-Voice Enhancement in Head-Worn Microphone Arrays Using Switching-Adaptive Beamforming
标题:使用开关自适应射束形成缓解耳机遮挡以增强头戴麦克风阵列中的自有语音
链接:https://arxiv.org/abs/2507.09350

作者:ddelberg, Jung-Suk Lee, Saeed Bagheri Sereshki, Ali Aroudi, Vladimir Tourbabin, Daniel D. E. Wong
备注:Accepted for publication at WASPAA 2025
摘要:增强头戴式麦克风阵列的用户自己的语音是噪声环境中的重要任务,以允许更容易的语音通信和用户设备交互。然而,很少解决的挑战是当麦克风中的一个或多个被皮肤、衣服或头发遮挡时麦克风的传递函数的改变。基于波束形成的语音增强的根本问题是必须考虑到自身语音和噪声分量的(潜在快速)变化的传递函数以实现最佳性能。在本文中,我们解决了头戴式麦克风阵列中的麦克风被遮挡的问题。我们调查三种替代缓解方法,通过(i)传统的自适应波束形成,(ii)之间的先验估计的波束形成器系数的阻塞和未阻塞状态的切换,和(iii)一个混合的方法,使用切换自适应波束形成器。在与真实世界的录音和模拟闭塞的评估,我们证明了不同的方法在降噪,自己的声音失真和鲁棒性对语音活动检测错误方面的优势。
摘要:Enhancing the user's own-voice for head-worn microphone arrays is an important task in noisy environments to allow for easier speech communication and user-device interaction. However, a rarely addressed challenge is the change of the microphones' transfer functions when one or more of the microphones gets occluded by skin, clothes or hair. The underlying problem for beamforming-based speech enhancement is the (potentially rapidly) changing transfer functions of both the own-voice and the noise component that have to be accounted for to achieve optimal performance. In this paper, we address the problem of an occluded microphone in a head-worn microphone array. We investigate three alternative mitigation approaches by means of (i) conventional adaptive beamforming, (ii) switching between a-priori estimates of the beamformer coefficients for the occluded and unoccluded state, and (iii) a hybrid approach using a switching-adaptive beamformer. In an evaluation with real-world recordings and simulated occlusion, we demonstrate the advantages of the different approaches in terms of noise reduction, own-voice distortion and robustness against voice activity detection errors.


【12】ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
标题:ZipVoice-Dialogue:通过流量匹配生成非自回归口语对话
链接:https://arxiv.org/abs/2507.09318

作者:Wei Kang, Liyong Guo, Zengwei Yao, Fangjun Kuang, Weiji Zhuang, Zhaoqing Li, Zhifeng Han, Dong Zhang, Xin Zhang, Xingchen Song, Long Lin, Daniel Povey
摘要:生成口语对话比独白文本到语音(TTS)更具挑战性,因为需要逼真的话轮转换和不同的扬声器音色。现有的口语对话生成模型,是自回归的,遭受缓慢和不稳定的推理。为了克服这些限制,我们引入ZipVoice-Dialog,一个非自回归的zero-shot口语对话生成模型建立在流匹配。主要设计包括:1)用于精确的说话者话轮转换的说话者话轮嵌入; 2)用于稳定的语音-文本对齐的课程学习策略; 3)用于实现立体声对话生成的专门策略。此外,认识到缺乏开源的大规模口语对话数据集,我们策划了OpenDialog,这是一个来自野外语音数据的6. 8 k小时口语对话数据集。此外,我们建立了一个基准,以全面评估各种模型。实验结果表明,ZipVoice-Dialog在可懂度、话轮转换准确率、说话人相似度和推理速度等方面都有较好的表现。我们的代码、模型检查点、演示示例和OpenDialog数据集都可以在https://github.com/k2-fsa/ZipVoice上公开获取。
摘要:Generating spoken dialogue is more challenging than monologue text-to-speech (TTS) due to the need for realistic turn-taking and distinct speaker timbres. Existing spoken dialogue generation models, being auto-regressive, suffer from slow and unstable inference. To overcome these limitations, we introduce ZipVoice-Dialog, a non-autoregressive zero-shot spoken dialogue generation model built upon flow matching. Key designs include: 1) speaker-turn embeddings for precise speaker turn-taking; 2) a curriculum learning strategy for stable speech-text alignment; 3) specialized strategies to enable stereo dialogue generation. Additionally, recognizing the lack of open-source large-scale spoken dialogue datasets, we curated OpenDialog, a 6.8k-hour spoken dialogue dataset from in-the-wild speech data. Furthermore, we established a benchmark to comprehensively evaluate various models. Experimental results demonstrate that ZipVoice-Dialog achieves superior performance in intelligibility, speaker turn-taking accuracy, speaker similarity, and inference speed. Our codes, model checkpoints, demo samples, and the OpenDialog dataset are all publicly available at https://github.com/k2-fsa/ZipVoice.


【13】Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
标题:我们真的可以重新利用多说话人ASB数据库来进行说话人分区化吗?
链接:https://arxiv.org/abs/2507.09226

作者:iguchi, Naohiro Tawara, Takanori Ashihara, Atsushi Ando, Marc Delcroix
摘要:神经说话人日记化被广泛用于语音感知说话人日记化,但它需要大量的多说话人数据集进行训练。为了满足这种数据需求,通常通过组合多个语料库来构建大型数据集,包括最初为多说话人自动语音识别(ASR)设计的语料库。然而,ASR数据集通常具有松散定义的段边界,这些边界与日志化基准的更严格约定不一致。在这项工作中,我们表明,这种边界松散显着影响的diarization错误率,降低评估的可靠性。我们还发现,在具有不同边界精度的数据上训练的模型往往会学习特定于数据集的松散性,从而导致域外数据集的泛化能力较差。通过强制对齐使用标准化紧密边界进行训练不仅可以提高日志化性能,特别是在流媒体场景中,而且还可以在与简单的后处理相结合时提高ASR性能。
摘要:Neural speaker diarization is widely used for overlap-aware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large datasets are often constructed by combining multiple corpora, including those originally designed for multi-speaker automatic speech recognition (ASR). However, ASR datasets often feature loosely defined segment boundaries that do not align with the stricter conventions of diarization benchmarks. In this work, we show that such boundary looseness significantly impacts the diarization error rate, reducing evaluation reliability. We also reveal that models trained on data with varying boundary precision tend to learn dataset-specific looseness, leading to poor generalization across out-of-domain datasets. Training with standardized tight boundaries via forced alignment improves not only diarization performance, especially in streaming scenarios, but also ASR performance when combined with simple post-processing.


【14】Large Language Models and Non-Negative Matrix Factorization for Bioacoustic Signal Decomposition
标题:生物声学信号分解的大语言模型和非负矩阵分解
链接:https://arxiv.org/abs/2507.09161

作者:orabi, Shahram Shirani, James P. Reilly
备注:Presented at Queen's University Biological Station Seminars of Graduate Research in Ontario, Lake Shift Dissertation Camp (QUBS '25)
摘要:大型语言模型已经显示出从非结构化数据中提取意义的非凡能力,提供了超越传统数值方法的解释生物医学信号的新方法。在这项研究中,我们提出了一个矩阵分解框架的生物声学信号分析,这是加强了大型语言模型。重点是分离临床记录中通常重叠的生物声信号,使用矩阵分解将混合物分解为可解释的成分。然后,将大型语言模型应用于分离的信号,以将不同的声学模式与潜在的医学状况(例如心律紊乱或呼吸异常)相关联。从应用于临床人体模型的数字听诊器获得记录,以确保受控和高保真的采集环境。这种混合方法不需要标记的数据或源类型的先验知识,并且它为临床决策支持提供了一个更可解释和可访问的框架。该方法表明集成到未来的智能诊断工具的承诺。
摘要:Large language models have shown a remarkable ability to extract meaning from unstructured data, offering new ways to interpret biomedical signals beyond traditional numerical methods. In this study, we present a matrix factorization framework for bioacoustic signal analysis which is enhanced by large language models. The focus is on separating bioacoustic signals that commonly overlap in clinical recordings, using matrix factorization to decompose the mixture into interpretable components. A large language model is then applied to the separated signals to associate distinct acoustic patterns with potential medical conditions such as cardiac rhythm disturbances or respiratory abnormalities. Recordings were obtained from a digital stethoscope applied to a clinical manikin to ensure a controlled and high-fidelity acquisition environment. This hybrid approach does not require labeled data or prior knowledge of source types, and it provides a more interpretable and accessible framework for clinical decision support. The method demonstrates promise for integration into future intelligent diagnostic tools.


【15】SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
标题:SemAlignVC:使用语义对齐增强Zero-Shot音色转换
链接:https://arxiv.org/abs/2507.09070

作者:hta, Yingru Liu, Zhenyu Tang, Kainan Peng, Vimal Manohar, Shun Zhang, Mike Seltzer, Qing He, Mingbo Ma
备注:6 pages, 2 figures, Accepted at the ISCA Speech Synthesis Workshop (SSW) 2025
摘要:Zero-shot语音转换(VC)在保留语言和非语言内容的同时,以目标说话人的声音合成语音。然而,音色泄漏源扬声器的特质持续存在仍然是一个挑战,特别是在神经编解码器和基于LLM的VC,量化表示纠缠扬声器的身份与内容。我们介绍SemAlignVC,一个架构,旨在防止音色泄漏使用SemAlign,一种新的方法,对齐文本和音频表示,以确保扬声器独立的语义编码。这种解纠缠表示条件的自回归Transformer高保真转换没有显式的扬声器嵌入。实验表明,SemAlignVC显着降低了音色泄漏,在扬声器音色相似性,可懂度和自然度方面优于基线,使其成为一个强大的,保护隐私的,可推广的VC解决方案。音频样本可访问https://shivammehta25.github.io/SemAlignVC/
摘要:Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/


【16】AudioMAE++: learning better masked audio representations with SwiGLU FFNs
标题:AudioMAE++:使用SwiGLU FFN学习更好的掩蔽音频表示
链接:https://arxiv.org/abs/2507.10464

作者:adav, Sergios Theodoridis, Zheng-Hua Tan
备注:TO APPEAR AT IEEE MLSP 2025
摘要:在音频频谱图补丁上训练的掩蔽自动编码器(MAE)已经成为学习自监督音频表示的一种重要方法。虽然最近的几篇论文已经评估了在音频数据上训练MAE的关键方面,但这些方法中的大多数仍然利用普通的Transformer构建块,而Transformer社区已经看到了更新的架构进步的稳定集成。在这项工作中,我们提出了AudioMAE++,一个改进的音频掩蔽自动编码器与两个这样的增强,即马卡龙风格的Transformer块与门控线性单元。当在AudioSet数据集上进行预训练时,所提出的AudioMAE++模型在10个不同的下游任务上优于现有的基于MAE的方法,在音频分类和基于语音的基准测试上表现出出色的性能。所提出的AudioMAE++模型还表现出出色的缩放特性,性能优于直接可比的标准MAE基线,参数多达4倍。
摘要:Masked Autoencoders (MAEs) trained on audio spectrogram patches have emerged as a prominent approach for learning self-supervised audio representations. While several recent papers have evaluated key aspects of training MAEs on audio data, the majority of these approaches still leverage vanilla transformer building blocks, whereas the transformer community has seen steady integration of newer architectural advancements. In this work, we propose AudioMAE++, a revamped audio masked autoencoder with two such enhancements, namely macaron-style transformer blocks with gated linear units. When pretrained on the AudioSet dataset, the proposed AudioMAE++ models outperform existing MAE based approaches on 10 diverse downstream tasks, demonstrating excellent performance on audio classification and speech-based benchmarks. The proposed AudioMAE++ models also demonstrate excellent scaling characteristics, outperforming directly comparable standard MAE baselines with up to 4x more parameters.


【17】Radif corpus: a symbolic dataset for non-metric iranian classical music
标题:Radif文集:非度量伊朗古典音乐的符号数据集
链接:https://arxiv.org/abs/2507.10456

作者:nani, Sean O Leary, James McDermott
摘要:非格律音乐是伊朗古典音乐的核心。Dastgahi音乐是伊朗艺术音乐和某些民间传统的基础理论体系。在伊朗古典音乐的核心在于radif,一个基本的曲目,组织旋律材料的核心性能和教学。   在这项研究中,我们介绍了第一个数字语料库代表完整的非韵律radif剧目,涵盖了所有13个现有的组成部分,这一剧目。我们为228首音乐提供了描述音符、音符持续时间、音程和层次结构的文件(总共约281分钟)和数据电子表格。我们忠实地代表音调,包括四分之一色调,和非公制方面。此外,我们还提供了支持的基本统计数据,以及对语料库的复杂性和相似性的度量。   我们的语料库为伊朗古典音乐的计算研究提供了一个平台。研究人员可以将其用于研究旋律模式,调查即兴风格,或用于音乐信息检索,音乐理论和计算(民族)音乐学的其他任务。
摘要:Non-metric music forms the core of the repertoire in Iranian classical music. Dastgahi music serves as the underlying theoretical system for both Iranian art music and certain folk traditions. At the heart of Iranian classical music lies the radif, a foundational repertoire that organizes melodic material central to performance and pedagogy.   In this study, we introduce the first digital corpus representing the complete non-metrical radif repertoire, covering all 13 existing components of this repertoire. We provide MIDI files (about 281 minutes in total) and data spreadsheets describing notes, note durations, intervals, and hierarchical structures for 228 pieces of music. We faithfully represent the tonality including quarter-tones, and the non-metric aspect. Furthermore, we provide supporting basic statistics, and measures of complexity and similarity over the corpus.   Our corpus provides a platform for computational studies of Iranian classical music. Researchers might employ it in studying melodic patterns, investigating improvisational styles, or for other tasks in music information retrieval, music theory, and computational (ethno)musicology.


【18】Evaluating Fake Music Detection Performance Under Audio Augmentations
标题:评估音频增强下的假音乐检测性能
链接:https://arxiv.org/abs/2507.10447

作者:oka, Tomasz Wężowicz, Dominik Sidorczuk, Mateusz Modrzejewski
备注:ISMIR 2025 LBD, 2 pages + bibliography, 1 figure
摘要:随着生成音频模型的快速发展,区分人类创作和生成的音乐变得越来越具有挑战性。作为响应,已经提出了用于检测假音乐的模型。在这项工作中,我们探讨了这种系统的鲁棒性下的音频增强。为了评估模型的泛化能力,我们构建了一个由使用多个系统生成的真实音乐和合成音乐组成的数据集。然后,我们应用一系列音频转换,并分析它们如何影响分类准确性。我们测试了最近最先进的音乐deepfake检测模型在存在音频增强的情况下的性能。该模型的性能显着下降,即使与光增强的引入。
摘要:With the rapid advancement of generative audio models, distinguishing between human-composed and generated music is becoming increasingly challenging. As a response, models for detecting fake music have been proposed. In this work, we explore the robustness of such systems under audio augmentations. To evaluate model generalization, we constructed a dataset consisting of both real and synthetic music generated using several systems. We then apply a range of audio transformations and analyze how they affect classification accuracy. We test the performance of a recent state-of-the-art musical deepfake detection model in the presence of audio augmentations. The performance of the model decreases significantly even with the introduction of light augmentations.


【19】DQLoRA: A Lightweight Domain-Aware Denoising ASR via Adapter-guided Distillation
标题:DQLoRA:一种轻量级领域感知,通过适配器引导蒸馏降噪ASB
链接:https://arxiv.org/abs/2507.10313

作者:
摘要:我们展示了一个DQLoRA的演示,这是一个适配器引导的蒸馏框架,用于在低资源和噪声条件下进行鲁棒的语音识别。我们的方法采用冻结的Whisper模型作为教师来提供语义监督,以及配备基于QLoRA的适配器的轻量级Wav2Vec2学生。训练是在使用DNS风格噪声增强的FLEURS数据集上进行的。通过联合最小化CTC损失和基于KL的蒸馏损失来优化学生,从而在保持识别准确性的同时实现有效的自适应。
摘要:We present a demo of DQLoRA, an Adapter-Guided Distillation framework for robust speech recognition under low-resource and noisy conditions. Our method employs a frozen Whisper model as the teacher to provide semantic supervision, and a lightweight Wav2Vec2 student equipped with QLoRA-based Adapters. Training is conducted on the FLEURS dataset augmented with DNS-style noise. The student is optimized by jointly minimizing CTC loss and KL-based distillation loss, enabling efficient adaptation while preserving recognition accuracy.


【20】DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
标题:DualDub:通过语音和背景音频联合合成的视频到音轨生成
链接:https://arxiv.org/abs/2507.10109

作者:an, Xinfa Zhu, Haohe Liu, Zhixian Zhao, Zihao Chen, Chaofan Ding, Xinhan Di, Junjie Zheng, Lei Xie
摘要:虽然最近的视频到音频(V2A)模型可以从视觉输入生成逼真的背景音频,但它们在很大程度上忽略了语音,这是许多视频配乐的重要组成部分。本文提出了一个新的任务,视频到配乐(V2ST)的生成,其目的是在一个统一的框架内联合产生同步的背景音频和语音。为了解决V2ST问题,我们引入了DualDub,这是一个基于多模态语言模型的统一框架,它集成了多模态编码器、跨模态对齐器和双解码头,用于同时生成背景音频和语音。具体来说,我们提出的跨模态对准器采用因果和非因果的注意机制,以提高同步和声学和谐。此外,为因应资料缺乏的问题,本研究设计了一套课程学习策略,逐步建立多模态能力。最后,我们介绍了DualBench,这是V2ST评估的第一个基准,具有精心策划的测试集和全面的指标。实验结果表明,DualDub实现了最先进的性能,生成高质量和良好同步的音轨与语音和背景音频。
摘要:While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST) generation, which aims to jointly produce synchronized background audio and speech within a unified framework. To tackle V2ST, we introduce DualDub, a unified framework built on a multimodal language model that integrates a multimodal encoder, a cross-modal aligner, and dual decoding heads for simultaneous background audio and speech generation. Specifically, our proposed cross-modal aligner employs causal and non-causal attention mechanisms to improve synchronization and acoustic harmony. Besides, to handle data scarcity, we design a curriculum learning strategy that progressively builds the multimodal capability. Finally, we introduce DualBench, the first benchmark for V2ST evaluation with a carefully curated test set and comprehensive metrics. Experimental results demonstrate that DualDub achieves state-of-the-art performance, generating high-quality and well-synchronized soundtracks with both speech and background audio.


【21】The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents
标题:声音背后的人:通过多模式大型语言模型代理揭开音频私人属性分析的神秘面纱
链接:https://arxiv.org/abs/2507.10016

作者:, Kaixiang Yao, Xinfeng Li, Dong Yang, Haoyang Li, Xiaofeng Wang, Wei Dong
备注:22 pages, 4 figures
摘要:我们的研究揭示了与多模态大型语言模型(MLLM)相关的一种新的隐私风险:从音频数据中推断敏感个人属性的能力-我们称之为音频私有属性分析的技术。这种能力构成了重大威胁,因为音频可以在没有直接交互或可见性的情况下被秘密捕获。此外,与图像和文本相比,音频具有独特的特征,例如音调和音高,可以用于更详细的分析。然而,在理解MLLM采用的来自音频的私有属性剖析中存在两个关键挑战:(1)缺乏具有敏感属性注释的音频基准数据集,以及(2)当前MLLM直接从音频推断这些属性的能力有限。为了解决这些挑战,我们引入了AP^2,这是一个音频基准数据集,由从真实世界数据中收集和组成的两个子集组成,并且都用敏感属性标签进行了注释。此外,我们还提出了Gifts,这是一个混合多代理框架,它利用了音频语言模型(ALMs)和大型语言模型(LLM)的互补优势来增强推理能力。Gifts采用LLM来指导ALM推断敏感属性,然后从法医学角度分析和巩固ALM的推断,克服现有ALM在生成长上下文响应时的严重幻觉。我们的评估表明,礼物显着优于基线方法在推断敏感属性。最后,我们研究了模型级和数据级防御策略,以减轻音频私有属性分析的风险。我们的工作验证了使用MLLM进行基于音频的隐私攻击的可行性,强调了强大防御的必要性,并提供了一个数据集和框架,以促进未来的研究。
摘要:Our research uncovers a novel privacy risk associated with multimodal large language models (MLLMs): the ability to infer sensitive personal attributes from audio data -- a technique we term audio private attribute profiling. This capability poses a significant threat, as audio can be covertly captured without direct interaction or visibility. Moreover, compared to images and text, audio carries unique characteristics, such as tone and pitch, which can be exploited for more detailed profiling. However, two key challenges exist in understanding MLLM-employed private attribute profiling from audio: (1) the lack of audio benchmark datasets with sensitive attribute annotations and (2) the limited ability of current MLLMs to infer such attributes directly from audio. To address these challenges, we introduce AP^2, an audio benchmark dataset that consists of two subsets collected and composed from real-world data, and both are annotated with sensitive attribute labels. Additionally, we propose Gifts, a hybrid multi-agent framework that leverages the complementary strengths of audio-language models (ALMs) and large language models (LLMs) to enhance inference capabilities. Gifts employs an LLM to guide the ALM in inferring sensitive attributes, then forensically analyzes and consolidates the ALM's inferences, overcoming severe hallucinations of existing ALMs in generating long-context responses. Our evaluations demonstrate that Gifts significantly outperforms baseline approaches in inferring sensitive attributes. Finally, we investigate model-level and data-level defense strategies to mitigate the risks of audio private attribute profiling. Our work validates the feasibility of audio-based privacy attacks using MLLMs, highlighting the need for robust defenses, and provides a dataset and framework to facilitate future research.


【22】ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
标题:ASTAR-NTU解决方案AudioMOS挑战赛2025 Track 1
链接:https://arxiv.org/abs/2507.09904

作者:tter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei, Jeremy H.M. Wong, Nancy F. Chen, Hung-yi Lee
备注:Under Review - Submitted to AudioMOS Challenge 2025 - ASRU 2025
摘要:文本到音乐系统的评估受到成本和收集专家进行评估的可用性的限制。AudioMOS 2025 Challenge音轨1被创建为自动预测音乐印象(MI)以及提示和生成的音乐作品之间的文本对齐(TA)。本文报告了我们的获奖系统,该系统使用双分支架构,并将预训练的MuQ和RoBERTa模型作为音频和文本编码器。交叉注意机制融合了音频和文本表示。对于训练,我们将MI和TA预测重新构建为分类任务。为了结合MOS分数的顺序性质,使用高斯内核将独热标签转换为软分布。在官方测试集上,用该方法训练的单个模型实现了MI的系统级斯皮尔曼等级相关系数(SRCC)为0.991,TA的SRCC为0.952,对应于MI SRCC中的21.21%和TA SRCC中的31.47%的挑战基线的相对改善。
摘要:Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically predict music impression (MI) as well as text alignment (TA) between the prompt and the generated musical piece. This paper reports our winning system, which uses a dual-branch architecture with pre-trained MuQ and RoBERTa models as audio and text encoders. A cross-attention mechanism fuses the audio and text representations. For training, we reframe the MI and TA prediction as a classification task. To incorporate the ordinal nature of MOS scores, one-hot labels are converted to a soft distribution using a Gaussian kernel. On the official test set, a single model trained with this method achieves a system-level Spearman's Rank Correlation Coefficient (SRCC) of 0.991 for MI and 0.952 for TA, corresponding to a relative improvement of 21.21\% in MI SRCC and 31.47\% in TA SRCC over the challenge baseline.


【23】SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
标题:SpeakerVid-5 M:用于视听二元交互式人类生成的大规模高质量数据集
链接:https://arxiv.org/abs/2507.09862

作者:Zhang, Zhaoyang Li, Duomin Wang, Jiahe Zhang, Deyu Zhou, Zixin Yin, Xili Dai, Gang Yu, Xiu Li
摘要:大规模模型的快速发展促进了数字人类领域的重大突破。这些先进的方法提供了高保真的解决方案,化身驱动和渲染,导致学术界关注的下一个主要挑战:视听二元交互式虚拟人。为了促进这一新兴领域的研究,我们提出了SpeakerVid-5 M数据集,这是第一个为视听二元交互式虚拟人生成而设计的大规模,高质量的数据集。SpeakerVid-5 M总共超过8,743小时,包含超过520万个人体肖像视频剪辑。它涵盖了不同的规模和互动类型,包括一元谈话,倾听和二元对话。至关重要的是,数据集是沿着两个关键维度构建的:交互类型和数据质量。首先,它被分为四种类型(对话分支,单分支,倾听分支和多轮分支)的基础上的互动场景。其次,它被分层为一个大规模的预训练子集和一个精心策划的高质量子集,用于监督微调(SFT)。这种双重结构适应了各种各样的2D虚拟人任务。此外,我们还提供了一个基于自回归(AR)的视频聊天基线,并在此数据上进行了训练,同时还提供了一组专用的指标和测试数据,作为未来工作的基准VidChatBench。数据集和相应的数据处理代码都将公开发布。项目页面:https://dorniwang.github.io/SpeakerVid-5M/
摘要:The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the next major challenge: audio-visual dyadic interactive virtual human. To facilitate research in this emerging area, we present SpeakerVid-5M dataset, the first large-scale, high-quality dataset designed for audio-visual dyadic interactive virtual human generation. Totaling over 8,743 hours, SpeakerVid-5M contains more than 5.2 million video clips of human portraits. It covers diverse scales and interaction types, including monadic talking, listening, and dyadic conversations. Crucially, the dataset is structured along two key dimensions: interaction type and data quality. First, it is categorized into four types (dialogue branch, single branch, listening branch and multi-turn branch) based on the interaction scenario. Second, it is stratified into a large-scale pre-training subset and a curated, high-quality subset for Supervised Fine-Tuning (SFT). This dual structure accommodates a wide array of 2D virtual human tasks. In addition, we provide an autoregressive (AR)-based video chat baseline trained on this data, accompanied by a dedicated set of metrics and test data to serve as a benchmark VidChatBench for future work. Both the dataset and the corresponding data processing code will be publicly released. Project page: https://dorniwang.github.io/SpeakerVid-5M/


【24】Knowing When to Quit: Probabilistic Early Exits for Speech Separation
标题:知道何时退出:可能因言语分离而提前退出
链接:https://arxiv.org/abs/2507.09768

作者:kær Olsen. Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen, Rasmus Malik Høegh Lindrup, Bjørn Sand Jensen, Morten Mørup
摘要:近年来,基于深度学习的单通道语音分离有了很大的改进,这在很大程度上是由越来越多的计算和参数高效的神经网络架构驱动的。然而,大多数这样的架构被设计为具有固定的计算和参数预算,并且因此不能扩展到变化的计算需求或资源,这限制了它们在嵌入式和异构设备(诸如移动电话和耳机)中的使用。为了实现这样的用例,我们设计了一个能够提前退出的语音分离神经网络架构,我们提出了一个不确定性感知的概率框架来联合建模干净的语音信号和误差方差,我们使用它来推导概率提前退出条件所需的信噪比。我们在语音分离和增强任务上评估了我们的方法,并且我们表明,单个早期退出模型可以与在许多计算和参数预算下训练的最先进的模型竞争。我们的框架能够实现语音分离网络的细粒度动态计算缩放,同时实现最先进的性能和可解释的退出条件。
摘要:In recent years, deep learning-based single-channel speech separation has improved considerably, in large part driven by increasingly compute- and parameter-efficient neural network architectures. Most such architectures are, however, designed with a fixed compute and parameter budget, and consequently cannot scale to varying compute demands or resources, which limits their use in embedded and heterogeneous devices such as mobile phones and hearables. To enable such use-cases we design a neural network architecture for speech separation capable of early-exit, and we propose an uncertainty-aware probabilistic framework to jointly model the clean speech signal and error variance which we use to derive probabilistic early-exit conditions in terms of desired signal-to-noise ratios. We evaluate our methods on both speech separation and enhancement tasks, and we show that a single early-exit model can be competitive with state-of-the-art models trained at many compute and parameter budgets. Our framework enables fine-grained dynamic compute-scaling of speech separation networks while achieving state-of-the-art performance and interpretable exit conditions.


【25】MB-RIRs: a Synthetic Room Impulse Response Dataset with Frequency-Dependent Absorption Coefficients
标题:MB-RIR:具有频率相关吸收系数的合成房间脉冲响应数据集
链接:https://arxiv.org/abs/2507.09750

作者:ó, Joanna Luberadzka, Umut Sayin, Xavier Serra
备注:Accepted to WASPAA25
摘要:我们研究了四种策略对改善单声道语音增强(SE)的合成房间脉冲响应(RIR)数据集的生态有效性的影响。在传统的基于图像源方法(ISM)鞋盒RIR的基础上,我们实现了三个特征:多波段吸收系数、源方向性和接收器方向性。我们还考虑了SoundSpaces数据集中基于网格的RIR。然后,我们为每个RIR数据集训练DeepFilternet 3模型,并客观和主观地评估真实RIR测试集的性能。我们发现使用频率相关吸声系数(MB-RIR)的RIR在真实RIR上评估时可以获得+0.51dB的SDR和+8.9的MUSHRA评分。MB-RIR数据集公开提供免费下载。
摘要:We investigate the effects of four strategies for improving the ecological validity of synthetic room impulse response (RIR) datasets for monoaural Speech Enhancement (SE). We implement three features on top of the traditional image source method-based (ISM) shoebox RIRs: multiband absorption coefficients, source directivity and receiver directivity. We additionally consider mesh-based RIRs from the SoundSpaces dataset. We then train a DeepFilternet3 model for each RIR dataset and evaluate the performance on a test set of real RIRs both objectively and subjectively. We find that RIRs which use frequency-dependent acoustic absorption coefficients (MB-RIRs) can obtain +0.51dB of SDR and a +8.9 MUSHRA score when evaluated on real RIRs. The MB-RIRs dataset is publicly available for free download.


【26】THAI Speech Emotion Recognition (THAI-SER) corpus
标题:泰语语音情感识别(THAI-SER)语料库
链接:https://arxiv.org/abs/2507.09618

作者:Wongpithayadisai, Chompakorn Chaksangchaichot, Soravitt Sangnark, Patawee Prakrankamanant, Krit Gangwanpongpun, Siwa Boonpunmongkol, Premmarin Milindasuta, Dangkamon Na-Pombejra, Sarana Nutanong, Ekapol Chuangsuwanich
摘要:我们提出了第一个相当大的语料库泰国语音情感识别,THAI-SER,包含41小时36分钟(27,854话语),从100个录音在不同的录音环境:变焦和两个工作室设置。这些录音包括脚本和即兴表演,由200名专业演员(112名女性和88名男性,年龄在18至55岁之间)表演,并由专业导演执导。有五种主要的情绪:中性,愤怒,快乐,悲伤和沮丧,在记录话语时分配给演员。使用众包用情感类别注释话语。为了控制注释过程的质量,我们还设计了一个广泛的过滤和质量控制方案,以确保大多数协议得分保持在0.71以上。我们使用两个指标来评估我们的注释语料库:注释者间的可靠性和人类识别的准确性。使用Krippendorff的alpha计算注释者间的可靠性得分,其中我们的语料库在过滤后达到了0.692的alpha得分,高于推荐的0.667。对于人类识别准确性,我们的语料库在过滤后得分高达0.772。我们还提供了在语料库内和跨语料库设置上评估的语料库上训练的模型的结果。语料库是公开的知识共享BY-SA 4.0下,以及我们的实验代码。
摘要:We present the first sizeable corpus of Thai speech emotion recognition, THAI-SER, containing 41 hours and 36 minutes (27,854 utterances) from 100 recordings made in different recording environments: Zoom and two studio setups. The recordings contain both scripted and improvised sessions, acted by 200 professional actors (112 females and 88 males, aged 18 to 55) and were directed by professional directors. There are five primary emotions: neutral, angry, happy, sad, and frustrated, assigned to the actors when recording utterances. The utterances are annotated with an emotional category using crowdsourcing. To control the annotation process's quality, we also design an extensive filtering and quality control scheme to ensure that the majority agreement score remains above 0.71. We evaluate our annotated corpus using two metrics: inter-annotator reliability and human recognition accuracy. Inter-annotator reliability score was calculated using Krippendorff's alpha, where our corpus, after filtering, achieved an alpha score of 0.692, higher than a recommendation of 0.667. For human recognition accuracy, our corpus scored up to 0.772 post-filtering. We also provide the results of the model trained on the corpus evaluated on both in-corpus and cross-corpus setups. The corpus is publicly available under a Creative Commons BY-SA 4.0, as well as our codes for the experiments.


【27】Ensemble Confidence Calibration for Sound Event Detection in Open-environment
标题:开放环境下声事件检测的包围置信度标定
链接:https://arxiv.org/abs/2507.09606

作者:Chen, Han Yin
摘要:声音事件检测(SED)在具有明确事件类别的受控环境中取得了很大进展。然而,现实世界的应用程序往往发生在开放的环境中。在这种情况下,目前的方法往往会产生过于自信的预测,缺乏适当的方法来衡量不确定性。这限制了他们在新情况下的适应能力和良好表现。为了解决这个问题,我们是第一个使用集成方法在SED,以提高对域外(OOD)输入的鲁棒性。我们提出了一种称为基于能量的开放世界Softmax(EOW-Softmax)的置信度校准方法,该方法可以帮助系统更好地处理未知场景中的不确定性。我们进一步将EOW-Softmax应用于声音发生和重叠检测(SOD),通过调整预测。通过这种方式,模型变得更具适应性,同时保持其检测重叠事件的能力。实验表明,我们的方法提高了在开放环境中的性能。它减少过度自信,提高处理OOD情况的能力。
摘要:Sound event detection (SED) has made strong progress in controlled environments with clear event categories. However, real-world applications often take place in open environments. In such cases, current methods often produce predictions with too much confidence and lack proper ways to measure uncertainty. This limits their ability to adapt and perform well in new situations. To solve this problem, we are the first to use ensemble methods in SED to improve robustness against out-of-domain (OOD) inputs. We propose a confidence calibration method called Energy-based Open-World Softmax (EOW-Softmax), which helps the system better handle uncertainty in unknown scenes. We further apply EOW-Softmax to sound occurrence and overlap detection (SOD) by adjusting the prediction. In this way, the model becomes more adaptable while keeping its ability to detect overlapping events. Experiments show that our method improves performance in open environments. It reduces overconfidence and increases the ability to handle OOD situations.


【28】SC-TSE: Speaker Consistency-Aware Target Speaker Extraction
标题:SC-PSE:说话者一致性感知目标说话者提取
链接:https://arxiv.org/abs/2507.09510

作者:nbin Qi, Yanzhang Xie, Xiang Xie
备注:Accept to Interspeech2025
摘要:目标说话人提取(TSE)使用参考线索从混合语音中提取目标语音。在依赖于音频线索的TSE系统中,从注册语音中嵌入说话人对性能至关重要。然而,这些嵌入可能遭受说话人身份混淆。与以往的研究侧重于提高说话人嵌入提取不同,本文从说话人一致性的角度提高了TSE的性能。在本文中,我们提出了一个说话人一致性意识的目标说话人提取方法,结合了基于质心的说话人一致性损失。该方法通过确保登记的语音和提取的语音之间的说话者一致性来增强TSE性能。此外,我们将条件损失抑制集成到训练过程中。实验结果验证了我们提出的方法在提高TSE性能的有效性。在线提供演讲演示。\脚注{https:sc-tse.netlify.app/
摘要:Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled speech is crucial to performance. However, these embeddings may suffer from speaker identity confusion. Unlike previous studies that focus on improving speaker embedding extraction, we improve TSE performance from the perspective of speaker consistency. In this paper, we propose a speaker consistency-aware target speaker extraction method that incorporates a centroid-based speaker consistency loss. This approach enhances TSE performance by ensuring speaker consistency between the enrolled and extracted speech. In addition, we integrate conditional loss suppression into the training process. The experimental results validate the effectiveness of our proposed methods in advancing the TSE performance. A speech demo is available online.\footnote{https://sc-tse.netlify.app/


【29】Acoustic Wave Modeling Using 2D FDTD: Applications in Unreal Engine For Dynamic Sound Rendering
标题:使用2D MTD的声波建模:在动态声音渲染的虚幻引擎中的应用
链接:https://arxiv.org/abs/2507.09376

作者:amsurya
备注:Accepted to the 50th International Computer Music Conference (ICMC), 2025
摘要:准确的声音传播仿真对于在虚拟应用中提供沉浸式体验至关重要,但声学建模的行业方法通常无法解释声波现象的全部范围。本文提出了一种新的二维时域有限差分(FDTD)框架,在虚幻引擎中将声音传播模拟为基于波的模型,重点是捕获低频波现象,在生成的脉冲响应中嵌入遮挡、衍射、反射和干涉。该过程首先通过自上而下的投影将场景几何形状离散化为2D网格,从中导出障碍物遮罩和边界条件。基于Python的FDTD求解器在源位置注入正弦扫描,虚拟四声道麦克风阵列在预定义的收听者位置记录压力场响应。压力响应的去卷积产生多通道脉冲响应,这些脉冲响应保持空间方向性,然后将其集成到虚幻引擎的音频管道中进行动态播放。基准测试证实了与分析预期的一致性,论文概述了旨在实现商业可行性的混合扩展。
摘要:Accurate sound propagation simulation is essential for delivering immersive experiences in virtual applications, yet industry methods for acoustic modeling often do not account for the full breadth of acoustic wave phenomena. This paper proposes a novel two-dimensional (2D) finite-difference time-domain (FDTD) framework that simulates sound propagation as a wave-based model in Unreal Engine, with an emphasis on capturing lower frequency wave phenomena, embedding occlusion, diffraction, reflection and interference in generated impulse responses. The process begins by discretizing the scene geometry into a 2D grid via a top-down projection from which obstacle masks and boundary conditions are derived. A Python-based FDTD solver injects a sine sweep at a source position, and virtual quadraphonic microphone arrays record pressure field responses at pre-defined listener positions. De-convolution of the pressure responses yields multi-channel impulse responses that retain spatial directionality which are then integrated into Unreal Engine's audio pipeline for dynamic playback. Benchmark tests confirm agreement with analytical expectations, and the paper outlines hybrid extensions aimed at commercial viability.


【30】BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus
标题:BENYO-S2 ST-Corpus-1:英语到约鲁巴语的双语直接Speech-to-Speech翻译库
链接:https://arxiv.org/abs/2507.09342

作者:Adetiba, Abdultaofeek Abayomi, Raymond J. Kala, Ayodele H. Ifijeh, Oluwatobi E. Dare, Olabode Idowu-Bismark, Gabriel O. Sobola, Joy N. Adetiba, Monsurat Adepeju Lateef, Heather Cole-Lewis
摘要:语音到语音翻译(S2 ST)数据集严重短缺,用于高资源到低资源的语言对,如英语到约鲁巴语。因此,在这项研究中,我们策划了双语英语到约鲁巴语语音到语音翻译语料库版本1(BENYO-S2 ST-Corpus-1)。语料库是基于一个混合架构,我们开发的大规模直接S2 ST语料库创建以降低成本。为了实现这一目标,我们利用了非语音到语音的标准约鲁巴语(SY)实时音频和成绩单在YORULECT语料库以及相应的标准英语(SE)成绩单。YORULECT语料库是小规模的(1,504)样本,它没有配对的英语音频。因此,我们使用预训练的AI模型(即Facebook MMS)生成SE音频。我们还开发了一种名为AcoustAug的音频增强算法,基于三个潜在的声学特征,从两种语言的原始音频生成增强音频。BENYO-S2 ST-Corpus-1每种语言有12,032个音频样本,总共有24,064个样本大小。两种语言的音频总时长为41.20小时。这个尺寸是相当重要的。除了构建S2 ST模型之外,BENYO-S2 ST-Corpus-1还可以用于构建预训练模型或改进现有模型。使用创建的语料库和Coqui框架构建预训练的约鲁巴语TTS模型(命名为YoruTTS-0.5)作为概念证明。YoruTTS-0.5在1,000个epoch之后给出了63.54的F0 RMSE值,这表明与参考实时音频具有中等的基本音高相似性。最终,研究人员和开发人员可以利用本研究中的语料库架构来管理多语言高资源到低资源非洲语言的数据集。这将弥合高资源和低资源语言对之间翻译的巨大数字鸿沟。BENYO-S2 ST-Corpus-1和YoruTTS-0.5可在(https://bit.ly/40bGMwi)公开获得。
摘要:There is a major shortage of Speech-to-Speech Translation (S2ST) datasets for high resource-to-low resource language pairs such as English-to-Yoruba. Thus, in this study, we curated the Bilingual English-to-Yoruba Speech-to-Speech Translation Corpus Version 1 (BENYO-S2ST-Corpus-1). The corpus is based on a hybrid architecture we developed for large-scale direct S2ST corpus creation at reduced cost. To achieve this, we leveraged non speech-to-speech Standard Yoruba (SY) real-time audios and transcripts in the YORULECT Corpus as well as the corresponding Standard English (SE) transcripts. YORULECT Corpus is small scale(1,504) samples, and it does not have paired English audios. Therefore, we generated the SE audios using pre-trained AI models (i.e. Facebook MMS). We also developed an audio augmentation algorithm named AcoustAug based on three latent acoustic features to generate augmented audios from the raw audios of the two languages. BENYO-S2ST-Corpus-1 has 12,032 audio samples per language, which gives a total of 24,064 sample size. The total audio duration for the two languages is 41.20 hours. This size is quite significant. Beyond building S2ST models, BENYO-S2ST-Corpus-1 can be used to build pretrained models or improve existing ones. The created corpus and Coqui framework were used to build a pretrained Yoruba TTS model (named YoruTTS-0.5) as a proof of concept. The YoruTTS-0.5 gave a F0 RMSE value of 63.54 after 1,000 epochs, which indicates moderate fundamental pitch similarity with the reference real-time audio. Ultimately, the corpus architecture in this study can be leveraged by researchers and developers to curate datasets for multilingual high-resource-to-low-resource African languages. This will bridge the huge digital divides in translations among high and low-resource language pairs. BENYO-S2ST-Corpus-1 and YoruTTS-0.5 are publicly available at (https://bit.ly/40bGMwi).


【31】Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning
标题:具有隐性和显式声学特征条件反射的伦巴第说话风格的声音转换
链接:https://arxiv.org/abs/2507.09310

作者:Woszczyk, Manuel Sam Ribeiro, Thomas Merritt, Daniel Korzekwa
备注:Presented at Clarity Challenge 2023
摘要:朗伯说话风格的文本到语音(TTS)系统可以提高语音的整体清晰度,适用于听力损失和嘈杂环境。然而,训练这些模型需要大量的数据,并且由于扬声器和噪声的可变性以及令人疲劳的记录条件,Lombard效应的记录具有挑战性。语音转换(VC)已被证明是一种有用的增强技术,训练TTS系统的记录数据的情况下,从目标说话人在目标说话风格。在本文中,我们关注的是伦巴第语风格的迁移。我们的目标是转换扬声器身份,同时保留定义伦巴第语风格的声学属性。我们比较语音转换模型与隐式和显式的声学特征条件。我们观察到,我们提出的隐式条件反射策略实现了可懂度增益相比,模型条件显式的声学特征,同时还保持扬声器的相似性。
摘要:Text-to-Speech (TTS) systems in Lombard speaking style can improve the overall intelligibility of speech, useful for hearing loss and noisy conditions. However, training those models requires a large amount of data and the Lombard effect is challenging to record due to speaker and noise variability and tiring recording conditions. Voice conversion (VC) has been shown to be a useful augmentation technique to train TTS systems in the absence of recorded data from the target speaker in the target speaking style. In this paper, we are concerned with Lombard speaking style transfer. Our goal is to convert speaker identity while preserving the acoustic attributes that define the Lombard speaking style. We compare voice conversion models with implicit and explicit acoustic feature conditioning. We observe that our proposed implicit conditioning strategy achieves an intelligibility gain comparable to the model conditioned on explicit acoustic features, while also preserving speaker similarity.


【32】ClaritySpeech: Dementia Obfuscation in Speech
标题:ClaritySpeech:言语中的痴呆症混淆
链接:https://arxiv.org/abs/2507.09282

作者:Woszczyk, Ranya Aloufi, Soteris Demetriou
备注:Accepted at Interspeech 2025
摘要:痴呆症是一种神经退行性疾病,会改变言语模式,造成沟通障碍,并引发隐私问题。目前的语音技术,如自动语音转录(ASR),与痴呆症和非典型语音斗争,进一步挑战可访问性。本文提出了一种新的痴呆模糊语音框架ClaritySpeech,集成ASR,文本模糊,和zero-shot文本到语音(TTS),以纠正受痴呆影响的语音,同时在低数据环境中保持说话人身份,无需微调。结果显示,在ADReSS和ADReSSo的各种对抗性设置和模式(音频,文本,融合)中,平均F1得分分别下降了16%和10%,保持了50%的说话人相似性。我们还发现,我们的系统提高了WER(从0.73到0.08 ADReSS和0.15 ADReSSo)和语音质量从1.65到~2.15,提高了隐私和可访问性。
摘要:Dementia, a neurodegenerative disease, alters speech patterns, creating communication barriers and raising privacy concerns. Current speech technologies, such as automatic speech transcription (ASR), struggle with dementia and atypical speech, further challenging accessibility. This paper presents a novel dementia obfuscation in speech framework, ClaritySpeech, integrating ASR, text obfuscation, and zero-shot text-to-speech (TTS) to correct dementia-affected speech while preserving speaker identity in low-data environments without fine-tuning. Results show a 16% and 10% drop in mean F1 score across various adversarial settings and modalities (audio, text, fusion) for ADReSS and ADReSSo, respectively, maintaining 50% speaker similarity. We also find that our system improves WER (from 0.73 to 0.08 for ADReSS and 0.15 for ADReSSo) and speech quality from 1.65 to ~2.15, enhancing privacy and accessibility.


【33】Towards Spatial Audio Understanding via Question Answering
标题:基于问题生成的空间音频理解
链接:https://arxiv.org/abs/2507.09195

作者:rathy Sudarsanam, Archontis Politis
摘要:在本文中,我们介绍了一种新的框架,通过问答(QA)范式的一阶立体混响(FOA)信号的空间音频理解,旨在扩展的范围内的声音事件定位和检测(SELD)对空间场景的理解和推理。首先,我们使用基于规则的方法为STARSS 23数据集策划和发布细粒度的时空文本描述,并使用基于大型语言模型(LLM)的改写进一步增强语言多样性。我们还介绍了与STARSS23场景对齐的QA数据集,涵盖了事件存在,本地化,空间和时间关系等各个方面。为了增加语言的多样性,我们再次利用LLM为每个问题生成多个改写。最后,我们开发了一个基线空间音频QA模型,该模型以FOA信号和自然语言问题为输入,并提供关于场景中声音事件的各种发生,时间和空间关系的答案,该模型被制定为分类任务。尽管仅使用场景级问答监督进行训练,但我们的模型实现了与使用帧级时空注释训练的完全监督的声音事件定位和检测模型相当的性能。结果突出了空间音频理解的语言指导方法的潜力,并为将语言监督整合到空间场景分析中开辟了新的方向。
摘要:In this paper, we introduce a novel framework for spatial audio understanding of first-order ambisonic (FOA) signals through a question answering (QA) paradigm, aiming to extend the scope of sound event localization and detection (SELD) towards spatial scene understanding and reasoning. First, we curate and release fine-grained spatio-temporal textual descriptions for the STARSS23 dataset using a rule-based approach, and further enhance linguistic diversity using large language model (LLM)-based rephrasing. We also introduce a QA dataset aligned with the STARSS23 scenes, covering various aspects such as event presence, localization, spatial, and temporal relationships. To increase language variety, we again leverage LLMs to generate multiple rephrasings per question. Finally, we develop a baseline spatial audio QA model that takes FOA signals and natural language questions as input and provides answers regarding various occurrences, temporal, and spatial relationships of sound events in the scene formulated as a classification task. Despite being trained solely with scene-level question answering supervision, our model achieves performance that is comparable to a fully supervised sound event localization and detection model trained with frame-level spatiotemporal annotations. The results highlight the potential of language-guided approaches for spatial audio understanding and open new directions for integrating linguistic supervision into spatial scene analysis.


【34】Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
标题:LoRA专家的混合具有多模式和多粒度LLM生成错误纠正,用于强调语音识别
链接:https://arxiv.org/abs/2507.09116

作者:Mu, Kun Wei, Pengcheng Guo, Lei Xie
备注:IEEE Transactions on Audio, Speech and Language Processing
摘要:尽管ASR有了很大的改进,但在面对不利条件(如说话者口音)时,性能往往会下降。生成式纠错(GER)利用LLM丰富的语言知识和卓越的推理能力,显著优于典型的LM方法。然而,它在口音语音场景中缺乏特异性。在这项研究中,我们利用GER,以提高准确性的转录预测,解决口音语音识别的两个主要特征。为了充分利用发音信息,我们提出了多模态GER,它集成了语音模态的发音信息,以及多粒度GER,它包含了与发音相关的细粒度音素级信息。这两种方法使LLM能够利用口音语音的发音信息和来自单词级别假设的语义信息,通过LoRA微调进行准确的转录预测。一方面,我们采用三阶段训练策略来为每个口音训练单独的多模态GER模型,以获得单口音LoRA专家。通过采用我们提出的HDMoLE方法,该方法在LoRA专家的混合物中结合了分层路由和动态阈值,我们有效地将多个单口音LoRA专家合并在单个多模态GER中,以克服口音多样性带来的挑战。另一方面,多粒度GER利用由HDMoLE模型生成的N-最佳词级和音素级假设来预测最终的带口音语音传输。多口音英语数据集上的实验结果证明了我们提出的方法的有效性。与Whisper-large-v3基线相比,我们的方法实现了67.35%的显著相对WER降低。
摘要:Despite substantial improvements in ASR, performance tends to degrade when faced with adverse conditions such as speaker accents. Generative error correction (GER) leverages the rich linguistic knowledge and exceptional reasoning ability of LLMs, significantly outperforming typical LM methods. However, it lacks specificity in accented speech scenarios. In this study, we leverage GER to improve the accuracy of transcription predictions by addressing the two primary features of accented speech recognition. To fully leverage pronunciation information, we propose the multi-modal GER, which integrates pronunciation information from the speech modality, and the multi-granularity GER, which incorporates fine-grained phoneme-level information related to pronunciation. These two methods enable the LLM to utilize the pronunciation information of accented speech and the semantic information from word-level hypotheses for accurate transcription predictions through LoRA fine-tuning. On the one hand, we employ a three-stage training strategy to train separate multi-modal GER models for each accent to obtain mono-accent LoRA experts. By adopting our proposed HDMoLE method, which incorporates hierarchical routing and dynamic thresholds within the mixture of LoRA experts, we effectively merge multiple mono-accent LoRA experts within a single multi-modal GER to overcome the challenges posed by accent diversity. On the other hand, multi-granularity GER leverages the N-best word-level and phoneme-level hypotheses generated by the HDMoLE model to predict the final accented speech transcriptions. Experimental results on the multi-accent English dataset demonstrate the efficacy of our proposed methods. Our methods achieve a remarkable relative WER reduction of 67.35% compared to the Whisper-large-v3 baseline.


【35】Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
标题:更少的压力,更多的隐私:空中交通管制员神经化语音的压力检测
链接:https://arxiv.org/abs/2507.08882

作者:swanathan, Alexander Blatt, Konrad Hagemann, Dietrich Klakow
备注:8 pages, 2 figures, 4 tables, publication identification number   (URN)- urn:nbn:de:101:1-2022122008393409239462, see archived online   publication- https://d-nb.info/127614606X/34 & Katalogeintrag:   https://d-nb.info/127614606X/
摘要:空中交通管制(ATC)要求在时间压力下进行多任务处理,错误的后果很高。这可能会引起压力。压力检测是保持ATC高安全标准的关键。然而,处理ATC语音数据需要隐私限制,例如通用数据保护条例(GDPR)法。分析ATC语音数据是遵守这些限制的一种方法。在本文中,不同的架构,匿名ATCO语音的压力检测进行评估。我们最好的网络在模拟和实际压力下的语音(SUSAS)数据集的匿名版本上达到了93.6%的压力检测准确率,在我们的匿名ATC模拟数据集上达到了80.1%的准确率。这表明,隐私不一定是构建性能良好的基于深度学习的模型的障碍。
摘要:Air traffic control (ATC) demands multi-tasking under time pressure with high consequences of an error. This can induce stress. Detecting stress is a key point in maintaining the high safety standards of ATC. However, processing ATC voice data entails privacy restrictions, e.g. the General Data Protection Regulation (GDPR) law. Anonymizing the ATC voice data is one way to comply with these restrictions. In this paper, different architectures for stress detection for anonymized ATCO speech are evaluated. Our best networks reach a stress detection accuracy of 93.6% on an anonymized version of the Speech Under Simulated and Actual Stress (SUSAS) dataset and an accuracy of 80.1% on our anonymized ATC simulation dataset. This shows that privacy does not have to be an impediment in building well-performing deep-learning-based models.


机器翻译由腾讯交互翻译提供,仅供参考