今日论文合集:cs.SD语音17篇,eess.AS音频处理22篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
标题:ISA-Bench:大型音频语言模型的基准指令灵敏度
链接:https://arxiv.org/abs/2510.23558

作者:Bohan Li, Wenbin Huang, Yuhang Qiu, Yiwei Guo, Hankun Wang, Zhihan Li, Jing Peng, Ziyang Ma, Xie Chen, Kai Yu
备注:submitted to icassp 2026
摘要:大型音频语言模型(LALM)将声学感知与大型语言模型(LLM)相结合,以从音频中提取和理解各种信息,引起了学术界和工业界的浓厚兴趣。然而,现有的LALM对指令的措辞方式高度敏感,影响(i)遵守率和(ii)任务性能。然而,没有任何现有的基准对这种敏感性进行系统和全面的评价。我们介绍ISA-Bench,一个动态的基准评估指令灵敏度LALM沿三个轴:指令描述,输出格式,任务组成。我们使用ISA-Bench评估最近的开源和专有LALM,在受控指令变化下分析合规性和准确性。实验结果表明,即使是最先进的LALM遭受显着的指令敏感性,导致性能下降的基本音频理解任务。为了缓解这个问题,我们在一个专门构建的复杂的自适应变量数据集上对Qwen 2-Audio进行了微调,从而显著提高了自适应跟踪性能。然而,这也导致了非平凡的灾难性遗忘:当暴露于新的指令风格时,模型失去了一些以前掌握的任务能力。我们的基准测试为评估和提高LALM中的指令灵敏度提供了标准化的基础,强调了在现实世界的管道中对预防性强大的音频理解的需求。
摘要:Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted intense interest from both academic and industrial communities. However, existing LALMs are highly sensitive to how instructions are phrased, affecting both (i) instruction-following rates and (ii) task performance. Yet, no existing benchmarks offer a systematic and comprehensive evaluation of this sensitivity. We introduce ISA-Bench, a dynamic benchmark evaluating instruction sensitivity for LALMs along three axes: instruction description, output format, and task composition. We assess recent open-source and proprietary LALMs using ISA-Bench, profiling both compliance and accuracy under controlled instruction variations. Experimental results reveal that even state-of-the-art LALMs suffer significant instruction sensitivity, leading to degraded performance on fundamental audio understanding tasks. To mitigate this issue, we fine-tune Qwen2-Audio on a specifically constructed complex instruction-variant dataset, achieving a marked improvement in instruction-following performance. However, this also induces nontrivial catastrophic forgetting: the model loses some previously mastered task capabilities when exposed to new instruction styles. Our benchmark provides a standardized basis for assessing and improving instruction sensitivity in LALMs, underscoring the need for instruction-robust audio understanding in real-world pipelines.


【2】Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
标题:基于隐式正则化的音频一致性自编码器线性学习
链接:https://arxiv.org/abs/2510.23530

作者:Bernardo Torres, Manuel Moussallam, Gabriel Meseguer-Brocal
摘要:音频自动编码器学习有用的压缩音频表示,但它们的非线性潜在空间阻止了直观的代数操作,如混合或缩放。我们介绍了一种简单的训练方法,通过使用数据增强来诱导高压缩一致性自动编码器(CAE)中的线性,从而诱导同质性(标量增益的等方差)和可加性(解码器保留加法),而不改变模型的架构或损失函数。当用我们的方法训练时,CAE在编码器和解码器中都表现出线性行为,同时保持重建保真度。我们通过简单的潜在算法测试我们的学习空间对音乐源合成和分离的实际效用。这项工作提出了一种简单的技术,用于构建结构化的潜在空间,使更直观和有效的音频处理。
摘要:Audio autoencoders learn useful, compressed audio representations, but their non-linear latent spaces prevent intuitive algebraic manipulation such as mixing or scaling. We introduce a simple training methodology to induce linearity in a high-compression Consistency Autoencoder (CAE) by using data augmentation, thereby inducing homogeneity (equivariance to scalar gain) and additivity (the decoder preserves addition) without altering the model's architecture or loss function. When trained with our method, the CAE exhibits linear behavior in both the encoder and decoder while preserving reconstruction fidelity. We test the practical utility of our learned space on music source composition and separation via simple latent arithmetic. This work presents a straightforward technique for constructing structured latent spaces, enabling more intuitive and efficient audio processing.


【3】Arabic Little STT: Arabic Children Speech Recognition Dataset
标题:阿拉伯语Little STT:阿拉伯语儿童语音识别数据集
链接:https://arxiv.org/abs/2510.23319

作者:Mouhand Alkadri, Dania Desouki, Khloud Al Jallad
摘要:人工智能(AI)系统的性能从根本上取决于高质量的训练数据。然而,像阿拉伯语这样的低资源语言遭受严重的数据稀缺。此外,缺乏儿童专用语音语料库是一个重要的差距,构成了重大挑战。为了解决这一差距,我们提出了我们创建的数据集,阿拉伯语小STT,在教室里记录的黎凡特阿拉伯语儿童语音的数据集,包含来自288名儿童(6 - 13岁)的355个话语。我们进一步对Whisper(一种最先进的自动语音识别(ASR)模型)进行了系统的评估,并将其性能与成人阿拉伯语基准进行了比较。我们对八种Whisper变体的评估表明,即使是性能最好的模型(Large_v3)也很难在儿童语音上实现0.66的单词错误率(WER),与成人数据集上低于0.20的WER形成鲜明对比。这些结果与其他关于英语演讲的研究一致。研究结果强调,迫切需要专门的儿童语音基准和包容性的培训数据在ASR的发展。强调此类数据必须受到严格的道德和隐私框架的管理,以保护敏感的儿童信息。我们希望这项研究为今后为讲阿拉伯语的儿童提供公平的语音技术的工作迈出了第一步。我们希望我们公开的数据集丰富了ASR数据集中儿童的人口统计学表示。
摘要:The performance of Artificial Intelligence (AI) systems fundamentally depends on high-quality training data. However, low-resource languages like Arabic suffer from severe data scarcity. Moreover, the absence of child-specific speech corpora is an essential gap that poses significant challenges. To address this gap, we present our created dataset, Arabic Little STT, a dataset of Levantine Arabic child speech recorded in classrooms, containing 355 utterances from 288 children (ages 6 - 13). We further conduct a systematic assessment of Whisper, a state-of-the-art automatic speech recognition (ASR) model, on this dataset and compare its performance with adult Arabic benchmarks. Our evaluation across eight Whisper variants reveals that even the best-performing model (Large_v3) struggles significantly, achieving a 0.66 word error rate (WER) on child speech, starkly contrasting with its sub 0.20 WER on adult datasets. These results align with other research on English speech. Results highlight the critical need for dedicated child speech benchmarks and inclusive training data in ASR development. Emphasizing that such data must be governed by strict ethical and privacy frameworks to protect sensitive child information. We hope that this study provides an initial step for future work on equitable speech technologies for Arabic-speaking children. We hope that our publicly available dataset enrich the children's demographic representation in ASR datasets.


【4】Low-Resource Audio Codec (LRAC): 2025 Challenge Description
标题:低资源音频编解码器(LRAC):2025年挑战描述
链接:https://arxiv.org/abs/2510.23312

作者:Kamil Wojcicki, Yusuf Ziya Isik, Laura Lechler, Mansur Yesilbursa, Ivana Balić, Wolfgang Mack, Rafał Łaganowski, Guoqing Zhang, Yossi Adi, Minje Kim, Shinji Watanabe
摘要:虽然最近的神经音频编解码器在超低比特率下提供优于传统方法的语音质量,但它们的实际采用受到与低资源操作和对声学失真的鲁棒性相关的障碍的阻碍。边缘部署场景要求编解码器在严格的计算约束下运行,同时保持低延迟和比特率。背景噪声和混响的存在进一步需要对这种退化有弹性的设计。神经编解码器在这些约束下的性能及其与语音增强的集成在很大程度上仍未得到解决。为了促进这一领域的进展,我们推出了2025年低资源音频编解码器挑战赛,其目标是为资源受限的应用开发神经和混合编解码器。参与者将获得一个标准化的培训数据集、两个基线系统和一个全面的评估框架。预计该挑战将产生适用于编解码器设计和相关下游音频任务的有价值的见解。
摘要:While recent neural audio codecs deliver superior speech quality at ultralow bitrates over traditional methods, their practical adoption is hindered by obstacles related to low-resource operation and robustness to acoustic distortions. Edge deployment scenarios demand codecs that operate under stringent compute constraints while maintaining low latency and bitrate. The presence of background noise and reverberation further necessitates designs that are resilient to such degradations. The performance of neural codecs under these constraints and their integration with speech enhancement remain largely unaddressed. To catalyze progress in this area, we introduce the 2025 Low-Resource Audio Codec Challenge, which targets the development of neural and hybrid codecs for resource-constrained applications. Participants are supported with a standardized training dataset, two baseline systems, and a comprehensive evaluation framework. The challenge is expected to yield valuable insights applicable to both codec design and related downstream audio tasks.


【5】TwinShift: Benchmarking Audio Deepfake Detection across Synthesizer and Speaker Shifts
标题:TwinChange:跨合成器和扬声器转换的音频Deepfake检测基准
链接:https://arxiv.org/abs/2510.23096

作者:Jiyoung Hong, Yoonseo Chung, Seungyeon Oh, Juntae Kim, Jiyoung Lee, Sookyung Kim, Hyunsoo Cho
备注:Submitted to ICASSP 2026
摘要:音频deepfake构成了越来越大的威胁,已经被用于欺诈和错误信息。一个关键的挑战是确保检测器对看不见的合成方法和不同的扬声器保持鲁棒性,因为生成技术发展迅速。尽管有很强的基准测试结果,但目前的系统很难推广到限制现实世界可靠性的新条件。为了解决这个问题,我们引入了TWINCITY,这是一个明确设计用于在严格看不见的条件下评估检测鲁棒性的基准。我们的基准测试由六个不同的合成系统构建,每个系统与不相交的扬声器集配对,从而可以严格评估当生成模型和扬声器身份发生变化时检测器的泛化能力。通过广泛的实验,我们表明,TWINESTRUCTURE揭示了重要的鲁棒性差距,发现被忽视的局限性,并提供原则性的指导开发ADD系统。可以在https://github.com/intheMeantime/TWINSHIFT上访问TWINDB2基准测试。
摘要:Audio deepfakes pose a growing threat, already exploited in fraud and misinformation. A key challenge is ensuring detectors remain robust to unseen synthesis methods and diverse speakers, since generation techniques evolve quickly. Despite strong benchmark results, current systems struggle to generalize to new conditions limiting real-world reliability. To address this, we introduce TWINSHIFT, a benchmark explicitly designed to evaluate detection robustness under strictly unseen conditions. Our benchmark is constructed from six different synthesis systems, each paired with disjoint sets of speakers, allowing for a rigorous assessment of how well detectors generalize when both the generative model and the speaker identity change. Through extensive experiments, we show that TWINSHIFT reveals important robustness gaps, uncover overlooked limitations, and provide principled guidance for developing ADD systems. The TWINSHIFT benchmark can be accessed at https://github.com/intheMeantime/TWINSHIFT.


【6】SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
标题:SAO-Instruct:使用自然语言指令进行自由格式音频编辑
链接:https://arxiv.org/abs/2510.22795

作者:Michael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi, Changho Choi, Roger Wattenhofer
备注:Accepted at NeurIPS 2025
摘要:生成模型在从简短的文本描述合成高保真音频方面取得了重大进展。然而,使用自然语言编辑现有音频在很大程度上仍然没有得到充分探索。当前的方法要么需要对所编辑的音频的完整描述,要么受限于缺乏灵活性的预定义编辑指令。在这项工作中,我们介绍了SAO-Instruct,一个基于Stable Audio Open的模型,能够使用任何自由形式的自然语言指令编辑音频片段。为了训练我们的模型,我们创建了一个音频编辑三元组(输入音频,编辑指令,输出音频)的数据集,使用提示,DDPM反转和手动编辑管道。虽然我们的模型部分基于合成数据进行了训练,但它可以很好地推广到真实的野外音频片段和看不见的编辑指令。我们证明了SAO-Instruct在客观指标上取得了有竞争力的性能,并在主观听力研究中优于其他音频编辑方法。为了鼓励未来的研究,我们发布了我们的代码和模型权重。
摘要:Generative models have made significant progress in synthesizing high-fidelity audio from short textual descriptions. However, editing existing audio using natural language has remained largely underexplored. Current approaches either require the complete description of the edited audio or are constrained to predefined edit instructions that lack flexibility. In this work, we introduce SAO-Instruct, a model based on Stable Audio Open capable of editing audio clips using any free-form natural language instruction. To train our model, we create a dataset of audio editing triplets (input audio, edit instruction, output audio) using Prompt-to-Prompt, DDPM inversion, and a manual editing pipeline. Although partially trained on synthetic data, our model generalizes well to real in-the-wild audio clips and unseen edit instructions. We demonstrate that SAO-Instruct achieves competitive performance on objective metrics and outperforms other audio editing approaches in a subjective listening study. To encourage future research, we release our code and model weights.


【7】Evaluating Multimodal Large Language Models on Core Music Perception Tasks
标题:评估核心音乐感知任务的多模式大型语言模型
链接:https://arxiv.org/abs/2510.22455

作者:Brandon James Carone, Iran R. Roman, Pablo Ripollés
备注:Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables
摘要:多模态大型语言模型(LLM)通过将听力与乐谱阅读相结合的评估来声称“音乐理解”。我们在三个核心音乐技能上对三个SOTA LLM(Gemini 2.5 Pro,Gemini 2.5 Flash和Qwen2.5-Omni)进行基准测试:切分音评分,换位检测和和弦质量识别。此外,我们分离出三个可变性来源:(i)感知限制(音频与音频输入),(ii)暴露于示例(零与Few-Shot操作),以及(iii)推理策略(独立,CoT,LogicLM)。对于后者,我们适应LogicLM,一个框架结合LLM与符号求解器进行结构化推理,音乐。结果显示了一个明显的感知差距:模型在音频上的表现接近天花板,但在音频上的准确性下降。推理和Few-Shot提示只能提供最小的收益。这对于性能达到饱和的音频来说是意料之中的,但对于音频来说更令人惊讶的是,LogicLM尽管有近乎完美的音频精度,但仍然非常脆弱。在所有型号中,Gemini Pro在大多数条件下都能实现最高性能。总的来说,当前的系统在符号(symbols,简写为symbol)上推理得很好,但还不能可靠地从音频中“收听”。我们的方法和数据集使感知推理边界明确,并为构建强大的音频优先音乐系统提供可操作的指导。
摘要:Multimodal Large Language Models (LLMs) claim "musical understanding" via evaluations that conflate listening with score reading. We benchmark three SOTA LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni) across three core music skills: Syncopation Scoring, Transposition Detection, and Chord Quality Identification. Moreover, we separate three sources of variability: (i) perceptual limitations (audio vs. MIDI inputs), (ii) exposure to examples (zero- vs. few-shot manipulations), and (iii) reasoning strategies (Standalone, CoT, LogicLM). For the latter we adapt LogicLM, a framework combining LLMs with symbolic solvers to perform structured reasoning, to music. Results reveal a clear perceptual gap: models perform near ceiling on MIDI but show accuracy drops on audio. Reasoning and few-shot prompting offer minimal gains. This is expected for MIDI, where performance reaches saturation, but more surprising for audio, where LogicLM, despite near-perfect MIDI accuracy, remains notably brittle. Among models, Gemini Pro achieves the highest performance across most conditions. Overall, current systems reason well over symbols (MIDI) but do not yet "listen" reliably from audio. Our method and dataset make the perception-reasoning boundary explicit and offer actionable guidance for building robust, audio-first music systems.


【8】PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
标题:通过潜在整流匹配产生多模态房间脉冲响应
链接:https://arxiv.org/abs/2510.22439

作者:Ali Vosoughi, Yongyi Zang, Qihui Yang, Nathan Peak, Randal Leistikow, Chenliang Xu
备注:9 pages, 2 figures, 4 tables
摘要:房间脉冲响应(RIR)生成仍然是创建沉浸式虚拟声学环境的关键挑战。目前的方法有两个基本的限制:全波段RIR数据集的稀缺性和现有的模型无法从不同的输入模态生成声学准确的响应。我们提出了一个两阶段的生成框架来解决这些挑战。我们的方法结合了一个变分自动编码器,上采样带限RIR全频带质量(48 kHz),和一个条件扩散Transformer模型的基础上整流匹配,产生RIR从自然语言的描述。经验评估表明,与现有方法相比,RIR产生的感知质量和声学精度更高,平均RT 60误差为8.8%,而广泛使用的基线为-37%,并产生更真实的室内声学参数。我们的方法可以在虚拟现实,建筑声学和音频制作中实现实际应用,其中灵活,高质量的RIR合成至关重要。
摘要:Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: the scarcity of full-band RIR datasets and the inability of existing models to generate acoustically accurate responses from diverse input modalities. We present PromptReverb, a two-stage generative framework that addresses these challenges. Our approach combines a variational autoencoder that upsamples band-limited RIRs to full-band quality (48 kHz), and a conditional diffusion transformer model based on rectified flow matching that generates RIRs from descriptions in natural language. Empirical evaluation demonstrates that PromptReverb produces RIRs with superior perceptual quality and acoustic accuracy compared to existing methods, achieving 8.8% mean RT60 error compared to -37% for widely used baselines and yielding more realistic room-acoustic parameters. Our method enables practical applications in virtual reality, architectural acoustics, and audio production where flexible, high-quality RIR synthesis is essential.


【9】FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
标题:FOA令牌器:空间一致性损失的一阶立体声音响的低比特率神经编解码器
链接:https://arxiv.org/abs/2510.22241

作者:Parthasaarathy Sudarsanam, Sebastian Braun, Hannes Gamper
备注:Submitted to ICASSP 2026
摘要:神经音频编解码器已被广泛研究用于单声道和立体声信号,但空间音频仍然在很大程度上未被探索。我们提出了第一个用于一阶立体混响(FOA)的离散神经空间音频编解码器。在WavTokenizer架构的基础上,我们将其扩展为支持四通道FOA信号,并引入了一种新的空间一致性损失,以在高度压缩的表示下保留重构信号中的方向线索。我们的编解码器压缩4通道FOA音频在24 kHz到75离散令牌每秒,对应于0.9 kbps的比特率。对模拟混响混合物、非混响干净语音和具有真实房间脉冲响应的FOA混合物的评估显示出准确的重建,在三种条件下,平均角度误差分别为13.76{\deg}、3.96{\deg}和25.83{\deg}。此外,从我们的编解码器派生的离散潜在表示为下游空间音频任务提供了有用的功能,如STARSS 23真实录音的声音事件定位和检测所示。
摘要:Neural audio codecs have been widely studied for mono and stereo signals, but spatial audio remains largely unexplored. We present the first discrete neural spatial audio codec for first-order ambisonics (FOA). Building on the WavTokenizer architecture, we extend it to support four-channel FOA signals and introduce a novel spatial consistency loss to preserve directional cues in the reconstructed signals under a highly compressed representation. Our codec compresses 4-channel FOA audio at 24 kHz into 75 discrete tokens per second, corresponding to a bit rate of 0.9 kbps. Evaluations on simulated reverberant mixtures, non-reverberant clean speech, and FOA mixtures with real room impulse responses show accurate reconstruction, with mean angular errors of 13.76{\deg}, 3.96{\deg}, and 25.83{\deg}, respectively, across the three conditions. In addition, discrete latent representations derived from our codec provide useful features for downstream spatial audio tasks, as demonstrated on sound event localization and detection with STARSS23 real recordings.


【10】M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
标题:M-inf:基于inf的非自回归ASB的多尺度对齐
链接:https://arxiv.org/abs/2510.22172

作者:Ruixiang Mao, Xiangnan Ma, Qing Yang, Ziming Zhu, Yucheng Qiao, Yuan Ge, Tong Xiao, Shengxiang Gao, Zhengtao Yu, Jingbo Zhu
摘要:连续积分触发(CIF)机制为非自回归(NAR)语音识别提供了有效的对齐。该机制创建了从声学特征到目标标记的平滑和单调的映射,实现了与其他NAR方法相比具有竞争力的普通话性能。然而,如果没有更细粒度的指导,它的稳定性在某些语言(如英语和法语)中会下降。在本文中,我们提出了多尺度CIF(M-CIF),它通过集成字符和音素级别的监督逐步蒸馏到子字表示,从而提高了强大的声学文本对齐进行多级对齐。实验表明,与Paraformer基线相比,M-CIF降低了WER,特别是在德语中的CommonVoice上降低了4.21%,在法语中降低了3.05%。为了进一步研究这些收益,我们定义语音混淆错误(PE)和空间相关的分割错误(SE)作为评估指标。这些指标在不同的M-CIF设置的分析表明,音素和字符层是必不可少的增强渐进CIF对齐。
摘要:The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mapping from acoustic features to target tokens, achieving performance on Mandarin competitive with other NAR approaches. However, without finer-grained guidance, its stability degrades in some languages such as English and French. In this paper, we propose Multi-scale CIF (M-CIF), which performs multi-level alignment by integrating character and phoneme level supervision progressively distilled into subword representations, thereby enhancing robust acoustic-text alignment. Experiments show that M-CIF reduces WER compared to the Paraformer baseline, especially on CommonVoice by 4.21% in German and 3.05% in French. To further investigate these gains, we define phonetic confusion errors (PE) and space-related segmentation errors (SE) as evaluation metrics. Analysis of these metrics across different M-CIF settings reveals that the phoneme and character layers are essential for enhancing progressive CIF alignment.


【11】Streaming Generation for Music Accompaniment
标题:音乐伴奏者的流媒体生成
链接:https://arxiv.org/abs/2510.22105

作者:Yusong Wu, Mason Wang, Heidi Lei, Stephen Brade, Lancelot Blanchard, Shih-Lun Wu, Aaron Courville, Anna Huang
摘要:音乐生成模型可以在给定完整音频输入的情况下生成高保真连贯伴奏,但仅限于编辑和基于循环的工作流程。我们研究实时音频到音频伴奏:当模型听到输入音频流时(例如,歌手唱歌),它还必须同时实时生成相干伴随流(例如,吉他伴奏)。在这项工作中,我们提出了一个模型设计,考虑不可避免的系统延迟在实际部署中有两个设计变量:未来的可见性$t_f$,输出回放时间和最新的输入时间之间的偏移量用于调节,和输出块持续时间$k$,每次调用发出的帧的数量。我们训练Transformer解码器跨越网格的$(t_f,k)$,并显示两个一致的权衡:增加有效的$t_f$通过减少新近间隔来提高一致性,但需要更快的推理以保持在延迟预算内;增加$k$提高吞吐量,但由于更新速率降低而导致伴奏降级。最后,我们观察到,天真的最大似然流训练是不够的连贯伴奏,未来的背景是不可用的,激励先进的预期和代理的目标现场干扰。
摘要:Music generation models can produce high-fidelity coherent accompaniment given complete audio input, but are limited to editing and loop-based workflows. We study real-time audio-to-audio accompaniment: as a model hears an input audio stream (e.g., a singer singing), it has to also simultaneously generate in real-time a coherent accompanying stream (e.g., a guitar accompaniment). In this work, we propose a model design considering inevitable system delays in practical deployment with two design variables: future visibility $t_f$, the offset between the output playback time and the latest input time used for conditioning, and output chunk duration $k$, the number of frames emitted per call. We train Transformer decoders across a grid of $(t_f,k)$ and show two consistent trade-offs: increasing effective $t_f$ improves coherence by reducing the recency gap, but requires faster inference to stay within the latency budget; increasing $k$ improves throughput but results in degraded accompaniment due to a reduced update rate. Finally, we observe that naive maximum-likelihood streaming training is insufficient for coherent accompaniment where future context is not available, motivating advanced anticipatory and agentic objectives for live jamming.


【12】GuitarFlow: Realistic Electric Guitar Synthesis From Tablatures via Flow Matching and Style Transfer
标题:GuitarFlow:通过流量匹配和风格转移从表格中进行逼真的电吉他合成
链接:https://arxiv.org/abs/2510.21872

作者:Jackson Loth, Pedro Sarmento, Mark Sandler, Mathieu Barthet
备注:To be published in Proceedings of the 17th International Symposium on Computer Music and Multidisciplinary Research (CMMR)
摘要:近年来,使用人工智能(AI)在音频领域生成音乐取得了稳步进展。然而,对于某些乐器,特别是吉他,可控乐器合成在表现力方面仍然有限。我们介绍GuitarFlow,一个专门为电吉他合成设计的模型。生成过程是使用表格,一个无处不在的和直观的吉他特定的符号格式指导。指谱格式很容易表现出吉他特有的演奏技巧(例如压音、静音弦和连奏),而这些技巧在其他常见的音乐符号格式(如连奏)中则更难表现。我们的模型依赖于一个中间步骤,首先使用一个简单的基于样本的虚拟乐器将指谱渲染为音频,然后使用Flow Matching执行风格转换,以便将虚拟乐器音频转换为更逼真的声音示例。这使得模型能够快速进行训练并执行推理,所需的训练数据少于6小时。我们提供了客观评估指标的结果以及听力测试,其中我们显示出从图表生成的吉他音频的真实感有显着提高。
摘要:Music generation in the audio domain using artificial intelligence (AI) has witnessed steady progress in recent years. However for some instruments, particularly the guitar, controllable instrument synthesis remains limited in expressivity. We introduce GuitarFlow, a model designed specifically for electric guitar synthesis. The generative process is guided using tablatures, an ubiquitous and intuitive guitar-specific symbolic format. The tablature format easily represents guitar-specific playing techniques (e.g. bends, muted strings and legatos), which are more difficult to represent in other common music notation formats such as MIDI. Our model relies on an intermediary step of first rendering the tablature to audio using a simple sample-based virtual instrument, then performing style transfer using Flow Matching in order to transform the virtual instrument audio into more realistic sounding examples. This results in a model that is quick to train and to perform inference, requiring less than 6 hours of training data. We present the results of objective evaluation metrics, together with a listening test, in which we show significant improvement in the realism of the generated guitar audio from tablatures.


【13】Quantifying Multimodal Imbalance: A GMM-Guided Adaptive Loss for Audio-Visual Learning
标题:量化多模式失衡:GMM引导的视听学习适应性损失
链接:https://arxiv.org/abs/2510.21797

作者:Zhaocheng Liu, Zhiwen Yu, Xiaoqing Liu
摘要:目前解决多模态不平衡的主流方法主要集中在基于架构的修改和优化上,往往忽略了对模态间不平衡程度的定量分析。为了解决这一差距,我们的工作引入了一种新的方法来定量分析多模态不平衡,这反过来又为样本级自适应损失函数的设计提供了信息。我们首先将“模态差距”定义为不同模态的Softmax分数之间的差异(例如,音频和视频)用于地面实况类预测。模态间隙分布的分析表明,它可以有效地建模的双峰高斯混合模型(GMM)。这两个组成部分被发现分别对应于“模态平衡”和“模态不平衡”的数据样本。随后,我们应用贝叶斯定理计算每个样本属于这两个不同分布的后验概率,并在此基础上设计了一个新的自适应损失函数,该函数具有三个目标:(1)最小化总体模态间隙;(2)鼓励不平衡样本分布向平衡分布移动;(3)对不平衡样本施加更大的惩罚权重。实验结果表明,该方法在CREMA-D和AVE数据集上分别达到了SOTA(state-of-the-art)性能,准确率分别为80.65美元和70.90美元。这验证了我们提出的方法的有效性。
摘要:Current mainstream approaches to addressing multimodal imbalance primarily focus on architectural modifications and optimization-based, often overlooking a quantitative analysis of the imbalance degree between modalities. To address this gap, our work introduces a novel method for the quantitative analysis of multi-modal imbalance, which in turn informs the design of a sample-level adaptive loss function.We begin by defining the "Modality Gap" as the difference between the Softmax scores of different modalities (e.g., audio and visual) for the ground-truth class prediction. Analysis of the Modality Gap distribution reveals that it can be effectively modeled by a bimodal Gaussian Mixture Model (GMM). These two components are found to correspond respectively to "modality-balanced" and "modality-imbalanced" data samples. Subsequently, we apply Bayes' theorem to compute the posterior probability of each sample belonging to these two distinct distributions.Informed by this quantitative analysis, we design a novel adaptive loss function with three objectives: (1) to minimize the overall Modality Gap; (2) to encourage the imbalanced sample distribution to shift towards the balanced one; and (3) to apply greater penalty weights to imbalanced samples. We employ a two-stage training strategy consisting of a warm-up phase followed by an adaptive training phase.Experimental results demonstrate that our approach achieves state-of-the-art (SOTA) performance on the public CREMA-D and AVE datasets, attaining accuracies of $80.65\%$ and $70.90\%$, respectively. This validates the effectiveness of our proposed methodology.


【14】SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
标题:SoulX-播客:迈向具有方言和副语言多样性的现实长篇播客
链接:https://arxiv.org/abs/2510.23541

作者:Hanke Xie, Haopeng Lin, Wenxiao Cao, Dake Guo, Wenjie Tian, Jun Wu, Hanlin Wen, Ruixuan Shang, Hongmei Liu, Zhiqi Jiang, Yuepeng Jiang, Wenxi Chen, Ruiqi Yan, Jiale Qian, Yichao Yan, Shunshun Yin, Ming Tao, Xie Chen, Lei Xie, Xinsheng Wang
摘要:文语转换(TTS)合成的最新进展显着提高了语音的表现力和自然度。然而,大多数现有的系统是为单扬声器合成量身定制的,并且在生成连贯的多扬声器会话语音方面存在不足。本技术报告介绍了SoulX-Podcast,这是一个专为播客风格的多回合,多扬声器对话语音生成而设计的系统,同时也在传统的TTS任务中实现了最先进的性能。   为了满足多轮口语对话更高的自然度要求,SoulX-Podcast集成了一系列的语言控制,支持普通话和英语,以及几种中国方言,包括四川话,河南话和广东话,从而实现更个性化的播客风格的语音生成。实验结果表明,SoulX-Podcast可以连续产生超过90分钟的对话,具有稳定的扬声器音色和平滑的扬声器过渡。此外,说话者表现出语境自适应韵律,反映自然的节奏和语调变化的对话进展。在多个评估指标中,SoulX-Podcast在独白TTS和多轮对话语音合成方面都达到了最先进的性能。
摘要:Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent multi-speaker conversational speech. This technical report presents SoulX-Podcast, a system designed for podcast-style multi-turn, multi-speaker dialogic speech generation, while also achieving state-of-the-art performance in conventional TTS tasks.   To meet the higher naturalness demands of multi-turn spoken dialogue, SoulX-Podcast integrates a range of paralinguistic controls and supports both Mandarin and English, as well as several Chinese dialects, including Sichuanese, Henanese, and Cantonese, enabling more personalized podcast-style speech generation. Experimental results demonstrate that SoulX-Podcast can continuously produce over 90 minutes of conversation with stable speaker timbre and smooth speaker transitions. Moreover, speakers exhibit contextually adaptive prosody, reflecting natural rhythm and intonation changes as dialogues progress. Across multiple evaluation metrics, SoulX-Podcast achieves state-of-the-art performance in both monologue TTS and multi-turn conversational speech synthesis.


【15】LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization
标题:LibriConvo:模拟阅读文学中的对话以进行ZR和Dialization
链接:https://arxiv.org/abs/2510.23320

作者:Máté Gedeon, Péter Mihajlik
备注:Submitted to LREC 2026
摘要:我们介绍LibriConvo,一个基于说话者感知会话模拟(SASC)的模拟多说话者会话数据集,旨在支持说话者日志化和自动语音识别(ASR)系统的训练和评估。与以前的资源主要依赖于语义上不连贯的话语和令人难以置信的时间间隙不同,LibriConvo确保了语义的连贯性和现实的会话时间。我们的管道利用CallHome与外部VAD进行可靠的边界,应用压缩以减少不自然的长时间沉默,并按书籍组织LibriTTS话语以保持上下文一致性。通过一种新的房间脉冲响应选择程序,增强了声学现实主义,该程序通过空间可扩展性、平衡现实主义和多样性来排列扬声器麦克风配置。该数据集包括240.1小时的1,496个对话,830个独特的发言者,以发言者不相交的方式进行分割,以进行稳健的评估。基线显示,sortformer模型在日志化方面优于pyannote管道,而经过微调的Fast Conformer-CTC XLarge与Serialized Output Training在ASR方面达到了7.29\%的WER,超过了zero-shot Whisper-large-v3。LibriConvo为推进多说话者语音处理研究提供了宝贵的资源,具有逼真的会话动态和受控的实验条件。
摘要:We introduce LibriConvo, a simulated multi-speaker conversational dataset based on speaker-aware conversation simulation (SASC), designed to support training and evaluation of speaker diarization and automatic speech recognition (ASR) systems. Unlike prior resources that mostly rely on semantically disconnected utterances and implausible temporal gaps, LibriConvo ensures semantic coherence and realistic conversational timing. Our pipeline leverages CallHome with external VAD for reliable boundaries, applies compression to reduce unnaturally long silences, and organizes LibriTTS utterances by book to maintain contextual consistency. Acoustic realism is enhanced via a novel room impulse response selection procedure that ranks speaker-microphone configurations by spatial plausibility, balancing realism and diversity. The dataset comprises 240.1 hours across 1,496 dialogues with 830 unique speakers, split in a speaker-disjoint manner for robust evaluation. Baselines show that the sortformer model outperforms the pyannote pipeline in diarization, while a fine-tuned Fast Conformer-CTC XLarge with Serialized Output Training achieves 7.29\% WER for ASR, surpassing zero-shot Whisper-large-v3. LibriConvo provides a valuable resource for advancing multi-speaker speech processing research with realistic conversational dynamics and controlled experimental conditions.


【16】Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMS
标题:使用LLMS缓解视听语音识别中的注意力下降和大量激活
链接:https://arxiv.org/abs/2510.22603

作者:Anand, Umberto Cappellazzo, Stavros Petridis, Maja Pantic
备注:The code is available at this https URL
摘要:大语言模型(LLM)最近已经发展了听觉语音识别(ASR),视觉语音识别(VSR)和视听语音识别(AVSR)。然而,对其微调下的内部动态的理解仍然有限。在自然语言处理中,最近的工作揭示了注意力下沉,吸引不成比例的高注意力的令牌,以及相关的大规模激活,其中下沉令牌的一些特征在LLM中表现出巨大的激活。在这项工作中,我们是第一个研究这些现象在多模态语音识别。通过对视听LLM的详细分析,我们不仅在BOS令牌上,而且在ASR、VSR和AVSR的中间低语义令牌上识别出注意力汇和大规模激活。我们发现,大规模的激活起源于MLP层,并对应于所有汇令牌的固定特征索引。我们进一步表明,中间水槽令牌表现出高余弦相似性的BOS令牌,从而放大注意力和激活。基于这些见解,我们引入了一个简单的去相关损失,降低了BOS和其他令牌之间的余弦相似性,有效地减少了中间汇和大规模激活。此外,我们的方法在高视听特征下采样下提高了字错误率(WER),同时在较低的下采样率下保持稳定。
摘要:Large language models (LLMs) have recently advanced auditory speech recognition (ASR), visual speech recognition (VSR), and audio-visual speech recognition (AVSR). However, understanding of their internal dynamics under fine-tuning remains limited. In natural language processing, recent work has revealed attention sinks, tokens that attract disproportionately high attention, and associated massive activations in which some features of sink tokens exhibit huge activation in LLMs. In this work, we are the first to study these phenomena in multimodal speech recognition. Through a detailed analysis of audio-visual LLMs, we identify attention sinks and massive activations not only at the BOS token but also at intermediate low-semantic tokens across ASR, VSR, and AVSR. We show that massive activations originate in the MLP layers and correspond to fixed feature indices across all sink tokens. We further show that intermediate sink tokens exhibit high cosine similarity to the BOS token, thereby amplifying attention and activation. Building on these insights, we introduce a simple decorrelation loss that reduces cosine similarity between BOS and other tokens, effectively mitigating intermediate sinks and massive activations. Furthermore, our method improves word error rate (WER) under high audio-visual feature downsampling while remaining stable at lower downsampling rates.


【17】DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
标题:DialoSpeech:使用LLM和流匹配的双扬声器对话生成
链接:https://arxiv.org/abs/2510.08373

作者:Hanke Xie, Dake Guo, Chengyou Wang, Yue Li, Wenjie Tian, Xinfa Zhu, Xinsheng Wang, Xiulin Li, Guanqiong Miao, Bo Liu, Lei Xie
摘要:文本到语音(TTS)合成的最新进展,特别是那些利用大型语言模型(LLM),显着提高了表现力和自然度。然而,生成类似人类的交互式对话语音仍然具有挑战性。目前的系统面临的限制,由于缺乏双轨数据和困难,在实现自然,上下文连贯性,和intermingdynamics,如话轮转换,重叠的语音,和扬声器的一致性,在多轮对话。为了解决这些挑战,我们提出了DialoSpeech,一个双轨架构,结合了一个大的语言模型与分块流匹配的表达,人性化的对话语音合成。DialoSpeech生成自然的多话轮对话,具有连贯的说话人话轮和自然重叠,支持中文和英文以及跨语言语音合成。我们引入了一个数据处理管道来构建双轨对话数据集,促进可扩展的训练和实验验证。实验表明,我们的模型优于基线,为生成类似人类的口语对话提供了解决方案。音频样本可在https://tiamojames.github.io/DialoSpeech上获得
摘要:Recent advances in text-to-speech (TTS) synthesis, particularly those leveraging large language models (LLMs), have significantly improved expressiveness and naturalness. However, generating human-like, interactive dialogue speech remains challenging. Current systems face limitations due to the scarcity of dual-track data and difficulties in achieving naturalness, contextual coherence, and interactional dynamics, such as turn-taking, overlapping speech, and speaker consistency, in multi-turn conversations. To address these challenges, we propose DialoSpeech, a dual-track architecture combining a large language model with Chunked Flow Matching for expressive, human-like dialogue speech synthesis. DialoSpeech generates natural multi-turn conversations with coherent speaker turns and natural overlaps, supporting both Chinese and English and cross-lingual speech synthesis. We introduce a data processing pipeline to construct dual-track dialogue datasets, facilitating scalable training and experimental validation. Experiments show that our model outperforms baselines, offering a solution for generating human-like spoken dialogues. Audio samples are available at https://tiamojames.github.io/DialoSpeech


eess.AS音频处理


【1】SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
标题:SoulX-播客:迈向具有方言和副语言多样性的现实长篇播客
链接:https://arxiv.org/abs/2510.23541

作者:Hanke Xie, Haopeng Lin, Wenxiao Cao, Dake Guo, Wenjie Tian, Jun Wu, Hanlin Wen, Ruixuan Shang, Hongmei Liu, Zhiqi Jiang, Yuepeng Jiang, Wenxi Chen, Ruiqi Yan, Jiale Qian, Yichao Yan, Shunshun Yin, Ming Tao, Xie Chen, Lei Xie, Xinsheng Wang
摘要:文语转换(TTS)合成的最新进展显着提高了语音的表现力和自然度。然而,大多数现有的系统是为单扬声器合成量身定制的,并且在生成连贯的多扬声器会话语音方面存在不足。本技术报告介绍了SoulX-Podcast,这是一个专为播客风格的多回合,多扬声器对话语音生成而设计的系统,同时也在传统的TTS任务中实现了最先进的性能。   为了满足多轮口语对话更高的自然度要求,SoulX-Podcast集成了一系列的语言控制,支持普通话和英语,以及几种中国方言,包括四川话,河南话和广东话,从而实现更个性化的播客风格的语音生成。实验结果表明,SoulX-Podcast可以连续产生超过90分钟的对话,具有稳定的扬声器音色和平滑的扬声器过渡。此外,说话者表现出语境自适应韵律,反映自然的节奏和语调变化的对话进展。在多个评估指标中,SoulX-Podcast在独白TTS和多轮对话语音合成方面都达到了最先进的性能。
摘要:Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent multi-speaker conversational speech. This technical report presents SoulX-Podcast, a system designed for podcast-style multi-turn, multi-speaker dialogic speech generation, while also achieving state-of-the-art performance in conventional TTS tasks.   To meet the higher naturalness demands of multi-turn spoken dialogue, SoulX-Podcast integrates a range of paralinguistic controls and supports both Mandarin and English, as well as several Chinese dialects, including Sichuanese, Henanese, and Cantonese, enabling more personalized podcast-style speech generation. Experimental results demonstrate that SoulX-Podcast can continuously produce over 90 minutes of conversation with stable speaker timbre and smooth speaker transitions. Moreover, speakers exhibit contextually adaptive prosody, reflecting natural rhythm and intonation changes as dialogues progress. Across multiple evaluation metrics, SoulX-Podcast achieves state-of-the-art performance in both monologue TTS and multi-turn conversational speech synthesis.


【2】Evaluation of Spherical Wavelet Framework in Comparsion with Ambisonics
标题:球形子波框架与立体声模拟的比较评价
链接:https://arxiv.org/abs/2510.23403

作者:Ş. Ekmen, H. Lee
备注:13 pages, 8 figures. Submitted to IEEE TASLP
摘要:最近,提出了球面小波框架(Spherical Wavelet Framework,简称SPW),通过利用高度局部化的基函数来结合高保真度立体声和基于对象的音频(Object-Based Audio,简称OBA)的优点。降噪可以增强甜蜜点区域并减少定位模糊,同时仍然能够实现完整声场的稀疏表示,使存储和传输更有效。初步的矢量分析和听力测试的ESTA已经显示出可喜的结果,但是,这些发现仅限于非常具体的条件,不包括感知指标。本研究更详细地研究了高保真度立体声,并将其与高保真度立体声进行了比较。使用IACC,ITD和ILD估计,以及生态上有效的声源的听力测试进行比较。各种复制布局:评估了规则多面体、t-设计和Lebedev网格及其相应的高保真度立体声响复制阶数和通道计数。结果表明,在整体空间和音色保真度方面,与Ambisonics相比,它与参考更相似;然而,它在很大程度上取决于球体的细分。此外,它不能原生地表示以连续方向到达的波。提出了可能的解决办法。
摘要:Recently, the Spherical Wavelet Framework (SWF) was proposed to combine the benefits of Ambisonics and Object-Based Audio (OBA) by utilising highly localised basis functions. SWF can enhance the sweet-spot area and reduce localisation blur while still enabling a sparse representation of the complete sound field, making storage and transmission more efficient. Initial vector analysis and listening test of SWF have shown promising results; however, these findings are limited to very specific conditions and do not include perceptual metrics. The present study investigates SWF in greater detail, comparing it with Ambisonics. The comparison was carried out using IACC, ITD, and ILD estimations, as well as listening tests with ecologically valid sound sources. Various reproduction layouts: regular polyhedron, t-design, and Lebedev grid with their corresponding Ambisonics orders and channel counts were evaluated. Results indicate that SWF is rated significantly more similar to the reference than Ambisonics is, in terms of overall spatial and timbral fidelity; however, it is considerably dependent on the subdivison of the sphere. Moreover, it cannot natively represent a wave arriving at a continuous direction. Possible solutions are proposed.


【3】LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization
标题:LibriConvo:模拟阅读文学中的对话以进行ZR和Dialization
链接:https://arxiv.org/abs/2510.23320

作者:Máté Gedeon, Péter Mihajlik
备注:Submitted to LREC 2026
摘要:我们介绍LibriConvo,一个基于说话者感知会话模拟(SASC)的模拟多说话者会话数据集,旨在支持说话者日志化和自动语音识别(ASR)系统的训练和评估。与以前的资源主要依赖于语义上不连贯的话语和令人难以置信的时间间隙不同,LibriConvo确保了语义的连贯性和现实的会话时间。我们的管道利用CallHome与外部VAD进行可靠的边界,应用压缩以减少不自然的长时间沉默,并按书籍组织LibriTTS话语以保持上下文一致性。通过一种新的房间脉冲响应选择程序,增强了声学现实主义,该程序通过空间可扩展性、平衡现实主义和多样性来排列扬声器麦克风配置。该数据集包括240.1小时的1,496个对话,830个独特的发言者,以发言者不相交的方式进行分割,以进行稳健的评估。基线显示,sortformer模型在日志化方面优于pyannote管道,而经过微调的Fast Conformer-CTC XLarge与Serialized Output Training在ASR方面达到了7.29\%的WER,超过了zero-shot Whisper-large-v3。LibriConvo为推进多说话者语音处理研究提供了宝贵的资源,具有逼真的会话动态和受控的实验条件。
摘要:We introduce LibriConvo, a simulated multi-speaker conversational dataset based on speaker-aware conversation simulation (SASC), designed to support training and evaluation of speaker diarization and automatic speech recognition (ASR) systems. Unlike prior resources that mostly rely on semantically disconnected utterances and implausible temporal gaps, LibriConvo ensures semantic coherence and realistic conversational timing. Our pipeline leverages CallHome with external VAD for reliable boundaries, applies compression to reduce unnaturally long silences, and organizes LibriTTS utterances by book to maintain contextual consistency. Acoustic realism is enhanced via a novel room impulse response selection procedure that ranks speaker-microphone configurations by spatial plausibility, balancing realism and diversity. The dataset comprises 240.1 hours across 1,496 dialogues with 830 unique speakers, split in a speaker-disjoint manner for robust evaluation. Baselines show that the sortformer model outperforms the pyannote pipeline in diarization, while a fine-tuned Fast Conformer-CTC XLarge with Serialized Output Training achieves 7.29\% WER for ASR, surpassing zero-shot Whisper-large-v3. LibriConvo provides a valuable resource for advancing multi-speaker speech processing research with realistic conversational dynamics and controlled experimental conditions.


【4】Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
标题:通过习得声学嵌入和反馈延迟网络匹配回响语音
链接:https://arxiv.org/abs/2510.23158

作者:Philipp Götz, Gloria Dal Santo, Sebastian J. Schlecht, Vesa Välimäki, Emanuël A.P. Habets
备注:Submitted to ICASSP 2026
摘要:混响传达关于环境的关键声学线索,支持空间意识和沉浸感。对于听觉增强现实(AAR)系统,实时生成感知上合理的混响仍然是一个关键的挑战,特别是当显式声学测量不可用时。我们通过将人工混响参数的盲估计制定为混响信号匹配任务来解决这个问题,利用学习的室内声学先验。此外,我们提出了一个反馈延迟网络(FDN)的结构,再现频率相关的衰减时间和直接混响比的目标空间。对领先的自动FDN调谐方法的实验评估表明,在估计的室内声学参数和人工混响语音的感知可扩展性的改善。这些结果突出了我们的方法在AAR应用中高效,感知一致的混响渲染的潜力。
摘要:Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key challenge, especially when explicit acoustic measurements are unavailable. We address this by formulating blind estimation of artificial reverberation parameters as a reverberant signal matching task, leveraging a learned room-acoustic prior. Furthermore, we propose a feedback delay network (FDN) structure that reproduces both frequency-dependent decay times and the direct-to-reverberation ratio of a target space. Experimental evaluation against a leading automatic FDN tuning method demonstrates improvements in estimated room-acoustic parameters and perceptual plausibility of artificial reverberant speech. These results highlight the potential of our approach for efficient, perceptually consistent reverberation rendering in AAR applications.


【5】Treble10: A high-quality dataset for far-field speech recognition, dereverberation, and enhancement
标题:Treble10:用于远场语音识别、去回响和增强的高质量数据集
链接:https://arxiv.org/abs/2510.23141

作者:Sarabeth S. Mullins, Georg Götz, Eric Bezzam, Steven Zheng, Daniel Gert Nielsen
摘要:准确的远场语音数据集对于自动语音识别(ASR)、去混响、语音增强和源分离等任务至关重要。然而,目前的数据集是有限的声学现实主义和可扩展性之间的权衡。测量语料库提供了真实的物理,但昂贵,低覆盖率,很少包括配对的干净和混响数据。相比之下,大多数基于模拟的数据集依赖于简化的几何声学,因此无法再现控制复杂环境中声音传播的衍射、散射和干涉等关键物理现象。我们介绍Treble10,一个大规模的,物理上精确的房间声学数据集。Treble10包含超过3000个宽带房间脉冲响应(RIR),在10个家具齐全的真实世界的房间中模拟,使用Treble SDK中实现的混合模拟范例,结合了基于波的几何声学求解器。该数据集提供了六个互补子集,包括单声道、8阶高保真度立体声和6声道设备RIR,以及与LibriSpeech话语配对的预卷积混响语音场景。所有信号均在32 kHz下进行模拟,准确模拟低频波效应和高频反射。Treble10弥合了测量和模拟之间的现实差距,为远场语音任务提供了可重现的、物理接地的评估和大规模数据增强。该数据集通过Hugging Face Hub公开提供,旨在作为下一代模拟驱动音频研究的基准和模板。
摘要:Accurate far-field speech datasets are critical for tasks such as automatic speech recognition (ASR), dereverberation, speech enhancement, and source separation. However, current datasets are limited by the trade-off between acoustic realism and scalability. Measured corpora provide faithful physics but are expensive, low-coverage, and rarely include paired clean and reverberant data. In contrast, most simulation-based datasets rely on simplified geometrical acoustics, thus failing to reproduce key physical phenomena like diffraction, scattering, and interference that govern sound propagation in complex environments. We introduce Treble10, a large-scale, physically accurate room-acoustic dataset. Treble10 contains over 3000 broadband room impulse responses (RIRs) simulated in 10 fully furnished real-world rooms, using a hybrid simulation paradigm implemented in the Treble SDK that combines a wave-based and geometrical acoustics solver. The dataset provides six complementary subsets, spanning mono, 8th-order Ambisonics, and 6-channel device RIRs, as well as pre-convolved reverberant speech scenes paired with LibriSpeech utterances. All signals are simulated at 32 kHz, accurately modelling low-frequency wave effects and high-frequency reflections. Treble10 bridges the realism gap between measurement and simulation, enabling reproducible, physically grounded evaluation and large-scale data augmentation for far-field speech tasks. The dataset is openly available via the Hugging Face Hub, and is intended as both a benchmark and a template for next-generation simulation-driven audio research.


【6】Adapting Speech Foundation Models with Large Language Models for Unified Speech Recognition
标题:将语音基础模型与大型语言模型相适应以实现统一语音识别
链接:https://arxiv.org/abs/2510.22961

作者:Jing-Xuan Zhang, Genshun Wan, Jin Li, Jianqing Gao
备注:submitted to Pattern Recognition
摘要:统一语音识别的目标是在一个单一的模型框架内执行听觉,视觉和视听语音识别。虽然语音基础模型(SFM)在听觉任务中表现出显着的性能,其适应多模态的情况下仍然未充分探索。本文提出了UASR-LLM,一种新的框架,通过利用大型语言模型(LLM)作为文本解码器,将冻结的SFM适应于统一的VSR,ASR和AVSR任务。我们的方法通过视觉注入模块将视觉表示引入到多个SFM层中,从而实现多模态输入处理和统一的隐藏表示。增强的SIM通过前馈适配器与仅解码器的LLM连接,其中级联表示和指令提示引导语音转录。我们实现了一个两阶段的训练策略:视觉注入预训练,然后是语音识别微调。SFM参数在整个训练过程中保持冻结,最初只优化了视觉注入模块,随后使用LoRA参数对LLM进行微调。实验结果表明,在干净和嘈杂的条件下,在VSR,ASR和AVSR任务的最先进的基线优越的性能。消融研究证实了各种SFM和LLM的泛化,验证了所提出的训练策略。
摘要:Unified speech recognition aims to perform auditory, visual, and audiovisual speech recognition within a single model framework. While speech foundation models (SFMs) have demonstrated remarkable performance in auditory tasks, their adaptation to multimodal scenarios remains underexplored. This paper presents UASR-LLM, a novel framework that adapts frozen SFMs to unified VSR, ASR, and AVSR tasks by leveraging large language models (LLMs) as text decoders. Our approach introduces visual representations into multiple SFM layers through visual injection modules, enabling multimodal input processing and unified hidden representations. The augmented SFMs connect with decoder-only LLMs via a feed-forward adaptor, where concatenated representations and instruction prompts guide speech transcription. We implement a twostage training strategy: visual injection pretraining followed by speech recognition finetuning. SFM parameters remain frozen throughout training, with only visual injection modules optimized initially, and LLMs finetuned using LoRA parameters subsequently. Experimental results demonstrate superior performance over state-of-the-art baselines across VSR, ASR, and AVSR tasks under both clean and noisy conditions. Ablation studies confirm generalization across various SFMs and LLMs, validating the proposed training strategy.


【7】DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching
标题:迪夫节奏2:通过块流匹配高效、高保真的歌曲生成
链接:https://arxiv.org/abs/2510.22950

作者:Yuepeng Jiang, Huakang Chen, Ziqian Ning, Jixun Yao, Zerui Han, Di Wu, Meng Meng, Jian Luan, Zhonghua Fu, Lei Xie
摘要:制作完整长度的高质量歌曲是一项挑战,因为它需要在文本和音乐模态之间以及在音乐模态本身内部保持长期的一致性。现有的非自回归(NAR)框架虽然能够产生高质量的歌曲,但经常难以实现歌词和人声之间的一致性。同时,迎合不同的音乐偏好需要从人的反馈中强化学习(RLHF)。然而,现有的方法往往依赖于合并多个模型在多偏好优化,这导致显着的性能下降。为了解决这些挑战,我们引入了DiffRhythm 2,这是一个端到端的框架,专为高保真,可控的歌曲生成而设计。为了解决歌词对齐问题,DiffRhythm 2采用了基于块流匹配的半自回归架构。这种设计能够使歌词与演唱的声音忠实一致,而不依赖于外部标签和约束,同时保持NAR模型的高生成质量和效率。为了使该框架在计算上易于处理长序列,我们实现了一个音乐变分自动编码器(VAE),该编码器实现了5 Hz的低帧率,同时仍然能够实现高保真音频重建。另外,为了克服多偏好优化在RLHF中的局限性,提出了交叉对偏好优化算法。这种方法有效地缓解了通常与模型合并相关的性能下降,允许在不同的人类偏好中进行更鲁棒的优化。通过引入随机块表示对齐丢失,我们进一步增强了音乐性和结构连贯性。
摘要:Generating full-length, high-quality songs is challenging, as it requires maintaining long-term coherence both across text and music modalities and within the music modality itself. Existing non-autoregressive (NAR) frameworks, while capable of producing high-quality songs, often struggle with the alignment between lyrics and vocal. Concurrently, catering to diverse musical preferences necessitates reinforcement learning from human feedback (RLHF). However, existing methods often rely on merging multiple models during multi-preference optimization, which results in significant performance degradation. To address these challenges, we introduce DiffRhythm 2, an end-to-end framework designed for high-fidelity, controllable song generation. To tackle the lyric alignment problem, DiffRhythm 2 employs a semi-autoregressive architecture based on block flow matching. This design enables faithful alignment of lyrics to singing vocals without relying on external labels and constraints, all while preserving the high generation quality and efficiency of NAR models. To make this framework computationally tractable for long sequences, we implement a music variational autoencoder (VAE) that achieves a low frame rate of 5 Hz while still enabling high-fidelity audio reconstruction. In addition, to overcome the limitations of multi-preference optimization in RLHF, we propose cross-pair preference optimization. This method effectively mitigates the performance drop typically associated with model merging, allowing for more robust optimization across diverse human preferences. We further enhance musicality and structural coherence by introducing stochastic block representation alignment loss.


【8】SRP-PHAT-NET: A Reliability-Driven DNN for Reverberant Speaker Localization
标题:SRP-PHAT-NET:一个可靠性驱动的DNN,用于回响者本地化
链接:https://arxiv.org/abs/2510.22682

作者:Bar Shaybet, Vladimir Tourbabin, Boaz Rafaely
备注:In submission process to the IEEE Transactions on Audio, Speech and Language Processing, 2025
摘要:混响环境下的准确波达方向(DOA)估计仍然是空间音频应用的一个基本挑战。虽然深度学习方法在这种情况下表现出了强大的性能,但它们通常缺乏评估其预测可靠性的机制-这是现实世界部署的基本功能。在这项工作中,我们提出了SRP-PHAT-NET,这是一个深度神经网络框架,它利用SRP-PHAT方向图作为空间特征,并引入了内置的可靠性估计。为了实现有意义的可靠性评分,模型使用以真实方向为中心的高斯加权标签进行训练。我们系统地分析了标签平滑对准确性和可靠性的影响,证明了高斯核宽度的选择可以根据特定应用的要求进行调整。实验结果表明,选择性地使用高置信度预测可以显著提高定位精度,突出了将可靠性集成到基于深度学习的DOA估计中的实际好处。
摘要:Accurate Direction-of-Arrival (DOA) estimation in reverberant environments remains a fundamental challenge for spatial audio applications. While deep learning methods have shown strong performance in such conditions, they typically lack a mechanism to assess the reliability of their predictions - an essential feature for real-world deployment. In this work, we present the SRP-PHAT-NET, a deep neural network framework that leverages SRP-PHAT directional maps as spatial features and introduces a built-in reliability estimation. To enable meaningful reliability scoring, the model is trained using Gaussian-weighted labels centered around the true direction. We systematically analyze the influence of label smoothing on accuracy and reliability, demonstrating that the choice of Gaussian kernel width can be tuned to application-specific requirements. Experimental results show that selectively using high-confidence predictions yields significantly improved localization accuracy, highlighting the practical benefits of integrating reliability into deep learning-based DOA estimation.


【9】HyBeam: Hybrid Microphone-Beamforming Array-Agnostic Speech Enhancement for Wearables
标题:HyBeam:可穿戴设备的混合麦克风-射束成形阵列-不可知语音增强
链接:https://arxiv.org/abs/2510.22637

作者:Yuval Bar Ilan (1), Boaz Rafaely (1), Vladimir Tourbabin (2) ((1) School of Electrical and Computer Engineering, Ben-Gurion University of the Negev, Beer-Sheva, Israel (2) Reality Labs Research, Meta, Redmond, WA, USA)
摘要:语音增强是信号处理中的一个基本挑战,特别是在各种声学条件和麦克风设置需要鲁棒性时。深度学习方法在语音增强方面取得了成功,但通常假设固定的阵列几何形状,限制了它们在移动,嵌入式和可穿戴设备中的使用。现有的阵列不可知方法通常依赖于原始麦克风信号或波束形成器输出,但两者在变化的几何形状下都有缺点。我们介绍HyBeam,这是一个混合框架,它使用低频的原始麦克风信号和高频的波束形成器信号,利用它们的互补优势,同时保持高度的阵列不可知性。不同房间和可穿戴阵列配置的模拟表明,HyBeam在PESQ、STOI和SI-SDR方面始终超过仅麦克风和仅波束形成器的基线。频带分析表明,混合方法利用波束形成器的方向性在高频和麦克风线索在低频,优于任何一种方法单独在所有频段。
摘要:Speech enhancement is a fundamental challenge in signal processing, particularly when robustness is required across diverse acoustic conditions and microphone setups. Deep learning methods have been successful for speech enhancement, but often assume fixed array geometries, limiting their use in mobile, embedded, and wearable devices. Existing array-agnostic approaches typically rely on either raw microphone signals or beamformer outputs, but both have drawbacks under changing geometries. We introduce HyBeam, a hybrid framework that uses raw microphone signals at low frequencies and beamformer signals at higher frequencies, exploiting their complementary strengths while remaining highly array-agnostic. Simulations across diverse rooms and wearable array configurations demonstrate that HyBeam consistently surpasses microphone-only and beamformer-only baselines in PESQ, STOI, and SI-SDR. A bandwise analysis shows that the hybrid approach leverages beamformer directivity at high frequencies and microphone cues at low frequencies, outperforming either method alone across all bands.


【10】Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMS
标题:使用LLMS缓解视听语音识别中的注意力下降和大量激活
链接:https://arxiv.org/abs/2510.22603

作者:Anand, Umberto Cappellazzo, Stavros Petridis, Maja Pantic
备注:The code is available at this https URL
摘要:大语言模型(LLM)最近已经发展了听觉语音识别(ASR),视觉语音识别(VSR)和视听语音识别(AVSR)。然而,对其微调下的内部动态的理解仍然有限。在自然语言处理中,最近的工作揭示了注意力下沉,吸引不成比例的高注意力的令牌,以及相关的大规模激活,其中下沉令牌的一些特征在LLM中表现出巨大的激活。在这项工作中,我们是第一个研究这些现象在多模态语音识别。通过对视听LLM的详细分析,我们不仅在BOS令牌上,而且在ASR、VSR和AVSR的中间低语义令牌上识别出注意力汇和大规模激活。我们发现,大规模的激活起源于MLP层,并对应于所有汇令牌的固定特征索引。我们进一步表明,中间水槽令牌表现出高余弦相似性的BOS令牌,从而放大注意力和激活。基于这些见解,我们引入了一个简单的去相关损失,降低了BOS和其他令牌之间的余弦相似性,有效地减少了中间汇和大规模激活。此外,我们的方法在高视听特征下采样下提高了字错误率(WER),同时在较低的下采样率下保持稳定。
摘要:Large language models (LLMs) have recently advanced auditory speech recognition (ASR), visual speech recognition (VSR), and audio-visual speech recognition (AVSR). However, understanding of their internal dynamics under fine-tuning remains limited. In natural language processing, recent work has revealed attention sinks, tokens that attract disproportionately high attention, and associated massive activations in which some features of sink tokens exhibit huge activation in LLMs. In this work, we are the first to study these phenomena in multimodal speech recognition. Through a detailed analysis of audio-visual LLMs, we identify attention sinks and massive activations not only at the BOS token but also at intermediate low-semantic tokens across ASR, VSR, and AVSR. We show that massive activations originate in the MLP layers and correspond to fixed feature indices across all sink tokens. We further show that intermediate sink tokens exhibit high cosine similarity to the BOS token, thereby amplifying attention and activation. Building on these insights, we introduce a simple decorrelation loss that reduces cosine similarity between BOS and other tokens, effectively mitigating intermediate sinks and massive activations. Furthermore, our method improves word error rate (WER) under high audio-visual feature downsampling while remaining stable at lower downsampling rates.


【11】UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
标题:UltraVoice:扩展口语对话模型的细粒度风格控制语音对话
链接:https://arxiv.org/abs/2510.22588

作者:Wenming Tu, Guanrou Yang, Ruiqi Yan, Wenxi Chen, Ziyang Ma, Yipeng Kang, Kai Yu, Xie Chen, Zilong Zheng
备注:23 pages, 4 figures
摘要:口语对话模型目前缺乏细粒度语音风格控制的能力,这是一种类似人类交互的关键能力,通常被忽视,而倾向于推理和问答等纯功能性功能。为了解决这个问题,我们引入了UltraVoice,这是第一个为多个细粒度语音风格控制而设计的大规模语音对话数据集。UltraVoice包含超过830小时的演讲对话,提供六个关键演讲风格维度的指导:情感,速度,音量,口音,语言和复合风格。对UltraVoice上的SLAM-Omni和VocalNet等领先模型进行微调,可显著增强其细粒度的语音风格可控性,而不会降低核心会话能力。具体来说,我们的微调模型实现了29.12-42.33%的平均意见得分(MOS)和14.61-40.09个百分点的指令遵循率(IFR)的多维控制任务设计的UltraVoice。此外,在URO-Bench基准测试中,我们的微调模型在核心理解、推理和会话能力方面表现出了实质性的进步,在基本设置上平均提高了+10.84%,在专业设置上平均提高了+7.87%。此外,该数据集的实用性扩展到训练可控的文本到语音(TTS)模型,强调其高质量和表达性语音合成的广泛适用性。完整的数据集和模型检查点可在https://github.com/bigai-nlco/UltraVoice上获得。
摘要:Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question answering. To address this limitation, we introduce UltraVoice, the first large-scale speech dialogue dataset engineered for multiple fine-grained speech style control. Encompassing over 830 hours of speech dialogues, UltraVoice provides instructions across six key speech stylistic dimensions: emotion, speed, volume, accent, language, and composite styles. Fine-tuning leading models such as SLAM-Omni and VocalNet on UltraVoice significantly enhances their fine-grained speech stylistic controllability without degrading core conversational abilities. Specifically, our fine-tuned models achieve improvements of 29.12-42.33% in Mean Opinion Score (MOS) and 14.61-40.09 percentage points in Instruction Following Rate (IFR) on multi-dimensional control tasks designed in the UltraVoice. Moreover, on the URO-Bench benchmark, our fine-tuned models demonstrate substantial gains in core understanding, reasoning, and conversational abilities, with average improvements of +10.84% on the Basic setting and +7.87% on the Pro setting. Furthermore, the dataset's utility extends to training controllable Text-to-Speech (TTS) models, underscoring its high quality and broad applicability for expressive speech synthesis. The complete dataset and model checkpoints are available at: https://github.com/bigai-nlco/UltraVoice.


【12】Empowering Multimodal Respiratory Sound Classification with Counterfactual Adversarial Debiasing for Out-of-Distribution Robustness
标题:通过反事实对抗去偏置来增强多模式呼吸声分类,以实现分布外的鲁棒性
链接:https://arxiv.org/abs/2510.22263

作者:Heejoon Koo, Miika Toikkanen, Yoon Tae Kim, Soo Yong Kim, June-Woo Kim
备注:3 figures, 4 Tables, and 5 pages
摘要:多模态呼吸音分类通过将生物声学信号与患者元数据相结合,为早期肺部疾病检测提供了希望。然而,目前的方法仍然容易受到来自诸如年龄、性别或采集设备等属性的虚假相关性的影响,这阻碍了它们的推广,特别是在跨临床站点的分布变化下。为此,我们提出了一个反事实对抗性去偏框架。首先,我们采用基于因果图的反事实去偏策略来抑制患者元数据的非因果依赖关系。其次,我们引入对抗性去偏置来学习元数据不敏感的表示并减少元数据特定的偏见。第三,我们设计了反事实元数据增强,以进一步减轻虚假的相关性,并加强元数据不变的表示。通过这样做,我们的方法在分布内和分布变化下的评估中始终优于强基线。该代码可在https://github.com/RSC-Toolkit/BTS-CARD上获得。
摘要:Multimodal respiratory sound classification offers promise for early pulmonary disease detection by integrating bioacoustic signals with patient metadata. Nevertheless, current approaches remain vulnerable to spurious correlations from attributes such as age, sex, or acquisition device, which hinder their generalization, especially under distribution shifts across clinical sites. To this end, we propose a counterfactual adversarial debiasing framework. First, we employ a causal graph-based counterfactual debiasing strategy to suppress non-causal dependencies from patient metadata. Second, we introduce adversarial debiasing to learn metadata-insensitive representations and reduce metadata-specific biases. Third, we design counterfactual metadata augmentation to mitigate spurious correlations further and strengthen metadata-invariant representations. By doing so, our method consistently outperforms strong baselines in evaluations under both in-distribution and distribution shifts. The code is available at https://github.com/RSC-Toolkit/BTS-CARD.


【13】Binaural Signal Matching with Wearable Arrays for Near-Field Sources and Directional Focus
标题:近场源和定向聚焦的可穿戴阵列的双耳信号匹配
链接:https://arxiv.org/abs/2510.22258

作者:Sapir Goldring, Zamir Ben Hur, David Lou Alon, Chad McKell, Sebastian Prepelita, Boaz Rafaely
摘要:本文研究了双耳信号匹配(BSM)方法的近场声音再现使用可穿戴眼镜安装麦克风阵列的性能。BSM是一种灵活的,信号独立的方法,用于任意阵列的双耳渲染,但其传统的配方假设远场源。在我们以前的工作中,我们提出了一个近场扩展的BSM(NF-BSM),它结合了距离相关的建模,并显示出使用分析数据的远场BSM的性能有所改善,尽管对于非常接近阵列的源,退化仍然存在。在这项研究中,我们通过使用近场头部相关的传递函数(HRTF)和声学传递函数(ATF)的阵列,占听众头部旋转和评估双耳线索,如耳间水平和时间差(ILD和ITD)的现实模拟数据扩展分析。一个关键的贡献是引入了视场(FoV)加权,旨在强调感知相关的方向,并提高在具有挑战性的条件下的鲁棒性。仿真和听力测试的结果证实,NF-BSM在近场场景中优于传统的远场BSM,并且所提出的NF-FoV-BSM方法在所有测试方法中实现了最佳的感知和客观质量,特别是在近源距离和头部旋转下。这些发现突出了近场源中远场模型的局限性,并表明将源距离和方向加权结合起来可以显着提高可穿戴空间音频系统的双耳再现性能。
摘要:This paper investigates the performance of Binaural Signal Matching (BSM) methods for near-field sound reproduction using a wearable glasses-mounted microphone array. BSM is a flexible, signal-independent approach for binaural rendering with arbitrary arrays, but its conventional formulation assumes far-field sources. In our previous work, we proposed a near-field extension of BSM (NF-BSM) that incorporates distance-dependent modeling and showed improved performance over far-field BSM using analytic data, though degradation persisted for sources very close to the array. In this study, we extend that analysis by using realistic simulated data of near-field Head-Related Transfer Functions (HRTFs) and Acoustic Transfer Functions (ATFs) of the array, accounting for listener head rotation and evaluating binaural cues such as interaural level and time differences (ILD and ITD). A key contribution is the introduction of a Field of View (FoV) weighting, designed to emphasize perceptually relevant directions and improve robustness under challenging conditions. Results from both simulation and a listening test confirm that NF-BSM outperforms traditional far-field BSM in near-field scenarios, and that the proposed NF-FoV-BSM method achieves the best perceptual and objective quality among all tested methods, particularly at close source distances and under head rotation. These findings highlight the limitations for far-field models in near-field sources and demonstrate that incorporating source distance and directional weighting can significantly improve binaural reproduction performance for wearable spatial audio systems.


【14】Bridging the Perceptual - Statistical Gap in Dysarthria Assessment: Why Machine Learning Still Falls Short
标题:弥合味觉障碍评估中的统计差距:为什么机器学习仍然落后
链接:https://arxiv.org/abs/2510.22237

作者:Krishna Gurugubelli
摘要:由于其潜在的临床影响,自动构音障碍检测和语音严重程度评估已经吸引了大量的研究关注。尽管声学建模和深度学习取得了快速进展,但模型仍然无法达到人类专家的性能。这份手稿对这一差距背后的原因进行了全面分析,强调了我们称之为“感知统计差距”的概念分歧。我们详细介绍了人类专家的感知过程,调查机器学习的表示和方法,回顾现有的文献特征集和建模策略,并提出了一个理论分析的标签噪声和评分员间的差异所施加的限制。我们进一步概述了缩小差距的实用策略,感知动机特征,自我监督预训练,ASR通知目标,多模态融合,人在回路训练和可解释性方法。最后,我们提出了与临床目标相一致的实验方案和评估指标,以指导未来的临床可靠和可解释的构音障碍评估工具的研究。
摘要:Automated dysarthria detection and severity assessment from speech have attracted significant research attention due to their potential clinical impact. Despite rapid progress in acoustic modeling and deep learning, models still fall short of human expert performance. This manuscript provides a comprehensive analysis of the reasons behind this gap, emphasizing a conceptual divergence we term the ``perceptual-statistical gap''. We detail human expert perceptual processes, survey machine learning representations and methods, review existing literature on feature sets and modeling strategies, and present a theoretical analysis of limits imposed by label noise and inter-rater variability. We further outline practical strategies to narrow the gap, perceptually motivated features, self-supervised pretraining, ASR-informed objectives, multimodal fusion, human-in-the-loop training, and explainability methods. Finally, we propose experimental protocols and evaluation metrics aligned with clinical goals to guide future research toward clinically reliable and interpretable dysarthria assessment tools.


【15】A Unified Framework for Direction and Diffuseness Estimation Using Tight-Frame Microphone Arrays
标题:使用紧框麦克风阵列进行方向和扩散度估计的统一框架
链接:https://arxiv.org/abs/2510.22183

作者:Akira Omoto
备注:36 pages including 14 files
摘要:这项工作提出了一个统一的框架,估计声场方向和扩散使用实际的麦克风阵列具有不同的空间配置。在基于协方差的扩散模型的基础上,我们制定了一种仅速度协方差方法,该方法能够在异构阵列几何形状上进行一致的扩散评估,而无需模式白化或球谐分解。三种阵列类型-A格式阵列,刚性球阵列,和一个新提出的紧帧阵列-建模和比较,通过模拟和基于测量的实验。结果表明,紧帧配置实现了近各向同性的方向性采样和再现扩散特性相媲美的高阶球面阵列,同时保持紧凑的物理结构。我们进一步研究的准确性的到达方向估计声强度的基础上在同一框架内。这些研究结果连接理论的扩散分析与可实施的阵列设计,并支持强大的,宽带的空间声场表征方法的发展。
摘要:This work presents a unified framework for estimating both sound-field direction and diffuseness using practical microphone arrays with different spatial configurations. Building on covariance-based diffuseness models, we formulate a velocity-only covariance approach that enables consistent diffuseness evaluation across heterogeneous array geometries without requiring mode whitening or spherical-harmonic decomposition. Three array types -- an A-format array, a rigid-sphere array, and a newly proposed tight-frame array -- are modeled and compared through both simulations and measurement-based experiments. The results show that the tight-frame configuration achieves near-isotropic directional sampling and reproduces diffuseness characteristics comparable to those of higher-order spherical arrays, while maintaining a compact physical structure. We further examine the accuracy of direction-of-arrival estimation based on acoustic intensity within the same framework. These findings connect theoretical diffuseness analysis with implementable array designs and support the development of robust, broadband methods for spatial-sound-field characterization.


【16】ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
标题:ISA-Bench:大型音频语言模型的基准指令灵敏度
链接:https://arxiv.org/abs/2510.23558

作者:Bohan Li, Wenbin Huang, Yuhang Qiu, Yiwei Guo, Hankun Wang, Zhihan Li, Jing Peng, Ziyang Ma, Xie Chen, Kai Yu
备注:submitted to icassp 2026
摘要:大型音频语言模型(LALM)将声学感知与大型语言模型(LLM)相结合,以从音频中提取和理解各种信息,引起了学术界和工业界的浓厚兴趣。然而,现有的LALM对指令的措辞方式高度敏感,影响(i)遵守率和(ii)任务性能。然而,没有任何现有的基准对这种敏感性进行系统和全面的评价。我们介绍ISA-Bench,一个动态的基准评估指令灵敏度LALM沿三个轴:指令描述,输出格式,任务组成。我们使用ISA-Bench评估最近的开源和专有LALM,在受控指令变化下分析合规性和准确性。实验结果表明,即使是最先进的LALM遭受显着的指令敏感性,导致性能下降的基本音频理解任务。为了缓解这个问题,我们在一个专门构建的复杂的自适应变量数据集上对Qwen 2-Audio进行了微调,从而显著提高了自适应跟踪性能。然而,这也导致了非平凡的灾难性遗忘:当暴露于新的指令风格时,模型失去了一些以前掌握的任务能力。我们的基准测试为评估和提高LALM中的指令灵敏度提供了标准化的基础,强调了在现实世界的管道中对预防性强大的音频理解的需求。
摘要:Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted intense interest from both academic and industrial communities. However, existing LALMs are highly sensitive to how instructions are phrased, affecting both (i) instruction-following rates and (ii) task performance. Yet, no existing benchmarks offer a systematic and comprehensive evaluation of this sensitivity. We introduce ISA-Bench, a dynamic benchmark evaluating instruction sensitivity for LALMs along three axes: instruction description, output format, and task composition. We assess recent open-source and proprietary LALMs using ISA-Bench, profiling both compliance and accuracy under controlled instruction variations. Experimental results reveal that even state-of-the-art LALMs suffer significant instruction sensitivity, leading to degraded performance on fundamental audio understanding tasks. To mitigate this issue, we fine-tune Qwen2-Audio on a specifically constructed complex instruction-variant dataset, achieving a marked improvement in instruction-following performance. However, this also induces nontrivial catastrophic forgetting: the model loses some previously mastered task capabilities when exposed to new instruction styles. Our benchmark provides a standardized basis for assessing and improving instruction sensitivity in LALMs, underscoring the need for instruction-robust audio understanding in real-world pipelines.


【17】Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
标题:基于隐式正则化的音频一致性自编码器线性学习
链接:https://arxiv.org/abs/2510.23530

作者:Bernardo Torres, Manuel Moussallam, Gabriel Meseguer-Brocal
摘要:音频自动编码器学习有用的压缩音频表示,但它们的非线性潜在空间阻止了直观的代数操作,如混合或缩放。我们介绍了一种简单的训练方法,通过使用数据增强来诱导高压缩一致性自动编码器(CAE)中的线性,从而诱导同质性(标量增益的等方差)和可加性(解码器保留加法),而不改变模型的架构或损失函数。当用我们的方法训练时,CAE在编码器和解码器中都表现出线性行为,同时保持重建保真度。我们通过简单的潜在算法测试我们的学习空间对音乐源合成和分离的实际效用。这项工作提出了一种简单的技术,用于构建结构化的潜在空间,使更直观和有效的音频处理。
摘要:Audio autoencoders learn useful, compressed audio representations, but their non-linear latent spaces prevent intuitive algebraic manipulation such as mixing or scaling. We introduce a simple training methodology to induce linearity in a high-compression Consistency Autoencoder (CAE) by using data augmentation, thereby inducing homogeneity (equivariance to scalar gain) and additivity (the decoder preserves addition) without altering the model's architecture or loss function. When trained with our method, the CAE exhibits linear behavior in both the encoder and decoder while preserving reconstruction fidelity. We test the practical utility of our learned space on music source composition and separation via simple latent arithmetic. This work presents a straightforward technique for constructing structured latent spaces, enabling more intuitive and efficient audio processing.


【18】Low-Resource Audio Codec (LRAC): 2025 Challenge Description
标题:低资源音频编解码器(LRAC):2025年挑战描述
链接:https://arxiv.org/abs/2510.23312

作者:Kamil Wojcicki, Yusuf Ziya Isik, Laura Lechler, Mansur Yesilbursa, Ivana Balić, Wolfgang Mack, Rafał Łaganowski, Guoqing Zhang, Yossi Adi, Minje Kim, Shinji Watanabe
摘要:虽然最近的神经音频编解码器在超低比特率下提供优于传统方法的语音质量,但它们的实际采用受到与低资源操作和对声学失真的鲁棒性相关的障碍的阻碍。边缘部署场景要求编解码器在严格的计算约束下运行,同时保持低延迟和比特率。背景噪声和混响的存在进一步需要对这种退化有弹性的设计。神经编解码器在这些约束下的性能及其与语音增强的集成在很大程度上仍未得到解决。为了促进这一领域的进展,我们推出了2025年低资源音频编解码器挑战赛,其目标是为资源受限的应用开发神经和混合编解码器。参与者将获得一个标准化的培训数据集、两个基线系统和一个全面的评估框架。预计该挑战将产生适用于编解码器设计和相关下游音频任务的有价值的见解。
摘要:While recent neural audio codecs deliver superior speech quality at ultralow bitrates over traditional methods, their practical adoption is hindered by obstacles related to low-resource operation and robustness to acoustic distortions. Edge deployment scenarios demand codecs that operate under stringent compute constraints while maintaining low latency and bitrate. The presence of background noise and reverberation further necessitates designs that are resilient to such degradations. The performance of neural codecs under these constraints and their integration with speech enhancement remain largely unaddressed. To catalyze progress in this area, we introduce the 2025 Low-Resource Audio Codec Challenge, which targets the development of neural and hybrid codecs for resource-constrained applications. Participants are supported with a standardized training dataset, two baseline systems, and a comprehensive evaluation framework. The challenge is expected to yield valuable insights applicable to both codec design and related downstream audio tasks.


【19】Evaluating Multimodal Large Language Models on Core Music Perception Tasks
标题:评估核心音乐感知任务的多模式大型语言模型
链接:https://arxiv.org/abs/2510.22455

作者:Brandon James Carone, Iran R. Roman, Pablo Ripollés
备注:Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables
摘要:多模态大型语言模型(LLM)通过将听力与乐谱阅读相结合的评估来声称“音乐理解”。我们在三个核心音乐技能上对三个SOTA LLM(Gemini 2.5 Pro,Gemini 2.5 Flash和Qwen2.5-Omni)进行基准测试:切分音评分,换位检测和和弦质量识别。此外,我们分离出三个可变性来源:(i)感知限制(音频与音频输入),(ii)暴露于示例(零与Few-Shot操作),以及(iii)推理策略(独立,CoT,LogicLM)。对于后者,我们适应LogicLM,一个框架结合LLM与符号求解器进行结构化推理,音乐。结果显示了一个明显的感知差距:模型在音频上的表现接近天花板,但在音频上的准确性下降。推理和Few-Shot提示只能提供最小的收益。这对于性能达到饱和的音频来说是意料之中的,但对于音频来说更令人惊讶的是,LogicLM尽管有近乎完美的音频精度,但仍然非常脆弱。在所有型号中,Gemini Pro在大多数条件下都能实现最高性能。总的来说,当前的系统在符号(symbols,简写为symbol)上推理得很好,但还不能可靠地从音频中“收听”。我们的方法和数据集使感知推理边界明确,并为构建强大的音频优先音乐系统提供可操作的指导。
摘要:Multimodal Large Language Models (LLMs) claim "musical understanding" via evaluations that conflate listening with score reading. We benchmark three SOTA LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni) across three core music skills: Syncopation Scoring, Transposition Detection, and Chord Quality Identification. Moreover, we separate three sources of variability: (i) perceptual limitations (audio vs. MIDI inputs), (ii) exposure to examples (zero- vs. few-shot manipulations), and (iii) reasoning strategies (Standalone, CoT, LogicLM). For the latter we adapt LogicLM, a framework combining LLMs with symbolic solvers to perform structured reasoning, to music. Results reveal a clear perceptual gap: models perform near ceiling on MIDI but show accuracy drops on audio. Reasoning and few-shot prompting offer minimal gains. This is expected for MIDI, where performance reaches saturation, but more surprising for audio, where LogicLM, despite near-perfect MIDI accuracy, remains notably brittle. Among models, Gemini Pro achieves the highest performance across most conditions. Overall, current systems reason well over symbols (MIDI) but do not yet "listen" reliably from audio. Our method and dataset make the perception-reasoning boundary explicit and offer actionable guidance for building robust, audio-first music systems.


【20】GuitarFlow: Realistic Electric Guitar Synthesis From Tablatures via Flow Matching and Style Transfer
标题:GuitarFlow:通过流量匹配和风格转移从表格中进行逼真的电吉他合成
链接:https://arxiv.org/abs/2510.21872

作者:Jackson Loth, Pedro Sarmento, Mark Sandler, Mathieu Barthet
备注:To be published in Proceedings of the 17th International Symposium on Computer Music and Multidisciplinary Research (CMMR)
摘要:近年来,使用人工智能(AI)在音频领域生成音乐取得了稳步进展。然而,对于某些乐器,特别是吉他,可控乐器合成在表现力方面仍然有限。我们介绍GuitarFlow,一个专门为电吉他合成设计的模型。生成过程是使用表格,一个无处不在的和直观的吉他特定的符号格式指导。指谱格式很容易表现出吉他特有的演奏技巧(例如压音、静音弦和连奏),而这些技巧在其他常见的音乐符号格式(如连奏)中则更难表现。我们的模型依赖于一个中间步骤,首先使用一个简单的基于样本的虚拟乐器将指谱渲染为音频,然后使用Flow Matching执行风格转换,以便将虚拟乐器音频转换为更逼真的声音示例。这使得模型能够快速进行训练并执行推理,所需的训练数据少于6小时。我们提供了客观评估指标的结果以及听力测试,其中我们显示出从图表生成的吉他音频的真实感有显着提高。
摘要:Music generation in the audio domain using artificial intelligence (AI) has witnessed steady progress in recent years. However for some instruments, particularly the guitar, controllable instrument synthesis remains limited in expressivity. We introduce GuitarFlow, a model designed specifically for electric guitar synthesis. The generative process is guided using tablatures, an ubiquitous and intuitive guitar-specific symbolic format. The tablature format easily represents guitar-specific playing techniques (e.g. bends, muted strings and legatos), which are more difficult to represent in other common music notation formats such as MIDI. Our model relies on an intermediary step of first rendering the tablature to audio using a simple sample-based virtual instrument, then performing style transfer using Flow Matching in order to transform the virtual instrument audio into more realistic sounding examples. This results in a model that is quick to train and to perform inference, requiring less than 6 hours of training data. We present the results of objective evaluation metrics, together with a listening test, in which we show significant improvement in the realism of the generated guitar audio from tablatures.


【21】Quantifying Multimodal Imbalance: A GMM-Guided Adaptive Loss for Audio-Visual Learning
标题:量化多模式失衡:GMM引导的视听学习适应性损失
链接:https://arxiv.org/abs/2510.21797

作者:Zhaocheng Liu, Zhiwen Yu, Xiaoqing Liu
摘要:目前解决多模态不平衡的主流方法主要集中在基于架构的修改和优化上,往往忽略了对模态间不平衡程度的定量分析。为了解决这一差距,我们的工作引入了一种新的方法来定量分析多模态不平衡,这反过来又为样本级自适应损失函数的设计提供了信息。我们首先将“模态差距”定义为不同模态的Softmax分数之间的差异(例如,音频和视频)用于地面实况类预测。模态间隙分布的分析表明,它可以有效地建模的双峰高斯混合模型(GMM)。这两个组成部分被发现分别对应于“模态平衡”和“模态不平衡”的数据样本。随后,我们应用贝叶斯定理计算每个样本属于这两个不同分布的后验概率,并在此基础上设计了一个新的自适应损失函数,该函数具有三个目标:(1)最小化总体模态间隙;(2)鼓励不平衡样本分布向平衡分布移动;(3)对不平衡样本施加更大的惩罚权重。实验结果表明,该方法在CREMA-D和AVE数据集上分别达到了SOTA(state-of-the-art)性能,准确率分别为80.65美元和70.90美元。这验证了我们提出的方法的有效性。
摘要:Current mainstream approaches to addressing multimodal imbalance primarily focus on architectural modifications and optimization-based, often overlooking a quantitative analysis of the imbalance degree between modalities. To address this gap, our work introduces a novel method for the quantitative analysis of multi-modal imbalance, which in turn informs the design of a sample-level adaptive loss function.We begin by defining the "Modality Gap" as the difference between the Softmax scores of different modalities (e.g., audio and visual) for the ground-truth class prediction. Analysis of the Modality Gap distribution reveals that it can be effectively modeled by a bimodal Gaussian Mixture Model (GMM). These two components are found to correspond respectively to "modality-balanced" and "modality-imbalanced" data samples. Subsequently, we apply Bayes' theorem to compute the posterior probability of each sample belonging to these two distinct distributions.Informed by this quantitative analysis, we design a novel adaptive loss function with three objectives: (1) to minimize the overall Modality Gap; (2) to encourage the imbalanced sample distribution to shift towards the balanced one; and (3) to apply greater penalty weights to imbalanced samples. We employ a two-stage training strategy consisting of a warm-up phase followed by an adaptive training phase.Experimental results demonstrate that our approach achieves state-of-the-art (SOTA) performance on the public CREMA-D and AVE datasets, attaining accuracies of $80.65\%$ and $70.90\%$, respectively. This validates the effectiveness of our proposed methodology.


【22】Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
标题:超越SVR触摸音调:使用LLM的客户意图路由
链接:https://arxiv.org/abs/2510.21715

作者:Sergio Rojas-Galeano
备注:Accepted for publication in the Proceedings of the Workshop on Engineering Applications 2025 (WEA 2025)
摘要:人们普遍对刻板的按键式交互式语音应答(IVR)系统感到失望,这突显了对更直接、更直观的语言交互的需求。虽然语音技术是必要的,但关键的挑战在于将用户的意图从用户的措辞路由到IVR菜单路径,这是一项大型语言模型(LLM)显示出强大潜力的任务。然而,进展受到数据稀缺的限制,因为真正的IVR结构和交互通常是专有的。我们提出了一种新的基于LLM的方法来解决这一差距。使用三种不同的模型,我们合成了一个现实的23节点IVR结构,生成了920个用户意图(230个基础和690个增强),并执行路由任务。我们评估两个提示设计:描述性的层次菜单和扁平化的路径表示,在基础和增强数据集。结果表明,扁平化路径始终产生更高的准确率,在基本数据集上达到89.13%,而描述性格式为81.30%,而增强则引入了语言噪声,略微降低了性能。混淆矩阵分析进一步表明,低性能的路线可能不仅反映了模型的局限性,但也冗余的菜单设计。总的来说,我们的研究结果表明,LLM可以通过更流畅,更无缝的用户体验实现IVR路由-将客户服务提前一步。
摘要:Widespread frustration with rigid touch-tone Interactive Voice Response (IVR) systems for customer service underscores the need for more direct and intuitive language interaction. While speech technologies are necessary, the key challenge lies in routing intents from user phrasings to IVR menu paths, a task where Large Language Models (LLMs) show strong potential. Progress, however, is limited by data scarcity, as real IVR structures and interactions are often proprietary. We present a novel LLM-based methodology to address this gap. Using three distinct models, we synthesized a realistic 23-node IVR structure, generated 920 user intents (230 base and 690 augmented), and performed the routing task. We evaluate two prompt designs: descriptive hierarchical menus and flattened path representations, across both base and augmented datasets. Results show that flattened paths consistently yield higher accuracy, reaching 89.13% on the base dataset compared to 81.30% with the descriptive format, while augmentation introduces linguistic noise that slightly reduces performance. Confusion matrix analysis further suggests that low-performing routes may reflect not only model limitations but also redundancies in menu design. Overall, our findings demonstrate proof-of-concept that LLMs can enable IVR routing through a smoother, more seamless user experience -- moving customer service one step ahead of touch-tone menus.


机器翻译由腾讯交互翻译提供,仅供参考