今日论文合集:cs.SD语音37篇,eess.AS音频处理35篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 VoxAging: Continuously Tracking Speaker Aging with a Large-Scale  Longitudinal Dataset in English and Mandarin
标题: VoxAging:使用英语和普通话大规模纵向数据集连续跟踪说话者老化
链接:https://arxiv.org/abs/2505.21445
作者: Zhiqi Ai,  Meixuan Bao,  Zhiyong Chen,  Zhi Yang,  Xinnuo Li,  Shugong Xu 
备注:5 pages, 4 figures, Accepted by Interspeech 2025
摘要:说话人老化会对说话人确认系统的性能产生不利影响。然而,由于数据收集方面的挑战,特别是缺乏持续和大规模的个人纵向数据,扬声器老化的研究仍然很困难。在本文中,我们介绍了VoxAging,这是一个大规模的纵向数据集,收集了293名说话者(226名英语使用者和67名普通话使用者)多年的数据,最长的时间跨度达到17年(约900周)。每隔一周记录一次每个发言者的数据。研究了说话人老化现象及其对高级说话人确认系统的影响,分析了个体说话人老化过程,探讨了年龄、性别等因素对说话人老化研究的影响。
摘要:The performance of speaker verification systems is adversely affected by speaker aging. However, due to challenges in data collection, particularly the lack of sustained and large-scale longitudinal data for individuals, research on speaker aging remains difficult. In this paper, we present VoxAging, a large-scale longitudinal dataset collected from 293 speakers (226 English speakers and 67 Mandarin speakers) over several years, with the longest time span reaching 17 years (approximately 900 weeks). For each speaker, the data were recorded at weekly intervals. We studied the phenomenon of speaker aging and its effects on advanced speaker verification systems, analyzed individual speaker aging processes, and explored the impact of factors such as age group and gender on speaker aging research.


【2】 Towards Robust Automated Perceptual Voice Quality Assessment with Deep  Learning

标题: 利用深度学习实现稳健的自动感知语音质量评估
链接:https://arxiv.org/abs/2505.21356
作者: Whenty Ariyanti,  Kuan-Yu Chen,  Sabato Marco Siniscalchi,  Hsin-Min Wang,  Yu Tsao 
摘要:目的:感知嗓音质量评估通过提供对发声功能的标准化评价,在嗓音疾病的诊断和监测中起着至关重要的作用。传统上,这个过程依赖于专家评分员使用标准量表,如声音的共识听觉感知评估(CAPE-V)和等级,粗糙度,呼吸,虚弱和紧张(GRBAS)。然而,这些指标本质上是主观的,容易受到评分者之间的变化,激发了对自动化和客观评估方法的需求。研究方法:我们提出了语音质量评估网络(VOQANet),这是一个基于深度学习的框架,具有注意力机制,利用语音基础模型(SFM)从原始语音中捕获高级声学和韵律信息。为了增强鲁棒性和可解释性,我们提出了VOQANet+,它集成了手工制作的声学特征,如抖动,闪烁和谐波噪声比(HNR)与SFM嵌入。结果:基于句子的输入比基于元音的输入产生更强的性能,特别是在病人的水平。VOQANet在RMSE和PCC方面始终优于基线方法,而VOQANet+在噪声条件下表现更好并保持鲁棒性。结论:将SFM嵌入与域信息声学特征相结合,提高了可解释性和弹性。重要性:VOQANet+显示出在现实世界和远程医疗环境中部署的强大潜力,通过可解释和抗噪声的解决方案解决了主观感知评估的局限性。
摘要:Objective: Perceptual voice quality assessment plays a critical role in diagnosing and monitoring voice disorders by providing standardized evaluation of vocal function. Traditionally, this process relies on expert raters utilizing standard scales, such as the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and Grade, Roughness, Breathiness, Asthenia, and Strain (GRBAS). However, these metrics are inherently subjective and susceptible to inter-rater variability, motivating the need for automated and objective assessment methods. Methods: We propose Voice Quality Assessment Network (VOQANet), a deep learning-based framework with an attention mechanism that leverages a Speech Foundation Model (SFM) to capture high-level acoustic and prosodic information from raw speech. To enhance robustness and interpretability, we present VOQANet+, which integrates handcrafted acoustic features such as jitter, shimmer, and harmonics-to-noise ratio (HNR) with SFM embeddings. Results: Sentence-based input yields stronger performance than vowel-based input, especially at the patient level. VOQANet consistently outperforms baseline methods in RMSE and PCC, while VOQANet+ performs even better and maintains robustness under noisy conditions. Conclusion: Combining SFM embeddings with domain-informed acoustic features improves interpretability and resilience. Significance: VOQANet+ shows strong potential for deployment in real-world and telehealth settings, addressing the limitations of subjective perceptual assessments with an interpretable and noise-resilient solution.


【3】 Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using  Co-training and Stochastic Precision

标题: 迈向一位ASO:使用联合训练和随机精度的极低位适形器量化
链接:https://arxiv.org/abs/2505.21245
作者: Zhaoqing Li,  Haoning Xu,  Zengrui Jin,  Lingwei Meng,  Tianzi Wang,  Huimeng Wang,  Youjun Chen,  Mingyu Cui,  Shujie Hu,  Xunying Liu 
备注:Accepted by Interspeech2025
摘要:随着现代语音系统规模的迅速增加,模型压缩已经成为一种新兴的需求。在本文中,我们研究模型权重量化,直接减少内存占用,以适应计算资源受限的应用程序。我们提出了新的方法来执行极低比特(即,2-比特和1比特)量化的Conformer自动语音识别系统使用多精度模型协同训练、随机精度和张量式可学习缩放因子来减轻量化引起的性能损失。所提出的方法可以实现性能无损的2位和1位量化的Conformer ASR系统训练的300小时开关板和960小时LibriSpeech语料库。最大的整体性能无损压缩比的16.2和16.6倍,实现了没有统计上显着增加的字错误率(WER)在全精度基线系统,分别。
摘要:Model compression has become an emerging need as the sizes of modern speech systems rapidly increase. In this paper, we study model weight quantization, which directly reduces the memory footprint to accommodate computationally resource-constrained applications. We propose novel approaches to perform extremely low-bit (i.e., 2-bit and 1-bit) quantization of Conformer automatic speech recognition systems using multiple precision model co-training, stochastic precision, and tensor-wise learnable scaling factors to alleviate quantization incurred performance loss. The proposed methods can achieve performance-lossless 2-bit and 1-bit quantization of Conformer ASR systems trained with the 300-hr Switchboard and 960-hr LibriSpeech corpus. Maximum overall performance-lossless compression ratios of 16.2 and 16.6 times are achieved without a statistically significant increase in the word error rate (WER) over the full precision baseline systems, respectively.


【4】 Unfolding A Few Structures for The Many: Memory-Efficient Compression of  Conformer and Speech Foundation Models

标题: 为许多人展开一些结构:适形器和语音基础模型的内存高效压缩
链接:https://arxiv.org/abs/2505.21237
作者: Zhaoqing Li,  Haoning Xu,  Xurong Xie,  Zengrui Jin,  Tianzi Wang,  Xunying Liu 
备注:Accepted by Interspeech2025
摘要:本文提出了一种新的内存有效的模型压缩方法的Conformer ASR和语音基础系统。我们的方法具有独特的“从小到大”设计。包含几个Conformer或Transformer块的紧凑“种子”模型经过多次训练和展开,以模拟具有不同逻辑深度的较大未压缩模型的性能。种子模型和许多展开路径在单个展开周期内联合训练。在自蒸馏过程中使用最大展开和最小种子模型之间的KL发散,以最小化它们的性能差异。实验结果表明,我们的可折叠模型产生的ASR性能与单独构建的Conformer和wav 2 vec 2/HuBERT语音基础模型在各种深度配置下相当,同时只需要最少的内存和存储。构象和wav 2 vec 2模型的参数分别减少了35%和30%,而性能没有损失。
摘要:This paper presents a novel memory-efficient model compression approach for Conformer ASR and speech foundation systems. Our approach features a unique "small-to-large" design. A compact "seed" model containing a few Conformer or Transformer blocks is trained and unfolded many times to emulate the performance of larger uncompressed models with different logical depths. The seed model and many unfolded paths are jointly trained within a single unfolding cycle. The KL-divergence between the largest unfolded and smallest seed models is used in a self-distillation process to minimize their performance disparity. Experimental results show that our foldable model produces ASR performance comparable to individually constructed Conformer and wav2vec2/HuBERT speech foundation models under various depth configurations, while requiring only minimal memory and storage. Conformer and wav2vec2 models with a reduction of 35% and 30% parameters are obtained without loss of performance, respectively.


【5】 Universal Speech Enhancement with Regression and Generative Mamba

标题: 使用回归和生成曼巴的通用语音增强
链接:https://arxiv.org/abs/2505.21198
作者: Rong Chao,  Rauf Nasretdinov,  Yu-Chiang Frank Wang,  Ante Jukić,  Szu-Wei Fu,  Yu Tsao 
备注:Accepted to Interspeech 2025
摘要:Interspeech 2025 URGENT Challenge旨在通过统一各种条件下的语音增强任务,包括七种不同的失真类型和五种语言,来推进通用、稳健和可推广的语音增强。我们提出了通用语音增强Mamba(USEamba),这是一种状态空间语音增强模型,旨在处理远程序列建模、时频结构化处理和采样频率独立特征提取。我们的方法主要依赖于基于回归的建模,它在大多数失真中表现良好。然而,对于数据包丢失和带宽扩展,必须推断丢失的内容,所提出的USEMAMBA的生成变体证明更有效。尽管仅在完整训练数据的一个子集上进行训练,但USEMAMBA在盲测阶段在Track 1中获得了第二名,在各种条件下表现出了很强的泛化能力。
摘要:The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions.


【6】 Topological Deep Learning for Speech Data

标题: 语音数据的布局深度学习
链接:https://arxiv.org/abs/2505.21173
作者: Zhiwang Yu 
备注:21 pages, 15 figures
摘要:拓扑数据分析(TDA)为深度学习提供了新的数学工具。受Carlsson等人的启发,这项研究设计了拓扑感知卷积核,显著改善了语音识别网络。理论上,通过研究正交群作用于核函数,我们建立了矩阵空间的纤维束分解,从而实现了新的滤波器生成方法。实际上,我们提出的正交特征(OF)层在音素识别方面取得了优异的性能,特别是在低噪声的情况下,同时表现出跨域的适应性。这项工作揭示了TDA在神经网络优化方面的潜力,为神经学-深度学习跨学科研究开辟了新的途径。
摘要:Topological data analysis (TDA) offers novel mathematical tools for deep learning. Inspired by Carlsson et al., this study designs topology-aware convolutional kernels that significantly improve speech recognition networks. Theoretically, by investigating orthogonal group actions on kernels, we establish a fiber-bundle decomposition of matrix spaces, enabling new filter generation methods. Practically, our proposed Orthogonal Feature (OF) layer achieves superior performance in phoneme recognition, particularly in low-noise scenarios, while demonstrating cross-domain adaptability. This work reveals TDA's potential in neural network optimization, opening new avenues for mathematics-deep learning interdisciplinary studies.


【7】 Model as Loss: A Self-Consistent Training Paradigm

标题: 模型即损失:自我一致的训练范式
链接:https://arxiv.org/abs/2505.21156
作者: Saisamarth Rajesh Phaye,  Milos Cernak,  Andrew Harper 
备注:Accepted in Interspeech 2025
摘要:用于语音增强的常规方法依赖于手工制作的损失函数(例如,时域或频域损失)或深特征损失(例如,使用WavLM或wav2vec),这通常无法捕获对于最佳性能至关重要的细微信号特性。为了解决这个问题,我们提出了Model as Loss,这是一种新颖的训练范式,它利用来自同一模型的编码器作为损失函数来指导训练。   损失模型范式利用编码器的特定于任务的特征空间,优化解码器以产生与干净信号的感知和任务相关特征一致的输出。通过使用编码器的学习功能作为损失函数,该框架强制执行干净的参考语音和增强的模型输出之间的自一致性。我们的方法在标准语音增强基准测试中优于预训练的深度特征损失,为域内和域外数据集提供更好的感知质量和鲁棒的泛化。
摘要:Conventional methods for speech enhancement rely on handcrafted loss functions (e.g., time or frequency domain losses) or deep feature losses (e.g., using WavLM or wav2vec), which often fail to capture subtle signal properties essential for optimal performance. To address this, we propose Model as Loss, a novel training paradigm that utilizes the encoder from the same model as a loss function to guide the training.   The Model as Loss paradigm leverages the encoder's task-specific feature space, optimizing the decoder to produce output consistent with perceptual and task-relevant characteristics of the clean signal. By using the encoder's learned features as a loss function, this framework enforces self-consistency between the clean reference speech and the enhanced model output. Our approach outperforms pre-trained deep feature losses on standard speech enhancement benchmarks, offering better perceptual quality and robust generalization to both in-domain and out-of-domain datasets.


【8】 Assessment of L2 Oral Proficiency using Speech Large Language Models

标题: 使用言语大语言模型评估二语口语能力
链接:https://arxiv.org/abs/2505.21148
作者: Rao Ma,  Mengjie Qian,  Siyuan Tang,  Stefano Bannò,  Kate M. Knill,  Mark J.F. Gales 
备注:submitted to Interspeech
摘要:随着二语人口的不断增长,对口语评估自动评分系统的需求也越来越大。从历史上看,统计模型,文本编码器和自监督语音模型已被用于此任务。然而,级联系统遭受信息丢失,而E2E分级机也有局限性。随着多模态大语言模型(LLM)的最新进展,我们的目标是探索其作为二语口语水平评分的潜力,并克服这些问题。在这项工作中,我们比较了使用回归和分类目标的各种训练策略。我们的研究结果表明,语音LLM优于所有以前的竞争基线,在两个数据集上实现了卓越的性能。此外,受过训练的评分员在跨部分或跨任务评估中表现出较强的概括能力,这得益于LLM预培训期间获得的音频理解知识。
摘要:The growing population of L2 English speakers has increased the demand for developing automatic graders for spoken language assessment (SLA). Historically, statistical models, text encoders, and self-supervised speech models have been utilised for this task. However, cascaded systems suffer from the loss of information, while E2E graders also have limitations. With the recent advancements of multi-modal large language models (LLMs), we aim to explore their potential as L2 oral proficiency graders and overcome these issues. In this work, we compare various training strategies using regression and classification targets. Our results show that speech LLMs outperform all previous competitive baselines, achieving superior performance on two datasets. Furthermore, the trained grader demonstrates strong generalisation capabilities in the cross-part or cross-task evaluation, facilitated by the audio understanding knowledge acquired during LLM pre-training.


【9】 Leveraging LLM and Self-Supervised Training Models for Speech  Recognition in Chinese Dialects: A Comparative Analysis

标题: 利用LLM和自我监督训练模型进行中文方言语音识别:比较分析
链接:https://arxiv.org/abs/2505.21138
作者: Tianyi Xu,  Hongjie Chen,  Wang Qing,  Lv Hang,  Jian Kang,  Li Jie,  Zhennan Lin,  Yongxiang Li,  Xie Lei 
摘要:大规模的训练语料库显著提高了ASR模型的性能。不幸的是,由于数据相对稀缺,中国口音和方言对于大多数ASR模型来说仍然是一个挑战。自监督学习的最新进展表明,自监督预训练与大型语言模型(LLM)相结合,可以有效地提高低资源场景中的ASR性能。我们的目的是调查这种范式对汉语方言的有效性。具体来说,我们在300,000小时的未标记方言和口音语音数据上预训练Data2vec2模型,并在40,000小时的监督数据集上进行对齐训练。然后,我们系统地研究了各种投影仪和LLM对普通话,方言和口音的语音识别性能的影响,在这个范例。我们的方法在多个方言数据集上实现了SOTA结果,包括Kespeech。我们将开源我们的工作,以促进可重复的研究
摘要:Large-scale training corpora have significantly improved the performance of ASR models. Unfortunately, due to the relative scarcity of data, Chinese accents and dialects remain a challenge for most ASR models. Recent advancements in self-supervised learning have shown that self-supervised pre- training, combined with large language models (LLM), can effectively enhance ASR performance in low-resource scenarios. We aim to investigate the effectiveness of this paradigm for Chinese dialects. Specifically, we pre-train a Data2vec2 model on 300,000 hours of unlabeled dialect and accented speech data and do alignment training on a supervised dataset of 40,000 hours. Then, we systematically examine the impact of various projectors and LLMs on Mandarin, dialect, and accented speech recognition performance under this paradigm. Our method achieved SOTA results on multiple dialect datasets, including Kespeech. We will open-source our work to promote reproducible research


【10】 Scaling and Prompting for Improved End-to-End Spoken Grammatical Error  Correction

标题: 缩放和绘图以改进端到端口语语法错误纠正
链接:https://arxiv.org/abs/2505.21137
作者: Mengjie Qian,  Rao Ma,  Stefano Bannò,  Kate M. Knill,  Mark J.F. Gales 
备注:submitted to Interspeech
摘要:口语语法错误纠正和反馈对于二语学习者、教师和考生都是至关重要的。传统的SGEC系统依赖于由ASR、用于不流利检测(DD)和去除的模块以及用于GEC的模块组成的级联流水线。随着端到端(E2E)语音基础模型的兴起,我们研究了它们在SGEC和反馈生成中的有效性。这项工作引入了一个伪标记过程来解决有限的标记数据的挑战,将训练数据的大小从77小时扩展到大约2500小时,从而提高了性能。此外,我们提示了一个基于E2E Whisper的SGEC模型,该模型具有流畅的transmittance,SGEC性能略有改善,在反馈生成方面有更显着的收益。最后,我们评估了增加模型大小的影响,揭示了虽然伪标记数据不会为更大的Whisper模型带来性能增益,但使用提示进行训练是有益的。
摘要:Spoken Grammatical Error Correction (SGEC) and Feedback (SGECF) are crucial for second language learners, teachers and test takers. Traditional SGEC systems rely on a cascaded pipeline consisting of an ASR, a module for disfluency detection (DD) and removal and one for GEC. With the rise of end-to-end (E2E) speech foundation models, we investigate their effectiveness in SGEC and feedback generation. This work introduces a pseudo-labelling process to address the challenge of limited labelled data, expanding the training data size from 77 hours to approximately 2500 hours, leading to improved performance. Additionally, we prompt an E2E Whisper-based SGEC model with fluent transcriptions, showing a slight improvement in SGEC performance, with more significant gains in feedback generation. Finally, we assess the impact of increasing model size, revealing that while pseudo-labelled data does not yield performance gain for a larger Whisper model, training with prompts proves beneficial.


【11】 Text-Queried Audio Source Separation via Hierarchical Modeling

标题: 通过分层建模的文本查询音频源分离
链接:https://arxiv.org/abs/2505.21025
作者: Xinlei Yin,  Xiulian Peng,  Xue Jiang,  Zhiwei Xiong,  Yan Lu 
摘要:目标音频源分离与自然语言查询提出了一个有前途的范例提取任意的音频事件,通过任意的文本描述。现有的方法主要面临两个挑战,即在盲学习的单阶段架构中联合建模声学-文本对齐和语义感知分离的困难,以及依赖大规模准确标记的训练数据来补偿低效的跨模态学习和分离。为了解决这些挑战,我们提出了一个分层分解框架,HSM-TSS,它将任务分解为全局-局部语义引导的特征分离和结构保留的声学重建。我们的方法引入了一个双阶段的语义分离机制,在不同的全球和本地语义特征空间。我们首先通过与文本查询对齐的全局语义特征空间执行全局语义分离。Q-Audio架构用于对齐音频和文本模态,用作预训练的全局语义编码器。在预测的全局特征的条件下,我们然后对保留时频结构的AudioMAE特征执行第二阶段局部语义分离,然后进行声学重建。我们还提出了一个指令处理流水线,将任意文本查询解析为结构化操作,提取或删除,再加上音频描述,实现灵活的声音操作。我们的方法通过数据高效的训练实现了最先进的分离性能,同时在复杂的听觉场景中保持了与查询的出色语义一致性。
摘要:Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing methods mainly face two challenges, the difficulty in jointly modeling acoustic-textual alignment and semantic-aware separation within a blindly-learned single-stage architecture, and the reliance on large-scale accurately-labeled training data to compensate for inefficient cross-modal learning and separation. To address these challenges, we propose a hierarchical decomposition framework, HSM-TSS, that decouples the task into global-local semantic-guided feature separation and structure-preserving acoustic reconstruction. Our approach introduces a dual-stage mechanism for semantic separation, operating on distinct global and local semantic feature spaces. We first perform global-semantic separation through a global semantic feature space aligned with text queries. A Q-Audio architecture is employed to align audio and text modalities, serving as pretrained global-semantic encoders. Conditioned on the predicted global feature, we then perform the second-stage local-semantic separation on AudioMAE features that preserve time-frequency structures, followed by acoustic reconstruction. We also propose an instruction processing pipeline to parse arbitrary text queries into structured operations, extraction or removal, coupled with audio descriptions, enabling flexible sound manipulation. Our method achieves state-of-the-art separation performance with data-efficient training while maintaining superior semantic consistency with queries in complex auditory scenes.


【12】 ClearSphere: Multi-Earphone Synergy for Enhanced Conversational Clarity

标题: ClearGlobe:多耳机协同作用,增强对话清晰度
链接:https://arxiv.org/abs/2505.21004
作者: Lixing He 
摘要:在会议等人群密集的场所,背景噪音、重叠的声音和生动的互动使人难以进行清晰的对话。这种情况通常会导致被称为“鸡尾酒会耳聋"的现象。“我们介绍了ClearSphere,这是一个协作系统,可以通过多耳机在对话级别增强语音。实时会话增强需要对会话中的所有成员进行整体建模,并需要一种有效的方法来从混合中提取语音。ClearSphere通过两个关键贡献将声学传感器系统和最先进的深度学习用于目标语音提取:1)对话驱动的网络协议,以及2)强大的目标对话提取模型。我们的网络协议实现了耳机设备之间的移动、无基础设施协调。我们的会话提取模型可以利用中继音频在带宽效率的方式。ClearSphere在真实世界的实验和模拟中进行了评估。结果表明,我们的会话网络获得超过90\%的准确率在组形成,提高了高达8.8 dB的语音质量超过国家的最先进的基线,并展示了实时性能的移动终端。在一项有20名参与者的用户研究中,ClearSphere的评分比基线高得多,具有良好的可用性。
摘要:In crowded places such as conferences, background noise, overlapping voices, and lively interactions make it difficult to have clear conversations. This situation often worsens the phenomenon known as "cocktail party deafness." We present ClearSphere, the collaborative system that enhances speech at the conversation level with multi-earphones. Real-time conversation enhancement requires a holistic modeling of all the members in the conversation, and an effective way to extract the speech from the mixture. ClearSphere bridges the acoustic sensor system and state-of-the-art deep learning for target speech extraction by making two key contributions: 1) a conversation-driven network protocol, and 2) a robust target conversation extraction model. Our networking protocol enables mobile, infrastructure-free coordination among earphone devices. Our conversation extraction model can leverage the relay audio in a bandwidth-efficient way. ClearSphere is evaluated in both real-world experiments and simulations. Results show that our conversation network obtains more than 90\% accuracy in group formation, improves the speech quality by up to 8.8 dB over state-of-the-art baselines, and demonstrates real-time performance on a mobile device. In a user study with 20 participants, ClearSphere has a much higher score than baseline with good usability.


【13】 MelodySim: Measuring Melody-aware Music Similarity for Plagiarism  Detection

标题: MelodySim:测量旋律感知音乐相似性以检测抄袭
链接:https://arxiv.org/abs/2505.20979
作者: Tongyu Lu,  Charlotta-Marlena Geist,  Jan Melechovsky,  Abhinaba Roy,  Dorien Herremans 
摘要:我们提出了MelodySim,一个旋律感知的音乐相似性模型和数据集的剽窃检测。首先,我们介绍了一种新的方法来构建一个数据集,重点是旋律相似性。通过增强Slakh 2100;现有的数据集,我们生成每首作品的变化,同时通过修改,如音符分割,琶音,次要轨道脱落(不包括低音)和重新乐器保留旋律。一项用户研究证实,阳性对确实包含相似的旋律,而其他曲目则发生了显着变化。其次,我们开发了一个分段旋律相似性检测模型,该模型使用MERT编码器并应用三元组神经网络来捕获旋律相似性。由此产生的决策矩阵突出了可能发生剽窃的地方。我们的模型在MelodySim测试集上实现了高精度。
摘要:We propose MelodySim, a melody-aware music similarity model and dataset for plagiarism detection. First, we introduce a novel method to construct a dataset with focus on melodic similarity. By augmenting Slakh2100; an existing MIDI dataset, we generate variations of each piece while preserving the melody through modifications such as note splitting, arpeggiation, minor track dropout (excluding bass), and re-instrumentation. A user study confirms that positive pairs indeed contain similar melodies, with other musical tracks significantly changed. Second, we develop a segment-wise melodic-similarity detection model that uses a MERT encoder and applies a triplet neural network to capture melodic similarity. The resultant decision matrix highlights where plagiarism might occur. Our model achieves high accuracy on the MelodySim test set.


【14】 Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization

标题: 高效的麦克风容错3D声源定位
链接:https://arxiv.org/abs/2505.20961
作者: Yiyuan Yang,  Shitong Xu,  Niki Trigoni,  Andrew Markham 
备注:Accepted by Interspeech 2025 Conference
摘要:声源定位是在复杂环境中确定声源位置的关键技术。然而,现有的方法面临的挑战,如高计算成本和精确的校准要求,限制其部署在动态或资源受限的环境。本文介绍了一种新的3D SSL框架,它使用稀疏交叉注意,预训练和自适应信号相干性度量,以实现准确和计算效率的定位与较少的输入麦克风。该框架还对不可靠甚至未知的麦克风位置输入具有容错能力,确保其在现实世界场景中的适用性。初步实验表明,它的可扩展性多源定位,而不需要额外的硬件。这项工作通过平衡模型的性能和效率并提高其对真实世界场景的鲁棒性来推进SSL。
摘要:Sound source localization (SSL) is a critical technology for determining the position of sound sources in complex environments. However, existing methods face challenges such as high computational costs and precise calibration requirements, limiting their deployment in dynamic or resource-constrained environments. This paper introduces a novel 3D SSL framework, which uses sparse cross-attention, pretraining, and adaptive signal coherence metrics, to achieve accurate and computationally efficient localization with fewer input microphones. The framework is also fault-tolerant to unreliable or even unknown microphone position inputs, ensuring its applicability in real-world scenarios. Preliminary experiments demonstrate its scalability for multi-source localization without requiring additional hardware. This work advances SSL by balancing the model's performance and efficiency and improving its robustness for real-world scenarios.


【15】 Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound  Event Detection

标题: 基于不一致-多样性混合主动学习的生物声事件检测
链接:https://arxiv.org/abs/2505.20956
作者: Shiqi Zhang,  Tuomas Virtanen 
备注:5 pages, 1 figure, accepted by EUSIPCO 2025
摘要:生物声学声音事件检测(BioSED)对于生物多样性保护至关重要,但在模型开发和训练过程中面临着实际挑战:有限的注释数据,稀疏事件,物种多样性和类别不平衡。为了在有限的标签预算下有效地解决这些挑战,我们采用了失配优先最远遍历(MFFT),这是一种集成委员会投票分歧和多样性分析的主动学习方法。我们还改进了现有的BioSED数据集,专门用于评估主动学习算法。实验结果表明,MFFT在冷启动时实现了68%的mAP,在热启动时实现了71%的mAP(接近于75%的全监督mAP),同时仅使用2.3%的注释。值得注意的是,MFFT在冷启动场景和稀有物种方面表现出色,这对于监测濒危物种至关重要,证明了其实用价值。
摘要:Bioacoustic sound event detection (BioSED) is crucial for biodiversity conservation but faces practical challenges during model development and training: limited amounts of annotated data, sparse events, species diversity, and class imbalance. To address these challenges efficiently with a limited labeling budget, we apply the mismatch-first farthest-traversal (MFFT), an active learning method integrating committee voting disagreement and diversity analysis. We also refine an existing BioSED dataset specifically for evaluating active learning algorithms. Experimental results demonstrate that MFFT achieves a mAP of 68% when cold-starting and 71% when warm-starting (which is close to the fully-supervised mAP of 75%) while using only 2.3% of the annotations. Notably, MFFT excels in cold-start scenarios and with rare species, which are critical for monitoring endangered species, demonstrating its practical value.


【16】 Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing

标题: Dub-S2 ST:无缝配音的无文本语音翻译
链接:https://arxiv.org/abs/2505.20899
作者: Jeongsoo Choi,  Jaehun Kim,  Joon Son Chung 
摘要:本文介绍了一个跨语言配音系统,语音从一种语言翻译到另一种语言,同时保留关键特征,如持续时间,扬声器的身份,和说话速度。尽管现有的语音翻译方法具有很强的翻译质量,但它们往往忽略了语音模式的传递,导致与源语音的不匹配,并限制了它们对配音应用的适用性。为了解决这个问题,我们提出了一个离散的基于扩散的语音到单元的翻译模型,具有显式的持续时间控制,使时间对齐的翻译。然后,我们合成语音的基础上预测的单位和源身份与条件流匹配模型。此外,我们引入了一个基于单元的速度自适应机制,指导翻译模型以与源一致的速率生成语音,而不依赖于任何文本。大量的实验表明,我们的框架生成自然和流畅的翻译,符合原始语音的持续时间和说话节奏,同时实现有竞争力的翻译性能。
摘要:This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong translation quality of existing speech translation approaches, they often overlook the transfer of speech patterns, leading to mismatches with source speech and limiting their suitability for dubbing applications. To address this, we propose a discrete diffusion-based speech-to-unit translation model with explicit duration control, enabling time-aligned translation. We then synthesize speech based on the predicted units and source identity with a conditional flow matching model. Additionally, we introduce a unit-based speed adaptation mechanism that guides the translation model to produce speech at a rate consistent with the source, without relying on any text. Extensive experiments demonstrate that our framework generates natural and fluent translations that align with the original speech's duration and speaking pace, while achieving competitive translation performance.


【17】 Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction  and Style Direction Adjustment for Expressive Text-to-Speech

标题: Spotlight TTS:通过语音感知的风格提取和风格方向调整来突出风格,用于表达性文本到语音
链接:https://arxiv.org/abs/2505.20868
作者: Nam-Gyu Kim,  Deok-Hyeon Cho,  Seung-Bin Kim,  Seong-Whan Lee 
备注:Submitted to Interspeech
摘要:表达性文本到语音(TTS)的最新进展介绍了各种方法的基础上提取参考语音的风格嵌入。然而,合成高质量的表达语音仍然具有挑战性。我们提出了Spotlight TTS,它专门强调风格,通过风格感知的风格提取和风格方向调整。浊音感知风格提取关注与风格高度相关的浊音区域,同时保持不同语音区域之间的连续性以提高表现力。我们调整了提取的风格的方向,以最佳地整合到TTS模型中,从而提高了语音质量。实验结果表明,聚光灯TTS实现了卓越的表现力,整体语音质量和风格转移能力的基线模型相比。我们的音频样本是公开的。
摘要:Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose Spotlight-TTS, which exclusively emphasizes style via voiced-aware style extraction and style direction adjustment. Voiced-aware style extraction focuses on voiced regions highly related to style while maintaining continuity across different speech regions to improve expressiveness. We adjust the direction of the extracted style for optimal integration into the TTS model, which improves speech quality. Experimental results demonstrate that Spotlight-TTS achieves superior performance compared to baseline models in terms of expressiveness, overall speech quality, and style transfer capability. Our audio samples are publicly available.


【18】 VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing  Voice Conversion

标题: VibE-SRC:利用高频F0轮廓进行颤音提取,用于歌唱声音转换
链接:https://arxiv.org/abs/2505.20794
作者: Joon-Seung Choi,  Dong-Min Byun,  Hyung-Seok Oh,  Seong-Whan Lee 
备注:Proceedings of Interspeech 2025
摘要:掌握歌唱风格是获得富有表现力和自然的歌唱声音的关键。在各种风格因素中,颤音在传达情感和增强音乐深度方面起着关键作用。然而,建模颤音仍然具有挑战性,由于其动态性质,使其难以控制在歌唱声音转换。为了解决这个问题,我们提出了VibESVC,一个可控的歌声转换模型,明确提取和操纵颤音,使用离散小波变换。与以前隐式建模颤音的方法不同,我们的方法将F0轮廓分解为频率分量,从而实现精确的传输。这允许颤音控制增强的灵活性。实验结果表明,VibE-SVC在保持说话人相似性的同时,有效地变换了演唱风格。主观和客观评价都证实了高质量的转换。
摘要:Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propose VibESVC, a controllable singing voice conversion model that explicitly extracts and manipulates vibrato using discrete wavelet transform. Unlike previous methods that model vibrato implicitly, our approach decomposes the F0 contour into frequency components, enabling precise transfer. This allows vibrato control for enhanced flexibility. Experimental results show that VibE-SVC effectively transforms singing styles while preserving speaker similarity. Both subjective and objective evaluations confirm high-quality conversion.


【19】 Can Large Language Models Predict Audio Effects Parameters from Natural  Language?

标题: 大型语言模型可以从自然语言预测音效参数吗?
链接:https://arxiv.org/abs/2505.20770
作者: Seungheon Doh,  Junghyun Koo,  Marco A. Martínez-Ramírez,  Wei-Hsiang Liao,  Juhan Nam,  Yuki Mitsufuji 
备注:Submitted to WASPAA 2025
摘要:在音乐制作中,通过自然语言操纵音频效果(Fx)参数有可能减少非专家的技术障碍。我们提出了LLM2Fx,这是一个利用大型语言模型(LLM)直接从文本描述中预测Fx参数的框架,而不需要特定于任务的训练或微调。我们的方法通过将自然语言描述映射到相应的Fx参数来进行均衡和混响,从而解决文本效果参数预测(Text2Fx)任务。我们证明,LLM可以产生Fx参数在一个zero-shot的方式,阐明在音乐制作中的音色语义和音频效果之间的关系。为了提高性能,我们介绍了三种类型的上下文示例:音频数字信号处理(DSP)功能,DSP功能代码,和Few-Shot的例子。我们的研究结果表明,基于LLM的Fx参数生成优于以前的优化方法,在将自然语言描述翻译为适当的Fx设置方面提供了有竞争力的性能。此外,LLM可以作为音频制作的文本驱动界面,为更直观和更易于访问的音乐制作工具铺平了道路。
摘要:In music production, manipulating audio effects (Fx) parameters through natural language has the potential to reduce technical barriers for non-experts. We present LLM2Fx, a framework leveraging Large Language Models (LLMs) to predict Fx parameters directly from textual descriptions without requiring task-specific training or fine-tuning. Our approach address the text-to-effect parameter prediction (Text2Fx) task by mapping natural language descriptions to the corresponding Fx parameters for equalization and reverberation. We demonstrate that LLMs can generate Fx parameters in a zero-shot manner that elucidates the relationship between timbre semantics and audio effects in music production. To enhance performance, we introduce three types of in-context examples: audio Digital Signal Processing (DSP) features, DSP function code, and few-shot examples. Our results demonstrate that LLM-based Fx parameter generation outperforms previous optimization approaches, offering competitive performance in translating natural language descriptions to appropriate Fx settings. Furthermore, LLMs can serve as text-driven interfaces for audio production, paving the way for more intuitive and accessible music production tools.


【20】 Foundation Model Hidden Representations for Heart Rate Estimation from  Auscultation

标题: 用于听诊心率估计的基础模型隐藏表示
链接:https://arxiv.org/abs/2505.20745
作者: Jingping Nie,  Dung T. Tran,  Karan Thakkar,  Vasudha Kowtha,  John Huang,  Carlos Avendano,  Erdrin Azemi,  Vikramjit Mitra 
备注:5 pages, Interspeech 2025 conference
摘要:听诊,特别是心音,是一种提供重要生命体征信息的非侵入性技术。最近,已经提出了自监督声学表示基础模型(FM),以提供对基于声学的生命体征的见解。然而,很少有人探索听诊在这些预先训练的FM表示中编码的程度。在这项工作中,使用公开可用的心音图(PCG)数据集和心率(HR)估计模型,我们对六种声学表示FM进行了逐层调查:HuBERT,wav 2 vec 2,wavLM,Whisper,对比听觉音频预训练(CLAP)和内部CLAP模型。此外,我们实现了Nie等人的基线方法,2024(依赖于声学特征),并表明总体而言,来自预训练基础模型(FM)的表示向量提供了与基线相当的性能。值得注意的是,使用来自内部CLAP模型的音频编码器的表示的HR估计优于从基线获得的结果,尽管存在域失配,但在各种训练/验证/测试分割上实现了较低的平均绝对误差(MAE)。
摘要:Auscultation, particularly heart sound, is a non-invasive technique that provides essential vital sign information. Recently, self-supervised acoustic representation foundation models (FMs) have been proposed to offer insights into acoustics-based vital signs. However, there has been little exploration of the extent to which auscultation is encoded in these pre-trained FM representations. In this work, using a publicly available phonocardiogram (PCG) dataset and a heart rate (HR) estimation model, we conduct a layer-wise investigation of six acoustic representation FMs: HuBERT, wav2vec2, wavLM, Whisper, Contrastive Language-Audio Pretraining (CLAP), and an in-house CLAP model. Additionally, we implement the baseline method from Nie et al., 2024 (which relies on acoustic features) and show that overall, representation vectors from pre-trained foundation models (FMs) offer comparable performance to the baseline. Notably, HR estimation using the representations from the audio encoder of the in-house CLAP model outperforms the results obtained from the baseline, achieving a lower mean absolute error (MAE) across various train/validation/test splits despite the domain mismatch.


【21】 Uni-VERSA: Versatile Speech Assessment with a Unified Network

标题: Uni-VERSA:具有统一网络的多功能语音评估
链接:https://arxiv.org/abs/2505.20741
作者: Jiatong Shi,  Hye-Jin Shim,  Shinji Watanabe 
备注:Accepted by Interspeech
摘要:主观听力测试仍然是语音质量评估的黄金标准,但成本高,变化大,难以衡量。相比之下,现有的客观指标,如PESQ,F0相关性,和DNSMOS,通常只捕捉语音质量的特定方面。为了解决这些限制,我们引入了Uni-VERSA,这是一个统一的网络,可以同时预测各种客观指标,包括自然度,可懂度,说话人特征,韵律和噪声,用于语音信号的综合评估。我们正式的框架,评估协议,并在语音增强,合成和质量控制的应用。基于URGENT 24挑战的基准测试,以及利用自我监督表示的基线,表明Uni-VERSA为单方面评估方法提供了一种可行的替代方案。此外,它与人类的感知密切相关,使其成为未来语音质量评估的一种有前途的方法。
摘要:Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only specific aspects of speech quality. To address these limitations, we introduce Uni-VERSA, a unified network that simultaneously predicts various objective metrics, encompassing naturalness, intelligibility, speaker characteristics, prosody, and noise, for a comprehensive evaluation of speech signals. We formalize its framework, evaluation protocol, and applications in speech enhancement, synthesis, and quality control. A benchmark based on the URGENT24 challenge, along with a baseline leveraging self-supervised representations, demonstrates that Uni-VERSA provides a viable alternative to single-aspect evaluation methods. Moreover, it aligns closely with human perception, making it a promising approach for future speech quality assessment.


【22】 Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent  Speech in Low-Resource Indian Languages

标题: 菲尔·赫拉·费尔(Phir Hera Fairy):一位英国童话演员是低资源印度语言流利演讲的忠实伪造者
链接:https://arxiv.org/abs/2505.20693
作者: Praveen Srinivasa Varadhan,  Srija Anand,  Soma Siddhartha,  Mitesh M.Khapra 
摘要:当一个英国童话作家对印度语言进行微调时会发生什么?我们评估英语F5-TTS模型如何适应11种印度语言,测量多语流利性,语音克隆,风格克隆和代码混合。我们比较:(i)从头开始训练,(ii)在印度数据上微调英语F5,以及(iii)在印度和英语数据上微调以防止遗忘。仅使用印度数据进行微调被证明是最有效的,并且由此产生的IN-F5是一种接近人类的多语言;这使得一种语言的使用者(例如,Odia)流利地用另一种语言说话(例如,印地语)。我们的研究结果表明,英语预培训艾滋病低资源TTS在达到人类平等。为了帮助其他低资源语言的进步,我们研究了数据约束的设置,并得出了一个计算最优策略。最后,我们展示了IN-F5可以通过合成数据生成,使用零资源TTS的人在回路方法合成看不见的语言,如Bhojpuri和Tulu。
摘要:What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing. We compare: (i) training from scratch, (ii) fine-tuning English F5 on Indian data, and (iii) fine-tuning on both Indian and English data to prevent forgetting. Fine-tuning with only Indian data proves most effective and the resultant IN-F5 is a near-human polyglot; that enables speakers of one language (e.g., Odia) to fluently speak in another (e.g., Hindi). Our results show English pretraining aids low-resource TTS in reaching human parity. To aid progress in other low-resource languages, we study data-constrained setups and arrive at a compute optimal strategy. Finally, we show IN-F5 can synthesize unseen languages like Bhojpuri and Tulu using a human-in-the-loop approach for zero-resource TTS via synthetic data generation.


【23】 Music's Multimodal Complexity in AVQA: Why We Need More than General  Multimodal LLMs

标题: AVQA中音乐的多模式复杂性:为什么我们需要的不仅仅是一般的多模式LLM
链接:https://arxiv.org/abs/2505.20638
作者: Wenhao You,  Xingjian Diao,  Chunhui Zhang,  Keyi Kong,  Weiyi Wu,  Zhongyu Ouyang,  Chiyu Ma,  Tingxuan Wu,  Noah Wei,  Zong Ke,  Ming Cheng,  Soroush Vosoughi,  Jiang Gui 
摘要:虽然最近的多模态大型语言模型在一般多模态任务方面表现出令人印象深刻的能力,但像音乐这样的专业领域需要量身定制的方法。音乐视听问题问答(Music AVQA)特别强调了这一点,其连续,密集分层的视听内容,复杂的时间动态以及对特定领域知识的迫切需求带来了独特的挑战。通过对音乐AVQA数据集和方法的系统分析,本立场文件确定了专门的输入处理,包含专用时空设计的架构以及特定于音乐的建模策略对于该领域的成功至关重要。我们的研究为研究人员提供了有价值的见解,突出了有效的设计模式经验联系到强大的性能,提出了具体的未来方向,将音乐先验,并旨在建立一个强大的基础,推进多模态音乐的理解。这项工作旨在激发更广泛的关注和进一步的研究,并得到不断更新的匿名GitHub相关论文库的支持:https://github.com/xid32/Survey4MusicAVQA。
摘要:While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Audio-Visual Question Answering (Music AVQA) particularly underscores this, presenting unique challenges with its continuous, densely layered audio-visual content, intricate temporal dynamics, and the critical need for domain-specific knowledge. Through a systematic analysis of Music AVQA datasets and methods, this position paper identifies that specialized input processing, architectures incorporating dedicated spatial-temporal designs, and music-specific modeling strategies are critical for success in this domain. Our study provides valuable insights for researchers by highlighting effective design patterns empirically linked to strong performance, proposing concrete future directions for incorporating musical priors, and aiming to establish a robust foundation for advancing multimodal musical understanding. This work is intended to inspire broader attention and further research, supported by a continuously updated anonymous GitHub repository of relevant papers: https://github.com/xid32/Survey4MusicAVQA.


【24】 Training Articulatory Inversion Models for Inter-Speaker Consistency

标题: 训练发音倒置模型以实现说话者间一致性
链接:https://arxiv.org/abs/2505.20529
作者: Charles McGhee,  Mark J.F. Gales,  Kate M. Knill 
摘要:声学到发音反转(AAI)试图模拟从语音到发音的逆映射。仅仅从语音中准确地预测发音可能是不可能的,因为说话者可以选择不同的发音形式,而似乎不需要参考他们的声道结构。然而,一旦说话者选择了一种发音形式,他们的产出变化最小。AAI最近的工作提出了将自监督学习(SSL)模型适应于单说话者数据集,声称这些单说话者模型提供了一个通用的发音模板。在本文中,我们调查是否SSL适应模型训练的单和多扬声器数据产生发音目标,这是一致的跨扬声器身份的英语和俄语。我们这样做,通过使用一种新的评价方法,提取发音目标,使用最小对集。我们还提出了一种训练方法,可以提高说话人之间的一致性,只用语音数据。
摘要:Acoustic-to-Articulatory Inversion (AAI) attempts to model the inverse mapping from speech to articulation. Exact articulatory prediction from speech alone may be impossible, as speakers can choose different forms of articulation seemingly without reference to their vocal tract structure. However, once a speaker has selected an articulatory form, their productions vary minimally. Recent works in AAI have proposed adapting Self-Supervised Learning (SSL) models to single-speaker datasets, claiming that these single-speaker models provide a universal articulatory template. In this paper, we investigate whether SSL-adapted models trained on single and multi-speaker data produce articulatory targets which are consistent across speaker identities for English and Russian. We do this through the use of a novel evaluation method which extracts articulatory targets using minimal pair sets. We also present a training method which can improve inter-speaker consistency using only speech data.


【25】 ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

标题: ArVoice:用于阿拉伯语语音合成的多说话人数据集
链接:https://arxiv.org/abs/2505.20506
作者: Hawau Olamide Toyin,  Rufael Marew,  Humaid Alblooshi,  Samar M. Magdy,  Hanan Aldarmaki 
备注:Accepted at INTERSPEECH 2025 The dataset is available at this https URL
摘要:我们介绍了ArVoice,一个多扬声器现代标准阿拉伯语(MSA)语音语料库,带有变音音符,用于多扬声器语音合成,并可用于其他任务,如基于语音的变音符号恢复,语音转换和deepfake检测。ArVoice包括:(1)由六位不同人口统计学特征的语音人才组成的新的专业录制集,(2)阿拉伯语语音语料库的修改子集;以及(3)来自两个商业系统的高质量合成语音。完整的语料库包括11种声音的83.52小时的语音;大约10小时由7个说话者的人类声音组成。我们训练三个开源TTS和两个语音转换系统来说明数据集的用例。语料库可供研究使用。
摘要:We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice comprises: (1) a new professionally recorded set from six voice talents with diverse demographics, (2) a modified subset of the Arabic Speech Corpus; and (3) high-quality synthetic speech from two commercial systems. The complete corpus consists of a total of 83.52 hours of speech across 11 voices; around 10 hours consist of human voices from 7 speakers. We train three open-source TTS and two voice conversion systems to illustrate the use cases of the dataset. The corpus is available for research use.


【26】 PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems

标题: PSRB:评估波斯ASR系统的综合基准
链接:https://arxiv.org/abs/2505.21230
作者: Nima Sedghiyeh,  Sara Sadeghi,  Reza Khodadadi,  Farzin Kashani,  Omid Aghdaei,  Somayeh Rahimi,  Mohammad Sadegh Safari 
备注:25 pages, 7 figures
摘要:虽然自动语音识别(ASR)系统已经成为现代技术的一个组成部分,但其评估仍然具有挑战性,特别是对于波斯语等低资源语言。本文介绍了波斯语语音识别基准(PSRB),一个全面的基准,旨在解决这一差距,结合不同的语言和声学条件。我们评估了10个ASR系统,包括最先进的商业和开源模型,以检查性能变化和固有的偏见。此外,我们进行了深入的分析波斯语ASR transmittance,确定关键的错误类型,并提出了一个新的度量,权重替代错误。该指标通过减少微小和部分错误的影响,从而提高性能评估的精确度,增强了评估的鲁棒性。我们的研究结果表明,虽然ASR模型通常在标准波斯语上表现良好,但它们在区域口音,儿童语音和特定语言挑战方面表现不佳。这些结果强调了微调和整合多样化的、有代表性的训练数据集的必要性,以减轻偏差并提高整体ASR性能。PSRB为推进波斯语的ASR研究提供了宝贵的资源,并作为开发其他低资源语言基准的框架。PSRB数据集的一个子集可在https://huggingface.co/datasets/PartAI/PSRB上公开获得。
摘要:Although Automatic Speech Recognition (ASR) systems have become an integral part of modern technology, their evaluation remains challenging, particularly for low-resource languages such as Persian. This paper introduces Persian Speech Recognition Benchmark(PSRB), a comprehensive benchmark designed to address this gap by incorporating diverse linguistic and acoustic conditions. We evaluate ten ASR systems, including state-of-the-art commercial and open-source models, to examine performance variations and inherent biases. Additionally, we conduct an in-depth analysis of Persian ASR transcriptions, identifying key error types and proposing a novel metric that weights substitution errors. This metric enhances evaluation robustness by reducing the impact of minor and partial errors, thereby improving the precision of performance assessment. Our findings indicate that while ASR models generally perform well on standard Persian, they struggle with regional accents, children's speech, and specific linguistic challenges. These results highlight the necessity of fine-tuning and incorporating diverse, representative training datasets to mitigate biases and enhance overall ASR performance. PSRB provides a valuable resource for advancing ASR research in Persian and serves as a framework for developing benchmarks in other low-resource languages. A subset of the PSRB dataset is publicly available at https://huggingface.co/datasets/PartAI/PSRB.


【27】 Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and  Machine Learning Approaches

标题: 使用视听和机器学习方法对ALS言语障碍进行多模式评估
链接:https://arxiv.org/abs/2505.21093
作者: Francesco Pierotti,  Andrea Bandini 
备注:Submitted to Interspeech

摘要:肌萎缩侧索硬化症患者的言语分析是临床医生评估延髓功能障碍的有力工具。然而,目前在临床实践中使用的方法包括主观评价或昂贵的仪器。本研究探讨了结合视听分析和机器学习来预测临床医生进行的言语障碍评估的不同方法。使用从语音任务的音频和视频记录中提取的声学和运动学特征的小数据集,我们训练和测试了一些回归模型。使用具有多峰特征的极端提升机回归器实现了最佳性能,这导致在5至25的范围内的均方根误差为0.93。结果表明,

整合音视频分析增强了言语障碍评估,提供了一个客观的工具,早期检测和监测延髓功能障碍,也在家庭环境。

摘要:The analysis of speech in individuals with amyotrophic lateral sclerosis is a powerful tool to support clinicians in the assessment of bulbar dysfunction. However, current methods used in clinical practice consist of subjective evaluations or expensive instrumentation. This study investigates different approaches combining audio-visual analysis and machine learning to predict the speech impairment evaluation performed by clinicians. Using a small dataset of acoustic and kinematic features extracted from audio and video recordings of speech tasks, we trained and tested some regression models. The best performance was achieved using the extreme boosting machine regressor with multimodal features, which resulted in a root mean squared error of 0.93 on a scale ranging from 5 to 25. Results suggest that integrating audio-video analysis enhances speech impairment assessment, providing an objective tool for early detection and monitoring of bulbar dysfunction, also in home settings.


【28】 Study of Lightweight Transformer Architectures for Single-Channel Speech  Enhancement

标题: 用于单通道语音增强的轻量级Transformer结构研究
链接:https://arxiv.org/abs/2505.21057
作者: Haixin Zhao,  Nilesh Madhu 
备注:Accepted by EUSIPCO 2025
摘要:在语音增强中,实现最先进的(SotA)性能,同时遵守边缘设备上的计算约束仍然是一个艰巨的挑战。集成堆叠的时间和频谱建模的网络有效地利用了改进的架构,如Transformers;然而,它们不可避免地导致大量的计算复杂性和模型扩展。通过对基于变换器的时间和频谱建模的系统消融分析,我们证明了采用流线型频率-时间-频率(FTF)堆叠Transformers的架构有效地学习了因果背景下的全局依赖关系,同时避免了相当大的计算需求。在训练中利用鉴别器进一步提高了学习效率和增强,而不会在推理过程中引入额外的复杂性。提出的轻量级,因果,基于transformer的对抗训练架构(LCT-GAN)在当代轻量级模型中的工具指标上产生SoTA性能,但开销要小得多。与DeepFilterNet 2相比,LCT-GAN只需要6%的参数,复杂度和性能相似。与CCFNet+(Lite)相比,LCT-GAN节省了9%的参数和10%的乘法累加运算,同时提高了性能。此外,LCT-GAN在广泛使用的测试数据集上的表现甚至优于更复杂的常见基线模型。
摘要:In speech enhancement, achieving state-of-the-art (SotA) performance while adhering to the computational constraints on edge devices remains a formidable challenge. Networks integrating stacked temporal and spectral modelling effectively leverage improved architectures such as transformers; however, they inevitably incur substantial computational complexity and model expansion. Through systematic ablation analysis on transformer-based temporal and spectral modelling, we demonstrate that the architecture employing streamlined Frequency-Time-Frequency (FTF) stacked transformers efficiently learns global dependencies within causal context, while avoiding considerable computational demands. Utilising discriminators in training further improves learning efficacy and enhancement without introducing additional complexity during inference. The proposed lightweight, causal, transformer-based architecture with adversarial training (LCT-GAN) yields SoTA performance on instrumental metrics among contemporary lightweight models, but with far less overhead. Compared to DeepFilterNet2, the LCT-GAN only requires 6% of the parameters, at similar complexity and performance. Against CCFNet+(Lite), LCT-GAN saves 9% in parameters and 10% in multiply-accumulate operations yet yielding improved performance. Further, the LCT-GAN even outperforms more complex, common baseline models on widely used test datasets.


【29】 REWIND: Speech Time Reversal for Enhancing Speaker Representations in  Diffusion-based Voice Conversion

标题: REWIND:用于增强基于扩散的语音转换中说话者表示的语音时间限制器
链接:https://arxiv.org/abs/2505.20756
作者: Ishan D. Biyani,  Nirmesh J. Shah,  Ashishkumar P. Gudmalwar,  Pankaj Wasnik,  Rajiv R. Shah 
备注:Accepted in INTERSPEECH 2025
摘要:语音时间反转是指在时间上反转整个语音信号,使其向后播放的过程。这样的信号是完全无法理解的,因为音素和音节的基本结构被破坏了。然而,尽管失去了语言内容,它们仍然保留了能够感知说话者识别的音调模式。在本文中,我们提出利用从时间反转语音中学习的说话人表示作为增强策略来增强说话人表示。值得注意的是,语音转换(VC)中的说话者和语言解纠缠对于准确保留说话者独特的声音特征同时最大限度地减少语言内容的干扰至关重要。所提出的方法的有效性进行评估的背景下,最先进的扩散为基础的VC模型。实验结果表明,该方法显着提高说话人相似性相关的分数,同时保持高的语音质量。
摘要:Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables are destroyed. However, they still retain tonal patterns that enable perceptual speaker identification despite losing linguistic content. In this paper, we propose leveraging speaker representations learned from time reversed speech as an augmentation strategy to enhance speaker representation. Notably, speaker and language disentanglement in voice conversion (VC) is essential to accurately preserve a speaker's unique vocal traits while minimizing interference from linguistic content. The effectiveness of the proposed approach is evaluated in the context of state-of-the-art diffusion-based VC models. Experimental results indicate that the proposed approach significantly improves speaker similarity-related scores while maintaining high speech quality.


【30】 PromptEVC: Controllable Emotional Voice Conversion with Natural Language  Prompts

标题: EVC:可控的情感语音转换与自然语言的转换
链接:https://arxiv.org/abs/2505.20678
作者: Tianhua Qi,  Shiyan Wang,  Cheng Lu,  Tengfei Song,  Hao Yang,  Zhanglin Wu,  Wenming Zheng 
备注:Accepted to INTERSPEECH2025
摘要:可控情绪语音转换(EVC)的目的是操纵情绪表达,以增加合成语音的多样性。现有的方法通常依赖于预定义的标签,参考音频或预先指定的因素值,通常忽略了情绪感知和表达的个体差异。在本文中,我们介绍了利用自然语言提示精确和灵活的情感控制的EVC。为了将文本描述与情感语音连接起来,我们提出了情感描述符和提示映射器来生成细粒度的情感嵌入,并与参考嵌入一起训练。为了增强自然性,我们提出了一个韵律建模和控制管道,根据语言内容和情感线索调整节奏。此外,扬声器编码器被合并以保持身份。实验结果表明,该方法在情感转换、强度控制、混合情感合成和韵律处理等方面优于现有的可控EVC方法。语音样本可在https://jeremychee4.github.io/PromptEVC/上获得。
摘要:Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecified factor values, often overlooking individual differences in emotion perception and expression. In this paper, we introduce PromptEVC that utilizes natural language prompts for precise and flexible emotion control. To bridge text descriptions with emotional speech, we propose emotion descriptor and prompt mapper to generate fine-grained emotion embeddings, trained jointly with reference embeddings. To enhance naturalness, we present a prosody modeling and control pipeline that adjusts the rhythm based on linguistic content and emotional cues. Additionally, a speaker encoder is incorporated to preserve identity. Experimental results demonstrate that PromptEVC outperforms state-of-the-art controllable EVC methods in emotion conversion, intensity control, mixed emotion synthesis, and prosody manipulation. Speech samples are available at https://jeremychee4.github.io/PromptEVC/.


【31】 Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual  Speaker Extraction

标题: 基于即插即用的人脸注意力鲁棒视听说话人提取
链接:https://arxiv.org/abs/2505.20635
作者: Zexu Pan,  Shengkui Zhao,  Tingting Wang,  Kun Zhou,  Yukun Ma,  Chong Zhang,  Bin Ma 
备注:Interspeech 2025
摘要:视听说话人提取将目标说话人的语音从以视觉提示为条件的混合语音信号中分离出来,通常使用目标说话人的面部记录。然而,在现实世界的场景中,其他共同出现的面孔往往出现在屏幕上,提供有价值的扬声器活动线索的场景。在这项工作中,我们引入了一个即插即用的说话者间注意力模块来处理这些灵活数量的共同出现的面孔,从而在复杂的多人环境中实现更准确的说话者提取。我们将我们的模块集成到两个突出的模型中:AV-DPRNN和最先进的AV-TFGridNet。在不同数据集上进行的广泛实验,包括高度重叠的VoxCeleb 2和稀疏重叠的MISP,表明我们的方法始终优于基线。此外,LRS 2和LRS 3上的跨数据集评估证实了我们方法的鲁棒性和通用性。
摘要:Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are often present on-screen, providing valuable speaker activity cues in the scene. In this work, we introduce a plug-and-play inter-speaker attention module to process these flexible numbers of co-occurring faces, allowing for more accurate speaker extraction in complex multi-person environments. We integrate our module into two prominent models: the AV-DPRNN and the state-of-the-art AV-TFGridNet. Extensive experiments on diverse datasets, including the highly overlapped VoxCeleb2 and sparsely overlapped MISP, demonstrate that our approach consistently outperforms baselines. Furthermore, cross-dataset evaluations on LRS2 and LRS3 confirm the robustness and generalizability of our method.


【32】 Techniques for Quantum-Computing-Aided Algorithmic Composition:  Experiments in Rhythm, Timbre, Harmony, and Space

标题: 量子计算辅助数学作曲技术:节奏、音色、和声和空间实验
链接:https://arxiv.org/abs/2505.20565
作者: Christopher Dobrian,  Omar Costa Hamido 
摘要:量子计算可以用于计算机辅助音乐创作,以控制不同结构层次的音乐的各种属性。本文介绍了应用量子模拟模型组成的决策,模拟量子粒子跟踪产生噪声为基础的音色,使用基态矢量旋转,以引起颗粒谐波纹理的概率行为的变化,以及利用量子测量误差,造成嘈杂的扰动空间声音路径。我们描述了这些技术的基本概念,我们提供了算法和软件制定他们,我们提供的例子,展示他们在计算机生成的音乐的实现。
摘要:Quantum computing can be employed in computer-aided music composition to control various attributes of the music at different structural levels. This article describes the application of quantum simulation to model compositional decision making, the simulation of quantum particle tracking to produce noise-based timbres, the use of basis state vector rotation to cause changing probabilistic behaviors in granular harmonic textures, and the exploitation of quantum measurement error to cause noisy perturbations of spatial soundpaths. We describe the concepts fundamental to these techniques, we provide algorithms and software enacting them, and we provide examples demonstrating their implementation in computer-generated music.


【33】 Effect of laboratory conditions on the perception of virtual stages for  music

标题: 实验室条件对音乐虚拟舞台感知的影响
链接:https://arxiv.org/abs/2505.20552
作者: Ernesto Accolti 
摘要:这份手稿提出了支持定制听力室中增强声学实验的关键初步发现,解决了在这些高度敏感的设置中确保感知有效性和实验严谨性的关键挑战。这种验证确保了我们提出的方法是合理的,保证了未来结果的可靠性,并为后续的感知研究奠定了基础,并为虚拟声学研究中的实验室设计制定了可靠的指导方针。三个不同的房间的声学条件对音乐的虚拟舞台的感知效果的初步研究:一个消声室,一个定制的听力舱与不足的声音吸收,和另一个定制的听力舱与可实现的声音吸收。本研究的目的是评估这些不同的条件对音乐虚拟舞台的感知的影响。结果表明,消声室和可实现吸声的听力室的总声音和虚拟声音之间的差异低于刚可察觉的差异,这意味着虚拟声音不会被感知到比它应该的更大。相比之下,吸声不足的听力室具有高于刚可察觉的差异的差异,这意味着虚拟声音被感知为比它应该的声音更大。这项研究提供了一个初步的验证所提出的方法来评估定制的听力箱在舞台声学实验的声学条件。未来的工作将包括对结果进行更全面的分析,包括不同声源的影响。
摘要:This manuscript presents initial findings critical for supporting augmented acoustics experiments in custom-made hearing booths, addressing a key challenge in ensuring perceptual validity and experimental rigor in these highly sensitive setups. This validation ensures our proposed methodology is sound, guarantees the reliability of future results, and lays the foundational groundwork for subsequent perceptual studies and the development of robust guidelines for laboratory design in virtual acoustics research. A preliminary study on the effect of the acoustical conditions of three different rooms on the perception of virtual stages for music is presented: an anechoic room, a custom-made hearing booth with insufficient sound absorption, and another custom-made hearing booth with achievable sound absorption. The goal of this study is to assess the impact of these different conditions on the perception of virtual stages for music. The results show that the anechoic room and the hearing booth with achievable sound absorption have a difference between the total sound and the virtual sound below the just-noticeable difference, which means that the virtual sound is not perceived louder than it should. In contrast, the hearing booth with insufficient sound absorption has a difference above the just-noticeable difference, which means that the virtual sound is perceived louder than it should. This study provides a preliminary validation of the proposed methodology for assessing the acoustical conditions of custom-made hearing booths in stage acoustics experiments. Future work will include a more comprehensive analysis of the results, including the effect of different sound sources.


【34】 ReverbFX: A Dataset of Room Impulse Responses Derived from Reverb Effect  Plugins for Singing Voice Dereverberation

标题: ReverbFX:从用于歌唱声音去回响的回响效应插件衍生的房间脉冲响应数据集
链接:https://arxiv.org/abs/2505.20533
作者: Julius Richter,  Till Svajda,  Timo Gerkmann 
备注:Submitted to ITG Conference on Speech Communication
摘要:我们提出了ReverbFX,一个新的房间脉冲响应(RIR)数据集,专为歌声去混响研究。与基于真实录制的RIR的现有数据集不同,ReverbFX具有从音乐制作中常用的各种混响音频效果插件捕获的各种RIR集合。我们使用提议的数据集进行了全面的实验,以基准测试受人工混响影响的歌唱录音去混响的挑战。我们使用ReverbFX训练了两个最先进的生成模型,并证明了使用插件派生的RIR训练的模型在人工混响场景中优于使用真实RIR训练的模型。
摘要:We present ReverbFX, a new room impulse response (RIR) dataset designed for singing voice dereverberation research. Unlike existing datasets based on real recorded RIRs, ReverbFX features a diverse collection of RIRs captured from various reverb audio effect plugins commonly used in music production. We conduct comprehensive experiments using the proposed dataset to benchmark the challenge of dereverberation of singing voice recordings affected by artificial reverbs. We train two state-of-the-art generative models using ReverbFX and demonstrate that models trained with plugin-derived RIRs outperform those trained on realistic RIRs in artificial reverb scenarios.


【35】 In-context learning capabilities of Large Language Models to detect  suicide risk among adolescents from speech transcripts

标题: 大型语言模型的背景学习能力,可从言语记录中检测青少年自杀风险
链接:https://arxiv.org/abs/2505.20491
作者: Filomene Roquefort,  Alexandre Ducorroy,  Rachid Riad 
备注:Accepted to Interspeech 2025
摘要:青少年自杀风险的早期检测至关重要,但受到当前评估的可扩展性挑战的阻碍。本文介绍了我们的方法,第一次言语健康挑战(SW1),其目的是通过言语分析评估中国青少年的自杀风险。由于语音匿名化的限制,我们专注于语言特征,利用大型语言模型(LLM)进行基于转录的分类。使用DSPy系统提示工程,我们开发了一个强大的上下文学习方法,优于传统的微调语言和声学标记。我们的系统在180多个提交中获得了第三和第四名,仅使用成绩单的准确率为0.68(F1=0.7)。消融分析显示,增加提示示例可改善性能(p=0.003),不同模型类型和尺寸的效果不同。这些发现推进了自动自杀风险评估,并证明了LLM在心理健康应用中的价值。
摘要:Early suicide risk detection in adolescents is critical yet hindered by scalability challenges of current assessments. This paper presents our approach to the first SpeechWellness Challenge (SW1), which aims to assess suicide risk in Chinese adolescents through speech analysis. Due to speech anonymization constraints, we focused on linguistic features, leveraging Large Language Models (LLMs) for transcript-based classification. Using DSPy for systematic prompt engineering, we developed a robust in-context learning approach that outperformed traditional fine-tuning on both linguistic and acoustic markers. Our systems achieved third and fourth places among 180+ submissions, with 0.68 accuracy (F1=0.7) using only transcripts. Ablation analyses showed that increasing prompt example improved performance (p=0.003), with varying effects across model types and sizes. These findings advance automated suicide risk assessment and demonstrate LLMs' value in mental health applications.


【36】 Robust fine-tuning of speech recognition models via model merging:  application to disordered speech

标题: 通过模型合并对语音识别模型进行稳健微调:应用于无序语音
链接:https://arxiv.org/abs/2505.20477
作者: Alexandre Ducorroy,  Rachid Riad 
备注:Accepted to Interspeech 2025
摘要:自动语音识别(ASR)已经与语音基础模型(SFM)的进步,但性能下降构音障碍语音由于可变性和有限的数据。这项研究作为提交给语音可访问性挑战的一部分,探讨了模型合并,以提高ASR泛化使用耳语作为基础SFM。我们比较了微调与单轨迹合并,合并来自一个微调路径的模型,以及多运行合并,合并独立训练的模型。我们最好的多运行合并方法实现了WER相对于经典微调的12%的相对减少,以及长形式音频的16.2%的相对减少,这是构音障碍ASR的主要损失因素。合并越来越多的模型带来了持续的收益,在低数据状态下保持有效,并在模型架构中推广。这些结果强调了模型合并作为一种易于复制的自适应方法,可以持续改善ASR,而无需额外的推理成本或超参数调整。
摘要:Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility challenge, explored model merging to improve ASR generalization using Whisper as the base SFM. We compared fine-tuning with single-trajectory merging, combining models from one fine-tuning path, and multi-run merging, merging independently trained models. Our best multi-run merging approach achieved a 12% relative decrease of WER over classic fine-tuning, and a 16.2% relative decrease on long-form audios, a major loss contributor in dysarthric ASR. Merging more and more models led to continuous gains, remained effective in low-data regimes, and generalized across model architectures. These results highlight model merging as an easily replicable adaptation method that consistently improves ASR without additional inference cost or hyperparameter tuning.


【37】 Towards Emotionally Consistent Text-Based Speech Editing: Introducing  EmoCorrector and The ECD-TSE Dataset

标题: 实现语音一致的基于文本的语音编辑:引入MIDI纠正器和ECD-PSE数据集
链接:https://arxiv.org/abs/2505.20341
作者: Rui Liu,  Pu Gao,  Jiatian Xi,  Berrak Sisman,  Carlos Busso,  Haizhou Li 
备注:INTERSPEECH2025. Code and audio examples: this https URL
摘要:基于文本的语音编辑(TSE)仅使用文本修改语音,消除了重新录制。然而,现有的TSE方法,主要集中在合成语音段的内容准确性和声学一致性,往往忽略了文本变化所引入的情感变化或不一致的问题。为了解决这个问题,我们提出了一种新的TSE后校正方案-。Corrector通过提取编辑文本的情感特征,检索具有匹配情感的语音样本,并合成与所需情感一致的语音,同时保留说话者的身份和质量,从而利用检索增强生成(RAG)。为了支持TSE中情绪一致性建模的训练和评估,我们开创了TSE(ECD-TSE)的基准情绪校正数据集。ECD-TSE的突出方面是它包含了$<$text,speech$>$配对数据,具有不同的文本变化和一系列情感表达。通过对ECD-TSE的主观和客观实验以及综合分析,证实了该算法在解决当前TSE方法中情感不一致性缺陷的同时,显著增强了预期情感的表达。代码和音频示例可在https://github.com/AI-S2-Lab/EmoCorrector上获得。
摘要:Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consistency of synthetic speech segments, and often overlook the emotional shifts or inconsistency issues introduced by text changes. To address this issue, we propose EmoCorrector, a novel post-correction scheme for TSE. EmoCorrector leverages Retrieval-Augmented Generation (RAG) by extracting the edited text's emotional features, retrieving speech samples with matching emotions, and synthesizing speech that aligns with the desired emotion while preserving the speaker's identity and quality. To support the training and evaluation of emotional consistency modeling in TSE, we pioneer the benchmarking Emotion Correction Dataset for TSE (ECD-TSE). The prominent aspect of ECD-TSE is its inclusion of $<$text, speech$>$ paired data featuring diverse text variations and a range of emotional expressions. Subjective and objective experiments and comprehensive analysis on ECD-TSE confirm that EmoCorrector significantly enhances the expression of intended emotion while addressing emotion inconsistency limitations in current TSE methods. Code and audio examples are available at https://github.com/AI-S2-Lab/EmoCorrector.


eess.AS音频处理


【1】 PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems

标题: PSRB:评估波斯ASR系统的综合基准
链接:https://arxiv.org/abs/2505.21230
作者: Nima Sedghiyeh,  Sara Sadeghi,  Reza Khodadadi,  Farzin Kashani,  Omid Aghdaei,  Somayeh Rahimi,  Mohammad Sadegh Safari 
备注:25 pages, 7 figures
摘要:虽然自动语音识别(ASR)系统已经成为现代技术的一个组成部分,但其评估仍然具有挑战性,特别是对于波斯语等低资源语言。本文介绍了波斯语语音识别基准(PSRB),一个全面的基准,旨在解决这一差距,结合不同的语言和声学条件。我们评估了10个ASR系统,包括最先进的商业和开源模型,以检查性能变化和固有的偏见。此外,我们进行了深入的分析波斯语ASR transmittance,确定关键的错误类型,并提出了一个新的度量,权重替代错误。该指标通过减少微小和部分错误的影响,从而提高性能评估的精确度,增强了评估的鲁棒性。我们的研究结果表明,虽然ASR模型通常在标准波斯语上表现良好,但它们在区域口音,儿童语音和特定语言挑战方面表现不佳。这些结果强调了微调和整合多样化的、有代表性的训练数据集的必要性,以减轻偏差并提高整体ASR性能。PSRB为推进波斯语的ASR研究提供了宝贵的资源,并作为开发其他低资源语言基准的框架。PSRB数据集的一个子集可在https://huggingface.co/datasets/PartAI/PSRB上公开获得。
摘要:Although Automatic Speech Recognition (ASR) systems have become an integral part of modern technology, their evaluation remains challenging, particularly for low-resource languages such as Persian. This paper introduces Persian Speech Recognition Benchmark(PSRB), a comprehensive benchmark designed to address this gap by incorporating diverse linguistic and acoustic conditions. We evaluate ten ASR systems, including state-of-the-art commercial and open-source models, to examine performance variations and inherent biases. Additionally, we conduct an in-depth analysis of Persian ASR transcriptions, identifying key error types and proposing a novel metric that weights substitution errors. This metric enhances evaluation robustness by reducing the impact of minor and partial errors, thereby improving the precision of performance assessment. Our findings indicate that while ASR models generally perform well on standard Persian, they struggle with regional accents, children's speech, and specific linguistic challenges. These results highlight the necessity of fine-tuning and incorporating diverse, representative training datasets to mitigate biases and enhance overall ASR performance. PSRB provides a valuable resource for advancing ASR research in Persian and serves as a framework for developing benchmarks in other low-resource languages. A subset of the PSRB dataset is publicly available at https://huggingface.co/datasets/PartAI/PSRB.


【2】 Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and  Machine Learning Approaches

标题: 使用视听和机器学习方法对ALS言语障碍进行多模式评估
链接:https://arxiv.org/abs/2505.21093
作者: Francesco Pierotti,  Andrea Bandini 
备注:Submitted to Interspeech
摘要:肌萎缩侧索硬化症患者的言语分析是临床医生评估延髓功能障碍的有力工具。然而,目前在临床实践中使用的方法包括主观评价或昂贵的仪器。本研究探讨了结合视听分析和机器学习来预测临床医生进行的言语障碍评估的不同方法。我们使用从语音任务的音频和视频记录中提取的声学和运动学特征的小数据集,训练和测试了一些回归模型。使用具有多峰特征的极端提升机回归器实现了最佳性能,这导致在5至25的范围内的均方根误差为0.93。结果表明,整合音视频分析增强了言语障碍评估,提供了一个客观的工具,早期检测和监测延髓功能障碍,也在家庭环境。
摘要:The analysis of speech in individuals with amyotrophic lateral sclerosis is a powerful tool to support clinicians in the assessment of bulbar dysfunction. However, current methods used in clinical practice consist of subjective evaluations or expensive instrumentation. This study investigates different approaches combining audio-visual analysis and machine learning to predict the speech impairment evaluation performed by clinicians. Using a small dataset of acoustic and kinematic features extracted from audio and video recordings of speech tasks, we trained and tested some regression models. The best performance was achieved using the extreme boosting machine regressor with multimodal features, which resulted in a root mean squared error of 0.93 on a scale ranging from 5 to 25. Results suggest that integrating audio-video analysis enhances speech impairment assessment, providing an objective tool for early detection and monitoring of bulbar dysfunction, also in home settings.


【3】 Study of Lightweight Transformer Architectures for Single-Channel Speech  Enhancement

标题: 用于单通道语音增强的轻量级Transformer结构研究
链接:https://arxiv.org/abs/2505.21057
作者: Haixin Zhao,  Nilesh Madhu 
备注:Accepted by EUSIPCO 2025
摘要:在语音增强中,实现最先进的(SotA)性能,同时遵守边缘设备上的计算约束仍然是一个艰巨的挑战。集成堆叠的时间和频谱建模的网络有效地利用了改进的架构,如Transformers;然而,它们不可避免地导致大量的计算复杂性和模型扩展。通过对基于变换器的时间和频谱建模的系统消融分析,我们证明了采用流线型频率-时间-频率(FTF)堆叠Transformers的架构有效地学习了因果背景下的全局依赖关系,同时避免了相当大的计算需求。在训练中利用鉴别器进一步提高了学习效率和增强,而不会在推理过程中引入额外的复杂性。提出的轻量级,因果,基于transformer的对抗训练架构(LCT-GAN)在当代轻量级模型中的工具指标上产生SoTA性能,但开销要小得多。与DeepFilterNet 2相比,LCT-GAN只需要6%的参数,复杂度和性能相似。与CCFNet+(Lite)相比,LCT-GAN节省了9%的参数和10%的乘法累加运算,同时提高了性能。此外,LCT-GAN在广泛使用的测试数据集上的表现甚至优于更复杂的常见基线模型。
摘要:In speech enhancement, achieving state-of-the-art (SotA) performance while adhering to the computational constraints on edge devices remains a formidable challenge. Networks integrating stacked temporal and spectral modelling effectively leverage improved architectures such as transformers; however, they inevitably incur substantial computational complexity and model expansion. Through systematic ablation analysis on transformer-based temporal and spectral modelling, we demonstrate that the architecture employing streamlined Frequency-Time-Frequency (FTF) stacked transformers efficiently learns global dependencies within causal context, while avoiding considerable computational demands. Utilising discriminators in training further improves learning efficacy and enhancement without introducing additional complexity during inference. The proposed lightweight, causal, transformer-based architecture with adversarial training (LCT-GAN) yields SoTA performance on instrumental metrics among contemporary lightweight models, but with far less overhead. Compared to DeepFilterNet2, the LCT-GAN only requires 6% of the parameters, at similar complexity and performance. Against CCFNet+(Lite), LCT-GAN saves 9% in parameters and 10% in multiply-accumulate operations yet yielding improved performance. Further, the LCT-GAN even outperforms more complex, common baseline models on widely used test datasets.


【4】 REWIND: Speech Time Reversal for Enhancing Speaker Representations in  Diffusion-based Voice Conversion

标题: REWIND:用于增强基于扩散的语音转换中说话者表示的语音时间限制器
链接:https://arxiv.org/abs/2505.20756
作者: Ishan D. Biyani,  Nirmesh J. Shah,  Ashishkumar P. Gudmalwar,  Pankaj Wasnik,  Rajiv R. Shah 
备注:Accepted in INTERSPEECH 2025
摘要:语音时间反转是指在时间上反转整个语音信号,使其向后播放的过程。这样的信号是完全无法理解的,因为音素和音节的基本结构被破坏了。然而,尽管失去了语言内容,它们仍然保留了能够感知说话者识别的音调模式。在本文中,我们提出利用从时间反转语音中学习的说话人表示作为增强策略来增强说话人表示。值得注意的是,语音转换(VC)中的说话者和语言解纠缠对于准确保留说话者独特的声音特征同时最大限度地减少语言内容的干扰至关重要。所提出的方法的有效性进行评估的背景下,最先进的扩散为基础的VC模型。实验结果表明,该方法显着提高说话人相似性相关的分数,同时保持高的语音质量。
摘要:Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables are destroyed. However, they still retain tonal patterns that enable perceptual speaker identification despite losing linguistic content. In this paper, we propose leveraging speaker representations learned from time reversed speech as an augmentation strategy to enhance speaker representation. Notably, speaker and language disentanglement in voice conversion (VC) is essential to accurately preserve a speaker's unique vocal traits while minimizing interference from linguistic content. The effectiveness of the proposed approach is evaluated in the context of state-of-the-art diffusion-based VC models. Experimental results indicate that the proposed approach significantly improves speaker similarity-related scores while maintaining high speech quality.


【5】 PromptEVC: Controllable Emotional Voice Conversion with Natural Language  Prompts

标题: EVC:可控的情感语音转换与自然语言的转换
链接:https://arxiv.org/abs/2505.20678
作者: Tianhua Qi,  Shiyan Wang,  Cheng Lu,  Tengfei Song,  Hao Yang,  Zhanglin Wu,  Wenming Zheng 
备注:Accepted to INTERSPEECH2025
摘要:可控情绪语音转换(EVC)的目的是操纵情绪表达,以增加合成语音的多样性。现有的方法通常依赖于预定义的标签,参考音频或预先指定的因素值,通常忽略了情绪感知和表达的个体差异。在本文中,我们介绍了利用自然语言提示精确和灵活的情感控制的EVC。为了将文本描述与情感语音连接起来,我们提出了情感描述符和提示映射器来生成细粒度的情感嵌入,并与参考嵌入一起训练。为了增强自然性,我们提出了一个韵律建模和控制管道,根据语言内容和情感线索调整节奏。此外,扬声器编码器被合并以保持身份。实验结果表明,该方法在情感转换、强度控制、混合情感合成和韵律处理等方面优于现有的可控EVC方法。语音样本可在https://jeremychee4.github.io/PromptEVC/上获得。
摘要:Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecified factor values, often overlooking individual differences in emotion perception and expression. In this paper, we introduce PromptEVC that utilizes natural language prompts for precise and flexible emotion control. To bridge text descriptions with emotional speech, we propose emotion descriptor and prompt mapper to generate fine-grained emotion embeddings, trained jointly with reference embeddings. To enhance naturalness, we present a prosody modeling and control pipeline that adjusts the rhythm based on linguistic content and emotional cues. Additionally, a speaker encoder is incorporated to preserve identity. Experimental results demonstrate that PromptEVC outperforms state-of-the-art controllable EVC methods in emotion conversion, intensity control, mixed emotion synthesis, and prosody manipulation. Speech samples are available at https://jeremychee4.github.io/PromptEVC/.


【6】 Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual  Speaker Extraction

标题: 基于即插即用的人脸注意力鲁棒视听说话人提取
链接:https://arxiv.org/abs/2505.20635
作者: Zexu Pan,  Shengkui Zhao,  Tingting Wang,  Kun Zhou,  Yukun Ma,  Chong Zhang,  Bin Ma 
备注:Interspeech 2025
摘要:视听说话人提取将目标说话人的语音从以视觉提示为条件的混合语音信号中分离出来,通常使用目标说话人的面部记录。然而,在现实世界的场景中,其他共同出现的面孔往往出现在屏幕上,提供有价值的扬声器活动线索的场景。在这项工作中,我们引入了一个即插即用的说话者间注意力模块来处理这些灵活数量的共同出现的面孔,从而在复杂的多人环境中实现更准确的说话者提取。我们将我们的模块集成到两个突出的模型中:AV-DPRNN和最先进的AV-TFGridNet。在不同数据集上进行的广泛实验,包括高度重叠的VoxCeleb 2和稀疏重叠的MISP,表明我们的方法始终优于基线。此外,LRS 2和LRS 3上的跨数据集评估证实了我们方法的鲁棒性和通用性。
摘要:Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are often present on-screen, providing valuable speaker activity cues in the scene. In this work, we introduce a plug-and-play inter-speaker attention module to process these flexible numbers of co-occurring faces, allowing for more accurate speaker extraction in complex multi-person environments. We integrate our module into two prominent models: the AV-DPRNN and the state-of-the-art AV-TFGridNet. Extensive experiments on diverse datasets, including the highly overlapped VoxCeleb2 and sparsely overlapped MISP, demonstrate that our approach consistently outperforms baselines. Furthermore, cross-dataset evaluations on LRS2 and LRS3 confirm the robustness and generalizability of our method.


【7】 Effect of laboratory conditions on the perception of virtual stages for  music

标题: 实验室条件对音乐虚拟舞台感知的影响
链接:https://arxiv.org/abs/2505.20552
作者: Ernesto Accolti 
摘要:这份手稿提出了支持定制听力室中增强声学实验的关键初步发现,解决了在这些高度敏感的设置中确保感知有效性和实验严谨性的关键挑战。这种验证确保了我们提出的方法是合理的,保证了未来结果的可靠性,并为后续的感知研究和虚拟声学研究实验室设计的稳健指南的制定奠定了基础。三个不同的房间的声学条件对音乐的虚拟舞台的感知效果的初步研究:一个消声室,一个定制的听力舱与不足的声音吸收,和另一个定制的听力舱与可实现的声音吸收。本研究的目的是评估这些不同的条件对音乐虚拟舞台的感知的影响。结果表明,消声室和可实现吸声的听力室的总声音和虚拟声音之间的差异低于刚可察觉的差异,这意味着虚拟声音不会被感知到比它应该的更大。相比之下,吸声不足的听力室具有高于刚可察觉的差异的差异,这意味着虚拟声音被感知为比它应该的声音更大。这项研究提供了一个初步的验证所提出的方法来评估定制的听力箱在舞台声学实验的声学条件。未来的工作将包括对结果进行更全面的分析,包括不同声源的影响。
摘要:This manuscript presents initial findings critical for supporting augmented acoustics experiments in custom-made hearing booths, addressing a key challenge in ensuring perceptual validity and experimental rigor in these highly sensitive setups. This validation ensures our proposed methodology is sound, guarantees the reliability of future results, and lays the foundational groundwork for subsequent perceptual studies and the development of robust guidelines for laboratory design in virtual acoustics research. A preliminary study on the effect of the acoustical conditions of three different rooms on the perception of virtual stages for music is presented: an anechoic room, a custom-made hearing booth with insufficient sound absorption, and another custom-made hearing booth with achievable sound absorption. The goal of this study is to assess the impact of these different conditions on the perception of virtual stages for music. The results show that the anechoic room and the hearing booth with achievable sound absorption have a difference between the total sound and the virtual sound below the just-noticeable difference, which means that the virtual sound is not perceived louder than it should. In contrast, the hearing booth with insufficient sound absorption has a difference above the just-noticeable difference, which means that the virtual sound is perceived louder than it should. This study provides a preliminary validation of the proposed methodology for assessing the acoustical conditions of custom-made hearing booths in stage acoustics experiments. Future work will include a more comprehensive analysis of the results, including the effect of different sound sources.


【8】 ReverbFX: A Dataset of Room Impulse Responses Derived from Reverb Effect  Plugins for Singing Voice Dereverberation

标题: ReverbFX:从用于歌唱声音去回响的回响效应插件衍生的房间脉冲响应数据集
链接:https://arxiv.org/abs/2505.20533
作者: Julius Richter,  Till Svajda,  Timo Gerkmann 
备注:Submitted to ITG Conference on Speech Communication
摘要:我们提出了ReverbFX,一个新的房间脉冲响应(RIR)数据集,专为歌声去混响研究。与基于真实录制的RIR的现有数据集不同,ReverbFX具有从音乐制作中常用的各种混响音频效果插件捕获的各种RIR集合。我们使用提议的数据集进行了全面的实验,以基准测试受人工混响影响的歌唱录音去混响的挑战。我们使用ReverbFX训练了两个最先进的生成模型,并证明了使用插件派生的RIR训练的模型在人工混响场景中优于使用真实RIR训练的模型。
摘要:We present ReverbFX, a new room impulse response (RIR) dataset designed for singing voice dereverberation research. Unlike existing datasets based on real recorded RIRs, ReverbFX features a diverse collection of RIRs captured from various reverb audio effect plugins commonly used in music production. We conduct comprehensive experiments using the proposed dataset to benchmark the challenge of dereverberation of singing voice recordings affected by artificial reverbs. We train two state-of-the-art generative models using ReverbFX and demonstrate that models trained with plugin-derived RIRs outperform those trained on realistic RIRs in artificial reverb scenarios.


【9】 In-context learning capabilities of Large Language Models to detect  suicide risk among adolescents from speech transcripts

标题: 大型语言模型的背景学习能力,可从言语记录中检测青少年自杀风险
链接:https://arxiv.org/abs/2505.20491
作者: Filomene Roquefort,  Alexandre Ducorroy,  Rachid Riad 
备注:Accepted to Interspeech 2025
摘要:青少年自杀风险的早期检测至关重要,但受到当前评估的可扩展性挑战的阻碍。本文介绍了我们的方法,第一次言语健康挑战(SW1),其目的是通过言语分析评估中国青少年的自杀风险。由于语音匿名化的限制,我们专注于语言特征,利用大型语言模型(LLM)进行基于转录的分类。使用DSPy系统提示工程,我们开发了一个强大的上下文学习方法,优于传统的微调语言和声学标记。我们的系统在180多个提交中获得了第三和第四名,仅使用成绩单的准确率为0.68(F1=0.7)。消融分析显示,增加提示示例可改善性能(p=0.003),不同模型类型和尺寸的效果不同。这些发现推进了自动自杀风险评估,并证明了LLM在心理健康应用中的价值。
摘要:Early suicide risk detection in adolescents is critical yet hindered by scalability challenges of current assessments. This paper presents our approach to the first SpeechWellness Challenge (SW1), which aims to assess suicide risk in Chinese adolescents through speech analysis. Due to speech anonymization constraints, we focused on linguistic features, leveraging Large Language Models (LLMs) for transcript-based classification. Using DSPy for systematic prompt engineering, we developed a robust in-context learning approach that outperformed traditional fine-tuning on both linguistic and acoustic markers. Our systems achieved third and fourth places among 180+ submissions, with 0.68 accuracy (F1=0.7) using only transcripts. Ablation analyses showed that increasing prompt example improved performance (p=0.003), with varying effects across model types and sizes. These findings advance automated suicide risk assessment and demonstrate LLMs' value in mental health applications.


【10】 Robust fine-tuning of speech recognition models via model merging:  application to disordered speech

标题: 通过模型合并对语音识别模型进行稳健微调:应用于无序语音
链接:https://arxiv.org/abs/2505.20477
作者: Alexandre Ducorroy,  Rachid Riad 
备注:Accepted to Interspeech 2025
摘要:自动语音识别(ASR)已经与语音基础模型(SFM)的进步,但性能下降构音障碍语音由于可变性和有限的数据。这项研究作为提交给语音可访问性挑战的一部分,探讨了模型合并,以提高ASR泛化使用耳语作为基础SFM。我们比较了微调与单轨迹合并,合并来自一个微调路径的模型,以及多运行合并,合并独立训练的模型。我们最好的多运行合并方法实现了WER相对于经典微调的12%的相对减少,以及长形式音频的16.2%的相对减少,这是构音障碍ASR的主要损失因素。合并越来越多的模型带来了持续的收益,在低数据状态下保持有效,并在模型架构中推广。这些结果强调了模型合并作为一种易于复制的自适应方法,可以持续改善ASR,而无需额外的推理成本或超参数调整。
摘要:Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility challenge, explored model merging to improve ASR generalization using Whisper as the base SFM. We compared fine-tuning with single-trajectory merging, combining models from one fine-tuning path, and multi-run merging, merging independently trained models. Our best multi-run merging approach achieved a 12% relative decrease of WER over classic fine-tuning, and a 16.2% relative decrease on long-form audios, a major loss contributor in dysarthric ASR. Merging more and more models led to continuous gains, remained effective in low-data regimes, and generalized across model architectures. These results highlight model merging as an easily replicable adaptation method that consistently improves ASR without additional inference cost or hyperparameter tuning.


【11】 Towards Emotionally Consistent Text-Based Speech Editing: Introducing  EmoCorrector and The ECD-TSE Dataset

标题: 实现语音一致的基于文本的语音编辑:引入MIDI纠正器和ECD-PSE数据集
链接:https://arxiv.org/abs/2505.20341
作者: Rui Liu,  Pu Gao,  Jiatian Xi,  Berrak Sisman,  Carlos Busso,  Haizhou Li 
备注:INTERSPEECH2025. Code and audio examples: this https URL
摘要:基于文本的语音编辑(TSE)仅使用文本修改语音,消除了重新录制。然而,现有的TSE方法,主要集中在合成语音段的内容准确性和声学一致性,往往忽略了文本变化所引入的情感变化或不一致的问题。为了解决这个问题,我们提出了一种新的TSE后校正方案-。Corrector通过提取编辑文本的情感特征,检索具有匹配情感的语音样本,并合成与所需情感一致的语音,同时保留说话者的身份和质量,从而利用检索增强生成(RAG)。为了支持TSE中情绪一致性建模的训练和评估,我们开创了TSE(ECD-TSE)的基准情绪校正数据集。ECD-TSE的突出方面是它包含了$<$text,speech$>$配对数据,具有不同的文本变化和一系列情感表达。通过对ECD-TSE的主观和客观实验以及综合分析,证实了该算法在解决当前TSE方法中情感不一致性缺陷的同时,显著增强了预期情感的表达。代码和音频示例可在https://github.com/AI-S2-Lab/EmoCorrector上获得。
摘要:Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consistency of synthetic speech segments, and often overlook the emotional shifts or inconsistency issues introduced by text changes. To address this issue, we propose EmoCorrector, a novel post-correction scheme for TSE. EmoCorrector leverages Retrieval-Augmented Generation (RAG) by extracting the edited text's emotional features, retrieving speech samples with matching emotions, and synthesizing speech that aligns with the desired emotion while preserving the speaker's identity and quality. To support the training and evaluation of emotional consistency modeling in TSE, we pioneer the benchmarking Emotion Correction Dataset for TSE (ECD-TSE). The prominent aspect of ECD-TSE is its inclusion of $<$text, speech$>$ paired data featuring diverse text variations and a range of emotional expressions. Subjective and objective experiments and comprehensive analysis on ECD-TSE confirm that EmoCorrector significantly enhances the expression of intended emotion while addressing emotion inconsistency limitations in current TSE methods. Code and audio examples are available at https://github.com/AI-S2-Lab/EmoCorrector.


【12】 Towards Robust Automated Perceptual Voice Quality Assessment with Deep  Learning

标题: 利用深度学习实现稳健的自动感知语音质量评估
链接:https://arxiv.org/abs/2505.21356
作者: Whenty Ariyanti,  Kuan-Yu Chen,  Sabato Marco Siniscalchi,  Hsin-Min Wang,  Yu Tsao 
摘要:目的:感知嗓音质量评估通过提供对发声功能的标准化评估,在嗓音疾病的诊断和监测中起着至关重要的作用。传统上,这个过程依赖于专家评分员使用标准量表,如声音的共识听觉感知评估(CAPE-V)和等级,粗糙度,呼吸,虚弱和紧张(GRBAS)。然而,这些指标本质上是主观的,容易受到评分者之间的变化,激发了对自动化和客观评估方法的需求。研究方法:我们提出了语音质量评估网络(VOQANet),这是一个基于深度学习的框架,具有注意力机制,利用语音基础模型(SFM)从原始语音中捕获高级声学和韵律信息。为了增强鲁棒性和可解释性,我们提出了VOQANet+,它集成了手工制作的声学特征,如抖动,闪烁和谐波噪声比(HNR)与SFM嵌入。结果:基于句子的输入比基于元音的输入产生更强的性能,特别是在病人的水平。VOQANet在RMSE和PCC方面始终优于基线方法,而VOQANet+在噪声条件下表现更好并保持鲁棒性。结论:将SFM嵌入与域信息声学特征相结合,提高了可解释性和弹性。重要性:VOQANet+显示出在现实世界和远程医疗环境中部署的强大潜力,通过可解释和抗噪声的解决方案解决了主观感知评估的局限性。
摘要:Objective: Perceptual voice quality assessment plays a critical role in diagnosing and monitoring voice disorders by providing standardized evaluation of vocal function. Traditionally, this process relies on expert raters utilizing standard scales, such as the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and Grade, Roughness, Breathiness, Asthenia, and Strain (GRBAS). However, these metrics are inherently subjective and susceptible to inter-rater variability, motivating the need for automated and objective assessment methods. Methods: We propose Voice Quality Assessment Network (VOQANet), a deep learning-based framework with an attention mechanism that leverages a Speech Foundation Model (SFM) to capture high-level acoustic and prosodic information from raw speech. To enhance robustness and interpretability, we present VOQANet+, which integrates handcrafted acoustic features such as jitter, shimmer, and harmonics-to-noise ratio (HNR) with SFM embeddings. Results: Sentence-based input yields stronger performance than vowel-based input, especially at the patient level. VOQANet consistently outperforms baseline methods in RMSE and PCC, while VOQANet+ performs even better and maintains robustness under noisy conditions. Conclusion: Combining SFM embeddings with domain-informed acoustic features improves interpretability and resilience. Significance: VOQANet+ shows strong potential for deployment in real-world and telehealth settings, addressing the limitations of subjective perceptual assessments with an interpretable and noise-resilient solution.


【13】 Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using  Co-training and Stochastic Precision

标题: 迈向一位ASO:使用联合训练和随机精度的极低位适形器量化
链接:https://arxiv.org/abs/2505.21245
作者: Zhaoqing Li,  Haoning Xu,  Zengrui Jin,  Lingwei Meng,  Tianzi Wang,  Huimeng Wang,  Youjun Chen,  Mingyu Cui,  Shujie Hu,  Xunying Liu 
备注:Accepted by Interspeech2025
摘要:随着现代语音系统规模的迅速增加,模型压缩已经成为一种新兴的需求。在本文中,我们研究模型权重量化,直接减少内存占用,以适应计算资源受限的应用程序。我们提出了新的方法来执行极低比特(即,2-比特和1比特)量化的Conformer自动语音识别系统使用多精度模型协同训练、随机精度和张量式可学习缩放因子来减轻量化引起的性能损失。所提出的方法可以实现性能无损的2位和1位量化的Conformer ASR系统训练的300小时开关板和960小时LibriSpeech语料库。最大的整体性能无损压缩比的16.2和16.6倍,实现了没有统计上显着增加的字错误率(WER)在全精度基线系统,分别。
摘要:Model compression has become an emerging need as the sizes of modern speech systems rapidly increase. In this paper, we study model weight quantization, which directly reduces the memory footprint to accommodate computationally resource-constrained applications. We propose novel approaches to perform extremely low-bit (i.e., 2-bit and 1-bit) quantization of Conformer automatic speech recognition systems using multiple precision model co-training, stochastic precision, and tensor-wise learnable scaling factors to alleviate quantization incurred performance loss. The proposed methods can achieve performance-lossless 2-bit and 1-bit quantization of Conformer ASR systems trained with the 300-hr Switchboard and 960-hr LibriSpeech corpus. Maximum overall performance-lossless compression ratios of 16.2 and 16.6 times are achieved without a statistically significant increase in the word error rate (WER) over the full precision baseline systems, respectively.


【14】 Unfolding A Few Structures for The Many: Memory-Efficient Compression of  Conformer and Speech Foundation Models

标题: 为许多人展开一些结构:适形器和语音基础模型的内存高效压缩
链接:https://arxiv.org/abs/2505.21237
作者: Zhaoqing Li,  Haoning Xu,  Xurong Xie,  Zengrui Jin,  Tianzi Wang,  Xunying Liu 
备注:Accepted by Interspeech2025
摘要:本文提出了一种新的内存有效的模型压缩方法的Conformer ASR和语音基础系统。我们的方法具有独特的“从小到大”设计。包含几个Conformer或Transformer块的紧凑“种子”模型经过多次训练和展开,以模拟具有不同逻辑深度的较大未压缩模型的性能。种子模型和许多展开路径在单个展开周期内联合训练。在自蒸馏过程中使用最大展开和最小种子模型之间的KL发散,以最小化它们的性能差异。实验结果表明,我们的可折叠模型产生的ASR性能与单独构建的Conformer和wav 2 vec 2/HuBERT语音基础模型在各种深度配置下相当,同时只需要最少的内存和存储。构象和wav 2 vec 2模型的参数分别减少了35%和30%,而性能没有损失。
摘要:This paper presents a novel memory-efficient model compression approach for Conformer ASR and speech foundation systems. Our approach features a unique "small-to-large" design. A compact "seed" model containing a few Conformer or Transformer blocks is trained and unfolded many times to emulate the performance of larger uncompressed models with different logical depths. The seed model and many unfolded paths are jointly trained within a single unfolding cycle. The KL-divergence between the largest unfolded and smallest seed models is used in a self-distillation process to minimize their performance disparity. Experimental results show that our foldable model produces ASR performance comparable to individually constructed Conformer and wav2vec2/HuBERT speech foundation models under various depth configurations, while requiring only minimal memory and storage. Conformer and wav2vec2 models with a reduction of 35% and 30% parameters are obtained without loss of performance, respectively.


【15】 Universal Speech Enhancement with Regression and Generative Mamba

标题: 使用回归和生成曼巴的通用语音增强
链接:https://arxiv.org/abs/2505.21198
作者: Rong Chao,  Rauf Nasretdinov,  Yu-Chiang Frank Wang,  Ante Jukić,  Szu-Wei Fu,  Yu Tsao 
备注:Accepted to Interspeech 2025
摘要:Interspeech 2025 URGENT Challenge旨在通过统一各种条件下的语音增强任务,包括七种不同的失真类型和五种语言,来推进通用、稳健和可推广的语音增强。我们提出了通用语音增强Mamba(USEMAMBA),一个状态空间语音增强模型,旨在处理长距离序列建模,时频结构化处理,和采样频率无关的特征提取。我们的方法主要依赖于基于回归的建模,它在大多数失真中表现良好。然而,对于数据包丢失和带宽扩展,必须推断丢失的内容,所提出的USEMAMBA的生成变体证明更有效。尽管仅在完整训练数据的一个子集上进行训练,但USEMAMBA在盲测阶段在Track 1中获得了第二名,在各种条件下表现出了很强的泛化能力。
摘要:The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions.


【16】 Topological Deep Learning for Speech Data

标题: 语音数据的布局深度学习
链接:https://arxiv.org/abs/2505.21173
作者: Zhiwang Yu 
备注:21 pages, 15 figures
摘要:拓扑数据分析(TDA)为深度学习提供了新的数学工具。受Carlsson等人的启发,这项研究设计了拓扑感知卷积核,显著改善了语音识别网络。理论上,通过研究正交群作用于核函数,我们建立了矩阵空间的纤维束分解,从而实现了新的滤波器生成方法。实际上,我们提出的正交特征(OF)层在音素识别方面取得了优异的性能,特别是在低噪声的情况下,同时表现出跨域的适应性。这项工作揭示了TDA在神经网络优化方面的潜力,为神经学-深度学习跨学科研究开辟了新的途径。
摘要:Topological data analysis (TDA) offers novel mathematical tools for deep learning. Inspired by Carlsson et al., this study designs topology-aware convolutional kernels that significantly improve speech recognition networks. Theoretically, by investigating orthogonal group actions on kernels, we establish a fiber-bundle decomposition of matrix spaces, enabling new filter generation methods. Practically, our proposed Orthogonal Feature (OF) layer achieves superior performance in phoneme recognition, particularly in low-noise scenarios, while demonstrating cross-domain adaptability. This work reveals TDA's potential in neural network optimization, opening new avenues for mathematics-deep learning interdisciplinary studies.


【17】 Model as Loss: A Self-Consistent Training Paradigm

标题: 模型即损失:自我一致的训练范式
链接:https://arxiv.org/abs/2505.21156
作者: Saisamarth Rajesh Phaye,  Milos Cernak,  Andrew Harper 
备注:Accepted in Interspeech 2025
摘要:用于语音增强的常规方法依赖于手工制作的损失函数(例如,时域或频域损失)或深特征损失(例如,使用WavLM或wav2vec),这通常无法捕获对于最佳性能至关重要的细微信号特性。为了解决这个问题,我们提出了Model as Loss,这是一种新的训练范式,它利用来自同一模型的编码器作为损失函数来指导训练。   损失模型范式利用编码器的特定于任务的特征空间,优化解码器以产生与干净信号的感知和任务相关特征一致的输出。通过使用编码器的学习功能作为损失函数,该框架强制执行干净的参考语音和增强的模型输出之间的自一致性。我们的方法在标准语音增强基准测试中优于预训练的深度特征损失,为域内和域外数据集提供更好的感知质量和鲁棒的泛化。
摘要:Conventional methods for speech enhancement rely on handcrafted loss functions (e.g., time or frequency domain losses) or deep feature losses (e.g., using WavLM or wav2vec), which often fail to capture subtle signal properties essential for optimal performance. To address this, we propose Model as Loss, a novel training paradigm that utilizes the encoder from the same model as a loss function to guide the training.   The Model as Loss paradigm leverages the encoder's task-specific feature space, optimizing the decoder to produce output consistent with perceptual and task-relevant characteristics of the clean signal. By using the encoder's learned features as a loss function, this framework enforces self-consistency between the clean reference speech and the enhanced model output. Our approach outperforms pre-trained deep feature losses on standard speech enhancement benchmarks, offering better perceptual quality and robust generalization to both in-domain and out-of-domain datasets.


【18】 Assessment of L2 Oral Proficiency using Speech Large Language Models

标题: 使用言语大语言模型评估二语口语能力
链接:https://arxiv.org/abs/2505.21148
作者: Rao Ma,  Mengjie Qian,  Siyuan Tang,  Stefano Bannò,  Kate M. Knill,  Mark J.F. Gales 
备注:submitted to Interspeech
摘要:随着二语人口的不断增长,对口语评估自动评分系统的需求也越来越大。从历史上看,统计模型、文本编码器和自监督语音模型已被用于这项任务。然而,级联系统遭受信息丢失,而E2E分级机也有局限性。随着多模态大语言模型(LLM)的最新进展,我们的目标是探索其作为二语口语水平评分的潜力,并克服这些问题。在这项工作中,我们比较了使用回归和分类目标的各种训练策略。我们的研究结果表明,语音LLM优于所有以前的竞争基线,在两个数据集上实现了卓越的性能。此外,受过训练的评分员在跨部分或跨任务评估中表现出较强的概括能力,这得益于LLM预培训期间获得的音频理解知识。
摘要:The growing population of L2 English speakers has increased the demand for developing automatic graders for spoken language assessment (SLA). Historically, statistical models, text encoders, and self-supervised speech models have been utilised for this task. However, cascaded systems suffer from the loss of information, while E2E graders also have limitations. With the recent advancements of multi-modal large language models (LLMs), we aim to explore their potential as L2 oral proficiency graders and overcome these issues. In this work, we compare various training strategies using regression and classification targets. Our results show that speech LLMs outperform all previous competitive baselines, achieving superior performance on two datasets. Furthermore, the trained grader demonstrates strong generalisation capabilities in the cross-part or cross-task evaluation, facilitated by the audio understanding knowledge acquired during LLM pre-training.


【19】 Leveraging LLM and Self-Supervised Training Models for Speech  Recognition in Chinese Dialects: A Comparative Analysis

标题: 利用LLM和自我监督训练模型进行中文方言语音识别:比较分析
链接:https://arxiv.org/abs/2505.21138
作者: Tianyi Xu,  Hongjie Chen,  Wang Qing,  Lv Hang,  Jian Kang,  Li Jie,  Zhennan Lin,  Yongxiang Li,  Xie Lei 
摘要:大规模的训练语料库显著提高了ASR模型的性能。不幸的是,由于数据相对稀缺,中国口音和方言仍然是大多数ASR模型的挑战。自监督学习的最新进展表明,自监督预训练与大型语言模型(LLM)相结合,可以有效地提高低资源场景中的ASR性能。我们的目的是调查这种范式对汉语方言的有效性。具体来说,我们在300,000小时的未标记方言和口音语音数据上预训练Data2vec2模型,并在40,000小时的监督数据集上进行对齐训练。然后,我们系统地研究了各种投影仪和LLM对普通话,方言和口音的语音识别性能的影响,在这个范例。我们的方法在多个方言数据集上实现了SOTA结果,包括Kespeech。我们将开源我们的工作,以促进可重复的研究
摘要:Large-scale training corpora have significantly improved the performance of ASR models. Unfortunately, due to the relative scarcity of data, Chinese accents and dialects remain a challenge for most ASR models. Recent advancements in self-supervised learning have shown that self-supervised pre- training, combined with large language models (LLM), can effectively enhance ASR performance in low-resource scenarios. We aim to investigate the effectiveness of this paradigm for Chinese dialects. Specifically, we pre-train a Data2vec2 model on 300,000 hours of unlabeled dialect and accented speech data and do alignment training on a supervised dataset of 40,000 hours. Then, we systematically examine the impact of various projectors and LLMs on Mandarin, dialect, and accented speech recognition performance under this paradigm. Our method achieved SOTA results on multiple dialect datasets, including Kespeech. We will open-source our work to promote reproducible research


【20】 Scaling and Prompting for Improved End-to-End Spoken Grammatical Error  Correction

标题: 缩放和绘图以改进端到端口语语法错误纠正
链接:https://arxiv.org/abs/2505.21137
作者: Mengjie Qian,  Rao Ma,  Stefano Bannò,  Kate M. Knill,  Mark J.F. Gales 
备注:submitted to Interspeech
摘要:口语语法错误纠正和反馈对于二语学习者、教师和考生都是至关重要的。传统的SGEC系统依赖于由ASR、用于不流利检测(DD)和去除的模块以及用于GEC的模块组成的级联流水线。随着端到端(E2E)语音基础模型的兴起,我们研究了它们在SGEC和反馈生成中的有效性。这项工作引入了一个伪标记过程来解决有限的标记数据的挑战,将训练数据的大小从77小时扩展到大约2500小时,从而提高了性能。此外,我们提示了一个基于E2 E Whisper的SGEC模型,具有流畅的转录,显示SGEC性能略有改善,反馈生成方面有更显着的改进。最后,我们评估了增加模型大小的影响,揭示了虽然伪标记数据不会为更大的Whisper模型带来性能增益,但使用提示进行训练是有益的。
摘要:Spoken Grammatical Error Correction (SGEC) and Feedback (SGECF) are crucial for second language learners, teachers and test takers. Traditional SGEC systems rely on a cascaded pipeline consisting of an ASR, a module for disfluency detection (DD) and removal and one for GEC. With the rise of end-to-end (E2E) speech foundation models, we investigate their effectiveness in SGEC and feedback generation. This work introduces a pseudo-labelling process to address the challenge of limited labelled data, expanding the training data size from 77 hours to approximately 2500 hours, leading to improved performance. Additionally, we prompt an E2E Whisper-based SGEC model with fluent transcriptions, showing a slight improvement in SGEC performance, with more significant gains in feedback generation. Finally, we assess the impact of increasing model size, revealing that while pseudo-labelled data does not yield performance gain for a larger Whisper model, training with prompts proves beneficial.


【21】 Text-Queried Audio Source Separation via Hierarchical Modeling

标题: 通过分层建模的文本查询音频源分离
链接:https://arxiv.org/abs/2505.21025
作者: Xinlei Yin,  Xiulian Peng,  Xue Jiang,  Zhiwei Xiong,  Yan Lu 
摘要:目标音频源分离与自然语言查询提出了一个有前途的范例提取任意音频事件,通过任意的文本描述。现有的方法主要面临两个挑战,即在盲学习的单阶段架构中联合建模声学-文本对齐和语义感知分离的困难,以及依赖大规模准确标记的训练数据来补偿低效的跨模态学习和分离。为了解决这些挑战,我们提出了一个分层分解框架,HSM-TSS,它将任务分解为全局-局部语义引导的特征分离和结构保留的声学重建。我们的方法引入了一个双阶段的语义分离机制,在不同的全球和本地语义特征空间。我们首先通过与文本查询对齐的全局语义特征空间执行全局语义分离。Q-Audio架构用于对齐音频和文本模态,用作预训练的全局语义编码器。在预测的全局特征的条件下,我们然后对保留时频结构的AudioMAE特征执行第二阶段局部语义分离,然后进行声学重建。我们还提出了一个指令处理流水线,将任意文本查询解析为结构化操作,提取或删除,再加上音频描述,实现灵活的声音操作。我们的方法通过数据高效的训练实现了最先进的分离性能,同时在复杂的听觉场景中保持了与查询的出色语义一致性。
摘要:Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing methods mainly face two challenges, the difficulty in jointly modeling acoustic-textual alignment and semantic-aware separation within a blindly-learned single-stage architecture, and the reliance on large-scale accurately-labeled training data to compensate for inefficient cross-modal learning and separation. To address these challenges, we propose a hierarchical decomposition framework, HSM-TSS, that decouples the task into global-local semantic-guided feature separation and structure-preserving acoustic reconstruction. Our approach introduces a dual-stage mechanism for semantic separation, operating on distinct global and local semantic feature spaces. We first perform global-semantic separation through a global semantic feature space aligned with text queries. A Q-Audio architecture is employed to align audio and text modalities, serving as pretrained global-semantic encoders. Conditioned on the predicted global feature, we then perform the second-stage local-semantic separation on AudioMAE features that preserve time-frequency structures, followed by acoustic reconstruction. We also propose an instruction processing pipeline to parse arbitrary text queries into structured operations, extraction or removal, coupled with audio descriptions, enabling flexible sound manipulation. Our method achieves state-of-the-art separation performance with data-efficient training while maintaining superior semantic consistency with queries in complex auditory scenes.


【22】 MelodySim: Measuring Melody-aware Music Similarity for Plagiarism  Detection

标题: MelodySim:测量旋律感知音乐相似性以检测抄袭
链接:https://arxiv.org/abs/2505.20979
作者: Tongyu Lu,  Charlotta-Marlena Geist,  Jan Melechovsky,  Abhinaba Roy,  Dorien Herremans 
摘要:我们提出了MelodySim,一个旋律感知的音乐相似性模型和数据集的剽窃检测。首先,我们介绍了一种新的方法来构建一个数据集,重点是旋律相似性。通过增强Slakh 2100;现有的数据集,我们生成每首作品的变化,同时通过修改,如音符分割,琶音,次要轨道脱落(不包括低音)和重新乐器保留旋律。一项用户研究证实,阳性对确实包含相似的旋律,而其他曲目则发生了显着变化。其次,我们开发了一个分段旋律相似性检测模型,该模型使用MERT编码器并应用三元组神经网络来捕获旋律相似性。由此产生的决策矩阵突出了可能发生剽窃的地方。我们的模型在MelodySim测试集上达到了很高的精度。
摘要:We propose MelodySim, a melody-aware music similarity model and dataset for plagiarism detection. First, we introduce a novel method to construct a dataset with focus on melodic similarity. By augmenting Slakh2100; an existing MIDI dataset, we generate variations of each piece while preserving the melody through modifications such as note splitting, arpeggiation, minor track dropout (excluding bass), and re-instrumentation. A user study confirms that positive pairs indeed contain similar melodies, with other musical tracks significantly changed. Second, we develop a segment-wise melodic-similarity detection model that uses a MERT encoder and applies a triplet neural network to capture melodic similarity. The resultant decision matrix highlights where plagiarism might occur. Our model achieves high accuracy on the MelodySim test set.


【23】 Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization

标题: 高效的麦克风容错3D声源定位
链接:https://arxiv.org/abs/2505.20961
作者: Yiyuan Yang,  Shitong Xu,  Niki Trigoni,  Andrew Markham 
备注:Accepted by Interspeech 2025 Conference
摘要:声源定位是在复杂环境中确定声源位置的关键技术。然而,现有的方法面临的挑战,如高计算成本和精确的校准要求,限制其部署在动态或资源受限的环境。本文介绍了一种新的3D SSL框架,它使用稀疏交叉注意,预训练和自适应信号相干性度量,以实现准确和计算效率的定位与较少的输入麦克风。该框架还对不可靠甚至未知的麦克风位置输入具有容错能力,确保其在现实世界场景中的适用性。初步实验表明,它的可扩展性多源定位,而不需要额外的硬件。这项工作通过平衡模型的性能和效率并提高其对真实世界场景的鲁棒性来推进SSL。
摘要:Sound source localization (SSL) is a critical technology for determining the position of sound sources in complex environments. However, existing methods face challenges such as high computational costs and precise calibration requirements, limiting their deployment in dynamic or resource-constrained environments. This paper introduces a novel 3D SSL framework, which uses sparse cross-attention, pretraining, and adaptive signal coherence metrics, to achieve accurate and computationally efficient localization with fewer input microphones. The framework is also fault-tolerant to unreliable or even unknown microphone position inputs, ensuring its applicability in real-world scenarios. Preliminary experiments demonstrate its scalability for multi-source localization without requiring additional hardware. This work advances SSL by balancing the model's performance and efficiency and improving its robustness for real-world scenarios.


【24】 Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound  Event Detection

标题: 基于不一致-多样性混合主动学习的生物声事件检测
链接:https://arxiv.org/abs/2505.20956
作者: Shiqi Zhang,  Tuomas Virtanen 
备注:5 pages, 1 figure, accepted by EUSIPCO 2025
摘要:生物声学声音事件检测(BioSED)对于生物多样性保护至关重要,但在模型开发和训练过程中面临着实际挑战:有限的注释数据,稀疏事件,物种多样性和类别不平衡。为了在有限的标签预算下有效地解决这些挑战,我们采用了失配优先最远遍历(MFFT),这是一种集成委员会投票分歧和多样性分析的主动学习方法。我们还改进了现有的BioSED数据集,专门用于评估主动学习算法。实验结果表明,MFFT在冷启动时实现了68%的mAP,在热启动时实现了71%的mAP(接近于75%的全监督mAP),同时仅使用2.3%的注释。值得注意的是,MFFT在冷启动场景和稀有物种方面表现出色,这对于监测濒危物种至关重要,证明了其实用价值。
摘要:Bioacoustic sound event detection (BioSED) is crucial for biodiversity conservation but faces practical challenges during model development and training: limited amounts of annotated data, sparse events, species diversity, and class imbalance. To address these challenges efficiently with a limited labeling budget, we apply the mismatch-first farthest-traversal (MFFT), an active learning method integrating committee voting disagreement and diversity analysis. We also refine an existing BioSED dataset specifically for evaluating active learning algorithms. Experimental results demonstrate that MFFT achieves a mAP of 68% when cold-starting and 71% when warm-starting (which is close to the fully-supervised mAP of 75%) while using only 2.3% of the annotations. Notably, MFFT excels in cold-start scenarios and with rare species, which are critical for monitoring endangered species, demonstrating its practical value.


【25】 Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing

标题: Dub-S2 ST:无缝配音的无文本语音翻译
链接:https://arxiv.org/abs/2505.20899
作者: Jeongsoo Choi,  Jaehun Kim,  Joon Son Chung 
摘要:本文介绍了一个跨语言配音系统,语音从一种语言翻译到另一种语言,同时保留关键特征,如持续时间,扬声器的身份,和说话速度。尽管现有的语音翻译方法具有很强的翻译质量,但它们往往忽略了语音模式的传递,导致与源语音的不匹配,并限制了它们对配音应用的适用性。为了解决这个问题,我们提出了一个离散的基于扩散的语音到单元的翻译模型,具有显式的持续时间控制,使时间对齐的翻译。然后,我们合成语音的基础上预测的单位和源身份与条件流匹配模型。此外,我们引入了一个基于单元的速度自适应机制,指导翻译模型以与源一致的速率生成语音,而不依赖于任何文本。大量的实验表明,我们的框架生成自然和流畅的翻译,符合原始语音的持续时间和说话节奏,同时实现有竞争力的翻译性能。
摘要:This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong translation quality of existing speech translation approaches, they often overlook the transfer of speech patterns, leading to mismatches with source speech and limiting their suitability for dubbing applications. To address this, we propose a discrete diffusion-based speech-to-unit translation model with explicit duration control, enabling time-aligned translation. We then synthesize speech based on the predicted units and source identity with a conditional flow matching model. Additionally, we introduce a unit-based speed adaptation mechanism that guides the translation model to produce speech at a rate consistent with the source, without relying on any text. Extensive experiments demonstrate that our framework generates natural and fluent translations that align with the original speech's duration and speaking pace, while achieving competitive translation performance.


【26】 Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction  and Style Direction Adjustment for Expressive Text-to-Speech

标题: Spotlight TTS:通过语音感知的风格提取和风格方向调整来突出风格,用于表达性文本到语音
链接:https://arxiv.org/abs/2505.20868
作者: Nam-Gyu Kim,  Deok-Hyeon Cho,  Seung-Bin Kim,  Seong-Whan Lee 
备注:Submitted to Interspeech
摘要:表达性文本到语音(TTS)的最新进展介绍了各种方法的基础上提取参考语音的风格嵌入。然而,合成高质量的表达语音仍然具有挑战性。我们提出了Spotlight TTS,它专门强调风格,通过风格感知的风格提取和风格方向调整。浊音感知风格提取关注与风格高度相关的浊音区域,同时保持不同语音区域之间的连续性以提高表现力。我们调整了提取的风格的方向,以最佳地整合到TTS模型中,从而提高了语音质量。实验结果表明,聚光灯TTS实现了卓越的表现力,整体语音质量和风格转移能力的基线模型相比。我们的音频样本是公开的。
摘要:Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose Spotlight-TTS, which exclusively emphasizes style via voiced-aware style extraction and style direction adjustment. Voiced-aware style extraction focuses on voiced regions highly related to style while maintaining continuity across different speech regions to improve expressiveness. We adjust the direction of the extracted style for optimal integration into the TTS model, which improves speech quality. Experimental results demonstrate that Spotlight-TTS achieves superior performance compared to baseline models in terms of expressiveness, overall speech quality, and style transfer capability. Our audio samples are publicly available.


【27】 VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing  Voice Conversion

标题: VibE-SRC:利用高频F0轮廓进行颤音提取,用于歌唱声音转换
链接:https://arxiv.org/abs/2505.20794
作者: Joon-Seung Choi,  Dong-Min Byun,  Hyung-Seok Oh,  Seong-Whan Lee 
备注:Proceedings of Interspeech 2025
摘要:掌握歌唱风格是获得富有表现力和自然的歌唱声音的关键。在各种风格因素中,颤音在传达情感和增强音乐深度方面起着关键作用。然而,建模颤音仍然具有挑战性,由于其动态性质,使其难以控制在歌唱声音转换。为了解决这个问题,我们提出了VibESVC,一个可控的歌声转换模型,明确提取和操纵颤音,使用离散小波变换。与以前隐式建模颤音的方法不同,我们的方法将F0轮廓分解为频率分量,从而实现精确的传输。这允许颤音控制增强的灵活性。实验结果表明,VibE-SVC在保持说话人相似性的同时,有效地变换了演唱风格。主观和客观评估都证实了高质量的转换。
摘要:Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propose VibESVC, a controllable singing voice conversion model that explicitly extracts and manipulates vibrato using discrete wavelet transform. Unlike previous methods that model vibrato implicitly, our approach decomposes the F0 contour into frequency components, enabling precise transfer. This allows vibrato control for enhanced flexibility. Experimental results show that VibE-SVC effectively transforms singing styles while preserving speaker similarity. Both subjective and objective evaluations confirm high-quality conversion.


【28】 Can Large Language Models Predict Audio Effects Parameters from Natural  Language?

标题: 大型语言模型可以从自然语言预测音效参数吗?
链接:https://arxiv.org/abs/2505.20770
作者: Seungheon Doh,  Junghyun Koo,  Marco A. Martínez-Ramírez,  Wei-Hsiang Liao,  Juhan Nam,  Yuki Mitsufuji 
备注:Submitted to WASPAA 2025
摘要:在音乐制作中,通过自然语言操纵音频效果(Fx)参数有可能减少非专家的技术障碍。我们提出了LLM2Fx,这是一个利用大型语言模型(LLM)的框架,可以直接从文本描述中预测Fx参数,而不需要特定于任务的训练或微调。我们的方法通过将自然语言描述映射到相应的Fx参数来进行均衡和混响,从而解决文本效果参数预测(Text2Fx)任务。我们证明,LLM可以产生Fx参数在一个zero-shot的方式,阐明在音乐制作中的音色语义和音频效果之间的关系。为了提高性能,我们介绍了三种类型的上下文示例:音频数字信号处理(DSP)功能,DSP功能代码,和Few-Shot的例子。我们的研究结果表明,基于LLM的Fx参数生成优于以前的优化方法,在将自然语言描述翻译为适当的Fx设置方面提供了有竞争力的性能。此外,LLM可以作为音频制作的文本驱动界面,为更直观和更易于访问的音乐制作工具铺平了道路。
摘要:In music production, manipulating audio effects (Fx) parameters through natural language has the potential to reduce technical barriers for non-experts. We present LLM2Fx, a framework leveraging Large Language Models (LLMs) to predict Fx parameters directly from textual descriptions without requiring task-specific training or fine-tuning. Our approach address the text-to-effect parameter prediction (Text2Fx) task by mapping natural language descriptions to the corresponding Fx parameters for equalization and reverberation. We demonstrate that LLMs can generate Fx parameters in a zero-shot manner that elucidates the relationship between timbre semantics and audio effects in music production. To enhance performance, we introduce three types of in-context examples: audio Digital Signal Processing (DSP) features, DSP function code, and few-shot examples. Our results demonstrate that LLM-based Fx parameter generation outperforms previous optimization approaches, offering competitive performance in translating natural language descriptions to appropriate Fx settings. Furthermore, LLMs can serve as text-driven interfaces for audio production, paving the way for more intuitive and accessible music production tools.


【29】 Foundation Model Hidden Representations for Heart Rate Estimation from  Auscultation

标题: 用于听诊心率估计的基础模型隐藏表示
链接:https://arxiv.org/abs/2505.20745
作者: Jingping Nie,  Dung T. Tran,  Karan Thakkar,  Vasudha Kowtha,  John Huang,  Carlos Avendano,  Erdrin Azemi,  Vikramjit Mitra 
备注:5 pages, Interspeech 2025 conference
摘要:听诊,特别是心音,是一种提供重要生命体征信息的非侵入性技术。最近,已经提出了自监督声学表示基础模型(FM),以提供对基于声学的生命体征的见解。然而,很少有人探索听诊在这些预先训练的FM表示中编码的程度。在这项工作中,使用公开可用的心音图(PCG)数据集和心率(HR)估计模型,我们对六种声学表示FM进行了逐层调查:HuBERT,wav 2 vec 2,wavLM,Whisper,对比听觉音频预训练(CLAP)和内部CLAP模型。此外,我们实现了Nie等人的基线方法,2024(依赖于声学特征),并表明总体而言,来自预训练基础模型(FM)的表示向量提供了与基线相当的性能。值得注意的是,使用来自内部CLAP模型的音频编码器的表示的HR估计优于从基线获得的结果,尽管存在域失配,但在各种训练/验证/测试分割上实现了较低的平均绝对误差(MAE)。
摘要:Auscultation, particularly heart sound, is a non-invasive technique that provides essential vital sign information. Recently, self-supervised acoustic representation foundation models (FMs) have been proposed to offer insights into acoustics-based vital signs. However, there has been little exploration of the extent to which auscultation is encoded in these pre-trained FM representations. In this work, using a publicly available phonocardiogram (PCG) dataset and a heart rate (HR) estimation model, we conduct a layer-wise investigation of six acoustic representation FMs: HuBERT, wav2vec2, wavLM, Whisper, Contrastive Language-Audio Pretraining (CLAP), and an in-house CLAP model. Additionally, we implement the baseline method from Nie et al., 2024 (which relies on acoustic features) and show that overall, representation vectors from pre-trained foundation models (FMs) offer comparable performance to the baseline. Notably, HR estimation using the representations from the audio encoder of the in-house CLAP model outperforms the results obtained from the baseline, achieving a lower mean absolute error (MAE) across various train/validation/test splits despite the domain mismatch.


【30】 Uni-VERSA: Versatile Speech Assessment with a Unified Network

标题: Uni-VERSA:具有统一网络的多功能语音评估
链接:https://arxiv.org/abs/2505.20741
作者: Jiatong Shi,  Hye-Jin Shim,  Shinji Watanabe 
备注:Accepted by Interspeech
摘要:主观听力测试仍然是语音质量评估的黄金标准,但成本高,变化大,难以衡量。相比之下,现有的客观指标,如PESQ,F0相关性,和DNSMOS,通常只捕捉语音质量的特定方面。为了解决这些限制,我们引入了Uni-VERSA,这是一个统一的网络,可以同时预测各种客观指标,包括自然度,可懂度,说话人特征,韵律和噪声,用于语音信号的综合评估。我们正式的框架,评估协议,并在语音增强,合成和质量控制的应用。基于URGENT 24挑战的基准测试,以及利用自我监督表示的基线,表明Uni-VERSA为单方面评估方法提供了一种可行的替代方案。此外,它与人类的感知密切相关,使其成为未来语音质量评估的一种有前途的方法。
摘要:Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only specific aspects of speech quality. To address these limitations, we introduce Uni-VERSA, a unified network that simultaneously predicts various objective metrics, encompassing naturalness, intelligibility, speaker characteristics, prosody, and noise, for a comprehensive evaluation of speech signals. We formalize its framework, evaluation protocol, and applications in speech enhancement, synthesis, and quality control. A benchmark based on the URGENT24 challenge, along with a baseline leveraging self-supervised representations, demonstrates that Uni-VERSA provides a viable alternative to single-aspect evaluation methods. Moreover, it aligns closely with human perception, making it a promising approach for future speech quality assessment.


【31】 Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent  Speech in Low-Resource Indian Languages

标题: 菲尔·赫拉·费尔(Phir Hera Fairy):一位英国童话演员是低资源印度语言流利演讲的忠实伪造者
链接:https://arxiv.org/abs/2505.20693
作者: Praveen Srinivasa Varadhan,  Srija Anand,  Soma Siddhartha,  Mitesh M.Khapra 
摘要:当一个英国童话作家对印度语言进行微调时会发生什么?我们评估英语F5-TTS模型如何适应11种印度语言,测量多语流利性,语音克隆,风格克隆和代码混合。我们比较:(i)从头开始训练,(ii)在印度数据上微调英语F5,以及(iii)在印度和英语数据上微调以防止遗忘。仅使用印度数据进行微调被证明是最有效的,并且由此产生的IN-F5是一种接近人类的多语言;这使得一种语言的使用者(例如,Odia)流利地用另一种语言说话(例如,印地语)。我们的研究结果表明,英语预培训艾滋病低资源TTS在达到人类平等。为了帮助其他低资源语言的进步,我们研究了数据约束的设置,并得出了一个计算最优策略。最后,我们展示了IN-F5可以通过合成数据生成,使用零资源TTS的人在回路方法合成看不见的语言,如Bhojpuri和Tulu。
摘要:What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing. We compare: (i) training from scratch, (ii) fine-tuning English F5 on Indian data, and (iii) fine-tuning on both Indian and English data to prevent forgetting. Fine-tuning with only Indian data proves most effective and the resultant IN-F5 is a near-human polyglot; that enables speakers of one language (e.g., Odia) to fluently speak in another (e.g., Hindi). Our results show English pretraining aids low-resource TTS in reaching human parity. To aid progress in other low-resource languages, we study data-constrained setups and arrive at a compute optimal strategy. Finally, we show IN-F5 can synthesize unseen languages like Bhojpuri and Tulu using a human-in-the-loop approach for zero-resource TTS via synthetic data generation.


【32】 Music's Multimodal Complexity in AVQA: Why We Need More than General  Multimodal LLMs

标题: AVQA中音乐的多模式复杂性:为什么我们需要的不仅仅是一般的多模式LLM
链接:https://arxiv.org/abs/2505.20638
作者: Wenhao You,  Xingjian Diao,  Chunhui Zhang,  Keyi Kong,  Weiyi Wu,  Zhongyu Ouyang,  Chiyu Ma,  Tingxuan Wu,  Noah Wei,  Zong Ke,  Ming Cheng,  Soroush Vosoughi,  Jiang Gui 
摘要:虽然最近的多模态大型语言模型在一般多模态任务方面表现出令人印象深刻的能力,但像音乐这样的专业领域需要量身定制的方法。音乐视听问题问答(Music AVQA)特别强调了这一点,其连续,密集分层的视听内容,复杂的时间动态以及对特定领域知识的迫切需求带来了独特的挑战。通过对音乐AVQA数据集和方法的系统分析,本立场文件确定了专门的输入处理,包含专用时空设计的架构以及特定于音乐的建模策略对于该领域的成功至关重要。我们的研究为研究人员提供了有价值的见解,突出了有效的设计模式经验联系到强大的性能,提出了具体的未来方向,将音乐先验,并旨在建立一个强大的基础,推进多模态音乐的理解。这项工作旨在激发更广泛的关注和进一步的研究,并得到不断更新的匿名GitHub相关论文库的支持:https://github.com/xid32/Survey4MusicAVQA。
摘要:While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Audio-Visual Question Answering (Music AVQA) particularly underscores this, presenting unique challenges with its continuous, densely layered audio-visual content, intricate temporal dynamics, and the critical need for domain-specific knowledge. Through a systematic analysis of Music AVQA datasets and methods, this position paper identifies that specialized input processing, architectures incorporating dedicated spatial-temporal designs, and music-specific modeling strategies are critical for success in this domain. Our study provides valuable insights for researchers by highlighting effective design patterns empirically linked to strong performance, proposing concrete future directions for incorporating musical priors, and aiming to establish a robust foundation for advancing multimodal musical understanding. This work is intended to inspire broader attention and further research, supported by a continuously updated anonymous GitHub repository of relevant papers: https://github.com/xid32/Survey4MusicAVQA.


【33】 Techniques for Quantum-Computing-Aided Algorithmic Composition:  Experiments in Rhythm, Timbre, Harmony, and Space

标题: 量子计算辅助数学作曲技术:节奏、音色、和声和空间实验
链接:https://arxiv.org/abs/2505.20565
作者: Christopher Dobrian,  Omar Costa Hamido 
摘要:量子计算可以用于计算机辅助音乐创作,以控制不同结构层次的音乐的各种属性。本文介绍了应用量子模拟模型组成的决策,模拟量子粒子跟踪产生噪声为基础的音色,使用基态矢量旋转,以引起颗粒谐波纹理的概率行为的变化,以及利用量子测量误差,造成嘈杂的扰动空间声音路径。我们描述了这些技术的基本概念,我们提供了算法和软件制定他们,我们提供的例子,展示他们在计算机生成的音乐的实现。
摘要:Quantum computing can be employed in computer-aided music composition to control various attributes of the music at different structural levels. This article describes the application of quantum simulation to model compositional decision making, the simulation of quantum particle tracking to produce noise-based timbres, the use of basis state vector rotation to cause changing probabilistic behaviors in granular harmonic textures, and the exploitation of quantum measurement error to cause noisy perturbations of spatial soundpaths. We describe the concepts fundamental to these techniques, we provide algorithms and software enacting them, and we provide examples demonstrating their implementation in computer-generated music.


【34】 Training Articulatory Inversion Models for Inter-Speaker Consistency

标题: 训练发音倒置模型以实现说话者间一致性
链接:https://arxiv.org/abs/2505.20529
作者: Charles McGhee,  Mark J.F. Gales,  Kate M. Knill 
摘要:声学到发音反转(AAI)试图模拟从语音到发音的逆映射。仅仅从语音中准确地预测发音可能是不可能的,因为说话者可以选择不同的发音形式,而似乎不需要参考他们的声道结构。然而,一旦说话者选择了一种发音形式,他们的产出变化最小。AAI最近的工作提出了将自监督学习(SSL)模型适应于单说话者数据集,声称这些单说话者模型提供了一个通用的发音模板。在本文中,我们调查是否SSL适应模型训练的单和多扬声器数据产生发音目标,这是一致的跨扬声器身份的英语和俄语。我们这样做,通过使用一种新的评价方法,提取发音目标,使用最小对集。我们还提出了一种训练方法,可以提高说话人之间的一致性,只用语音数据。
摘要:Acoustic-to-Articulatory Inversion (AAI) attempts to model the inverse mapping from speech to articulation. Exact articulatory prediction from speech alone may be impossible, as speakers can choose different forms of articulation seemingly without reference to their vocal tract structure. However, once a speaker has selected an articulatory form, their productions vary minimally. Recent works in AAI have proposed adapting Self-Supervised Learning (SSL) models to single-speaker datasets, claiming that these single-speaker models provide a universal articulatory template. In this paper, we investigate whether SSL-adapted models trained on single and multi-speaker data produce articulatory targets which are consistent across speaker identities for English and Russian. We do this through the use of a novel evaluation method which extracts articulatory targets using minimal pair sets. We also present a training method which can improve inter-speaker consistency using only speech data.


【35】 ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

标题: ArVoice:用于阿拉伯语语音合成的多说话人数据集
链接:https://arxiv.org/abs/2505.20506
作者: Hawau Olamide Toyin,  Rufael Marew,  Humaid Alblooshi,  Samar M. Magdy,  Hanan Aldarmaki 
备注:Accepted at INTERSPEECH 2025 The dataset is available at this https URL
摘要:我们介绍ArVoice,这是一个带有变音符号转录的多扬声器现代标准阿拉伯语(MSA)语音语料库,用于多扬声器语音合成,并且可以用于其他任务,例如基于语音的变音符号恢复、语音转换和深度伪造检测。ArVoice包括:(1)由六位不同人口统计学特征的语音人才组成的新的专业录制集,(2)阿拉伯语语音语料库的修改子集;以及(3)来自两个商业系统的高质量合成语音。完整的语料库包括11种声音的83.52小时的语音;大约10小时由7个说话者的人类声音组成。我们训练三个开源TTS和两个语音转换系统来说明数据集的用例。语料库可供研究使用。
摘要:We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration, voice conversion, and deepfake detection. ArVoice comprises: (1) a new professionally recorded set from six voice talents with diverse demographics, (2) a modified subset of the Arabic Speech Corpus; and (3) high-quality synthetic speech from two commercial systems. The complete corpus consists of a total of 83.52 hours of speech across 11 voices; around 10 hours consist of human voices from 7 speakers. We train three open-source TTS and two voice conversion systems to illustrate the use cases of the dataset. The corpus is available for research use.


机器翻译由腾讯交互翻译提供,仅供参考