今日论文合集:cs.SD语音10篇,eess.AS音频处理12篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】ECHO: Frequency-aware Hierarchical Encoding for Variable-length Signal
标题:ECHO:可变长度信号的频率感知分层编码
链接:https://arxiv.org/abs/2508.14689

作者:Yucong Zhang, Juan Liu, Ming Li
摘要:预先训练的基础模型在视觉和语言方面取得了显著的成功,但它们在通用机器信号建模(包括声学、振动和其他工业传感器数据)方面的潜力仍有待开发。使用基于子带的编码器的现有方法已经取得了有竞争力的结果,但受到固定输入长度和缺乏显式频率位置编码的限制。在这项工作中,我们提出了一种新的基础模型,集成了先进的频带分离架构与相对频率位置嵌入,使精确的频谱定位在任意采样配置。该模型支持任意长度的输入,而无需填充或分割,从而产生简洁的嵌入,同时保留时间和频谱保真度。我们在SIREN(https://github.com/yuanzh/SIREN)上评估了我们的方法,SIREN是一个新引入的机器信号编码的大规模基准,它统一了多个数据集,包括所有DCASE任务2挑战(2020-2025)和广泛使用的工业信号语料库。实验结果表明,在异常检测和故障识别一致的最先进的性能,证实了所提出的模型的有效性和泛化能力。我们在https://github.com/yucongzh/ECHO上开源了ECHO。
摘要:Pre-trained foundation models have demonstrated remarkable success in vision and language, yet their potential for general machine signal modeling-covering acoustic, vibration, and other industrial sensor data-remains under-explored. Existing approach using sub-band-based encoders has achieved competitive results but are limited by fixed input lengths, and the absence of explicit frequency positional encoding. In this work, we propose a novel foundation model that integrates an advanced band-split architecture with relative frequency positional embeddings, enabling precise spectral localization across arbitrary sampling configurations. The model supports inputs of arbitrary length without padding or segmentation, producing a concise embedding that retains both temporal and spectral fidelity. We evaluate our method on SIREN (https://github.com/yucongzh/SIREN), a newly introduced large-scale benchmark for machine signal encoding that unifies multiple datasets, including all DCASE task 2 challenges (2020-2025) and widely-used industrial signal corpora. Experimental results demonstrate consistent state-of-the-art performance in anomaly detection and fault identification, confirming the effectiveness and generalization capability of the proposed model. We open-sourced ECHO on https://github.com/yucongzh/ECHO.


【2】BioSonix: Can Physics-Based Sonification Perceptualize Tissue Deformations From Tool Interactions?
标题:BioSonix:基于物理的超声可以从工具相互作用中感知组织变形吗?
链接:https://arxiv.org/abs/2508.14688

作者:Veronica Ruozzi, Sasan Matinfar, Laura Schütz, Benedikt Wiestler, Alberto Redaelli, Emiliano Votta, Nassir Navab
备注:V. Ruozzi and S. Matinfar contributed equally to this work
摘要:在外科手术中感知工具与可变形结构的相互作用仍然具有挑战性,因为单峰可视化技术通常由于诸如遮挡和有限的深度感知的约束而无法捕获这些相互作用的复杂性。本文提出了一种新的方法来增强工具导航在混合现实环境中提供听觉表示的工具组织动力学,特别是与软组织的相互作用。BioSonix是一种基于物理学的设计框架,它利用3D空间中的组织位移来计算编码组织特性(如刚度和密度)的声音模型的激励力。采用生物力学模拟来模拟工具-组织相互作用引起的颗粒位移,为该方法建立了坚实的基础。一种优化方法被用来定义配置捕获不同的互动场景与不同的工具轨迹。实验进行了验证的声音位移映射的准确性。此外,还进行了两项用户研究:第一项研究涉及两名临床专业人员(神经放射科医生和心脏病学家),他们证实了该方法的影响并实现了高任务准确性;第二项研究包括22名生物医学专家,他们在组织分化和靶向任务中表现出高分辨准确性。结果显示,工具-组织动力学与其相应的听觉特征之间存在很强的相关性,突出了这些声音表征增强对复杂相互作用的直观理解的潜力。
摘要:Perceptualizing tool interactions with deformable structures in surgical procedures remains challenging, as unimodal visualization techniques often fail to capture the complexity of these interactions due to constraints such as occlusion and limited depth perception. This paper presents a novel approach to augment tool navigation in mixed reality environments by providing auditory representations of tool-tissue dynamics, particularly for interactions with soft tissue. BioSonix, a physics-informed design framework, utilizes tissue displacements in 3D space to compute excitation forces for a sound model encoding tissue properties such as stiffness and density. Biomechanical simulations were employed to model particle displacements resulting from tool-tissue interactions, establishing a robust foundation for the method. An optimization approach was used to define configurations for capturing diverse interaction scenarios with varying tool trajectories. Experiments were conducted to validate the accuracy of the sound-displacement mappings. Additionally, two user studies were performed: the first involved two clinical professionals (a neuroradiologist and a cardiologist), who confirmed the method's impact and achieved high task accuracy; the second included 22 biomedical experts, who demonstrated high discrimination accuracy in tissue differentiation and targeting tasks. The results revealed a strong correlation between tool-tissue dynamics and their corresponding auditory profiles, highlighting the potential of these sound representations to enhance the intuitive understanding of complex interactions.


【3】Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
标题:Mamba2 Meets Silence:稀疏区域的鲁棒声源分离
链接:https://arxiv.org/abs/2508.14556

作者:Euiyeon Kim, Yong-Hoon Choi
摘要:我们介绍了一种新的音乐源分离模型,专为准确的声乐隔离。与基于transformer的方法不同,这种方法通常无法捕获间歇性出现的声音,我们的模型利用了最近的状态空间模型Mamba 2,以更好地捕获长期的时间依赖性。为了有效地处理长输入序列,我们结合了带分裂策略与双路径架构。实验表明,我们的方法优于最新的最先进的模型,实现了11.03 dB的cSDR-迄今为止最好的报告-并提供了大量的收益在uSDR。此外,该模型在不同的输入长度和声音发生模式下表现出稳定和一致的性能。这些结果证明了基于Mamba的模型用于高分辨率音频处理的有效性,并为音频研究中更广泛的应用开辟了新的方向。
摘要:We introduce a new music source separation model tailored for accurate vocal isolation. Unlike Transformer-based approaches, which often fail to capture intermittently occurring vocals, our model leverages Mamba2, a recent state space model, to better capture long-range temporal dependencies. To handle long input sequences efficiently, we combine a band-splitting strategy with a dual-path architecture. Experiments show that our approach outperforms recent state-of-the-art models, achieving a cSDR of 11.03 dB-the best reported to date-and delivering substantial gains in uSDR. Moreover, the model exhibits stable and consistent performance across varying input lengths and vocal occurrence patterns. These results demonstrate the effectiveness of Mamba-based models for high-resolution audio processing and open up new directions for broader applications in audio research.


【4】EmoTale: An Enacted Speech-emotion Dataset in Danish
标题:收件箱故事:丹麦语已制定的言语情感数据集
链接:https://arxiv.org/abs/2508.14548

作者:Maja J. Hjuler, Harald V. Skat-Rørdam, Line H. Clemmensen, Sneha Das
备注:To appear in the proceedings of ASRU 2025
摘要:虽然对于常用语言存在多个情感语音语料库,但对于较小的(口语)语言(如丹麦语)缺乏功能数据集。据我们所知,丹麦情绪语音(DES),发表于1997年,是丹麦情绪语音的唯一的其他数据库。我们目前的童话,语料库包括丹麦语和英语的语音录音与他们相关的制定情绪注释。我们证明了数据集的有效性,通过调查和提出其预测能力,使用语音情感识别(SER)模型。我们使用自监督语音模型(SSLM)嵌入和openSMILE特征提取器为ESTALE和参考数据集开发SER模型。我们发现嵌入优于手工制作的功能。最好的模型实现了64.1%的未加权平均召回率(UAR)的EQUIPTale语料库使用留一说话人交叉验证,DES上的性能相当。
摘要:While multiple emotional speech corpora exist for commonly spoken languages, there is a lack of functional datasets for smaller (spoken) languages, such as Danish. To our knowledge, Danish Emotional Speech (DES), published in 1997, is the only other database of Danish emotional speech. We present EmoTale; a corpus comprising Danish and English speech recordings with their associated enacted emotion annotations. We demonstrate the validity of the dataset by investigating and presenting its predictive power using speech emotion recognition (SER) models. We develop SER models for EmoTale and the reference datasets using self-supervised speech model (SSLM) embeddings and the openSMILE feature extractor. We find the embeddings superior to the hand-crafted features. The best model achieves an unweighted average recall (UAR) of 64.1% on the EmoTale corpus using leave-one-speaker-out cross-validation, comparable to the performance on DES.


【5】EffiFusion-GAN: Efficient Fusion Generative Adversarial Network for Speech Enhancement
标题:EfiFusion-GAN:用于语音增强的高效融合生成对抗网络
链接:https://arxiv.org/abs/2508.14525

作者:Bin Wen, Tien-Ping Tan
摘要:我们介绍EffiFusion-GAN(高效融合生成对抗网络),这是一种轻量级但功能强大的语音增强模型。该模型在多尺度块内集成了深度可分离卷积,以有效地捕获不同的声学特征。具有双重归一化和残差细化的增强注意力机制进一步提高了训练稳定性和收敛性。此外,动态修剪应用于减少模型的大小,同时保持性能,使框架适合资源受限的环境。在公共VoiceBank+DEMAND数据集上的实验评估表明,EffiFusion-GAN的PESQ得分为3.45,优于相同参数设置下的现有模型。
摘要:We introduce EffiFusion-GAN (Efficient Fusion Generative Adversarial Network), a lightweight yet powerful model for speech enhancement. The model integrates depthwise separable convolutions within a multi-scale block to capture diverse acoustic features efficiently. An enhanced attention mechanism with dual normalization and residual refinement further improves training stability and convergence. Additionally, dynamic pruning is applied to reduce model size while maintaining performance, making the framework suitable for resource-constrained environments. Experimental evaluation on the public VoiceBank+DEMAND dataset shows that EffiFusion-GAN achieves a PESQ score of 3.45, outperforming existing models under the same parameter settings.


【6】Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
标题:心理健康和神经退行性疾病的开放语音生物标志物数据集的系统公平性评估
链接:https://arxiv.org/abs/2508.14089

作者:Ishaan Mahapatra, Nihar R. Mahapatra
备注:To appear in the Proceedings of the 28th International Conference on Text, Speech and Dialogue (TSD 2025), Erlangen, Germany, August 25-28, 2025
摘要:语音生物标志物-人类产生的声音信号,如语音,咳嗽和呼吸-是有前途的工具,可扩展,非侵入性检测和监测心理健康和神经退行性疾病。然而,它们的临床应用仍然受到公开数据集质量不一致和可用性有限的限制。为了解决这一差距,我们提出了第一个系统的FAIR(可发现,可解释,可互操作,可重复使用)评估27个公开的声音生物标志物数据集集中在这些疾病领域。使用FAIR数据成熟度模型和结构化的优先级加权评分方法,我们评估了子原则,原则和复合水平的公平性。我们的分析显示,可查找性一直很高,但在可访问性、互操作性和可重用性方面存在很大的可变性和弱点。精神健康数据集在FAIR评分中表现出更大的变异性,而神经退行性疾病数据集则稍微一致。存储库的选择也显着影响公平分数。为了提高数据集质量和临床实用性,我们建议采用结构化的、特定领域的元数据标准,优先考虑符合FAIR的存储库,并定期应用结构化的FAIR评估框架。这些发现为改善数据集的互操作性和重用提供了可操作的指导,从而加速了语音生物标志物技术的临床转化。
摘要:Voice biomarkers--human-generated acoustic signals such as speech, coughing, and breathing--are promising tools for scalable, non-invasive detection and monitoring of mental health and neurodegenerative diseases. Yet, their clinical adoption remains constrained by inconsistent quality and limited usability of publicly available datasets. To address this gap, we present the first systematic FAIR (Findable, Accessible, Interoperable, Reusable) evaluation of 27 publicly available voice biomarker datasets focused on these disease areas. Using the FAIR Data Maturity Model and a structured, priority-weighted scoring method, we assessed FAIRness at subprinciple, principle, and composite levels. Our analysis revealed consistently high Findability but substantial variability and weaknesses in Accessibility, Interoperability, and Reusability. Mental health datasets exhibited greater variability in FAIR scores, while neurodegenerative datasets were slightly more consistent. Repository choice also significantly influenced FAIRness scores. To enhance dataset quality and clinical utility, we recommend adopting structured, domain-specific metadata standards, prioritizing FAIR-compliant repositories, and routinely applying structured FAIR evaluation frameworks. These findings provide actionable guidance to improve dataset interoperability and reuse, thereby accelerating the clinical translation of voice biomarker technologies.


【7】Long-Context Speech Synthesis with Context-Aware Memory
标题:具有上下文感知记忆的长上下文语音合成
链接:https://arxiv.org/abs/2508.14713

作者:Zhipeng Li, Xiaofen Xing, Jingyuan Xing, Hangrui Hu, Heng Lu, Xiangmin Xu
备注:Accepted by Interspeech25
摘要:在长文本语音合成中,当前的方法通常在句子级将文本转换成语音,并将结果连接以形成伪段落级语音。这些方法忽略了段落的上下文连贯性,导致长式演讲的自然度降低,风格和音色不一致。为了解决这些问题,我们提出了一个基于上下文感知记忆(CAM)的长上下文文语转换(TTS)模型。CAM模块集成并检索长期存储器和本地上下文细节,从而实现动态存储器更新和长段落内的传输,以指导句子级语音合成。此外,前缀掩码通过在保持单向生成的同时实现对前缀令牌的双向关注来增强上下文内学习能力。实验结果表明,该方法优于基线和国家的最先进的长上下文方法的韵律表现力,连贯性和上下文推理成本跨段落级语音。
摘要:In long-text speech synthesis, current approaches typically convert text to speech at the sentence-level and concatenate the results to form pseudo-paragraph-level speech. These methods overlook the contextual coherence of paragraphs, leading to reduced naturalness and inconsistencies in style and timbre across the long-form speech. To address these issues, we propose a Context-Aware Memory (CAM)-based long-context Text-to-Speech (TTS) model. The CAM block integrates and retrieves both long-term memory and local context details, enabling dynamic memory updates and transfers within long paragraphs to guide sentence-level speech synthesis. Furthermore, the prefix mask enhances the in-context learning ability by enabling bidirectional attention on prefix tokens while maintaining unidirectional generation. Experimental results demonstrate that the proposed method outperforms baseline and state-of-the-art long-context methods in terms of prosody expressiveness, coherence and context inference cost across paragraph-level speech.


【8】Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
标题:通过神经可区分的DSP声码器改进提高资源效率的语音增强
链接:https://arxiv.org/abs/2508.14709

作者: Heitor R. Guimaraes, Ke Tan, Juan Azcarreta, Jesus Alvarez, Prabhav Agrawal, Ashutosh Pandey, Buye Xu
备注:Accepted to the 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
摘要:由于设备上的有限计算资源,在诸如智能眼镜之类的可穿戴设备中部署语音增强(SE)系统具有挑战性。尽管深度学习方法已经取得了高质量的结果,但其计算成本限制了其在嵌入式平台上的可行性。这项工作提出了一个有效的端到端SE框架,利用差分数字信号处理(DDSP)声码器的高质量的语音合成。首先,一个紧凑的神经网络预测噪声语音的增强声学特征:频谱包络,基频(F0)和周期性。这些特征被馈送到DDSP声码器以合成增强波形。该系统通过STFT和对抗性损失进行端到端训练,从而实现在特征和波形级别的直接优化。实验结果表明,我们的方法在不显著增加计算的情况下,在强基线上将可懂度和质量提高了4%(STOI)和19%(DNSMOS),使其非常适合实时应用。
摘要:Deploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their computational cost limits their feasibility on embedded platforms. This work presents an efficient end-to-end SE framework that leverages a Differentiable Digital Signal Processing (DDSP) vocoder for high-quality speech synthesis. First, a compact neural network predicts enhanced acoustic features from noisy speech: spectral envelope, fundamental frequency (F0), and periodicity. These features are fed into the DDSP vocoder to synthesize the enhanced waveform. The system is trained end-to-end with STFT and adversarial losses, enabling direct optimization at the feature and waveform levels. Experimental results show that our method improves intelligibility and quality by 4% (STOI) and 19% (DNSMOS) over strong baselines without significantly increasing computation, making it well-suited for real-time applications.


【9】A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References
标题:有噪参考语音分离中尺度不变信失真比的研究
链接:https://arxiv.org/abs/2508.14623

作者:Simon Dahl Jepsen, Mads Græsbøll Christensen, Jesper Rindom Jensen
备注:Accepted for IEEE ASRU 2025, Workshop on Automatic Speech Recognition and Understanding. Copyright (c) 2025 IEEE. 8 pages, 6 figures, 2 tables
摘要:本文探讨了使用尺度不变的信号失真比(SI-SDR)作为评估和训练目标的监督语音分离的影响,当训练参考包含噪声,是事实上的基准WSJ 0 - 2 Mix的情况下。具有噪声基准的SI-SDR的推导表明,噪声限制了可实现的SI-SDR,或者导致分离输出中的不期望噪声。为了解决这个问题,提出了一种方法来增强参考和增加与WHAM!,旨在训练避免学习有噪参考的模型。在这些增强的数据集上训练的两个模型使用非侵入式NISQA.v2度量进行评估。结果表明,在分离的语音噪声减少,但建议处理参考可能会引入文物,限制整体质量的收益。在WSJ 0 - 2 Mix和Libri 2 Mix测试集上,发现SI-SDR和感知噪声之间存在负相关性,强调了推导得出的结论。
摘要:This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation.


【10】Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
标题:采用短上下文发言人嵌入的多个发言人实现低延迟跟踪
链接:https://arxiv.org/abs/2508.14115

作者:Taous Iatariene, Alexandre Guérin, Romain Serizel (MULTISPEECH)
备注:None
摘要:说话人嵌入是一种很有前途的身份相关特征,可以通过利用其空间预测,即通过执行身份重新分配来提高跟踪系统的身份分配性能。普通说话人嵌入提取器通常与短时间上下文和重叠语音斗争,这需要长时间的身份重新分配来利用更长的时间上下文。然而,这增加了跟踪系统错误的可能性,这反过来又对身份重新分配产生负面影响。为了解决这个问题,我们提出了一种基于知识蒸馏(KD)的训练方法,用于从两个说话人混合物中提取短上下文说话人嵌入。我们利用感兴趣的扬声器的空间信息,使用波束成形,以减少重叠。我们研究了在固定大小的块上执行身份重新分配的可行性,即,块式身份重新分配,以实现基于低延迟扬声器嵌入的跟踪系统。结果表明,我们的蒸馏模型是有效的短上下文嵌入提取和更强大的重叠。虽然,块重新分配的结果表明,需要进一步的工作,以更有效地处理同步语音。
摘要:Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding extractors usually struggle with short temporal contexts and overlapping speech, which imposes long-term identity reassignment to exploit longer temporal contexts. However, this increases the probability of tracking system errors, which in turn impacts negatively on identity reassignment. To address this, we propose a Knowledge Distillation (KD) based training approach for short context speaker embedding extraction from two speaker mixtures. We leverage the spatial information of the speaker of interest using beamforming to reduce overlap. We study the feasibility of performing identity reassignment over blocks of fixed size, i.e., blockwise identity reassignment, to go towards a low-latency speaker embedding based tracking system. Results demonstrate that our distilled models are effective at short-context embedding extraction and more robust to overlap. Although, blockwise reassignment results indicate that further work is needed to handle simultaneous speech more effectively.


eess.AS音频处理


【1】PadAug: Robust Speaker Verification with Simple Waveform-Level Silence Padding
标题:PadAug:通过简单的波级静音填充进行稳健的说话人验证
链接:https://arxiv.org/abs/2508.14732

作者:Zijun Huang, Chengdong Liang, Jiadi Yao, Xiao-Lei Zhang
摘要:非语音段的存在往往会导致说话人确认的性能下降。现有的系统通常使用语音激活检测作为预处理步骤,以切断长的沉默段。然而,短的沉默段,特别是那些语音段之间,仍然是说话人确认的一个问题。为了解决这个问题,在本文中,我们提出了一个简单的波级数据增强方法,\textit{PadAug},其目的是提高系统的鲁棒性沉默段。\textit{PadAug}的核心思想是在波形级别将静音段与语音段连接起来进行模型训练。由于其简单性,它可以直接应用于当前最先进的架构。实验结果证明了所提出的\textit{PadAug}的有效性。例如,将\textit{PadAug}应用于ResNet 34在voxceleb数据集上实现了5.0\%的相对相等错误率降低。此外,基于\textit{PadAug}的系统对测试数据中不同长度和比例的静音段具有鲁棒性。
摘要:The presence of non-speech segments in utterances often leads to the performance degradation of speaker verification. Existing systems usually use voice activation detection as a preprocessing step to cut off long silence segments. However, short silence segments, particularly those between speech segments, still remain a problem for speaker verification. To address this issue, in this paper, we propose a simple wave-level data augmentation method, \textit{PadAug}, which aims to enhance the system's robustness to silence segments. The core idea of \textit{PadAug} is to concatenate silence segments with speech segments at the waveform level for model training. Due to its simplicity, it can be directly applied to the current state-of-the art architectures. Experimental results demonstrate the effectiveness of the proposed \textit{PadAug}. For example, applying \textit{PadAug} to ResNet34 achieves a relative equal error rate reduction of 5.0\% on the voxceleb dataset. Moreover, the \textit{PadAug} based systems are robust to different lengths and proportions of silence segments in the test data.


【2】Long-Context Speech Synthesis with Context-Aware Memory
标题:具有上下文感知记忆的长上下文语音合成
链接:https://arxiv.org/abs/2508.14713

作者:Zhipeng Li, Xiaofen Xing, Jingyuan Xing, Hangrui Hu, Heng Lu, Xiangmin Xu
备注:Accepted by Interspeech25
摘要:在长文本语音合成中,当前的方法通常在句子级将文本转换成语音,并将结果连接以形成伪段落级语音。这些方法忽略了段落的上下文连贯性,导致长式演讲的自然度降低,风格和音色不一致。为了解决这些问题,我们提出了一个基于上下文感知记忆(CAM)的长上下文文语转换(TTS)模型。CAM模块集成并检索长期存储器和本地上下文细节,从而实现动态存储器更新和长段落内的传输,以指导句子级语音合成。此外,前缀掩码通过在保持单向生成的同时实现对前缀令牌的双向关注来增强上下文内学习能力。实验结果表明,该方法优于基线和国家的最先进的长上下文方法的韵律表现力,连贯性和上下文推理成本跨段落级语音。
摘要:In long-text speech synthesis, current approaches typically convert text to speech at the sentence-level and concatenate the results to form pseudo-paragraph-level speech. These methods overlook the contextual coherence of paragraphs, leading to reduced naturalness and inconsistencies in style and timbre across the long-form speech. To address these issues, we propose a Context-Aware Memory (CAM)-based long-context Text-to-Speech (TTS) model. The CAM block integrates and retrieves both long-term memory and local context details, enabling dynamic memory updates and transfers within long paragraphs to guide sentence-level speech synthesis. Furthermore, the prefix mask enhances the in-context learning ability by enabling bidirectional attention on prefix tokens while maintaining unidirectional generation. Experimental results demonstrate that the proposed method outperforms baseline and state-of-the-art long-context methods in terms of prosody expressiveness, coherence and context inference cost across paragraph-level speech.


【3】Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
标题:通过神经可区分的DSP声码器改进提高资源效率的语音增强
链接:https://arxiv.org/abs/2508.14709

作者: Simon Dahl Jepsen. Guimaraes, Ke Tan, Juan Azcarreta, Jesus Alvarez, Prabhav Agrawal, Ashutosh Pandey, Buye Xu
备注:Accepted to the 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
摘要:由于设备上的有限计算资源,在诸如智能眼镜之类的可穿戴设备中部署语音增强(SE)系统具有挑战性。尽管深度学习方法已经取得了高质量的结果,但其计算成本限制了其在嵌入式平台上的可行性。这项工作提出了一个有效的端到端SE框架,利用差分数字信号处理(DDSP)声码器的高质量的语音合成。首先,一个紧凑的神经网络预测噪声语音的增强声学特征:频谱包络,基频(F0)和周期性。这些特征被馈送到DDSP声码器以合成增强波形。该系统通过STFT和对抗性损失进行端到端训练,从而实现在特征和波形级别的直接优化。实验结果表明,我们的方法在不显著增加计算的情况下,在强基线上将可懂度和质量提高了4%(STOI)和19%(DNSMOS),使其非常适合实时应用。
摘要:Deploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their computational cost limits their feasibility on embedded platforms. This work presents an efficient end-to-end SE framework that leverages a Differentiable Digital Signal Processing (DDSP) vocoder for high-quality speech synthesis. First, a compact neural network predicts enhanced acoustic features from noisy speech: spectral envelope, fundamental frequency (F0), and periodicity. These features are fed into the DDSP vocoder to synthesize the enhanced waveform. The system is trained end-to-end with STFT and adversarial losses, enabling direct optimization at the feature and waveform levels. Experimental results show that our method improves intelligibility and quality by 4% (STOI) and 19% (DNSMOS) over strong baselines without significantly increasing computation, making it well-suited for real-time applications.


【4】A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References
标题:有噪参考语音分离中尺度不变信失真比的研究
链接:https://arxiv.org/abs/2508.14623

作者:Simon Dahl Jepsen, Mads Græsbøll Christensen, Jesper Rindom Jensen
备注:Accepted for IEEE ASRU 2025, Workshop on Automatic Speech Recognition and Understanding. Copyright (c) 2025 IEEE. 8 pages, 6 figures, 2 tables
摘要:本文探讨了使用尺度不变的信号失真比(SI-SDR)作为评估和训练目标的监督语音分离的影响,当训练参考包含噪声,是事实上的基准WSJ 0 - 2 Mix的情况下。具有噪声基准的SI-SDR的推导表明,噪声限制了可实现的SI-SDR,或者导致分离输出中的不期望噪声。为了解决这个问题,提出了一种方法来增强参考和增加与WHAM!,旨在训练避免学习噪声参考的模型。在这些增强的数据集上训练的两个模型使用非侵入式NISQA.v2度量进行评估。结果表明,在分离的语音噪声减少,但建议处理参考可能会引入文物,限制整体质量的收益。在WSJ 0 - 2 Mix和Libri 2 Mix测试集上,发现SI-SDR和感知噪声之间存在负相关性,强调了推导得出的结论。
摘要:This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation.


【5】EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
标题:LLM:用于语音情感识别的LLM的参数有效自适应
链接:https://arxiv.org/abs/2508.14130

作者:Hugo Thimonier, Antony Perzo, Renaud Seguier
摘要:从语音中识别情感是一项具有挑战性的任务,需要捕获语言和非语言线索,在人机交互和心理健康监测中具有重要应用。最近的工作突出了大型语言模型(LLM)在唯一的自然语言区域之外执行任务的能力。特别是,最近的方法已经研究了通过使用预先训练的骨干和不同的融合机制将LLM与其他数据模态耦合。这项工作提出了一种新的方法,微调LLM与音频和文本表示的情感预测。我们的方法首先使用音频特征提取器提取音频特征,然后通过可学习的接口模块将其映射到LLM的表示空间。LLM将(1)变换后的音频特征,(2)自然语言形式的附加特征(例如,抄本),以及(3)描述情绪预测任务的文本提示。为了有效地使LLM适应这种多模态任务,我们采用了低秩自适应(LoRA),从而实现了参数有效的微调。标准的情感识别基准的实验结果表明,我们的模型优于所有,但一个现有的语音-文本LLM在文献中,而需要不到一半的竞争方法的参数。这突出了我们的方法的有效性,在集成多模态输入语音为基础的情感理解,同时保持显着的计算效率。
摘要:Emotion recognition from speech is a challenging task that requires capturing both linguistic and paralinguistic cues, with critical applications in human-computer interaction and mental health monitoring. Recent works have highlighted the ability of Large Language Models (LLMs) to perform tasks outside of the sole natural language area. In particular, recent approaches have investigated coupling LLMs with other data modalities by using pre-trained backbones and different fusion mechanisms. This work proposes a novel approach that fine-tunes an LLM with audio and text representations for emotion prediction. Our method first extracts audio features using an audio feature extractor, which are then mapped into the LLM's representation space via a learnable interfacing module. The LLM takes as input (1) the transformed audio features, (2) additional features in the form of natural language (e.g., the transcript), and (3) a textual prompt describing the emotion prediction task. To efficiently adapt the LLM to this multimodal task, we employ Low-Rank Adaptation (LoRA), enabling parameter-efficient fine-tuning. Experimental results on standard emotion recognition benchmarks demonstrate that our model outperforms all but one existing Speech-Text LLMs in the literature, while requiring less than half the parameters of competing approaches. This highlights our approach's effectiveness in integrating multi-modal inputs for speech-based emotion understanding while maintaining significant computational efficiency.


【6】Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
标题:采用短上下文发言人嵌入的多个发言人实现低延迟跟踪
链接:https://arxiv.org/abs/2508.14115

作者:Taous Iatariene, Alexandre Guérin, Romain Serizel (MULTISPEECH)
备注:None
摘要:说话人嵌入是一种很有前途的身份相关特征,可以通过利用其空间预测,即通过执行身份重新分配来提高跟踪系统的身份分配性能。普通说话人嵌入提取器通常与短时间上下文和重叠语音斗争,这需要长时间的身份重新分配来利用更长的时间上下文。然而,这增加了跟踪系统错误的可能性,这反过来又对身份重新分配产生负面影响。为了解决这个问题,我们提出了一种基于知识蒸馏(KD)的训练方法,用于从两个说话人混合物中提取短上下文说话人嵌入。我们利用感兴趣的扬声器的空间信息,使用波束成形,以减少重叠。我们研究了在固定大小的块上执行身份重新分配的可行性,即,块式身份重新分配,以实现基于低延迟扬声器嵌入的跟踪系统。结果表明,我们的蒸馏模型是有效的短上下文嵌入提取和更强大的重叠。虽然,块重新分配的结果表明,需要进一步的工作,以更有效地处理同步语音。
摘要:Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding extractors usually struggle with short temporal contexts and overlapping speech, which imposes long-term identity reassignment to exploit longer temporal contexts. However, this increases the probability of tracking system errors, which in turn impacts negatively on identity reassignment. To address this, we propose a Knowledge Distillation (KD) based training approach for short context speaker embedding extraction from two speaker mixtures. We leverage the spatial information of the speaker of interest using beamforming to reduce overlap. We study the feasibility of performing identity reassignment over blocks of fixed size, i.e., blockwise identity reassignment, to go towards a low-latency speaker embedding based tracking system. Results demonstrate that our distilled models are effective at short-context embedding extraction and more robust to overlap. Although, blockwise reassignment results indicate that further work is needed to handle simultaneous speech more effectively.


【7】MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
标题:Mahahttps:多语言文本到语音合成的统一框架
链接:https://arxiv.org/abs/2508.14049

作者:Jaskaran Singh, Amartya Roy Chowdhury, Raghav Prabhakar, Varshul C. W
摘要:目前的文本到语音转换模型带来了多语言的挑战,其中大多数模型传统上都集中在英语和欧洲语言上,从而损害了向更多人提供信息的潜力。为了解决这一差距,我们推出了MahaTTS-v2,这是一种多语言多说话者文本到语音(TTS)系统,具有出色的印度语多语言表达能力。该模型已经在大约2万小时的数据上进行了训练,这些数据专门针对印度语言。我们的方法利用Wav2Vec2.0标记进行语义提取,并使用语言模型(LM)进行文本到语义建模。此外,我们使用了一个条件流模型(CFM)的语义melspectogram生成。实验结果表明了该方法的有效性。我们的代码可在https://github.com/dubverse-ai/MahaTTSv2上获得
摘要:Current Text-to-Speech models pose a multilingual challenge, where most of the models traditionally focus on English and European languages, thereby hurting the potential to provide access to information to many more people. To address this gap, we introduce MahaTTS-v2 a Multilingual Multi-speaker Text-To-Speech (TTS) system that has excellent multilingual expressive capabilities in Indic languages. The model has been trained on around 20K hours of data specifically focused on Indian languages. Our approach leverages Wav2Vec2.0 tokens for semantic extraction, and a Language Model (LM) for text-to-semantic modeling. Additionally, we have used a Conditional Flow Model (CFM) for semantics to melspectogram generation. The experimental results indicate the effectiveness of the proposed approach over other frameworks. Our code is available at https://github.com/dubverse-ai/MahaTTSv2


【8】RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
标题:RAG-Boost:检索增强代增强的基于LLM的语音识别
链接:https://arxiv.org/abs/2508.14048

作者:Pengcheng Wang, Sheng Li, Takahiro Shinozaki
备注:accepted at Interspeech2025 MLC-SLM Challenge workshop (task I system description)
摘要:在本文中,我们提出了RAG-Boost(ST-ShinozakiLab任务I系统),它增强了MLC-SLM挑战赛(任务I)的基线基于LLM的ASR系统,并在飞行中使用检索增强生成(RAG)模块。每个部分ASR假设查询音频文本对和领域术语的向量存储,并且检索到的结果与实时ASR假设融合以修复识别错误。融合的假设被传递到LLM,产生改进的响应。
摘要:In this paper, we propose RAG-Boost (ST-ShinozakiLab Task I system), which enhances the baseline LLM-based ASR system of the MLC-SLM Challenge (task I) with a retrieval-augmented generation (RAG) module on the fly. Each partial ASR hypothesis queries a vector store of audio-text pairs and domain terms, and the retrieved results are fused with the live ASR hypotheses to fix recognition errors. The fused hypotheses are passed to the LLM, yielding improved responses.


【9】BioSonix: Can Physics-Based Sonification Perceptualize Tissue Deformations From Tool Interactions?
标题:BioSonix:基于物理的超声可以从工具相互作用中感知组织变形吗?
链接:https://arxiv.org/abs/2508.14688

作者:Veronica Ruozzi, Sasan Matinfar, Laura Schütz, Benedikt Wiestler, Alberto Redaelli, Emiliano Votta, Nassir Navab
备注:V. Ruozzi and S. Matinfar contributed equally to this work
摘要:在外科手术中感知工具与可变形结构的相互作用仍然具有挑战性,因为单峰可视化技术通常由于诸如遮挡和有限的深度感知的约束而无法捕获这些相互作用的复杂性。本文提出了一种新的方法来增强工具导航在混合现实环境中提供听觉表示的工具组织动力学,特别是与软组织的相互作用。BioSonix是一种基于物理学的设计框架,它利用3D空间中的组织位移来计算编码组织特性(如刚度和密度)的声音模型的激励力。采用生物力学模拟来模拟工具-组织相互作用引起的颗粒位移,为该方法建立了坚实的基础。一种优化方法被用来定义配置捕捉不同的互动场景与不同的工具轨迹。实验进行了验证的声音位移映射的准确性。此外,还进行了两项用户研究:第一项研究涉及两名临床专业人员(神经放射科医生和心脏病学家),他们证实了该方法的影响并实现了高任务准确性;第二项研究包括22名生物医学专家,他们在组织分化和靶向任务中表现出高分辨准确性。结果显示,工具-组织动力学与其相应的听觉特征之间存在很强的相关性,突出了这些声音表征增强对复杂相互作用的直观理解的潜力。
摘要:Perceptualizing tool interactions with deformable structures in surgical procedures remains challenging, as unimodal visualization techniques often fail to capture the complexity of these interactions due to constraints such as occlusion and limited depth perception. This paper presents a novel approach to augment tool navigation in mixed reality environments by providing auditory representations of tool-tissue dynamics, particularly for interactions with soft tissue. BioSonix, a physics-informed design framework, utilizes tissue displacements in 3D space to compute excitation forces for a sound model encoding tissue properties such as stiffness and density. Biomechanical simulations were employed to model particle displacements resulting from tool-tissue interactions, establishing a robust foundation for the method. An optimization approach was used to define configurations for capturing diverse interaction scenarios with varying tool trajectories. Experiments were conducted to validate the accuracy of the sound-displacement mappings. Additionally, two user studies were performed: the first involved two clinical professionals (a neuroradiologist and a cardiologist), who confirmed the method's impact and achieved high task accuracy; the second included 22 biomedical experts, who demonstrated high discrimination accuracy in tissue differentiation and targeting tasks. The results revealed a strong correlation between tool-tissue dynamics and their corresponding auditory profiles, highlighting the potential of these sound representations to enhance the intuitive understanding of complex interactions.


【10】Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
标题:Mamba2 Meets Silence:稀疏区域的鲁棒声源分离
链接:https://arxiv.org/abs/2508.14556

作者:Euiyeon Kim, Yong-Hoon Choi
摘要:我们介绍了一种新的音乐源分离模型,专为准确的声乐隔离。与基于transformer的方法不同,这种方法通常无法捕获间歇性出现的声音,我们的模型利用了最近的状态空间模型Mamba 2,以更好地捕获长期的时间依赖性。为了有效地处理长输入序列,我们结合了带分裂策略与双路径架构。实验表明,我们的方法优于最新的最先进的模型,实现了11.03 dB的cSDR-迄今为止最好的报告-并提供了大量的收益在uSDR。此外,该模型在不同的输入长度和声音发生模式下表现出稳定和一致的性能。这些结果证明了基于Mamba的模型用于高分辨率音频处理的有效性,并为音频研究中更广泛的应用开辟了新的方向。
摘要:We introduce a new music source separation model tailored for accurate vocal isolation. Unlike Transformer-based approaches, which often fail to capture intermittently occurring vocals, our model leverages Mamba2, a recent state space model, to better capture long-range temporal dependencies. To handle long input sequences efficiently, we combine a band-splitting strategy with a dual-path architecture. Experiments show that our approach outperforms recent state-of-the-art models, achieving a cSDR of 11.03 dB-the best reported to date-and delivering substantial gains in uSDR. Moreover, the model exhibits stable and consistent performance across varying input lengths and vocal occurrence patterns. These results demonstrate the effectiveness of Mamba-based models for high-resolution audio processing and open up new directions for broader applications in audio research.


【11】EmoTale: An Enacted Speech-emotion Dataset in Danish
标题:收件箱故事:丹麦语已制定的言语情感数据集
链接:https://arxiv.org/abs/2508.14548

作者:Maja J. Hjuler, Harald V. Skat-Rørdam, Line H. Clemmensen, Sneha Das
备注:To appear in the proceedings of ASRU 2025
摘要:虽然对于常用语言存在多个情感语音语料库,但对于较小的(口语)语言(如丹麦语)缺乏功能数据集。据我们所知,丹麦情绪语音(DES),发表于1997年,是丹麦情绪语音的唯一的其他数据库。我们目前的童话,语料库包括丹麦语和英语的语音录音与他们相关的制定情绪注释。我们证明了数据集的有效性,通过调查和提出其预测能力,使用语音情感识别(SER)模型。我们使用自监督语音模型(SSLM)嵌入和openSMILE特征提取器为ESTALE和参考数据集开发SER模型。我们发现嵌入优于手工制作的功能。最好的模型实现了64.1%的未加权平均召回率(UAR)的EQUIPTale语料库使用留一说话人交叉验证,DES上的性能相当。
摘要:While multiple emotional speech corpora exist for commonly spoken languages, there is a lack of functional datasets for smaller (spoken) languages, such as Danish. To our knowledge, Danish Emotional Speech (DES), published in 1997, is the only other database of Danish emotional speech. We present EmoTale; a corpus comprising Danish and English speech recordings with their associated enacted emotion annotations. We demonstrate the validity of the dataset by investigating and presenting its predictive power using speech emotion recognition (SER) models. We develop SER models for EmoTale and the reference datasets using self-supervised speech model (SSLM) embeddings and the openSMILE feature extractor. We find the embeddings superior to the hand-crafted features. The best model achieves an unweighted average recall (UAR) of 64.1% on the EmoTale corpus using leave-one-speaker-out cross-validation, comparable to the performance on DES.


【12】Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
标题:心理健康和神经退行性疾病的开放语音生物标志物数据集的系统公平性评估
链接:https://arxiv.org/abs/2508.14089

作者:Ishaan Mahapatra, Nihar R. Mahapatra
备注:To appear in the Proceedings of the 28th International Conference on Text, Speech and Dialogue (TSD 2025), Erlangen, Germany, August 25-28, 2025
摘要:语音生物标志物-人类产生的声音信号,如语音,咳嗽和呼吸-是有前途的工具,可扩展,非侵入性检测和监测心理健康和神经退行性疾病。然而,它们的临床应用仍然受到公开数据集质量不一致和可用性有限的限制。为了解决这一差距,我们提出了第一个系统的FAIR(可发现,可解释,可互操作,可重复使用)评估27个公开的声音生物标志物数据集集中在这些疾病领域。使用FAIR数据成熟度模型和结构化的优先级加权评分方法,我们评估了子原则,原则和复合水平的公平性。我们的分析显示,可查找性一直很高,但在可访问性、互操作性和可重用性方面存在很大的可变性和弱点。精神健康数据集在FAIR评分中表现出更大的变异性,而神经退行性疾病数据集则稍微一致。存储库的选择也显着影响公平分数。为了提高数据集质量和临床实用性,我们建议采用结构化的、特定领域的元数据标准,优先考虑符合FAIR的存储库,并定期应用结构化的FAIR评估框架。这些发现为改善数据集的互操作性和重用提供了可操作的指导,从而加速了语音生物标志物技术的临床转化。
摘要:Voice biomarkers--human-generated acoustic signals such as speech, coughing, and breathing--are promising tools for scalable, non-invasive detection and monitoring of mental health and neurodegenerative diseases. Yet, their clinical adoption remains constrained by inconsistent quality and limited usability of publicly available datasets. To address this gap, we present the first systematic FAIR (Findable, Accessible, Interoperable, Reusable) evaluation of 27 publicly available voice biomarker datasets focused on these disease areas. Using the FAIR Data Maturity Model and a structured, priority-weighted scoring method, we assessed FAIRness at subprinciple, principle, and composite levels. Our analysis revealed consistently high Findability but substantial variability and weaknesses in Accessibility, Interoperability, and Reusability. Mental health datasets exhibited greater variability in FAIR scores, while neurodegenerative datasets were slightly more consistent. Repository choice also significantly influenced FAIRness scores. To enhance dataset quality and clinical utility, we recommend adopting structured, domain-specific metadata standards, prioritizing FAIR-compliant repositories, and routinely applying structured FAIR evaluation frameworks. These findings provide actionable guidance to improve dataset interoperability and reuse, thereby accelerating the clinical translation of voice biomarker technologies.


机器翻译由腾讯交互翻译提供,仅供参考