今日论文合集:cs.SD语音16篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
标题:CS-FLEURS:大规模多语言和代码交换语音数据集
链接:https://arxiv.org/abs/2509.14161

作者:Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Lodagala, William Chen, Olga Iakovenko, Bashar Talafha, Amir Hussein, Alexander Polok, Kalvin Chang, Dominik Klement, Sara Althubaiti, Puyuan Peng, Matthew Wiesner, Thamar Solorio, Ahmed Ali, Sanjeev Khudanpur, Shinji Watanabe, Chih-Chen Chen, Zhen Wu, Karim Benharrak, Anuj Diwan, Samuele Cornell, Eunjung Yeo, Kwanghee Choi, Carlos Carvalho, Karen Rosero
摘要:我们提出了CS-FLEURS,这是一个新的数据集,用于开发和评估高资源语言之外的代码切换语音识别和翻译系统。CS-FLEURS由4个测试集组成,涵盖52种语言的113个独特的代码转换语言对:1)14 X-英语语言对集合,其具有阅读合成生成的代码转换句子的真实语音,2)16 X-英语语言对集合,其具有生成的文本到语音,3)60 {阿拉伯语,普通话,印地语,西班牙语}-X语言对测试集,具有生成的文本到语音,以及4)45个X-英语低资源语言对测试集,具有连接的文本到语音。除了四个测试集外,CS-FLEURS还提供了一个训练集,其中包含16个X英语语言对的128小时生成文本到语音数据。我们希望CS-FLEURS有助于拓宽未来语码转换语音研究的范围。数据集链接:https://huggingface.co/datasets/byan/cs-fleurs。
摘要:We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 unique code-switched language pairs across 52 languages: 1) a 14 X-English language pair set with real voices reading synthetically generated code-switched sentences, 2) a 16 X-English language pair set with generative text-to-speech 3) a 60 {Arabic, Mandarin, Hindi, Spanish}-X language pair set with the generative text-to-speech, and 4) a 45 X-English lower-resourced language pair test set with concatenative text-to-speech. Besides the four test sets, CS-FLEURS also provides a training set with 128 hours of generative text-to-speech data across 16 X-English language pairs. Our hope is that CS-FLEURS helps to broaden the scope of future code-switched speech research. Dataset link: https://huggingface.co/datasets/byan/cs-fleurs.


【2】AnyAccomp: Generalizable Accompaniment Generation via Quantized Melodic Bottleneck
标题:AnyAccopp:通过量化旋律瓶颈的可概括伴奏生成
链接:https://arxiv.org/abs/2509.14052

作者:Junan Zhang, Yunjia Zhang, Xueyao Zhang, Zhizheng Wu
备注:Demo audio and code: this https URL
摘要:歌唱伴奏生成(SAG)是为给定的干净的声乐输入生成器乐的过程。然而,现有的SAG技术使用源分离的人声作为输入和过拟合分离工件。这会造成关键的训练-测试不匹配,导致在干净的真实声音输入上失败。我们介绍AnyAccomp,一个框架,解决了这个问题,从依赖于源代码的文物解耦伴奏生成。AnyAccomp首先采用量化的旋律瓶颈,使用色度图和VQ-VAE来提取核心旋律的离散和音色不变的表示。随后的流匹配模型,然后生成这些强大的代码的伴奏条件。实验表明,AnyAccomp在单独的声乐基准上实现了具有竞争力的性能,同时在干净的录音室声乐和独奏乐器曲目的泛化测试集上显著优于基线。这证明了泛化的质的飞跃,实现了乐器的强大伴奏-现有模型完全失败的任务-并为更多功能的音乐共同创作工具铺平了道路。演示音频和代码:https://anyaccomp.github.io
摘要:Singing Accompaniment Generation (SAG) is the process of generating instrumental music for a given clean vocal input. However, existing SAG techniques use source-separated vocals as input and overfit to separation artifacts. This creates a critical train-test mismatch, leading to failure on clean, real-world vocal inputs. We introduce AnyAccomp, a framework that resolves this by decoupling accompaniment generation from source-dependent artifacts. AnyAccomp first employs a quantized melodic bottleneck, using a chromagram and a VQ-VAE to extract a discrete and timbre-invariant representation of the core melody. A subsequent flow-matching model then generates the accompaniment conditioned on these robust codes. Experiments show AnyAccomp achieves competitive performance on separated-vocal benchmarks while significantly outperforming baselines on generalization test sets of clean studio vocals and, notably, solo instrumental tracks. This demonstrates a qualitative leap in generalization, enabling robust accompaniment for instruments - a task where existing models completely fail - and paving the way for more versatile music co-creation tools. Demo audio and code: https://anyaccomp.github.io


【3】Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
标题:资源受限设备上基于CNN的音频标记模型的综合评估
链接:https://arxiv.org/abs/2509.14049

作者:Jordi Grau-Haro, Ruben Ribes-Serrano, Javier Naranjo-Alcazar, Marta Garcia-Ballesteros, Pedro Zuccarello
备注:Accepted at Computing Conference 2026, London, UK
摘要:卷积神经网络(CNN)在音频标记任务中表现出卓越的性能。然而,在资源受限的设备(如Raspberry Pi)上部署这些模型会带来与计算效率和热管理相关的挑战。在本文中,对Raspberry Pi上用于音频标记的多卷积神经网络(CNN)架构进行了全面评估,包括来自预训练音频神经网络(PANN)框架的所有1D和2D模型,适用于音频分类的基于ConvNeXT的模型,以及MobileNetV3架构。此外,两个PANN衍生的网络,CNN9和CNN13,最近提出的,也进行了评估。为了提高跨不同硬件平台的部署效率和可移植性,所有模型都转换为开放神经网络交换(ONNX)格式。与以前专注于单个模型的工作不同,我们的分析涵盖了更广泛的架构,并涉及连续24小时的推理会话来评估性能稳定性。我们的实验表明,通过适当的模型选择和优化,可以保持一致的推理延迟,并在较长的时间内有效地管理热行为。这些发现为在现实世界的边缘计算场景中部署音频标记模型提供了有价值的见解。
摘要:Convolutional Neural Networks (CNNs) have demonstrated exceptional performance in audio tagging tasks. However, deploying these models on resource-constrained devices like the Raspberry Pi poses challenges related to computational efficiency and thermal management. In this paper, a comprehensive evaluation of multiple convolutional neural network (CNN) architectures for audio tagging on the Raspberry Pi is conducted, encompassing all 1D and 2D models from the Pretrained Audio Neural Networks (PANNs) framework, a ConvNeXt-based model adapted for audio classification, as well as MobileNetV3 architectures. In addition, two PANNs-derived networks, CNN9 and CNN13, recently proposed, are also evaluated. To enhance deployment efficiency and portability across diverse hardware platforms, all models are converted to the Open Neural Network Exchange (ONNX) format. Unlike previous works that focus on a single model, our analysis encompasses a broader range of architectures and involves continuous 24-hour inference sessions to assess performance stability. Our experiments reveal that, with appropriate model selection and optimization, it is possible to maintain consistent inference latency and manage thermal behavior effectively over extended periods. These findings provide valuable insights for deploying audio tagging models in real-world edge computing scenarios.


【4】RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing
标题:RFM编辑:用于文本引导音频编辑的纠正流匹配
链接:https://arxiv.org/abs/2509.14003

作者:Liting Gao, Yi Yuan, Yaru Chen, Yuelan Cheng, Zhenbo Li, Juan Wen, Shubin Zhang, Wenwu Wang
摘要:扩散模型在文本到音频生成方面取得了显着进展。然而,文本引导的音频编辑仍处于早期阶段。该任务的重点是修改音频信号中的目标内容,同时保留其余内容,因此需要根据文本提示进行精确定位和忠实编辑。现有的基于训练和zero-shot的方法依赖于完整的字幕或昂贵的优化,通常难以进行复杂的编辑或缺乏实用性。在这项工作中,我们提出了一种新的端到端的高效整流匹配为基础的音频编辑的扩散框架,并构建了一个数据集,具有重叠的多事件音频,以支持复杂场景中的训练和基准测试。实验表明,我们的模型实现了忠实的语义对齐,而不需要辅助字幕或掩码,同时保持竞争力的编辑质量跨指标。
摘要:Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest, thus demanding precise localization and faithful editing according to the text prompt. Existing training-based and zero-shot methods that rely on full-caption or costly optimization often struggle with complex editing or lack practicality. In this work, we propose a novel end-to-end efficient rectified flow matching-based diffusion framework for audio editing, and construct a dataset featuring overlapping multi-event audio to support training and benchmarking in complex scenarios. Experiments show that our model achieves faithful semantic alignment without requiring auxiliary captions or masks, while maintaining competitive editing quality across metrics.


【5】Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound Detection
标题:基于噪声监督的对比学习和扰动的异常声音检测
链接:https://arxiv.org/abs/2509.13853

作者:Shun Huang, Zhihua Fang, Liang He
备注:Accept ICASSP 2025
摘要:无监督异常声音检测的目的是通过仅使用正常音频数据训练模型来检测未知的异常声音。尽管自我监督方法取得了进步,但在处理来自不同机器的相同类型的样本时频繁的假警报问题仍然没有得到解决。本文介绍了一种新的训练技术,称为一阶段监督对比学习(OS-SCL),它显着地解决了这个问题,扰动嵌入空间中的功能,并采用一个阶段的噪声监督对比学习方法。在DCASE 2020挑战任务2中,仅使用Log-Mel特征,其实现了94.64\% AUC、88.42\% pAUC和89.24\% mAUC。此外,一个名为TFgram的时频特征,这是从原始音频提取。该功能有效地捕获了异常声音检测的关键信息,最终达到95.71\% AUC,90.23\% pAUC和91.23\% mAUC。源代码可在以下位置获得:\underline{www.github.com/huangswt/OS-SCL}。
摘要:Unsupervised anomalous sound detection aims to detect unknown anomalous sounds by training a model using only normal audio data. Despite advancements in self-supervised methods, the issue of frequent false alarms when handling samples of the same type from different machines remains unresolved. This paper introduces a novel training technique called one-stage supervised contrastive learning (OS-SCL), which significantly addresses this problem by perturbing features in the embedding space and employing a one-stage noisy supervised contrastive learning approach. On the DCASE 2020 Challenge Task 2, it achieved 94.64\% AUC, 88.42\% pAUC, and 89.24\% mAUC using only Log-Mel features. Additionally, a time-frequency feature named TFgram is proposed, which is extracted from raw audio. This feature effectively captures critical information for anomalous sound detection, ultimately achieving 95.71\% AUC, 90.23\% pAUC, and 91.23\% mAUC. The source code is available at: \underline{www.github.com/huangswt/OS-SCL}.


【6】Neural Speech Separation with Parallel Amplitude and Phase Spectrum Estimation
标题:并行幅度和相谱估计的神经语音分离
链接:https://arxiv.org/abs/2509.13825

作者:Fei Liu, Yang Ai, Zhen-Hua Ling
备注:Accepted by APSIPA2025
摘要:本文提出了一种新的并行幅度和相位谱估计的神经语音分离模型APSS。与大多数现有的语音分离方法不同,APSS通过显式地估计相位谱来实现更完整和准确的分离。具体地说,APSS首先从混合语音信号中提取幅度谱和相位谱。随后,提取的幅度和相位谱由特征组合器融合成联合表示,然后由具有时频Transformers的深度处理器进一步处理,以捕获时间和频谱依赖性。最后,利用并行的幅度和相位分离器,APSS从所得到的特征中估计每个扬声器的相应频谱,然后通过逆短时傅立叶变换(iSTFT)将其组合以重构分离的语音信号。实验结果表明,APSS优于时域分离方法和基于隐相位估计的时频方法。此外,APSS在多个数据集上取得了稳定和有竞争力的结果,突出了其强大的泛化能力和实用性。
摘要:This paper proposes APSS, a novel neural speech separation model with parallel amplitude and phase spectrum estimation. Unlike most existing speech separation methods, the APSS distinguishes itself by explicitly estimating the phase spectrum for more complete and accurate separation. Specifically, APSS first extracts the amplitude and phase spectra from the mixed speech signal. Subsequently, the extracted amplitude and phase spectra are fused by a feature combiner into joint representations, which are then further processed by a deep processor with time-frequency Transformers to capture temporal and spectral dependencies. Finally, leveraging parallel amplitude and phase separators, the APSS estimates the respective spectra for each speaker from the resulting features, which are then combined via inverse short-time Fourier transform (iSTFT) to reconstruct the separated speech signals. Experimental results indicate that APSS surpasses both time-domain separation methods and implicit-phase-estimation-based time-frequency approaches. Also, APSS achieves stable and competitive results on multiple datasets, highlighting its strong generalization capability and practical applicability.


【7】Invisible Ears at Your Fingertips: Acoustic Eavesdropping via Mouse Sensors
标题:指尖看不见的耳朵:通过鼠标传感器落下的声音发射器
链接:https://arxiv.org/abs/2509.13581

作者:Mohamad Habib Fakih, Rahul Dharmaji, Youssef Mahmoud, Halima Bouzidi, Mohammad Abdullah Al Faruque
备注:Appearing in the Annual Computer Security Applications Conference (ACSAC 2025)
摘要:现代光学鼠标传感器具有先进的精度和高响应能力,但却存在一个经常被忽视的漏洞:它们可能被用于侧信道攻击。本文介绍了Mic-E-Mouse,这是有史以来第一次针对高性能光学鼠标传感器的侧信道攻击,以秘密窃听用户。我们证明,音频信号可以引起微妙的表面振动检测鼠标的光学传感器。值得注意的是,流行操作系统上的用户空间软件可以收集和广播这种敏感的侧信道,允许攻击者访问原始鼠标数据,而无需直接的系统级权限。最初,从鼠标数据中提取的振动信号由于非均匀采样、非线性频率响应和显著量化而质量差。为了克服这些限制,Mic-E-Mouse采用了一个复杂的端到端数据过滤管道,该管道结合了维纳过滤、恢复校正和创新的仅编码器频谱图神经过滤技术。我们评估了攻击在不同条件下的效果,包括说话音量、鼠标轮询率和DPI、表面材料、说话者语言和环境噪声。在受控环境中,Mic-E-Mouse可将语音重建的信噪比(SNR)提高高达+19 dB。此外,我们的研究结果表明,在AudioMNIST和VCTK数据集上的语音识别准确率约为42%至61%。我们所有的代码和数据集都可以在https://sites.google.com/view/mic-e-mouse上公开访问。
摘要:Modern optical mouse sensors, with their advanced precision and high responsiveness, possess an often overlooked vulnerability: they can be exploited for side-channel attacks. This paper introduces Mic-E-Mouse, the first-ever side-channel attack that targets high-performance optical mouse sensors to covertly eavesdrop on users. We demonstrate that audio signals can induce subtle surface vibrations detectable by a mouse's optical sensor. Remarkably, user-space software on popular operating systems can collect and broadcast this sensitive side channel, granting attackers access to raw mouse data without requiring direct system-level permissions. Initially, the vibration signals extracted from mouse data are of poor quality due to non-uniform sampling, a non-linear frequency response, and significant quantization. To overcome these limitations, Mic-E-Mouse employs a sophisticated end-to-end data filtering pipeline that combines Wiener filtering, resampling corrections, and an innovative encoder-only spectrogram neural filtering technique. We evaluate the attack's efficacy across diverse conditions, including speaking volume, mouse polling rate and DPI, surface materials, speaker languages, and environmental noise. In controlled environments, Mic-E-Mouse improves the signal-to-noise ratio (SNR) by up to +19 dB for speech reconstruction. Furthermore, our results demonstrate a speech recognition accuracy of roughly 42% to 61% on the AudioMNIST and VCTK datasets. All our code and datasets are publicly accessible on https://sites.google.com/view/mic-e-mouse.


【8】Field of View Enhanced Signal Dependent Binauralization with Mixture of Experts Framework for Continuous Source Motion
标题:视野增强信号相关的双耳化,采用混合专家框架实现连续源运动
链接:https://arxiv.org/abs/2509.13548

作者:Manan Mittal, Thomas Deppisch, Joseph Forrer, Chris Le Sueur, Zamir Ben-Hur, David Lou Along, Daniel D.E. Wong
备注:5 pages, 3 figures
摘要:我们提出了一种新的混合专家框架的视场增强双耳信号匹配。我们的方法可以实现动态空间音频渲染,适应连续的说话者运动,允许用户强调或抑制来自选定方向的声音,同时保留自然的双耳线索。与依赖于显式到达方向估计或在高保真度立体声域中操作的传统方法不同,我们的信号依赖框架使用隐式定位以在线方式组合多个双耳滤波器。这允许实时跟踪和增强移动声源,支持增强和虚拟现实中的语音聚焦、降噪和世界锁定音频等应用。该方法与阵列几何形状无关,为下一代消费音频设备中的空间音频捕获和个性化回放提供了灵活的解决方案。
摘要:We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress sounds from selected directions while preserving natural binaural cues. Unlike traditional methods that rely on explicit direction-of-arrival estimation or operate in the Ambisonics domain, our signal-dependent framework combines multiple binaural filters in an online manner using implicit localization. This allows for real-time tracking and enhancement of moving sound sources, supporting applications such as speech focus, noise reduction, and world-locked audio in augmented and virtual reality. The method is agnostic to array geometry offering a flexible solution for spatial audio capture and personalized playback in next-generation consumer audio devices.


【9】A Domain Knowledge Informed Approach for Anomaly Detection of Electric Vehicle Interior Sounds
标题:电动汽车车内声音异常检测领域知识知情方法
链接:https://arxiv.org/abs/2509.13390

作者:Deepti Kunte, Bram Cornelis, Claudio Colangeli, Karl Janssens, Brecht Van Baelen, Konstantinos Gryllias
备注:Submitted to: Mechanical Systems and Signal Processing
摘要:汽车驾驶室声音异常的检测对于确保车辆质量和保持乘客舒适性至关重要。在许多现实环境中,由于标记错误数据的稀缺性或完全不存在,该任务更适合作为无监督学习问题而不是监督情况。在这种无监督设置中,模型只在健康样本上进行训练,并将异常检测为与正常行为的偏差。然而,在缺乏用于验证的标记错误样本以及常用度量的有限可靠性(例如验证重建误差)的情况下,有效的模型选择仍然是一个重大挑战。为了克服这些限制,提出了一种基于领域知识的模型选择方法,其中通过健康谱图的结构化扰动设计的代理异常用于验证集以支持模型选择。在高保真电动汽车数据集上评估了所提出的方法,该数据集包括五种代表性故障类型的健康和故障车厢声音,即,不平衡、调制、呜呜声、风和脉宽调制。该数据集使用先进的声音合成技术生成,并通过专家评审团评估进行验证,已公开提供,以促进进一步的研究。五个故障情况下的实验评估表明,使用代理异常的最佳模型的选择,显着优于传统的模型选择策略。
摘要:The detection of anomalies in automotive cabin sounds is critical for ensuring vehicle quality and maintaining passenger comfort. In many real-world settings, this task is more appropriately framed as an unsupervised learning problem rather than the supervised case due to the scarcity or complete absence of labeled faulty data. In such an unsupervised setting, the model is trained exclusively on healthy samples and detects anomalies as deviations from normal behavior. However, in the absence of labeled faulty samples for validation and the limited reliability of commonly used metrics, such as validation reconstruction error, effective model selection remains a significant challenge. To overcome these limitations, a domain-knowledge-informed approach for model selection is proposed, in which proxy-anomalies engineered through structured perturbations of healthy spectrograms are used in the validation set to support model selection. The proposed methodology is evaluated on a high-fidelity electric vehicle dataset comprising healthy and faulty cabin sounds across five representative fault types viz., Imbalance, Modulation, Whine, Wind, and Pulse Width Modulation. This dataset, generated using advanced sound synthesis techniques, and validated via expert jury assessments, has been made publicly available to facilitate further research. Experimental evaluations on the five fault cases demonstrate the selection of optimal models using proxy-anomalies, significantly outperform conventional model selection strategies.


【10】A Lightweight Fourier-based Network for Binaural Speech Enhancement with Spatial Cue Preservation
标题:具有空间线索保留的基于傅里叶的轻量级网络用于双耳语音增强
链接:https://arxiv.org/abs/2509.14076

作者:Xikun Lu, Yujian Ma, Xianquan Jiang, Xuelong Wang, Jinqiu Sang
备注:Submitted to ICASSP 2026
摘要:双耳语音增强面临着一个严峻的权衡挑战,其中最先进的性能是通过计算密集型架构实现的,而轻量级解决方案往往以显著的性能下降为代价。为了弥合这一差距,我们提出了全局自适应傅立叶网络(GAF-Net),这是一种轻量级的深度复杂网络,旨在建立性能和计算效率之间的平衡。GAF-Net架构由三个部分组成。首先,结合短时傅立叶变换和伽玛特征的双特征编码器增强了声学表示的鲁棒性。第二,通道独立的全局自适应傅立叶调制器有效地捕获长期的时间依赖性,同时保留空间线索。最后,实现了动态门控机制以减少处理伪影。实验结果表明,GAF-Net实现了有竞争力的性能,特别是在双耳线索(ILD和IPD错误)和客观可懂度(MBSTOI)方面,具有更少的参数和计算成本。这些结果证实,GAF-Net提供了一种可行的方法来实现高保真双耳处理资源受限的设备。
摘要:Binaural speech enhancement faces a severe trade-off challenge, where state-of-the-art performance is achieved by computationally intensive architectures, while lightweight solutions often come at the cost of significant performance degradation. To bridge this gap, we propose the Global Adaptive Fourier Network (GAF-Net), a lightweight deep complex network that aims to establish a balance between performance and computational efficiency. The GAF-Net architecture consists of three components. First, a dual-feature encoder combining short-time Fourier transform and gammatone features enhances the robustness of acoustic representation. Second, a channel-independent globally adaptive Fourier modulator efficiently captures long-term temporal dependencies while preserving the spatial cues. Finally, a dynamic gating mechanism is implemented to reduce processing artifacts. Experimental results show that GAF-Net achieves competitive performance, particularly in terms of binaural cues (ILD and IPD error) and objective intelligibility (MBSTOI), with fewer parameters and computational cost. These results confirm that GAF-Net provides a feasible way to achieve high-fidelity binaural processing on resource-constrained devices.


【11】Lightweight Implicit Neural Network for Binaural Audio Synthesis
标题:用于双耳音频合成的轻量级隐式神经网络
链接:https://arxiv.org/abs/2509.14069

作者:Xikun Lu, Fang Liu, Weizhi Shi, Jinqiu Sang
备注:Submitted to ICASSP 2026
摘要:高保真双耳音频合成对于沉浸式聆听至关重要,但现有方法需要大量的计算资源,限制了其边缘设备应用。为了解决这个问题,我们提出了轻量级隐式神经网络(LINN),一个新的两阶段框架。LINN首先使用时域扭曲生成初始估计,然后通过隐式双耳校正器(IBC)模块进行细化。IBC是一种隐式神经网络,可直接预测振幅和相位校正,从而形成高度紧凑的模型架构。实验结果表明,LINN实现了统计上可比的感知质量的最佳性能的基线模型,同时显着提高计算效率。与现有最有效的方法相比,LINN实现了72.7%的参数减少和显着减少的计算操作(MAC)。这表明我们的方法有效地解决了合成质量和计算效率之间的权衡,为高保真边缘设备空间音频应用提供了一种新的解决方案。
摘要:High-fidelity binaural audio synthesis is crucial for immersive listening, but existing methods require extensive computational resources, limiting their edge-device application. To address this, we propose the Lightweight Implicit Neural Network (LINN), a novel two-stage framework. LINN first generates initial estimates using a time-domain warping, which is then refined by an Implicit Binaural Corrector (IBC) module. IBC is an implicit neural network that predicts amplitude and phase corrections directly, resulting in a highly compact model architecture. Experimental results show that LINN achieves statistically comparable perceptual quality to the best-performing baseline model while significantly improving computational efficiency. Compared to the most efficient existing method, LINN achieves a 72.7% reduction in parameters and significantly fewer compute operations (MACs). This demonstrates that our approach effectively addresses the trade-off between synthesis quality and computational efficiency, providing a new solution for high-fidelity edge-device spatial audio applications.


【12】Network representations reveal structured uncertainty in music
标题:网络表示揭示了音乐中的结构化不确定性
链接:https://arxiv.org/abs/2509.14053

作者: Lluc Bono Rosselló, Robert Jankowski, Hugues Bersini, Marián Boguñá, M. Ángeles Serrano
摘要:音乐作为一种结构化但感知丰富的体验,可以被建模为一个网络,以揭示人类如何编码和处理听觉信息。虽然基于网络的音乐表征越来越普遍,但特征选择对结构属性和认知对齐的影响仍然没有得到充分研究。在这项研究中,我们评估了八个网络模型,每个模型都是从钢琴作品的符号表示中构建的,使用音高,八度音阶,持续时间和间隔的不同组合,旨在代表文献中的现有方法。通过比较这些模型,通过拓扑度量,熵分析,和分歧方面推断的认知表示,我们评估了他们的结构和感知效率。我们的研究结果表明,更简单的,特定于特征的模型更好地匹配人类的感知,而复杂的,多维的表示引入认知效率低下。这些结果支持了人类依赖于模块化、并行认知网络的观点--这种架构与预测处理和自由能最小化理论一致。此外,我们发现,音乐网络的结构组织,以引导注意力转向过渡,这是不确定的和可推断的。由此产生的结构将不确定性集中在一些经常访问的节点上,创造出在稳定和不可预测区域之间交替的局部熵梯度,从而实现了定义音乐体验的张力和释放的表达动态。这些发现表明,网络结构使音乐中的不确定性组织变得可观察,为期望的模式化流动如何塑造感知提供了新的见解,并通过网络科学的镜头为研究音乐结构如何在流派,文化和历史时期演变开辟了新的方向。
摘要:Music, as a structured yet perceptually rich experience, can be modeled as a network to uncover how humans encode and process auditory information. While network-based representations of music are increasingly common, the impact of feature selection on structural properties and cognitive alignment remains underexplored. In this study, we evaluated eight network models, each constructed from symbolic representations of piano compositions using distinct combinations of pitch, octave, duration, and interval, designed to be representative of existing approaches in the literature. By comparing these models through topological metrics, entropy analysis, and divergence with respect to inferred cognitive representations, we assessed both their structural and perceptual efficiency. Our findings reveal that simpler, feature-specific models better match human perception, whereas complex, multidimensional representations introduce cognitive inefficiencies. These results support the view that humans rely on modular, parallel cognitive networks--an architecture consistent with theories of predictive processing and free energy minimization. Moreover, we find that musical networks are structurally organized to guide attention toward transitions that are both uncertain and inferable. The resulting structure concentrates uncertainty in a few frequently visited nodes, creating local entropy gradients that alternate between stable and unpredictable regions, thereby enabling the expressive dynamics of tension and release that define the musical experience. These findings show that network structures make the organization of uncertainty in music observable, offering new insight into how patterned flows of expectation shape perception, and open new directions for studying how musical structures evolve across genres, cultures, and historical periods through the lens of network science.


【13】DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
标题:DSpAST:使用大型语言模型的空间音频推理的分离表示
链接:https://arxiv.org/abs/2509.13927

作者:Kevin Wilkinghoff, Zheng-Hua Tan
摘要:对于具有大型语言模型的空间音频的推理需要空间音频编码器作为声学前端以获得音频嵌入用于进一步处理。这样的编码器需要捕获检测声音事件类型所需的所有信息,以及它们对应的源的方向和距离。用单个音频编码器完成这一任务是非常苛刻的,因为这些任务中的每一个所需的信息大多是相互独立的。因此,使用单个编码器获得的性能通常比使用特定于任务的音频编码器时更差。在这项工作中,我们提出了DSpAST,一种新的音频编码器的基础上SpatialAST,学习空间音频的解纠缠表示,而只有0.2%的额外参数。利用空间音频推理系统BAT对SpatialSoundQA进行实验,结果表明DSpAST的性能明显优于SpatialAST.
摘要:Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the type of sound events, as well as the direction and distance of their corresponding sources. Accomplishing this with a single audio encoder is demanding as the information required for each of these tasks is mostly independent of each other. As a result, the performance obtained with a single encoder is often worse than when using task-specific audio encoders. In this work, we present DSpAST, a novel audio encoder based on SpatialAST that learns disentangled representations of spatial audio while having only 0.2% additional parameters. Experiments on SpatialSoundQA with the spatial audio reasoning system BAT demonstrate that DSpAST significantly outperforms SpatialAST.


【14】Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
标题:可推广音频深度伪造检测中的低级别适配器专家混合
链接:https://arxiv.org/abs/2509.13878

作者:Janne Laakkonen, Ivan Kukanov, Ville Hautamäki
备注:6 pages, 3 figures, 1 table
摘要:像Wav 2 Vec 2这样的基础模型擅长语音任务中的表示学习,包括音频deepfake检测。然而,在对一组固定的真实和欺骗的音频片段进行微调后,它们往往无法推广到训练中没有出现的新的deepfake方法。为了解决这个问题,我们提出了一种混合LoRA专家的方法,将多个低秩适配器(LoRA)集成到模型的注意力层中。路由机制有选择地激活专业专家,增强对不断发展的deepfake攻击的适应性。实验结果表明,我们的方法优于标准的微调,在域内和域外的情况下,减少相等的错误率相对于基线模型。值得注意的是,我们最好的MoE-LoRA模型将平均域外EER从8.55\%降低到6.08\%,证明了其在实现可推广的音频深度伪造检测方面的有效性。
摘要:Foundation models such as Wav2Vec2 excel at representation learning in speech tasks, including audio deepfake detection. However, after being fine-tuned on a fixed set of bonafide and spoofed audio clips, they often fail to generalize to novel deepfake methods not represented in training. To address this, we propose a mixture-of-LoRA-experts approach that integrates multiple low-rank adapters (LoRA) into the model's attention layers. A routing mechanism selectively activates specialized experts, enhancing adaptability to evolving deepfake attacks. Experimental results show that our method outperforms standard fine-tuning in both in-domain and out-of-domain scenarios, reducing equal error rates relative to baseline models. Notably, our best MoE-LoRA model lowers the average out-of-domain EER from 8.55\% to 6.08\%, demonstrating its effectiveness in achieving generalizable audio deepfake detection.


【15】Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
标题:多语言对话语音语言模型挑战总结:数据集、任务、基线和方法
链接:https://arxiv.org/abs/2509.13785

作者:Bingshen Mu, Pengcheng Guo, Zhaokai Sun, Shuai Wang, Hexin Liu, Mingchen Shao, Lei Xie, Eng Siong Chng, Longshuai Xiao, Qiangze Feng, Daliang Wang
摘要:本文总结了Interspeech 2025多语言会话语音语言模型(MLC-SLM)挑战,旨在推进构建有效的多语言会话语音LLM(SLLM)的探索。我们详细描述了MLC-SLM挑战赛的任务设置、已发布的总计约1,604小时的真实世界多语言会话语音数据集以及参与者的基线系统。MLC-SLM挑战赛吸引了来自13个国家的78支队伍参加,两项任务共获得489项有效排行榜结果和14份技术报告。我们根据参与者提交的意见,提炼出构建多语言会话SLLM的宝贵见解,旨在为社区的发展做出贡献。
摘要:This paper summarizes the Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) challenge, which aims to advance the exploration of building effective multilingual conversational speech LLMs (SLLMs). We provide a detailed description of the task settings for the MLC-SLM challenge, the released real-world multilingual conversational speech dataset totaling approximately 1,604 hours, and the baseline systems for participants. The MLC-SLM challenge attracts 78 teams from 13 countries to participate, with 489 valid leaderboard results and 14 technical reports for the two tasks. We distill valuable insights on building multilingual conversational SLLMs based on submissions from participants, aiming to contribute to the advancement of the community.


【16】Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
标题:利用DS SCNet和跨数据库自适应增强与说话者无关的发音障碍语音严重程度分类
链接:https://arxiv.org/abs/2509.13442

作者:Arnab Kumar Roy, Hemant Kumar Kathania, Paban Sapkota
备注:Speaker-independent experiments on classification of dysarthric speech severity
摘要:构音障碍性言语严重程度分类对于运动性言语障碍患者的客观临床评估和进展监测至关重要。虽然先前的方法已经解决了这个任务,但在说话者无关(SID)场景中实现鲁棒的泛化仍然具有挑战性。这项工作介绍了DSSCNet,这是一种新型的深度神经架构,它结合了卷积,挤压激励(SE)和残差网络,帮助它从mel频谱图中提取构音障碍语音的判别表示。SE块的添加选择性地集中在构音障碍语音的重要特征上,从而最小化损失并增强整体模型性能。我们还提出了一个跨语料库微调框架的严重性分类,适应基于检测的迁移学习方法。DSSCNet在两个基准构音障碍语音语料库上进行评估:TORGO和UA-语音,使用说话者独立评估协议:One-Speaker-Per-Severity(OSPS)和Leave-One-Speaker-Out(LOSO)协议。DSSCNet在TORGO和UA-Speech上的OSPS和LOSO设置下的准确率分别为56.84%和62.62%和63.47%和64.18%,优于现有的最先进的方法。经过微调,性能大幅提高,DSSCNet在OSPS中的TORGO和UA语音上的准确率分别高达75.80%和68.25%,在LOSO中分别高达77.76%和79.44%。这些结果证明了DSSCNet在不同构音障碍语音数据集上进行细粒度严重程度分类的有效性和通用性。
摘要:Dysarthric speech severity classification is crucial for objective clinical assessment and progress monitoring in individuals with motor speech disorders. Although prior methods have addressed this task, achieving robust generalization in speaker-independent (SID) scenarios remains challenging. This work introduces DSSCNet, a novel deep neural architecture that combines Convolutional, Squeeze-Excitation (SE), and Residual network, helping it extract discriminative representations of dysarthric speech from mel spectrograms. The addition of SE block selectively focuses on the important features of the dysarthric speech, thereby minimizing loss and enhancing overall model performance. We also propose a cross-corpus fine-tuning framework for severity classification, adapted from detection-based transfer learning approaches. DSSCNet is evaluated on two benchmark dysarthric speech corpora: TORGO and UA-Speech under speaker-independent evaluation protocols: One-Speaker-Per-Severity (OSPS) and Leave-One-Speaker-Out (LOSO) protocols. DSSCNet achieves accuracies of 56.84% and 62.62% under OSPS and 63.47% and 64.18% under LOSO setting on TORGO and UA-Speech respectively outperforming existing state-of-the-art methods. Upon fine-tuning, the performance improves substantially, with DSSCNet achieving up to 75.80% accuracy on TORGO and 68.25% on UA-Speech in OSPS, and up to 77.76% and 79.44%, respectively, in LOSO. These results demonstrate the effectiveness and generalizability of DSSCNet for fine-grained severity classification across diverse dysarthric speech datasets.


eess.AS音频处理


【1】Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMs
标题:读听:使用文本描述和LLM的Zero-Shot发音评估
链接:https://arxiv.org/abs/2509.14187

作者:Yu-Wen Chen, Melody Ma, Julia Hirschberg
备注:EMNLP 2025 MainConference
摘要:自动发音评估通常由在音频分数对上训练的声学模型来执行。虽然有效,但这些系统只提供数字分数,没有帮助学习者理解错误所需的信息。与此同时,大型语言模型(LLM)已被证明在支持语言学习方面是有效的,但其评估发音的潜力尚未开发。在这项工作中,我们介绍了TextPA,一个zero-shot,基于文本的发音评估方法。TextPA利用人类可读的语音信号表示,将其输入LLM以评估发音准确性和流畅性,同时还提供分配分数背后的推理。最后,音素序列匹配评分方法被用来细化的准确性分数。我们的工作突出了一个以前被忽视的发音评估方向。我们利用书面文本中嵌入的丰富的发音知识,而不是依赖于有监督的训练和音频样本。实验结果表明,我们的方法是成本效益和竞争力的性能。此外,TextPA通过提供互补的视角,显着提高了传统音频分数训练模型在域外数据上的性能。
摘要:Automatic pronunciation assessment is typically performed by acoustic models trained on audio-score pairs. Although effective, these systems provide only numerical scores, without the information needed to help learners understand their errors. Meanwhile, large language models (LLMs) have proven effective in supporting language learning, but their potential for assessing pronunciation remains unexplored. In this work, we introduce TextPA, a zero-shot, Textual description-based Pronunciation Assessment approach. TextPA utilizes human-readable representations of speech signals, which are fed into an LLM to assess pronunciation accuracy and fluency, while also providing reasoning behind the assigned scores. Finally, a phoneme sequence match scoring method is used to refine the accuracy scores. Our work highlights a previously overlooked direction for pronunciation assessment. Instead of relying on supervised training with audio-score examples, we exploit the rich pronunciation knowledge embedded in written text. Experimental results show that our approach is both cost-efficient and competitive in performance. Furthermore, TextPA significantly improves the performance of conventional audio-score-trained models on out-of-domain data by offering a complementary perspective.


【2】SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification
标题:SV混合器:用轻量级MLP取代Transformer编码器,用于说话人验证中的自我监督模型压缩
链接:https://arxiv.org/abs/2509.14136

作者:Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Kyo-won Koo, Seung-bin Kim, Jisoo Son, Ha-Jin Yu
备注:8 pages, 5 figures, accepted at IEEE ASRU 2025
摘要:自监督学习(SSL)已经将说话人验证的准确性推向了最先进的水平,但大多数SSL编码器中使用的Transformer主干阻碍了设备上和实时部署。先前的压缩工作修剪了层深度或宽度,但仍然继承了自我关注的二次成本。我们提出了SV-Mixer,这是第一个完全基于MLP的SSL蒸馏学生编码器。SV-Mixer用三个轻量级模块取代了Transformer:用于多分辨率时间特征的多尺度混合,用于帧到话语上下文的局部全局混合,以及用于频谱子空间的组通道混合。从WavLM中提炼出来的SV-Mixer比Transformer学生的性能高出14.6%,同时将参数和GMAC削减了一半以上,并且在75%的压缩率下,它与教师的性能非常接近。我们的研究结果表明,无需注意的SSL学生可以通过硬件友好的足迹提供教师级的准确性,为强大的设备上扬声器验证打开大门。
摘要:Self-supervised learning (SSL) has pushed speaker verification accuracy close to state-of-the-art levels, but the Transformer backbones used in most SSL encoders hinder on-device and real-time deployment. Prior compression work trims layer depth or width yet still inherits the quadratic cost of self-attention. We propose SV-Mixer, the first fully MLP-based student encoder for SSL distillation. SV-Mixer replaces Transformer with three lightweight modules: Multi-Scale Mixing for multi-resolution temporal features, Local-Global Mixing for frame-to-utterance context, and Group Channel Mixing for spectral subspaces. Distilled from WavLM, SV-Mixer outperforms a Transformer student by 14.6% while cutting parameters and GMACs by over half, and at 75% compression, it closely matches the teacher's performance. Our results show that attention-free SSL students can deliver teacher-level accuracy with hardware-friendly footprints, opening the door to robust on-device speaker verification.


【3】A Lightweight Fourier-based Network for Binaural Speech Enhancement with Spatial Cue Preservation
标题:具有空间线索保留的基于傅里叶的轻量级网络用于双耳语音增强
链接:https://arxiv.org/abs/2509.14076

作者:Xikun Lu, Yujian Ma, Xianquan Jiang, Xuelong Wang, Jinqiu Sang
备注:Submitted to ICASSP 2026
摘要:双耳语音增强面临着一个严峻的权衡挑战,其中最先进的性能是通过计算密集型架构实现的,而轻量级解决方案往往以显著的性能下降为代价。为了弥合这一差距,我们提出了全局自适应傅立叶网络(GAF-Net),这是一种轻量级的深度复杂网络,旨在建立性能和计算效率之间的平衡。GAF-Net架构由三个部分组成。首先,结合短时傅立叶变换和伽玛特征的双特征编码器增强了声学表示的鲁棒性。第二,通道独立的全局自适应傅立叶调制器有效地捕获长期的时间依赖性,同时保留空间线索。最后,实现了动态门控机制以减少处理伪影。实验结果表明,GAF-Net实现了有竞争力的性能,特别是在双耳线索(ILD和IPD错误)和客观可懂度(MBSTOI)方面,具有更少的参数和计算成本。这些结果证实,GAF-Net提供了一种可行的方法来实现高保真双耳处理资源受限的设备。
摘要:Binaural speech enhancement faces a severe trade-off challenge, where state-of-the-art performance is achieved by computationally intensive architectures, while lightweight solutions often come at the cost of significant performance degradation. To bridge this gap, we propose the Global Adaptive Fourier Network (GAF-Net), a lightweight deep complex network that aims to establish a balance between performance and computational efficiency. The GAF-Net architecture consists of three components. First, a dual-feature encoder combining short-time Fourier transform and gammatone features enhances the robustness of acoustic representation. Second, a channel-independent globally adaptive Fourier modulator efficiently captures long-term temporal dependencies while preserving the spatial cues. Finally, a dynamic gating mechanism is implemented to reduce processing artifacts. Experimental results show that GAF-Net achieves competitive performance, particularly in terms of binaural cues (ILD and IPD error) and objective intelligibility (MBSTOI), with fewer parameters and computational cost. These results confirm that GAF-Net provides a feasible way to achieve high-fidelity binaural processing on resource-constrained devices.


【4】Lightweight Implicit Neural Network for Binaural Audio Synthesis
标题:用于双耳音频合成的轻量级隐式神经网络
链接:https://arxiv.org/abs/2509.14069

作者:Xikun Lu, Fang Liu, Weizhi Shi, Jinqiu Sang
备注:Submitted to ICASSP 2026
摘要:高保真双耳音频合成对于沉浸式聆听至关重要,但现有方法需要大量的计算资源,限制了其边缘设备应用。为了解决这个问题,我们提出了轻量级隐式神经网络(LINN),一个新的两阶段框架。LINN首先使用时域扭曲生成初始估计,然后通过隐式双耳校正器(IBC)模块进行细化。IBC是一种隐式神经网络,可直接预测振幅和相位校正,从而形成高度紧凑的模型架构。实验结果表明,LINN实现了统计上可比的感知质量的最佳性能的基线模型,同时显着提高计算效率。与现有最有效的方法相比,LINN实现了72.7%的参数减少和显着减少的计算操作(MAC)。这表明我们的方法有效地解决了合成质量和计算效率之间的权衡,为高保真边缘设备空间音频应用提供了一种新的解决方案。
摘要:High-fidelity binaural audio synthesis is crucial for immersive listening, but existing methods require extensive computational resources, limiting their edge-device application. To address this, we propose the Lightweight Implicit Neural Network (LINN), a novel two-stage framework. LINN first generates initial estimates using a time-domain warping, which is then refined by an Implicit Binaural Corrector (IBC) module. IBC is an implicit neural network that predicts amplitude and phase corrections directly, resulting in a highly compact model architecture. Experimental results show that LINN achieves statistically comparable perceptual quality to the best-performing baseline model while significantly improving computational efficiency. Compared to the most efficient existing method, LINN achieves a 72.7% reduction in parameters and significantly fewer compute operations (MACs). This demonstrates that our approach effectively addresses the trade-off between synthesis quality and computational efficiency, providing a new solution for high-fidelity edge-device spatial audio applications.


【5】Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
标题:你听到我的意思了吗?量化教学引导表达性文本到语音系统中的教学感知差距
链接:https://arxiv.org/abs/2509.13989

作者:Yi-Cheng Lin, Huang-Cheng Chou, Tzu-Chieh Wei, Kuan-Yu Chen, Hung-yi Lee
备注:Submission to ICASSP 2026
摘要:指令引导的文本到语音(ITTS)使用户能够通过自然语言提示控制语音生成,提供比传统TTS更直观的界面。然而,用户风格指令和听众感知之间的一致性在很大程度上仍未被探索。这项工作首先提出了一个感性的分析ITTS可控性在两个表达层面(程度副词和分级的情感强度),并收集人的评级扬声器年龄和单词级的强调属性。为了全面揭示发音-感知差距,我们提供了一个大规模人类评估的数据集,名为表达性语音控制(E-VOC)语料库。此外,我们发现(1)gpt-4 o-mini-tts是最可靠的ITTS模型,在声学维度上,指令和生成的话语之间具有很好的一致性。(2)5个分析的ITTS系统倾向于生成成人的声音,即使指令要求使用儿童或老年人的声音。(3)细粒度控制仍然是一个主要的挑战,这表明大多数ITTS系统在解释略有不同的属性指令方面有很大的改进空间。
摘要:Instruction-guided text-to-speech (ITTS) enables users to control speech generation through natural language prompts, offering a more intuitive interface than traditional TTS. However, the alignment between user style instructions and listener perception remains largely unexplored. This work first presents a perceptual analysis of ITTS controllability across two expressive dimensions (adverbs of degree and graded emotion intensity) and collects human ratings on speaker age and word-level emphasis attributes. To comprehensively reveal the instruction-perception gap, we provide a data collection with large-scale human evaluations, named Expressive VOice Control (E-VOC) corpus. Furthermore, we reveal that (1) gpt-4o-mini-tts is the most reliable ITTS model with great alignment between instruction and generated utterances across acoustic dimensions. (2) The 5 analyzed ITTS systems tend to generate Adult voices even when the instructions ask to use child or Elderly voices. (3) Fine-grained control remains a major challenge, indicating that most ITTS systems have substantial room for improvement in interpreting slightly different attribute instructions.


【6】DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
标题:DSpAST:使用大型语言模型的空间音频推理的分离表示
链接:https://arxiv.org/abs/2509.13927

作者:Kevin Wilkinghoff, Zheng-Hua Tan
摘要:对于具有大型语言模型的空间音频的推理需要空间音频编码器作为声学前端以获得音频嵌入用于进一步处理。这样的编码器需要捕获检测声音事件类型所需的所有信息,以及它们对应的源的方向和距离。用单个音频编码器完成这一任务是非常苛刻的,因为这些任务中的每一个所需的信息大多是相互独立的。因此,使用单个编码器获得的性能通常比使用特定于任务的音频编码器时更差。在这项工作中,我们提出了DSpAST,一种新的音频编码器的基础上SpatialAST,学习空间音频的解纠缠表示,而只有0.2%的额外参数。利用空间音频推理系统BAT对SpatialSoundQA进行实验,结果表明DSpAST的性能明显优于SpatialAST.
摘要:Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the type of sound events, as well as the direction and distance of their corresponding sources. Accomplishing this with a single audio encoder is demanding as the information required for each of these tasks is mostly independent of each other. As a result, the performance obtained with a single encoder is often worse than when using task-specific audio encoders. In this work, we present DSpAST, a novel audio encoder based on SpatialAST that learns disentangled representations of spatial audio while having only 0.2% additional parameters. Experiments on SpatialSoundQA with the spatial audio reasoning system BAT demonstrate that DSpAST significantly outperforms SpatialAST.


【7】Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
标题:可推广音频深度伪造检测中的低级别适配器专家混合
链接:https://arxiv.org/abs/2509.13878

作者:Janne Laakkonen, Ivan Kukanov, Ville Hautamäki
备注:6 pages, 3 figures, 1 table
摘要:像Wav 2 Vec 2这样的基础模型擅长语音任务中的表示学习,包括音频deepfake检测。然而,在对一组固定的真实和欺骗的音频片段进行微调后,它们往往无法推广到训练中没有出现的新的deepfake方法。为了解决这个问题,我们提出了一种混合LoRA专家的方法,将多个低秩适配器(LoRA)集成到模型的注意力层中。路由机制有选择地激活专业专家,增强对不断发展的deepfake攻击的适应性。实验结果表明,我们的方法优于标准的微调,在域内和域外的情况下,减少相等的错误率相对于基线模型。值得注意的是,我们最好的MoE-LoRA模型将平均域外EER从8.55\%降低到6.08\%,证明了其在实现可推广的音频深度伪造检测方面的有效性。
摘要:Foundation models such as Wav2Vec2 excel at representation learning in speech tasks, including audio deepfake detection. However, after being fine-tuned on a fixed set of bonafide and spoofed audio clips, they often fail to generalize to novel deepfake methods not represented in training. To address this, we propose a mixture-of-LoRA-experts approach that integrates multiple low-rank adapters (LoRA) into the model's attention layers. A routing mechanism selectively activates specialized experts, enhancing adaptability to evolving deepfake attacks. Experimental results show that our method outperforms standard fine-tuning in both in-domain and out-of-domain scenarios, reducing equal error rates relative to baseline models. Notably, our best MoE-LoRA model lowers the average out-of-domain EER from 8.55\% to 6.08\%, demonstrating its effectiveness in achieving generalizable audio deepfake detection.


【8】Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
标题:多语言对话语音语言模型挑战总结:数据集、任务、基线和方法
链接:https://arxiv.org/abs/2509.13785

作者:Bingshen Mu, Pengcheng Guo, Zhaokai Sun, Shuai Wang, Hexin Liu, Mingchen Shao, Lei Xie, Eng Siong Chng, Longshuai Xiao, Qiangze Feng, Daliang Wang
摘要:本文总结了Interspeech 2025多语言会话语音语言模型(MLC-SLM)挑战,旨在推进构建有效的多语言会话语音LLM(SLLM)的探索。我们详细描述了MLC-SLM挑战赛的任务设置、已发布的总计约1,604小时的真实世界多语言会话语音数据集以及参与者的基线系统。MLC-SLM挑战赛吸引了来自13个国家的78支队伍参加,两项任务共获得489项有效排行榜结果和14份技术报告。我们根据参与者提交的意见,提炼出构建多语言会话SLLM的宝贵见解,旨在为社区的发展做出贡献。
摘要:This paper summarizes the Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) challenge, which aims to advance the exploration of building effective multilingual conversational speech LLMs (SLLMs). We provide a detailed description of the task settings for the MLC-SLM challenge, the released real-world multilingual conversational speech dataset totaling approximately 1,604 hours, and the baseline systems for participants. The MLC-SLM challenge attracts 78 teams from 13 countries to participate, with 489 valid leaderboard results and 14 technical reports for the two tasks. We distill valuable insights on building multilingual conversational SLLMs based on submissions from participants, aiming to contribute to the advancement of the community.


【9】Self-Guided Target Sound Extraction and Classification Through Universal Sound Separation Model and Multiple Clues
标题:通过通用声音分离模型和多线索自引导目标声音提取和分类
链接:https://arxiv.org/abs/2509.13741

作者:Younghoo Kwon, Dongheon Lee, Dohwan Kim, Jung-Woo Choi
备注:5 pages, 2 figures, submitted to DCASE workshop 2025
摘要:本文介绍了一个多阶段的自我导向框架,旨在解决DCASE 2025任务4挑战中的声音场景(S5)任务的空间语义分割。该框架集成了专注于三个不同任务的模型:通用声音分离(USS),单标签分类(SC)和目标声音提取(TSE)。最初,USS将复杂的音频混合分解为单独的源波形。然后,这些分离的波形中的每一个都由SC块处理,生成两条关键信息:波形本身及其对应的类标签。这些作为TSE阶段的输入,TSE阶段隔离与此信息匹配的源。由于这些输入是在系统内产生的,因此自动识别提取目标,消除了外部指导的必要性。提取的波形可以循环回分类任务,创建一个迭代细化的循环,逐步提高可分性和标记准确性。因此,我们称我们的框架是一个多阶段的自我指导系统,由于这些自包含的特点。在官方评估数据集上,该系统实现了11.00 dB的类感知信号失真比改善(CA-SDRi)和55.8%的标签预测准确率,分别比ResUNetK基线高出4.4 dB和4.3%,并在所有提交中获得第一名。
摘要:This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This framework integrates models focused on three distinct tasks: Universal Sound Separation (USS), Single-label Classification (SC), and Target Sound Extraction (TSE). Initially, USS breaks down a complex audio mixture into separate source waveforms. Each of these separated waveforms is then processed by a SC block, generating two critical pieces of information: the waveform itself and its corresponding class label. These serve as inputs for the TSE stage, which isolates the source that matches this information. Since these inputs are produced within the system, the extraction target is identified autonomously, removing the necessity for external guidance. The extracted waveform can be looped back into the classification task, creating a cycle of iterative refinement that progressively enhances both separability and labeling accuracy. We thus call our framework a multi-stage self-guided system due to these self-contained characteristics. On the official evaluation dataset, the proposed system achieves an 11.00 dB increase in class-aware signal-to-distortion ratio improvement (CA-SDRi) and a 55.8\% accuracy in label prediction, outperforming the ResUNetK baseline by 4.4 dB and 4.3\%, respectively, and achieving first place among all submissions.


【10】A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
标题:具有知识蒸馏的高质量、低复杂性的可流化神经语音编解码器
链接:https://arxiv.org/abs/2509.13670

作者:En-Wei Zhang, Hui-Peng Du, Xiao-Hang Jiang, Yang Ai, Zhen-Hua Ling
备注:Accepted by APSIPA ASC 2025
摘要:虽然许多当前的神经语音编解码器实现了令人印象深刻的重建语音质量,但它们往往忽略了延迟和复杂性的考虑,限制了它们在实时语音通信和高效语音压缩等下游任务中的实际部署。在我们之前的工作中,我们提出了StreamCodec,它通过利用模型因果关系和标量-矢量组合量化策略来实现可流式语音编码,但其重建质量和复杂度仍有改进的空间。因此,本文提出了一种改进的迭代StreamCodec,命名为StreamCodec 2。StreamCodec 2通过采用完全因果架构并减少卷积通道,支持流式和轻量级语音编码。为了补偿模型因果化和修剪造成的语音质量下降,我们引入了一个非因果的,高复杂度的教师编解码器,通过知识蒸馏来指导StreamCodec 2的训练。实验结果表明,我们提出的StreamCodec 2,与知识蒸馏策略训练,可以实现高质量的语音重建,同时保持低延迟(仅20 ms),低计算复杂度(仅910 MFLOPs),低模型复杂度(仅5.4 M参数)。
摘要:While many current neural speech codecs achieve impressive reconstructed speech quality, they often neglect latency and complexity considerations, limiting their practical deployment in downstream tasks such as real-time speech communication and efficient speech compression. In our previous work, we proposed StreamCodec, which enables streamable speech coding by leveraging model causalization and a scalar-vector-combined quantization strategy, but its reconstructed quality and complexity still have room for improvement. Therefore, this paper proposes an improved iteration of StreamCodec, named StreamCodec2. The StreamCodec2 supports streamable and lightweight speech coding by adopting a fully causal architecture and reducing the convolutional channels. To compensate for the speech quality degradation caused by model causalization and pruning, we introduce a non-causal, high-complexity teacher codec to guide the training of StreamCodec2 through knowledge distillation. Experimental results demonstrate that our proposed StreamCodec2, trained with the knowledge distillation strategy, can achieve high-quality speech reconstruction while maintaining low latency (only 20 ms), low computational complexity (only 910 MFLOPs), and low model complexity (only 5.4 M parameters).


【11】A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
标题:具有显式幅度和相预测的蒸馏低延迟神经声码器
链接:https://arxiv.org/abs/2509.13667

作者:Hui-Peng Du, Yang Ai, Zhen-Hua Ling
备注:Accepted by APSIPA ASC 2025
摘要:大多数主流神经声码器主要关注语音质量和生成速度,而忽略了延迟,这是实时应用中的关键因素。过度的延迟会导致用户交互的明显延迟,严重降低用户体验,并使此类系统无法实时使用。因此,本文提出了DLL-APNet,一种蒸馏低延迟神经声码器,它首先从输入梅尔频谱图中显式预测幅度和相位频谱,然后通过逆短时傅立叶变换(iSTFT)重建语音波形。DLL-APNet声码器利用因果卷积将信息的利用限制在当前和历史上下文中,从而有效地最大限度地减少延迟。为了减轻因果约束引起的语音质量下降,提出了一种知识蒸馏策略,其中预先训练的非因果教师声码器指导因果学生DLL-APNet声码器的中间特征生成。实验结果表明,建议DLL-APNet声码器产生更高质量的语音比其他因果声码器,而需要更少的计算资源。此外,所提出的DLL-APNet声码器实现了与主流非因果神经声码器相当的语音质量,验证了其提供高感知质量和低延迟的能力。
摘要:The majority of mainstream neural vocoders primarily focus on speech quality and generation speed, while overlooking latency, which is a critical factor in real-time applications. Excessive latency leads to noticeable delays in user interaction, severely degrading the user experience and rendering such systems impractical for real-time use. Therefore, this paper proposes DLL-APNet, a Distilled Low-Latency neural vocoder which first predicts the Amplitude and Phase spectra explicitly from input mel spectrogram and then reconstructs the speech waveform via inverse short-time Fourier transform (iSTFT). The DLL-APNet vocoder leverages causal convolutions to constrain the utilization of information to current and historical contexts, effectively minimizing latency. To mitigate speech quality degradation caused by causal constraints, a knowledge distillation strategy is proposed, where a pre-trained non-causal teacher vocoder guides intermediate feature generation of the causal student DLL-APNet vocoder. Experimental results demonstrate that the proposed DLL-APNet vocoder produces higher-quality speech than other causal vocoders, while requiring fewer computational resources. Furthermore, the proposed DLL-APNet vocoder achieves speech quality on par with mainstream non-causal neural vocoders, validating its ability to deliver both high perceptual quality and low latency.


【12】Assessing Data Replication in Symbolic Music via Adapted Structural Similarity Index Measure
标题:通过调整的结构相似性指数衡量评估象征性音乐中的数据复制
链接:https://arxiv.org/abs/2509.13658

作者:Shulei Ji, Zihao Wang, Le Ma, Jiaxing Yu, Kejun Zhang
摘要:人工智能生成的音乐可能会无意中复制训练数据中的样本,从而引发对剽窃的担忧。相似性度量可以量化这种复制,从而为音乐生成模型提供监督和指导。现有的符号音乐相似性度量方法主要针对旋律重复,在评价具有丰富织体和表现力特征的复杂音乐时存在空白。为了解决这一差距,我们介绍SSIMuse,结构相似性指数测量(SSIM)的第一个适应从图像到象征性的音乐。具体来说,我们代表象征性的音乐,像钢琴卷在二进制和速度为基础的形式。在这些表征的基础上,我们在音乐背景下重新定义并适当修改SSIM组件,以开发两种变体,即,SSIMuse-B和SSIMuse-V,分别用于评估数据复制的组合和动态性能。对来自多个数据集的合成样本进行的受控实验表明,SSIMuse可以可靠地检测至少一个bar粒度的精确复制。SSIMuse允许对音乐生成中的复制进行公开评估,并引起人们对其更广泛的伦理,社会,法律和经济影响的关注。该代码可在https://github.com/Tayjsl97/SSIMuse上获得。
摘要:AI-generated music may inadvertently replicate samples from the training data, raising concerns of plagiarism. Similarity measures can quantify such replication, thereby offering supervision and guidance for music generation models. Existing similarity measure methods for symbolic music mainly target melody repetition, leaving a gap in assessing complex music with rich textures and expressive performance characteristics. To address this gap, we introduce SSIMuse, the first adaptation of the Structural Similarity Index Measure (SSIM) from images to symbolic music. Specifically, we represent symbolic music as image-like piano rolls in binary and velocity-based forms. Build upon these representations, we reinterprete and suitably modify the SSIM components in the musical context to develop two variants, i.e., SSIMuse-B and SSIMuse-V, for evaluating data replication in composition and dynamic performance, respectively. Controlled experiments on synthetic samples from multiple datasets show that SSIMuse can reliably detect exact replication at a granularity of at least one bar. SSIMuse enables open evaluation of replication in music generation and draws attention to its broader ethical, social, legal, and economic implications. The code is available at https://github.com/Tayjsl97/SSIMuse.


【13】Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
标题:利用DS SCNet和跨数据库自适应增强与说话者无关的发音障碍语音严重程度分类
链接:https://arxiv.org/abs/2509.13442

作者:Arnab Kumar Roy, Hemant Kumar Kathania, Paban Sapkota
备注:Speaker-independent experiments on classification of dysarthric speech severity
摘要:构音障碍性言语严重程度分类对于运动性言语障碍患者的客观临床评估和进展监测至关重要。虽然先前的方法已经解决了这个任务,但在说话者无关(SID)场景中实现鲁棒的泛化仍然具有挑战性。这项工作介绍了DSSCNet,这是一种新型的深度神经架构,它结合了卷积,挤压激励(SE)和残差网络,帮助它从mel频谱图中提取构音障碍语音的判别表示。SE块的添加选择性地集中在构音障碍语音的重要特征上,从而最小化损失并增强整体模型性能。我们还提出了一个跨语料库微调框架的严重性分类,适应基于检测的迁移学习方法。DSSCNet在两个基准构音障碍语音语料库上进行评估:TORGO和UA-语音,使用说话者独立评估协议:One-Speaker-Per-Severity(OSPS)和Leave-One-Speaker-Out(LOSO)协议。DSSCNet在TORGO和UA-Speech上的OSPS和LOSO设置下的准确率分别为56.84%和62.62%和63.47%和64.18%,优于现有的最先进的方法。经过微调,性能大幅提高,DSSCNet在OSPS中的TORGO和UA语音上的准确率分别高达75.80%和68.25%,在LOSO中分别高达77.76%和79.44%。这些结果证明了DSSCNet在不同构音障碍语音数据集上进行细粒度严重程度分类的有效性和通用性。
摘要:Dysarthric speech severity classification is crucial for objective clinical assessment and progress monitoring in individuals with motor speech disorders. Although prior methods have addressed this task, achieving robust generalization in speaker-independent (SID) scenarios remains challenging. This work introduces DSSCNet, a novel deep neural architecture that combines Convolutional, Squeeze-Excitation (SE), and Residual network, helping it extract discriminative representations of dysarthric speech from mel spectrograms. The addition of SE block selectively focuses on the important features of the dysarthric speech, thereby minimizing loss and enhancing overall model performance. We also propose a cross-corpus fine-tuning framework for severity classification, adapted from detection-based transfer learning approaches. DSSCNet is evaluated on two benchmark dysarthric speech corpora: TORGO and UA-Speech under speaker-independent evaluation protocols: One-Speaker-Per-Severity (OSPS) and Leave-One-Speaker-Out (LOSO) protocols. DSSCNet achieves accuracies of 56.84% and 62.62% under OSPS and 63.47% and 64.18% under LOSO setting on TORGO and UA-Speech respectively outperforming existing state-of-the-art methods. Upon fine-tuning, the performance improves substantially, with DSSCNet achieving up to 75.80% accuracy on TORGO and 68.25% on UA-Speech in OSPS, and up to 77.76% and 79.44%, respectively, in LOSO. These results demonstrate the effectiveness and generalizability of DSSCNet for fine-grained severity classification across diverse dysarthric speech datasets.


【14】TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
标题:TICI:用于语音上下文学习的文本嵌入KNN释放大型多模式模型的语音识别能力
链接:https://arxiv.org/abs/2509.13395

作者:Haolong Zheng, Yekaterina Yegorova, Mark Hasegawa-Johnson
摘要:语音基础模型最近已经证明了执行语音上下文学习(SICL)的能力。选择有效的背景下的例子是至关重要的SICL的性能,但选择方法仍然未充分探索。在这项工作中,我们提出了SICL(TICL)的文本嵌入KNN,这是一个简单的管道,它使用语义上下文来增强现成的大型多模态模型的语音识别能力,而无需微调。在具有挑战性的自动语音识别任务,包括口音英语,多语种语音和儿童的语音,我们的方法使模型能够超过zero-shot性能高达84.7%的相对WER减少。我们进行消融研究,以显示我们的方法的鲁棒性和效率。
摘要:Speech foundation models have recently demonstrated the ability to perform Speech In-Context Learning (SICL). Selecting effective in-context examples is crucial for SICL performance, yet selection methodologies remain underexplored. In this work, we propose Text-Embedding KNN for SICL (TICL), a simple pipeline that uses semantic context to enhance off-the-shelf large multimodal models' speech recognition ability without fine-tuning. Across challenging automatic speech recognition tasks, including accented English, multilingual speech, and children's speech, our method enables models to surpass zero-shot performance with up to 84.7% relative WER reduction. We conduct ablation studies to show the robustness and efficiency of our method.


【15】CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
标题:CS-FLEURS:大规模多语言和代码交换语音数据集
链接:https://arxiv.org/abs/2509.14161

作者:Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Lodagala, William Chen, Olga Iakovenko, Bashar Talafha, Amir Hussein, Alexander Polok, Kalvin Chang, Dominik Klement, Sara Althubaiti, Puyuan Peng, Matthew Wiesner, Thamar Solorio, Ahmed Ali, Sanjeev Khudanpur, Shinji Watanabe, Chih-Chen Chen, Zhen Wu, Karim Benharrak, Anuj Diwan, Samuele Cornell, Eunjung Yeo, Kwanghee Choi, Carlos Carvalho, Karen Rosero
摘要:我们提出了CS-FLEURS,这是一个新的数据集,用于开发和评估高资源语言之外的代码切换语音识别和翻译系统。CS-FLEURS由4个测试集组成,涵盖52种语言的113个独特的代码转换语言对:1)14 X-英语语言对集合,其具有阅读合成生成的代码转换句子的真实语音,2)16 X-英语语言对集合,其具有生成的文本到语音,3)60 {阿拉伯语,普通话,印地语,西班牙语}-X语言对测试集,具有生成的文本到语音,以及4)45个X-英语低资源语言对测试集,具有连接的文本到语音。除了四个测试集之外,CS-FLEURS还提供了一个训练集,其中包含16个X-English语言对的128小时生成文本到语音数据。我们希望CS-FLEURS有助于拓宽未来语码转换语音研究的范围。数据集链接:https://huggingface.co/datasets/byan/cs-fleurs。
摘要:We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 unique code-switched language pairs across 52 languages: 1) a 14 X-English language pair set with real voices reading synthetically generated code-switched sentences, 2) a 16 X-English language pair set with generative text-to-speech 3) a 60 {Arabic, Mandarin, Hindi, Spanish}-X language pair set with the generative text-to-speech, and 4) a 45 X-English lower-resourced language pair test set with concatenative text-to-speech. Besides the four test sets, CS-FLEURS also provides a training set with 128 hours of generative text-to-speech data across 16 X-English language pairs. Our hope is that CS-FLEURS helps to broaden the scope of future code-switched speech research. Dataset link: https://huggingface.co/datasets/byan/cs-fleurs.


【16】Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
标题:Canary-1B-v2和Parakeet-TDT-0. 6 B-v3:用于多语言ASB和AST的高效高性能模型
链接:https://arxiv.org/abs/2509.14128

作者:Monica Sekoyan, Nithin Rao Koluguri, Nune Tadevosyan, Piotr Zelasko, Travis Bartley, Nick Karpov, Jagadeesh Balam, Boris Ginsburg
备注:Mini Version of it Submitted to ICASSP 2026
摘要:本报告介绍了Canary-1B-v2,一种用于自动语音识别(ASR)和语音到文本翻译(AST)的快速,强大的多语言模型。内置FastConformer编码器和Transformer解码器,支持25种语言,主要是欧洲语言。该模型在170万小时的总数据样本上进行了训练,包括Granary和NeMo ASR Set 3.0,并添加了非语音音频以减少ASR和AST的幻觉。我们描述了它的两个阶段的预训练和微调过程与动态数据平衡,以及nGPT编码器的实验。结果显示,nGPT在海量数据中表现良好,而FastConformer在微调后表现出色。对于时间戳,Canary-1B-v2使用NeMo Forced Aligner(NFA)和辅助CTC模型,为ASR和AST提供可靠的段级时间戳。评估显示,Canary-1B-v2在英语ASR上的性能优于Whisper-large-v3,同时速度快10倍,并提供了具有竞争力的多语言ASR和AST性能,以对抗更大的模型,如无障碍M4 T-v2-large和基于LLM的系统。我们还发布了Parakeet-TDT-0. 6 B-v3,它是v2的后续版本,仅用600 M参数就可以在相同的25种语言中提供多语言ASR。
摘要:This report introduces Canary-1B-v2, a fast, robust multilingual model for Automatic Speech Recognition (ASR) and Speech-to-Text Translation (AST). Built with a FastConformer encoder and Transformer decoder, it supports 25 languages primarily European. The model was trained on 1.7M hours of total data samples, including Granary and NeMo ASR Set 3.0, with non-speech audio added to reduce hallucinations for ASR and AST. We describe its two-stage pre-training and fine-tuning process with dynamic data balancing, as well as experiments with an nGPT encoder. Results show nGPT scales well with massive data, while FastConformer excels after fine-tuning. For timestamps, Canary-1B-v2 uses the NeMo Forced Aligner (NFA) with an auxiliary CTC model, providing reliable segment-level timestamps for ASR and AST. Evaluations show Canary-1B-v2 outperforms Whisper-large-v3 on English ASR while being 10x faster, and delivers competitive multilingual ASR and AST performance against larger models like Seamless-M4T-v2-large and LLM-based systems. We also release Parakeet-TDT-0.6B-v3, a successor to v2, offering multilingual ASR across the same 25 languages with just 600M parameters.


【17】A Domain Knowledge Informed Approach for Anomaly Detection of Electric Vehicle Interior Sounds
标题:电动汽车车内声音异常检测领域知识知情方法
链接:https://arxiv.org/abs/2509.13390

作者:Deepti Kunte, Bram Cornelis, Claudio Colangeli, Karl Janssens, Brecht Van Baelen, Konstantinos Gryllias
备注:Submitted to: Mechanical Systems and Signal Processing
摘要:汽车驾驶室声音异常的检测对于确保车辆质量和保持乘客舒适性至关重要。在许多现实环境中,由于标记错误数据的稀缺性或完全不存在,该任务更适合作为无监督学习问题而不是监督情况。在这种无监督设置中,模型只在健康样本上进行训练,并将异常检测为与正常行为的偏差。然而,在缺乏用于验证的标记错误样本以及常用度量的有限可靠性(例如验证重建误差)的情况下,有效的模型选择仍然是一个重大挑战。为了克服这些限制,提出了一种基于领域知识的模型选择方法,其中通过健康谱图的结构化扰动设计的代理异常用于验证集以支持模型选择。在高保真电动汽车数据集上评估了所提出的方法,该数据集包括五种代表性故障类型的健康和故障车厢声音,即,不平衡、调制、呜呜声、风和脉宽调制。该数据集使用先进的声音合成技术生成,并通过专家评审团评估进行验证,已公开提供,以促进进一步的研究。五个故障情况下的实验评估表明,使用代理异常的最佳模型的选择,显着优于传统的模型选择策略。
摘要:The detection of anomalies in automotive cabin sounds is critical for ensuring vehicle quality and maintaining passenger comfort. In many real-world settings, this task is more appropriately framed as an unsupervised learning problem rather than the supervised case due to the scarcity or complete absence of labeled faulty data. In such an unsupervised setting, the model is trained exclusively on healthy samples and detects anomalies as deviations from normal behavior. However, in the absence of labeled faulty samples for validation and the limited reliability of commonly used metrics, such as validation reconstruction error, effective model selection remains a significant challenge. To overcome these limitations, a domain-knowledge-informed approach for model selection is proposed, in which proxy-anomalies engineered through structured perturbations of healthy spectrograms are used in the validation set to support model selection. The proposed methodology is evaluated on a high-fidelity electric vehicle dataset comprising healthy and faulty cabin sounds across five representative fault types viz., Imbalance, Modulation, Whine, Wind, and Pulse Width Modulation. This dataset, generated using advanced sound synthesis techniques, and validated via expert jury assessments, has been made publicly available to facilitate further research. Experimental evaluations on the five fault cases demonstrate the selection of optimal models using proxy-anomalies, significantly outperform conventional model selection strategies.


机器翻译由腾讯交互翻译提供,仅供参