微信公众号:arXiv_Daily
cs.SD语音
【1】A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
标题:语音信号心理稳定性分类的新迁移学习方法
链接:https://arxiv.org/abs/2601.16793
摘要:这项研究提出了一种新的迁移学习方法和数据增强技术,用于使用人类语音信号进行精神稳定性分类,并解决了与有限数据可用性相关的挑战。卷积神经网络(CNN)已经被用于分析从语音记录生成的频谱图图像。三个CNN架构,VGG 16,InceptionV3和DenseNet121,在三个实验阶段进行了评估:非增强数据,增强数据和迁移学习。这种提出的迁移学习方法涉及在增强数据集上预训练模型,并在非增强数据集上对其进行微调,同时确保严格的数据分离以防止数据泄漏。结果表明,与基线方法相比,分类性能有显着改善。在三种CNN架构中,使用所提出的迁移学习方法,DenseNet121实现了94%的最高准确率和99%的AUC得分。这一发现强调了数据增强和迁移学习相结合的有效性,可以使用语音频谱图增强基于CNN的心理稳定性分类,为心理健康诊断提供了一种有前途的非侵入性工具。
摘要:This study presents a novel transfer learning approach and data augmentation technique for mental stability classification using human voice signals and addresses the challenges associated with limited data availability. Convolutional neural networks (CNNs) have been employed to analyse spectrogram images generated from voice recordings. Three CNN architectures, VGG16, InceptionV3, and DenseNet121, were evaluated across three experimental phases: training on non-augmented data, augmented data, and transfer learning. This proposed transfer learning approach involves pre-training models on the augmented dataset and fine-tuning them on the non-augmented dataset while ensuring strict data separation to prevent data leakage. The results demonstrate significant improvements in classification performance compared to the baseline approach. Among three CNN architectures, DenseNet121 achieved the highest accuracy of 94% and an AUC score of 99% using the proposed transfer learning approach. This finding highlights the effectiveness of combining data augmentation and transfer learning to enhance CNN-based classification of mental stability using voice spectrograms, offering a promising non-invasive tool for mental health diagnostics.
【2】E2E-AEC: Implementing an end-to-end neural network learning approach for acoustic echo cancellation
标题:E2 E-AEC:实施端到端神经网络学习方法用于声学回声消除
链接:https://arxiv.org/abs/2601.16774
备注:This paper has been accepted by ICASSP2026
摘要:我们提出了一种新型的基于神经网络的端到端声学回声消除(E2 E-AEC)方法,能够进行流推理,该方法可以有效地运行,而不依赖于传统的线性AEC(LAEC)技术和时间延迟估计。我们的方法包括几个关键策略:首先,我们引入并改进渐进式学习,以逐步增强回声抑制。其次,我们的模型通过使用预先训练的基于LAEC的模型进行初始化,利用从LAEC培训中获得的见解来进行知识转移。第三,我们优化了注意力机制,在注意力权重上应用了损失函数,以实现参考信号和麦克风信号之间的精确时间对准。最后,我们将语音活动检测,以提高语音质量和改善回声消除掩蔽网络输出时,近端语音缺席。通过在公共数据集上进行的实验验证了该方法的有效性。
摘要:We propose a novel neural network-based end-to-end acoustic echo cancellation (E2E-AEC) method capable of streaming inference, which operates effectively without reliance on traditional linear AEC (LAEC) techniques and time delay estimation. Our approach includes several key strategies: First, we introduce and refine progressive learning to gradually enhance echo suppression. Second, our model employs knowledge transfer by initializing with a pre-trained LAECbased model, harnessing the insights gained from LAEC training. Third, we optimize the attention mechanism with a loss function applied on attention weights to achieve precise time alignment between the reference and microphone signals. Lastly, we incorporate voice activity detection to enhance speech quality and improve echo removal by masking the network output when near-end speech is absent. The effectiveness of our approach is validated through experiments conducted on public datasets.
【3】I Guess That's Why They Call it the Blues: Causal Analysis for Audio Classifiers
标题:我想这就是为什么他们称之为蓝调:音频分类器的因果分析
链接:https://arxiv.org/abs/2601.16675
摘要:众所周知,音频分类器通常依赖于非音乐相关特征和虚假相关性来对音频进行分类。因此,音频分类器很容易被操纵或混淆,导致错误的分类。虽然导致错误分类并不难,但直到现在,分类器所依赖的特征集还没有得到很好的理解。 在本文中,我们介绍了一种新的方法,使用因果推理发现的频率空间的功能是足够的,必要的一个给定的分类。我们描述了该算法在工具FreqReX中的实现,并提供了一些标准基准数据集上的实验结果。我们的实验表明,因果充分和必要的子集允许我们通过非常轻微地改变输入以各种方式操纵模型的输出。也就是说,改变到240,000个频率中的一个频率会导致58%的时间的分类改变,并且该改变可以小到几乎听不见。这些结果表明,因果分析是有用的理解音频分类器的推理过程,并可以用来成功地操纵他们的输出。
摘要:It is well-known that audio classifiers often rely on non-musically relevant features and spurious correlations to classify audio. Hence audio classifiers are easy to manipulate or confuse, resulting in wrong classifications. While inducing a misclassification is not hard, until now the set of features that the classifiers rely on was not well understood. In this paper we introduce a new method that uses causal reasoning to discover features of the frequency space that are sufficient and necessary for a given classification. We describe an implementation of this algorithm in the tool FreqReX and provide experimental results on a number of standard benchmark datasets. Our experiments show that causally sufficient and necessary subsets allow us to manipulate the outputs of the models in a variety of ways by changing the input very slightly. Namely, a change to one out of 240,000 frequencies results in a change in classification 58% of the time, and the change can be so small that it is practically inaudible. These results show that causal analysis is useful for understanding the reasoning process of audio classifiers and can be used to successfully manipulate their outputs.
【4】Omni-directional attention mechanism based on Mamba for speech separation
标题:基于Mamba的全方位注意力机制在语音分离中的应用
链接:https://arxiv.org/abs/2601.16603
摘要:Mamba是一种选择性状态空间模型(SSM),它是Transformers语音建模的有效替代方案,可以实现线性复杂度的长序列处理。虽然在语音分离中是有效的,但现有的方法,无论是在时域还是时频域中,通常在使用Mamba处理它们之前,将输入沿着单个维度分解为短的一维序列,这将其限制为局部1D建模,并限制了其在2D频谱图中捕获全局依赖关系的能力。在这项工作中,我们提出了一个有效的全方位的注意力(OA)机制,建立在单向曼巴,从十个不同的方向上的频谱图模型的全球依赖关系。我们将所提出的机制扩展为两个基线分离模型,并在三个公共数据集上进行评估。实验结果表明,我们的方法在保持线性复杂度的同时,始终比基线实现了显着的性能提升,优于现有的最先进(SOTA)系统。
摘要:Mamba, a selective state-space model (SSM), has emerged as an efficient alternative to Transformers for speech modeling, enabling long-sequence processing with linear complexity. While effective in speech separation, existing approaches, whether in the time or time-frequency domain, typically decompose the input along a single dimension into short one-dimensional sequences before processing them with Mamba, which restricts it to local 1D modeling and limits its ability to capture global dependencies across the 2D spectrogram. In this work, we propose an efficient omni-directional attention (OA) mechanism built upon unidirectional Mamba, which models global dependencies from ten different directions on the spectrogram. We expand the proposed mechanism into two baseline separation models and evaluate on three public datasets. Experimental results show that our approach consistently achieves significant performance gains over the baselines while preserving linear complexity, outperforming existing state-of-the-art (SOTA) systems.
【5】CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
标题:CORD:通过加权政策跨模式蒸馏弥合音频文本推理差距
链接:https://arxiv.org/abs/2601.16547
备注:13 pages, 4 figures
摘要:大型音频语言模型(LALM)已经获得了重大的研究兴趣。尽管建立在基于文本的大型语言模型(LLM)上,但LALM经常表现出知识和推理能力的下降。我们假设,这种限制源于当前的训练范式的失败,有效地弥合特征表示空间内的声学语义差距。为了应对这一挑战,我们提出了CORD,一个统一的对齐框架,执行在线跨模态自蒸馏。具体来说,它在一个统一的模型中将音频条件推理与文本条件推理相结合。利用文本模态作为内部教师,CORD在整个音频推出过程中执行多粒度对齐。在令牌级别,它采用具有重要性感知权重的策略反向KL发散来优先考虑早期和语义关键的令牌。在序列级,CORD引入了基于判断的全局奖励,通过组相对策略优化(GRPO)来优化完整的推理轨迹。多个基准测试的实证结果表明,CORD始终增强了音频条件推理,并仅用80 k合成训练样本就大大弥合了音频-文本性能差距,验证了我们的策略,多级跨模态对齐方法的有效性和数据效率。
摘要:Large Audio Language Models (LALMs) have garnered significant research interest. Despite being built upon text-based large language models (LLMs), LALMs frequently exhibit a degradation in knowledge and reasoning capabilities. We hypothesize that this limitation stems from the failure of current training paradigms to effectively bridge the acoustic-semantic gap within the feature representation space. To address this challenge, we propose CORD, a unified alignment framework that performs online cross-modal self-distillation. Specifically, it aligns audio-conditioned reasoning with its text-conditioned counterpart within a unified model. Leveraging the text modality as an internal teacher, CORD performs multi-granularity alignment throughout the audio rollout process. At the token level, it employs on-policy reverse KL divergence with importance-aware weighting to prioritize early and semantically critical tokens. At the sequence level, CORD introduces a judge-based global reward to optimize complete reasoning trajectories via Group Relative Policy Optimization (GRPO). Empirical results across multiple benchmarks demonstrate that CORD consistently enhances audio-conditioned reasoning and substantially bridges the audio-text performance gap with only 80k synthetic training samples, validating the efficacy and data efficiency of our on-policy, multi-level cross-modal alignment approach.
【6】Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
标题:模特们像我们一样倾听吗?探索音频LLM和自然主义脑电波的代表性一致性
链接:https://arxiv.org/abs/2601.16540
摘要:音频大语言模型(Audio LLM)在整合语音感知和语言理解方面表现出强大的能力。然而,在自然主义听力过程中,它们的内部表征是否与人类神经动力学保持一致,在很大程度上尚未探索。在这项工作中,我们系统地研究了12个开源音频LLM和2个数据集上的脑电图(EEG)信号之间的逐层表示对齐。具体来说,我们采用8个相似性度量,如斯皮尔曼为基础的表征相似性分析(RSA),表征内的句子表征几何。我们的分析揭示了3个关键发现:(1)我们观察到一个等级依赖分裂,其中模型等级在不同的相似性度量中变化很大;(2)我们识别出时空对齐模式,其特征在于深度依赖对齐峰值和250-500 ms时间窗口内RSA的显著增加,与N400相关的神经动力学一致;(3)我们发现了一种情感分离,即使用三模态邻域一致性(TNC)标准识别的负韵律降低了几何相似性,同时增强了基于协方差的依赖性。这些发现为音频LLM的表征机制提供了新的神经生物学见解。
摘要:Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal representations align with human neural dynamics during naturalistic listening remains largely unexplored. In this work, we systematically examine layer-wise representational alignment between 12 open-source Audio LLMs and Electroencephalogram (EEG) signals across 2 datasets. Specifically, we employ 8 similarity metrics, such as Spearman-based Representational Similarity Analysis (RSA), to characterize within-sentence representational geometry. Our analysis reveals 3 key findings: (1) we observe a rank-dependence split, in which model rankings vary substantially across different similarity metrics; (2) we identify spatio-temporal alignment patterns characterized by depth-dependent alignment peaks and a pronounced increase in RSA within the 250-500 ms time window, consistent with N400-related neural dynamics; (3) we find an affective dissociation whereby negative prosody, identified using a proposed Tri-modal Neighborhood Consistency (TNC) criterion, reduces geometric similarity while enhancing covariance-based dependence. These findings provide new neurobiological insights into the representational mechanisms of Audio LLMs.
【7】The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
标题:CMU-AIST提交ICME 2025音频编码器挑战赛
链接:https://arxiv.org/abs/2601.16273
摘要:本技术报告介绍了我们提交的ICME 2025音频编码器挑战赛。我们提交的系统是建立在BEAT上的,BEAT是一种基于掩码语音令牌预测的音频编码器。我们使用来自各种语音,音乐和声音语料库的74,000小时的数据扩展了BEAT模型,并将其架构扩展到3亿个参数。我们使用语音密集和平衡的预训练混合物进行实验,以研究不同领域对最终性能的影响。我们提交的系统由大生12亿模型和两个定制的放大BEAT模型组成,这些模型是在上述预训练数据混合物上训练的。我们还提出了一种简单的集成技术,保留了组成模型的最佳功能,并超越了基线和大圣1.2B。对于开放科学,我们通过huggingface在https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND和https://huggingface.co/shikhar7ssu/OpenBEATs-ICME上公开发布我们训练的检查点。
摘要:This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encoder. We extend the BEATs model using 74,000 hours of data derived from various speech, music, and sound corpora and scale its architecture upto 300 million parameters. We experiment with speech-heavy and balanced pre-training mixtures to study the impact of different domains on final performance. Our submitted system consists of an ensemble of the Dasheng 1.2 billion model with two custom scaled-up BEATs models trained on the aforementioned pre-training data mixtures. We also propose a simple ensembling technique that retains the best capabilities of constituent models and surpasses both the baseline and Dasheng 1.2B. For open science, we publicly release our trained checkpoints via huggingface at https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND and https://huggingface.co/shikhar7ssu/OpenBEATs-ICME.
【8】Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
标题:个性化语音增强中嵌入细化的对比知识提炼
链接:https://arxiv.org/abs/2601.16235
摘要:个性化语音增强(PSE)在从干扰语音中提取已知目标语音方面表现出令人信服的结果。相应的系统通常将目标语音的表示并入增强系统内,该表示是从具有上游模型的目标语音的登记剪辑中提取的。这些模型通常是沉重的,因为说话人嵌入的质量直接影响PSE性能。然而,预先生成的嵌入不能解释在推理时间期间目标语音的变化。在本文中,我们建议使用一个微小的扬声器编码器执行对飞行细化的扬声器嵌入。我们首先介绍了一种新的对比知识蒸馏方法,以训练一个150 k参数的编码器从复杂的嵌入。然后,我们使用此编码器的增强系统在推理过程中,并表明所提出的方法大大提高了PSE的性能,同时保持低的计算负载。
摘要:Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding's quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-thefly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load.
【9】SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
标题:SoundBreak:三峰模型上纯音频对抗攻击的系统研究
链接:https://arxiv.org/abs/2601.16231
摘要:集成了音频、视觉和语言的多模态基础模型在推理和生成任务上实现了强大的性能,但它们对对抗性操作的鲁棒性仍然知之甚少。我们研究了一个现实的和未充分探索的威胁模型:无针对性的,只有音频的对抗性攻击的三模态音频-视频-语言模型。我们分析了针对多模态处理不同阶段的六个互补攻击目标,包括音频编码器表示,跨模态注意力,隐藏状态和输出可能性。在三个最先进的模型和多个基准测试中,我们表明,仅音频扰动可以引起严重的多模态故障,攻击成功率高达96%。我们进一步表明,攻击可以在低感知失真(LPIPS <= 0.08,SI-SNR >= 0)下成功,并且从扩展优化中受益更多,而不是增加数据规模。跨模型和编码器的可移植性仍然有限,而语音识别系统(如Whisper)主要响应扰动幅度,在严重失真的情况下实现>97%的攻击成功率。这些结果暴露了以前被忽视的单一模态攻击表面在多模态系统和激励防御,强制执行跨模态的一致性。
摘要:Multimodal foundation models that integrate audio, vision, and language achieve strong performance on reasoning and generation tasks, yet their robustness to adversarial manipulation remains poorly understood. We study a realistic and underexplored threat model: untargeted, audio-only adversarial attacks on trimodal audio-video-language models. We analyze six complementary attack objectives that target different stages of multimodal processing, including audio encoder representations, cross-modal attention, hidden states, and output likelihoods. Across three state-of-the-art models and multiple benchmarks, we show that audio-only perturbations can induce severe multimodal failures, achieving up to 96% attack success rate. We further show that attacks can be successful at low perceptual distortions (LPIPS <= 0.08, SI-SNR >= 0) and benefit more from extended optimization than increased data scale. Transferability across models and encoders remains limited, while speech recognition systems such as Whisper primarily respond to perturbation magnitude, achieving >97% attack success under severe distortion. These results expose a previously overlooked single-modality attack surface in multimodal systems and motivate defenses that enforce cross-modal consistency.
【10】Auditory Attention Decoding without Spatial Information: A Diotic EEG Study
标题:无空间信息的听觉注意力解码:一项神经脑电研究
链接:https://arxiv.org/abs/2601.16442
摘要:听觉注意解码(AAD)通过解码脑电信号(如EEG)来识别多说话人环境中的注意语音流。这项技术对于实现解决鸡尾酒会问题的智能助听器和促进客观测听系统至关重要。现有的AAD研究主要利用双耳分听环境,其中不同的语音信号被呈现给左耳和右耳,使得模型能够对方向性注意而不是语音内容进行分类。然而,这种空间依赖性限制了对真实世界场景的适用性,例如“鸡尾酒会”情况,其中扬声器重叠或动态移动。为了解决这一挑战,我们提出了一个AAD框架diotic环境中,相同的语音混合物呈现给双耳,消除空间线索。我们的方法将EEG和语音信号映射到一个共享的潜在空间,使用独立的编码器。我们使用wav2vec 2.0提取语音特征,并使用2层1D卷积神经网络(CNN)对其进行编码,同时采用BrainNetwork架构进行EEG编码。该模型通过计算EEG和语音表征之间的余弦相似度来识别所关注的语音。我们评估我们的方法在diotic EEG数据集上,并达到72.70%的准确率,这比最先进的基于方向的AAD方法高出22.58%。
摘要:Auditory attention decoding (AAD) identifies the attended speech stream in multi-speaker environments by decoding brain signals such as electroencephalography (EEG). This technology is essential for realizing smart hearing aids that address the cocktail party problem and for facilitating objective audiometry systems. Existing AAD research mainly utilizes dichotic environments where different speech signals are presented to the left and right ears, enabling models to classify directional attention rather than speech content. However, this spatial reliance limits applicability to real-world scenarios, such as the "cocktail party" situation, where speakers overlap or move dynamically. To address this challenge, we propose an AAD framework for diotic environments where identical speech mixtures are presented to both ears, eliminating spatial cues. Our approach maps EEG and speech signals into a shared latent space using independent encoders. We extract speech features using wav2vec 2.0 and encode them with a 2-layer 1D convolutional neural network (CNN), while employing the BrainNetwork architecture for EEG encoding. The model identifies the attended speech by calculating the cosine similarity between EEG and speech representations. We evaluate our method on a diotic EEG dataset and achieve 72.70% accuracy, which is 22.58% higher than the state-of-the-art direction-based AAD method.
【11】TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
标题:TidyVoice:一个基于普通语音的多语种说话人确认数据集
链接:https://arxiv.org/abs/2601.16358
备注:Accepted at ICASSP 2026
摘要:强大的多语言说话人识别系统的开发受到缺乏大规模,公开可用和多语言数据集的阻碍,特别是对于反欺骗等应用程序至关重要的阅读语音风格。为了解决这一差距,我们引入了TidyVoice数据集,该数据集来自Mozilla Common Voice语料库,在减轻了所提供的客户端ID中固有的扬声器异质性之后。TidyVoice目前包含来自超过212,000名单语使用者(Tidy-M)和约4,500名多语言使用者(Tidy-X)的训练和测试数据,我们从中得出两种不同的条件。Tidy-M条件包含来自81种语言的单语说话者的目标和非目标试验。Tidy-X条件包含来自多语种说话者的目标和非目标试验,包括相同和跨语言试验。我们采用两种ResNet模型架构,通过对我们全面的Tidy-M分区进行微调,实现了0.35%的EER。此外,我们表明,这种微调增强了模型的泛化能力,提高了CANDOR语料库中看不见的会话面试数据的性能。完整的数据集,评估试验和我们的模型都是公开发布的,为社区提供新的资源。
摘要:The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-speech style crucial for applications like anti-spoofing. To address this gap, we introduce the TidyVoice dataset derived from the Mozilla Common Voice corpus after mitigating its inherent speaker heterogeneity within the provided client IDs. TidyVoice currently contains training and test data from over 212,000 monolingual speakers (Tidy-M) and around 4,500 multilingual speakers (Tidy-X) from which we derive two distinct conditions. The Tidy-M condition contains target and non-target trials from monolingual speakers across 81 languages. The Tidy-X condition contains target and non-target trials from multilingual speakers in both same- and cross-language trials. We employ two architectures of ResNet models, achieving a 0.35% EER by fine-tuning on our comprehensive Tidy-M partition. Moreover, we show that this fine-tuning enhances the model's generalization, improving performance on unseen conversational interview data from the CANDOR corpus. The complete dataset, evaluation trials, and our models are publicly released to provide a new resource for the community.
【12】EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
标题:EdgeSpot:用于关键词发现的高效且高性能的Few-Shot模型
链接:https://arxiv.org/abs/2601.16316
备注:Accepted to be presented in IEEE ICASSP 2026
摘要:我们为边缘设备引入了一种高效的Few-Shot关键字定位模型EdgeSpot,该模型将基于BC-ResNet的声学骨干的优化版本与可训练的每通道能量归一化前端和轻量级的时间自注意力配对。在训练过程中,通过采用自我监督的教师模型来利用知识蒸馏,并使用子中心ArcFace损失进行优化。这项研究表明,EdgeSpot模型在固定的误报率(FAR)下始终比强大的BC-ResNet基线提供更好的准确性。最大的变体EdgeSpot-4将1% FAR下的10次射击精度从73.7%提高到82.0%,这只需要29.4M MAC和128 k参数。
摘要:We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters.
【13】Test-Time Adaptation for Speech Emotion Recognition
标题:语音情感识别的测试时间自适应
链接:https://arxiv.org/abs/2601.16240
备注:Accepted by 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:语音情感识别(SER)系统的实际效用是破坏了他们的脆弱性域的变化,如扬声器的变化,行为和自然主义的情绪之间的区别,和跨语料库的变化。虽然域自适应和微调被广泛研究,它们需要源数据或标记的目标数据,这往往是不可用的或提高隐私问题在SER。测试时自适应(TTA)通过调整模型在推理仅使用未标记的目标数据来弥合这一差距。然而,已经主要被设计用于图像分类和语音识别,TTA用于减轻SER中的独特域移位的功效尚未被研究。在本文中,我们提出了第一个系统的评估和比较,涵盖11 TTA方法在三个代表SER任务。结果表明,反向传播自由TTA方法是最有前途的。相反,熵最小化和伪标签通常会失败,因为它们的核心假设是一个单一的,自信的地面真理标签与情感表达的固有模糊性不相容。此外,没有任何一种方法是普遍适用的,其有效性高度依赖于分配的变化和任务。
摘要:The practical utility of Speech Emotion Recognition (SER) systems is undermined by their fragility to domain shifts, such as speaker variability, the distinction between acted and naturalistic emotions, and cross-corpus variations. While domain adaptation and fine-tuning are widely studied, they require either source data or labelled target data, which are often unavailable or raise privacy concerns in SER. Test-time adaptation (TTA) bridges this gap by adapting models at inference using only unlabeled target data. Yet, having been predominantly designed for image classification and speech recognition, the efficacy of TTA for mitigating the unique domain shifts in SER has not been investigated. In this paper, we present the first systematic evaluation and comparison covering 11 TTA methods across three representative SER tasks. The results indicate that backpropagation-free TTA methods are the most promising. Conversely, entropy minimization and pseudo-labeling generally fail, as their core assumption of a single, confident ground-truth label is incompatible with the inherent ambiguity of emotional expression. Further, no single method universally excels, and its effectiveness is highly dependent on the distributional shifts and tasks.
【14】Zero-Shot Speech LLMs for Multi-Aspect Evaluation of L2 Speech: Challenges and Opportunities
标题:用于L2言语多方面评估的Zero-Shot言语LLM:挑战和机遇
链接:https://arxiv.org/abs/2601.16230
备注:This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:准确评估L2英语发音对语言学习至关重要,因为它提供个性化的反馈,并确保对个人进步的公平评估。然而,自动化评分仍然具有挑战性,由于复杂的水平流畅性,韵律和完整性。本文评估了Qwen 2-Audio-7 B-Instruct的zero-shot性能,这是一种基于5,000个Speechocean 762语音的语音调谐LLM。该模型生成准确性,流畅性,韵律和完整性的评分,在+-2公差范围内与人类评分高度一致,特别是对于高质量的语音。然而,它倾向于过度预测低质量的语音分数,并且在错误检测中缺乏精度。这些研究结果表明,语音LLM在可扩展的发音评估中具有强大的潜力,并建议通过增强提示,校准和语音整合来推进计算机辅助发音训练。
摘要:An accurate assessment of L2 English pronunciation is crucial for language learning, as it provides personalized feedback and ensures a fair evaluation of individual progress. However, automated scoring remains challenging due to the complexity of sentence-level fluency, prosody, and completeness. This paper evaluates the zero-shot performance of Qwen2-Audio-7B-Instruct, an instruction-tuned speech-LLM, on 5,000 Speechocean762 utterances. The model generates rubric-aligned scores for accuracy, fluency, prosody, and completeness, showing strong agreement with human ratings within +-2 tolerance, especially for high-quality speech. However, it tends to overpredict low-quality speech scores and lacks precision in error detection. These findings demonstrate the strong potential of speech LLMs in scalable pronunciation assessment and suggest future improvements through enhanced prompting, calibration, and phonetic integration to advance Computer-Assisted Pronunciation Training.
【15】ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
标题:ES 4 R:基于前置情感模型的语音编码用于同理心反应生成
链接:https://arxiv.org/abs/2601.16225
摘要:移情言语对话不仅需要理解语言内容,还需要感知丰富的非语言信息,如韵律,音调和情感强度的情感理解。现有的语音到语音大型语言模型要么依赖于ASR转录,要么使用编码器来提取潜在的表示,这往往会削弱多轮对话中的情感信息和上下文连贯性。为了解决这个问题,我们提出了\textbf{ES 4 R},一个基于语音的移情反应生成框架。我们的核心创新在于在语音编码之前明确地建模结构化的情感上下文,而不是依赖于编码器的内隐学习或外显的情感监督。具体来说,我们引入了双级别注意力机制来捕获回合级别的情感状态和对话级别的情感动态。然后,通过语音引导的跨模态注意,将所产生的情感表征与文本语义相结合,以产生移情反应。对于语音输出,我们采用基于能量的策略选择和风格融合来实现移情语音合成。ES 4 R在自动和人工评估中始终优于强大的基线,并在不同的LLM主干中保持稳健。
摘要:Empathetic speech dialogue requires not only understanding linguistic content but also perceiving rich paralinguistic information such as prosody, tone, and emotional intensity for affective understandings. Existing speech-to-speech large language models either rely on ASR transcription or use encoders to extract latent representations, often weakening affective information and contextual coherence in multi-turn dialogues. To address this, we propose \textbf{ES4R}, a framework for speech-based empathetic response generation. Our core innovation lies in explicitly modeling structured affective context before speech encoding, rather than relying on implicit learning by the encoder or explicit emotion supervision. Specifically, we introduce a dual-level attention mechanism to capture turn-level affective states and dialogue-level affective dynamics. The resulting affective representations are then integrated with textual semantics through speech-guided cross-modal attention to generate empathetic responses. For speech output, we employ energy-based strategy selection and style fusion to achieve empathetic speech synthesis. ES4R consistently outperforms strong baselines in both automatic and human evaluations and remains robust across different LLM backbones.
【1】FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
标题:FlowSE-GRPO:通过在线强化学习训练流匹配语音增强
链接:https://arxiv.org/abs/2601.16483
备注:Accepted by ICASSP 2026
摘要:生成式语音增强通过对以噪声输入为条件的干净语音的分布进行建模,为传统的判别方法提供了一种有前途的替代方案。通过强化学习(RL)进行的训练后对齐有效地将生成模型与自然语言处理等领域的人类偏好和下游指标对齐,但其在语音增强中的使用仍然有限,特别是对于在线RL。以前的工作探索离线方法,如直接偏好优化(DPO);在线方法,如组相对策略优化(GRPO)仍然在很大程度上未被研究。在本文中,我们首次成功地将在线GRPO集成到流匹配语音增强框架中,从而实现了有效的训练后对齐,以感知和面向任务的指标,只需很少的更新步骤。与先前GRPO在大型语言模型上的工作不同,我们将算法适应语音的连续性,时间序列性和流匹配生成模型的动态性。我们表明,优化一个单一的奖励产生快速的度量增益,但往往会导致奖励黑客降低音频保真度,尽管更高的分数。为了缓解这一问题,我们提出了一个多指标奖励优化策略,平衡竞争目标,大大减少过拟合和提高整体性能。我们的实验验证了在线GRPO的语音增强,并为基于RL的生成音频模型的后训练提供了实际指导。
摘要:Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-training alignment via reinforcement learning (RL) effectively aligns generative models with human preferences and downstream metrics in domains such as natural language processing, but its use in speech enhancement remains limited, especially for online RL. Prior work explores offline methods like Direct Preference Optimization (DPO); online methods such as Group Relative Policy Optimization (GRPO) remain largely uninvestigated. In this paper, we present the first successful integration of online GRPO into a flow-matching speech enhancement framework, enabling efficient post-training alignment to perceptual and task-oriented metrics with few update steps. Unlike prior GRPO work on Large Language Models, we adapt the algorithm to the continuous, time-series nature of speech and to the dynamics of flow-matching generative models. We show that optimizing a single reward yields rapid metric gains but often induces reward hacking that degrades audio fidelity despite higher scores. To mitigate this, we propose a multi-metric reward optimization strategy that balances competing objectives, substantially reducing overfitting and improving overall performance. Our experiments validate online GRPO for speech enhancement and provide practical guidance for RL-based post-training of generative audio models.
【2】Auditory Attention Decoding without Spatial Information: A Diotic EEG Study
标题:无空间信息的听觉注意力解码:一项神经脑电研究
链接:https://arxiv.org/abs/2601.16442
摘要:听觉注意解码(AAD)通过解码脑电信号(如EEG)来识别多说话人环境中的注意语音流。这项技术对于实现解决鸡尾酒会问题的智能助听器和促进客观测听系统至关重要。现有的AAD研究主要利用双耳分听环境,其中不同的语音信号被呈现给左耳和右耳,使得模型能够对方向性注意而不是语音内容进行分类。然而,这种空间依赖性限制了对真实世界场景的适用性,例如“鸡尾酒会”情况,其中扬声器重叠或动态移动。为了解决这一挑战,我们提出了一个AAD框架diotic环境中,相同的语音混合物呈现给双耳,消除空间线索。我们的方法将EEG和语音信号映射到一个共享的潜在空间,使用独立的编码器。我们使用wav2vec 2.0提取语音特征,并使用2层1D卷积神经网络(CNN)对其进行编码,同时采用BrainNetwork架构进行EEG编码。该模型通过计算EEG和语音表征之间的余弦相似度来识别所关注的语音。我们评估我们的方法在diotic EEG数据集上,并达到72.70%的准确率,这比最先进的基于方向的AAD方法高出22.58%。
摘要:Auditory attention decoding (AAD) identifies the attended speech stream in multi-speaker environments by decoding brain signals such as electroencephalography (EEG). This technology is essential for realizing smart hearing aids that address the cocktail party problem and for facilitating objective audiometry systems. Existing AAD research mainly utilizes dichotic environments where different speech signals are presented to the left and right ears, enabling models to classify directional attention rather than speech content. However, this spatial reliance limits applicability to real-world scenarios, such as the "cocktail party" situation, where speakers overlap or move dynamically. To address this challenge, we propose an AAD framework for diotic environments where identical speech mixtures are presented to both ears, eliminating spatial cues. Our approach maps EEG and speech signals into a shared latent space using independent encoders. We extract speech features using wav2vec 2.0 and encode them with a 2-layer 1D convolutional neural network (CNN), while employing the BrainNetwork architecture for EEG encoding. The model identifies the attended speech by calculating the cosine similarity between EEG and speech representations. We evaluate our method on a diotic EEG dataset and achieve 72.70% accuracy, which is 22.58% higher than the state-of-the-art direction-based AAD method.
【3】TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
标题:TidyVoice:一个基于普通语音的多语种说话人确认数据集
链接:https://arxiv.org/abs/2601.16358
备注:Accepted at ICASSP 2026
摘要:强大的多语言说话人识别系统的开发受到缺乏大规模,公开可用和多语言数据集的阻碍,特别是对于反欺骗等应用程序至关重要的阅读语音风格。为了解决这一差距,我们引入了TidyVoice数据集,该数据集来自Mozilla Common Voice语料库,在减轻了所提供的客户端ID中固有的扬声器异质性之后。TidyVoice目前包含来自超过212,000名单语使用者(Tidy-M)和约4,500名多语言使用者(Tidy-X)的训练和测试数据,我们从中得出两种不同的条件。Tidy-M条件包含来自81种语言的单语说话者的目标和非目标试验。Tidy-X条件包含来自多语种说话者的目标和非目标试验,包括相同和跨语言试验。我们采用两种ResNet模型架构,通过对我们全面的Tidy-M分区进行微调,实现了0.35%的EER。此外,我们表明,这种微调增强了模型的泛化能力,提高了CANDOR语料库中看不见的会话面试数据的性能。完整的数据集,评估试验和我们的模型都是公开发布的,为社区提供新的资源。
摘要:The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-speech style crucial for applications like anti-spoofing. To address this gap, we introduce the TidyVoice dataset derived from the Mozilla Common Voice corpus after mitigating its inherent speaker heterogeneity within the provided client IDs. TidyVoice currently contains training and test data from over 212,000 monolingual speakers (Tidy-M) and around 4,500 multilingual speakers (Tidy-X) from which we derive two distinct conditions. The Tidy-M condition contains target and non-target trials from monolingual speakers across 81 languages. The Tidy-X condition contains target and non-target trials from multilingual speakers in both same- and cross-language trials. We employ two architectures of ResNet models, achieving a 0.35% EER by fine-tuning on our comprehensive Tidy-M partition. Moreover, we show that this fine-tuning enhances the model's generalization, improving performance on unseen conversational interview data from the CANDOR corpus. The complete dataset, evaluation trials, and our models are publicly released to provide a new resource for the community.
【4】EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
标题:EdgeSpot:用于关键词发现的高效且高性能的Few-Shot模型
链接:https://arxiv.org/abs/2601.16316
备注:Accepted to be presented in IEEE ICASSP 2026
摘要:我们为边缘设备引入了一种高效的Few-Shot关键字定位模型EdgeSpot,该模型将基于BC-ResNet的声学骨干的优化版本与可训练的每通道能量归一化前端和轻量级的时间自注意力配对。在训练过程中,通过采用自我监督的教师模型来利用知识蒸馏,并使用子中心ArcFace损失进行优化。这项研究表明,EdgeSpot模型在固定的误报率(FAR)下始终比强大的BC-ResNet基线提供更好的准确性。最大的变体EdgeSpot-4将1% FAR下的10次射击精度从73.7%提高到82.0%,这只需要29.4M MAC和128 k参数。
摘要:We introduce an efficient few-shot keyword spotting model for edge devices, EdgeSpot, that pairs an optimized version of a BC-ResNet-based acoustic backbone with a trainable Per-Channel Energy Normalization frontend and lightweight temporal self-attention. Knowledge distillation is utilized during training by employing a self-supervised teacher model, optimized with Sub-center ArcFace loss. This study demonstrates that the EdgeSpot model consistently provides better accuracy at a fixed false-alarm rate (FAR) than strong BC-ResNet baselines. The largest variant, EdgeSpot-4, improves the 10-shot accuracy at 1% FAR from 73.7% to 82.0%, which requires only 29.4M MACs with 128k parameters.
【5】Test-Time Adaptation for Speech Emotion Recognition
标题:语音情感识别的测试时间自适应
链接:https://arxiv.org/abs/2601.16240
备注:Accepted by 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:语音情感识别(SER)系统的实际效用是破坏了他们的脆弱性域的变化,如扬声器的变化,行为和自然主义的情绪之间的区别,和跨语料库的变化。虽然域自适应和微调被广泛研究,它们需要源数据或标记的目标数据,这往往是不可用的或提高隐私问题在SER。测试时自适应(TTA)通过调整模型在推理仅使用未标记的目标数据来弥合这一差距。然而,已经主要被设计用于图像分类和语音识别,TTA用于减轻SER中的独特域移位的功效尚未被研究。在本文中,我们提出了第一个系统的评估和比较,涵盖11 TTA方法在三个代表SER任务。结果表明,无反向传播的TTA方法是最有前途的。相反,熵最小化和伪标签通常会失败,因为它们的核心假设是一个单一的,自信的地面真理标签与情感表达的固有模糊性不相容。此外,没有任何一种方法是普遍适用的,其有效性高度依赖于分配的变化和任务。
摘要:The practical utility of Speech Emotion Recognition (SER) systems is undermined by their fragility to domain shifts, such as speaker variability, the distinction between acted and naturalistic emotions, and cross-corpus variations. While domain adaptation and fine-tuning are widely studied, they require either source data or labelled target data, which are often unavailable or raise privacy concerns in SER. Test-time adaptation (TTA) bridges this gap by adapting models at inference using only unlabeled target data. Yet, having been predominantly designed for image classification and speech recognition, the efficacy of TTA for mitigating the unique domain shifts in SER has not been investigated. In this paper, we present the first systematic evaluation and comparison covering 11 TTA methods across three representative SER tasks. The results indicate that backpropagation-free TTA methods are the most promising. Conversely, entropy minimization and pseudo-labeling generally fail, as their core assumption of a single, confident ground-truth label is incompatible with the inherent ambiguity of emotional expression. Further, no single method universally excels, and its effectiveness is highly dependent on the distributional shifts and tasks.
【6】Zero-Shot Speech LLMs for Multi-Aspect Evaluation of L2 Speech: Challenges and Opportunities
标题:用于L2言语多方面评估的Zero-Shot言语LLM:挑战和机遇
链接:https://arxiv.org/abs/2601.16230
备注:This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)
摘要:准确评估L2英语发音对语言学习至关重要,因为它提供个性化的反馈,并确保对个人进步的公平评估。然而,自动化评分仍然具有挑战性,由于复杂的水平流畅性,韵律和完整性。本文评估了Qwen 2-Audio-7 B-Instruct的zero-shot性能,这是一种基于5,000个Speechocean 762语音的语音调谐LLM。该模型生成准确性,流畅性,韵律和完整性的评分,在+-2公差范围内与人类评分高度一致,特别是对于高质量的语音。然而,它倾向于过度预测低质量的语音分数,并且在错误检测中缺乏精度。这些研究结果表明,语音LLM在可扩展的发音评估中具有强大的潜力,并建议通过增强提示,校准和语音整合来推进计算机辅助发音训练。
摘要:An accurate assessment of L2 English pronunciation is crucial for language learning, as it provides personalized feedback and ensures a fair evaluation of individual progress. However, automated scoring remains challenging due to the complexity of sentence-level fluency, prosody, and completeness. This paper evaluates the zero-shot performance of Qwen2-Audio-7B-Instruct, an instruction-tuned speech-LLM, on 5,000 Speechocean762 utterances. The model generates rubric-aligned scores for accuracy, fluency, prosody, and completeness, showing strong agreement with human ratings within +-2 tolerance, especially for high-quality speech. However, it tends to overpredict low-quality speech scores and lacks precision in error detection. These findings demonstrate the strong potential of speech LLMs in scalable pronunciation assessment and suggest future improvements through enhanced prompting, calibration, and phonetic integration to advance Computer-Assisted Pronunciation Training.
【7】ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
标题:ES 4 R:基于前置情感模型的语音编码用于同理心反应生成
链接:https://arxiv.org/abs/2601.16225
摘要:移情言语对话不仅需要理解语言内容,还需要感知丰富的非语言信息,如韵律,音调和情感强度的情感理解。现有的语音到语音大型语言模型要么依赖于ASR转录,要么使用编码器来提取潜在的表示,这往往会削弱多轮对话中的情感信息和上下文连贯性。为了解决这个问题,我们提出了\textbf{ES 4 R},一个基于语音的移情反应生成框架。我们的核心创新在于在语音编码之前明确地建模结构化的情感上下文,而不是依赖于编码器的内隐学习或外显的情感监督。具体来说,我们引入了一个双层次的注意力机制来捕捉轮层次的情感状态和对话层次的情感动态。然后,通过语音引导的跨模态注意力将由此产生的情感表示与文本语义整合,以产生同理心反应。对于语音输出,我们采用基于能量的策略选择和风格融合来实现移情语音合成。ES 4 R在自动和人工评估中始终优于强大的基线,并在不同的LLM主干中保持稳健。
摘要:Empathetic speech dialogue requires not only understanding linguistic content but also perceiving rich paralinguistic information such as prosody, tone, and emotional intensity for affective understandings. Existing speech-to-speech large language models either rely on ASR transcription or use encoders to extract latent representations, often weakening affective information and contextual coherence in multi-turn dialogues. To address this, we propose \textbf{ES4R}, a framework for speech-based empathetic response generation. Our core innovation lies in explicitly modeling structured affective context before speech encoding, rather than relying on implicit learning by the encoder or explicit emotion supervision. Specifically, we introduce a dual-level attention mechanism to capture turn-level affective states and dialogue-level affective dynamics. The resulting affective representations are then integrated with textual semantics through speech-guided cross-modal attention to generate empathetic responses. For speech output, we employ energy-based strategy selection and style fusion to achieve empathetic speech synthesis. ES4R consistently outperforms strong baselines in both automatic and human evaluations and remains robust across different LLM backbones.
【8】A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
标题:语音信号心理稳定性分类的新迁移学习方法
链接:https://arxiv.org/abs/2601.16793
摘要:这项研究提出了一种新的迁移学习方法和数据增强技术,用于使用人类语音信号进行精神稳定性分类,并解决了与有限数据可用性相关的挑战。卷积神经网络(CNN)已经被用于分析从语音记录生成的频谱图图像。三个CNN架构,VGG 16,InceptionV3和DenseNet121,在三个实验阶段进行了评估:非增强数据,增强数据和迁移学习。这种提出的迁移学习方法涉及在增强数据集上预训练模型,并在非增强数据集上对其进行微调,同时确保严格的数据分离以防止数据泄漏。结果表明,与基线方法相比,分类性能有显着改善。在三种CNN架构中,使用所提出的迁移学习方法,DenseNet121实现了94%的最高准确率和99%的AUC得分。这一发现强调了数据增强和迁移学习相结合的有效性,可以使用语音频谱图增强基于CNN的心理稳定性分类,为心理健康诊断提供了一种有前途的非侵入性工具。
摘要:This study presents a novel transfer learning approach and data augmentation technique for mental stability classification using human voice signals and addresses the challenges associated with limited data availability. Convolutional neural networks (CNNs) have been employed to analyse spectrogram images generated from voice recordings. Three CNN architectures, VGG16, InceptionV3, and DenseNet121, were evaluated across three experimental phases: training on non-augmented data, augmented data, and transfer learning. This proposed transfer learning approach involves pre-training models on the augmented dataset and fine-tuning them on the non-augmented dataset while ensuring strict data separation to prevent data leakage. The results demonstrate significant improvements in classification performance compared to the baseline approach. Among three CNN architectures, DenseNet121 achieved the highest accuracy of 94% and an AUC score of 99% using the proposed transfer learning approach. This finding highlights the effectiveness of combining data augmentation and transfer learning to enhance CNN-based classification of mental stability using voice spectrograms, offering a promising non-invasive tool for mental health diagnostics.
【9】E2E-AEC: Implementing an end-to-end neural network learning approach for acoustic echo cancellation
标题:E2 E-AEC:实施端到端神经网络学习方法用于声学回声消除
链接:https://arxiv.org/abs/2601.16774
备注:This paper has been accepted by ICASSP2026
摘要:我们提出了一种新型的基于神经网络的端到端声学回声消除(E2 E-AEC)方法,能够进行流推理,该方法可以有效地运行,而不依赖于传统的线性AEC(LAEC)技术和时间延迟估计。我们的方法包括几个关键策略:首先,我们引入并改进渐进式学习,以逐步增强回声抑制。其次,我们的模型通过使用预先训练的基于LAEC的模型进行初始化,利用从LAEC培训中获得的见解来进行知识转移。第三,我们优化了注意力机制,在注意力权重上应用了损失函数,以实现参考信号和麦克风信号之间的精确时间对准。最后,我们将语音活动检测,以提高语音质量和改善回声消除掩蔽网络输出时,近端语音缺席。通过在公共数据集上进行的实验验证了该方法的有效性。
摘要:We propose a novel neural network-based end-to-end acoustic echo cancellation (E2E-AEC) method capable of streaming inference, which operates effectively without reliance on traditional linear AEC (LAEC) techniques and time delay estimation. Our approach includes several key strategies: First, we introduce and refine progressive learning to gradually enhance echo suppression. Second, our model employs knowledge transfer by initializing with a pre-trained LAECbased model, harnessing the insights gained from LAEC training. Third, we optimize the attention mechanism with a loss function applied on attention weights to achieve precise time alignment between the reference and microphone signals. Lastly, we incorporate voice activity detection to enhance speech quality and improve echo removal by masking the network output when near-end speech is absent. The effectiveness of our approach is validated through experiments conducted on public datasets.
【10】I Guess That's Why They Call it the Blues: Causal Analysis for Audio Classifiers
标题:我想这就是为什么他们称之为蓝调:音频分类器的因果分析
链接:https://arxiv.org/abs/2601.16675
摘要:众所周知,音频分类器通常依赖于非音乐相关特征和虚假相关性来对音频进行分类。因此,音频分类器很容易被操纵或混淆,导致错误的分类。虽然导致错误分类并不难,但直到现在,分类器所依赖的特征集还没有得到很好的理解。 在本文中,我们介绍了一种新的方法,使用因果推理发现的频率空间的功能是足够的,必要的一个给定的分类。我们描述了该算法在工具FreqReX中的实现,并提供了一些标准基准数据集上的实验结果。我们的实验表明,因果充分和必要的子集允许我们通过非常轻微地改变输入以各种方式操纵模型的输出。也就是说,改变到240,000个频率中的一个频率会导致58%的时间的分类改变,并且该改变可以小到几乎听不见。这些结果表明,因果分析是有用的理解音频分类器的推理过程,并可以用来成功地操纵他们的输出。
摘要:It is well-known that audio classifiers often rely on non-musically relevant features and spurious correlations to classify audio. Hence audio classifiers are easy to manipulate or confuse, resulting in wrong classifications. While inducing a misclassification is not hard, until now the set of features that the classifiers rely on was not well understood. In this paper we introduce a new method that uses causal reasoning to discover features of the frequency space that are sufficient and necessary for a given classification. We describe an implementation of this algorithm in the tool FreqReX and provide experimental results on a number of standard benchmark datasets. Our experiments show that causally sufficient and necessary subsets allow us to manipulate the outputs of the models in a variety of ways by changing the input very slightly. Namely, a change to one out of 240,000 frequencies results in a change in classification 58% of the time, and the change can be so small that it is practically inaudible. These results show that causal analysis is useful for understanding the reasoning process of audio classifiers and can be used to successfully manipulate their outputs.
【11】Omni-directional attention mechanism based on Mamba for speech separation
标题:基于Mamba的全方位注意力机制在语音分离中的应用
链接:https://arxiv.org/abs/2601.16603
摘要:Mamba是一种选择性状态空间模型(SSM),它是Transformers语音建模的有效替代方案,可以实现线性复杂度的长序列处理。虽然在语音分离中是有效的,但现有的方法,无论是在时域还是时频域中,通常在使用Mamba处理它们之前,将输入沿着单个维度分解为短的一维序列,这将其限制为局部1D建模,并限制了其在2D频谱图中捕获全局依赖关系的能力。在这项工作中,我们提出了一个有效的全方位的注意力(OA)机制,建立在单向曼巴,从十个不同的方向上的频谱图模型的全球依赖关系。我们将所提出的机制扩展为两个基线分离模型,并在三个公共数据集上进行评估。实验结果表明,我们的方法在保持线性复杂度的同时,始终比基线实现了显着的性能提升,优于现有的最先进(SOTA)系统。
摘要:Mamba, a selective state-space model (SSM), has emerged as an efficient alternative to Transformers for speech modeling, enabling long-sequence processing with linear complexity. While effective in speech separation, existing approaches, whether in the time or time-frequency domain, typically decompose the input along a single dimension into short one-dimensional sequences before processing them with Mamba, which restricts it to local 1D modeling and limits its ability to capture global dependencies across the 2D spectrogram. In this work, we propose an efficient omni-directional attention (OA) mechanism built upon unidirectional Mamba, which models global dependencies from ten different directions on the spectrogram. We expand the proposed mechanism into two baseline separation models and evaluate on three public datasets. Experimental results show that our approach consistently achieves significant performance gains over the baselines while preserving linear complexity, outperforming existing state-of-the-art (SOTA) systems.
【12】CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
标题:CORD:通过加权政策跨模式蒸馏弥合音频文本推理差距
链接:https://arxiv.org/abs/2601.16547
备注:13 pages, 4 figures
摘要:大型音频语言模型(LALM)已经获得了重大的研究兴趣。尽管建立在基于文本的大型语言模型(LLM)上,但LALM经常表现出知识和推理能力的下降。我们假设,这种限制源于当前的训练范式的失败,有效地弥合特征表示空间内的声学语义差距。为了应对这一挑战,我们提出了CORD,一个统一的对齐框架,执行在线跨模态自蒸馏。具体来说,它在一个统一的模型中将音频条件推理与文本条件推理相结合。利用文本模态作为内部教师,CORD在整个音频推出过程中执行多粒度对齐。在令牌级别,它采用具有重要性感知权重的策略反向KL发散来优先考虑早期和语义关键的令牌。在序列级,CORD引入了基于判断的全局奖励,通过组相对策略优化(GRPO)来优化完整的推理轨迹。多个基准测试的实证结果表明,CORD始终增强了音频条件推理,并仅用80 k合成训练样本就大大弥合了音频-文本性能差距,验证了我们的策略,多级跨模态对齐方法的有效性和数据效率。
摘要:Large Audio Language Models (LALMs) have garnered significant research interest. Despite being built upon text-based large language models (LLMs), LALMs frequently exhibit a degradation in knowledge and reasoning capabilities. We hypothesize that this limitation stems from the failure of current training paradigms to effectively bridge the acoustic-semantic gap within the feature representation space. To address this challenge, we propose CORD, a unified alignment framework that performs online cross-modal self-distillation. Specifically, it aligns audio-conditioned reasoning with its text-conditioned counterpart within a unified model. Leveraging the text modality as an internal teacher, CORD performs multi-granularity alignment throughout the audio rollout process. At the token level, it employs on-policy reverse KL divergence with importance-aware weighting to prioritize early and semantically critical tokens. At the sequence level, CORD introduces a judge-based global reward to optimize complete reasoning trajectories via Group Relative Policy Optimization (GRPO). Empirical results across multiple benchmarks demonstrate that CORD consistently enhances audio-conditioned reasoning and substantially bridges the audio-text performance gap with only 80k synthetic training samples, validating the efficacy and data efficiency of our on-policy, multi-level cross-modal alignment approach.
【13】Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
标题:模特们像我们一样倾听吗?探索音频LLM和自然主义脑电波的代表性一致性
链接:https://arxiv.org/abs/2601.16540
摘要:音频大语言模型(Audio LLM)在整合语音感知和语言理解方面表现出强大的能力。然而,在自然主义听力过程中,它们的内部表征是否与人类神经动力学保持一致,在很大程度上尚未探索。在这项工作中,我们系统地研究了12个开源音频LLM和2个数据集上的脑电图(EEG)信号之间的逐层表示对齐。具体来说,我们采用8个相似性度量,如斯皮尔曼为基础的表征相似性分析(RSA),表征内的句子表征几何。我们的分析揭示了3个关键发现:(1)我们观察到一个等级依赖分裂,其中模型等级在不同的相似性度量中变化很大;(2)我们识别出时空对齐模式,其特征在于深度依赖对齐峰值和250-500 ms时间窗口内RSA的显著增加,与N400相关的神经动力学一致;(3)我们发现了一种情感分离,即使用三模态邻域一致性(TNC)标准识别的负韵律降低了几何相似性,同时增强了基于协方差的依赖性。这些发现为音频LLM的表征机制提供了新的神经生物学见解。
摘要:Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal representations align with human neural dynamics during naturalistic listening remains largely unexplored. In this work, we systematically examine layer-wise representational alignment between 12 open-source Audio LLMs and Electroencephalogram (EEG) signals across 2 datasets. Specifically, we employ 8 similarity metrics, such as Spearman-based Representational Similarity Analysis (RSA), to characterize within-sentence representational geometry. Our analysis reveals 3 key findings: (1) we observe a rank-dependence split, in which model rankings vary substantially across different similarity metrics; (2) we identify spatio-temporal alignment patterns characterized by depth-dependent alignment peaks and a pronounced increase in RSA within the 250-500 ms time window, consistent with N400-related neural dynamics; (3) we find an affective dissociation whereby negative prosody, identified using a proposed Tri-modal Neighborhood Consistency (TNC) criterion, reduces geometric similarity while enhancing covariance-based dependence. These findings provide new neurobiological insights into the representational mechanisms of Audio LLMs.
【14】The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
标题:CMU-AIST提交ICME 2025音频编码器挑战赛
链接:https://arxiv.org/abs/2601.16273
摘要:本技术报告介绍了我们提交的ICME 2025音频编码器挑战赛。我们提交的系统是建立在BEAT上的,BEAT是一种基于掩码语音令牌预测的音频编码器。我们使用来自各种语音,音乐和声音语料库的74,000小时的数据扩展了BEAT模型,并将其架构扩展到3亿个参数。我们使用语音密集和平衡的预训练混合物进行实验,以研究不同领域对最终性能的影响。我们提交的系统由大生12亿模型和两个定制的放大BEAT模型组成,这些模型是在上述预训练数据混合物上训练的。我们还提出了一种简单的集成技术,保留了组成模型的最佳功能,并超越了基线和大圣1.2B。对于开放科学,我们通过huggingface在https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND和https://huggingface.co/shikhar7ssu/OpenBEATs-ICME上公开发布我们训练的检查点。
摘要:This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encoder. We extend the BEATs model using 74,000 hours of data derived from various speech, music, and sound corpora and scale its architecture upto 300 million parameters. We experiment with speech-heavy and balanced pre-training mixtures to study the impact of different domains on final performance. Our submitted system consists of an ensemble of the Dasheng 1.2 billion model with two custom scaled-up BEATs models trained on the aforementioned pre-training data mixtures. We also propose a simple ensembling technique that retains the best capabilities of constituent models and surpasses both the baseline and Dasheng 1.2B. For open science, we publicly release our trained checkpoints via huggingface at https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND and https://huggingface.co/shikhar7ssu/OpenBEATs-ICME.
【15】Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
标题:个性化语音增强中嵌入细化的对比知识提炼
链接:https://arxiv.org/abs/2601.16235
摘要:个性化语音增强(PSE)在从干扰语音中提取已知目标语音方面表现出令人信服的结果。相应的系统通常将目标语音的表示并入增强系统内,该表示是从具有上游模型的目标语音的登记剪辑中提取的。这些模型通常是沉重的,因为说话人嵌入的质量直接影响PSE性能。然而,预先生成的嵌入不能解释在推理时间期间目标语音的变化。在本文中,我们建议使用一个微小的扬声器编码器执行对飞行细化的扬声器嵌入。我们首先介绍了一种新的对比知识蒸馏方法,以训练一个150 k参数的编码器从复杂的嵌入。然后,我们使用此编码器的增强系统在推理过程中,并表明所提出的方法大大提高了PSE的性能,同时保持低的计算负载。
摘要:Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding's quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-thefly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load.
【16】SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
标题:SoundBreak:三峰模型上纯音频对抗攻击的系统研究
链接:https://arxiv.org/abs/2601.16231
摘要:集成了音频、视觉和语言的多模态基础模型在推理和生成任务上实现了强大的性能,但它们对对抗性操作的鲁棒性仍然知之甚少。我们研究了一个现实的和未充分探索的威胁模型:无针对性的,只有音频的对抗性攻击的三模态音频-视频-语言模型。我们分析了针对多模态处理不同阶段的六个互补攻击目标,包括音频编码器表示,跨模态注意力,隐藏状态和输出可能性。在三个最先进的模型和多个基准测试中,我们表明,仅音频扰动可以引起严重的多模态故障,攻击成功率高达96%。我们进一步表明,攻击可以在低感知失真(LPIPS <= 0.08,SI-SNR >= 0)下成功,并且从扩展优化中受益更多,而不是增加数据规模。跨模型和编码器的可移植性仍然有限,而语音识别系统(如Whisper)主要响应扰动幅度,在严重失真的情况下实现>97%的攻击成功率。这些结果暴露了以前被忽视的单一模态攻击表面在多模态系统和激励防御,强制执行跨模态的一致性。
摘要:Multimodal foundation models that integrate audio, vision, and language achieve strong performance on reasoning and generation tasks, yet their robustness to adversarial manipulation remains poorly understood. We study a realistic and underexplored threat model: untargeted, audio-only adversarial attacks on trimodal audio-video-language models. We analyze six complementary attack objectives that target different stages of multimodal processing, including audio encoder representations, cross-modal attention, hidden states, and output likelihoods. Across three state-of-the-art models and multiple benchmarks, we show that audio-only perturbations can induce severe multimodal failures, achieving up to 96% attack success rate. We further show that attacks can be successful at low perceptual distortions (LPIPS <= 0.08, SI-SNR >= 0) and benefit more from extended optimization than increased data scale. Transferability across models and encoders remains limited, while speech recognition systems such as Whisper primarily respond to perturbation magnitude, achieving >97% attack success under severe distortion. These results expose a previously overlooked single-modality attack surface in multimodal systems and motivate defenses that enforce cross-modal consistency.
机器翻译由腾讯交互翻译提供,仅供参考
