微信公众号:arXiv_Daily
cs.SD语音
【1】AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
标题:AC-Foley:具有声学传输的参考音频引导视频到音频合成
链接:https://arxiv.org/abs/2603.15597
备注:Accepted at ICLR 2026. 15 pages, 5 figures
摘要:现有的视频到音频(V2 A)生成方法主要依赖于文本提示以及视觉信息来合成音频。然而,两个关键的瓶颈仍然存在:训练数据中的语义粒度差距,例如将声学上不同的声音合并在粗糙的标签下,以及描述微观声学特征时的文本模糊性。这些瓶颈使得很难使用文本控制模式执行细粒度的声音合成。为了解决这些限制,我们提出了AC-Foley,一种音频调节的V2 A模型,它直接利用参考音频来实现对生成的声音的精确和细粒度控制。这种方法能够实现细粒度的声音合成、音色转移、zero-shot声音生成和改进的音频质量。通过直接调节音频信号,我们的方法绕过了文本描述的语义模糊性,同时能够精确地操纵声学属性。从经验上讲,AC-Foley在参考音频条件下实现了Foley生成的最先进性能,同时即使没有音频条件,也能与最先进的视频到音频方法保持竞争力。
摘要:Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating acoustically distinct sounds under coarse labels, and textual ambiguity in describing micro-acoustic features. These bottlenecks make it difficult to perform fine-grained sound synthesis using text-controlled modes. To address these limitations, we propose AC-Foley, an audio-conditioned V2A model that directly leverages reference audio to achieve precise and fine-grained control over generated sounds. This approach enables fine-grained sound synthesis, timbre transfer, zero-shot sound generation, and improved audio quality. By directly conditioning on audio signals, our approach bypasses the semantic ambiguities of text descriptions while enabling precise manipulation of acoustic attributes. Empirically, AC-Foley achieves state-of-the-art performance for Foley generation when conditioned on reference audio, while remaining competitive with state-of-the-art video-to-audio methods even without audio conditioning.
【2】Music Genre Classification: A Comparative Analysis of Classical Machine Learning and Deep Learning Approaches
标题:音乐流派分类:经典机器学习和深度学习方法的比较分析
链接:https://arxiv.org/abs/2603.15440
备注:8 pages
摘要:自动音乐流派分类是音乐信息检索(MIR)中的一个长期挑战;非西方音乐传统的工作仍然很少。尼泊尔音乐包含了丰富的文化和不同的音乐类型-从Lok Dohori的呼唤和回应二重唱到Deuda的节奏诗和Tamang Selo的独特旋律-这些都没有被现有的分类系统所解决。在本文中,我们构建了一个新的数据集,包含大约8,000个标记的30秒音频片段,跨越8个尼泊尔音乐流派,并对两种范式的9个分类模型进行了系统的比较。五种经典的机器学习分类器(Logistic Regression,SVM,KNN,Random Forest和XGBoost)在通过Librosa提取的51个手工制作的音频特征上进行训练,而四种深度学习架构(CNN,RNN,并行CNN-RNN和顺序CNN,然后是RNN)在维度为640 x 128的Mel频谱图上进行操作。我们的实验表明,序列卷积递归神经网络(CRNN)-其中卷积层馈入LSTM-实现了84%的最高准确度,大大优于最佳经典模型(Logistic Regression和XGBoost,均为71%)和所有其他深度架构。我们为每个模型提供每类精度,召回率,F1分数,混淆矩阵和ROC分析,并提供基于文化的错误分类模式解释,反映了尼泊尔音乐传统的真正重叠。
摘要:Automatic music genre classification is a long-standing challenge in Music Information Retrieval (MIR); work on non-Western music traditions remains scarce. Nepali music encompasses culturally rich and acoustically diverse genres--from the call-and-response duets of Lok Dohori to the rhythmic poetry of Deuda and the distinctive melodies of Tamang Selo--that have not been addressed by existing classification systems. In this paper, we construct a novel dataset of approximately 8,000 labeled 30-second audio clips spanning eight Nepali music genres and conduct a systematic comparison of nine classification models across two paradigms. Five classical machine learning classifiers (Logistic Regression, SVM, KNN, Random Forest, and XGBoost) are trained on 51 hand-crafted audio features extracted via Librosa, while four deep learning architectures (CNN, RNN, parallel CNN-RNN, and sequential CNN followed by RNN) operate on Mel spectrograms of dimension 640 x 128. Our experiments reveal that the sequential Convolutional Recurrent Neural Network (CRNN)--in which convolutional layers feed into an LSTM--achieves the highest accuracy of 84%, substantially outperforming both the best classical models (Logistic Regression and XGBoost, both at 71%) and all other deep architectures. We provide per-class precision, recall, F1-score, confusion matrices, and ROC analysis for every model, and offer a culturally grounded interpretation of misclassification patterns that reflects genuine overlaps in Nepal's musical traditions.
【3】NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
标题:NV-Bench:用于表达性文本到语音生成的非言语发声合成基准
链接:https://arxiv.org/abs/2603.15352
摘要:虽然最近的文本到语音(TTS)系统越来越多地集成非语言发声(NV),但它们的评估缺乏标准化的指标和可靠的地面实况参考。为了弥合这一差距,我们提出了NV-Bench,第一个基准接地的功能分类法,将NVs作为沟通行为,而不是声学文物。NV-Bench包含1,651个多语言的野外话语,带有配对的人类参考音频,在14个NV类别中平衡。我们引入了一个双维度的评估协议:(1)指令对齐,利用提出的语言字符错误率(PCER)来评估可控性,(2)声学保真度,测量分布差距,以真实的录音,以评估声学的真实性。我们评估了不同的TTS模型,并开发了两个基线。实验结果表明,我们的客观指标和人类感知之间的强相关性,建立NV-Bench作为一个标准化的评估框架。
摘要:While recent text-to-speech (TTS) systems increasingly integrate nonverbal vocalizations (NVs), their evaluations lack standardized metrics and reliable ground-truth references. To bridge this gap, we propose NV-Bench, the first benchmark grounded in a functional taxonomy that treats NVs as communicative acts rather than acoustic artifacts. NV-Bench comprises 1,651 multi-lingual, in-the-wild utterances with paired human reference audio, balanced across 14 NV categories. We introduce a dual-dimensional evaluation protocol: (1) Instruction Alignment, utilizing the proposed paralinguistic character error rate (PCER) to assess controllability, (2) Acoustic Fidelity, measuring the distributional gap to real recordings to assess acoustic realism. We evaluate diverse TTS models and develop two baselines. Experimental results demonstrate a strong correlation between our objective metrics and human perception, establishing NV-Bench as a standardized evaluation framework.
【4】Two-Stage Adaptation for Non-Normative Speech Recognition: Revisiting Speaker-Independent Initialization for Personalization
标题:非规范语音识别的两阶段自适应:重新审视用于个性化的与说话人无关的自适应算法
链接:https://arxiv.org/abs/2603.15261
备注:submitted to Interspeech 2026
摘要:个性化的自动语音识别(ASR)系统的非规范性语音,如构音障碍和失语症的语音,是具有挑战性的。虽然特定于说话者的微调(SS-FT)被广泛使用,但它通常直接从通用的预训练模型初始化。在这种不匹配的情况下,说话者无关自适应是否提供了更强的初始化先验仍然不清楚。在这项工作中,我们提出了一个两阶段的自适应框架,包括扬声器独立微调(SI-FT)的多扬声器非规范性数据,其次是SS-FT,并通过控制比较与直接SS-FT在相同的每个扬声器条件下进行评估。使用Whisper-Large-v3和Qwen 3-ASR对AphasiaBank和UA语音进行的实验,以及对典型语音数据集TED-LIUM v3和FLEURS的评估表明,两阶段自适应始终提高了个性化,同时保持了可管理的域外(OOD)权衡。
摘要:Personalizing automatic speech recognition (ASR) systems for non-normative speech, such as dysarthric and aphasic speech, is challenging. While speaker-specific fine-tuning (SS-FT) is widely used, it is typically initialized directly from a generic pre-trained model. Whether speaker-independent adaptation provides a stronger initialization prior under such mismatch remains unclear. In this work, we propose a two-stage adaptation framework consisting of speaker-independent fine-tuning (SI-FT) on multi-speaker non-normative data followed by SS-FT, and evaluate it through a controlled comparison with direct SS-FT under identical per-speaker conditions. Experiments on AphasiaBank and UA-Speech with Whisper-Large-v3 and Qwen3-ASR, alongside evaluation on typical-speech datasets TED-LIUM v3 and FLEURS, show that two-stage adaptation consistently improves personalization while maintaining manageable out-of-domain (OOD) trade-offs.
【5】ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
标题:ReactMotion:根据说话者的言论生成反应性语音动作
链接:https://arxiv.org/abs/2603.15083
备注:42 pages, 11 tables, 8 figures
摘要:在本文中,我们介绍了一个新的任务,反应性运动生成从扬声器话语,其目的是生成自然的听者身体运动,适当地响应扬声器的话语。然而,由于人类反应固有的非确定性,对这种非语言听众行为的建模仍然是探索不足和具有挑战性的。为了促进这项任务,我们提出了ReactMotionNet,这是一个大规模的数据集,它将说话者的话语与多个候选听者的运动配对,并以不同程度的适当性进行注释。这种数据集设计明确地捕捉了监听者行为的一对多性质,并提供了超越单一地面实况运动的监督。在此数据集设计的基础上,我们开发了面向偏好的评估协议,用于评估反应的适当性,其中传统的运动指标专注于输入运动对齐忽略。我们进一步提出了ReactMotion,这是一个统一的生成框架,它联合建模文本,音频,情感和运动,并使用基于偏好的目标进行训练,以鼓励适当和多样化的听众响应。大量的实验表明,ReactMotion优于检索基线和级联的基于LLM的管道,生成更自然,多样和适当的侦听器运动。
摘要:In this paper, we introduce a new task, Reactive Listener Motion Generation from Speaker Utterance, which aims to generate naturalistic listener body motions that appropriately respond to a speaker's utterance. However, modeling such nonverbal listener behaviors remains underexplored and challenging due to the inherently non-deterministic nature of human reactions. To facilitate this task, we present ReactMotionNet, a large-scale dataset that pairs speaker utterances with multiple candidate listener motions annotated with varying degrees of appropriateness. This dataset design explicitly captures the one-to-many nature of listener behavior and provides supervision beyond a single ground-truth motion. Building on this dataset design, we develop preference-oriented evaluation protocols tailored to evaluate reactive appropriateness, where conventional motion metrics focusing on input-motion alignment ignore. We further propose ReactMotion, a unified generative framework that jointly models text, audio, emotion, and motion, and is trained with preference-based objectives to encourage both appropriate and diverse listener responses. Extensive experiments show that ReactMotion outperforms retrieval baselines and cascaded LLM-based pipelines, generating more natural, diverse, and appropriate listener motions.
【6】PhonemeDF: A Synthetic Speech Dataset for Audio Deepfake Detection and Naturalness Evaluation
标题:PhonemEDF:用于音频Deepfake检测和自然度评估的合成语音数据集
链接:https://arxiv.org/abs/2603.15037
备注:11 pages, 6 figures, 9 tables. Accepted at the 15th Language Resources and Evaluation Conference (LREC 2026), Palma, Spain
摘要:人工智能(AI)生成的语音越来越复杂,这给音频深度伪造检测带来了新的挑战。文语转换(TTS)和语音转换(VC)技术可以创建具有自然度和可懂度的高度令人信服的合成语音。这对语音生物识别安全和旨在打击口头错误信息传播的系统构成了严重威胁,其中合成语音可能被用于传播虚假或恶意内容。虽然对人工智能生成语音的兴趣有所增加,但用于在音素层面评估自然度的资源仍然有限。在这项工作中,我们通过呈现音素级DeepFake数据集(PhonemEDF)来解决这一差距,该数据集包括在音素级分割的并行真实和合成语音。真实的语音样本来自LibriSpeech的一个子集,而合成样本使用四个TTS和三个VC系统生成。对于每个系统,音素对齐的TextGrid文件使用Montreal Forced Aligner(MFA)获得。我们计算真实和合成音素分布之间的Kullback-Leibler分歧(KLD)来量化保真度,并根据与自然语音的相似性建立排名。我们的研究结果显示,真实和合成音素分布的KLD与经过训练以区分它们的分类器的性能之间存在明显的相关性,这表明KLD可以作为最具鉴别力的音素的指标,用于deepfake检测。
摘要:The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic speech with naturalness and intelligibility. This poses serious threats to voice biometric security and to systems designed to combat the spread of spoken misinformation, where synthetic voices may be used to disseminate false or malicious content. While interest in AI-generated speech has increased, resources for evaluating naturalness at the phoneme level remain limited. In this work, we address this gap by presenting the Phoneme-Level DeepFake dataset (PhonemeDF), comprising parallel real and synthetic speech segmented at the phoneme level. Real speech samples are derived from a subset of LibriSpeech, while synthetic samples are generated using four TTS and three VC systems. For each system, phoneme-aligned TextGrid files are obtained using the Montreal Forced Aligner (MFA). We compute the Kullback-Leibler divergence (KLD) between real and synthetic phoneme distributions to quantify fidelity and establish a ranking based on similarity to natural speech. Our findings show a clear correlation between the KLD of real and synthetic phoneme distributions and the performance of classifiers trained to distinguish them, suggesting that KLD can serve as an indicator of the most discriminative phonemes for deepfake detection.
【7】Cepstral Smoothing of Binary Masks for Convolutive Blind Separation of Speech Mixtures
标题:混合语音卷积盲分离中二值掩码的倒谱平滑
链接:https://arxiv.org/abs/2603.14983
摘要:在本文中,我们提出了一种新的分离系统,从两个麦克风录音提取两个语音信号。我们的系统结合了盲源分离技术与二进制时频掩模的倒谱平滑。最后一步由两个步骤组成。首先,从分离的BSS算法的输出信号中估计两个二进制掩码。在第二步中,倒谱平滑应用于这些频谱掩模,以减少通常由时频掩蔽产生的音乐噪声。实验进行了两个人工混合的语音信号,使用模拟的房间模型和两个真实的录音。评价结果是有希望的,并显示了我们的系统的有效性。
摘要:In this paper, we propose a novel separation system for extracting two speech signals from two microphone recordings. Our system combines the blind source separation technique with cepstral smoothing of binary time-frequency masks. The last is composed of two steps. First, the two binary masks are estimated from the separated output signals of BSS algorithm. In the second step, a cepstral smoothing is applied of these spectral masks in order to reduce musical noise typically produced by time-frequency masking. Experiments were carried out with both artificially mixed speech signals using simulated room model and two real recordings. The evaluation results are promising and have shown the effectiveness of our system.
【8】WhispSynth: Scaling Multilingual Whisper Corpus through Real Data Curation and A Novel Pitch-free Generative Framework
标题:WhispSynth:通过真实数据处理和新型无音调生成框架扩展多语言Whisper Corpus
链接:https://arxiv.org/abs/2603.14853
备注:Under Review
摘要:耳语的产生受到数据收集难度的限制。由于耳语音具有较低的声振幅,因此高保真记录具有挑战性。在本文中,我们介绍了WhispSynth,一个大规模的多语种语料库,通过一种新的高保真生成框架构建。具体来说,我们提出了一个管道集成差分数字信号处理(DDSP)为基础的无音高的方法与文本到语音(TTS)模型。这个框架完善了一个全面的资源集合,包括我们新构建的WhispNJU数据集,从479个扬声器中提取了118小时的高保真耳语。与标准的合成或嘈杂的真实数据不同,我们的数据引擎忠实地保留了源语音音色和语言内容,同时确保声学一致性,为文本到耳语研究提供了坚实的基础。实验结果表明,WhispSynth具有显着更高的质量比现有的语料库。此外,我们的CosyWhisper,与WhispSynth调谐,实现与地面真实样本同等的语音自然度。正式实施和相关资源可在https://github.com/tan90xx/cosywhisper上获取。
摘要:Whisper generation is constrained by the difficulty of data collection. Because whispered speech has low acoustic amplitude, high-fidelity recording is challenging. In this paper, we introduce WhispSynth, a large-scale multilingual corpus constructed via a novel high-fidelity generative framework. Specifically, we propose a pipeline integrating Differentiable Digital Signal Processing (DDSP)-based pitch-free method with Text-to-Speech (TTS) models. This framework refines a comprehensive collection of resources, including our newly constructed WhispNJU dataset, into 118 hours of high-fidelity whispered speech from 479 speakers. Unlike standard synthetic or noisy real data, our data engine faithfully preserves source vocal timbre and linguistic content while ensuring acoustic consistency, providing a robust foundation for text-to-whisper research. Experimental results demonstrate that WhispSynth exhibits significantly higher quality than existing corpora. Moreover, our CosyWhisper, tuned with WhispSynth, achieves speech naturalness on par with ground-truth samples. The official implementation and related resources are available at https://github.com/tan90xx/cosywhisper.
【9】VorTEX: Various overlap ratio for Target speech EXtraction
标题:VorTEK:目标语音提取的各种重叠率
链接:https://arxiv.org/abs/2603.14803
备注:arXiv Preprint
摘要:目标语音提取(TSE)的目的是从混合语音中恢复出目标说话人的语音。虽然最近的文本提示方法已经显示出希望,但大多数方法都假设完全重叠的混合物,限制了对实际重叠率行为的深入了解。我们介绍了VorTEX(目标语音提取的各种重叠率),这是一种文本提示的TSE架构,具有解耦自适应多分支(DAM)融合块,将主要提取与辅助正则化路径分离。为了实现受控分析,我们构建了PORTE,这是一个覆盖重叠率从0%到100%的双说话人数据集。我们还提出了能量抑制比(SuRE),这是一种诊断指标,可以检测常规措施无法捕获的抑制行为。实验表明,现有模型在重叠下表现出抑制或残留干扰,而VorTEX在20-100%重叠范围内实现了最高的分离保真度(例如,5.50 20%时为2.04 dB,100%时为2.04 dB),同时保持零SuRE,表明提取稳健,无抑制驱动的伪影。
摘要:Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overlapped mixtures, limiting insight into behavior across realistic overlap ratios. We introduce VorTEX (Various overlap ratio for Target speech EXtraction), a text-prompted TSE architecture with a Decoupled Adaptive Multi-branch (DAM) Fusion block that separates primary extraction from auxiliary regularization pathways. To enable controlled analysis, we construct PORTE, a two-speaker dataset spanning overlap ratios from 0% to 100%. We further propose Suppression Ratio on Energy (SuRE), a diagnostic metric that detects suppression behavior not captured by conventional measures. Experiments show that existing models exhibit suppression or residual interference under overlap, whereas VorTEX achieves the highest separation fidelity across 20-100% overlap (e.g., 5.50 dB at 20% and 2.04 dB at 100%) while maintaining zero SuRE, indicating robust extraction without suppression-driven artifacts.
【10】Investigating the Impact of Speech Enhancement on Audio Deepfake Detection in Noisy Environments
标题:研究语音增强对噪音环境中音频深度伪造检测的影响
链接:https://arxiv.org/abs/2603.14767
摘要:逻辑访问(LA)攻击,也称为音频深度伪造攻击,使用文本到语音(TTS)或语音转换(VC)方法来生成伪造的语音数据。这可能对自动说话人验证(ASV)系统构成严重威胁,因为入侵者可以使用此类攻击绕过语音生物识别安全。在这项研究中,我们研究了语音质量和音频欺骗检测系统性能之间的相关性(即,LA任务)。为此,两个增强算法的性能进行评估的基础上两个感知语音质量的措施,即感知评价语音质量(PESQ)和语音混响调制比(SRMR),并在其对音频欺骗检测系统的影响。我们采用了ASVspoof 2019挑战赛中提供的LA数据集,并使用不同的信噪比(SNR)水平破坏其测试集,同时保持训练数据不变。采用增强的方法来减弱噪声对语音的不利影响,并比较了语音增强生成对抗网络(SEGAN)和度量优化生成对抗网络Plus(MetricGAN+)两种模型的性能。虽然我们预期语音质量将与语音应用的性能很好地相关,但是如果引入不想要的伪像或从语音信号中移除相关信息,则其也可能对下游任务产生副作用。我们的研究结果证实了这一假设,因为我们发现,导致最高语音质量分数的增强算法MetricGAN+在音频欺骗检测任务中提供了最低的等误率(EER),而具有最低语音质量分数的增强方法SEGAN导致了最低的EER,从而在LA任务中获得了更好的性能。
摘要:Logical Access (LA) attacks, also known as audio deepfake attacks, use Text-to-Speech (TTS) or Voice Conversion (VC) methods to generate spoofed speech data. This can represent a serious threat to Automatic Speaker Verification (ASV) systems, as intruders can use such attacks to bypass voice biometric security. In this study, we investigate the correlation between speech quality and the performance of audio spoofing detection systems (i.e., LA task). For that, the performance of two enhancement algorithms is evaluated based on two perceptual speech quality measures, namely Perceptual Evaluation of Speech Quality (PESQ) and Speech-to-Reverberation Modulation Ratio (SRMR), and in respect to their impact on the audio spoofing detection system. We adopted the LA dataset, provided in the ASVspoof 2019 Challenge, and corrupted its test set with different Signal-to-Noise Ratio (SNR) levels, while leaving the training data untouched. Enhancement was applied to attenuate the detrimental effects of noisy speech, and the performances of two models, Speech Enhancement Generative Adversarial Network (SEGAN) and Metric-Optimized Generative Adversarial Network Plus (MetricGAN+), were compared. Although we expect that speech quality will correlate well with speech applications' performance, it can also have as a side effect on downstream tasks if unwanted artifacts are introduced or relevant information is removed from the speech signal. Our results corroborate with this hypothesis, as we found that the enhancement algorithm leading to the highest speech quality scores, MetricGAN+, provided the lowest Equal Error Rate (EER) on the audio spoofing detection task, whereas the enhancement method with the lowest speech quality scores, SEGAN, led to the lowest EER, thus leading to better performance on the LA task.
【11】Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models
标题:推动隐藏状态:大型音频语言模型中思想链推理的免训练模型引导
链接:https://arxiv.org/abs/2603.14636
备注:6 pages, 4 figures, 2 tables
摘要:思想链(CoT)提示已经扩展到大型音频语言模型(LALM)以引发推理,但在没有训练的情况下提高其有效性仍然具有挑战性。我们研究了推理时间模型转向作为一种无需训练的方法来改善LALM推理。我们介绍了三种策略,使用不同的信息来源,并评估他们在四个LALM和四个基准。结果显示,一般的准确性增益高达4.4%,超过CoT提示。值得注意的是,我们确定了一个跨模态传输,其中来自少数文本样本的导向向量有效地引导基于语音的推理,表现出高数据效率。我们还研究了超参数敏感性,以了解这些方法的鲁棒性。我们的研究结果定位模型转向作为一个实际的方向,加强LALM推理。
摘要:Chain-of-thought (CoT) prompting has been extended to large audio-language models (LALMs) to elicit reasoning, yet enhancing its effectiveness without training remains challenging. We study inference-time model steering as a training-free approach to improve LALM reasoning. We introduce three strategies using diverse information sources and evaluate them across four LALMs and four benchmarks. Results show general accuracy gains up to 4.4% over CoT prompting. Notably, we identify a cross-modal transfer where steering vectors derived from few text samples effectively guide speech-based reasoning, demonstrating high data efficiency. We also examine hyperparameter sensitivity to understand the robustness of these approaches. Our findings position model steering as a practical direction for strengthening LALM reasoning.
【12】PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark
标题:PARSA-Bench:全面的波斯语音频模型基准
链接:https://arxiv.org/abs/2603.14456
备注:Submitted to Interspeech 2026
摘要:波斯语通过其古典诗歌,传统音乐和普遍的代码转换提出了独特的音频理解挑战-现有基准没有捕获。我们介绍了PARSA-Bench(波斯语音频推理和语音评估基准),这是第一个评估波斯语和文化的大型音频语言模型的基准,包括16个任务和8,000多个样本,涉及语音理解,语言分析和文化音频理解。新引入了十个任务,包括诗歌韵律和风格检测、传统波斯音乐理解和语码转换检测。纯文本基线的表现一直优于音频基线,这表明模型可能不会利用音频特定信息,而不仅仅是转录本身。基于文化的任务暴露出一种性质上不同的失败模式:所有模型在vazn检测上都表现出近乎随机的机会,无论规模如何,这表明韵律感知仍然超出了当前模型的范围。该数据集可在https://huggingface.co/datasets/MohammadJRanjbar/PARSA-Bench上公开获取
摘要:Persian poses unique audio understanding challenges through its classical poetry, traditional music, and pervasive code-switching - none captured by existing benchmarks. We introduce PARSA-Bench (Persian Audio Reasoning and Speech Assessment Benchmark), the first benchmark for evaluating large audio-language models on Persian language and culture, comprising 16 tasks and over 8,000 samples across speech understanding, paralinguistic analysis, and cultural audio understanding. Ten tasks are newly introduced, including poetry meter and style detection, traditional Persian music understanding, and code-switching detection. Text-only baselines consistently outperform audio counterparts, suggesting models may not leverage audio-specific information beyond what transcription alone provides. Culturally-grounded tasks expose a qualitatively distinct failure mode: all models perform near random chance on vazn detection regardless of scale, suggesting prosodic perception remains beyond the reach of current models. The dataset is publicly available at https://huggingface.co/datasets/MohammadJRanjbar/PARSA-Bench
【13】Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
标题:Affectron:具有情感和上下文一致的非言语发声的情感语音合成
链接:https://arxiv.org/abs/2603.14432
摘要:非言语发声(NV),如笑声和叹息,是情感语音合成中情感线索表达的核心。然而,由于NV数据有限和缺乏明确的监督,在开放环境中学习多样化和上下文一致的NV仍然具有挑战性。出于这一挑战,我们提出Affectron作为情感和上下文对齐NV生成的框架。建立在一个小规模的开放和解耦的语料库,Affectron引入了一个NV增强的训练策略,扩大了NV类型和插入位置的分布。我们进一步将NV结构掩蔽纳入纯口头语音预训练的语音主干中,以实现多样化和自然的NV合成。实验结果表明,Affectron产生更多的表现力和多样化的NV比基线系统,同时保持自然的口头语音流。
摘要:Nonverbal vocalizations (NVs), such as laughter and sighs, are central to the expression of affective cues in emotional speech synthesis. However, learning diverse and contextually aligned NVs remains challenging in open settings due to limited NV data and the lack of explicit supervision. Motivated by this challenge, we propose Affectron as a framework for affective and contextually aligned NV generation. Built on a small-scale open and decoupled corpus, Affectron introduces an NV-augmented training strategy that expands the distribution of NV types and insertion locations. We further incorporate NV structural masking into a speech backbone pre-trained on purely verbal speech to enable diverse and natural NV synthesis. Experimental results demonstrate that Affectron produces more expressive and diverse NVs than baseline systems while preserving the naturalness of the verbal speech stream.
【14】CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
标题:CodecMOS-Accent:跨英语口音的神经编解码器重新合成和TTC语音的MOS基准
链接:https://arxiv.org/abs/2603.14328
备注:Preprint
摘要:我们提出了CodecMOS-Accent数据集,这是一个平均意见评分(MOS)基准,旨在评估神经音频编解码器(NAC)模型和基于大语言模型(LLM)的文本到语音(TTS)模型,特别是在非标准语音(如口音语音)中。该数据集包括来自24个系统的4,000个编解码器再合成和TTS样本,具有32个扬声器,跨越10个口音。进行了大规模的主观测试,从25名听众中收集了19,600条注释,涉及三个维度:自然度,说话人相似性和口音相似性。该数据集不仅代表了对最近语音合成系统性能的最新研究,而且揭示了一些见解,包括说话者和口音相似性之间的紧密关系,客观指标的预测能力,以及当听众与说话者共享相同口音时的感知偏差。预计该数据集将促进对NAC和重音TTS进行更多以人为中心的评估的研究。
摘要:We present the CodecMOS-Accent dataset, a mean opinion score (MOS) benchmark designed to evaluate neural audio codec (NAC) models and the large language model (LLM)-based text-to-speech (TTS) models trained upon them, especially across non-standard speech like accented speech. The dataset comprises 4,000 codec resynthesis and TTS samples from 24 systems, featuring 32 speakers spanning ten accents. A large-scale subjective test was conducted to collect 19,600 annotations from 25 listeners across three dimensions: naturalness, speaker similarity, and accent similarity. This dataset does not only represent an up-to-date study of recent speech synthesis system performance but reveals insights including a tight relationship between speaker and accent similarity, the predictive power of objective metrics, and a perceptual bias when listeners share the same accent with the speaker. This dataset is expected to foster research on more human-centric evaluation for NAC and accented TTS.
【15】DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
标题:DiFlowDubber:通过跨模式对齐和同步实现自动视频配音的离散流匹配
链接:https://arxiv.org/abs/2603.14267
备注:Accepted at CVPR 2026 Findings
摘要:视频配音在电影制作、多媒体创作和辅助语音技术中有着广泛的应用。现有的方法要么直接在有限的配音数据集上进行训练,要么采用两阶段流水线来适应预先训练的文本到语音(TTS)模型,这通常难以产生富有表现力的韵律、丰富的声学特征和精确的同步。为了解决这些问题,我们提出了DiFlowDubber,它具有一个新的两阶段训练框架,可以有效地将知识从预训练的TTS模型转移到视频驱动的配音,并具有离散流匹配生成骨干。具体来说,我们设计了一个FaPro模块,从面部表情中捕捉全球韵律和风格线索,并利用这些信息来指导后续语音属性的建模。为了确保精确的语音嘴唇同步,我们引入了一个同步器模块,它可以弥合文本,视频和语音之间的模态差距,从而改善跨模态对齐并生成与嘴唇运动时间同步的语音。在两个主要基准数据集上的实验表明,DiFlowDubber在多个指标上优于以前的方法。
摘要:Video dubbing has broad applications in filmmaking, multimedia creation, and assistive speech technology. Existing approaches either train directly on limited dubbing datasets or adopt a two-stage pipeline that adapts pre-trained text-to-speech (TTS) models, which often struggle to produce expressive prosody, rich acoustic characteristics, and precise synchronization. To address these issues, we propose DiFlowDubber with a novel two-stage training framework that effectively transfers knowledge from a pre-trained TTS model to video-driven dubbing, with a discrete flow matching generative backbone. Specifically, we design a FaPro module that captures global prosody and stylistic cues from facial expressions and leverages this information to guide the modeling of subsequent speech attributes. To ensure precise speech-lip synchronization, we introduce a Synchronizer module that bridges the modality gap among text, video, and speech, thereby improving cross-modal alignment and generating speech that is temporally synchronized with lip movements. Experiments on two primary benchmark datasets demonstrate that DiFlowDubber outperforms previous methods across multiple metrics.
【16】Semi-Automatic Flute Robot and Its Acoustic Sensing
标题:半自动长笛机器人及其声学传感
链接:https://arxiv.org/abs/2603.14180
备注:This paper was submitted to a journal and received thorough reviews with high marks from the experts. Despite addressing three rounds of major revisions, it was ultimately rejected due to an unreasonable reviewer. We are uploading it here as a preprint
摘要:长笛演奏需要掌握复杂的指法组合和音域依赖的口部控制,特别是低音域生产的喷射偏移调整。现有的触觉和半自动系统不能通过机械致动同时解决这两个方面。据我们所知,没有现有的系统完全自动化指法,同时机械地帮助低登记音的生产,而不需要口控制。我们开发了一款带有自动指法机制的半自动长笛机器人:十四个伺服电机通过基于线和齿条齿轮的驱动器驱动所有键,以响应气流输入,使表演者能够仅通过气流产生完整的音乐作品。喷射偏移辅助机构在低音域通道期间通过校准的22 °旋转头部接头,将喷射偏移向低音域配置移动,而无需修改乐器或吹奏口。基本频率估计确认正确的音高生产在整个色范围(C4-C7)和音乐表演。所有键和杠杆动作均在77.50~ms内完成,对应节拍能力超过标准要求。和声分析($Δ SPL} = \mathrm {SPL}_2 - \mathrm {SPL}_3$)显示,激活时所有低音域音符的$Δ$SPL一致增加,与预期的喷射偏移一致。头关节旋转完成40.00毫秒。这些结果表明,在受控条件下,集成自动化指法和寄存器相关的射流偏移辅助的机械可行性。
摘要:Flute performance requires mastery of complex fingering combinations and register-dependent embouchure control, particularly jet offset adjustment for low-register production. Existing haptic and semi-automated systems do not address both aspects simultaneously through mechanical actuation. To our knowledge, no prior system fully automates fingering while mechanically assisting low-register tone production without requiring embouchure control. We developed a semi-automatic flute robot with an automatic fingering mechanism: fourteen servo motors actuate all keys via wire-based and rack-and-pinion drives in response to MIDI input, enabling performers to produce complete musical pieces through airflow alone. A jet offset assist mechanism rotates the head joint by a calibrated $22^\circ$ during low-register passages, shifting the jet offset toward a low-register configuration without modifying the instrument or embouchure. Fundamental frequency estimation confirmed correct pitch production across the chromatic range (C4--C7) and during musical performance. All key and lever movements were completed within 77.50~ms, corresponding to tempo capacity exceeding standard requirements. Harmonic analysis ($Δ\mathrm{SPL} = \mathrm{SPL}_2 - \mathrm{SPL}_3$) showed a consistent increase in $Δ$SPL for all low-register notes when activated, consistent with the intended jet offset shift. Head joint rotation completed within 40.00~ms. These results demonstrate mechanical feasibility of integrating automated fingering and register-dependent jet offset assistance under controlled conditions.
【17】Probing neural audio codecs for distinctions among English nuclear tunes
标题:探索神经音频编解码器以了解英语核曲调之间的区别
链接:https://arxiv.org/abs/2603.14035
备注:5 pages; 1 table; 3 figures. Accepted as conference paper at Speech Prosody 2026
摘要:最先进的口语对话模型(Défossez et al. 2024; Schalkwyk et al. 2025)使用神经音频编解码器将音频信号“标记化”为矢量潜在表示的低频流,每个矢量潜在表示都使用矢量码本的层次结构进行量化。Transformer层允许这些表示反映一些依赖于时间和上下文的模式。我们对来自Cole等人(2023)的标记音频数据进行训练探针,以测试表征英语短语结尾(核)语调曲调的音高轨迹是否在这些模式中。结果如下:线性探针训练的未量化的潜在的或一些相关的码字产生以上的机会的准确性,在区分八个语音指定的核曲调单调音高口音(最高平均测试精度(TATA):0.31)和这些曲调的五个集群,是强大的人类语音生产和感知(TATA:0.45)。更高的准确性(TATA:0.74-0.89),分别用于问题和断言的上升与下降曲调的类别之间的二元区分。关于曲调的信息在所有码本中传播,这使得文献中发现的“语义”和“声学”码本之间的区别受到质疑。使用非线性探测器可以提高精度,但五个聚类之间的区分仍然远远不能达到人类的性能,这表明当前编解码器存在根本性的局限性。
摘要:State-of-the-art spoken dialogue models (Défossez et al. 2024; Schalkwyk et al. 2025) use neural audio codecs to "tokenize" audio signals into a lower-frequency stream of vectorial latent representations, each quantized using a hierarchy of vector codebooks. A transformer layer allows these representations to reflect some time- and context-dependent patterns. We train probes on labeled audio data from Cole et al. (2023) to test whether the pitch trajectories that characterize English phrase-final (nuclear) intonational tunes are among these patterns. Results: Linear probes trained on the unquantized latents or some of the associated codewords yield above-chance accuracy in distinguishing eight phonologically specified nuclear tunes with monotonal pitch accents (top average test accuracy (TATA): 0.31) and the five clusters of these tunes that are robust in human speech production and perception (TATA: 0.45). Greater accuracy (TATAs: 0.74-0.89) is attained for binary distinctions between classes of rising vs. falling tunes, respectively used for questions and assertions. Information about tunes is spread among all codebooks, which calls into question a distinction between 'semantic' and 'acoustic' codebooks found in the literature. Accuracies improve with nonlinear probes, but discrimination among the five clusters remains far from human performance, suggesting a fundamental limitation of current codecs.
【18】What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
标题:什么才算真实?语音恢复和语音质量转换对Deepfake检测提出新挑战
链接:https://arxiv.org/abs/2603.14033
备注:5 pages, 4 figures, 3 tables. Submitted to Interspeech 2026
摘要:音频反欺骗系统通常被公式化为区分真实语音和欺骗语音的二进制分类器。这种假设在分层生成处理下失败,其中良性转换引入了被错误分类为欺骗的分布偏移。我们表明,发音修改语音转换和语音恢复被视为外的分布,尽管保留扬声器的真实性。使用分离真实,转换,欺骗和转换欺骗语音的多类设置,我们通过自监督学习(SSL)嵌入和声学相关来分析模型行为。良性的转换在SSL空间中引起漂移,压缩真实和欺骗的语音并降低分类器的可分离性。将反欺骗重新定义为多类问题,提高了对良性变化的鲁棒性,同时保留了欺骗检测,这表明二进制系统对原始语音的分布而不是真实性本身进行建模。
摘要:Audio anti-spoofing systems are typically formulated as binary classifiers distinguishing bona fide from spoofed speech. This assumption fails under layered generative processing, where benign transformations introduce distributional shifts that are misclassified as spoofing. We show that phonation-modifying voice conversion and speech restoration are treated as out-of-distribution despite preserving speaker authenticity. Using a multi-class setup separating bona fide, converted, spoofed, and converted-spoofed speech, we analyse model behaviour through self-supervised learning (SSL) embeddings and acoustic correlates. The benign transformations induce a drift in the SSL space, compressing bona fide and spoofed speech and reducing classifier separability. Reformulating anti-spoofing as a multi-class problem improves robustness to benign shifts while preserving spoof detection, suggesting binary systems model the distribution of raw speech rather than authenticity itself.
【19】LightBeam: An Accurate and Memory-Efficient CTC Decoder for Speech Neuroprostheses
标题:LightBeam:用于语音神经假体的准确且内存高效的CSC解码器
链接:https://arxiv.org/abs/2603.14002
备注:4 pages, 2 figures
摘要:一种有希望的恢复构音障碍和构音障碍患者沟通的途径是语音神经假体,它直接从皮层神经活动中解码语音。两个基准,Brain-to-Text '24和'25,发布了构音障碍患者的颅内记录,以及使用联结主义时间分类(CTC)训练的基线算法。尽管在这些基准测试上有重大创新,但所有领先的已发布工作都依赖于基于WFST的CTC解码器,需要${\sim}$320 GB的RAM。这些内存要求限制了患者和研究人员的可访问性。在这里,我们提出了LightBeam,一个基于非WFST的CTC解码器,只需要${\sim}$10 GB的RAM,并在两个基准测试中实现了最先进的性能。LightBeam通过延迟融合将LLM集成到波束搜索过程中来实现这一点,从而避免了之前使用大N-gram LM的需要。LightBeam是用Python实现的,是开源的。
摘要:A promising pathway for restoring communication in patients with dysarthria and anarthria is speech neuroprostheses, which directly decode speech from cortical neural activity. Two benchmarks, Brain-to-Text '24 and '25, released intracranial recordings from patients with dysarthria along with a baseline algorithm trained with Connectionist Temporal Classification (CTC). Despite significant innovation on these benchmarks, all leading published prior work relies on a WFST-based CTC decoder that requires ${\sim}$320 GB of RAM. These memory requirements limit accessibility for both patients and researchers. Here, we propose LightBeam, a non-WFST based CTC decoder that requires only ${\sim}$10 GB of RAM and achieves state-of-the-art performance on both benchmarks. LightBeam achieves this by integrating an LLM into the beam-search process via delayed fusion, obviating the prior need for using a large N-gram LM. LightBeam is implemented in Python and is open-source.
【20】LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
标题:用于视听语音增强的LLM引导强化学习
链接:https://arxiv.org/abs/2603.13952
备注:6 pages, 4 figures, submitted to Interspeech 2026
摘要:在现有的视听语音增强(AVSE)方法中,诸如尺度不变信噪比(SI-SNR)和均方误差(MSE)的目标被广泛使用;然而,它们通常与感知质量相关性差并且提供有限的可解释性用于优化。这项工作提出了一个基于强化学习的AVSE框架,该框架具有基于大型语言模型(LLM)的可解释奖励模型。音频LLM生成增强语音的自然语言描述,这些描述由情感分析模型转换为1-5的评级分数,作为用于微调预训练AVSE模型的PPO奖励。与标量度量相比,LLM生成的反馈在语义上是丰富的,并且明确地描述了语音质量的改进。在第四届COG-MHEAR AVSE挑战赛(AVSEC-4)数据集上的实验表明,该方法在PESQ,STOI,神经质量指标和主观听力测试中优于监督基线和基于DNSMOS的RL基线。
摘要:In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, they often correlate poorly with perceptual quality and provide limited interpretability for optimization. This work proposes a reinforcement learning-based AVSE framework with a Large Language Model (LLM)-based interpretable reward model. An audio LLM generates natural language descriptions of enhanced speech, which are converted by a sentiment analysis model into a 1-5 rating score serving as the PPO reward for fine-tuning a pretrained AVSE model. Compared with scalar metrics, LLM-generated feedback is semantically rich and explicitly describes improvements in speech quality. Experiments on the 4th COG-MHEAR AVSE Challenge (AVSEC-4) dataset show that the proposed method outperforms a supervised baseline and a DNSMOS-based RL baseline in PESQ, STOI, neural quality metrics, and subjective listening tests.
【21】Distributed Acoustic Sensing for Urban Traffic Monitoring: Spatio-Temporal Attention in Recurrent Neural Networks
标题:用于城市交通监测的分布式声学传感:回归神经网络中的时空注意力
链接:https://arxiv.org/abs/2603.13903
摘要:有效的城市交通监控对于改善流动性、提高安全性和支持可持续城市至关重要。分布式声学传感(DAS)通过将现有的光纤基础设施转换为密集的振动传感器阵列,实现了大规模的交通观测。然而,建模的DAS数据的高分辨率时空结构的可靠的交通事件识别仍然具有挑战性。这项研究提出了一个真实世界的基于DAS的交通监控实验在格拉纳达,西班牙,车辆穿过垂直于道路部署的光纤。递归神经网络(RNN)被用来建模事件内和事件间的时间依赖性。空间和时间注意力机制被系统地集成在RNN架构中,以分析它们对识别性能、参数效率和可解释性的影响。结果表明,一个适当的和互补的注意力模块的位置提高了准确性和模型复杂性之间的平衡。注意力热图通过突出信息丰富的空间位置和时间段,提供了对分类决策的物理上有意义的解释。此外,所提出的SA-双TA配置展示了空间可转移性,成功地识别交通事件在不同的传感位置在训练过程中使用的,只有适度的性能下降。这些研究结果支持开发可扩展和可解释的基于DAS的交通监测系统,能够在异构的城市传感条件下运行。
摘要:Effective urban traffic monitoring is essential for improving mobility, enhancing safety, and supporting sustainable cities. Distributed Acoustic Sensing (DAS) enables large-scale traffic observation by transforming existing fiber-optic infrastructure into dense arrays of vibration sensors. However, modeling the high-resolution spatio-temporal structure of DAS data for reliable traffic event recognition remains challenging. This study presents a real-world DAS-based traffic monitoring experiment conducted in Granada, Spain, where vehicles cross a fiber deployed perpendicular to the roadway. Recurrent neural networks (RNNs) are employed to model intra- and inter-event temporal dependencies. Spatial and temporal attention mechanisms are systematically integrated within the RNN architecture to analyze their impact on recognition performance, parameter efficiency, and interpretability. Results show that an appropriate and complementary placement of attention modules improves the balance between accuracy and model complexity. Attention heatmaps provide physically meaningful interpretations of classification decisions by highlighting informative spatial locations and temporal segments. Furthermore, the proposed SA-bi-TA configuration demonstrates spatial transferability, successfully recognizing traffic events at sensing locations different from those used during training, with only moderate performance degradation. These findings support the development of scalable and interpretable DAS-based traffic monitoring systems capable of operating under heterogeneous urban sensing conditions.
【22】Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs
标题:警报器的低语:语音驱动的LLM的听不见的近超声波越狱
链接:https://arxiv.org/abs/2603.13847
备注:USENIX Security'26 Camera-ready
摘要:语音驱动的大型语言模型(LLM)越来越多地通过语音接口访问,通过开放的声学通道引入了新的安全风险。我们提出了Sirens' Whisper(SWhisper),这是第一个实用的框架,用于在现实的黑盒条件下使用商品硬件对语音驱动的LLM进行基于隐蔽的攻击。SWhisper通过将任意目标基带音频编码为近超声波波形,在声学传输和麦克风非线性之后忠实地解调,从而在商品设备上实现稳健、听不见的传输,包括长而结构化的音频。这是通过一种简单而有效的方法来实现的,该方法可以跨设备和环境建模非线性信道特性,并结合轻量级信道反转预补偿。基于这种高保真隐蔽通道,我们设计了一种语音感知越狱生成方法,确保语音驱动界面下的可理解性,简洁性和可传输性。商业和开源语音驱动的LLM的实验都证明了强大的黑盒有效性。在商业模型上,SWhisper实现了高达0.94的非拒绝(NR)和0.925的特定令人信服(SC)。一项受控的用户研究进一步表明,注入的越狱音频在感知上与人类听众的仅背景播放无法区分。虽然越狱是一个案例研究,但底层的隐蔽声学通道可以实现更广泛的高保真注入攻击和命令扩展攻击。
摘要:Speech-driven large language models (LLMs) are increasingly accessed through speech interfaces, introducing new security risks via open acoustic channels. We present Sirens' Whisper (SWhisper), the first practical framework for covert prompt-based attacks against speech-driven LLMs under realistic black-box conditions using commodity hardware. SWhisper enables robust, inaudible delivery of arbitrary target baseband audio-including long and structured prompts-on commodity devices by encoding it into near-ultrasound waveforms that demodulate faithfully after acoustic transmission and microphone nonlinearity. This is achieved through a simple yet effective approach to modeling nonlinear channel characteristics across devices and environments, combined with lightweight channel-inversion pre-compensation. Building on this high-fidelity covert channel, we design a voice-aware jailbreak generation method that ensures intelligibility, brevity, and transferability under speech-driven interfaces. Experiments across both commercial and open-source speech-driven LLMs demonstrate strong black-box effectiveness. On commercial models, SWhisper achieves up to 0.94 non-refusal (NR) and 0.925 specific-convincing (SC). A controlled user study further shows that the injected jailbreak audio is perceptually indistinguishable from background-only playback for human listeners. Although jailbreaks serve as a case study, the underlying covert acoustic channel enables a broader class of high-fidelity prompt-injection and commandexecution attacks.
【23】Evaluating Semantic Fragility in Text-to-Audio Generation Systems Under Controlled Prompt Perturbations
标题:受控提示扰动下文本到音频生成系统中的语义脆弱性评估
链接:https://arxiv.org/abs/2603.13824
备注:8 pages, 4 figures, Under ICCC'26 review
摘要:文本到音频生成的最新进展使模型能够将自然语言描述转换为不同的音乐输出。然而,这些系统的鲁棒性下语义等价的提示变化仍然在很大程度上未被探索。微小的语言变化可能会导致生成的音频发生实质性变化,从而引发对实际使用中可靠性的担忧。 在这项研究中,我们评估的文本到音频系统的语义脆弱性控制提示扰动。我们选择了MusicGen-small、MusicGen-large和Stable Audio 2.5作为代表模型,并在最小词汇替换(MLS)、强度转换(IS)和结构改写(SR)下对它们进行了评估。建议的数据集包含75个提示组,旨在保留语义意图,同时引入本地化的语言变化。生成的输出通过互补的光谱,时间和语义相似性的措施进行比较,使跨多个代表性水平的鲁棒性分析。 实验结果表明,较大的模型实现了更好的语义一致性,MusicGen-large在MLS下达到0.77的余弦相似度,在IS下达到0.82。然而,声学和时间分析显示,持续的分歧在所有的模型,即使嵌入相似性仍然很高。这些研究结果表明,脆弱性主要出现在语义声学实现,而不是多模态嵌入对齐。我们的研究引入了一个控制框架,用于评估文本到音频生成的鲁棒性,并强调了生成音频系统中多级稳定性评估的必要性。
摘要:Recent advances in text-to-audio generation enable models to translate natural-language descriptions into diverse musical output. However, the robustness of these systems under semantically equivalent prompt variations remains largely unexplored. Small linguistic changes may lead to substantial variation in generated audio, raising concerns about reliability in practical use. In this study, we evaluate the semantic fragility of text-to-audio systems under controlled prompt perturbations. We selected MusicGen-small, MusicGen-large, and Stable Audio 2.5 as representative models, and we evaluated them under Minimal Lexical Substitution (MLS), Intensity Shifts (IS), and Structural Rephrasing (SR). The proposed dataset contains 75 prompt groups designed to preserve semantic intent while introducing localized linguistic variation. Generated outputs are compared through complementary spectral, temporal, and semantic similarity measures, enabling robustness analysis across multiple representational levels. Experimental results show that larger models achieve improved semantic consistency, with MusicGen-large reaching cosine similarities of 0.77 under MLS and 0.82 under IS. However, acoustic and temporal analyses reveal persistent divergence across all models, even when embedding similarity remains high. These findings indicate that fragility arises primarily during semantic-to-acoustic realization rather than multi-modal embedding alignment. Our study introduces a controlled framework for evaluating robustness in text-to-audio generation and highlights the need for multi-level stability assessment in generative audio systems.
【24】Causal Tracing of Audio-Text Fusion in Large Audio Language Models
标题:大型音频语言模型中音频文本融合的因果追踪
链接:https://arxiv.org/abs/2603.13768
备注:Submitted to Interspeech 2026
摘要:尽管大型音频语言模型(LALM)在各种任务中表现出色,但它们如何以及在何处将声学特征与文本上下文相结合仍不清楚。我们适应因果跟踪调查的内部信息流的LALM在音频理解。通过对DeSTA、Qwen和Voxtral进行逐层和逐标记分析,我们评估了各个隐藏状态的因果效应。逐层分析确定了不同的融合策略,从DeSTA中的渐进式融合到Qwen中的突然后期融合。令牌分析表明,最终的序列令牌作为一个信息瓶颈,网络决定性地从音频中检索相关信息。我们还在中间标记位置观察到类似注意力的查询机制,该机制触发模型提取与任务相关的音频上下文。这些发现清楚地描述了LAM内多模态集成发生的时间和地点。
摘要:Despite the strong performance of large audio language models (LALMs) in various tasks, exactly how and where they integrate acoustic features with textual context remains unclear. We adapt causal tracing to investigate the internal information flow of LALMs during audio comprehension. By conducting layer-wise and token-wise analyses across DeSTA, Qwen, and Voxtral, we evaluate the causal effects of individual hidden states. Layer-wise analysis identifies different fusion strategies, from progressive integration in DeSTA to abrupt late-stage fusion in Qwen. Token-wise analysis shows that the final sequence token acts as an informational bottleneck where the network decisively retrieves relevant information from the audio. We also observe an attention-like query mechanism at intermediate token positions that triggers the model to pull task-relevant audio context. These findings provide a clear characterization of when and where multi-modal integration occurs within LALMs.
【25】Multimodal Emotion Regression with Multi-Objective Optimization and VAD-Aware Audio Modeling for the 10th ABAW EMI Track
标题:第十届ABAW EMI曲目的多模式情感回归和多目标优化和VAR感知音频建模
链接:https://arxiv.org/abs/2603.13760
摘要:我们参加了第10届ABAW挑战赛,专注于Hume-Vidmimic 2数据集上的情绪模仿强度(EMI)估计。该任务旨在预测六个连续的情绪维度:钦佩,娱乐,决心,移情疼痛,兴奋和喜悦。通过对预训练的高级特征进行系统的多模态探索,我们发现,在我们的预训练特征设置下,直接特征串联的性能优于我们测试的更复杂的融合策略。这一实证发现促使我们设计了一种基于三个核心原则的系统方法:(i)通过特征级级联保留模态特定属性;(ii)通过多目标优化提高训练稳定性和度量对齐;以及(iii)用VAD启发的潜在先验丰富声学表示。我们的最终框架集成了基于级联的多模态融合,共享的六维回归头,多目标优化MSE,Pearson相关性和辅助分支监督,EMA用于参数稳定,以及VAD启发的声学分支潜在先验。在官方验证集上,该方案实现了我们的最佳平均Pearson相关系数0.478567。
摘要:We participated in the 10th ABAW Challenge, focusing on the Emotional Mimicry Intensity (EMI) Estimation track on the Hume-Vidmimic2 dataset. This task aims to predict six continuous emotion dimensions: Admiration, Amusement, Determination, Empathic Pain, Excitement, and Joy. Through systematic multimodal exploration of pretrained high-level features, we found that, under our pretrained feature setting, direct feature concatenation outperformed the more complex fusion strategies we tested. This empirical finding motivated us to design a systematic approach built upon three core principles: (i) preserving modality-specific attributes through feature-level concatenation; (ii) improving training stability and metric alignment via multi-objective optimization; and (iii) enriching acoustic representations with a VAD-inspired latent prior. Our final framework integrates concatenation-based multimodal fusion, a shared six-dimensional regression head, multi-objective optimization with MSE, Pearson-correlation, and auxiliary branch supervision, EMA for parameter stabilization, and a VAD-inspired latent prior for the acoustic branch. On the official validation set, the proposed scheme achieved our best mean Pearson Correlation Coefficient of 0.478567.
【26】Sub-Band Spectral Matching with Localized Score Aggregation for Robust Anomalous Sound Detection
标题:具有局部分数聚集的子带频谱匹配用于鲁棒异常声音检测
链接:https://arxiv.org/abs/2603.13749
备注:Manuscript under review
摘要:在嘈杂的声学环境中检测细微的偏差是异常声音检测(ASD)的核心。常见的免训练ASD流水线在时间上将帧级表示汇集到频带保留特征向量中,并使用单个最近邻匹配对异常进行评分。然而,这种全局匹配可能会通过两种效应扩大正常分数方差。首先,当正常声音表现出频带变化时,单个全局邻居迫使所有频带共享相同的参考,增加了频带电平失配。其次,基于余弦的匹配是能量耦合的,允许一些高能量带在正常能量波动下主导分数计算,并进一步增加方差。我们提出了BEAM,它存储在一个内存库的时间池的子带向量,检索每个子带的邻居,并均匀地聚合分数,以减少正常的分数变化和提高辨别力。我们进一步引入了一个无参数的自适应融合,以更好地处理不同的时间动态子带响应。在多个DCASE任务2基准测试上的实验显示出强大的性能,无需特定于任务的训练,对噪声和域偏移的鲁棒性,以及与编码器微调相结合时的互补增益。
摘要:Detecting subtle deviations in noisy acoustic environments is central to anomalous sound detection (ASD). A common training-free ASD pipeline temporally pools frame-level representations into a band-preserving feature vector and scores anomalies using a single nearest-neighbor match. However, this global matching can inflate normal-score variance through two effects. First, when normal sounds exhibit band-wise variability, a single global neighbor forces all bands to share the same reference, increasing band-level mismatch. Second, cosine-based matching is energy-coupled, allowing a few high-energy bands to dominate score computation under normal energy fluctuations and further increase variance. We propose BEAM, which stores temporally pooled sub-band vectors in a memory bank, retrieves neighbors per sub-band, and uniformly aggregates scores to reduce normal-score variability and improve discriminability. We further introduce a parameter-free adaptive fusion to better handle diverse temporal dynamics in sub-band responses. Experiments on multiple DCASE Task 2 benchmarks show strong performance without task-specific training, robustness to noise and domain shifts, and complementary gains when combined with encoder fine-tuning.
【27】$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
标题:$tau $-Voice:在现实世界域名上对全方位语音代理进行基准测试
链接:https://arxiv.org/abs/2603.13686
摘要:全双工语音代理--同时听和说的系统--正迅速从研究转向生产。然而,现有的评价孤立地处理对话动态和任务完成情况。我们引入$τ$-voice,一个评估语音代理在现实世界复杂性的基础任务上的基准:代理必须导航复杂的多回合对话,坚持域策略,并与环境交互。该框架将$τ^2$-bench扩展为一个新的语音代理基准测试,结合了复杂接地任务的可验证完成,全双工交互和逼真的音频-可以直接比较语音和文本性能。可控且逼真的语音用户模拟器提供了多样化的口音、逼真的音频环境和丰富的话轮转换动态;通过将模拟与挂钟时间解耦,用户模拟器可以在没有实时约束的情况下使用最强大的LLM。我们评估了278个任务的任务完成情况(pass@1)和语音交互质量:(推理)达到85%,语音代理在干净的条件下仅达到31- 51%,在有噪音和不同口音的现实条件下仅达到26- 38%-保留30- 45%的文本能力;定性分析证实,79- 90%的失败源于代理人的行为,这表明观察到的失败主要反映了我们的评估设置下的代理人的行为。$τ$-voice提供了一个可重复的测试平台,用于测量自然、会话和可靠的语音代理的进展。
摘要:Full-duplex voice agents--systems that listen and speak simultaneously--are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We introduce $τ$-voice, a benchmark for evaluating voice agents on grounded tasks with real-world complexity: agents must navigate complex multi-turn conversations, adhere to domain policies, and interact with the environment. The framework extends $τ^2$-bench into a novel voice agent benchmark combining verifiable completion of complex grounded tasks, full-duplex interaction, and realistic audio--enabling direct comparison between voice and text performance. A controllable and realistic voice user simulator provides diverse accents, realistic audio environments, and rich turn-taking dynamics; by decoupling simulation from wall-clock time, the user simulator can use the most capable LLM without real-time constraints. We evaluate task completion (pass@1) and voice interaction quality across 278 tasks: while GPT-5 (reasoning) achieves 85%, voice agents reach only 31--51% under clean conditions and 26--38% under realistic conditions with noise and diverse accents--retaining only 30--45% of text capability; qualitative analysis confirms 79--90% of failures stem from agent behavior, suggesting that observed failures primarily reflect agent behavior under our evaluation setup. $τ$-voice provides a reproducible testbed for measuring progress toward voice agents that are natural, conversational, and reliable.
【28】Evaluating Compositional Structure in Audio Representations
标题:评估音频表示中的成分结构
链接:https://arxiv.org/abs/2603.13685
备注:Accepted to ICASSP 2026
摘要:我们提出了一个基准评估音频表示的组合性。音频合成性是指根据组成源和属性来表示声音场景,并将它们系统地组合在一起。虽然中央听觉感知,这一属性是在很大程度上缺席目前的评估协议。我们的框架通过两个任务来适应从视觉和语言到音频的想法:A-COAT,它测试加法变换下的一致性,以及A-TRE,它探测属性级原语的可重构性。这两项任务都得到了大型合成数据集的支持,这些数据集具有声学属性的受控变化,提供了音频嵌入中组成结构的第一个基准。
摘要:We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attributes, and combining them systematically. While central to auditory perception, this property is largely absent from current evaluation protocols. Our framework adapts ideas from vision and language to audio through two tasks: A-COAT, which tests consistency under additive transformations, and A-TRE, which probes reconstructibility from attribute-level primitives. Both tasks are supported by large synthetic datasets with controlled variation in acoustic attributes, providing the first benchmark of compositional structure in audio embeddings.
【29】A Hierarchical End-of-Turn Model with Primary Speaker Segmentation for Real-Time Conversational AI
标题:用于实时对话人工智能的具有主要发言人分割的分层回合结束模型
链接:https://arxiv.org/abs/2603.13379
备注:Accepted for presentation at the IEEE Conference on Artificial Intelligence
摘要:我们提出了一个基于语音的会话AI的实时前端,通过将主说话人分割与分层结束回合(EOT)检测相结合,在两个说话人的场景中实现自然的话轮转换。为了在多说话者环境中稳健地操作,系统连续地识别和跟踪主要用户,确保下游EOT决策不被背景对话混淆。跟踪的活动片段被馈送到分层的因果EOT模型,该模型通过独立地分析来自主要说话者和机器人的每个说话者的语音特征来预测即时会话状态。同时,该模型预计不久的将来的状态($t{+}10/20/30 $\,ms)通过概率预测,知道对话伙伴的讲话。特定任务的知识蒸馏将wav 2 vec ~2.0表示(768\,D)压缩为紧凑的基于MFCC的学生(32\,D)以进行有效部署。该系统在多类帧级F1为82%,在反向检测时F1为70.6%,在二进制帧级F1为69.3%。其他任务。在端到端的转向检测基准测试中,我们的模型达到了87.7%的召回率,Smart Turn~v3的平均检测延迟为58.9\%,同时保持36\,ms的中值检测延迟,而不是800- 1300\,ms。尽管仅使用1.14\,M参数,但所提出的模型匹配或超过基于变压器的基线,同时大幅降低延迟和内存占用,使其适合边缘部署。
摘要:We present a real-time front-end for voice-based conversational AI to enable natural turn-taking in two-speaker scenarios by combining primary speaker segmentation with hierarchical End-of-Turn (EOT) detection. To operate robustly in multi-speaker environments, the system continuously identifies and tracks the primary user, ensuring that downstream EOT decisions are not confounded by background conversations. The tracked activity segments are fed to a hierarchical, causal EOT model that predicts the immediate conversational state by independently analyzing per-speaker speech features from both the primary speaker and the bot. Simultaneously, the model anticipates near-future states ($t{+}10/20/30$\,ms) through probabilistic predictions that are aware of the conversation partner's speech. Task-specific knowledge distillation compresses wav2vec~2.0 representations (768\,D) into a compact MFCC-based student (32\,D) for efficient deployment. The system achieves 82\% multi-class frame-level F1 and 70.6\% F1 on Backchannel detection, with 69.3\% F1 on a binary Final vs.\ Others task. On an end-to-end turn-detection benchmark, our model reaches 87.7\% recall vs.\ 58.9\% for Smart Turn~v3 while keeping a median detection latency of 36\,ms versus 800--1300\,ms. Despite using only 1.14\,M parameters, the proposed model matches or exceeds transformer-based baselines while substantially reducing latency and memory footprint, making it suitable for edge deployment.
【30】Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings
标题:来自多部位听诊记录的患者级多模式问题回答
链接:https://arxiv.org/abs/2603.13362
摘要:听诊是一种重要的诊断工具,但其实用性往往受到主观解释的限制。虽然通用的音频语言模型(ALM)在一般领域表现出色,但它们在生理信号的细微差别方面表现不佳。我们提出了一个框架,通过门控交叉注意直接将多站点听诊记录与冻结的大语言模型(LLM)嵌入空间对齐。通过利用LLM的潜在世界知识,我们的方法超越了孤立的分类,走向全面的,患者级的评估。在CaReSound基准测试中,我们的模型达到了最先进的0.865 F1宏和0.952 BERTScore。我们证明,轻量级的,特定于域的编码器的竞争对手大规模的ALM和多站点聚合提供空间冗余,减轻时间截断。医学声学与文本基础的这种对齐为桥接信号处理和临床评估提供了可扩展的路径。
摘要:Auscultation is a vital diagnostic tool, yet its utility is often limited by subjective interpretation. While general-purpose Audio-Language Models (ALMs) excel in general domains, they struggle with the nuances of physiological signals. We propose a framework that aligns multi-site auscultation recordings directly with a frozen Large Language Model (LLM) embedding space via gated cross-attention. By leveraging the LLM's latent world knowledge, our approach moves beyond isolated classification toward holistic, patient-level assessment. On the CaReSound benchmark, our model achieves a state-of-the-art 0.865 F1-macro and 0.952 BERTScore. We demonstrate that lightweight, domain-specific encoders rival large-scale ALMs and that multi-site aggregation provides spatial redundancy that mitigates temporal truncation. This alignment of medical acoustics with text foundations offers a scalable path for bridging signal processing and clinical assessment.
【31】Evaluation of Audio Language Models for Fairness, Safety, and Security
标题:公平性、安全性和保障性的音频语言模型评估
链接:https://arxiv.org/abs/2603.13262
摘要:音频大语言模型(ALLM)最近通过将语音处理与大语言模型相结合来推进口语交互。然而,现有的公平性,安全性和安全性(FSS)的评估仍然是支离破碎的,主要是因为ALLM在声学信息的表示方式和语义推理发生的地方有根本的不同。这些差异很少被明确表达出来。因此,评估往往合并结构不同的系统,模糊模型设计和观察到的FSS行为之间的关系。在这项工作中,我们介绍了ALLM的结构分类法(系统级和代表性),该分类法沿着两个轴对系统进行分类:音频输入表示的形式(例如,离散对连续)和语义推理的轨迹(例如,级联、多模式或音频原生)。在分类的基础上,我们提出了一个统一的评估框架,该框架评估了语义不变性下的语言变化,拒绝和毒性行为下的不安全提示,以及对抗性音频扰动的鲁棒性。我们将这个框架应用到两个有代表性的系统中,观察到音频和文本输入之间在拒绝率、攻击成功率和毒性方面的系统差异。我们的研究结果表明,FSS的行为是紧密耦合的声学信息是如何集成到语义推理,强调需要结构感知评估的音频语言模型。
摘要:Audio large language models (ALLMs) have recently advanced spoken interaction by integrating speech processing with large language models. However, existing evaluations of fairness, safety, and security (FSS) remain fragmented, largely because ALLMs differ fundamentally in how acoustic information is represented and where semantic reasoning occurs. Differences that are rarely made explicit. As a result, evaluations often conflate structurally distinct systems, obscuring the relationship between model design and observed FSS behavior. In this work, we introduce a structural taxonomy (system-level and representational) of ALLMs that categorizes systems along two axes: the form of audio input representation (e.g., discrete vs. continuous) and the locus of semantic reasoning (e.g., cascaded, multimodal, or audio-native). Building on the taxonomy, we propose a unified evaluation framework that assesses semantic invariance under paralinguistic variation, refusal and toxicity behavior under unsafe prompts, and robustness to adversarial audio perturbations. We apply this framework to two representative systems and observe systematic differences in refusal rates, attack success, and toxicity between audio and text inputs. Our findings demonstrate that FSS behavior is tightly coupled to how acoustic information is integrated into semantic reasoning, underscoring the need for structure-aware evaluation of audio language models.
【32】Controllable Accent Normalization via Discrete Diffusion
标题:通过离散扩散实现可控口音规范化
链接:https://arxiv.org/abs/2603.14275
备注:Submitted for review to Interspeech 2026
摘要:现有的口音规范化方法通常不提供对口音强度的控制,但许多应用程序(如语言学习和配音)需要可调的口音保留。我们提出了DLM-AN,一个可控的口音规范化系统,建立在自监督语音令牌上的掩蔽离散扩散。一个共同的令牌预测识别源令牌,可能编码母语发音;这些令牌有选择地重复使用,以初始化反向扩散过程。这为控制重音强度提供了一种简单而有效的机制:重用更多的标记可以保留更多的原始重音。DLM-AN还集成了一个流量匹配持续时间比预测器,可以自动调整总持续时间,以更好地匹配原生节律。在多口音英语数据上的实验表明,DLM-AN在所有比较系统中实现了最低的单词错误率,同时提供了有竞争力的口音减少和平滑,可解释的口音强度控制。
摘要:Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. A Common Token Predictor identifies source tokens that likely encode native pronunciation; these tokens are selectively reused to initialize the reverse diffusion process. This provides a simple yet effective mechanism for controlling accent strength: reusing more tokens preserves more of the original accent. DLM-AN further incorporates a flow-matching Duration Ratio Predictor that automatically adjusts the total duration to better match the native rhythm. Experiments on multi-accent English data show that DLM-AN achieves the lowest word error rate among all compared systems while delivering competitive accent reduction and smooth, interpretable accent strength control.
【33】Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
标题:通过三级公式和TLR进行集成欺骗稳健的自动说话人验证
链接:https://arxiv.org/abs/2603.13780
备注:Submitted to Interspeech 2026; put on arxiv based on requirement from Interspeech: "Interspeech no longer enforces an anonymity period for submissions." and "For authors that prefer to upload their paper online, a note indicating that the paper was submitted for review to Interspeech should be included in the posting."
摘要:抗欺骗的自动说话人确认(SASV)旨在将自动说话人确认(ASV)和对抗(CM)相结合。一个流行的解决方案是融合独立的ASV和CM分数。为了更好地建模SASV,一些框架将ASV和CM集成在单个网络中。然而,这些解决方案通常是基于双编码器的,提供有限的可解释性,并且在没有重新训练的情况下不能容易地适应新的评估参数。在此基础上,我们提出了一个统一的端到端的框架,通过三个类的配方,使对数似然比(LLR)推断类logits更可解释的决策管道。实验结果表明,与ASVSpoof5上的现有方法相比,SpoofCeleb上的结果更好。可视化和分析也证明了三类重构提供了更好的可解释性。
摘要:Spoofing-robust automatic speaker verification (SASV) aims to integrate automatic speaker verification (ASV) and countermeasure (CM). A popular solution is fusion of independent ASV and CM scores. To better modeling SASV, some frameworks integrate ASV and CM within a single network. However, these solutions are typically bi-encoder based, offer limited interpretability, and cannot be readily adapted to new evaluation parameters without retraining. Based on this, we propose a unified end-to-end framework via a three-class formulation that enables log-likelihood ratio (LLR) inference from class logits for a more interpretable decision pipeline. Experiments show comparable performance to existing methods on ASVSpoof5 and better results on SpoofCeleb. The visualization and analysis also prove that the three-class reformulation provides more interpretability.
【34】VoXtream2: Full-stream TTS with dynamic speaking rate control
标题:VoXtream 2:具有动态语速控制的全流RTS
链接:https://arxiv.org/abs/2603.13518
备注:10 pages, 9 figures, Submitted to Interspeech 2026
摘要:交互式系统的全流文本到语音(TTS)必须以最小的延迟开始说话,同时随着文本的增量到达而保持可控。我们提出了VoXtream2,一个zero-shot全码流TTS模型与动态说话速率控制,可以更新中间话语的飞行。VoXtream2将持续时间状态的分布匹配机制与调节信号的无分类器指导相结合,以提高可控性和合成质量。非文本屏蔽启用非文本音频提示,从而消除了对提示转录的需要。在标准的zero-shot基准测试和专用的说话速率测试集上,VoXtream2实现了与公共基线相比具有竞争力的客观和主观结果,尽管模型较小,训练数据较少。在全流模式下,它在消费级GPU上的运行速度比实时快4倍,首包延迟为74 ms。
摘要:Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a zero-shot full-stream TTS model with dynamic speaking-rate control that can be updated mid-utterance on the fly. VoXtream2 combines a distribution matching mechanism over duration states with classifier-free guidance across conditioning signals to improve controllability and synthesis quality. Prompt-text masking enables textless audio prompting, removing the need for prompt transcription. Across standard zero-shot benchmarks and a dedicated speaking-rate test set, VoXtream2 achieves competitive objective and subjective results against public baselines despite a smaller model and less training data. In full-stream mode, it runs 4 times faster than real time with 74 ms first-packet latency on a consumer GPU.
【35】BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
标题:BrainWhisperer:利用大规模ASR模型进行神经语音解码
链接:https://arxiv.org/abs/2603.13321
摘要:从皮质内记录中解码连续语音是脑机接口(BCI)的一个核心挑战,对于那些说话能力受损的人来说,这具有变革的潜力。虽然最近的微电极阵列(MEA)解码器实现了令人印象深刻的准确性,但它们的性能从根本上受到现有数据集的小尺寸的限制,它们仍然容易受到会话间变异性的影响,并且它们在参与者之间进行概括的能力仍未得到探索。我们介绍BrainWhisperer,一种神经语音解码器,它将高分辨率MEA记录与大型预训练自动语音识别(ASR)模型集成在一起。基于可解释性研究结果表明,Whisper的编码器学习音素选择性表示与本地化的注意,我们训练一个定制版本的Whisper,修改处理神经功能,使用一个混合目标,结合CTC损失音素-预测从第三编码器层-和交叉熵损失的单词令牌。我们引入了领域信息修改,包括窗口化自我注意力以捕获发音连续性,分层月/日特定的低秩预测以解决非平稳性,以及特定于主题的嵌入器,从而实现跨主题培训。在公开可用的MEA数据集上进行评价(Card等人),BrainWhisperer匹配或优于先前的最先进的解码器。至关重要的是,跨数据集训练即使在没有微调的情况下也能提高单个数据集的性能,表现出前所未有的泛化能力。该模型支持双重解码路径:基于音素的高精度路径,具有外部语言模型重新评分,以及快速直接文本生成路径,能够以最小的硬件要求进行低于100 ms的推理。
摘要:Decoding continuous speech from intracortical recordings is a central challenge for brain-computer interfaces (BCIs), with transformative potential for individuals with conditions that impair their ability to speak. While recent microelectrode array (MEA) decoders achieve impressive accuracy, their performance is fundamentally limited by the small size of existing datasets, they remain brittle to session-to-session variability, and their ability to generalize across participants remains unexplored. We introduce BrainWhisperer, a neural speech decoder that integrates high-resolution MEA recordings with a large pretrained automatic speech recognition (ASR) model. Building on interpretability findings showing that Whisper's encoder learns phoneme-selective representations with localized attention, we train a customized version of Whisper, modified to process neural features, using a hybrid objective that combines CTC loss on phonemes--predicted from the third encoder layer--and cross-entropy loss on word tokens. We introduce domain-informed modifications including windowed self-attention to capture articulatory continuity, hierarchical month/day-specific low-rank projections to address non-stationarity, and subject-specific embedders enabling cross-subject training. Evaluated on a publicly available MEA dataset (Card et al.), BrainWhisperer matches or outperforms prior state-of-the-art decoders. Critically, cross-dataset training improves performance even on individual datasets without fine-tuning, demonstrating unprecedented generalization. The model supports dual decoding paths: a high-accuracy phoneme-based path with external language model rescoring, and a fast direct text generation path enabling sub-100ms inference with minimal hardware requirements.
【1】spINAch: A Diachronic Corpus of French Broadcast Speech Controlled for Speakers' Age and Gender
标题:spINAch:针对演讲者年龄和性别进行控制的法语广播语音历时数据库
链接:https://arxiv.org/abs/2603.15516
备注:16 pages, 3 figures, to be published in the Fifteenth International Conference on Language Resources and Evaluation (LREC 2026)
摘要:我们介绍了spINAch,这是一个来自广播和电视档案的法语演讲的大型历时语料库,按演讲者的性别,年龄(20-95岁)进行平衡,从1955年到2015年跨越60年。该数据集包括来自2000多名演讲者的320多小时录音。建立语料库的方法进行了说明,集中在声学方面收集的样本的质量。这些数据被自动转录和语音对齐,以便在音素水平上进行研究。超过300万的口语元音已经被分析,提出了他们的基频和共振峰。该语料库可供社区用于研究目的,对于通过性别和年龄的代表来描述巴黎法语的演变是有价值的。所提出的分析还表明,语料库的历时性质允许观察各种语音现象,如语音音高随时间的演变(这并不因性别而异,在我们的数据)和中和/a/-/$a$/反对在巴黎法语在此期间。
摘要:We present spINAch, a large diachronic corpus of French speech from radio and television archives, balanced by speakers' gender, age (20-95 years old), and spanning 60 years from 1955 to 2015. The dataset includes over 320 hours of recordings from more than two thousand speakers. The methodology for building the corpus is described, focusing on the quality of collected samples in acoustic terms. The data were automatically transcribed and phonetically aligned to allow studies at a phonemic level. More than 3 million oral vowels have been analyzed to propose their fundamental frequency and formants. The corpus, available to the community for research purposes, is valuable for describing the evolution of Parisian French through the representation of gender and age. The presented analyses also demonstrate that the diachronic nature of the corpus allows the observation of various phonetic phenomena, such as the evolution of voice pitch over time (which does not differ by gender in our data) and the neutralization of the /a/-/$a$/ opposition in Parisian French during this period.
【2】Neural Network-Based Time-Frequency-Bin-Wise Linear Combination of Beamformers for Underdetermined Target Source Extraction
标题:基于神经网络的时间-频率-逐bin线性组合束形成器用于欠定目标源提取
链接:https://arxiv.org/abs/2603.15288
备注:Accepted by ICASSP 2026
摘要:从欠定混合信号中提取目标源是波束形成方法的一个挑战。最近提出的时频分仓切换(TFS)和线性组合(TFLC)策略通过在每个时频(TF)仓中组合多个波束形成器并选择使输出功率最小化的组合权重来减轻这种情况。然而,为每个TF仓独立地做出该决定可能会削弱时间-频谱相干性,导致不连续性并因此降低提取性能。在本文中,我们提出了一种新的基于神经网络的时频分线性组合(NN-TFLC)框架,构建最小功率无失真响应(MPDR)波束形成器没有显式的噪声协方差估计。该网络对混合和波束形成器输出进行编码,并通过交叉注意机制预测时间和频谱相干线性组合权重。在具有多个波束形成器的双麦克风混合中,NN-TFLC-MPDR始终优于TFS/TFLC-MPDR,并与基于最小方差无失真响应(MVDR)波束形成器构建的TFS/TFLC实现竞争性性能,需要噪声先验。
摘要:Extracting a target source from underdetermined mixtures is challenging for beamforming approaches. Recently proposed time-frequency-bin-wise switching (TFS) and linear combination (TFLC) strategies mitigate this by combining multiple beamformers in each time-frequency (TF) bin and choosing combination weights that minimize the output power. However, making this decision independently for each TF bin can weaken temporal-spectral coherence, causing discontinuities and consequently degrading extraction performance. In this paper, we propose a novel neural network-based time-frequency-bin-wise linear combination (NN-TFLC) framework that constructs minimum power distortionless response (MPDR) beamformers without explicit noise covariance estimation. The network encodes the mixture and beamformer outputs, and predicts temporally and spectrally coherent linear combination weights via a cross-attention mechanism. On dual-microphone mixtures with multiple interferers, NN-TFLC-MPDR consistently outperforms TFS/TFLC-MPDR and achieves competitive performance with TFS/TFLC built on the minimum variance distortionless response (MVDR) beamformers that require noise priors.
【3】How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
标题:注意力如何影响情感:言语情感识别注意力机制的比较研究
链接:https://arxiv.org/abs/2603.15120
摘要:语音情感识别在促进人机交互中起着关键作用。注意机制由于其能够捕捉长距离依赖关系和强调显著信息的能力,已经成为建模情感语音的主要方法。然而,标准的自我关注遭受二次计算和存储器的复杂性,限制了其可扩展性。在这项工作中,我们提出了一个系统的基准优化的注意力机制SER,包括RetNet,LightNet,GSA,FOX和KDA。在两个MSP-Podcast基准测试版本上的实验表明,虽然标准的自我注意力在测试集上实现了最强的识别性能,但有效的注意力变体显着提高了可扩展性,将推理延迟和内存使用量降低了一个数量级。这些结果突出了准确性和效率之间的关键权衡,为设计可扩展的SER系统提供了实用的见解。
摘要:Speech Emotion Recognition (SER) plays a key role in advancing human-computer interaction. Attention mechanisms have become the dominant approach for modeling emotional speech due to their ability to capture long-range dependencies and emphasize salient information. However, standard self-attention suffers from quadratic computational and memory complexity, limiting its scalability. In this work, we present a systematic benchmark of optimized attention mechanisms for SER, including RetNet, LightNet, GSA, FoX, and KDA. Experiments on both MSP-Podcast benchmark versions show that while standard self-attention achieves the strongest recognition performance across test sets, efficient attention variants dramatically improve scalability, reducing inference latency and memory usage by up to an order of magnitude. These results highlight a critical trade-off between accuracy and efficiency, providing practical insights for designing scalable SER systems.
【4】LLMs and Speech: Integration vs. Combination
标题:法学硕士和演讲:融合与结合
链接:https://arxiv.org/abs/2603.15045
备注:Submitted to Interspeech 2026
摘要:在这项工作中,我们研究如何最好地利用预训练的LLM自动语音识别。具体来说,我们比较了声学模型(AM)与LLM(“语音LLM”)的紧密集成与通过浅层融合将AM和LLM相结合的传统方式。为了紧密集成,我们提供了不同标签单元的效果,微调策略,LLM大小和预训练数据,注意力界面,编码器下采样,文本提示和长度归一化。此外,我们研究联合识别与CTC模型,以减轻幻觉的语音LLM和目前有效的优化,这种联合识别。对于浅融合,我们研究了微调LLM对使用不同标签单元的transmittance的影响,并将AM假设重新评分与AM和LLM分数的标签或延迟融合的单次识别进行比较。我们对Librispeech和Loquacious进行训练,并在HuggingFace ASR排行榜上评估我们的模型。
摘要:In this work, we study how to best utilize pre-trained LLMs for automatic speech recognition. Specifically, we compare the tight integration of an acoustic model (AM) with the LLM ("speech LLM") to the traditional way of combining AM and LLM via shallow fusion. For tight integration, we provide ablations on the effect of different label units, fine-tuning strategies, LLM sizes and pre-training data, attention interfaces, encoder downsampling, text prompts, and length normalization. Additionally, we investigate joint recognition with a CTC model to mitigate hallucinations of speech LLMs and present effective optimizations for this joint recognition. For shallow fusion, we investigate the effect of fine-tuning the LLM on the transcriptions using different label units, and we compare rescoring AM hypotheses to single-pass recognition with label-wise or delayed fusion of AM and LLM scores. We train on Librispeech and Loquacious and evaluate our models on the HuggingFace ASR leaderboard.
【5】Deep Filter Estimation from Inter-Frame Correlations for Monaural Speech Dereverberation
标题:基于帧间相关性的深度过滤器估计单耳语音去回响
链接:https://arxiv.org/abs/2603.14986
备注:Submitted for review to Interspeech
摘要:由于混响和目标信号之间的高度相关性,远距离麦克风场景中的语音去混响仍然具有挑战性,通常导致在现实环境中的泛化能力差。我们提出了IF-CorrNet,一个相关滤波器架构,旨在对声学变化的鲁棒性。与直接估计复频谱的传统黑盒映射方法不同,IF-CorrNet明确利用帧间STFT相关性来估计每个时频点的多帧深度滤波器。通过将学习目标从直接映射转移到滤波器估计,网络有效地约束了解空间,从而简化了训练过程并减轻了对合成数据的过拟合。在REVERB Challenge数据集上的实验结果表明,IF-CorrNet在RealData上的SRMR度量中实现了实质性的增益,证实了其在实际非合成环境中抑制混响和噪声的鲁棒性。
摘要:Speech dereverberation in distant-microphone scenarios remains challenging due to the high correlation between reverberation and target signals, often leading to poor generalization in real-world environments. We propose IF-CorrNet, a correlation-to-filter architecture designed for robustness against acoustic variability. Unlike conventional black-box mapping methods that directly estimate complex spectra, IF-CorrNet explicitly exploits inter-frame STFT correlations to estimate multi-frame deep filters for each time-frequency bin. By shifting the learning objective from direct mapping to filter estimation, the network effectively constrains the solution space, which simplifies the training process and mitigates overfitting to synthetic data. Experimental results on the REVERB Challenge dataset demonstrate that IF-CorrNet achieves a substantial gain in the SRMR metric on RealData, confirming its robustness in suppressing reverberation and noise in practical, non-synthetic environments.
【6】Spectrogram features for audio and speech analysis
标题:用于音频和语音分析的频谱图功能
链接:https://arxiv.org/abs/2603.14917
备注:30 pages
摘要:基于频谱图的表示已经成为深度学习音频分析系统的主要特征空间,并且经常被用于语音分析。最初,基于频谱图的表示的主要动力是它们能够将声音呈现为时频平面中的二维信号,这不仅为分析声音提供了可解释的物理基础,而且还解锁了广泛的机器学习技术的使用,例如卷积神经网络,这些技术已被开发用于图像处理。频谱图是一个矩阵,其特征在于其二维的分辨率和跨度,以及每个元素的表示和缩放。研究人员在许多应用领域探索了这三个特征的许多可能性,不同的设置显示出对各种任务的亲和力。本文回顾了基于谱图的表示的使用,并调查了最新的问题,前端特征表示选择如何与后端分类器架构结盟,以完成不同的任务。
摘要:Spectrogram-based representations have grown to dominate the feature space for deep learning audio analysis systems, and are often adopted for speech analysis also. Initially, the primary motivator for spectrogram-based representations was their ability to present sound as a two dimensional signal in the time-frequency plane, which not only provides an interpretable physical basis for analysing sound, but also unlocks the use of a wide range of machine learning techniques such as convolutional neural networks, that had been developed for image processing. A spectrogram is a matrix characterised by the resolution and span of its two dimensions, as well as by the representation and scaling of each element. Many possibilities for these three characteristics have been explored by researchers across numerous application areas, with different settings showing affinity for various tasks. This paper reviews the use of spectrogram-based representations and surveys the state-of-the-art to question how front-end feature representation choice allies with back-end classifier architecture for different tasks.
【7】Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
标题:建模和基准口语对话具有情态和口语性的奖励
链接:https://arxiv.org/abs/2603.14889
摘要:端到端口语对话系统的快速发展要求超越单纯的文本语义,以纳入语言的细微差别和人类对话的自发性。然而,目前的方法与两个关键的差距斗争:模态差距,涉及韵律和情感,口语化差距,区分书面文字从自然语音。为了解决这些挑战,我们引入了SDiaReward,这是一个在SDiaReward-Dataset上训练的端到端多回合奖励模型,SDiaReward是一个明确针对这些差距的情节级偏好对的新集合。它直接对完整的多轮语音片段进行操作,并通过成对偏好监督进行优化,从而能够在单个评估器中对模态和口语进行联合评估。我们进一步建立了ESDR-Bench,这是一个用于稳健的剧集级别评估的分层基准。实验表明,SDiaReward实现了最先进的成对偏好准确性,显著优于通用音频LLM。进一步的分析表明,SDiaReward捕捉相对会话表达超越表面的合成线索,提高跨域和记录条件的概括。代码、数据和演示可在https://sdiareward.github.io/上获得。
摘要:The rapid evolution of end-to-end spoken dialogue systems demands transcending mere textual semantics to incorporate paralinguistic nuances and the spontaneous nature of human conversation. However, current methods struggle with two critical gaps: the modality gap, involving prosody and emotion, and the colloquialness gap, distinguishing written scripts from natural speech. To address these challenges, we introduce SDiaReward, an end-to-end multi-turn reward model trained on SDiaReward-Dataset, a novel collection of episode-level preference pairs explicitly targeting these gaps. It operates directly on full multi-turn speech episodes and is optimized with pairwise preference supervision, enabling joint assessment of modality and colloquialness in a single evaluator. We further establish ESDR-Bench, a stratified benchmark for robust episode-level evaluation. Experiments demonstrate that SDiaReward achieves state-of-the-art pairwise preference accuracy, significantly outperforming general-purpose audio LLMs. Further analysis suggests that SDiaReward captures relative conversational expressiveness beyond superficial synthesis cues, improving generalization across domains and recording conditions. Code, data, and demos are available at https://sdiareward.github.io/.
【8】SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation
标题:SoulX-Dupplug:用于实时高保真语音对话的即插即用流媒体状态预测模块
链接:https://arxiv.org/abs/2603.14877
备注:submitted to Interspeech 2026, under review
摘要:口语对话系统的最新进展使得人们越来越关注类似人类的全双工语音交互。然而,我们对这一领域的全面回顾揭示了一些挑战,包括难以获得训练数据、灾难性遗忘和有限的可扩展性。在这项工作中,我们提出了SoulX-Duplug,即插即用流状态预测模块全双工口语对话系统。通过联合执行流式ASR,SoulX-Duplug显式地利用文本信息来识别用户意图,有效地充当语义VAD。为了促进公平的评估,我们引入了SoulX-Duplug-Eval,扩展了广泛使用的基准,提高了双语覆盖率。实验结果表明,SoulX-Duplug能够实现低延迟流媒体对话状态控制,基于它构建的系统在整体回合管理和延迟性能方面优于现有全双工模型。我们有开源的SoulX-Duplug和SoulX-Duplug-Eval。
摘要:Recent advances in spoken dialogue systems have brought increased attention to human-like full-duplex voice interactions. However, our comprehensive review of this field reveals several challenges, including the difficulty in obtaining training data, catastrophic forgetting, and limited scalability. In this work, we propose SoulX-Duplug, a plug-and-play streaming state prediction module for full-duplex spoken dialogue systems. By jointly performing streaming ASR, SoulX-Duplug explicitly leverages textual information to identify user intent, effectively serving as a semantic VAD. To promote fair evaluation, we introduce SoulX-Duplug-Eval, extending widely used benchmarks with improved bilingual coverage. Experimental results show that SoulX-Duplug enables low-latency streaming dialogue state control, and the system built upon it outperforms existing full-duplex models in overall turn management and latency performance. We have open-sourced SoulX-Duplug and SoulX-Duplug-Eval.
【9】Controllable Accent Normalization via Discrete Diffusion
标题:通过离散扩散实现可控口音规范化
链接:https://arxiv.org/abs/2603.14275
备注:Submitted for review to Interspeech 2026
摘要:现有的口音规范化方法通常不提供对口音强度的控制,但许多应用程序(如语言学习和配音)需要可调的口音保留。我们提出了DLM-AN,一个可控的口音规范化系统,建立在自监督语音令牌上的掩蔽离散扩散。一个共同的令牌预测识别源令牌,可能编码母语发音;这些令牌有选择地重复使用,以初始化反向扩散过程。这为控制重音强度提供了一种简单而有效的机制:重用更多的标记可以保留更多的原始重音。DLM-AN还集成了一个流量匹配持续时间比预测器,可以自动调整总持续时间,以更好地匹配原生节律。在多口音英语数据上的实验表明,DLM-AN在所有比较系统中实现了最低的单词错误率,同时提供了有竞争力的口音减少和平滑,可解释的口音强度控制。
摘要:Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. A Common Token Predictor identifies source tokens that likely encode native pronunciation; these tokens are selectively reused to initialize the reverse diffusion process. This provides a simple yet effective mechanism for controlling accent strength: reusing more tokens preserves more of the original accent. DLM-AN further incorporates a flow-matching Duration Ratio Predictor that automatically adjusts the total duration to better match the native rhythm. Experiments on multi-accent English data show that DLM-AN achieves the lowest word error rate among all compared systems while delivering competitive accent reduction and smooth, interpretable accent strength control.
【10】Beyond Two-stage Diffusion TTS: Joint Structure and Content Refinement via Jump Diffusion
标题:超越两阶段扩散TTC:通过跳跃扩散的联合结构和内容细化
链接:https://arxiv.org/abs/2603.14032
备注:5 pages, 5 figures. Audio samples available at https://anonymousinterpseech.github.io/TTS_Demo/
摘要:扩散和流动匹配TTS面临着离散时间结构和连续频谱建模之间的紧张关系。两阶段模型在固定路线上扩散,通常会崩溃为平均韵律;单阶段模型避免明确的持续时间,但会遭受路线不稳定性。我们提出了一个跳跃扩散框架,离散跳跃模型的时间结构和连续扩散细化一个过程中的频谱内容。即使在其一次性退化形式中,我们的框架在LJSpeech上使用改进的UTMOSv 2实现了3.37%的WER与4.38%的Grad-TTS。完整的迭代UDD变体进一步实现了自适应韵律,在分布外的慢速语音中自动插入自然停顿,而不是均匀地拉伸。音频样本可在https://anonymousinterpseech.github.io/TTS_Demo/上获得。
摘要:Diffusion and flow matching TTS faces a tension between discrete temporal structure and continuous spectral modeling. Two-stage models diffuse on fixed alignments, often collapsing to mean prosody; single-stage models avoid explicit durations but suffer alignment instability. We propose a jump-diffusion framework where discrete jumps model temporal structure and continuous diffusion refines spectral content within one process. Even in its one-shot degenerate form, our framework achieves 3.37% WER vs. 4.38% for Grad-TTS with improved UTMOSv2 on LJSpeech. The full iterative UDD variant further enables adaptive prosody, autonomously inserting natural pauses in out-of-distribution slow speech rather than stretching uniformly. Audio samples are available at https://anonymousinterpseech.github.io/TTS_Demo/.
【11】Evaluating Pretrained General-Purpose Audio Representations for Music Genre Classification
标题:评估预训练的通用音频表示以进行音乐流派分类
链接:https://arxiv.org/abs/2603.13871
备注:Accepted and presented at the International Conference on Pattern Recognition and Machine Intelligence (PReMI), 2025
摘要:这项研究调查了自我监督学习嵌入的使用,特别是BYOL-A,结合深度神经网络分类器进行音乐流派分类。我们的实验表明,BYOL-A嵌入优于其他预训练模型,如PANN和VGGish,在GTZAN数据集上的准确率为81.5%,在FMA-Small上的准确率为64.3%。所提出的DNN分类器比线性分类器提高了10-16%的性能。我们探索了对比和三重损失以及优化损失权重的多任务训练的效果,达到了最高的准确率。为了解决跨数据集的挑战,我们将GTZAN和FMA-Small合并到一个统一的18类标签空间中进行联合训练,导致GTZAN的性能略有下降,但FMA-Small的结果相当。在这项工作中开发的脚本是公开的。
摘要:This study investigates the use of self-supervised learning embeddings, particularly BYOL-A, in conjunction with a deep neural network classifier for Music Genre Classification. Our experiments demonstrate that BYOL-A embeddings outperform other pre-trained models, such as PANNs and VGGish, achieving an accuracy of 81.5% on the GTZAN dataset and 64.3% on FMA-Small. The proposed DNN classifier improved performance by 10-16% over linear classifiers. We explore the effects of contrastive and triplet loss and multitask training with optimized loss weights, achieving the highest accuracy. To address cross dataset challenges, we combined GTZAN and FMA-Small into a unified 18-class label space for joint training, resulting in slight performance drops on GTZAN but comparable results on FMA-Small. The scripts developed in this work are publicly available.
【12】Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
标题:通过三级公式和TLR进行集成欺骗稳健的自动说话人验证
链接:https://arxiv.org/abs/2603.13780
备注:Submitted to Interspeech 2026; put on arxiv based on requirement from Interspeech: "Interspeech no longer enforces an anonymity period for submissions." and "For authors that prefer to upload their paper online, a note indicating that the paper was submitted for review to Interspeech should be included in the posting."
摘要:抗欺骗的自动说话人确认(SASV)旨在将自动说话人确认(ASV)和对抗(CM)相结合。一个流行的解决方案是融合独立的ASV和CM分数。为了更好地建模SASV,一些框架将ASV和CM集成在单个网络中。然而,这些解决方案通常是基于双编码器的,提供有限的可解释性,并且在没有重新训练的情况下不能容易地适应新的评估参数。在此基础上,我们通过三类公式提出了一个统一的端到端框架,该框架能够从类逻辑中进行对数似然比(LLR)推断,以获得更具可解释性的决策管道。实验结果表明,与ASVSpoof5上的现有方法相比,SpoofCeleb上的结果更好。可视化和分析也证明了三类重构提供了更好的可解释性。
摘要:Spoofing-robust automatic speaker verification (SASV) aims to integrate automatic speaker verification (ASV) and countermeasure (CM). A popular solution is fusion of independent ASV and CM scores. To better modeling SASV, some frameworks integrate ASV and CM within a single network. However, these solutions are typically bi-encoder based, offer limited interpretability, and cannot be readily adapted to new evaluation parameters without retraining. Based on this, we propose a unified end-to-end framework via a three-class formulation that enables log-likelihood ratio (LLR) inference from class logits for a more interpretable decision pipeline. Experiments show comparable performance to existing methods on ASVSpoof5 and better results on SpoofCeleb. The visualization and analysis also prove that the three-class reformulation provides more interpretability.
【13】VoXtream2: Full-stream TTS with dynamic speaking rate control
标题:VoXtream 2:具有动态语速控制的全流RTS
链接:https://arxiv.org/abs/2603.13518
备注:10 pages, 9 figures, Submitted to Interspeech 2026
摘要:交互式系统的全流文本到语音(TTS)必须以最小的延迟开始说话,同时随着文本的增量到达而保持可控。我们提出了VoXtream2,一个zero-shot全码流TTS模型与动态说话速率控制,可以更新中间话语的飞行。VoXtream2将持续时间状态的分布匹配机制与调节信号的无分类器指导相结合,以提高可控性和合成质量。非文本屏蔽启用非文本音频提示,从而消除了对提示转录的需要。在标准的zero-shot基准测试和专用的说话速率测试集上,VoXtream2实现了与公共基线相比具有竞争力的客观和主观结果,尽管模型较小,训练数据较少。在全流模式下,它在消费级GPU上的运行速度比实时快4倍,首包延迟为74 ms。
摘要:Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a zero-shot full-stream TTS model with dynamic speaking-rate control that can be updated mid-utterance on the fly. VoXtream2 combines a distribution matching mechanism over duration states with classifier-free guidance across conditioning signals to improve controllability and synthesis quality. Prompt-text masking enables textless audio prompting, removing the need for prompt transcription. Across standard zero-shot benchmarks and a dedicated speaking-rate test set, VoXtream2 achieves competitive objective and subjective results against public baselines despite a smaller model and less training data. In full-stream mode, it runs 4 times faster than real time with 74 ms first-packet latency on a consumer GPU.
【14】Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
标题:了解SSL模型用于音频deepfake模型属性的优势和劣势
链接:https://arxiv.org/abs/2603.13488
备注:Accepted for publication at ICASSP 2026
摘要:音频deepfake模型归因旨在通过识别负责生成给定音频样本的源模型来减少合成语音的滥用,从而实现问责制并通知供应商。这项任务具有挑战性,但自监督学习(SSL)衍生的声学特征已经展示了最先进的归因能力,但推动其成功的潜在因素及其区分能力的限制仍不清楚。在本文中,我们系统地研究了SSL派生的特征如何捕获音频deepfakes中的架构签名。通过控制音频生成过程的多个维度,我们揭示了模型检查点,文本提示,声码器或扬声器身份的微妙扰动如何影响归因。我们的研究结果为基于SSL的deepfake归因的鲁棒性,偏见和局限性提供了新的见解,突出了其在现实场景中的优势和弱点。
摘要:Audio deepfake model attribution aims to mitigate the misuse of synthetic speech by identifying the source model responsible for generating a given audio sample, enabling accountability and informing vendors. The task is challenging, but self-supervised learning (SSL)-derived acoustic features have demonstrated state-of-the-art attribution capabilities, yet the underlying factors driving their success and the limits of their discriminative power remain unclear. In this paper, we systematically investigate how SSL-derived features capture architectural signatures in audio deepfakes. By controlling multiple dimensions of the audio generation process we reveal how subtle perturbations in model checkpoints, text prompts, vocoders, or speaker identity influence attribution. Our results provide new insights into the robustness, biases, and limitations of SSL-based deepfake attribution, highlighting both its strengths and vulnerabilities in realistic scenarios.
【15】BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
标题:BrainWhisperer:利用大规模ASR模型进行神经语音解码
链接:https://arxiv.org/abs/2603.13321
摘要:从皮质内记录中解码连续语音是脑机接口(BCI)的一个核心挑战,对于那些说话能力受损的人来说,这具有变革的潜力。虽然最近的微电极阵列(MEA)解码器实现了令人印象深刻的准确性,但它们的性能从根本上受到现有数据集的小尺寸的限制,它们仍然容易受到会话间变异性的影响,并且它们在参与者之间进行概括的能力仍未得到探索。我们介绍BrainWhisperer,一种神经语音解码器,它将高分辨率MEA记录与大型预训练自动语音识别(ASR)模型集成在一起。基于可解释性研究结果表明,Whisper的编码器学习音素选择性表示与本地化的注意,我们训练一个定制版本的Whisper,修改处理神经功能,使用一个混合目标,结合CTC损失音素-预测从第三编码器层-和交叉熵损失的单词令牌。我们引入了域信息修改,包括窗口化自我注意力以捕获发音连续性,分层月/日特定的低秩预测以解决非平稳性,以及特定于主题的嵌入器,从而实现跨主题训练。在公开可用的MEA数据集上进行评价(Card等人),BrainWhisperer匹配或优于先前的最先进的解码器。至关重要的是,跨数据集训练即使在没有微调的情况下也能提高单个数据集的性能,表现出前所未有的泛化能力。该模型支持双重解码路径:基于音素的高精度路径,具有外部语言模型重新评分,以及快速直接文本生成路径,能够以最小的硬件要求进行低于100 ms的推理。
摘要:Decoding continuous speech from intracortical recordings is a central challenge for brain-computer interfaces (BCIs), with transformative potential for individuals with conditions that impair their ability to speak. While recent microelectrode array (MEA) decoders achieve impressive accuracy, their performance is fundamentally limited by the small size of existing datasets, they remain brittle to session-to-session variability, and their ability to generalize across participants remains unexplored. We introduce BrainWhisperer, a neural speech decoder that integrates high-resolution MEA recordings with a large pretrained automatic speech recognition (ASR) model. Building on interpretability findings showing that Whisper's encoder learns phoneme-selective representations with localized attention, we train a customized version of Whisper, modified to process neural features, using a hybrid objective that combines CTC loss on phonemes--predicted from the third encoder layer--and cross-entropy loss on word tokens. We introduce domain-informed modifications including windowed self-attention to capture articulatory continuity, hierarchical month/day-specific low-rank projections to address non-stationarity, and subject-specific embedders enabling cross-subject training. Evaluated on a publicly available MEA dataset (Card et al.), BrainWhisperer matches or outperforms prior state-of-the-art decoders. Critically, cross-dataset training improves performance even on individual datasets without fine-tuning, demonstrating unprecedented generalization. The model supports dual decoding paths: a high-accuracy phoneme-based path with external language model rescoring, and a fast direct text generation path enabling sub-100ms inference with minimal hardware requirements.
【16】AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
标题:AC-Foley:具有声学传输的参考音频引导视频到音频合成
链接:https://arxiv.org/abs/2603.15597
备注:Accepted at ICLR 2026. 15 pages, 5 figures
摘要:现有的视频到音频(V2 A)生成方法主要依赖于文本提示以及视觉信息来合成音频。然而,两个关键的瓶颈仍然存在:训练数据中的语义粒度差距,例如将声学上不同的声音合并在粗糙的标签下,以及描述微观声学特征时的文本模糊性。这些瓶颈使得很难使用文本控制模式执行细粒度的声音合成。为了解决这些限制,我们提出了AC-Foley,一种音频调节的V2 A模型,它直接利用参考音频来实现对生成的声音的精确和细粒度控制。这种方法能够实现细粒度的声音合成、音色转移、zero-shot声音生成和改进的音频质量。通过直接调节音频信号,我们的方法绕过了文本描述的语义模糊性,同时能够精确地操纵声学属性。从经验上讲,AC-Foley在参考音频条件下实现了Foley生成的最先进性能,同时即使没有音频条件,也能与最先进的视频到音频方法保持竞争力。
摘要:Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating acoustically distinct sounds under coarse labels, and textual ambiguity in describing micro-acoustic features. These bottlenecks make it difficult to perform fine-grained sound synthesis using text-controlled modes. To address these limitations, we propose AC-Foley, an audio-conditioned V2A model that directly leverages reference audio to achieve precise and fine-grained control over generated sounds. This approach enables fine-grained sound synthesis, timbre transfer, zero-shot sound generation, and improved audio quality. By directly conditioning on audio signals, our approach bypasses the semantic ambiguities of text descriptions while enabling precise manipulation of acoustic attributes. Empirically, AC-Foley achieves state-of-the-art performance for Foley generation when conditioned on reference audio, while remaining competitive with state-of-the-art video-to-audio methods even without audio conditioning.
【17】Music Genre Classification: A Comparative Analysis of Classical Machine Learning and Deep Learning Approaches
标题:音乐流派分类:经典机器学习和深度学习方法的比较分析
链接:https://arxiv.org/abs/2603.15440
备注:8 pages
摘要:自动音乐流派分类是音乐信息检索(MIR)中的一个长期挑战;非西方音乐传统的工作仍然很少。尼泊尔音乐包含了丰富的文化和不同的音乐类型-从Lok Dohori的呼唤和回应二重唱到Deuda的节奏诗和Tamang Selo的独特旋律-这些都没有被现有的分类系统所解决。在本文中,我们构建了一个新的数据集,包含大约8,000个标记的30秒音频片段,跨越8个尼泊尔音乐流派,并对两种范式的9个分类模型进行了系统的比较。五种经典的机器学习分类器(Logistic Regression,SVM,KNN,Random Forest和XGBoost)在通过Librosa提取的51个手工制作的音频特征上进行训练,而四种深度学习架构(CNN,RNN,并行CNN-RNN和顺序CNN,然后是RNN)在维度为640 x 128的Mel频谱图上进行操作。我们的实验表明,序列卷积递归神经网络(CRNN)-其中卷积层馈入LSTM-实现了84%的最高准确度,大大优于最佳经典模型(Logistic Regression和XGBoost,均为71%)和所有其他深度架构。我们为每个模型提供每类精度,召回率,F1分数,混淆矩阵和ROC分析,并提供基于文化的错误分类模式解释,反映了尼泊尔音乐传统的真正重叠。
摘要:Automatic music genre classification is a long-standing challenge in Music Information Retrieval (MIR); work on non-Western music traditions remains scarce. Nepali music encompasses culturally rich and acoustically diverse genres--from the call-and-response duets of Lok Dohori to the rhythmic poetry of Deuda and the distinctive melodies of Tamang Selo--that have not been addressed by existing classification systems. In this paper, we construct a novel dataset of approximately 8,000 labeled 30-second audio clips spanning eight Nepali music genres and conduct a systematic comparison of nine classification models across two paradigms. Five classical machine learning classifiers (Logistic Regression, SVM, KNN, Random Forest, and XGBoost) are trained on 51 hand-crafted audio features extracted via Librosa, while four deep learning architectures (CNN, RNN, parallel CNN-RNN, and sequential CNN followed by RNN) operate on Mel spectrograms of dimension 640 x 128. Our experiments reveal that the sequential Convolutional Recurrent Neural Network (CRNN)--in which convolutional layers feed into an LSTM--achieves the highest accuracy of 84%, substantially outperforming both the best classical models (Logistic Regression and XGBoost, both at 71%) and all other deep architectures. We provide per-class precision, recall, F1-score, confusion matrices, and ROC analysis for every model, and offer a culturally grounded interpretation of misclassification patterns that reflects genuine overlaps in Nepal's musical traditions.
【18】NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
标题:NV-Bench:用于表达性文本到语音生成的非言语发声合成基准
链接:https://arxiv.org/abs/2603.15352
摘要:虽然最近的文本到语音(TTS)系统越来越多地集成非语言发声(NV),但它们的评估缺乏标准化的指标和可靠的地面实况参考。为了弥合这一差距,我们提出了NV-Bench,第一个基准接地的功能分类法,将NVs作为沟通行为,而不是声学文物。NV-Bench包含1,651个多语言的野外话语,带有配对的人类参考音频,在14个NV类别中平衡。我们引入了一个双维度的评估协议:(1)指令对齐,利用提出的语言字符错误率(PCER)来评估可控性,(2)声学保真度,测量分布差距,以真实的录音,以评估声学的真实性。我们评估了不同的TTS模型,并开发了两个基线。实验结果表明,我们的客观指标和人类感知之间的强相关性,建立NV-Bench作为一个标准化的评估框架。
摘要:While recent text-to-speech (TTS) systems increasingly integrate nonverbal vocalizations (NVs), their evaluations lack standardized metrics and reliable ground-truth references. To bridge this gap, we propose NV-Bench, the first benchmark grounded in a functional taxonomy that treats NVs as communicative acts rather than acoustic artifacts. NV-Bench comprises 1,651 multi-lingual, in-the-wild utterances with paired human reference audio, balanced across 14 NV categories. We introduce a dual-dimensional evaluation protocol: (1) Instruction Alignment, utilizing the proposed paralinguistic character error rate (PCER) to assess controllability, (2) Acoustic Fidelity, measuring the distributional gap to real recordings to assess acoustic realism. We evaluate diverse TTS models and develop two baselines. Experimental results demonstrate a strong correlation between our objective metrics and human perception, establishing NV-Bench as a standardized evaluation framework.
【19】Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models
标题:推动隐藏状态:大型音频语言模型中思想链推理的免训练模型引导
链接:https://arxiv.org/abs/2603.14636
备注:6 pages, 4 figures, 2 tables
摘要:思想链(CoT)提示已经扩展到大型音频语言模型(LALM)以引发推理,但在没有训练的情况下提高其有效性仍然具有挑战性。我们研究了推理时间模型转向作为一种无需训练的方法来改善LALM推理。我们介绍了三种策略,使用不同的信息来源,并评估他们在四个LALM和四个基准。结果显示,一般的准确性增益高达4.4%,超过CoT提示。值得注意的是,我们确定了一个跨模态传输,其中来自少数文本样本的导向向量有效地引导基于语音的推理,表现出高数据效率。我们还研究了超参数敏感性,以了解这些方法的鲁棒性。我们的研究结果定位模型转向作为一个实际的方向,加强LALM推理。
摘要:Chain-of-thought (CoT) prompting has been extended to large audio-language models (LALMs) to elicit reasoning, yet enhancing its effectiveness without training remains challenging. We study inference-time model steering as a training-free approach to improve LALM reasoning. We introduce three strategies using diverse information sources and evaluate them across four LALMs and four benchmarks. Results show general accuracy gains up to 4.4% over CoT prompting. Notably, we identify a cross-modal transfer where steering vectors derived from few text samples effectively guide speech-based reasoning, demonstrating high data efficiency. We also examine hyperparameter sensitivity to understand the robustness of these approaches. Our findings position model steering as a practical direction for strengthening LALM reasoning.
【20】CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
标题:CodecMOS-Accent:跨英语口音的神经编解码器重新合成和TTC语音的MOS基准
链接:https://arxiv.org/abs/2603.14328
备注:Preprint
摘要:我们提出了CodecMOS-Accent数据集,这是一个平均意见评分(MOS)基准,旨在评估神经音频编解码器(NAC)模型和基于大语言模型(LLM)的文本到语音(TTS)模型,特别是在非标准语音(如口音语音)中。该数据集包括来自24个系统的4,000个编解码器再合成和TTS样本,具有32个扬声器,跨越10个口音。进行了大规模的主观测试,从25名听众中收集了19,600条注释,涉及三个维度:自然度,说话人相似性和口音相似性。该数据集不仅代表了对最近语音合成系统性能的最新研究,而且揭示了一些见解,包括说话者和口音相似性之间的紧密关系,客观指标的预测能力,以及当听众与说话者共享相同口音时的感知偏差。预计该数据集将促进对NAC和重音TTS进行更多以人为中心的评估的研究。
摘要:We present the CodecMOS-Accent dataset, a mean opinion score (MOS) benchmark designed to evaluate neural audio codec (NAC) models and the large language model (LLM)-based text-to-speech (TTS) models trained upon them, especially across non-standard speech like accented speech. The dataset comprises 4,000 codec resynthesis and TTS samples from 24 systems, featuring 32 speakers spanning ten accents. A large-scale subjective test was conducted to collect 19,600 annotations from 25 listeners across three dimensions: naturalness, speaker similarity, and accent similarity. This dataset does not only represent an up-to-date study of recent speech synthesis system performance but reveals insights including a tight relationship between speaker and accent similarity, the predictive power of objective metrics, and a perceptual bias when listeners share the same accent with the speaker. This dataset is expected to foster research on more human-centric evaluation for NAC and accented TTS.
【21】What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
标题:什么才算真实?语音恢复和语音质量转换对Deepfake检测提出新挑战
链接:https://arxiv.org/abs/2603.14033
备注:5 pages, 4 figures, 3 tables. Submitted to Interspeech 2026
摘要:音频反欺骗系统通常被公式化为区分真实语音和欺骗语音的二进制分类器。这种假设在分层生成处理下失败,其中良性转换引入了被错误分类为欺骗的分布偏移。我们表明,发音修改语音转换和语音恢复被视为外的分布,尽管保留扬声器的真实性。使用分离真实,转换,欺骗和转换欺骗语音的多类设置,我们通过自监督学习(SSL)嵌入和声学相关来分析模型行为。良性的转换在SSL空间中引起漂移,压缩真实和欺骗的语音并降低分类器的可分离性。将反欺骗重新定义为多类问题,提高了对良性变化的鲁棒性,同时保留了欺骗检测,这表明二进制系统对原始语音的分布而不是真实性本身进行建模。
摘要:Audio anti-spoofing systems are typically formulated as binary classifiers distinguishing bona fide from spoofed speech. This assumption fails under layered generative processing, where benign transformations introduce distributional shifts that are misclassified as spoofing. We show that phonation-modifying voice conversion and speech restoration are treated as out-of-distribution despite preserving speaker authenticity. Using a multi-class setup separating bona fide, converted, spoofed, and converted-spoofed speech, we analyse model behaviour through self-supervised learning (SSL) embeddings and acoustic correlates. The benign transformations induce a drift in the SSL space, compressing bona fide and spoofed speech and reducing classifier separability. Reformulating anti-spoofing as a multi-class problem improves robustness to benign shifts while preserving spoof detection, suggesting binary systems model the distribution of raw speech rather than authenticity itself.
【22】LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
标题:用于视听语音增强的LLM引导强化学习
链接:https://arxiv.org/abs/2603.13952
备注:6 pages, 4 figures, submitted to Interspeech 2026
摘要:在现有的视听语音增强(AVSE)方法中,诸如尺度不变信噪比(SI-SNR)和均方误差(MSE)的目标被广泛使用;然而,它们通常与感知质量相关性差并且提供有限的可解释性用于优化。这项工作提出了一个基于强化学习的AVSE框架,该框架具有基于大型语言模型(LLM)的可解释奖励模型。音频LLM生成增强语音的自然语言描述,这些描述由情感分析模型转换为1-5的评级分数,作为用于微调预训练AVSE模型的PPO奖励。与标量度量相比,LLM生成的反馈在语义上是丰富的,并且明确地描述了语音质量的改进。在第四届COG-MHEAR AVSE挑战赛(AVSEC-4)数据集上的实验表明,该方法在PESQ,STOI,神经质量指标和主观听力测试中优于监督基线和基于DNSMOS的RL基线。
摘要:In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, they often correlate poorly with perceptual quality and provide limited interpretability for optimization. This work proposes a reinforcement learning-based AVSE framework with a Large Language Model (LLM)-based interpretable reward model. An audio LLM generates natural language descriptions of enhanced speech, which are converted by a sentiment analysis model into a 1-5 rating score serving as the PPO reward for fine-tuning a pretrained AVSE model. Compared with scalar metrics, LLM-generated feedback is semantically rich and explicitly describes improvements in speech quality. Experiments on the 4th COG-MHEAR AVSE Challenge (AVSEC-4) dataset show that the proposed method outperforms a supervised baseline and a DNSMOS-based RL baseline in PESQ, STOI, neural quality metrics, and subjective listening tests.
【23】Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings
标题:来自多部位听诊记录的患者级多模式问题回答
链接:https://arxiv.org/abs/2603.13362
摘要:听诊是一种重要的诊断工具,但其实用性往往受到主观解释的限制。虽然通用的音频语言模型(ALM)在一般领域表现出色,但它们在生理信号的细微差别方面表现不佳。我们提出了一个框架,通过门控交叉注意直接将多站点听诊记录与冻结的大语言模型(LLM)嵌入空间对齐。通过利用LLM的潜在世界知识,我们的方法超越了孤立的分类,走向全面的,患者级的评估。在CaReSound基准测试中,我们的模型达到了最先进的0.865 F1宏和0.952 BERTScore。我们证明,轻量级的,特定于域的编码器的竞争对手大规模的ALM和多站点聚合提供空间冗余,减轻时间截断。医学声学与文本基础的这种对齐为桥接信号处理和临床评估提供了可扩展的路径。
摘要:Auscultation is a vital diagnostic tool, yet its utility is often limited by subjective interpretation. While general-purpose Audio-Language Models (ALMs) excel in general domains, they struggle with the nuances of physiological signals. We propose a framework that aligns multi-site auscultation recordings directly with a frozen Large Language Model (LLM) embedding space via gated cross-attention. By leveraging the LLM's latent world knowledge, our approach moves beyond isolated classification toward holistic, patient-level assessment. On the CaReSound benchmark, our model achieves a state-of-the-art 0.865 F1-macro and 0.952 BERTScore. We demonstrate that lightweight, domain-specific encoders rival large-scale ALMs and that multi-site aggregation provides spatial redundancy that mitigates temporal truncation. This alignment of medical acoustics with text foundations offers a scalable path for bridging signal processing and clinical assessment.
机器翻译由腾讯交互翻译提供,仅供参考
