微信公众号:arXiv_Daily
cs.SD语音
【1】Adapting Speech Language Model to Singing Voice Synthesis
标题:语音语言模型适应歌唱声音合成
链接:https://arxiv.org/abs/2512.14657
备注:Accepted by NeurIPS 2025 workshop AI for Music
摘要:语音语言模型(Speech Language Models,SLM)是近年来出现的一种统一的语音相关任务处理方法,包括文本到语音转换(TTS)、语音增强(SE)和自动语音识别(ASR)。然而,大规模预训练SLM的泛化能力仍然没有得到充分的探索。在这项工作中,我们适应了1.7B参数TTS预训练SLM的歌声合成(SVS),只使用135小时的合成歌唱语料库,ACE-OpenPop。在ESPNet-SpeechLM的基础上,我们的配方涉及以下过程:(1)乐谱条件和演唱波形的标记化,(2)多流语言模型标记预测,(3)基于条件流匹配的梅尔声谱图生成。(4)一种梅尔到波声码器。实验结果表明,我们的适应SLM推广以及SVS和领先的离散令牌为基础的SVS模型实现性能相当。
摘要:Speech Language Models (SLMs) have recently emerged as a unified paradigm for addressing a wide range of speech-related tasks, including text-to-speech (TTS), speech enhancement (SE), and automatic speech recognition (ASR). However, the generalization capability of large-scale pre-trained SLMs remains underexplored. In this work, we adapt a 1.7B parameter TTS pretrained SLM for singing voice synthesis (SVS), using only a 135-hour synthetic singing corpus, ACE-Opencpop. Building upon the ESPNet-SpeechLM, our recipe involves the following procedure: (1) tokenization of music score conditions and singing waveforms, (2) multi-stream language model token prediction, (3) conditional flow matching-based mel-spectrogram generation. (4) a mel-to-wave vocoder. Experimental results demonstrate that our adapted SLM generalizes well to SVS and achieves performance comparable to leading discrete token-based SVS models.
【2】Robust Training of Singing Voice Synthesis Using Prior and Posterior Uncertainty
标题:使用先验和后验不确定性的歌唱声音合成的稳健训练
链接:https://arxiv.org/abs/2512.14653
备注:Accepted by ASRU 2025
摘要:歌唱声音合成(SVS)近年来取得了显著的进步。然而,与语音和一般音频数据相比,公开可用的唱歌数据集仍然有限。在实践中,这种数据稀缺往往会导致长尾场景中的性能下降,例如不平衡的音高分布或罕见的演唱风格。为了缓解这些挑战,我们提出了基于不确定性的优化来改进端到端SVS模型的训练过程。首先,我们在对抗训练中引入可微数据增强,它以样本方式操作以增加先验不确定性。其次,我们引入了一个帧级不确定性预测模块,用于估计后验不确定性,使模型能够为低置信度段分配更多的学习能力。在中国和日本的Opencpop和Ofuton-P上的实证结果表明,我们的方法在各个方面都提高了性能。
摘要:Singing voice synthesis (SVS) has seen remarkable advancements in recent years. However, compared to speech and general audio data, publicly available singing datasets remain limited. In practice, this data scarcity often leads to performance degradation in long-tail scenarios, such as imbalanced pitch distributions or rare singing styles. To mitigate these challenges, we propose uncertainty-based optimization to improve the training process of end-to-end SVS models. First, we introduce differentiable data augmentation in the adversarial training, which operates in a sample-wise manner to increase the prior uncertainty. Second, we incorporate a frame-level uncertainty prediction module that estimates the posterior uncertainty, enabling the model to allocate more learning capacity to low-confidence segments. Empirical results on the Opencpop and Ofuton-P, across Chinese and Japanese, demonstrate that our approach improves performance in various perspectives.
【3】MuseCPBench: an Empirical Study of Music Editing Methods through Music Context Preservation
标题:MuseCPBench:通过音乐语境保存进行音乐编辑方法的实证研究
链接:https://arxiv.org/abs/2512.14629
摘要:音乐编辑在现代音乐制作中起着至关重要的作用,在电影,广播和游戏开发中都有应用。音乐生成模型的最新进展使不同的编辑任务,如音色转移,乐器替换和流派转换。然而,许多现有的作品忽略了他们的能力,以保持音乐方面,应该保持不变,在编辑的属性,我们定义为音乐上下文保存(MCP)的评估。虽然一些研究确实考虑了MCP,但它们采用了不一致的评估协议和指标,导致不可靠和不公平的比较。为了解决这一差距,我们引入了第一个MCP评估基准MuseCPBench,它涵盖了四类音乐方面,并能够在五个代表性的音乐编辑基线之间进行全面比较。通过对音乐方面、方法和模型的系统分析,我们发现了当前音乐编辑方法中一致的保存差距,并提供了有见地的解释。我们希望我们的研究结果为开发更有效和可靠的具有强大MCP能力的音乐编辑策略提供实际指导
摘要:Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game development. Recent advances in music generation models have enabled diverse editing tasks such as timbre transfer, instrument substitution, and genre transformation. However, many existing works overlook the evaluation of their ability to preserve musical facets that should remain unchanged during editing a property we define as Music Context Preservation (MCP). While some studies do consider MCP, they adopt inconsistent evaluation protocols and metrics, leading to unreliable and unfair comparisons. To address this gap, we introduce the first MCP evaluation benchmark, MuseCPBench, which covers four categories of musical facets and enables comprehensive comparisons across five representative music editing baselines. Through systematic analysis along musical facets, methods, and models, we identify consistent preservation gaps in current music editing methods and provide insightful explanations. We hope our findings offer practical guidance for developing more effective and reliable music editing strategies with strong MCP capability
【4】Sound and Music Biases in Deep Music Transcription Models: A Systematic Analysis
标题:深度音乐转录模型中的声音和音乐偏见:系统分析
链接:https://arxiv.org/abs/2512.14602
备注:pre-print of the upcoming EURASIP JASM journal article
摘要:自动音乐转录(AMT)--将音乐音频转换为音符表示的任务--在深度学习系统的推动下取得了快速进展。由于丰富注释的音乐数据集有限,AMT的大部分进展都集中在古典钢琴音乐上,甚至是一些非常具体的数据集。这些系统是否能有效地推广到其他音乐环境仍然是一个悬而未决的问题。补充最近关于声音分布变化的研究(例如,录音条件),在这项工作中,我们调查的音乐层面-具体而言,在体裁,动态和复调水平的变化。为此,我们引入MDS语料库,包括三个不同的子集-(1)类型,(2)随机,(3)MAEtest-模拟不同的轴分布的转变。我们评估了几个国家的最先进的AMT系统的MDS语料库使用传统的信息检索和音乐的性能指标的性能。我们广泛的评估隔离并暴露了特定分布变化下不同程度的性能下降。特别是,我们测量了由于声音而导致的20个百分点的音符级F1性能下降,以及由于类型而导致的14个百分点。一般来说,我们发现,动态估计证明更容易受到音乐的变化比发病预测。基于音乐的评估指标,特别是那些捕捉和声结构的指标,有助于识别潜在的影响因素。此外,随机生成的,非音乐序列的实验揭示了极端的音乐分布变化下的系统性能的明显局限性。总而言之,这些发现提供了新的证据,证明语料库偏见问题在深层AMT系统中的持续影响。
摘要:Automatic Music Transcription (AMT) -- the task of converting music audio into note representations -- has seen rapid progress, driven largely by deep learning systems. Due to the limited availability of richly annotated music datasets, much of the progress in AMT has been concentrated on classical piano music, and even a few very specific datasets. Whether these systems can generalize effectively to other musical contexts remains an open question. Complementing recent studies on distribution shifts in sound (e.g., recording conditions), in this work we investigate the musical dimension -- specifically, variations in genre, dynamics, and polyphony levels. To this end, we introduce the MDS corpus, comprising three distinct subsets -- (1) Genre, (2) Random, and (3) MAEtest -- to emulate different axes of distribution shift. We evaluate the performance of several state-of-the-art AMT systems on the MDS corpus using both traditional information-retrieval and musically-informed performance metrics. Our extensive evaluation isolates and exposes varying degrees of performance degradation under specific distribution shifts. In particular, we measure a note-level F1 performance drop of 20 percentage points due to sound, and 14 due to genre. Generally, we find that dynamics estimation proves more vulnerable to musical variation than onset prediction. Musically informed evaluation metrics, particularly those capturing harmonic structure, help identify potential contributing factors. Furthermore, experiments with randomly generated, non-musical sequences reveal clear limitations in system performance under extreme musical distribution shifts. Altogether, these findings offer new evidence of the persistent impact of the Corpus Bias problem in deep AMT systems.
【5】Linguists should learn to love speech-based deep learning models
标题:语言学家应该学会热爱基于语音的深度学习模型
链接:https://arxiv.org/abs/2512.14506
备注:Commentary on Futrell, R., & Mahowald, K. arXiv:2501.17047 (in press). How Linguistics Learned to Stop Worrying and Love the Language Models. Behavioural and Brain Sciences
摘要:Futrell和Mahowald提出了一个有用的框架,将面向技术的深度学习系统和面向解释的语言学理论联系起来。不幸的是,目标文章对基于生成文本的LLM的关注从根本上限制了与语言学的富有成效的互动,因为许多关于人类语言的有趣问题都超出了书面文本的范围。我们认为,基于音频的深度学习模型可以而且应该发挥关键作用。
摘要:Futrell and Mahowald present a useful framework bridging technology-oriented deep learning systems and explanation-oriented linguistic theories. Unfortunately, the target article's focus on generative text-based LLMs fundamentally limits fruitful interactions with linguistics, as many interesting questions on human language fall outside what is captured by written text. We argue that audio-based deep learning models can and should play a crucial role.
【6】GLM-TTS Technical Report
标题:GLM-TTC技术报告
链接:https://arxiv.org/abs/2512.14291
摘要:这项工作提出了GLM-TTS,一个生产级的TTS系统设计的效率,可控性和高保真语音生成。GLM-TTS遵循两阶段架构,由文本到令牌自回归模型和令牌到波形扩散模型组成。仅使用10万小时的训练数据,GLM-TTS就在多个开源基准测试中实现了最先进的性能。为了满足生产要求,GLM-TTS通过具有基频约束的优化语音标记器和基于GRPO的多奖励强化学习框架来提高语音质量,该框架联合优化发音,说话人相似性和表达韵律。同时,该系统通过基于参数高效的LoRA语音定制和提供精确发音控制的混合音素-文本输入方案实现了高效和可控的部署。我们的代码可在https://github.com/zai-org/GLM-TTS上获得。实时语音合成演示可通过Z.ai(audio.z.ai)、知普清颜app/web(chatglm.cn)提供。
摘要:This work proposes GLM-TTS, a production-level TTS system designed for efficiency, controllability, and high-fidelity speech generation. GLM-TTS follows a two-stage architecture, consisting of a text-to-token autoregressive model and a token-to-waveform diffusion model. With only 100k hours of training data, GLM-TTS achieves state-of-the-art performance on multiple open-source benchmarks. To meet production requirements, GLM-TTS improves speech quality through an optimized speech tokenizer with fundamental frequency constraints and a GRPO-based multi-reward reinforcement learning framework that jointly optimizes pronunciation, speaker similarity, and expressive prosody. In parallel, the system enables efficient and controllable deployment via parameter-efficient LoRA-based voice customization and a hybrid phoneme-text input scheme that provides precise pronunciation control. Our code is available at https://github.com/zai-org/GLM-TTS. Real-time speech synthesis demos are provided via Z.ai (audio.z.ai), the Zhipu Qingyan app/web (chatglm.cn).
【7】Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting
标题:用于鲁棒性口语检测和关键词发现的联合多模式对比学习
链接:https://arxiv.org/abs/2512.14115
摘要:声学词嵌入(AWE)提高了语音检索任务的效率,如口语词检测(STD)和关键词定位(KWS)。然而,现有的方法存在局限性,包括单模态监督,音频-音频和音频-文本对齐的不相交优化,以及需要特定于任务的模型。为了解决这些缺点,我们提出了一个联合多模态对比学习框架,该框架将声学和跨模态监督统一在共享的嵌入空间中。我们的方法同时优化了:(i)音频-文本对比学习,灵感来自CLAP损失,以对齐音频和文本表示,以及(ii)音频-音频对比学习,通过深度单词识别(DWD)损失,以增强类内紧凑性和类间分离。所提出的方法优于现有的AWE基线的单词歧视任务,同时灵活地支持STD和KWS。据我们所知,这是第一个此类全面办法。
摘要:Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer from limitations, including unimodal supervision, disjoint optimization of audio-audio and audio-text alignment, and the need for task-specific models. To address these shortcomings, we propose a joint multimodal contrastive learning framework that unifies both acoustic and cross-modal supervision in a shared embedding space. Our approach simultaneously optimizes: (i) audio-text contrastive learning, inspired by the CLAP loss, to align audio and text representations and (ii) audio-audio contrastive learning, via Deep Word Discrimination (DWD) loss, to enhance intra-class compactness and inter-class separation. The proposed method outperforms existing AWE baselines on word discrimination task while flexibly supporting both STD and KWS. To our knowledge, this is the first comprehensive approach of its kind.
【8】Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study
标题:多语言和连续反向渠道预测:跨语言研究
链接:https://arxiv.org/abs/2512.14085
备注:This paper has been accepted for presentation at International Workshop on Spoken Dialogue Systems Technology 2026 (IWSDS 2026) and represents the author's version of the work
摘要:我们提出了一个多语种的,连续的反向通道预测模型,日语,英语和汉语,并使用它来研究跨语言的时间行为。该模型是基于transformer的,在框架级别上运行,在大约300小时的二元对话中与辅助任务联合训练。在所有三种语言中,多语言模型匹配或超过单语言基线,这表明它既学习了语言通用线索,又学习了语言特定的时间模式。两种语言培训的Zero-shot迁移仍然有限,突出了实质性的跨语言差异。扰动分析揭示了不同的线索使用:日本人更依赖于短期的语言信息,而英语和汉语更敏感的沉默时间和韵律变化;多语言培训鼓励共享但适应性强的表示,并减少过度依赖音高在中国。上下文长度研究进一步表明,日语对较短的上下文相对稳健,而汉语则明显受益于较长的上下文。最后,我们将训练好的模型集成到实时处理软件中,演示仅CPU推理。总之,这些研究结果提供了一个统一的模型和经验证据,说明不同语言之间的反向通道时间是如何不同的,为设计更自然、更有文化意识的口语对话系统提供了信息。
摘要:We present a multilingual, continuous backchannel prediction model for Japanese, English, and Chinese, and use it to investigate cross-linguistic timing behavior. The model is Transformer-based and operates at the frame level, jointly trained with auxiliary tasks on approximately 300 hours of dyadic conversations. Across all three languages, the multilingual model matches or surpasses monolingual baselines, indicating that it learns both language-universal cues and language-specific timing patterns. Zero-shot transfer with two-language training remains limited, underscoring substantive cross-lingual differences. Perturbation analyses reveal distinct cue usage: Japanese relies more on short-term linguistic information, whereas English and Chinese are more sensitive to silence duration and prosodic variation; multilingual training encourages shared yet adaptable representations and reduces overreliance on pitch in Chinese. A context-length study further shows that Japanese is relatively robust to shorter contexts, while Chinese benefits markedly from longer contexts. Finally, we integrate the trained model into a real-time processing software, demonstrating CPU-only inference. Together, these findings provide a unified model and empirical evidence for how backchannel timing differs across languages, informing the design of more natural, culturally-aware spoken dialogue systems.
【9】Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition
标题:Memo2496:专家注释数据集和双视图自适应框架,用于音乐情感识别
链接:https://arxiv.org/abs/2512.13998
摘要:由于有限的高质量注释数据集和解决跨音轨特征漂移的困难,音乐情感识别器(MER)研究面临挑战。这项工作提出了两个主要的贡献,以解决这些问题。Memo 2496是一个大规模的数据集,提供了2496首带有连续效价唤醒标签的器乐曲目,由30位经过认证的音乐专家进行注释。注释质量是通过极端情绪样本和0.25的一致性阈值,通过在效价唤醒空间的欧几里得距离测量校准。此外,介绍了双视角自适应音乐情感识别器(DAMER)。DAMER集成了三个协同模块:双流注意力融合(DSAF)通过交叉注意机制促进Mel频谱图和耳蜗图之间的标记级双向交互;渐进置信标签(PCL)采用基于温度的时间表和使用Jensen Shannon散度的一致性量化生成可靠的伪标签;和风格锚定记忆学习(SAML)保持对比记忆队列以减轻跨轨迹特征漂移。在Memo 2496、1000 songs和PMEmo数据集上的大量实验证明了DAMER的最先进性能,分别将唤醒维度准确度提高了3.43%、2.25%和0.17%。消融研究和可视化分析验证了每个模块的贡献。数据集和源代码都是公开的。
摘要:Music Emotion Recogniser (MER) research faces challenges due to limited high-quality annotated datasets and difficulties in addressing cross-track feature drift. This work presents two primary contributions to address these issues. Memo2496, a large-scale dataset, offers 2496 instrumental music tracks with continuous valence arousal labels, annotated by 30 certified music specialists. Annotation quality is ensured through calibration with extreme emotion exemplars and a consistency threshold of 0.25, measured by Euclidean distance in the valence arousal space. Furthermore, the Dual-view Adaptive Music Emotion Recogniser (DAMER) is introduced. DAMER integrates three synergistic modules: Dual Stream Attention Fusion (DSAF) facilitates token-level bidirectional interaction between Mel spectrograms and cochleagrams via cross attention mechanisms; Progressive Confidence Labelling (PCL) generates reliable pseudo labels employing curriculum-based temperature scheduling and consistency quantification using Jensen Shannon divergence; and Style Anchored Memory Learning (SAML) maintains a contrastive memory queue to mitigate cross-track feature drift. Extensive experiments on the Memo2496, 1000songs, and PMEmo datasets demonstrate DAMER's state-of-the-art performance, improving arousal dimension accuracy by 3.43%, 2.25%, and 0.17%, respectively. Ablation studies and visualisation analyses validate each module's contribution. Both the dataset and source code are publicly available.
【10】Ensemble-Guided Distillation for Compact and Robust Acoustic Scene Classification on Edge Devices
标题:在边缘设备上进行紧凑且鲁棒的声学场景分类的集合引导蒸馏
链接:https://arxiv.org/abs/2512.13905
摘要:我们提出了一个紧凑的、可量化的声学场景分类(ASC)框架,该框架将高效的学生网络与经验丰富的教师集合和知识蒸馏相结合。学生骨干使用堆叠的深度可分离的“扩展深度投影”块和全局响应归一化来稳定训练并提高对设备和噪声变化的鲁棒性,而全局池头产生类logits以进行有效的边缘推断。为了注入更丰富的归纳偏差,我们组装了一组不同的教师模型,并学习了两个互补的融合头:z1,它使用学生风格的骨干预测每个教师的混合权重,z2,一个轻量级的MLP,它执行每个类的logit融合。学生是从合奏通过温度缩放的软目标结合硬标签,使其能够近似合奏的决策几何与一个单一的紧凑的模型。在TAU Urban Acoustic Scenes 2022 Mobile基准上进行评估,我们的方法在匹配的边缘部署约束下在TAU数据集上实现了最先进的(SOTA)结果,证明了移动ASC的强大性能和实用性。
摘要:We present a compact, quantization-ready acoustic scene classification (ASC) framework that couples an efficient student network with a learned teacher ensemble and knowledge distillation. The student backbone uses stacked depthwise-separable "expand-depthwise-project" blocks with global response normalization to stabilize training and improve robustness to device and noise variability, while a global pooling head yields class logits for efficient edge inference. To inject richer inductive bias, we assemble a diverse set of teacher models and learn two complementary fusion heads: z1, which predicts per-teacher mixture weights using a student-style backbone, and z2, a lightweight MLP that performs per-class logit fusion. The student is distilled from the ensemble via temperature-scaled soft targets combined with hard labels, enabling it to approximate the ensemble's decision geometry with a single compact model. Evaluated on the TAU Urban Acoustic Scenes 2022 Mobile benchmark, our approach achieves state-of-the-art (SOTA) results on the TAU dataset under matched edge-deployment constraints, demonstrating strong performance and practicality for mobile ASC.
【11】Privacy-Enhancing Infant Cry Classification with Federated Transformers and Denoising Regularization
标题:基于联邦变换器和去噪正则化的隐私增强婴儿哭声分类
链接:https://arxiv.org/abs/2512.13880
备注:This paper was accepted for presentation and presented at the 2025 International Conference on Computer Engineering, Network, and Intelligent Multimedia (CENIM 2025)
摘要:婴儿哭声分类可以帮助早期评估婴儿的需求。然而,此类解决方案的部署受到音频数据隐私问题、对背景噪音的敏感性以及录制环境中的域转移的限制。我们提出了一个端到端的婴儿哭声分析管道,它集成了一个去噪自动编码器(DAE),一个卷积标记器和一个使用通信高效的联邦学习(FL)训练的Transformer编码器。该系统执行设备上去噪、自适应分割、事后校准和基于能量的分布外(OOD)抑制。联合训练采用了正则化的控制变量更新与8位适配器增量下的安全聚合。使用带有ESC-50噪声覆盖的Baby Chillanto和Donate-a-Cry数据集,该模型实现了0.938的宏F1得分,0.962的AUC和0.032的预期校准误差(ECE),同时将每轮客户端上传从大约36到42 MB减少到3.3 MB。NVIDIA Jetson Nano(4 GB,TensorRT FP 16)上的实时边缘推理可实现96 ms/秒的声谱图帧。这些结果展示了一个实用的路径,以保护隐私,噪声鲁棒性和通信效率的婴儿哭分类适合联邦部署。
摘要:Infant cry classification can aid early assessment of infant needs. However, deployment of such solutions is limited by privacy concerns around audio data, sensitivity to background noise, and domain shift across recording environments. We present an end-to-end infant cry analysis pipeline that integrates a denoising autoencoder (DAE), a convolutional tokenizer, and a Transformer encoder trained using communication-efficient federated learning (FL). The system performs on-device denoising, adaptive segmentation, post hoc calibration, and energy-based out-of-distribution (OOD) abstention. Federated training employs a regularized control variate update with 8-bit adapter deltas under secure aggregation. Using the Baby Chillanto and Donate-a-Cry datasets with ESC-50 noise overlays, the model achieves a macro F1 score of 0.938, an AUC of 0.962, and an Expected Calibration Error (ECE) of 0.032, while reducing per-round client upload from approximately 36 to 42 MB to 3.3 MB. Real-time edge inference on an NVIDIA Jetson Nano (4 GB, TensorRT FP16) achieves 96 ms per one-second spectrogram frame. These results demonstrate a practical path toward privacy-preserving, noise-robust, and communication-efficient infant cry classification suitable for federated deployment.
【12】Toward Noise-Aware Audio Deepfake Detection: Survey, SNR-Benchmarks, and Practical Recipes
标题:迈向噪音感知音频深度伪造检测:调查、SNR基准和实用食谱
链接:https://arxiv.org/abs/2512.13744
备注:6 pages
摘要:Deepfake音频检测在强大的预训练编码器(例如,WavLM、Wav2Vec2、MMS)。然而,在现实的捕获条件下的性能-背景噪声(家庭/办公室/运输),房间混响和消费者通道-通常落后于清洁实验室的结果。我们调查和评估了最先进的音频deepfake检测模型的鲁棒性,并提出了一个可重现的框架,该框架将MS-SNSD噪声与ASVspoof 2021 DF话语混合,以在受控信噪比(SNR)下进行评估。SNR是语音中广泛使用的噪声严重性的测量代理;它允许我们从接近干净(35 dB)到非常嘈杂(-5 dB)进行扫描,以量化优雅的降级。我们研究了预训练编码器(WavLM,Wav 2 Vec 2,MMS)的多条件训练和固定SNR测试,报告准确性,ROC-AUC和二进制和四类(真实性x腐败)任务的EER。在我们的实验中,微调在10-0 dB SNR下跨主干将EER降低了10-15个百分点。
摘要:Deepfake audio detection has progressed rapidly with strong pre-trained encoders (e.g., WavLM, Wav2Vec2, MMS). However, performance in realistic capture conditions - background noise (domestic/office/transport), room reverberation, and consumer channels - often lags clean-lab results. We survey and evaluate robustness for state-of-the-art audio deepfake detection models and present a reproducible framework that mixes MS-SNSD noises with ASVspoof 2021 DF utterances to evaluate under controlled signal-to-noise ratios (SNRs). SNR is a measured proxy for noise severity used widely in speech; it lets us sweep from near-clean (35 dB) to very noisy (-5 dB) to quantify graceful degradation. We study multi-condition training and fixed-SNR testing for pretrained encoders (WavLM, Wav2Vec2, MMS), reporting accuracy, ROC-AUC, and EER on binary and four-class (authenticity x corruption) tasks. In our experiments, finetuning reduces EER by 10-15 percentage points at 10-0 dB SNR across backbones.
【1】Segmental Attention Decoding With Long Form Acoustic Encodings
标题:基于长形式声编码的分段注意解码
链接:https://arxiv.org/abs/2512.14652
备注:5 pages, 1 fig
摘要:我们解决了基于注意力的编码器-解码器(AED)模型与长形式声学编码的根本不兼容性。在分割的话语上训练的AED模型通过利用超出片段边界的有限声学上下文来学习编码绝对帧位置,但是在解码这些线索消失的长形式片段时无法泛化。由于交叉注意中键和值的排列不变性,该模型失去了对声学编码进行排序的能力。我们提出四项修改:(1)将显式绝对位置编码注入到每个解码片段的交叉注意中,(2)利用扩展的声学上下文进行长形式训练以消除隐式绝对位置编码,(3)片段级联以覆盖训练期间所需的各种分段,以及(4)语义分段以将AED解码片段与训练片段对齐。我们表明,这些修改关闭连续和分段声学编码之间的准确性差距,使自回归使用的注意力解码器。
摘要:We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to encode absolute frame positions by exploiting limited acoustic context beyond segment boundaries, but fail to generalize when decoding long-form segments where these cues vanish. The model loses ability to order acoustic encodings due to permutation invariance of keys and values in cross-attention. We propose four modifications: (1) injecting explicit absolute positional encodings into cross-attention for each decoded segment, (2) long-form training with extended acoustic context to eliminate implicit absolute position encoding, (3) segment concatenation to cover diverse segmentations needed during training, and (4) semantic segmentation to align AED-decoded segments with training segments. We show these modifications close the accuracy gap between continuous and segmented acoustic encodings, enabling auto-regressive use of the attention decoder.
【2】Investigating the impact of stereo processing -- a study for extending the Open Dataset of Audio Quality (ODAQ)
标题:调查立体声处理的影响--扩展音频质量开放数据集(ODAQ)的研究
链接:https://arxiv.org/abs/2512.14259
备注:Presented at the Audio Engineering Society (AES) 159th Convention, October 2025, Paper number 365, see https://aes2.org/publications/elibrary-page/?id=23039
摘要:在本文中,我们提出了一个初步的研究扩展开放数据集的音频质量(ODAQ)对立体声处理的影响。单声道文物从ODAQ适应与左右(LR)和中侧(MS)立体声处理的组合,跨刺激,包括独奏乐器,典型的宽立体声混音和硬摇摄混音。听力测试在不同的演示环境-与MS和LR条件的直接比较和没有-进行收集主观数据超出单声道文物,同时也仔细检查听力测试方法。ODAQ数据集扩展了新材料以及16位专家听众的主观分数。听力测试结果表明,刺激的空间特性以及呈现上下文的实质性影响。值得注意的是,LR和MS之间的几个显着差异只发生在直接比较时。研究结果表明,听者主要评估时,空间特征是一致的,立体声图像,只有当音色质量是相似的。额外的单声道锚点的评级在不同的立体声特征中总体上是一致的,在MUSHRA量表上平均为65,进一步证实了听众优先考虑音色而不是空间印象。
摘要:In this paper, we present an initial study for extending Open Dataset of Audio Quality (ODAQ) towards the impact of stereo processing. Monaural artifacts from ODAQ were adapted in combinations with left-right (LR) and mid-side (MS) stereo processing, across stimuli including solo instruments, typical wide stereo mixes and and hard-panned mixes. Listening tests in different presentation context -- with and without direct comparison of MS and LR conditions -- were conducted to collect subjective data beyond monaural artifacts while also scrutinizing the listening test methodology. The ODAQ dataset is extended with new material along with subjective scores from 16 expert listeners. The listening test results show substantial influences of the stimuli's spatial characteristics as well as the presentation context. Notably, several significant disparities between LR and MS only occur when presented in direct comparison. The findings suggest that listeners primarily assess timbral impairments when spatial characteristics are consistent and focus on stereo image only when timbral quality is similar. The rating of an additional mono anchor was overall consistent across different stereo characteristics, averaging at 65 on the MUSHRA scale, further corroborating that listeners prioritize timbral over spatial impressions.
【3】Scalable Frameworks for Real-World Audio-Visual Speech Recognition
标题:用于现实世界视听语音识别的可扩展框架
链接:https://arxiv.org/abs/2512.14083
备注:PhD Dissertation
摘要:视听语音识别(AVSR)系统的实际部署从根本上受到现实环境中的显著性能下降的挑战,其特征在于不可预测的声学噪声和视觉干扰。本文认为,一个系统的,分层的方法是必不可少的,以克服这些挑战,实现强大的可扩展性的表示,架构和系统级别。在表示层面上,我们研究了建立一个统一模型的方法,该模型可以学习对各种现实世界腐败具有固有鲁棒性的视听特征,从而能够在没有专门模块的情况下推广到新环境。为了解决架构可扩展性问题,我们探索如何有效地扩展模型容量,同时确保多模态输入的自适应和可靠使用,开发一个根据输入特征智能分配计算资源的框架。最后,在系统层面,我们提出了通过与大规模基础模型的模块化集成来扩展系统功能的方法,利用其强大的认知和生成能力来最大限度地提高最终识别精度。通过在这三个层次上系统地提供解决方案,本文旨在构建下一代,鲁棒的,可扩展的AVSR系统,在现实世界中的应用具有高可靠性。
摘要:The practical deployment of Audio-Visual Speech Recognition (AVSR) systems is fundamentally challenged by significant performance degradation in real-world environments, characterized by unpredictable acoustic noise and visual interference. This dissertation posits that a systematic, hierarchical approach is essential to overcome these challenges, achieving the robust scalability at the representation, architecture, and system levels. At the representation level, we investigate methods for building a unified model that learns audio-visual features inherently robust to diverse real-world corruptions, thereby enabling generalization to new environments without specialized modules. To address architectural scalability, we explore how to efficiently expand model capacity while ensuring the adaptive and reliable use of multimodal inputs, developing a framework that intelligently allocates computational resources based on the input characteristics. Finally, at the system level, we present methods to expand the system's functionality through modular integration with large-scale foundation models, leveraging their powerful cognitive and generative capabilities to maximize final recognition accuracy. By systematically providing solutions at each of these three levels, this dissertation aims to build a next-generation, robust, and scalable AVSR system with high reliability in real-world applications.
【4】Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
标题:口语对话总结:一个描述丰富的对话数据集,用于口语对话总结
链接:https://arxiv.org/abs/2512.14687
备注:12 pages, 2 figures
摘要:最近的音频语言模型可以跟随长对话。然而,情感感知或口语对话摘要的研究受到缺乏语音,摘要和语言线索的数据的限制。我们介绍Spoken DialogSum,这是第一个将原始会话音频与事实摘要,情感丰富的摘要以及扬声器年龄,性别和情感的话语级别标签对齐的语料库。数据集分为两个阶段:首先,LLM用Switchboard风格的填充器和反向通道重写DialogSum脚本,然后用情感、音高和语速标记每个话语。第二,一个富有表现力的TTS引擎从标记的脚本中合成语音,与非语言标签对齐。Spoken DialogSum包括13,460个情感多样的对话,每个对话都有一个事实和一个以情感为中心的摘要。该数据集可在https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/在线获得。基线显示,Audio-LLM相对于级联ASR-LLM系统将情感摘要ROUGE-L提高了28%,证实了端到端语音建模的价值。
摘要:Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech, summaries, and paralinguistic cues. We introduce Spoken DialogSum, the first corpus aligning raw conversational audio with factual summaries, emotion-rich summaries, and utterance-level labels for speaker age, gender, and emotion. The dataset is built in two stages: first, an LLM rewrites DialogSum scripts with Switchboard-style fillers and back-channels, then tags each utterance with emotion, pitch, and speaking rate. Second, an expressive TTS engine synthesizes speech from the tagged scripts, aligned with paralinguistic labels. Spoken DialogSum comprises 13,460 emotion-diverse dialogues, each paired with both a factual and an emotion-focused summary. The dataset is available online at https://fatfat-emosum.github.io/EmoDialog-Sum-Audio-Samples/. Baselines show that an Audio-LLM raises emotional-summary ROUGE-L by 28% relative to a cascaded ASR-LLM system, confirming the value of end-to-end speech modeling.
【5】Linguists should learn to love speech-based deep learning models
标题:语言学家应该学会热爱基于语音的深度学习模型
链接:https://arxiv.org/abs/2512.14506
备注:Commentary on Futrell, R., & Mahowald, K. arXiv:2501.17047 (in press). How Linguistics Learned to Stop Worrying and Love the Language Models. Behavioural and Brain Sciences
摘要:Futrell和Mahowald提出了一个有用的框架,将面向技术的深度学习系统和面向解释的语言学理论联系起来。不幸的是,目标文章对基于生成文本的LLM的关注从根本上限制了与语言学的富有成效的互动,因为许多关于人类语言的有趣问题都超出了书面文本的范围。我们认为,基于音频的深度学习模型可以而且应该发挥关键作用。
摘要:Futrell and Mahowald present a useful framework bridging technology-oriented deep learning systems and explanation-oriented linguistic theories. Unfortunately, the target article's focus on generative text-based LLMs fundamentally limits fruitful interactions with linguistics, as many interesting questions on human language fall outside what is captured by written text. We argue that audio-based deep learning models can and should play a crucial role.
机器翻译由腾讯交互翻译提供,仅供参考
