微信公众号:arXiv_Daily
cs.SD语音
【1】Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
标题:通过大规模多模式通信学习突破视听感知的前沿
链接:https://arxiv.org/abs/2512.19687
摘要:我们介绍Perception Encoder Audiovisual,PE-AV,这是一个新的编码器系列,用于通过缩放对比学习进行音频和视频理解训练。PE-AV建立在PE基础上,为将表示扩展到音频做出了几项关键贡献,并原生支持跨音频-视频,音频-文本和视频-文本模式的联合嵌入。PE-AV的统一跨模态嵌入实现了语音检索等新任务,并在标准音频和视频基准测试中树立了新的艺术水平。我们通过构建一个强大的视听数据引擎来解锁这一点,该引擎可以为O(100 M)音频视频对合成高质量的字幕,从而实现跨模式的大规模监督。我们的音频数据包括语音,音乐和一般的声音效果,避免了以前工作中常见的单域限制。我们利用10对对比的目标,显示缩放跨模态和字幕类型对加强对齐,提高zero-shot性能。我们进一步开发PE-A-Frame,通过使用帧级对比目标微调PE-AV,为声音事件检测等任务实现细粒度的音频帧到文本对齐。
摘要:We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend representations to audio, and natively support joint embeddings across audio-video, audio-text, and video-text modalities. PE-AV's unified cross-modal embeddings enable novel tasks such as speech retrieval, and set a new state of the art across standard audio and video benchmarks. We unlock this by building a strong audiovisual data engine that synthesizes high-quality captions for O(100M) audio-video pairs, enabling large-scale supervision consistent across modalities. Our audio data includes speech, music, and general sound effects-avoiding single-domain limitations common in prior work. We exploit ten pairwise contrastive objectives, showing that scaling cross-modality and caption-type pairs strengthens alignment and improves zero-shot performance. We further develop PE-A-Frame by fine-tuning PE-AV with frame-level contrastive objectives, enabling fine-grained audio-frame-to-text alignment for tasks such as sound event detection.
【2】DeepGESI: A Non-Intrusive Objective Evaluation Model for Predicting Speech Intelligibility in Hearing-Impaired Listeners
标题:DeepGESI:一种用于预测听力受损听众语音可理解性的非侵入性客观评估模型
链接:https://arxiv.org/abs/2512.19374
摘要:语音清晰度评估对于许多与语音相关的应用是必不可少的。然而,大多数客观的可懂度度量是侵入性的,因为除了用于评估的降级或处理的信号之外,它们还需要干净的参考语音。此外,诸如STOI的现有度量主要是针对正常听力收听者设计的,并且它们对听力受损的语音可懂度的预测准确性仍然有限。另一方面,GESI(Gammachirp包络相似性指数)可用于评估听力受损听众的可懂度,但它也是侵入性的,因为它取决于参考信号。这一要求限制了其在现实世界场景中的适用性。 为了克服这一限制,本研究提出了DeepGESI,这是一种基于深度学习的非侵入式模型,能够准确有效地预测听障听众的语音清晰度,而不需要任何清晰的参考语音。实验结果表明,在第二次清晰度预测挑战(CPC 2)数据集的测试条件下,DeepGESI预测的GESI分数与实际GESI分数表现出很强的相关性。此外,所提出的模型实现了更快的预测速度相比,传统的方法。
摘要:Speech intelligibility assessment is essential for many speech-related applications. However, most objective intelligibility metrics are intrusive, as they require clean reference speech in addition to the degraded or processed signal for evaluation. Furthermore, existing metrics such as STOI are primarily designed for normal hearing listeners, and their predictive accuracy for hearing impaired speech intelligibility remains limited. On the other hand, the GESI (Gammachirp Envelope Similarity Index) can be used to estimate intelligibility for hearing-impaired listeners, but it is also intrusive, as it depends on reference signals. This requirement limits its applicability in real-world scenarios. To overcome this limitation, this study proposes DeepGESI, a non-intrusive deep learning-based model capable of accurately and efficiently predicting the speech intelligibility of hearing-impaired listeners without requiring any clean reference speech. Experimental results demonstrate that, under the test conditions of the 2nd Clarity Prediction Challenge(CPC2) dataset, the GESI scores predicted by DeepGESI exhibit a strong correlation with the actual GESI scores. In addition, the proposed model achieves a substantially faster prediction speed compared to conventional methods.
【3】JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
标题:JoyVoice:拟人化多说话者对话合成的长上下文条件反射
链接:https://arxiv.org/abs/2512.19090
摘要:大型语音生成模型正从单说话人、短句合成向多说话人、长对话生成发展。目前的长形式语音生成模型主要限于二元,回合为基础的互动。为了解决这个问题,我们引入了JoyVoice,这是一种新颖的拟人基础模型,旨在灵活,无边界合成多达八个扬声器。与传统的级联系统不同,JoyVoice采用统一的E2 E-Transformer-DiT架构,该架构将自回归隐藏表示直接用于扩散输入,从而实现整体端到端优化。我们进一步提出了一个MM-令牌器工作在12.5 Hz的低比特率,它集成了多任务语义和MMSE损失,有效地模拟语义和声学信息。此外,该模型通过大规模数据扰动集成了强大的文本前端处理。实验表明,JoyVoice在多语种生成(中、英、日、韩)和zero-shot语音克隆方面取得了较好的效果。JoyVoice在Seed-TTS-Eval Benchmark和多扬声器长形式对话语音克隆任务中均取得了顶级结果,展示了卓越的音频质量和泛化能力。它实现了显着的改善,韵律连续性的长形式的语音,节奏丰富的多扬声器会话,语言自然,除了卓越的可懂度。我们鼓励读者在https://jea-speech.github.io/JoyVoice上收听演示
摘要:Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice
【4】Speaker Recognition -- Wavelet Packet Based Multiresolution Feature Extraction Approach
标题:说话人识别--基于子波包的多分辨率特征提取方法
链接:https://arxiv.org/abs/2512.18902
备注:This paper was originally written in Summer 2013 and previously made available on Figshare. The present submission is uploaded for archival and citation purposes
摘要:本文提出了一种新的基于小波包的特征提取方法,用于文本无关的说话人识别。该方法利用Mel倒谱系数(MFCC)和小波包变换(WPT)相结合的方法进行特征提取,混合特征技术利用MFCC模拟人耳的优点,结合WPT的多分辨率特性和抗噪性。为了验证本文提出的方法对于文本无关说话人识别和验证的有效性,我们分别使用高斯混合模型(GMM)和隐马尔可夫模型(HMM)作为分类器。在voxforge语音语料库和CSTR US KED Timit数据库上进行了测试。加入不同信噪比水平的标准噪声信号后,对该方法进行了噪声鲁棒性评价。实验结果表明,该方法在说话人识别和说话人确认两方面都取得了较好的效果。
摘要:This paper proposes a novel Wavelet Packet based feature extraction approach for the task of text independent speaker recognition. The features are extracted by using the combination of Mel Frequency Cepstral Coefficient (MFCC) and Wavelet Packet Transform (WPT).Hybrid Features technique uses the advantage of human ear simulation offered by MFCC combining it with multi-resolution property and noise robustness of WPT. To check the validity of the proposed approach for the text independent speaker identification and verification we have used the Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM) respectively as the classifiers. The proposed paradigm is tested on voxforge speech corpus and CSTR US KED Timit database. The paradigm is also evaluated after adding standard noise signal at different level of SNRs for evaluating the noise robustness. Experimental results show that better results are achieved for the tasks of both speaker identification as well as speaker verification.
【5】Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation
标题:节奏作为稳定提示:音乐到3D舞蹈一代的节奏和节拍专家的分层混合
链接:https://arxiv.org/abs/2512.18804
摘要:音乐到3D舞蹈生成的目的是从音乐中合成现实和节奏同步的人类舞蹈。虽然现有的方法通常依赖于附加的流派标签来进一步改进舞蹈生成,但是这样的标签通常是嘈杂的、粗糙的、不可用的或不足以捕捉真实世界音乐的多样性,这可能导致节奏不对准或风格漂移。相比之下,我们观察到节奏,反映音乐节奏和节奏的核心属性,在数据集和流派之间保持相对一致,通常在60到200 BPM之间。基于这一发现,我们提出了TempoMoE,一个层次化的节奏感知专家混合模块,增强了扩散模型及其节奏感知。TempoMoE将运动专家组织成不同节奏范围的节奏结构组,多尺度节拍专家捕捉精细和长距离的节奏动态。分层节奏自适应路由从音乐特征中动态选择和融合专家,从而实现灵活的节奏对齐生成,而无需手动流派标签。大量的实验表明,TempoMoE在舞蹈质量和节奏对齐方面达到了最先进的效果。
摘要:Music to 3D dance generation aims to synthesize realistic and rhythmically synchronized human dance from music. While existing methods often rely on additional genre labels to further improve dance generation, such labels are typically noisy, coarse, unavailable, or insufficient to capture the diversity of real-world music, which can result in rhythm misalignment or stylistic drift. In contrast, we observe that tempo, a core property reflecting musical rhythm and pace, remains relatively consistent across datasets and genres, typically ranging from 60 to 200 BPM. Based on this finding, we propose TempoMoE, a hierarchical tempo-aware Mixture-of-Experts module that enhances the diffusion model and its rhythm perception. TempoMoE organizes motion experts into tempo-structured groups for different tempo ranges, with multi-scale beat experts capturing fine- and long-range rhythmic dynamics. A Hierarchical Rhythm-Adaptive Routing dynamically selects and fuses experts from music features, enabling flexible, rhythm-aligned generation without manual genre labels. Extensive experiments demonstrate that TempoMoE achieves state-of-the-art results in dance quality and rhythm alignment.
【6】Reliable Audio Deepfake Detection in Variable Conditions via Quantum-Kernel SVMs
标题:通过量子核支持器在可变条件下进行可靠的音频深度伪造检测
链接:https://arxiv.org/abs/2512.18797
备注:This paper is accepted in ICDM 2025-MLC workshop
摘要:当标记数据稀缺且记录条件变化时,检测合成语音具有挑战性。现有的端到端深度模型通常过拟合或无法泛化,虽然内核方法可以保持竞争力,但它们的性能在很大程度上取决于所选择的内核。在这里,我们证明了在音频deepfake检测中使用量子内核可以降低误报率,而不会增加模型大小。量子特征映射将数据嵌入到高维希尔伯特空间中,使表达相似性度量和紧凑分类器的使用成为可能。基于这一动机,我们使用相同的mel谱图预处理和四个语料库(ASVspoof 2019 LA,ASVspoof 5(2024),ADD 23和In-the-Wild集)的分层5倍交叉验证将量子核SVM(QSVM)与经典SVM进行比较。QSVM的等误率(EER)始终较低:ASVspoof 5(2024)为0.183 vs. 0.299,ADD 23为0.081 vs. 0.188,ASVspoof 2019为0.346 vs. 0.399,野外为0.355 vs. 0.413。在EER工作点(FPR等于FNR),这些分别对应于0.116(38.8%)、0.107(56.9%)、0.053(13.3%)和0.058(14.0%)的绝对假阳性减少。我们还报告了跨交叉验证折叠和基于边缘的类分离措施的结果是如何一致的,对两个模型使用相同的设置。唯一的修改是内核;特征和SVM保持不变,没有引入额外的可训练参数,量子内核在传统计算机上计算。
摘要:Detecting synthetic speech is challenging when labeled data are scarce and recording conditions vary. Existing end-to-end deep models often overfit or fail to generalize, and while kernel methods can remain competitive, their performance heavily depends on the chosen kernel. Here, we show that using a quantum kernel in audio deepfake detection reduces falsepositive rates without increasing model size. Quantum feature maps embed data into high-dimensional Hilbert spaces, enabling the use of expressive similarity measures and compact classifiers. Building on this motivation, we compare quantum-kernel SVMs (QSVMs) with classical SVMs using identical mel-spectrogram preprocessing and stratified 5-fold cross-validation across four corpora (ASVspoof 2019 LA, ASVspoof 5 (2024), ADD23, and an In-the-Wild set). QSVMs achieve consistently lower equalerror rates (EER): 0.183 vs. 0.299 on ASVspoof 5 (2024), 0.081 vs. 0.188 on ADD23, 0.346 vs. 0.399 on ASVspoof 2019, and 0.355 vs. 0.413 In-the-Wild. At the EER operating point (where FPR equals FNR), these correspond to absolute false-positiverate reductions of 0.116 (38.8%), 0.107 (56.9%), 0.053 (13.3%), and 0.058 (14.0%), respectively. We also report how consistent the results are across cross-validation folds and margin-based measures of class separation, using identical settings for both models. The only modification is the kernel; the features and SVM remain unchanged, no additional trainable parameters are introduced, and the quantum kernel is computed on a conventional computer.
【7】Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform
标题:Spark:一种基于离散小波变换的文本到语音扩散模型水印
链接:https://arxiv.org/abs/2512.18791
摘要:文语转换(TTS)扩散模型产生高质量的语音,这对模型的知识产权保护和合法使用的语音跟踪提出了挑战。音频水印是一种很有前途的解决方案。然而,由于各种TTS扩散模型之间的结构差异,现有的水印方法往往是针对特定模型设计的,降低了音频质量,限制了它们的实用性。为了解决这个难题,本文提出了一种通用的TTS扩散模型的水印方案,称为Spark。这是通过设计一个轻量级的水印嵌入框架,在所有TTS扩散模型共享的共同的反向扩散范例。为了减轻对音频质量的影响,Spark利用离散小波变换(DWT)将水印嵌入到音频的相对稳定的低频区域中,这确保了水印-音频的无缝集成,并且在反向扩散过程中不易被去除。大量的实验进行评估的音频质量和水印性能在各种模拟现实世界的攻击场景。实验结果表明,Smark在音频质量和水印提取精度方面都取得了优异的性能。
摘要:Text-to-Speech (TTS) diffusion models generate high-quality speech, which raises challenges for the model intellectual property protection and speech tracing for legal use. Audio watermarking is a promising solution. However, due to the structural differences among various TTS diffusion models, existing watermarking methods are often designed for a specific model and degrade audio quality, which limits their practical applicability. To address this dilemma, this paper proposes a universal watermarking scheme for TTS diffusion models, termed Smark. This is achieved by designing a lightweight watermark embedding framework that operates in the common reverse diffusion paradigm shared by all TTS diffusion models. To mitigate the impact on audio quality, Smark utilizes the discrete wavelet transform (DWT) to embed watermarks into the relatively stable low-frequency regions of the audio, which ensures seamless watermark-audio integration and is resistant to removal during the reverse diffusion process. Extensive experiments are conducted to evaluate the audio quality and watermark performance in various simulated real-world attack scenarios. The experimental results show that Smark achieves superior performance in both audio quality and watermark extraction accuracy.
【8】X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
标题:X-Talk:模块化语音对话系统被低估的潜力
链接:https://arxiv.org/abs/2512.18706
备注:14 pages
摘要:我们提出了X-Talk,这是一个开源框架,支持LLM驱动的语音到语音(S2 S)系统的解耦,模块化设计。虽然主流趋势倾向于端到端(E2 E)建模以优化信息流,但这些“全模型”通常难以平衡单个网络中复杂语音任务的竞争目标。X-Talk通过展示系统优化的级联流水线可以在不牺牲模块灵活性的情况下实现亚秒级延迟来挑战这种模式。我们的框架无缝集成了专门的前端组件(例如,VAD,语音增强)和多种理解模型(例如,ASR,情感和环境声音分析)与LLM功能,如检索增强生成(RAG)和工具使用。通过振兴级联方法,X-Talk强调了模块化S2 S系统被低估的潜力,并为未来的研究和应用提供了坚实的基础。
摘要:We present X-Talk, an open-source framework that champions a decoupled, modular design for LLM-driven speech-to-speech (S2S) systems. While the dominant trend favors end-to-end (E2E) modeling to optimize information flow, these "omni-models" often struggle to balance the competing objectives of complex speech tasks within a single network. X-Talk challenges this paradigm by demonstrating that a systematically optimized cascaded pipeline can achieve sub-second latency without sacrificing modular flexibility. Our framework seamlessly integrates specialized front-end components (e.g., VAD, speech enhancement) and diverse understanding models (e.g., ASR, emotion, and environmental sound analysis) with LLM capabilities like retrieval-augmented generation (RAG) and tool use. By revitalizing the cascaded approach, X-Talk highlights the underestimated potential of modular S2S systems and provides a robust foundation for future research and applications.
【9】Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
标题:TTC中的任务载体:迈向情感表达方言语音合成
链接:https://arxiv.org/abs/2512.18699
摘要:文语转换(TTS)技术的最新进展在自然度和可懂度方面取得了显著的进步。在这些成就的基础上,研究越来越多地转向增强生成语音的表达能力,如方言和情感TTS。然而,结合方言和情感的跨风格合成仍然具有挑战性,并且在很大程度上未被探索,主要是由于带有情感标签的方言数据的稀缺。为了解决这个问题,我们提出了层次表达向量(HE-Vector),一个两阶段的情感方言TTS方法。在第一阶段,我们构建不同的任务向量来独立地对方言和情感风格进行建模,然后通过调整它们的权重来增强单一风格的合成,这种方法我们称为表达向量(E-Vector)。第二阶段,我们分层整合这些向量,实现可控的情感表达方言合成,而不需要联合标记的数据,对应于分层表达向量(HE-Vector)。实验结果表明,HE-Vectors在方言语音合成中取得了优异的性能,在zero-shot环境下能够合成具有情感表达的方言语音。
摘要:Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted toward enhancing the expressiveness of generated speech, such as dialectal and emotional TTS. However, cross-style synthesis combining both dialect and emotion remains challenging and largely unexplored, mainly due to the scarcity of dialectal data with emotional labels. To address this, we propose Hierarchical Expressive Vector (HE-Vector), a two-stage method for Emotional Dialectal TTS. In the first stage, we construct different task vectors to model dialectal and emotional styles independently, and then enhance single-style synthesis by adjusting their weights, a method we refer to as Expressive Vector (E-Vector). For the second stage, we hierarchically integrate these vectors to achieve controllable emotionally expressive dialect synthesis without requiring jointly labeled data, corresponding to Hierarchical Expressive Vector (HE-Vector). Experimental results demonstrate that HE-Vectors achieve superior performance in dialect synthesis, and promising results in synthesizing emotionally expressive dialectal speech in a zero-shot setting.
【10】Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
标题:可解释转换器-CNN融合用于抗噪语音情感识别
链接:https://arxiv.org/abs/2512.18298
摘要:语音情感识别(SER)系统的性能往往会下降时,暴露于不可预测的声音干扰在现实世界中发现的环境。此外,深度学习模型的不透明性阻碍了它们在信任敏感应用程序中的采用。为了弥合这一差距,我们提出了一个混合变换器-CNN框架,将Wav 2 Vec 2.0的上下文建模与一维卷积神经网络的谱稳定性统一起来。我们的双流架构处理原始波形以捕获长距离时间依赖性,同时通过自定义的Attentive Temporal Pooling机制提取抗噪声频谱特征(MFCC,ZCR,RMSE)。我们对四个不同的基准数据集进行了广泛的验证:RAVDESS,TESS,SAVEE和CREMA-D。为了严格测试鲁棒性,我们使用来自SAS-KIIT数据集的真实噪声分布对模型进行了非平稳声学干扰。所提出的框架在所有数据集上表现出卓越的泛化能力和最先进的准确性,在现实环境干扰下显著优于单分支基线。此外,我们通过将SHAP和Score-CAM集成到评估管道中来解决“黑盒”问题。这些工具提供了精细的视觉解释,揭示了模型如何在时间和光谱线索之间战略性地转移注意力,以在复杂的环境噪声中保持可靠性。
摘要:Speech Emotion Recognition (SER) systems often degrade in performance when exposed to the unpredictable acoustic interference found in real-world environments. Additionally, the opacity of deep learning models hinders their adoption in trust-sensitive applications. To bridge this gap, we propose a Hybrid Transformer-CNN framework that unifies the contextual modeling of Wav2Vec 2.0 with the spectral stability of 1D-Convolutional Neural Networks. Our dual-stream architecture processes raw waveforms to capture long-range temporal dependencies while simultaneously extracting noise-resistant spectral features (MFCC, ZCR, RMSE) via a custom Attentive Temporal Pooling mechanism. We conducted extensive validation across four diverse benchmark datasets: RAVDESS, TESS, SAVEE, and CREMA-D. To rigorously test robustness, we subjected the model to non-stationary acoustic interference using real-world noise profiles from the SAS-KIIT dataset. The proposed framework demonstrates superior generalization and state-of-the-art accuracy across all datasets, significantly outperforming single-branch baselines under realistic environmental interference. Furthermore, we address the ``black-box" problem by integrating SHAP and Score-CAM into the evaluation pipeline. These tools provide granular visual explanations, revealing how the model strategically shifts attention between temporal and spectral cues to maintain reliability in the presence of complex environmental noise.
【11】AutoSchA: Automatic Hierarchical Music Representations via Multi-Relational Node Isolation
标题:AutoSchA:通过多关系节点隔离的自动分层音乐表示
链接:https://arxiv.org/abs/2512.18232
摘要:层次表示为分析许多音乐流派提供了强大而有原则的方法。这种表现在音乐理论中得到了广泛的研究,例如通过申克分析(SchA)。然而,分层音乐分析是高度成本密集型的;分析一段音乐需要训练有素的专家花费大量的时间和精力。以计算机可读格式表示层次分析是另一个挑战。鉴于分层深度学习的最新发展和越来越多的计算机可读数据,将此类工作扩展为自动分层表示框架大有希望。因此,本文介绍了一种新的方法,AutoSchA,它扩展了图神经网络(GNNs)的分层音乐分析的最新发展。AutoSchA具有三个关键贡献:1)用于分层音乐表示的新的图学习框架,2)基于节点隔离的新的图池机制,直接优化学习的池分配,以及3)集成自动分层音乐分析的最新架构。我们表明,在一系列的实验中,AutoSchA执行人类专家在分析巴洛克赋格主题。
摘要:Hierarchical representations provide powerful and principled approaches for analyzing many musical genres. Such representations have been broadly studied in music theory, for instance via Schenkerian analysis (SchA). Hierarchical music analyses, however, are highly cost-intensive; the analysis of a single piece of music requires a great deal of time and effort from trained experts. The representation of hierarchical analyses in a computer-readable format is a further challenge. Given recent developments in hierarchical deep learning and increasing quantities of computer-readable data, there is great promise in extending such work for an automatic hierarchical representation framework. This paper thus introduces a novel approach, AutoSchA, which extends recent developments in graph neural networks (GNNs) for hierarchical music analysis. AutoSchA features three key contributions: 1) a new graph learning framework for hierarchical music representation, 2) a new graph pooling mechanism based on node isolation that directly optimizes learned pooling assignments, and 3) a state-of-the-art architecture that integrates such developments for automatic hierarchical music analysis. We show, in a suite of experiments, that AutoSchA performs comparably to human experts when analyzing Baroque fugue subjects.
【12】A Data-Centric Approach to Generalizable Speech Deepfake Detection
标题:一种以数据为中心的可推广语音深度伪造检测方法
链接:https://arxiv.org/abs/2512.18210
摘要:在语音深度伪造检测(SDD)中实现鲁棒的泛化仍然是一个主要挑战,因为模型通常无法检测到看不见的伪造方法。虽然研究集中在以模型为中心和以算法为中心的解决方案上,但数据组合的影响往往没有得到充分的探索。本文提出了一种以数据为中心的方法,从两个实际的角度分析SDD数据景观:构建单个数据集和聚合多个数据集。为了解决第一个角度,我们进行了大规模的实证研究,以表征SDD的数据缩放律,量化源和发电机多样性的影响。为了解决第二个问题,我们提出了多样性优化采样策略(DOSS),这是一个原则性的框架,用于混合异构数据,有两种实现:DOSS选择(修剪)和DOSS权重(重新加权)。我们的实验表明,DOSS-Select优于天真的聚合基线,而使用的总可用数据的3%。此外,我们的最终模型使用最佳DOSS-Weight策略在12 k小时的策展数据池上进行训练,实现了最先进的性能,在公共基准测试和各种商业API的新挑战集上表现优于大规模基线,具有更高的数据和模型效率。
摘要:Achieving robust generalization in speech deepfake detection (SDD) remains a primary challenge, as models often fail to detect unseen forgery methods. While research has focused on model-centric and algorithm-centric solutions, the impact of data composition is often underexplored. This paper proposes a data-centric approach, analyzing the SDD data landscape from two practical perspectives: constructing a single dataset and aggregating multiple datasets. To address the first perspective, we conduct a large-scale empirical study to characterize the data scaling laws for SDD, quantifying the impact of source and generator diversity. To address the second, we propose the Diversity-Optimized Sampling Strategy (DOSS), a principled framework for mixing heterogeneous data with two implementations: DOSS-Select (pruning) and DOSS-Weight (re-weighting). Our experiments show that DOSS-Select outperforms the naive aggregation baseline while using only 3% of the total available data. Furthermore, our final model, trained on a 12k-hour curated data pool using the optimal DOSS-Weight strategy, achieves state-of-the-art performance, outperforming large-scale baselines with greater data and model efficiency on both public benchmarks and a new challenge set of various commercial APIs.
【13】Influence of string register locations on vibratos among violoncellists
标题:弦乐套位位置对大提琴演奏家振动的影响
链接:https://arxiv.org/abs/2512.18162
摘要:这项研究分析了颤音如何随着手指位置沿大提琴琴弦变化而变化。通过对94个片段的分析,我们发现将手指移向琴桥会大大增加颤音的深度(p =0.6902,p=1.408\cdot 10^{-14}$)。然而,表演者的物理手指振幅同时减小($ρ=-0.6391 $,$p=4.172\cdot 10^{-12}$)。这表明演奏者在较高的位置减少手指运动,但不足以抵消更大的音高偏差,揭示了补偿性颤音行为的存在和限制。
摘要:This study analyzes how vibrato changes with finger position along the cello string. Examining 94 excerpts, we found moving the finger toward the bridge strongly increases acoustic vibrato depth ($ρ=0.6902$, $p=1.408\cdot 10^{-14}$). However, the performer's physical finger amplitude simultaneously decreases ($ρ=-0.6391$, $p=4.172\cdot 10^{-12}$). This shows players reduce finger motion in higher positions, but not enough to counteract the greater pitch deviation there, revealing both the presence and limits of compensatory vibrato behavior.
【14】Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
标题:让模型学会感受:符号音乐情感识别的模式引导调性注入
链接:https://arxiv.org/abs/2512.17946
备注:Accepted by AAAI 2026
摘要:音乐情感识别是符号音乐理解中的一个关键任务。最近的方法通过微调大规模预训练模型(例如,MIDIBERT,符号音乐理解的基准),以映射音乐语义的情感标签。虽然这些模型有效地捕捉分布的音乐语义,他们往往忽略了音调结构,特别是音乐模式,根据音乐心理学的情感感知中发挥着至关重要的作用。在本文中,我们调查的代表性能力MIDIBERT,并确定其局限性,在捕捉模式的情感协会。为了解决这个问题,我们提出了一种模式引导增强(MoGE)策略,该策略将对模式的心理见解纳入模型中。具体来说,我们首先进行模式增强分析,这表明MIDIBERT未能有效地编码情感模式的相关性。然后,我们确定MIDIBERT内的情感相关性最低的层,并引入一个模式引导的逐行线性调制注入(MoFi)框架来注入显式模式特征,从而增强模型在情感表示和推理方面的能力。在EMOPIA和VGSTO数据集上的大量实验表明,我们的模式注入策略显着提高了SMER性能,分别达到75.2%和59.1%的准确率。这些结果验证了模式引导建模在符号音乐情感识别中的有效性。
摘要:Music emotion recognition is a key task in symbolic music understanding (SMER). Recent approaches have shown promising results by fine-tuning large-scale pre-trained models (e.g., MIDIBERT, a benchmark in symbolic music understanding) to map musical semantics to emotional labels. While these models effectively capture distributional musical semantics, they often overlook tonal structures, particularly musical modes, which play a critical role in emotional perception according to music psychology. In this paper, we investigate the representational capacity of MIDIBERT and identify its limitations in capturing mode-emotion associations. To address this issue, we propose a Mode-Guided Enhancement (MoGE) strategy that incorporates psychological insights on mode into the model. Specifically, we first conduct a mode augmentation analysis, which reveals that MIDIBERT fails to effectively encode emotion-mode correlations. We then identify the least emotion-relevant layer within MIDIBERT and introduce a Mode-guided Feature-wise linear modulation injection (MoFi) framework to inject explicit mode features, thereby enhancing the model's capability in emotional representation and inference. Extensive experiments on the EMOPIA and VGMIDI datasets demonstrate that our mode injection strategy significantly improves SMER performance, achieving accuracies of 75.2% and 59.1%, respectively. These results validate the effectiveness of mode-guided modeling in symbolic music emotion recognition.
【15】chatter: a Python library for applying information theory and AI/ML models to animal communication
标题:chatter:将信息理论和AI/ML模型应用于动物交流的Python库
链接:https://arxiv.org/abs/2512.17935
摘要:动物交流的研究通常涉及将单位分类为类型(例如鸣禽的音节或座头鲸的音符)。虽然这种方法在许多情况下是有用的,但它必然会忽略实际通信系统中存在的复杂性和细微差别。Chatter是一个新的Python库,用于使用信息论和现代机器学习技术分析连续潜在空间中的动物通信。它在分类学上是不可知的,已经用鸟类、蝙蝠、鲸鱼和灵长类动物的发声进行了测试。通过利用各种不同的架构,包括变分自动编码器和Vision Transformers,chatter将声音序列表示为高维潜在空间中的轨迹,从而绕过了手动或自动分类单元的需要。该库提供了一个端到端的工作流程-从预处理和分割到模型训练和特征提取-使研究人员能够量化声音序列的复杂性,可预测性,相似性和新颖性。
摘要:The study of animal communication often involves categorizing units into types (e.g. syllables in songbirds, or notes in humpback whales). While this approach is useful in many cases, it necessarily flattens the complexity and nuance present in real communication systems. chatter is a new Python library for analyzing animal communication in continuous latent space using information theory and modern machine learning techniques. It is taxonomically agnostic, and has been tested with the vocalizations of birds, bats, whales, and primates. By leveraging a variety of different architectures, including variational autoencoders and vision transformers, chatter represents vocal sequences as trajectories in high-dimensional latent space, bypassing the need for manual or automatic categorization of units. The library provides an end-to-end workflow -- from preprocessing and segmentation to model training and feature extraction -- that enables researchers to quantify the complexity, predictability, similarity, and novelty of vocal sequences.
【16】Real-Time Streamable Generative Speech Restoration with Flow Matching
标题:具有流匹配的实时可流生成语音恢复
链接:https://arxiv.org/abs/2512.19442
备注:This work has been submitted to the IEEE for possible publication
摘要:近年来,基于扩散的生成模型对语音处理领域产生了巨大的影响,表现出很高的语音自然度,并产生了一个新的研究方向。然而,它们在实时通信中的应用仍然落后,因为它们的计算量很大,涉及多次调用大型DNN。 在这里,我们提出了Stream.FM,一个基于帧因果流的生成模型,算法延迟为32毫秒(ms),总延迟为48 ms,为实时通信中的生成语音处理铺平了道路。我们提出了一个缓冲流推理方案和一个优化的DNN架构,展示了学习的几步数值求解器如何在固定的计算预算下提高输出质量,探索模型权重压缩以找到计算/质量权衡的有利点,并为语音增强任务贡献一个总延迟为24 ms的模型变体。 我们的工作超越了理论上的限制,表明高质量的流生成语音处理可以在当今可用的消费者GPU上实现。FM可以以流的方式解决各种语音处理任务:语音增强、去混响、编解码器后滤波、带宽扩展、STFT相位恢复和Mel声码。正如我们通过全面评估和MUSHRA听力测试所验证的那样,Stream.FM建立了最先进的生成流语音恢复,与非流变体相比,质量仅出现合理的降低,并且在生成流语音增强方面优于我们最近的工作(扩散缓冲区),同时以较低的延迟运行。
摘要:Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their application in real-time communication is, however, still lagging behind due to their computation-heavy nature involving multiple calls of large DNNs. Here, we present Stream.FM, a frame-causal flow-based generative model with an algorithmic latency of 32 milliseconds (ms) and a total latency of 48 ms, paving the way for generative speech processing in real-time communication. We propose a buffered streaming inference scheme and an optimized DNN architecture, show how learned few-step numerical solvers can boost output quality at a fixed compute budget, explore model weight compression to find favorable points along a compute/quality tradeoff, and contribute a model variant with 24 ms total latency for the speech enhancement task. Our work looks beyond theoretical latencies, showing that high-quality streaming generative speech processing can be realized on consumer GPUs available today. Stream.FM can solve a variety of speech processing tasks in a streaming fashion: speech enhancement, dereverberation, codec post-filtering, bandwidth extension, STFT phase retrieval, and Mel vocoding. As we verify through comprehensive evaluations and a MUSHRA listening test, Stream.FM establishes a state-of-the-art for generative streaming speech restoration, exhibits only a reasonable reduction in quality compared to a non-streaming variant, and outperforms our recent work (Diffusion Buffer) on generative streaming speech enhancement while operating at a lower latency.
【17】Sonified Quantum Seizures. Sonification of time series in epileptic seizures and simulation of seizures via quantum modelling
标题:声波量子癫痫发作。癫痫发作时间序列的声化以及通过量子建模模拟癫痫发作
链接:https://arxiv.org/abs/2512.19272
备注:Presented at ISQCMC '25: 3rd International Symposium on Quantum Computing and Musical Creativity
摘要:我们应用声化策略和量子计算来分析癫痫发作。我们首先从选择的通道(从真实的ECoG数据)声化信号,获得复调序列。然后,我们提出了两种量子方法来模拟一个类似的发作,我们sonify的结果。声化的比较可以提示真实数据和模拟之间的相似性和差异,有助于改进\textit{in silico}模型。这是一种开创性的方法,展示了量子计算和超声的结合如何拓宽真实数据调查的视角,并帮助定义了一个新的测试平台,用于分析和预测癫痫发作。
摘要:We apply sonification strategies and quantum computing to the analysis of an episode of seizure. We first sonify the signal from a selection of channels (from real ECoG data), obtaining a polyphonic sequence. Then, we propose two quantum approaches to simulate a similar episode of seizure, and we sonify the results. The comparison of sonifications can give hints on similarities and discrepancies between real data and simulations, helping refine the \textit{in silico} model. This is a pioneering approach, showing how the combination of quantum computing and sonification can broaden the perspective of real-data investigation, and helping define a new test bench for analysis and prediction of seizures.
【18】Phoneme-based speech recognition driven by large language models and sampling marginalization
标题:大型语言模型和抽样边缘化驱动的基于音素的语音识别
链接:https://arxiv.org/abs/2512.18371
备注:Published at NCMMSC 2025, in Chinese language
摘要:近年来,基于大语言模型的音素-字形(LLM-P2 G)方法在语音识别任务中表现出了优异的性能,成为替代传统WFST解码方法的可行方向。该框架通过音素预测和文本生成的两阶段建模,兼顾了识别精度和系统可扩展性。然而,现有的LLM-P2 G采用Top-K Marginalized(TKM)训练策略,其候选音素序列依赖于波束搜索生成,存在路径多样性不足、训练效率低、资源开销大等问题。为此,提出了一种采样边缘化训练策略(Sampling-K Marginalized,SKM),用随机采样代替波束搜索生成候选路径,提高了边缘化建模和训练效率。在波兰和德国数据集上进行了实验,结果表明,SKM在保持模型复杂度的同时,进一步提高了模型的学习收敛速度和识别性能。与使用投影仪结合大语言模型(SpeechLLM)的语音识别方法的比较实验也表明,SKM驱动的LLM-P2 G在识别准确性和结构简单性方面具有更多优势。研究结果验证了该方法在跨语言语音识别系统中的实用价值和应用潜力。
摘要:Recently, the Large Language Model-based Phoneme-to-Grapheme (LLM-P2G) method has shown excellent performance in speech recognition tasks and has become a feasible direction to replace the traditional WFST decoding method. This framework takes into account both recognition accuracy and system scalability through two-stage modeling of phoneme prediction and text generation. However, the existing LLM-P2G adopts the Top-K Marginalized (TKM) training strategy, and its candidate phoneme sequences rely on beam search generation, which has problems such as insufficient path diversity, low training efficiency, and high resource overhead. To this end, this paper proposes a sampling marginalized training strategy (Sampling-K Marginalized, SKM), which replaces beam search with random sampling to generate candidate paths, improving marginalized modeling and training efficiency. Experiments were conducted on Polish and German datasets, and the results showed that SKM further improved the model learning convergence speed and recognition performance while maintaining the complexity of the model. Comparative experiments with a speech recognition method that uses a projector combined with a large language model (SpeechLLM) also show that the SKM-driven LLM-P2G has more advantages in recognition accuracy and structural simplicity. The study verified the practical value and application potential of this method in cross-language speech recognition systems.
【19】MEGState: Phoneme Decoding from Magnetoencephalography Signals
标题:MEGState:脑磁图信号的音素解码
链接:https://arxiv.org/abs/2512.17978
备注:Accepted for presentation at LibriBrain Competition, NeurIPS 2025
摘要:从非侵入性神经记录中解码语言上有意义的表示仍然是神经语音解码的核心挑战。在现有的神经成像方式,脑磁图(MEG)提供了一个安全和可重复的手段映射语音相关的皮层动力学,但其低信噪比和高时间维度继续阻碍鲁棒解码。在这项工作中,我们介绍MEGState,一种新的架构,从MEG信号的音素解码,捕捉细粒度的听觉刺激引起的皮层反应。在LibriBrain数据集上进行的大量实验表明,MEGState在多个评估指标上始终优于基线模型。这些发现突出了基于MEG的音素解码作为非侵入性脑机接口语音的可扩展途径的潜力。
摘要:Decoding linguistically meaningful representations from non-invasive neural recordings remains a central challenge in neural speech decoding. Among available neuroimaging modalities, magnetoencephalography (MEG) provides a safe and repeatable means of mapping speech-related cortical dynamics, yet its low signal-to-noise ratio and high temporal dimensionality continue to hinder robust decoding. In this work, we introduce MEGState, a novel architecture for phoneme decoding from MEG signals that captures fine-grained cortical responses evoked by auditory stimuli. Extensive experiments on the LibriBrain dataset demonstrate that MEGState consistently surpasses baseline model across multiple evaluation metrics. These findings highlight the potential of MEG-based phoneme decoding as a scalable pathway toward non-invasive brain-computer interfaces for speech.
【20】LIWhiz: A Non-Intrusive Lyric Intelligibility Prediction System for the Cadenza Challenge
标题:LIWhiz:Cadenza挑战赛的非侵入性抒情可理解性预测系统
链接:https://arxiv.org/abs/2512.17937
摘要:我们提出了LIWhiz,一个非侵入式的抒情可懂度预测系统提交给ICASSP 2026华彩乐段挑战赛。LIWhiz利用Whisper进行强大的特征提取和可训练的后端进行分数预测。在Cadenza歌词可懂度预测(CLIP)评估集上进行测试,LIWhiz在基于STOI的基线上实现了22.4%的相对均方根误差降低,从而大大改善了归一化互相关。
摘要:We present LIWhiz, a non-intrusive lyric intelligibility prediction system submitted to the ICASSP 2026 Cadenza Challenge. LIWhiz leverages Whisper for robust feature extraction and a trainable back-end for score prediction. Tested on the Cadenza Lyric Intelligibility Prediction (CLIP) evaluation set, LIWhiz achieves a 22.4% relative root mean squared error reduction over the STOI-based baseline, yielding a substantial improvement in normalized cross-correlation.
【21】Continual Learning for Acoustic Event Classification
标题:声学事件分类的持续学习
链接:https://arxiv.org/abs/2512.17932
备注:Master project report
摘要:考虑到对计算资源的限制(例如,模型大小、运行内存)。为了缓解这个问题,我们提出了两个新的多样性意识的增量学习方法的口语关键词发现和环境声音分类。我们的方法通过测量每个样本的分类不确定性来选择用于训练的历史数据。对于口语关键词定位应用程序,建议的RK方法引入了一个多样性感知采样器,通过计算分类不确定性从历史和传入的关键词中选择一个不同的集合。因此,RK方法可以在不忘记先前知识的情况下渐进地学习新任务。此外,RK方法还提出了数据增强和知识蒸馏损失函数,以实现边缘设备上的高效内存管理。对于环境声音分类应用程序,我们通过观察数据的分类概率如何随添加到分类器嵌入的并行扰动而波动来测量不确定性。通过这种方式,与向原始数据添加扰动相比,可以显著降低计算成本。实验结果表明,所提出的RK方法在Google Speech Command数据集上的平均准确率比最佳基线提高了4.2%,所需内存更少。在DCASE 2019 Task 1和ESC-50数据集上的实验结果表明,我们提出的方法在分类精度和计算效率方面优于基线持续学习方法,表明我们的方法可以有效地增量学习新类,而不会出现灾难性遗忘问题。
摘要:Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device acoustic event classification given the restrictions on computation resources (e.g., model size, running memory). To alleviate such an issue, we propose two novel diversity-aware incremental learning method for Spoken Keyword Spotting and Environmental Sound Classification. Our method selects the historical data for the training by measuring the per-sample classification uncertainty. For the Spoken Keyword Spotting application, the proposed RK approach introduces a diversity-aware sampler to select a diverse set from historical and incoming keywords by calculating classification uncertainty. As a result, the RK approach can incrementally learn new tasks without forgetting prior knowledge. Besides, the RK approach also proposes data augmentation and knowledge distillation loss function for efficient memory management on the edge device. For the Environmental Sound Classification application, we measure the uncertainty by observing how the classification probability of data fluctuates against the parallel perturbations added to the classifier embedding. In this way, the computation cost can be significantly reduced compared with adding perturbation to the raw data. Experimental results show that the proposed RK approach achieves 4.2% absolute improvement in terms of average accuracy over the best baseline on Google Speech Command dataset with less required memory. Experimental results on the DCASE 2019 Task 1 and ESC-50 dataset show that our proposed method outperforms baseline continual learning methods on classification accuracy and computational efficiency, indicating our method can efficiently and incrementally learn new classes without the catastrophic forgetting problem for on-device environmental sound classification
【1】Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization
标题:通过多码本载体量化进行知识提炼增强完全扩展的端到端语音识别
链接:https://arxiv.org/abs/2512.18967
备注:Accepted to ASRU 2025
摘要:传统的自动语音识别(ASR)模型通常将输出作为缺少标点符号和大写的规范化文本,从而需要后处理模型来增强可读性。然而,由于级联系统设计,这种方法引入了额外的复杂性和延迟。为了应对这一挑战,开发能够直接预测标点符号和大写字母的端到端(E2 E)ASR模型的趋势日益明显,尽管这一领域仍有待探索。在本文中,我们提出了一个增强的完全格式化的E2 E ASR模型,利用知识蒸馏(KD)通过多码本矢量量化(MVQ)。实验结果表明,我们的模型显着优于以前的作品在单词错误率(WER),无论有没有标点符号和大写,并在标点符号错误率(PER)。在LibriSpeech-PC测试干净和测试其他子集上的评估表明,我们的模型达到了最先进的结果。
摘要:Conventional automatic speech recognition (ASR) models typically produce outputs as normalized texts lacking punctuation and capitalization, necessitating post-processing models to enhance readability. This approach, however, introduces additional complexity and latency due to the cascaded system design. In response to this challenge, there is a growing trend to develop end-to-end (E2E) ASR models capable of directly predicting punctuation and capitalization, though this area remains underexplored. In this paper, we propose an enhanced fully formatted E2E ASR model that leverages knowledge distillation (KD) through multi-codebook vector quantization (MVQ). Experimental results demonstrate that our model significantly outperforms previous works in word error rate (WER) both with and without punctuation and capitalization, and in punctuation error rate (PER). Evaluations on the LibriSpeech-PC test-clean and test-other subsets show that our model achieves state-of-the-art results.
【2】MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
标题:MeanFlow-PSE:使用Mean Flow的一步生成目标说话人提取
链接:https://arxiv.org/abs/2512.18572
备注:6 pages, 2 figures, 2 tables
摘要:目标说话人提取(TSE)的目的是使用辅助信息(如参考话语)从多说话人混合语音中分离出所需说话人的语音。虽然扩散和流匹配模型的最新进展提高了TSE性能,但这些方法通常需要多步采样,这限制了它们在低延迟设置中的实用性。在这项工作中,我们提出了MeanFlow-TSE,这是一个用平均流目标训练的一步生成TSE框架,可以在没有迭代细化的情况下快速和高质量地生成。建立在AD-FlowTSE范例上,我们的方法定义了由混合比(MR)控制的背景和目标源之间的流。Libri 2 Mix语料库上的实验表明,我们的方法优于现有的基于扩散和流匹配的TSE模型在分离质量和感知指标,而只需要一个单一的推理步骤。这些结果表明,平均流引导的一步生成提供了一个有效的和高效的替代实时目标说话人提取。代码可在https://github.com/rikishimizu/MeanFlow-TSE上获得。
摘要:Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved TSE performance, these methods typically require multi-step sampling, which limits their practicality in low-latency settings. In this work, we propose MeanFlow-TSE, a one-step generative TSE framework trained with mean-flow objectives, enabling fast and high-quality generation without iterative refinement. Building on the AD-FlowTSE paradigm, our method defines a flow between the background and target source that is governed by the mixing ratio (MR). Experiments on the Libri2Mix corpus show that our approach outperforms existing diffusion- and flow-matching-based TSE models in separation quality and perceptual metrics while requiring only a single inference step. These results demonstrate that mean-flow-guided one-step generation offers an effective and efficient alternative for real-time target speaker extraction. Code is available at https://github.com/rikishimizu/MeanFlow-TSE.
【3】Phoneme-based speech recognition driven by large language models and sampling marginalization
标题:大型语言模型和抽样边缘化驱动的基于音素的语音识别
链接:https://arxiv.org/abs/2512.18371
备注:Published at NCMMSC 2025, in Chinese language
摘要:近年来,基于大语言模型的音素-字形(LLM-P2 G)方法在语音识别任务中表现出了优异的性能,成为替代传统WFST解码方法的可行方向。该框架通过音素预测和文本生成的两阶段建模,兼顾了识别精度和系统可扩展性。然而,现有的LLM-P2 G采用Top-K Marginalized(TKM)训练策略,其候选音素序列依赖于波束搜索生成,存在路径多样性不足、训练效率低、资源开销大等问题。为此,提出了一种采样边缘化训练策略(Sampling-K Marginalized,SKM),用随机采样代替波束搜索生成候选路径,提高了边缘化建模和训练效率。在波兰和德国数据集上进行了实验,结果表明,SKM在保持模型复杂度的同时,进一步提高了模型的学习收敛速度和识别性能。与使用投影仪结合大语言模型(SpeechLLM)的语音识别方法的比较实验也表明,SKM驱动的LLM-P2 G在识别准确性和结构简单性方面具有更多优势。研究结果验证了该方法在跨语言语音识别系统中的实用价值和应用潜力。
摘要:Recently, the Large Language Model-based Phoneme-to-Grapheme (LLM-P2G) method has shown excellent performance in speech recognition tasks and has become a feasible direction to replace the traditional WFST decoding method. This framework takes into account both recognition accuracy and system scalability through two-stage modeling of phoneme prediction and text generation. However, the existing LLM-P2G adopts the Top-K Marginalized (TKM) training strategy, and its candidate phoneme sequences rely on beam search generation, which has problems such as insufficient path diversity, low training efficiency, and high resource overhead. To this end, this paper proposes a sampling marginalized training strategy (Sampling-K Marginalized, SKM), which replaces beam search with random sampling to generate candidate paths, improving marginalized modeling and training efficiency. Experiments were conducted on Polish and German datasets, and the results showed that SKM further improved the model learning convergence speed and recognition performance while maintaining the complexity of the model. Comparative experiments with a speech recognition method that uses a projector combined with a large language model (SpeechLLM) also show that the SKM-driven LLM-P2G has more advantages in recognition accuracy and structural simplicity. The study verified the practical value and application potential of this method in cross-language speech recognition systems.
【4】What Does the Speaker Embedding Encode?
标题:扬声器嵌入编码是什么?
链接:https://arxiv.org/abs/2512.18286
备注:This paper was accepted by Interspeech 2017. However, no public version is currently available, as the original link provided by ISCA is no longer accessible. The version uploaded herein has undergone automatic English polishing using GPT (Expanded for better calarity)
摘要:开发良好的说话人嵌入在语音社区中引起了极大的兴趣,i-vector和d-vector等表示在各种任务中表现出卓越的性能。尽管它们被广泛采用,但一个基本问题仍然没有被探索:这些嵌入实际上编码了什么属性?为了解决这一差距,我们对三种主要的说话人嵌入方法进行了全面分析:i-vector,d-vector和基于RNN/LSTM的序列向量(s-vector)。通过精心设计的分类任务,我们系统地研究了他们在多个维度上的编码能力,包括说话者身份,性别,语速,文本内容,词序和通道信息。我们的分析揭示了每种嵌入类型的不同优势和局限性:i-向量擅长说话人识别,但编码有限的顺序信息; s-向量有效地捕获文本内容和词序,但与说话人身份斗争; d-向量表现出平衡的性能,但通过平均丢失顺序信息。基于这些见解,我们提出了一种新的多任务学习框架,它集成了i-vector和s-vector,从而产生了一种新的说话人嵌入(i-s-vector),结合了它们的互补优势。在RSR 2015上的实验结果表明,与i-vector基线相比,本文提出的i-s-vector在内容不匹配测试中实现了超过50%的EER降低,验证了本文方法的有效性。
摘要:Developing a good speaker embedding has received tremendous interest in the speech community, with representations such as i-vector and d-vector demonstrating remarkable performance across various tasks. Despite their widespread adoption, a fundamental question remains largely unexplored: what properties are actually encoded in these embeddings? To address this gap, we conduct a comprehensive analysis of three prominent speaker embedding methods: i-vector, d-vector, and RNN/LSTM-based sequence-vector (s-vector). Through carefully designed classification tasks, we systematically investigate their encoding capabilities across multiple dimensions, including speaker identity, gender, speaking rate, text content, word order, and channel information. Our analysis reveals distinct strengths and limitations of each embedding type: i-vector excels at speaker discrimination but encodes limited sequential information; s-vector captures text content and word order effectively but struggles with speaker identity; d-vector shows balanced performance but loses sequential information through averaging. Based on these insights, we propose a novel multi-task learning framework that integrates i-vector and s-vector, resulting in a new speaker embedding (i-s-vector) that combines their complementary advantages. Experimental results on RSR2015 demonstrate that the proposed i-s-vector achieves more than 50% EER reduction compared to the i-vector baseline on content mismatch trials, validating the effectiveness of our approach.
【5】TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
标题:TICI+:儿童语音识别的语音上下文学习案例研究
链接:https://arxiv.org/abs/2512.18263
备注:Published at IEEE ASRU 2025 Satellite Workshop-AI for Children's Speech and Language
摘要:儿童的语音识别仍然具有挑战性,由于大量的声学和语言的变化,有限的标记数据,并从成人语音显着差异。语音基础模型可以通过语音上下文学习(SICL)来解决这些挑战,允许适应新的领域而无需微调。然而,SICL的有效性取决于如何选择上下文中的例子。我们扩展了现有的基于检索的方法,文本嵌入KNN的SICL(TICL),引入一个声学重新排序步骤,以创建TICL+。这个扩展优先考虑那些在语义上和声学上都与测试输入一致的例子。在四个儿童语音语料库上的实验表明,TICL+实现了高达53.3%的相对字错误率降低超过zero-shot性能和37.6%超过基线TICL,突出了结合语义和声学信息的儿童语音中的鲁棒性,可扩展的ASR的价值。
摘要:Children's speech recognition remains challenging due to substantial acoustic and linguistic variability, limited labeled data, and significant differences from adult speech. Speech foundation models can address these challenges through Speech In-Context Learning (SICL), allowing adaptation to new domains without fine-tuning. However, the effectiveness of SICL depends on how in-context examples are selected. We extend an existing retrieval-based method, Text-Embedding KNN for SICL (TICL), introducing an acoustic reranking step to create TICL+. This extension prioritizes examples that are both semantically and acoustically aligned with the test input. Experiments on four children's speech corpora show that TICL+ achieves up to a 53.3% relative word error rate reduction over zero-shot performance and 37.6% over baseline TICL, highlighting the value of combining semantic and acoustic information for robust, scalable ASR in children's speech.
【6】SAM Audio: Segment Anything in Audio
标题:萨姆音频:在音频中分割任何内容
链接:https://arxiv.org/abs/2512.18099
摘要:通用音频源分离是多模态AI系统的关键能力,可以感知和推理声音。尽管近年来取得了重大进展,但现有的分离模型要么是特定于领域的,专为固定类别(如语音或音乐)而设计,要么是可控性有限的,仅支持单一的提示方式(如文本)。在这项工作中,我们提出了SAM音频,通用音频分离的基础模型,统一在一个单一的框架内的文本,视觉和时间跨度提示。SAM Audio构建在扩散Transformer架构上,通过对涵盖语音、音乐和一般声音的大规模音频数据进行流匹配进行训练,可以灵活地分离由语言、视觉掩码或时间跨度描述的目标源。该模型在各种基准测试中实现了最先进的性能,包括在野外和专业制作的音频中的一般声音,语音,音乐和乐器分离,大大优于以前的通用和专用系统。此外,我们引入了一个新的现实世界的分离基准与人类标记的多模态提示和一个无参考的评价模型,与人类的判断密切相关。
摘要:General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed categories such as speech or music, or limited in controllability, supporting only a single prompting modality such as text. In this work, we present SAM Audio, a foundation model for general audio separation that unifies text, visual, and temporal span prompting within a single framework. Built on a diffusion transformer architecture, SAM Audio is trained with flow matching on large-scale audio data spanning speech, music, and general sounds, and can flexibly separate target sources described by language, visual masks, or temporal spans. The model achieves state-of-the-art performance across a diverse suite of benchmarks, including general sound, speech, music, and musical instrument separation in both in-the-wild and professionally produced audios, substantially outperforming prior general-purpose and specialized systems. Furthermore, we introduce a new real-world separation benchmark with human-labeled multimodal prompts and a reference-free evaluation model that correlates strongly with human judgment.
【7】LIWhiz: A Non-Intrusive Lyric Intelligibility Prediction System for the Cadenza Challenge
标题:LIWhiz:Cadenza挑战赛的非侵入性抒情可理解性预测系统
链接:https://arxiv.org/abs/2512.17937
摘要:我们提出了LIWhiz,一个非侵入式的抒情可懂度预测系统提交给ICASSP 2026华彩乐段挑战赛。LIWhiz利用Whisper进行强大的特征提取和可训练的后端进行分数预测。在Cadenza歌词可懂度预测(CLIP)评估集上进行测试,LIWhiz在基于STOI的基线上实现了22.4%的相对均方根误差降低,从而大大改善了归一化互相关。
摘要:We present LIWhiz, a non-intrusive lyric intelligibility prediction system submitted to the ICASSP 2026 Cadenza Challenge. LIWhiz leverages Whisper for robust feature extraction and a trainable back-end for score prediction. Tested on the Cadenza Lyric Intelligibility Prediction (CLIP) evaluation set, LIWhiz achieves a 22.4% relative root mean squared error reduction over the STOI-based baseline, yielding a substantial improvement in normalized cross-correlation.
【8】Continual Learning for Acoustic Event Classification
标题:声学事件分类的持续学习
链接:https://arxiv.org/abs/2512.17932
备注:Master project report
摘要:考虑到对计算资源的限制(例如,模型大小、运行内存)。为了缓解这个问题,我们提出了两个新的多样性意识的增量学习方法的口语关键词发现和环境声音分类。我们的方法通过测量每个样本的分类不确定性来选择用于训练的历史数据。对于口语关键词定位应用程序,建议的RK方法引入了一个多样性感知采样器,通过计算分类不确定性从历史和传入的关键词中选择一个不同的集合。因此,RK方法可以在不忘记先前知识的情况下渐进地学习新任务。此外,RK方法还提出了数据增强和知识蒸馏损失函数,以实现边缘设备上的高效内存管理。对于环境声音分类应用程序,我们通过观察数据的分类概率如何随添加到分类器嵌入的并行扰动而波动来测量不确定性。通过这种方式,与向原始数据添加扰动相比,可以显著降低计算成本。实验结果表明,所提出的RK方法在Google Speech Command数据集上的平均准确率比最佳基线提高了4.2%,所需内存更少。在DCASE 2019 Task 1和ESC-50数据集上的实验结果表明,我们提出的方法在分类精度和计算效率方面优于基线持续学习方法,表明我们的方法可以有效地增量学习新类,而不会出现灾难性遗忘问题。
摘要:Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device acoustic event classification given the restrictions on computation resources (e.g., model size, running memory). To alleviate such an issue, we propose two novel diversity-aware incremental learning method for Spoken Keyword Spotting and Environmental Sound Classification. Our method selects the historical data for the training by measuring the per-sample classification uncertainty. For the Spoken Keyword Spotting application, the proposed RK approach introduces a diversity-aware sampler to select a diverse set from historical and incoming keywords by calculating classification uncertainty. As a result, the RK approach can incrementally learn new tasks without forgetting prior knowledge. Besides, the RK approach also proposes data augmentation and knowledge distillation loss function for efficient memory management on the edge device. For the Environmental Sound Classification application, we measure the uncertainty by observing how the classification probability of data fluctuates against the parallel perturbations added to the classifier embedding. In this way, the computation cost can be significantly reduced compared with adding perturbation to the raw data. Experimental results show that the proposed RK approach achieves 4.2% absolute improvement in terms of average accuracy over the best baseline on Google Speech Command dataset with less required memory. Experimental results on the DCASE 2019 Task 1 and ESC-50 dataset show that our proposed method outperforms baseline continual learning methods on classification accuracy and computational efficiency, indicating our method can efficiently and incrementally learn new classes without the catastrophic forgetting problem for on-device environmental sound classification
【9】MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
标题:MauBERT:少数镜头声学单位发现的通用语音感应偏差
链接:https://arxiv.org/abs/2512.19612
摘要:本文介绍了MauBERT,HuBERT的多语言扩展,利用发音功能进行强大的跨语言语音表示学习。我们继续进行HuBERT预训练,并基于55种语言的语音到发音特征映射进行监督。我们的模型从多语言数据中学习,以预测发音特征或音素,从而产生捕获多语言语音特性的独立于语言的表示。通过全面的ABX可辨别性测试,我们表明MauBERT模型比最先进的多语言自监督学习模型产生更多的上下文不变表示。此外,这些模型可以有效地适应看不见的语言和随意的语音,并进行最少的自我监督微调(10小时的语音)。这为在自监督语音模型中注入语言归纳偏差建立了一种有效的方法。
摘要:This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MauBERT models produce more context-invariant representations than state-of-the-art multilingual self-supervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models.
【10】JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
标题:JoyVoice:拟人化多说话者对话合成的长上下文条件反射
链接:https://arxiv.org/abs/2512.19090
摘要:大型语音生成模型正从单说话人、短句合成向多说话人、长对话生成发展。目前的长形式语音生成模型主要限于二元,回合为基础的互动。为了解决这个问题,我们引入了JoyVoice,这是一种新颖的拟人基础模型,旨在灵活,无边界合成多达八个扬声器。与传统的级联系统不同,JoyVoice采用统一的E2 E-Transformer-DiT架构,该架构将自回归隐藏表示直接用于扩散输入,从而实现整体端到端优化。我们进一步提出了一个MM-令牌器工作在12.5 Hz的低比特率,它集成了多任务语义和MMSE损失,有效地模拟语义和声学信息。此外,该模型通过大规模数据扰动集成了强大的文本前端处理。实验表明,JoyVoice在多语种生成(中、英、日、韩)和zero-shot语音克隆方面取得了较好的效果。JoyVoice在Seed-TTS-Eval Benchmark和多扬声器长形式对话语音克隆任务中均取得了顶级结果,展示了卓越的音频质量和泛化能力。它实现了显着的改善,韵律连续性的长形式的语音,节奏丰富的多扬声器会话,语言自然,除了卓越的可懂度。我们鼓励读者在https://jea-speech.github.io/JoyVoice上收听演示
摘要:Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice
【11】chatter: a Python library for applying information theory and AI/ML models to animal communication
标题:chatter:将信息理论和AI/ML模型应用于动物交流的Python库
链接:https://arxiv.org/abs/2512.17935
摘要:动物交流的研究通常涉及将单位分类为类型(例如鸣禽的音节或座头鲸的音符)。虽然这种方法在许多情况下是有用的,但它必然会忽略实际通信系统中存在的复杂性和细微差别。Chatter是一个新的Python库,用于使用信息论和现代机器学习技术分析连续潜在空间中的动物通信。它在分类学上是不可知的,已经用鸟类、蝙蝠、鲸鱼和灵长类动物的发声进行了测试。通过利用各种不同的架构,包括变分自动编码器和Vision Transformers,chatter将声音序列表示为高维潜在空间中的轨迹,从而绕过了手动或自动分类单元的需要。该库提供了一个端到端的工作流程-从预处理和分割到模型训练和特征提取-使研究人员能够量化声音序列的复杂性,可预测性,相似性和新颖性。
摘要:The study of animal communication often involves categorizing units into types (e.g. syllables in songbirds, or notes in humpback whales). While this approach is useful in many cases, it necessarily flattens the complexity and nuance present in real communication systems. chatter is a new Python library for analyzing animal communication in continuous latent space using information theory and modern machine learning techniques. It is taxonomically agnostic, and has been tested with the vocalizations of birds, bats, whales, and primates. By leveraging a variety of different architectures, including variational autoencoders and vision transformers, chatter represents vocal sequences as trajectories in high-dimensional latent space, bypassing the need for manual or automatic categorization of units. The library provides an end-to-end workflow -- from preprocessing and segmentation to model training and feature extraction -- that enables researchers to quantify the complexity, predictability, similarity, and novelty of vocal sequences.
机器翻译由腾讯交互翻译提供,仅供参考
