微信公众号:arXiv_Daily
cs.SD语音
【1】ADNAC: Audio Denoiser using Neural Audio Codec
标题:ADNEC:使用神经音频编解码器的音频降噪器
链接:https://arxiv.org/abs/2511.01773
备注:Accepted and presented at the 13th International Conference on Speech Technology and Human-Computer Dialogue (SpeD), Cluj-Napoca, Romania, October 19-22, 2025. 4 pages, 1 figure. IEEE Catalog Number: CFP2555H-USB, ISBN: 979-8-3315-7485-7
摘要:音频去噪在信号处理中至关重要,可增强恢复音乐录音等应用的清晰度和保真度。本文提出了一种概念验证,用于适应最先进的神经音频编解码器,描述音频编解码器(DAC),用于音乐去噪。这项工作克服了传统架构(如U-Nets)的局限性,通过在从不同来源构建的大规模定制合成数据集上训练模型。训练由结合时域、频谱和信号级保真度度量的多目标损失函数指导。最终,本文的目的是提出一个高保真,生成音频恢复的方法。
摘要:Audio denoising is critical in signal processing, enhancing intelligibility and fidelity for applications like restoring musical recordings. This paper presents a proof-of-concept for adapting a state-of-the-art neural audio codec, the Descript Audio Codec (DAC), for music denoising. This work overcomes the limitations of traditional architectures like U-Nets by training the model on a large-scale, custom-synthesized dataset built from diverse sources. Training is guided by a multi objective loss function that combines time-domain, spectral, and signal-level fidelity metrics. Ultimately, this paper aims to present a PoC for high-fidelity, generative audio restoration.
【2】The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity
标题:钥匙里的幽灵:人类与人工智能音乐协同创造的丰富演示
链接:https://arxiv.org/abs/2511.01663
摘要:虽然音乐创作的生成模型越来越有能力,但音乐家对它们的采用受到文本提示的阻碍,这是一种与乐器演奏的具体化,响应性本质脱节的异步工作流程。为了解决这个问题,我们引入了Aria-Duet,这是一个交互式系统,可以促进人类钢琴家和Aria之间的实时音乐二重奏,Aria是一个最先进的生成模型,使用Yamaha吉他作为共享的物理接口。该框架实现了一个回合合作:用户执行,信号切换,模型生成一个连贯的连续执行钢琴上的声学。除了描述实现这种低延迟交互的技术架构之外,我们还从音乐学的角度分析了系统的输出,发现该模型可以保持风格语义并开发连贯的短语思想,证明这种体现系统可以参与音乐复杂的对话,并为人类-人工智能共同创造开辟了一条有前途的新道路。
摘要:While generative models for music composition are increasingly capable, their adoption by musicians is hindered by text-prompting, an asynchronous workflow disconnected from the embodied, responsive nature of instrumental performance. To address this, we introduce Aria-Duet, an interactive system facilitating a real-time musical duet between a human pianist and Aria, a state-of-the-art generative model, using a Yamaha Disklavier as a shared physical interface. The framework enables a turn-taking collaboration: the user performs, signals a handover, and the model generates a coherent continuation performed acoustically on the piano. Beyond describing the technical architecture enabling this low-latency interaction, we analyze the system's output from a musicological perspective, finding the model can maintain stylistic semantics and develop coherent phrasal ideas, demonstrating that such embodied systems can engage in musically sophisticated dialogue and open a promising new path for human-AI co-creation.
【3】Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
标题:演讲-戏剧:演讲角色扮演中人性化基准的框架
链接:https://arxiv.org/abs/2511.01261
备注:67 pages
摘要:角色扮演已经成为生成模型的关键测试平台,从纯文本对话扩展到多模式交互。将角色扮演扩展到语音捕捉韵律,情感和交付,但也提出了新的评估挑战。当前的流水线通常使用音频大语言模型(ALLM)作为zero-shot法官,其错过了非语言线索,将多个方面折叠成粗略的分数,并且依赖于不能反映真实世界角色的合成语音参考。我们提出Speech-DRAME,这是一个统一的框架,在三个层面上做出贡献:(i)Speech-DRAME-EvalBench,具有双语人类注释数据和用于训练和测试语音评估模型(SEM)的协议的评估基准,(ii)DRAME-Eval,微调的评估模型,其显著优于zero-shot和Few-Shot ALLM,以及(iii)Speech-DRAME-RoleBench,一个语音角色扮演基准,利用DRAME-Eval作为自动判断来比较语音基础模型(SFM)。Speech-DRAME区分了两种互补的评估策略:原型评估,一种自上而下的方法,测量对广泛角色原型的遵守情况,以及现实主义评估,一种基于真实人类语言的自下而上的方法,强调细微的角色质量。与zero-shot ALLM判断相比,DRAME-Eval与人类评分的一致性更高(原型的Pearson相关系数为0.480至0.629,现实主义的Pearson相关系数为0.390至0.625)。通过整合透明的基准资源,建模方法和系统级评估,Speech-DRAME为评估口语角色扮演提供了第一个全面的,可复制的基础。
摘要:Role-play has become a key testbed for generative models, expanding from text-only dialogue to multimodal interaction. Extending role-play to speech captures prosody, emotion, and delivery, but also poses new evaluation challenges. Current pipelines often use audio large language models (ALLMs) as zero-shot judges, which miss paralinguistic cues, collapse multiple aspects into coarse scores, and rely on synthetic speech references that fail to reflect real-world roles. We present Speech-DRAME, a unified framework that contributes at three levels: (i) Speech-DRAME-EvalBench, an evaluation benchmark with bilingual human-annotated data and protocols for training and testing speech evaluation models (SEMs), (ii) DRAME-Eval, a fine-tuned evaluation model, which substantially outperforms zero-shot and few-shot ALLMs, and (iii) Speech-DRAME-RoleBench, a speech role-play benchmark that leverages DRAME-Eval as an automatic judge to compare speech foundation models (SFMs). Speech-DRAME distinguishes between two complementary evaluation strategies: Archetype Evaluation, a top-down approach measuring adherence to broad role archetypes, and Realism Evaluation, a bottom-up approach grounded in real human speech that emphasizes nuanced role quality. Compared to zero-shot ALLM judges, DRAME-Eval achieves stronger agreement with human ratings (Pearson correlation from 0.480 to 0.629 in archetypes, and 0.390 to 0.625 in realism). By integrating transparent benchmark resources, modeling approaches, and system-level evaluation, Speech-DRAME provides the first comprehensive, reproducible foundation for assessing spoken role-play.
【4】Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
标题:具有大型音频语言模型的反馈驱动检索增强音频生成
链接:https://arxiv.org/abs/2511.01091
摘要:我们提出了一个通用的反馈驱动的检索增强生成(RAG)的方法,利用大型音频语言模型(LALM),以解决文本到音频(TTA)生成的特定声音事件的缺失或不完善的合成。与以前的基于RAG的TTA方法不同,通常从头开始训练专门的模型,我们利用LALM来分析音频生成输出,检索预训练模型难以从外部数据库生成的概念,并将检索到的信息纳入生成过程。实验结果表明,我们的方法不仅增强了LALM识别丢失的声音事件的能力,而且在不同的模型中提供了改进,优于现有的RAG专用方法。
摘要:We propose a general feedback-driven retrieval-augmented generation (RAG) approach that leverages Large Audio Language Models (LALMs) to address the missing or imperfect synthesis of specific sound events in text-to-audio (TTA) generation. Unlike previous RAG-based TTA methods that typically train specialized models from scratch, we utilize LALMs to analyze audio generation outputs, retrieve concepts that pre-trained models struggle to generate from an external database, and incorporate the retrieved information into the generation process. Experimental results show that our method not only enhances the ability of LALMs to identify missing sound events but also delivers improvements across different models, outperforming existing RAG-specialized approaches.
【5】Rhythm in the Air: Vision-based Real-Time Music Generation through Gestures
标题:空气中的节奏:通过手势生成基于视觉的实时音乐
链接:https://arxiv.org/abs/2511.00793
备注:8 pages, 7 figures
摘要:手势识别是人机交互(HCI)的重要组成部分,有助于用户和计算机系统之间的无缝互连,而无需物理接触。本文介绍了一种创新的应用程序,基于视觉的动态手势识别(VDGR)的实时音乐创作,通过手势。为了实现这个应用程序,我们生成了一个自定义的手势数据集,其中包括超过15000个样本在21个类,结合7个音符,每个音符表现在三个不同的音高水平。为了有效地处理适度的训练数据量,并准确地识别和优先考虑复杂的手势序列的音乐创作,我们开发了一个多层基于注意力的门控递归单元(MLA-GRU)模型,其中门控递归单元(GRU)用于从观察到的序列中学习时间模式,并采用注意力层来关注音乐相关的手势片段。我们的实证研究表明,MLA-GRU显着超过经典的GRU模型,实现了96.83%的显着准确性相比,基线的86.7%。此外,我们的方法具有优越的效率和处理速度,这是至关重要的交互式应用程序。使用我们提出的系统,我们相信人们将以一种新的、令人兴奋的方式与音乐互动。它不仅提升了人机交互体验,还突出了MLA-GRU在需要快速精确手势识别的场景中的有效性。
摘要:Gesture recognition is an essential component of human-computer interaction (HCI), facilitating seamless interconnectivity between users and computer systems without physical touch. This paper introduces an innovative application of vision-based dynamic gesture recognition (VDGR) for real-time music composition through gestures. To implement this application, we generate a custom gesture dataset that encompasses over 15000 samples across 21 classes, incorporating 7 musical notes each manifesting at three distinct pitch levels. To effectively deal with the modest volume of training data and to accurately discern and prioritize complex gesture sequences for music creation, we develop a multi-layer attention-based gated recurrent unit (MLA-GRU) model, in which gated recurrent unit (GRU) is used to learn temporal patterns from the observed sequence and an attention layer is employed to focus on musically pertinent gesture segments. Our empirical studies demonstrate that MLA-GRU significantly surpasses the classical GRU model, achieving a remarkable accuracy of 96.83% compared to the baseline's 86.7%. Moreover, our approach exhibits superior efficiency and processing speed, which are crucial for interactive applications. Using our proposed system, we believe that people will interact with music in a new and exciting way. It not only advances HCI experiences but also highlights MLA-GRU's effectiveness in scenarios demanding swift and precise gesture recognition.
【6】More Than A Shortcut: A Hyperbolic Approach To Early-Exit Networks
标题:不仅仅是收件箱:提前退出网络的双曲线方法
链接:https://arxiv.org/abs/2511.00641
摘要:在资源受限的设备上部署准确的事件检测面临着性能和计算成本之间的权衡的挑战。虽然早期退出(EE)网络通过自适应计算提供了一种解决方案,但它们通常无法实施连贯的层次结构,从而限制了其早期预测的可靠性。为了解决这个问题,我们提出了双曲早期退出网络(HypEE),这是一种在双曲空间中学习EE表示的新框架。我们的核心贡献是一个具有新的蕴涵损失的分层训练目标,它强制执行部分排序约束,以确保更深的网络层在几何上细化更浅的网络层的表示。多个音频事件检测任务和骨干架构的实验表明,HypEE显着优于标准的欧几里得EE基线,特别是在最早的,最计算关键的出口。学习的几何结构还提供了一种原则性的不确定性度量,从而实现了一种新的触发机制,使整个系统比传统的EE和标准骨干模型更有效,更准确,而无需提前退出。
摘要:Deploying accurate event detection on resource-constrained devices is challenged by the trade-off between performance and computational cost. While Early-Exit (EE) networks offer a solution through adaptive computation, they often fail to enforce a coherent hierarchical structure, limiting the reliability of their early predictions. To address this, we propose Hyperbolic Early-Exit networks (HypEE), a novel framework that learns EE representations in the hyperbolic space. Our core contribution is a hierarchical training objective with a novel entailment loss, which enforces a partial-ordering constraint to ensure that deeper network layers geometrically refine the representations of shallower ones. Experiments on multiple audio event detection tasks and backbone architectures show that HypEE significantly outperforms standard Euclidean EE baselines, especially at the earliest, most computationally-critical exits. The learned geometry also provides a principled measure of uncertainty, enabling a novel triggering mechanism that makes the overall system both more efficient and more accurate than a conventional EE and standard backbone models without early-exits.
【7】Physics-Informed Neural Networks for Speech Production
标题:用于语音生成的物理信息神经网络
链接:https://arxiv.org/abs/2511.00428
备注:11 pages, 10 figures
摘要:基于声带和声道物理模型的言语产生分析是声带行为研究和语言学研究的基础。提出了一种基于物理信息神经网络的语音生成分析方法。该网络直接在声褶振动和声道声学的控制方程上训练。声音折叠碰撞引入不可微性和消失梯度,对PINN具有挑战性的现象。然而,我们证明,引入一个可微的近似函数,使PINN框架内的声乐倍振动的分析。自激声带振动的周期通常是未知的。我们表明,通过处理的周期作为一个可学习的网络参数,周期解可以得到。此外,通过实现声门流和声道声学之间的耦合作为硬约束,声门-声道相互作用实现没有额外的损失项。我们证实了该方法的有效性,通过正向和反向分析,表明声门流量,声带振动状态,声门下压力可以同时估计从语音信号。值得注意的是,相同的网络架构可以应用于正向和反向分析,突出了这种方法的多功能性。该方法继承了PINNs的优点,包括无网格计算和自然纳入非线性,因此具有广泛的应用前景。
摘要:The analysis of speech production based on physical models of the vocal folds and vocal tract is essential for studies on vocal-fold behavior and linguistic research. This paper proposes a speech production analysis method using physics-informed neural networks (PINNs). The networks are trained directly on the governing equations of vocal-fold vibration and vocal-tract acoustics. Vocal-fold collisions introduce nondifferentiability and vanishing gradients, challenging phenomena for PINNs. We demonstrate, however, that introducing a differentiable approximation function enables the analysis of vocal-fold vibrations within the PINN framework. The period of self-excited vocal-fold vibration is generally unknown. We show that by treating the period as a learnable network parameter, a periodic solution can be obtained. Furthermore, by implementing the coupling between glottal flow and vocal-tract acoustics as a hard constraint, glottis-tract interaction is achieved without additional loss terms. We confirmed the method's validity through forward and inverse analyses, demonstrating that the glottal flow rate, vocal-fold vibratory state, and subglottal pressure can be simultaneously estimated from speech signals. Notably, the same network architecture can be applied to both forward and inverse analyses, highlighting the versatility of this approach. The proposed method inherits the advantages of PINNs, including mesh-free computation and the natural incorporation of nonlinearities, and thus holds promise for a wide range of applications.
【8】Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study
标题:使用轻量级和基于转换器的模型进行语音情绪检测:比较和消融研究
链接:https://arxiv.org/abs/2511.00402
摘要:语音情感识别在具有移情能力的人机交互系统的开发中起着至关重要的作用。本文通过对CREMA-D数据集中的六种核心情绪进行分类,对基于轻量级transformer的模型DistilHuBERT和PaSST进行了比较分析。我们使用MFCC特征将其性能与传统的CNN-LSTM基线模型进行基准测试。DistilHuBERT表现出卓越的准确性(70.64%)和F1评分(70.36%),同时保持非常小的模型大小(0.02 MB),优于PaSST和基线。此外,我们对PaSST、线性、MLP和注意力池头的三种变体进行了消融研究,以了解分类头架构对模型性能的影响。我们的研究结果表明,PaSST与MLP头产生最好的性能在其变体,但仍然低于蒸馏休伯特。在情绪类别中,愤怒始终是最准确检测的,而厌恶仍然是最具挑战性的。这些发现表明,像DistilHuBERT这样的轻量级Transformers为边缘设备上的实时语音情感识别提供了一个令人信服的解决方案。该代码可从以下网址获得:https://github.com/luckymaduabuchi/Emotion-detection-。
摘要:Emotion recognition from speech plays a vital role in the development of empathetic human-computer interaction systems. This paper presents a comparative analysis of lightweight transformer-based models, DistilHuBERT and PaSST, by classifying six core emotions from the CREMA-D dataset. We benchmark their performance against a traditional CNN-LSTM baseline model using MFCC features. DistilHuBERT demonstrates superior accuracy (70.64%) and F1 score (70.36%) while maintaining an exceptionally small model size (0.02 MB), outperforming both PaSST and the baseline. Furthermore, we conducted an ablation study on three variants of the PaSST, Linear, MLP, and Attentive Pooling heads, to understand the effect of classification head architecture on model performance. Our results indicate that PaSST with an MLP head yields the best performance among its variants but still falls short of DistilHuBERT. Among the emotion classes, angry is consistently the most accurately detected, while disgust remains the most challenging. These findings suggest that lightweight transformers like DistilHuBERT offer a compelling solution for real-time speech emotion recognition on edge devices. The code is available at: https://github.com/luckymaduabuchi/Emotion-detection-.
【9】LongCat-Flash-Omni Technical Report
标题:LongCat-Flash-Omni技术报告
链接:https://arxiv.org/abs/2511.00279
摘要:我们引入LongCat-Flash-Omni,这是一个最先进的开源全模态模型,拥有5600亿个参数,擅长实时视听交互。LongCat-Flash-Omni采用了一种启发式的渐进式训练策略,从简单的模态序列建模任务过渡到越来越复杂的模态序列建模任务,在保持强大的单峰能力的同时,实现了全面的多模态能力。基于LongCat-Flash,它采用了一个高性能的快捷连接混合专家(MoE)架构与零计算专家,LongCat-Flash-Omni集成了高效的多模态感知和语音重建模块。尽管其560 B参数的巨大规模(27 B激活),LongCat-Flash-Omni实现了低延迟的实时视听交互。对于训练基础设施,我们开发了一个模态解耦的并行方案,专门用于管理大规模多模态训练中固有的数据和模型异构性。这种创新的方法通过保持纯文本训练所实现的吞吐量的90%以上来展示卓越的效率。广泛的评估表明,LongCat-Flash-Omni在开源模型中的全模态基准测试中达到了最先进的性能。此外,它在各种特定于模态的任务中提供了极具竞争力的结果,包括文本、图像和视频理解,以及音频理解和生成。我们提供了模型架构设计,培训程序和数据策略的全面概述,并开源模型,以促进社区未来的研究和开发。
摘要:We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strategy that transitions from simpler to increasingly complex modality sequence modeling tasks, LongCat-Flash-Omni attains comprehensive multimodal capabilities while maintaining strong unimodal capability. Building upon LongCat-Flash, which adopts a high-performance Shortcut-connected Mixture-of-Experts (MoE) architecture with zero-computation experts, LongCat-Flash-Omni integrates efficient multimodal perception and speech reconstruction modules. Despite its immense size of 560B parameters (with 27B activated), LongCat-Flash-Omni achieves low-latency real-time audio-visual interaction. For training infrastructure, we developed a modality-decoupled parallelism scheme specifically designed to manage the data and model heterogeneity inherent in large-scale multimodal training. This innovative approach demonstrates exceptional efficiency by sustaining over 90% of the throughput achieved by text-only training. Extensive evaluations show that LongCat-Flash-Omni achieves state-of-the-art performance on omni-modal benchmarks among open-source models. Furthermore, it delivers highly competitive results across a wide range of modality-specific tasks, including text, image, and video understanding, as well as audio understanding and generation. We provide a comprehensive overview of the model architecture design, training procedures, and data strategies, and open-source the model to foster future research and development in the community.
【10】Leveraging Language Information for Target Language Extraction
标题:利用语言信息进行目标语言提取
链接:https://arxiv.org/abs/2511.01652
备注:Accepted to APSIPA ASC 2025
摘要:目标语言提取旨在从包含多个讲不同语言的说话者的混合波形中提取特定语言的语音。人类的听觉系统擅长用特定语言的知识来执行这项任务。然而,常规提取系统的性能受到缺乏这种先验知识的限制。语音预训练模型可以从大规模的野外语料库中捕获丰富的语言和语音表示,可以为这些系统提供这种缺失的语言知识。在这项工作中,我们提出了一个新的端到端框架,以利用语音预训练模型的语言知识。这些知识用于指导抽取模型更好地捕捉目标语言特征,从而提高抽取质量。为了证明我们提出的方法的有效性,我们构建了第一个公开的多语言数据集的目标语言提取。实验结果表明,我们的方法实现了1.22 dB和1.12 dB的SI-SNR的改善,分别为英语和德语提取,从包含两种语言的混合物。
摘要:Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory system is adept at performing this task with the knowledge of the particular language. However, the performance of the conventional extraction systems is limited by the lack of this prior knowledge. Speech pre-trained models, which capture rich linguistic and phonetic representations from large-scale in-the-wild corpora, can provide this missing language knowledge to these systems. In this work, we propose a novel end-to-end framework to leverage language knowledge from speech pre-trained models. This knowledge is used to guide the extraction model to better capture the target language characteristics, thereby improving extraction quality. To demonstrate the effectiveness of our proposed approach, we construct the first publicly available multilingual dataset for Target Language Extraction. Experimental results show that our method achieves improvements of 1.22 dB and 1.12 dB in SI-SNR for English and German extraction, respectively, from mixtures containing both languages.
【11】MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
标题:Multi-Bench:评估口语对话模型情商能力的多回合互动基准
链接:https://arxiv.org/abs/2511.00850
备注:Submitted to ICASSP 2026
摘要:口语对话模型(SDM)发展迅速,但他们的能力,以维持真正的互动多轮对话仍然没有得到充分的探索,因为大多数基准集中在单轮交流。我们介绍多台,第一个基准明确设计用于评估SDM在多轮互动对话,强调情商。Multi-Bench采用层次结构,其中基本轨道用于情感理解和推理,高级轨道用于情感支持和应用。它包括五个精心设计的任务和大约3.2K个样本,从情感识别到复杂的推理和交互式对话,由可重复的评估框架支持。我们评估了六个代表性的SDM的八个子集的多台。结果表明,虽然目前的SDM在基本的理解任务上取得了良好的表现,但在高级多轮交互式对话和推理相关的任务中,特别是在情感意识和应用方面,它们仍有改进的空间。
摘要:Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on single-turn exchanges. We introduce Multi-Bench, the first benchmark explicitly designed to evaluate SDMs in multi-turn interactive dialogue with an emphasis on emotional intelligence. Multi-Bench employs a hierarchical structure with a basic track for emotion understanding and reasoning and an advanced track for emotion support and application. It comprises five carefully designed tasks and about 3.2K samples, ranging from emotion recognition to complex reasoning and interactive dialogue, supported by a reproducible evaluation framework. We evaluate six representative SDMs on eight subsets of Multi-Bench. Results show that while current SDMs achieve good performance on basic understanding tasks, they still have room for improvement in advanced multi-turn interactive dialogue and reasoning-related tasks, particularly in emotion awareness and application.
【12】NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
标题:NaturalVoices:用于语音转换的大规模、自发和情感播客数据集
链接:https://arxiv.org/abs/2511.00256
备注:Under review for IEEE Transactions on Affective Computing
摘要:日常言语传达的远不止文字,它反映了我们是谁,我们的感受以及我们互动的环境。然而,大多数现有的语音数据集都是行动的,规模有限,无法捕捉现实生活中交流的丰富表达。随着大型神经网络的兴起,出现了几个大规模的语音语料库,并在各种语音处理任务中被广泛采用。然而,语音转换(VC)领域仍然缺乏大规模的,表现力,和现实生活中的语音资源,适合建模自然韵律和情感。为了填补这一空白,我们发布了NaturalVoices(NV),这是第一个专门为情感感知语音转换而设计的大规模自发播客数据集。它包括5,049小时的自发播客录音,并自动注释情感(分类和基于属性),语音质量,转录,扬声器身份和声音事件。该数据集捕捉了数千名演讲者,不同主题和自然演讲风格的表达情感变化。我们还提供了一个具有模块化注释工具和灵活过滤的开源管道,使研究人员能够为各种VC任务构建自定义子集。实验表明,NaturalVoices支持开发强大的和可推广的VC模型,能够产生自然,富有表现力的语音,同时揭示了当前架构在应用于大规模自发数据时的局限性。这些结果表明,NaturalVoices既是一种宝贵的资源,也是推进语音转换领域的一个具有挑战性的基准。数据集可从以下网址获得:https://huggingface.co/JHU-SmileLab
摘要:Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive richness of real-life communication. With the rise of large neural networks, several large-scale speech corpora have emerged and been widely adopted across various speech processing tasks. However, the field of voice conversion (VC) still lacks large-scale, expressive, and real-life speech resources suitable for modeling natural prosody and emotion. To fill this gap, we release NaturalVoices (NV), the first large-scale spontaneous podcast dataset specifically designed for emotion-aware voice conversion. It comprises 5,049 hours of spontaneous podcast recordings with automatic annotations for emotion (categorical and attribute-based), speech quality, transcripts, speaker identity, and sound events. The dataset captures expressive emotional variation across thousands of speakers, diverse topics, and natural speaking styles. We also provide an open-source pipeline with modular annotation tools and flexible filtering, enabling researchers to construct customized subsets for a wide range of VC tasks. Experiments demonstrate that NaturalVoices supports the development of robust and generalizable VC models capable of producing natural, expressive speech, while revealing limitations of current architectures when applied to large-scale spontaneous data. These results suggest that NaturalVoices is both a valuable resource and a challenging benchmark for advancing the field of voice conversion. Dataset is available at: https://huggingface.co/JHU-SmileLab
【1】Leveraging Language Information for Target Language Extraction
标题:利用语言信息进行目标语言提取
链接:https://arxiv.org/abs/2511.01652
备注:Accepted to APSIPA ASC 2025
摘要:目标语言提取旨在从包含多个讲不同语言的说话者的混合波形中提取特定语言的语音。人类的听觉系统擅长用特定语言的知识来执行这项任务。然而,常规提取系统的性能受到缺乏这种先验知识的限制。语音预训练模型可以从大规模的野外语料库中捕获丰富的语言和语音表示,可以为这些系统提供这种缺失的语言知识。在这项工作中,我们提出了一个新的端到端框架,以利用语音预训练模型的语言知识。这些知识用于指导抽取模型更好地捕捉目标语言特征,从而提高抽取质量。为了证明我们提出的方法的有效性,我们构建了第一个公开的多语言数据集的目标语言提取。实验结果表明,我们的方法实现了1.22 dB和1.12 dB的SI-SNR的改善,分别为英语和德语提取,从包含两种语言的混合物。
摘要:Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory system is adept at performing this task with the knowledge of the particular language. However, the performance of the conventional extraction systems is limited by the lack of this prior knowledge. Speech pre-trained models, which capture rich linguistic and phonetic representations from large-scale in-the-wild corpora, can provide this missing language knowledge to these systems. In this work, we propose a novel end-to-end framework to leverage language knowledge from speech pre-trained models. This knowledge is used to guide the extraction model to better capture the target language characteristics, thereby improving extraction quality. To demonstrate the effectiveness of our proposed approach, we construct the first publicly available multilingual dataset for Target Language Extraction. Experimental results show that our method achieves improvements of 1.22 dB and 1.12 dB in SI-SNR for English and German extraction, respectively, from mixtures containing both languages.
【2】AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
标题:AudioNet:用于检索类似音频事件的监督深度哈希
链接:https://arxiv.org/abs/2511.01372
备注:None
摘要:这项工作提出了一个监督深度哈希方法检索类似的音频事件。所提出的方法名为AudioNet,是一种基于深度学习的系统,用于使用音频示例作为查询来有效地散列和检索类似的音频事件。AudioNet通过为相似的音频事件生成二进制哈希码,在多个标准数据集上实现了高检索性能,在该领域建立了新的基准,并突出了其与其他哈希方法相比的功效和有效性。通过在标准数据集上的综合实验,我们的研究在评估相似音频事件的检索性能方面具有开创性意义。提出了一种新的损失函数,它结合了加权对比和加权成对损失以及哈希码平衡,以提高音频事件检索的效率。该方法采用离散梯度传播,这使得梯度在反向传播过程中通过离散变量传播。这使得网络能够使用标准的基于梯度的优化算法来优化离散哈希码,该算法通常用于连续变量。实验结果表明,该方法具有良好的检索性能,即使在处理不平衡的数据集。本研究中进行的系统分析进一步支持了所提出的方法在多个数据集检索性能方面的显着优势。这项工作中提出的研究结果为未来使用深度音频嵌入有效检索类似音频事件的研究奠定了基础。
摘要:This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing and retrieval of similar audio events using an audio example as a query. AudioNet achieves high retrieval performance on multiple standard datasets by generating binary hash codes for similar audio events, setting new benchmarks in the field, and highlighting its efficacy and effectiveness compare to other hashing methods. Through comprehensive experiments on standard datasets, our research represents a pioneering effort in evaluating the retrieval performance of similar audio events. A novel loss function is proposed which incorporates weighted contrastive and weighted pairwise loss along with hashcode balancing to improve the efficiency of audio event retrieval. The method adopts discrete gradient propagation, which allows gradients to be propagated through discrete variables during backpropagation. This enables the network to optimize the discrete hash codes using standard gradient-based optimization algorithms, which are typically used for continuous variables. The proposed method showcases promising retrieval performance, as evidenced by the experimental results, even when dealing with imbalanced datasets. The systematic analysis conducted in this study further supports the significant benefits of the proposed method in retrieval performance across multiple datasets. The findings presented in this work establish a baseline for future studies on the efficient retrieval of similar audio events using deep audio embeddings.
【3】Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
标题:迈向一般听觉智力:机器听力和口语的大型多模式模型
链接:https://arxiv.org/abs/2511.01299
备注:22 pages, 11 figures
摘要:在大型语言模型(LLM)和人工通用智能(AGI)时代,计算机听觉必须超越传统范式,充分利用基础模型的能力,实现更全面的理解,更自然的生成和更人性化的交互。音频作为一种富含语义、情感和上下文线索的模态,在实现自然的、具身的机器智能方面起着至关重要的作用。这项调查提供了一个全面的审查最近的进展,在集成音频到LLM,重点是四个关键领域:音频理解,音频生成,基于语音的交互和视听理解。我们分析了LLM如何重塑音频感知和推理,使系统能够在更深的语义层面上理解声音,生成富有表现力的音频输出,并进行类似人类的语音交互。此外,我们探讨了音频和视觉模态的融合如何增强情境意识和跨模态推理,推动多模态智能的边界。这项调查不仅综合了现有的研究,还确定了构建音频原生AGI系统的关键挑战和未来方向,这些系统能够像人类一样自然地通过声音感知,理解和交互。
摘要:In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities of foundation models, towards more comprehensive understanding, more natural generation and more human-like interaction. Audio, as a modality rich in semantic, emotional, and contextual cues, plays a vital role in achieving naturalistic and embodied machine intelligence. This survey provides a comprehensive review of recent progress in integrating audio into LLMs, with a focus on four key areas: audio comprehension, audio generation, speech-based interaction, and audio-visual understanding. We analyze how LLMs are reshaping audio perception and reasoning, enabling systems to understand sound at a deeper semantic level, generate expressive audio outputs, and engage in human-like spoken interaction. Furthermore, we explore how the fusion of audio and visual modalities enhances situational awareness and cross-modal reasoning, pushing the boundaries of multimodal intelligence. This survey not only synthesizes existing research but also identifies critical challenges and future directions for building audio-native AGI systems capable of perceiving, understanding, and interacting through sound as naturally as humans do.
【4】WhisperVC: Target Speaker-Controllable Mandarin Whisper-to-Speech Conversion
标题:WhisperVC:目标说话人可控的普通话语音转换
链接:https://arxiv.org/abs/2511.01056
摘要:耳语语音缺乏声带激励,并表现出降低的能量和移动的共振峰频率,使自然和可理解的语音重建极具挑战性。为了解决这个问题,我们提出了一个三阶段的框架,用于普通话耳语到语音(W2 S)转换。阶段1采用基于OpenAI Whisper-large~V3模型的微调内容编码器和基于Conformer的变分自动编码器,采用软DTW对齐来学习域不变和时间一致的表示。第二阶段介绍了确定性的长度-通道对齐器和无时长限制的FastSpeech~2模型,该模型以说话人嵌入为条件,以实现可控的音色和稳定的韵律。阶段3在预测的梅尔频谱图上微调HiFi-GAN声码器以合成高保真波形。在AISHELL 6-Whisper语料库上的实验表明,WhisperVC在保持说话人相似度(\textbf{cosine~ 0. 76})的同时,实现了接近地面实况的质量(\textbf{DNSMOS~ 3. 11},\textbf{UTMOS~ 2. 52},\textbf{CER~ 18. 67\%}),并且在仅耳语推理下具有鲁棒性。
摘要:Whispered speech lacks vocal-fold excitation and exhibits reduced energy and shifted formant frequencies, making natural and intelligible voice reconstruction highly challenging. To address this issue, we propose \emph{WhisperVC}, a three-stage framework for Mandarin whisper-to-speech (W2S) conversion. Stage~1 employs a fine-tuned Content Encoder based on the OpenAI Whisper-large~V3 model and a Conformer-based variational autoencoder with soft-DTW alignment to learn domain-invariant and temporally consistent representations. Stage~2 introduces a deterministic Length--Channel Aligner and a duration-free FastSpeech~2 model conditioned on speaker embeddings for controllable timbre and stable prosody. Stage~3 fine-tunes a HiFi-GAN vocoder on predicted mel-spectrograms to synthesize high-fidelity waveforms. Experiments on the AISHELL6-Whisper corpus demonstrate that WhisperVC achieves near ground-truth quality (\textbf{DNSMOS~3.11}, \textbf{UTMOS~2.52}, \textbf{CER~18.67\%}), while maintaining speaker similarity (\textbf{cosine~0.76}) and robust performance under whisper-only inference.
【5】MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
标题:Multi-Bench:评估口语对话模型情商能力的多回合互动基准
链接:https://arxiv.org/abs/2511.00850
备注:Submitted to ICASSP 2026
摘要:口语对话模型(SDM)发展迅速,但他们的能力,以维持真正的互动多轮对话仍然没有得到充分的探索,因为大多数基准集中在单轮交流。我们介绍多台,第一个基准明确设计用于评估SDM在多轮互动对话,强调情商。Multi-Bench采用分层结构,具有用于情感理解和推理的基本轨道以及用于情感支持和应用的高级轨道。它包括五个精心设计的任务和大约3.2K个样本,从情感识别到复杂的推理和交互式对话,由可重复的评估框架支持。我们评估了六个代表性的SDM的八个子集的多台。结果表明,虽然目前的SDM在基本的理解任务上取得了良好的表现,但在高级多轮交互式对话和推理相关的任务中,特别是在情感意识和应用方面,它们仍有改进的空间。
摘要:Spoken Dialogue Models (SDMs) have advanced rapidly, yet their ability to sustain genuinely interactive multi-turn conversations remains underexplored, as most benchmarks focus on single-turn exchanges. We introduce Multi-Bench, the first benchmark explicitly designed to evaluate SDMs in multi-turn interactive dialogue with an emphasis on emotional intelligence. Multi-Bench employs a hierarchical structure with a basic track for emotion understanding and reasoning and an advanced track for emotion support and application. It comprises five carefully designed tasks and about 3.2K samples, ranging from emotion recognition to complex reasoning and interactive dialogue, supported by a reproducible evaluation framework. We evaluate six representative SDMs on eight subsets of Multi-Bench. Results show that while current SDMs achieve good performance on basic understanding tasks, they still have room for improvement in advanced multi-turn interactive dialogue and reasoning-related tasks, particularly in emotion awareness and application.
【6】NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
标题:NaturalVoices:用于语音转换的大规模、自发和情感播客数据集
链接:https://arxiv.org/abs/2511.00256
备注:Under review for IEEE Transactions on Affective Computing
摘要:日常言语传达的远不止文字,它反映了我们是谁,我们的感受以及我们互动的环境。然而,大多数现有的语音数据集都是行动的,规模有限,无法捕捉现实生活中交流的丰富表达。随着大型神经网络的兴起,出现了几个大规模的语音语料库,并在各种语音处理任务中被广泛采用。然而,语音转换(VC)领域仍然缺乏大规模的,表现力,和现实生活中的语音资源,适合建模自然韵律和情感。为了填补这一空白,我们发布了NaturalVoices(NV),这是第一个专门为情感感知语音转换而设计的大规模自发播客数据集。它包括5,049小时的自发播客录音,并自动注释情感(分类和基于属性),语音质量,转录,扬声器身份和声音事件。该数据集捕捉了数千名演讲者,不同主题和自然演讲风格的表达情感变化。我们还提供了一个具有模块化注释工具和灵活过滤的开源管道,使研究人员能够为各种VC任务构建自定义子集。实验表明,NaturalVoices支持开发强大的和可推广的VC模型,能够产生自然,富有表现力的语音,同时揭示了当前架构在应用于大规模自发数据时的局限性。这些结果表明,NaturalVoices既是一个宝贵的资源,也是推进语音转换领域的一个具有挑战性的基准。数据集可从以下网址获得:https://huggingface.co/JHU-SmileLab
摘要:Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive richness of real-life communication. With the rise of large neural networks, several large-scale speech corpora have emerged and been widely adopted across various speech processing tasks. However, the field of voice conversion (VC) still lacks large-scale, expressive, and real-life speech resources suitable for modeling natural prosody and emotion. To fill this gap, we release NaturalVoices (NV), the first large-scale spontaneous podcast dataset specifically designed for emotion-aware voice conversion. It comprises 5,049 hours of spontaneous podcast recordings with automatic annotations for emotion (categorical and attribute-based), speech quality, transcripts, speaker identity, and sound events. The dataset captures expressive emotional variation across thousands of speakers, diverse topics, and natural speaking styles. We also provide an open-source pipeline with modular annotation tools and flexible filtering, enabling researchers to construct customized subsets for a wide range of VC tasks. Experiments demonstrate that NaturalVoices supports the development of robust and generalizable VC models capable of producing natural, expressive speech, while revealing limitations of current architectures when applied to large-scale spontaneous data. These results suggest that NaturalVoices is both a valuable resource and a challenging benchmark for advancing the field of voice conversion. Dataset is available at: https://huggingface.co/JHU-SmileLab
【7】Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
标题:演讲-戏剧:演讲角色扮演中人性化基准的框架
链接:https://arxiv.org/abs/2511.01261
备注:67 pages
摘要:角色扮演已经成为生成模型的关键测试平台,从纯文本对话扩展到多模式交互。将角色扮演扩展到语音捕捉韵律,情感和交付,但也提出了新的评估挑战。当前的流水线通常使用音频大语言模型(ALLM)作为zero-shot法官,其错过了非语言线索,将多个方面折叠成粗略的分数,并且依赖于不能反映真实世界角色的合成语音参考。我们提出Speech-DRAME,这是一个统一的框架,在三个层面上做出贡献:(i)Speech-DRAME-EvalBench,具有双语人类注释数据和用于训练和测试语音评估模型(SEM)的协议的评估基准,(ii)DRAME-Eval,微调的评估模型,其显著优于zero-shot和Few-Shot ALLM,以及(iii)Speech-DRAME-RoleBench,一个语音角色扮演基准,利用DRAME-Eval作为自动判断来比较语音基础模型(SFM)。Speech-DRAME区分了两种互补的评估策略:原型评估,一种自上而下的方法,测量对广泛角色原型的遵守情况,以及现实主义评估,一种基于真实人类语言的自下而上的方法,强调细微的角色质量。与zero-shot ALLM判断相比,DRAME-Eval与人类评分的一致性更高(原型的Pearson相关系数为0.480至0.629,现实主义的Pearson相关系数为0.390至0.625)。通过整合透明的基准资源,建模方法和系统级评估,Speech-DRAME为评估口语角色扮演提供了第一个全面的,可复制的基础。
摘要:Role-play has become a key testbed for generative models, expanding from text-only dialogue to multimodal interaction. Extending role-play to speech captures prosody, emotion, and delivery, but also poses new evaluation challenges. Current pipelines often use audio large language models (ALLMs) as zero-shot judges, which miss paralinguistic cues, collapse multiple aspects into coarse scores, and rely on synthetic speech references that fail to reflect real-world roles. We present Speech-DRAME, a unified framework that contributes at three levels: (i) Speech-DRAME-EvalBench, an evaluation benchmark with bilingual human-annotated data and protocols for training and testing speech evaluation models (SEMs), (ii) DRAME-Eval, a fine-tuned evaluation model, which substantially outperforms zero-shot and few-shot ALLMs, and (iii) Speech-DRAME-RoleBench, a speech role-play benchmark that leverages DRAME-Eval as an automatic judge to compare speech foundation models (SFMs). Speech-DRAME distinguishes between two complementary evaluation strategies: Archetype Evaluation, a top-down approach measuring adherence to broad role archetypes, and Realism Evaluation, a bottom-up approach grounded in real human speech that emphasizes nuanced role quality. Compared to zero-shot ALLM judges, DRAME-Eval achieves stronger agreement with human ratings (Pearson correlation from 0.480 to 0.629 in archetypes, and 0.390 to 0.625 in realism). By integrating transparent benchmark resources, modeling approaches, and system-level evaluation, Speech-DRAME provides the first comprehensive, reproducible foundation for assessing spoken role-play.
【8】Ultralow-power standoff acoustic leak detection
标题:超低功耗隔离声泄漏检测
链接:https://arxiv.org/abs/2511.00348
备注:5 pages, 4 figures
摘要:一个自动化的,远距离声学泄漏检测计划已经设计,建造和测试。它融合了玻璃破碎和烟雾探测的原理,以警告来自加压管道的泄漏。在超过10 m的间隔距离处,已可靠地检测到以0.15 l/min流速流动的模拟漏水。该装置还可有效识别位于墙壁、门、地板和天花板等表面后面的泄漏。预期的应用是作为一个自治的,电池供电的,远程无线节点。所有信号处理和分析都在边缘进行,无需将音频数据传输到云端。传感器状态按需传送,只有几个字节的信息,需要最小的带宽。功耗范围为20- 200微瓦,具体取决于环境噪声量和所需的传感器延迟。为了获得最佳的灵敏度和可靠性,硬件的工作频率远远高于人类对话的范围,使窃听成为不可能。开发已经完成了水从加压管道泄漏,但传感器的概念可以有效地用于检测气体泄漏。
摘要:An automated, standoff acoustic leak detection scheme has been designed, built, and tested. It merges the principles of glass breakage and smoke detection to alert for the presence of leaks emanating from pressurized plumbing. A simulated water leak flowing at 0.15 l/min has been reliably detected at a standoff distance of more than 10 m. The device is also effective at identifying the presence of leaks located behind surfaces such as walls, doors, floors, and ceilings. The anticipated application is as an autonomous, battery-powered, remote wireless node. All signal processing and analysis takes place on the edge with no need to stream audio data to the cloud. Sensor status is conveyed on-demand with only a few bytes of information, requiring minimal bandwidth. Power consumption is the range of 20--200 micro-Watts, depending on the amount of environmental noise and desired sensor latency. To attain optimum sensitivity and reliability, the hardware operates at acoustic frequencies well above the range of human conversations, making eavesdropping impossible. Development has been done with water escaping from pressurized plumbing, but the sensor concept can be used effectively to detect gas leaks.
机器翻译由腾讯交互翻译提供,仅供参考
