今日论文合集:cs.SD语音12篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Domain-Incremental Continual Learning for Robust and Efficient Keyword Spotting in Resource Constrained Systems
标题:领域增量连续学习在资源受限系统中实现稳健有效的关键词发现
链接:https://arxiv.org/abs/2601.16158

作者:Prakash Dhungana,Sayed Ahmad Salehi
备注:12 pages, 8 figures, and 3 tables
摘要:在边缘设备上部署的具有小尺寸模型的关键字定位(KWS)系统面临着显著的准确性和鲁棒性挑战,这是由于不同的噪声和记录条件引起的域偏移。为了解决这个问题,我们提出了一个全面的框架,旨在适应新的领域,同时保持计算效率的持续学习。建议的管道集成了一个双输入卷积神经网络,利用梅尔频率倒谱系数(MFCC)和梅尔频谱图功能,支持多级去噪过程,涉及离散小波变换和谱减法技术,加上模型和原型更新块。与限制更新到特定层的先前方法不同,我们的方法更新了完整的量化模型,由于紧凑的模型架构而成为可能。在运行时使用类原型和置信度驱动过滤选择输入样本的子集,然后将其伪标记并与排练缓冲区相结合以进行增量模型再训练。噪声测试数据集上的实验结果表明,该框架的有效性,达到99.63%的准确率在干净的数据和保持强大的性能(超过94%的准确率)在不同的噪声环境中,即使在-10 dB的信噪比.所提出的框架工作证实,将有效的去噪与基于原型的持续学习相结合,使KWS模型能够在资源受限的动态环境中自主和鲁棒地运行。
摘要:Keyword Spotting (KWS) systems with small footprint models deployed on edge devices face significant accuracy and robustness challenges due to domain shifts caused by varying noise and recording conditions. To address this, we propose a comprehensive framework for continual learning designed to adapt to new domains while maintaining computational efficiency. The proposed pipeline integrates a dual-input Convolutional Neural Network, utilizing both Mel Frequency Cepstral Coefficients (MFCC) and Mel-spectrogram features, supported by a multi-stage denoising process, involving discrete wavelet transform and spectral subtraction techniques, plus model and prototype update blocks. Unlike prior methods that restrict updates to specific layers, our approach updates the complete quantized model, made possible due to compact model architecture. A subset of input samples are selected during runtime using class prototypes and confidence-driven filtering, which are then pseudo-labeled and combined with rehearsal buffer for incremental model retraining. Experimental results on noisy test dataset demonstrate the framework's effectiveness, achieving 99.63\% accuracy on clean data and maintaining robust performance (exceeding 94\% accuracy) across diverse noisy environments, even at -10 dB Signal-to-Noise Ratio. The proposed framework work confirms that integrating efficient denoising with prototype-based continual learning enables KWS models to operate autonomously and robustly in resource-constrained, dynamic environments.


【2】Pay (Cross) Attention to the Melody: Curriculum Masking for Single-Encoder Melodic Harmonization
标题:(交叉)关注旋律:单一编码器旋律协调的课程掩蔽
链接:https://arxiv.org/abs/2601.16150

作者:Maximos Kaliakatsos-Papakostas,Dimos Makris,Konstantinos Soiledis,Konstantinos-Theodoros Tsamis,Vassilis Katsouros,Emilios Cambouropoulos
摘要:旋律协调,即为给定旋律生成和声的任务,仍然是计算音乐生成中的核心挑战。最近的单编码器Transformer方法将和声视为掩蔽序列建模问题,但受离散扩散启发的现有训练课程通常会导致旋律与和声之间的弱(交叉)注意力。这导致了对旋律线索的有限利用,特别是在域外语境中。在这项工作中,我们引入了一个训练课程,FF(全到全),它保持所有的和声标记掩盖了几个训练步骤,然后在训练过程中逐步揭开整个序列,以加强旋律和声的相互作用。我们系统地评估这种方法对以前的课程在多个实验轴,包括时间量化(四分之一与十六分之一音符),酒吧级与时间签名条件反射,旋律表示(全范围与音高类),和推理时间揭露策略。模型在HookTheory数据集上进行训练,并在域内和爵士乐标准的策划集合上进行评估,使用一套全面的指标来评估和弦进行结构,和声旋律对齐和节奏连贯性。结果表明,所提出的FF课程始终优于基线几乎所有的指标,特别是强大的增益域外的评价,其中谐波适应性,以新颖的旋律队列是至关重要的。我们进一步发现,四分之一音符量化,交织的酒吧令牌,和音高类的旋律表示是有利的FF设置。我们的研究结果强调了培训课程的重要性,使有效的旋律条件反射,并建议全面全面揭露提供了一个强大的战略,单一的编码器协调。
摘要:Melodic harmonization, the task of generating harmonic accompaniments for a given melody, remains a central challenge in computational music generation. Recent single encoder transformer approaches have framed harmonization as a masked sequence modeling problem, but existing training curricula inspired by discrete diffusion often result in weak (cross) attention between melody and harmony. This leads to limited exploitation of melodic cues, particularly in out-of-domain contexts. In this work, we introduce a training curriculum, FF (full-to-full), which keeps all harmony tokens masked for several training steps before progressively unmasking entire sequences during training to strengthen melody-harmony interactions. We systematically evaluate this approach against prior curricula across multiple experimental axes, including temporal quantization (quarter vs. sixteenth note), bar-level vs. time-signature conditioning, melody representation (full range vs. pitch class), and inference-time unmasking strategies. Models are trained on the HookTheory dataset and evaluated both in-domain and on a curated collection of jazz standards, using a comprehensive set of metrics that assess chord progression structure, harmony-melody alignment, and rhythmic coherence. Results demonstrate that the proposed FF curriculum consistently outperforms baselines in nearly all metrics, with particularly strong gains in out-of-domain evaluations where harmonic adaptability to novel melodic queues is crucial. We further find that quarter-note quantization, intertwining of bar tokens, and pitch-class melody representations are advantageous in the FF setting. Our findings highlight the importance of training curricula in enabling effective melody conditioning and suggest that full-to-full unmasking offers a robust strategy for single encoder harmonization.


【3】Distillation-based Layer Dropping (DLD) Effective End-to-end Framework for Dynamic Speech Networks
标题:基于蒸馏的层丢弃(DLD)动态语音网络的有效端到端框架
链接:https://arxiv.org/abs/2601.16117

作者:Abdul Hannan,Daniele Falavigna,Shah Nawaz,Mubashir Noman,Markus Schedl,Alessio Brutti
备注:Accepted at ICASSP 2026
摘要:边缘设备在受限和变化的资源设置中运行,需要能够适应可用资源限制的动态架构。为了满足这样的需求,层丢弃($\mathcal{LD}$)方法通常用于通过跳过网络的部分来将静态模型转换为动态模型,同时降低整体计算复杂度。然而,现有的$\mathcal{LD}$方法极大地影响了动态模型的低和高丢弃情况下的性能,恶化了性能计算的权衡。为此,我们提出了一个基于蒸馏的层丢弃(DLD)框架,该框架以端到端的方式有效地结合了知识蒸馏和$\mathcal{LD}$的功能,从而实现了动态语音网络的最新性能。综合实验利用著名的语音识别方法,包括conformer和WavLM,在三个公共基准测试证明了我们的框架的有效性,减少字的错误率由$9.32\%$和$2.25\%$的高,没有下降的情况下,与$33.3\%$减少训练时间。
摘要:Edge devices operate in constrained and varying resource settings, requiring dynamic architectures that can adapt to limitations of the available resources. To meet such demands, layer dropping ($\mathcal{LD}$) approach is typically used to transform static models into dynamic ones by skipping parts of the network along with reducing overall computational complexity. However, existing $\mathcal{LD}$ methods greatly impact the dynamic model's performance for low and high dropping cases, deteriorating the performance-computation trade-off. To this end, we propose a distillation-based layer dropping (DLD) framework that effectively combines the capabilities of knowledge distillation and $\mathcal{LD}$ in an end-to-end fashion, thereby achieving state-of-the-art performance for dynamic speech networks. Comprehensive experimentation utilizing well-known speech recognition methods, including conformer and WavLM, on three public benchmarks demonstrates the effectiveness of our framework, reducing the word error rate by $9.32\%$ and $2.25\%$ for high and no dropping cases with $33.3\%$ reduction in training time.


【4】PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation
标题:PF-D2 M:通用舞蹈到音乐生成的无姿势扩散模型
链接:https://arxiv.org/abs/2601.15872

作者:Jaekwon Im,Natalia Polouliakh,Taketo Akama
备注:4 pages, 2 figures
摘要:Dance-to-Music生成旨在生成与舞蹈动作一致的音乐。现有的方法通常依赖于从单个人类舞者和有限的舞蹈到音乐数据集提取的身体运动特征,这限制了它们对涉及多个舞者和非人类舞者的真实世界场景的性能和适用性。在本文中,我们提出了PF-D2 M,一个通用的基于扩散的舞蹈到音乐生成模型,结合了从舞蹈视频中提取的视觉特征。PF-D2 M采用渐进式训练策略进行训练,可有效解决数据稀缺和泛化挑战。客观和主观的评估都表明,PF-D2 M在舞蹈音乐对齐和音乐质量方面达到了最先进的性能。
摘要:Dance-to-music generation aims to generate music that is aligned with dance movements. Existing approaches typically rely on body motion features extracted from a single human dancer and limited dance-to-music datasets, which restrict their performance and applicability to real-world scenarios involving multiple dancers and non-human dancers. In this paper, we propose PF-D2M, a universal diffusion-based dance-to-music generation model that incorporates visual features extracted from dance videos. PF-D2M is trained with a progressive training strategy that effectively addresses data scarcity and generalization challenges. Both objective and subjective evaluations show that PF-D2M achieves state-of-the-art performance in dance-music alignment and music quality.


【5】U3-xi: Pushing the Boundaries of Speaker Recognition via Incorporating Uncertainty
标题:U3-xi:通过消除不确定性突破说话人识别的界限
链接:https://arxiv.org/abs/2601.15719

作者:Junjie Li,Kong Aik Lee
摘要:通常通过聚合帧级表示的序列来获得话语级说话人嵌入。然而,在现实世界的情况下,各个帧不仅编码说话者相关的信息,而且还编码各种干扰因素。因此,不同的帧对自动说话人确认系统的最终话语级说话人表示的贡献不相等。为了解决这个问题,我们建议估计每个帧的固有不确定性,并相应地分配自适应权重,其中具有较高不确定性的帧受到较低的关注。基于这一思想,我们提出了U3-xi,一个全面的框架,旨在产生更可靠和可解释的不确定性估计的扬声器嵌入。具体来说,我们介绍了几种策略的不确定性监督。首先,我们通过随机方差损失提出了说话者级别的不确定性监督,其中话语嵌入与其对应的说话者质心之间的距离作为不确定性学习的伪地面真值。其次,我们通过在训练过程中将预测的不确定性注入sof tmax量表来纳入全球层面的不确定性监督。这种自适应缩放机制根据样本难度调整决策边界的清晰度,提供全局指导。第三,我们重新设计的不确定性估计模块集成了Transformer编码器与多视图自注意,使模型能够捕捉丰富的本地和远程的时间依赖性。综合实验表明,U3-xi是模型不可知的,可以无缝地应用于各种扬声器编码器。特别是,当应用于ECAPA-TDNN时,它在EER和minDCF方面分别实现了VoxCeleb 1测试集的21.1%和15.57%的相对改善。
摘要:An utterance-level speaker embedding is typically obtained by aggregating a sequence of frame-level representations. However, in real-world scenarios, individual frames encode not only speaker-relevant information but also various nuisance factors. As a result, different frames contribute unequally to the final utterance-level speaker representation for Automatic Speaker Verification systems. To address this issue, we propose to estimate the inherent uncertainty of each frame and assign adaptive weights accordingly, where frames with higher uncertainty receive lower attention. Based on this idea, we present U3-xi, a comprehensive framework designed to produce more reliable and interpretable uncertainty estimates for speaker embeddings. Specifically, we introduce several strategies for uncertainty supervision. First, we propose speaker-level uncertainty supervision via a Stochastic Variance Loss, where the distance between an utterance embedding and its corresponding speaker centroid serves as a pseudo ground truth for uncertainty learning. Second, we incorporate global-level uncertainty supervision by injecting the predicted uncertainty into the sof tmax scale during training. This adaptive scaling mechanism adjusts the sharpness of the decision boundary according to sample difficulty, providing global guidance. Third, we redesign the uncertainty estimation module by integrating a Transformer encoder with multi-view self-attention, enabling the model to capture rich local and long-range temporal dependencies. Comprehensive experiments demonstrate that U3-xi is model-agnostic and can be seamlessly applied to various speaker encoders. In particular, when applied to ECAPA-TDNN, it achieves 21.1% and 15.57% relative improvements on the VoxCeleb1 test sets in terms of EER and minDCF, respectively.


【6】Bridging the Perception Gap: A Lightweight Coarse-to-Fine Architecture for Edge Audio Systems
标题:弥合感知差距:边缘音频系统的轻量级从粗到精架构
链接:https://arxiv.org/abs/2601.15676

作者:Hengfan Zhang,Yueqian Lin,Hai Helen Li,Yiran Chen
备注:10 pages, 3 figures, 2 tables. Preprint
摘要:在边缘基础设施上部署音频语言模型(Audio-LLM)暴露了感知深度和计算效率之间的持续紧张关系。轻量级本地模型往往会产生被动感知--通用摘要会错过多步音频推理所需的细微证据--而不加选择的云卸载会导致不可接受的延迟、带宽成本和隐私风险。我们提出了CoFi-Agent(Tool-Augmented Coarse-to-Fine Agent),这是一种针对边缘服务器和网关的混合架构。它执行快速局部感知,并仅在检测到不确定性时触发条件取证细化。CoFi-Agent在本地7 B Audio-LLM上运行初始单遍,然后云控制器门控困难的情况并为设备上工具发布轻量级计划,例如临时重新监听和本地ASR。在MMAR基准测试中,CoFi-Agent将准确率从27.20%提高到53.60%,同时比始终在线的调查管道实现了更好的准确性-效率权衡。总的来说,CoFi-Agent通过在实际系统约束下启用工具的有条件边缘云协作来弥合感知差距。
摘要:Deploying Audio-Language Models (Audio-LLMs) on edge infrastructure exposes a persistent tension between perception depth and computational efficiency. Lightweight local models tend to produce passive perception - generic summaries that miss the subtle evidence required for multi-step audio reasoning - while indiscriminate cloud offloading incurs unacceptable latency, bandwidth cost, and privacy risk. We propose CoFi-Agent (Tool-Augmented Coarse-to-Fine Agent), a hybrid architecture targeting edge servers and gateways. It performs fast local perception and triggers conditional forensic refinement only when uncertainty is detected. CoFi-Agent runs an initial single-pass on a local 7B Audio-LLM, then a cloud controller gates difficult cases and issues lightweight plans for on-device tools such as temporal re-listening and local ASR. On the MMAR benchmark, CoFi-Agent improves accuracy from 27.20% to 53.60%, while achieving a better accuracy-efficiency trade-off than an always-on investigation pipeline. Overall, CoFi-Agent bridges the perception gap via tool-enabled, conditional edge-cloud collaboration under practical system constraints.


【7】EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
标题:描述思想者:用于可解释语音情感推理的韵律感知强化学习
链接:https://arxiv.org/abs/2601.15668

作者:Dingdong Wang,Shujie Liu,Tianhua Zhang,Youjun Chen,Jinyu Li,Helen Meng
摘要:言语中的情感信息在多模态感知中起着独特的作用。然而,目前的语音大语言模型(SpeechLLM),类似于传统的语音情感识别(SER)系统,仍然把情感理解作为一个简单的分类问题。这提供了有限的可解释性的预测,而离开LLM的表达和推理能力未得到充分利用。在这项工作中,我们迈出了第一步,通过强化学习(RL)将SER重新表述为深度推理问题。我们提出了一种基于细粒度声学线索的可解释解释的情感预测方法,它旨在生成准确的情感预测。为了实现这一点,我们首先构建了一个带有思想链注释和详细标题的情感推理数据集--pactionCoT-35 K。其次,我们观察到,目前的SpeechLLM表现出弱韵律感知,而韵律线索构成的基本信号,解释情绪。为了解决这一问题,我们开发了韵律增强的基础模型PrestitionThinker-Base,并证明韵律增强可以提高情感理解。第三,我们引入了用于RL的具有渐进信任感知推理奖励的组相对策略优化(GRPO-PTR)。与标准GRPO仅依赖基于规则的结果奖励不同,GRPO-PTR逐步引入推理奖励,通过反映推理与结果一致性的可信度权重动态调整推理奖励,并采用基于多维标准的奖励模型评估整体推理质量。在情感准确性和解释质量方面,EmotionThinker优于以前最先进的评估模型,将SER推向可解释的多模态推理。项目页面:https://github.com/dingdongwang/EmotionThinker
摘要:Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion recognition (SER) systems, still treat emotion understanding as a simple classification problem. This provides limited interpretability of predictions, while leaving the LLMs' expressive and reasoning capabilities underutilized. In this work, we take the first step to reformulate SER as a deep reasoning problem through reinforcement learning (RL). We propose EmotionThinker, which is designed to generate accurate emotion predictions with interpretable explanations grounded in fine-grained acoustic cues. To achieve this, we first construct EmotionCoT-35K, an emotional reasoning dataset with Chain-of-Thought annotations and detailed captions. Second, we observe that current SpeechLLMs exhibit weak prosody perception, whereas prosodic cues constitute fundamental signals for interpreting emotions. To address this, we develop the prosody-enhanced foundation model EmotionThinker-Base, and demonstrate that prosody enhancement improves emotion understanding. Third, we introduce Group-Relative-Policy-Optimization with Progressive-Trust-aware-Reasoning-Reward (GRPO-PTR) for RL. Different from standard GRPO, which relies only on rule-based outcome rewards, GRPO-PTR progressively introduces reasoning reward, dynamically adjusts it with a trustworthiness weight reflecting the alignment between reasoning and outcome, and evaluates the overall reasoning quality with a reward model based on multi-dimensional criteria. EmotionThinker outperforms previous state-of-the-art evaluation models both in emotion accuracy and explanation quality, advancing SER toward interpretable multimodal reasoning. Project page: https://github.com/dingdongwang/EmotionThinker


【8】Qwen3-TTS Technical Report
标题:Qwen 3-TTC技术报告
链接:https://arxiv.org/abs/2601.15621

作者:Hangrui Hu,Xinfa Zhu,Ting He,Dake Guo,Bin Zhang,Xiong Wang,Zhifang Guo,Ziyue Jiang,Hongkun Hao,Zishan Guo,Xinyu Zhang,Pei Zhang,Baosong Yang,Jin Xu,Jingren Zhou,Junyang Lin
备注:https://github.com/QwenLM/Qwen3-TTS
摘要:在本报告中,我们介绍了Qwen 3-TTS系列,这是一系列先进的多语言,可控,强大和流式文本到语音模型。Qwen 3-TTS支持最先进的3秒语音克隆和基于预处理的控制,允许创建全新的语音和对输出语音的细粒度操作。Qwen 3-TTS采用双轨LM架构进行实时合成,并配备两个语音标记器:1)Qwen-TTS-Tokenizer-25 Hz是一个强调语义内容的单码本编解码器,可与Qwen-Audio无缝集成,并通过逐块DiT实现流式波形重构。2)Qwen-TTS-Tokenizer-12 Hz实现了极高的比特率降低和超低延迟流传输,通过其12.5 Hz、16层多码本设计和轻量级因果ConvNet实现了即时的第一个数据包发射(97\,\mathrm {ms}$)。广泛的实验表明,在不同的客观和主观基准(例如,TTS多语言测试集、InstructTTSEval和我们的长语音测试集)。为了促进社区的研究和开发,我们在Apache 2.0许可证下发布了tokenizer和模型。
摘要:In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control, allowing both the creation of entirely novel voices and fine-grained manipulation over the output speech. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS adopts a dual-track LM architecture for real-time synthesis, coupled with two speech tokenizers: 1) Qwen-TTS-Tokenizer-25Hz is a single-codebook codec emphasizing semantic content, which offers seamlessly integration with Qwen-Audio and enables streaming waveform reconstruction via a block-wise DiT. 2) Qwen-TTS-Tokenizer-12Hz achieves extreme bitrate reduction and ultra-low-latency streaming, enabling immediate first-packet emission ($97\,\mathrm{ms}$) through its 12.5 Hz, 16-layer multi-codebook design and a lightweight causal ConvNet. Extensive experiments indicate state-of-the-art performance across diverse objective and subjective benchmark (e.g., TTS multilingual test set, InstructTTSEval, and our long speech test set). To facilitate community research and development, we release both tokenizers and models under the Apache 2.0 license.


【9】DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
标题:DeepASMR:基于LLM的Zero-ShotASMR语音生成,适合任何语音的任何人
链接:https://arxiv.org/abs/2601.15596

作者:Leying Zhang,Tingxiao Zhou,Haiyang Sun,Mengxiao Bi,Yanmin Qian
摘要:虽然现代文本到语音(TTS)系统实现了阅读风格语音的高保真度,但它们很难生成自主感觉经络反应(ASMR),这是一种对放松至关重要的专门的低强度语音风格。固有的挑战包括ASMR的微妙,往往无声的特点和要求zero-shot扬声器自适应。在本文中,我们介绍了DeepASMR,这是第一个为zero-shot ASMR生成而设计的框架。我们证明,一个扬声器的普通,阅读风格的语音的一个简短的片段是足以合成高保真ASMR在他们的声音,消除了需要从目标扬声器的耳语训练数据。方法上,我们首先确定,离散语音令牌提供了一个软分解的ASMR风格从扬声器音色。利用这一洞察力,我们提出了一个两阶段的管道,其中包括一个大语言模型(LLM)的内容风格的编码和流匹配的声学解码器的音色重建。此外,我们贡献了DeepASMR-DB,一个全面的670小时的英汉多说话人ASMR语音语料库,并介绍了一种新的评估协议,集成了客观指标,人类听力测试,基于LLM的评分和清音语音分析。大量的实验证实,DeepASMR在任何语音的ASMR生成中实现了最先进的自然度和风格保真度,同时保持了正常语音合成的竞争力。
摘要:While modern Text-to-Speech (TTS) systems achieve high fidelity for read-style speech, they struggle to generate Autonomous Sensory Meridian Response (ASMR), a specialized, low-intensity speech style essential for relaxation. The inherent challenges include ASMR's subtle, often unvoiced characteristics and the demand for zero-shot speaker adaptation. In this paper, we introduce DeepASMR, the first framework designed for zero-shot ASMR generation. We demonstrate that a single short snippet of a speaker's ordinary, read-style speech is sufficient to synthesize high-fidelity ASMR in their voice, eliminating the need for whispered training data from the target speaker. Methodologically, we first identify that discrete speech tokens provide a soft factorization of ASMR style from speaker timbre. Leveraging this insight, we propose a two-stage pipeline incorporating a Large Language Model (LLM) for content-style encoding and a flow-matching acoustic decoder for timbre reconstruction. Furthermore, we contribute DeepASMR-DB, a comprehensive 670-hour English-Chinese multi-speaker ASMR speech corpus, and introduce a novel evaluation protocol integrating objective metrics, human listening tests, LLM-based scoring and unvoiced speech analysis. Extensive experiments confirm that DeepASMR achieves state-of-the-art naturalness and style fidelity in ASMR generation for anyone of any voice, while maintaining competitive performance on normal speech synthesis.


【10】Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
标题:超越预算:通过逻辑空间集成(LOGIC)对语音LLM进行高效且稳健的上下文偏置
链接:https://arxiv.org/abs/2601.15397

作者:Peidong Wang
摘要:在文化变迁、不断发展的趋势和个性化用户数据的推动下,新实体的快速出现对现有的语音大语言模型(Speech LLM)提出了重大挑战。虽然这些模型在一般的会话任务中表现出色,但它们的静态训练知识限制了它们识别特定领域术语的能力,例如联系人姓名,播放列表或技术术语。现有的解决方案主要依赖于提示,其具有较差的可扩展性:随着实体列表的增长,提示遇到上下文窗口限制、增加的推理延迟以及“中间丢失”现象。另一种方法,生成错误纠正(GEC),试图通过后处理重写转录,但经常遭受“过度纠正”,引入从未说过的实体的幻觉。   在这项工作中,我们介绍了逻辑(逻辑空间集成上下文偏置),一个有效的和强大的框架,直接在解码层。与提示不同,LOGIC将上下文注入从输入处理中分离出来,确保相对于提示长度的恒定时间复杂度。使用Phi-4-MM模型在11个多语言环境中进行的大量实验表明,LOGIC在实体WER中平均实现了9%的相对减少,而误报率增加了0.30%。
摘要:The rapid emergence of new entities -- driven by cultural shifts, evolving trends, and personalized user data -- poses a significant challenge for existing Speech Large Language Models (Speech LLMs). While these models excel at general conversational tasks, their static training knowledge limits their ability to recognize domain-specific terms such as contact names, playlists, or technical jargon. Existing solutions primarily rely on prompting, which suffers from poor scalability: as the entity list grows, prompting encounters context window limitations, increased inference latency, and the "lost-in-the-middle" phenomenon. An alternative approach, Generative Error Correction (GEC), attempts to rewrite transcripts via post-processing but frequently suffers from "over-correction", introducing hallucinations of entities that were never spoken.   In this work, we introduce LOGIC (Logit-Space Integration for Contextual Biasing), an efficient and robust framework that operates directly in the decoding layer. Unlike prompting, LOGIC decouples context injection from input processing, ensuring constant-time complexity relative to prompt length. Extensive experiments using the Phi-4-MM model across 11 multilingual locales demonstrate that LOGIC achieves an average 9% relative reduction in Entity WER with a negligible 0.30% increase in False Alarm Rate.


【11】Abusive music and song transformation using GenAI and LLMs
标题:使用GenAI和LLM进行辱骂音乐和歌曲转换
链接:https://arxiv.org/abs/2601.15348

作者:Jiyang Choi,Rohitash Chandra
摘要:反复接触音乐和歌曲内容中的暴力和辱骂内容会影响听众的情绪和行为,可能使攻击行为正常化或强化有害的陈规定型观念。在这项研究中,我们探索使用生成人工智能(GenAI)和大型语言模型(LLM)来自动转换流行音乐中的辱骂性词语(声乐交付)和抒情内容。我们的方法不是简单地静音或替换一个单词,而是改变语气,强度和情绪,因此不仅仅改变歌词,而是如何表达。我们提出了一个比较分析的四个选定的英文歌曲和他们的转换对应,通过声学和情感为基础的镜头评估变化。我们的研究结果表明,Gen-AI显着降低了声音的侵略性,声学分析显示谐波噪声比,倒频谱峰值突出度和微光的改善。情绪分析减少了63.3- 85.6%的艺术家之间的侵略,与合唱部分的重大改进(高达88.6%的减少)。转换后的版本保持了音乐的连贯性,同时减少了有害内容,为传统的内容审查提供了一个有希望的替代方案,避免触发“禁果”效应,即审查内容变得更有吸引力,仅仅是因为它受到限制。这种方法展示了GenAI在保护艺术表达的同时创造更安全的聆听体验的潜力。
摘要:Repeated exposure to violence and abusive content in music and song content can influence listeners' emotions and behaviours, potentially normalising aggression or reinforcing harmful stereotypes. In this study, we explore the use of generative artificial intelligence (GenAI) and Large Language Models (LLMs) to automatically transform abusive words (vocal delivery) and lyrical content in popular music. Rather than simply muting or replacing a single word, our approach transforms the tone, intensity, and sentiment, thus not altering just the lyrics, but how it is expressed. We present a comparative analysis of four selected English songs and their transformed counterparts, evaluating changes through both acoustic and sentiment-based lenses. Our findings indicate that Gen-AI significantly reduces vocal aggressiveness, with acoustic analysis showing improvements in Harmonic to Noise Ratio, Cepstral Peak Prominence, and Shimmer. Sentiment analysis reduced aggression by 63.3-85.6\% across artists, with major improvements in chorus sections (up to 88.6\% reduction). The transformed versions maintained musical coherence while mitigating harmful content, offering a promising alternative to traditional content moderation that avoids triggering the "forbidden fruit" effect, where the censored content becomes more appealing simply because it is restricted. This approach demonstrates the potential for GenAI to create safer listening experiences while preserving artistic expression.


【12】A Stabilized Hybrid Active Noise Control Algorithm of GFANC and FxNLMS with Online Clustering
标题:一种具有在线分簇的稳定的GF非国大和FxNRMS混合主动噪音控制算法
链接:https://arxiv.org/abs/2601.15889

作者:Zhengding Luo,Haozhe Ma,Boxiang Wang,Ziyi Yang,Dongyuan Shi,Woon-Seng Gan
备注:Accepted by 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:滤波-x归一化最小均方(FxNLMS)算法存在收敛速度慢和发散风险的问题,尽管它在充分自适应后可以实现低稳态误差。相比之下,生成式固定滤波器有源噪声控制(GFANC)方法提供了快速的响应速度,但其缺乏适应性可能导致大的稳态误差。本文提出了一种混合GFANC-FxNLMS算法,利用这两种方法的互补优势。在混合GFANC-FxNLMS算法中,GFANC提供帧级控制滤波器作为FxNLMS的初始化,而FxNLMS以采样率执行连续自适应。GFANC生成的滤波器的微小变化可能会重复重新初始化FxNLMS,中断其自适应过程并使系统不稳定。在线聚类模块的引入,以避免不必要的重新初始化,提高系统的稳定性。仿真结果表明,该算法具有响应速度快、稳态误差小、稳定性高等优点,且只需要一个预训练的宽带滤波器。
摘要:The Filtered-x Normalized Least Mean Square (FxNLMS) algorithm suffers from slow convergence and a risk of divergence, although it can achieve low steady-state errors after sufficient adaptation. In contrast, the Generative Fixed-Filter Active Noise Control (GFANC) method offers fast response speed, but its lack of adaptability may lead to large steady-state errors. This paper proposes a hybrid GFANC-FxNLMS algorithm to leverage the complementary advantages of both approaches. In the hybrid GFANC-FxNLMS algorithm, GFANC provides a frame-level control filter as an initialization for FxNLMS, while FxNLMS performs continuous adaptation at the sampling rate. Small variations in the GFANC-generated filter may repeatedly reinitialize FxNLMS, interrupting its adaptation process and destabilizing the system. An online clustering module is introduced to avoid unnecessary re-initializations and improve system stability. Simulation results show that the proposed algorithm achieves fast response, very low steady-state error, and high stability, requiring only one pre-trained broadband filter.


eess.AS音频处理


【1】Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
标题:用于动态环境中多通道日志化和会议增强的光谱和空间模型的松耦合
链接:https://arxiv.org/abs/2601.16077

作者:Adrian Meise,Tobias Cord-Landwehr,Christoph Boeddeker,Marc Delcroix,Tomohiro Nakatani,Reinhold Haeb-Umbach
备注:Accepted at ICASSP 2026
摘要:麦克风阵列的声音捕获开辟了利用空间的可能性,除了频谱,信息日记和信号增强,两个重要的任务,在会议转录。然而,如果说话者移动,则不存在空间中的位置到说话者的一对一映射。在这里,我们解决这个问题,提出了一种新的联合空间和频谱混合模型,其两个子模型是松散耦合的说话人和位置指数之间的关系建模概率。因此,可以联合利用空间和频谱信息,同时允许说话者从不同位置说话。LibriCSS数据集与模拟扬声器位置变化的实验表明,紧密耦合的子系统有很大的改善。
摘要:Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping of positions in space to speakers if speakers move. Here, we address this by proposing a novel joint spatial and spectral mixture model, whose two submodels are loosely coupled by modeling the relationship between speaker and position index probabilistically. Thus, spatial and spectral information can be jointly exploited, while at the same time allowing for speakers speaking from different positions. Experiments on the LibriCSS data set with simulated speaker position changes show great improvements over tightly coupled subsystems.


【2】Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs
标题:基于Timbre-Aware LLM的直接语音翻译可扩展到多语言对
链接:https://arxiv.org/abs/2601.16023

作者:Lalaram Arya,Mrinmoy Bhattacharjee,Adarsh C. R.,S. R. Mahadeva Prasanna
备注:13 pages
摘要:直接语音到语音翻译(S2 ST)由于其将语音从一种语言翻译成另一种语言的能力而受到越来越多的关注,同时减少了传统级联流水线中固有的错误传播和延迟。然而,现有的直接S2 ST系统继续面临着显着的挑战,包括不稳定的语义声学对齐时,并行语音数据是稀缺的,难以保持扬声器的身份,有限的多语种可扩展性。在这项工作中,我们介绍DS 2ST-LM,一个可扩展的,单阶段的直接S2 ST框架,利用多语言的大型语言模型(LLM)。该架构集成了Whisper语音编码器、可学习投影模块、Qwen 2 -0.5B LLM和音色控制声码器。我们构建了GigaS 2S-1000,一个1000小时的双语语料库,通过扩展GigaST数据集与高保真合成目标语音,并表明这种合成数据在一定程度上弥补了数据稀缺。我们研究了两种语义令牌生成策略:语音派生的S3令牌和由预训练的LLM生成的文本派生令牌,并分析了它们对训练稳定性和语义一致性的影响。我们进一步评估了三种投影架构(线性,Conv 1D-Linear和Q-Former),并观察到虽然高容量投影仪收敛速度更快,但简单的线性投影仪实现了更高的性能。大量的实验表明,DS 2ST-LM在词汇(BLEU,METEOR)和语义(BLEURT,COMET)指标上优于传统的级联和ST(Qwen-Audio)+ TTS基线,同时扩展到多种语言对,包括法语,西班牙语,德语,印地语,孟加拉语和乌尔都语。此外,我们将音色感知的语音合成,以保持扬声器信息,使DS 2ST-LM超过以前的直接S2 ST系统在扬声器的相似性和感知自然度。
摘要:Direct Speech-to-Speech Translation (S2ST) has gained increasing attention for its ability to translate speech from one language to another, while reducing error propagation and latency inherent in traditional cascaded pipelines. However, existing direct S2ST systems continue to face notable challenges, including instability in semantic-acoustic alignment when parallel speech data is scarce, difficulty in preserving speaker identity, and limited multilingual scalability. In this work, we introduce DS2ST-LM, a scalable, single-stage direct S2ST framework leveraging a multilingual Large Language Model (LLM). The architecture integrates a Whisper speech encoder, a learnable projection module, a Qwen2-0.5B LLM, and a timbre-controlled vocoder. We construct GigaS2S-1000, a 1000-hour bilingual corpus by extending the GigaST dataset with high-fidelity synthetic target speech, and show that this synthetic data alleviates data scarcity to some extent. We investigate two semantic token generation strategies: speech-derived S3 tokens and text-derived tokens generated by a pre-trained LLM, and analyze their impact on training stability and semantic consistency. We further evaluate three projection architectures (Linear, Conv1D-Linear, and Q-Former) and observe that while higher-capacity projectors converge faster, the simple Linear projector achieves higher performance. Extensive experiments demonstrate that DS2ST-LM outperforms traditional cascaded and ST (Qwen-Audio) + TTS baselines across both lexical (BLEU, METEOR) and semantic (BLEURT, COMET) metrics, while extending to multiple language pairs, including French, Spanish, German, Hindi, Bengali, and Urdu. Furthermore, we incorporate timbre-aware speech synthesis to preserve speaker information, enabling DS2ST-LM to surpass prior direct S2ST systems in both speaker similarity and perceptual naturalness.


【3】A Stabilized Hybrid Active Noise Control Algorithm of GFANC and FxNLMS with Online Clustering
标题:一种具有在线分簇的稳定的GF非国大和FxNRMS混合主动噪音控制算法
链接:https://arxiv.org/abs/2601.15889

作者:Zhengding Luo,Haozhe Ma,Boxiang Wang,Ziyi Yang,Dongyuan Shi,Woon-Seng Gan
备注:Accepted by 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)
摘要:滤波-x归一化最小均方(FxNLMS)算法存在收敛速度慢和发散风险的问题,尽管它在充分自适应后可以实现低稳态误差。相比之下,生成式固定滤波器有源噪声控制(GFANC)方法提供了快速的响应速度,但其缺乏适应性可能导致大的稳态误差。本文提出了一种混合GFANC-FxNLMS算法,利用这两种方法的互补优势。在混合GFANC-FxNLMS算法中,GFANC提供帧级控制滤波器作为FxNLMS的初始化,而FxNLMS以采样率执行连续自适应。GFANC生成的滤波器的微小变化可能会重复重新初始化FxNLMS,中断其自适应过程并使系统不稳定。在线聚类模块的引入,以避免不必要的重新初始化,提高系统的稳定性。仿真结果表明,该算法具有响应速度快、稳态误差小、稳定性高等优点,且只需要一个预训练的宽带滤波器。
摘要:The Filtered-x Normalized Least Mean Square (FxNLMS) algorithm suffers from slow convergence and a risk of divergence, although it can achieve low steady-state errors after sufficient adaptation. In contrast, the Generative Fixed-Filter Active Noise Control (GFANC) method offers fast response speed, but its lack of adaptability may lead to large steady-state errors. This paper proposes a hybrid GFANC-FxNLMS algorithm to leverage the complementary advantages of both approaches. In the hybrid GFANC-FxNLMS algorithm, GFANC provides a frame-level control filter as an initialization for FxNLMS, while FxNLMS performs continuous adaptation at the sampling rate. Small variations in the GFANC-generated filter may repeatedly reinitialize FxNLMS, interrupting its adaptation process and destabilizing the system. An online clustering module is introduced to avoid unnecessary re-initializations and improve system stability. Simulation results show that the proposed algorithm achieves fast response, very low steady-state error, and high stability, requiring only one pre-trained broadband filter.


【4】Distributed Multichannel Active Noise Control with Asynchronous Communication
标题:采用同步通信的分布式多通道主动噪音控制
链接:https://arxiv.org/abs/2601.15653

作者:Junwei Ji,Dongyuan Shi,Boxiang Wang,Ziyi Yang,Haowen Li,Woon-Seng Gan
摘要:分布式多通道有源噪声控制(DMCANC)通过将集中控制的计算负载分配给多个低成本节点,在大空间区域内提供有效的降噪。然而,传统的DMCANC方法通常假设同步通信并且需要频繁的数据交换,从而导致高通信开销。为了提高效率和适应性,这项工作提出了一种异步通信策略,其中每个节点执行权重约束的滤波x LMS(WCFxLMS)算法,并独立地请求通信时,其本地降噪性能下降。根据请求,其他节点发送其本地控制滤波器与WCFxLMS中的中心点之间的权重差,然后将其积分以更新控制滤波器和中心点。这种设计使节点能够异步操作,同时保持合作行为。仿真结果表明,所提出的异步通信DMCANC(ACDMCANC)系统保持有效的降噪与显着降低通信负载,提供了改进的异构网络的可扩展性。
摘要:Distributed multichannel active noise control (DMCANC) offers effective noise reduction across large spatial areas by distributing the computational load of centralized control to multiple low-cost nodes. Conventional DMCANC methods, however, typically assume synchronous communication and require frequent data exchange, resulting in high communication overhead. To enhance efficiency and adaptability, this work proposes an asynchronous communication strategy where each node executes a weight-constrained filtered-x LMS (WCFxLMS) algorithm and independently requests communication only when its local noise reduction performance degrades. Upon request, other nodes transmit the weight difference between their local control filter and the center point in WCFxLMS, which are then integrated to update both the control filter and the center point. This design enables nodes to operate asynchronously while preserving cooperative behavior. Simulation results demonstrate that the proposed asynchronous communication DMCANC (ACDMCANC) system maintains effective noise reduction with significantly reduced communication load, offering improved scalability for heterogeneous networks.


【5】DynamicSound simulator for simulating moving sources and microphone arrays
标题:DynamicSound模拟器,用于模拟移动源和麦克风阵列
链接:https://arxiv.org/abs/2601.15433

作者:Luca Barbisan,Marco Levorato,Fabrizio Riente
摘要:开发用于声音分类、检测和定位的算法需要大量灵活且真实的音频数据,尤其是在利用现代机器学习和波束成形技术时。然而,大多数现有的声学模拟器都是为室内环境量身定制的,并且仅限于静态声源,这使得它们不适合涉及移动源,移动麦克风或长距离传播的场景。本文介绍了DynamicSound一个开源的声学仿真框架,用于从一个或多个声源生成多通道音频,并可以在三维空间中连续移动它们,并通过任意配置的麦克风阵列记录。所提出的模型明确考虑有限的声音传播延迟,多普勒效应,距离相关的衰减,空气吸收,从平面表面的一阶反射,产生时间上一致的空间音频信号。与传统的单声道或立体声模拟器不同,所提出的系统为任意数量的虚拟麦克风合成音频,准确地再现麦克风间的时间延迟,电平差和由环境引起的频谱着色。与现有开源工具的比较评估表明,所生成的信号在不同的源位置和声学条件下保持了较高的空间保真度。通过在受控和可重复的条件下生成逼真的多声道音频,所提出的开放式框架为现代空间音频和声源定位算法的开发、训练和评估提供了灵活且可再现的工具。
摘要:Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing acoustic simulators are tailored for indoor environments and are limited to static sound sources, making them unsuitable for scenarios involving moving sources, moving microphones, or long-distance propagation. This paper presents DynamicSound an open-source acoustic simulation framework for generating multichannel audio from one or more sound sources with the possibility to move them continuously in three-dimensional space and recorded by arbitrarily configured microphone arrays. The proposed model explicitly accounts for finite sound propagation delays, Doppler effects, distance-dependent attenuation, air absorption, and first-order reflections from planar surfaces, yielding temporally consistent spatial audio signals. Unlike conventional mono or stereo simulators, the proposed system synthesizes audio for an arbitrary number of virtual microphones, accurately reproducing inter-microphone time delays, level differences, and spectral coloration induced by the environment. Comparative evaluations with existing open-source tools demonstrate that the generated signals preserve high spatial fidelity across varying source positions and acoustic conditions. By enabling the generation of realistic multichannel audio under controlled and repeatable conditions, the proposed open framework provides a flexible and reproducible tool for the development, training, and evaluation of modern spatial audio and sound-source localization algorithms.


【6】PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation
标题:PF-D2 M:通用舞蹈到音乐生成的无姿势扩散模型
链接:https://arxiv.org/abs/2601.15872

作者:Jaekwon Im,Natalia Polouliakh,Taketo Akama
备注:4 pages, 2 figures
摘要:Dance-to-Music生成旨在生成与舞蹈动作一致的音乐。现有的方法通常依赖于从单个人类舞者和有限的舞蹈到音乐数据集提取的身体运动特征,这限制了它们对涉及多个舞者和非人类舞者的真实世界场景的性能和适用性。在本文中,我们提出了PF-D2 M,一个通用的基于扩散的舞蹈到音乐生成模型,结合了从舞蹈视频中提取的视觉特征。PF-D2 M采用渐进式训练策略进行训练,可有效解决数据稀缺和泛化挑战。客观和主观的评估都表明,PF-D2 M在舞蹈音乐对齐和音乐质量方面达到了最先进的性能。
摘要:Dance-to-music generation aims to generate music that is aligned with dance movements. Existing approaches typically rely on body motion features extracted from a single human dancer and limited dance-to-music datasets, which restrict their performance and applicability to real-world scenarios involving multiple dancers and non-human dancers. In this paper, we propose PF-D2M, a universal diffusion-based dance-to-music generation model that incorporates visual features extracted from dance videos. PF-D2M is trained with a progressive training strategy that effectively addresses data scarcity and generalization challenges. Both objective and subjective evaluations show that PF-D2M achieves state-of-the-art performance in dance-music alignment and music quality.


【7】Bridging the Perception Gap: A Lightweight Coarse-to-Fine Architecture for Edge Audio Systems
标题:弥合感知差距:边缘音频系统的轻量级从粗到精架构
链接:https://arxiv.org/abs/2601.15676

作者:Hengfan Zhang,Yueqian Lin,Hai Helen Li,Yiran Chen
备注:10 pages, 3 figures, 2 tables. Preprint
摘要:在边缘基础设施上部署音频语言模型(Audio-LLM)暴露了感知深度和计算效率之间的持续紧张关系。轻量级本地模型往往会产生被动感知--通用摘要会错过多步音频推理所需的细微证据--而不加选择的云卸载会导致不可接受的延迟、带宽成本和隐私风险。我们提出了CoFi-Agent(Tool-Augmented Coarse-to-Fine Agent),这是一种针对边缘服务器和网关的混合架构。它执行快速局部感知,并仅在检测到不确定性时触发条件取证细化。CoFi-Agent在本地7 B Audio-LLM上运行初始单遍,然后云控制器门控困难的情况并为设备上工具发布轻量级计划,例如临时重新监听和本地ASR。在MMAR基准测试中,CoFi-Agent将准确率从27.20%提高到53.60%,同时比始终在线的调查管道实现了更好的准确性-效率权衡。总体而言,CoFi-Agent在实际系统限制下通过工具支持的、有条件的边缘云协作来弥合感知差距。
摘要:Deploying Audio-Language Models (Audio-LLMs) on edge infrastructure exposes a persistent tension between perception depth and computational efficiency. Lightweight local models tend to produce passive perception - generic summaries that miss the subtle evidence required for multi-step audio reasoning - while indiscriminate cloud offloading incurs unacceptable latency, bandwidth cost, and privacy risk. We propose CoFi-Agent (Tool-Augmented Coarse-to-Fine Agent), a hybrid architecture targeting edge servers and gateways. It performs fast local perception and triggers conditional forensic refinement only when uncertainty is detected. CoFi-Agent runs an initial single-pass on a local 7B Audio-LLM, then a cloud controller gates difficult cases and issues lightweight plans for on-device tools such as temporal re-listening and local ASR. On the MMAR benchmark, CoFi-Agent improves accuracy from 27.20% to 53.60%, while achieving a better accuracy-efficiency trade-off than an always-on investigation pipeline. Overall, CoFi-Agent bridges the perception gap via tool-enabled, conditional edge-cloud collaboration under practical system constraints.


【8】Qwen3-TTS Technical Report
标题:Qwen 3-TTC技术报告
链接:https://arxiv.org/abs/2601.15621

作者:Hangrui Hu,Xinfa Zhu,Ting He,Dake Guo,Bin Zhang,Xiong Wang,Zhifang Guo,Ziyue Jiang,Hongkun Hao,Zishan Guo,Xinyu Zhang,Pei Zhang,Baosong Yang,Jin Xu,Jingren Zhou,Junyang Lin
备注:https://github.com/QwenLM/Qwen3-TTS
摘要:在本报告中,我们介绍了Qwen 3-TTS系列,这是一系列先进的多语言,可控,强大和流式文本到语音模型。Qwen 3-TTS支持最先进的3秒语音克隆和基于预处理的控制,允许创建全新的语音和对输出语音的细粒度操作。Qwen 3-TTS采用双轨LM架构进行实时合成,并配备两个语音标记器:1)Qwen-TTS-Tokenizer-25 Hz是一个强调语义内容的单码本编解码器,可与Qwen-Audio无缝集成,并通过逐块DiT实现流式波形重构。2)Qwen-TTS-Tokenizer-12 Hz实现了极高的比特率降低和超低延迟流传输,通过其12.5 Hz、16层多码本设计和轻量级因果ConvNet实现了即时的第一个数据包发射(97\,\mathrm {ms}$)。广泛的实验表明,在不同的客观和主观基准(例如,TTS多语言测试集、InstructTTSEval和我们的长语音测试集)。为了促进社区的研究和开发,我们在Apache 2.0许可证下发布了tokenizer和模型。
摘要:In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control, allowing both the creation of entirely novel voices and fine-grained manipulation over the output speech. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS adopts a dual-track LM architecture for real-time synthesis, coupled with two speech tokenizers: 1) Qwen-TTS-Tokenizer-25Hz is a single-codebook codec emphasizing semantic content, which offers seamlessly integration with Qwen-Audio and enables streaming waveform reconstruction via a block-wise DiT. 2) Qwen-TTS-Tokenizer-12Hz achieves extreme bitrate reduction and ultra-low-latency streaming, enabling immediate first-packet emission ($97\,\mathrm{ms}$) through its 12.5 Hz, 16-layer multi-codebook design and a lightweight causal ConvNet. Extensive experiments indicate state-of-the-art performance across diverse objective and subjective benchmark (e.g., TTS multilingual test set, InstructTTSEval, and our long speech test set). To facilitate community research and development, we release both tokenizers and models under the Apache 2.0 license.


【9】DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
标题:DeepASMR:基于LLM的Zero-ShotASMR语音生成,适合任何语音的任何人
链接:https://arxiv.org/abs/2601.15596

作者:Leying Zhang,Tingxiao Zhou,Haiyang Sun,Mengxiao Bi,Yanmin Qian
摘要:虽然现代文本到语音(TTS)系统实现了阅读风格语音的高保真度,但它们很难生成自主感觉经络反应(ASMR),这是一种对放松至关重要的专门的低强度语音风格。固有的挑战包括ASMR的微妙,往往无声的特点和要求zero-shot扬声器自适应。在本文中,我们介绍了DeepASMR,这是第一个为zero-shot ASMR生成而设计的框架。我们证明,一个扬声器的普通,阅读风格的语音的一个简短的片段是足以合成高保真ASMR在他们的声音,消除了需要从目标扬声器的耳语训练数据。方法上,我们首先确定,离散语音令牌提供了一个软分解的ASMR风格从扬声器音色。利用这一洞察力,我们提出了一个两阶段的管道,其中包括一个大语言模型(LLM)的内容风格的编码和流匹配的声学解码器的音色重建。此外,我们贡献了DeepASMR-DB,一个全面的670小时的英汉多说话人ASMR语音语料库,并介绍了一种新的评估协议,集成了客观指标,人类听力测试,基于LLM的评分和清音语音分析。大量的实验证实,DeepASMR在任何语音的ASMR生成中实现了最先进的自然度和风格保真度,同时保持了正常语音合成的竞争力。
摘要:While modern Text-to-Speech (TTS) systems achieve high fidelity for read-style speech, they struggle to generate Autonomous Sensory Meridian Response (ASMR), a specialized, low-intensity speech style essential for relaxation. The inherent challenges include ASMR's subtle, often unvoiced characteristics and the demand for zero-shot speaker adaptation. In this paper, we introduce DeepASMR, the first framework designed for zero-shot ASMR generation. We demonstrate that a single short snippet of a speaker's ordinary, read-style speech is sufficient to synthesize high-fidelity ASMR in their voice, eliminating the need for whispered training data from the target speaker. Methodologically, we first identify that discrete speech tokens provide a soft factorization of ASMR style from speaker timbre. Leveraging this insight, we propose a two-stage pipeline incorporating a Large Language Model (LLM) for content-style encoding and a flow-matching acoustic decoder for timbre reconstruction. Furthermore, we contribute DeepASMR-DB, a comprehensive 670-hour English-Chinese multi-speaker ASMR speech corpus, and introduce a novel evaluation protocol integrating objective metrics, human listening tests, LLM-based scoring and unvoiced speech analysis. Extensive experiments confirm that DeepASMR achieves state-of-the-art naturalness and style fidelity in ASMR generation for anyone of any voice, while maintaining competitive performance on normal speech synthesis.


机器翻译由腾讯交互翻译提供,仅供参考