今日论文合集:CS.SD语音与音频 | 共 10 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


快速导航

1. 语音识别与关键词检测 2 篇

2. 语音合成与声音生成 3 篇

3. 说话人识别、验证与分离 2 篇

4. 数据集、基准与评测 1 篇

5. 其他/综合语音音频 2 篇


1. 语音识别与关键词检测 | 2 篇

1. Sparse Weight and Edge Circuit Discovery in Transformer-based Acoustic Models

基于Transformer的声学模型中的稀疏权重与边电路发现

AI 总结:本研究将电路发现框架DiscoGP扩展到语音编码器,发现极紧凑电路性能媲美完整模型,并推出内存高效变体,拓展了机制解释的应用范围。

链接:https://arxiv.org/abs/2609.10645

机构:University of Toronto(多伦多大学)

作者:Jiankun Wei, Ewan Dunbar, Gerald Penn

英文摘要:Transformer-based foundation models are powerful but opaque, motivating Mechanistic Interpretation methods to uncover the black-box by identifying small computation subgraphs responsible for a task. DiscoGP is a joint weight-and-edge circuit discovery framework originally developed for text decoders. We extend DiscoGP to speech encoders and present, to our knowledge, the first circuit discovery study for modern speech foundation models. Across HuBERT and Wav2Vec 2.0 on several speech classification tasks, we find that the discovered circuits are extremely compact, yet often match or even exceed the performance of the full pretrained encoder with the same downstream head. Through ablations, we show that these circuits reflect pretrained computation rather than random structure or task-head artifacts. We also introduce a memory-efficient DiscoGP variant that reduces the GPU memory cost of edge-circuit discovery at runtime from quartic to cubic. Overall, our results broaden Mechanistic Interpretation beyond text decoders and show that circuit-level analysis can reveal both explanatory structure and unexpected functional behavior in speech encoders.


2. X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

X-AuT:基于跨尺度蒸馏的语音大语言模型渐进式音频编码器压缩

AI 总结:X-AuT通过渐进式层剪枝和跨尺度蒸馏压缩语音LLM的音频编码器,在Qwen3-ASR上以更少参数降低错误率,优于直接剪枝。

链接:https://arxiv.org/abs/2609.11412

机构:XPeng Inc.(小鹏汽车)

作者:Haojun Zhang, Yi Zou, Min Chen, Qize Yu, Lianrui Fan, Xini Ding, Hao Li, Shuchang Zhou, Xianming Liu, Shiyu Huang

英文摘要:Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese--English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 18$\rightarrow$14 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: this https URL


2. 语音合成与声音生成 | 3 篇


3. Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

面向细粒度情感与时长控制的自然语言后训练零样本文本到语音合成

AI 总结:针对现有TTS系统难以实现单句内细粒度情感与时长控制的问题,提出统一后训练框架,结合监督微调与强化学习,实现自然语言控制的片段级情感和时长调节,显著提升可控性并保持语音质量。

链接:https://arxiv.org/abs/2609.11523

机构:Nankai University(南开大学)

作者:Lianru Gao, Yujie Guo, Yong Qin

英文摘要:Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditioning, making such fine-grained control difficult to achieve. In light of this, and inspired by the success of post-training in large language models, we propose a unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration. Supervised fine-tuning establishes instruction-conditioned speech generation, while reinforcement learning with group relative policy optimization refines control accuracy using emotion and duration rewards alongside content and speaker preservation objectives. By reusing the pretrained architecture, our approach avoids additional inference-time control modules. Experiments demonstrate significantly improved fine-grained controllability while maintaining speech intelligibility and speaker identity, highlighting post-training as a practical approach to extending existing speech synthesis models.


4. ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

ZipCodec:超低帧率流式语音编码

AI 总结:ZipCodec通过WavLM蒸馏、Transformer架构和球形量化,在6.25Hz超低帧率下实现高质量流式语音编码,显著优于现有编解码器。

链接:https://arxiv.org/abs/2609.11642

机构:Concordia University(康考迪亚大学); Mila-Quebec AI Institute(Mila-魁北克人工智能研究所); Université Laval(拉瓦尔大学)

作者:Luca Della Libera, Cem Subakan, Mirco Ravanelli

英文摘要:Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation with a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware streaming decoder. Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate. Despite its 842M parameters, ZipCodec achieves real-time single-stream inference on a consumer-grade CPU. Demo samples, code and checkpoints are available at this https URL.


5. Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

连续时间声学建模:基于神经受控微分方程

AI 总结:本文提出用神经受控微分方程实现连续时间声学建模,使隐藏状态随音素内容和时长演化,提升情感强度排序一致性并保持表达质量。

链接:https://arxiv.org/abs/2609.11725

机构:University of Sheffield(谢菲尔德大学)

作者:Mattias Cross, Minghui Zhao, Anton Ragni

英文摘要:Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use of duration typically changes only where and how often latent states appear, not the values of the states themselves. This paper proposes a continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs). We formulate the phone representation as a temporally parameterised control path and use a neural acoustic vector field to produce a continuous-time hidden state whose values evolve with phonetic content and duration-derived timing. The resulting trajectory can be sampled at discrete points and integrated into a standard acoustic decoder pipeline. Objective results contrast CDEs and typical recurrent models. Subjective results suggest that CDE-based models evaluating one phone per step can improve rank-order agreement between synthesised and reference emotion intensity while maintaining comparable emotion-expression quality to a strong baseline. Additional experiments with half-phone step-sizes suggest that temporal resolution changes the trade-off between style tracking and absolute calibration. These results position CDEs as a promising design space for continuous-time and duration-aware style-sensitive TTS.


3. 说话人识别、验证与分离 | 2 篇


6. Xiaomi-CocktailASR-1 Technical Report

Xiaomi-CocktailASR-1 技术报告

AI 总结:针对多说话人场景中鸡尾酒会问题,提出基于大语言模型的端到端目标说话人识别架构Xiaomi-CocktailASR-1,利用参考语音作为声纹提示直接转录目标语音,无需语音分离,具备弃权能力和思维链推理,在多个基准上达到最优性能。

链接:https://arxiv.org/abs/2609.11274

机构:Xiaomi Inc.(小米公司)

作者:Yiru Zhang, Hang Su, Lichun Fan, Ying Zeng, Chang Liu, Yifeng Wang, Yuquan Liang, Tao Li, Lian Li, Wenhao Yang, Jian Luan, Cong Zou, Heng Qu

英文摘要:Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voiceprint prompts, it directly transcribes the target speaker's speech without requiring speech separation. Xiaomi-CocktailASR-1 maintains competitive performance in single-speaker scenarios, comparable to mainstream ASR models. It also features a negative sample rejection capability, outputting empty text when the target speaker is absent from the mixed speech. Additionally, Xiaomi-CocktailASR-1 supports a Chain-of-Thought (CoT) reasoning mode to provide explicit reasoning steps. Extensive experiments on various synthetic and real-world multispeaker benchmarks demonstrate that Xiaomi-CocktailASR-1 achieves state-of-the-art performance, effectively addressing the cocktail party problem through a unified architecture that balances multispeaker and single-speaker recognition accuracy, along with rejection capability.


7. EConv-TasNet: Efficient Conv-TasNet for Effective Speech Separation

EConv-TasNet:高效Conv-TasNet用于有效语音分离

AI 总结:提出eConv-TasNet,通过分组早期拆分和多组特征聚合模块,在减少22.4%参数、提升18.9%推理速度的同时,SI-SNRi提高14.0%-28.0%,实现高效语音分离。

链接:https://arxiv.org/abs/2609.11342

机构:Novatek Microelectronics Corporation(联咏科技股份有限公司)

作者:Pei-Chun Chang, Chuan-Yi Liu

英文摘要:Conv-TasNet has served as a strong baseline for time-domain speech separation, and many studies have extended it with advanced architectures such as dual-path networks, U-Nets, and attention mechanisms. However, these methods often introduce high computational cost and complexity, limiting their deployment in resource-constrained scenarios. To address this issue, we propose eConv-TasNet, an efficient variant of Conv-TasNet that improves both effectiveness and efficiency without relying on resource-intensive modules. The proposed model consists of a group-wise early-splitting (GES) module and a multi-group feature aggregation (MGFA) module. GES generates discriminative speaker embeddings at intermediate stages, while MGFA progressively aggregates these group-level representations for refined mask estimation. Experimental results show that eConv-TasNet reduces model size by 22.4%, accelerates inference by 18.9%, and improves SI-SNRi by 14.0%-28.0% across three public benchmarks. Moreover, it achieves competitive performance compared with state-of-the-art methods while requiring significantly fewer parameters and lower inference cost. These results demonstrate a favorable efficiency-effectiveness trade-off for edge deployment.


4. 数据集、基准与评测 | 1 篇

8. Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data

Project Qualia:从会话共现数据中恢复体验性音乐结构

AI 总结:Project Qualia通过分析大规模收听会话数据,训练Song2Vec模型并采用艺术家残差方法,成功从嵌入中恢复了独立于艺术家身份的体验性音乐结构,验证了该结构的可学习性。

链接:https://arxiv.org/abs/2609.10862

机构:DePaul University(德保罗大学); Hampton University(汉普顿大学)

作者:Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige

英文摘要:This report presents results from Project Qualia, an ongoing effort to determine whether experiential similarity between songs, a structure not captured by genre or metadata taxonomies, can be recovered from real listening behavior. We constructed a large-scale dataset of listening sessions, comprising 1.29 billion scrobbles collected from 9,396 users via the this http URL API and reduced through a preprocessing pipeline to 531.6 million training scrobbles across 28.6 million sessions. On this corpus, we trained a skip-gram Word2Vec model (Song2Vec), treating each session as a sentence and each track as a token. As anticipated, the resulting embedding space was dominated by artist identity, a consequence of single-artist runs within sessions. To test for a subtler, artist-independent signal, we developed an artist-residual procedure: subtracting each artist's centroid from its tracks' embeddings and evaluating whether the remainder retained structure. Mean cross-artist cosine similarity fell from 0.2487 in raw embedding space to 0.0005 in residual space, yet 4,577 cross-artist track pairs retained cosine similarity $\ge 0.70$ in residual space, forming coherent genre- and era-based clusters, including trip-hop, 1990s grunge, 2020 mainstream pop, and cross-composer classical piano pairs at cosine similarity up to 0.95. These results confirm that the training data contains experiential structure independent of artist identity, establishing an empirical basis for an architecture designed to learn this experiential layer directly.


5. 其他/综合语音音频 | 2 篇

9. Learned Continuous Synthesis of Quadratic Difference Tone Spectra

二次差音频谱的学习连续合成

AI 总结:本文提出一种基于神经网络的连续合成方法,学习二次差音频谱的逆映射,解决先前数值方法的不连续和难控制问题,并实现实时版本,适用于音乐应用。

链接:https://arxiv.org/abs/2609.10913

机构:Universitat Pompeu Fabra(庞培法布拉大学); University of Birmingham(伯明翰大学); Pontificia Universidad Católica de Chile(智利天主教大学)

作者:Esteban Gutiérrez, Behzad Haki, Christopher Haworth, Xavier Serra, Rodrigo Cádiz

英文摘要:Quadratic difference tones (QDTs) are a species of auditory distortion product in which a "phantom" pure tone, absent from the acoustic signal, is clearly audible to listeners. Exploiting this phenomenon, one can synthesize harmonically rich tones for musical purposes, a technique called Quadratic Difference Tone Spectrum (QDTS) synthesis. Previous works have introduced numerical methods to synthesize QDTS based on the distortion function, which links a target QDTS and an overtone-structured carrier signal. While accurate, these methods were stochastic and discontinuous, making them difficult to control for musical purposes and effectively limiting them to stationary signals. This paper proposes a neural network-based approach that learns an approximate inverse of the distortion mapping in an autoencoder-like configuration, producing a continuous approximation that addresses prior limitations. Experimental results show that, although slightly less numerically precise, the method is sufficient for perceptual and musical applications. We also implement a real-time version in Max and evaluate its performance. Various sound examples demonstrate its expressive and musical potential. The source code, audio examples, tutorials, and software accompanying this work are available at this https URL


10. Copying Versus Randomization in Lempel-Ziv Music Synthesis

Lempel-Ziv音乐合成中的复制与随机化

AI 总结:利用Lempel-Ziv压缩生成音乐音符,通过调整字典平均序列长度来控制复制与随机化倾向。

链接:https://arxiv.org/abs/2609.11353

作者:Nadav Mishan, Ram Zamir

英文摘要:We utilize Lempel-Ziv universal compression for music note generation. We control the algorithm's tendency to over-copy or under-copy training data by manipulating the average sequence length saved in the dictionary.