今日论文合集:CS.SD语音与音频 | 共 7 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音合成与声音生成 1 篇

2. 说话人识别、验证与分离 1 篇

3. 语音增强、降噪与音频修复 1 篇

4. 音乐信息检索与音乐生成 1 篇

5. 数据集、基准与评测 1 篇

6. 其他/综合语音音频 2 篇

1. 语音合成与声音生成 | 1 篇

1. A Quantized Native Runtime for On-Device Semantic Audio Generation

用于设备端语义音频生成的量化原生运行时

AI 总结:研究针对设备端语义音频生成,提出无依赖原生运行时aria,通过量化研究降低内存需求,经实验证明8位精度无质量损失,4位有小成本但能满足特定设备运行大模型,且生成速度快,还具备激活引导功能,为设备端语义音频提供实用基础。

链接:https://arxiv.org/abs/2607.08526

作者:Matteo Spanio, Antonio Rodà

英文摘要:Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. We present \textit{aria}, a dependency-free native runtime that runs the complete text-to-music pipeline of Stable Audio~3 (SA3) on ordinary GPUs, CPU-only machines, and a Raspberry~Pi~5, with no Python or deep-learning framework underneath. Our main contribution is a study of quantization: running the model at lower numerical precision to fit tight memory budgets, saving memory in place rather than adding to it. Because the runtime owns every internal tensor, it also exposes activation steering, a low-cost way to steer what the model generates. We judge the quality cost with three independent measures of the output (prompt adherence, overall audio quality, taste preservation), each compared against the ordinary variation between random seeds. Eight-bit precision shows no measurable quality loss on any measure while sharply cutting memory, and it is the fastest mode on the GPU; four-bit adds a small, bounded cost but shrinks the footprint enough to run the $1.2$-billion-parameter model on an $8$\,GB Pi. Against the official implementation, aria matches or exceeds generation speed and starts about seven times faster. A case study of the steering interface generates music carrying taste associations (\emph{sonic seasoning}), with genuine but bounded control for a subset of attributes. These results make a compact, quantized runtime with built-in control a practical basis for on-device semantic audio in Internet-of-Sounds settings. The \textit{aria} runtime is released at this https URL.

2. 说话人识别、验证与分离 | 1 篇

2. PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

PS4:用于真实目标说话人提取的代理监督联合训练

AI 总结:针对真实对话混合语音中目标说话人提取模型训练难题,提出PS4框架。构建大规模语料库,采用代理监督联合训练策略,用四个互补目标微调基于BSRNN的模型,在REAL-T挑战中总体排名第二,取得良好效果。

链接:https://arxiv.org/abs/2607.08111

作者:Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng

英文摘要:Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational mixtures, with two main contributions. First, we construct a large-scale corpus of 71,771 training samples derived from four public datasets, covering both Chinese and English scenarios. Each sample contains an overlapping speech mixture, per-speaker enrollment audio, a ground-truth transcript, and frame-level voice activity labels. Second, we propose a proxy-supervised joint training strategy that fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross-entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. Starting from a publicly available pre-trained checkpoint, only the BSRNN separator is updated during fine-tuning. On the REAL-T challenge leaderboard, PS4 ranks 2nd overall, achieving the best speaker similarity and timing F1 among all submitted systems.

3. 语音增强、降噪与音频修复 | 1 篇

3. It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement

只需少量即可实现 TANGO:一种用于双耳语音增强的量化分布式模型

AI 总结:研究基于神经网络的多通道语音增强系统在资源受限设备上的低精度推理,以 TANGO 为对象,评估其神经组件量化方法,利用空间滤波阶段对量化误差的补偿能力简化模型,结合多种技术减小模型大小和计算复杂度并保持性能。

链接:https://arxiv.org/abs/2607.08645

作者:Zahra Benslimane, Pierre Chouteau, Martyna Poreba, Fabrice Auzanneau, Michal Szczepanski, Fabian Chersi, Romain Serizel

英文摘要:Neural network-based multichannel speech enhancement systems achieve strong enhancement performance, but their computational and memory requirements limit deployment on resource-constrained devices. This paper investigates low-precision inference for TANGO, a hybrid distributed binaural speech enhancement system combining neural mask estimation with spatial filtering. We evaluate post-training quantization and quantization-aware training for the neural components, and analyze how quantization errors in the mask estimators propagate through the downstream spatial filtering stage. Our analysis shows that, although quantization degrades intermediate mask estimates, the spatial filtering stage compensates for most quantization-induced errors. Leveraging this robustness, we simplify TANGO into MN-TANGO, reducing both model size and computational complexity while maintaining comparable final performance. By combining INT8 weight-and-activation quantization with ERB compression and grouped recurrent layers, the most compact MN-TANGO reaches 4.65 MMAC/s and 0.177 MB.

4. 音乐信息检索与音乐生成 | 1 篇

4. MuScriptor: An Open Model for Multi-Instrument Music Transcription

MuScriptor:一种用于多乐器音乐转录的开放模型

AI 总结:研究多乐器音乐转录问题,核心方法是分析合成数据预训练有效性,结合真实音乐音频微调与强化学习后训练,还引入乐器存在条件定制转录,主要贡献是发布MuScriptor开放模型用于多种音乐录音。

链接:https://arxiv.org/abs/2607.08168

作者:Simon Rouard, Michael Krause, Axel Roebel, Carl-Johann Simon-Gabriel, Alexandre Défossez

英文摘要:Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.

5. 数据集、基准与评测 | 1 篇

5. MulTTiPop: A Multitrack Transcription Dataset for Pop Music

MulTTiPop:一个用于流行音乐的多轨转录数据集

AI 总结:介绍用于评估自动音乐转录模型的MulTTiPop数据集,通过对Lakh MIDI和TheoryTab数据集歌曲片段基于元数据匹配等方式收集,评估模型发现有改进空间,最佳模型起始F1得分为38%。

链接:https://arxiv.org/abs/2607.08756

作者:Nathan Pruyne, Benjamin Stoler, William Chen, Chien-yu Huang, Shinji Watanabe, Chris Donahue

英文摘要:We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 38% Onset F1. More details and sound examples of MulTTiPop are available at this https URL.

6. 其他/综合语音音频 | 2 篇

6. A Self-Supervised Approach for Minimal-Annotation Hydroacoustic Data Exploration

一种用于最少标注水声数据探索的自监督方法

AI 总结:针对被动水声监测中人工标注成本高、监督检测方法需大量标注数据的问题,提出自监督探索管道,用预训练的MAE提取特征,经聚合、聚类识别水声模式,应用于实际数据集,评估显示该方法有实用价值。

链接:https://arxiv.org/abs/2607.07733

机构:Laboratoire de Géologie, Ecole Normale Supérieure/CNRS UMR 8538, PSL Research University(地质实验室,巴黎高等师范学校/法国国家科学研究中心联合研究单位8538,巴黎文理研究大学); Université de Paris, Institut de physique du globe de Paris, CNRS(巴黎大学,巴黎地球物理研究所,法国国家科学研究中心)

作者:Pierre-Yves Raumer, Axel Marmoret, Dorian Cazau, Anatole Gros-Martial, Richard Dreo, Maelle Torterotot, Sara Bazin, Flore Samaran, Jean-Yves Royer

英文摘要:Passive hydroacoustic monitoring often generates large volumes of continuous recordings that are only partially exploited due to the cost of manual annotation. Supervised detection methods perform well but require large labeled datasets, seldom available for rare signals or understudied environments. This work proposes a self-supervised exploration pipeline to address this limitation in low-frequency settings. A Masked AutoEncoder (MAE) is pre-trained on a reconstruction pretext task, then used to extract patch-level representations from spectrograms. Within each spectrogram, adjacent informative patches are aggregated into event-level embeddings, enabling the disentanglement of overlapping events. These embeddings are then clustered at the dataset scale using the dimension reduction algorithm UMAP and the clustering algorithm HDBSCAN to identify hydroacoustic patterns. The pipeline was applied to a multi-year hydroacoustic dataset collected near Mayotte Island, Indian Ocean, containing marine mammal vocalizations, seismo-volcanic signals, and anthropogenic noise. The 317 clusters were manually mapped to 15 hydroacoustic classes or noise in less than one hour. The method was evaluated in two ways. Quantitatively, when used as a classifier, it achieved performance comparable to two existing detectors. Qualitatively, it recovered known seasonal patterns of marine mammal acoustic activity. It also identified patterns of previously unstudied signals, thereby demonstrating its practical value.

7. Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

端到端音频模型中频率表示的结构瓶颈

AI 总结:研究端到端音频模型中频率表示的结构瓶颈,发现现有步长卷积编码器存在两个瓶颈,引入Gabor潜因子分解进行干预,可降低滤波器带宽,保持重建保真度,提高可控性和可解释性。

链接:https://arxiv.org/abs/2607.08545

机构:Yale University(耶鲁大学)

作者:Nicole Cosme-Clifford

英文摘要:End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.