今日论文合集:cs.SD语音11,eess.AS音频处理11篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】AUDRON: A Deep Learning Framework with Fused Acoustic Signatures for Drone Type Recognition
标题:AUDRON:用于无人机类型识别的具有融合声学特征的深度学习框架
链接:https://arxiv.org/abs/2512.20407

作者:Rajdeep Chatterjee,Sudip Chakrabarty,Trishaani Acharjee,Deepanjali Mishra
备注:Presented at the 2025 IEEE 22nd India Council International Conference (INDICON). 6 pages, 3 figures
摘要:无人机(UAV),通常称为无人机,越来越多地用于各种领域,包括物流,农业,监视和国防。虽然这些系统提供了许多好处,但它们的滥用引起了安全和安保问题,因此必须建立有效的检测机制。声学传感为视觉或基于雷达的检测提供了一种低成本和非侵入性的替代方案,因为无人机螺旋桨会产生独特的声音模式。这项研究介绍了AUDRON(基于音频的无人机识别网络),这是一种用于无人机声音检测的混合深度学习框架,采用了Mel频率倒谱系数(MFCC),用卷积神经网络(CNN)处理的短时傅里叶变换(STFT)频谱图,用于时间建模的递归层和基于自动编码器的表示。分类级融合是在分类之前对互补信息进行融合。实验评估表明,AUDRON能够有效区分无人机的声学特征和背景噪声,实现高精度,同时在不同条件下保持通用性。AUDRON在二进制和多类分类中的准确率分别达到98.51%和97.11%。结果突出了将多个特征表示与深度学习相结合以进行可靠的声学无人机检测的优势,表明该框架在视觉或雷达感知可能有限的安全和监视应用中的部署潜力。
摘要:Unmanned aerial vehicles (UAVs), commonly known as drones, are increasingly used across diverse domains, including logistics, agriculture, surveillance, and defense. While these systems provide numerous benefits, their misuse raises safety and security concerns, making effective detection mechanisms essential. Acoustic sensing offers a low-cost and non-intrusive alternative to vision or radar-based detection, as drone propellers generate distinctive sound patterns. This study introduces AUDRON (AUdio-based Drone Recognition Network), a hybrid deep learning framework for drone sound detection, employing a combination of Mel-Frequency Cepstral Coefficients (MFCC), Short-Time Fourier Transform (STFT) spectrograms processed with convolutional neural networks (CNNs), recurrent layers for temporal modeling, and autoencoder-based representations. Feature-level fusion integrates complementary information before classification. Experimental evaluation demonstrates that AUDRON effectively differentiates drone acoustic signatures from background noise, achieving high accuracy while maintaining generalizability across varying conditions. AUDRON achieves 98.51 percent and 97.11 percent accuracy in binary and multiclass classification. The results highlight the advantage of combining multiple feature representations with deep learning for reliable acoustic drone detection, suggesting the framework's potential for deployment in security and surveillance applications where visual or radar sensing may be limited.


【2】EnvSSLAM-FFN: Lightweight Layer-Fused System for ESDD 2026 Challenge
标题:EnvSSLAM-FFN:适用于ESDD 2026挑战赛的轻量级分层融合系统
链接:https://arxiv.org/abs/2512.20369

作者:Xiaoxuan Guo,Hengyan Huang,Jiayi Zhou,Renhe Sun,Jian Liu,Haonan Cheng,Long Ye,Qin Zhang
备注:ESDD 2026 Challenge Technical Report
摘要:生成音频模型的最新进展使高保真环境声音合成成为可能,引起了对音频安全的严重关注。因此,ESDD 2026挑战赛解决了在看不见的发生器(轨道1)和黑盒低资源检测(轨道2)条件下的环境声音深度伪造检测。我们提出了EnvSSLAM-FFN,它集成了一个冻结的SSLAM自监督编码器与一个轻量级的FFN后端。为了在严重数据不平衡的情况下有效地捕获欺骗伪影,我们融合了第4-9层的中间SSLAM表示,并采用了类加权的训练目标。实验结果表明,该系统始终优于官方基线上的两个轨道,实现测试等误差率(EER)的1.20%和1.05%,分别。
摘要:Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under unseen generators (Track 1) and black-box low-resource detection (Track 2) conditions. We propose EnvSSLAM-FFN, which integrates a frozen SSLAM self-supervised encoder with a lightweight FFN back-end. To effectively capture spoofing artifacts under severe data imbalance, we fuse intermediate SSLAM representations from layers 4-9 and adopt a class-weighted training objective. Experimental results show that the proposed system consistently outperforms the official baselines on both tracks, achieving Test Equal Error Rates (EERs) of 1.20% and 1.05%, respectively.


【3】MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model
标题:MMEDIT:通过音频语言模型进行多类型音频编辑的统一框架
链接:https://arxiv.org/abs/2512.20339

作者:Ye Tao,Xuenan Xu,Wen Wu,Shuai Wang,Mengyue Wu,Chao Zhang
备注:Under review
摘要:文本引导的音频编辑旨在修改特定的声学事件,同时严格保留非目标内容。尽管最近取得了一些进展,但现有的办法从根本上说仍然有限。无训练方法通常会受到扩散反演引起的信号退化的影响,而基于训练的方法虽然实现了更高的生成质量,但受到高质量配对数据和任务公式的严重限制,这些数据和任务公式仅覆盖编辑操作的一个狭窄子集。此外,标准架构通常将文本和音频处理解耦,限制了将指令与特定声学上下文对齐的能力。   为了应对这些挑战,我们提出了MMEdit,一个音频语言模型驱动的框架,用于统一的音频编辑。我们系统地扩展了任务定义,以涵盖全面的编辑操作,包括添加,替换,删除,重新排序和属性修改。此外,我们设计了一个可扩展的数据合成管道,以构建具有细粒度事件级注释的大规模配对数据集。为了捕捉复杂的编辑语义,我们将Qwen 2-Audio编码器与基于MMDiT的生成器集成在一起,从而实现精确的跨模态对齐和本地化编辑。   实验结果表明,我们的方法实现了卓越的编辑定位精度,鲁棒的指令遵循,和高保真的非编辑区域。
摘要:Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally limited. Training-free methods often suffer from signal degradation caused by diffusion inversion, while training-based methods, although achieving higher generation quality, are severely constrained by the scarcity of high-quality paired data and task formulations that cover only a narrow subset of editing operations. In addition, standard architectures typically decouple text and audio processing, limiting the ability to align instructions with specific acoustic contexts.   To address these challenges, we propose MMEdit, an audio-language-model-driven framework for unified audio editing. We systematically extend task definitions to cover a comprehensive range of editing operations, including addition, replacement, removal, reordering, and attribute modification. Furthermore, we design a scalable data synthesis pipeline to construct large-scale paired datasets with fine-grained event-level annotations. To capture complex editing semantics, we integrate a Qwen2-Audio encoder with an MMDiT-based generator, enabling precise cross-modal alignment and localized editing.   Experimental results demonstrate that our method achieves superior editing localization accuracy, robust instruction following, and high fidelity in non-edited regions.


【4】SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
标题:SpidR:在没有监督的情况下学习口语模型的快速稳定的语言单位
链接:https://arxiv.org/abs/2512.20308

作者:Maxime Poli,Mahi Luthra,Youssef Benchekroun,Yosuke Higuchi,Martin Gleize,Jiayi Shen,Robin Algayres,Yu-An Chung,Mido Assran,Juan Pino,Emmanuel Dupoux
备注:30 pages, 16 figures
摘要:语言建模和语音表征学习的并行进展提出了直接从语音学习语言而无需文本中间体的前景。这需要直接从语音中提取语义表示。我们的贡献有三方面。首先,我们介绍了SpidR,这是一种自监督语音表示模型,可以有效地学习具有高度可访问语音信息的表示,这使得它特别适合无文本口语建模。它使用掩蔽的预测目标结合自蒸馏和在线聚类在原始波形上进行训练。学生模型的中间层学习预测来自教师中间层的作业。与以前的方法相比,这种学习目标使在线聚类过程稳定,从而产生更高质量的码本。SpidR在下游语言建模基准测试(sWUGGY、sBLIMP、tSC)中的表现优于wav2vec 2.0、HuBERT、WavLM和DinoSR。其次,我们系统地评估跨模型和层的语音单元质量(ABX,PNMI)和语言建模性能之间的相关性,验证这些指标作为可靠的代理。最后,与HuBERT相比,SpidR显著减少了预训练时间,只需要在16个GPU上进行一天的预训练,而不是一周。这种加速是通过预训练方法和高效的代码库实现的,它允许更快的迭代和更容易的实验。我们在https://github.com/facebookresearch/spidr上开源了训练代码和模型检查点。
摘要:The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This requires extracting semantic representations directly from speech. Our contributions are threefold. First, we introduce SpidR, a self-supervised speech representation model that efficiently learns representations with highly accessible phonetic information, which makes it particularly suited for textless spoken language modeling. It is trained on raw waveforms using a masked prediction objective combined with self-distillation and online clustering. The intermediate layers of the student model learn to predict assignments derived from the teacher's intermediate layers. This learning objective stabilizes the online clustering procedure compared to previous approaches, resulting in higher quality codebooks. SpidR outperforms wav2vec 2.0, HuBERT, WavLM, and DinoSR on downstream language modeling benchmarks (sWUGGY, sBLIMP, tSC). Second, we systematically evaluate across models and layers the correlation between speech unit quality (ABX, PNMI) and language modeling performance, validating these metrics as reliable proxies. Finally, SpidR significantly reduces pretraining time compared to HuBERT, requiring only one day of pretraining on 16 GPUs, instead of a week. This speedup is enabled by the pretraining method and an efficient codebase, which allows faster iteration and easier experimentation. We open-source the training code and model checkpoints at https://github.com/facebookresearch/spidr.


【5】Aliasing-Free Neural Audio Synthesis
标题:无混淆神经音频合成
链接:https://arxiv.org/abs/2512.20211

作者:Yicheng Gu,Junan Zhang,Chaoren Wang,Jerry Li,Zhizheng Wu,Lauri Juvela
备注:Submitted to TASLP
摘要:神经声码器和编解码器从声学表示重建波形,这直接影响音频质量。在现有的方法中,基于上采样的时域模型在推理速度和合成质量方面都很优越,达到了最先进的性能。尽管它们在产生感知自然的声音方面取得了成功,但由于设计不充分的模型架构带来的混叠伪影,它们的合成保真度仍然有限。特别是,无约束的非线性激活会产生超过奈奎斯特频率的无限数量的谐波,从而导致“折回”混叠伪影。广泛使用的上采样层ConvTranspose复制镜像的低频部分以填充空的高频区域,从而导致“镜像”混叠伪影。同时,其固有的周期性和镜像直流偏置的组合也带来了“音调伪影”,导致恒定频率振铃。本文旨在从信号处理的角度来解决这些问题。具体来说,我们对激活函数应用过采样和反导数反锯齿以获得其反锯齿形式,并将有问题的ConvTranspose层替换为rescent以避免“音调伪影”并消除锯齿分量。基于我们提出的抗锯齿模块,我们引入了Pupu-Vocoder和Pupu-Codec,并发布了高质量的预训练检查点,以方便音频生成研究。我们建立了一个测试信号基准来说明抗锯齿模块的有效性,并对语音,歌声,音乐和音频进行实验,以验证我们提出的模型。实验结果证实,我们的轻量级Pupu-Vocoder和Pupu-Codec模型可以很容易地超过现有的系统唱歌的声音,音乐和音频,同时实现语音相当的性能。
摘要:Neural vocoders and codecs reconstruct waveforms from acoustic representations, which directly impact the audio quality. Among existing methods, upsampling-based time-domain models are superior in both inference speed and synthesis quality, achieving state-of-the-art performance. Still, despite their success in producing perceptually natural sound, their synthesis fidelity remains limited due to the aliasing artifacts brought by the inadequately designed model architectures. In particular, the unconstrained nonlinear activation generates an infinite number of harmonics that exceed the Nyquist frequency, resulting in ``folded-back'' aliasing artifacts. The widely used upsampling layer, ConvTranspose, copies the mirrored low-frequency parts to fill the empty high-frequency region, resulting in ``mirrored'' aliasing artifacts. Meanwhile, the combination of its inherent periodicity and the mirrored DC bias also brings ``tonal artifact,'' resulting in constant-frequency ringing. This paper aims to solve these issues from a signal processing perspective. Specifically, we apply oversampling and anti-derivative anti-aliasing to the activation function to obtain its anti-aliased form, and replace the problematic ConvTranspose layer with resampling to avoid the ``tonal artifact'' and eliminate aliased components. Based on our proposed anti-aliased modules, we introduce Pupu-Vocoder and Pupu-Codec, and release high-quality pre-trained checkpoints to facilitate audio generation research. We build a test signal benchmark to illustrate the effectiveness of the anti-aliased modules, and conduct experiments on speech, singing voice, music, and audio to validate our proposed models. Experimental results confirm that our lightweight Pupu-Vocoder and Pupu-Codec models can easily outperform existing systems on singing voice, music, and audio, while achieving comparable performance on speech.


【6】Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
标题:光谱还是空间?在具有挑战性的数据条件下利用两者进行说话人提取
链接:https://arxiv.org/abs/2512.20165

作者:Aviad Eisenberg,Sharon Gannot,Shlomo E. Chazan
摘要:本文提出了一种鲁棒的多通道说话人提取算法,旨在处理不准确的参考信息。虽然现有的方法通常只依赖于空间或频谱线索来识别目标说话人,我们的方法集成了两个信息源,以提高鲁棒性。我们的方法的一个关键方面是强调稳定性,即使其中一个功能降级或误导,也能确保可靠的性能。给定一个嘈杂的混合物和两个潜在的不可靠的线索,一个专用的网络被训练来动态地平衡它们的贡献,或者在必要时忽略信息量较少的一个。我们评估系统在具有挑战性的条件下,通过模拟的推断时间误差,使用一个简单的到达方向(DOA)估计和嘈杂的频谱注册过程。实验结果表明,该模型成功地提取所需的扬声器,即使在存在大量的参考不准确。
摘要:This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker, our method integrates both sources of information to enhance robustness. A key aspect of our approach is its emphasis on stability, ensuring reliable performance even when one of the features is degraded or misleading. Given a noisy mixture and two potentially unreliable cues, a dedicated network is trained to dynamically balance their contributions-or disregard the less informative one when necessary. We evaluate the system under challenging conditions by simulating inference-time errors using a simple direction of arrival (DOA) estimator and a noisy spectral enrollment process. Experimental results demonstrate that the proposed model successfully extracts the desired speaker even in the presence of substantial reference inaccuracies.


【7】Fun-Audio-Chat Technical Report
标题:趣味音频聊天技术报告
链接:https://arxiv.org/abs/2512.20156

作者:Qian Chen,Luyao Cheng,Chong Deng,Xiangang Li,Jiaqing Liu,Chao-Hong Tan,Wen Wang,Junhao Xu,Jieping Ye,Qinglin Zhang,Qiquan Zhang,Jingren Zhou
备注:21 pages, https://github.com/FunAudioLLM/Fun-Audio-Chat
摘要:联合语音-文本模型的最新进展显示出无缝语音交互的巨大潜力。然而,现有的模型面临着严峻的挑战:语音标记(25 Hz)和文本标记(~ 3 Hz)之间的时间分辨率不匹配稀释了语义信息,导致高计算成本,并导致灾难性的遗忘文本LLM知识。我们引入Fun-Audio-Chat,一个大型音频语言模型,通过我们以前的工作DrVoice的两个创新来解决这些限制。首先,双分辨率语音表示(DRSR):共享LLM以有效的5 Hz(通过令牌分组)处理音频,而语音精炼头以25 Hz生成高质量令牌,平衡效率(约50% GPU减少)和质量。第二,核心鸡尾酒训练,一个两阶段的微调与中间合并,减轻灾难性的遗忘。然后,我们应用多任务DPO训练来增强鲁棒性,音频理解,跟随和语音同理心。这种多阶段的后期训练使Fun-Audio-Chat能够保留文本LLM知识,同时获得强大的音频理解,推理和生成。与最近需要大规模音频文本预训练的LALM不同,Fun-Audio-Chat利用了预训练模型和广泛的后训练。Fun-Audio-Chat 8B和MoE 30 B-A3 B在语音到文本和语音到语音任务上实现了具有竞争力的性能,在口语QA基准测试中在类似规模的模型中排名第一。他们还在音频理解、语音功能调用、指令遵循和语音同理心方面取得了竞争优势。我们开发了Fun-Audio-Chat-Duplex,这是一种全双工变体,在口语QA和全双工交互方面具有强大的性能。我们开源了Fun-Audio-Chat-8B的训练和推理代码,并提供了一个交互式演示。
摘要:Recent advancements in joint speech-text models show great potential for seamless voice interactions. However, existing models face critical challenges: temporal resolution mismatch between speech tokens (25Hz) and text tokens (~3Hz) dilutes semantic information, incurs high computational costs, and causes catastrophic forgetting of text LLM knowledge. We introduce Fun-Audio-Chat, a Large Audio Language Model addressing these limitations via two innovations from our previous work DrVoice. First, Dual-Resolution Speech Representations (DRSR): the Shared LLM processes audio at efficient 5Hz (via token grouping), while the Speech Refined Head generates high-quality tokens at 25Hz, balancing efficiency (~50% GPU reduction) and quality. Second, Core-Cocktail Training, a two-stage fine-tuning with intermediate merging that mitigates catastrophic forgetting. We then apply Multi-Task DPO Training to enhance robustness, audio understanding, instruction-following and voice empathy. This multi-stage post-training enables Fun-Audio-Chat to retain text LLM knowledge while gaining powerful audio understanding, reasoning, and generation. Unlike recent LALMs requiring large-scale audio-text pre-training, Fun-Audio-Chat leverages pre-trained models and extensive post-training. Fun-Audio-Chat 8B and MoE 30B-A3B achieve competitive performance on Speech-to-Text and Speech-to-Speech tasks, ranking top among similar-scale models on Spoken QA benchmarks. They also achieve competitive to superior performance on Audio Understanding, Speech Function Calling, Instruction-Following and Voice Empathy. We develop Fun-Audio-Chat-Duplex, a full-duplex variant with strong performance on Spoken QA and full-duplex interactions. We open-source Fun-Audio-Chat-8B with training and inference code, and provide an interactive demo.


【8】DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual Segmentation
标题:DDAVS:用于视听分割的解开音频语义和延迟双向对齐
链接:https://arxiv.org/abs/2512.20117

作者:Jingqi Tian,Yiheng Du,Haoji Zhang,Yuji Wang,Isaac Ning Lee,Xulong Bai,Tianrui Zhu,Jingxuan Niu,Yansong Tang
备注:https://trilarflagz.github.io/DDAVS-page/
摘要:视听分割(AVS)的目的是通过联合利用听觉和视觉信息,在像素级定位发声对象。然而,现有的方法往往遭受多源纠缠和视听错位,这导致偏向更响亮或更大的对象,而忽略了更弱,更小,或共同出现的源。为了解决这些挑战,我们提出了DDAVS,一个解开音频语义和延迟双向对齐框架。为了减轻多源纠缠,DDAVS采用可学习的查询提取音频语义和锚定他们在一个结构化的语义空间从音频原型记忆库。这通过对比学习进一步优化,以增强可辨别性和鲁棒性。为了缓解视听错位,DDAVS引入了延迟模态交互的双重交叉注意,提高了多模态对齐的鲁棒性。AVS对象和VPO基准测试的大量实验表明,DDAVS始终优于现有的方法,在单源,多源和多实例场景中表现出强大的性能。这些结果验证了我们的框架在具有挑战性的现实世界的视听分割条件下的有效性和泛化能力。项目页面:https://trilarflagz.github.io/DDAVS-page/
摘要:Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by jointly leveraging auditory and visual information. However, existing methods often suffer from multi-source entanglement and audio-visual misalignment, which lead to biases toward louder or larger objects while overlooking weaker, smaller, or co-occurring sources. To address these challenges, we propose DDAVS, a Disentangled Audio Semantics and Delayed Bidirectional Alignment framework. To mitigate multi-source entanglement, DDAVS employs learnable queries to extract audio semantics and anchor them within a structured semantic space derived from an audio prototype memory bank. This is further optimized through contrastive learning to enhance discriminability and robustness. To alleviate audio-visual misalignment, DDAVS introduces dual cross-attention with delayed modality interaction, improving the robustness of multimodal alignment. Extensive experiments on the AVS-Objects and VPO benchmarks demonstrate that DDAVS consistently outperforms existing approaches, exhibiting strong performance across single-source, multi-source, and multi-instance scenarios. These results validate the effectiveness and generalization ability of our framework under challenging real-world audio-visual segmentation conditions. Project page: https://trilarflagz.github.io/DDAVS-page/


【9】OASI: Objective-Aware Surrogate Initialization for Multi-Objective Bayesian Optimization in TinyML Keyword Spotting
标题:OATI:TinyML关键词发现中多目标Bayesian优化的感知代理程序
链接:https://arxiv.org/abs/2512.19739

作者:Soumen Garai,Suman Samui
备注:Baseline version
摘要:语音助手利用关键字识别(KWS)来实现高效、隐私友好的激活。然而,在超低功耗TinyML设备(通常闪存小于2$ MB)上实现精确的KWS模型需要在精确性与严格的资源约束之间实现微妙的平衡。多目标贝叶斯优化(MOBO)是管理这种权衡的理想选择,但它高度依赖于初始化,特别是在预算黑盒设置下。现有的方法通常回到原始的、特别的采样例程(例如,拉丁超立方体采样(LHS),Sobol序列或随机搜索),既不适应帕累托前沿,也不经过严格的统计比较。为了解决这个问题,我们提出了一种新的初始化策略,利用多目标模拟退火(MOSA)来生成一组高性能和多样化的配置,明确平衡精度和模型大小的种子帕累托集。在TinyML KWS设置中进行评估,OASI优于LHS,Sobol和Random初始化,在多次运行中实现了最高的超卷(0.0627)和最低的代间距离(0.0),计算时间仅略有增加(1934 s vs. 1500 s)。使用Kruskal-Wallis检验($H = 5.40$,$p = 0.144$,$η^2 = 0.0007$)和Dunn事后检验的非参数统计分析证实了OASI的优越一致性,尽管在$α=0.05$阈值方面存在非显著性总体差异。
摘要:Voice assistants utilize Keyword Spotting (KWS) to enable efficient, privacy-friendly activation. However, realizing accurate KWS models on ultra-low-power TinyML devices (often with less than $<2$ MB of flash memory) necessitates a delicate balance between accuracy with strict resource constraints. Multi-objective Bayesian Optimization (MOBO) is an ideal candidate for managing such a trade-off but is highly initialization-dependent, especially under the budgeted black-box setting. Existing methods typically fall back to naive, ad-hoc sampling routines (e.g., Latin Hypercube Sampling (LHS), Sobol sequences, or Random search) that are adapted to neither the Pareto front nor undergo rigorous statistical comparison. To address this, we propose Objective-Aware Surrogate Initialization (OASI), a novel initialization strategy that leverages Multi-Objective Simulated Annealing (MOSA) to generate a seed Pareto set of high-performing and diverse configurations that explicitly balance accuracy and model size. Evaluated in a TinyML KWS setting, OASI outperforms LHS, Sobol, and Random initialization, achieving the highest hypervolume (0.0627) and the lowest generational distance (0.0) across multiple runs, with only a modest increase in computation time (1934 s vs. $\sim$1500 s). A non-parametric statistical analysis using the Kruskal-Wallis test ($H = 5.40$, $p = 0.144$, $η^2 = 0.0007$) and Dunn's post-hoc test confirms OASI's superior consistency despite the non-significant overall difference with respect to the $α=0.05$ threshold.


【10】QuarkAudio Technical Report
标题:QuarkAudio技术报告
链接:https://arxiv.org/abs/2512.20151

作者:Chengwei Liu,Haoyin Yan,Shaofei Xue,Xiaotao Liang,Xiaofu Chen,Bin Gong,Zheng Xue,Gang Song
摘要:许多现有的音频处理和生成模型依赖于特定于任务的架构,导致分散的开发工作和有限的可扩展性。因此,有希望设计一个能够处理多个任务的统一框架,同时提供强大的指令和音频理解以及高质量的音频生成。这需要一个兼容的范例设计,一个强大的骨干,和一个高保真的音频重建模块。为了满足这些要求,本技术报告介绍了QuarkAudio,这是一种基于解码器的自回归(AR)LM生成框架,可统一多个任务。该框架包括一个统一的离散音频标记器,H-Codec,它将自监督学习(SSL)表示纳入标记和重建过程。我们进一步提出了几个改进的H-编解码器,如动态帧速率机制和扩展音频采样率为48 kHz。QuarkAudio通过使用特定于任务的条件信息作为仅解码器LM的条件序列来统一任务,并以AR方式预测离散目标音频令牌。该框架支持广泛的音频处理和生成任务,包括语音恢复(SR)、目标说话人提取(TSE)、语音分离(SS)、语音转换(VC)和语言查询音频源分离(LASS)。此外,我们将下游任务扩展到由自然语言指令引导的通用自由形式音频编辑(包括语音语义编辑和音频事件编辑)。实验结果表明,H-Codec以低帧速率实现了高质量的音频重建,提高了下游音频生成的效率和性能,QuarkAudio在多个任务中提供了与最先进的特定任务或多任务系统竞争或相当的性能。
摘要:Many existing audio processing and generation models rely on task-specific architectures, resulting in fragmented development efforts and limited extensibility. It is therefore promising to design a unified framework capable of handling multiple tasks, while providing robust instruction and audio understanding and high-quality audio generation. This requires a compatible paradigm design, a powerful backbone, and a high-fidelity audio reconstruction module. To meet these requirements, this technical report introduces QuarkAudio, a decoder-only autoregressive (AR) LM-based generative framework that unifies multiple tasks. The framework includes a unified discrete audio tokenizer, H-Codec, which incorporates self-supervised learning (SSL) representations into the tokenization and reconstruction process. We further propose several improvements to H-Codec, such as a dynamic frame-rate mechanism and extending the audio sampling rate to 48 kHz. QuarkAudio unifies tasks by using task-specific conditional information as the conditioning sequence of the decoder-only LM, and predicting discrete target audio tokens in an AR manner. The framework supports a wide range of audio processing and generation tasks, including speech restoration (SR), target speaker extraction (TSE), speech separation (SS), voice conversion (VC), and language-queried audio source separation (LASS). In addition, we extend downstream tasks to universal free-form audio editing guided by natural language instructions (including speech semantic editing and audio event editing). Experimental results show that H-Codec achieves high-quality audio reconstruction with a low frame rate, improving both the efficiency and performance of downstream audio generation, and that QuarkAudio delivers competitive or comparable performance to state-of-the-art task-specific or multi-task systems across multiple tasks.


【11】ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
标题:ASH:音频文本检索的自适应自我改进知识框架
链接:https://arxiv.org/abs/2512.19703

作者:Siyuan Fu,Xuchen Guo,Mingjun Liu,Hongxiang Li,Boyin Tan,Gongxi Zhu,Xianwei Zhuang,Jinghan Ru,Yuxin Xie,Yuguo Yin
摘要:音频文本检索(ATR)的主导范式依赖于基于小批量的对比学习。然而,这个过程本质上受到我们形式化为梯度局部性瓶颈(GLB)的限制,GLB在结构上阻止模型利用批外知识,从而损害细粒度和长尾学习。虽然外部知识增强方法可以缓解GLB,但我们发现了一个关键的未解决的副作用:表示漂移失配(RDM),其中静态知识库逐渐与不断发展的模型不一致,将指导变成噪音。为了解决这一双重挑战,我们提出了自适应自我改进知识(ASK)框架,一个模型不可知的,即插即用的解决方案。ASK通过多粒度知识注入打破了GLB,通过动态知识精化系统地减轻了RDM,并引入了一种新的自适应可靠性加权方案,以确保一致的知识有助于优化。在两个基准数据集上的实验结果证明了我们提出的ASK框架的有效性。
摘要:The dominant paradigm for Audio-Text Retrieval (ATR) relies on mini-batch-based contrastive learning. This process, however, is inherently limited by what we formalize as the Gradient Locality Bottleneck (GLB), which structurally prevents models from leveraging out-of-batch knowledge and thus impairs fine-grained and long-tail learning. While external knowledge-enhanced methods can alleviate the GLB, we identify a critical, unaddressed side effect: the Representation-Drift Mismatch (RDM), where a static knowledge base becomes progressively misaligned with the evolving model, turning guidance into noise. To address this dual challenge, we propose the Adaptive Self-improving Knowledge (ASK) framework, a model-agnostic, plug-and-play solution. ASK breaks the GLB via multi-grained knowledge injection, systematically mitigates RDM through dynamic knowledge refinement, and introduces a novel adaptive reliability weighting scheme to ensure consistent knowledge contributes to optimization. Experimental results on two benchmark datasets with superior, state-of-the-art performance justify the efficacy of our proposed ASK framework.


eess.AS音频处理


【1】LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
标题:LP-CFM:感知不变性感知条件流匹配语音建模
链接:https://arxiv.org/abs/2512.20314

作者:Doyeop Kwak,Youngjoon Jang,Joon Son Chung
摘要:本文的目标是通过结合幅度缩放和时间偏移等感知不变性来提供语音建模的新视角。传统的生成公式通常将每个数据集样本视为目标分布的固定代表。然而,从生成的角度来看,这样的样本只是真实语音分布中许多感知上等价的变体之一。为了解决这个问题,我们提出了线性投影条件流匹配(LP-CFM),它将目标建模为沿着感知等价变量的投影对齐的细长高斯。我们进一步引入矢量校准采样(Vector Calibrated Sampling,简称CFM)来保持采样过程与线投影路径的一致性。在不同模型大小、数据规模和采样步长的神经声码实验中,该方法始终优于传统的最优传输CFM,在低资源和少步长的情况下具有特别强的增益。这些结果突出了LP-CFM和EMPs提供更强大和感知接地生成语音建模的潜力。
摘要:The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional generative formulations often treat each dataset sample as a fixed representative of the target distribution. From a generative standpoint, however, such samples are only one among many perceptually equivalent variants within the true speech distribution. To address this, we propose Linear Projection Conditional Flow Matching (LP-CFM), which models targets as projection-aligned elongated Gaussians along perceptually equivalent variants. We further introduce Vector Calibrated Sampling (VCS) to keep the sampling process aligned with the line-projection path. In neural vocoding experiments across model sizes, data scales, and sampling steps, the proposed approach consistently improves over the conventional optimal transport CFM, with particularly strong gains in low-resource and few-step scenarios. These results highlight the potential of LP-CFM and VCS to provide more robust and perceptually grounded generative modeling of speech.


【2】QuarkAudio Technical Report
标题:QuarkAudio技术报告
链接:https://arxiv.org/abs/2512.20151

作者:Chengwei Liu,Haoyin Yan,Shaofei Xue,Xiaotao Liang,Xiaofu Chen,Bin Gong,Zheng Xue,Gang Song
摘要:许多现有的音频处理和生成模型依赖于特定于任务的架构,导致分散的开发工作和有限的可扩展性。因此,有希望设计一个能够处理多个任务的统一框架,同时提供强大的指令和音频理解以及高质量的音频生成。这需要一个兼容的范例设计,一个强大的骨干,和一个高保真的音频重建模块。为了满足这些要求,本技术报告介绍了QuarkAudio,这是一种基于解码器的自回归(AR)LM生成框架,可统一多个任务。该框架包括一个统一的离散音频标记器,H-Codec,它将自监督学习(SSL)表示纳入标记和重建过程。我们进一步提出了几个改进的H-编解码器,如动态帧速率机制和扩展音频采样率为48 kHz。QuarkAudio通过使用特定于任务的条件信息作为仅解码器LM的条件序列来统一任务,并以AR方式预测离散目标音频令牌。该框架支持广泛的音频处理和生成任务,包括语音恢复(SR)、目标说话人提取(TSE)、语音分离(SS)、语音转换(VC)和语言查询音频源分离(LASS)。此外,我们将下游任务扩展到由自然语言指令引导的通用自由形式音频编辑(包括语音语义编辑和音频事件编辑)。实验结果表明,H-Codec以低帧速率实现了高质量的音频重建,提高了下游音频生成的效率和性能,QuarkAudio在多个任务中提供了与最先进的特定任务或多任务系统竞争或相当的性能。
摘要:Many existing audio processing and generation models rely on task-specific architectures, resulting in fragmented development efforts and limited extensibility. It is therefore promising to design a unified framework capable of handling multiple tasks, while providing robust instruction and audio understanding and high-quality audio generation. This requires a compatible paradigm design, a powerful backbone, and a high-fidelity audio reconstruction module. To meet these requirements, this technical report introduces QuarkAudio, a decoder-only autoregressive (AR) LM-based generative framework that unifies multiple tasks. The framework includes a unified discrete audio tokenizer, H-Codec, which incorporates self-supervised learning (SSL) representations into the tokenization and reconstruction process. We further propose several improvements to H-Codec, such as a dynamic frame-rate mechanism and extending the audio sampling rate to 48 kHz. QuarkAudio unifies tasks by using task-specific conditional information as the conditioning sequence of the decoder-only LM, and predicting discrete target audio tokens in an AR manner. The framework supports a wide range of audio processing and generation tasks, including speech restoration (SR), target speaker extraction (TSE), speech separation (SS), voice conversion (VC), and language-queried audio source separation (LASS). In addition, we extend downstream tasks to universal free-form audio editing guided by natural language instructions (including speech semantic editing and audio event editing). Experimental results show that H-Codec achieves high-quality audio reconstruction with a low frame rate, improving both the efficiency and performance of downstream audio generation, and that QuarkAudio delivers competitive or comparable performance to state-of-the-art task-specific or multi-task systems across multiple tasks.


【3】SpatialNet with Binaural Loss Function for Correcting Binaural Signal Matching Outputs under Head Rotations
标题:具有双耳损失功能的SpatialNet用于在头部旋转下纠正双耳信号匹配输出
链接:https://arxiv.org/abs/2512.20122

作者:Dor Shamay,Boaz Rafaely
摘要:随着虚拟现实耳机、智能眼镜和头戴式耳机等设备的兴起,双耳再现越来越受到关注。用这些系统实现精确的双耳信号是具有挑战性的,因为它们通常采用具有有限空间分辨率的任意麦克风阵列。双耳信号与幅度最小二乘匹配(BSM-MagLS)方法的开发是为了解决早期BSM公式的局限性,提高在高频和头部旋转下的再现性。然而,随着头部旋转的增加,其准确性仍然会降低,导致空间和音色伪影,特别是当虚拟听众的耳朵远离最近的麦克风时。在这项工作中,我们提出将深度学习与BSM-MagLS相结合,以减轻这些退化。采用基于SpatialNet网络的后处理框架,利用其有效处理空间信息的能力,并由来自人类双耳听力理论模型的信号电平损失和感知动机双耳损失指导。该方法的有效性进行了研究,在一个模拟研究与六个麦克风半圆形阵列,显示其能力,以执行强大的头部旋转。这些发现进一步研究了在不同的混响声学环境和头部旋转的听力实验,表明所提出的框架有效地减轻BSM-MagLS退化,并提供了强大的校正在大量的头部旋转。
摘要:Binaural reproduction is gaining increasing attention with the rise of devices such as virtual reality headsets, smart glasses, and head-tracked headphones. Achieving accurate binaural signals with these systems is challenging, as they often employ arbitrary microphone arrays with limited spatial resolution. The Binaural Signals Matching with Magnitude Least-Squares (BSM-MagLS) method was developed to address limitations of earlier BSM formulations, improving reproduction at high frequencies and under head rotation. However, its accuracy still degrades as head rotation increases, resulting in spatial and timbral artifacts, particularly when the virtual listener's ear moves farther from the nearest microphones. In this work, we propose the integration of deep learning with BSM-MagLS to mitigate these degradations. A post-processing framework based on the SpatialNet network is employed, leveraging its ability to process spatial information effectively and guided by both signal-level loss and a perceptually motivated binaural loss derived from a theoretical model of human binaural hearing. The effectiveness of the approach is investigated in a simulation study with a six-microphone semicircular array, showing its ability to perform robustly across head rotations. These findings are further studied in a listening experiment across different reverberant acoustic environments and head rotations, demonstrating that the proposed framework effectively mitigates BSM-MagLS degradations and provides robust correction across substantial head rotations.


【4】ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
标题:ASH:音频文本检索的自适应自我改进知识框架
链接:https://arxiv.org/abs/2512.19703

作者:Siyuan Fu,Xuchen Guo,Mingjun Liu,Hongxiang Li,Boyin Tan,Gongxi Zhu,Xianwei Zhuang,Jinghan Ru,Yuxin Xie,Yuguo Yin
摘要:音频文本检索(ATR)的主导范式依赖于基于小批量的对比学习。然而,这个过程本质上受到我们形式化为梯度局部性瓶颈(GLB)的限制,GLB在结构上阻止模型利用批外知识,从而损害细粒度和长尾学习。虽然外部知识增强方法可以缓解GLB,但我们发现了一个关键的未解决的副作用:表示漂移失配(RDM),其中静态知识库逐渐与不断发展的模型不一致,将指导变成噪音。为了解决这一双重挑战,我们提出了自适应自我改进知识(ASK)框架,一个模型不可知的,即插即用的解决方案。ASK通过多粒度知识注入打破了GLB,通过动态知识精化系统地减轻了RDM,并引入了一种新的自适应可靠性加权方案,以确保一致的知识有助于优化。在两个基准数据集上的实验结果证明了我们提出的ASK框架的有效性。
摘要:The dominant paradigm for Audio-Text Retrieval (ATR) relies on mini-batch-based contrastive learning. This process, however, is inherently limited by what we formalize as the Gradient Locality Bottleneck (GLB), which structurally prevents models from leveraging out-of-batch knowledge and thus impairs fine-grained and long-tail learning. While external knowledge-enhanced methods can alleviate the GLB, we identify a critical, unaddressed side effect: the Representation-Drift Mismatch (RDM), where a static knowledge base becomes progressively misaligned with the evolving model, turning guidance into noise. To address this dual challenge, we propose the Adaptive Self-improving Knowledge (ASK) framework, a model-agnostic, plug-and-play solution. ASK breaks the GLB via multi-grained knowledge injection, systematically mitigates RDM through dynamic knowledge refinement, and introduces a novel adaptive reliability weighting scheme to ensure consistent knowledge contributes to optimization. Experimental results on two benchmark datasets with superior, state-of-the-art performance justify the efficacy of our proposed ASK framework.


【5】EnvSSLAM-FFN: Lightweight Layer-Fused System for ESDD 2026 Challenge
标题:EnvSSLAM-FFN:适用于ESDD 2026挑战赛的轻量级分层融合系统
链接:https://arxiv.org/abs/2512.20369

作者:Xiaoxuan Guo,Hengyan Huang,Jiayi Zhou,Renhe Sun,Jian Liu,Haonan Cheng,Long Ye,Qin Zhang
备注:ESDD 2026 Challenge Technical Report
摘要:生成音频模型的最新进展使高保真环境声音合成成为可能,引起了对音频安全的严重关注。因此,ESDD 2026挑战赛解决了在看不见的发生器(轨道1)和黑盒低资源检测(轨道2)条件下的环境声音深度伪造检测。我们提出了EnvSSLAM-FFN,它集成了一个冻结的SSLAM自监督编码器与一个轻量级的FFN后端。为了在严重数据不平衡的情况下有效地捕获欺骗伪影,我们融合了第4-9层的中间SSLAM表示,并采用了类加权的训练目标。实验结果表明,该系统始终优于官方基线上的两个轨道,实现测试等误差率(EER)的1.20%和1.05%,分别。
摘要:Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under unseen generators (Track 1) and black-box low-resource detection (Track 2) conditions. We propose EnvSSLAM-FFN, which integrates a frozen SSLAM self-supervised encoder with a lightweight FFN back-end. To effectively capture spoofing artifacts under severe data imbalance, we fuse intermediate SSLAM representations from layers 4-9 and adopt a class-weighted training objective. Experimental results show that the proposed system consistently outperforms the official baselines on both tracks, achieving Test Equal Error Rates (EERs) of 1.20% and 1.05%, respectively.


【6】SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
标题:SpidR:在没有监督的情况下学习口语模型的快速稳定的语言单位
链接:https://arxiv.org/abs/2512.20308

作者:Maxime Poli,Mahi Luthra,Youssef Benchekroun,Yosuke Higuchi,Martin Gleize,Jiayi Shen,Robin Algayres,Yu-An Chung,Mido Assran,Juan Pino,Emmanuel Dupoux
备注:30 pages, 16 figures
摘要:语言建模和语音表征学习的并行进展提出了直接从语音学习语言而无需文本中间体的前景。这需要直接从语音中提取语义表示。我们的贡献是三方面的。首先,我们介绍了SpidR,这是一种自监督语音表示模型,可以有效地学习具有高度可访问语音信息的表示,这使得它特别适合无文本口语建模。它使用掩蔽的预测目标结合自蒸馏和在线聚类在原始波形上进行训练。学生模型的中间层学习预测来自教师中间层的作业。与以前的方法相比,这种学习目标使在线聚类过程稳定,从而产生更高质量的码本。SpidR在下游语言建模基准测试(sWUGGY、sBLIMP、tSC)中的表现优于wav2vec 2.0、HuBERT、WavLM和DinoSR。其次,我们系统地评估跨模型和层的语音单元质量(ABX,PNMI)和语言建模性能之间的相关性,验证这些指标作为可靠的代理。最后,与HuBERT相比,SpidR显著减少了预训练时间,只需要在16个GPU上进行一天的预训练,而不是一周。这种加速是通过预训练方法和高效的代码库实现的,它允许更快的迭代和更容易的实验。我们在https://github.com/facebookresearch/spidr上开源了训练代码和模型检查点。
摘要:The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This requires extracting semantic representations directly from speech. Our contributions are threefold. First, we introduce SpidR, a self-supervised speech representation model that efficiently learns representations with highly accessible phonetic information, which makes it particularly suited for textless spoken language modeling. It is trained on raw waveforms using a masked prediction objective combined with self-distillation and online clustering. The intermediate layers of the student model learn to predict assignments derived from the teacher's intermediate layers. This learning objective stabilizes the online clustering procedure compared to previous approaches, resulting in higher quality codebooks. SpidR outperforms wav2vec 2.0, HuBERT, WavLM, and DinoSR on downstream language modeling benchmarks (sWUGGY, sBLIMP, tSC). Second, we systematically evaluate across models and layers the correlation between speech unit quality (ABX, PNMI) and language modeling performance, validating these metrics as reliable proxies. Finally, SpidR significantly reduces pretraining time compared to HuBERT, requiring only one day of pretraining on 16 GPUs, instead of a week. This speedup is enabled by the pretraining method and an efficient codebase, which allows faster iteration and easier experimentation. We open-source the training code and model checkpoints at https://github.com/facebookresearch/spidr.


【7】TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
标题:TAVID:文本驱动的视听交互对话生成
链接:https://arxiv.org/abs/2512.20296

作者:Ji-Hoon Kim,Junseok Ahn,Doyeop Kwak,Joon Son Chung,Shinji Watanabe
备注:Project page: https://mm.kaist.ac.kr/projects/TAVID
摘要:本文的目标是联合合成交互式视频和会话语音从文本和参考图像。随着构建类人会话系统的最终目标,最近的研究已经探索了说话或倾听头生成以及会话语音生成。然而,这些作品通常被孤立地研究,忽视了人类对话的多模态性质,这涉及到紧密耦合的视听交互。在本文中,我们介绍了TAVID,一个统一的框架,生成交互式的面孔和会话语音同步的方式。TAVID通过两个跨模态映射器(即,运动映射器和扬声器映射器),其使得能够在音频和视觉模态之间双向交换补充信息。我们评估我们的系统在四个方面:说话的脸现实主义,听头反应,二元互动流畅性,和语音质量。大量的实验证明了我们的方法在所有这些方面的有效性。
摘要:The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or listening head generation as well as conversational speech generation. However, these works are typically studied in isolation, overlooking the multimodal nature of human conversation, which involves tightly coupled audio-visual interactions. In this paper, we introduce TAVID, a unified framework that generates both interactive faces and conversational speech in a synchronized manner. TAVID integrates face and speech generation pipelines through two cross-modal mappers (i.e., a motion mapper and a speaker mapper), which enable bidirectional exchange of complementary information between the audio and visual modalities. We evaluate our system across four dimensions: talking face realism, listening head responsiveness, dyadic interaction fluency, and speech quality. Extensive experiments demonstrate the effectiveness of our approach across all these aspects.


【8】Aliasing-Free Neural Audio Synthesis
标题:无混淆神经音频合成
链接:https://arxiv.org/abs/2512.20211

作者:Yicheng Gu,Junan Zhang,Chaoren Wang,Jerry Li,Zhizheng Wu,Lauri Juvela
备注:Submitted to TASLP
摘要:神经声码器和编解码器从声学表示重建波形,这直接影响音频质量。在现有的方法中,基于上采样的时域模型在推理速度和合成质量方面都很优越,达到了最先进的性能。尽管它们在产生感知自然的声音方面取得了成功,但由于设计不充分的模型架构带来的混叠伪影,它们的合成保真度仍然有限。特别是,无约束的非线性激活会产生超过奈奎斯特频率的无限数量的谐波,从而导致“折回”混叠伪影。广泛使用的上采样层ConvTranspose复制镜像的低频部分以填充空的高频区域,从而导致“镜像”混叠伪影。同时,其固有的周期性和镜像直流偏置的组合也带来了“音调伪影”,导致恒定频率振铃。本文旨在从信号处理的角度来解决这些问题。具体来说,我们对激活函数应用过采样和反导数反锯齿以获得其反锯齿形式,并将有问题的ConvTranspose层替换为rescent以避免“音调伪影”并消除锯齿分量。基于我们提出的抗锯齿模块,我们引入了Pupu-Vocoder和Pupu-Codec,并发布了高质量的预训练检查点,以方便音频生成研究。我们建立了一个测试信号基准来说明抗锯齿模块的有效性,并对语音,歌声,音乐和音频进行实验,以验证我们提出的模型。实验结果证实,我们的轻量级Pupu-Vocoder和Pupu-Codec模型可以很容易地超过现有的系统唱歌的声音,音乐和音频,同时实现语音相当的性能。
摘要:Neural vocoders and codecs reconstruct waveforms from acoustic representations, which directly impact the audio quality. Among existing methods, upsampling-based time-domain models are superior in both inference speed and synthesis quality, achieving state-of-the-art performance. Still, despite their success in producing perceptually natural sound, their synthesis fidelity remains limited due to the aliasing artifacts brought by the inadequately designed model architectures. In particular, the unconstrained nonlinear activation generates an infinite number of harmonics that exceed the Nyquist frequency, resulting in ``folded-back'' aliasing artifacts. The widely used upsampling layer, ConvTranspose, copies the mirrored low-frequency parts to fill the empty high-frequency region, resulting in ``mirrored'' aliasing artifacts. Meanwhile, the combination of its inherent periodicity and the mirrored DC bias also brings ``tonal artifact,'' resulting in constant-frequency ringing. This paper aims to solve these issues from a signal processing perspective. Specifically, we apply oversampling and anti-derivative anti-aliasing to the activation function to obtain its anti-aliased form, and replace the problematic ConvTranspose layer with resampling to avoid the ``tonal artifact'' and eliminate aliased components. Based on our proposed anti-aliased modules, we introduce Pupu-Vocoder and Pupu-Codec, and release high-quality pre-trained checkpoints to facilitate audio generation research. We build a test signal benchmark to illustrate the effectiveness of the anti-aliased modules, and conduct experiments on speech, singing voice, music, and audio to validate our proposed models. Experimental results confirm that our lightweight Pupu-Vocoder and Pupu-Codec models can easily outperform existing systems on singing voice, music, and audio, while achieving comparable performance on speech.


【9】Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
标题:光谱还是空间?在具有挑战性的数据条件下利用两者进行说话人提取
链接:https://arxiv.org/abs/2512.20165

作者:Aviad Eisenberg,Sharon Gannot,Shlomo E. Chazan
摘要:本文提出了一种鲁棒的多通道说话人提取算法,旨在处理不准确的参考信息。虽然现有的方法通常只依赖于空间或频谱线索来识别目标说话人,我们的方法集成了两个信息源,以提高鲁棒性。我们的方法的一个关键方面是强调稳定性,即使其中一个功能降级或误导,也能确保可靠的性能。给定一个嘈杂的混合物和两个潜在的不可靠的线索,一个专用的网络被训练来动态地平衡它们的贡献,或者在必要时忽略信息量较少的一个。我们评估系统在具有挑战性的条件下,通过模拟的推断时间误差,使用一个简单的到达方向(DOA)估计和嘈杂的频谱注册过程。实验结果表明,该模型成功地提取所需的扬声器,即使在存在大量的参考不准确。
摘要:This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker, our method integrates both sources of information to enhance robustness. A key aspect of our approach is its emphasis on stability, ensuring reliable performance even when one of the features is degraded or misleading. Given a noisy mixture and two potentially unreliable cues, a dedicated network is trained to dynamically balance their contributions-or disregard the less informative one when necessary. We evaluate the system under challenging conditions by simulating inference-time errors using a simple direction of arrival (DOA) estimator and a noisy spectral enrollment process. Experimental results demonstrate that the proposed model successfully extracts the desired speaker even in the presence of substantial reference inaccuracies.


【10】Fun-Audio-Chat Technical Report
标题:趣味音频聊天技术报告
链接:https://arxiv.org/abs/2512.20156

作者:Qian Chen,Luyao Cheng,Chong Deng,Xiangang Li,Jiaqing Liu,Chao-Hong Tan,Wen Wang,Junhao Xu,Jieping Ye,Qinglin Zhang,Qiquan Zhang,Jingren Zhou
备注:21 pages, https://github.com/FunAudioLLM/Fun-Audio-Chat
摘要:联合语音-文本模型的最新进展显示出无缝语音交互的巨大潜力。然而,现有的模型面临着严峻的挑战:语音标记(25 Hz)和文本标记(~ 3 Hz)之间的时间分辨率不匹配稀释了语义信息,导致高计算成本,并导致灾难性的遗忘文本LLM知识。我们引入Fun-Audio-Chat,一个大型音频语言模型,通过我们以前的工作DrVoice的两个创新来解决这些限制。首先,双分辨率语音表示(DRSR):共享LLM以有效的5 Hz(通过令牌分组)处理音频,而语音精炼头以25 Hz生成高质量令牌,平衡效率(约50% GPU减少)和质量。第二,核心鸡尾酒训练,一个两阶段的微调与中间合并,减轻灾难性的遗忘。然后,我们应用多任务DPO训练来增强鲁棒性,音频理解,跟随和语音同理心。这种多阶段的后期训练使Fun-Audio-Chat能够保留文本LLM知识,同时获得强大的音频理解,推理和生成。与最近需要大规模音频文本预训练的LALM不同,Fun-Audio-Chat利用了预训练模型和广泛的后训练。Fun-Audio-Chat 8B和MoE 30 B-A3 B在语音到文本和语音到语音任务上实现了具有竞争力的性能,在口语QA基准测试中在类似规模的模型中排名第一。他们还在音频理解、语音功能调用、指令遵循和语音同理心方面取得了竞争优势。我们开发了Fun-Audio-Chat-Duplex,这是一种全双工变体,在口语QA和全双工交互方面具有强大的性能。我们开源了Fun-Audio-Chat-8B的训练和推理代码,并提供了一个交互式演示。
摘要:Recent advancements in joint speech-text models show great potential for seamless voice interactions. However, existing models face critical challenges: temporal resolution mismatch between speech tokens (25Hz) and text tokens (~3Hz) dilutes semantic information, incurs high computational costs, and causes catastrophic forgetting of text LLM knowledge. We introduce Fun-Audio-Chat, a Large Audio Language Model addressing these limitations via two innovations from our previous work DrVoice. First, Dual-Resolution Speech Representations (DRSR): the Shared LLM processes audio at efficient 5Hz (via token grouping), while the Speech Refined Head generates high-quality tokens at 25Hz, balancing efficiency (~50% GPU reduction) and quality. Second, Core-Cocktail Training, a two-stage fine-tuning with intermediate merging that mitigates catastrophic forgetting. We then apply Multi-Task DPO Training to enhance robustness, audio understanding, instruction-following and voice empathy. This multi-stage post-training enables Fun-Audio-Chat to retain text LLM knowledge while gaining powerful audio understanding, reasoning, and generation. Unlike recent LALMs requiring large-scale audio-text pre-training, Fun-Audio-Chat leverages pre-trained models and extensive post-training. Fun-Audio-Chat 8B and MoE 30B-A3B achieve competitive performance on Speech-to-Text and Speech-to-Speech tasks, ranking top among similar-scale models on Spoken QA benchmarks. They also achieve competitive to superior performance on Audio Understanding, Speech Function Calling, Instruction-Following and Voice Empathy. We develop Fun-Audio-Chat-Duplex, a full-duplex variant with strong performance on Spoken QA and full-duplex interactions. We open-source Fun-Audio-Chat-8B with training and inference code, and provide an interactive demo.


【11】DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual Segmentation
标题:DDAVS:用于视听分割的解开音频语义和延迟双向对齐
链接:https://arxiv.org/abs/2512.20117

作者:Jingqi Tian,Yiheng Du,Haoji Zhang,Yuji Wang,Isaac Ning Lee,Xulong Bai,Tianrui Zhu,Jingxuan Niu,Yansong Tang
备注:https://trilarflagz.github.io/DDAVS-page/
摘要:视听分割(AVS)的目的是通过联合利用听觉和视觉信息,在像素级定位发声对象。然而,现有的方法往往遭受多源纠缠和视听错位,这导致偏向更响亮或更大的对象,而忽略了更弱,更小,或共同出现的源。为了解决这些挑战,我们提出了DDAVS,一个解开音频语义和延迟双向对齐框架。为了减轻多源纠缠,DDAVS采用可学习的查询提取音频语义和锚定他们在一个结构化的语义空间从音频原型记忆库。这通过对比学习进一步优化,以增强可辨别性和鲁棒性。为了缓解视听错位,DDAVS引入了延迟模态交互的双重交叉注意,提高了多模态对齐的鲁棒性。AVS对象和VPO基准测试的大量实验表明,DDAVS始终优于现有的方法,在单源,多源和多实例场景中表现出强大的性能。这些结果验证了我们的框架在具有挑战性的现实世界的视听分割条件下的有效性和泛化能力。项目页面:https://trilarflagz.github.io/DDAVS-page/
摘要:Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by jointly leveraging auditory and visual information. However, existing methods often suffer from multi-source entanglement and audio-visual misalignment, which lead to biases toward louder or larger objects while overlooking weaker, smaller, or co-occurring sources. To address these challenges, we propose DDAVS, a Disentangled Audio Semantics and Delayed Bidirectional Alignment framework. To mitigate multi-source entanglement, DDAVS employs learnable queries to extract audio semantics and anchor them within a structured semantic space derived from an audio prototype memory bank. This is further optimized through contrastive learning to enhance discriminability and robustness. To alleviate audio-visual misalignment, DDAVS introduces dual cross-attention with delayed modality interaction, improving the robustness of multimodal alignment. Extensive experiments on the AVS-Objects and VPO benchmarks demonstrate that DDAVS consistently outperforms existing approaches, exhibiting strong performance across single-source, multi-source, and multi-instance scenarios. These results validate the effectiveness and generalization ability of our framework under challenging real-world audio-visual segmentation conditions. Project page: https://trilarflagz.github.io/DDAVS-page/


机器翻译由腾讯交互翻译提供,仅供参考