今日论文合集:cs.SD语音14篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】FOCAL: A Novel Benchmarking Technique for Multi-modal Agents
标题:FOCAL:一种新颖的多模式代理基准技术
链接:https://arxiv.org/abs/2601.07367

作者:Aditya Choudhary,Anupam Purwar
备注:We present a framework for evaluation of Multi-modal Agents consisting of Voice-to-voice model components viz. Text to Speech (TTS), Retrieval Augmented Generation (RAG) and Speech-to-text (STT)
摘要:随着推理能力、使用MCP服务器和音频语言模型(ALM)的工具调用的最新进展,多模态代理(具有语音和文本支持)的开发和集成已经走到了行业的最前沿。语音代理的级联管道仍然在行业中发挥着核心作用,因为它们具有由LLM促进的卓越推理能力。但是,级联流水线通常会通过流水线传播错误。我们提出了一个框架,FOCAL基准端到端的推理,组件明智的错误传播和错误分析的自动化以及人工辅助测试的多模态代理(语音到语音+文本输入)。我们还分享了两个新的指标,即推理和语义分数,以评估代理在语音模式下进行有意义的对话的功效。
摘要:With the recent advancements in reasoning capa- bilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voice and text support) has come to the industry forefront. Cascading pipelines for voice agents still play a central role in the industry owing to their superior reasoning capabilities facilitated by LLMs. Although, cascading pipelines often present error propagation through the pipeline. We propose a framework, FOCAL to benchmark end-to-end reasoning, component-wise error propagation and error analysis for automated as well as human-assisted testing of multi-modal agents (voice to voice + text input). We also share two novel metrics viz. Reasoning and Semantic scores to evaluate efficacy of the agent in having meaningful conversations in voice mode.


【2】SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models
标题:SEE:用于量化大型音频语言模型中的噪音干扰的信号嵌入能量
链接:https://arxiv.org/abs/2601.07331

作者:Yuanhe Zhang,Jiayu Tian,Yibo Zhang,Shilinlu Yan,Liang Lin,Zhenhong Zhou,Li Sun,Sen Su
摘要:大型音频语言模型(LALM)在实时场景中得到了广泛的应用,例如车载助手和在线会议理解。在实践中,音频输入经常被设备和环境噪声破坏,导致性能下降。然而,现有的LALM对噪声的研究缺乏定量分析,主要依赖于直觉和经验观察,从而无法理解实际的鲁棒性。为了解决这个问题,我们引入了信号嵌入能量(SEE),这是一种用于量化噪声强度对LALM输入的影响的方法,可以区分现实部署中的LALM鲁棒性。SEE引入了一种基于结构化激活子空间的视角,该结构化激活子空间来自模型的内部表示,比原始音频特征更准确地捕捉其对噪声的感知。在整个实验中,SEE表现出与LALM性能的强相关性,达到0.98的相关性。令人惊讶的是,传统的音频去噪方法对LALM只有轻微的效果,在某些情况下,甚至会增加SEE并损害性能。这表明以语音为中心的去噪目标与现代LALM的噪声灵敏度之间存在不匹配。因此,我们提出了一种源自SEE的缓解策略来对LALM输入进行降噪,其性能优于现有的降噪方法。本文介绍了一种新的度量噪声量化LALM,在现实世界的部署提供指导的鲁棒性改进。
摘要:Large Audio Language Models (LALMs) have been widely applied in real-time scenarios, such as in-car assistants and online meeting comprehension. In practice, audio inputs are often corrupted by device and environmental noise, leading to performance degradation. However, existing LALM studies on noise lack quantitative analysis and rely mainly on intuition and empirical observation, thus failing to understand practical robustness. To address this issue, we introduce Signal Embedding Energy (SEE), a method for quantifying the impact of noise intensity on LALM inputs, enabling the differentiation of LALM robustness in real-world deployments. SEE introduces a perspective based on structured activation subspaces derived from the model's internal representations, which more accurately captures its perception of noise than raw audio features. Across experiments, SEE exhibits a strong correlation with LALM performance, achieving a correlation of 0.98. Surprisingly, traditional audio denoising methods are only marginally effective for LALMs, and, in some cases, even increase SEE and impair performance. This suggests a mismatch between speech-centric denoising objectives and the noise sensitivity of modern LALMs. Therefore, we propose a mitigation strategy derived from SEE to denoise LALM inputs, outperforming existing denoising methods. This paper introduces a novel metric for noise quantification in LALMs, providing guidance for robustness improvements in real-world deployments.


【3】ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge Evaluation Plan
标题:ESDD 2:环境意识语音和声音Deepfake检测挑战评估计划
链接:https://arxiv.org/abs/2601.07303

作者:Xueping Zhang,Han Yin,Yang Xiao,Lin Zhang,Ting Dang
摘要:在真实世界环境中记录的音频通常包含前景语音和背景环境声音的混合。随着文本到语音、语音转换和其他生成模型的快速发展,现在可以独立地修改任何一个组件。这种组件级的操作更难检测,因为剩余的未更改组件可能会误导为整个deepfake音频设计的系统,并且它们通常对人类听众来说听起来更自然。为了解决这个问题,我们提出了CompSpoofV2数据集和分离增强的联合学习框架。CompSpoofV2是一个为组件级音频反欺骗而设计的大规模策划数据集,包含超过250k音频样本,总持续时间约为283小时。基于CompSpoofV2和分离增强的联合学习框架,我们推出了环境感知语音和声音Deepfake检测挑战赛(ESDD2),专注于组件级欺骗,其中语音和环境声音都可以被操纵或合成,从而创建更具挑战性和现实感的检测场景。该挑战赛将与IEEE 2026年多媒体和博览会国际会议(ICME 2026)一起举行。
摘要:Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be modified independently. Such component-level manipulations are harder to detect, as the remaining unaltered component can mislead the systems designed for whole deepfake audio, and they often sound more natural to human listeners. To address this gap, we have proposed CompSpoofV2 dataset and a separation-enhanced joint learning framework. CompSpoofV2 is a large-scale curated dataset designed for component-level audio anti-spoofing, which contains over 250k audio samples, with a total duration of approximately 283 hours. Based on the CompSpoofV2 and the separation-enhanced joint learning framework, we launch the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2), focusing on component-level spoofing, where both speech and environmental sounds may be manipulated or synthesized, creating a more challenging and realistic detection scenario. The challenge will be held in conjunction with the IEEE International Conference on Multimedia and Expo 2026 (ICME 2026).


【4】Directional Selective Fixed-Filter Active Noise Control Based on a Convolutional Neural Network in Reverberant Environments
标题:回响环境下基于卷积神经网络的定向选择性固定过滤器主动噪音控制
链接:https://arxiv.org/abs/2601.06981

作者:Boxiang Wang,Zhengding Luo,Haowen Li,Dongyuan Shi,Junwei Ji,Ziyi Yang,Woon-Seng Gan
摘要:选择性固定滤波器有源噪声控制(SFANC)是一种能够抑制具有变化频率特性的噪声的新方法。与传统的自适应算法相比,它具有更快的响应速度和更高的计算效率。然而,空间因素,特别是噪声源位置的影响,往往被忽视。现有的一些研究已经探索了噪声源的到达方向(DoA)对ANC性能的影响,但它们大多限于自由场条件,并且没有考虑更复杂的室内混响环境。为了解决这一问题,本文提出了一种基于学习的方向SFANC方法,该方法结合了混响环境中噪声源的DoA。在该框架中,多个参考信号由卷积神经网络(CNN)处理以估计噪声源的方位角和仰角,以及识别用于有效噪声消除的最适当的控制滤波器。与传统的自适应算法相比,所提出的方法实现了更短的响应时间,即使在混响的存在下,优越的降噪。
摘要:Selective fixed-filter active noise control (SFANC) is a novel approach capable of mitigating noise with varying frequency characteristics. It offers faster response and greater computational efficiency compared to traditional adaptive algorithms. However, spatial factors, particularly the influence of the noise source location, are often overlooked. Some existing studies have explored the impact of the direction-of-arrival (DoA) of the noise source on ANC performance, but they are mostly limited to free-field conditions and do not consider the more complex indoor reverberant environments. To address this gap, this paper proposes a learning-based directional SFANC method that incorporates the DoA of the noise source in reverberant environments. In this framework, multiple reference signals are processed by a convolutional neural network (CNN) to estimate the azimuth and elevation angles of the noise source, as well as to identify the most appropriate control filter for effective noise cancellation. Compared to traditional adaptive algorithms, the proposed approach achieves superior noise reduction with shorter response times, even in the presence of reverberations.


【5】MoEScore: Mixture-of-Experts-Based Text-Audio Relevance Score Prediction for Text-to-Audio System Evaluation
标题:MoEscore:用于文本到音频系统评估的基于专家混合的文本音频相关性分数预测
链接:https://arxiv.org/abs/2601.06829

作者:Bochao Sun,Yang Xiao,Han Yin
摘要:生成模型的最新进展使得现代文本到音频(TTA)系统能够合成具有高感知质量的音频。然而,TTA系统往往很难保持与输入文本的语义一致性,导致声音事件,时间结构或上下文关系的不匹配。在TTA中评估语义保真度仍然是一个重大的挑战。传统的方法主要依赖于主观的人类听力测试,这是耗时的。为了解决这个问题,我们提出了一个基于具有顺序交叉注意力(SeqCoAttn)的混合专家(MoE)架构的客观评估器。我们的模型在XACLE挑战赛中获得了第一名,在测试数据集上的SRCC为0.6402(比挑战基线提高了30.6%)。代码可从以下网址获得:https://github.com/S-Orion/MOESCORE。
摘要:Recent advances in generative models have enabled modern Text-to-Audio (TTA) systems to synthesize audio with high perceptual quality. However, TTA systems often struggle to maintain semantic consistency with the input text, leading to mismatches in sound events, temporal tructures, or contextual relationships. Evaluating semantic fidelity in TTA remains a significant challenge. Traditional methods primarily rely on subjective human listening tests, which is time-consuming. To solve this, we propose an objective evaluator based on a Mixture of Experts (MoE) architecture with Sequential Cross-Attention (SeqCoAttn). Our model achieves the first rank in the XACLE Challenge, with an SRCC of 0.6402 (an improvement of 30.6% over the challenge baseline) on the test dataset. Code is available at: https://github.com/S-Orion/MOESCORE.


【6】Representing Sounds as Neural Amplitude Fields: A Benchmark of Coordinate-MLPs and A Fourier Kolmogorov-Arnold Framework
标题:将声音表示为神经振幅场:坐标MLP基准和傅立叶Kolmogorov-Arnold框架
链接:https://arxiv.org/abs/2601.06406

作者:Linfei Li,Lin Zhang,Zhong Wang,Fengyi Zhang,Zelin Li,Ying Shen
备注:Accepted by AAAI 2025. Code: https://github.com/lif314/Fourier-ASR
摘要:虽然基于坐标MLP的隐式神经表示在表示辐射场、3D形状和图像方面表现出色,但它们在音频信号中的应用仍然有待探索。为了填补这一空白,我们调查了现有的隐式神经表征,从中我们提取了3种类型的位置编码和16种常用的激活函数。通过组合设计,我们建立了第一个基准的坐标MLP在音频信号表示。我们的基准测试表明,坐标MLP需要复杂的超参数调整和频率相关的初始化,限制了它们的鲁棒性。为了解决这些问题,我们提出了傅立叶ASR,一个新的框架的基础上傅立叶级数定理和Kolmogorov-Arnold表示定理。Fourier-ASR引入了傅立叶-柯尔莫哥洛夫-阿诺德网络(Fourier-KAN),它利用周期性和强非线性来表示音频信号,无需额外的位置编码。此外,频率自适应学习策略(FaLS),提出了提高傅立叶-KAN的收敛,通过捕获高频分量和防止过拟合的低频信号。在自然语音和音乐数据集上进行的大量实验表明:(1)Coordinate-MLP中设计良好的位置编码和激活函数可以有效地提高音频表示质量;(2)Fourier-ASR可以鲁棒地表示复杂的音频信号,而无需大量的超参数调整。展望未来,隐式音频表示的连续性和无限分辨率使我们的研究非常有前途的任务,如音频压缩,合成和生成。源代码将公开发布,以确保可重复性。该代码可在https://github.com/lif314/Fourier-ASR上获得。
摘要:Although Coordinate-MLP-based implicit neural representations have excelled in representing radiance fields, 3D shapes, and images, their application to audio signals remains underexplored. To fill this gap, we investigate existing implicit neural representations, from which we extract 3 types of positional encoding and 16 commonly used activation functions. Through combinatorial design, we establish the first benchmark for Coordinate-MLPs in audio signal representations. Our benchmark reveals that Coordinate-MLPs require complex hyperparameter tuning and frequency-dependent initialization, limiting their robustness. To address these issues, we propose Fourier-ASR, a novel framework based on the Fourier series theorem and the Kolmogorov-Arnold representation theorem. Fourier-ASR introduces Fourier Kolmogorov-Arnold Networks (Fourier-KAN), which leverage periodicity and strong nonlinearity to represent audio signals, eliminating the need for additional positional encoding. Furthermore, a Frequency-adaptive Learning Strategy (FaLS) is proposed to enhance the convergence of Fourier-KAN by capturing high-frequency components and preventing overfitting of low-frequency signals. Extensive experiments conducted on natural speech and music datasets reveal that: (1) well-designed positional encoding and activation functions in Coordinate-MLPs can effectively improve audio representation quality; and (2) Fourier-ASR can robustly represent complex audio signals without extensive hyperparameter tuning. Looking ahead, the continuity and infinite resolution of implicit audio representations make our research highly promising for tasks such as audio compression, synthesis, and generation. The source code will be released publicly to ensure reproducibility. The code is available at https://github.com/lif314/Fourier-ASR.


【7】An Intelligent AI glasses System with Multi-Agent Architecture for Real-Time Voice Processing and Task Execution
标题:具有多代理架构的智能人工智能眼镜系统,用于实时语音处理和任务执行
链接:https://arxiv.org/abs/2601.06235

作者:Sheng-Kai Chen,Jyh-Horng Wu,Ching-Yao Lin,Yen-Ting Lin
备注:Published in NCS 2025 (Paper No. N0180)
摘要:本文提出了一种集成实时语音处理、人工智能(AI)代理和跨网络流媒体功能的AI眼镜系统。该系统采用双代理架构,其中代理01处理自动语音识别(ASR),代理02通过本地大型语言模型(LLM),模型上下文协议(MCP)工具和检索增强生成(RAG)管理AI处理。该系统支持实时RTSP流,用于语音和视频数据传输,眼动跟踪数据收集,以及通过RabbitMQ消息传递执行远程任务。实现演示了成功的语音命令处理与多语言支持和跨平台的任务执行能力。
摘要:This paper presents an AI glasses system that integrates real-time voice processing, artificial intelligence(AI) agents, and cross-network streaming capabilities. The system employs dual-agent architecture where Agent 01 handles Automatic Speech Recognition (ASR) and Agent 02 manages AI processing through local Large Language Models (LLMs), Model Context Protocol (MCP) tools, and Retrieval-Augmented Generation (RAG). The system supports real-time RTSP streaming for voice and video data transmission, eye tracking data collection, and remote task execution through RabbitMQ messaging. Implementation demonstrates successful voice command processing with multilingual support and cross-platform task execution capabilities.


【8】AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
标题:Deliveros:通过自生成的免指令调优将LLM扩展到语音
链接:https://arxiv.org/abs/2601.06086

作者:Yiwen Shao,Wei Liu,Jiahong Li,Tianzi Wang,Kun Wei,Meng Yu,Dong Yu
备注:Technical Report
摘要:将大型语言模型(LLM)扩展到语音领域最近受到了极大的关注。一种典型的方法是通过投影模块将预训练的LLM与音频编码器连接起来,并在大规模的特定于任务的预处理调整数据集上训练生成的模型。然而,为特定需求管理这种预防调整数据是耗时的,并且以这种方式训练的模型通常不能很好地推广到看不见的任务。在这项工作中,我们首先制定了语音LLM的最强泛化是在使用自生成无指令调谐(SIFT)进行训练时实现的,其中监督信号由冻结的LLM使用语音的文本表示作为输入生成。我们提出的SIFT范式消除了收集特定任务的问题-答案对的需要,并产生了理论上最好的概括看不见的任务。在此范例的基础上,我们引入了AZeroS(Auden Zero-instruction-tuned Speech-LLM),它是在来自公开语料库的语音文本对上训练的,包括大约25,000小时的带有ASR成绩单的语音和3,000小时的带有非语言标签的语音。该模型基于Qwen2.5- 7 B-Instruct构建,仅更新两个轻量级投影模块(每个模块2380万个参数),同时保持LLM和音频编码器冻结。尽管训练成本最低,数据规模适中,但AZeroS在语义和语言基准测试方面都达到了最先进的性能,包括语音基准测试、AIR-Bench Foundation(语音)和AIR-Bench Chat(语音)。
摘要:Extending large language models (LLMs) to the speech domain has recently gained significant attention. A typical approach connects a pretrained LLM with an audio encoder through a projection module and trains the resulting model on large-scale, task-specific instruction-tuning datasets. However, curating such instruction-tuning data for specific requirements is time-consuming, and models trained in this manner often generalize poorly to unseen tasks. In this work, we first formulate that the strongest generalization of a speech-LLM is achieved when it is trained with Self-Generated Instruction-Free Tuning (SIFT), in which supervision signals are generated by a frozen LLM using textual representations of speech as input. Our proposed SIFT paradigm eliminates the need for collecting task-specific question-answer pairs and yields the theoretically best generalization to unseen tasks. Building upon this paradigm, we introduce AZeroS (Auden Zero-instruction-tuned Speech-LLM), which is trained on speech-text pairs derived from publicly available corpora, including approximately 25,000 hours of speech with ASR transcripts and 3,000 hours of speech with paralinguistic labels. Built upon Qwen2.5-7B-Instruct, the model updates only two lightweight projection modules (23.8 million parameters each), while keeping both the LLM and audio encoders frozen. Despite the minimal training cost and modest data scale, AZeroS achieves state-of-the-art performance on both semantic and paralinguistic benchmarks, including VoiceBench, AIR-Bench Foundation (Speech), and AIR-Bench Chat (Speech).


【9】The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
标题:ICASP 2026自动歌曲美学评估挑战赛
链接:https://arxiv.org/abs/2601.07237

作者:Guobin Ma,Yuxuan Xia,Jixun Yao,Huixin Xue,Hexin Liu,Shuai Wang,Hao Liu,Lei Xie
备注:Official summary paper for the ICASSP 2026 ASAE Challenge
摘要:本文总结了ICASSP 2026自动歌曲美学评估(ASAE)挑战赛,该挑战赛专注于预测AI生成歌曲的主观美学评分。该挑战包括两个轨道:轨道1的目标是预测整体音乐性得分,而轨道2侧重于预测五个细粒度的美学得分。这一挑战引起了研究界的强烈兴趣,并收到了学术界和工业界的许多意见书。表现最好的系统大大超过了官方基线,表明在将客观指标与人类审美偏好相结合方面取得了实质性进展。这些成果为现代音乐生成系统建立了标准化的基准,并推进了与人类一致的评估方法。
摘要:This paper summarizes the ICASSP 2026 Automatic Song Aesthetics Evaluation (ASAE) Challenge, which focuses on predicting the subjective aesthetic scores of AI-generated songs. The challenge consists of two tracks: Track 1 targets the prediction of the overall musicality score, while Track 2 focuses on predicting five fine-grained aesthetic scores. The challenge attracted strong interest from the research community and received numerous submissions from both academia and industry. Top-performing systems significantly surpassed the official baseline, demonstrating substantial progress in aligning objective metrics with human aesthetic preferences. The outcomes establish a standardized benchmark and advance human-aligned evaluation methodologies for modern music generation systems.


【10】Dereverberation Filter by Deconvolution with Frequency Bin Specific Faded Impulse Response
标题:利用特定频率段的衰落冲激响应去卷积进行去混滤
链接:https://arxiv.org/abs/2601.06662

作者:Stefan Ciba
备注:8 pages, 3 figures, github repository with code and audio
摘要:这项工作介绍了一个强大的单通道逆滤波器去混响的非理想的录音,在真实的音频验证。所开发的方法侧重于离散脉冲响应的计算和修改,以便从已知的数字单通道记录设置和房间特性(如早期反射和混响)中过滤特性。目标是更干燥和更清晰的信号重建,理想情况下是直接路径信号。从倒谱域计算时域脉冲响应,并通过频谱中的频率单元特定指数衰减来衰减。衰减率是通过使用记录的输出和测试信号之间的混响时间比的盲估计获得的每个频率箱。修改的脉冲响应通过去卷积对记录的音频信号进行滤波。盲估计是众所周知的,并且由于其对噪声和非理想性的鲁棒性而突出。直接路径信号的估计是许多应用的关键。
摘要:This work introduces a robust single-channel inverse filter for dereverberation of non-ideal recordings, validated on real audio. The developed method focuses on the calculation and modification of a discrete impulse response in order to filter the characteristics from a known digital single channel recording setup and room characteristics such as early reflections and reverberations. The aim is a dryer and clearer signal reconstruction, which ideally would be the direct-path signal. The time domain impulse response is calculated from the cepstral domain and faded by means of frequency bin specific exponential decay in the spectrum. The decay rates are obtained by using the blind estimates of reverberation time ratio between recorded output and test signals for each frequency bin. The modified impulse response does filter a recorded audio-signal by deconvolution. The blind estimation is well known and stands out for its robustness to noise and non-idealities. Estimation of a direct path signal is key to many applications.


【11】Stereo Audio Rendering for Personal Sound Zones Using a Binaural Spatially Adaptive Neural Network (BSANN)
标题:使用双耳空间自适应神经网络(BSANN)的个人声音区域的立体声音频渲染
链接:https://arxiv.org/abs/2601.06621

作者:Hao Jiang,Edgar Choueiri
备注:Submitted to IEEE Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:提出了一种用于个人声音区域(PSZ)的双耳渲染框架,以使多个头部跟踪的收听者能够接收完全独立的立体声音频节目。目前的PSZ系统通常依赖于单声道渲染,因此无法分别控制左耳和右耳,这限制了空间成像的质量和准确性。所提出的方法采用双耳空间自适应神经网络(BSANN)来生成耳朵优化的扬声器滤波器,该滤波器在多个听众的每个耳朵处重建所需的声场。该框架集成了消声测量的扬声器频率响应,分析建模的换能器方向性,和刚性球头相关的传递函数(HRTF),以提高声学精度和空间渲染保真度。显式有源串扰消除(XTC)级进一步改善了三维空间感知。实验表明,在测量的客观性能指标,包括区间隔离(IZI),程序间隔离(IPI),和串扰消除(XTC),与对数频率加权值为10.23/10.03 dB(IZI),11.11/9.16 dB(IPI),和10.55/11.13 dB(XTC),分别超过100- 20000 Hz的显着增益。耳控、精确声学建模和集成有源XTC的组合使用产生了一种统一的渲染方法,可提供更好的隔离性能、更强的房间不对称鲁棒性以及真实声学环境中更忠实的空间再现。
摘要:A binaural rendering framework for personal sound zones (PSZs) is proposed to enable multiple head-tracked listeners to receive fully independent stereo audio programs. Current PSZ systems typically rely on monophonic rendering and therefore cannot control the left and right ears separately, which limits the quality and accuracy of spatial imaging. The proposed method employs a Binaural Spatially Adaptive Neural Network (BSANN) to generate ear-optimized loudspeaker filters that reconstruct the desired acoustic field at each ear of multiple listeners. The framework integrates anechoically measured loudspeaker frequency responses, analytically modeled transducer directivity, and rigid-sphere head-related transfer functions (HRTFs) to enhance acoustic accuracy and spatial rendering fidelity. An explicit active crosstalk cancellation (XTC) stage further improves three-dimensional spatial perception. Experiments show significant gains in measured objective performance metrics, including inter-zone isolation (IZI), inter-program isolation (IPI), and crosstalk cancellation (XTC), with log-frequency-weighted values of 10.23/10.03 dB (IZI), 11.11/9.16 dB (IPI), and 10.55/11.13 dB (XTC), respectively, over 100-20,000 Hz. The combined use of ear-wise control, accurate acoustic modeling, and integrated active XTC produces a unified rendering method that delivers greater isolation performance, increased robustness to room asymmetry, and more faithful spatial reproduction in real acoustic environments.


【12】Lightweight Resolution-Aware Audio Deepfake Detection via Cross-Scale Attention and Consistency Learning
标题:通过跨尺度注意力和一致性学习的轻量级分辨率感知音频Deepfake检测
链接:https://arxiv.org/abs/2601.06560

作者:K. A. Shahriar
摘要:由于语音合成和语音转换技术的快速发展,音频deepfake检测变得越来越具有挑战性,特别是在信道失真,重放攻击和真实世界录音条件下。本文提出了一种分辨率感知的音频deepfake检测框架,该框架通过跨尺度注意力和一致性学习来显式建模和对齐多分辨率频谱表示。与传统的单分辨率或隐式特征融合方法不同,所提出的方法在互补的时间-频率尺度上执行协议。该框架在三个代表性的基准上进行了评估:ASVspoof 2019(LA和PA),Fake or Real(FoR)数据集和扬声器不相交协议下的In-the-Wild Audio Deepfake数据集。该方法在ASVspoof LA(EER 0.16%)上实现了近乎完美的性能,在ASVspoof PA(EER 5.09%),FoR重新录制的音频(EER 4.54%)和野外deepfakes(AUC 0.98,EER 4.81%)上实现了强大的鲁棒性,在具有挑战性的条件下显著优于单分辨率和非注意基线。所提出的模型仍然是轻量级和高效的,只需要159k的参数和小于1~GFLOP每个推理,使其适合于实际部署。全面的消融研究证实了跨尺度注意力和一致性学习的关键贡献,而基于梯度的可解释性分析表明,该模型在不同的欺骗条件下学习分辨率一致和语义上有意义的光谱线索。这些结果表明,显式交叉分辨率建模为下一代音频deepfake检测系统提供了原则性,鲁棒性和可扩展性的基础。
摘要:Audio deepfake detection has become increasingly challenging due to rapid advances in speech synthesis and voice conversion technologies, particularly under channel distortions, replay attacks, and real-world recording conditions. This paper proposes a resolution-aware audio deepfake detection framework that explicitly models and aligns multi-resolution spectral representations through cross-scale attention and consistency learning. Unlike conventional single-resolution or implicit feature-fusion approaches, the proposed method enforces agreement across complementary time--frequency scales. The proposed framework is evaluated on three representative benchmarks: ASVspoof 2019 (LA and PA), the Fake-or-Real (FoR) dataset, and the In-the-Wild Audio Deepfake dataset under a speaker-disjoint protocol. The method achieves near-perfect performance on ASVspoof LA (EER 0.16%), strong robustness on ASVspoof PA (EER 5.09%), FoR rerecorded audio (EER 4.54%), and in-the-wild deepfakes (AUC 0.98, EER 4.81%), significantly outperforming single-resolution and non-attention baselines under challenging conditions. The proposed model remains lightweight and efficient, requiring only 159k parameters and less than 1~GFLOP per inference, making it suitable for practical deployment. Comprehensive ablation studies confirm the critical contributions of cross-scale attention and consistency learning, while gradient-based interpretability analysis reveals that the model learns resolution-consistent and semantically meaningful spectral cues across diverse spoofing conditions. These results demonstrate that explicit cross-resolution modeling provides a principled, robust, and scalable foundation for next-generation audio deepfake detection systems.


【13】FastSLM: Hierarchical Frame Q-Former for Effective Speech Modality Adaptation
标题:FastLAM:用于有效语音情态自适应的分层帧Q-形成器
链接:https://arxiv.org/abs/2601.06199

作者:Junseok Lee,Sangyong Lee,Chang-Jae Chun
摘要:大型语言模型(LLM)的最新进展已经证明了人类专家级的能力,这引发了人们对它们实现人工通用智能(AGI)潜力的极大兴趣。特别是,通过开发多模式LLM(MLLM),使LLM适应各种模式(包括视觉,视频和语音)的势头越来越大。然而,现有的语音语言模型(SLM)的研究在很大程度上忽略了成本效益的适应策略,利用LLM在语音域。在本文中,我们提出了FastSLM,一个轻量级的,但有效的SLM设计有效的理解和推理长形式的语音。为了解决高帧速率语音特征与LLM对齐的挑战,我们引入了分层帧查询Transformer(HFQ-Former),它在捕获局部和全局上下文的同时压缩帧级语音特征。此外,我们提出了一种新的三阶段训练策略,可以在广泛的语音相关任务中增强泛化能力。实验结果表明,FastSLM实现了竞争力的性能相比,现有的国家的最先进的模型,尽管操作具有显着较低的FLOP和参数计数,同时表示语音只有1.67令牌每秒。源代码和模型检查点可以在https://huggingface.co/okestro-ai-lab/FastSLM上找到。
摘要:Recent advances in large language models (LLMs) have demonstrated human-expert-level capabilities, driving significant interest in their potential for achieving artificial general intelligence (AGI). In particular, there is growing momentum in adapting LLMs to various modalities, including vision, video, and speech, through the development of multimodal LLMs (MLLMs). However, existing speech-language model (SLM) research has largely overlooked cost-effective adaptation strategies for leveraging LLMs in the speech domain. In this paper, we propose FastSLM, a lightweight yet efficient SLM designed for effective understanding and reasoning over long-form speech. To address the challenge of aligning high-frame-rate speech features with LLMs, we introduce the Hierarchical Frame Querying Transformer (HFQ-Former), which compresses frame-level speech features while capturing both local and global context. Furthermore, we present a novel three-stage training strategy that enhances generalization across a wide range of speech-related tasks. Experimental results demonstrate that FastSLM achieves competitive performance compared to existing state-of-the-art models, despite operating with significantly lower FLOPs and parameter counts, while representing speech with only 1.67 tokens per second. The source code and model checkpoints are available at https://huggingface.co/okestro-ai-lab/FastSLM.


【14】Auditory Filter Behavior and Updated Estimated Constants
标题:听觉过滤器行为和更新的估计常数
链接:https://arxiv.org/abs/2601.06094

作者:Samiya A Alkhairy
备注:19 pages, 36 equations, 10 figures, 2 tables, submitted
摘要:Gammatone系列的滤波器通常用于模拟听觉信号处理,但用于模拟人类听觉的滤波器常数值主要设置为基于几十年前收集的历史心理声学数据的值。在这里,我们远离这个长期存在的惯例,并估计滤波器常数使用一系列最近报道的滤波器特性(如质量因子和质量因子与峰值群延迟之间的比率)在一个基于特性的框架内,澄清滤波器的行为是如何相关的基础常数。使用锐化滤波器近似,捕获共享的峰值区域的行为在某些类别的过滤器,我们分析的行为范围时,利用过滤器的自由度的全部,而不是固定的过滤器顺序或指数的历史规定的值。滤波器行为的特征在于使用基于幅度和基于相位的特性及其比率,其揭示了哪些特性对于约束滤波器常数是有用的,哪些特性仅是弱约束的。我们表明,这些见解和估计方法扩展到多个可实现的过滤器类的Gammatone家庭和应用它们,连同最近的生理和心理声学的观察,推导出的限制和估计人类听觉过滤器的过滤器常数。更广泛地说,这个框架支持设计的听觉滤波器与任意的特征级规格,并使系统的评估如何在过滤器特性的变化影响听觉模型,感知的结果,和技术,依赖于听觉滤波器组。
摘要:Filters from the Gammatone family are often used to model auditory signal processing, but the filter constant values used to mimic human hearing are largely set to values based on historical psychoacoustic data collected several decades ago. Here, we move away from this long-standing convention, and estimate filter constants using a range of more recent reported filter characteristics (such as quality factors and ratios between quality factors and peak group delay) within a characteristics-based framework that clarifies how filter behavior is related to the underlying constants. Using a sharp-filter approximation that captures shared peak-region behavior across certain classes of filters, we analyze the range of behaviors accessible when the full degrees of freedom of the filter are utilized rather than fixing the filter order or exponent to historically prescribed values. Filter behavior is characterized using magnitude-based and phase-based characteristics and their ratios, which reveal which characteristics are informative for constraining filter constants and which are only weakly constraining. We show that these insights and estimation methods extend to multiple realizable filter classes from the Gammatone family and apply them, together with recent physiological and psychoacoustic observations, to derive constraints on and estimates for filter constants for human auditory filters. More broadly, this framework supports the design of auditory filters with arbitrary characteristic-level specifications and enables systematic assessment of how variations in filter characteristics influence auditory models, perceptual findings, and technologies that rely on auditory filterbanks.


eess.AS音频处理


【1】Directional reflection modeling via wavenumber-domain reflection coefficient for 3D acoustic field simulation
标题:通过波数域反射系数进行定向反射建模用于3D声学场模拟
链接:https://arxiv.org/abs/2601.07481

作者:Satoshi Hoshika,Takahiro Iwami,Akira Omoto
备注:Submitted to Proceedings of Meetings on Acoustics (PoMA)
摘要:这项研究提出了一个框架,将波数域声反射系数到声场分析,以表征方向相关的材料反射和散射现象。反射系数被定义为每个传播方向的入射波和反射波之间的振幅比,并且从入射声场和反射声场的空间傅立叶变换来估计。所得到的波数域反射系数被转换成声导纳表示,该表示与诸如边界元法(BEM)的数值方法直接兼容,从而能够模拟超出简单镜面分量的反射。不同于传统的扩展反应模型,所提出的方法避免了显式建模的材料内部。这显著降低了计算成本,同时允许直接使用测量数据、经验模型或用户定义的定向反射特性。作者先前通过二维声场模拟证明了所提出的公式的有效性,其中确认了方向相关反射行为的准确再现。在目前的工作中,该框架扩展到三维分析,证明其适用于更现实和复杂的声学环境。所提出的方法提供了一个实用和灵活的工具,模拟方向相关的声反射和散射,在建筑声学,材料表征和噪声控制的潜在应用。
摘要:This study proposes a framework for incorporating wavenumber-domain acoustic reflection coefficients into sound field analysis to characterize direction-dependent material reflection and scattering phenomena. The reflection coefficient is defined as the amplitude ratio between incident and reflected waves for each propagation direction and is estimated from spatial Fourier transforms of the incident and reflected sound fields. The resulting wavenumber-domain reflection coefficients are converted into an acoustic admittance representation that is directly compatible with numerical methods such as the Boundary Element Method (BEM), enabling simulation of reflections beyond simple specular components. Unlike conventional extended reaction models, the proposed approach avoids explicit modeling of the material interior. This significantly reduces computational cost while allowing direct use of measured data, empirical models, or user-defined directional reflection characteristics. The validity of the proposed formulation was previously demonstrated by the authors through two-dimensional sound field simulations, in which accurate reproduction of direction-dependent reflection behavior was confirmed. In the present work, the framework is extended to three-dimensional analysis, demonstrating its applicability to more realistic and complex acoustic environments. The proposed approach provides a practical and flexible tool for simulating direction-dependent acoustic reflections and scattering, with potential applications in architectural acoustics, material characterization, and noise control.


【2】The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
标题:ICASP 2026自动歌曲美学评估挑战赛
链接:https://arxiv.org/abs/2601.07237

作者:Guobin Ma,Yuxuan Xia,Jixun Yao,Huixin Xue,Hexin Liu,Shuai Wang,Hao Liu,Lei Xie
备注:Official summary paper for the ICASSP 2026 ASAE Challenge
摘要:本文总结了ICASSP 2026自动歌曲美学评估(ASAE)挑战赛,该挑战赛专注于预测AI生成歌曲的主观美学评分。该挑战包括两个轨道:轨道1的目标是预测整体音乐性得分,而轨道2侧重于预测五个细粒度的美学得分。这一挑战引起了研究界的强烈兴趣,并收到了学术界和工业界的许多意见书。表现最好的系统大大超过了官方基线,表明在将客观指标与人类审美偏好相结合方面取得了实质性进展。这些成果为现代音乐生成系统建立了标准化的基准,并推进了与人类一致的评估方法。
摘要:This paper summarizes the ICASSP 2026 Automatic Song Aesthetics Evaluation (ASAE) Challenge, which focuses on predicting the subjective aesthetic scores of AI-generated songs. The challenge consists of two tracks: Track 1 targets the prediction of the overall musicality score, while Track 2 focuses on predicting five fine-grained aesthetic scores. The challenge attracted strong interest from the research community and received numerous submissions from both academia and industry. Top-performing systems significantly surpassed the official baseline, demonstrating substantial progress in aligning objective metrics with human aesthetic preferences. The outcomes establish a standardized benchmark and advance human-aligned evaluation methodologies for modern music generation systems.


【3】Bridging Attribution and Open-Set Detection using Graph-Augmented Instance Learning in Synthetic Speech
标题:在合成语音中使用图增强实例学习实现归因和开集检测的桥梁
链接:https://arxiv.org/abs/2601.07064

作者:Mohd Mujtaba Akhtar,Girish,Farhan Sheth,Muskaan Singh
备注:Accepted to EACL 2026
摘要:我们提出了一个统一的框架,不仅归因于合成语音的来源,但也用于检测语音合成器在训练过程中没有遇到的。这就需要超越简单检测的方法来支持详细的取证分析和开集泛化。为了解决这个问题,我们引入信号,一个混合框架,结合语音基础模型(SFM)与基于图形的建模和开放集感知推理。我们的框架集成了图神经网络(GNNs)和k最近邻(KNN)分类器,使其能够捕获话语之间有意义的关系,并识别不属于任何已知生成器的语音。它在生成器类原型上构建了一个查询条件图,使GNN能够推理候选生成器之间的关系,而KNN分支通过基于置信度的阈值来支持开集检测。我们使用DiffSSD数据集评估SIGNAL,该数据集提供了来自开源和商业扩散TTS系统的真实语音和合成音频的多样化组合。为了进一步评估泛化,我们还在SingFake基准上进行了测试。我们的研究结果表明,SIGNAL始终提高了这两项任务的性能,基于Mamba的嵌入提供了特别强大的结果。据我们所知,这是第一个统一基于图的学习和开集检测的研究,用于追踪合成语音的起源。
摘要:We propose a unified framework for not only attributing synthetic speech to its source but also for detecting speech generated by synthesizers that were not encountered during training. This requires methods that move beyond simple detection to support both detailed forensic analysis and open-set generalization. To address this, we introduce SIGNAL, a hybrid framework that combines speech foundation models (SFMs) with graph-based modeling and open-set-aware inference. Our framework integrates Graph Neural Networks (GNNs) and a k-Nearest Neighbor (KNN) classifier, allowing it to capture meaningful relationships between utterances and recognize speech that doesn`t belong to any known generator. It constructs a query-conditioned graph over generator class prototypes, enabling the GNN to reason over relationships among candidate generators, while the KNN branch supports open-set detection via confidence-based thresholding. We evaluate SIGNAL using the DiffSSD dataset, which offers a diverse mix of real speech and synthetic audio from both open-source and commercial diffusion-based TTS systems. To further assess generalization, we also test on the SingFake benchmark. Our results show that SIGNAL consistently improves performance across both tasks, with Mamba-based embeddings delivering especially strong results. To the best of our knowledge, this is the first study to unify graph-based learning and open-set detection for tracing synthetic speech back to its origin.


【4】DIVINE: Coordinating Multimodal Disentangled Representations for Oro-Facial Neurological Disorder Assessment
标题:DIVINE:协调口腔面部神经系统疾病评估的多模式分解表示
链接:https://arxiv.org/abs/2601.07014

作者:Mohd Mujtaba Akhtar,Girish,Muskaan Singh
备注:Accepted to EACL 2026
摘要:在这项研究中,我们提出了一个多模态框架,通过捕捉声音和面部线索来预测神经面部疾病。我们假设,明确地解开共享和模态特定的表示在多模态基础模型嵌入可以提高临床的可解释性和推广。为了验证这一假设,我们提出了DIVINE一个完全分离的多模式框架,该框架对从最先进的(SOTA)音频和视频基础模型中提取的表示进行操作,包括分层变分瓶颈,稀疏门控融合和可学习的症状令牌。DIVINE在多任务学习设置中运行,以联合预测诊断类别(健康对照,ALS,中风)和严重程度(轻度,中度,重度)。该模型使用同步的音频和视频输入进行训练,并在完整(音频-视频)以及单模态(仅音频和仅视频)测试条件下在Toronto NeuroFace数据集上进行评估。我们提出的方法DIVINE实现了SOTA结果,DeepSeek-VL 2和TRILLsson的组合达到了98.26%的准确率和97.51%的F1分数。在模态约束的情况下,该框架表现良好,在仅使用视频或仅使用音频输入进行测试时表现出很强的泛化能力。与单模态模型和基线融合技术相比,它始终产生卓越的性能。据我们所知,DIVINE是第一个结合了跨模态解纠缠、自适应融合和多任务学习的框架,可以使用同步语音和面部视频全面评估神经系统疾病。
摘要:In this study, we present a multimodal framework for predicting neuro-facial disorders by capturing both vocal and facial cues. We hypothesize that explicitly disentangling shared and modality-specific representations within multimodal foundation model embeddings can enhance clinical interpretability and generalization. To validate this hypothesis, we propose DIVINE a fully disentangled multimodal framework that operates on representations extracted from state-of-the-art (SOTA) audio and video foundation models, incorporating hierarchical variational bottlenecks, sparse gated fusion, and learnable symptom tokens. DIVINE operates in a multitask learning setup to jointly predict diagnostic categories (Healthy Control,ALS, Stroke) and severity levels (Mild, Moderate, Severe). The model is trained using synchronized audio and video inputs and evaluated on the Toronto NeuroFace dataset under full (audio-video) as well as single-modality (audio- only and video-only) test conditions. Our proposed approach, DIVINE achieves SOTA result, with the DeepSeek-VL2 and TRILLsson combination reaching 98.26% accuracy and 97.51% F1-score. Under modality-constrained scenarios, the framework performs well, showing strong generalization when tested with video-only or audio-only inputs. It consistently yields superior performance compared to unimodal models and baseline fusion techniques. To the best of our knowledge, DIVINE is the first framework that combines cross-modal disentanglement, adaptive fusion, and multitask learning to comprehensively assess neurological disorders using synchronized speech and facial video.


【5】TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
标题:TagSpeech:端到端多扬声器ASB和具有细粒度时间基础的拨号
链接:https://arxiv.org/abs/2601.06896

作者:Mingyue Huo,Yiwen Shao,Yuheng Zhang
摘要:我们提出了TagSpeech,一个统一的基于LLM的框架,利用时间锚接地联合多扬声器ASR和日记。该框架建立在两个关键设计之上:(1)通过串行化输出训练(SOT)进行微调的解耦语义和说话者流,以学习话轮转换动态;(2)交织时间锚机制,不仅支持细粒度时间戳预测,还充当语义理解和说话者跟踪之间的同步信号。与以前主要关注说话者属性ASR或隐式日记化的作品相比,TagSpeech解决了细粒度说话者内容对齐的挑战,并以端到端的方式明确地建模“谁说了什么,什么时候说了什么”。在AMI和AliMeeting基准测试上的实验表明,我们的方法在强大的端到端基线(包括Qwen-Omni和Gemini)上实现了Diarization Error Rate(DER)的一致改进,特别是在处理复杂的语音重叠方面。此外,TagSpeech采用了一种参数高效的训练范式,其中LLM骨干被冻结,并且只训练轻量级投影仪,从而以低计算成本获得强大的性能。
摘要:We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and (2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models "who spoke what and when" in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost.


【6】Dereverberation Filter by Deconvolution with Frequency Bin Specific Faded Impulse Response
标题:利用特定频率段的衰落冲激响应去卷积进行去混滤
链接:https://arxiv.org/abs/2601.06662

作者:Stefan Ciba
备注:8 pages, 3 figures, github repository with code and audio
摘要:这项工作介绍了一个强大的单通道逆滤波器去混响的非理想的录音,在真实的音频验证。所开发的方法侧重于离散脉冲响应的计算和修改,以便从已知的数字单通道记录设置和房间特性(如早期反射和混响)中过滤特性。目标是更干燥和更清晰的信号重建,理想情况下是直接路径信号。从倒谱域计算时域脉冲响应,并通过频谱中的频率单元特定指数衰减来衰减。衰减率是通过使用记录的输出和测试信号之间的混响时间比的盲估计获得的每个频率箱。修改的脉冲响应通过去卷积对记录的音频信号进行滤波。盲估计是众所周知的,并且由于其对噪声和非理想性的鲁棒性而突出。直接路径信号的估计是许多应用的关键。
摘要:This work introduces a robust single-channel inverse filter for dereverberation of non-ideal recordings, validated on real audio. The developed method focuses on the calculation and modification of a discrete impulse response in order to filter the characteristics from a known digital single channel recording setup and room characteristics such as early reflections and reverberations. The aim is a dryer and clearer signal reconstruction, which ideally would be the direct-path signal. The time domain impulse response is calculated from the cepstral domain and faded by means of frequency bin specific exponential decay in the spectrum. The decay rates are obtained by using the blind estimates of reverberation time ratio between recorded output and test signals for each frequency bin. The modified impulse response does filter a recorded audio-signal by deconvolution. The blind estimation is well known and stands out for its robustness to noise and non-idealities. Estimation of a direct path signal is key to many applications.


【7】Stereo Audio Rendering for Personal Sound Zones Using a Binaural Spatially Adaptive Neural Network (BSANN)
标题:使用双耳空间自适应神经网络(BSANN)的个人声音区域的立体声音频渲染
链接:https://arxiv.org/abs/2601.06621

作者:Hao Jiang,Edgar Choueiri
备注:Submitted to IEEE Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:提出了一种用于个人声音区域(PSZ)的双耳渲染框架,以使多个头部跟踪的收听者能够接收完全独立的立体声音频节目。目前的PSZ系统通常依赖于单声道渲染,因此无法分别控制左耳和右耳,这限制了空间成像的质量和准确性。所提出的方法采用双耳空间自适应神经网络(BSANN)来生成耳朵优化的扬声器滤波器,该滤波器在多个听众的每个耳朵处重建所需的声场。该框架集成了消声测量的扬声器频率响应,分析建模的换能器方向性,和刚性球头相关的传递函数(HRTF),以提高声学精度和空间渲染保真度。显式有源串扰消除(XTC)级进一步改善了三维空间感知。实验表明,在测量的客观性能指标,包括区间隔离(IZI),程序间隔离(IPI),和串扰消除(XTC),与对数频率加权值为10.23/10.03 dB(IZI),11.11/9.16 dB(IPI),和10.55/11.13 dB(XTC),分别超过100- 20000 Hz的显着增益。耳控、精确声学建模和集成有源XTC的组合使用产生了一种统一的渲染方法,可提供更好的隔离性能、更强的房间不对称鲁棒性以及真实声学环境中更忠实的空间再现。
摘要:A binaural rendering framework for personal sound zones (PSZs) is proposed to enable multiple head-tracked listeners to receive fully independent stereo audio programs. Current PSZ systems typically rely on monophonic rendering and therefore cannot control the left and right ears separately, which limits the quality and accuracy of spatial imaging. The proposed method employs a Binaural Spatially Adaptive Neural Network (BSANN) to generate ear-optimized loudspeaker filters that reconstruct the desired acoustic field at each ear of multiple listeners. The framework integrates anechoically measured loudspeaker frequency responses, analytically modeled transducer directivity, and rigid-sphere head-related transfer functions (HRTFs) to enhance acoustic accuracy and spatial rendering fidelity. An explicit active crosstalk cancellation (XTC) stage further improves three-dimensional spatial perception. Experiments show significant gains in measured objective performance metrics, including inter-zone isolation (IZI), inter-program isolation (IPI), and crosstalk cancellation (XTC), with log-frequency-weighted values of 10.23/10.03 dB (IZI), 11.11/9.16 dB (IPI), and 10.55/11.13 dB (XTC), respectively, over 100-20,000 Hz. The combined use of ear-wise control, accurate acoustic modeling, and integrated active XTC produces a unified rendering method that delivers greater isolation performance, increased robustness to room asymmetry, and more faithful spatial reproduction in real acoustic environments.


【8】Lightweight Resolution-Aware Audio Deepfake Detection via Cross-Scale Attention and Consistency Learning
标题:通过跨尺度注意力和一致性学习的轻量级分辨率感知音频Deepfake检测
链接:https://arxiv.org/abs/2601.06560

作者:K. A. Shahriar
摘要:由于语音合成和语音转换技术的快速发展,音频deepfake检测变得越来越具有挑战性,特别是在信道失真,重放攻击和真实世界录音条件下。本文提出了一种分辨率感知的音频deepfake检测框架,该框架通过跨尺度注意力和一致性学习来显式建模和对齐多分辨率频谱表示。与传统的单分辨率或隐式特征融合方法不同,所提出的方法在互补的时间-频率尺度上执行协议。该框架在三个代表性的基准上进行了评估:ASVspoof 2019(LA和PA),Fake or Real(FoR)数据集和扬声器不相交协议下的In-the-Wild Audio Deepfake数据集。该方法在ASVspoof LA(EER 0.16%)上实现了近乎完美的性能,在ASVspoof PA(EER 5.09%),FoR重新录制的音频(EER 4.54%)和野外deepfakes(AUC 0.98,EER 4.81%)上实现了强大的鲁棒性,在具有挑战性的条件下显著优于单分辨率和非注意基线。所提出的模型仍然是轻量级和高效的,只需要159 k的参数和小于1~GFLOP每个推理,使其适合于实际部署。全面的消融研究证实了跨尺度注意力和一致性学习的关键贡献,而基于梯度的可解释性分析表明,该模型在不同的欺骗条件下学习分辨率一致和语义上有意义的光谱线索。这些结果表明,显式交叉分辨率建模为下一代音频deepfake检测系统提供了原则性,鲁棒性和可扩展性的基础。
摘要:Audio deepfake detection has become increasingly challenging due to rapid advances in speech synthesis and voice conversion technologies, particularly under channel distortions, replay attacks, and real-world recording conditions. This paper proposes a resolution-aware audio deepfake detection framework that explicitly models and aligns multi-resolution spectral representations through cross-scale attention and consistency learning. Unlike conventional single-resolution or implicit feature-fusion approaches, the proposed method enforces agreement across complementary time--frequency scales. The proposed framework is evaluated on three representative benchmarks: ASVspoof 2019 (LA and PA), the Fake-or-Real (FoR) dataset, and the In-the-Wild Audio Deepfake dataset under a speaker-disjoint protocol. The method achieves near-perfect performance on ASVspoof LA (EER 0.16%), strong robustness on ASVspoof PA (EER 5.09%), FoR rerecorded audio (EER 4.54%), and in-the-wild deepfakes (AUC 0.98, EER 4.81%), significantly outperforming single-resolution and non-attention baselines under challenging conditions. The proposed model remains lightweight and efficient, requiring only 159k parameters and less than 1~GFLOP per inference, making it suitable for practical deployment. Comprehensive ablation studies confirm the critical contributions of cross-scale attention and consistency learning, while gradient-based interpretability analysis reveals that the model learns resolution-consistent and semantically meaningful spectral cues across diverse spoofing conditions. These results demonstrate that explicit cross-resolution modeling provides a principled, robust, and scalable foundation for next-generation audio deepfake detection systems.


【9】FastSLM: Hierarchical Frame Q-Former for Effective Speech Modality Adaptation
标题:FastLAM:用于有效语音情态自适应的分层帧Q-形成器
链接:https://arxiv.org/abs/2601.06199

作者:Junseok Lee,Sangyong Lee,Chang-Jae Chun
摘要:大型语言模型(LLM)的最新进展已经证明了人类专家级的能力,这引发了人们对它们实现人工通用智能(AGI)潜力的极大兴趣。特别是,通过开发多模式LLM(MLLM),使LLM适应各种模式(包括视觉,视频和语音)的势头越来越大。然而,现有的语音语言模型(SLM)的研究在很大程度上忽略了成本效益的适应策略,利用LLM在语音域。在本文中,我们提出了FastSLM,一个轻量级的,但有效的SLM设计有效的理解和推理长形式的语音。为了解决高帧速率语音特征与LLM对齐的挑战,我们引入了分层帧查询Transformer(HFQ-Former),它在捕获局部和全局上下文的同时压缩帧级语音特征。此外,我们提出了一种新的三阶段训练策略,可以在广泛的语音相关任务中增强泛化能力。实验结果表明,FastSLM实现了竞争力的性能相比,现有的国家的最先进的模型,尽管操作具有显着较低的FLOP和参数计数,同时表示语音只有1.67令牌每秒。源代码和模型检查点可以在https://huggingface.co/okestro-ai-lab/FastSLM上找到。
摘要:Recent advances in large language models (LLMs) have demonstrated human-expert-level capabilities, driving significant interest in their potential for achieving artificial general intelligence (AGI). In particular, there is growing momentum in adapting LLMs to various modalities, including vision, video, and speech, through the development of multimodal LLMs (MLLMs). However, existing speech-language model (SLM) research has largely overlooked cost-effective adaptation strategies for leveraging LLMs in the speech domain. In this paper, we propose FastSLM, a lightweight yet efficient SLM designed for effective understanding and reasoning over long-form speech. To address the challenge of aligning high-frame-rate speech features with LLMs, we introduce the Hierarchical Frame Querying Transformer (HFQ-Former), which compresses frame-level speech features while capturing both local and global context. Furthermore, we present a novel three-stage training strategy that enhances generalization across a wide range of speech-related tasks. Experimental results demonstrate that FastSLM achieves competitive performance compared to existing state-of-the-art models, despite operating with significantly lower FLOPs and parameter counts, while representing speech with only 1.67 tokens per second. The source code and model checkpoints are available at https://huggingface.co/okestro-ai-lab/FastSLM.


【10】Auditory Filter Behavior and Updated Estimated Constants
标题:听觉过滤器行为和更新的估计常数
链接:https://arxiv.org/abs/2601.06094

作者:Samiya A Alkhairy
备注:19 pages, 36 equations, 10 figures, 2 tables, submitted
摘要:Gammatone系列的滤波器通常用于模拟听觉信号处理,但用于模拟人类听觉的滤波器常数值主要设置为基于几十年前收集的历史心理声学数据的值。在这里,我们远离这个长期存在的惯例,并估计滤波器常数使用一系列最近报道的滤波器特性(如质量因子和质量因子与峰值群延迟之间的比率)在一个基于特性的框架内,澄清滤波器的行为是如何相关的基础常数。使用锐化滤波器近似,捕获共享的峰值区域的行为在某些类别的过滤器,我们分析的行为范围时,利用过滤器的自由度的全部,而不是固定的过滤器顺序或指数的历史规定的值。滤波器行为的特征在于使用基于幅度和基于相位的特性及其比率,其揭示了哪些特性对于约束滤波器常数是有用的,哪些特性仅是弱约束的。我们表明,这些见解和估计方法扩展到多个可实现的过滤器类的Gammatone家庭和应用它们,连同最近的生理和心理声学的观察,推导出的限制和估计人类听觉过滤器的过滤器常数。更广泛地说,这个框架支持设计的听觉滤波器与任意的特征级规格,并使系统的评估如何在过滤器特性的变化影响听觉模型,感知的结果,和技术,依赖于听觉滤波器组。
摘要:Filters from the Gammatone family are often used to model auditory signal processing, but the filter constant values used to mimic human hearing are largely set to values based on historical psychoacoustic data collected several decades ago. Here, we move away from this long-standing convention, and estimate filter constants using a range of more recent reported filter characteristics (such as quality factors and ratios between quality factors and peak group delay) within a characteristics-based framework that clarifies how filter behavior is related to the underlying constants. Using a sharp-filter approximation that captures shared peak-region behavior across certain classes of filters, we analyze the range of behaviors accessible when the full degrees of freedom of the filter are utilized rather than fixing the filter order or exponent to historically prescribed values. Filter behavior is characterized using magnitude-based and phase-based characteristics and their ratios, which reveal which characteristics are informative for constraining filter constants and which are only weakly constraining. We show that these insights and estimation methods extend to multiple realizable filter classes from the Gammatone family and apply them, together with recent physiological and psychoacoustic observations, to derive constraints on and estimates for filter constants for human auditory filters. More broadly, this framework supports the design of auditory filters with arbitrary characteristic-level specifications and enables systematic assessment of how variations in filter characteristics influence auditory models, perceptual findings, and technologies that rely on auditory filterbanks.


【11】Directional Selective Fixed-Filter Active Noise Control Based on a Convolutional Neural Network in Reverberant Environments
标题:回响环境下基于卷积神经网络的定向选择性固定过滤器主动噪音控制
链接:https://arxiv.org/abs/2601.06981

作者:Boxiang Wang,Zhengding Luo,Haowen Li,Dongyuan Shi,Junwei Ji,Ziyi Yang,Woon-Seng Gan
摘要:选择性固定滤波器有源噪声控制(SFANC)是一种能够抑制具有变化频率特性的噪声的新方法。与传统的自适应算法相比,它具有更快的响应速度和更高的计算效率。然而,空间因素,特别是噪声源位置的影响,往往被忽视。现有的一些研究已经探索了噪声源的到达方向(DoA)对ANC性能的影响,但它们大多限于自由场条件,并且没有考虑更复杂的室内混响环境。为了解决这一问题,本文提出了一种基于学习的方向SFANC方法,该方法结合了混响环境中噪声源的DoA。在该框架中,多个参考信号由卷积神经网络(CNN)处理以估计噪声源的方位角和仰角,以及识别用于有效噪声消除的最适当的控制滤波器。与传统的自适应算法相比,所提出的方法实现了更短的响应时间,即使在混响的存在下,优越的降噪。
摘要:Selective fixed-filter active noise control (SFANC) is a novel approach capable of mitigating noise with varying frequency characteristics. It offers faster response and greater computational efficiency compared to traditional adaptive algorithms. However, spatial factors, particularly the influence of the noise source location, are often overlooked. Some existing studies have explored the impact of the direction-of-arrival (DoA) of the noise source on ANC performance, but they are mostly limited to free-field conditions and do not consider the more complex indoor reverberant environments. To address this gap, this paper proposes a learning-based directional SFANC method that incorporates the DoA of the noise source in reverberant environments. In this framework, multiple reference signals are processed by a convolutional neural network (CNN) to estimate the azimuth and elevation angles of the noise source, as well as to identify the most appropriate control filter for effective noise cancellation. Compared to traditional adaptive algorithms, the proposed approach achieves superior noise reduction with shorter response times, even in the presence of reverberations.


【12】Variational decomposition autoencoding improves disentanglement of latent representations
标题:变分分解自动编码改进了潜在表示的解纠缠
链接:https://arxiv.org/abs/2601.06844

作者:Ioannis Ziogas,Aamna Al Shehhi,Ahsan H. Khandoker,Leontios J. Hadjileontiadis
备注:Supplementary information file at: https://drive.google.com/drive/folders/1sZl2AcCtRK-1oav7XZSaxlu0Cq0-3MMs?usp=sharing
摘要:理解复杂、非平稳、高维时间演化信号的结构是科学数据分析的核心挑战。在许多领域,如语音和生物医学信号处理,学习解纠缠和可解释的表示的能力是揭示潜在的生成机制的关键。传统的无监督表示学习方法,包括变分自编码器(VAE),往往难以捕捉这些数据中固有的时间和频谱多样性。在这里,我们介绍变分分解自编码(VDA),一个框架,扩展VAE通过将一个强大的结构偏向信号分解。VDA通过变分分解自编码器(DecVAE)来实例化,即,仅编码器神经网络,其组合信号分解模型、对比自监督任务和变分先验近似以学习与时频特性对准的多个潜在子空间。我们证明了DecVAE对模拟数据和三个公开可用的科学数据集的有效性,包括语音识别,构音障碍严重程度评估和情感语音分类。我们的研究结果表明,DecVAE在解纠缠质量、跨任务泛化和潜在编码的可解释性方面超过了最先进的基于VAE的方法。这些发现表明,分解感知架构可以作为从动态信号中提取结构化表示的强大工具,在临床诊断,人机交互和自适应神经技术中具有潜在的应用。
摘要:Understanding the structure of complex, nonstationary, high-dimensional time-evolving signals is a central challenge in scientific data analysis. In many domains, such as speech and biomedical signal processing, the ability to learn disentangled and interpretable representations is critical for uncovering latent generative mechanisms. Traditional approaches to unsupervised representation learning, including variational autoencoders (VAEs), often struggle to capture the temporal and spectral diversity inherent in such data. Here we introduce variational decomposition autoencoding (VDA), a framework that extends VAEs by incorporating a strong structural bias toward signal decomposition. VDA is instantiated through variational decomposition autoencoders (DecVAEs), i.e., encoder-only neural networks that combine a signal decomposition model, a contrastive self-supervised task, and variational prior approximation to learn multiple latent subspaces aligned with time-frequency characteristics. We demonstrate the effectiveness of DecVAEs on simulated data and three publicly available scientific datasets, spanning speech recognition, dysarthria severity evaluation, and emotional speech classification. Our results demonstrate that DecVAEs surpass state-of-the-art VAE-based methods in terms of disentanglement quality, generalization across tasks, and the interpretability of latent encodings. These findings suggest that decomposition-aware architectures can serve as robust tools for extracting structured representations from dynamic signals, with potential applications in clinical diagnostics, human-computer interaction, and adaptive neurotechnologies.


【13】AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
标题:Deliveros:通过自生成的免指令调优将LLM扩展到语音
链接:https://arxiv.org/abs/2601.06086

作者:Yiwen Shao,Wei Liu,Jiahong Li,Tianzi Wang,Kun Wei,Meng Yu,Dong Yu
备注:Technical Report
摘要:将大型语言模型(LLM)扩展到语音领域最近受到了极大的关注。一种典型的方法是通过投影模块将预训练的LLM与音频编码器连接起来,并在大规模的特定于任务的预处理调整数据集上训练生成的模型。然而,为特定需求管理这种预防-调整数据是耗时的,并且以这种方式训练的模型通常对看不见的任务泛化能力很差。在这项工作中,我们首先制定了语音LLM的最强泛化是在使用自生成无指令调谐(SIFT)进行训练时实现的,其中监督信号由冻结的LLM使用语音的文本表示作为输入生成。我们提出的SIFT范式消除了收集特定任务的问题-答案对的需要,并产生了理论上最好的概括看不见的任务。在此范例的基础上,我们引入了AZeroS(Auden Zero-instruction-tuned Speech-LLM),它是在来自公开语料库的语音文本对上训练的,包括大约25,000小时的带有ASR成绩单的语音和3,000小时的带有非语言标签的语音。该模型基于Qwen2.5- 7 B-Instruct构建,仅更新两个轻量级投影模块(每个模块2380万个参数),同时保持LLM和音频编码器冻结。尽管训练成本最低,数据规模适中,但AZeroS在语义和语言基准测试方面都达到了最先进的性能,包括语音基准测试、AIR-Bench Foundation(语音)和AIR-Bench Chat(语音)。
摘要:Extending large language models (LLMs) to the speech domain has recently gained significant attention. A typical approach connects a pretrained LLM with an audio encoder through a projection module and trains the resulting model on large-scale, task-specific instruction-tuning datasets. However, curating such instruction-tuning data for specific requirements is time-consuming, and models trained in this manner often generalize poorly to unseen tasks. In this work, we first formulate that the strongest generalization of a speech-LLM is achieved when it is trained with Self-Generated Instruction-Free Tuning (SIFT), in which supervision signals are generated by a frozen LLM using textual representations of speech as input. Our proposed SIFT paradigm eliminates the need for collecting task-specific question-answer pairs and yields the theoretically best generalization to unseen tasks. Building upon this paradigm, we introduce AZeroS (Auden Zero-instruction-tuned Speech-LLM), which is trained on speech-text pairs derived from publicly available corpora, including approximately 25,000 hours of speech with ASR transcripts and 3,000 hours of speech with paralinguistic labels. Built upon Qwen2.5-7B-Instruct, the model updates only two lightweight projection modules (23.8 million parameters each), while keeping both the LLM and audio encoders frozen. Despite the minimal training cost and modest data scale, AZeroS achieves state-of-the-art performance on both semantic and paralinguistic benchmarks, including VoiceBench, AIR-Bench Foundation (Speech), and AIR-Bench Chat (Speech).


机器翻译由腾讯交互翻译提供,仅供参考