今日论文合集:cs.SD语音14篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models
标题:超越文本跟踪:音频语言模型中的可修复仲裁逆转
链接:https://arxiv.org/abs/2606.05161
作者:Yichen Gao, Yiqun Zhang, Zijing Wang, Yujia Li, Heng Guo, Xi Wu, Xiaocui Yang, Shi Feng, Yifei Zhang, Daling Wang
摘要:
摘要:


【2】Audio Interaction Model

标题:音频交互模型
链接:https://arxiv.org/abs/2606.05121
作者:Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, Yue Liao, Ziyang Ma, Dongchao Yang, Mingbao Lin, Deheng Ye, Shuicheng Yan, Chunyan Miao
备注:Next generation of LALMs, work in progress
摘要:
摘要:


【3】FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

标题:FoeGlass:简单的上下文学习足以满足Red团队音频Deepfake检测器的需求
链接:https://arxiv.org/abs/2606.05101
作者:Sepehr Dehdashtian, Jacob H Seidman, Vishnu N Boddeti, Gaurav Bharaj
备注:Accepted at ICML 2026
摘要:
摘要:


【4】SURF: Separation via Unsupervised Remixing Flow

标题:SURF:通过无监督的重新混合流程进行分离
链接:https://arxiv.org/abs/2606.04921
作者:Henry Li,Robin Scheibler,Efthymios Tzinis,Matt Shannon,Arnaud Doucet,John R. Hershey
备注:Accepted at ICML 2026
摘要:单通道源分离的目标是重建$K$源给定他们的混合。在有监督的环境中,大量的干净的源数据是可用的,这个具有挑战性的,不适定的问题已经成功地解决了生成扩散和基于流的先验模型。然而,对这些干净源样本的访问通常是有限的,即使可用,监督模型也容易受到域偏移的影响。为了弥补这一差距,我们提出了通过无监督再混合流(SURF)进行分离,这是一种用于源分离的无监督流匹配方法,可直接从观察到的混合物中学习。该方法依赖于最先进的监督流匹配和基于回归的自监督技术的新组合。在高层次上,从教师模型开始,我们利用“重新混合”步骤来引导从教师的估计中学习学生流模型。我们提供了通过这种方法优化的目标的见解,并绘制了一个新的连接到唤醒-睡眠算法。对图像和音频基准的实证评估表明,SURF建立了一个新的国家的最先进的,显着优于现有的无监督方法。查看我们的演示页面以获取示例。https://google.github.io/df-conformer/surf/
摘要:The goal of single-channel source separation is to reconstruct $K$ sources given their mixture. In supervised settings where vast amounts of clean source data are available, this challenging, ill-posed problem has been addressed successfully by generative diffusion and flow-based prior models. However, access to such clean source samples is often limited, and even when available, supervised models are vulnerable to domain shifts. To bridge this gap, we present Separation via Unsupervised Remixing Flow (SURF), an unsupervised flow matching approach for source separation that learns directly from observed mixtures. This method relies on a novel combination of state-of-the-art supervised flow matching and regression-based self-supervised techniques. At a high level, starting from a teacher model, we utilize a "remixing" step to bootstrap the learning of a student flow model from the teacher's estimates. We provide insights into the objectives optimized by this approach and draw a novel connection to the Wake-Sleep algorithm. Empirical evaluations on image and audio benchmarks demonstrate that SURF establishes a new state-of-the-art, significantly outperforming existing unsupervised methods. See our demo page for examples. https://google.github.io/df-conformer/surf/


【5】Drift-Augmented Scoring: Text-Derived Noise Robustness for Zero-Shot Audio-Language Classification

标题:漂移增强评分:Zero-Shot音频语言分类的文本派生噪音稳健性
链接:https://arxiv.org/abs/2606.04844
作者:Tu Vo,Sheir Zaheer,Chan Y. Park
摘要:CLAP等对比音频语言模型支持zero-shot音频分类:通过将声音的嵌入与文本提示嵌入相匹配来标记声音,而没有标记的音频。这种匹配在声学噪声下失效,在标准基准上,在0 dB SNR下,精度和mAP下降12 - 30个百分点。我们提出了漂移增强评分(DAS),一个小的每类奖金添加到余弦分数。当噪声音频嵌入在类的噪声条件文本提示预测的方向上漂移时,奖金奖励类。它仅从文本中派生,计算一次并缓存,并在推理时为每个类添加单个内积,没有梯度,也没有测试时批处理。在LAION CLAP主干上,我们将DAS与Acevedo等人的四种变体进行比较。的并发方法上UrbanSound8K和完整的FSD50K评估集,混合每个剪辑与城市声学场景噪声在一个范围内的信噪比。DAS在每个测试条件下都能提高指标:UrbanSound8K上的准确度为+2.60至+5.75,FSD50K上的准确度为+1.50至+1.74 mAP。
摘要:Contrastive audio-language models such as CLAP enable zero-shot audio classification: a sound is labelled by matching its embedding to text prompt embeddings, with no labelled audio. This matching breaks down under acoustic noise, where accuracy and mAP fall by 12-30 percentage points at 0 dB SNR on standard benchmarks. We propose Drift Augmented Scoring (DAS), a small per-class bonus added to the cosine score. The bonus rewards a class when the noisy audio embedding drifts in the direction that the class's noise-conditioned text prompts predict. It is derived from text alone, computed once and cached, and adds a single inner product per class at inference, with no gradients and no test-time batch. On a LAION CLAP backbone, we compare DAS against the four variants of Acevedo et al.'s concurrent method on UrbanSound8K and the full FSD50K eval set, mixing each clip with urban acoustic scene noise across a range of SNRs. DAS improves the metric on every test condition: by +2.60 to +5.75 accuracy points on UrbanSound8K and +1.50 to +1.74 mAP points on FSD50K.


【6】SHB-AE: Spherical harmonic beamforming based Ambisonics encoding and upscaling method for smartphone microphone array

标题:SHB-AE:智能手机麦克风阵列基于球调和束的立体声合成编码和放大方法
链接:https://arxiv.org/abs/2606.04584
作者:Yuhuan You, Yufan Qian, Tianshu Qu, Bin Wang, Xueyang Lv
备注:Accepted for presentation at AES Europe 2025 Convention (AES 158th Convention), Warsaw, Poland, May 22-24, 2025
摘要:
摘要:


【7】Flow-HOA: Generative Joint Optimization for Ambisonics Encoding via Flow Matching

标题:Flow-HOA:通过流匹配的立体声立体声编码生成联合优化
链接:https://arxiv.org/abs/2606.04570
作者:Yuhuan You, Yufan Qian, Tianshu Qu, Bin Wang, Xueyang Lv
备注:Accepted for presentation at AES Europe 2026 Convention (AES 160th Convention), Copenhagen, Denmark, May 28-30, 2026
摘要:
摘要:


【8】A Second-Order Cepstral Signature of Contact-Vibration Sounds Reproduced by Laptop Loudspeakers: A Synthetic Case Study

标题:笔记本电脑扬声器再现的接触振动声音的二阶倒谱特征:综合案例研究
链接:https://arxiv.org/abs/2606.04475
作者:Jim Salsman
备注:11 pages, 4 tables, 5 figures, 8 references
摘要:
摘要:


【9】CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

标题:CleanCodec:通过感知引导编码实现高效和鲁棒的语音令牌化
链接:https://arxiv.org/abs/2606.04418
作者:Eugene Kwek, Feng Liu, Rui Zhang, Wenpeng Yin
摘要:
摘要:


【10】Gauss Circle Lattices with Geometric Convolutions for Synthesizing High Dimensional Image-Source Room Impulse Responses

标题:具有几何卷积的高斯圆格用于合成多维像源室脉冲响应
链接:https://arxiv.org/abs/2606.04358
作者:Yuancheng Luo
备注:Accepted for publication at the 29th International Conference on Digital Audio Effects 2026
摘要:
摘要:


【11】Feasibility of Time-Domain DNN-Based Speech Enhancement on Embedded FPGA for Hearing Aid

标题:嵌入式DSP助听器中基于时分DNN的语音增强的可行性
链接:https://arxiv.org/abs/2606.04221
作者:Feyisayo Olalere, Umut Altin, Kiki van der Heijden, Marcel van Gerven
备注:13 pages
摘要:
摘要:


【12】DetectZoo: A Unified Toolkit for AI-Generated Content Detection Across Text, Audio, and Image Modalities

标题:DetectZoo:用于跨文本、音频和图像模式的人工智能生成内容检测的统一工具包
链接:https://arxiv.org/abs/2606.04205
作者:Sajad Ebrahimi, Nima Jamali, Bardia Shirsalimian, Kelly McConvey, Wentao Zhang, Jalehsadat Mahdavimoghaddam, Maksym Taranukhin, Maura Grossman, Vered Shwartz, Yuntian Deng, Ebrahim Bagheri
摘要:
摘要:


【13】The Differentiable Auditory Loop (DAL): An ML Framework for Hyper-Personalized Hearing Aids

标题:可区分听觉环(DAL):超个性化助听器的ML框架
链接:https://arxiv.org/abs/2606.04103
作者:Alejandro Ballesta Rosen, Jason Mikiel-Hunter, Julian Maclaren, Jack Collins, Richard F. Lyon, Simon Carlile
摘要:
摘要:


【14】Channel-Oriented Design for EEG-to-Music Reconstruction

标题:面向队列的脑电到音乐重建设计
链接:https://arxiv.org/abs/2606.04040
作者:Jiaxin Qing, Junwei Lu, Lexin Li
摘要:
摘要:


【15】Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

标题:阅读您所听到的:无参考假设评估与声学差异
链接:https://arxiv.org/abs/2606.04680
作者:Zhihan Li, Hankun Wang, Yiwei Guo, Bohan Li, Xie Chen, Kai Yu
备注:Submitted to Interspeech 2026. 6 pages, 4 figures
摘要:
摘要:


【16】Masked Wavelet Scattering Transform Neural Field for Sound Field Reconstruction

标题:掩蔽子波散射变换神经场用于声学场重建
链接:https://arxiv.org/abs/2606.04370
作者:Xinmeng Luan, Samuel A. Verburg, Efren Fernandez-Grande, Gary Scavone
备注:5 pages, 2 figures, conference
摘要:
摘要:


【17】Representation Matters in Randomized Smoothing for Audio Classification

标题:音频分类随机平滑中的表示很重要
链接:https://arxiv.org/abs/2606.04210
作者:Jong-Ik Park, Shreyas Chaudhari, José M. F. Moura, Carlee Joe-Wong
摘要:
摘要:


eess.AS音频处理


【1】Differentiable Articulatory Copy-Synthesis of Biphonic Singing
标题:差异性的发音复制--双音歌唱的合成
链接:https://arxiv.org/abs/2606.04943
作者:Mateo Cámara, María Pilar Daza-Llin, Fernando Marcos-Macías, José Luis Blanco
备注:Accepted to DAFx 2026
摘要:
摘要:


【2】UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning

标题:UAT:用于音频生成、编辑和字幕的统一音频文本传播
链接:https://arxiv.org/abs/2606.04939
作者:Hui Wang, Yifan Yang, Zeyue Tian, Yuhang Jia, Jinghua Zhao, Long Zhou, Bing Han, Cheng Liu, Jiaming Zhou, Geng Tu, Yong Qin
摘要:
摘要:


【3】Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

标题:阅读您所听到的:无参考假设评估与声学差异
链接:https://arxiv.org/abs/2606.04680
作者:Zhihan Li, Hankun Wang, Yiwei Guo, Bohan Li, Xie Chen, Kai Yu
备注:Submitted to Interspeech 2026. 6 pages, 4 figures
摘要:
摘要:


【4】Masked Wavelet Scattering Transform Neural Field for Sound Field Reconstruction

标题:掩蔽子波散射变换神经场用于声学场重建
链接:https://arxiv.org/abs/2606.04370
作者:Xinmeng Luan, Samuel A. Verburg, Efren Fernandez-Grande, Gary Scavone
备注:5 pages, 2 figures, conference
摘要:
摘要:


【5】Representation Matters in Randomized Smoothing for Audio Classification

标题:音频分类随机平滑中的表示很重要
链接:https://arxiv.org/abs/2606.04210
作者:Jong-Ik Park, Shreyas Chaudhari, José M. F. Moura, Carlee Joe-Wong
摘要:
摘要:


【6】Audio Interaction Model

标题:音频交互模型
链接:https://arxiv.org/abs/2606.05121
作者:Zhifei Xie,Zihang Liu,Ze An,Xiaobin Hu,Yue Liao,Ziyang Ma,Dongchao Yang,Mingbao Lin,Deheng Ye,Shuicheng Yan,Chunyan Miao
备注:Next generation of LALMs, work in progress
摘要:音频是一种固有的交互方式,但今天的大型音频语言模型(LALM)是离线的,每个流式音频模型只处理单个任务,如流式ASR或语音聊天。现在是时候将它们统一到一个在线LALM中了:这个模型通过一个始终在线的感知-决定-响应循环,实时收听声音、环境和指令,并在飞行中做出反应。我们将这种制度形式化为音频交互模型,并使用音频交互实现它,音频交互是一个统一的流模型,它保留离线任务执行,同时添加在线通用音频指令,从对话到完整的语音聊天,决定何时从流的语义中做出响应。为了实现这一点,我们提出了SoundFlow,这是一个框架,它通过流媒体原生数据构建,理解感知训练和异步低延迟推理来实现从数据到训练再到部署的端到端的感知-决定-响应循环。我们进一步构建了StreamAudio-2 M,一个包含7个基本能力和28个子任务的2.6M条目的流媒体语料库,以及用于评估主动音频干预的Proactive-Sound-Bench。在8个基准测试中,Audio-Interaction在主流音频任务上保持了具有竞争力的性能,同时解锁了离线LALM无法访问的功能,包括实时ASR、流式音频指令跟踪和主动帮助。
摘要:Audio is an inherently interactive modality, yet today's Large Audio Language Models (LALMs) are offline, and streaming audio models each handle only a single task such as streaming ASR or voice chatting. It is time to unify them into one online LALM: a model that, through an always-on perceive-decide-respond loop, listens to sound, environment, and instructions in real time and reacts on the fly. We formalize this regime as the Audio Interaction Model, and realize it with Audio-Interaction, a unified streaming model that retains offline task execution while adding online general audio instruction following, from dialogue to full voice chatting, deciding when to respond from the semantics of the stream. To enable this, we propose SoundFlow, a framework that instantiates the perceive-decide-respond loop end to end, from data to training to deployment, through streaming-native data construction, comprehension-aware training, and asynchronous low-latency inference for stable real-time interaction. We further construct StreamAudio-2M, a 2.6M-item streaming corpus spanning 7 fundamental abilities and 28 sub-tasks, and Proactive-Sound-Bench for evaluating proactive audio intervention. Across 8 benchmarks, Audio-Interaction preserves competitive performance on mainstream audio tasks while unlocking capabilities inaccessible to offline LALMs, including real-time ASR, streaming audio instruction following, and proactive help.


【7】SURF: Separation via Unsupervised Remixing Flow

标题:SURF:通过无监督的重新混合流程进行分离
链接:https://arxiv.org/abs/2606.04921
作者:Henry Li,Robin Scheibler,Efthymios Tzinis,Matt Shannon,Arnaud Doucet,John R. Hershey
备注:Accepted at ICML 2026
摘要:单通道源分离的目标是重建$K$源给定他们的混合。在有监督的环境中,大量的干净的源数据是可用的,这个具有挑战性的,不适定的问题已经成功地解决了生成扩散和基于流的先验模型。然而,对这些干净源样本的访问通常是有限的,即使可用,监督模型也容易受到域偏移的影响。为了弥补这一差距,我们提出了通过无监督再混合流(SURF)进行分离,这是一种用于源分离的无监督流匹配方法,可直接从观察到的混合物中学习。该方法依赖于最先进的监督流匹配和基于回归的自监督技术的新组合。在高级别上,从教师模型开始,我们利用“重新混合”步骤来从教师的估计中引导学生流模型的学习。我们提供了通过这种方法优化的目标的见解,并绘制了一个新的连接到唤醒-睡眠算法。对图像和音频基准的实证评估表明,SURF建立了一个新的国家的最先进的,显着优于现有的无监督方法。查看我们的演示页面以获取示例。https://google.github.io/df-conformer/surf/
摘要:The goal of single-channel source separation is to reconstruct $K$ sources given their mixture. In supervised settings where vast amounts of clean source data are available, this challenging, ill-posed problem has been addressed successfully by generative diffusion and flow-based prior models. However, access to such clean source samples is often limited, and even when available, supervised models are vulnerable to domain shifts. To bridge this gap, we present Separation via Unsupervised Remixing Flow (SURF), an unsupervised flow matching approach for source separation that learns directly from observed mixtures. This method relies on a novel combination of state-of-the-art supervised flow matching and regression-based self-supervised techniques. At a high level, starting from a teacher model, we utilize a "remixing" step to bootstrap the learning of a student flow model from the teacher's estimates. We provide insights into the objectives optimized by this approach and draw a novel connection to the Wake-Sleep algorithm. Empirical evaluations on image and audio benchmarks demonstrate that SURF establishes a new state-of-the-art, significantly outperforming existing unsupervised methods. See our demo page for examples. https://google.github.io/df-conformer/surf/


【8】Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026

标题:多语言长篇演讲指导如下:KIT提交给IWIT2026
链接:https://arxiv.org/abs/2606.04730
作者:Enes Yavuz Ugan,Maike Züfle,Yuka Ko,Supriti Sinhamahapatra,Fabian Retkowski,Seymanur Akti,Jan Niehues,Alexander Waibel
备注:9 pages main paper, IWSLT 2026 Instruction Following track
摘要:随着大语言模型的出现,单任务和基于标记的多任务模型已经发展成为基于推理的系统,该系统从自然语言提示中隐式地推断任务和目标语言。这一趋势反映在IWITOS的指令跟踪中,今年引入了新的任务,包括未知的惊喜任务,对已知任务的过度拟合提出了真正的挑战。我们目前的KIT提交的长期和短期指令以下的轨道在不受约束的设置。我们的方法结合了一个通用的数据增强管道,通过片段拼接,基于LLM的标签生成和跨语言翻译将短格式语料库转换为长格式训练数据,在六个任务和四种语言中产生超过100万个实例。我们进一步表明,基于可能性的重新排序,而ASR非常有效,系统地降低语义任务,通过虚假地选择从分段音频处理而不是整体的长形式推理产生的候选人,通过结合可能性与最小贝叶斯风险解码解决的故障模式。
摘要:With the advent of Large Language Models, single-task and token-based multi-task models have evolved into instruction-based systems that infer task and target language implicitly from natural language prompts. This trend is reflected in IWSLT's Instruction Following Track, which this year introduced new tasks including an unknown surprise task, posing a genuine challenge against overfitting to known tasks. We present KIT's submission to the Long and Short Instruction Following tracks in the unconstrained setting. Our approach combines a general data augmentation pipeline that converts short-form corpora into long-form training data through segment concatenation, LLM-based label generation, and cross-lingual translation, yielding over 1M instances across six tasks and four languages. We further show that likelihood-based re-ranking, while highly effective for ASR, systematically degrades semantic tasks by spuriously selecting candidates generated from segmented audio processing rather than holistic long-form inference, a failure mode resolved by combining likelihood with Minimum Bayes Risk decoding.


【9】Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention

标题:言语LLM推理中的实体绑定失败:诊断和思想链干预
链接:https://arxiv.org/abs/2606.04474
作者:Ming-Hao Hsu,Xiaohai Tian,Jun Zhang,Zhizheng Wu
摘要:语音大语言模型(SLLM)在复杂推理方面表现不佳。我们发现,这种模态差距是不是一个统一的认知缺陷。评估三种不同的SLLM,我们显示语音到文本(S2 T)匹配或超过文本到文本(T2 T)的空间,句法和事实任务。然而,在需要实体跟踪的逻辑任务中,S2 T的准确性会下降到偶然性。我们将这种局部化的退化诊断为实体绑定失败:连续的语音特征会导致模型在隐式推理过程中失去精确的实体属性关联。为了解决这个问题,我们提出了一种智能感知思想链(EA-CoT),迫使SLLM在推理之前显式枚举实体并将其绑定到声明。引人注目的是,EA-CoT弥补了这一差距,即使在口语名称被错误识别的情况下,也能获得高达24.4%的绝对准确率提升。消融确认这些收益完全源于明确的语义绑定,重新定义的差距作为一个可解决的瓶颈。
摘要:Speech Large Language Models (SLLMs) underperform their text counterparts on complex reasoning. We reveal that this modality gap is not a uniform cognitive deficit. Evaluating three diverse SLLMs, we show speech-to-text (S2T) matches or exceeds text-to-text (T2T) on spatial, syntactic, and factual tasks. However, on logical tasks requiring entity tracking, S2T accuracy collapses to chance. We diagnose this localized degradation as an entity binding failure: continuous speech features cause models to lose precise entity-property associations during implicit reasoning. To resolve this, we propose Entity-Aware Chain-of-Thought (EA-CoT), forcing SLLMs to explicitly enumerate entities and bind them to claims before reasoning. Strikingly, EA-CoT bridges the gap, even when spoken names are misrecognized, yielding up to a 24.4% absolute accuracy improvement. Ablations confirm these gains stem entirely from explicit semantic binding, reframing the gap as a resolvable bottleneck.


【10】CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

标题:CleanCodec:通过感知引导编码实现高效和鲁棒的语音令牌化
链接:https://arxiv.org/abs/2606.04418
作者:Eugene Kwek,Feng Liu,Rui Zhang,Wenpeng Yin
摘要:神经音频编解码器是语音处理管道的关键组件,将音频压缩成离散令牌用于下游建模。然而,现有的编解码器难以平衡重建质量与令牌效率,通常以语言和声学上有意义的内容为代价对诸如背景噪声和记录伪像的感知无关信息进行编码。我们将音频标记化重新定义为选择性信息瓶颈问题,并提出CleanCodec,一种去噪音频编解码器,它学会只对感知重要的特征进行编码,并丢弃不可感知的信息。CleanCodec每秒仅需12.5个标记,就能实现最先进的标记化效率,在说话人相似性和语音清晰度方面大大优于现有的编解码器。对下游文本到语音和语音转换任务的评估进一步证明了性能的提高和高达17倍的推理速度,突出了显着的效率提升。
摘要:Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content. We reframe audio tokenization as a selective information bottleneck problem and propose CleanCodec, a denoising audio codec which learns to encode only perceptually important features and discard imperceptible information. At just 12.5 tokens per second, CleanCodec achieves state-of-the-art tokenization efficiency, substantially outperforming existing codecs in speaker similarity and speech intelligibility. Evaluations on downstream text-to-speech and voice conversion tasks further demonstrate improved performance and up to 17x faster inference, highlighting significant efficiency gains.


【11】Gauss Circle Lattices with Geometric Convolutions for Synthesizing High Dimensional Image-Source Room Impulse Responses

标题:具有几何卷积的高斯圆格用于合成多维像源室脉冲响应
链接:https://arxiv.org/abs/2606.04358
作者:Yuancheng Luo
备注:Accepted for publication at the 29th International Conference on Digital Audio Effects 2026
摘要:镜像声源模型(ISM)是一种广泛采用的方法,可以有效地模拟镜面反射假设下的房间声脉冲响应(RIR)。源和接收器之间的声学路径被追踪到从房间的边界平面上的连续反射计算出的格点。矩形房间的图像源的总数是多项式的RIR的持续时间或距离$k$等效,与度等于房间尺寸的数量$N$。因此,直接ISM模拟计算上限为O \left(k^N \right)$,并且仅考虑$N \leq 3$的情况以便于处理和实际应用。这项工作提出了一种替代的计算方法,通过将ISM格点计数减少到经典的高斯圆问题(GCP),将整数坐标和房间尺寸的渐近计算界降低到O \left(N k^2 \log k \right)。我们扩展的晶格计数模型的频率依赖性和反射加权图像源在更高的维度,通过卷积算子的连续维度之间的相关解决方案。两个结构实现RIR,随着时间频率控制,错误和运行时分析,和RIR统计。
摘要:The image-source model (ISM) is a widely adopted method for efficiently simulating acoustic room impulse responses (RIRs) under specular reflection assumptions. Acoustic paths between source and receiver are traced to lattice points computed from successive reflections over bounding planes of the room. Rectangular rooms bound the total number of image-sources to be polynomial in the RIR's duration or distance $k$ equivalent, with degree equal the number of room dimensions $N$. Direct ISM simulations are therefore compute upper-bound by $O \left ( k^N \right )$, and consider only cases of $N \leq 3$ for tractability and real-world applications. This work proposes an alternative computational method that lowers the asymptotic compute bound to $O \left ( N k^2 \log k \right )$ for integer coordinates and room dimensions via reducing ISM lattice point counting to the classic Gauss circle problem (GCP). We extend the lattice counting model to frequency-dependent and reflection weighted image-sources in higher dimensions, relating solutions between successive dimensions via the convolution operator. Two constructions for realizing RIRs are presented, along with time-frequency controls, error and run-time analysis, and RIR statistics.


【12】Feasibility of Time-Domain DNN-Based Speech Enhancement on Embedded FPGA for Hearing Aid

标题:嵌入式DSP助听器中基于时分DNN的语音增强的可行性
链接:https://arxiv.org/abs/2606.04221
作者:Feyisayo Olalere,Umut Altin,Kiki van der Heijden,Marcel van Gerven
备注:13 pages
摘要:助听器施加了严格的延迟和功率限制,当前基于DNN的语音增强系统在嵌入式硬件上难以满足这些限制。我们通过在AMD-Xilinx Kria KV 260上使用轻量级SuDoRM-RF++架构部署语音分离和去噪来表征这一差距,并在FP 32和16位定点精度下对每个任务进行评估。在这些配置中,第一样本延迟跟踪片上参数缓存,而不是算术吞吐量,将数据移动确定为主要瓶颈。精度降低一半的模型内存占用,而不影响客观的语音质量。定点去噪加速器实现了9.7 ms的第一样本延迟,满足10 ms的临床阈值,而语音分离达到16.0 ms。这些测量为基于嵌入式DNN的语音增强建立了具体的资源要求,并量化了助听器部署的剩余差距。
摘要:Hearing aids impose strict latency and power constraints that current DNN-based speech enhancement systems struggle to meet on embedded hardware. We characterize this gap by deploying both speech separation and denoising using the lightweight SuDoRM-RF++ architecture on the AMD-Xilinx Kria KV260, evaluated at FP32 and 16-bit fixed-point precision for each task. Across these configurations, first-sample latency tracks with on-chip parameter caching rather than arithmetic throughput, identifying data movement as the primary bottleneck. Precision reduction halves the model memory footprint without compromising objective speech quality. The fixed-point denoising accelerator achieves a first-sample latency of 9.7~ms, meeting the 10~ms clinical threshold, while speech separation reaches 16.0~ms. These measurements establish concrete resource requirements for embedded DNN-based speech enhancement and quantify the remaining gap to hearing aid deployment.


【13】The Differentiable Auditory Loop (DAL): An ML Framework for Hyper-Personalized Hearing Aids

标题:可区分听觉环(DAL):超个性化助听器的ML框架
链接:https://arxiv.org/abs/2606.04103
作者:Alejandro Ballesta Rosen,Jason Mikiel-Hunter,Julian Maclaren,Jack Collins,Richard F. Lyon,Simon Carlile
摘要:传统的助听器依赖于固定的、依赖于频率的放大和压缩来管理降低的灵敏度,这通常不能在复杂的环境中提供足够的听力支持,例如多个扬声器的情况(“鸡尾酒会”问题)。为了更全面地解决听力损失的潜在编码功能障碍,我们引入了可区分听觉环路(DAL),这是一种用于个性化助听器设计和验配的新开源框架。我们的第一个DAL实现结合了CARFAC,这是一种人类耳蜗功能的可微分模型,我们将其移植到JAX,以优化深度神经网络,将受损的听觉神经活动模式与正常听力参考相匹配。为了构建具有所需的细粒度频谱-时间信号处理的助听器,我们采用了SEANet,一种波形到波形完全卷积的UNet生成器。我们通过比较适合正常听力的CARFAC模型的输出与适合匹配每个受试者的个体听力障碍的CARFAC模型的输出来微调网络。使用从各自的CARFAC神经活动模式(NAP)输出和稳定的听觉图像(SAIs),后者提供了一个二维表示,捕捉听觉神经输出中的相位不敏感的时间结构的损失函数进行比较。通过梯度下降,SEANet模型学习对输入进行降噪,并补偿受损CARFAC模型模拟的听力损失。在神经表征和信号保真度指标方面,DAL优化的SEANet模型优于测试的主助听器(MHA)基线。DAL框架为基于模型的、机器学习驱动的助听器信号处理个性化提供了一条切实可行的途径。接下来的步骤包括硬件部署,以实现真实世界的临床测试。
摘要:Conventional hearing aids rely on fixed, frequency-dependent amplification and compression to manage reduced sensitivity, which often fails to provide sufficient listening support in complex environments, such as situations with multiple speakers (the ``cocktail party'' problem). To more comprehensively address the underlying encoding dysfunctions of hearing loss, we introduce the Differentiable Auditory Loop (DAL), a new open-source framework for personalized hearing aid design and fitting. Our first implementation of DAL incorporates CARFAC, a differentiable model of human cochlear function, which we ported to JAX, to optimize a deep neural network to match impaired auditory neural activity patterns with a normal-hearing reference. To build a hearing aid with the fine-grained spectro-temporal signal processing required, we adopt SEANet, a waveform-to-waveform fully convolutional UNet generator. We fine-tune the network by comparing the outputs of a CARFAC model fitted to normal hearing with that of a CARFAC model fitted to match each subject's individual hearing impairment. The comparison is done using loss functions derived from the respective CARFAC neural activity pattern (NAP) outputs and stabilized auditory images (SAIs), the latter providing a 2D representation that captures phase-insensitive temporal structure in the auditory nerve output. Through gradient descent, the SEANet model learns to both denoise the input and compensate for the hearing loss modelled by the impaired CARFAC model. Across neural-representation and signal-fidelity metrics, the DAL-optimized SEANet model outperformed the tested master hearing aid (MHA) baselines. The DAL framework provides a practical path toward model-based, machine-learning-driven personalization of hearing aid signal processing. Next steps include hardware deployment to enable real-world clinical testing.


【14】Channel-Oriented Design for EEG-to-Music Reconstruction

标题:面向队列的脑电到音乐重建设计
链接:https://arxiv.org/abs/2606.04040
作者:Jiaxin Qing,Junwei Lu,Lexin Li
摘要:脑机接口旨在从神经信号中解码自然刺激,但迄今为止,大多数进展都集中在视觉和语言上。在这篇文章中,我们研究了一个更具挑战性但探索较少的设置,EEG到音乐重建,其中信号很弱,分布,并且非常容易受到噪声和信道变化的影响。我们的中心发现是,早期通道混合破坏弱,但歧视性的EEG信号。为了解决这个问题,我们提出了一个面向渠道的设计,有三个关键组成部分。具体而言,通道标记化将每个电极视为显式标记以保留空间定位的神经证据,通道多视图自蒸馏强制跨时间作物和随机通道子集的一致性以学习鲁棒和分布式表示,通道数据增强引入结构化通道丢弃以提高对噪声,伪影和缺失电极的不变性。这些组件一起保存了跨通道的微弱但信息丰富的信号,并实现了与语义音乐表示空间的稳定对齐。我们将这种面向通道的设计集成在用于EEG到音乐重建的编码-对齐-解码管道中。从理论上讲,我们的特点时,保留通道级结构,导致改进的对齐。从经验上讲,我们与一系列最先进的基线进行了比较,并证明了一致和显着的性能增益。
摘要:Brain-computer interfaces aim to decode naturalistic stimuli from neural signals, yet most progress to date has focused on vision and language. In this article, we study a more challenging but far less explored setting, EEG-to-music reconstruction, where signals are weak, distributed, and highly susceptible to noise and channel variability. Our central finding is that early channel mixing destroys weak but discriminative EEG signals. To address this, we propose a channel-oriented design with three key components. Specifically, channel-wise tokenization treats each electrode as an explicit token to retain spatially localized neural evidence, channel-wise multi-view self-distillation enforces consistency across temporal crops and random channel subsets to learn robust and distributed representations, and channel-wise data augmentation introduces structured channel dropout to improve invariance to noise, artifacts, and missing electrodes. Together, these components preserve weak yet informative signals across channels and enable stable alignment to a semantic music representation space. We integrate this channel-oriented design within an encoding-alignment-decoding pipeline for EEG-to-music reconstruction. Theoretically, we characterize when preserving channel-level structure leads to improved alignment. Empirically, we compare with a range of state-of-the-art baselines and demonstrate consistent and significant performance gains.


机器翻译由腾讯交互翻译提供,仅供参考