今日论文合集:cs.SD语音13篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
标题:SALM:具有用于理解和编辑的结构化嵌入的空间音频语言模型
链接:https://arxiv.org/abs/2507.16724

作者:Jinbo Hu, Yin Cao, Ming Wu, Feiran Yang, Jun Yang
备注:5 pages, 1 figure
摘要:空间音频理解对于准确感知和解释声学环境至关重要。然而,现有的音频-语言模型难以处理空间音频和感知空间声学场景。我们介绍了空间音频语言模型(SALM),一个新的框架,桥梁空间音频和语言通过多模态对比学习。SALM由文本编码器和双分支音频编码器组成,通过结构化音频嵌入将空间声音分解为语义和空间分量。SALM的主要特征包括空间和文本表示的无缝对齐、空间和语义信息的单独和联合提取、zero-shot方向分类和对空间音频编辑的鲁棒支持。实验结果表明,SALM有效地捕获和对齐跨模态表示。此外,它还支持高级编辑功能,例如使用基于文本的嵌入来更改定向音频。
摘要:Spatial audio understanding is essential for accurately perceiving and interpreting acoustic environments. However, existing audio-language models struggle with processing spatial audio and perceiving spatial acoustic scenes. We introduce the Spatial Audio Language Model (SALM), a novel framework that bridges spatial audio and language via multi-modal contrastive learning. SALM consists of a text encoder and a dual-branch audio encoder, decomposing spatial sound into semantic and spatial components through structured audio embeddings. Key features of SALM include seamless alignment of spatial and text representations, separate and joint extraction of spatial and semantic information, zero-shot direction classification and robust support for spatial audio editing. Experimental results demonstrate that SALM effectively captures and aligns cross-modal representations. Furthermore, it supports advanced editing capabilities, such as altering directional audio using text-based embeddings.


【2】FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
标题:FISHER:多模式工业信号综合表示的基础模型
链接:http://arxiv.org/pdf/2507.16696v1

作者:Pingyi Fan, Anbai Jiang, Shuwei Zhang, Zhiqiang Lv, Bing Han, Xinhu Zheng, Wenrui Liang, Junjie Li, Wei-Qiang Zhang, Yanmin Qian, Xie Chen, Cheng Lu, Jia Liu
备注:11 pages, 6 figures
摘要:随着SCADA系统的快速部署,如何有效地分析工业信号并检测异常状态是行业的迫切需求。由于这些信号的显着异质性,我们总结为M5问题,以前的工作只专注于小的子问题,并采用专门的模型,未能利用模态之间的协同作用和强大的标度律。然而,我们认为,M5信号可以建模在一个统一的方式,由于内在的相似性。因此,我们提出了FISHER,一个多模态工业信号增强表示的基础模型。为了支持任意采样率,FISHER将采样率的增量视为子带信息的级联。具体而言,FISHER以STFT子带为建模单元,采用师生SSL框架进行预训练。我们还开发了RMIS基准,该基准评估了M5工业信号在多个健康管理任务上的表现。与顶级SSL型号相比,FISHER展示了多功能和出色的功能,一般性能增益高达5.03%,以及更有效的扩展曲线。我们还研究了下游任务的缩放律,并为未来的工作提供了潜在的途径。FISHER现已在https: github.com jianganbai FISHER上开源。
摘要:With the rapid deployment of SCADA systems, how to effectively analyze industrial signals and detect abnormal states is an urgent need for the industry. Due to the significant heterogeneity of these signals, which we summarize as the M5 problem, previous works only focus on small sub-problems and employ specialized models, failing to utilize the synergies between modalities and the powerful scaling law. However, we argue that the M5 signals can be modeled in a unified manner due to the intrinsic similarity. As a result, we propose FISHER, a Foundation model for multi-modal Industrial Signal compreHEnsive Representation. To support arbitrary sampling rates, FISHER considers the increment of sampling rate as the concatenation of sub-band information. Specifically, FISHER takes the STFT sub-band as the modeling unit and adopts a teacher student SSL framework for pre-training. We also develop the RMIS benchmark, which evaluates the representations of M5 industrial signals on multiple health management tasks. Compared with top SSL models, FISHER showcases versatile and outstanding capabilities with a general performance gain up to 5.03%, along with much more efficient scaling curves. We also investigate the scaling law on downstream tasks and derive potential avenues for future works. FISHER is now open-sourced on https: github.com jianganbai FISHER.


【3】Step-Audio 2 Technical Report

标题:Step-音频2技术报告
链接:https://arxiv.org/abs/2507.16632

作者:StepFun Audio Team
摘要:本文介绍了Step-Audio~2,一个端到端的多模态大型语言模型,专为工业级音频理解和语音会话而设计。通过集成潜在音频编码器和以推理为中心的强化学习(RL),Step-Audio 2在自动语音识别(ASR)和音频理解方面取得了令人满意的性能。为了促进真正的端到端语音对话,Step-Audio 2将离散音频令牌的生成纳入语言建模中,显着增强了其对说话风格和情感等非语言信息的响应能力。为了有效地利用现实世界数据中丰富的文本和声学知识,Step-Audio 2集成了检索增强生成(RAG),并能够调用外部工具,如Web搜索以减轻幻觉和音频搜索以切换音色。经过数百万小时的语音和音频数据训练,Step-Audio 2在不同的对话场景中提供智能和表现力。评估结果表明,与其他开源和商业解决方案相比,Step-Audio 2在各种音频理解和对话基准测试中实现了最先进的性能。请访问https://github.com/stepfun-ai/Step-Audio2了解更多信息。
摘要:This paper presents Step-Audio~2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent audio encoder and reasoning-centric reinforcement learning (RL), Step-Audio 2 achieves promising performance in automatic speech recognition (ASR) and audio understanding. To facilitate genuine end-to-end speech conversation, Step-Audio 2 incorporates the generation of discrete audio tokens into language modeling, significantly enhancing its responsiveness to paralinguistic information such as speaking styles and emotions. To effectively leverage the rich textual and acoustic knowledge in real-world data, Step-Audio 2 integrates retrieval-augmented generation (RAG) and is able to call external tools such as web search to mitigate hallucination and audio search to switch timbres. Trained on millions of hours of speech and audio data, Step-Audio 2 delivers intelligence and expressiveness across diverse conversational scenarios. Evaluation results demonstrate that Step-Audio 2 achieves state-of-the-art performance on various audio understanding and conversational benchmarks compared to other open-source and commercial solutions. Please visit https://github.com/stepfun-ai/Step-Audio2 for more information.


【4】TTMBA: Towards Text To Multiple Sources Binaural Audio Generation

标题:TTMBA:迈向文本到多源双耳音频生成
链接:https://arxiv.org/abs/2507.16564

作者:Yuxuan He, Xiaoran Yang, Ningning Pan, Gongping Huang
备注:5 pages,3 figures,2 tables
摘要:大多数现有的文本到音频(TTA)生成方法产生单声道输出,忽略了沉浸式听觉体验的基本空间信息。为了解决这个问题,我们提出了一种具有时间和空间控制的文本到多源双耳音频生成(TTMBA)的级联方法。首先,预训练的大型语言模型(LLM)将文本分割成结构化格式,其中包含每个声音事件的时间和空间细节。接下来,预训练的单声道音频生成网络为每个事件创建具有不同持续时间的多个单声道音频。使用基于来自LLM的空间数据的双耳渲染神经网络将这些单声道音频转换为双耳音频。最后,双耳音频按其开始时间排列,从而产生多声道双耳音频。实验结果表明,该方法在音频生成质量和空间感知精度方面的优越性。
摘要:Most existing text-to-audio (TTA) generation methods produce mono outputs, neglecting essential spatial information for immersive auditory experiences. To address this issue, we propose a cascaded method for text-to-multisource binaural audio generation (TTMBA) with both temporal and spatial control. First, a pretrained large language model (LLM) segments the text into a structured format with time and spatial details for each sound event. Next, a pretrained mono audio generation network creates multiple mono audios with varying durations for each event. These mono audios are transformed into binaural audios using a binaural rendering neural network based on spatial data from the LLM. Finally, the binaural audios are arranged by their start times, resulting in multisource binaural audio. Experimental results demonstrate the superiority of the proposed method in terms of both audio generation quality and spatial perceptual accuracy.


【5】Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries

标题:检测任何声音:使用多模式收件箱的开放词汇声音事件检测
链接:http://arxiv.org/pdf/2507.16343v1

作者:Pengfei Cai, Yan Song, Qing Gu, Nan Jiang, Haoyu Song, Ian McLoughlin
备注:Accepted by MM 2025
摘要:现有的大多数声音事件检测算法都是在闭集的假设下工作的,其检测能力局限于预定义的类别。虽然最近的努力已经探索了语言驱动的zero-shot SED利用音频语言模型,其性能仍然远远不能令人满意,由于缺乏细粒度的对齐和跨模态特征融合。在这项工作中,我们提出了检测任何声音模型(DASM),一个基于查询的框架开放词汇SED引导多模态查询。DASM制定SED作为一个帧级检索任务,其中音频功能匹配对来自文本或音频提示的查询向量。为了支持这一提法,DASM引入了一个双流解码器,明确地将事件识别和时间定位:跨模态事件解码器执行查询特征融合,并确定在剪辑级的声音事件的存在,而上下文网络模型的帧级本地化的时间依赖性。此外,提出了一种推理时间注意力掩蔽策略,以利用基本类和新类之间的语义关系,大大提高了对新类的泛化能力。在AudioSet Strong数据集上的实验表明,DASM有效地平衡了定位精度与对新类的泛化,在开放词汇设置(+ 7.8 PSDS)和闭集设置(+ 6.9 PSDS)中优于基于CLAP的方法。此外,在DESED上的交叉数据集zero-shot评估中,DASM的PSDS 1得分为42.2,甚至超过了监督CRNN基线。项目页面可在https: cai525.github.io Transformer4SED demo_page DASM 上找到。
摘要:Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting audio-language models, their performance is still far from satisfactory due to the lack of fine-grained alignment and cross-modal feature fusion. In this work, we propose the Detect Any Sound Model (DASM), a query-based framework for open-vocabulary SED guided by multi-modal queries. DASM formulates SED as a frame-level retrieval task, where audio features are matched against query vectors derived from text or audio prompts. To support this formulation, DASM introduces a dual-stream decoder that explicitly decouples event recognition and temporal localization: a cross-modality event decoder performs query-feature fusion and determines the presence of sound events at the clip-level, while a context network models temporal dependencies for frame-level localization. Additionally, an inference-time attention masking strategy is proposed to leverage semantic relations between base and novel classes, substantially enhancing generalization to novel classes. Experiments on the AudioSet Strong dataset demonstrate that DASM effectively balances localization accuracy with generalization to novel classes, outperforming CLAP-based methods in open-vocabulary setting (+ 7.8 PSDS) and the baseline in the closed-set setting (+ 6.9 PSDS). Furthermore, in cross-dataset zero-shot evaluation on DESED, DASM achieves a PSDS1 score of 42.2, even exceeding the supervised CRNN baseline. The project page is available at https: cai525.github.io Transformer4SED demo_page DASM .


【6】Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation

标题:通过丰富标记的合成声景增强进行稳健的生物声学检测

链接:http://arxiv.org/pdf/2507.16235v1

作者:Kaspar Soltero, Tadeu Siqueira, Stefanie Gutschmidt
备注:12 pages, 4 figures

摘要:被动声监测(PAM)分析通常会受到创建标记训练数据所需的大量手动工作的阻碍。这项研究引入了一个合成数据框架,从非常有限的源材料中生成大量标记丰富的训练数据,提高了生物声学检测模型的鲁棒性。我们的框架通过将干净的背景噪声与孤立的目标发声(小猫头鹰)相结合来合成逼真的音景,在合成过程中自动生成动态标签,如边界框。根据这些数据进行微调的模型可以很好地推广到现实世界的音景,即使在源发声的多样性大幅减少的情况下,性能仍然很高,这表明该模型在没有过度拟合的情况下学习了一般化的特征。这表明,合成数据生成是一种非常有效的策略,用于从小源数据集训练鲁棒的生物声学检测器。该方法大大减少了人工标记工作,克服了计算生物声学的关键瓶颈,提高了生态评估能力。

摘要:Passive Acoustic Monitoring (PAM) analysis is often hindered by the intensive manual effort needed to create labelled training data. This study introduces a synthetic data framework to generate large volumes of richly labelled training data from very limited source material, improving the robustness of bioacoustic detection models. Our framework synthesises realistic soundscapes by combining clean background noise with isolated target vocalisations (little owl), automatically generating dynamic labels like bounding boxes during synthesis. A model fine-tuned on this data generalised well to real-world soundscapes, with performance remaining high even when the diversity of source vocalisations was drastically reduced, indicating the model learned generalised features without overfitting. This demonstrates that synthetic data generation is a highly effective strategy for training robust bioacoustic detectors from small source datasets. The approach significantly reduces manual labelling effort, overcoming a key bottleneck in computational bioacoustics and enhancing ecological assessment capabilities.


【7】LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech

标题:LENS-DF:长形式有噪语音的Deepfake检测和时间定位

链接:http://arxiv.org/pdf/2507.16220v1

作者:Xuechen Liu, Wanying Ge, Xin Wang, Junichi Yamagishi

备注:Accepted by IEEE International Joint Conference on Biometrics (IJCB) 2025, Osaka, Japan

摘要:这项研究介绍了LENS-DF,这是一种新颖而全面的配方,用于在复杂和现实的音频条件下训练和评估音频深度假检测和时间定位。配方的生成部分以可控的方式输出来自输入数据集的具有几个关键特征的音频,例如较长的持续时间、嘈杂的条件和包含多个扬声器。相应的检测和定位协议使用模型。我们进行了基于自监督学习前端和简单后端的实验。结果表明,使用LENS-DF生成的数据训练的模型始终优于通过传统配方训练的模型,证明了LENS-DF在强大的音频深度伪造检测和定位方面的有效性和实用性。我们还对引入的变化进行消融研究,调查它们对该领域现实挑战的影响和相关性。

摘要:This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio conditions. The generation part of the recipe outputs audios from the input dataset with several critical characteristics, such as longer duration, noisy conditions, and containing multiple speakers, in a controllable fashion. The corresponding detection and localization protocol uses models. We conduct experiments based on self-supervised learning front-end and simple back-end. The results indicate that models trained using data generated with LENS-DF consistently outperform those trained via conventional recipes, demonstrating the effectiveness and usefulness of LENS-DF for robust audio deepfake detection and localization. We also conduct ablation studies on the variations introduced, investigating their impact on and relevance to realistic challenges in the field.


【8】LABNet: A Lightweight Attentive Beamforming Network for Ad-hoc Multichannel Microphone Invariant Real-Time Speech Enhancement

标题:LABNet:一种用于自组织多通道麦克风不变实时语音增强的轻量级专注性束形成网络
链接:https://arxiv.org/abs/2507.16190

作者:Haoyin Yan, Jie Zhang, Chengqian Jiang, Shuang Zhang
摘要:多通道语音增强(SE)的目的是恢复干净的语音噪声测量利用时空信号的功能。在ad-hoc阵列条件下,麦克风不变性(MI)要求系统处理不同的麦克风数量和阵列几何形状。从实践的角度来看,多通道录音不可避免地增加了边缘设备应用的计算负担,突出了轻量级和高效部署的必要性。在这项工作中,我们提出了一个轻量级的注意波束形成网络(LABNet)集成MI在一个低复杂度的实时SE系统。我们设计了一个三阶段的框架,有效的渠道内建模和渠道间的互动。开发了一个跨通道注意模块,有选择地聚合来自每个通道的特征。实验结果表明,我们的LABNet实现了令人印象深刻的性能与超轻的资源开销,同时保持MI,表明自组织阵列处理的巨大潜力。
摘要:Multichannel speech enhancement (SE) aims to restore clean speech from noisy measurements by leveraging spatiotemporal signal features. In ad-hoc array conditions, microphone invariance (MI) requires systems to handle different microphone numbers and array geometries. From a practical perspective, multichannel recordings inevitably increase the computational burden for edge-device applications, highlighting the necessity of lightweight and efficient deployments. In this work, we propose a lightweight attentive beamforming network (LABNet) to integrate MI in a low-complexity real-time SE system. We design a three-stage framework for efficient intra-channel modeling and inter-channel interaction. A cross-channel attention module is developed to aggregate features from each channel selectively. Experimental results demonstrate our LABNet achieves impressive performance with ultra-light resource overhead while maintaining the MI, indicating great potential for ad-hoc array processing.


【9】SDBench: A Comprehensive Benchmark Suite for Speaker Diarization

标题:SDBench:扬声器扩展的全面基准套件
链接:https://arxiv.org/abs/2507.16136

作者:Berkin Durmus ,  Blaise Munyampirwa ,  Eduardo Pacheco ,  Atila Orhon ,  Andrey Leonov
摘要:即使是最先进的说话人日记系统,在不同的数据集上也表现出很高的错误率差异,代表了许多用例和领域。此外,跨系统进行比较需要仔细应用最佳实践,例如数据集分割和指标定义,以实现苹果对苹果的比较。我们提出了SDBench(Speaker Diarization Benchmark),这是一个开源的基准测试套件,集成了13个不同的数据集,内置工具,用于对各种设备和服务器端系统的扬声器日志性能进行一致和细粒度的分析。SDBench能够随着时间的推移进行可重复的评估并轻松集成新系统。为了证明SDBench的功效,我们构建了SpeakerKit,这是一个基于Pyannote v3构建的以推理效率为中心的系统。SDBench能够快速执行消融研究,导致SpeakerKit比Pyannote v3快9.6倍,同时实现了相当的错误率。我们对6个最先进的系统进行了基准测试,包括Deepgram,AWS Transcribe和Pyannote AI API,揭示了准确性和速度之间的重要权衡。
摘要:Even state-of-the-art speaker diarization systems exhibit high variance in error rates across different datasets, representing numerous use cases and domains. Furthermore, comparing across systems requires careful application of best practices such as dataset splits and metric definitions to allow for apples-to-apples comparison. We propose SDBench (Speaker Diarization Benchmark), an open-source benchmark suite that integrates 13 diverse datasets with built-in tooling for consistent and fine-grained analysis of speaker diarization performance for various on-device and server-side systems. SDBench enables reproducible evaluation and easy integration of new systems over time. To demonstrate the efficacy of SDBench, we built SpeakerKit, an inference efficiency-focused system built on top of Pyannote v3. SDBench enabled rapid execution of ablation studies that led to SpeakerKit being 9.6x faster than Pyannote v3 while achieving comparable error rates. We benchmark 6 state-of-the-art systems including Deepgram, AWS Transcribe, and Pyannote AI API, revealing important trade-offs between accuracy and speed.


【10】A new XML conversion process for mensural music encoding : CMME_to_MEI (via Verovio)

链接:http://arxiv.org/pdf/2507.15991v1

作者:David Fiala (CESR), Laurent Pugin, Marnix van Berchum (KNAW), Martha Thomae (NOVA), Kévin Roger (CESR, UL, CRULH)
Journal-ref:Music Encoding Conference 2025, City St. George's, University of London, Jun 2025, Londres, United Kingdom

摘要:Ricercar实验室--图尔大学文艺复兴高级研究中心的音乐学研究团队--在法国数字基础设施Biblissima的支持下,决定开放获取,这是一个包含约3500个15世纪C的XML文件的大型语料库。音乐.该语料库由德国音乐学家Clemens Goldberg制作,他自2010年起编码了34个主要的15 th-c音乐内容。音乐手稿和其他补充文件,以便在他的基金会网站上提供Du Fay,Binchois,Okeghem,Busnoys和大多数主要同时代人的完整作品集的PDF文件,重点是他们的世俗输出。该语料库以名为CMME(计算机化定量音乐编辑)的XML格式编码,该格式是Theodor Dumitrescu在21世纪初专门为定量音乐设计的,以及自那时以来一直没有更新的编辑和出版工具。本文重点介绍为这些CMME文件开发一组转换工具,以满足更先进的音乐编码标准,即MEI。2024年9月,在巴黎孔多塞校区举办了一场研讨会,聚集了在定量乐谱、XML格式和编程方面拥有广泛知识的专家。直接在开源渲染库Verovio中开发了一个转换器,允许从CMME到MEI mensural的转换。之后转换为MEI CMN,可以将这些文件加载到常见的雕刻软件中,如MuseScore,信息损失最小。通过将CMME-XML直接导入Verovio,现有CMME文件的语料库获得了新的生命。此外,由于独立的CMME编辑器仍然可以正常工作,并且没有替代品可用于本地MEI,因此转换器为编码和编辑定量音乐提供了新的管道。
摘要:The Ricercar Lab - the musicological research team at the Center for advanced Studies in the Renaissance at the University of Tours - has decided to make available in open access, thanks to the support of the French digital infrastructure Biblissima, a large corpus of about 3500 XML files of 15th-c. music. This corpus was produced by the German musicologist Clemens Goldberg who encoded since 2010 onwards the musical content of 34 major 15th-c. music manuscripts and other complementary files, in order to offer on his foundation's website PDF files of complete collections of works by Du Fay, Binchois, Okeghem, Busnoys and most of their major contemporaries, focusing on their secular output. This corpus was encoded in an XML format named CMME (Computerized Mensural Music Editing), specifically conceived for mensural music by Theodor Dumitrescu in the 2000s, together with editorial and publication tools which have not been updated since then. This article focuses on the development of a set of conversion tools for these CMME files to meet more up-to-date standards of music encoding, namely MEI. A workshop was organised in September 2024 at the Campus Condorcet in Paris, gathering experts with a wide range of knowledge on mensural music notation, XML formats and programming. A converter was developped directly in the open-source rendering library Verovio, allowing the conversion from CMME to MEI mensural. A conversion to MEI CMN was implemented afterwards, enabling to load these files in common engraving softwares such as MuseScore with minimal loss of information. With the availability of a direct import of CMME-XML into Verovio, the corpus of existing CMME files gets a new life. Furthermore, since the stand-alone CMME editor still works fine and no alternative is available yet for native MEI, the converter offers a new pipeline for encoding and editing mensural music.


【11】Nonlinear Framework for Speech Bandwidth Extension

标题:语音带宽扩展的非线性框架
链接:https://arxiv.org/abs/2507.15970

作者:Tarikul Islam Tamiti, Nursad Mamun, Anomadarshi Barua
摘要:恢复因带宽限制而丢失的高频分量对于从电信到有限资源上的高保真音频等应用至关重要。我们引入了NDSI-BWE,这是一种新的对抗性带宽扩展(BWE)框架,它利用了四个受非线性动力学系统启发的新鉴别器来捕获不同的时间行为:多分辨率李雅普诺夫鉴别器(MRLD),用于通过捕获确定性混沌来确定对初始条件的灵敏度;多尺度递归鉴别器(MS-RD),用于自相似递归动力学;用于长范围慢变尺度不变关系的多尺度去趋势分形分析鉴别器(MSDFA),用于捕获隐藏的潜在空间关系的多分辨率庞加莱图鉴别器(MR-PPD),用于循环模式的多周期鉴别器(MPD),多分辨率幅度鉴别器(MRAD)和多分辨率相位鉴别器(MRPD),用于捕获复杂的幅度-相位过渡统计。通过在每个鉴别器中的卷积块的核心处使用深度卷积,NDSI-BWE实现了八倍的参数减少。这七个鉴别器引导基于复值ConformerNeXt的genetor,其具有基于双流Lattice-Net的架构,用于同时细化幅度和相位。生成器利用了基于Transformer的Conformer的全局依赖性建模和ConvNeXt块的局部时间建模能力。通过六个客观评估指标和五个人工评判的主观文本,NDSI-BWE在BWE中建立了一个新的SoTA。
摘要:Recovering high-frequency components lost to bandwidth constraints is crucial for applications ranging from telecommunications to high-fidelity audio on limited resources. We introduce NDSI-BWE, a new adversarial Band Width Extension (BWE) framework that leverage four new discriminators inspired by nonlinear dynamical system to capture diverse temporal behaviors: a Multi-Resolution Lyapunov Discriminator (MRLD) for determining sensitivity to initial conditions by capturing deterministic chaos, a Multi-Scale Recurrence Discriminator (MS-RD) for self-similar recurrence dynamics, a Multi-Scale Detrended Fractal Analysis Discriminator (MSDFA) for long range slow variant scale invariant relationship, a Multi-Resolution Poincar\'e Plot Discriminator (MR-PPD) for capturing hidden latent space relationship, a Multi-Period Discriminator (MPD) for cyclical patterns, a Multi-Resolution Amplitude Discriminator (MRAD) and Multi-Resolution Phase Discriminator (MRPD) for capturing intricate amplitude-phase transition statistics. By using depth-wise convolution at the core of the convolutional block with in each discriminators, NDSI-BWE attains an eight-times parameter reduction. These seven discriminators guide a complex-valued ConformerNeXt based genetor with a dual stream Lattice-Net based architecture for simultaneous refinement of magnitude and phase. The genertor leverage the transformer based conformer's global dependency modeling and ConvNeXt block's local temporal modeling capability. Across six objective evaluation metrics and subjective based texts comprises of five human judges, NDSI-BWE establishes a new SoTA in BWE.


【12】An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
标题:一种在大型语言模型(LLM)驱动的应用程序背景下测量自动语音识别(ASB)模型性能的方法
链接:https://arxiv.org/abs/2507.16456

作者:Sujith Pulikodan, Sahapthan K, Prasanta Kumar Ghosh, Visruth Sanka, Nihar Desai
备注:Accepted at INTERSPEECH 2025
摘要:自动语音识别(ASR)在人机交互中起着至关重要的作用,并作为广泛应用的接口。传统上,ASR性能是使用字错误率(WER)来评估的,字错误率是量化生成的transname中的插入、删除和替换的数量的度量。然而,随着越来越多地采用大型和强大的大型语言模型(LLM)作为各种应用程序中的核心处理组件,下游任务中不同类型的ASR错误的重要性值得进一步探索。在这项工作中,我们分析了LLM纠正ASR引入的错误的能力,并提出了一种新的措施来评估LLM驱动的应用程序的ASR性能。
摘要:Automatic Speech Recognition (ASR) plays a crucial role in human-machine interaction and serves as an interface for a wide range of applications. Traditionally, ASR performance has been evaluated using Word Error Rate (WER), a metric that quantifies the number of insertions, deletions, and substitutions in the generated transcriptions. However, with the increasing adoption of large and powerful Large Language Models (LLMs) as the core processing component in various applications, the significance of different types of ASR errors in downstream tasks warrants further exploration. In this work, we analyze the capabilities of LLMs to correct errors introduced by ASRs and propose a new measure to evaluate ASR performance for LLM-powered applications.


【13】Interpretable Embeddings of Speech Enhance and Explain Brain Encoding Performance of Audio Models
标题:语音的可解释嵌入增强和解释音频模型的大脑编码性能
链接:http://arxiv.org/pdf/2507.16080v1

作者:Riki Shimizu, Richard J. Antonello, Chandan Singh, Nima Mesgarani
备注:7pages, 4 figures
摘要:自监督语音模型(SSMs)越来越被誉为比基于传统手工特征的模型更强大的人类语音感知计算模型。然而,由于它们的表征本质上是黑箱的,因此仍然不清楚是什么驱动它们与大脑反应的一致性。为了解决这个问题,我们从六个可解释的特征家族中构建了线性编码模型:梅尔频谱图,Gabor滤波器组特征,语音存在,语音,句法和语义特征,以及来自三个最先进的SSM(Whisper,HuBERT,WavLM)的上下文嵌入,量化每个特征类捕获的共享和唯一的神经方差。与普遍的假设相反,我们的可解释模型预测皮层电图(ECoG)的反应,以语音更准确地比任何SSM。此外,用可解释的特征增强SSM表示产生了最好的整体神经预测,显著优于单独的任何一类。进一步的方差划分分析揭示了以前未解决的组件SSM表示,有助于他们的神经对齐:1。尽管通常假设SSM的后续层丢弃低水平声学信息,但这些模型压缩并优先保留对语音的神经编码至关重要的频带(100-1000 Hz)。2.与之前的说法相反,SSM编码与大脑相关的语义信息,这些信息不能被简化为较低级别的特征,并随着上下文长度和模型大小的增加而改善。这些结果突出了使用精炼的,可解释的功能在理解语音感知的重要性。
摘要:Self-supervised speech models (SSMs) are increasingly hailed as more powerful computational models of human speech perception than models based on traditional hand-crafted features. However, since their representations are inherently black-box, it remains unclear what drives their alignment with brain responses. To remedy this, we built linear encoding models from six interpretable feature families: mel-spectrogram, Gabor filter bank features, speech presence, phonetic, syntactic, and semantic Question-Answering features, and contextualized embeddings from three state-of-the-art SSMs (Whisper, HuBERT, WavLM), quantifying the shared and unique neural variance captured by each feature class. Contrary to prevailing assumptions, our interpretable model predicted electrocorticography (ECoG) responses to speech more accurately than any SSM. Moreover, augmenting SSM representations with interpretable features yielded the best overall neural predictions, significantly outperforming either class alone. Further variance-partitioning analyses revealed previously unresolved components of SSM representations that contribute to their neural alignment: 1. Despite the common assumption that later layers of SSMs discard low-level acoustic information, these models compress and preferentially retain frequency bands critical for neural encoding of speech (100-1000 Hz). 2. Contrary to previous claims, SSMs encode brain-relevant semantic information that cannot be reduced to lower-level features, improving with context length and model size. These results highlight the importance of using refined, interpretable features in understanding speech perception.


eess.AS音频处理


【1】An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
标题:一种在大型语言模型(LLM)驱动的应用程序背景下测量自动语音识别(ASB)模型性能的方法
链接:https://arxiv.org/abs/2507.16456

作者:Sujith Pulikodan, Sahapthan K, Prasanta Kumar Ghosh, Visruth Sanka, Nihar Desai
备注:Accepted at INTERSPEECH 2025
摘要:自动语音识别(ASR)在人机交互中起着至关重要的作用,并作为广泛应用的接口。传统上,ASR性能是使用字错误率(WER)来评估的,字错误率是量化生成的transname中的插入、删除和替换的数量的度量。然而,随着越来越多地采用大型和强大的大型语言模型(LLM)作为各种应用程序中的核心处理组件,下游任务中不同类型的ASR错误的重要性值得进一步探索。在这项工作中,我们分析了LLM纠正ASR引入的错误的能力,并提出了一种新的措施来评估LLM驱动的应用程序的ASR性能。
摘要:Automatic Speech Recognition (ASR) plays a crucial role in human-machine interaction and serves as an interface for a wide range of applications. Traditionally, ASR performance has been evaluated using Word Error Rate (WER), a metric that quantifies the number of insertions, deletions, and substitutions in the generated transcriptions. However, with the increasing adoption of large and powerful Large Language Models (LLMs) as the core processing component in various applications, the significance of different types of ASR errors in downstream tasks warrants further exploration. In this work, we analyze the capabilities of LLMs to correct errors introduced by ASRs and propose a new measure to evaluate ASR performance for LLM-powered applications.


【2】Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
标题:通过窗口交叉注意力的分布式同步设备语音增强
链接:https://arxiv.org/abs/2507.16104

作者:Gene-Ping Yang, Sebastian Braun
备注:None
摘要:配备麦克风的个人设备数量的增加提供了极大的灵活性和潜力,将其用作动态会议环境中的自组织麦克风阵列。然而,大多数现有方法都是针对时间同步的麦克风设置而设计的,这在现实世界的会议场景中可能不成立,其中时间延迟和时钟漂移在设备之间会有所不同。在这种情况下,我们发现transform-average-concatenate(TAC),一个流行的神经多麦克风处理模块,不足以处理时间异步麦克风。作为回应,我们提出了一个窗口交叉注意模块能够动态对齐所有麦克风之间的功能。该模块对麦克风的排列和数量都是不变的,并且可以很容易地集成到现有的模型中。此外,我们提出了一个最佳的训练目标,多说话人环境。我们评估了我们的方法在多麦克风嘈杂的混响设置未知的时间延迟和每个麦克风的时钟漂移。实验结果表明,我们的方法在iFaSNet和CRUSE模型上的性能优于TAC,提供了更快的收敛和更好的学习,证明了窗口交叉注意模块在异步麦克风设置中的有效性。
摘要:The increasing number of microphone-equipped personal devices offers great flexibility and potential using them as ad-hoc microphone arrays in dynamic meeting environments. However, most existing approaches are designed for time-synchronized microphone setups, a condition that may not hold in real-world meeting scenarios, where time latency and clock drift vary across devices. Under such conditions, we found transform-average-concatenate (TAC), a popular module for neural multi-microphone processing, insufficient in handling time-asynchronous microphones. In response, we propose a windowed cross-attention module capable of dynamically aligning features between all microphones. This module is invariant to both the permutation and the number of microphones and can be easily integrated into existing models. Furthermore, we propose an optimal training target for multi-talker environments. We evaluated our approach in a multi-microphone noisy reverberant setup with unknown time latency and clock drift of each microphone. Experimental results show that our method outperforms TAC on both iFaSNet and CRUSE models, offering faster convergence and improved learning, demonstrating the efficacy of the windowed cross-attention module for asynchronous microphone setups.


【3】SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
标题:SALM:具有用于理解和编辑的结构化嵌入的空间音频语言模型
链接:https://arxiv.org/abs/2507.16724

作者:Jinbo Hu, Yin Cao, Ming Wu, Feiran Yang, Jun Yang
备注:5 pages, 1 figure
摘要:空间音频理解对于准确感知和解释声学环境至关重要。然而,现有的音频-语言模型难以处理空间音频和感知空间声学场景。我们介绍了空间音频语言模型(SALM),一个新的框架,桥梁空间音频和语言通过多模态对比学习。SALM由文本编码器和双分支音频编码器组成,通过结构化音频嵌入将空间声音分解为语义和空间分量。SALM的主要特征包括空间和文本表示的无缝对齐、空间和语义信息的单独和联合提取、zero-shot方向分类和对空间音频编辑的鲁棒支持。实验结果表明,SALM有效地捕获和对齐跨模态表示。此外,它还支持高级编辑功能,例如使用基于文本的嵌入来更改定向音频。
摘要:Spatial audio understanding is essential for accurately perceiving and interpreting acoustic environments. However, existing audio-language models struggle with processing spatial audio and perceiving spatial acoustic scenes. We introduce the Spatial Audio Language Model (SALM), a novel framework that bridges spatial audio and language via multi-modal contrastive learning. SALM consists of a text encoder and a dual-branch audio encoder, decomposing spatial sound into semantic and spatial components through structured audio embeddings. Key features of SALM include seamless alignment of spatial and text representations, separate and joint extraction of spatial and semantic information, zero-shot direction classification and robust support for spatial audio editing. Experimental results demonstrate that SALM effectively captures and aligns cross-modal representations. Furthermore, it supports advanced editing capabilities, such as altering directional audio using text-based embeddings.


【4】Step-Audio 2 Technical Report
标题:Step-音频2技术报告
链接:https://arxiv.org/abs/2507.16632

作者:StepFun Audio Team
摘要:本文介绍了Step-Audio~2,一个端到端的多模态大型语言模型,专为工业级音频理解和语音会话而设计。通过集成潜在音频编码器和以推理为中心的强化学习(RL),Step-Audio 2在自动语音识别(ASR)和音频理解方面取得了令人满意的性能。为了促进真正的端到端语音对话,Step-Audio 2将离散音频令牌的生成纳入语言建模中,显着增强了其对说话风格和情感等非语言信息的响应能力。为了有效地利用现实世界数据中丰富的文本和声学知识,Step-Audio 2集成了检索增强生成(RAG),并能够调用外部工具,如Web搜索以减轻幻觉和音频搜索以切换音色。经过数百万小时的语音和音频数据训练,Step-Audio 2在不同的对话场景中提供智能和表现力。评估结果表明,与其他开源和商业解决方案相比,Step-Audio 2在各种音频理解和对话基准测试中实现了最先进的性能。请访问https://github.com/stepfun-ai/Step-Audio2了解更多信息。
摘要:This paper presents Step-Audio~2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent audio encoder and reasoning-centric reinforcement learning (RL), Step-Audio 2 achieves promising performance in automatic speech recognition (ASR) and audio understanding. To facilitate genuine end-to-end speech conversation, Step-Audio 2 incorporates the generation of discrete audio tokens into language modeling, significantly enhancing its responsiveness to paralinguistic information such as speaking styles and emotions. To effectively leverage the rich textual and acoustic knowledge in real-world data, Step-Audio 2 integrates retrieval-augmented generation (RAG) and is able to call external tools such as web search to mitigate hallucination and audio search to switch timbres. Trained on millions of hours of speech and audio data, Step-Audio 2 delivers intelligence and expressiveness across diverse conversational scenarios. Evaluation results demonstrate that Step-Audio 2 achieves state-of-the-art performance on various audio understanding and conversational benchmarks compared to other open-source and commercial solutions. Please visit https://github.com/stepfun-ai/Step-Audio2 for more information.


【5】TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
标题:TTMBA:迈向文本到多源双耳音频生成
链接:https://arxiv.org/abs/2507.16564

作者:Yuxuan He, Xiaoran Yang, Ningning Pan, Gongping Huang
备注:5 pages,3 figures,2 tables
摘要:大多数现有的文本到音频(TTA)生成方法产生单声道输出,忽略了沉浸式听觉体验的基本空间信息。为了解决这个问题,我们提出了一种具有时间和空间控制的文本到多源双耳音频生成(TTMBA)的级联方法。首先,预训练的大型语言模型(LLM)将文本分割成结构化格式,其中包含每个声音事件的时间和空间细节。接下来,预训练的单声道音频生成网络为每个事件创建具有不同持续时间的多个单声道音频。使用基于来自LLM的空间数据的双耳渲染神经网络将这些单声道音频转换为双耳音频。最后,双耳音频按其开始时间排列,从而产生多声道双耳音频。实验结果表明,该方法在音频生成质量和空间感知精度方面的优越性。
摘要:Most existing text-to-audio (TTA) generation methods produce mono outputs, neglecting essential spatial information for immersive auditory experiences. To address this issue, we propose a cascaded method for text-to-multisource binaural audio generation (TTMBA) with both temporal and spatial control. First, a pretrained large language model (LLM) segments the text into a structured format with time and spatial details for each sound event. Next, a pretrained mono audio generation network creates multiple mono audios with varying durations for each event. These mono audios are transformed into binaural audios using a binaural rendering neural network based on spatial data from the LLM. Finally, the binaural audios are arranged by their start times, resulting in multisource binaural audio. Experimental results demonstrate the superiority of the proposed method in terms of both audio generation quality and spatial perceptual accuracy.


【6】LABNet: A Lightweight Attentive Beamforming Network for Ad-hoc Multichannel Microphone Invariant Real-Time Speech Enhancement
标题:LABNet:一种用于自组织多通道麦克风不变实时语音增强的轻量级专注性束形成网络
链接:https://arxiv.org/abs/2507.16190

作者:Haoyin Yan, Jie Zhang, Chengqian Jiang, Shuang Zhang
摘要:多通道语音增强(SE)的目的是恢复干净的语音噪声测量利用时空信号的功能。在ad-hoc阵列条件下,麦克风不变性(MI)要求系统处理不同的麦克风数量和阵列几何形状。从实践的角度来看,多通道录音不可避免地增加了边缘设备应用的计算负担,突出了轻量级和高效部署的必要性。在这项工作中,我们提出了一个轻量级的注意波束形成网络(LABNet)集成MI在一个低复杂度的实时SE系统。我们设计了一个三阶段的框架,有效的渠道内建模和渠道间的互动。开发了一个跨通道注意模块,有选择地聚合来自每个通道的特征。实验结果表明,我们的LABNet实现了令人印象深刻的性能与超轻的资源开销,同时保持MI,表明自组织阵列处理的巨大潜力。
摘要:Multichannel speech enhancement (SE) aims to restore clean speech from noisy measurements by leveraging spatiotemporal signal features. In ad-hoc array conditions, microphone invariance (MI) requires systems to handle different microphone numbers and array geometries. From a practical perspective, multichannel recordings inevitably increase the computational burden for edge-device applications, highlighting the necessity of lightweight and efficient deployments. In this work, we propose a lightweight attentive beamforming network (LABNet) to integrate MI in a low-complexity real-time SE system. We design a three-stage framework for efficient intra-channel modeling and inter-channel interaction. A cross-channel attention module is developed to aggregate features from each channel selectively. Experimental results demonstrate our LABNet achieves impressive performance with ultra-light resource overhead while maintaining the MI, indicating great potential for ad-hoc array processing.


【7】SDBench: A Comprehensive Benchmark Suite for Speaker Diarization
标题:SDBench:扬声器扩展的全面基准套件
链接:https://arxiv.org/abs/2507.16136

作者:Berkin Durmus ,  Blaise Munyampirwa ,  Eduardo Pacheco ,  Atila Orhon ,  Andrey Leonov
摘要:即使是最先进的说话人日记系统,在不同的数据集上也表现出很高的错误率差异,代表了许多用例和领域。此外,跨系统进行比较需要仔细应用最佳实践,例如数据集分割和指标定义,以实现苹果对苹果的比较。我们提出了SDBench(Speaker Diarization Benchmark),这是一个开源的基准测试套件,集成了13个不同的数据集,内置工具,用于对各种设备和服务器端系统的扬声器日志性能进行一致和细粒度的分析。SDBench能够随着时间的推移进行可重复的评估并轻松集成新系统。为了证明SDBench的功效,我们构建了SpeakerKit,这是一个基于Pyannote v3构建的以推理效率为中心的系统。SDBench能够快速执行消融研究,导致SpeakerKit比Pyannote v3快9.6倍,同时实现了相当的错误率。我们对6个最先进的系统进行了基准测试,包括Deepgram,AWS Transcribe和Pyannote AI API,揭示了准确性和速度之间的重要权衡。
摘要:Even state-of-the-art speaker diarization systems exhibit high variance in error rates across different datasets, representing numerous use cases and domains. Furthermore, comparing across systems requires careful application of best practices such as dataset splits and metric definitions to allow for apples-to-apples comparison. We propose SDBench (Speaker Diarization Benchmark), an open-source benchmark suite that integrates 13 diverse datasets with built-in tooling for consistent and fine-grained analysis of speaker diarization performance for various on-device and server-side systems. SDBench enables reproducible evaluation and easy integration of new systems over time. To demonstrate the efficacy of SDBench, we built SpeakerKit, an inference efficiency-focused system built on top of Pyannote v3. SDBench enabled rapid execution of ablation studies that led to SpeakerKit being 9.6x faster than Pyannote v3 while achieving comparable error rates. We benchmark 6 state-of-the-art systems including Deepgram, AWS Transcribe, and Pyannote AI API, revealing important trade-offs between accuracy and speed.


【8】Nonlinear Framework for Speech Bandwidth Extension
标题:语音带宽扩展的非线性框架
链接:https://arxiv.org/abs/2507.15970

作者:Tarikul Islam Tamiti, Nursad Mamun, Anomadarshi Barua
摘要:恢复因带宽限制而丢失的高频分量对于从电信到有限资源上的高保真音频等应用至关重要。我们引入了NDSI-BWE,这是一种新的对抗性带宽扩展(BWE)框架,它利用了四个受非线性动力学系统启发的新鉴别器来捕获不同的时间行为:多分辨率李雅普诺夫鉴别器(MRLD),用于通过捕获确定性混沌来确定对初始条件的灵敏度;多尺度递归鉴别器(MS-RD),用于自相似递归动力学;用于长范围慢变尺度不变关系的多尺度去趋势分形分析鉴别器(MSDFA),用于捕获隐藏的潜在空间关系的多分辨率庞加莱图鉴别器(MR-PPD),用于循环模式的多周期鉴别器(MPD),多分辨率幅度鉴别器(MRAD)和多分辨率相位鉴别器(MRPD),用于捕获复杂的幅度-相位过渡统计。通过在每个鉴别器中的卷积块的核心处使用深度卷积,NDSI-BWE实现了八倍的参数减少。这七个鉴别器引导基于复值ConformerNeXt的genetor,其具有基于双流Lattice-Net的架构,用于同时细化幅度和相位。生成器利用了基于Transformer的Conformer的全局依赖性建模和ConvNeXt块的局部时间建模能力。通过六个客观评估指标和五个人工评判的主观文本,NDSI-BWE在BWE中建立了一个新的SoTA。
摘要:Recovering high-frequency components lost to bandwidth constraints is crucial for applications ranging from telecommunications to high-fidelity audio on limited resources. We introduce NDSI-BWE, a new adversarial Band Width Extension (BWE) framework that leverage four new discriminators inspired by nonlinear dynamical system to capture diverse temporal behaviors: a Multi-Resolution Lyapunov Discriminator (MRLD) for determining sensitivity to initial conditions by capturing deterministic chaos, a Multi-Scale Recurrence Discriminator (MS-RD) for self-similar recurrence dynamics, a Multi-Scale Detrended Fractal Analysis Discriminator (MSDFA) for long range slow variant scale invariant relationship, a Multi-Resolution Poincar\'e Plot Discriminator (MR-PPD) for capturing hidden latent space relationship, a Multi-Period Discriminator (MPD) for cyclical patterns, a Multi-Resolution Amplitude Discriminator (MRAD) and Multi-Resolution Phase Discriminator (MRPD) for capturing intricate amplitude-phase transition statistics. By using depth-wise convolution at the core of the convolutional block with in each discriminators, NDSI-BWE attains an eight-times parameter reduction. These seven discriminators guide a complex-valued ConformerNeXt based genetor with a dual stream Lattice-Net based architecture for simultaneous refinement of magnitude and phase. The genertor leverage the transformer based conformer's global dependency modeling and ConvNeXt block's local temporal modeling capability. Across six objective evaluation metrics and subjective based texts comprises of five human judges, NDSI-BWE establishes a new SoTA in BWE.


机器翻译由腾讯交互翻译提供,仅供参考