今日论文合集:cs.SD语音14篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
标题:魔兽长凳:通过海洋哺乳动物发声评估音频语言模型中的细粒度声学感知
链接:https://arxiv.org/abs/2508.20976

作者:Jaeyeon Kim, Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, Gunhee Kim
备注:Preprint. Project page: this https URL
摘要:大型音频语言模型(LALM)将语言理解扩展到听觉领域,但它们执行低级别听力(如音高和持续时间检测)的能力仍有待研究。然而,低水平的倾听对于现实世界的非分布任务至关重要,在这些任务中,模型必须基于细粒度的声学线索来推理不熟悉的声音。为了解决这一差距,我们引入了世界鲸鱼基准(WoW-Bench),以评估低层次的听觉感知和认知使用海洋哺乳动物发声。WoW-bench由一个用于对新声音进行分类的感知基准和一个受Bloom分类法启发的认知基准组成,用于评估记忆,理解,应用和分析声音事件的能力。对于认知基准测试,我们还引入了干扰因素问题,以评估模型是否真正通过倾听而不是依赖其他方法来解决问题。最先进的LALM的实验显示,其性能远低于人类水平,这表明LALM需要更强的听觉基础。
摘要:Large audio language models (LALMs) extend language understanding into the auditory domain, yet their ability to perform low-level listening, such as pitch and duration detection, remains underexplored. However, low-level listening is critical for real-world, out-of-distribution tasks where models must reason about unfamiliar sounds based on fine-grained acoustic cues. To address this gap, we introduce the World-of-Whale benchmark (WoW-Bench) to evaluate low-level auditory perception and cognition using marine mammal vocalizations. WoW-bench is composed of a Perception benchmark for categorizing novel sounds and a Cognition benchmark, inspired by Bloom's taxonomy, to assess the abilities to remember, understand, apply, and analyze sound events. For the Cognition benchmark, we additionally introduce distractor questions to evaluate whether models are truly solving problems through listening rather than relying on other heuristics. Experiments with state-of-the-art LALMs show performance far below human levels, indicating a need for stronger auditory grounding in LALMs.


【2】Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
标题:通过特征提取从双耳音频中学习稳健的空间表示
链接:https://arxiv.org/abs/2508.20914

作者:Holger Severin Bovbjerg (1), Jan Østergaard (1), Jesper Jensen (1, 2), Shinji Watanabe (3), Zheng-Hua Tan ((1) Aalborg University (2) Eriksholm Research Centre, (3) Carnegie Mellon University)
备注:To appear in Proc. WASPAA 2025, October 12-15, 2025, Tahoe, US. Copyright (c) 2025 IEEE. 5 pages, 2 figures, 2 tables
摘要:最近,深度表示学习在多个音频任务中表现出了强大的性能。然而,其用于从多声道音频学习空间表示的用途还未得到充分探索。我们研究了使用基于特征蒸馏的预训练阶段来学习双耳语音的鲁棒空间表示,而不需要数据标签。在这个框架中,从干净的双耳语音样本计算空间特征,以形成预测标签。然后使用神经网络从相应的增强语音预测这些干净的特征。在预训练之后,我们丢弃空间特征预测器,并使用学习的编码器权重来初始化DoA估计模型,我们对DoA估计进行微调。我们的实验表明,与完全监督模型和经典信号处理方法相比,预训练模型在对到达方向估计进行微调后,在噪声和混响环境中表现出更好的性能。
摘要:Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is underexplored. We investigate the use of a pretraining stage based on feature distillation to learn a robust spatial representation of binaural speech without the need for data labels. In this framework, spatial features are computed from clean binaural speech samples to form prediction labels. These clean features are then predicted from corresponding augmented speech using a neural network. After pretraining, we throw away the spatial feature predictor and use the learned encoder weights to initialize a DoA estimation model which we fine-tune for DoA estimation. Our experiments demonstrate that the pretrained models show improved performance in noisy and reverberant environments after fine-tuning for direction-of-arrival estimation, when compared to fully supervised models and classic signal processing methods.


【3】SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization
标题:SincQDR-VAR:利用可学习过滤器和排名感知优化的噪音稳健语音活动检测框架
链接:https://arxiv.org/abs/2508.20885

作者:Chien-Chun Wang, En-Lun Yu, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen
备注:Accepted to IEEE ASRU 2025
摘要:语音活动检测(VAD)对于语音驱动的应用程序至关重要,但在嘈杂和资源有限的环境中仍然远远不够完美。现有的方法往往缺乏对噪声的鲁棒性,并且它们的逐帧分类损失仅与VAD的评估度量松散耦合。为了解决这些挑战,我们提出了SincQDR-VAD,一个紧凑而强大的框架,它结合了Sinc-extractor前端与一个新的二次视差排名损失。Sinc-extractor使用可学习的带通滤波器来捕获抗噪声的频谱特征,而排名损失优化了语音和非语音帧之间的成对得分顺序,以改善接收器工作特性曲线(AUROC)下的面积。在代表性基准数据集上进行的一系列实验表明,我们的框架大大提高了AUROC和F2-Score,同时与现有技术相比仅使用了69%的参数,证实了其效率和实际可行性。
摘要:Voice activity detection (VAD) is essential for speech-driven applications, but remains far from perfect in noisy and resource-limited environments. Existing methods often lack robustness to noise, and their frame-wise classification losses are only loosely coupled with the evaluation metric of VAD. To address these challenges, we propose SincQDR-VAD, a compact and robust framework that combines a Sinc-extractor front-end with a novel quadratic disparity ranking loss. The Sinc-extractor uses learnable bandpass filters to capture noise-resistant spectral features, while the ranking loss optimizes the pairwise score order between speech and non-speech frames to improve the area under the receiver operating characteristic curve (AUROC). A series of experiments conducted on representative benchmark datasets show that our framework considerably improves both AUROC and F2-Score, while using only 69% of the parameters compared to prior arts, confirming its efficiency and practical viability.


【4】OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
标题:OLMoASB:用于训练稳健语音识别模型的开放模型和数据
链接:https://arxiv.org/abs/2508.20869

作者:Huong Ngo, Matt Deitke, Martijn Bartelds, Sarah Pratt, Josh Gardner, Matt Jordan, Ludwig Schmidt
备注:17 pages, 7 figures
摘要:训练数据规模和质量的改进带来了重大进步,但其在语音识别中的影响仍然没有得到充分研究。在本文中,我们提出了一个大规模的数据集,OLMoASR池,和一系列的模型,OLMoASR,研究和开发强大的zero-shot语音识别模型。从OLMoASR-Pool开始,收集了300万小时的英语音频和1700万份抄本,我们设计了文本启发式过滤器来删除低质量或错误的数据。我们的策展管道产生了一个新的数据集,其中包含100万小时的高质量音频-转录对,我们称之为OLMoASR-Mix。我们使用OLMoASR-Mix来训练OLMoASR-Mix模型套件,参数范围从39 M(tiny.en)到1.5 B(large.en)。在所有模型规模中,OLMoASR在短格式和长格式语音识别基准测试中的平均性能与OpenAI的Whisper相当。值得注意的是,OLMoASR-medium.en达到了12.8\%和11.0\%的单词错误率(WER),与Whisper最大的纯英语模型Whisper-medium.en的12.4\%和10.5\% WER分别用于短形式和长形式识别(在等效参数计数下)。OLMoASR池,OLMoASR模型,以及过滤,训练和评估代码将公开提供,以进一步研究鲁棒语音处理。
摘要:Improvements in training data scale and quality have led to significant advances, yet its influence in speech recognition remains underexplored. In this paper, we present a large-scale dataset, OLMoASR-Pool, and series of models, OLMoASR, to study and develop robust zero-shot speech recognition models. Beginning from OLMoASR-Pool, a collection of 3M hours of English audio and 17M transcripts, we design text heuristic filters to remove low-quality or mistranscribed data. Our curation pipeline produces a new dataset containing 1M hours of high-quality audio-transcript pairs, which we call OLMoASR-Mix. We use OLMoASR-Mix to train the OLMoASR-Mix suite of models, ranging from 39M (tiny.en) to 1.5B (large.en) parameters. Across all model scales, OLMoASR achieves comparable average performance to OpenAI's Whisper on short and long-form speech recognition benchmarks. Notably, OLMoASR-medium.en attains a 12.8\% and 11.0\% word error rate (WER) that is on par with Whisper's largest English-only model Whisper-medium.en's 12.4\% and 10.5\% WER for short and long-form recognition respectively (at equivalent parameter count). OLMoASR-Pool, OLMoASR models, and filtering, training and evaluation code will be made publicly available to further research on robust speech processing.


【5】Exploring Machine Learning and Language Models for Multimodal Depression Detection
标题:探索用于多模式抑郁检测的机器学习和语言模型
链接:https://arxiv.org/abs/2508.20805

作者:Javier Si Zhao Hong, Timothy Zoe Delaya, Sherwyn Chan Yin Kit, Pai Chet Ng, Xiaoxiao Miao
备注:This paper has been accepted by APCIPA ASC 2025
摘要:本文介绍了我们对第一个多模态个性感知抑郁症检测挑战的方法,重点是使用机器学习和深度学习模型进行多模态抑郁症检测。我们探索并比较了XGBoost、基于transformer的架构和大型语言模型(LLM)在音频、视频和文本特征上的性能。我们的研究结果突出了每种类型的模型在跨模态捕获抑郁相关信号方面的优势和局限性,为心理健康预测的有效多模态表示策略提供了见解。
摘要:This paper presents our approach to the first Multimodal Personality-Aware Depression Detection Challenge, focusing on multimodal depression detection using machine learning and deep learning models. We explore and compare the performance of XGBoost, transformer-based architectures, and large language models (LLMs) on audio, video, and text features. Our results highlight the strengths and limitations of each type of model in capturing depression-related signals across modalities, offering insights into effective multimodal representation strategies for mental health prediction.


【6】Speech Emotion Recognition via Entropy-Aware Score Selection
标题:通过熵感知分数选择的语音情感识别
链接:https://arxiv.org/abs/2508.20796

作者:ChenYi Chua, JunKai Wong, Chengxin Chen, Xiaoxiao Miao
备注:The paper has been accepted by APCIPA ASC 2025
摘要:在本文中,我们提出了一个多模态语音情感识别框架,利用熵感知得分选择相结合的语音和文本预测。所提出的方法集成了一个主要管道,该管道由基于wav2vec2.0的声学模型和一个次要管道组成,该管道由使用Roberta-XLM的情感分析模型组成,并通过Whisper-large-v3生成transmittance。我们提出了一种基于熵和变熵阈值的后期评分融合方法,以克服主要管道预测的置信度约束。情感映射策略将三个情感类别转换为四个目标情感类别,从而实现多模态预测的连贯集成。IEMOCAP和MSP-IMPROV数据集上的结果表明,该方法比传统的单模态系统提供了实用和可靠的增强。
摘要:In this paper, we propose a multimodal framework for speech emotion recognition that leverages entropy-aware score selection to combine speech and textual predictions. The proposed method integrates a primary pipeline that consists of an acoustic model based on wav2vec2.0 and a secondary pipeline that consists of a sentiment analysis model using RoBERTa-XLM, with transcriptions generated via Whisper-large-v3. We propose a late score fusion approach based on entropy and varentropy thresholds to overcome the confidence constraints of primary pipeline predictions. A sentiment mapping strategy translates three sentiment categories into four target emotion classes, enabling coherent integration of multimodal predictions. The results on the IEMOCAP and MSP-IMPROV datasets show that the proposed method offers a practical and reliable enhancement over traditional single-modality systems.


【7】Unified Multi-task Learning for Voice-Based Detection of Diverse Clinical Conditions
标题:统一多任务学习用于基于语音的各种临床状况检测
链接:https://arxiv.org/abs/2508.20717

作者: Ran Piao, Yuan Lu, Hareld Kemps, Tong Xia, Aaqib Saeed
摘要:基于语音的健康评估为可扩展的非侵入性疾病筛查提供了前所未有的机会,但现有方法通常专注于单一条件,无法利用语音中嵌入的丰富,多方面的信息。我们提出了MARVEL(基于语音的健康分析的多任务声学表示),这是一个具有隐私意识的多任务学习框架,可以同时检测九种不同的神经系统,呼吸系统和语音障碍,仅使用派生的声学特征,无需原始音频传输。我们的双分支架构采用专用编码器,特定任务的头共享一个共同的声学骨干,实现有效的跨条件知识转移。在大规模Bridge 2AI-Voice v2.0数据集上进行评估,MARVEL的总体AUROC为0.78,在神经系统疾病方面表现出色(AUROC = 0.89),特别是阿尔茨海默病/轻度认知障碍(AUROC = 0.97)。我们的框架始终优于单模态基线5-19%,并在9项任务中的7项上超过了最先进的自我监督模型,而相关性分析表明,学习的表示与已建立的声学特征表现出有意义的相似性,表明模型的内部表示与临床识别的声学模式一致。通过证明单个统一模型可以有效地筛选不同的条件,这项工作为资源受限和远程医疗环境中可部署的基于语音的诊断奠定了基础。
摘要:Voice-based health assessment offers unprecedented opportunities for scalable, non-invasive disease screening, yet existing approaches typically focus on single conditions and fail to leverage the rich, multi-faceted information embedded in speech. We present MARVEL (Multi-task Acoustic Representations for Voice-based Health Analysis), a privacy-conscious multitask learning framework that simultaneously detects nine distinct neurological, respiratory, and voice disorders using only derived acoustic features, eliminating the need for raw audio transmission. Our dual-branch architecture employs specialized encoders with task-specific heads sharing a common acoustic backbone, enabling effective cross-condition knowledge transfer. Evaluated on the large-scale Bridge2AI-Voice v2.0 dataset, MARVEL achieves an overall AUROC of 0.78, with exceptional performance on neurological disorders (AUROC = 0.89), particularly for Alzheimer's disease/mild cognitive impairment (AUROC = 0.97). Our framework consistently outperforms single-modal baselines by 5-19% and surpasses state-of-the-art self-supervised models on 7 of 9 tasks, while correlation analysis reveals that the learned representations exhibit meaningful similarities with established acoustic features, indicating that the model's internal representations are consistent with clinically recognized acoustic patterns. By demonstrating that a single unified model can effectively screen for diverse conditions, this work establishes a foundation for deployable voice-based diagnostics in resource-constrained and remote healthcare settings.


【8】Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music
标题:Amadeus:具有双向属性建模的符号音乐自回归模型
链接:https://arxiv.org/abs/2508.20665

作者:Hongju Su, Ke Li, Lan Yang, Honggang Zhang, Yi-Zhe Song
备注:Under review
摘要:现有的最先进的符号音乐生成模型主要采用自回归或分层自回归架构,将符号音乐建模为具有单向时间依赖性的属性令牌序列,假设这些属性之间具有固定的严格依赖结构。然而,我们观察到,在这些模型中使用不同的属性作为初始令牌会带来相当的性能。这表明音符的属性本质上是一个并发的无序集合,而不是一个时间依赖的序列。基于这一认识,我们介绍了一个新颖的符号音乐生成框架Amadeus。Amadeus采用两级架构:音符序列的自回归模型和属性的双向离散扩散模型。为了提高性能,我们提出了音乐潜在空间可辨别性增强策略(MLSDES),结合对比学习约束,放大中间音乐表示的可辨别性。条件信息增强模块(CIEM)同时通过注意力机制增强音符潜在向量表示,从而实现更精确的音符解码。我们进行了广泛的实验无条件和文本条件生成任务。Amadeus在多个指标上显著优于SOTA模型,同时实现了至少4倍的速度提升。此外,我们证明了训练免费,细粒度的笔记属性控制的可行性,使用我们的模型。为了探索Amadeus架构的性能上限,我们编译了迄今为止最大的开源符号音乐数据集AMD(Amadeus符号音乐数据集),支持预训练和微调。
摘要:Existing state-of-the-art symbolic music generation models predominantly adopt autoregressive or hierarchical autoregressive architectures, modelling symbolic music as a sequence of attribute tokens with unidirectional temporal dependencies, under the assumption of a fixed, strict dependency structure among these attributes. However, we observe that using different attributes as the initial token in these models leads to comparable performance. This suggests that the attributes of a musical note are, in essence, a concurrent and unordered set, rather than a temporally dependent sequence. Based on this insight, we introduce Amadeus, a novel symbolic music generation framework. Amadeus adopts a two-level architecture: an autoregressive model for note sequences and a bidirectional discrete diffusion model for attributes. To enhance performance, we propose Music Latent Space Discriminability Enhancement Strategy(MLSDES), incorporating contrastive learning constraints that amplify discriminability of intermediate music representations. The Conditional Information Enhancement Module (CIEM) simultaneously strengthens note latent vector representation via attention mechanisms, enabling more precise note decoding. We conduct extensive experiments on unconditional and text-conditioned generation tasks. Amadeus significantly outperforms SOTA models across multiple metrics while achieving at least 4$\times$ speed-up. Furthermore, we demonstrate training-free, fine-grained note attribute control feasibility using our model. To explore the upper performance bound of the Amadeus architecture, we compile the largest open-source symbolic music dataset to date, AMD (Amadeus MIDI Dataset), supporting both pre-training and fine-tuning.


【9】Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement
标题:具有条件流匹配的流直器实现准确的语音增强
链接:https://arxiv.org/abs/2508.20584

作者:Mattias Cross, Anton Ragni
备注:preprint, accepted
摘要:当前基于流的生成式语音增强方法学习曲线概率路径,该曲线概率路径对干净语音和有噪语音之间的映射进行建模。尽管有令人印象深刻的表现,弯曲的概率路径的含义是未知的。诸如薛定谔桥的方法专注于弯曲路径,其中时间依赖的梯度和方差不促进直线路径。机器学习研究的发现表明,直路径,如条件流匹配,更容易训练,并提供更好的泛化。在本文中,我们量化的路径直线度对语音增强质量的影响。我们报告的实验与薛定谔桥,在那里我们表明,某些配置导致更直的路径。相反,我们提出了独立的条件流匹配的语音增强,模型之间的噪声和干净的语音的直线路径。我们的经验表明,一个时间无关的方差比梯度对样本质量有更大的影响。虽然条件流匹配提高了几个语音质量指标,它需要多个推理步骤。我们通过推断训练的基于流的模型来纠正这一点,就好像它是直接预测的一样。我们的工作表明,直的时间无关的概率路径提高生成语音增强曲线依赖于时间的路径。
摘要:Current flow-based generative speech enhancement methods learn curved probability paths which model a mapping between clean and noisy speech. Despite impressive performance, the implications of curved probability paths are unknown. Methods such as Schrodinger bridges focus on curved paths, where time-dependent gradients and variance do not promote straight paths. Findings in machine learning research suggest that straight paths, such as conditional flow matching, are easier to train and offer better generalisation. In this paper we quantify the effect of path straightness on speech enhancement quality. We report experiments with the Schrodinger bridge, where we show that certain configurations lead to straighter paths. Conversely, we propose independent conditional flow-matching for speech enhancement, which models straight paths between noisy and clean speech. We demonstrate empirically that a time-independent variance has a greater effect on sample quality than the gradient. Although conditional flow matching improves several speech quality metrics, it requires multiple inference steps. We rectify this with a one-step solution by inferring the trained flow-based model as if it was directly predictive. Our work suggests that straighter time-independent probability paths improve generative speech enhancement over curved time-dependent paths.


【10】MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
标题:Mo塔斯:从TTS增强语音中MoE引导的特征选择,用于增强多模式阿尔茨海默氏症早期筛查
链接:https://arxiv.org/abs/2508.20513

作者:Yongqi Shao, Binxin Mei, Cong Tan, Hong Huo, Tao Fang
摘要:通过言语进行阿尔茨海默病(AD)的早期筛查是一种很有前途的非侵入性方法。然而,诸如有限的数据和缺乏细粒度的自适应特征选择等挑战通常会阻碍性能。为了解决这些问题,我们提出了MoTAS,这是一个旨在提高AD筛查效率的强大框架。MoTAS利用文本到语音(TTS)增强来增加数据量,并采用专家混合(MoE)机制来改进多模态特征选择,共同增强模型泛化。该过程从自动语音识别(ASR)开始,以获得准确的翻译。TTS然后用于合成语音,丰富了数据集。在提取声学和文本嵌入之后,MoE机制动态地选择信息量最大的特征,优化特征融合以改进分类。在ADReSSo数据集上进行评估,MoTAS达到了85.71%的领先准确率,优于现有的基线。消融研究进一步验证了TTS增强和MoE在提高分类性能方面的各自贡献。这些发现突出了MoTAS在现实世界AD筛查场景中的实用价值,特别是在数据有限的情况下。
摘要:Early screening for Alzheimer's Disease (AD) through speech presents a promising non-invasive approach. However, challenges such as limited data and the lack of fine-grained, adaptive feature selection often hinder performance. To address these issues, we propose MoTAS, a robust framework designed to enhance AD screening efficiency. MoTAS leverages Text-to-Speech (TTS) augmentation to increase data volume and employs a Mixture of Experts (MoE) mechanism to improve multimodal feature selection, jointly enhancing model generalization. The process begins with automatic speech recognition (ASR) to obtain accurate transcriptions. TTS is then used to synthesize speech that enriches the dataset. After extracting acoustic and text embeddings, the MoE mechanism dynamically selects the most informative features, optimizing feature fusion for improved classification. Evaluated on the ADReSSo dataset, MoTAS achieves a leading accuracy of 85.71\%, outperforming existing baselines. Ablation studies further validate the individual contributions of TTS augmentation and MoE in boosting classification performance. These findings highlight the practical value of MoTAS in real-world AD screening scenarios, particularly in data-limited settings.


【11】Automatic Inspection Based on Switch Sounds of Electric Point Machines
标题:基于电动转辙机开关声音的自动检测
链接:https://arxiv.org/abs/2508.20870

作者:Ayano Shibata, Toshiki Gunji, Mitsuaki Tsuda, Takashi Endo, Kota Dohi, Tomoya Nishida, Satoko Nomoto
备注:Accepted at ASPECT 2025
摘要:自2018年以来,东日本铁路公司和日立制作所一直致力于用基于物联网的监控取代人工检查。目的是节省设备检查所需的时间,并提供适当的预防性维护。作为目视检查的替代方案,很难替代电气特性监测,并且引入新的高性能传感器成本高昂。2019年,我们在“NS”电动转辙机中安装了摄像头和麦克风,以减少设备故障造成的停机时间,从而实现对锁片状况的远程监控。提出了一种基于声音信息的道岔转辙错误检测方法,并取得了预期的试验结果。所提出的方法将使实时检测设备故障成为可能,从而减少对目视检查的需求。本文介绍了我们的技术研究成果,旨在使用声音自动化检查电子转辙机,特别是从2019年开始关注“开关声音”。
摘要:Since 2018, East Japan Railway Company and Hitachi, Ltd. have been working to replace human inspections with IoT-based monitoring. The purpose is Labor-saving required for equipment inspections and provide appropriate preventive maintenance. As an alternative to visual inspection, it has been difficult to substitute electrical characteristic monitoring, and the introduction of new high-performance sensors has been costly. In 2019, we implemented cameras and microphones in an ``NS'' electric point machines to reduce downtime from equipment failures, allowing for remote monitoring of lock-piece conditions. This method for detecting turnout switching errors based on sound information was proposed, and the expected test results were obtained. The proposed method will make it possible to detect equipment failures in real time, thereby reducing the need for visual inspections. This paper presents the results of our technical studies aimed at automating the inspection of electronic point machines using sound, specifically focusing on ``switch sound'' beginning in 2019.


【12】CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
标题:CodecBench:声学和语义评估的综合基准
链接:https://arxiv.org/abs/2508.20660

作者:Ruifan Deng, Yitian Gong, Qinghui Gao, Luozhijie Jin, Qinyuan Cheng, Zhaoye Fei, Shimin Li, Xipeng Qiu
摘要:随着多模态大型语言模型(LLM)的兴起,音频编解码器在将音频编码为离散令牌方面发挥着越来越重要的作用,从而能够将音频集成到基于文本的LLM中。当前的音频编解码器捕获两种类型的信息:声学和语义。随着语音语言模型中音频编解码器应用场景的多样化,需要对越来越复杂的信息进行建模,并适应不同的语境,如多说话人、背景噪声或更丰富的语言信息等。然而,现有编解码器自身的评估受到过于简单的度量和场景的限制,并且现有的音频编解码器基准不是针对复杂的应用场景设计的,这限制了声学和语义能力在复杂数据集上的评估性能。我们引入CodecBench,这是一个全面的评估数据集,可以从四个数据域的声学和语义角度评估音频编解码器的性能。通过这个基准,我们的目标是确定当前的限制,突出未来的研究方向,并促进音频编解码器的发展。这些代码可在https://github.com/RayYuki/CodecBench上获得。
摘要:With the rise of multimodal large language models (LLMs), audio codec plays an increasingly vital role in encoding audio into discrete tokens, enabling integration of audio into text-based LLMs. Current audio codec captures two types of information: acoustic and semantic. As audio codec is applied to diverse scenarios in speech language model , it needs to model increasingly complex information and adapt to varied contexts, such as scenarios with multiple speakers, background noise, or richer paralinguistic information. However, existing codec's own evaluation has been limited by simplistic metrics and scenarios, and existing benchmarks for audio codec are not designed for complex application scenarios, which limits the assessment performance on complex datasets for acoustic and semantic capabilities. We introduce CodecBench, a comprehensive evaluation dataset to assess audio codec performance from both acoustic and semantic perspectives across four data domains. Through this benchmark, we aim to identify current limitations, highlight future research directions, and foster advances in the development of audio codec. The codes are available at https://github.com/RayYuki/CodecBench.


【13】Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
标题:通过多扬声器编码器统一数字化、分离和ASB
链接:https://arxiv.org/abs/2508.20474

作者:Muhammad Shakeel, Yui Sudo, Yifan Peng, Chyi-Jiunn Lin, Shinji Watanabe
备注:Accepted to IEEE ASRU 2025
摘要:本文提出了一种统一的多扬声器编码器(UME),一种新的架构,共同学习表示扬声器diarization(SD),语音分离(SS),和多扬声器自动语音识别(ASR)任务使用共享的语音基础编码器。我们利用UME多层的隐藏表示作为残差加权和编码(RWSE)来有效地使用来自不同语义级别的信息,从而有助于任务之间自下而上的对齐。这种联合训练方法捕获了任务之间的内在相互依赖性,提高了重叠语音数据的整体性能。我们的评估表明,UME在LibriMix评估集上大大改善了专用于SD,SS和多扬声器ASR的单任务基线。值得注意的是,对于SD,UME优于先前的研究,Libri2Mix和Libri3Mix评价集的日志错误率分别为1.37%和2.29%。
摘要:This paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a shared speech foundational encoder. We leverage the hidden representations from multiple layers of UME as a residual weighted-sum encoding (RWSE) to effectively use information from different semantic levels, contributing to bottom-up alignment between tasks. This joint training approach captures the inherent interdependencies among the tasks, enhancing overall performance on overlapping speech data. Our evaluations demonstrate that UME substantially improves over the single-task baselines dedicated to SD, SS, and multi-speaker ASR on LibriMix evaluation sets. Notably, for SD, UME outperforms the previous studies, achieving diarization error rates of 1.37% and 2.29% on Libri2Mix and Libri3Mix evaluation sets, respectively.


【14】Live Vocal Extraction from K-pop Performances
标题:从韩国流行音乐表演中提取现场声乐
链接:https://arxiv.org/abs/2508.20273

作者:Yujin Kim, Richa Namballa, Magdalena Fuentes
备注:2 pages + references, 1 figure, Extended Abstracts for the Late-Breaking Demo Session of the 26th International Society for Music Information Retrieval Conference
摘要:K-pop的全球成功得益于其充满活力的表演和充满活力的粉丝参与。受K-pop粉丝文化的启发,我们提出了一种从表演中自动提取现场人声的方法。我们使用源分离,互相关和幅度缩放的组合,自动删除预先录制的人声和乐器从现场表演。我们的初步工作介绍了现场人声分离的任务,并提供了一个基础,为今后的研究在这一主题。
摘要:K-pop's global success is fueled by its dynamic performances and vibrant fan engagement. Inspired by K-pop fan culture, we propose a methodology for automatically extracting live vocals from performances. We use a combination of source separation, cross-correlation, and amplitude scaling to automatically remove pre-recorded vocals and instrumentals from a live performance. Our preliminary work introduces the task of live vocal separation and provides a foundation for future research in this topic.


eess.AS音频处理


【1】Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System
标题:用于稳健音频深度伪造检测的多语言数据集集成策略:安全挑战系统
链接:https://arxiv.org/abs/2508.20983

作者:Hashim Ali, Surya Subramani, Lekha Bollinani, Nithin Sai Adupa, Sali El-Loh, Hafiz Malik
摘要:SAFE Challenge评估了三个任务的合成语音检测:未修改的音频,带有压缩伪影的处理音频,以及旨在逃避检测的清洗音频。我们系统地探索了自监督学习(SSL)前端、训练数据组成和音频长度配置,以实现强大的深度伪造检测。我们基于AASIST的方法将WavLM大型前端与RawBoost增强相结合,在包含256,600个样本的多语言数据集上进行训练,这些样本涵盖9种语言和来自CodecFake,MLAAD v5,SpoofCeleb,Famous Figures和MAILABS的70多个TTS系统。通过对不同SSL前端、三个训练数据版本和两种音频长度的广泛实验,我们在任务1(未修改音频检测)和任务3(清洗音频检测)中均获得第二名,证明了强大的泛化能力和鲁棒性。
摘要:The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning (SSL) front-ends, training data compositions, and audio length configurations for robust deepfake detection. Our AASIST-based approach incorporates WavLM large frontend with RawBoost augmentation, trained on a multilingual dataset of 256,600 samples spanning 9 languages and over 70 TTS systems from CodecFake, MLAAD v5, SpoofCeleb, Famous Figures, and MAILABS. Through extensive experimentation with different SSL front-ends, three training data versions, and two audio lengths, we achieved second place in both Task 1 (unmodified audio detection) and Task 3 (laundered audio detection), demonstrating strong generalization and robustness.


【2】Automatic Inspection Based on Switch Sounds of Electric Point Machines
标题:基于电动转辙机开关声音的自动检测
链接:https://arxiv.org/abs/2508.20870

作者:Ayano Shibata, Toshiki Gunji, Mitsuaki Tsuda, Takashi Endo, Kota Dohi, Tomoya Nishida, Satoko Nomoto
备注:Accepted at ASPECT 2025
摘要:自2018年以来,东日本铁路公司和日立制作所一直致力于用基于物联网的监控取代人工检查。目的是节省设备检查所需的时间,并提供适当的预防性维护。作为目视检查的替代方案,很难替代电气特性监测,并且引入新的高性能传感器成本高昂。2019年,我们在“NS”电动转辙机中安装了摄像头和麦克风,以减少设备故障造成的停机时间,从而实现对锁片状况的远程监控。提出了一种基于声音信息的道岔转辙错误检测方法,并取得了预期的试验结果。所提出的方法将使实时检测设备故障成为可能,从而减少对目视检查的需求。本文介绍了我们的技术研究成果,旨在使用声音自动化检查电子转辙机,特别是从2019年开始关注“开关声音”。
摘要:Since 2018, East Japan Railway Company and Hitachi, Ltd. have been working to replace human inspections with IoT-based monitoring. The purpose is Labor-saving required for equipment inspections and provide appropriate preventive maintenance. As an alternative to visual inspection, it has been difficult to substitute electrical characteristic monitoring, and the introduction of new high-performance sensors has been costly. In 2019, we implemented cameras and microphones in an ``NS'' electric point machines to reduce downtime from equipment failures, allowing for remote monitoring of lock-piece conditions. This method for detecting turnout switching errors based on sound information was proposed, and the expected test results were obtained. The proposed method will make it possible to detect equipment failures in real time, thereby reducing the need for visual inspections. This paper presents the results of our technical studies aimed at automating the inspection of electronic point machines using sound, specifically focusing on ``switch sound'' beginning in 2019.


【3】Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
标题:利用区分性潜在表示进行条件化基于GAN的语音增强
链接:https://arxiv.org/abs/2508.20859

作者:Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
备注:This manuscript has been submitted to IEEE Transactions on Audio, Speech and Language Processing
摘要:基于生成对抗网络(GANs)和扩散模型的生成语音增强方法在各种语音增强任务中显示出有希望的结果。然而,它们在非常低的信噪比(SNR)场景中的性能仍然没有得到充分的探索和限制,因为这些条件对判别和生成最先进的方法都构成了重大挑战。为了解决这个问题,我们提出了一种方法,该方法利用从区分性语音增强模型中提取的潜在特征作为通用条件特征来改进基于GAN的语音增强。所提出的方法被称为DisCoGAN,它展示了比基线模型更好的性能,特别是在低SNR情况下,同时在高SNR条件下和真实世界记录中也保持了竞争力或优越的性能。我们还对传统的基于GAN的架构进行了全面评估,包括经过端到端训练的GANs、作为第一处理阶段的GANs、滤波后的GANs以及低SNR条件下的判别模型。我们表明,DisCoGAN始终优于现有的方法。最后,我们提出了一个消融研究,探讨了DiscoGAN中的各个组件的贡献,并分析了歧视性条件反射方法对整体性能的影响。
摘要:Generative speech enhancement methods based on generative adversarial networks (GANs) and diffusion models have shown promising results in various speech enhancement tasks. However, their performance in very low signal-to-noise ratio (SNR) scenarios remains under-explored and limited, as these conditions pose significant challenges to both discriminative and generative state-of-the-art methods. To address this, we propose a method that leverages latent features extracted from discriminative speech enhancement models as generic conditioning features to improve GAN-based speech enhancement. The proposed method, referred to as DisCoGAN, demonstrates performance improvements over baseline models, particularly in low-SNR scenarios, while also maintaining competitive or superior performance in high-SNR conditions and on real-world recordings. We also conduct a comprehensive evaluation of conventional GAN-based architectures, including GANs trained end-to-end, GANs as a first processing stage, and post-filtering GANs, as well as discriminative models under low-SNR conditions. We show that DisCoGAN consistently outperforms existing methods. Finally, we present an ablation study that investigates the contributions of individual components within DisCoGAN and analyzes the impact of the discriminative conditioning method on overall performance.


【4】A Solution of Ultra Wideband Based High-resolution and Lossless Audio Transmission
标题:基于超宽带的高分辨率、无损音频传输解决方案
链接:https://arxiv.org/abs/2508.20782

作者:Fengyun Zhang
备注:5 pages
摘要:本文概述了当前无线音频传输面临的挑战,并强调了现有技术在数据带宽、数据压缩、延迟和设备间兼容性方面的局限性。为了解决这些缺点,它提出了一种高分辨率,无损音频传输方案,利用超宽带(UWB)技术。UWB提供了必要的带宽,以超低延迟实现卓越的音质,使其成为实时音频应用的理想选择,并解决了视听用例中的同步问题,从而成为一种有前途的解决方案。此外,UWB的独特功能不仅限于高分辨率音频,还允许在增强和虚拟现实应用中进行精确的位置跟踪。
摘要:This paper provides an overview of the current challenges in wireless audio transmission and highlights the limitations of existing technologies regarding data bandwidth, data compression, latency, and inter-device compatibility. To address these shortcomings, it proposes a high-resolution, lossless audio transmission scheme utilizing ultra wideband (UWB) technology. UWB emerges as a promising solution by offering the necessary bandwidth to enable exceptional sound quality with ultra-low latency, making it ideal for real-time audio applications and addressing synchronization concerns in audio-visual use cases. Additionally, UWB's unique capabilities extend beyond high-resolution audio, allowing for precise location tracking in augmented and virtual reality applications.


【5】Online incremental learning for audio classification using a pretrained audio model
标题:使用预训练的音频模型进行音频分类的在线增量学习
链接:https://arxiv.org/abs/2508.20732

作者: Manjunath Mulimani, Annamaria Mesaros
备注:Accepted to IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:增量学习的目的是在不忘记以前学习过的任务的情况下顺序地学习新任务。大多数现有的音频增量学习方法都专注于在初始任务上从头开始训练模型,并且使用相同的模型来学习即将到来的增量任务。该模型经过多次迭代训练,以适应每个新任务,并使用一些特定的方法来减少对旧任务的遗忘。在这项工作中,我们提出了一种方法,用于使用由预训练模型产生的可概括的音频嵌入来开发一个在线增量学习器,该学习器可以随着时间的推移解决连续的音频分类任务。具体来说,我们在预训练模型的音频嵌入和分类器之间注入了一个具有非线性激活函数的层;该层扩展了嵌入的维度,并有效地捕获了声音类的独特特征。我们的方法通过任何任务的训练样本在单个前向传递(在线)中调整模型,同时最大限度地减少对旧任务的遗忘。我们在两个增量学习设置中展示了所提出的方法的性能:一个是使用ESC-50的类增量学习,另一个是来自TAU Urban Acoustic Scenes 2019数据集的不同城市的域增量学习;对于这两种情况,所提出的方法都优于其他方法。
摘要:Incremental learning aims to learn new tasks sequentially without forgetting the previously learned ones. Most of the existing incremental learning methods for audio focus on training the model from scratch on the initial task, and the same model is used to learn upcoming incremental tasks. The model is trained for several iterations to adapt to each new task, using some specific approaches to reduce the forgetting of old tasks. In this work, we propose a method for using generalizable audio embeddings produced by a pre-trained model to develop an online incremental learner that solves sequential audio classification tasks over time. Specifically, we inject a layer with a nonlinear activation function between the pre-trained model's audio embeddings and the classifier; this layer expands the dimensionality of the embeddings and effectively captures the distinct characteristics of sound classes. Our method adapts the model in a single forward pass (online) through the training samples of any task, with minimal forgetting of old tasks. We demonstrate the performance of the proposed method in two incremental learning setups: one class-incremental learning using ESC-50 and one domain-incremental learning of different cities from the TAU Urban Acoustic Scenes 2019 dataset; for both cases, the proposed approach outperforms other methods.


【6】Sound event detection with audio-text models and heterogeneous temporal annotations
标题:使用音频文本模型和异类时间注释的声音事件检测
链接:https://arxiv.org/abs/2508.20703

作者:Manu Harju, Annamaria Mesaros
备注:Accepted to IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:基于音频和相关元数据生成合成字幕的最新进展允许使用自然语言中包含的信息作为其他音频任务的输入。在本文中,我们提出了一种新的方法来指导一个声音事件检测系统与自由形式的文本。我们使用机器生成的标题作为强标签的补充信息进行训练,并使用不同类型的文本输入来评估系统。此外,我们还研究了一种情况,即只有部分训练数据具有强标签,其余部分只有时间上的弱标签。我们的研究结果表明,合成字幕提高了性能,在这两种情况下相比,通常用于声音事件检测的CRNN架构。在50个高度不平衡类的数据集上,当使用强标签训练时,PSDS-1得分从0.223增加到0.277,当一半的训练数据只有弱标签时,PSDS-1得分从0.166增加到0.218。
摘要:Recent advances in generating synthetic captions based on audio and related metadata allow using the information contained in natural language as input for other audio tasks. In this paper, we propose a novel method to guide a sound event detection system with free-form text. We use machine-generated captions as complementary information to the strong labels for training, and evaluate the systems using different types of textual inputs. In addition, we study a scenario where only part of the training data has strong labels, and the rest of it only has temporally weak labels. Our findings show that synthetic captions improve the performance in both cases compared to the CRNN architecture typically used for sound event detection. On a dataset of 50 highly unbalanced classes, the PSDS-1 score increases from 0.223 to 0.277 when trained with strong labels, and from 0.166 to 0.218 when half of the training data has only weak labels.


【7】CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
标题:CodecBench:声学和语义评估的综合基准
链接:https://arxiv.org/abs/2508.20660

作者:Ruifan Deng, Yitian Gong, Qinghui Gao, Luozhijie Jin, Qinyuan Cheng, Zhaoye Fei, Shimin Li, Xipeng Qiu
摘要:随着多模态大型语言模型(LLM)的兴起,音频编解码器在将音频编码为离散令牌方面发挥着越来越重要的作用,从而能够将音频集成到基于文本的LLM中。当前的音频编解码器捕获两种类型的信息:声学和语义。随着语音语言模型中音频编解码器应用场景的多样化,需要对越来越复杂的信息进行建模,并适应不同的语境,如多说话人、背景噪声或更丰富的语言信息等。然而,现有编解码器自身的评估受到过于简单的度量和场景的限制,并且现有的音频编解码器基准不是针对复杂的应用场景设计的,这限制了声学和语义能力在复杂数据集上的评估性能。我们引入CodecBench,这是一个全面的评估数据集,可以从四个数据域的声学和语义角度评估音频编解码器的性能。通过这个基准,我们的目标是确定当前的限制,突出未来的研究方向,并促进音频编解码器的发展。这些代码可在https://github.com/RayYuki/CodecBench上获得。
摘要:With the rise of multimodal large language models (LLMs), audio codec plays an increasingly vital role in encoding audio into discrete tokens, enabling integration of audio into text-based LLMs. Current audio codec captures two types of information: acoustic and semantic. As audio codec is applied to diverse scenarios in speech language model , it needs to model increasingly complex information and adapt to varied contexts, such as scenarios with multiple speakers, background noise, or richer paralinguistic information. However, existing codec's own evaluation has been limited by simplistic metrics and scenarios, and existing benchmarks for audio codec are not designed for complex application scenarios, which limits the assessment performance on complex datasets for acoustic and semantic capabilities. We introduce CodecBench, a comprehensive evaluation dataset to assess audio codec performance from both acoustic and semantic perspectives across four data domains. Through this benchmark, we aim to identify current limitations, highlight future research directions, and foster advances in the development of audio codec. The codes are available at https://github.com/RayYuki/CodecBench.


【8】Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
标题:通过多扬声器编码器统一数字化、分离和ASB
链接:https://arxiv.org/abs/2508.20474

作者:Muhammad Shakeel, Yui Sudo, Yifan Peng, Chyi-Jiunn Lin, Shinji Watanabe
备注:Accepted to IEEE ASRU 2025
摘要:本文提出了一种统一的多扬声器编码器(UME),一种新的架构,共同学习表示扬声器diarization(SD),语音分离(SS),和多扬声器自动语音识别(ASR)任务使用共享的语音基础编码器。我们利用UME多层的隐藏表示作为残差加权和编码(RWSE)来有效地使用来自不同语义级别的信息,从而有助于任务之间自下而上的对齐。这种联合训练方法捕获了任务之间的内在相互依赖性,提高了重叠语音数据的整体性能。我们的评估表明,UME在LibriMix评估集上大大改善了专用于SD,SS和多扬声器ASR的单任务基线。值得注意的是,对于SD,UME优于先前的研究,Libri2Mix和Libri3Mix评价集的日志错误率分别为1.37%和2.29%。
摘要:This paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a shared speech foundational encoder. We leverage the hidden representations from multiple layers of UME as a residual weighted-sum encoding (RWSE) to effectively use information from different semantic levels, contributing to bottom-up alignment between tasks. This joint training approach captures the inherent interdependencies among the tasks, enhancing overall performance on overlapping speech data. Our evaluations demonstrate that UME substantially improves over the single-task baselines dedicated to SD, SS, and multi-speaker ASR on LibriMix evaluation sets. Notably, for SD, UME outperforms the previous studies, achieving diarization error rates of 1.37% and 2.29% on Libri2Mix and Libri3Mix evaluation sets, respectively.


【9】Live Vocal Extraction from K-pop Performances
标题:从韩国流行音乐表演中提取现场声乐
链接:https://arxiv.org/abs/2508.20273

作者:Yujin Kim, Richa Namballa, Magdalena Fuentes
备注:2 pages + references, 1 figure, Extended Abstracts for the Late-Breaking Demo Session of the 26th International Society for Music Information Retrieval Conference
摘要:K-pop的全球成功得益于其充满活力的表演和充满活力的粉丝参与。受K-pop粉丝文化的启发,我们提出了一种从表演中自动提取现场人声的方法。我们使用源分离,互相关和幅度缩放的组合,自动删除预先录制的人声和乐器从现场表演。我们的初步工作介绍了现场人声分离的任务,并提供了一个基础,为今后的研究在这一主题。
摘要:K-pop's global success is fueled by its dynamic performances and vibrant fan engagement. Inspired by K-pop fan culture, we propose a methodology for automatically extracting live vocals from performances. We use a combination of source separation, cross-correlation, and amplitude scaling to automatically remove pre-recorded vocals and instrumentals from a live performance. Our preliminary work introduces the task of live vocal separation and provides a foundation for future research in this topic.


【10】WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
标题:魔兽长凳:通过海洋哺乳动物发声评估音频语言模型中的细粒度声学感知
链接:https://arxiv.org/abs/2508.20976

作者:Jaeyeon Kim,  Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, Gunhee Kim
备注:Preprint. Project page: this https URL
摘要:大型音频语言模型(LALM)将语言理解扩展到听觉领域,但它们执行低级别听力(如音高和持续时间检测)的能力仍有待研究。然而,低水平的倾听对于现实世界的非分布任务至关重要,在这些任务中,模型必须基于细粒度的声学线索来推理不熟悉的声音。为了解决这一差距,我们引入了世界鲸鱼基准(WoW-Bench),以评估低层次的听觉感知和认知使用海洋哺乳动物发声。WoW-bench由一个用于对新声音进行分类的感知基准和一个受Bloom分类法启发的认知基准组成,用于评估记忆,理解,应用和分析声音事件的能力。对于认知基准测试,我们还引入了干扰因素问题,以评估模型是否真正通过倾听而不是依赖其他方法来解决问题。最先进的LALM的实验显示,其性能远低于人类水平,这表明LALM需要更强的听觉基础。
摘要:Large audio language models (LALMs) extend language understanding into the auditory domain, yet their ability to perform low-level listening, such as pitch and duration detection, remains underexplored. However, low-level listening is critical for real-world, out-of-distribution tasks where models must reason about unfamiliar sounds based on fine-grained acoustic cues. To address this gap, we introduce the World-of-Whale benchmark (WoW-Bench) to evaluate low-level auditory perception and cognition using marine mammal vocalizations. WoW-bench is composed of a Perception benchmark for categorizing novel sounds and a Cognition benchmark, inspired by Bloom's taxonomy, to assess the abilities to remember, understand, apply, and analyze sound events. For the Cognition benchmark, we additionally introduce distractor questions to evaluate whether models are truly solving problems through listening rather than relying on other heuristics. Experiments with state-of-the-art LALMs show performance far below human levels, indicating a need for stronger auditory grounding in LALMs.


【11】Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
标题:通过特征提取从双耳音频中学习稳健的空间表示
链接:https://arxiv.org/abs/2508.20914

作者:Holger Severin Bovbjerg (1), Jan Østergaard (1), Jesper Jensen (1, 2), Shinji Watanabe (3), Zheng-Hua Tan ((1) Aalborg University (2) Eriksholm Research Centre, (3) Carnegie Mellon University)
备注:To appear in Proc. WASPAA 2025, October 12-15, 2025, Tahoe, US. Copyright (c) 2025 IEEE. 5 pages, 2 figures, 2 tables
摘要:最近,深度表示学习在多个音频任务中表现出了强大的性能。然而,其用于从多声道音频学习空间表示的用途还未得到充分探索。我们研究了使用基于特征蒸馏的预训练阶段来学习双耳语音的鲁棒空间表示,而不需要数据标签。在这个框架中,空间特征是从干净的双耳语音样本中计算出来的,以形成预测标签。然后使用神经网络从相应的增强语音预测这些干净的特征。在预训练之后,我们丢弃空间特征预测器,并使用学习的编码器权重来初始化DoA估计模型,我们对DoA估计进行微调。我们的实验表明,与完全监督模型和经典信号处理方法相比,预训练模型在对到达方向估计进行微调后,在噪声和混响环境中表现出更好的性能。
摘要:Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is underexplored. We investigate the use of a pretraining stage based on feature distillation to learn a robust spatial representation of binaural speech without the need for data labels. In this framework, spatial features are computed from clean binaural speech samples to form prediction labels. These clean features are then predicted from corresponding augmented speech using a neural network. After pretraining, we throw away the spatial feature predictor and use the learned encoder weights to initialize a DoA estimation model which we fine-tune for DoA estimation. Our experiments demonstrate that the pretrained models show improved performance in noisy and reverberant environments after fine-tuning for direction-of-arrival estimation, when compared to fully supervised models and classic signal processing methods.


【12】OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
标题:OLMoASB:用于训练稳健语音识别模型的开放模型和数据
链接:https://arxiv.org/abs/2508.20869

作者:Huong Ngo, Matt Deitke, Martijn Bartelds, Sarah Pratt, Josh Gardner, Matt Jordan, Ludwig Schmidt
备注:17 pages, 7 figures
摘要:训练数据规模和质量的改进带来了重大进步,但其在语音识别中的影响仍然没有得到充分研究。在本文中,我们提出了一个大规模的数据集,OLMoASR池,和一系列的模型,OLMoASR,研究和开发强大的zero-shot语音识别模型。从OLMoASR-Pool开始,收集了300万小时的英语音频和1700万份抄本,我们设计了文本启发式过滤器来删除低质量或错误的数据。我们的策展管道生成了一个包含100万小时高质量音频转录对的新数据集,我们称之为OLMoASR-Mix。我们使用OLMoASR-Mix来训练OLMoASR-Mix模型套件,参数范围从39 M(tiny.en)到1.5 B(large.en)。在所有模型规模中,OLMoASR在短格式和长格式语音识别基准测试中的平均性能与OpenAI的Whisper相当。值得注意的是,OLMoASR-medium.en达到了12.8\%和11.0\%的单词错误率(WER),与Whisper最大的纯英语模型Whisper-medium.en的12.4\%和10.5\% WER分别用于短形式和长形式识别(在等效参数计数下)。OLMoASR池,OLMoASR模型,以及过滤,训练和评估代码将公开提供,以进一步研究鲁棒语音处理。
摘要:Improvements in training data scale and quality have led to significant advances, yet its influence in speech recognition remains underexplored. In this paper, we present a large-scale dataset, OLMoASR-Pool, and series of models, OLMoASR, to study and develop robust zero-shot speech recognition models. Beginning from OLMoASR-Pool, a collection of 3M hours of English audio and 17M transcripts, we design text heuristic filters to remove low-quality or mistranscribed data. Our curation pipeline produces a new dataset containing 1M hours of high-quality audio-transcript pairs, which we call OLMoASR-Mix. We use OLMoASR-Mix to train the OLMoASR-Mix suite of models, ranging from 39M (tiny.en) to 1.5B (large.en) parameters. Across all model scales, OLMoASR achieves comparable average performance to OpenAI's Whisper on short and long-form speech recognition benchmarks. Notably, OLMoASR-medium.en attains a 12.8\% and 11.0\% word error rate (WER) that is on par with Whisper's largest English-only model Whisper-medium.en's 12.4\% and 10.5\% WER for short and long-form recognition respectively (at equivalent parameter count). OLMoASR-Pool, OLMoASR models, and filtering, training and evaluation code will be made publicly available to further research on robust speech processing.


【13】Towards Inclusive Communication: A Unified LLM-Based Framework for Sign Language, Lip Movements, and Audio Understanding
标题:迈向包容性沟通:基于LLM的统一手语、嘴唇运动和音频理解框架
链接:https://arxiv.org/abs/2508.20476

作者:Jeong Hun Yeo, Hyeongseop Rha, Sungjune Park, Junil Won, Yong Man Ro
备注:Code available at: this https URL
摘要:音频是人类交流的主要形式,并推动了自动语音识别(ASR)技术的成功。然而,聋人或听力有障碍的人仍然无法使用这种系统。视觉替代品,如手语和唇读提供了有效的替代品,而手语翻译(手语翻译)和视觉语音识别(视觉语音识别)的最新进展改善了无音频交流。然而,对这些模式的研究基本上是孤立的,将其纳入一个统一框架的探索仍然不足。在本文中,我们介绍了第一个统一的框架,能够处理各种组合的手语,嘴唇运动,和音频口语文本生成。我们专注于三个主要目标:(i)设计一个统一的,模态不可知的架构,能够有效地处理异质输入;(ii)探索模态之间的协同作用,特别是嘴唇运动的非手动提示在手语理解中的作用;(iii)实现性能等同于或优于专门用于个人任务的最先进的模型。在此框架的基础上,我们实现的性能与特定于任务的最先进的模型相当或更好,包括VSR,ASR和AVSR。此外,我们的分析表明,明确建模嘴唇运动作为一个单独的模态显着提高了识别性能。
摘要:Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such systems remain inherently inaccessible to individuals who are deaf or hard of hearing. Visual alternatives such as sign language and lip reading offer effective substitutes, and recent advances in Sign Language Translation (SLT) and Visual Speech Recognition (VSR) have improved audio-less communication. Yet, these modalities have largely been studied in isolation, and their integration within a unified framework remains underexplored. In this paper, we introduce the first unified framework capable of handling diverse combinations of sign language, lip movements, and audio for spoken-language text generation. We focus on three main objectives: (i) designing a unified, modality-agnostic architecture capable of effectively processing heterogeneous inputs; (ii) exploring the underexamined synergy among modalities, particularly the role of lip movements as non-manual cues in sign language comprehension; and (iii) achieving performance on par with or superior to state-of-the-art models specialized for individual tasks. Building on this framework, we achieve performance on par with or better than task-specific state-of-the-art models across SLT, VSR, ASR, and AVSR. Furthermore, our analysis reveals that explicitly modeling lip movements as a separate modality significantly improves SLT performance.


机器翻译由腾讯交互翻译提供,仅供参考