今日论文合集:cs.SD语音11篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】 Benchmarking machine learning for bowel sound pattern classification  from tabular features to pretrained models

标题:从表格特征到预训练模型的肠道声音模式分类机器学习基准
链接:https://arxiv.org/abs/2502.15607
作者:Zahra Mansour,  Verena Uslar,  Dirk Weyhe,  Danilo Hollosi,  Nils Strodthoff
备注:9 pages, 6 figures and 1 table
摘要:电子听诊器和可穿戴记录传感器的发展为肠鸣音(BS)信号的自动分析打开了大门。这使得能够对肠鸣音模式、它们的相互关系以及它们与不同病理的相关性进行数据驱动的分析。这项工作利用了从16名健康受试者收集的BS数据集,根据四个已建立的BS模式进行注释。该数据集用于评估机器学习模型检测和/或分类BS模式的性能。所考虑的模型的选择包括使用表格特征的模型、基于频谱图的卷积神经网络以及在大型音频数据集上预训练的模型。结果突出了预训练模型的明显优势,特别是在检测样本较少的类别时,使用HuBERT模型区分BS与非BS的AUC为0.89,使用Wav 2 Vec 2.0模型区分肠鸣模式的AUC为0.89。这些结果为改善对肠鸣音的理解以及未来机器学习驱动的胃肠道检查诊断应用铺平了道路
摘要:The development of electronic stethoscopes and wearable recording sensorsopened the door to the automated analysis of bowel sound (BS) signals. Thisenables a data-driven analysis of bowel sound patterns, their interrelations,and their correlation to different pathologies. This work leverages a BSdataset collected from 16 healthy subjects that was annotated according to fourestablished BS patterns. This dataset is used to evaluate the performance ofmachine learning models to detect and/or classify BS patterns. The selection ofconsidered models covers models using tabular features, convolutional neuralnetworks based on spectrograms and models pre-trained on large audio datasets.The results highlight the clear superiority of pre-trained models, particularlyin detecting classes with few samples, achieving an AUC of 0.89 indistinguishing BS from non-BS using a HuBERT model and an AUC of 0.89 indifferentiating bowel sound patterns using a Wav2Vec 2.0 model. These resultspave the way for an improved understanding of bowel sounds in general andfuture machine-learning-driven diagnostic applications for gastrointestinalexaminations

【2】 KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio  Generation
标题:KAD:不再时尚!一种有效且高效的音频生成评估指标
链接:https://arxiv.org/abs/2502.15602
作者:Yoonjin Chung,  Pilsun Eu,  Junwon Lee,  Keunwoo Choi,  Juhan Nam,  Ben Sangbae Chon
摘要:虽然被广泛采用用于评估所生成的音频信号,但Fr 'echet音频距离(FAD)受到显著的限制,包括依赖于高斯假设、对样本大小的敏感性和高计算复杂性。作为一种替代方案,我们介绍了内核音频距离(KAD),一种新的,分布自由,无偏,计算效率高的度量的基础上最大均值离散(MMD)。通过分析和实证验证,我们证明了KAD的优势:(1)更快的收敛速度,更小的样本量,使有限的数据进行可靠的评估;(2)更低的计算成本,可扩展的GPU加速;(3)更强的对齐与人类的感知判断。通过利用先进的嵌入和特征内核,KAD捕捉真实音频和生成音频之间的细微差别。KAD在kadtk工具包中开源,为评估生成音频模型提供了一个高效,可靠和感知一致的基准。
摘要:Although being widely adopted for evaluating generated audio signals, theFr\'echet Audio Distance (FAD) suffers from significant limitations, includingreliance on Gaussian assumptions, sensitivity to sample size, and highcomputational complexity. As an alternative, we introduce the Kernel AudioDistance (KAD), a novel, distribution-free, unbiased, and computationallyefficient metric based on Maximum Mean Discrepancy (MMD). Through analysis andempirical validation, we demonstrate KAD's advantages: (1) faster convergencewith smaller sample sizes, enabling reliable evaluation with limited data; (2)lower computational cost, with scalable GPU acceleration; and (3) strongeralignment with human perceptual judgments. By leveraging advanced embeddingsand characteristic kernels, KAD captures nuanced differences between real andgenerated audio. Open-sourced in the kadtk toolkit, KAD provides an efficient,reliable, and perceptually aligned benchmark for evaluating generative audiomodels.

【3】 Advancing User-Voice Interaction: Exploring Emotion-Aware Voice  Assistants Through a Role-Swapping Approach
标题:推进用户语音交互:通过角色交换方法探索具有语音意识的语音助理
链接:https://arxiv.org/abs/2502.15367
作者:Yong Ma,  Yuchong Zhang,  Di Fu,  Stephanie Zubicueta Portales,  Danica Kragic,  Morten Fjeld
备注:19 pages, 6 figures
摘要:随着语音助理(VA)越来越多地融入日常生活,对能够识别用户情绪并做出适当响应的情绪感知系统的需求也在增长。虽然语音情感识别(SER)和情感分析已经取得了重大进展,有效地解决用户的情绪,特别是消极的,仍然是一个挑战。这项研究使用角色交换方法探索了VA交互中的人类情感反应策略,参与者调节AI情感,而不是接受预先编程的反应。通过语音特征分析和自然语言处理(NLP),我们研究了各种情感场景中的声学和语言模式。结果表明,当参与者与负面情绪线索接触时,他们倾向于中性或积极的情绪反应,突出了情绪调节和降级的自然趋势。关键的声学指标,如均方根(RMS),过零率(ZCR),和抖动被确定为敏感的情绪状态,而情感极性和词汇多样性(TTR)区分积极和消极的反应。这些研究结果提供了宝贵的见解,发展适应性,上下文感知的VA能够提供同理心,文化敏感,用户一致的反应。通过了解人类如何自然地调节人工智能交互中的情绪,这项研究有助于设计更直观和情感智能的语音助手,增强用户对人工智能交互的信任和参与度。
摘要:As voice assistants (VAs) become increasingly integrated into daily life, theneed for emotion-aware systems that can recognize and respond appropriately touser emotions has grown. While significant progress has been made in speechemotion recognition (SER) and sentiment analysis, effectively addressing useremotions-particularly negative ones-remains a challenge. This study exploreshuman emotional response strategies in VA interactions using a role-swappingapproach, where participants regulate AI emotions rather than receivingpre-programmed responses. Through speech feature analysis and natural languageprocessing (NLP), we examined acoustic and linguistic patterns across variousemotional scenarios. Results show that participants favor neutral or positiveemotional responses when engaging with negative emotional cues, highlighting anatural tendency toward emotional regulation and de-escalation. Key acousticindicators such as root mean square (RMS), zero-crossing rate (ZCR), and jitterwere identified as sensitive to emotional states, while sentiment polarity andlexical diversity (TTR) distinguished between positive and negative responses.These findings provide valuable insights for developing adaptive, context-awareVAs capable of delivering empathetic, culturally sensitive, and user-alignedresponses. By understanding how humans naturally regulate emotions in AIinteractions, this research contributes to the design of more intuitive andemotionally intelligent voice assistants, enhancing user trust and engagementin human-AI interactions.

【4】 Offload Rethinking by Cloud Assistance for Efficient Environmental Sound  Recognition on LPWANs
标题:通过云协助卸载重新思考,以实现LPWAN上的高效环境声音识别
链接:https://arxiv.org/abs/2502.15285
作者:Le Zhang,  Quanling Zhao,  Run Wang,  Shirley Bian,  Onat Gungor,  Flavio Ponzina,  Tajana Rosing
摘要:基于学习的环境声音识别已成为生物研究和城市规模传感系统中超低功耗环境监测的重要方法。这些系统通常在有限的资源下运行,并且通常由偏远地区采集的能源供电。由于资源限制,最近在设备上的声音识别方面的努力受到低准确性的影响,而云卸载策略受到高通信成本的阻碍。在这项工作中,我们介绍了ORCA,一种新型的资源高效的云辅助环境声音识别系统,在低功耗广域网(LPWAN)上运行的无电池设备上,针对广域音频传感应用。我们提出了一种云辅助策略,可以弥补设备上推理的低准确性,同时最大限度地减少云卸载的通信成本。通过利用基于自注意力的云子光谱特征选择方法来促进有效的设备上推理,ORCA解决了LPWAN上资源受限云卸载的三个关键挑战:1)高通信成本和低数据速率,2)动态无线信道条件,以及3)不可靠的卸载。我们在一个能量收集无电池微控制器上实现了ORCA,并在现实世界的城市声音测试平台上对其进行了评估。我们的研究结果表明,ORCA优于国家的最先进的方法,高达80\倍$的能源节省和220\倍$的延迟减少,同时保持相当的准确性。
摘要:Learning-based environmental sound recognition has emerged as a crucialmethod for ultra-low-power environmental monitoring in biological research andcity-scale sensing systems. These systems usually operate under limitedresources and are often powered by harvested energy in remote areas. Recentefforts in on-device sound recognition suffer from low accuracy due to resourceconstraints, whereas cloud offloading strategies are hindered by highcommunication costs. In this work, we introduce ORCA, a novelresource-efficient cloud-assisted environmental sound recognition system onbatteryless devices operating over the Low-Power Wide-Area Networks (LPWANs),targeting wide-area audio sensing applications. We propose a cloud assistancestrategy that remedies the low accuracy of on-device inference while minimizingthe communication costs for cloud offloading. By leveraging aself-attention-based cloud sub-spectral feature selection method to facilitateefficient on-device inference, ORCA resolves three key challenges forresource-constrained cloud offloading over LPWANs: 1) high communication costsand low data rates, 2) dynamic wireless channel conditions, and 3) unreliableoffloading. We implement ORCA on an energy-harvesting batterylessmicrocontroller and evaluate it in a real world urban sound testbed. Ourresults show that ORCA outperforms state-of-the-art methods by up to $80\times$ in energy savings and $220 \times$ in latency reduction whilemaintaining comparable accuracy.

【5】 Retrieval-Augmented Speech Recognition Approach for Domain Challenges
标题:应对领域挑战的检索增强语音识别方法
链接:https://arxiv.org/abs/2502.15264
作者:Peng Shen,  Xugang Lu,  Hisashi Kawai
摘要:语音识别系统经常面临由于域不匹配的挑战,特别是在现实世界中的应用程序中,由于数据可访问性和机密性的限制,特定于域的数据是不可用的。受大型语言模型(LLM)检索增强生成(RAG)技术的启发,本文提出了一种基于LLM的检索增强语音识别方法,该方法在推理阶段引入特定领域的文本数据,以提高识别性能。在训练阶段,我们的模型不是依赖于特定于域的文本数据,而是经过训练来学习如何利用LLM解码器提示中提供的文本信息来提高语音识别性能。受益于RAG检索机制的优势,我们的方法有效地访问本地可用的特定领域的文件,确保一个方便和有效的过程来解决域不匹配的问题。在CSJ数据库上进行的实验表明,该方法显着提高了语音识别的准确性,并在CSJ数据集上实现了最先进的结果,即使不依赖于完整的训练数据。
摘要:Speech recognition systems often face challenges due to domain mismatch,particularly in real-world applications where domain-specific data isunavailable because of data accessibility and confidentiality constraints.Inspired by Retrieval-Augmented Generation (RAG) techniques for large languagemodels (LLMs), this paper introduces a LLM-based retrieval-augmented speechrecognition method that incorporates domain-specific textual data at theinference stage to enhance recognition performance. Rather than relying ondomain-specific textual data during the training phase, our model is trained tolearn how to utilize textual information provided in prompts for LLM decoder toimprove speech recognition performance. Benefiting from the advantages of theRAG retrieval mechanism, our approach efficiently accesses locally availabledomain-specific documents, ensuring a convenient and effective process forsolving domain mismatch problems. Experiments conducted on the CSJ databasedemonstrate that the proposed method significantly improves speech recognitionaccuracy and achieves state-of-the-art results on the CSJ dataset, even withoutrelying on the full training data.

【6】 ESPnet-SpeechLM: An Open Speech Language Model Toolkit
标题:ESPnet-SpeechLM:开放语音语言模型工具包
链接:https://arxiv.org/abs/2502.15218
作者:Jinchuan Tian,  Jiatong Shi,  William Chen,  Siddhant Arora,  Yoshiki Masuyama,  Takashi Maekaku,  Yihan Wu,  Junyi Peng,  Shikhar Bharadwaj,  Yiwen Zhao,  Samuele Cornell,  Yifan Peng,  Xiang Yue,  Chao-Han Huck Yang,  Graham Neubig,  Shinji Watanabe
摘要:我们提出ESPnet-SpeechLM,一个开放的工具包,旨在民主化的语音语言模型(SpeechLM)和语音驱动的代理应用程序的发展。该工具包通过将语音处理任务框定为通用的顺序建模问题,包括数据预处理,预训练,推理和任务评估的内聚工作流来简化语音处理任务。使用ESPnet-SpeechLM,用户可以轻松定义任务模板和配置关键设置,从而实现无缝和简化的SpeechLM开发。该工具包通过为工作流的每个阶段提供高度可配置的模块,确保了灵活性、效率和可扩展性。为了说明其功能,我们提供了多个用例,演示如何使用ESPnet-SpeechLM构建具有竞争力的SpeechLM,包括在不同基准测试中对文本和语音任务进行预训练的1. 7 B参数模型。该工具包及其配方是完全透明和可复制的:https://github.com/espnet/espnet/tree/speechlm。
摘要:We present ESPnet-SpeechLM, an open toolkit designed to democratize thedevelopment of speech language models (SpeechLMs) and voice-driven agenticapplications. The toolkit standardizes speech processing tasks by framing themas universal sequential modeling problems, encompassing a cohesive workflow ofdata preprocessing, pre-training, inference, and task evaluation. WithESPnet-SpeechLM, users can easily define task templates and configure keysettings, enabling seamless and streamlined SpeechLM development. The toolkitensures flexibility, efficiency, and scalability by offering highlyconfigurable modules for every stage of the workflow. To illustrate itscapabilities, we provide multiple use cases demonstrating how competitiveSpeechLMs can be constructed with ESPnet-SpeechLM, including a 1.7B-parametermodel pre-trained on both text and speech tasks, across diverse benchmarks. Thetoolkit and its recipes are fully transparent and reproducible at:https://github.com/espnet/espnet/tree/speechlm.

【7】 Improving Streaming Speech Recognition With Time-Shifted Contextual  Attention And Dynamic Right Context Masking
标题:利用时移上下文注意力和动态正确上下文掩蔽改进流语音识别
链接:https://arxiv.org/abs/2502.15158
作者:Khanh Le,  Duc Chau
备注:INTERSPEECH 2024
摘要:基于组块的推理是开发实时流语音识别的一种流行方法,因其简单和高效而受到重视。但是,由于它将模型的焦点限制为仅关注历史和当前块上下文,因此在需要考虑未来上下文的场景中可能会导致性能下降。针对这一点,我们提出了一种新的方法,具有时移上下文注意力(TSCA)和动态右上下文(DRC)掩蔽。我们的方法显示了一个相对字的错误率降低10至13.9%的Librispeech数据集与TSCA提供的上下文中的未来信息的列入。此外,我们提出了一个流自动语音识别管道,便于TSCA的集成与最小的用户感知的延迟,同时还实现批处理能力,使其适用于各种应用程序。
摘要:Chunk-based inference stands out as a popular approach in developingreal-time streaming speech recognition, valued for its simplicity andefficiency. However, because it restricts the model's focus to only the historyand current chunk context, it may result in performance degradation inscenarios that demand consideration of future context. Addressing this, wepropose a novel approach featuring Time-Shifted Contextual Attention (TSCA) andDynamic Right Context (DRC) masking. Our method shows a relative word errorrate reduction of 10 to 13.9% on the Librispeech dataset with the inclusion ofin-context future information provided by TSCA. Moreover, we present astreaming automatic speech recognition pipeline that facilitates theintegration of TSCA with minimal user-perceived latency, while also enablingbatch processing capability, making it practical for various applications.

【8】 Fundamental Survey on Neuromorphic Based Audio Classification
标题:基于神经形态的音频分类的基础研究
链接:https://arxiv.org/abs/2502.15056
作者:Amlan Basu,  Pranav Chaudhari,  Gaetano Di Caterina
备注:24 Pages, 1 Table
摘要:音频分类在各种应用中至关重要,包括监控、医疗监测和环境分析。传统方法通常依赖于复杂的信号处理算法和手工制作的功能,这可能无法完全捕捉音频模式的复杂性。神经形态计算,灵感来自于人类大脑的架构和功能,为音频分类任务提供了一个有前途的替代方案。这项调查提供了一个详尽的检查,目前国家的最先进的神经形态为基础的音频分类。它深入研究了神经形态系统的关键组件,如尖峰神经网络(SNN),忆阻器和神经形态硬件平台,突出了它们在音频分类中的优势。此外,该调查还探讨了神经形态音频分类中采用的各种方法和策略,包括基于事件的处理,基于尖峰的学习和生物启发的特征提取。它探讨了这些方法如何解决传统音频分类方法的局限性,特别是在能源效率,实时处理和对环境噪声的鲁棒性方面。此外,本文对不同的神经形态音频分类模型和基准进行了比较分析,评估了它们的性能指标,计算效率和可扩展性。通过为研究人员、工程师和从业者提供全面的指南,本调查旨在刺激不断发展的神经形态音频分类领域的进一步创新和进步。
摘要:Audio classification is paramount in a variety of applications includingsurveillance, healthcare monitoring, and environmental analysis. Traditionalmethods frequently depend on intricate signal processing algorithms andmanually crafted features, which may fall short in fully capturing thecomplexities of audio patterns. Neuromorphic computing, inspired by thearchitecture and functioning of the human brain, presents a promisingalternative for audio classification tasks. This survey provides an exhaustiveexamination of the current state-of-the-art in neuromorphic-based audioclassification. It delves into the crucial components of neuromorphic systems,such as Spiking Neural Networks (SNNs), memristors, and neuromorphic hardwareplatforms, highlighting their advantages in audio classification. Furthermore,the survey explores various methodologies and strategies employed inneuromorphic audio classification, including event-based processing,spike-based learning, and bio-inspired feature extraction. It examines howthese approaches address the limitations of traditional audio classificationmethods, particularly in terms of energy efficiency, real-time processing, androbustness to environmental noise. Additionally, the paper conducts acomparative analysis of different neuromorphic audio classification models andbenchmarks, evaluating their performance metrics, computational efficiency, andscalability. By providing a comprehensive guide for researchers, engineers andpractitioners, this survey aims to stimulate further innovation andadvancements in the evolving field of neuromorphic audio classification.

【9】 NOTA: Multimodal Music Notation Understanding for Visual Large Language  Model
标题:NOTA:视觉大语言模型的多模式音乐记法理解
链接:https://arxiv.org/abs/2502.14893
作者:Mingni Tang,  Jiajia Li,  Lu Yang,  Zhiqiang Zhang,  Jinghao Tian,  Zuchao Li,  Lefei Zhang,  Ping Wang
摘要:符号音乐以两种不同的形式表示:二维的、视觉上直观的乐谱图像和一维的、标准化的文本注释序列。虽然大型语言模型在音乐中显示出非凡的潜力,但目前的研究主要集中在单峰符号序列文本上。现有的通用领域视觉语言模型仍然缺乏对乐谱的理解能力。认识到这一差距,我们提出了NOTA,第一个大规模的综合多模态音乐记谱数据集。它由来自世界3个地区的1,019,237条记录组成,包含3项任务。基于数据集,我们训练了NotaGPT,这是一个音乐符号可视化大语言模型。具体来说,我们涉及一个预对准训练阶段,用于乐谱图像中描绘的音符与ABC符号中的文本表示之间的跨模态对准。随后的训练阶段侧重于基础音乐信息提取,然后是音乐符号分析的训练。实验结果表明,我们的NotaGPT-7 B在音乐理解方面取得了显着的改善,展示了NOTA和训练管道的有效性。我们的数据集在https://huggingface.co/datasets/MYTH-Lab/NOTA-dataset上开源。
摘要:Symbolic music is represented in two distinct forms: two-dimensional,visually intuitive score images, and one-dimensional, standardized textannotation sequences. While large language models have shown extraordinarypotential in music, current research has primarily focused on unimodal symbolsequence text. Existing general-domain visual language models still lack theability of music notation understanding. Recognizing this gap, we propose NOTA,the first large-scale comprehensive multimodal music notation dataset. Itconsists of 1,019,237 records, from 3 regions of the world, and contains 3tasks. Based on the dataset, we trained NotaGPT, a music notation visual largelanguage model. Specifically, we involve a pre-alignment training phase forcross-modal alignment between the musical notes depicted in music score imagesand their textual representation in ABC notation. Subsequent training phasesfocus on foundational music information extraction, followed by training onmusic notation analysis. Experimental results demonstrate that our NotaGPT-7Bachieves significant improvement on music understanding, showcasing theeffectiveness of NOTA and the training pipeline. Our datasets are open-sourcedat https://huggingface.co/datasets/MYTH-Lab/NOTA-dataset.

【10】 Audio signal interpolation using optimal transportation of spectrograms
标题:使用频谱图的最佳传输的音频信号内插
链接:https://arxiv.org/abs/2502.15430
作者:David Valdivia,  Marien Renaud,  Elsa Cazelles,  Cédric Févotte
摘要:我们提出了一种新的方法,用于生成一个人工音频信号,在给定的源和目标声音之间进行插值。我们的方法依赖于计算Wasserstein重心的源和目标频谱图,其次是相位重建和反演。与以前的作品相比,我们的新方法认为,全球范围内的频谱图,并没有在一个时间帧到帧的基础上操作。另一个贡献是赋予运输成本矩阵一个特定的结构,禁止能源沿时间轴的远程位移,并通过利用不平衡的运输框架,使最佳运输成为可能。从音频角度来看,提出的成本矩阵是有意义的,并且还可以减少计算负载。合成音符和真实环境声音的结果说明了我们的新方法的潜力。
摘要:We present a novel approach for generating an artificial audio signal thatinterpolates between given source and target sounds. Our approach relies on thecomputation of Wasserstein barycenters of the source and target spectrograms,followed by phase reconstruction and inversion. In contrast with previousworks, our new method considers the spectrograms globally and does not operateon a temporal frame-to-frame basis. An other contribution is to endow thetransportation cost matrix with a specific structure that prohibits remotedisplacements of energy along the time axis, and for which optimal transport ismade possible by leveraging the unbalanced transport framework. The proposedcost matrix makes sense from the audio perspective and also allows to reducethe computation load. Results with synthetic musical notes and realenvironmental sounds illustrate the potential of our novel approach.

【11】 Enhancing Speech Large Language Models with Prompt-Aware Mixture of  Audio Encoders
标题:使用预算感知的音频编码器混合增强语音大型语言模型
链接:https://arxiv.org/abs/2502.15178
作者:Weiqiao Shan,  Yuang Li,  Yuhao Zhang,  Yingfeng Luo,  Chen Xu,  Xiaofeng Zhao,  Long Meng,  Yunfei Lu,  Min Zhang,  Hao Yang,  Tong Xiao,  Jingbo Zhu
备注:12 pages,4 figures, 7 tables
摘要:将音频编码器与大型语言模型(LLM)连接,使LLM能够执行各种音频理解任务,例如自动语音识别(ASR)和音频字幕(AC)。大多数研究集中在训练适配器层以生成用于LLM的统一音频特征。然而,不同的任务可能需要强调语义或声学方面的不同特征,使得特定于任务的音频特征更可取。在本文中,我们提出了语音感知混合(PaM),以增强使用多个音频编码器的语音LLM。我们的方法涉及使用不同的专家提取不同的功能的提示,指示不同的任务。实验表明,与PaM,只有一个语音LLM超过所有单编码器语音LLM实现的ASR,扬声器数量验证,和AC任务的最佳性能。PaM也优于其他特征融合基线,例如串联和平均。
摘要:Connecting audio encoders with large language models (LLMs) allows the LLM toperform various audio understanding tasks, such as automatic speech recognition(ASR) and audio captioning (AC). Most research focuses on training an adapterlayer to generate a unified audio feature for the LLM. However, different tasksmay require distinct features that emphasize either semantic or acousticaspects, making task-specific audio features more desirable. In this paper, wepropose Prompt-aware Mixture (PaM) to enhance the Speech LLM that uses multipleaudio encoders. Our approach involves using different experts to extractdifferent features based on the prompt that indicates different tasks.Experiments demonstrate that with PaM, only one Speech LLM surpasses the bestperformances achieved by all single-encoder Speech LLMs on ASR, Speaker NumberVerification, and AC tasks. PaM also outperforms other feature fusionbaselines, such as concatenation and averaging.

eess.AS音频处理

【1】 Audio signal interpolation using optimal transportation of spectrograms
标题:使用频谱图的最佳传输的音频信号内插
链接:https://arxiv.org/abs/2502.15430
作者:David Valdivia,  Marien Renaud,  Elsa Cazelles,  Cédric Févotte
摘要:我们提出了一种新的方法,用于生成一个人工音频信号,在给定的源和目标声音之间进行插值。我们的方法依赖于计算Wasserstein重心的源和目标频谱图,其次是相位重建和反演。与以前的作品相比,我们的新方法认为,全球范围内的频谱图,并没有在一个时间帧到帧的基础上操作。另一个贡献是赋予运输成本矩阵一个特定的结构,禁止能源沿时间轴的远程位移,并通过利用不平衡的运输框架,使最佳运输成为可能。从音频角度来看,提出的成本矩阵是有意义的,并且还可以减少计算负载。合成音符和真实环境声音的结果说明了我们的新方法的潜力。
摘要:We present a novel approach for generating an artificial audio signal thatinterpolates between given source and target sounds. Our approach relies on thecomputation of Wasserstein barycenters of the source and target spectrograms,followed by phase reconstruction and inversion. In contrast with previousworks, our new method considers the spectrograms globally and does not operateon a temporal frame-to-frame basis. An other contribution is to endow thetransportation cost matrix with a specific structure that prohibits remotedisplacements of energy along the time axis, and for which optimal transport ismade possible by leveraging the unbalanced transport framework. The proposedcost matrix makes sense from the audio perspective and also allows to reducethe computation load. Results with synthetic musical notes and realenvironmental sounds illustrate the potential of our novel approach.

【2】 Enhancing Speech Large Language Models with Prompt-Aware Mixture of  Audio Encoders
标题:使用预算感知的音频编码器混合增强语音大型语言模型
链接:https://arxiv.org/abs/2502.15178
作者:Weiqiao Shan,  Yuang Li,  Yuhao Zhang,  Yingfeng Luo,  Chen Xu,  Xiaofeng Zhao,  Long Meng,  Yunfei Lu,  Min Zhang,  Hao Yang,  Tong Xiao,  Jingbo Zhu
备注:12 pages,4 figures, 7 tables
摘要:将音频编码器与大型语言模型(LLM)连接,使LLM能够执行各种音频理解任务,例如自动语音识别(ASR)和音频字幕(AC)。大多数研究集中在训练适配器层以生成用于LLM的统一音频特征。然而,不同的任务可能需要强调语义或声学方面的不同特征,使得特定于任务的音频特征更可取。在本文中,我们提出了语音感知混合(PaM),以增强使用多个音频编码器的语音LLM。我们的方法涉及使用不同的专家提取不同的功能的提示,指示不同的任务。实验表明,与PaM,只有一个语音LLM超过所有单编码器语音LLM实现的ASR,扬声器数量验证,和AC任务的最佳性能。PaM也优于其他特征融合基线,例如串联和平均。
摘要:Connecting audio encoders with large language models (LLMs) allows the LLM toperform various audio understanding tasks, such as automatic speech recognition(ASR) and audio captioning (AC). Most research focuses on training an adapterlayer to generate a unified audio feature for the LLM. However, different tasksmay require distinct features that emphasize either semantic or acousticaspects, making task-specific audio features more desirable. In this paper, wepropose Prompt-aware Mixture (PaM) to enhance the Speech LLM that uses multipleaudio encoders. Our approach involves using different experts to extractdifferent features based on the prompt that indicates different tasks.Experiments demonstrate that with PaM, only one Speech LLM surpasses the bestperformances achieved by all single-encoder Speech LLMs on ASR, Speaker NumberVerification, and AC tasks. PaM also outperforms other feature fusionbaselines, such as concatenation and averaging.

【3】 Benchmarking machine learning for bowel sound pattern classification  from tabular features to pretrained models
标题:从表格特征到预训练模型的肠道声音模式分类机器学习基准
链接:https://arxiv.org/abs/2502.15607
作者:Zahra Mansour,  Verena Uslar,  Dirk Weyhe,  Danilo Hollosi,  Nils Strodthoff
备注:9 pages, 6 figures and 1 table
摘要:电子听诊器和可穿戴记录传感器的发展为肠鸣音(BS)信号的自动分析打开了大门。这使得能够对肠鸣音模式、它们的相互关系以及它们与不同病理的相关性进行数据驱动的分析。这项工作利用了从16名健康受试者收集的BS数据集,根据四个已建立的BS模式进行注释。该数据集用于评估机器学习模型检测和/或分类BS模式的性能。所考虑的模型的选择包括使用表格特征的模型、基于频谱图的卷积神经网络以及在大型音频数据集上预训练的模型。结果突出了预训练模型的明显优势,特别是在检测样本较少的类别时,使用HuBERT模型区分BS与非BS的AUC为0.89,使用Wav 2 Vec 2.0模型区分肠鸣模式的AUC为0.89。这些结果为改善对肠鸣音的理解以及未来机器学习驱动的胃肠道检查诊断应用铺平了道路
摘要:The development of electronic stethoscopes and wearable recording sensorsopened the door to the automated analysis of bowel sound (BS) signals. Thisenables a data-driven analysis of bowel sound patterns, their interrelations,and their correlation to different pathologies. This work leverages a BSdataset collected from 16 healthy subjects that was annotated according to fourestablished BS patterns. This dataset is used to evaluate the performance ofmachine learning models to detect and/or classify BS patterns. The selection ofconsidered models covers models using tabular features, convolutional neuralnetworks based on spectrograms and models pre-trained on large audio datasets.The results highlight the clear superiority of pre-trained models, particularlyin detecting classes with few samples, achieving an AUC of 0.89 indistinguishing BS from non-BS using a HuBERT model and an AUC of 0.89 indifferentiating bowel sound patterns using a Wav2Vec 2.0 model. These resultspave the way for an improved understanding of bowel sounds in general andfuture machine-learning-driven diagnostic applications for gastrointestinalexaminations

【4】 KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio  Generation
标题:KAD:不再时尚!一种有效且高效的音频生成评估指标
链接:https://arxiv.org/abs/2502.15602
作者:Yoonjin Chung,  Pilsun Eu,  Junwon Lee,  Keunwoo Choi,  Juhan Nam,  Ben Sangbae Chon
摘要:虽然被广泛采用用于评估所生成的音频信号,但Fr 'echet音频距离(FAD)受到显著的限制,包括依赖于高斯假设、对样本大小的敏感性和高计算复杂性。作为一种替代方案,我们介绍了内核音频距离(KAD),一种新的,分布自由,无偏,计算效率高的度量的基础上最大均值离散(MMD)。通过分析和实证验证,我们证明了KAD的优势:(1)更快的收敛速度,更小的样本量,使有限的数据进行可靠的评估;(2)更低的计算成本,可扩展的GPU加速;(3)更强的对齐与人类的感知判断。通过利用先进的嵌入和特征内核,KAD捕捉真实音频和生成音频之间的细微差别。KAD在kadtk工具包中开源,为评估生成音频模型提供了一个高效,可靠和感知一致的基准。
摘要:Although being widely adopted for evaluating generated audio signals, theFr\'echet Audio Distance (FAD) suffers from significant limitations, includingreliance on Gaussian assumptions, sensitivity to sample size, and highcomputational complexity. As an alternative, we introduce the Kernel AudioDistance (KAD), a novel, distribution-free, unbiased, and computationallyefficient metric based on Maximum Mean Discrepancy (MMD). Through analysis andempirical validation, we demonstrate KAD's advantages: (1) faster convergencewith smaller sample sizes, enabling reliable evaluation with limited data; (2)lower computational cost, with scalable GPU acceleration; and (3) strongeralignment with human perceptual judgments. By leveraging advanced embeddingsand characteristic kernels, KAD captures nuanced differences between real andgenerated audio. Open-sourced in the kadtk toolkit, KAD provides an efficient,reliable, and perceptually aligned benchmark for evaluating generative audiomodels.

【5】 Advancing User-Voice Interaction: Exploring Emotion-Aware Voice  Assistants Through a Role-Swapping Approach
标题:推进用户语音交互:通过角色交换方法探索具有语音意识的语音助理
链接:https://arxiv.org/abs/2502.15367
作者:Yong Ma,  Yuchong Zhang,  Di Fu,  Stephanie Zubicueta Portales,  Danica Kragic,  Morten Fjeld
备注:19 pages, 6 figures
摘要:随着语音助理(VA)越来越多地融入日常生活,对能够识别用户情绪并做出适当响应的情绪感知系统的需求也在增长。虽然语音情感识别(SER)和情感分析已经取得了重大进展,有效地解决用户的情绪,特别是消极的,仍然是一个挑战。这项研究使用角色交换方法探索了VA交互中的人类情感反应策略,参与者调节AI情感,而不是接受预先编程的反应。通过语音特征分析和自然语言处理(NLP),我们研究了各种情感场景中的声学和语言模式。结果表明,当参与者与负面情绪线索接触时,他们倾向于中性或积极的情绪反应,突出了情绪调节和降级的自然趋势。关键的声学指标,如均方根(RMS),过零率(ZCR),和抖动被确定为敏感的情绪状态,而情感极性和词汇多样性(TTR)区分积极和消极的反应。这些研究结果提供了宝贵的见解,发展适应性,上下文感知的VA能够提供同理心,文化敏感,用户一致的反应。通过了解人类如何自然地调节人工智能交互中的情绪,这项研究有助于设计更直观和情感智能的语音助手,增强用户对人工智能交互的信任和参与度。
摘要:As voice assistants (VAs) become increasingly integrated into daily life, theneed for emotion-aware systems that can recognize and respond appropriately touser emotions has grown. While significant progress has been made in speechemotion recognition (SER) and sentiment analysis, effectively addressing useremotions-particularly negative ones-remains a challenge. This study exploreshuman emotional response strategies in VA interactions using a role-swappingapproach, where participants regulate AI emotions rather than receivingpre-programmed responses. Through speech feature analysis and natural languageprocessing (NLP), we examined acoustic and linguistic patterns across variousemotional scenarios. Results show that participants favor neutral or positiveemotional responses when engaging with negative emotional cues, highlighting anatural tendency toward emotional regulation and de-escalation. Key acousticindicators such as root mean square (RMS), zero-crossing rate (ZCR), and jitterwere identified as sensitive to emotional states, while sentiment polarity andlexical diversity (TTR) distinguished between positive and negative responses.These findings provide valuable insights for developing adaptive, context-awareVAs capable of delivering empathetic, culturally sensitive, and user-alignedresponses. By understanding how humans naturally regulate emotions in AIinteractions, this research contributes to the design of more intuitive andemotionally intelligent voice assistants, enhancing user trust and engagementin human-AI interactions.

【6】 Offload Rethinking by Cloud Assistance for Efficient Environmental Sound  Recognition on LPWANs
标题:通过云协助卸载重新思考,以实现LPWAN上的高效环境声音识别
链接:https://arxiv.org/abs/2502.15285
作者:Le Zhang,  Quanling Zhao,  Run Wang,  Shirley Bian,  Onat Gungor,  Flavio Ponzina,  Tajana Rosing
摘要:基于学习的环境声音识别已成为生物研究和城市规模传感系统中超低功耗环境监测的重要方法。这些系统通常在有限的资源下运行,并且通常由偏远地区采集的能源供电。由于资源限制,最近在设备上的声音识别方面的努力受到低准确性的影响,而云卸载策略受到高通信成本的阻碍。在这项工作中,我们介绍了ORCA,一种新型的资源高效的云辅助环境声音识别系统,在低功耗广域网(LPWAN)上运行的无电池设备上,针对广域音频传感应用。我们提出了一种云辅助策略,可以弥补设备上推理的低准确性,同时最大限度地减少云卸载的通信成本。通过利用基于自注意力的云子光谱特征选择方法来促进有效的设备上推理,ORCA解决了LPWAN上资源受限云卸载的三个关键挑战:1)高通信成本和低数据速率,2)动态无线信道条件,以及3)不可靠的卸载。我们在一个能量收集无电池微控制器上实现了ORCA,并在现实世界的城市声音测试平台上对其进行了评估。我们的研究结果表明,ORCA优于国家的最先进的方法,高达80\倍$的能源节省和220\倍$的延迟减少,同时保持相当的准确性。
摘要:Learning-based environmental sound recognition has emerged as a crucialmethod for ultra-low-power environmental monitoring in biological research andcity-scale sensing systems. These systems usually operate under limitedresources and are often powered by harvested energy in remote areas. Recentefforts in on-device sound recognition suffer from low accuracy due to resourceconstraints, whereas cloud offloading strategies are hindered by highcommunication costs. In this work, we introduce ORCA, a novelresource-efficient cloud-assisted environmental sound recognition system onbatteryless devices operating over the Low-Power Wide-Area Networks (LPWANs),targeting wide-area audio sensing applications. We propose a cloud assistancestrategy that remedies the low accuracy of on-device inference while minimizingthe communication costs for cloud offloading. By leveraging aself-attention-based cloud sub-spectral feature selection method to facilitateefficient on-device inference, ORCA resolves three key challenges forresource-constrained cloud offloading over LPWANs: 1) high communication costsand low data rates, 2) dynamic wireless channel conditions, and 3) unreliableoffloading. We implement ORCA on an energy-harvesting batterylessmicrocontroller and evaluate it in a real world urban sound testbed. Ourresults show that ORCA outperforms state-of-the-art methods by up to $80\times$ in energy savings and $220 \times$ in latency reduction whilemaintaining comparable accuracy.

【7】 Retrieval-Augmented Speech Recognition Approach for Domain Challenges
标题:应对领域挑战的检索增强语音识别方法
链接:https://arxiv.org/abs/2502.15264
作者:Peng Shen,  Xugang Lu,  Hisashi Kawai
摘要:语音识别系统经常面临由于域不匹配的挑战,特别是在现实世界中的应用程序中,由于数据可访问性和机密性的限制,特定于域的数据是不可用的。受大型语言模型(LLM)检索增强生成(RAG)技术的启发,本文提出了一种基于LLM的检索增强语音识别方法,该方法在推理阶段引入特定领域的文本数据,以提高识别性能。在训练阶段,我们的模型不是依赖于特定于域的文本数据,而是经过训练来学习如何利用LLM解码器提示中提供的文本信息来提高语音识别性能。受益于RAG检索机制的优势,我们的方法有效地访问本地可用的特定领域的文件,确保一个方便和有效的过程来解决域不匹配的问题。在CSJ数据库上进行的实验表明,该方法显着提高了语音识别的准确性,并在CSJ数据集上实现了最先进的结果,即使不依赖于完整的训练数据。
摘要:Speech recognition systems often face challenges due to domain mismatch,particularly in real-world applications where domain-specific data isunavailable because of data accessibility and confidentiality constraints.Inspired by Retrieval-Augmented Generation (RAG) techniques for large languagemodels (LLMs), this paper introduces a LLM-based retrieval-augmented speechrecognition method that incorporates domain-specific textual data at theinference stage to enhance recognition performance. Rather than relying ondomain-specific textual data during the training phase, our model is trained tolearn how to utilize textual information provided in prompts for LLM decoder toimprove speech recognition performance. Benefiting from the advantages of theRAG retrieval mechanism, our approach efficiently accesses locally availabledomain-specific documents, ensuring a convenient and effective process forsolving domain mismatch problems. Experiments conducted on the CSJ databasedemonstrate that the proposed method significantly improves speech recognitionaccuracy and achieves state-of-the-art results on the CSJ dataset, even withoutrelying on the full training data.

【8】 ESPnet-SpeechLM: An Open Speech Language Model Toolkit
标题:ESPnet-SpeechLM:开放语音语言模型工具包
链接:https://arxiv.org/abs/2502.15218
作者:Jinchuan Tian,  Jiatong Shi,  William Chen,  Siddhant Arora,  Yoshiki Masuyama,  Takashi Maekaku,  Yihan Wu,  Junyi Peng,  Shikhar Bharadwaj,  Yiwen Zhao,  Samuele Cornell,  Yifan Peng,  Xiang Yue,  Chao-Han Huck Yang,  Graham Neubig,  Shinji Watanabe
摘要:我们提出ESPnet-SpeechLM,一个开放的工具包,旨在民主化的语音语言模型(SpeechLM)和语音驱动的代理应用程序的发展。该工具包通过将语音处理任务框定为通用的顺序建模问题,包括数据预处理,预训练,推理和任务评估的内聚工作流来简化语音处理任务。使用ESPnet-SpeechLM,用户可以轻松定义任务模板和配置关键设置,从而实现无缝和简化的SpeechLM开发。该工具包通过为工作流的每个阶段提供高度可配置的模块,确保了灵活性、效率和可扩展性。为了说明其功能,我们提供了多个用例,演示如何使用ESPnet-SpeechLM构建具有竞争力的SpeechLM,包括在不同基准测试中对文本和语音任务进行预训练的1. 7 B参数模型。该工具包及其配方是完全透明和可复制的:https://github.com/espnet/espnet/tree/speechlm。
摘要:We present ESPnet-SpeechLM, an open toolkit designed to democratize thedevelopment of speech language models (SpeechLMs) and voice-driven agenticapplications. The toolkit standardizes speech processing tasks by framing themas universal sequential modeling problems, encompassing a cohesive workflow ofdata preprocessing, pre-training, inference, and task evaluation. WithESPnet-SpeechLM, users can easily define task templates and configure keysettings, enabling seamless and streamlined SpeechLM development. The toolkitensures flexibility, efficiency, and scalability by offering highlyconfigurable modules for every stage of the workflow. To illustrate itscapabilities, we provide multiple use cases demonstrating how competitiveSpeechLMs can be constructed with ESPnet-SpeechLM, including a 1.7B-parametermodel pre-trained on both text and speech tasks, across diverse benchmarks. Thetoolkit and its recipes are fully transparent and reproducible at:https://github.com/espnet/espnet/tree/speechlm.

【9】 Improving Streaming Speech Recognition With Time-Shifted Contextual  Attention And Dynamic Right Context Masking
标题:利用时移上下文注意力和动态正确上下文掩蔽改进流语音识别
链接:https://arxiv.org/abs/2502.15158
作者:Khanh Le,  Duc Chau
备注:INTERSPEECH 2024
摘要:基于组块的推理是开发实时流语音识别的一种流行方法,因其简单和高效而受到重视。但是,由于它将模型的焦点限制为仅关注历史和当前块上下文,因此在需要考虑未来上下文的场景中可能会导致性能下降。针对这一点,我们提出了一种新的方法,具有时移上下文注意力(TSCA)和动态右上下文(DRC)掩蔽。我们的方法显示了一个相对字的错误率降低10至13.9%的Librispeech数据集与TSCA提供的上下文中的未来信息的列入。此外,我们提出了一个流自动语音识别管道,便于TSCA的集成与最小的用户感知的延迟,同时还实现批处理能力,使其适用于各种应用程序。
摘要:Chunk-based inference stands out as a popular approach in developingreal-time streaming speech recognition, valued for its simplicity andefficiency. However, because it restricts the model's focus to only the historyand current chunk context, it may result in performance degradation inscenarios that demand consideration of future context. Addressing this, wepropose a novel approach featuring Time-Shifted Contextual Attention (TSCA) andDynamic Right Context (DRC) masking. Our method shows a relative word errorrate reduction of 10 to 13.9% on the Librispeech dataset with the inclusion ofin-context future information provided by TSCA. Moreover, we present astreaming automatic speech recognition pipeline that facilitates theintegration of TSCA with minimal user-perceived latency, while also enablingbatch processing capability, making it practical for various applications.

【10】 Fundamental Survey on Neuromorphic Based Audio Classification
标题:基于神经形态的音频分类的基础研究
链接:https://arxiv.org/abs/2502.15056
作者:Amlan Basu,  Pranav Chaudhari,  Gaetano Di Caterina
备注:24 Pages, 1 Table
摘要:音频分类在各种应用中至关重要,包括监控、医疗监测和环境分析。传统方法通常依赖于复杂的信号处理算法和手工制作的功能,这可能无法完全捕捉音频模式的复杂性。神经形态计算,灵感来自于人类大脑的架构和功能,为音频分类任务提供了一个有前途的替代方案。这项调查提供了一个详尽的检查,目前国家的最先进的神经形态为基础的音频分类。它深入研究了神经形态系统的关键组件,如尖峰神经网络(SNN),忆阻器和神经形态硬件平台,突出了它们在音频分类中的优势。此外,该调查还探讨了神经形态音频分类中采用的各种方法和策略,包括基于事件的处理,基于尖峰的学习和生物启发的特征提取。它探讨了这些方法如何解决传统音频分类方法的局限性,特别是在能源效率,实时处理和对环境噪声的鲁棒性方面。此外,本文对不同的神经形态音频分类模型和基准进行了比较分析,评估了它们的性能指标,计算效率和可扩展性。通过为研究人员、工程师和从业者提供全面的指南,本调查旨在刺激不断发展的神经形态音频分类领域的进一步创新和进步。
摘要:Audio classification is paramount in a variety of applications includingsurveillance, healthcare monitoring, and environmental analysis. Traditionalmethods frequently depend on intricate signal processing algorithms andmanually crafted features, which may fall short in fully capturing thecomplexities of audio patterns. Neuromorphic computing, inspired by thearchitecture and functioning of the human brain, presents a promisingalternative for audio classification tasks. This survey provides an exhaustiveexamination of the current state-of-the-art in neuromorphic-based audioclassification. It delves into the crucial components of neuromorphic systems,such as Spiking Neural Networks (SNNs), memristors, and neuromorphic hardwareplatforms, highlighting their advantages in audio classification. Furthermore,the survey explores various methodologies and strategies employed inneuromorphic audio classification, including event-based processing,spike-based learning, and bio-inspired feature extraction. It examines howthese approaches address the limitations of traditional audio classificationmethods, particularly in terms of energy efficiency, real-time processing, androbustness to environmental noise. Additionally, the paper conducts acomparative analysis of different neuromorphic audio classification models andbenchmarks, evaluating their performance metrics, computational efficiency, andscalability. By providing a comprehensive guide for researchers, engineers andpractitioners, this survey aims to stimulate further innovation andadvancements in the evolving field of neuromorphic audio classification.

【11】 NOTA: Multimodal Music Notation Understanding for Visual Large Language  Model
标题:NOTA:视觉大语言模型的多模式音乐记法理解
链接:https://arxiv.org/abs/2502.14893
作者:Mingni Tang,  Jiajia Li,  Lu Yang,  Zhiqiang Zhang,  Jinghao Tian,  Zuchao Li,  Lefei Zhang,  Ping Wang
摘要:符号音乐以两种不同的形式表示:二维的、视觉上直观的乐谱图像和一维的、标准化的文本注释序列。虽然大型语言模型在音乐中显示出非凡的潜力,但目前的研究主要集中在单峰符号序列文本上。现有的通用领域视觉语言模型仍然缺乏乐谱理解能力。认识到这一差距,我们提出了NOTA,第一个大规模的综合多模态音乐记谱数据集。它由来自世界3个地区的1,019,237条记录组成,包含3项任务。基于数据集,我们训练了NotaGPT,这是一个音乐符号可视化大语言模型。具体来说,我们涉及一个预对准训练阶段,用于乐谱图像中描绘的音符与ABC符号中的文本表示之间的跨模态对准。随后的训练阶段侧重于基础音乐信息提取,然后是音乐符号分析的训练。实验结果表明,我们的NotaGPT-7 B在音乐理解方面取得了显着的改善,展示了NOTA和训练管道的有效性。我们的数据集在https://huggingface.co/datasets/MYTH-Lab/NOTA-dataset上开源。
摘要:Symbolic music is represented in two distinct forms: two-dimensional,visually intuitive score images, and one-dimensional, standardized textannotation sequences. While large language models have shown extraordinarypotential in music, current research has primarily focused on unimodal symbolsequence text. Existing general-domain visual language models still lack theability of music notation understanding. Recognizing this gap, we propose NOTA,the first large-scale comprehensive multimodal music notation dataset. Itconsists of 1,019,237 records, from 3 regions of the world, and contains 3tasks. Based on the dataset, we trained NotaGPT, a music notation visual largelanguage model. Specifically, we involve a pre-alignment training phase forcross-modal alignment between the musical notes depicted in music score imagesand their textual representation in ABC notation. Subsequent training phasesfocus on foundational music information extraction, followed by training onmusic notation analysis. Experimental results demonstrate that our NotaGPT-7Bachieves significant improvement on music understanding, showcasing theeffectiveness of NOTA and the training pipeline. Our datasets are open-sourcedat https://huggingface.co/datasets/MYTH-Lab/NOTA-dataset.

机器翻译由腾讯交互翻译提供,仅供参考