今日论文合集:cs.SD语音16篇,eess.AS音频处理19篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音
【1】Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition  and Translation
标题:低等级和稀疏模型融合用于多语言语音识别和翻译
链接:https://arxiv.org/abs/2502.17380
作者:Qiuming Zhao,  Guangzhi Sun,  Chao Zhang,  Mingxing Xu,  Thomas Fang Zheng
备注:13 pages, submitted to ACL 2025
摘要:语言多样性在语音到文本(S2 T)任务中提出了重大挑战,例如自动语音识别和翻译。传统的多任务训练方法旨在通过联合优化各种语言的多个语音识别和翻译任务来解决这个问题。虽然建立在这些策略基础上的Whisper等模型表现出了强大的性能,但它们仍然面临着高计算成本、语言干扰、次优训练配置和有限的可扩展性等问题。为了克服这些挑战,我们引入了LoRS-Merging(低秩和稀疏模型合并),这是一种新技术,旨在有效地集成在不同语言或任务上训练的模型,同时保持性能并减少计算开销。LoRS-Merging结合了低秩和稀疏修剪,以保留基本结构,同时消除冗余参数,减轻语言和任务干扰,并增强可扩展性。一系列语言的实验结果表明,LoRS-Merging显著优于传统的多语言多任务训练基线。我们的研究结果表明,模型合并,特别是LoRS合并,是对S2 T应用程序的传统多语言训练策略的可扩展和有效的补充。
摘要:Language diversity presents a significant challenge in speech-to-text (S2T)tasks, such as automatic speech recognition and translation. Traditionalmulti-task training approaches aim to address this by jointly optimizingmultiple speech recognition and translation tasks across various languages.While models like Whisper, built on these strategies, demonstrate strongperformance, they still face issues of high computational cost, languageinterference, suboptimal training configurations, and limited extensibility. Toovercome these challenges, we introduce LoRS-Merging (low-rank and sparse modelmerging), a novel technique designed to efficiently integrate models trained ondifferent languages or tasks while preserving performance and reducingcomputational overhead. LoRS-Merging combines low-rank and sparse pruning toretain essential structures while eliminating redundant parameters, mitigatinglanguage and task interference, and enhancing extensibility. Experimentalresults across a range of languages demonstrate that LoRS-Merging significantlyoutperforms conventional multi-lingual multi-task training baselines. Ourfindings suggest that model merging, particularly LoRS-Merging, is a scalableand effective complement to traditional multi-lingual training strategies forS2T applications.

【2】 Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning  Whisper on the JASMIN-CGN Corpus
标题:通过对JASMIN-CGN Corpus上的Whisper进行微调来提高荷兰语语音识别的包容性
链接:https://arxiv.org/abs/2502.17284
作者:Golshid Shekoufandeh,  Paul Boersma,  Antal van den Bosch
备注:ISCA/ITG Workshop on Diversity in Large Speech and Language Models
摘要:我们测试和研究的微调版本的耳语模型对儿童,老年人和非母语的荷兰语语音从JASMIN-CGN语料库的语音识别的变化。我们的主要目标是评估说话者的年龄和语言背景如何影响Whisper的表现。Whisper在对特定年龄和语言背景的亚群进行微调时,可以实现不同的单词错误率(WER)。微调性能明显优于zero-shot性能,实现了81%的本地儿童,72%的非本地儿童,67%的非本地成年人,65%的本地老年人的WER相对减少。我们的研究结果强调了在儿童、老年人和非母语者等代表性不足的亚群中训练像Whisper这样的语音识别模型的重要性。
摘要:We test and study the variation in speech recognition of fine-tuned versionsof the Whisper model on child, elderly and non-native Dutch speech from theJASMIN-CGN corpus. Our primary goal is to evaluate how speakers' age andlinguistic background influence Whisper's performance. Whisper achieves varyingWord Error Rates (WER) when fine-tuned on subpopulations of specific ages andlinguistic backgrounds. Fine-tuned performance is remarkably better thanzero-shot performance, achieving a relative reduction in WER of 81% for nativechildren, 72% for non-native children, 67% for non-native adults, and 65% fornative elderly people. Our findings underscore the importance of trainingspeech recognition models like Whisper on underrepresented subpopulations suchas children, the elderly, and non-native speakers.

【3】 Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
标题:白川音频:端到端语音交互的统一框架
链接:https://arxiv.org/abs/2502.17239
作者:Tianpeng Li,  Jun Liu,  Tao Zhang,  Yuanbo Fang,  Da Pan,  Mingrui Wang,  Zheng Liang,  Zehuan Li,  Mingan Lin,  Guosheng Dong,  Jianhua Xu,  Haoze Sun,  Zenan Zhou,  Weipeng Chen
摘要:我们介绍了百川音频,一个端到端的音频大语言模型,无缝集成音频理解和生成。它具有文本引导对齐语音生成机制,实现了具有理解和生成功能的实时语音交互。Baichuan-Audio利用预训练的ASR模型,然后以12.5 Hz的帧速率对语音进行多码本离散化。这种多码本设置确保语音令牌保留语义和声学信息。为了进一步增强建模,采用独立的音频头来处理音频令牌,有效地捕获它们的独特特征。为了减轻预训练期间的智能损失并保留LLM的原始功能,我们提出了一种两阶段预训练策略,该策略在增强音频建模的同时保持语言理解。在对齐之后,该模型在基于语音的实时对话中表现出色,并表现出出色的问答能力,证明了其通用性和效率。该模型在实时口语对话中表现出优越的性能,并具有较强的问答能力。我们的代码、模型和训练数据可在https://github.com/baichuan-inc/Baichuan-Audio上获得
摘要:We introduce Baichuan-Audio, an end-to-end audio large language model thatseamlessly integrates audio understanding and generation. It features atext-guided aligned speech generation mechanism, enabling real-time speechinteraction with both comprehension and generation capabilities. Baichuan-Audioleverages a pre-trained ASR model, followed by multi-codebook discretization ofspeech at a frame rate of 12.5 Hz. This multi-codebook setup ensures thatspeech tokens retain both semantic and acoustic information. To further enhancemodeling, an independent audio head is employed to process audio tokens,effectively capturing their unique characteristics. To mitigate the loss ofintelligence during pre-training and preserve the original capabilities of theLLM, we propose a two-stage pre-training strategy that maintains languageunderstanding while enhancing audio modeling. Following alignment, the modelexcels in real-time speech-based conversation and exhibits outstandingquestion-answering capabilities, demonstrating its versatility and efficiency.The proposed model demonstrates superior performance in real-time spokendialogue and exhibits strong question-answering abilities. Our code, model andtraining data are available at https://github.com/baichuan-inc/Baichuan-Audio

【4】 Supervised contrastive learning from weakly-labeled audio segments for  musical version matching
标题:从弱标记音频片段进行有监督的对比学习以进行音乐版本匹配
链接:https://arxiv.org/abs/2502.16936
作者:Joan Serrà,  R. Oguz Araz,  Dmitry Bogdanov,  Yuki Mitsufuji
备注:15 pages, 6 figures, 7 tables; includes Appendix
摘要:检测音乐版本(同一作品的不同演奏)是一项具有重要应用的具有挑战性的任务。由于地面实况性质,现有方法在音轨级匹配音乐版本(例如,整首歌)。但是,大多数应用程序需要在段级别(例如,20s块)。此外,现有的方法诉诸于分类和三重损失,忽视了可能带来有意义的改进的最近的损失。在本文中,我们提出了一种从弱注释片段中学习的方法,以及一种比研究良好的替代品表现更好的对比损失变体。前者是基于成对段距离减少,而后者修改现有的损失解耦,超参数和几何考虑。有了这两个要素,我们不仅在标准赛道级评价中取得了最先进的成绩,而且在分段级评价中也取得了突破性的成绩。我们相信,由于这里所解决的挑战的一般性,所提出的方法可以在音频或音乐版本匹配之外的领域中找到实用性。
摘要:Detecting musical versions (different renditions of the same piece) is achallenging task with important applications. Because of the ground truthnature, existing approaches match musical versions at the track level (e.g.,whole song). However, most applications require to match them at the segmentlevel (e.g., 20s chunks). In addition, existing approaches resort toclassification and triplet losses, disregarding more recent losses that couldbring meaningful improvements. In this paper, we propose a method to learn fromweakly annotated segments, together with a contrastive loss variant thatoutperforms well-studied alternatives. The former is based on pairwise segmentdistance reductions, while the latter modifies an existing loss followingdecoupling, hyper-parameter, and geometric considerations. With these twoelements, we do not only achieve state-of-the-art results in the standardtrack-level evaluation, but we also obtain a breakthrough performance in asegment-level evaluation. We believe that, due to the generality of thechallenges addressed here, the proposed methods may find utility in domainsbeyond audio or musical version matching.

【5】 ENACT-Heart -- ENsemble-based Assessment Using CNN and Transformer on  Heart Sounds
标题:ENACT-Heart --基于EN的评估使用CNN和Transformer来评估心弦
链接:https://arxiv.org/abs/2502.16914
作者:Jiho Han,  Adnan Shaout
备注:Accepted but not published in Global Digital Health Knowledge Exchange & Empowerment Conference (gDigiHealth.KEE)
摘要:本研究探讨了Vision Transformer(ViT)原理在音频分析中的应用,特别是专注于心音。本文介绍了ENACT-Heart -一种新的集成方法,该方法通过混合专家(MoE)框架利用卷积神经网络(CNN)和ViT的互补优势,实现了97.52%的显着分类准确率。这超过了ViT(93.88%)和CNN(95.45%)的单独贡献,证明了在心血管健康监测中提高诊断准确性的潜力。这些结果表明集成方法在提高心血管健康监测和诊断的分类性能方面的潜力。
摘要:This study explores the application of Vision Transformer (ViT) principles inaudio analysis, specifically focusing on heart sounds. This paper introducesENACT-Heart - a novel ensemble approach that leverages the complementarystrengths of Convolutional Neural Networks (CNN) and ViT through a Mixture ofExperts (MoE) framework, achieving a remarkable classification accuracy of97.52%. This outperforms the individual contributions of ViT (93.88%) and CNN(95.45%), demonstrating the potential for enhanced diagnostic accuracy incardiovascular health monitoring. These results demonstrate the potential ofensemble methods in enhancing classification performance for cardiovascularhealth monitoring and diagnosis.

【6】 AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
标题:AAD-LLM:神经注意力驱动的听觉场景理解
链接:https://arxiv.org/abs/2502.16794
作者:Xilin Jiang,  Sukru Samet Dindar,  Vishal Choudhari,  Stephan Bickel,  Ashesh Mehta,  Guy M McKhann,  Adeen Flinker,  Daniel Friedman,  Nima Mesgarani
摘要:听觉基础模型,包括听觉大语言模型(LLM),平等地处理所有声音输入,独立于听者感知。然而,人类的听觉感知具有内在的选择性:在复杂的听觉场景中,听者专注于特定的说话者,而忽略其他人。现有的模型不包含这种选择性,限制了它们生成感知对齐响应的能力。为了解决这个问题,我们介绍了意图知情的听觉场景理解(II-ASU)和听觉注意力驱动的LLM(AAD-LLM),一个原型系统,集成大脑信号来推断听者的注意力。AAD-LLM通过结合颅内脑电图(iEEG)记录来解码收听者正在关注的扬声器并相应地改进响应,从而扩展了听觉LLM。该模型首先从神经活动中预测出席的扬声器,然后在此推断的注意状态上产生响应。我们评估AAD-LLM的扬声器描述,语音转录和提取,并在多说话者的情况下回答问题,客观和主观的评级显示与听众的意图改善对齐。通过向意图感知听觉AI迈出第一步,这项工作探索了一种新的范式,在这种范式中,听众感知告知机器听力,为未来以听众为中心的听觉系统铺平了道路。演示和代码可用:https://aad-llm.github.io。
摘要:Auditory foundation models, including auditory large language models (LLMs),process all sound inputs equally, independent of listener perception. However,human auditory perception is inherently selective: listeners focus on specificspeakers while ignoring others in complex auditory scenes. Existing models donot incorporate this selectivity, limiting their ability to generateperception-aligned responses. To address this, we introduce Intention-InformedAuditory Scene Understanding (II-ASU) and present Auditory Attention-Driven LLM(AAD-LLM), a prototype system that integrates brain signals to infer listenerattention. AAD-LLM extends an auditory LLM by incorporating intracranialelectroencephalography (iEEG) recordings to decode which speaker a listener isattending to and refine responses accordingly. The model first predicts theattended speaker from neural activity, then conditions response generation onthis inferred attentional state. We evaluate AAD-LLM on speaker description,speech transcription and extraction, and question answering in multitalkerscenarios, with both objective and subjective ratings showing improvedalignment with listener intention. By taking a first step towardintention-aware auditory AI, this work explores a new paradigm where listenerperception informs machine listening, paving the way for futurelistener-centered auditory systems. Demo and code available:https://aad-llm.github.io.

【7】 Target Speaker Extraction through Comparing Noisy Positive and Negative  Audio Enrollments
标题:通过比较有噪音的积极和消极音频注册来提取目标说话人
链接:https://arxiv.org/abs/2502.16611
作者:Shitong Xu,  Yiyuan Yang,  Niki Trigoni,  Andrew Markham
备注:16 pages, 5 figures, appendix included
摘要:目标说话人提取的重点是从包含多个说话人的音频混合中分离出特定说话人的声音。为了提供有关目标说话者身份的信息,之前的工作利用干净的音频示例作为条件输入。然而,这种干净的音频示例并不总是容易获得的(例如,在鸡尾酒会上获得陌生人的声音的干净的音频示例而不远离嘈杂的环境是不切实际的)。有限的先前研究已经探索了从嘈杂的音频示例中提取目标说话者的特征,所述嘈杂的音频示例可以包括来自干扰说话者的重叠语音。在这项工作中,我们专注于目标扬声器提取时,多个扬声器存在于注册阶段,通过利用音频段之间的差异,其中目标扬声器说话(积极注册)和段,他们不是(消极注册)。实验结果表明,我们的模型架构和专用的预训练方法的有效性所提出的任务。我们的方法在建议的应用程序设置中实现了最先进的性能,并在具有挑战性和现实的场景中表现出很强的通用性。
摘要:Target speaker extraction focuses on isolating a specific speaker's voicefrom an audio mixture containing multiple speakers. To provide informationabout the target speaker's identity, prior works have utilized clean audioexamples as conditioning inputs. However, such clean audio examples are notalways readily available (e.g. It is impractical to obtain a clean audioexample of a stranger's voice at a cocktail party without stepping away fromthe noisy environment). Limited prior research has explored extracting thetarget speaker's characteristics from noisy audio examples, which may includeoverlapping speech from disturbing speakers. In this work, we focus on targetspeaker extraction when multiple speakers are present during the enrollmentstage, through leveraging differences between audio segments where the targetspeakers are speaking (Positive Enrollments) and segments where they are not(Negative Enrollments). Experiments show the effectiveness of our modelarchitecture and the dedicated pretraining method for the proposed task. Ourmethod achieves state-of-the-art performance in the proposed applicationsettings and demonstrates strong generalizability across challenging andrealistic scenarios.

【8】 Audio-FLAN: A Preliminary Release
标题:音频FLAN:初步发布
链接:https://arxiv.org/abs/2502.16584
作者:Liumeng Xue,  Ziya Zhou,  Jiahao Pan,  Zixuan Li,  Shuai Fan,  Yinghao Ma,  Sitong Cheng,  Dongchao Yang,  Haohan Guo,  Yujia Xiao,  Xinsheng Wang,  Zixuan Shen,  Chuanbo Zhu,  Xinshen Zhang,  Tianchi Liu,  Ruibin Yuan,  Zeyue Tian,  Haohe Liu,  Emmanouil Benetos,  Ge Zhang,  Yike Guo,  Wei Xue
摘要:音频标记化的最新进展显着增强了音频功能到大型语言模型(LLM)中的集成。然而,音频理解和生成通常被视为不同的任务,阻碍了真正统一的音频语言模型的发展。虽然指令调整在改善文本和视觉的泛化和zero-shot学习方面取得了显着的成功,但其在音频方面的应用在很大程度上尚未探索。一个主要障碍是缺乏统一音频理解和生成的综合数据集。为了解决这个问题,我们引入了Audio-FLAN,这是一个大规模的语音调整数据集,涵盖了语音,音乐和声音领域的80个不同任务,拥有超过1亿个实例。Audio-FLAN为统一的音频语言模型奠定了基础,该模型可以无缝地处理理解(例如,转录,理解)和生成(例如,语音、音乐、声音)任务。Audio-FLAN数据集在HuggingFace和GitHub上提供,并将不断更新。
摘要:Recent advancements in audio tokenization have significantly enhanced theintegration of audio capabilities into large language models (LLMs). However,audio understanding and generation are often treated as distinct tasks,hindering the development of truly unified audio-language models. Whileinstruction tuning has demonstrated remarkable success in improvinggeneralization and zero-shot learning across text and vision, its applicationto audio remains largely unexplored. A major obstacle is the lack ofcomprehensive datasets that unify audio understanding and generation. Toaddress this, we introduce Audio-FLAN, a large-scale instruction-tuning datasetcovering 80 diverse tasks across speech, music, and sound domains, with over100 million instances. Audio-FLAN lays the foundation for unifiedaudio-language models that can seamlessly handle both understanding (e.g.,transcription, comprehension) and generation (e.g., speech, music, sound) tasksacross a wide range of audio domains in a zero-shot manner. The Audio-FLANdataset is available on HuggingFace and GitHub and will be continuouslyupdated.

【9】 Improving Speech Enhancement by Cross- and Sub-band Processing with  State Space Model
标题:利用状态空间模型通过跨带和子带处理改进语音增强
链接:https://arxiv.org/abs/2502.16207
作者:Jizhen Li,  Weiping Tu,  Yuhong Yang,  Xinmeng Xu,  Yiqun Zhang,  Yanzhen Ren
摘要:最近,以Mamba为代表的状态空间模型(SSM)在包括语音增强在内的长期序列建模任务中表现出了显著的性能。然而,由于子带特征的实质性差异,将相同的SSM应用于所有子带限制了其推断能力。此外,当处理时频表示的每个时间帧时,SSM可能会忘记某些低能量的高频信息,使得高频带中的结构恢复具有挑战性。为此,我们提出了交叉和子带曼巴(CSMamba)。为了帮助SSM灵活地处理不同的子带特征,我们提出了一个带分割块,该块根据它们的信息相似性将全带分割成四个具有不同宽度的子带。然后,我们为每个子带分配独立的权重,从而减少了SSM的推理负担。此外,为了减轻遗忘的低能量信息的高频段的SSM,我们引入了一个频谱恢复块,从多个角度增强了跨带特征的表示。在DNS Challenge 2021数据集上的实验结果表明,CSMamba在三个客观评价指标中以较少的参数表现出优于几种最先进的(SOTA)语音增强方法。
摘要:Recently, the state space model (SSM) represented by Mamba has shownremarkable performance in long-term sequence modeling tasks, including speechenhancement. However, due to substantial differences in sub-band features,applying the same SSM to all sub-bands limits its inference capability.Additionally, when processing each time frame of the time-frequencyrepresentation, the SSM may forget certain high-frequency information of lowenergy, making the restoration of structure in the high-frequency bandschallenging. For this reason, we propose Cross- and Sub-band Mamba (CSMamba).To assist the SSM in handling different sub-band features flexibly, we proposea band split block that splits the full-band into four sub-bands with differentwidths based on their information similarity. We then allocate independentweights to each sub-band, thereby reducing the inference burden on the SSM.Furthermore, to mitigate the forgetting of low-energy information in thehigh-frequency bands by the SSM, we introduce a spectrum restoration block thatenhances the representation of the cross-band features from multipleperspectives. Experimental results on the DNS Challenge 2021 datasetdemonstrate that CSMamba outperforms several state-of-the-art (SOTA) speechenhancement methods in three objective evaluation metrics with fewerparameters.

【10】 Mind the Gap! Static and Interactive Evaluations of Large Audio Models
标题:注意差距!大型音频模型的静态和交互式评估
链接:https://arxiv.org/abs/2502.15919
作者:Minzhi Li,  William Barr Held,  Michael J Ryan,  Kunat Pipatanakul,  Potsawee Manakul,  Hao Zhu,  Diyi Yang
摘要:随着人工智能聊天机器人变得无处不在,语音交互提供了一种引人注目的方式,可以为语义和社交信号提供快速,高带宽的通信。这推动了对大型音频模型(LAM)的研究,以推动语音原生体验。然而,将LAM开发与用户目标相结合需要清楚地了解用户需求和偏好,以建立可靠的进度指标。本研究通过引入一种交互式方法来评估LAM并从484名参与者中收集7,500个LAM交互来应对这些挑战。通过对用户查询的主题建模,我们确定了音频接口的主要用例。然后,我们分析用户偏好排名和定性反馈,以确定哪些模型最符合用户需求。最后,我们评估了静态基准测试如何预测交互性能-我们的分析显示,没有一个单独的基准测试与交互结果密切相关(所有基准测试的$\tau \leq 0.33$)。虽然组合多个粗粒度特征会产生适度的预测能力($R^2$=$0.30$),但20个口语问答和年龄预测数据集中只有两个显示出显著的正相关性。这表明显然需要开发与用户偏好更好相关的LAM评估。
摘要:As AI chatbots become ubiquitous, voice interaction presents a compelling wayto enable rapid, high-bandwidth communication for both semantic and socialsignals. This has driven research into Large Audio Models (LAMs) to powervoice-native experiences. However, aligning LAM development with user goalsrequires a clear understanding of user needs and preferences to establishreliable progress metrics. This study addresses these challenges by introducingan interactive approach to evaluate LAMs and collecting 7,500 LAM interactionsfrom 484 participants. Through topic modeling of user queries, we identifyprimary use cases for audio interfaces. We then analyze user preferencerankings and qualitative feedback to determine which models best align withuser needs. Finally, we evaluate how static benchmarks predict interactiveperformance - our analysis reveals no individual benchmark strongly correlateswith interactive results ($\tau \leq 0.33$ for all benchmarks). While combiningmultiple coarse-grained features yields modest predictive power ($R^2$=$0.30$),only two out of twenty datasets on spoken question answering and age predictionshow significantly positive correlations. This suggests a clear need to developLAM evaluations that better correlate with user preferences.

【11】 Deriving Representative Structure from Music Corpora
标题:从音乐库中获取代表性结构
链接:https://arxiv.org/abs/2502.15849
作者:Ilana Shapiro,  Ruanqianqian (Lisa) Huang,  Zachary Novack,  Cheng-i Wang,  Hao-Wen Dong,  Taylor Berg-Kirkpatrick,  Shlomo Dubnov,  Sorin Lerner
备注:12 pages, 8 figures, 7 tables
摘要:西方音乐是一个天生的层次系统的结构相互作用的水平,从细粒度的旋律,以高层次的形式。为了分析音乐作品的整体性和多粒度,我们提出了一个统一的,层次化的元表示的音乐结构称为结构时间图(STG)。对于单个作品来说,STG是一种数据结构,它定义了逐渐精细的结构音乐特征的层次结构以及它们之间的时间关系。我们使用的STG,使一种新的方法来获得一个有代表性的结构摘要的音乐语料库,我们正式作为一个双重NP-难组合优化问题扩展的广义中值图问题。我们的方法首先应用模拟退火开发一个措施的两个音乐作品之间的结构距离植根于图同构。然后,我们的方法相结合的SMT求解器与嵌套模拟退火的结构距离的正式保证,产生一个结构合理的,代表性的质心STG为整个语料库的STG从个别作品。为了评估我们的方法,我们进行实验验证,结构距离准确区分音乐作品,并推导出质心准确地在结构上表征他们的语料库。
摘要:Western music is an innately hierarchical system of interacting levels ofstructure, from fine-grained melody to high-level form. In order to analyzemusic compositions holistically and at multiple granularities, we propose aunified, hierarchical meta-representation of musical structure called thestructural temporal graph (STG). For a single piece, the STG is a datastructure that defines a hierarchy of progressively finer structural musicalfeatures and the temporal relationships between them. We use the STG to enablea novel approach for deriving a representative structural summary of a musiccorpus, which we formalize as a dually NP-hard combinatorial optimizationproblem extending the Generalized Median Graph problem. Our approach firstapplies simulated annealing to develop a measure of structural distance betweentwo music pieces rooted in graph isomorphism. Our approach then combines theformal guarantees of SMT solvers with nested simulated annealing overstructural distances to produce a structurally sound, representative centroidSTG for an entire corpus of STGs from individual pieces. To evaluate ourapproach, we conduct experiments verifying that structural distance accuratelydifferentiates between music pieces, and that derived centroids accuratelystructurally characterize their corpora.

【12】 LACTOSE: Linear Array of Conditions, TOpologies with Separated  Error-backpropagation -- The Differentiable "IF" Conditional for  Differentiable Digital Signal Processing
链接:https://arxiv.org/abs/2502.15829
作者:Christopher Johann Clarke
备注:6 pages
摘要:将条件语句作为神经网络图的一部分存在困难(例如,如果input $> x$,则将输入传递给网络$N$)。这是由于不能通过分支条件反向传播。条件线性阵列、误差分离反向传播拓扑(LACTOSE)算法解决了这个问题,并允许有条件地使用监督学习模型的可用机器学习层。本文将LACTOSE算法应用于DDSP的一个简单的使用中,然而,重点是为DDSP的使用开发“if”条件。LACTOSE算法为每个用户指定的数值范围存储训练参数,并在预测期间动态加载参数。
摘要:There has been difficulty utilising conditional statements as part of theneural network graph (e.g. if input $> x$, pass input to network $N$). This isdue to the inability to backpropagate through branching conditions. The LinearArray of Conditions, TOpologies with Separated Error-backpropagation (LACTOSE)Algorithm addresses this issue and allows the conditional use of availablemachine learning layers for supervised learning models. In this paper, theLACTOSE algorithm is applied to a simple use of DDSP, however, the main pointis the development of the "if" conditional for DDSP use. The LACTOSE algorithmstores trained parameters for each user-specified numerical range and loads theparameters dynamically during prediction.

【13】 Slamming: Training a Speech Language Model on One GPU in a Day
标题:猛烈抨击:一天内在一个图形处理器上训练语音语言模型
链接:https://arxiv.org/abs/2502.15814
作者:Gallil Maimon,  Avishai Elmakies,  Yossi Adi
摘要:我们介绍了Slam,这是一种在24小时内在单个学术GPU上训练高质量语音语言模型(SLM)的方法。我们通过对模型初始化和架构、合成训练数据、使用合成数据进行偏好优化以及调整所有其他组件的实证分析来实现这一点。我们根据经验证明,这种训练方法也可以很好地扩展,更多的计算可以以一小部分的计算成本获得与领先的SLM相当的结果。我们希望这些见解将使可持续土地管理培训和研究更容易获得。在SLM缩放律的背景下,我们的结果远远优于预测的计算最佳性能,对SLM的可行性给予了乐观的看法。请参阅代码、数据、模型、示例-https://pages.cs.huji.ac.il/adiyoss-lab/slamming。
摘要:We introduce Slam, a recipe for training high-quality Speech Language Models(SLMs) on a single academic GPU in 24 hours. We do so through empiricalanalysis of model initialisation and architecture, synthetic training data,preference optimisation with synthetic data and tweaking all other components.We empirically demonstrate that this training recipe also scales well with morecompute getting results on par with leading SLMs in a fraction of the computecost. We hope these insights will make SLM training and research moreaccessible. In the context of SLM scaling laws, our results far outperformpredicted compute optimal performance, giving an optimistic view to SLMfeasibility. See code, data, models, samples at -https://pages.cs.huji.ac.il/adiyoss-lab/slamming .

【14】 voc2vec: A Foundation Model for Non-Verbal Vocalization
标题:voc2vec:非言语发声的基础模型
链接:https://arxiv.org/abs/2502.16298
作者:Alkis Koudounas,  Moreno La Quatra,  Marco Sabato Siniscalchi,  Elena Baralis
备注:Accepted at ICASSP 2025
摘要:语音基础模型在语音相关任务中表现出了卓越的能力。然而,这些模型经常与非语言音频数据(诸如发声、婴儿哭声等)作斗争,这对于各种现实世界的应用是至关重要的。音频基础模型可以很好地处理非语音数据,但也无法捕捉非语言人类声音的细微差别。在这项工作中,我们的目标是克服上述缺点,并提出一种新的基础模型,称为voc2vec,专门为非语言人类数据设计,利用专门的开源非语言音频数据集。我们采用了10个数据集的集合,涵盖了大约125小时的非语言音频。实验结果表明,voc2vec在非言语发声分类中是有效的,其性能优于传统的语音和音频基础模型。此外,voc2vec在六个不同的基准数据集上的表现一直优于强基线,即OpenSmile和emotion2vec。据作者所知,voc2vec是第一个用于发声任务的通用表示模型。
摘要:Speech foundation models have demonstrated exceptional capabilities inspeech-related tasks. Nevertheless, these models often struggle with non-verbalaudio data, such as vocalizations, baby crying, etc., which are critical forvarious real-world applications. Audio foundation models well handle non-speechdata but also fail to capture the nuanced features of non-verbal human sounds.In this work, we aim to overcome the above shortcoming and propose a novelfoundation model, termed voc2vec, specifically designed for non-verbal humandata leveraging exclusively open-source non-verbal audio datasets. We employ acollection of 10 datasets covering around 125 hours of non-verbal audio.Experimental results prove that voc2vec is effective in non-verbal vocalizationclassification, and it outperforms conventional speech and audio foundationmodels. Moreover, voc2vec consistently outperforms strong baselines, namelyOpenSmile and emotion2vec, on six different benchmark datasets. To the best ofthe authors' knowledge, voc2vec is the first universal representation model forvocalization tasks.

【15】 Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
标题:使用神经音频编解码器的连续嵌入的语音增强
链接:https://arxiv.org/abs/2502.16240
作者:Haoyang Li,  Jia Qi Yip,  Tianyu Fan,  Eng Siong Chng
备注:Accepted to ICASSP 2025
摘要:神经音频编解码器(NAC)模型的最新进展激发了它们在各种语音处理任务中的使用,包括语音增强(SE)。在这项工作中,我们提出了一种新的,有效的SE方法,利用预训练的NAC编码器的预量化输出。与之前基于NAC的SE方法不同,该方法使用语言模型(LM)处理离散语音令牌,我们在预训练NAC的连续嵌入空间内执行SE,该空间沿时间维度高度压缩以实现高效表示。我们的轻量级SE模型通过嵌入级损失进行了优化,提供了与在较大数据集上训练的SE基线相当的结果,其实时系数显著降低至0.005。此外,我们的方法实现了3.94的低GMAC,在模拟的基于云的音频传输环境中,与Sepformer相比,复杂性降低了18倍。这项工作突出了一个新的,高效的基于NAC的SE解决方案,特别适合于云应用程序,其中NAC用于在传输前压缩音频。  版权所有20 XX IEEE。允许个人使用本材料。在任何当前或将来的媒体上进行所有其他用途必须获得IEEE的许可,包括出于广告或促销目的重印/重新发布本材料、创建新的集体作品、转售或重新分发到服务器或列表、或在其他作品中重新使用本作品的任何受版权保护的组件。
摘要:Recent advancements in Neural Audio Codec (NAC) models have inspired theiruse in various speech processing tasks, including speech enhancement (SE). Inthis work, we propose a novel, efficient SE approach by leveraging thepre-quantization output of a pretrained NAC encoder. Unlike prior NAC-based SEmethods, which process discrete speech tokens using Language Models (LMs), weperform SE within the continuous embedding space of the pretrained NAC, whichis highly compressed along the time dimension for efficient representation. Ourlightweight SE model, optimized through an embedding-level loss, deliversresults comparable to SE baselines trained on larger datasets, with asignificantly lower real-time factor of 0.005. Additionally, our methodachieves a low GMAC of 3.94, reducing complexity 18-fold compared to Sepformerin a simulated cloud-based audio transmission environment. This work highlightsa new, efficient NAC-based SE solution, particularly suitable for cloudapplications where NAC is used to compress audio before transmission. Copyright 20XX IEEE. Personal use of this material is permitted. Permissionfrom IEEE must be obtained for all other uses, in any current or future media,including reprinting/republishing this material for advertising or promotionalpurposes, creating new collective works, for resale or redistribution toservers or lists, or reuse of any copyrighted component of this work in otherworks.

【16】 Multizone sound field reproduction with  direction-of-arrival-distribution-based regularization and its application to  binaural-centered mode-matching
标题:基于到达方向分布的规则化的多区场再现及其在双耳中心模式匹配中的应用
链接:https://arxiv.org/abs/2502.16213
作者:Ryo Matsuda,  Makoto Otani
备注:Presented at International Congress on Acoustics (ICA) 2022
摘要:在高阶高保真度立体声响复制(声场再现的框架)中,次级源驱动信号通常通过正则化模式匹配来获得。作者提出了一种正则化技术的基础上的到达方向(DoA)分布的波前在主声场。这种基于DoA分布的正则化使得能够抑制在远离主源方向的方向上的次级源的过大驱动信号增益。这提高了远离再现中心的区域处的再现精度。首先,本研究将基于DOA分布的正则化应用于基于加法定理的多区域声场再现。此外,规则化的多区域声场再现被扩展到双耳中心模式匹配(BCMM),该模式匹配产生两个再现点,每个耳朵一个,以避免由于较高频率下的最佳点缩小而导致再现准确度下降。自由场和双耳模拟进行了数值研究的有效性DoA分布为基础的正则化的多区域声场再现和BCMM。
摘要:In higher-order Ambisonics, a framework for sound field reproduction,secondary-source driving signals are generally obtained by regularized modematching. The authors have proposed a regularization technique based ondirection-of-arrival (DoA) distribution of wavefronts in the primary soundfield. Such DoA-distribution-based regularization enables a suppression ofexcessively large driving signal gains for secondary sources that are in thedirections far from the primary source direction. This improves thereproduction accuracy at regions away from the reproduction center. First, thisstudy applies the DoA-distribution-based regularization to a multizone soundfield reproduction based on the addition theorem. Furthermore, the regularizedmultizone sound field reproduction is extended to a binaural-centered modematching (BCMM), which produces two reproduction points, one at each ear, toavoid a degraded reproduction accuracy due to a shrinking sweet spot at higherfrequencies. Free-field and binaural simulations were numerically performed toexamine the effectiveness of the DoA-distribution-based regularization on themultizone sound field reproduction and the BCMM.

eess.AS音频处理

【1】 Balancing Speech Understanding and Generation Using Continual  Pre-training for Codec-based Speech LLM
标题:使用基于编解码的语音LLM的连续预训练平衡语音理解和生成
链接:https://arxiv.org/abs/2502.16897
作者:Jiatong Shi,  Chunlei Zhang,  Jinchuan Tian,  Junrui Ni,  Hao Zhang,  Shinji Watanabe,  Dong Yu
摘要:最近的努力已经将文本LLM扩展到语音域。然而,一个关键的挑战仍然存在,即在将声学丰富的基于编解码器的表示集成到最初基于文本训练的模型中时,平衡语音理解和生成,同时避免灾难性遗忘。在这项工作中,我们提出了一种新的方法,利用连续预训练(CPT)的预训练的文本LLM创建一个基于编解码器的语音语言模型。这种策略减轻了文本和语音之间的模态差距,保留了原始模型的语言推理,同时实现了高保真语音合成。我们通过跨多个任务的广泛实验验证了我们的方法,包括自动语音识别,文本到语音,语音到文本翻译和语音到语音翻译(S2ST),证明我们的模型实现了卓越的TTS性能,特别是第一个基于神经编解码器的端到端S2ST系统。
摘要:Recent efforts have extended textual LLMs to the speech domain. Yet, a keychallenge remains, which is balancing speech understanding and generation whileavoiding catastrophic forgetting when integrating acoustically rich codec-basedrepresentations into models originally trained on text. In this work, wepropose a novel approach that leverages continual pre-training (CPT) on apre-trained textual LLM to create a codec-based speech language model. Thisstrategy mitigates the modality gap between text and speech, preserving thelinguistic reasoning of the original model while enabling high-fidelity speechsynthesis. We validate our approach with extensive experiments across multipletasks, including automatic speech recognition, text-to-speech, speech-to-texttranslation, and speech-to-speech translation (S2ST), demonstrating that ourmodel achieves superior TTS performance and, notably, the first end-to-end S2STsystem based on neural codecs.

【2】 voc2vec: A Foundation Model for Non-Verbal Vocalization
标题:voc2vec:非言语发声的基础模型
链接:https://arxiv.org/abs/2502.16298
作者:Alkis Koudounas,  Moreno La Quatra,  Marco Sabato Siniscalchi,  Elena Baralis
备注:Accepted at ICASSP 2025
摘要:语音基础模型在语音相关任务中表现出了卓越的能力。然而,这些模型经常与非语言音频数据(诸如发声、婴儿哭声等)作斗争,这对于各种现实世界的应用是至关重要的。音频基础模型可以很好地处理非语音数据,但也无法捕捉非语言人类声音的细微差别。在这项工作中,我们的目标是克服上述缺点,并提出一种新的基础模型,称为voc2vec,专门为非语言人类数据设计,利用专门的开源非语言音频数据集。我们采用了10个数据集的集合,涵盖了大约125小时的非语言音频。实验结果表明,voc2vec在非言语发声分类中是有效的,其性能优于传统的语音和音频基础模型。此外,voc2vec在六个不同的基准数据集上的表现一直优于强基线,即OpenSmile和emotion2vec。据作者所知,voc2vec是第一个用于发声任务的通用表示模型。
摘要:Speech foundation models have demonstrated exceptional capabilities inspeech-related tasks. Nevertheless, these models often struggle with non-verbalaudio data, such as vocalizations, baby crying, etc., which are critical forvarious real-world applications. Audio foundation models well handle non-speechdata but also fail to capture the nuanced features of non-verbal human sounds.In this work, we aim to overcome the above shortcoming and propose a novelfoundation model, termed voc2vec, specifically designed for non-verbal humandata leveraging exclusively open-source non-verbal audio datasets. We employ acollection of 10 datasets covering around 125 hours of non-verbal audio.Experimental results prove that voc2vec is effective in non-verbal vocalizationclassification, and it outperforms conventional speech and audio foundationmodels. Moreover, voc2vec consistently outperforms strong baselines, namelyOpenSmile and emotion2vec, on six different benchmark datasets. To the best ofthe authors' knowledge, voc2vec is the first universal representation model forvocalization tasks.

【3】 Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
标题:使用神经音频编解码器的连续嵌入的语音增强
链接:https://arxiv.org/abs/2502.16240
作者:Haoyang Li,  Jia Qi Yip,  Tianyu Fan,  Eng Siong Chng
备注:Accepted to ICASSP 2025
摘要:神经音频编解码器(NAC)模型的最新进展激发了它们在各种语音处理任务中的使用,包括语音增强(SE)。在这项工作中,我们提出了一种新的,有效的SE方法,利用预训练的NAC编码器的预量化输出。与之前基于NAC的SE方法不同,该方法使用语言模型(LM)处理离散语音令牌,我们在预训练NAC的连续嵌入空间内执行SE,该空间沿时间维度高度压缩以实现高效表示。我们的轻量级SE模型通过嵌入级损失进行了优化,提供了与在较大数据集上训练的SE基线相当的结果,其实时系数显著降低至0.005。此外,我们的方法实现了3.94的低GMAC,在模拟的基于云的音频传输环境中,与Sepformer相比,复杂性降低了18倍。这项工作突出了一个新的,高效的基于NAC的SE解决方案,特别适合于云应用程序,其中NAC用于在传输前压缩音频。  版权所有20 XX IEEE。允许个人使用本材料。在任何当前或将来的媒体上进行所有其他用途必须获得IEEE的许可,包括出于广告或促销目的重印/重新发布本材料、创建新的集体作品、转售或重新分发到服务器或列表、或在其他作品中重新使用本作品的任何受版权保护的组件。
摘要:Recent advancements in Neural Audio Codec (NAC) models have inspired theiruse in various speech processing tasks, including speech enhancement (SE). Inthis work, we propose a novel, efficient SE approach by leveraging thepre-quantization output of a pretrained NAC encoder. Unlike prior NAC-based SEmethods, which process discrete speech tokens using Language Models (LMs), weperform SE within the continuous embedding space of the pretrained NAC, whichis highly compressed along the time dimension for efficient representation. Ourlightweight SE model, optimized through an embedding-level loss, deliversresults comparable to SE baselines trained on larger datasets, with asignificantly lower real-time factor of 0.005. Additionally, our methodachieves a low GMAC of 3.94, reducing complexity 18-fold compared to Sepformerin a simulated cloud-based audio transmission environment. This work highlightsa new, efficient NAC-based SE solution, particularly suitable for cloudapplications where NAC is used to compress audio before transmission. Copyright 20XX IEEE. Personal use of this material is permitted. Permissionfrom IEEE must be obtained for all other uses, in any current or future media,including reprinting/republishing this material for advertising or promotionalpurposes, creating new collective works, for resale or redistribution toservers or lists, or reuse of any copyrighted component of this work in otherworks.

【4】 Multizone sound field reproduction with  direction-of-arrival-distribution-based regularization and its application to  binaural-centered mode-matching
标题:基于到达方向分布的规则化的多区场再现及其在双耳中心模式匹配中的应用
链接:https://arxiv.org/abs/2502.16213
作者:Ryo Matsuda,  Makoto Otani
备注:Presented at International Congress on Acoustics (ICA) 2022
摘要:在高阶高保真度立体声响复制(声场再现的框架)中,次级源驱动信号通常通过正则化模式匹配来获得。作者提出了一种正则化技术的基础上的到达方向(DoA)分布的波前在主声场。这种基于DoA分布的正则化使得能够抑制在远离主源方向的方向上的次级源的过大驱动信号增益。这提高了远离再现中心的区域处的再现精度。首先,本研究将基于DOA分布的正则化应用于基于加法定理的多区域声场再现。此外,规则化的多区域声场再现被扩展到双耳中心模式匹配(BCMM),该模式匹配产生两个再现点,每个耳朵一个,以避免由于较高频率下的最佳点缩小而导致再现准确度下降。自由场和双耳模拟进行了数值研究的有效性DoA分布为基础的正则化的多区域声场再现和BCMM。
摘要:In higher-order Ambisonics, a framework for sound field reproduction,secondary-source driving signals are generally obtained by regularized modematching. The authors have proposed a regularization technique based ondirection-of-arrival (DoA) distribution of wavefronts in the primary soundfield. Such DoA-distribution-based regularization enables a suppression ofexcessively large driving signal gains for secondary sources that are in thedirections far from the primary source direction. This improves thereproduction accuracy at regions away from the reproduction center. First, thisstudy applies the DoA-distribution-based regularization to a multizone soundfield reproduction based on the addition theorem. Furthermore, the regularizedmultizone sound field reproduction is extended to a binaural-centered modematching (BCMM), which produces two reproduction points, one at each ear, toavoid a degraded reproduction accuracy due to a shrinking sweet spot at higherfrequencies. Free-field and binaural simulations were numerically performed toexamine the effectiveness of the DoA-distribution-based regularization on themultizone sound field reproduction and the BCMM.

【5】 Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition  and Translation
标题:低等级和稀疏模型融合用于多语言语音识别和翻译
链接:https://arxiv.org/abs/2502.17380
作者:Qiuming Zhao,  Guangzhi Sun,  Chao Zhang,  Mingxing Xu,  Thomas Fang Zheng
备注:13 pages, submitted to ACL 2025
摘要:语言多样性在语音到文本(S2 T)任务中提出了重大挑战,例如自动语音识别和翻译。传统的多任务训练方法旨在通过联合优化各种语言的多个语音识别和翻译任务来解决这个问题。虽然建立在这些策略基础上的Whisper等模型表现出了强大的性能,但它们仍然面临着高计算成本、语言干扰、次优训练配置和有限的可扩展性等问题。为了克服这些挑战,我们引入了LoRS-Merging(低秩和稀疏模型合并),这是一种新技术,旨在有效地集成在不同语言或任务上训练的模型,同时保持性能并减少计算开销。LoRS-Merging结合了低秩和稀疏修剪,以保留基本结构,同时消除冗余参数,减轻语言和任务干扰,并增强可扩展性。一系列语言的实验结果表明,LoRS-Merging显著优于传统的多语言多任务训练基线。我们的研究结果表明,模型合并,特别是LoRS-Merging,是对S2 T应用程序传统多语言培训策略的可扩展且有效的补充。
摘要:Language diversity presents a significant challenge in speech-to-text (S2T)tasks, such as automatic speech recognition and translation. Traditionalmulti-task training approaches aim to address this by jointly optimizingmultiple speech recognition and translation tasks across various languages.While models like Whisper, built on these strategies, demonstrate strongperformance, they still face issues of high computational cost, languageinterference, suboptimal training configurations, and limited extensibility. Toovercome these challenges, we introduce LoRS-Merging (low-rank and sparse modelmerging), a novel technique designed to efficiently integrate models trained ondifferent languages or tasks while preserving performance and reducingcomputational overhead. LoRS-Merging combines low-rank and sparse pruning toretain essential structures while eliminating redundant parameters, mitigatinglanguage and task interference, and enhancing extensibility. Experimentalresults across a range of languages demonstrate that LoRS-Merging significantlyoutperforms conventional multi-lingual multi-task training baselines. Ourfindings suggest that model merging, particularly LoRS-Merging, is a scalableand effective complement to traditional multi-lingual training strategies forS2T applications.

【6】 Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning  Whisper on the JASMIN-CGN Corpus
标题:通过对JASMIN-CGN Corpus上的Whisper进行微调来提高荷兰语语音识别的包容性
链接:https://arxiv.org/abs/2502.17284
作者:Golshid Shekoufandeh,  Paul Boersma,  Antal van den Bosch
备注:ISCA/ITG Workshop on Diversity in Large Speech and Language Models
摘要:我们测试和研究的微调版本的耳语模型对儿童,老年人和非母语的荷兰语语音从JASMIN-CGN语料库的语音识别的变化。我们的主要目标是评估说话者的年龄和语言背景如何影响Whisper的表现。Whisper在对特定年龄和语言背景的亚群进行微调时,可以实现不同的单词错误率(WER)。微调性能明显优于zero-shot性能,实现了81%的本地儿童,72%的非本地儿童,67%的非本地成年人,65%的本地老年人的WER相对减少。我们的研究结果强调了在儿童、老年人和非母语者等代表性不足的亚群中训练像Whisper这样的语音识别模型的重要性。
摘要:We test and study the variation in speech recognition of fine-tuned versionsof the Whisper model on child, elderly and non-native Dutch speech from theJASMIN-CGN corpus. Our primary goal is to evaluate how speakers' age andlinguistic background influence Whisper's performance. Whisper achieves varyingWord Error Rates (WER) when fine-tuned on subpopulations of specific ages andlinguistic backgrounds. Fine-tuned performance is remarkably better thanzero-shot performance, achieving a relative reduction in WER of 81% for nativechildren, 72% for non-native children, 67% for non-native adults, and 65% fornative elderly people. Our findings underscore the importance of trainingspeech recognition models like Whisper on underrepresented subpopulations suchas children, the elderly, and non-native speakers.

【7】 Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
标题:白川音频:端到端语音交互的统一框架
链接:https://arxiv.org/abs/2502.17239
作者:Tianpeng Li,  Jun Liu,  Tao Zhang,  Yuanbo Fang,  Da Pan,  Mingrui Wang,  Zheng Liang,  Zehuan Li,  Mingan Lin,  Guosheng Dong,  Jianhua Xu,  Haoze Sun,  Zenan Zhou,  Weipeng Chen
摘要:我们介绍了百川音频,一个端到端的音频大语言模型,无缝集成音频理解和生成。它具有文本引导对齐语音生成机制,实现了具有理解和生成功能的实时语音交互。百川音频利用预训练的ASR模型,然后以12.5 Hz的帧率对语音进行多码本离散化。这种多码本设置确保语音令牌保留语义和声学信息。为了进一步增强建模,采用独立的音频头来处理音频令牌,有效地捕获它们的独特特征。为了减轻预训练期间的智能损失并保留LLM的原始功能,我们提出了一种两阶段预训练策略,该策略在增强音频建模的同时保持语言理解。在对齐之后,该模型在基于语音的实时对话中表现出色,并表现出出色的问答能力,证明了其通用性和效率。该模型在实时口语对话中表现出优越的性能,并具有较强的问答能力。我们的代码、模型和训练数据可在https://github.com/baichuan-inc/Baichuan-Audio上获取
摘要:We introduce Baichuan-Audio, an end-to-end audio large language model thatseamlessly integrates audio understanding and generation. It features atext-guided aligned speech generation mechanism, enabling real-time speechinteraction with both comprehension and generation capabilities. Baichuan-Audioleverages a pre-trained ASR model, followed by multi-codebook discretization ofspeech at a frame rate of 12.5 Hz. This multi-codebook setup ensures thatspeech tokens retain both semantic and acoustic information. To further enhancemodeling, an independent audio head is employed to process audio tokens,effectively capturing their unique characteristics. To mitigate the loss ofintelligence during pre-training and preserve the original capabilities of theLLM, we propose a two-stage pre-training strategy that maintains languageunderstanding while enhancing audio modeling. Following alignment, the modelexcels in real-time speech-based conversation and exhibits outstandingquestion-answering capabilities, demonstrating its versatility and efficiency.The proposed model demonstrates superior performance in real-time spokendialogue and exhibits strong question-answering abilities. Our code, model andtraining data are available at https://github.com/baichuan-inc/Baichuan-Audio

【8】 Supervised contrastive learning from weakly-labeled audio segments for  musical version matching
标题:从弱标记音频片段进行有监督的对比学习以进行音乐版本匹配
链接:https://arxiv.org/abs/2502.16936
作者:Joan Serrà,  R. Oguz Araz,  Dmitry Bogdanov,  Yuki Mitsufuji
备注:15 pages, 6 figures, 7 tables; includes Appendix
摘要:检测音乐版本(同一作品的不同演奏)是一项具有重要应用的具有挑战性的任务。由于地面实况性质,现有方法在音轨级匹配音乐版本(例如,整首歌)。但是,大多数应用程序需要在段级别(例如,20s块)。此外,现有的方法诉诸于分类和三重损失,忽视了可能带来有意义的改进的最近的损失。在本文中,我们提出了一种从弱注释片段中学习的方法,以及一种比研究良好的替代品表现更好的对比损失变体。前者是基于成对段距离减少,而后者修改现有的损失解耦,超参数和几何考虑。有了这两个要素,我们不仅在标准赛道级评价中取得了最先进的成绩,而且在分段级评价中也取得了突破性的成绩。我们相信,由于这里所解决的挑战的一般性,所提出的方法可以在音频或音乐版本匹配之外的领域中找到实用性。
摘要:Detecting musical versions (different renditions of the same piece) is achallenging task with important applications. Because of the ground truthnature, existing approaches match musical versions at the track level (e.g.,whole song). However, most applications require to match them at the segmentlevel (e.g., 20s chunks). In addition, existing approaches resort toclassification and triplet losses, disregarding more recent losses that couldbring meaningful improvements. In this paper, we propose a method to learn fromweakly annotated segments, together with a contrastive loss variant thatoutperforms well-studied alternatives. The former is based on pairwise segmentdistance reductions, while the latter modifies an existing loss followingdecoupling, hyper-parameter, and geometric considerations. With these twoelements, we do not only achieve state-of-the-art results in the standardtrack-level evaluation, but we also obtain a breakthrough performance in asegment-level evaluation. We believe that, due to the generality of thechallenges addressed here, the proposed methods may find utility in domainsbeyond audio or musical version matching.

【9】 ENACT-Heart -- ENsemble-based Assessment Using CNN and Transformer on  Heart Sounds
标题:ENACT-Heart --基于EN的评估使用CNN和Transformer来评估心弦
链接:https://arxiv.org/abs/2502.16914
作者:Jiho Han,  Adnan Shaout
备注:Accepted but not published in Global Digital Health Knowledge Exchange & Empowerment Conference (gDigiHealth.KEE)
摘要:本研究探讨了Vision Transformer(ViT)原理在音频分析中的应用,特别是专注于心音。本文介绍了ENACT-Heart -一种新的集成方法,该方法通过混合专家(MoE)框架利用卷积神经网络(CNN)和ViT的互补优势,实现了97.52%的显着分类准确率。这超过了ViT(93.88%)和CNN(95.45%)的单独贡献,证明了在心血管健康监测中提高诊断准确性的潜力。这些结果表明集成方法在提高心血管健康监测和诊断的分类性能方面的潜力。
摘要:This study explores the application of Vision Transformer (ViT) principles inaudio analysis, specifically focusing on heart sounds. This paper introducesENACT-Heart - a novel ensemble approach that leverages the complementarystrengths of Convolutional Neural Networks (CNN) and ViT through a Mixture ofExperts (MoE) framework, achieving a remarkable classification accuracy of97.52%. This outperforms the individual contributions of ViT (93.88%) and CNN(95.45%), demonstrating the potential for enhanced diagnostic accuracy incardiovascular health monitoring. These results demonstrate the potential ofensemble methods in enhancing classification performance for cardiovascularhealth monitoring and diagnosis.

【10】 AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
标题:AAD-LLM:神经注意力驱动的听觉场景理解
链接:https://arxiv.org/abs/2502.16794
作者:Xilin Jiang,  Sukru Samet Dindar,  Vishal Choudhari,  Stephan Bickel,  Ashesh Mehta,  Guy M McKhann,  Adeen Flinker,  Daniel Friedman,  Nima Mesgarani
摘要:听觉基础模型,包括听觉大语言模型(LLM),平等地处理所有声音输入,独立于听者感知。然而,人类的听觉感知具有内在的选择性:在复杂的听觉场景中,听者专注于特定的说话者,而忽略其他人。现有的模型不包含这种选择性,限制了它们生成感知对齐响应的能力。为了解决这个问题,我们介绍了意图知情的听觉场景理解(II-ASU)和听觉注意力驱动的LLM(AAD-LLM),一个原型系统,集成大脑信号来推断听者的注意力。AAD-LLM通过结合颅内脑电图(iEEG)记录来解码收听者正在关注的扬声器并相应地改进响应,从而扩展了听觉LLM。该模型首先从神经活动中预测出席的扬声器,然后在此推断的注意状态上产生响应。我们评估AAD-LLM的扬声器描述,语音转录和提取,并在多说话者的情况下回答问题,客观和主观的评级显示与听众的意图改善对齐。通过向意图感知听觉AI迈出第一步,这项工作探索了一种新的范式,在这种范式中,听众感知告知机器听力,为未来以听众为中心的听觉系统铺平了道路。演示和代码可用:https://aad-llm.github.io。
摘要:Auditory foundation models, including auditory large language models (LLMs),process all sound inputs equally, independent of listener perception. However,human auditory perception is inherently selective: listeners focus on specificspeakers while ignoring others in complex auditory scenes. Existing models donot incorporate this selectivity, limiting their ability to generateperception-aligned responses. To address this, we introduce Intention-InformedAuditory Scene Understanding (II-ASU) and present Auditory Attention-Driven LLM(AAD-LLM), a prototype system that integrates brain signals to infer listenerattention. AAD-LLM extends an auditory LLM by incorporating intracranialelectroencephalography (iEEG) recordings to decode which speaker a listener isattending to and refine responses accordingly. The model first predicts theattended speaker from neural activity, then conditions response generation onthis inferred attentional state. We evaluate AAD-LLM on speaker description,speech transcription and extraction, and question answering in multitalkerscenarios, with both objective and subjective ratings showing improvedalignment with listener intention. By taking a first step towardintention-aware auditory AI, this work explores a new paradigm where listenerperception informs machine listening, paving the way for futurelistener-centered auditory systems. Demo and code available:https://aad-llm.github.io.

【11】 Target Speaker Extraction through Comparing Noisy Positive and Negative  Audio Enrollments
标题:通过比较有噪音的积极和消极音频注册来提取目标说话人
链接:https://arxiv.org/abs/2502.16611
作者:Shitong Xu,  Yiyuan Yang,  Niki Trigoni,  Andrew Markham
备注:16 pages, 5 figures, appendix included
摘要:目标说话人提取的重点是从包含多个说话人的音频混合中分离出特定说话人的声音。为了提供有关目标说话者身份的信息,之前的工作利用干净的音频示例作为条件输入。然而,这种干净的音频示例并不总是容易获得的(例如,在鸡尾酒会上获得陌生人的声音的干净的音频示例而不远离嘈杂的环境是不切实际的)。有限的先前研究已经探索了从嘈杂的音频示例中提取目标说话者的特征,所述嘈杂的音频示例可以包括来自干扰说话者的重叠语音。在这项工作中,我们专注于目标扬声器提取时,多个扬声器存在于注册阶段,通过利用音频段之间的差异,其中目标扬声器说话(积极注册)和段,他们不是(消极注册)。实验结果表明,我们的模型架构和专用的预训练方法的有效性所提出的任务。我们的方法在建议的应用程序设置中实现了最先进的性能,并在具有挑战性和现实的场景中表现出很强的通用性。
摘要:Target speaker extraction focuses on isolating a specific speaker's voicefrom an audio mixture containing multiple speakers. To provide informationabout the target speaker's identity, prior works have utilized clean audioexamples as conditioning inputs. However, such clean audio examples are notalways readily available (e.g. It is impractical to obtain a clean audioexample of a stranger's voice at a cocktail party without stepping away fromthe noisy environment). Limited prior research has explored extracting thetarget speaker's characteristics from noisy audio examples, which may includeoverlapping speech from disturbing speakers. In this work, we focus on targetspeaker extraction when multiple speakers are present during the enrollmentstage, through leveraging differences between audio segments where the targetspeakers are speaking (Positive Enrollments) and segments where they are not(Negative Enrollments). Experiments show the effectiveness of our modelarchitecture and the dedicated pretraining method for the proposed task. Ourmethod achieves state-of-the-art performance in the proposed applicationsettings and demonstrates strong generalizability across challenging andrealistic scenarios.

【12】 Audio-FLAN: A Preliminary Release
标题:音频FLAN:初步发布
链接:https://arxiv.org/abs/2502.16584
作者:Liumeng Xue,  Ziya Zhou,  Jiahao Pan,  Zixuan Li,  Shuai Fan,  Yinghao Ma,  Sitong Cheng,  Dongchao Yang,  Haohan Guo,  Yujia Xiao,  Xinsheng Wang,  Zixuan Shen,  Chuanbo Zhu,  Xinshen Zhang,  Tianchi Liu,  Ruibin Yuan,  Zeyue Tian,  Haohe Liu,  Emmanouil Benetos,  Ge Zhang,  Yike Guo,  Wei Xue
摘要:音频标记化的最新进展显着增强了音频功能到大型语言模型(LLM)中的集成。然而,音频理解和生成通常被视为不同的任务,阻碍了真正统一的音频语言模型的发展。虽然指令调整在改善文本和视觉的泛化和zero-shot学习方面取得了显着的成功,但其在音频方面的应用在很大程度上尚未探索。一个主要障碍是缺乏统一音频理解和生成的综合数据集。为了解决这个问题,我们引入了Audio-FLAN,这是一个大规模的语音调整数据集,涵盖了语音,音乐和声音领域的80个不同任务,拥有超过1亿个实例。Audio-FLAN为统一的音频语言模型奠定了基础,该模型可以无缝地处理理解(例如,转录,理解)和生成(例如,语音、音乐、声音)任务。Audio-FLAN数据集在HuggingFace和GitHub上提供,并将不断更新。
摘要:Recent advancements in audio tokenization have significantly enhanced theintegration of audio capabilities into large language models (LLMs). However,audio understanding and generation are often treated as distinct tasks,hindering the development of truly unified audio-language models. Whileinstruction tuning has demonstrated remarkable success in improvinggeneralization and zero-shot learning across text and vision, its applicationto audio remains largely unexplored. A major obstacle is the lack ofcomprehensive datasets that unify audio understanding and generation. Toaddress this, we introduce Audio-FLAN, a large-scale instruction-tuning datasetcovering 80 diverse tasks across speech, music, and sound domains, with over100 million instances. Audio-FLAN lays the foundation for unifiedaudio-language models that can seamlessly handle both understanding (e.g.,transcription, comprehension) and generation (e.g., speech, music, sound) tasksacross a wide range of audio domains in a zero-shot manner. The Audio-FLANdataset is available on HuggingFace and GitHub and will be continuouslyupdated.

【13】 Improving Speech Enhancement by Cross- and Sub-band Processing with  State Space Model
标题:利用状态空间模型通过跨带和子带处理改进语音增强
链接:https://arxiv.org/abs/2502.16207
作者:Jizhen Li,  Weiping Tu,  Yuhong Yang,  Xinmeng Xu,  Yiqun Zhang,  Yanzhen Ren
摘要:最近,以Mamba为代表的状态空间模型(SSM)在包括语音增强在内的长期序列建模任务中表现出了显著的性能。然而,由于子带特征的实质性差异,将相同的SSM应用于所有子带限制了其推断能力。此外,当处理时频表示的每个时间帧时,SSM可能会忘记某些低能量的高频信息,使得高频带中的结构恢复具有挑战性。为此,我们提出了交叉和子带曼巴(CSMamba)。为了帮助SSM灵活地处理不同的子带特征,我们提出了一个带分割块,该块根据它们的信息相似性将全带分割成四个具有不同宽度的子带。然后,我们为每个子带分配独立的权重,从而减少了SSM的推理负担。此外,为了减轻遗忘的低能量信息的高频段的SSM,我们引入了一个频谱恢复块,从多个角度增强了跨带特征的表示。在DNS Challenge 2021数据集上的实验结果表明,CSMamba在三个客观评价指标中以较少的参数表现出优于几种最先进的(SOTA)语音增强方法。
摘要:Recently, the state space model (SSM) represented by Mamba has shownremarkable performance in long-term sequence modeling tasks, including speechenhancement. However, due to substantial differences in sub-band features,applying the same SSM to all sub-bands limits its inference capability.Additionally, when processing each time frame of the time-frequencyrepresentation, the SSM may forget certain high-frequency information of lowenergy, making the restoration of structure in the high-frequency bandschallenging. For this reason, we propose Cross- and Sub-band Mamba (CSMamba).To assist the SSM in handling different sub-band features flexibly, we proposea band split block that splits the full-band into four sub-bands with differentwidths based on their information similarity. We then allocate independentweights to each sub-band, thereby reducing the inference burden on the SSM.Furthermore, to mitigate the forgetting of low-energy information in thehigh-frequency bands by the SSM, we introduce a spectrum restoration block thatenhances the representation of the cross-band features from multipleperspectives. Experimental results on the DNS Challenge 2021 datasetdemonstrate that CSMamba outperforms several state-of-the-art (SOTA) speechenhancement methods in three objective evaluation metrics with fewerparameters.

【14】 Understanding Zero-shot Rare Word Recognition Improvements Through LLM  Integration
标题:通过LLM集成了解零次稀有词识别改进
链接:https://arxiv.org/abs/2502.16142
作者:Haoxuan Wang
摘要:在这项研究中,我们研究了一个大的语言模型(LLM)与自动语音识别(ASR)系统的集成,特别是专注于提高罕见的单词识别性能。使用主要来自YouTube的190,000小时数据集,使用Whisper V3伪标签进行预处理,我们证明了LLM-ASR架构在大型数据集上训练后,在zero-shot稀有词识别任务中优于传统的Zipformer-Transducer模型。我们的分析表明,LLM有助于显着改善罕见字错误率(R-WER),而语音编码器主要决定整体转录性能(正交字错误率,O-WER,和归一化字错误率,N-WER)。通过广泛的消融研究,我们强调了适配器集成在调整语音编码器输出与LLM的语言能力的重要性。此外,我们强调了高质量标记数据在实现最佳性能方面的关键作用。这些发现为基于LLM的ASR架构之间的协同作用提供了有价值的见解,为基于LLM的大规模语音识别系统的未来发展铺平了道路。
摘要:In this study, we investigate the integration of a large language model (LLM)with an automatic speech recognition (ASR) system, specifically focusing onenhancing rare word recognition performance. Using a 190,000-hour datasetprimarily sourced from YouTube, pre-processed with Whisper V3 pseudo-labeling,we demonstrate that the LLM-ASR architecture outperforms traditionalZipformer-Transducer models in the zero-shot rare word recognition task, aftertraining on a large dataset. Our analysis reveals that the LLM contributessignificantly to improvements in rare word error rate (R-WER), while the speechencoder primarily determines overall transcription performance (OrthographicWord Error Rate, O-WER, and Normalized Word Error Rate, N-WER). Throughextensive ablation studies, we highlight the importance of adapter integrationin aligning speech encoder outputs with the LLM's linguistic capabilities.Furthermore, we emphasize the critical role of high-quality labeled data inachieving optimal performance. These findings provide valuable insights intothe synergy between LLM-based ASR architectures, paving the way for futureadvancements in large-scale LLM-based speech recognition systems.

【15】 Mind the Gap! Static and Interactive Evaluations of Large Audio Models
标题:注意差距!大型音频模型的静态和交互式评估
链接:https://arxiv.org/abs/2502.15919
作者:Minzhi Li,  William Barr Held,  Michael J Ryan,  Kunat Pipatanakul,  Potsawee Manakul,  Hao Zhu,  Diyi Yang
摘要:随着人工智能聊天机器人变得无处不在,语音交互提供了一种引人注目的方式,可以为语义和社交信号提供快速,高带宽的通信。这推动了对大型音频模型(LAM)的研究,以推动语音原生体验。然而,将LAM开发与用户目标相结合需要清楚地了解用户需求和偏好,以建立可靠的进度指标。本研究通过引入一种交互式方法来评估LAM并从484名参与者中收集7,500个LAM交互来应对这些挑战。通过对用户查询的主题建模,我们确定了音频接口的主要用例。然后,我们分析用户偏好排名和定性反馈,以确定哪些模型最符合用户需求。最后,我们评估了静态基准测试如何预测交互性能-我们的分析显示,没有一个单独的基准测试与交互结果密切相关(所有基准测试的$\tau \leq 0.33$)。虽然组合多个粗粒度特征会产生适度的预测能力($R^2$=$0.30$),但20个口语问答和年龄预测数据集中只有两个显示出显著的正相关性。这表明显然需要开发与用户偏好更好相关的LAM评估。
摘要:As AI chatbots become ubiquitous, voice interaction presents a compelling wayto enable rapid, high-bandwidth communication for both semantic and socialsignals. This has driven research into Large Audio Models (LAMs) to powervoice-native experiences. However, aligning LAM development with user goalsrequires a clear understanding of user needs and preferences to establishreliable progress metrics. This study addresses these challenges by introducingan interactive approach to evaluate LAMs and collecting 7,500 LAM interactionsfrom 484 participants. Through topic modeling of user queries, we identifyprimary use cases for audio interfaces. We then analyze user preferencerankings and qualitative feedback to determine which models best align withuser needs. Finally, we evaluate how static benchmarks predict interactiveperformance - our analysis reveals no individual benchmark strongly correlateswith interactive results ($\tau \leq 0.33$ for all benchmarks). While combiningmultiple coarse-grained features yields modest predictive power ($R^2$=$0.30$),only two out of twenty datasets on spoken question answering and age predictionshow significantly positive correlations. This suggests a clear need to developLAM evaluations that better correlate with user preferences.

【16】 Deriving Representative Structure from Music Corpora
标题:从音乐库中获取代表性结构
链接:https://arxiv.org/abs/2502.15849
作者:Ilana Shapiro,  Ruanqianqian (Lisa) Huang,  Zachary Novack,  Cheng-i Wang,  Hao-Wen Dong,  Taylor Berg-Kirkpatrick,  Shlomo Dubnov,  Sorin Lerner
备注:12 pages, 8 figures, 7 tables
摘要:西方音乐是一个天生的层次系统的结构相互作用的水平,从细粒度的旋律,以高层次的形式。为了分析音乐作品的整体性和多粒度,我们提出了一个统一的,层次化的元表示的音乐结构称为结构时间图(STG)。对于单个作品,STG是一种数据结构,它定义了一个逐步精细的结构音乐特征及其之间的时间关系的层次结构。我们使用的STG,使一种新的方法来获得一个有代表性的结构摘要的音乐语料库,我们正式作为一个双重NP-难组合优化问题扩展的广义中值图问题。我们的方法首先应用模拟退火开发一个措施的两个音乐作品之间的结构距离植根于图同构。然后,我们的方法相结合的SMT求解器与嵌套模拟退火的结构距离的正式保证,产生一个结构合理的,代表性的质心STG为整个语料库的STG从个别作品。为了评估我们的方法,我们进行实验验证,结构距离准确区分音乐作品,并推导出质心准确地在结构上表征他们的语料库。
摘要:Western music is an innately hierarchical system of interacting levels ofstructure, from fine-grained melody to high-level form. In order to analyzemusic compositions holistically and at multiple granularities, we propose aunified, hierarchical meta-representation of musical structure called thestructural temporal graph (STG). For a single piece, the STG is a datastructure that defines a hierarchy of progressively finer structural musicalfeatures and the temporal relationships between them. We use the STG to enablea novel approach for deriving a representative structural summary of a musiccorpus, which we formalize as a dually NP-hard combinatorial optimizationproblem extending the Generalized Median Graph problem. Our approach firstapplies simulated annealing to develop a measure of structural distance betweentwo music pieces rooted in graph isomorphism. Our approach then combines theformal guarantees of SMT solvers with nested simulated annealing overstructural distances to produce a structurally sound, representative centroidSTG for an entire corpus of STGs from individual pieces. To evaluate ourapproach, we conduct experiments verifying that structural distance accuratelydifferentiates between music pieces, and that derived centroids accuratelystructurally characterize their corpora.

【17】 LACTOSE: Linear Array of Conditions, TOpologies with Separated  Error-backpropagation -- The Differentiable "IF" Conditional for  Differentiable Digital Signal Processing
链接:https://arxiv.org/abs/2502.15829
作者:Christopher Johann Clarke
备注:6 pages
摘要:将条件语句作为神经网络图的一部分存在困难(例如,如果input $> x$,则将输入传递给网络$N$)。这是由于不能通过分支条件反向传播。条件线性阵列、误差分离反向传播拓扑(LACTOSE)算法解决了这个问题,并允许有条件地使用监督学习模型的可用机器学习层。本文将LACTOSE算法应用于DDSP的一个简单的使用中,然而,重点是为DDSP的使用开发“if”条件。LACTOSE算法为每个用户指定的数值范围存储训练参数,并在预测期间动态加载参数。
摘要:There has been difficulty utilising conditional statements as part of theneural network graph (e.g. if input $> x$, pass input to network $N$). This isdue to the inability to backpropagate through branching conditions. The LinearArray of Conditions, TOpologies with Separated Error-backpropagation (LACTOSE)Algorithm addresses this issue and allows the conditional use of availablemachine learning layers for supervised learning models. In this paper, theLACTOSE algorithm is applied to a simple use of DDSP, however, the main pointis the development of the "if" conditional for DDSP use. The LACTOSE algorithmstores trained parameters for each user-specified numerical range and loads theparameters dynamically during prediction.

【18】 Slamming: Training a Speech Language Model on One GPU in a Day
标题:猛烈抨击:一天内在一个图形处理器上训练语音语言模型
链接:https://arxiv.org/abs/2502.15814
作者:Gallil Maimon,  Avishai Elmakies,  Yossi Adi
摘要:我们介绍了Slam,这是一种在24小时内在单个学术GPU上训练高质量语音语言模型(SLM)的方法。我们通过对模型初始化和架构、合成训练数据、使用合成数据进行偏好优化以及调整所有其他组件的实证分析来实现这一点。我们根据经验证明,这种训练方法也可以很好地扩展,更多的计算可以以一小部分的计算成本获得与领先的SLM相当的结果。我们希望这些见解将使可持续土地管理培训和研究更容易获得。在SLM缩放律的背景下,我们的结果远远优于预测的计算最佳性能,对SLM的可行性给予了乐观的看法。请参阅代码,数据,模型,样本-https://pages.cs.huji.ac.il/adiyoss-lab/slamming.
摘要:We introduce Slam, a recipe for training high-quality Speech Language Models(SLMs) on a single academic GPU in 24 hours. We do so through empiricalanalysis of model initialisation and architecture, synthetic training data,preference optimisation with synthetic data and tweaking all other components.We empirically demonstrate that this training recipe also scales well with morecompute getting results on par with leading SLMs in a fraction of the computecost. We hope these insights will make SLM training and research moreaccessible. In the context of SLM scaling laws, our results far outperformpredicted compute optimal performance, giving an optimistic view to SLMfeasibility. See code, data, models, samples at -https://pages.cs.huji.ac.il/adiyoss-lab/slamming .

【19】 Benchmarking machine learning for bowel sound pattern classification  from tabular features to pretrained models
标题:从表格特征到预训练模型的肠道声音模式分类机器学习基准
链接:https://arxiv.org/abs/2502.15607
作者:Zahra Mansour,  Verena Uslar,  Dirk Weyhe,  Danilo Hollosi,  Nils Strodthoff
备注:9 pages, 6 figures and 1 table
摘要:电子听诊器和可穿戴记录传感器的发展为肠鸣音(BS)信号的自动分析打开了大门。这使得能够对肠鸣音模式、它们的相互关系以及它们与不同病理的相关性进行数据驱动的分析。这项工作利用了从16名健康受试者收集的BS数据集,根据四个已建立的BS模式进行注释。该数据集用于评估机器学习模型检测和/或分类BS模式的性能。所考虑的模型的选择包括使用表格特征的模型、基于频谱图的卷积神经网络以及在大型音频数据集上预训练的模型。结果突出了预训练模型的明显优势,特别是在检测样本较少的类别时,使用HuBERT模型区分BS与非BS的AUC为0.89,使用Wav 2 Vec 2.0模型区分肠鸣模式的AUC为0.89。这些结果为改善对肠鸣音的理解以及未来机器学习驱动的胃肠道检查诊断应用铺平了道路
摘要:The development of electronic stethoscopes and wearable recording sensorsopened the door to the automated analysis of bowel sound (BS) signals. Thisenables a data-driven analysis of bowel sound patterns, their interrelations,and their correlation to different pathologies. This work leverages a BSdataset collected from 16 healthy subjects that was annotated according to fourestablished BS patterns. This dataset is used to evaluate the performance ofmachine learning models to detect and/or classify BS patterns. The selection ofconsidered models covers models using tabular features, convolutional neuralnetworks based on spectrograms and models pre-trained on large audio datasets.The results highlight the clear superiority of pre-trained models, particularlyin detecting classes with few samples, achieving an AUC of 0.89 indistinguishing BS from non-BS using a HuBERT model and an AUC of 0.89 indifferentiating bowel sound patterns using a Wav2Vec 2.0 model. These resultspave the way for an improved understanding of bowel sounds in general andfuture machine-learning-driven diagnostic applications for gastrointestinalexaminations

机器翻译由腾讯交互翻译提供,仅供参考