今日论文合集:cs.SD语音8篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】A Multimodal Symphony: Integrating Taste and Sound through Generative AI
标题:多模式交响曲:通过生成人工智能整合味觉和声音
链接:https://arxiv.org/abs/2503.02823
作者:Matteo Spanio,  Massimiliano Zampini,  Antonio Rodà,  Franco Pierucci
备注:17 pages, 6 figures (2 + 2 figures with 2 subfigures each)
摘要:近几十年来,神经科学和心理学研究已经追踪了味觉和听觉感知之间的直接关系。本文探讨了能够将味觉信息转换为音乐的多模态生成模型,建立在这一基础研究的基础上。我们提供了一个简短的回顾,在这一领域的最先进的,突出的关键发现和方法。我们提出了一个实验,其中微调版本的生成音乐模型(MusicGEN)用于生成音乐的基础上提供的每个音乐作品的详细口味描述。结果令人鼓舞:根据参与者($n=111$)的评价,与非微调模型相比,微调模型产生更连贯地反映输入品味描述的音乐。这项研究代表了理解和发展AI,声音和味道之间的具体互动的重要一步,为生成AI领域开辟了新的可能性。我们在https://osf.io/xs5jy/上发布了我们的数据集,代码和预训练模型。
摘要:In recent decades, neuroscientific and psychological research has traceddirect relationships between taste and auditory perceptions. This articleexplores multimodal generative models capable of converting taste informationinto music, building on this foundational research. We provide a brief reviewof the state of the art in this field, highlighting key findings andmethodologies. We present an experiment in which a fine-tuned version of agenerative music model (MusicGEN) is used to generate music based on detailedtaste descriptions provided for each musical piece. The results are promising:according the participants' ($n=111$) evaluation, the fine-tuned model producesmusic that more coherently reflects the input taste descriptions compared tothe non-fine-tuned model. This study represents a significant step towardsunderstanding and developing embodied interactions between AI, sound, andtaste, opening new possibilities in the field of generative AI. We release ourdataset, code and pre-trained model at: https://osf.io/xs5jy/.

【2】 InSerter: Speech Instruction Following with Unsupervised Interleaved  Pre-training
标题:InSerter:语音指令遵循无监督交织预训练
链接:https://arxiv.org/abs/2503.02769
作者:Dingdong Wang,  Jin Xu,  Ruihang Chu,  Zhifang Guo,  Xiong Wang,  Jincenzi Wu,  Dongchao Yang,  Shengpeng Ji,  Junyang Lin
摘要:语音大语言模型(SpeechLLM)的最新进展引起了人们的广泛关注。尽管如此,目前的方法在遵守语音指令方面表现出次优的性能。值得注意的是,与直接文本形式的输入相比,处理语音形式的输入时,模型的智能显著降低。先前的工作试图通过诸如表示和行为对齐之类的技术来减轻语音和文本表示之间的这种语义不一致,这些技术涉及在后训练阶段对数据对进行细致的设计。在本文中,我们介绍了一种简单且可扩展的训练方法,称为InSerter,它代表Interleaved Speech-Text Representation Pre-training。InSerter旨在预训练大规模无监督语音文本序列,其中语音是使用文本到语音转换从大量文本语料库中随机选择的片段合成的。因此,该模型获得了生成与所提供的语音片段相对应的文本延续的能力,从而避免了对密集的数据设计努力的需要。为了系统地评估语音识别能力,我们引入了SpeechInstructBench,这是第一个专门为语音识别任务设计的综合基准。我们提出的InSerter在SpeechInstructBench中实现了SOTA性能,并在各种语音处理任务中表现出卓越或有竞争力的结果。
摘要:Recent advancements in speech large language models (SpeechLLMs) haveattracted considerable attention. Nonetheless, current methods exhibitsuboptimal performance in adhering to speech instructions. Notably, theintelligence of models significantly diminishes when processing speech-forminput as compared to direct text-form input. Prior work has attempted tomitigate this semantic inconsistency between speech and text representationsthrough techniques such as representation and behavior alignment, which involvethe meticulous design of data pairs during the post-training phase. In thispaper, we introduce a simple and scalable training method called InSerter,which stands for Interleaved Speech-Text Representation Pre-training. InSerteris designed to pre-train large-scale unsupervised speech-text sequences, wherethe speech is synthesized from randomly selected segments of an extensive textcorpus using text-to-speech conversion. Consequently, the model acquires theability to generate textual continuations corresponding to the provided speechsegments, obviating the need for intensive data design endeavors. Tosystematically evaluate speech instruction-following capabilities, we introduceSpeechInstructBench, the first comprehensive benchmark specifically designedfor speech-oriented instruction-following tasks. Our proposed InSerter achievesSOTA performance in SpeechInstructBench and demonstrates superior orcompetitive results across diverse speech processing tasks.

【3】 A Hypernetwork-Based Approach to KAN Representation of Audio Signals
标题:基于超网络的音频信号KAN表示方法
链接:https://arxiv.org/abs/2503.02585
作者:Patryk Marszałek,  Maciej Rut,  Piotr Kawa,  Piotr Syga
摘要:内隐神经表征(INR)在多媒体数据的高效编码方面取得了显著的成就,但其在音频信号中的应用仍然有限。这项研究介绍了Kolmogorov-Arnold网络(KAN),一种新的架构,使用可学习的激活函数,作为一个有效的INR模型的音频表示。KAN表现出优于以前的INR的感知性能,实现了1.5秒音频的最低对数频谱距离1.29和最高语音质量感知评估3.57。为了扩展KAN的实用程序,我们提出了FewSound,一个基于超网络的架构,增强INR参数更新。FewSound优于最先进的HyperSound,MSE提高了33.3%,SI-SNR提高了60.87%。这些结果表明,KAN作为一个强大的和适应性强的音频表示的可扩展性和集成到各种超网络框架的潜力。源代码可在https://github.com/gmum/fewsound.git上获得。
摘要:Implicit neural representations (INR) have gained prominence for efficientlyencoding multimedia data, yet their applications in audio signals remainlimited. This study introduces the Kolmogorov-Arnold Network (KAN), a novelarchitecture using learnable activation functions, as an effective INR modelfor audio representation. KAN demonstrates superior perceptual performance overprevious INRs, achieving the lowest Log-SpectralDistance of 1.29 and thehighest Perceptual Evaluation of Speech Quality of 3.57 for 1.5 s audio. Toextend KAN's utility, we propose FewSound, a hypernetwork-based architecturethat enhances INR parameter updates. FewSound outperforms the state-of-the-artHyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These resultsshow KAN as a robust and adaptable audio representation with the potential forscalability and integration into various hypernetwork frameworks. The sourcecode can be accessed at https://github.com/gmum/fewsound.git.

【4】 Aggregation Strategies for Efficient Annotation of Bioacoustic Sound  Events Using Active Learning
标题:使用主动学习有效注释生物声学声音事件的聚集策略
链接:https://arxiv.org/abs/2503.02422
作者:Richard Lindholm,  Oscar Marklund,  Olof Mogren,  John Martinsson
摘要:在声音事件检测(SED)应用中收集的大量音频数据需要高效的注释策略来实现监督学习。手动标记昂贵且耗时,使得主动学习(AL)成为减少注释工作的有前途的方法。我们引入了Top K Entropy,这是一种新的AL不确定性聚合策略,它优先考虑音频记录中最不确定的片段,而不是平均所有片段的不确定性。这种方法可以选择整个记录进行注释,从而提高稀疏数据场景中的效率。我们将Top K Entropy与随机采样和Mean Entropy进行了比较,并表明更少的标签可以导致相同的模型性能,特别是在具有稀疏声音事件的数据集中。对公园的声音记录与猫鼬,狗和婴儿哭泣的声音事件,代表现实世界的生物声学监测方案的音频混合进行评估。使用Top K Entropy进行主动学习,我们可以在只有8%标签的完全标签数据集上实现与训练相当的性能。前K熵优于平均熵,这表明最好让最不确定的片段表示音频文件的不确定性。研究结果突出了AL在音频和时间序列应用(包括生物声学)中可扩展注释的潜力。
摘要:The vast amounts of audio data collected in Sound Event Detection (SED)applications require efficient annotation strategies to enable supervisedlearning. Manual labeling is expensive and time-consuming, making ActiveLearning (AL) a promising approach for reducing annotation effort. We introduceTop K Entropy, a novel uncertainty aggregation strategy for AL that prioritizesthe most uncertain segments within an audio recording, instead of averaginguncertainty across all segments. This approach enables the selection of entirerecordings for annotation, improving efficiency in sparse data scenarios. Wecompare Top K Entropy to random sampling and Mean Entropy, and show that fewerlabels can lead to the same model performance, particularly in datasets withsparse sound events. Evaluations are conducted on audio mixtures of soundrecordings from parks with meerkat, dog, and baby crying sound events,representing real-world bioacoustic monitoring scenarios. Using Top K Entropyfor active learning, we can achieve comparable performance to training on thefully labeled dataset with only 8% of the labels. Top K Entropy outperformsMean Entropy, suggesting that it is best to let the most uncertain segmentsrepresent the uncertainty of an audio file. The findings highlight thepotential of AL for scalable annotation in audio and time-series applications,including bioacoustics.

【5】 Robust detection of overlapping bioacoustic sound events
标题:对重叠的生物声学声音事件的稳健检测
链接:https://arxiv.org/abs/2503.02389
作者:Louis Mahon,  Benjamin Hoffman,  Logan S James,  Maddie Cusimano,  Masato Hagiwara,  Sarah C Woolley,  Olivier Pietquin
摘要:我们提出了一种方法,用于准确地检测生物声学的声音事件,是强大的重叠事件,一个共同的问题,如行为学,生态学和保护领域。虽然标准方法采用基于帧的多标签方法,但我们引入了一种基于发作的检测方法,我们将其命名为Voxaboxen。它从计算机视觉中的对象检测方法中获得灵感,但同时利用了自监督音频编码器的最新进展。对于每个时间窗口,Voxaboxen预测它是否包含发声的开始以及发声的时间长度。它也做同样的反向,预测每个窗口是否包含一个发声的结束,以及多久之前开始的。然后使用图形匹配算法融合两组边界框。我们还发布了一个新的数据集,旨在衡量检测重叠发声的性能。这包括斑马雀的记录,这些记录用时间上很强的标签注释,并显示出频繁的重叠。我们在七个现有的数据集和我们的新数据集上测试Voxaboxen。我们比较Voxaboxen的自然基线和现有的声音事件检测方法,并证明SotA的结果。进一步的实验表明,改进是强大的频繁发声重叠。
摘要:We propose a method for accurately detecting bioacoustic sound events that isrobust to overlapping events, a common issue in domains such as ethology,ecology and conservation. While standard methods employ a frame-based,multi-label approach, we introduce an onset-based detection method which wename Voxaboxen. It takes inspiration from object detection methods in computervision, but simultaneously takes advantage of recent advances inself-supervised audio encoders. For each time window, Voxaboxen predictswhether it contains the start of a vocalization and how long the vocalizationis. It also does the same in reverse, predicting whether each window containsthe end of a vocalization, and how long ago it started. The two resulting setsof bounding boxes are then fused using a graph-matching algorithm. We alsorelease a new dataset designed to measure performance on detecting overlappingvocalizations. This consists of recordings of zebra finches annotated withtemporally-strong labels and showing frequent overlaps. We test Voxaboxen onseven existing data sets and on our new data set. We compare Voxaboxen tonatural baselines and existing sound event detection methods and demonstrateSotA results. Further experiments show that improvements are robust to frequentvocalization overlap.

【6】 Audio-Reasoner: Improving Reasoning Capability in Large Audio Language  Models
标题:音频推理者:提高大型音频语言模型中的推理能力
链接:https://arxiv.org/abs/2503.02318
作者:Zhifei Xie,  Mingbao Lin,  Zihang Liu,  Pengcheng Wu,  Shuicheng Yan,  Chunyan Miao
备注:Technical report, in process
摘要:多模态推理的最新进展在很大程度上忽略了音频模态。我们介绍了Audio-Reasoner,这是一个用于音频任务深度推理的大规模音频语言模型。我们精心策划了一个具有简单注释的大规模和多样化的多任务音频数据集。然后,我们利用闭源模型进行二次标记,QA生成,以及结构化的COT过程。这些数据集共同形成了一个高质量的推理数据集,其中包含120万个推理丰富的样本,我们将其命名为CoTA。遵循推理缩放原则,我们在CoTA上训练Audio-Reasoner,使其在音频推理中具有强大的逻辑能力。实验表明,在关键基准测试中,包括MMAU-mini(+25.42%),AIR-Bench chat/foundation(+14.57%/+10.13%)和MELD(+8.01%)的性能达到了最先进的水平。我们的研究结果强调了结构化CoT培训在推进音频推理方面的核心。
摘要:Recent advancements in multimodal reasoning have largely overlooked the audiomodality. We introduce Audio-Reasoner, a large-scale audio language model fordeep reasoning in audio tasks. We meticulously curated a large-scale anddiverse multi-task audio dataset with simple annotations. Then, we leverageclosed-source models to conduct secondary labeling, QA generation, along withstructured COT process. These datasets together form a high-quality reasoningdataset with 1.2 million reasoning-rich samples, which we name CoTA. Followinginference scaling principles, we train Audio-Reasoner on CoTA, enabling it toachieve great logical capabilities in audio reasoning. Experiments showstate-of-the-art performance across key benchmarks, including MMAU-mini(+25.42%), AIR-Bench chat/foundation(+14.57%/+10.13%), and MELD (+8.01%). Ourfindings stress the core of structured CoT training in advancing audioreasoning.

【7】 Nexus-O: An Omni-Perceptive And -Interactive Model for Language, Audio,  And Vision
标题:Nexus-O:语言、音频和视觉的全感知和交互模型
链接:https://arxiv.org/abs/2503.01879
作者:Che Liu,  Yingji Zhang,  Dong Zhang,  Weijie Zhang,  Chenggong Gong,  Haohan Li,  Yu Lu,  Shilin Zhou,  Yue Lu,  Ziliang Gan,  Ziao Wang,  Junwei Liao,  Haipang Wu,  Ji Liu,  André Freitas,  Qifan Wang,  Zenglin Xu,  Rongjuncheng Zhang,  Yong Dai
摘要:人类通过一系列的感官形式感知现实世界,包括听觉,视觉和语言能力。实现通用人工智能(AGI)的旅程需要开发能够模拟这些多方面感知能力并全面理解这些多样化数据的模型。为此,我们引入了\textbf{Nexus-O},这是一个行业级的\textbf{全感知和交互}模型,能够有效地处理任何组合的音频、图像、视频和文本数据,并以端到端的方式输出音频/文本。我们通过解决三个关键研究问题来系统地研究Nexus-O:首先,如何有效地设计和训练模型,以实现跨多种模态的三模态对齐、理解和推理能力?第二,可以实施什么方法来评估三模态模型的鲁棒性,确保在现实世界场景中的可靠性能和适用性?第三,可以采用什么策略来策划和获得高质量的真实场景语音数据集?对于第一个问题,我们基于视觉语言模型而不是语言模型来设计和预训练Nexus-O。通过在高质量的合成音频数据上对模型进行预训练,我们的模型能够进行三模态感知和交互。对于第二个问题,我们介绍了一个新的音频测试平台Nexus-O-audio,它包括各种自动语音识别(ASR)样本,跨越各种现实场景,如公司会议和直播。对于第三个问题,我们设计了语音数据合成管道,以获得高质量的语音训练数据集,覆盖各种真实场景。综合实验和三模态对齐在潜在空间的深入分析表明,我们的模型在下游任务的优势。
摘要:Human beings perceive the real world through a spectrum of sensorymodalities, encompassing auditory, visual, and linguistic faculties. Thejourney towards achieving Artificial General Intelligence (AGI) necessitatesthe development of models that can emulate these multifaceted perceptualcapabilities and comprehensively understand these diversified data. To thisend, we introduce \textbf{Nexus-O}, an industry-level \textbf{omni-perceptiveand -interactive} model capable of efficiently processing Audio, Image, Video,and Text data in any combination and output audio/text in an end-to-end way. Wesystematically investigate Nexus-O by addressing three key research questions:First, how can models be efficiently designed and trained to achieve tri-modalalignment, understanding and reasoning capabilities across multiple modalities?Second, what approaches can be implemented to evaluate tri-modal modelrobustness, ensuring reliable performance and applicability in real-worldscenarios? Third, what strategies can be employed to curate and obtainhigh-quality, real-life scenario speech datasets? For the first question, wedesign and pre-train Nexus-O based on the vision-language model, rather thanthe language model. By pre-training the model over high-quality synthetic audiodata, our model is capable of tri-modal perception and interaction. For thesecond question, we introduce a new audio testbed, Nexus-O-audio, comprisingdiverse Automatic Speech Recognition (ASR) samples, spanning various real-worldscenarios, such as corporate meetings and live stream. For the third question,we design the speech data synthesis pipeline to obtain high-quality speechtraining datasets, covering various real-world scenarios. Comprehensiveexperimentation and an in-depth analysis of tri-modal alignment over latentspace demonstrate the advantages of our model on downstream tasks.

【8】 CNN-based Robust Sound Source Localization with SRP-PHAT for the Extreme  Edge
标题:利用SRP-PHAT实现基于CNN的鲁棒性光源定位
链接:https://arxiv.org/abs/2503.02046
作者:Jun Yin,  Marian Verhelst
备注:None
摘要:针对噪声和混响环境的稳健声源定位越来越多地利用具有各种声学特征的深度神经网络。然而,最先进的研究主要集中在优化算法的准确性上,导致庞大的模型阻碍了边缘设备的部署。然而,边缘敦促实时低足迹声学推理应用,如助听器和机器人交互。因此,我们从使用SRP-PHAT功能的强大的基于CNN的模型出发,Cross 3D [16],为极端边缘追求高效而紧凑的模型架构。对于SRP特征表示和神经网络,我们分别提出了可扩展的LC-SRP-Edge和Cross 3D-Edge算法,这些算法都是针对较低的硬件开销进行优化的。与原始LC-SRP [19]相比,LC-SRP-Edge将sinc插值的复杂性和片内存储器开销减半。在多个SRP分辨率情况下,Cross 3D-Edge与Cross 3D基线相比节省了10.32~73.71%的计算复杂度和59.77~94.66%的神经网络权重。在精度-效率权衡方面,最平衡的版本(EM)总共只需要127.1 MFLOPS计算、3.71 MByte/s带宽和0.821 MByte片内存储器,同时在最先进的精度比较中仍保持竞争力。它在Rasberry Pi 4 B上实现了8.59 ms/帧的端到端延迟,比相应的基线快7.26倍。
摘要:Robust sound source localization for environments with noise andreverberation are increasingly exploiting deep neural networks fed with variousacoustic features. Yet, state-of-the-art research mainly focuses on optimizingalgorithmic accuracy, resulting in huge models preventing edge-devicedeployment. The edge, however, urges for real-time low-footprint acousticreasoning for applications such as hearing aids and robot interactions. Hence,we set off from a robust CNN-based model using SRP-PHAT features, Cross3D [16],to pursue an efficient yet compact model architecture for the extreme edge. Forboth the SRP feature representation and neural network, we propose respectivelyour scalable LC-SRP-Edge and Cross3D-Edge algorithms which are optimizedtowards lower hardware overhead. LC-SRP-Edge halves the complexity and on-chipmemory overhead for the sinc interpolation compared to the original LC-SRP[19]. Over multiple SRP resolution cases, Cross3D-Edge saves 10.32~73.71%computational complexity and 59.77~94.66% neural network weights against theCross3D baseline. In terms of the accuracy-efficiency tradeoff, the mostbalanced version (EM) requires only 127.1 MFLOPS computation, 3.71 MByte/sbandwidth, and 0.821 MByte on-chip memory in total, while still retainingcompetitiveness in state-of-the-art accuracy comparisons. It achieves 8.59ms/frame end-to-end latency on a Rasberry Pi 4B, which is 7.26x faster than thecorresponding baseline.

eess.AS音频处理

【1】 CNN-based Robust Sound Source Localization with SRP-PHAT for the Extreme  Edge
标题:利用SRP-PHAT实现基于CNN的鲁棒性光源定位
链接:https://arxiv.org/abs/2503.02046
作者:Jun Yin,  Marian Verhelst
备注:None
摘要:针对噪声和混响环境的稳健声源定位越来越多地利用具有各种声学特征的深度神经网络。然而,最先进的研究主要集中在优化算法的准确性上,导致庞大的模型阻碍了边缘设备的部署。然而,边缘敦促实时低足迹声学推理应用,如助听器和机器人交互。因此,我们从使用SRP-PHAT功能的强大的基于CNN的模型出发,Cross 3D [16],为极端边缘追求高效而紧凑的模型架构。对于SRP特征表示和神经网络,我们分别提出了可扩展的LC-SRP-Edge和Cross 3D-Edge算法,这些算法都是针对较低的硬件开销进行优化的。与原始LC-SRP [19]相比,LC-SRP-Edge将sinc插值的复杂性和片内存储器开销减半。在多个SRP分辨率情况下,Cross 3D-Edge与Cross 3D基线相比节省了10.32~73.71%的计算复杂度和59.77~94.66%的神经网络权重。在精度-效率权衡方面,最平衡的版本(EM)总共只需要127.1 MFLOPS计算、3.71 MByte/s带宽和0.821 MByte片内存储器,同时在最先进的精度比较中仍保持竞争力。它在Rasberry Pi 4 B上实现了8.59 ms/帧的端到端延迟,比相应的基线快7.26倍。
摘要:Robust sound source localization for environments with noise andreverberation are increasingly exploiting deep neural networks fed with variousacoustic features. Yet, state-of-the-art research mainly focuses on optimizingalgorithmic accuracy, resulting in huge models preventing edge-devicedeployment. The edge, however, urges for real-time low-footprint acousticreasoning for applications such as hearing aids and robot interactions. Hence,we set off from a robust CNN-based model using SRP-PHAT features, Cross3D [16],to pursue an efficient yet compact model architecture for the extreme edge. Forboth the SRP feature representation and neural network, we propose respectivelyour scalable LC-SRP-Edge and Cross3D-Edge algorithms which are optimizedtowards lower hardware overhead. LC-SRP-Edge halves the complexity and on-chipmemory overhead for the sinc interpolation compared to the original LC-SRP[19]. Over multiple SRP resolution cases, Cross3D-Edge saves 10.32~73.71%computational complexity and 59.77~94.66% neural network weights against theCross3D baseline. In terms of the accuracy-efficiency tradeoff, the mostbalanced version (EM) requires only 127.1 MFLOPS computation, 3.71 MByte/sbandwidth, and 0.821 MByte on-chip memory in total, while still retainingcompetitiveness in state-of-the-art accuracy comparisons. It achieves 8.59ms/frame end-to-end latency on a Rasberry Pi 4B, which is 7.26x faster than thecorresponding baseline.

【2】 A Multimodal Symphony: Integrating Taste and Sound through Generative AI
标题:多模式交响曲:通过生成人工智能整合味觉和声音
链接:https://arxiv.org/abs/2503.02823
作者:Matteo Spanio,  Massimiliano Zampini,  Antonio Rodà,  Franco Pierucci
备注:17 pages, 6 figures (2 + 2 figures with 2 subfigures each)
摘要:近几十年来,神经科学和心理学研究已经追踪了味觉和听觉感知之间的直接关系。本文探讨了能够将味觉信息转换为音乐的多模态生成模型,建立在这一基础研究的基础上。我们简要回顾了该领域的最新技术水平,重点介绍了关键发现和方法。我们提出了一个实验,其中微调版本的生成音乐模型(MusicGEN)用于生成音乐的基础上提供的每个音乐作品的详细口味描述。结果令人鼓舞:根据参与者($n=111$)的评价,与非微调模型相比,微调模型产生更连贯地反映输入品味描述的音乐。这项研究代表了理解和发展AI,声音和味道之间的具体互动的重要一步,为生成AI领域开辟了新的可能性。我们在https://osf.io/xs5jy/上发布了我们的数据集,代码和预训练模型。
摘要:In recent decades, neuroscientific and psychological research has traceddirect relationships between taste and auditory perceptions. This articleexplores multimodal generative models capable of converting taste informationinto music, building on this foundational research. We provide a brief reviewof the state of the art in this field, highlighting key findings andmethodologies. We present an experiment in which a fine-tuned version of agenerative music model (MusicGEN) is used to generate music based on detailedtaste descriptions provided for each musical piece. The results are promising:according the participants' ($n=111$) evaluation, the fine-tuned model producesmusic that more coherently reflects the input taste descriptions compared tothe non-fine-tuned model. This study represents a significant step towardsunderstanding and developing embodied interactions between AI, sound, andtaste, opening new possibilities in the field of generative AI. We release ourdataset, code and pre-trained model at: https://osf.io/xs5jy/.

【3】 InSerter: Speech Instruction Following with Unsupervised Interleaved  Pre-training
标题:InSerter:语音指令遵循无监督交织预训练
链接:https://arxiv.org/abs/2503.02769
作者:Dingdong Wang,  Jin Xu,  Ruihang Chu,  Zhifang Guo,  Xiong Wang,  Jincenzi Wu,  Dongchao Yang,  Shengpeng Ji,  Junyang Lin
摘要:语音大语言模型(SpeechLLM)的最新进展引起了人们的广泛关注。尽管如此,目前的方法在遵守语音指令方面表现出次优的性能。值得注意的是,与直接文本形式的输入相比,处理语音形式的输入时,模型的智能显著降低。先前的工作试图通过诸如表示和行为对齐之类的技术来减轻语音和文本表示之间的这种语义不一致,这些技术涉及在后训练阶段对数据对进行细致的设计。在本文中,我们介绍了一种简单且可扩展的训练方法,称为InSerter,它代表Interleaved Speech-Text Representation Pre-training。InSerter旨在预训练大规模无监督语音文本序列,其中语音是使用文本到语音转换从大量文本语料库中随机选择的片段合成的。因此,该模型获得了生成与所提供的语音片段相对应的文本延续的能力,从而避免了对密集的数据设计努力的需要。为了系统地评估语音识别能力,我们引入了SpeechInstructBench,这是第一个专门为语音识别任务设计的综合基准。我们提出的InSerter在SpeechInstructBench中实现了SOTA性能,并在各种语音处理任务中表现出卓越或有竞争力的结果。
摘要:Recent advancements in speech large language models (SpeechLLMs) haveattracted considerable attention. Nonetheless, current methods exhibitsuboptimal performance in adhering to speech instructions. Notably, theintelligence of models significantly diminishes when processing speech-forminput as compared to direct text-form input. Prior work has attempted tomitigate this semantic inconsistency between speech and text representationsthrough techniques such as representation and behavior alignment, which involvethe meticulous design of data pairs during the post-training phase. In thispaper, we introduce a simple and scalable training method called InSerter,which stands for Interleaved Speech-Text Representation Pre-training. InSerteris designed to pre-train large-scale unsupervised speech-text sequences, wherethe speech is synthesized from randomly selected segments of an extensive textcorpus using text-to-speech conversion. Consequently, the model acquires theability to generate textual continuations corresponding to the provided speechsegments, obviating the need for intensive data design endeavors. Tosystematically evaluate speech instruction-following capabilities, we introduceSpeechInstructBench, the first comprehensive benchmark specifically designedfor speech-oriented instruction-following tasks. Our proposed InSerter achievesSOTA performance in SpeechInstructBench and demonstrates superior orcompetitive results across diverse speech processing tasks.

【4】 A Hypernetwork-Based Approach to KAN Representation of Audio Signals
标题:基于超网络的音频信号KAN表示方法
链接:https://arxiv.org/abs/2503.02585
作者:Patryk Marszałek,  Maciej Rut,  Piotr Kawa,  Piotr Syga
摘要:内隐神经表征(INR)在多媒体数据的高效编码方面取得了显著的成就,但其在音频信号中的应用仍然有限。本研究介绍了Kolmogorov-Arnold网络(KAN),一种使用可学习激活函数的新型架构,作为音频表示的有效INR模型。KAN表现出优于以前的INR的感知性能,实现了1.5秒音频的最低对数频谱距离1.29和最高语音质量感知评估3.57。为了扩展KAN的实用程序,我们提出了FewSound,一个基于超网络的架构,增强INR参数更新。FewSound优于最先进的HyperSound,MSE提高了33.3%,SI-SNR提高了60.87%。这些结果表明,KAN作为一个强大的和适应性强的音频表示的可扩展性和集成到各种超网络框架的潜力。源代码可在https://github.com/gmum/fewsound.git上获得。
摘要:Implicit neural representations (INR) have gained prominence for efficientlyencoding multimedia data, yet their applications in audio signals remainlimited. This study introduces the Kolmogorov-Arnold Network (KAN), a novelarchitecture using learnable activation functions, as an effective INR modelfor audio representation. KAN demonstrates superior perceptual performance overprevious INRs, achieving the lowest Log-SpectralDistance of 1.29 and thehighest Perceptual Evaluation of Speech Quality of 3.57 for 1.5 s audio. Toextend KAN's utility, we propose FewSound, a hypernetwork-based architecturethat enhances INR parameter updates. FewSound outperforms the state-of-the-artHyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These resultsshow KAN as a robust and adaptable audio representation with the potential forscalability and integration into various hypernetwork frameworks. The sourcecode can be accessed at https://github.com/gmum/fewsound.git.

【5】 Aggregation Strategies for Efficient Annotation of Bioacoustic Sound  Events Using Active Learning
标题:使用主动学习有效注释生物声学声音事件的聚集策略
链接:https://arxiv.org/abs/2503.02422
作者:Richard Lindholm,  Oscar Marklund,  Olof Mogren,  John Martinsson
摘要:在声音事件检测(SED)应用中收集的大量音频数据需要高效的注释策略来实现监督学习。手动标记昂贵且耗时,使得主动学习(AL)成为减少注释工作的有前途的方法。我们引入了Top K Entropy,这是一种新的AL不确定性聚合策略,它优先考虑音频记录中最不确定的片段,而不是平均所有片段的不确定性。这种方法可以选择整个记录进行注释,从而提高稀疏数据场景中的效率。我们将Top K Entropy与随机采样和Mean Entropy进行了比较,并表明更少的标签可以导致相同的模型性能,特别是在具有稀疏声音事件的数据集中。对公园的声音记录与猫鼬,狗和婴儿哭泣的声音事件,代表现实世界的生物声学监测方案的音频混合进行评估。使用Top K Entropy进行主动学习,我们可以在只有8%标签的完全标签数据集上实现与训练相当的性能。前K熵优于平均熵,这表明最好让最不确定的片段表示音频文件的不确定性。研究结果突出了AL在音频和时间序列应用(包括生物声学)中可扩展注释的潜力。
摘要:The vast amounts of audio data collected in Sound Event Detection (SED)applications require efficient annotation strategies to enable supervisedlearning. Manual labeling is expensive and time-consuming, making ActiveLearning (AL) a promising approach for reducing annotation effort. We introduceTop K Entropy, a novel uncertainty aggregation strategy for AL that prioritizesthe most uncertain segments within an audio recording, instead of averaginguncertainty across all segments. This approach enables the selection of entirerecordings for annotation, improving efficiency in sparse data scenarios. Wecompare Top K Entropy to random sampling and Mean Entropy, and show that fewerlabels can lead to the same model performance, particularly in datasets withsparse sound events. Evaluations are conducted on audio mixtures of soundrecordings from parks with meerkat, dog, and baby crying sound events,representing real-world bioacoustic monitoring scenarios. Using Top K Entropyfor active learning, we can achieve comparable performance to training on thefully labeled dataset with only 8% of the labels. Top K Entropy outperformsMean Entropy, suggesting that it is best to let the most uncertain segmentsrepresent the uncertainty of an audio file. The findings highlight thepotential of AL for scalable annotation in audio and time-series applications,including bioacoustics.

【6】 Robust detection of overlapping bioacoustic sound events
标题:对重叠的生物声学声音事件的稳健检测
链接:https://arxiv.org/abs/2503.02389
作者:Louis Mahon,  Benjamin Hoffman,  Logan S James,  Maddie Cusimano,  Masato Hagiwara,  Sarah C Woolley,  Olivier Pietquin
摘要:我们提出了一种方法,用于准确地检测生物声学的声音事件,是强大的重叠事件,一个共同的问题,如行为学,生态学和保护领域。虽然标准方法采用基于帧的多标签方法,但我们引入了一种基于发作的检测方法,我们将其命名为Voxaboxen。它从计算机视觉中的对象检测方法中获得灵感,但同时利用了自监督音频编码器的最新进展。对于每个时间窗口,Voxaboxen预测它是否包含发声的开始以及发声的时间长度。它也做同样的反向,预测每个窗口是否包含一个发声的结束,以及多久之前开始的。然后使用图形匹配算法融合两组边界框。我们还发布了一个新的数据集,旨在衡量检测重叠发声的性能。这包括斑马雀的记录,这些记录用时间上很强的标签注释,并显示出频繁的重叠。我们在七个现有的数据集和我们的新数据集上测试Voxaboxen。我们比较Voxaboxen的自然基线和现有的声音事件检测方法,并证明SotA的结果。进一步的实验表明,这些改进对频繁的发声重叠具有鲁棒性。
摘要:We propose a method for accurately detecting bioacoustic sound events that isrobust to overlapping events, a common issue in domains such as ethology,ecology and conservation. While standard methods employ a frame-based,multi-label approach, we introduce an onset-based detection method which wename Voxaboxen. It takes inspiration from object detection methods in computervision, but simultaneously takes advantage of recent advances inself-supervised audio encoders. For each time window, Voxaboxen predictswhether it contains the start of a vocalization and how long the vocalizationis. It also does the same in reverse, predicting whether each window containsthe end of a vocalization, and how long ago it started. The two resulting setsof bounding boxes are then fused using a graph-matching algorithm. We alsorelease a new dataset designed to measure performance on detecting overlappingvocalizations. This consists of recordings of zebra finches annotated withtemporally-strong labels and showing frequent overlaps. We test Voxaboxen onseven existing data sets and on our new data set. We compare Voxaboxen tonatural baselines and existing sound event detection methods and demonstrateSotA results. Further experiments show that improvements are robust to frequentvocalization overlap.

【7】 Nexus-O: An Omni-Perceptive And -Interactive Model for Language, Audio,  And Vision
标题:Nexus-O:语言、音频和视觉的全感知和交互模型
链接:https://arxiv.org/abs/2503.01879
作者:Che Liu,  Yingji Zhang,  Dong Zhang,  Weijie Zhang,  Chenggong Gong,  Haohan Li,  Yu Lu,  Shilin Zhou,  Yue Lu,  Ziliang Gan,  Ziao Wang,  Junwei Liao,  Haipang Wu,  Ji Liu,  André Freitas,  Qifan Wang,  Zenglin Xu,  Rongjuncheng Zhang,  Yong Dai
摘要:人类通过一系列的感觉方式感知真实世界,包括听觉,视觉和语言能力。实现通用人工智能(AGI)的旅程需要开发能够模拟这些多方面感知能力并全面理解这些多样化数据的模型。为此,我们引入了\textbf{Nexus-O},这是一个行业级的\textbf{全感知和交互}模型,能够有效地处理任何组合的音频、图像、视频和文本数据,并以端到端的方式输出音频/文本。我们通过解决三个关键研究问题来系统地研究Nexus-O:首先,如何有效地设计和训练模型,以实现跨多种模态的三模态对齐、理解和推理能力?第二,可以实施什么方法来评估三模态模型的鲁棒性,确保在现实世界场景中的可靠性能和适用性?第三,可以采用什么策略来策划和获得高质量的真实场景语音数据集?对于第一个问题,我们基于视觉语言模型而不是语言模型来设计和预训练Nexus-O。通过使用高质量合成音频数据预训练模型,我们的模型能够进行三模式感知和交互。对于第二个问题,我们介绍了一个新的音频测试平台Nexus-O-audio,它包括各种自动语音识别(ASR)样本,跨越各种现实场景,如公司会议和直播。对于第三个问题,我们设计了语音数据合成管道,以获得高质量的语音训练数据集,覆盖各种真实场景。综合实验和三模态对齐在潜在空间的深入分析表明,我们的模型在下游任务的优势。
摘要:Human beings perceive the real world through a spectrum of sensorymodalities, encompassing auditory, visual, and linguistic faculties. Thejourney towards achieving Artificial General Intelligence (AGI) necessitatesthe development of models that can emulate these multifaceted perceptualcapabilities and comprehensively understand these diversified data. To thisend, we introduce \textbf{Nexus-O}, an industry-level \textbf{omni-perceptiveand -interactive} model capable of efficiently processing Audio, Image, Video,and Text data in any combination and output audio/text in an end-to-end way. Wesystematically investigate Nexus-O by addressing three key research questions:First, how can models be efficiently designed and trained to achieve tri-modalalignment, understanding and reasoning capabilities across multiple modalities?Second, what approaches can be implemented to evaluate tri-modal modelrobustness, ensuring reliable performance and applicability in real-worldscenarios? Third, what strategies can be employed to curate and obtainhigh-quality, real-life scenario speech datasets? For the first question, wedesign and pre-train Nexus-O based on the vision-language model, rather thanthe language model. By pre-training the model over high-quality synthetic audiodata, our model is capable of tri-modal perception and interaction. For thesecond question, we introduce a new audio testbed, Nexus-O-audio, comprisingdiverse Automatic Speech Recognition (ASR) samples, spanning various real-worldscenarios, such as corporate meetings and live stream. For the third question,we design the speech data synthesis pipeline to obtain high-quality speechtraining datasets, covering various real-world scenarios. Comprehensiveexperimentation and an in-depth analysis of tri-modal alignment over latentspace demonstrate the advantages of our model on downstream tasks.

机器翻译由腾讯交互翻译提供,仅供参考