今日论文合集:cs.SD语音15篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】A Reliable and Efficient Detection Pipeline for Rodent Ultrasonic  Vocalizations
标题:可靠有效的啮齿动物超声发声检测管道
链接:https://arxiv.org/abs/2503.18928
作者:Sabah Shahnoor Anis,  Devin M. Kellis,  Kris Ford Kaigler,  Marlene A. Wilson,  Christian O'Reilly
备注:Accepted for publication in the proceeding of the 7th International Conference on Advances in Signal Processing and Artificial Intelligence (ASPAI' 2025), 8-10 April 2025, Innsbruck, Austria
摘要:分析超声波发声(USVs)对于了解啮齿动物的情感状态和社会行为至关重要,但人工分析耗时且容易出错。自动USV检测系统已经被开发来解决这些挑战。然而,这些系统通常依赖于机器学习,无法有效地推广到新的数据集。为了解决这些缺点,我们引入了ContourUSV,这是一种用于从音频记录中检测USV的高效自动化系统。我们的流程包括频谱图生成、清理、预处理、轮廓检测、后处理以及针对手动注释的评估。为了确保鲁棒性和可靠性,我们使用现有的开放访问USV数据集(USVSEG)和我们与本文一起公开发布的第二个数据集将ContourUSV与三个最先进的系统进行了比较。平均而言,在两个数据集上,ContourUSV的性能优于其他三个系统,精度提高了1.51倍,召回率提高了1.17倍,F1得分提高了1.80倍,特异性提高了1.49倍,同时实现了117.07倍的平均加速。
摘要:Analyzing ultrasonic vocalizations (USVs) is crucial for understandingrodents' affective states and social behaviors, but the manual analysis istime-consuming and prone to errors. Automated USV detection systems have beendeveloped to address these challenges. Yet, these systems often rely on machinelearning and fail to generalize effectively to new datasets. To tackle theseshortcomings, we introduce ContourUSV, an efficient automated system fordetecting USVs from audio recordings. Our pipeline includes spectrogramgeneration, cleaning, pre-processing, contour detection, post-processing, andevaluation against manual annotations. To ensure robustness and reliability, wecompared ContourUSV with three state-of-the-art systems using an existingopen-access USV dataset (USVSEG) and a second dataset we are releasing publiclyalong with this paper. On average, across the two datasets, ContourUSVoutperformed the other three systems with a 1.51x improvement in precision,1.17x in recall, 1.80x in F1 score, and 1.49x in specificity while achieving anaverage speedup of 117.07x.

【2】 Seeing Speech and Sound: Distinguishing and Locating Audios in Visual  Scenes
标题:看到语音和声音:区分和定位视觉场景中的听众
链接:https://arxiv.org/abs/2503.18880
作者:Hyeonggon Ryu,  Seongyu Kim,  Joon Son Chung,  Arda Senocak
备注:CVPR 2025
摘要:我们提出了一个统一的模型,能够同时接地口语和非语音声音在视觉场景中,解决当前视听接地模型的关键限制。现有的方法通常限于独立地处理语音或非语音声音,或者最好一起但顺序地处理而不混合。这一限制使他们无法捕捉经常混合的真实世界音频源的复杂性。我们的方法引入了一个“混合和分离”的框架与视听对齐的目标,共同学习对应和使用混合音频解开。通过这些目标,我们的模型学习为每种音频类型产生不同的嵌入,从而实现混合音频源的有效解纠缠和接地。此外,我们创建了一个新的数据集来评估混合音频源的同时接地,证明我们的模型优于以前的方法。我们的方法在标准分割和跨模态检索任务中也取得了相当或更好的性能,突出了我们的混合和分离方法的好处。
摘要:We present a unified model capable of simultaneously grounding both spokenlanguage and non-speech sounds within a visual scene, addressing keylimitations in current audio-visual grounding models. Existing approaches aretypically limited to handling either speech or non-speech sounds independently,or at best, together but sequentially without mixing. This limitation preventsthem from capturing the complexity of real-world audio sources that are oftenmixed. Our approach introduces a 'mix-and-separate' framework with audio-visualalignment objectives that jointly learn correspondence and disentanglementusing mixed audio. Through these objectives, our model learns to producedistinct embeddings for each audio type, enabling effective disentanglement andgrounding across mixed audio sources. Additionally, we created a new dataset toevaluate simultaneous grounding of mixed audio sources, demonstrating that ourmodel outperforms prior methods. Our approach also achieves comparable orbetter performance in standard segmentation and cross-modal retrieval tasks,highlighting the benefits of our mix-and-separate approach.

【3】 CCMusic: An Open and Diverse Database for Chinese Music Information  Retrieval Research
标题:CCMusic:一个开放、多元化的中国音乐信息检索研究数据库
链接:https://arxiv.org/abs/2503.18802
作者:Monan Zhou,  Shenyang Xu,  Zhaorui Liu,  Zhaowen Wang,  Feng Yu,  Wei Li,  Baoqiang Han
备注:17 pages, 18 figures
摘要:数据在各个计算机相关领域都至关重要,包括音乐信息检索(MIR),这是一个连接计算机科学和音乐的跨学科领域。本文介绍了CCMusic,一个开放和多样化的数据库,包括多个数据集,专门为与中国音乐相关的任务设计,突出了我们对这个文化丰富的领域的关注。该数据库集成了已发布和未发布的数据集,并采取了数据清理、标签细化和数据结构统一等步骤,以确保数据一致性并创建随时可用的版本。我们使用专门为此目的开发的统一评估框架对所有数据集进行基准评估。这个公开可用的框架支持分类和检测任务,确保所有数据集的标准化和可重复结果。该数据库托管在HuggingFace和ModelScope这两个开放的多功能数据和模型托管平台上,确保易于访问和使用。
摘要:Data are crucial in various computer-related fields, including musicinformation retrieval (MIR), an interdisciplinary area bridging computerscience and music. This paper introduces CCMusic, an open and diverse databasecomprising multiple datasets specifically designed for tasks related to Chinesemusic, highlighting our focus on this culturally rich domain. The databaseintegrates both published and unpublished datasets, with steps taken such asdata cleaning, label refinement, and data structure unification to ensure dataconsistency and create ready-to-use versions. We conduct benchmark evaluationsfor all datasets using a unified evaluation framework developed specificallyfor this purpose. This publicly available framework supports bothclassification and detection tasks, ensuring standardized and reproducibleresults across all datasets. The database is hosted on HuggingFace andModelScope, two open and multifunctional data and model hosting platforms,ensuring ease of accessibility and usability.

【4】 Wireless Hearables With Programmable Speech AI Accelerators
标题:配备可编程语音AI加速器的无线可听设备
链接:https://arxiv.org/abs/2503.18698
作者:Malek Itani,  Tuochao Chen,  Arun Raghavan,  Gavriel Kohlberg,  Shyamnath Gollakota
摘要:传统观点认为,由于流式深度学习模型的高计算需求,设计具有设备上语音AI模型的超紧凑、电池受限的无线耳机具有挑战性。语音AI模型需要连续、实时的音频处理,施加了严格的计算和I/O约束。我们介绍了NeuralAids,这是一种用于无线耳机的完全设备上语音AI系统,可以在紧凑的电池受限设备上实现实时语音增强和去噪。我们的系统通过做出三项关键技术贡献,弥合了最先进的语音增强深度学习和低功耗AI硬件之间的差距:1)集成语音AI加速器的无线可听平台,用于高效的设备上流推理,2)优化的双路径神经网络,用于低延迟,高质量的语音增强,以及3)硬件-软件协同设计,其使用混合精度量化和量化感知训练来在严格的功率约束下实现实时性能。我们的系统实时处理6 ms的音频块,实现了5.54 ms的推理时间,同时消耗71.6 mW。在现实世界的评估中,包括一项有28名参与者的用户研究,我们的系统在语音质量和噪声抑制方面优于之前的设备模型,为下一代智能无线助听器铺平了道路,可以完全增强设备上的听力。
摘要:The conventional wisdom has been that designing ultra-compact,battery-constrained wireless hearables with on-device speech AI models ischallenging due to the high computational demands of streaming deep learningmodels. Speech AI models require continuous, real-time audio processing,imposing strict computational and I/O constraints. We present NeuralAids, afully on-device speech AI system for wireless hearables, enabling real-timespeech enhancement and denoising on compact, battery-constrained devices. Oursystem bridges the gap between state-of-the-art deep learning for speechenhancement and low-power AI hardware by making three key technicalcontributions: 1) a wireless hearable platform integrating a speech AIaccelerator for efficient on-device streaming inference, 2) an optimizeddual-path neural network designed for low-latency, high-quality speechenhancement, and 3) a hardware-software co-design that uses mixed-precisionquantization and quantization-aware training to achieve real-time performanceunder strict power constraints. Our system processes 6 ms audio chunks inreal-time, achieving an inference time of 5.54 ms while consuming 71.6 mW. Inreal-world evaluations, including a user study with 28 participants, our systemoutperforms prior on-device models in speech quality and noise suppression,paving the way for next-generation intelligent wireless hearables that canenhance hearing entirely on-device.

【5】 Music Similarity Representation Learning Focusing on Individual  Instruments with Source Separation and Human Preference
标题:音乐相似性表示学习专注于单个乐器,具有来源分离和人类偏好
链接:https://arxiv.org/abs/2503.18486
作者:Takehiro Imamura,  Yuka Hashizume,  Wen-Chin Huang,  Tomoki Toda
摘要:本文提出了基于单个乐器声音(InMSRL)的音乐相似性表示学习(MSRL),利用音乐源分离(MSS)和人类偏好,而不需要在推理过程中使用干净的乐器声音。我们提出了三种有效提高性能的方法。首先,我们介绍端到端微调(E2 E-FT)的级联方法,顺序执行MSS和音乐相似性特征提取。E2 E-FT允许模型最小化分离误差对特征提取的不利影响。其次,我们提出了多任务学习的直接方法,直接提取解开音乐相似性特征,使用一个单一的音乐相似性特征提取器。基于分离的音乐相似性特征提取的多任务学习和基于分离的音乐相似性特征重构的MSS进一步增强了乐器特征分离。第三,我们采用感知感知微调(PAFT)。PAFT利用人类偏好,允许模型执行与人类感知相似性对齐的InMSRL。我们进行了实验评估,并证明1)E2 E-FT for Cascade显着提高了InMSRL性能,2)Direct的多任务学习也有助于提高特征提取中的解纠缠性能,3)PAFT显着提高了感知InMSRL性能,4)Cascade with E2 E-FT and PAFT优于Direct with the multitask learning and PAFT。
摘要:This paper proposes music similarity representation learning (MSRL) based onindividual instrument sounds (InMSRL) utilizing music source separation (MSS)and human preference without requiring clean instrument sounds duringinference. We propose three methods that effectively improve performance.First, we introduce end-to-end fine-tuning (E2E-FT) for the Cascade approachthat sequentially performs MSS and music similarity feature extraction. E2E-FTallows the model to minimize the adverse effects of a separation error on thefeature extraction. Second, we propose multi-task learning for the Directapproach that directly extracts disentangled music similarity features using asingle music similarity feature extractor. Multi-task learning, which is basedon the disentangled music similarity feature extraction and MSS based onreconstruction with disentangled music similarity features, further enhancesinstrument feature disentanglement. Third, we employ perception-awarefine-tuning (PAFT). PAFT utilizes human preference, allowing the model toperform InMSRL aligned with human perceptual similarity. We conductexperimental evaluations and demonstrate that 1) E2E-FT for Cascadesignificantly improves InMSRL performance, 2) the multi-task learning forDirect is also helpful to improve disentanglement performance in the featureextraction, 3) PAFT significantly enhances the perceptual InMSRL performance,and 4) Cascade with E2E-FT and PAFT outperforms Direct with the multi-tasklearning and PAFT.

【6】 DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via  Personalizer-Guided Distillation
标题:DistusionTalker:通过个性化引导蒸馏实现高效、紧凑的语音驱动3D说话头
链接:https://arxiv.org/abs/2503.18159
作者:Peng Chen,  Xiaobao Wei,  Ming Lu,  Hui Chen,  Feng Tian
备注:Accepted by ICME2025
摘要:实时语音驱动的3D人脸动画在学术界和工业界一直很有吸引力。传统的方法主要集中在学习从语音到动画的确定性映射。最近的方法开始考虑语音驱动的3D人脸动画的不确定性事实,并采用扩散模型的任务。现有的基于扩散的方法可以提高人脸动画的多样性。然而,个性化的表达准确的唇语的说话风格仍然缺乏,此外,效率和紧凑性仍有待提高。在这项工作中,我们提出了DiffusionTalker通过个性化指导蒸馏来解决上述限制。在个性化方面,我们引入了一个对比个性化器,它学习身份和情感嵌入,以从音频中捕获说话风格。我们进一步提出了一个个性化增强器在蒸馏过程中,以提高嵌入对面部动画的影响。为了提高效率,我们使用迭代蒸馏来减少动画生成所需的步骤,并实现超过8倍的推理加速。为了实现紧凑性,我们将大型教师模型提取为较小的学生模型,将模型的存储量减少了86.4%,同时最大限度地减少了性能损失。经过提炼,用户可以从音频中获得他们的身份和情感嵌入,以快速创建反映特定说话风格的个性化动画。大量的实验表明,我们的方法优于国家的最先进的方法。该代码将在https://github.com/ChenVoid/DiffusionTalker上发布。
摘要:Real-time speech-driven 3D facial animation has been attractive in academiaand industry. Traditional methods mainly focus on learning a deterministicmapping from speech to animation. Recent approaches start to consider thenondeterministic fact of speech-driven 3D face animation and employ thediffusion model for the task. Existing diffusion-based methods can improve thediversity of facial animation. However, personalized speaking styles conveyingaccurate lip language is still lacking, besides, efficiency and compactnessstill need to be improved. In this work, we propose DiffusionTalker to addressthe above limitations via personalizer-guided distillation. In terms ofpersonalization, we introduce a contrastive personalizer that learns identityand emotion embeddings to capture speaking styles from audio. We furtherpropose a personalizer enhancer during distillation to enhance the influence ofembeddings on facial animation. For efficiency, we use iterative distillationto reduce the steps required for animation generation and achieve more than 8xspeedup in inference. To achieve compactness, we distill the large teachermodel into a smaller student model, reducing our model's storage by 86.4\%while minimizing performance loss. After distillation, users can derive theiridentity and emotion embeddings from audio to quickly create personalizedanimations that reflect specific speaking styles. Extensive experiments areconducted to demonstrate that our method outperforms state-of-the-art methods.The code will be released at: https://github.com/ChenVoid/DiffusionTalker.

【7】 Machine learning based animal emotion classification using audio signals
标题:使用音频信号的基于机器学习的动物情感分类
链接:https://arxiv.org/abs/2503.18138
作者:Mariia Slobodian,  Mykola Kozlenko
备注:5 pages, 3 figures. This paper was originally published in 2022 International Conference on Innovative Solutions in Software Engineering (ICISSE), available: https://zenodo.org/records/7514136
摘要:本文提出了机器学习方法的自动分类的狗的情绪状态的基础上处理和识别的音频信号。它为改善人机界面和开发更精确的工具从声学数据中分类情感提供了有用的信息。所提出的模型证明了一个整体的准确度值超过70%的音频信号记录一只狗。
摘要:This paper presents the machine learning approach to the automatedclassification of a dog's emotional state based on the processing andrecognition of audio signals. It offers helpful information for improvinghuman-machine interfaces and developing more precise tools for classifyingemotions from acoustic data. The presented model demonstrates an overallaccuracy value above 70% for audio signals recorded for one dog.

【8】 Anomaly Detection and Localization for Speech Deepfakes via Feature  Pyramid Matching
标题:通过特征金字塔匹配对语音Deepfakes进行异常检测和定位
链接:https://arxiv.org/abs/2503.18032
作者:Emma Coletta,  Davide Salvi,  Viola Negroni,  Daniele Ugo Leonzio,  Paolo Bestagini
摘要:人工智能驱动的生成模型的兴起使人们能够创建高度逼真的语音深度伪造-可以模仿目标说话者声音的合成音频信号-引发了严重的安全问题。现有的检测语音深度伪造的方法主要依赖于监督学习,这有两个关键的限制:对看不见的合成技术的推广有限,以及缺乏可解释性。在本文中,我们通过引入一种新的可解释的一类检测框架来解决这些问题,该框架将语音深度伪造检测重新定义为异常检测任务。我们的模型仅在真实语音上进行训练,以表征其分布,从而能够对合成生成的分布外样本进行分类。此外,我们的框架在推理过程中产生可解释的异常图,突出显示时域和频域的异常区域。这是通过一个学生-教师特征金字塔匹配系统来完成的,该系统通过离散缩放来增强,以提高在看不见的数据分布中的泛化能力。广泛的评估表明,与所考虑的基线相比,我们的方法具有优越的性能,验证了帧语音深度伪造检测作为异常检测问题的有效性。
摘要:The rise of AI-driven generative models has enabled the creation of highlyrealistic speech deepfakes - synthetic audio signals that can imitate targetspeakers' voices - raising critical security concerns. Existing methods fordetecting speech deepfakes primarily rely on supervised learning, which suffersfrom two critical limitations: limited generalization to unseen synthesistechniques and a lack of explainability. In this paper, we address these issuesby introducing a novel interpretable one-class detection framework, whichreframes speech deepfake detection as an anomaly detection task. Our model istrained exclusively on real speech to characterize its distribution, enablingthe classification of out-of-distribution samples as synthetically generated.Additionally, our framework produces interpretable anomaly maps duringinference, highlighting anomalous regions across both time and frequencydomains. This is done through a Student-Teacher Feature Pyramid Matchingsystem, enhanced with Discrepancy Scaling to improve generalizationcapabilities across unseen data distributions. Extensive evaluationsdemonstrate the superior performance of our approach compared to the consideredbaselines, validating the effectiveness of framing speech deepfake detection asan anomaly detection problem.

【9】 Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and  Speech Recognition
标题:通过将说话人分离和语音识别脱钩提高鲁棒的多说话者ASB
链接:https://arxiv.org/abs/2503.17886
作者:Yufeng Yang,  Hassan Taherian,  Vahid Ahmadi Kalkhorani,  DeLiang Wang
摘要:尽管随着深度学习的引入,自动语音识别(ASR)取得了巨大的成功,但在许多现实世界的多人场景中,其性能仍然不能令人满意。说话者分离在分离单个说话者方面表现出色,但作为前端,它引入了处理伪像,这些伪像会降低在干净语音上训练的ASR后端。因此,主流的鲁棒ASR系统训练关于噪声语音的后端以避免处理伪像。在这项工作中,我们建议将说话人分离前端和ASR后端的训练解耦,后者仅在干净的语音上训练。我们的解耦系统在Libri 2 Mix开发/测试集上实现了5.1%的字错误率(WER),显著优于其他多通话器ASR基线。其有效性也通过1通道和6通道SMS-WSJ上最先进的7.60%/5.74%WER得到了证明。此外,在记录的LibriCSS上,我们实现了2.92%的说话者归因WER。这些最先进的结果表明,解耦说话人分离和识别是一种有效的方法,以提高鲁棒的多说话人ASR。
摘要:Despite the tremendous success of automatic speech recognition (ASR) with theintroduction of deep learning, its performance is still unsatisfactory in manyreal-world multi-talker scenarios. Speaker separation excels in separatingindividual talkers but, as a frontend, it introduces processing artifacts thatdegrade the ASR backend trained on clean speech. As a result, mainstream robustASR systems train the backend on noisy speech to avoid processing artifacts. Inthis work, we propose to decouple the training of the speaker separationfrontend and the ASR backend, with the latter trained on clean speech only. Ourdecoupled system achieves 5.1% word error rates (WER) on the Libri2Mix dev/testsets, significantly outperforming other multi-talker ASR baselines. Itseffectiveness is also demonstrated with the state-of-the-art 7.60%/5.74% WERson 1-ch and 6-ch SMS-WSJ. Furthermore, on recorded LibriCSS, we achieve thespeaker-attributed WER of 2.92%. These state-of-the-art results suggest thatdecoupling speaker separation and recognition is an effective approach toelevate robust multi-talker ASR.

【10】 GSound-SIR: A Spatial Impulse Response Ray-Tracing and High-order  Ambisonic Auralization Python Toolkit
标题:GSound-Sir:空间脉冲响应光线追踪和高级立体声听觉化Python工具包
链接:https://arxiv.org/abs/2503.17866
作者:Yongyi Zang,  Qiuqiang Kong
摘要:准确、高效地模拟房间脉冲响应对于空间音频应用至关重要。然而,现有的声波射线追踪工具通常作为黑盒操作,并且仅输出脉冲响应(IR),提供对中间数据或空间保真度的有限访问。为了解决这些问题,本文提出了GSound-SIR,一个新的基于Python的室内声学仿真工具包,解决了这些限制。本文的贡献包括以下几点。首先,GSound-SIR可直接访问来自模拟的多达数百万个原始射线数据点,从而实现对声音传播路径的深入分析,这在以前的解决方案中是不可能的。其次,我们介绍了一种工具,将声波射线转换成高阶高保真度立体声脉冲响应合成,捕捉空间音频线索比标准技术具有更高的保真度。第三,为了提高效率,该工具包实现了基于能量的过滤算法,并且只能导出前X或前X %的射线。第四,我们建议将模拟结果存储为Parquet格式,以促进快速数据I/O并与数据分析工作流程无缝集成。这些功能使GSound-SIR成为室内声学研究的先进、高效和现代化的基础,为研究人员和开发人员提供了一个强大的空间音频探索新工具。我们在www.example.com上发布了Apache 2.0许可证下的库。
摘要:Accurate and efficient simulation of room impulse responses is crucial forspatial audio applications. However, existing acoustic ray-tracing tools oftenoperate as black boxes and only output impulse responses (IRs), providinglimited access to intermediate data or spatial fidelity. To address thoseproblems, this paper presents GSound-SIR, a novel Python-based toolkit for roomacoustics simulation that addresses these limitations. The contribution of thispaper includes the follows. First, GSound-SIR provides direct access to up tomillions of raw ray data points from simulations, enabling in-depth analysis ofsound propagation paths that was not possible with previous solutions. Second,we introduce a tool to convert acoustic rays into high-order Ambisonic impulseresponse synthesis, capturing spatial audio cues with greater fidelity thanstandard techniques. Third, to enhance efficiency, the toolkit implements anenergy-based filtering algorithm and can export only the top-X or top-X-% rays.Fourth, we propose to store the simulation results into Parquet formats,facilitating fast data I/O and seamless integration with data analysisworkflows. Together, these features make GSound-SIR an advanced, efficient, andmodern foundation for room acoustics research, providing researchers anddevelopers with a powerful new tool for spatial audio exploration. We releasethe library under Apache 2.0 License athttps://github.com/yongyizang/GSound-SIR.

【11】 LZMidi: Compression-Based Symbolic Music Generation
标题:LZMidi:基于压缩的象征音乐一代
链接:https://arxiv.org/abs/2503.17654
作者:Connor Ding,  Abhiram Gorle,  Sagnik Bhattacharya,  Divija Hasteer,  Naomi Sagan,  Tsachy Weissman
摘要:符号音乐生成的最新进展主要依赖于深度学习模型,如Transformers、GANs和扩散模型。虽然这些方法实现了高质量的结果,但它们需要大量的计算资源,限制了它们的可扩展性。我们介绍LZMidi,一个轻量级的符号音乐生成框架的基础上Lempel-Ziv(LZ 78)诱导顺序概率分配(SPA)。通过利用离散和连续的结构,我们的方法可以在标准CPU上以最小的训练和推理成本高效地生成音乐。从理论上讲,我们建立了普遍的收敛保证我们的方法,强调其可靠性和鲁棒性。与最先进的扩散模型相比,LZMidi实现了具有竞争力的Frechet音频距离(FAD),Wasserstein距离(WD)和Kullback-Leibler(KL)分数,同时显着降低了计算开销-高达30倍的训练速度和300倍的生成速度。我们的研究结果将LZMidi定位为基于压缩的学习的重大进步,突出了通用压缩技术如何有效地建模和生成结构化的序列数据,如符号音乐,具有实际的可扩展性和理论严谨性。
摘要:Recent advances in symbolic music generation primarily rely on deep learningmodels such as Transformers, GANs, and diffusion models. While these approachesachieve high-quality results, they require substantial computational resources,limiting their scalability. We introduce LZMidi, a lightweight symbolic musicgeneration framework based on a Lempel-Ziv (LZ78)-induced sequentialprobability assignment (SPA). By leveraging the discrete and sequentialstructure of MIDI data, our approach enables efficient music generation onstandard CPUs with minimal training and inference costs. Theoretically, weestablish universal convergence guarantees for our approach, underscoring itsreliability and robustness. Compared to state-of-the-art diffusion models,LZMidi achieves competitive Frechet Audio Distance (FAD), Wasserstein Distance(WD), and Kullback-Leibler (KL) scores, while significantly reducingcomputational overhead - up to 30x faster training and 300x faster generation.Our results position LZMidi as a significant advancement in compression-basedlearning, highlighting how universal compression techniques can efficientlymodel and generate structured sequential data, such as symbolic music, withpractical scalability and theoretical rigor.

【12】 Leveraging Audio Representations for Vibration-Based Crowd Monitoring in  Stadiums
标题:利用音频表示进行体育场中基于振动的人群监控
链接:https://arxiv.org/abs/2503.17646
作者:Yen Cheng Chang,  Jesse Codling,  Yiwen Dong,  Jiale Zhang,  Jiasi Chen,  Hae Young Noh,  Pei Zhang
摘要:体育场馆中的人群监控对于增强公共安全和改善观众体验非常重要。现有的方法主要依赖于摄像头和麦克风,这可能会造成严重的干扰,并经常引起隐私问题。在本文中,我们感测地板振动,这提供了一个更少的破坏性和更非侵入性的方式来感知人群,预测人群的行为。然而,由于基于振动的人群监测方法是新开发的,一个主要挑战是由于体育场馆是具有复杂身体活动的大型公共空间而缺乏训练数据。  在本文中,我们提出了ViLA(振动杠杆音频),一种基于振动的方法,通过使用未标记的跨模态数据进行预训练来减少对标记数据的依赖。ViLA首先以无监督的方式对音频数据进行预训练,然后使用最少量的域内振动数据进行微调。通过利用公开可用的音频数据集,ViLA从音频中学习波的行为,然后根据振动调整表示,减少对特定领域振动数据的依赖。我们的真实世界实验表明,与没有音频预训练的模型相比,使用公开可用的音频数据(YouTube8M)预训练振动模型可以实现高达5.8倍的错误减少。
摘要:Crowd monitoring in sports stadiums is important to enhance public safety andimprove the audience experience. Existing approaches mainly rely on cameras andmicrophones, which can cause significant disturbances and often raise privacyconcerns. In this paper, we sense floor vibration, which provides a lessdisruptive and more non-intrusive way of crowd sensing, to predict crowdbehavior. However, since the vibration-based crowd monitoring approach is newlydeveloped, one main challenge is the lack of training data due to sportsstadiums being large public spaces with complex physical activities. In this paper, we present ViLA (Vibration Leverage Audio), a vibration-basedmethod that reduces the dependency on labeled data by pre-training withunlabeled cross-modality data. ViLA is first pre-trained on audio data in anunsupervised manner and then fine-tuned with a minimal amount of in-domainvibration data. By leveraging publicly available audio datasets, ViLA learnsthe wave behaviors from audio and then adapts the representation to vibration,reducing the reliance on domain-specific vibration data. Our real-worldexperiments demonstrate that pre-training the vibration model using publiclyavailable audio data (YouTube8M) achieved up to a 5.8x error reduction comparedto the model without audio pre-training.

【13】 Measuring the Robustness of Audio Deepfake Detectors
标题:测量音频Deepfake检测器的稳健性
链接:https://arxiv.org/abs/2503.17577
作者:Xiang Li,  Pin-Yu Chen,  Wenqi Wei
摘要:Deepfakes已经成为生成AI在各种媒体类型(如图像、音频和视频)中普遍且迅速加剧的问题。其中,音频deepfake特别令人担忧,因为通过社交媒体和robocalls等平台可以轻松合成和分发高质量的语音。因此,检测音频deepfake在打击人工智能合成语音日益增长的滥用方面发挥着关键作用。然而,现实世界的场景通常会引入各种音频损坏,例如噪声、修改和压缩,这可能会显着影响检测性能。这项工作系统地评估了10种音频deepfake检测模型对16种常见损坏的鲁棒性,这些损坏分为噪声扰动,音频修改和压缩。使用传统的深度学习模型和最先进的基础模型,我们得出了四个独特的观察结果。首先,我们的研究结果表明,虽然大多数模型对噪声表现出很强的鲁棒性,但它们显然更容易受到修改和压缩的影响,特别是当应用神经编解码器时。其次,语音基础模型在大多数场景中的表现通常优于传统模型,这可能是由于其自监督学习范式和大规模预训练。第三,我们的研究结果表明,增加模型大小可以提高鲁棒性,尽管收益会递减。第四,我们展示了在训练过程中有针对性的数据增强如何增强模型对未知扰动的弹性。一项关于政治演讲deepfakes的案例研究强调了基础模型在现实世界条件下实现高准确性的有效性。这些发现强调了开发更强大的检测框架以确保实际部署环境中的可靠性的重要性。
摘要:Deepfakes have become a universal and rapidly intensifying concern ofgenerative AI across various media types such as images, audio, and videos.Among these, audio deepfakes have been of particular concern due to the ease ofhigh-quality voice synthesis and distribution via platforms such as socialmedia and robocalls. Consequently, detecting audio deepfakes plays a criticalrole in combating the growing misuse of AI-synthesized speech. However,real-world scenarios often introduce various audio corruptions, such as noise,modification, and compression, that may significantly impact detectionperformance. This work systematically evaluates the robustness of 10 audiodeepfake detection models against 16 common corruptions, categorized into noiseperturbation, audio modification, and compression. Using both traditional deeplearning models and state-of-the-art foundation models, we make four uniqueobservations. First, our findings show that while most models demonstratestrong robustness to noise, they are notably more vulnerable to modificationsand compression, especially when neural codecs are applied. Second, speechfoundation models generally outperform traditional models across mostscenarios, likely due to their self-supervised learning paradigm andlarge-scale pre-training. Third, our results show that increasing model sizeimproves robustness, albeit with diminishing returns. Fourth, we demonstratehow targeted data augmentation during training can enhance model resilience tounseen perturbations. A case study on political speech deepfakes highlights theeffectiveness of foundation models in achieving high accuracy under real-worldconditions. These findings emphasize the importance of developing more robustdetection frameworks to ensure reliability in practical deployment settings.

【14】 Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for  High Quality Data Expansion
标题:具有潜在空间扩展的音频增强视觉语言建模以实现高质量数据扩展
链接:https://arxiv.org/abs/2503.17551
作者:Yu Sun,  Yin Li,  Ruixiao Sun,  Chunhui Liu,  Fangming Zhou,  Ze Jin,  Linjie Wang,  Xiang Shen,  Zhuolin Hao,  Hongyu Xiong
摘要:基于transformer的多模态模型广泛用于工业规模的推荐、搜索和广告系统,用于内容理解和相关性排名。增强标记的训练数据质量和跨模态融合显著提高了模型性能,影响了质量观看率和广告收入等关键指标。高质量的注释对于推进内容建模至关重要,但传统的基于语义的主动学习(AL)方法面临着局限性:它们很难检测到过度自信的错误分类,并且在区分深度神经网络中语义相似的项目方面效果较差。此外,音频信息扮演着越来越重要的角色,特别是在短视频平台中,但大多数预训练的多模式架构主要关注文本和图像。虽然可以在所有三种模式中从头开始训练,但它牺牲了利用现有预训练的视觉语言(VL)和音频模型的好处。为了解决这些挑战,我们提出了基于kNN的潜在空间扩展(LSB)来提高AL效率和具有音频增强的视觉语言建模(VLMAE),这是一种将音频集成到VL模型中的中间融合方法。此系统部署在生产系统中,带来了显著的业务收益。
摘要:Transformer-based multimodal models are widely used in industrial-scalerecommendation, search, and advertising systems for content understanding andrelevance ranking. Enhancing labeled training data quality and cross-modalfusion significantly improves model performance, influencing key metrics suchas quality view rates and ad revenue. High-quality annotations are crucial foradvancing content modeling, yet traditional statistical-based active learning(AL) methods face limitations: they struggle to detect overconfidentmisclassifications and are less effective in distinguishing semanticallysimilar items in deep neural networks. Additionally, audio information plays anincreasing role, especially in short-video platforms, yet most pre-trainedmultimodal architectures primarily focus on text and images. While trainingfrom scratch across all three modalities is possible, it sacrifices thebenefits of leveraging existing pre-trained visual-language (VL) and audiomodels. To address these challenges, we propose kNN-based Latent SpaceBroadening (LSB) to enhance AL efficiency and Vision-Language Modeling withAudio Enhancement (VLMAE), a mid-fusion approach integrating audio into VLmodels. This system deployed in production systems, leading to significantbusiness gains.

【15】 A State-of-the-Art Review on Acoustic Preservation of Historical Worship  Spaces through Auralization
标题:通过听觉化保护历史崇拜空间的最新进展评论
链接:https://arxiv.org/abs/2503.18022
作者:Hannes Rosseel,  Toon van Waterschoot
备注:32 pages, 7 figures, 4 tables, Published in Signal Processing
摘要:历史崇拜空间是具有文化和精神价值的重要建筑地标。这些空间的声学特性在历史和当代的宗教礼拜仪式、仪式和仪式以及神圣音乐的表演中发挥着至关重要的作用。然而,这些空间的原始声学特性往往由于重新利用,翻新,自然灾害或随着时间的推移而恶化而面临风险。本文对声学信号的获取、分析和合成的研究现状进行了全面的综述,重点介绍了HWS。在比利时布鲁塞尔的拿骚教堂的一个例子的案例研究,展示了这些技术的历史崇拜空间声学的保存和可听化的应用。本文最后讨论了该领域的挑战和机遇,并概述了未来的研究方向。
摘要:Historical Worship Spaces (HWS) are significant architectural landmarks whichhold both cultural and spiritual value. The acoustic properties of these spacesplay a crucial role in historical and contemporary religious liturgies,rituals, and ceremonies, as well as in the performance of sacred music.However, the original acoustic characteristics of these spaces are often atrisk due to repurposing, renovations, natural disasters, or deterioration overtime. This paper presents a comprehensive review of the current state ofresearch on the acquisition, analysis, and synthesis of acoustics, with a focuson HWS. An example case study of the Nassau chapel in Brussels, Belgium, ispresented to demonstrate the application of these techniques for thepreservation and auralization of historical worship space acoustics. The paperconcludes with a discussion of the challenges and opportunities in the field,and outlines future research directions.

eess.AS音频处理

【1】 Joint Spectrogram Separation and TDOA Estimation using Optimal Transport
标题:使用最佳传输的联合谱图分离和TDOE估计
链接:https://arxiv.org/abs/2503.18600
作者:Linda Fabiani,  Sebastian J. Schlecht,  Isabel Haasler,  Filip Elvander
摘要:分离声源是语音增强和电信等应用中的常见挑战,其中区分重叠声音有助于减少干扰并提高信号质量。此外,在多通道系统中,正确的校准和同步对于精确分离和定位源信号至关重要。介绍了一种在时频域中进行盲源分离和信号到达时间差(TDOA)估计的方法。我们提出的方法有效地分离信号混合到其原始源频谱图,同时估计接收器之间的相对延迟,使用最优传输(OT)理论。通过利用OT问题的结构,我们将分离和延迟估计过程结合到一个统一的框架中,通过块坐标下降算法优化系统。我们分析了基于OT的估计器在各种噪声条件下的性能,并将其与传统的TDOA和源分离方法进行了比较。数值模拟结果表明,我们提出的方法可以实现不同的噪声场景中的TDOA和源分离任务的物理语音信号的准确性显着水平。
摘要:Separating sources is a common challenge in applications such as speechenhancement and telecommunications, where distinguishing between overlappingsounds helps reduce interference and improve signal quality. Additionally, inmultichannel systems, correct calibration and synchronization are essential toseparate and locate source signals accurately. This work introduces a methodfor blind source separation and estimation of the Time Difference of Arrival(TDOA) of signals in the time-frequency domain. Our proposed method effectivelyseparates signal mixtures into their original source spectrograms whilesimultaneously estimating the relative delays between receivers, using OptimalTransport (OT) theory. By exploiting the structure of the OT problem, wecombine the separation and delay estimation processes into a unified framework,optimizing the system through a block coordinate descent algorithm. We analyzethe performance of the OT-based estimator under various noise conditions andcompare it with conventional TDOA and source separation methods. Numericalsimulation results demonstrate that our proposed approach can achieve asignificant level of accuracy across diverse noise scenarios for physicalspeech signals in both TDOA and source separation tasks.

【2】 Target Speaker Selection for Neural Network Beamforming in Multi-Speaker  Scenarios
标题:多说话人场景中神经网络束形成的目标说话人选择
链接:https://arxiv.org/abs/2503.18590
作者:Luan Vinícius Fiorio,  Bruno Defraene,  Johan David,  Alex Young,  Frans Widdershoven,  Wim van Houtum,  Ronald M. Aarts
摘要:我们提出了一种用于训练端到端波束形成神经网络的说话人选择机制(SSM),这是基于最近的发现,即听众通常会以一定的下射角注视目标说话人。该机制允许神经网络模型在训练期间,在多说话者场景中,基于收听者和说话者的位置来学习向哪个说话者聚焦。然而,在推断期间仅需要音频信息。我们进行声学模拟证明的可行性和性能时,SSM是在训练中使用。结果表明,显着增加语音清晰度,质量和失真指标相比,最小方差无失真滤波器和相同的神经网络模型训练没有SSM。所提出的方法的成功是朝着解决鸡尾酒会问题迈出的重要一步。
摘要:We propose a speaker selection mechanism (SSM) for the training of anend-to-end beamforming neural network, based on recent findings that a listenerusually looks to the target speaker with a certain undershot angle. Themechanism allows the neural network model to learn toward which speaker tofocus, during training, in a multi-speaker scenario, based on the position oflistener and speakers. However, only audio information is necessary duringinference. We perform acoustic simulations demonstrating the feasibility andperformance when the SSM is employed in training. The results show significantincrease in speech intelligibility, quality, and distortion metrics whencompared to the minimum variance distortionless filter and the same neuralnetwork model trained without SSM. The success of the proposed method is asignificant step forward toward the solution of the cocktail party problem.

【3】 Unsupervised Variational Acoustic Clustering
标题:无监督变分声学聚集
链接:https://arxiv.org/abs/2503.18579
作者:Luan Vinícius Fiorio,  Bruno Defraene,  Johan David,  Frans Widdershoven,  Wim van Houtum,  Ronald M. Aarts
摘要:我们提出了一种无监督的变分声学聚类模型,用于在时频域中对音频数据进行聚类。该模型利用变分推理,扩展到自动编码器框架,高斯混合模型作为潜在空间的先验。专为音频应用而设计,我们引入了一个卷积递归变分自编码器优化高效的时间-频率处理。我们的实验结果表明,与传统方法相比,语音数字数据集在准确性和聚类性能方面有显着提高,展示了该模型捕获复杂音频模式的增强能力。
摘要:We propose an unsupervised variational acoustic clustering model forclustering audio data in the time-frequency domain. The model leveragesvariational inference, extended to an autoencoder framework, with a Gaussianmixture model as a prior for the latent space. Specifically designed for audioapplications, we introduce a convolutional-recurrent variational autoencoderoptimized for efficient time-frequency processing. Our experimental resultsconsidering a spoken digits dataset demonstrate a significant improvement inaccuracy and clustering performance compared to traditional methods, showcasingthe model's enhanced ability to capture complex audio patterns.

【4】 A State-of-the-Art Review on Acoustic Preservation of Historical Worship  Spaces through Auralization
标题:通过听觉化保护历史崇拜空间的最新进展评论
链接:https://arxiv.org/abs/2503.18022
作者:Hannes Rosseel,  Toon van Waterschoot
备注:32 pages, 7 figures, 4 tables, Published in Signal Processing
摘要:历史崇拜空间是具有文化和精神价值的重要建筑地标。这些空间的声学特性在历史和当代的宗教礼拜仪式、仪式和仪式以及神圣音乐的表演中发挥着至关重要的作用。然而,这些空间的原始声学特性往往由于重新利用,翻新,自然灾害或随着时间的推移而恶化而面临风险。本文对声学信号的获取、分析和合成的研究现状进行了全面的综述,重点介绍了HWS。在比利时布鲁塞尔的拿骚教堂的一个例子的案例研究,展示了这些技术的应用程序的历史崇拜空间声学的保存和可听化。本文最后讨论了该领域的挑战和机遇,并概述了未来的研究方向。
摘要:Historical Worship Spaces (HWS) are significant architectural landmarks whichhold both cultural and spiritual value. The acoustic properties of these spacesplay a crucial role in historical and contemporary religious liturgies,rituals, and ceremonies, as well as in the performance of sacred music.However, the original acoustic characteristics of these spaces are often atrisk due to repurposing, renovations, natural disasters, or deterioration overtime. This paper presents a comprehensive review of the current state ofresearch on the acquisition, analysis, and synthesis of acoustics, with a focuson HWS. An example case study of the Nassau chapel in Brussels, Belgium, ispresented to demonstrate the application of these techniques for thepreservation and auralization of historical worship space acoustics. The paperconcludes with a discussion of the challenges and opportunities in the field,and outlines future research directions.

【5】 Mixed-gradients Distributed Filtered Reference Least Mean Square  Algorithm -- A Robust Distributed Multichannel Active Noise Control Algorithm
标题:混合梯度分布式过滤参考最小均方算法--一种鲁棒的分布式多通道主动噪音控制算法
链接:https://arxiv.org/abs/2503.17634
作者:Junwei Ji,  Dongyuan Shi,  Woon-Seng Gan
备注:None
摘要:分布式多通道有源噪声控制(DMCANC),它利用多个单独的处理器,以实现与传统的集中式多通道有源噪声控制(MCANC)相比的全局降噪性能,由于其高计算效率而变得越来越有吸引力。然而,大多数当前DMCANC算法忽略跨节点串扰的影响,并强加没有通信限制的理想网络的假设,这是一个不切实际的假设。因此,这项工作提出了一个强大的DMCANC算法,采用补偿滤波器,以减轻串扰的影响。所提出的解决方案提高了DMCANC系统的灵活性和安全性,利用本地梯度,而不是本地控制滤波器来传达增强的信息,从而在混合梯度分布式滤波参考最小均方(MGDFxLMS)算法。性能调查表明,所提出的方法表现良好的集中式方法。此外,为了解决分布式网络中的通信延迟问题,实现了一种实用的策略,根据延迟的样本自动缩小步长值,以提高系统的弹性。数值仿真结果表明,所提出的自收缩步长MGDFxLMS(ASSS-MGDFxLMS)算法在不同通信时延下的有效性,具有一定的实用价值。
摘要:Distributed multichannel active noise control (DMCANC), which utilizesmultiple individual processors to achieve a global noise reduction performancecomparable to conventional centralized multichannel active noise control(MCANC), has become increasingly attractive due to its high computationalefficiency. However, the majority of current DMCANC algorithms disregard theimpact of crosstalk across nodes and impose the assumption of an ideal networkdevoid of communication limitations, which is an unrealistic assumption.Therefore, this work presents a robust DMCANC algorithm that employs thecompensating filter to mitigate the impact of crosstalk. The proposed solutionenhances the DMCANC system's flexibility and security by utilizing localgradients instead of local control filters to convey enhanced information,resulting in a mixed-gradients distributed filtered reference least mean square(MGDFxLMS) algorithm. The performance investigation demonstrates that theproposed approach performs well with the centralized method. Furthermore, toaddress the issue of communication delay in the distributed network, apractical strategy that auto-shrinks the step size value in response to thedelayed samples is implemented to improve the system's resilience. Thenumerical simulation results demonstrate the efficacy of the proposedauto-shrink step size MGDFxLMS (ASSS-MGDFxLMS) algorithm across variouscommunication delays, highlighting its practical value.

【6】 A Reliable and Efficient Detection Pipeline for Rodent Ultrasonic  Vocalizations
标题:可靠有效的啮齿动物超声发声检测管道
链接:https://arxiv.org/abs/2503.18928
作者:Sabah Shahnoor Anis,  Devin M. Kellis,  Kris Ford Kaigler,  Marlene A. Wilson,  Christian O'Reilly
备注:Accepted for publication in the proceeding of the 7th International Conference on Advances in Signal Processing and Artificial Intelligence (ASPAI' 2025), 8-10 April 2025, Innsbruck, Austria
摘要:分析超声波发声(USVs)对于了解啮齿动物的情感状态和社会行为至关重要,但人工分析耗时且容易出错。自动USV检测系统已经被开发来解决这些挑战。然而,这些系统通常依赖于机器学习,无法有效地推广到新的数据集。为了解决这些缺点,我们引入了ContourUSV,这是一种用于从音频记录中检测USV的高效自动化系统。我们的流程包括频谱图生成、清理、预处理、轮廓检测、后处理以及针对手动注释的评估。为了确保鲁棒性和可靠性,我们使用现有的开放访问USV数据集(USVSEG)和我们与本文一起公开发布的第二个数据集将ContourUSV与三个最先进的系统进行了比较。平均而言,在两个数据集上,ContourUSV的性能优于其他三个系统,精度提高了1.51倍,召回率提高了1.17倍,F1得分提高了1.80倍,特异性提高了1.49倍,同时实现了117.07倍的平均加速。
摘要:Analyzing ultrasonic vocalizations (USVs) is crucial for understandingrodents' affective states and social behaviors, but the manual analysis istime-consuming and prone to errors. Automated USV detection systems have beendeveloped to address these challenges. Yet, these systems often rely on machinelearning and fail to generalize effectively to new datasets. To tackle theseshortcomings, we introduce ContourUSV, an efficient automated system fordetecting USVs from audio recordings. Our pipeline includes spectrogramgeneration, cleaning, pre-processing, contour detection, post-processing, andevaluation against manual annotations. To ensure robustness and reliability, wecompared ContourUSV with three state-of-the-art systems using an existingopen-access USV dataset (USVSEG) and a second dataset we are releasing publiclyalong with this paper. On average, across the two datasets, ContourUSVoutperformed the other three systems with a 1.51x improvement in precision,1.17x in recall, 1.80x in F1 score, and 1.49x in specificity while achieving anaverage speedup of 117.07x.

【7】 Seeing Speech and Sound: Distinguishing and Locating Audios in Visual  Scenes
标题:看到语音和声音:区分和定位视觉场景中的听众
链接:https://arxiv.org/abs/2503.18880
作者:Hyeonggon Ryu,  Seongyu Kim,  Joon Son Chung,  Arda Senocak
备注:CVPR 2025
摘要:我们提出了一个统一的模型,能够同时接地口语和非语音声音在视觉场景中,解决当前视听接地模型的关键限制。现有的方法通常限于独立地处理语音或非语音声音,或者最好一起但顺序地处理而不混合。这一限制使他们无法捕捉经常混合的真实世界音频源的复杂性。我们的方法引入了一个“混合和分离”的框架与视听对齐的目标,共同学习对应和使用混合音频解开。通过这些目标,我们的模型学会为每种音频类型产生不同的嵌入,从而能够在混合音频源中有效地解开纠缠和接地。此外,我们创建了一个新的数据集来评估混合音频源的同时接地,证明我们的模型优于以前的方法。我们的方法在标准分割和跨模态检索任务中也取得了相当或更好的性能,突出了我们的混合和分离方法的好处。
摘要:We present a unified model capable of simultaneously grounding both spokenlanguage and non-speech sounds within a visual scene, addressing keylimitations in current audio-visual grounding models. Existing approaches aretypically limited to handling either speech or non-speech sounds independently,or at best, together but sequentially without mixing. This limitation preventsthem from capturing the complexity of real-world audio sources that are oftenmixed. Our approach introduces a 'mix-and-separate' framework with audio-visualalignment objectives that jointly learn correspondence and disentanglementusing mixed audio. Through these objectives, our model learns to producedistinct embeddings for each audio type, enabling effective disentanglement andgrounding across mixed audio sources. Additionally, we created a new dataset toevaluate simultaneous grounding of mixed audio sources, demonstrating that ourmodel outperforms prior methods. Our approach also achieves comparable orbetter performance in standard segmentation and cross-modal retrieval tasks,highlighting the benefits of our mix-and-separate approach.

【8】 Wireless Hearables With Programmable Speech AI Accelerators
标题:配备可编程语音AI加速器的无线可听设备
链接:https://arxiv.org/abs/2503.18698
作者:Malek Itani,  Tuochao Chen,  Arun Raghavan,  Gavriel Kohlberg,  Shyamnath Gollakota
摘要:None
摘要:The conventional wisdom has been that designing ultra-compact,battery-constrained wireless hearables with on-device speech AI models ischallenging due to the high computational demands of streaming deep learningmodels. Speech AI models require continuous, real-time audio processing,imposing strict computational and I/O constraints. We present NeuralAids, afully on-device speech AI system for wireless hearables, enabling real-timespeech enhancement and denoising on compact, battery-constrained devices. Oursystem bridges the gap between state-of-the-art deep learning for speechenhancement and low-power AI hardware by making three key technicalcontributions: 1) a wireless hearable platform integrating a speech AIaccelerator for efficient on-device streaming inference, 2) an optimizeddual-path neural network designed for low-latency, high-quality speechenhancement, and 3) a hardware-software co-design that uses mixed-precisionquantization and quantization-aware training to achieve real-time performanceunder strict power constraints. Our system processes 6 ms audio chunks inreal-time, achieving an inference time of 5.54 ms while consuming 71.6 mW. Inreal-world evaluations, including a user study with 28 participants, our systemoutperforms prior on-device models in speech quality and noise suppression,paving the way for next-generation intelligent wireless hearables that canenhance hearing entirely on-device.

【9】 Music Similarity Representation Learning Focusing on Individual  Instruments with Source Separation and Human Preference
标题:音乐相似性表示学习专注于单个乐器,具有来源分离和人类偏好
链接:https://arxiv.org/abs/2503.18486
作者:Takehiro Imamura,  Yuka Hashizume,  Wen-Chin Huang,  Tomoki Toda
摘要:本文提出了基于单个乐器声音(InMSRL)的音乐相似性表示学习(MSRL),利用音乐源分离(MSS)和人类偏好,而不需要在推理过程中使用干净的乐器声音。我们提出了三种有效提高性能的方法。首先,我们介绍端到端微调(E2 E-FT)的级联方法,顺序执行MSS和音乐相似性特征提取。E2 E-FT允许模型最小化分离误差对特征提取的不利影响。其次,我们提出了多任务学习的直接方法,直接提取解开音乐相似性特征,使用一个单一的音乐相似性特征提取器。基于分离的音乐相似性特征提取的多任务学习和基于分离的音乐相似性特征重构的MSS进一步增强了乐器特征分离。第三,我们采用感知感知微调(PAFT)。PAFT利用人类偏好,允许模型执行与人类感知相似性对齐的InMSRL。我们进行了实验评估,并证明1)E2 E-FT for Cascade显着提高了InMSRL性能,2)Direct的多任务学习也有助于提高特征提取中的解纠缠性能,3)PAFT显着提高了感知InMSRL性能,4)Cascade with E2 E-FT and PAFT优于Direct with the multitask learning and PAFT。
摘要:This paper proposes music similarity representation learning (MSRL) based onindividual instrument sounds (InMSRL) utilizing music source separation (MSS)and human preference without requiring clean instrument sounds duringinference. We propose three methods that effectively improve performance.First, we introduce end-to-end fine-tuning (E2E-FT) for the Cascade approachthat sequentially performs MSS and music similarity feature extraction. E2E-FTallows the model to minimize the adverse effects of a separation error on thefeature extraction. Second, we propose multi-task learning for the Directapproach that directly extracts disentangled music similarity features using asingle music similarity feature extractor. Multi-task learning, which is basedon the disentangled music similarity feature extraction and MSS based onreconstruction with disentangled music similarity features, further enhancesinstrument feature disentanglement. Third, we employ perception-awarefine-tuning (PAFT). PAFT utilizes human preference, allowing the model toperform InMSRL aligned with human perceptual similarity. We conductexperimental evaluations and demonstrate that 1) E2E-FT for Cascadesignificantly improves InMSRL performance, 2) the multi-task learning forDirect is also helpful to improve disentanglement performance in the featureextraction, 3) PAFT significantly enhances the perceptual InMSRL performance,and 4) Cascade with E2E-FT and PAFT outperforms Direct with the multi-tasklearning and PAFT.

【10】 Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and  Speech Recognition
标题:通过将说话人分离和语音识别脱钩提高鲁棒的多说话者ASB
链接:https://arxiv.org/abs/2503.17886
作者:Yufeng Yang,  Hassan Taherian,  Vahid Ahmadi Kalkhorani,  DeLiang Wang
摘要:尽管随着深度学习的引入,自动语音识别(ASR)取得了巨大的成功,但在许多现实世界的多人场景中,其性能仍然不能令人满意。说话者分离在分离单个说话者方面表现出色,但作为前端,它引入了处理伪像,这些伪像会降低在干净语音上训练的ASR后端。因此,主流的鲁棒ASR系统训练关于噪声语音的后端以避免处理伪像。在这项工作中,我们建议将说话人分离前端和ASR后端的训练解耦,后者仅在干净的语音上训练。我们的解耦系统在Libri 2 Mix开发/测试集上实现了5.1%的字错误率(WER),显著优于其他多通话器ASR基线。其有效性也通过1通道和6通道SMS-WSJ上最先进的7.60%/5.74%WER得到了证明。此外,在记录的LibriCSS上,我们实现了2.92%的说话者归因WER。这些最先进的结果表明,解耦说话人分离和识别是一种有效的方法,以提高鲁棒的多说话人ASR。
摘要:Despite the tremendous success of automatic speech recognition (ASR) with theintroduction of deep learning, its performance is still unsatisfactory in manyreal-world multi-talker scenarios. Speaker separation excels in separatingindividual talkers but, as a frontend, it introduces processing artifacts thatdegrade the ASR backend trained on clean speech. As a result, mainstream robustASR systems train the backend on noisy speech to avoid processing artifacts. Inthis work, we propose to decouple the training of the speaker separationfrontend and the ASR backend, with the latter trained on clean speech only. Ourdecoupled system achieves 5.1% word error rates (WER) on the Libri2Mix dev/testsets, significantly outperforming other multi-talker ASR baselines. Itseffectiveness is also demonstrated with the state-of-the-art 7.60%/5.74% WERson 1-ch and 6-ch SMS-WSJ. Furthermore, on recorded LibriCSS, we achieve thespeaker-attributed WER of 2.92%. These state-of-the-art results suggest thatdecoupling speaker separation and recognition is an effective approach toelevate robust multi-talker ASR.

【11】 Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for  High Quality Data Expansion
标题:具有潜在空间扩展的音频增强视觉语言建模以实现高质量数据扩展
链接:https://arxiv.org/abs/2503.17551
作者:Yu Sun,  Yin Li,  Ruixiao Sun,  Chunhui Liu,  Fangming Zhou,  Ze Jin,  Linjie Wang,  Xiang Shen,  Zhuolin Hao,  Hongyu Xiong
摘要:基于transformer的多模态模型广泛用于工业规模的推荐、搜索和广告系统,用于内容理解和相关性排名。增强标记的训练数据质量和跨模态融合显著提高了模型性能,影响了质量观看率和广告收入等关键指标。高质量的注释对于推进内容建模至关重要,但传统的基于语义的主动学习(AL)方法面临着局限性:它们很难检测到过度自信的错误分类,并且在区分深度神经网络中语义相似的项目方面效果较差。此外,音频信息扮演着越来越重要的角色,特别是在短视频平台中,但大多数预训练的多模式架构主要关注文本和图像。虽然可以在所有三种模式中从头开始训练,但它牺牲了利用现有预训练的视觉语言(VL)和音频模型的好处。为了解决这些挑战,我们提出了基于kNN的潜在空间扩展(LSB)来提高AL效率和具有音频增强的视觉语言建模(VLMAE),这是一种将音频集成到VL模型中的中间融合方法。此系统部署在生产系统中,带来了显著的业务收益。
摘要:Transformer-based multimodal models are widely used in industrial-scalerecommendation, search, and advertising systems for content understanding andrelevance ranking. Enhancing labeled training data quality and cross-modalfusion significantly improves model performance, influencing key metrics suchas quality view rates and ad revenue. High-quality annotations are crucial foradvancing content modeling, yet traditional statistical-based active learning(AL) methods face limitations: they struggle to detect overconfidentmisclassifications and are less effective in distinguishing semanticallysimilar items in deep neural networks. Additionally, audio information plays anincreasing role, especially in short-video platforms, yet most pre-trainedmultimodal architectures primarily focus on text and images. While trainingfrom scratch across all three modalities is possible, it sacrifices thebenefits of leveraging existing pre-trained visual-language (VL) and audiomodels. To address these challenges, we propose kNN-based Latent SpaceBroadening (LSB) to enhance AL efficiency and Vision-Language Modeling withAudio Enhancement (VLMAE), a mid-fusion approach integrating audio into VLmodels. This system deployed in production systems, leading to significantbusiness gains.

机器翻译由腾讯交互翻译提供,仅供参考