今日论文合集:cs.SD语音5篇,eess.AS音频处理2篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
标题: 释放具有多个声音源的自然音频的力量
链接:https://arxiv.org/abs/2504.17782
作者: Xize Cheng,  Slytherin Wang,  Zehan Wang,  Rongjie Huang,  Tao Jin,  Zhou Zhao 
备注:Work in Progress
摘要:通用声音分离旨在从混合音频中提取与不同事件对应的干净音轨,这对于人工听觉感知至关重要。然而,目前的方法严重依赖于人工混合音频进行训练,这限制了它们推广到现实环境中收集的自然混合音频的能力。为了克服这一限制,我们提出了ClearSep,这是一个创新的框架,它采用数据引擎将复杂的自然混合音频分解为多个独立的音轨,从而在现实世界的场景中实现有效的声音分离。我们引入了两个基于混音的评估指标来定量评估分离质量,并将这些指标作为阈值,在模型训练的同时迭代应用数据引擎,逐步优化分离性能。此外,我们提出了一系列的培训策略,针对这些分离的独立的轨道,使他们发挥最大的作用。大量实验表明,ClearSep在多个声音分离任务中实现了最先进的性能,突出了其在自然音频场景中推进声音分离的潜力。有关更多示例和详细结果,请访问我们的演示页面https://clearsep.github.io。
摘要:Universal sound separation aims to extract clean audio tracks corresponding to distinct events from mixed audio, which is critical for artificial auditory perception. However, current methods heavily rely on artificially mixed audio for training, which limits their ability to generalize to naturally mixed audio collected in real-world environments. To overcome this limitation, we propose ClearSep, an innovative framework that employs a data engine to decompose complex naturally mixed audio into multiple independent tracks, thereby allowing effective sound separation in real-world scenarios. We introduce two remix-based evaluation metrics to quantitatively assess separation quality and use these metrics as thresholds to iteratively apply the data engine alongside model training, progressively optimizing separation performance. In addition, we propose a series of training strategies tailored to these separated independent tracks to make the best use of them. Extensive experiments demonstrate that ClearSep achieves state-of-the-art performance across multiple sound separation tasks, highlighting its potential for advancing sound separation in natural audio scenarios. For more examples and detailed results, please visit our demo page at https://clearsep.github.io.


【2】 A Machine Learning Approach for Denoising and Upsampling HRTFs

标题: HRTI去噪和上采样的机器学习方法
链接:https://arxiv.org/abs/2504.17586
作者: Xuyi Hu,  Jian Li,  Lorenzo Picinali,  Aidan O. T. Hogg 
摘要:对逼真的虚拟沉浸式音频的需求持续增长,其中头部相关传递函数(HRTF)发挥着关键作用。HRTF捕捉声音如何到达我们的耳朵,反映了独特的解剖学特征并增强了空间感知。已经表明,个性化的HRTF提高了定位精度,但它们的测量仍然很耗时,并且需要无噪声的环境。虽然机器学习已经被证明可以减少所需的测量点,从而减少测量时间,但仍然需要受控的环境。本文提出了一种方法来解决这个问题,提出了一种新的技术,可以上采样稀疏,嘈杂的HRTF测量。所提出的方法结合了用于去噪的HRTF Denoisy U-Net和用于从三个测量点进行上采样的自动编码生成对抗网络(AE-GAN)。该方法实现了5.41 dB的对数谱失真(LSD)误差和0.0070的余弦相似性损失,证明了该方法在HRTF上采样的有效性。
摘要:The demand for realistic virtual immersive audio continues to grow, with Head-Related Transfer Functions (HRTFs) playing a key role. HRTFs capture how sound reaches our ears, reflecting unique anatomical features and enhancing spatial perception. It has been shown that personalized HRTFs improve localization accuracy, but their measurement remains time-consuming and requires a noise-free environment. Although machine learning has been shown to reduce the required measurement points and, thus, the measurement time, a controlled environment is still necessary. This paper proposes a method to address this constraint by presenting a novel technique that can upsample sparse, noisy HRTF measurements. The proposed approach combines an HRTF Denoisy U-Net for denoising and an Autoencoding Generative Adversarial Network (AE-GAN) for upsampling from three measurement points. The proposed method achieves a log-spectral distortion (LSD) error of 5.41 dB and a cosine similarity loss of 0.0070, demonstrating the method's effectiveness in HRTF upsampling.


【3】 Waveform-Logmel Audio Neural Networks for Respiratory Sound  Classification

标题: 用于呼吸声分类的Wave-Logmel音频神经网络
链接:https://arxiv.org/abs/2504.17156
作者: Jiadong Xie,  Yunlian Zhou,  Mingsheng Xu 
摘要:使用电子听诊器的听诊分析在呼吸系统疾病的临床诊断中越来越受到关注。最近,神经网络已经被应用于辅助呼吸声分类,并取得了成就。然而,由于异常呼吸音的缺乏,它仍然具有挑战性。在本文中,我们提出了一种新的架构,即波形Logmel音频神经网络(WLANN),它使用波形和logmel频谱图作为输入功能,并使用双向门控递归单元(Bi-GRU)的上下文模型融合的功能。将WLANN应用于SPRSound呼吸数据集的实验结果表明,该框架可以有效地区分病理性呼吸音类,优于以往的研究,灵敏度为90.3%,总分为93.6%。我们的研究证明了WLANN在呼吸系统疾病诊断中的高效性。
摘要:Auscultatory analysis using an electronic stethoscope has attracted increasing attention in the clinical diagnosis of respiratory diseases. Recently, neural networks have been applied to assist in respiratory sound classification with achievements. However, it remains challenging due to the scarcity of abnormal respiratory sound. In this paper, we propose a novel architecture, namely Waveform-Logmel audio neural networks (WLANN), which uses both waveform and log-mel spectrogram as the input features and uses Bidirectional Gated Recurrent Units (Bi-GRU) to context model the fused features. Experimental results of our WLANN applied to SPRSound respiratory dataset show that the proposed framework can effectively distinguish pathological respiratory sound classes, outperforming the previous studies, with 90.3% in sensitivity and 93.6% in total score. Our study demonstrates the high effectiveness of the WLANN in the diagnosis of respiratory diseases.


【4】 Multifaceted Evaluation of Audio-Visual Capability for MLLMs:  Effectiveness, Efficiency, Generalizability and Robustness

标题: MLLM视听能力的多方面评估:有效性、效率、概括性和稳健性
链接:https://arxiv.org/abs/2504.16936
作者: Yusheng Zhao,  Junyu Luo,  Xiao Luo,  Weizhi Zhang,  Zhiping Xiao,  Wei Ju,  Philip S. Yu,  Ming Zhang 
摘要:多模态大型语言模型(MLLM)最近在处理和理解来自不同模态的信息方面取得了巨大成功(例如,文本、音频和可视信号)。尽管这些模型越来越受欢迎,但仍然缺乏衡量这些模型视听能力的综合评估,特别是在不同的场景中(例如,分布变化和对抗性攻击)。在本文中,我们提出了一个多方面的评估的视听能力的MLLM,专注于四个关键方面:有效性,效率,概括性和鲁棒性。通过大量的实验,我们发现MLLM表现出很强的zero-shot和Few-Shot泛化能力,使他们能够在有限的数据下取得很好的性能。然而,它们的成功在很大程度上依赖于视觉模态,当视觉输入损坏或丢失时,视觉模态会损害性能。此外,虽然MLLM容易受到对抗性样本的影响,但与传统模型相比,它们表现出更强的鲁棒性。实验结果和我们的研究结果提供了对MLLM视听能力的见解,突出了需要改进的领域,并为未来的研究提供了指导。
摘要:Multi-modal large language models (MLLMs) have recently achieved great success in processing and understanding information from diverse modalities (e.g., text, audio, and visual signals). Despite their growing popularity, there remains a lack of comprehensive evaluation measuring the audio-visual capabilities of these models, especially in diverse scenarios (e.g., distribution shifts and adversarial attacks). In this paper, we present a multifaceted evaluation of the audio-visual capability of MLLMs, focusing on four key dimensions: effectiveness, efficiency, generalizability, and robustness. Through extensive experiments, we find that MLLMs exhibit strong zero-shot and few-shot generalization abilities, enabling them to achieve great performance with limited data. However, their success relies heavily on the vision modality, which impairs performance when visual input is corrupted or missing. Additionally, while MLLMs are susceptible to adversarial samples, they demonstrate greater robustness compared to traditional models. The experimental results and our findings provide insights into the audio-visual capabilities of MLLMs, highlighting areas for improvement and offering guidance for future research.


【5】 Unsupervised EEG-based decoding of absolute auditory attention with  canonical correlation analysis

标题: 基于无监督的绝对听觉注意力解码和典型相关分析
链接:https://arxiv.org/abs/2504.17724
作者: Nicolas Heintz,  Tom Francart,  Alexander Bertrand 
摘要:我们提出了一种完全无监督的算法,该算法可以从脑电图(EEG)记录中检测受试者何时主动倾听声音,而不是何时忽略声音。这个问题被称为绝对听觉注意解码(aAAD)。我们提出了一个无监督的判别CCA模型进行特征提取,并将其与一个无监督的分类器称为最小信息线性判别分析(MILDA)aAAD分类。值得注意的是,所提出的无监督算法的性能明显优于最先进的监督模型。一个关键的原因是,无监督算法可以成功地适应非平稳的测试数据,在一个较低的计算成本。这为使用EEG信号分析受试者的听觉注意力打开了大门,该模型可以自动调整自身以适应受试者,而无需事先进行艰苦的监督训练。
摘要:We propose a fully unsupervised algorithm that detects from encephalography (EEG) recordings when a subject actively listens to sound, versus when the sound is ignored. This problem is known as absolute auditory attention decoding (aAAD). We propose an unsupervised discriminative CCA model for feature extraction and combine it with an unsupervised classifier called minimally informed linear discriminant analysis (MILDA) for aAAD classification. Remarkably, the proposed unsupervised algorithm performs significantly better than a state-of-the-art supervised model. A key reason is that the unsupervised algorithm can successfully adapt to the non-stationary test data at a low computational cost. This opens the door to the analysis of the auditory attention of a subject using EEG signals with a model that automatically tunes itself to the subject without requiring an arduous supervised training session beforehand.


eess.AS音频处理

【1】 Generating Localized Audible Zones Using a Single-Channel Parametric  Loudspeaker
标题: 使用单通道参数扬声器生成局部可听区
链接:https://arxiv.org/abs/2504.17440
作者: Tao Zhuang,  Shaozhe Li,  Feng Niu,  Jia-Xin Zhong,  Jing Lu 
摘要:先进的声音区域控制(SZC)技术通常依赖于大规模的多声道扬声器阵列来创建高对比度的个人声音区域,使得单扬声器SZC似乎是不可能的。在这封信中,我们通过引入多载波参量扬声器(MCPL)来挑战这种模式,该扬声器仅使用单个扬声器即可实现SZC。在我们的方法中,不同的音频信号被调制到不同频率的单独的超声波载波上,并组合成一个单一的复合信号。该信号由单通道超声换能器发射,通过空气中的非线性解调,音频信号相互作用,虚拟地形成多通道输出。这种新的能力允许应用现有的SZC算法最初设计的多通道扬声器阵列。仿真验证了我们提出的单声道MCPL的有效性,证明了它作为传统多扬声器系统的一种有前途的替代方案的潜力,以实现高对比度SZC。我们的工作为简化SZC系统而不影响性能开辟了新的途径。
摘要:Advanced sound zone control (SZC) techniques typically rely on massive multi-channel loudspeaker arrays to create high-contrast personal sound zones, making single-loudspeaker SZC seem impossible. In this Letter, we challenge this paradigm by introducing the multi-carrier parametric loudspeaker (MCPL), which enables SZC using only a single loudspeaker. In our approach, distinct audio signals are modulated onto separate ultrasonic carrier waves at different frequencies and combined into a single composite signal. This signal is emitted by a single-channel ultrasonic transducer, and through nonlinear demodulation in air, the audio signals interact to virtually form multi-channel outputs. This novel capability allows the application of existing SZC algorithms originally designed for multi-channel loudspeaker arrays. Simulations validate the effectiveness of our proposed single-channel MCPL, demonstrating its potential as a promising alternative to traditional multi-loudspeaker systems for achieving high-contrast SZC. Our work opens new avenues for simplifying SZC systems without compromising performance.


【2】 Multifaceted Evaluation of Audio-Visual Capability for MLLMs:  Effectiveness, Efficiency, Generalizability and Robustness

标题: MLLM视听能力的多方面评估:有效性、效率、概括性和稳健性
链接:https://arxiv.org/abs/2504.16936
作者: Yusheng Zhao,  Junyu Luo,  Xiao Luo,  Weizhi Zhang,  Zhiping Xiao,  Wei Ju,  Philip S. Yu,  Ming Zhang 
摘要:多模态大型语言模型(MLLM)最近在处理和理解来自不同模态的信息方面取得了巨大成功(例如,文本、音频和可视信号)。尽管这些模型越来越受欢迎,但仍然缺乏衡量这些模型视听能力的综合评估,特别是在不同的场景中(例如,分布变化和对抗性攻击)。在本文中,我们提出了一个多方面的评估的视听能力的MLLM,专注于四个关键方面:有效性,效率,概括性和鲁棒性。通过大量的实验,我们发现MLLM表现出很强的zero-shot和Few-Shot泛化能力,使他们能够在有限的数据下取得很好的性能。然而,它们的成功在很大程度上依赖于视觉模态,当视觉输入损坏或丢失时,视觉模态会损害性能。此外,虽然MLLM容易受到对抗性样本的影响,但与传统模型相比,它们表现出更强的鲁棒性。实验结果和我们的研究结果提供了对MLLM视听能力的见解,突出了需要改进的领域,并为未来的研究提供了指导。
摘要:Multi-modal large language models (MLLMs) have recently achieved great success in processing and understanding information from diverse modalities (e.g., text, audio, and visual signals). Despite their growing popularity, there remains a lack of comprehensive evaluation measuring the audio-visual capabilities of these models, especially in diverse scenarios (e.g., distribution shifts and adversarial attacks). In this paper, we present a multifaceted evaluation of the audio-visual capability of MLLMs, focusing on four key dimensions: effectiveness, efficiency, generalizability, and robustness. Through extensive experiments, we find that MLLMs exhibit strong zero-shot and few-shot generalization abilities, enabling them to achieve great performance with limited data. However, their success relies heavily on the vision modality, which impairs performance when visual input is corrupted or missing. Additionally, while MLLMs are susceptible to adversarial samples, they demonstrate greater robustness compared to traditional models. The experimental results and our findings provide insights into the audio-visual capabilities of MLLMs, highlighting areas for improvement and offering guidance for future research.


机器翻译由腾讯交互翻译提供,仅供参考