今日论文合集:cs.SD语音5篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Whispering Under the Eaves: Protecting User Privacy Against Commercial  and LLM-powered Automatic Speech Recognition Systems
标题: 屋檐下低语:保护用户隐私免受商业和LLM支持的自动语音识别系统的侵害
链接:https://arxiv.org/abs/2504.00858
作者: Weifei Jin,  Yuxin Cao,  Junjie Su,  Derui Wang,  Yedi Zhang,  Minhui Xue,  Jie Hao,  Jin Song Dong,  Yixian Yang 
备注:Accept to USENIX Security 2025
摘要:自动语音识别(ASR)的广泛应用支持大规模语音监控,引发了对用户隐私的担忧。在本文中,我们专注于使用对抗性的例子,以减轻未经授权的语音隐私泄露挫败潜在的窃听者在语音通信。虽然音频对抗示例已经证明了误导ASR模型或逃避ASR监视的能力,但它们通常是通过时间密集型离线优化构建的,限制了它们在实时语音通信中的实用性。最近的工作通过生成通用对抗扰动(UAP)并增强其黑盒场景的可转移性来克服这一限制。然而,它们引入了过多的噪声,显著降低了音频质量并影响了人类的感知,从而限制了它们在实际场景中的有效性。为了解决这个限制,并保护实时用户的语音对ASR系统,我们提出了一个新的框架,AudioShield。这个框架的核心是潜在空间中的可转移普遍对抗扰动(LS-TUAP)的概念。通过将扰动转移到潜在空间,在很大程度上保持了音频质量。此外,我们提出了目标特征自适应,以提高转让的UAP嵌入目标文本特征的扰动。通过对4个商用ASR API(Google、Amazon、科大讯飞、阿里巴巴)、3个语音助手、2个基于LLM的ASR和1个基于NN的ASR的综合评测,证明了AudioShield相对于现有竞争对手的保护优势,客观和主观评测均表明AudioShield显著提高了音频质量。此外,AudioShield在实时端到端场景中也表现出很高的有效性,并对自适应对策表现出很强的弹性。
摘要:The widespread application of automatic speech recognition (ASR) supports large-scale voice surveillance, raising concerns about privacy among users. In this paper, we concentrate on using adversarial examples to mitigate unauthorized disclosure of speech privacy thwarted by potential eavesdroppers in speech communications. While audio adversarial examples have demonstrated the capability to mislead ASR models or evade ASR surveillance, they are typically constructed through time-intensive offline optimization, restricting their practicality in real-time voice communication. Recent work overcame this limitation by generating universal adversarial perturbations (UAPs) and enhancing their transferability for black-box scenarios. However, they introduced excessive noise that significantly degrades audio quality and affects human perception, thereby limiting their effectiveness in practical scenarios. To address this limitation and protect live users' speech against ASR systems, we propose a novel framework, AudioShield. Central to this framework is the concept of Transferable Universal Adversarial Perturbations in the Latent Space (LS-TUAP). By transferring the perturbations to the latent space, the audio quality is preserved to a large extent. Additionally, we propose target feature adaptation to enhance the transferability of UAPs by embedding target text features into the perturbations. Comprehensive evaluation on four commercial ASR APIs (Google, Amazon, iFlytek, and Alibaba), three voice assistants, two LLM-powered ASR and one NN-based ASR demonstrates the protection superiority of AudioShield over existing competitors, and both objective and subjective evaluations indicate that AudioShield significantly improves the audio quality. Moreover, AudioShield also shows high effectiveness in real-time end-to-end scenarios, and demonstrates strong resilience against adaptive countermeasures.


【2】 A Survey on Music Generation from Single-Modal, Cross-Modal, and  Multi-Modal Perspectives: Data, Methods, and Challenges
标题: 从单模式、跨模式和多模式角度对音乐生成的调查:数据、方法和挑战
链接:https://arxiv.org/abs/2504.00837

作者: Shuyu Li,  Shulei Ji,  Zihao Wang,  Songruoyao Wu,  Jiaxing Yu,  Kejun Zhang 
摘要:多模态音乐生成,使用多种模态,如图像,视频和文本以及乐谱和音频作为指导,是一个新兴的研究领域,具有广泛的应用。本文从模态的角度对音乐生成系统进行了分类。它涵盖了模态表示,多模态数据对齐,并利用它们来指导音乐生成。我们还讨论了当前的数据集和评估方法。这一领域的主要挑战包括有效的多模式集成、大规模综合数据集和系统化评估方法。最后,我们提供了一个展望未来的研究方向,重点是多模态融合,对齐,数据和评估。
摘要:Multi-modal music generation, using multiple modalities like images, video, and text alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music generation systems from the perspective of modalities. It covers modality representation, multi-modal data alignment, and their utilization to guide music generation. We also discuss current datasets and evaluation methods. Key challenges in this area include effective multi-modal integration, large-scale comprehensive datasets, and systematic evaluation methods. Finally, we provide an outlook on future research directions focusing on multi-modal fusion, alignment, data, and evaluation.


【3】 $C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker  Extraction
标题: $C^2$AV-TSE:上下文和置信度感知的视听目标说话人提取
链接:https://arxiv.org/abs/2504.00750

作者: Wenxuan Wu,  Xueyuan Chen,  Shuai Wang,  Jiadong Wang,  Lingwei Meng,  Xixin Wu,  Helen Meng,  Haizhou Li 
备注:Accepted by IEEE Journal of Selected Topics in Signal Processing (JSTSP)
摘要:视听目标说话人提取(AV-TSE)的目的是模仿人类的能力,以提高听觉感知使用视觉线索。虽然最近已经提出了许多模型,但它们中的大多数主要依靠声学特征内的局部依赖性来估计目标信号,未充分利用类似人类的能力来通过上下文信息推断不清楚的语音部分。这种限制不仅导致次优的性能,而且在整个话语的提取质量不一致,与一些段表现出质量差或干扰扬声器的抑制不足。为了缩小这一差距,我们提出了一种与模型无关的策略,称为Mask-And-Recover(MAR)。它集成了模态间和模态内的上下文相关性,以实现提取模块内的全局推理。此外,为了更好地针对每个样本中具有挑战性的部分,我们引入了细粒度置信度得分(FCS)模型来评估提取质量,并指导提取模块强调对低质量片段的改进。为了验证我们提出的与模型无关的训练范式的有效性,我们采用了六种流行的AV-TSE主干对VoxCeleb 2数据集进行评估,证明了各种指标的一致性能改进。
摘要:Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals by primarily relying on local dependencies within acoustic features, underutilizing the human-like capacity to infer unclear parts of speech through contextual information. This limitation results in not only suboptimal performance but also inconsistent extraction quality across the utterance, with some segments exhibiting poor quality or inadequate suppression of interfering speakers. To close this gap, we propose a model-agnostic strategy called the Mask-And-Recover (MAR). It integrates both inter- and intra-modality contextual correlations to enable global inference within extraction modules. Additionally, to better target challenging parts within each sample, we introduce a Fine-grained Confidence Score (FCS) model to assess extraction quality and guide extraction modules to emphasize improvement on low-quality segments. To validate the effectiveness of our proposed model-agnostic training paradigm, six popular AV-TSE backbones were adopted for evaluation on the VoxCeleb2 dataset, demonstrating consistent performance improvements across various metrics.


【4】 User authentication on earable devices via bone-conducted occlusion  sounds
标题: 通过骨导咬合声音在可佩戴设备上进行用户验证
链接:https://arxiv.org/abs/2504.00435

作者: Yadong Xie,  Fan Li,  Yue Wu,  Yu Wang 
备注:IEEE Transactions on Dependable and Secure Computing ( Volume: 21, Issue: 4, July-Aug. 2024)
摘要:随着移动设备的快速发展和敏感数据的迅速增加,人们对安全、方便的移动认证技术提出了更高的要求。除了传统的密码之外,许多移动设备具有基于生物特征的认证方法(例如,指纹、声纹和面部识别),但它们容易受到欺骗攻击。为了解决这个问题,我们研究了新的生物特征,这是基于牙齿咬合,并发现骨传导的声音的牙齿咬合收集在双耳管包含独特的功能,个人的骨骼和牙齿。基于此,我们提出了一种新的认证系统,TeethPass+,它使用耳塞来收集双耳管中的咬合声音来实现认证。首先,我们设计了一种基于频谱方差的事件检测方法来检测骨导声。然后,我们分析的时间-频率域的声音过滤掉运动噪声和提取用户的独特特征,从四个方面:牙齿结构,骨骼结构,咬合位置,咬合声音。最后,我们训练了一个Triplet网络来构造用户模板,并利用该模板完成认证。通过对53名志愿者的大量实验,验证了TeethPass+在不同环境下的性能。TeethPass+的准确率达到98.6%,并抵抗99.7%的欺骗攻击。
摘要:With the rapid development of mobile devices and the fast increase of sensitive data, secure and convenient mobile authentication technologies are desired. Except for traditional passwords, many mobile devices have biometric-based authentication methods (e.g., fingerprint, voiceprint, and face recognition), but they are vulnerable to spoofing attacks. To solve this problem, we study new biometric features which are based on the dental occlusion and find that the bone-conducted sound of dental occlusion collected in binaural canals contains unique features of individual bones and teeth. Motivated by this, we propose a novel authentication system, TeethPass+, which uses earbuds to collect occlusal sounds in binaural canals to achieve authentication. First, we design an event detection method based on spectrum variance to detect bone-conducted sounds. Then, we analyze the time-frequency domain of the sounds to filter out motion noises and extract unique features of users from four aspects: teeth structure, bone structure, occlusal location, and occlusal sound. Finally, we train a Triplet network to construct the user template, which is used to complete authentication. Through extensive experiments including 53 volunteers, the performance of TeethPass+ in different environments is verified. TeethPass+ achieves an accuracy of 98.6% and resists 99.7% of spoofing attacks.


【5】 Are you really listening? Boosting Perceptual Awareness in Music-QA  Benchmarks
标题: 你真的在听吗?提高音乐的感知意识-QA基准
链接:https://arxiv.org/abs/2504.00369

作者: Yongyi Zang,  Sean O'Brien,  Taylor Berg-Kirkpatrick,  Julian McAuley,  Zachary Novack 
摘要:大型音频语言模型(LALM),其中预训练的文本LLM与音频输入进行微调,在音乐理解方面取得了显着进展。然而,目前的评估方法表现出严重的局限性:在领先的音乐问题识别基准MuchoMusic上,没有音频感知功能的纯文本LLM实现了高达56.4%的惊人的高准确性,与大多数LALM相当或更高。此外,当呈现随机高斯噪声而不是实际音频时,LALM仍然表现出明显高于机会。这些发现表明,现有的基准主要评估推理能力,而不是音频感知。为了克服这一挑战,我们提出了RUListening:通过倾听,增强音乐QA基准感知评估的框架。我们引入了感知指数(PI),这是一个定量指标,通过分析纯文本语言模型的对数概率分布来衡量问题对音频感知的依赖。使用这个指标,我们生成合成的,具有挑战性的干扰,以创建QA对,需要真正的音频感知。当应用于MuchoMusic时,我们的过滤数据集成功地迫使模型依赖于感知信息-仅文本LLM在偶然水平上执行,而当音频输入被噪声取代时,LALM同样会恶化。这些结果验证了我们的框架在创建更准确地评估音频感知能力的基准方面的有效性。
摘要:Large Audio Language Models (LALMs), where pretrained text LLMs are finetuned with audio input, have made remarkable progress in music understanding. However, current evaluation methodologies exhibit critical limitations: on the leading Music Question Answering benchmark, MuchoMusic, text-only LLMs without audio perception capabilities achieve surprisingly high accuracy of up to 56.4%, on par or above most LALMs. Furthermore, when presented with random Gaussian noise instead of actual audio, LALMs still perform significantly above chance. These findings suggest existing benchmarks predominantly assess reasoning abilities rather than audio perception. To overcome this challenge, we present RUListening: Robust Understanding through Listening, a framework that enhances perceptual evaluation in Music-QA benchmarks. We introduce the Perceptual Index (PI), a quantitative metric that measures a question's reliance on audio perception by analyzing log probability distributions from text-only language models. Using this metric, we generate synthetic, challenging distractors to create QA pairs that necessitate genuine audio perception. When applied to MuchoMusic, our filtered dataset successfully forces models to rely on perceptual information-text-only LLMs perform at chance levels, while LALMs similarly deteriorate when audio inputs are replaced with noise. These results validate our framework's effectiveness in creating benchmarks that more accurately evaluate audio perception capabilities.


eess.AS音频处理


【1】 Expanding and Analyzing ODAQ -- the Open Dataset of Audio Quality
标题: 音频质量开放数据集ODAQ的扩展和分析
链接:https://arxiv.org/abs/2504.00742

作者: Sascha Dick,  Christoph Thompson,  Chih-Wei Wu,  Matteo Torcoli,  Pablo Delgado,  Phillip A. Williams,  Emanuel Habets 
备注:Accepted for presentation at the Audio Engineering Society (AES) 157th Convention, October 2024, New York, USA
摘要:音频质量开放数据集(ODAQ)最近被引入,以解决具有相应主观质量分数的公开可用音频数据集的稀缺性。该数据集在许可证下发布,包括使用六种不同的信号处理方法处理的音频材料,这些方法在五个质量级别上运行,以及相应的主观测试结果。为了扩大数据集,我们为大学生提供了听众培训,以进行进一步的主观测试,并获得了与以前的专家听众一致的结果。我们还展示了不同的训练方法如何影响绝对量表和锚点的使用。扩展后的数据集现在包括来自三个国际实验室的结果,共提供42个听众和10080个主观分数。本文详细介绍了扩展的细节,并进行了深入的分析。作为此分析的一部分,我们开始使用ODAQ作为基准来评估客观音频质量指标预测主观分数的能力
摘要:The Open Dataset of Audio Quality (ODAQ) was recently introduced to address the scarcity of openly available audio datasets with corresponding subjective quality scores. The dataset, released under permissive licenses, comprises audio material processed using six different signal processing methods operating at five quality levels, along with corresponding subjective test results. To expand the dataset, we provided listener training to university students to conduct further subjective tests and obtained results consistent with previous expert listeners. We also showed how different training approaches affect the use of absolute scales and anchors. The expanded dataset now comprises results from three international laboratories providing a total of 42 listeners and 10080 subjective scores. This paper provides the details of the expansion and an in-depth analysis. As part of this analysis, we initiate the use of ODAQ as a benchmark to evaluate objective audio quality metrics in their ability to predict subjective scores


【2】 How Cyclic Acoustic Patterns Influence ASMR Perception: A Signal  Processing Perspective
标题: 周期性声学模式如何影响ASMR感知:信号处理的角度
链接:https://arxiv.org/abs/2504.00621

作者: Zexin Fang,  Bin Han,  Henrik H. Sveen,  C. Clark Cao,  Hans D. Schotten 
备注:Submitted to IEEE Signal Processing Letters
摘要:自主感觉经络反应(ASMR)是近十年来研究的热点。虽然它的效果已经通过行为研究和神经生理学测量(如脑电图)和相关的生物信号分析进行了验证,但它的发展和触发仍然是一个争论的主题。以前的研究表明,它的触发器与循环模式高度相关:可预测的模式引入放松,而变化则保持吸引力。为了验证这一点并进一步了解声学特征对ASMR效应的影响,我们设计了三种不同的循环模式,具有单声道和立体声变化,同时控制其可预测性和随机性,并通过在线调查收集ASMR触发分数。然后,我们提取循环特征,并进行回归分析,寻求一个可解释的映射的循环特征和ASMR触发。我们发现,放松效果逐步积累,是独立的空间方向。循环模式显着影响心理和身体的影响,保持不变的时间。回归分析表明,平稳传播和能量密集的循环模式最有效地触发ASMR响应。
摘要:Autonomous Sensory Meridian Response (ASMR) has been remarkably popular in the recent decade. While its effect has been validated through behavioral studies and neuro-physiological measurements such as electroencephalography (EEG) and related bio-signal analyses, its development and triggers remain a subject of debate. Previous studies suggest that its triggers are highly linked with cyclic patterns: predictable patterns introduce relaxation while variations maintain intrigue. To validate this and further understand the impact of acoustic features on ASMR effects, we designed three distinct cyclic patterns with monophonic and stereophonic variations, while controlling their predictability and randomness, and collected ASMR triggering scores through online surveys. Then, we extracted cyclic features and carried out regression analysis, seeking an explainable mapping of cyclic features and ASMR triggers. We found that relaxing effects accumulate progressively and are independent of spatial orientation. Cyclic patterns significantly influence psychological and physical effects, which remain invariant with time. Regression analysis revealed that smoothly spread and energy-dense cyclic patterns most effectively trigger ASMR responses.


【3】 $C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker  Extraction
标题: $C^2$AV-TSE:上下文和置信度感知的视听目标说话人提取
链接:https://arxiv.org/abs/2504.00750

作者: Wenxuan Wu,  Xueyuan Chen,  Shuai Wang,  Jiadong Wang,  Lingwei Meng,  Xixin Wu,  Helen Meng,  Haizhou Li 
备注:Accepted by IEEE Journal of Selected Topics in Signal Processing (JSTSP)
摘要:视听目标说话人提取(AV-TSE)的目的是模仿人类的能力,以提高听觉感知使用视觉线索。虽然最近已经提出了许多模型,但它们中的大多数主要依靠声学特征内的局部依赖性来估计目标信号,未充分利用类似人类的能力来通过上下文信息推断不清楚的语音部分。这种限制不仅导致次优的性能,而且在整个话语的提取质量不一致,与一些段表现出质量差或干扰扬声器的抑制不足。为了缩小这一差距,我们提出了一种与模型无关的策略,称为Mask-And-Recover(MAR)。它集成了模态间和模态内的上下文相关性,以实现提取模块内的全局推理。此外,为了更好地针对每个样本中具有挑战性的部分,我们引入了细粒度置信度得分(FCS)模型来评估提取质量,并指导提取模块强调对低质量片段的改进。为了验证我们提出的与模型无关的训练范式的有效性,我们采用了六种流行的AV-TSE主干对VoxCeleb 2数据集进行评估,证明了各种指标的一致性能改进。
摘要:Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals by primarily relying on local dependencies within acoustic features, underutilizing the human-like capacity to infer unclear parts of speech through contextual information. This limitation results in not only suboptimal performance but also inconsistent extraction quality across the utterance, with some segments exhibiting poor quality or inadequate suppression of interfering speakers. To close this gap, we propose a model-agnostic strategy called the Mask-And-Recover (MAR). It integrates both inter- and intra-modality contextual correlations to enable global inference within extraction modules. Additionally, to better target challenging parts within each sample, we introduce a Fine-grained Confidence Score (FCS) model to assess extraction quality and guide extraction modules to emphasize improvement on low-quality segments. To validate the effectiveness of our proposed model-agnostic training paradigm, six popular AV-TSE backbones were adopted for evaluation on the VoxCeleb2 dataset, demonstrating consistent performance improvements across various metrics.


机器翻译由腾讯交互翻译提供,仅供参考