本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
【1】 Token Pruning in Audio Transformers: Optimizing Performance and Decoding Patch Importance
标题: 音频Transformer中的令牌修剪:优化性能和解码补丁重要性
链接:https://arxiv.org/abs/2504.01690
备注:This work has been submitted to the IEEE for possible publication. Source code is available at this https URL
摘要:Vision Transformers(ViT)在各种计算机视觉任务中实现了最先进的性能,但其高计算成本仍然是一个挑战。令牌修剪已经被提出来通过选择性地移除不太重要的令牌来降低这种成本。虽然通过丢弃非对象区域在视觉任务中是有效的,但将这种技术应用于音频任务提出了独特的挑战,因为在时频表示中区分相关区域和不相关区域不那么简单。在这项研究中,我们首次使用Mel频谱图将令牌修剪应用于基于ViT的音频分类模型,并分析了模型性能和计算成本之间的权衡:TopK令牌修剪可以将AudioMAE和AST的MAC操作减少30- 40%,分类准确性下降不到1%。我们的分析表明,虽然高强度标记对模型准确性有显着贡献,但低强度标记仍然很重要。特别是,它们在一般音频分类任务中比在语音特定任务中发挥更关键的作用。
摘要:Vision Transformers (ViTs) have achieved state-of-the-art performance across various computer vision tasks, but their high computational cost remains a challenge. Token pruning has been proposed to reduce this cost by selectively removing less important tokens. While effective in vision tasks by discarding non-object regions, applying this technique to audio tasks presents unique challenges, as distinguishing relevant from irrelevant regions in time-frequency representations is less straightforward. In this study, for the first time, we applied token pruning to ViT-based audio classification models using Mel-spectrograms and analyzed the trade-offs between model performance and computational cost: TopK token pruning can reduce MAC operations of AudioMAE and AST by 30-40%, with less than a 1% drop in classification accuracy. Our analysis reveals that while high-intensity tokens contribute significantly to model accuracy, low-intensity tokens remain important. In particular, they play a more critical role in general audio classification tasks than in speech-specific tasks.
【2】 AIM: Acoustic Inertial Measurement for Indoor Drone Localization and Tracking
标题: 目的:室内无人机定位和跟踪的声惯性测量
链接:https://arxiv.org/abs/2504.01297
备注:arXiv admin note: substantial text overlap with arXiv:2504.00445
摘要:我们提出了声学惯性测量(AIM),一种用于室内无人机定位和跟踪的技术。室内无人机定位和跟踪可以说是一个至关重要但尚未解决的挑战:在GPS拒绝的环境中,现有方法的适用性有限,特别是在非视线(NLoS)中,需要大量的环境仪器,或者需要对无人机进行相当大的硬件/软件更改。相比之下,AIM利用无人机的声学特性来估计它们的位置并导出它们的运动,即使在NLoS设置中也是如此。我们驯服的位置估计误差使用一个专用的卡尔曼滤波器和四分位距规则(IQR)。我们实现AIM使用现成的麦克风阵列,并评估其性能与商业无人机在不同的设置。结果表明,AIM的平均定位误差是46%低于商业UWB系统在复杂的室内场景,其中最先进的红外系统甚至不会工作,因为NLoS设置。我们进一步证明,AIM可以扩展到支持任意范围和布局的室内空间,而不会损失的准确性,通过部署分布式麦克风阵列。
摘要:We present Acoustic Inertial Measurement (AIM), a one-of-a-kind technique for indoor drone localization and tracking. Indoor drone localization and tracking are arguably a crucial, yet unsolved challenge: in GPS-denied environments, existing approaches enjoy limited applicability, especially in Non-Line of Sight (NLoS), require extensive environment instrumentation, or demand considerable hardware/software changes on drones. In contrast, AIM exploits the acoustic characteristics of the drones to estimate their location and derive their motion, even in NLoS settings. We tame location estimation errors using a dedicated Kalman filter and the Interquartile Range rule (IQR). We implement AIM using an off-the-shelf microphone array and evaluate its performance with a commercial drone under varied settings. Results indicate that the mean localization error of AIM is 46% lower than commercial UWB-based systems in complex indoor scenarios, where state-of-the-art infrared systems would not even work because of NLoS settings. We further demonstrate that AIM can be extended to support indoor spaces with arbitrary ranges and layouts without loss of accuracy by deploying distributed microphone arrays.
【3】 Multilingual and Multi-Accent Jailbreaking of Audio LLMs
标题: 多语言和多口音音频LL越狱
链接:https://arxiv.org/abs/2504.01094
备注:21 pages, 6 figures, 15 tables
摘要:大型音频语言模型(LALM)具有显著先进的音频理解能力,但会带来严重的安全风险,特别是通过音频越狱。虽然之前的工作主要集中在以英语为中心的攻击上,但我们暴露了一个更为严重的漏洞:对抗性的多语言和多口音音频越狱,其中语言和声学变化极大地放大了攻击的成功率。在本文中,我们介绍了Multi-AudioJail,这是第一个利用这些漏洞的系统框架,通过(1)一个新的受干扰的多语言/多口音音频越狱提示数据集,以及(2)一个分层评估管道,揭示了声学扰动(例如,混响、回声和耳语效果)与跨语言语音相互作用以使越狱成功率(JSR)激增高达+57.25个百分点(例如,对MERaLiON的肯尼亚口音攻击)。至关重要的是,我们的工作进一步揭示了多模态LLM本质上比单峰系统更脆弱:攻击者只需要利用最薄弱的环节(例如,非英语音频输入)来破坏整个模型,我们通过经验证明,多语言纯音频攻击的成功率比纯文本攻击高3.1倍。我们计划发布我们的数据集,以刺激对跨模式防御的研究,敦促社区随着LAM的发展以多模式解决这种不断扩大的攻击面。
摘要:Large Audio Language Models (LALMs) have significantly advanced audio understanding but introduce critical security risks, particularly through audio jailbreaks. While prior work has focused on English-centric attacks, we expose a far more severe vulnerability: adversarial multilingual and multi-accent audio jailbreaks, where linguistic and acoustic variations dramatically amplify attack success. In this paper, we introduce Multi-AudioJail, the first systematic framework to exploit these vulnerabilities through (1) a novel dataset of adversarially perturbed multilingual/multi-accent audio jailbreaking prompts, and (2) a hierarchical evaluation pipeline revealing that how acoustic perturbations (e.g., reverberation, echo, and whisper effects) interacts with cross-lingual phonetics to cause jailbreak success rates (JSRs) to surge by up to +57.25 percentage points (e.g., reverberated Kenyan-accented attack on MERaLiON). Crucially, our work further reveals that multimodal LLMs are inherently more vulnerable than unimodal systems: attackers need only exploit the weakest link (e.g., non-English audio inputs) to compromise the entire model, which we empirically show by multilingual audio-only attacks achieving 3.1x higher success rates than text-only attacks. We plan to release our dataset to spur research into cross-modal defenses, urging the community to address this expanding attack surface in multimodality as LALMs evolve.
【1】 Leveraging Embedding Techniques in Multimodal Machine Learning for Mental Illness Assessment
标题: 利用多模式机器学习中的嵌入技术进行精神疾病评估
链接:https://arxiv.org/abs/2504.01767
摘要:抑郁症和创伤后应激障碍等精神障碍在全球的流行率日益增加,需要客观和可扩展的诊断工具。传统的临床评估往往面临可及性、客观性和一致性方面的限制。本文研究了多模态机器学习解决这些挑战的潜力,利用文本,音频和视频数据中的互补信息。我们的方法涉及到各种数据预处理技术,包括新颖的组块和基于话语的格式化策略的全面分析。我们系统地评估了每种模态的一系列最先进的嵌入模型,并采用卷积神经网络(CNN)和双向LSTM网络(BiLSTM)进行特征提取。我们探索数据级,特征级和决策级融合技术,包括一个新的集成大语言模型(LLM)的预测。我们还研究了用支持向量机取代多层感知器分类器的影响。我们将分析扩展到使用PHQ-8和PCL-C评分和多类分类(考虑合并症)进行严重程度预测。我们的研究结果表明,基于话语的组块显着提高性能,特别是对于文本和音频模态。结合LLM预测的决策级融合达到了最高的准确度,抑郁症的平衡准确度为94.8%,PTSD检测为96.2%。CNN-BiLSTM架构与话语级组块的结合,再加上外部LLM的集成,为心理健康状况的检测和评估提供了一种强大而细致的方法。我们的研究结果强调了MMML在开发更准确、更容易获得和更个性化的心理健康工具方面的潜力。
摘要:The increasing global prevalence of mental disorders, such as depression and PTSD, requires objective and scalable diagnostic tools. Traditional clinical assessments often face limitations in accessibility, objectivity, and consistency. This paper investigates the potential of multimodal machine learning to address these challenges, leveraging the complementary information available in text, audio, and video data. Our approach involves a comprehensive analysis of various data preprocessing techniques, including novel chunking and utterance-based formatting strategies. We systematically evaluate a range of state-of-the-art embedding models for each modality and employ Convolutional Neural Networks (CNNs) and Bidirectional LSTM Networks (BiLSTMs) for feature extraction. We explore data-level, feature-level, and decision-level fusion techniques, including a novel integration of Large Language Model (LLM) predictions. We also investigate the impact of replacing Multilayer Perceptron classifiers with Support Vector Machines. We extend our analysis to severity prediction using PHQ-8 and PCL-C scores and multi-class classification (considering co-occurring conditions). Our results demonstrate that utterance-based chunking significantly improves performance, particularly for text and audio modalities. Decision-level fusion, incorporating LLM predictions, achieves the highest accuracy, with a balanced accuracy of 94.8% for depression and 96.2% for PTSD detection. The combination of CNN-BiLSTM architectures with utterance-level chunking, coupled with the integration of external LLM, provides a powerful and nuanced approach to the detection and assessment of mental health conditions. Our findings highlight the potential of MMML for developing more accurate, accessible, and personalized mental healthcare tools.
【2】 Spatial-Filter-Bank-Based Neural Method for Multichannel Speech Enhancement
标题: 基于空间过滤器库的多通道语音增强神经方法
链接:https://arxiv.org/abs/2504.01392
摘要:基于深度学习的多通道语音增强方法在麦克风阵列几何参数发生变化时,其性能往往会恶化。缓解此问题的传统方法通常涉及在多个麦克风阵列上进行训练,这可能是昂贵的。为了解决这一挑战,我们专注于均匀的圆形阵列,并提出使用空间滤波器组来提取几何参数近似不变的特征。然后,这些功能处理的两个阶段的一致性为基础的模型(TSCBM),以提高语音质量。实验结果表明,我们提出的方法可以在固定的麦克风阵列上进行训练,同时在应用过程中保持均匀圆形阵列的有效性能。
摘要:The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on multiple microphone arrays, which can be costly. To address this challenge, we focus on uniform circular arrays and propose the use of a spatial filter bank to extract features that are approximately invariant to geometric parameters. These features are then processed by a two-stage conformer-based model (TSCBM) to enhance speech quality. Experimental results demonstrate that our proposed method can be trained on a fixed microphone array while maintaining effective performance across uniform circular arrays with unseen geometric configurations during applications.
【3】 Chain of Correction for Full-text Speech Recognition with Large Language Models
标题: 大语言模型全文语音识别的纠正链
链接:https://arxiv.org/abs/2504.01519
摘要:使用大型语言模型(LLM)进行自动语音识别(ASR)的全文纠错已经受到越来越多的关注,因为它有可能在长上下文中纠正错误,并解决更广泛的错误类型,包括标点符号恢复和反向文本规范化。然而,许多挑战仍然存在,包括与稳定性、可控性、完整性和流畅性相关的问题。为了缓解这些挑战,本文提出了使用LLM进行全文纠错的纠错链(CoC),该纠错链使用预先识别的文本作为指导,在常规的多轮聊天格式中逐段纠正错误。CoC还使用预先识别的全文作为上下文,使模型能够更好地掌握全局语义并保持对整个内容的全面概述。利用开源的全文纠错数据集ChFT,我们对预训练的LLM进行了微调,以评估CoC框架的性能。实验结果表明,CoC有效地纠正了全文ASR输出中的错误,显著优于基线和基准系统。我们进一步分析了如何设置校正阈值来平衡校正不足和过度改写,在极长的ASR输出上外推CoC模型,并研究是否可以采用其他类型的信息来指导纠错过程。
摘要:Full-text error correction with Large Language Models (LLMs) for Automatic Speech Recognition (ASR) has gained increased attention due to its potential to correct errors across long contexts and address a broader spectrum of error types, including punctuation restoration and inverse text normalization. Nevertheless, many challenges persist, including issues related to stability, controllability, completeness, and fluency. To mitigate these challenges, this paper proposes the Chain of Correction (CoC) for full-text error correction with LLMs, which corrects errors segment by segment using pre-recognized text as guidance within a regular multi-turn chat format. The CoC also uses pre-recognized full text for context, allowing the model to better grasp global semantics and maintain a comprehensive overview of the entire content. Utilizing the open-sourced full-text error correction dataset ChFT, we fine-tune a pre-trained LLM to evaluate the performance of the CoC framework. Experimental results demonstrate that the CoC effectively corrects errors in full-text ASR outputs, significantly outperforming baseline and benchmark systems. We further analyze how to set the correction threshold to balance under-correction and over-rephrasing, extrapolate the CoC model on extremely long ASR outputs, and investigate whether other types of information can be employed to guide the error correction process.
【4】 AIM: Acoustic Inertial Measurement for Indoor Drone Localization and Tracking
标题: 目的:室内无人机定位和跟踪的声惯性测量
链接:https://arxiv.org/abs/2504.01297
备注:arXiv admin note: substantial text overlap with arXiv:2504.00445
摘要:我们提出了声学惯性测量(AIM),一种用于室内无人机定位和跟踪的技术。室内无人机定位和跟踪可以说是一个至关重要但尚未解决的挑战:在GPS拒绝的环境中,现有方法的适用性有限,特别是在非视线(NLoS)中,需要大量的环境仪器,或者需要对无人机进行相当大的硬件/软件更改。相比之下,AIM利用无人机的声学特性来估计它们的位置并导出它们的运动,即使在NLoS设置中也是如此。我们驯服的位置估计误差使用一个专用的卡尔曼滤波器和四分位距规则(IQR)。我们实现AIM使用现成的麦克风阵列,并评估其性能与商业无人机在不同的设置。结果表明,AIM的平均定位误差是46%低于商业UWB系统在复杂的室内场景,其中最先进的红外系统甚至不会工作,因为NLoS设置。我们进一步证明,AIM可以扩展到支持任意范围和布局的室内空间,而不会损失的准确性,通过部署分布式麦克风阵列。
摘要:We present Acoustic Inertial Measurement (AIM), a one-of-a-kind technique for indoor drone localization and tracking. Indoor drone localization and tracking are arguably a crucial, yet unsolved challenge: in GPS-denied environments, existing approaches enjoy limited applicability, especially in Non-Line of Sight (NLoS), require extensive environment instrumentation, or demand considerable hardware/software changes on drones. In contrast, AIM exploits the acoustic characteristics of the drones to estimate their location and derive their motion, even in NLoS settings. We tame location estimation errors using a dedicated Kalman filter and the Interquartile Range rule (IQR). We implement AIM using an off-the-shelf microphone array and evaluate its performance with a commercial drone under varied settings. Results indicate that the mean localization error of AIM is 46% lower than commercial UWB-based systems in complex indoor scenarios, where state-of-the-art infrared systems would not even work because of NLoS settings. We further demonstrate that AIM can be extended to support indoor spaces with arbitrary ranges and layouts without loss of accuracy by deploying distributed microphone arrays.
