本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
链接:https://arxiv.org/abs/2502.21097
摘要:在本文中,我们提出了一种深度学习方法,用于从表示为互谱矩阵的麦克风阵列数据中过滤掉环境噪声、反射或源方向性等影响。具体来说,我们专注于生成对抗网络(GAN)架构,旨在转换固定大小的交叉谱矩阵。这些模型使用为此目的开发的不同复杂性的声压模拟进行训练。基于在自动编码任务的超参数优化中应用这些方法的结果,我们训练优化的模型来执行五个不同的转换任务,这些任务来自我们声压模拟中固有的不同复杂性。
摘要:In this paper, we present a deep-learning method to filter out effects suchas ambient noise, reflections, or source directivity from microphone array datarepresented as cross-spectral matrices. Specifically, we focus on a generativeadversarial network (GAN) architecture designed to transform fixed-sizecross-spectral matrices. Theses models were trained using sound pressuresimulations of varying complexity developed for this purpose. Based on theresults from applying these methods in a hyperparameter optimization of anauto-encoding task, we trained the optimized model to perform five distincttransformation tasks derived from different complexities inherent in our soundpressure simulations.
标题:长时间无源声学监测中用于鲸鱼叫声检测和定位的弱监督多实例学习
链接:https://arxiv.org/abs/2502.20838
摘要:通过被动声学监测(PAM)进行海洋生态系统监测会产生大量数据,但深度学习通常需要精确的注释和简短的片段。我们介绍DSMIL-LocNet,一个多实例学习框架,用于鲸鱼呼叫检测和本地化,仅使用袋级标签。我们的双流模型处理2-30分钟的音频片段,利用频谱和时间特征以及基于注意力的实例选择。对南极鲸数据的测试表明,较长的上下文可以提高分类精度(F1:0.8-0.9),而中等的上下文可以确保定位精度(0.65-0.70)。这表明MIL可以增强可扩展的海洋监测。代码:https://github.com/Ragib-Amin-Nihal/DSMIL-Loc
摘要:Marine ecosystem monitoring via Passive Acoustic Monitoring (PAM) generatesvast data, but deep learning often requires precise annotations and shortsegments. We introduce DSMIL-LocNet, a Multiple Instance Learning framework forwhale call detection and localization using only bag-level labels. Ourdual-stream model processes 2-30 minute audio segments, leveraging spectral andtemporal features with attention-based instance selection. Tests on Antarcticwhale data show longer contexts improve classification (F1: 0.8-0.9) whilemedium instances ensure localization precision (0.65-0.70). This suggests MILcan enhance scalable marine monitoring. Code:https://github.com/Ragib-Amin-Nihal/DSMIL-Loc
标题:LiteASB:具有低等级逼近的高效自动语音识别
链接:https://arxiv.org/abs/2502.20583
摘要:现代自动语音识别(ASR)模型,如OpenAI的Whisper,依赖于深度编码器-解码器架构,由于计算强度高,其编码器是高效部署的关键瓶颈。我们介绍了LiteASR,这是一种用于ASR编码器的低秩压缩方案,可显着降低推理成本,同时保持转录准确性。我们的方法利用了在中间激活中观察到的强大的低秩属性:通过对小的校准数据集应用主成分分析(PCA),我们用一系列低秩矩阵乘法近似线性变换,并进一步优化自我注意力以在降低的维度上工作。评估结果表明,我们的方法可以将Whisper large-v3的编码器大小压缩50%以上,匹配Whisper medium的大小,具有更好的转录精度,从而建立一个新的Pareto最优的效率和性能边界。LiteASR的代码可在https://github.com/efeslab/LiteASR上获得。
摘要:Modern automatic speech recognition (ASR) models, such as OpenAI's Whisper,rely on deep encoder-decoder architectures, and their encoders are a criticalbottleneck for efficient deployment due to high computational intensity. Weintroduce LiteASR, a low-rank compression scheme for ASR encoders thatsignificantly reduces inference costs while maintaining transcription accuracy.Our approach leverages the strong low-rank properties observed in intermediateactivations: by applying principal component analysis (PCA) with a smallcalibration dataset, we approximate linear transformations with a chain oflow-rank matrix multiplications, and further optimize self-attention to work inthe reduced dimension. Evaluation results show that our method can compressWhisper large-v3's encoder size by over 50%, matching Whisper medium's sizewith better transcription accuracy, thereby establishing a new Pareto-optimalfrontier of efficiency and performance. The code of LiteASR is available athttps://github.com/efeslab/LiteASR.
标题:Deepen:音频Deepfake检测的渗透测试
链接:https://arxiv.org/abs/2502.20427
摘要:Deepfakes -操纵或伪造的音频和视频媒体-对个人,组织和整个社会构成重大安全风险。为了解决这些挑战,通常采用基于机器学习的分类器来检测deepfake内容。在本文中,我们通过系统的渗透测试方法来评估这些分类器的鲁棒性,我们将其称为Deepen。我们的方法在没有先验知识或访问目标deepfake检测模型的情况下运行。相反,它利用一组精心选择的信号处理修改(称为攻击)来评估模型漏洞。使用Deepen,我们分析了现实世界的生产系统和公开的学术模型检查点,证明了所有测试系统都存在弱点,并且可以通过简单的操作(如时间拉伸或回声添加)可靠地欺骗。此外,我们的研究结果表明,虽然一些攻击可以通过重新训练检测系统来减轻特定攻击的知识,但其他攻击仍然有效。我们将发布所有相关代码。
摘要:Deepfakes - manipulated or forged audio and video media - pose significantsecurity risks to individuals, organizations, and society at large. To addressthese challenges, machine learning-based classifiers are commonly employed todetect deepfake content. In this paper, we assess the robustness of suchclassifiers through a systematic penetration testing methodology, which weintroduce as DeePen. Our approach operates without prior knowledge of or accessto the target deepfake detection models. Instead, it leverages a set ofcarefully selected signal processing modifications - referred to as attacks -to evaluate model vulnerabilities. Using DeePen, we analyze both real-worldproduction systems and publicly available academic model checkpoints,demonstrating that all tested systems exhibit weaknesses and can be reliablydeceived by simple manipulations such as time-stretching or echo addition.Furthermore, our findings reveal that while some attacks can be mitigated byretraining detection systems with knowledge of the specific attack, othersremain persistently effective. We release all associated code.
标题:JiTTER:用于自监督声音事件检测的事件重建的Jigsaw时间Transformer
链接:https://arxiv.org/abs/2502.20857
摘要:声音事件检测(SED)已经显著地受益于自监督学习(SSL)方法,特别是用于SED的掩蔽音频Transformer(MAT-SED),其利用掩蔽块预测来重构丢失的音频段。然而,虽然在捕获全局依赖性方面是有效的,但掩蔽块预测破坏了瞬态声音事件并且缺乏对时间顺序的明确执行,使得其不太适合于细粒度事件边界检测。为了解决这些局限性,我们提出了JiTTER(拼图时间Transformer事件重建),SSL框架,旨在提高基于transformer的SED的时间建模。JiTTER引入了一种分层时间混洗重建策略,其中音频序列在块级和帧级都被随机混洗,迫使模型重建正确的时间顺序。这个预训练目标鼓励模型学习全局事件结构和细粒度的瞬态细节,提高其检测具有尖锐的发作-偏移特征的事件的能力。此外,我们在块洗牌过程中加入了噪声注入,提供了一种微妙的扰动机制,进一步规范了特征学习并增强了模型的鲁棒性。在DESED数据集上的实验结果表明,JiTTER优于MAT-SED,在PSDS中实现了5.89%的改进,突出了显式时态推理在基于SSL的SED中的有效性。我们的研究结果表明,结构化的时间重建任务,而不是简单的掩蔽预测,提供了一个更有效的预训练范式的声音事件表征学习。
摘要:Sound event detection (SED) has significantly benefited from self-supervisedlearning (SSL) approaches, particularly masked audio transformer for SED(MAT-SED), which leverages masked block prediction to reconstruct missing audiosegments. However, while effective in capturing global dependencies, maskedblock prediction disrupts transient sound events and lacks explicit enforcementof temporal order, making it less suitable for fine-grained event boundarydetection. To address these limitations, we propose JiTTER (Jigsaw TemporalTransformer for Event Reconstruction), an SSL framework designed to enhancetemporal modeling in transformer-based SED. JiTTER introduces a hierarchicaltemporal shuffle reconstruction strategy, where audio sequences are randomlyshuffled at both the block-level and frame-level, forcing the model toreconstruct the correct temporal order. This pretraining objective encouragesthe model to learn both global event structures and fine-grained transientdetails, improving its ability to detect events with sharp onset-offsetcharacteristics. Additionally, we incorporate noise injection during blockshuffle, providing a subtle perturbation mechanism that further regularizesfeature learning and enhances model robustness. Experimental results on theDESED dataset demonstrate that JiTTER outperforms MAT-SED, achieving a 5.89%improvement in PSDS, highlighting the effectiveness of explicit temporalreasoning in SSL-based SED. Our findings suggest that structured temporalreconstruction tasks, rather than simple masked prediction, offer a moreeffective pretraining paradigm for sound event representation learning.
标题:JiTTER:用于自监督声音事件检测的事件重建的Jigsaw时间Transformer
链接:https://arxiv.org/abs/2502.20857
摘要:声音事件检测(SED)已经显著地受益于自监督学习(SSL)方法,特别是用于SED的掩蔽音频Transformer(MAT-SED),其利用掩蔽块预测来重构丢失的音频段。然而,虽然在捕获全局依赖性方面是有效的,但掩蔽块预测破坏了瞬态声音事件并且缺乏对时间顺序的明确执行,使得其不太适合于细粒度事件边界检测。为了解决这些局限性,我们提出了JiTTER(拼图时间Transformer事件重建),SSL框架,旨在提高基于transformer的SED的时间建模。JiTTER引入了一种分层时间混洗重建策略,其中音频序列在块级和帧级都被随机混洗,迫使模型重建正确的时间顺序。该预训练目标鼓励模型学习全局事件结构和细粒度的瞬时细节,从而提高其检测具有尖锐的发作-偏移特征的事件的能力。此外,我们在块洗牌过程中加入了噪声注入,提供了一种微妙的扰动机制,进一步规范了特征学习并增强了模型的鲁棒性。在DESED数据集上的实验结果表明,JiTTER优于MAT-SED,在PSDS中实现了5.89%的改进,突出了显式时态推理在基于SSL的SED中的有效性。我们的研究结果表明,结构化的时间重建任务,而不是简单的掩蔽预测,提供了一个更有效的预训练范式的声音事件表征学习。
摘要:Sound event detection (SED) has significantly benefited from self-supervisedlearning (SSL) approaches, particularly masked audio transformer for SED(MAT-SED), which leverages masked block prediction to reconstruct missing audiosegments. However, while effective in capturing global dependencies, maskedblock prediction disrupts transient sound events and lacks explicit enforcementof temporal order, making it less suitable for fine-grained event boundarydetection. To address these limitations, we propose JiTTER (Jigsaw TemporalTransformer for Event Reconstruction), an SSL framework designed to enhancetemporal modeling in transformer-based SED. JiTTER introduces a hierarchicaltemporal shuffle reconstruction strategy, where audio sequences are randomlyshuffled at both the block-level and frame-level, forcing the model toreconstruct the correct temporal order. This pretraining objective encouragesthe model to learn both global event structures and fine-grained transientdetails, improving its ability to detect events with sharp onset-offsetcharacteristics. Additionally, we incorporate noise injection during blockshuffle, providing a subtle perturbation mechanism that further regularizesfeature learning and enhances model robustness. Experimental results on theDESED dataset demonstrate that JiTTER outperforms MAT-SED, achieving a 5.89%improvement in PSDS, highlighting the effectiveness of explicit temporalreasoning in SSL-based SED. Our findings suggest that structured temporalreconstruction tasks, rather than simple masked prediction, offer a moreeffective pretraining paradigm for sound event representation learning.
标题:使用生成式对抗网络对互谱矩阵进行基于深度学习的过滤
链接:https://arxiv.org/abs/2502.21097
摘要:在本文中,我们提出了一种深度学习方法,用于从表示为互谱矩阵的麦克风阵列数据中过滤掉环境噪声、反射或源方向性等影响。具体来说,我们专注于生成对抗网络(GAN)架构,旨在转换固定大小的交叉谱矩阵。这些模型使用为此目的开发的不同复杂性的声压模拟进行训练。基于在自动编码任务的超参数优化中应用这些方法的结果,我们训练优化的模型来执行五个不同的转换任务,这些任务来自我们声压模拟中固有的不同复杂性。
摘要:In this paper, we present a deep-learning method to filter out effects suchas ambient noise, reflections, or source directivity from microphone array datarepresented as cross-spectral matrices. Specifically, we focus on a generativeadversarial network (GAN) architecture designed to transform fixed-sizecross-spectral matrices. Theses models were trained using sound pressuresimulations of varying complexity developed for this purpose. Based on theresults from applying these methods in a hyperparameter optimization of anauto-encoding task, we trained the optimized model to perform five distincttransformation tasks derived from different complexities inherent in our soundpressure simulations.
标题:长时间无源声学监测中用于鲸鱼叫声检测和定位的弱监督多实例学习
链接:https://arxiv.org/abs/2502.20838
摘要:通过被动声学监测(PAM)进行海洋生态系统监测会产生大量数据,但深度学习通常需要精确的注释和简短的片段。我们介绍DSMIL-LocNet,一个多实例学习框架,用于鲸鱼呼叫检测和本地化,仅使用袋级标签。我们的双流模型处理2-30分钟的音频片段,利用频谱和时间特征以及基于注意力的实例选择。对南极鲸数据的测试表明,较长的上下文可以提高分类精度(F1:0.8-0.9),而中等的上下文可以确保定位精度(0.65-0.70)。这表明MIL可以增强可扩展的海洋监测。代码:https://github.com/Ragib-Amin-Nihal/DSMIL-Loc
摘要:Marine ecosystem monitoring via Passive Acoustic Monitoring (PAM) generatesvast data, but deep learning often requires precise annotations and shortsegments. We introduce DSMIL-LocNet, a Multiple Instance Learning framework forwhale call detection and localization using only bag-level labels. Ourdual-stream model processes 2-30 minute audio segments, leveraging spectral andtemporal features with attention-based instance selection. Tests on Antarcticwhale data show longer contexts improve classification (F1: 0.8-0.9) whilemedium instances ensure localization precision (0.65-0.70). This suggests MILcan enhance scalable marine monitoring. Code:https://github.com/Ragib-Amin-Nihal/DSMIL-Loc
标题:LiteASB:具有低等级逼近的高效自动语音识别
链接:https://arxiv.org/abs/2502.20583
摘要:现代自动语音识别(ASR)模型,如OpenAI的Whisper,依赖于深度编码器-解码器架构,由于计算强度高,其编码器是高效部署的关键瓶颈。我们介绍了LiteASR,这是一种用于ASR编码器的低秩压缩方案,可显着降低推理成本,同时保持转录准确性。我们的方法利用了在中间激活中观察到的强大的低秩属性:通过对小的校准数据集应用主成分分析(PCA),我们用一系列低秩矩阵乘法近似线性变换,并进一步优化自我注意力以在降低的维度上工作。评估结果表明,我们的方法可以将Whisper large-v3的编码器大小压缩50%以上,匹配Whisper medium的大小,具有更好的转录精度,从而建立一个新的Pareto最优的效率和性能边界。LiteASR的代码可在https://github.com/efeslab/LiteASR上获得。
摘要:Modern automatic speech recognition (ASR) models, such as OpenAI's Whisper,rely on deep encoder-decoder architectures, and their encoders are a criticalbottleneck for efficient deployment due to high computational intensity. Weintroduce LiteASR, a low-rank compression scheme for ASR encoders thatsignificantly reduces inference costs while maintaining transcription accuracy.Our approach leverages the strong low-rank properties observed in intermediateactivations: by applying principal component analysis (PCA) with a smallcalibration dataset, we approximate linear transformations with a chain oflow-rank matrix multiplications, and further optimize self-attention to work inthe reduced dimension. Evaluation results show that our method can compressWhisper large-v3's encoder size by over 50%, matching Whisper medium's sizewith better transcription accuracy, thereby establishing a new Pareto-optimalfrontier of efficiency and performance. The code of LiteASR is available athttps://github.com/efeslab/LiteASR.
标题:Deepen:音频Deepfake检测的渗透测试
链接:https://arxiv.org/abs/2502.20427
摘要:Deepfakes -操纵或伪造的音频和视频媒体-对个人,组织和整个社会构成重大安全风险。为了解决这些挑战,通常采用基于机器学习的分类器来检测deepfake内容。在本文中,我们通过系统的渗透测试方法来评估这些分类器的鲁棒性,我们将其称为Deepen。我们的方法在没有先验知识或访问目标deepfake检测模型的情况下运行。相反,它利用一组精心选择的信号处理修改(称为攻击)来评估模型漏洞。使用Deepen,我们分析了现实世界的生产系统和公开的学术模型检查点,证明了所有测试系统都存在弱点,并且可以通过简单的操作(如时间拉伸或回声添加)可靠地欺骗。此外,我们的研究结果表明,虽然一些攻击可以通过重新训练检测系统来减轻特定攻击的知识,但其他攻击仍然有效。我们将发布所有相关代码。
摘要:Deepfakes - manipulated or forged audio and video media - pose significantsecurity risks to individuals, organizations, and society at large. To addressthese challenges, machine learning-based classifiers are commonly employed todetect deepfake content. In this paper, we assess the robustness of suchclassifiers through a systematic penetration testing methodology, which weintroduce as DeePen. Our approach operates without prior knowledge of or accessto the target deepfake detection models. Instead, it leverages a set ofcarefully selected signal processing modifications - referred to as attacks -to evaluate model vulnerabilities. Using DeePen, we analyze both real-worldproduction systems and publicly available academic model checkpoints,demonstrating that all tested systems exhibit weaknesses and can be reliablydeceived by simple manipulations such as time-stretching or echo addition.Furthermore, our findings reveal that while some attacks can be mitigated byretraining detection systems with knowledge of the specific attack, othersremain persistently effective. We release all associated code.
