今日论文合集:cs.SD语音5篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】 Audio-Based Classification of Insect Species Using Machine Learning  Models: Cicada, Beetle, Termite, and Cricket

标题:使用机器学习模型对昆虫物种进行基于音频的分类:蝉、甲虫、白蚁和蟋蟀
链接:https://arxiv.org/abs/2502.13893
作者:Manas V Shetty,  Yoga Disha Sendhil Kumar
摘要:这个项目解决了昆虫物种分类的挑战:蝉,甲虫,白蚁和蟋蟀使用录音。准确的物种识别对于生态监测和害虫管理至关重要。我们采用XGBoost、随机森林和K最近邻(KNN)等机器学习模型来分析音频特征,包括Mel频率倒谱系数(MFCC)。这项工作的潜在新颖性在于将不同的音频特征和机器学习模型相结合,以解决昆虫分类问题,特别是专注于捕捉物种之间的微妙声学变化,这些变化在以前的研究中尚未得到充分利用。该数据集是从各种开放源编译的,我们预计将实现高分类精度,有助于改进自动昆虫检测系统。
摘要:This project addresses the challenge of classifying insect species: Cicada,Beetle, Termite, and Cricket using sound recordings. Accurate speciesidentification is crucial for ecological monitoring and pest management. Weemploy machine learning models such as XGBoost, Random Forest, and K NearestNeighbors (KNN) to analyze audio features, including Mel Frequency CepstralCoefficients (MFCC). The potential novelty of this work lies in the combinationof diverse audio features and machine learning models to tackle insectclassification, specifically focusing on capturing subtle acoustic variationsbetween species that have not been fully leveraged in previous research. Thedataset is compiled from various open sources, and we anticipate achieving highclassification accuracy, contributing to improved automated insect detectionsystems.

【2】 TALKPLAY: Multimodal Music Recommendation with Large Language Models
标题:TALKSYS:具有大型语言模型的多模式音乐推荐
链接:https://arxiv.org/abs/2502.13713
作者:Seungheon Doh,  Keunwoo Choi,  Juhan Nam
摘要:我们提出了TalkPlay,一个多模态音乐推荐系统,重新制定的推荐任务,大语言模型令牌生成。TalkPlay通过扩展的令牌词汇表来表示音乐,该词汇表对多种形式进行编码-音频,歌词,元数据,语义标签和播放列表共现。使用这些丰富的表示,该模型通过对音乐推荐对话的下一个令牌预测来学习生成推荐,这需要学习自然语言查询和响应以及音乐项的关联。换句话说,该公式将音乐推荐转换为自然语言理解任务,其中模型预测会话令牌的能力直接优化了查询项相关性。我们的方法消除了传统的解释对话管道的复杂性,实现了查询感知音乐推荐的端到端学习。在实验中,TalkPlay被成功训练,并在各个方面优于基线方法,展示了作为会话式音乐推荐器的强大上下文理解。
摘要:We present TalkPlay, a multimodal music recommendation system thatreformulates the recommendation task as large language model token generation.TalkPlay represents music through an expanded token vocabulary that encodesmultiple modalities - audio, lyrics, metadata, semantic tags, and playlistco-occurrence. Using these rich representations, the model learns to generaterecommendations through next-token prediction on music recommendationconversations, that requires learning the associations natural language queryand response, as well as music items. In other words, the formulationtransforms music recommendation into a natural language understanding task,where the model's ability to predict conversation tokens directly optimizesquery-item relevance. Our approach eliminates traditionalrecommendation-dialogue pipeline complexity, enabling end-to-end learning ofquery-aware music recommendations. In the experiment, TalkPlay is successfullytrained and outperforms baseline methods in various aspects, demonstratingstrong context understanding as a conversational music recommender.

【3】 Semi-supervised classification of bird vocalizations
标题:鸟类发声的半监督分类
链接:https://arxiv.org/abs/2502.13440
作者:Simen Hexeberg,  Mandar Chitre,  Matthias Hoffmann-Kuhnt,  Bing Wen Low
摘要:鸟类数量的变化可以表明生态系统的更广泛变化,使鸟类成为最重要的动物群体之一。将机器学习和被动声学相结合,可以在没有人类直接参与的情况下进行长时间的连续监测。然而,大多数现有的技术需要大量的专家标记的数据集进行训练,并且不能容易地检测繁忙音景中的时间重叠呼叫。我们提出了一个半监督的声学鸟检测器,旨在允许检测时间重叠的调用(当分离的频率)和使用几个标记的训练样本。该分类器是在社区记录的开源数据和来自新加坡的长时间音景录音的组合上进行训练和评估的。它实现了一个平均F0.5得分0.701跨越315类110种鸟类的保持测试集,平均11个标记的训练样本,每个类。它在103种鸟类的测试集上的表现优于最先进的BirdNET分类器,尽管标记的训练样本明显较少。该探测器进一步测试了144个麦克风小时的连续声景数据。新加坡丰富的音景使得抑制误报对原始连续数据流构成挑战。尽管如此,我们证明了在这种环境中用最少的标记训练数据实现高精度是可能的。
摘要:Changes in bird populations can indicate broader changes in ecosystems,making birds one of the most important animal groups to monitor. Combiningmachine learning and passive acoustics enables continuous monitoring overextended periods without direct human involvement. However, most existingtechniques require extensive expert-labeled datasets for training and cannoteasily detect time-overlapping calls in busy soundscapes. We propose asemi-supervised acoustic bird detector designed to allow both the detection oftime-overlapping calls (when separated in frequency) and the use of few labeledtraining samples. The classifier is trained and evaluated on a combination ofcommunity-recorded open-source data and long-duration soundscape recordingsfrom Singapore. It achieves a mean F0.5 score of 0.701 across 315 classes from110 bird species on a hold-out test set, with an average of 11 labeled trainingsamples per class. It outperforms the state-of-the-art BirdNET classifier on atest set of 103 bird species despite significantly fewer labeled trainingsamples. The detector is further tested on 144 microphone-hours of continuoussoundscape data. The rich soundscape in Singapore makes suppression of falsepositives a challenge on raw, continuous data streams. Nevertheless, wedemonstrate that achieving high precision in such environments with minimallabeled training data is possible.

【4】 MATS: An Audio Language Model under Text-only Supervision
标题:MATS:纯文本监督下的音频语言模型
链接:https://arxiv.org/abs/2502.13433
作者:Wen Wang,  Ruibing Hou,  Hong Chang,  Shiguang Shan,  Xilin Chen
备注:19 pages,11 figures
摘要:建立在强大的大型语言模型(LLM)基础上的大型音频语言模型(LALM)表现出了卓越的音频理解和推理能力。然而,LALM的训练需要大量的音频-语言对语料库,这需要大量的数据收集和训练资源。在本文中,我们提出了MATS,音频语言多模态LLM设计用于处理多个音频任务,仅使用文本监督。通过利用CLAP等预训练的音频语言对齐模型,我们开发了一种纯文本训练策略,将共享的音频语言潜在空间投影到LLM潜在空间中,赋予LLM音频理解能力,而无需在训练期间依赖音频数据。为了进一步弥合CLAP中音频和语言嵌入之间的模态差距,我们提出了带音频的强相关噪声文本(Santa)机制。Santa将音频嵌入映射到CLAP语言嵌入空间,同时保留来自音频输入的基本信息。大量的实验表明,MATS,尽管专门训练的文本数据,实现竞争力的性能相比,最近的LALM训练的大规模的音频语言对。
摘要:Large audio-language models (LALMs), built upon powerful Large LanguageModels (LLMs), have exhibited remarkable audio comprehension and reasoningcapabilities. However, the training of LALMs demands a large corpus ofaudio-language pairs, which requires substantial costs in both data collectionand training resources. In this paper, we propose MATS, an audio-languagemultimodal LLM designed to handle Multiple Audio task using solely Text-onlySupervision. By leveraging pre-trained audio-language alignment models such asCLAP, we develop a text-only training strategy that projects the sharedaudio-language latent space into LLM latent space, endowing the LLM with audiocomprehension capabilities without relying on audio data during training. Tofurther bridge the modality gap between audio and language embeddings withinCLAP, we propose the Strongly-related noisy text with audio (Santa) mechanism.Santa maps audio embeddings into CLAP language embedding space while preservingessential information from the audio input. Extensive experiments demonstratethat MATS, despite being trained exclusively on text data, achieves competitiveperformance compared to recent LALMs trained on large-scale audio-languagepairs.

【5】 Unsupervised CP-UNet Framework for Denoising DAS Data with Decay Noise
标题:无监督CP-UNet框架用于消除含衰变噪音的DAS数据
链接:https://arxiv.org/abs/2502.13395
作者:Tianye Huang,  Aopeng Li,  Xiang Li,  Jing Zhang,  Sijing Xian,  Qi Zhang,  Mingkong Lu,  Guodong Chen,  Liangming Xiong,  Xiangyun Hu
备注:13 pages, 8 figures
摘要:分布式声学传感器(DAS)技术利用光纤电缆来检测声学信号,提供经济高效的密集监测功能。它提供了几个优点,包括耐极端条件,抗电磁干扰和准确的检测。然而,DAS与地震检波器相比通常表现出较低的信噪比(S/N),并且易受各种噪声类型的影响,例如随机噪声、不稳定噪声、电平噪声和长周期噪声。这种降低的S/N可能会对包含反演和解释的数据分析产生负面影响。虽然人工智能已经展示了出色的去噪能力,但大多数现有方法都依赖于对标记数据的监督学习,这对标签的质量提出了严格的要求。为了解决这个问题,我们开发了一种基于上下文金字塔UNet(CP-UNet)的无标签无监督学习(UL)网络模型,以抑制DAS数据中的不稳定和随机噪声。CP-UNet在编码和解码过程中利用上下文金字塔模块来提取特征并重建DAS数据。为了增强浅层和深层特征之间的连接,我们在编码和解码部分都添加了连接模块(CM)。层归一化(LN)用于取代常用的批量归一化(BN),加速模型的收敛并防止训练过程中的梯度爆炸。Huber损失被采用作为我们的损失函数,其参数是实验确定的。我们应用网络的2-D合成和现场数据。与传统的去噪方法和最新的UL框架相比,我们所提出的方法表现出优越的降噪性能。
摘要:Distributed acoustic sensor (DAS) technology leverages optical fiber cablesto detect acoustic signals, providing cost-effective and dense monitoringcapabilities. It offers several advantages including resistance to extremeconditions, immunity to electromagnetic interference, and accurate detection.However, DAS typically exhibits a lower signal-to-noise ratio (S/N) compared togeophones and is susceptible to various noise types, such as random noise,erratic noise, level noise, and long-period noise. This reduced S/N cannegatively impact data analyses containing inversion and interpretation. Whileartificial intelligence has demonstrated excellent denoising capabilities, mostexisting methods rely on supervised learning with labeled data, which imposesstringent requirements on the quality of the labels. To address this issue, wedevelop a label-free unsupervised learning (UL) network model based onContext-Pyramid-UNet (CP-UNet) to suppress erratic and random noises in DASdata. The CP-UNet utilizes the Context Pyramid Module in the encoding anddecoding process to extract features and reconstruct the DAS data. To enhancethe connectivity between shallow and deep features, we add a Connected Module(CM) to both encoding and decoding section. Layer Normalization (LN) isutilized to replace the commonly employed Batch Normalization (BN),accelerating the convergence of the model and preventing gradient explosionduring training. Huber-loss is adopted as our loss function whose parametersare experimentally determined. We apply the network to both the 2-D syntheticand filed data. Comparing to traditional denoising methods and the latest ULframework, our proposed method demonstrates superior noise reductionperformance.

eess.AS音频处理

【1】 RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion  Models with Jointly Learned Prior
标题:RestoreGrad:使用条件去噪扩散模型和共同学习先验的信号恢复
链接:https://arxiv.org/abs/2502.13574
作者:Ching-Hua Lee,  Chouchang Yang,  Jaejin Cho,  Yashas Malur Saidutta,  Rakshith Sharma Srinivasa,  Yilin Shen,  Hongxia Jin
摘要:去噪扩散概率模型(DDPM)可以用于通过在退化信号上调节模型来从其退化观测恢复干净信号。退化信号本身是干净信号的污染版本;由于这种相关性,它们可能包含有关目标干净数据分布的某些有用信息。然而,现有的采用标准高斯作为先验分布,反过来又丢弃这样的信息,导致次优性能。在本文中,我们提出了改善条件DDPM信号恢复利用更丰富的信息之前,与扩散模型共同学习。该框架称为RestoreGrad,将DDPM无缝集成到变分自动编码器框架中,并利用退化信号和干净信号之间的相关性来编码更好的扩散先验。在语音和图像恢复任务中,我们表明RestoreGrad表现出更快的收敛速度(训练步骤减少5-10倍),以实现比现有DDPM基线更好的恢复信号质量,并提高了在推理时间内使用更少采样步骤的鲁棒性(减少2-2.5倍),倡导利用联合学习先验提高扩散过程效率的优势。
摘要:Denoising diffusion probabilistic models (DDPMs) can be utilized forrecovering a clean signal from its degraded observation(s) by conditioning themodel on the degraded signal. The degraded signals are themselves contaminatedversions of the clean signals; due to this correlation, they may encompasscertain useful information about the target clean data distribution. However,existing adoption of the standard Gaussian as the prior distribution in turndiscards such information, resulting in sub-optimal performance. In this paper,we propose to improve conditional DDPMs for signal restoration by leveraging amore informative prior that is jointly learned with the diffusion model. Theproposed framework, called RestoreGrad, seamlessly integrates DDPMs into thevariational autoencoder framework and exploits the correlation between thedegraded and clean signals to encode a better diffusion prior. On speech andimage restoration tasks, we show that RestoreGrad demonstrates fasterconvergence (5-10 times fewer training steps) to achieve better quality ofrestored signals over existing DDPM baselines, and improved robustness to usingfewer sampling steps in inference time (2-2.5 times fewer), advocating theadvantages of leveraging jointly learned prior for efficiency improvements inthe diffusion process.

【2】 Multi-channel Replay Speech Detection using an Adaptive Learnable  Beamformer
标题:使用自适应可学习束形成器的多通道重播语音检测
链接:https://arxiv.org/abs/2502.13473
作者:Michael Neri,  Tuomas Virtanen
备注:Submitted to IEEE Open Journal of Signal Processing
摘要:重放攻击属于针对语音控制系统的严重威胁类别,它利用录制和重放的语音对语音信号的容易访问性来授权对敏感数据的未经授权的访问。在这项工作中,我们提出了一个多通道的神经网络架构称为M-ALRAD的重放攻击检测的基础上的空间音频功能。这种方法将可学习的自适应波束形成器与卷积递归神经网络集成在一起,从而实现空间滤波和分类的联合优化。实验已经在ReMASC数据集上进行,ReMASC数据集是最先进的多通道重放语音检测数据集,包含具有不同阵列配置和四种环境的四个麦克风。ReMASC数据集上的结果表明,与最先进的方法相比,该方法具有优越性,并在具有挑战性的声学环境中取得了实质性的改进。此外,我们证明了我们的方法能够更好地推广到看不见的环境相对于以前的研究。
摘要:Replay attacks belong to the class of severe threats against voice-controlledsystems, exploiting the easy accessibility of speech signals by recorded andreplayed speech to grant unauthorized access to sensitive data. In this work,we propose a multi-channel neural network architecture called M-ALRAD for thedetection of replay attacks based on spatial audio features. This approachintegrates a learnable adaptive beamformer with a convolutional recurrentneural network, allowing for joint optimization of spatial filtering andclassification. Experiments have been carried out on the ReMASC dataset, whichis a state-of-the-art multi-channel replay speech detection datasetencompassing four microphones with diverse array configurations and fourenvironments. Results on the ReMASC dataset show the superiority of theapproach compared to the state-of-the-art and yield substantial improvementsfor challenging acoustic environments. In addition, we demonstrate that ourapproach is able to better generalize to unseen environments with respect toprior studies.

【3】 Adopting Whisper for Confidence Estimation
标题:采用Whisper进行置信度估计
链接:https://arxiv.org/abs/2502.13446
作者:Vaibhav Aggarwal,  Shabari S Nair,  Yash Verma,  Yash Jogi
备注:Accepted at IEEE ICASSP 2025
摘要:最近对语音识别系统的单词级置信度估计的研究主要集中在称为置信度估计模块(CEMs)的轻量级模型上,该模型依赖于从自动语音识别(ASR)输出中获得的手工设计的特征。相比之下,我们提出了一种新颖的端到端方法,该方法利用ASR模型本身(Whisper)来生成单词级别的置信度评分。具体来说,我们介绍了一种方法,其中的耳语模型进行微调,以产生标量的置信度分数给定的音频输入及其相应的假设成绩单。我们的实验表明,微调的Whisper-tiny模型,在大小上与强大的CEM基线相当,在域内数据集上实现了类似的性能,并在八个域外数据集上超过了CEM基线,而微调的Whisper-large模型在所有数据集上的性能都超过了CEM基线。
摘要:Recent research on word-level confidence estimation for speech recognitionsystems has primarily focused on lightweight models known as ConfidenceEstimation Modules (CEMs), which rely on hand-engineered features derived fromAutomatic Speech Recognition (ASR) outputs. In contrast, we propose a novelend-to-end approach that leverages the ASR model itself (Whisper) to generateword-level confidence scores. Specifically, we introduce a method in which theWhisper model is fine-tuned to produce scalar confidence scores given an audioinput and its corresponding hypothesis transcript. Our experiments demonstratethat the fine-tuned Whisper-tiny model, comparable in size to a strong CEMbaseline, achieves similar performance on the in-domain dataset and surpassesthe CEM baseline on eight out-of-domain datasets, whereas the fine-tunedWhisper-large model consistently outperforms the CEM baseline by a substantialmargin across all datasets.

【4】 Audio-Based Classification of Insect Species Using Machine Learning  Models: Cicada, Beetle, Termite, and Cricket
标题:使用机器学习模型对昆虫物种进行基于音频的分类:蝉、甲虫、白蚁和蟋蟀
链接:https://arxiv.org/abs/2502.13893
作者:Manas V Shetty,  Yoga Disha Sendhil Kumar
摘要:这个项目解决了昆虫物种分类的挑战:蝉,甲虫,白蚁和蟋蟀使用录音。准确的物种识别对于生态监测和害虫管理至关重要。我们采用XGBoost、随机森林和K最近邻(KNN)等机器学习模型来分析音频特征,包括Mel频率倒谱系数(MFCC)。这项工作的潜在新颖性在于将不同的音频特征和机器学习模型相结合,以解决昆虫分类问题,特别是专注于捕捉物种之间的微妙声学变化,这些变化在以前的研究中尚未得到充分利用。该数据集是从各种开放源编译的,我们预计将实现高分类精度,有助于改进自动昆虫检测系统。
摘要:This project addresses the challenge of classifying insect species: Cicada,Beetle, Termite, and Cricket using sound recordings. Accurate speciesidentification is crucial for ecological monitoring and pest management. Weemploy machine learning models such as XGBoost, Random Forest, and K NearestNeighbors (KNN) to analyze audio features, including Mel Frequency CepstralCoefficients (MFCC). The potential novelty of this work lies in the combinationof diverse audio features and machine learning models to tackle insectclassification, specifically focusing on capturing subtle acoustic variationsbetween species that have not been fully leveraged in previous research. Thedataset is compiled from various open sources, and we anticipate achieving highclassification accuracy, contributing to improved automated insect detectionsystems.

【5】 TALKPLAY: Multimodal Music Recommendation with Large Language Models
标题:TALKSYS:具有大型语言模型的多模式音乐推荐
链接:https://arxiv.org/abs/2502.13713
作者:Seungheon Doh,  Keunwoo Choi,  Juhan Nam
摘要:我们提出了TalkPlay,一个多模态音乐推荐系统,重新制定的推荐任务,大语言模型令牌生成。TalkPlay通过扩展的令牌词汇表来表示音乐,该词汇表对多种形式进行编码-音频,歌词,元数据,语义标签和播放列表共现。使用这些丰富的表示,该模型通过对音乐推荐对话的下一个令牌预测来学习生成推荐,这需要学习自然语言查询和响应以及音乐项的关联。换句话说,该公式将音乐推荐转换为自然语言理解任务,其中模型预测会话令牌的能力直接优化了查询项相关性。我们的方法消除了传统的解释对话管道的复杂性,实现了查询感知音乐推荐的端到端学习。在实验中,TalkPlay被成功训练,并在各个方面优于基线方法,展示了作为会话式音乐推荐器的强大上下文理解。
摘要:We present TalkPlay, a multimodal music recommendation system thatreformulates the recommendation task as large language model token generation.TalkPlay represents music through an expanded token vocabulary that encodesmultiple modalities - audio, lyrics, metadata, semantic tags, and playlistco-occurrence. Using these rich representations, the model learns to generaterecommendations through next-token prediction on music recommendationconversations, that requires learning the associations natural language queryand response, as well as music items. In other words, the formulationtransforms music recommendation into a natural language understanding task,where the model's ability to predict conversation tokens directly optimizesquery-item relevance. Our approach eliminates traditionalrecommendation-dialogue pipeline complexity, enabling end-to-end learning ofquery-aware music recommendations. In the experiment, TalkPlay is successfullytrained and outperforms baseline methods in various aspects, demonstratingstrong context understanding as a conversational music recommender.

【6】 Semi-supervised classification of bird vocalizations
标题:鸟类发声的半监督分类
链接:https://arxiv.org/abs/2502.13440
作者:Simen Hexeberg,  Mandar Chitre,  Matthias Hoffmann-Kuhnt,  Bing Wen Low
摘要:鸟类数量的变化可以表明生态系统的更广泛变化,使鸟类成为最重要的动物群体之一。将机器学习和被动声学相结合,可以在没有人类直接参与的情况下进行长时间的连续监测。然而,大多数现有的技术需要大量的专家标记的数据集进行训练,并且不能容易地检测繁忙音景中的时间重叠呼叫。我们提出了一个半监督的声学鸟检测器,旨在允许检测时间重叠的调用(当分离的频率)和使用几个标记的训练样本。该分类器是在社区记录的开源数据和来自新加坡的长时间音景录音的组合上进行训练和评估的。它实现了一个平均F0.5得分0.701跨越315类110种鸟类的保持测试集,平均11个标记的训练样本,每个类。它在103种鸟类的测试集上的表现优于最先进的BirdNET分类器,尽管标记的训练样本明显较少。该探测器进一步测试了144个麦克风小时的连续声景数据。新加坡丰富的音景使得抑制误报对原始连续数据流构成挑战。尽管如此,我们证明了在这种环境中用最少的标记训练数据实现高精度是可能的。
摘要:Changes in bird populations can indicate broader changes in ecosystems,making birds one of the most important animal groups to monitor. Combiningmachine learning and passive acoustics enables continuous monitoring overextended periods without direct human involvement. However, most existingtechniques require extensive expert-labeled datasets for training and cannoteasily detect time-overlapping calls in busy soundscapes. We propose asemi-supervised acoustic bird detector designed to allow both the detection oftime-overlapping calls (when separated in frequency) and the use of few labeledtraining samples. The classifier is trained and evaluated on a combination ofcommunity-recorded open-source data and long-duration soundscape recordingsfrom Singapore. It achieves a mean F0.5 score of 0.701 across 315 classes from110 bird species on a hold-out test set, with an average of 11 labeled trainingsamples per class. It outperforms the state-of-the-art BirdNET classifier on atest set of 103 bird species despite significantly fewer labeled trainingsamples. The detector is further tested on 144 microphone-hours of continuoussoundscape data. The rich soundscape in Singapore makes suppression of falsepositives a challenge on raw, continuous data streams. Nevertheless, wedemonstrate that achieving high precision in such environments with minimallabeled training data is possible.

【7】 MATS: An Audio Language Model under Text-only Supervision
标题:MATS:纯文本监督下的音频语言模型
链接:https://arxiv.org/abs/2502.13433
作者:Wen Wang,  Ruibing Hou,  Hong Chang,  Shiguang Shan,  Xilin Chen
备注:19 pages,11 figures
摘要:建立在强大的大型语言模型(LLM)基础上的大型音频语言模型(LALM)表现出了卓越的音频理解和推理能力。然而,LALM的训练需要大量的音频-语言对语料库,这需要大量的数据收集和训练资源。在本文中,我们提出了MATS,音频语言多模态LLM设计用于处理多个音频任务,仅使用文本监督。通过利用CLAP等预训练的音频语言对齐模型,我们开发了一种纯文本训练策略,将共享的音频语言潜在空间投影到LLM潜在空间中,赋予LLM音频理解能力,而无需在训练期间依赖音频数据。为了进一步弥合CLAP中音频和语言嵌入之间的模态差距,我们提出了带音频的强相关噪声文本(Santa)机制。Santa将音频嵌入映射到CLAP语言嵌入空间,同时保留来自音频输入的基本信息。大量的实验表明,MATS,尽管专门训练的文本数据,实现竞争力的性能相比,最近的LALM训练的大规模的音频语言对。
摘要:Large audio-language models (LALMs), built upon powerful Large LanguageModels (LLMs), have exhibited remarkable audio comprehension and reasoningcapabilities. However, the training of LALMs demands a large corpus ofaudio-language pairs, which requires substantial costs in both data collectionand training resources. In this paper, we propose MATS, an audio-languagemultimodal LLM designed to handle Multiple Audio task using solely Text-onlySupervision. By leveraging pre-trained audio-language alignment models such asCLAP, we develop a text-only training strategy that projects the sharedaudio-language latent space into LLM latent space, endowing the LLM with audiocomprehension capabilities without relying on audio data during training. Tofurther bridge the modality gap between audio and language embeddings withinCLAP, we propose the Strongly-related noisy text with audio (Santa) mechanism.Santa maps audio embeddings into CLAP language embedding space while preservingessential information from the audio input. Extensive experiments demonstratethat MATS, despite being trained exclusively on text data, achieves competitiveperformance compared to recent LALMs trained on large-scale audio-languagepairs.

【8】 Unsupervised CP-UNet Framework for Denoising DAS Data with Decay Noise
标题:无监督CP-UNet框架用于消除含衰变噪音的DAS数据
链接:https://arxiv.org/abs/2502.13395
作者:Tianye Huang,  Aopeng Li,  Xiang Li,  Jing Zhang,  Sijing Xian,  Qi Zhang,  Mingkong Lu,  Guodong Chen,  Liangming Xiong,  Xiangyun Hu
备注:13 pages, 8 figures
摘要:分布式声学传感器(DAS)技术利用光纤电缆来检测声学信号,提供经济高效的密集监测功能。它提供了几个优点,包括耐极端条件,抗电磁干扰和准确的检测。然而,DAS与地震检波器相比通常表现出较低的信噪比(S/N),并且易受各种噪声类型的影响,例如随机噪声、不稳定噪声、电平噪声和长周期噪声。这种降低的S/N可能会对包含反演和解释的数据分析产生负面影响。虽然人工智能已经展示了出色的去噪能力,但大多数现有方法都依赖于对标记数据的监督学习,这对标签的质量提出了严格的要求。为了解决这个问题,我们开发了一种基于上下文金字塔UNet(CP-UNet)的无标签无监督学习(UL)网络模型,以抑制DAS数据中的不稳定和随机噪声。CP-UNet在编码和解码过程中利用上下文金字塔模块来提取特征并重建DAS数据。为了增强浅层和深层特征之间的连接,我们在编码和解码部分都添加了连接模块(CM)。层归一化(LN)用于取代常用的批量归一化(BN),加速模型的收敛并防止训练过程中的梯度爆炸。Huber损失被采用作为我们的损失函数,其参数是实验确定的。我们应用网络的2-D合成和现场数据。与传统的去噪方法和最新的UL框架相比,我们所提出的方法表现出优越的降噪性能。
摘要:Distributed acoustic sensor (DAS) technology leverages optical fiber cablesto detect acoustic signals, providing cost-effective and dense monitoringcapabilities. It offers several advantages including resistance to extremeconditions, immunity to electromagnetic interference, and accurate detection.However, DAS typically exhibits a lower signal-to-noise ratio (S/N) compared togeophones and is susceptible to various noise types, such as random noise,erratic noise, level noise, and long-period noise. This reduced S/N cannegatively impact data analyses containing inversion and interpretation. Whileartificial intelligence has demonstrated excellent denoising capabilities, mostexisting methods rely on supervised learning with labeled data, which imposesstringent requirements on the quality of the labels. To address this issue, wedevelop a label-free unsupervised learning (UL) network model based onContext-Pyramid-UNet (CP-UNet) to suppress erratic and random noises in DASdata. The CP-UNet utilizes the Context Pyramid Module in the encoding anddecoding process to extract features and reconstruct the DAS data. To enhancethe connectivity between shallow and deep features, we add a Connected Module(CM) to both encoding and decoding section. Layer Normalization (LN) isutilized to replace the commonly employed Batch Normalization (BN),accelerating the convergence of the model and preventing gradient explosionduring training. Huber-loss is adopted as our loss function whose parametersare experimentally determined. We apply the network to both the 2-D syntheticand filed data. Comparing to traditional denoising methods and the latest ULframework, our proposed method demonstrates superior noise reductionperformance.

机器翻译由腾讯交互翻译提供,仅供参考