今日论文合集:cs.SD语音5篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 A Multi-task Learning Balanced Attention Convolutional Neural Network  Model for Few-shot Underwater Acoustic Target Recognition
标题: 一种多任务学习的平衡注意力卷积神经网络少炮水声目标识别模型
链接:https://arxiv.org/abs/2504.13102
作者: Wei Huang,  Shumeng Sun,  Junpeng Lu,  Zhenpeng Xu,  Zhengyang Xiu,  Hao Zhang 
摘要:水声目标识别对于保护海洋生物多样性和国防安全具有重要意义。深度学习的发展为UATR提供了新的机遇,但面临参考样本稀缺和复杂环境干扰带来的挑战。为了解决这些问题,我们提出了一个多任务平衡信道注意力卷积神经网络(MT-BCA-CNN)。该方法将通道注意机制与多任务学习策略相结合,构建共享的特征提取器和多任务分类器,共同优化目标分类和特征重构任务。通道注意机制动态地增强辨别性声学特征,诸如谐波结构,同时抑制噪声。在Watkins海洋生物数据集上的实验表明,MT-BCA-CNN在27类Few-Shot场景下达到了97%的分类准确率和95%的$F1$-score,显著优于传统的CNN和ACNN模型,以及流行的最先进的UATR方法。消融研究证实了多任务学习和注意机制的协同效益,而动态权重调整策略有效地平衡了任务的贡献。本文的工作为Few-Shot水声目标识别提供了一种有效的解决方案,推动了海洋生物声学和声纳信号处理的研究。
摘要:Underwater acoustic target recognition (UATR) is of great significance for the protection of marine diversity and national defense security. The development of deep learning provides new opportunities for UATR, but faces challenges brought by the scarcity of reference samples and complex environmental interference. To address these issues, we proposes a multi-task balanced channel attention convolutional neural network (MT-BCA-CNN). The method integrates a channel attention mechanism with a multi-task learning strategy, constructing a shared feature extractor and multi-task classifiers to jointly optimize target classification and feature reconstruction tasks. The channel attention mechanism dynamically enhances discriminative acoustic features such as harmonic structures while suppressing noise. Experiments on the Watkins Marine Life Dataset demonstrate that MT-BCA-CNN achieves 97\% classification accuracy and 95\% $F1$-score in 27-class few-shot scenarios, significantly outperforming traditional CNN and ACNN models, as well as popular state-of-the-art UATR methods. Ablation studies confirm the synergistic benefits of multi-task learning and attention mechanisms, while a dynamic weighting adjustment strategy effectively balances task contributions. This work provides an efficient solution for few-shot underwater acoustic recognition, advancing research in marine bioacoustics and sonar signal processing.


【2】 Can Masked Autoencoders Also Listen to Birds?

标题: 蒙面自动编码器也可以收听鸟类吗?
链接:https://arxiv.org/abs/2504.12880
作者: Lukas Rauch,  Ilyass Moummad,  René Heinrich,  Alexis Joly,  Bernhard Sick,  Christoph Scholz 
摘要:在AudioSet上预训练的Masked Autoencoders(MAE)无法捕获生物声学监测等专业领域的细粒度声学特征。鸟类声音分类对于评估环境健康至关重要,但通用模型无法充分解决其独特的声学挑战。为了解决这个问题,我们引入了Bird-MAE,这是一个在大规模BirdSet数据集上预训练的领域专用MAE。我们探索调整预训练,微调和利用冻结表示。Bird-MAE在所有BirdSet下游任务中实现了最先进的结果,与通用Audio-MAE基线相比,大大提高了多标签分类性能。此外,我们提出了原型探测,一个参数有效的方法,利用MAE的冻结表示。Bird-MAE的原型探测器在MAP中的性能比线性探测器高出37%,并在BirdSet上将差距缩小到平均约3%。
摘要:Masked Autoencoders (MAEs) pretrained on AudioSet fail to capture the fine-grained acoustic characteristics of specialized domains such as bioacoustic monitoring. Bird sound classification is critical for assessing environmental health, yet general-purpose models inadequately address its unique acoustic challenges. To address this, we introduce Bird-MAE, a domain-specialized MAE pretrained on the large-scale BirdSet dataset. We explore adjustments to pretraining, fine-tuning and utilizing frozen representations. Bird-MAE achieves state-of-the-art results across all BirdSet downstream tasks, substantially improving multi-label classification performance compared to the general-purpose Audio-MAE baseline. Additionally, we propose prototypical probing, a parameter-efficient method for leveraging MAEs' frozen representations. Bird-MAE's prototypical probes outperform linear probing by up to 37\% in MAP and narrow the gap to fine-tuning to approximately 3\% on average on BirdSet.


【3】 A Survey on Cross-Modal Interaction Between Music and Multimodal Data

标题: 音乐与多模式数据的跨模式交互研究
链接:https://arxiv.org/abs/2504.12796
作者: Sifei Li,  Mining Tan,  Feier Shen,  Minyan Luo,  Zijiao Yin,  Fan Tang,  Weiming Dong,  Changsheng Xu 
备注:34 pages, 7 figures
摘要:多模式学习推动了各个行业的创新,特别是在音乐领域。通过实现更直观的交互体验和增强沉浸感,它不仅降低了音乐的进入门槛,还增加了其整体吸引力。本调查旨在提供与音乐相关的多模态任务的全面综述,概述音乐如何有助于多模态学习,并为寻求扩大计算音乐边界的研究人员提供见解。与通常在语义或视觉上直观的文本和图像不同,音乐主要通过听觉感知与人类交互,使得其数据表示本质上不那么直观。因此,本文首先介绍了音乐的表示,并提供了音乐数据集的概述。随后,我们将音乐与多模态数据之间的跨模态交互分为三种类型:音乐驱动的跨模态交互、面向音乐的跨模态交互和双向音乐跨模态交互。对于每个类别,我们系统地跟踪相关子任务的发展,分析现有的局限性,并讨论新的趋势。此外,我们提供了一个全面的总结数据集和评估指标用于多模态任务相关的音乐,为未来的研究提供基准参考。最后,我们讨论了目前的挑战,在跨模态的互动涉及音乐,并提出了未来的研究方向。
摘要:Multimodal learning has driven innovation across various industries, particularly in the field of music. By enabling more intuitive interaction experiences and enhancing immersion, it not only lowers the entry barriers to the music but also increases its overall appeal. This survey aims to provide a comprehensive review of multimodal tasks related to music, outlining how music contributes to multimodal learning and offering insights for researchers seeking to expand the boundaries of computational music. Unlike text and images, which are often semantically or visually intuitive, music primarily interacts with humans through auditory perception, making its data representation inherently less intuitive. Therefore, this paper first introduces the representations of music and provides an overview of music datasets. Subsequently, we categorize cross-modal interactions between music and multimodal data into three types: music-driven cross-modal interactions, music-oriented cross-modal interactions, and bidirectional music cross-modal interactions. For each category, we systematically trace the development of relevant sub-tasks, analyze existing limitations, and discuss emerging trends. Furthermore, we provide a comprehensive summary of datasets and evaluation metrics used in multimodal tasks related to music, offering benchmark references for future research. Finally, we discuss the current challenges in cross-modal interactions involving music and propose potential directions for future research.


【4】 GOAT-TTS: LLM-based Text-To-Speech Generation Optimized via A  Dual-Branch Architecture

标题: GOAT-TTC:通过双分支架构优化的基于LLM的文本转语音生成
链接:https://arxiv.org/abs/2504.12339
作者: Yaodong Song,  Hongjie Chen,  Jie Lian,  Yuxin Zhang,  Guangmin Xia,  Zehan Li,  Genliang Zhao,  Jian Kang,  Yongxiang Li,  Jie Li 
摘要:None
摘要:While large language models (LLMs) have revolutionized text-to-speech (TTS) synthesis through discrete tokenization paradigms, current architectures exhibit fundamental tensions between three critical dimensions: 1) irreversible loss of acoustic characteristics caused by quantization of speech prompts; 2) stringent dependence on precisely aligned prompt speech-text pairs that limit real-world deployment; and 3) catastrophic forgetting of the LLM's native text comprehension during optimization for speech token generation. To address these challenges, we propose an LLM-based text-to-speech Generation approach Optimized via a novel dual-branch ArchiTecture (GOAT-TTS). Our framework introduces two key innovations: (1) The modality-alignment branch combines a speech encoder and projector to capture continuous acoustic embeddings, enabling bidirectional correlation between paralinguistic features (language, timbre, emotion) and semantic text representations without transcript dependency; (2) The speech-generation branch employs modular fine-tuning on top-k layers of an LLM for speech token prediction while freezing the bottom-k layers to preserve foundational linguistic knowledge. Moreover, multi-token prediction is introduced to support real-time streaming TTS synthesis. Experimental results demonstrate that our GOAT-TTS achieves performance comparable to state-of-the-art TTS models while validating the efficacy of synthesized dialect speech data.


【5】 Temporal Attention Pooling for Frequency Dynamic Convolution in Sound  Event Detection

标题: 声音事件检测中频率动态卷积的时间注意力池
链接:https://arxiv.org/abs/2504.12670
作者: Hyeonuk Nam,  Yong-Hwa Park 
摘要:深度学习的最新进展,特别是频率动态卷积(FDY conv),通过启用频率自适应特征提取,显着改善了声音事件检测(SED)。然而,FDY conv依赖于时间平均池化,这平等地对待所有时间帧,限制了其捕获瞬态声音事件(如闹钟、敲门声和语音爆破音)的能力。为了解决这个问题,我们提出了时间注意力池频率动态卷积(TFD conv),以取代时间平均池与时间注意力池(TAP)。TAP通过三种互补机制自适应地对时间特征进行加权:时间注意力池(TA)用于强调突出特征,速度注意力池(VA)用于捕获瞬态变化,以及传统的平均池用于对平稳信号的鲁棒性。消融研究表明,与FDY conv相比,TFD conv将平均PSDS 1提高了3.02%,参数计数仅增加了14.8%。Classwise ANOVA和Tukey HSD分析进一步表明,TFD conv显著增强了瞬态重事件的检测性能,优于现有的FDY conv模型。值得注意的是,TFD conv达到了0.456的最大PSDS 1分数,超过了之前最先进的SED系统。我们还探索了TAP与其他FDY conv变体的兼容性,包括扩张FDY conv(DFD conv),部分FDY conv(PFD conv)和多扩张FDY conv(MDFD conv)。其中,TAP与MDFD conv的整合达到了最好的结果,PSDS 1得分为0.459,验证了时间注意和多尺度频率适应的互补优势。这些研究结果建立TFD conv作为一个强大的和可推广的框架,提高瞬态灵敏度和整体功能的鲁棒性在SED。
摘要:Recent advances in deep learning, particularly frequency dynamic convolution (FDY conv), have significantly improved sound event detection (SED) by enabling frequency-adaptive feature extraction. However, FDY conv relies on temporal average pooling, which treats all temporal frames equally, limiting its ability to capture transient sound events such as alarm bells, door knocks, and speech plosives. To address this limitation, we propose temporal attention pooling frequency dynamic convolution (TFD conv) to replace temporal average pooling with temporal attention pooling (TAP). TAP adaptively weights temporal features through three complementary mechanisms: time attention pooling (TA) for emphasizing salient features, velocity attention pooling (VA) for capturing transient changes, and conventional average pooling for robustness to stationary signals. Ablation studies show that TFD conv improves average PSDS1 by 3.02% over FDY conv with only a 14.8% increase in parameter count. Classwise ANOVA and Tukey HSD analysis further demonstrate that TFD conv significantly enhances detection performance for transient-heavy events, outperforming existing FDY conv models. Notably, TFD conv achieves a maximum PSDS1 score of 0.456, surpassing previous state-of-the-art SED systems. We also explore the compatibility of TAP with other FDY conv variants, including dilated FDY conv (DFD conv), partial FDY conv (PFD conv), and multi-dilated FDY conv (MDFD conv). Among these, the integration of TAP with MDFD conv achieves the best result with a PSDS1 score of 0.459, validating the complementary strengths of temporal attention and multi-scale frequency adaptation. These findings establish TFD conv as a powerful and generalizable framework for enhancing both transient sensitivity and overall feature robustness in SED.


eess.AS音频处理

【1】 CST-former: Multidimensional Attention-based Transformer for Sound Event  Localization and Detection in Real Scenes

标题: CST-former:用于真实场景中声音事件定位和检测的多维基于注意力的Transformer
链接:https://arxiv.org/abs/2504.12870
作者: Yusun Shul,  Dayun Choi,  Jung-Woo Choi 
备注:12 pages, 10 figures, Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:声事件定位与检测(SELD)是利用多通道声信号对声事件进行分类和波达方向(DoA)识别的任务。为了有效地分类和定位,提出了一种通道-频谱-时间Transformer(CST-former)。CST-former在空间、光谱和时间域上采用多维注意机制,以扩大模型的能力,从而随着时间的推移学习事件检测和DoA估计所必需的域信息。在这项工作中,我们提出了一个增强版的CST前多尺度展开局部嵌入(MSULE)开发的捕获和聚合域信息在多个时间-频率尺度。此外,我们提出了微调和后处理技术,有利于在有限的训练数据集上进行SELD任务。深入的消融研究所提出的架构和详细的分析,建议的模块进行了验证SELD任务的多维注意力的有效性。通过STARSS 22和STARSS 23数据集上的实验进行实证验证,证明了CST前处理和后处理技术在不使用外部数据的情况下的显着性能。
摘要:Sound event localization and detection (SELD) is a task for the classification of sound events and the identification of direction of arrival (DoA) utilizing multichannel acoustic signals. For effective classification and localization, a channel-spectro-temporal transformer (CST-former) was suggested. CST-former employs multidimensional attention mechanisms across the spatial, spectral, and temporal domains to enlarge the model's capacity to learn the domain information essential for event detection and DoA estimation over time. In this work, we present an enhanced version of CST-former with multiscale unfolded local embedding (MSULE) developed to capture and aggregate domain information over multiple time-frequency scales. Also, we propose finetuning and post-processing techniques beneficial for conducting the SELD task over limited training datasets. In-depth ablation studies of the proposed architecture and detailed analysis on the proposed modules are carried out to validate the efficacy of multidimensional attentions on the SELD task. Empirical validation through experimentation on STARSS22 and STARSS23 datasets demonstrates the remarkable performance of CST-former and post-processing techniques without using external data.


【2】 EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text  Prompting

标题: DeliverVoice:基于LLM的情感文本到语音模型,具有自由式文本预处理
链接:https://arxiv.org/abs/2504.12867
作者: Guanrou Yang,  Chen Yang,  Qian Chen,  Ziyang Ma,  Wenxi Chen,  Wen Wang,  Tianrui Wang,  Yifan Yang,  Zhikang Niu,  Wenrui Liu,  Fan Yu,  Zhihao Du,  Zhifu Gao,  ShiLiang Zhang,  Xie Chen 
摘要:人类的语言不仅仅是信息的传递,它是一种深刻的情感交流和个人之间的联系。虽然文语转换(TTS)模型已经取得了巨大的进步,但它们仍然面临着控制生成的语音中的情感表达的挑战。在这项工作中,我们提出了一种新的情感可控的TTS模型,利用大语言模型(LLM),使细粒度的自由式自然语言情感控制,和音素增强变体设计,使模型输出音素令牌和音频令牌并行,以提高内容的一致性,灵感来自思想链(CoT)和模态的思想(CoM)技术。此外,我们还介绍了一个高质量的40小时英语情感数据集,它具有表达性的语音和细粒度的情感标签,并带有自然语言描述。在仅使用合成训练数据的英语语音测试集和使用我们内部数据的中文Secap测试集上,语音识别实现了最先进的性能。我们进一步研究了现有的情感评估指标的可靠性及其与人类感知偏好的一致性,并探索使用SOTA多模态LLM GPT-4 O音频和Gemini来评估情感语音。演示示例可在https://anonymous.4open.science/r/EmoVoice-DF55上获得。数据集、代码和检查点将被释放。
摘要:Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and modality-of-thought (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Demo samples are available at https://anonymous.4open.science/r/EmoVoice-DF55. Dataset, code, and checkpoints will be released.


【3】 Temporal Attention Pooling for Frequency Dynamic Convolution in Sound  Event Detection

标题: 声音事件检测中频率动态卷积的时间注意力池
链接:https://arxiv.org/abs/2504.12670
作者: Hyeonuk Nam,  Yong-Hwa Park 
摘要:深度学习的最新进展,特别是频率动态卷积(FDY conv),通过启用频率自适应特征提取,显着改善了声音事件检测(SED)。然而,FDY conv依赖于时间平均池化,这平等地对待所有时间帧,限制了其捕获瞬态声音事件(如闹钟、敲门声和语音爆破音)的能力。为了解决这个问题,我们提出了时间注意力池频率动态卷积(TFD conv),以取代时间平均池与时间注意力池(TAP)。TAP通过三种互补机制自适应地对时间特征进行加权:时间注意力池(TA)用于强调突出特征,速度注意力池(VA)用于捕获瞬态变化,以及传统的平均池用于对平稳信号的鲁棒性。消融研究表明,与FDY conv相比,TFD conv将平均PSDS 1提高了3.02%,参数计数仅增加了14.8%。Classwise ANOVA和Tukey HSD分析进一步表明,TFD conv显著增强了瞬态重事件的检测性能,优于现有的FDY conv模型。值得注意的是,TFD conv达到了0.456的最大PSDS 1分数,超过了之前最先进的SED系统。我们还探索了TAP与其他FDY conv变体的兼容性,包括扩张FDY conv(DFD conv),部分FDY conv(PFD conv)和多扩张FDY conv(MDFD conv)。其中,TAP与MDFD conv的整合达到了最好的结果,PSDS 1得分为0.459,验证了时间注意和多尺度频率适应的互补优势。这些研究结果建立TFD conv作为一个强大的和可推广的框架,提高瞬态灵敏度和整体功能的鲁棒性在SED。
摘要:Recent advances in deep learning, particularly frequency dynamic convolution (FDY conv), have significantly improved sound event detection (SED) by enabling frequency-adaptive feature extraction. However, FDY conv relies on temporal average pooling, which treats all temporal frames equally, limiting its ability to capture transient sound events such as alarm bells, door knocks, and speech plosives. To address this limitation, we propose temporal attention pooling frequency dynamic convolution (TFD conv) to replace temporal average pooling with temporal attention pooling (TAP). TAP adaptively weights temporal features through three complementary mechanisms: time attention pooling (TA) for emphasizing salient features, velocity attention pooling (VA) for capturing transient changes, and conventional average pooling for robustness to stationary signals. Ablation studies show that TFD conv improves average PSDS1 by 3.02% over FDY conv with only a 14.8% increase in parameter count. Classwise ANOVA and Tukey HSD analysis further demonstrate that TFD conv significantly enhances detection performance for transient-heavy events, outperforming existing FDY conv models. Notably, TFD conv achieves a maximum PSDS1 score of 0.456, surpassing previous state-of-the-art SED systems. We also explore the compatibility of TAP with other FDY conv variants, including dilated FDY conv (DFD conv), partial FDY conv (PFD conv), and multi-dilated FDY conv (MDFD conv). Among these, the integration of TAP with MDFD conv achieves the best result with a PSDS1 score of 0.459, validating the complementary strengths of temporal attention and multi-scale frequency adaptation. These findings establish TFD conv as a powerful and generalizable framework for enhancing both transient sensitivity and overall feature robustness in SED.


【4】 Benchmarking Audio Deepfake Detection Robustness in Real-world  Communication Scenarios

标题: 对现实世界通信场景中的音频Deepfake检测鲁棒性进行基准测试
链接:https://arxiv.org/abs/2504.12423
作者: Haohan Shi,  Xiyu Shi,  Safak Dogan,  Saif Alzubi,  Tianjin Huang,  Yunxiao Zhang 
备注:5 pages, 3 figures, submitted to EUSIPCO 2025
摘要:None
摘要:Existing Audio Deepfake Detection (ADD) systems often struggle to generalise effectively due to the significantly degraded audio quality caused by audio codec compression and channel transmission effects in real-world communication scenarios. To address this challenge, we developed a rigorous benchmark to evaluate ADD system performance under such scenarios. We introduced ADD-C, a new test dataset to evaluate the robustness of ADD systems under diverse communication conditions, including different combinations of audio codecs for compression and Packet Loss Rates (PLR). Benchmarking on three baseline ADD models with the ADD-C dataset demonstrated a significant decline in robustness under such conditions. A novel data augmentation strategy was proposed to improve the robustness of ADD systems. Experimental results demonstrated that the proposed approach increases the performance of ADD systems significantly with the proposed ADD-C dataset. Our benchmark can assist future efforts towards building practical and robustly generalisable ADD systems.


【5】 A Multi-task Learning Balanced Attention Convolutional Neural Network  Model for Few-shot Underwater Acoustic Target Recognition

标题: 一种多任务学习的平衡注意力卷积神经网络少炮水声目标识别模型
链接:https://arxiv.org/abs/2504.13102
作者: Wei Huang,  Shumeng Sun,  Junpeng Lu,  Zhenpeng Xu,  Zhengyang Xiu,  Hao Zhang 
摘要:水声目标识别对于保护海洋生物多样性和国防安全具有重要意义。深度学习的发展为UATR提供了新的机遇,但面临参考样本稀缺和复杂环境干扰带来的挑战。为了解决这些问题,我们提出了一个多任务平衡信道注意力卷积神经网络(MT-BCA-CNN)。该方法将通道注意机制与多任务学习策略相结合,构建共享的特征提取器和多任务分类器,共同优化目标分类和特征重构任务。通道注意机制动态地增强辨别性声学特征,诸如谐波结构,同时抑制噪声。在Watkins海洋生物数据集上的实验表明,MT-BCA-CNN在27类Few-Shot场景下达到了97%的分类准确率和95%的$F1$-score,显著优于传统的CNN和ACNN模型,以及流行的最先进的UATR方法。消融研究证实了多任务学习和注意机制的协同效益,而动态权重调整策略有效地平衡了任务的贡献。本文的工作为Few-Shot水声目标识别提供了一种有效的解决方案,推动了海洋生物声学和声纳信号处理的研究。
摘要:Underwater acoustic target recognition (UATR) is of great significance for the protection of marine diversity and national defense security. The development of deep learning provides new opportunities for UATR, but faces challenges brought by the scarcity of reference samples and complex environmental interference. To address these issues, we proposes a multi-task balanced channel attention convolutional neural network (MT-BCA-CNN). The method integrates a channel attention mechanism with a multi-task learning strategy, constructing a shared feature extractor and multi-task classifiers to jointly optimize target classification and feature reconstruction tasks. The channel attention mechanism dynamically enhances discriminative acoustic features such as harmonic structures while suppressing noise. Experiments on the Watkins Marine Life Dataset demonstrate that MT-BCA-CNN achieves 97\% classification accuracy and 95\% $F1$-score in 27-class few-shot scenarios, significantly outperforming traditional CNN and ACNN models, as well as popular state-of-the-art UATR methods. Ablation studies confirm the synergistic benefits of multi-task learning and attention mechanisms, while a dynamic weighting adjustment strategy effectively balances task contributions. This work provides an efficient solution for few-shot underwater acoustic recognition, advancing research in marine bioacoustics and sonar signal processing.


【6】 Can Masked Autoencoders Also Listen to Birds?

标题: 蒙面自动编码器也可以收听鸟类吗?
链接:https://arxiv.org/abs/2504.12880
作者: Lukas Rauch,  Ilyass Moummad,  René Heinrich,  Alexis Joly,  Bernhard Sick,  Christoph Scholz 
摘要:在AudioSet上预训练的Masked Autoencoders(MAE)无法捕获生物声学监测等专业领域的细粒度声学特征。鸟类声音分类对于评估环境健康至关重要,但通用模型无法充分解决其独特的声学挑战。为了解决这个问题,我们引入了Bird-MAE,这是一个在大规模BirdSet数据集上预训练的领域专用MAE。我们探索调整预训练,微调和利用冻结表示。Bird-MAE在所有BirdSet下游任务中实现了最先进的结果,与通用Audio-MAE基线相比,大大提高了多标签分类性能。此外,我们提出了原型探测,一个参数有效的方法,利用MAE的冻结表示。Bird-MAE的原型探测器在MAP中的性能比线性探测器高出37%,并在BirdSet上将差距缩小到平均约3%。
摘要:Masked Autoencoders (MAEs) pretrained on AudioSet fail to capture the fine-grained acoustic characteristics of specialized domains such as bioacoustic monitoring. Bird sound classification is critical for assessing environmental health, yet general-purpose models inadequately address its unique acoustic challenges. To address this, we introduce Bird-MAE, a domain-specialized MAE pretrained on the large-scale BirdSet dataset. We explore adjustments to pretraining, fine-tuning and utilizing frozen representations. Bird-MAE achieves state-of-the-art results across all BirdSet downstream tasks, substantially improving multi-label classification performance compared to the general-purpose Audio-MAE baseline. Additionally, we propose prototypical probing, a parameter-efficient method for leveraging MAEs' frozen representations. Bird-MAE's prototypical probes outperform linear probing by up to 37\% in MAP and narrow the gap to fine-tuning to approximately 3\% on average on BirdSet.


【7】 A Survey on Cross-Modal Interaction Between Music and Multimodal Data

标题: 音乐与多模式数据的跨模式交互研究
链接:https://arxiv.org/abs/2504.12796
作者: Sifei Li,  Mining Tan,  Feier Shen,  Minyan Luo,  Zijiao Yin,  Fan Tang,  Weiming Dong,  Changsheng Xu 
备注:34 pages, 7 figures
摘要:多模式学习推动了各个行业的创新,特别是在音乐领域。通过实现更直观的交互体验和增强沉浸感,它不仅降低了音乐的进入门槛,还增加了其整体吸引力。本调查旨在提供与音乐相关的多模态任务的全面综述,概述音乐如何有助于多模态学习,并为寻求扩大计算音乐边界的研究人员提供见解。与通常在语义或视觉上直观的文本和图像不同,音乐主要通过听觉感知与人类交互,使得其数据表示本质上不那么直观。因此,本文首先介绍了音乐的表示,并提供了音乐数据集的概述。随后,我们将音乐与多模态数据之间的跨模态交互分为三种类型:音乐驱动的跨模态交互、面向音乐的跨模态交互和双向音乐跨模态交互。对于每个类别,我们系统地跟踪相关子任务的发展,分析现有的局限性,并讨论新的趋势。此外,我们还提供了与音乐相关的多模态任务中使用的数据集和评估指标的全面总结,为未来的研究提供基准参考。最后,我们讨论了目前的挑战,在跨模态的互动涉及音乐,并提出了未来的研究方向。
摘要:Multimodal learning has driven innovation across various industries, particularly in the field of music. By enabling more intuitive interaction experiences and enhancing immersion, it not only lowers the entry barriers to the music but also increases its overall appeal. This survey aims to provide a comprehensive review of multimodal tasks related to music, outlining how music contributes to multimodal learning and offering insights for researchers seeking to expand the boundaries of computational music. Unlike text and images, which are often semantically or visually intuitive, music primarily interacts with humans through auditory perception, making its data representation inherently less intuitive. Therefore, this paper first introduces the representations of music and provides an overview of music datasets. Subsequently, we categorize cross-modal interactions between music and multimodal data into three types: music-driven cross-modal interactions, music-oriented cross-modal interactions, and bidirectional music cross-modal interactions. For each category, we systematically trace the development of relevant sub-tasks, analyze existing limitations, and discuss emerging trends. Furthermore, we provide a comprehensive summary of datasets and evaluation metrics used in multimodal tasks related to music, offering benchmark references for future research. Finally, we discuss the current challenges in cross-modal interactions involving music and propose potential directions for future research.


【8】 GOAT-TTS: LLM-based Text-To-Speech Generation Optimized via A  Dual-Branch Architecture

标题: GOAT-TTC:通过双分支架构优化的基于LLM的文本转语音生成
链接:https://arxiv.org/abs/2504.12339
作者: Yaodong Song,  Hongjie Chen,  Jie Lian,  Yuxin Zhang,  Guangmin Xia,  Zehan Li,  Genliang Zhao,  Jian Kang,  Yongxiang Li,  Jie Li 
摘要:虽然大型语言模型(LLM)通过离散标记化范例彻底改变了文本到语音(TTS)合成,但当前架构在三个关键维度之间表现出根本的紧张关系:1)由语音提示的量化引起的声学特性的不可逆损失; 2)严格依赖于限制现实世界部署的精确对齐的提示语音-文本对;以及3)在语音标记生成的优化期间LLM的本地文本理解的灾难性遗忘。为了解决这些挑战,我们提出了一种基于LLM的文本到语音生成方法优化通过一种新的双分支架构(GOAT-TTS)。我们的框架引入了两个关键的创新:(1)模态对齐分支结合了语音编码器和投影仪来捕获连续的声学嵌入,从而实现了语言特征之间的双向相关(语言,音色,情感)和语义文本表示没有转录依赖;(2)语音生成分支在LLM的顶部k层上采用模块化微调以用于语音令牌预测,同时冻结底部k层。k层来保存基础语言知识。此外,多令牌预测被引入到支持实时流TTS合成。实验结果表明,我们的GOAT-TTS实现性能媲美国家的最先进的TTS模型,同时验证合成方言语音数据的有效性。
摘要:While large language models (LLMs) have revolutionized text-to-speech (TTS) synthesis through discrete tokenization paradigms, current architectures exhibit fundamental tensions between three critical dimensions: 1) irreversible loss of acoustic characteristics caused by quantization of speech prompts; 2) stringent dependence on precisely aligned prompt speech-text pairs that limit real-world deployment; and 3) catastrophic forgetting of the LLM's native text comprehension during optimization for speech token generation. To address these challenges, we propose an LLM-based text-to-speech Generation approach Optimized via a novel dual-branch ArchiTecture (GOAT-TTS). Our framework introduces two key innovations: (1) The modality-alignment branch combines a speech encoder and projector to capture continuous acoustic embeddings, enabling bidirectional correlation between paralinguistic features (language, timbre, emotion) and semantic text representations without transcript dependency; (2) The speech-generation branch employs modular fine-tuning on top-k layers of an LLM for speech token prediction while freezing the bottom-k layers to preserve foundational linguistic knowledge. Moreover, multi-token prediction is introduced to support real-time streaming TTS synthesis. Experimental results demonstrate that our GOAT-TTS achieves performance comparable to state-of-the-art TTS models while validating the efficacy of synthesized dialect speech data.


机器翻译由腾讯交互翻译提供,仅供参考