微信公众号:arXiv_Daily
cs.SD语音
标题: Dragon:分配奖励优化扩散生成模型
链接:https://arxiv.org/abs/2504.15217
摘要:我们提出了分布式RewArds的生成优化(DRAGON),一个多功能的框架,微调媒体生成模型朝着预期的结果。与传统的人工反馈强化学习(RLHF)或直接偏好优化(DPO)等成对偏好方法相比,DRAGON更灵活。它可以优化评估单个示例或它们的分布的奖励函数,使其与广泛的实例,实例到分布和分布到分布的奖励兼容。利用这种多功能性,我们通过选择一个编码器和一组参考示例来构建新的奖励函数,以创建一个样本分布。当使用诸如CLAP的交叉模态编码器时,参考示例可以是不同的模态(例如,文本对音频)。然后,DRAGON收集在线和策略上的世代,对他们进行评分,以构建一个积极的示范集和一个消极集,并利用两个集之间的对比来最大化奖励。为了进行评估,我们对音频域文本到音乐扩散模型进行了微调,其中包含20种不同的奖励函数,包括自定义音乐美学模型,CLAP评分,Vendi多样性和Frechet音频距离(FAD)。我们进一步比较实例(每首歌)和全数据集FAD设置,同时消融多个FAD编码器和参考集。在所有20个目标奖励中,Dragon的平均胜率为81.45%。此外,基于范例集的奖励函数确实增强了世代,并且与基于模型的奖励相当。通过适当的样本集,DRAGON在没有对人类偏好注释进行训练的情况下实现了60.95%的人类投票音乐质量获胜率。因此,DRAGON展示了一种设计和优化奖励函数以提高人类感知质量的新方法。在https://ml-dragon.github.io/web上有声音的例子。
摘要:We present Distributional RewArds for Generative OptimizatioN (DRAGON), a versatile framework for fine-tuning media generation models towards a desired outcome. Compared with traditional reinforcement learning with human feedback (RLHF) or pairwise preference approaches such as direct preference optimization (DPO), DRAGON is more flexible. It can optimize reward functions that evaluate either individual examples or distributions of them, making it compatible with a broad spectrum of instance-wise, instance-to-distribution, and distribution-to-distribution rewards. Leveraging this versatility, we construct novel reward functions by selecting an encoder and a set of reference examples to create an exemplar distribution. When cross-modality encoders such as CLAP are used, the reference examples may be of a different modality (e.g., text versus audio). Then, DRAGON gathers online and on-policy generations, scores them to construct a positive demonstration set and a negative set, and leverages the contrast between the two sets to maximize the reward. For evaluation, we fine-tune an audio-domain text-to-music diffusion model with 20 different reward functions, including a custom music aesthetics model, CLAP score, Vendi diversity, and Frechet audio distance (FAD). We further compare instance-wise (per-song) and full-dataset FAD settings while ablating multiple FAD encoders and reference sets. Over all 20 target rewards, DRAGON achieves an 81.45% average win rate. Moreover, reward functions based on exemplar sets indeed enhance generations and are comparable to model-based rewards. With an appropriate exemplar set, DRAGON achieves a 60.95% human-voted music quality win rate without training on human preference annotations. As such, DRAGON exhibits a new approach to designing and optimizing reward functions for improving human-perceived quality. Sound examples at https://ml-dragon.github.io/web.
【2】 Histogram-based Parameter-efficient Tuning for Passive Sonar Classification
标题: 基于柱状图的被动声纳分类参数高效调整链接:https://arxiv.org/abs/2504.15214
备注:5 pages, 4 figures. Submitted to IEEE WASPAA 2025 for possible publication
摘要:参数高效迁移学习(PETL)方法使大型人工神经网络适应下游任务,而无需微调整个模型。然而,现有的加法方法,如适配器,有时很难捕捉中间特征嵌入的分布变化。我们提出了一种新的基于直方图的参数有效的调整(HPT)技术,捕获目标域的统计数据和调制的嵌入。在三个下游被动声纳数据集(ShipsEar,DeepShip,VTUAD)上的实验结果表明,HPT优于传统的适配器。值得注意的是,HPT在VTUAD上实现了91.8%对89.8%的准确性。此外,HPT训练速度更快,产生的特征表示更接近完全微调的模型。总的来说,HPT平衡了参数节省和性能,为现有适配器提供了一种分布式感知的替代方案,并为资源受限环境中的可扩展迁移学习提供了一个有前途的方向。该代码是公开的:www.example.com。
摘要:Parameter-efficient transfer learning (PETL) methods adapt large artificial neural networks to downstream tasks without fine-tuning the entire model. However, existing additive methods, such as adapters, sometimes struggle to capture distributional shifts in intermediate feature embeddings. We propose a novel histogram-based parameter-efficient tuning (HPT) technique that captures the statistics of the target domain and modulates the embeddings. Experimental results on three downstream passive sonar datasets (ShipsEar, DeepShip, VTUAD) demonstrate that HPT outperforms conventional adapters. Notably, HPT achieves 91.8% vs. 89.8% accuracy on VTUAD. Furthermore, HPT trains faster and yields feature representations closer to those of fully fine-tuned models. Overall, HPT balances parameter savings and performance, providing a distribution-aware alternative to existing adapters and shows a promising direction for scalable transfer learning in resource-constrained environments. The code is publicly available: https://github.com/Advanced-Vision-and-Learning-Lab/HLAST_DeepShip_ParameterEfficient.
【3】 Improving Sound Source Localization with Joint Slot Attention on Image and Audio
标题: 通过图像和音频的联合缝隙关注来改善光源定位链接:https://arxiv.org/abs/2504.15118
备注:Accepted to CVPR 2025
摘要:声源定位(SSL)是在图像中定位声源的任务。由于缺乏本地化标签,SSL的事实标准是将图像和音频分别表示为单个嵌入向量,并使用它们通过对比学习来学习SSL。为此,以前的工作采样一个本地图像特征作为图像嵌入,并聚合所有本地音频特征以获得音频嵌入,由于输入中存在与实际目标无关的噪声和背景,这远非最佳。我们提出了一种新的SSL方法,通过对图像和音频的联合槽注意力来解决这个慢性问题。具体而言,两个槽竞争性地参加图像和音频特征,以将它们分解为目标和非目标表示,并且仅图像和音频的目标表示用于对比学习。此外,我们引入跨模态注意力匹配,以进一步对齐图像和音频的局部特征。我们的方法在SSL的三个公共基准测试中几乎在所有设置中都取得了最好的成绩,并且在跨模态检索方面大大优于所有先前的工作。
摘要:Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embedding vector each, and use them to learn SSL via contrastive learning. To this end, previous work samples one of local image features as the image embedding and aggregates all local audio features to obtain the audio embedding, which is far from optimal due to the presence of noise and background irrelevant to the actual target in the input. We present a novel SSL method that addresses this chronic issue by joint slot attention on image and audio. To be specific, two slots competitively attend image and audio features to decompose them into target and off-target representations, and only target representations of image and audio are used for contrastive learning. Also, we introduce cross-modal attention matching to further align local features of image and audio. Our method achieved the best in almost all settings on three public benchmarks for SSL, and substantially outperformed all the prior work in cross-modal retrieval.
【4】 Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music Modeling
标题: Aria-SYS:用于象征性音乐建模的钢琴SYS文件数据集链接:https://arxiv.org/abs/2504.15071
备注:None
摘要:我们介绍了一个广泛的新的数据集的录音文件,通过转录钢琴演奏的录音到其组成的音符。我们使用的数据管道是多阶段的,采用语言模型,根据元数据从互联网上自动抓取和评分音频记录,然后使用音频分类器进行修剪和分割。由此产生的数据集包含超过100万个不同的音频文件,包括大约10万小时的转录音频。我们对我们的技术进行深入分析,提供统计见解,并通过提取我们还提供的元数据标签来调查内容。数据集可在https://github.com/loubbrad/aria-midi获得。
摘要:We introduce an extensive new dataset of MIDI files, created by transcribing audio recordings of piano performances into their constituent notes. The data pipeline we use is multi-stage, employing a language model to autonomously crawl and score audio recordings from the internet based on their metadata, followed by a stage of pruning and segmentation using an audio classifier. The resulting dataset contains over one million distinct MIDI files, comprising roughly 100,000 hours of transcribed audio. We provide an in-depth analysis of our techniques, offering statistical insights, and investigate the content by extracting metadata tags, which we also provide. Dataset available at https://github.com/loubbrad/aria-midi.
【5】 SOLIDO: A Robust Watermarking Method for Speech Synthesis via Low-Rank Adaptation
标题: SOIDO:一种通过低等级自适应进行语音合成的鲁棒水印方法链接:https://arxiv.org/abs/2504.15035
摘要:语音生成模型的加速发展带来了安全问题,包括模型侵权和未经授权的内容滥用。虽然现有的生成式水印技术已经提出了相应的解决方案,大多数方法需要大量的计算开销和训练成本。此外,一些方法在处理可变长度输入时具有鲁棒性方面的限制。为了解决这些挑战,我们提出了\textsc{SOLIDO},一种新的生成式水印方法,通过语音扩散模型的低秩自适应(LoRA),将参数有效的微调与语音水印相结合。具体地,水印编码器转换水印以与扩散模型的输入对准。为了实现对变长输入的精确水印提取,设计了基于深度可分离卷积的水印解码器。为了进一步提高语音生成性能和水印提取能力,我们提出了一种语音驱动的轻量级微调策略,通过LoRA减少了计算开销。实验结果表明,该方法在2000 bps的大容量下仍能保证水印语音的高保真度。此外,对常见的个人和复合语音攻击,我们的SOLIDO实现了最大的平均提取准确率分别为99.20%和98.43%。它在抵抗时间拉伸攻击方面超过其他最先进的方法近23%。
摘要:The accelerated advancement of speech generative models has given rise to security issues, including model infringement and unauthorized abuse of content. Although existing generative watermarking techniques have proposed corresponding solutions, most methods require substantial computational overhead and training costs. In addition, some methods have limitations in robustness when handling variable-length inputs. To tackle these challenges, we propose \textsc{SOLIDO}, a novel generative watermarking method that integrates parameter-efficient fine-tuning with speech watermarking through low-rank adaptation (LoRA) for speech diffusion models. Concretely, the watermark encoder converts the watermark to align with the input of diffusion models. To achieve precise watermark extraction from variable-length inputs, the watermark decoder based on depthwise separable convolution is designed for watermark recovery. To further enhance speech generation performance and watermark extraction capability, we propose a speech-driven lightweight fine-tuning strategy, which reduces computational overhead through LoRA. Comprehensive experiments demonstrate that the proposed method ensures high-fidelity watermarked speech even at a large capacity of 2000 bps. Furthermore, against common individual and compound speech attacks, our SOLIDO achieves a maximum average extraction accuracy of 99.20\% and 98.43\%, respectively. It surpasses other state-of-the-art methods by nearly 23\% in resisting time-stretching attacks.
【6】 Protecting Your Voice: Temporal-aware Robust Watermarking
标题: 保护你的声音:时间感知的鲁棒水印链接:https://arxiv.org/abs/2504.14832
摘要:生成模型的快速发展导致了真假模糊声音的合成。为了消除模糊性,在合成语音的频域特征中嵌入水印已成为一种常见的方法。然而,通过选择频域实现的鲁棒性通常以牺牲细粒度语音特征为代价,导致保真度损失。最大限度地提高时域特征的综合学习,以提高保真度,同时保持鲁棒性,我们开创了一种用于保护语音和歌唱声音的\textbf {\underline{t}}emporal-aware \textbf {\underline{r}}ob\textbf{\underline{u}}st wat\textbf{\underline{e}}rmarking(\true {True})方法。
摘要:The rapid advancement of generative models has led to the synthesis of real-fake ambiguous voices. To erase the ambiguity, embedding watermarks into the frequency-domain features of synthesized voices has become a common routine. However, the robustness achieved by choosing the frequency domain often comes at the expense of fine-grained voice features, leading to a loss of fidelity. Maximizing the comprehensive learning of time-domain features to enhance fidelity while maintaining robustness, we pioneer a \textbf{\underline{t}}emporal-aware \textbf{\underline{r}}ob\textbf{\underline{u}}st wat\textbf{\underline{e}}rmarking (\emph{True}) method for protecting the speech and singing voice.
【7】 DiffVox: A Differentiable Model for Capturing and Analysing Professional Effects Distributions
标题: 迪夫Vox:一种用于捕捉和分析专业效果分布的差异化模型链接:https://arxiv.org/abs/2504.14735
备注:Submitted to DAFx 2025
摘要:本研究介绍了一种新的和可解释的模型,DiffVox,在音乐制作中匹配的声乐效果。DiffVox是“可微分声乐Fx”的缩写,它将参数均衡、动态范围控制、延迟和混响与有效的可微分实现集成在一起,以实现基于梯度的参数估计优化。从两个数据集中检索声乐曲目,包括来自MedleyDB的70首曲目和来自私人收藏的365首曲目。参数相关性的分析突出了效果和参数之间的强关系,例如高通和低架滤波器通常一起作用以形成低端,并且延迟时间与延迟信号的强度相关。主成分分析揭示了与麦克亚当斯音色维度的联系,其中最关键的成分调节了感知的宽敞度,而次要成分影响了光谱亮度。统计测试证实了参数分布的非高斯性质,突出了声音效果空间的复杂性。这些关于参数分布的初步研究结果为今后的声乐效果建模和自动混音研究奠定了基础。我们的源代码和数据集可以在https://github.com/SonyResearch/diffvox上访问。
摘要:This study introduces a novel and interpretable model, DiffVox, for matching vocal effects in music production. DiffVox, short for ``Differentiable Vocal Fx", integrates parametric equalisation, dynamic range control, delay, and reverb with efficient differentiable implementations to enable gradient-based optimisation for parameter estimation. Vocal presets are retrieved from two datasets, comprising 70 tracks from MedleyDB and 365 tracks from a private collection. Analysis of parameter correlations highlights strong relationships between effects and parameters, such as the high-pass and low-shelf filters often behaving together to shape the low end, and the delay time correlates with the intensity of the delayed signals. Principal component analysis reveals connections to McAdams' timbre dimensions, where the most crucial component modulates the perceived spaciousness while the secondary components influence spectral brightness. Statistical testing confirms the non-Gaussian nature of the parameter distribution, highlighting the complexity of the vocal effects space. These initial findings on the parameter distributions set the foundation for future research in vocal effects modelling and automatic mixing. Our source code and datasets are accessible at https://github.com/SonyResearch/diffvox.
【8】 DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
标题: DialogueAgents:一种基于Agent的混合式多方对话语音合成框架链接:https://arxiv.org/abs/2504.14482
备注:Accepted by ICME 2025. Dataset and code are publicly available: [https://github.com/uirlx/DialogueAgents](this https URL)
摘要:语音合成对于人机交互至关重要,可以实现自然和直观的通信。然而,现有的数据集涉及高建设成本,由于手动注释,并遭受有限的字符多样性,上下文场景和情感表达。为了解决这些问题,我们提出了DialogueAgents,一种新的混合基于代理的语音合成框架,它集成了三个专门的代理-脚本作家,语音合成器,和对话评论家-协作生成对话。该框架以不同的字符池为基础,迭代地细化对话脚本,并基于语音评论合成语音,提高了合成对话的情感表达力和非语言特征。使用DialogueAgent,我们贡献了MultiTalk,这是一个双语,多方,多轮语音对话数据集,涵盖了不同的主题。大量的实验证明了我们的框架的有效性和高质量的MultiTalk数据集。我们发布了数据集和代码www.example.com,以促进未来对高级语音合成模型和定制数据生成的研究。
摘要:Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited character diversity, contextual scenarios, and emotional expressiveness. To address these issues, we propose DialogueAgents, a novel hybrid agent-based speech synthesis framework, which integrates three specialized agents -- a script writer, a speech synthesizer, and a dialogue critic -- to collaboratively generate dialogues. Grounded in a diverse character pool, the framework iteratively refines dialogue scripts and synthesizes speech based on speech review, boosting emotional expressiveness and paralinguistic features of the synthesized dialogues. Using DialogueAgent, we contribute MultiTalk, a bilingual, multi-party, multi-turn speech dialogue dataset covering diverse topics. Extensive experiments demonstrate the effectiveness of our framework and the high quality of the MultiTalk dataset. We release the dataset and code https://github.com/uirlx/DialogueAgents to facilitate future research on advanced speech synthesis models and customized data generation.
【9】 Transformation of audio embeddings into interpretable, concept-based representations
标题: 将音频嵌入转换为可解释的、基于概念的表示链接:https://arxiv.org/abs/2504.14076
备注:Accepted to International Joint Conference on Neural Networks (IJCNN) 2025
摘要:音频神经网络的进步已经在下游音频任务上建立了最先进的结果。然而,这些模型的黑盒结构使得难以解释其内部音频表示中编码的信息。在这项工作中,我们通过利用CLAP(一种将音频和文本带入共享嵌入空间的对比学习模型)来探索从这些神经网络中提取的音频嵌入的语义可解释性。我们实现了一个事后的方法来转换CLAP嵌入到基于概念的,稀疏表示与语义解释。定性和定量的评估表明,基于概念的表示优于或匹配的性能的原始音频嵌入下游任务,同时提供可解释性。此外,我们证明,微调的概念为基础的表示可以进一步提高其性能的下游任务。最后,我们发布了三个音频特定的词汇,基于概念的音频嵌入的可解释性。
摘要:Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio representations. In this work, we explore the semantic interpretability of audio embeddings extracted from these neural networks by leveraging CLAP, a contrastive learning model that brings audio and text into a shared embedding space. We implement a post-hoc method to transform CLAP embeddings into concept-based, sparse representations with semantic interpretability. Qualitative and quantitative evaluations show that the concept-based representations outperform or match the performance of original audio embeddings on downstream tasks while providing interpretability. Additionally, we demonstrate that fine-tuning the concept-based representations can further improve their performance on downstream tasks. Lastly, we publish three audio-specific vocabularies for concept-based interpretability of audio embeddings.
【10】 Evaluating Human-AI Interaction via Usability, User Experience and Acceptance Measures for MMM-C: A Creative AI System for Music Composition
标题: 通过MMM-C的可用性、用户体验和接受指标评估人机交互:音乐创作的创意人工智能系统链接:https://arxiv.org/abs/2504.14071
备注:10 pages, 6 figures, 1 table, first published at the 32nd International Joint Conference on Artificial Intelligence (IJCAI 2023), Macao, China
摘要:随着人工智能(AI)的兴起,人们对包括音乐在内的各种艺术领域的人类-AI共同创作越来越感兴趣,因为AI驱动的系统经常能够生成人类竞争性的人工制品。现在,这种系统对音乐实践的影响正在研究中。我们报告了对用户采用多轨音乐机(MMM)作为音乐作曲家的共同创作AI工具的全面评估。为了做到这一点,我们将MMM集成到Cubase中,Cubase是Steinberg的一个流行的数字音频工作站(Digital Audio Workstation,简称MMM),通过生成一个名为MMM-Cubase(MMM-C)的“单参数”插件接口,它可以实现人类-AI的共同合成。我们贡献了一个方法组合作为一个3部分的混合方法研究测量可用性,用户体验和技术接受的系统在两组专家级作曲家:业余爱好者和专业人士。结果显示积极的可用性和接受度得分。用户报告了使用该系统的新颖性、惊喜和易用性的体验,以及在生成音乐时界面的可控性和可预测性的限制。调查结果表明,两个用户组之间没有显着差异。
摘要:With the rise of artificial intelligence (AI), there has been increasing interest in human-AI co-creation in a variety of artistic domains including music as AI-driven systems are frequently able to generate human-competitive artifacts. Now, the implications of such systems for musical practice are being investigated. We report on a thorough evaluation of the user adoption of the Multi-Track Music Machine (MMM) as a co-creative AI tool for music composers. To do this, we integrate MMM into Cubase, a popular Digital Audio Workstation (DAW) by Steinberg, by producing a "1-parameter" plugin interface named MMM-Cubase (MMM-C), which enables human-AI co-composition. We contribute a methodological assemblage as a 3-part mixed method study measuring usability, user experience and technology acceptance of the system across two groups of expert-level composers: hobbyists and professionals. Results show positive usability and acceptance scores. Users report experiences of novelty, surprise and ease of use from using the system, and limitations on controllability and predictability of the interface when generating music. Findings indicate no significant difference between the two user groups.
【11】 Calliope: An Online Generative Music System for Symbolic Multi-Track Composition
标题: Calliope:用于象征性多轨作曲的在线生成音乐系统链接:https://arxiv.org/abs/2504.14058
备注:5 pages, 5 figures, first published at the 13th International Conference on Computational Creativity (ICCC 2022), Bozen-Bolzano, Italy
摘要:随着近年来人工智能的兴起,其在创意领域的应用迅速增加,包括音乐。存在许多将机器学习方法应用于计算机辅助音乐创作(CAC)问题的系统。Calliope是一个Web应用程序,它帮助用户在符号域中执行各种多轨道合成任务。用户可以上传(乐器数字接口)多音轨文件,可视化和编辑多音轨,并使用多音轨音乐机(MMM)生成部分(通过填充条)或完整的多音轨内容。新的音频摘要的生成可以批量完成,并且可以与主动回放监听相结合,以增强辅助合成工作流程。用户可以导出生成的音频素材或直接将音频回放从系统流式传输到他们最喜欢的数字音频工作站(Digital Audio Workstation,简称DTS)。我们提出了一个演示的系统,它的功能,生成参数,并描述了它提供的共同创造的工作流程。
摘要:With the rise of artificial intelligence in recent years, there has been a rapid increase in its application towards creative domains, including music. There exist many systems built that apply machine learning approaches to the problem of computer-assisted music composition (CAC). Calliope is a web application that assists users in performing a variety of multi-track composition tasks in the symbolic domain. The user can upload (Musical Instrument Digital Interface) MIDI files, visualize and edit MIDI tracks, and generate partial (via bar in-filling) or complete multi-track content using the Multi-Track Music Machine (MMM). Generation of new MIDI excerpts can be done in batch and can be combined with active playback listening for an enhanced assisted-composition workflow. The user can export generated MIDI materials or directly stream MIDI playback from the system to their favorite Digital Audio Workstation (DAW). We present a demonstration of the system, its features, generative parameters and describe the co-creative workflows that it affords.
【12】 Apollo: An Interactive Environment for Generating Symbolic Musical Phrases using Corpus-based Style Imitation
标题: Apollo:使用基于数据库的风格模仿生成象征性音乐短语的交互环境链接:https://arxiv.org/abs/2504.14055
备注:7 pages, 5 figures, Published as a paper at the 7th International Workshop on Musical Metacreation (MUME 2019), UNC Charlotte, North Carolina
摘要:随着机器智能和网络技术的最新发展,正在探索新的生成音乐系统,以使用网络上的机器学习技术进行辅助作曲。这些系统是为各种任务而构建的,例如旋律、和声或节奏生成、音乐插值、延续和风格模仿。在本文中,我们介绍阿波罗,一个互动的音乐应用程序生成传统的西方音乐的符号短语,使用基于语料库的风格模仿技术。除了能够建设和管理的符号音乐语料库,该系统使音乐艺术家和研究人员有可能产生新的音乐短语的风格建议语料库。该系统可作为桌面应用程序使用。生成的符号音乐材料以数字格式编码,可以出于各种目的导出或流式传输,包括将它们用作音乐项目的种子材料。我们提出了系统的设计,实现细节,讨论和总结与未来的工作系统。
摘要:With the recent developments in machine intelligence and web technologies, new generative music systems are being explored for assisted composition using machine learning techniques on the web. Such systems are built for various tasks such as melodic, harmonic or rhythm generation, music interpolation, continuation and style imitation. In this paper, we introduce Apollo, an interactive music application for generating symbolic phrases of conventional western music using corpus-based style imitation techniques. In addition to enabling the construction and management of symbolic musical corpora, the system makes it possible for music artists and researchers to generate new musical phrases in the style of the proposed corpus. The system is available as a desktop application. The generated symbolic music materials, encoded in the MIDI format, can be exported or streamed for various purposes including using them as seed material for musical projects. We present the system design, implementation details, discuss and conclude with future work for the system.
【13】 Mixer Metaphors: audio interfaces for non-musical applications
标题: Mixer Metaphors:非音乐应用的音频接口链接:https://arxiv.org/abs/2504.13944
备注:9 Pages
摘要:NIME会议传统上专注于音乐和音乐表达的界面。在本文中,我们扭转这一传统的问题,可以开发的音乐接口成功地适用于非音乐应用程序?为了帮助回答这个问题,我们设计并开发了一种新设备,它使用从模拟合成器和音频混合中借用的接口隐喻来物理控制大型语言模型的无形方面。我们比较了两个版本的设备,有和没有音频启发的增强,与一组艺术家谁使用每个版本超过一个星期的时间。我们的研究结果表明,使用类似音频的控件可以对LLM进行更直接,更直接和更具体的控制,允许用户创造性地尝试和使用非混音器设备。我们的项目展示了跨感官隐喻如何在设计新的技术界面时支持创造性思维和具体实践。
摘要:The NIME conference traditionally focuses on interfaces for music and musical expression. In this paper we reverse this tradition to ask, can interfaces developed for music be successfully appropriated to non-musical applications? To help answer this question we designed and developed a new device, which uses interface metaphors borrowed from analogue synthesisers and audio mixing to physically control the intangible aspects of a Large Language Model. We compared two versions of the device, with and without the audio-inspired augmentations, with a group of artists who used each version over a one week period. Our results show that the use of audio-like controls afforded more immediate, direct and embodied control over the LLM, allowing users to creatively experiment and play with the device over its non-mixer counterpart. Our project demonstrates how cross-sensory metaphors can support creative thinking and embodied practice when designing new technological interfaces.
【14】 OmniAudio: Generating Spatial Audio from 360-Degree Video
标题: OmniAudio:从360度视频生成空间音频链接:https://arxiv.org/abs/2504.14906
备注:Work in Progress
摘要:传统的视频到音频生成技术主要集中在视场(FoV)视频和非空间音频上,通常缺少在3D环境中准确表示声源所必需的空间线索。为了解决这一限制,我们引入了一个新的任务,360 V2 SA,从360度视频中生成空间音频,特别是产生一阶高保真度立体声(FOA)音频-一种用于表示3D空间音频的标准格式,可以捕获声音方向性并实现逼真的3D音频再现。首先,我们创建了一个新的数据集Sphere 360,这是一个专门为这项任务定制的数据集,它是从真实世界的数据中挑选出来的。我们还设计了一个高效的半自动管道,用于收集和清理配对的视频-音频数据。为了从360度视频中生成空间音频,我们提出了一个新的框架OmniAudio,它利用了空间音频数据(FOA格式)和大规模非空间数据的自监督预训练。此外,OmniAudio还具有双分支框架,该框架利用全景和FoV视频输入来从360度视频中捕获全面的本地和全球信息。实验结果表明,OmniAudio在Sphere 360上的客观和主观指标上都达到了最先进的性能。代码和数据集将在https://github.com/liuhuadai/OmniAudio上发布。演示页面可以在https://OmniAudio-360V2SA.github.io上找到。
摘要:Traditional video-to-audio generation techniques primarily focus on field-of-view (FoV) video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and FoV video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets will be released at https://github.com/liuhuadai/OmniAudio. The demo page is available at https://OmniAudio-360V2SA.github.io.
【15】 Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
标题: 使用神经声学场和检索增强预训练进行数据增强链接:https://arxiv.org/abs/2504.14409
备注:Presented at ICASSP 2025 GenDA Workshop
摘要:本报告详细介绍了MERL的房间脉冲响应(RIR)估计系统,该系统已提交给ICASSP 2025的生成数据增强研讨会,用于增强RIR数据(任务1)和改进扬声器距离估计(任务2)。我们首先在外部大规模数据集上预训练由房间几何形状调节的神经声场,其中提供了成对的RIR和几何形状。然后通过使用登记数据使神经声场适应每个目标房间,其中我们根据可用性利用所提供的房间几何形状或从外部数据集检索的几何形状。最后,我们预测任务1中指定的每对源和接收器位置的RIR,并使用这些RIR来训练任务2中的说话人距离估计模型。
摘要:This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimation (Task 2). We first pre-train a neural acoustic field conditioned by room geometry on an external large-scale dataset in which pairs of RIRs and the geometries are provided. The neural acoustic field is then adapted to each target room by using the enrollment data, where we leverage either the provided room geometries or geometries retrieved from the external dataset, depending on availability. Lastly, we predict the RIRs for each pair of source and receiver locations specified by Task 1, and use these RIRs to train the speaker distance estimation model in Task 2.
【1】 On feature representations for marmoset vocal communication analysis
标题: 绒猴声乐传播分析的特征表示链接:https://arxiv.org/abs/2504.14981
备注:None
摘要:对绒猴(Callithrix jacchus)发声的声学分析经常被用来理解人类语言的进化起源。目前,该分析主要以手动或半手动方式进行。因此,需要开发自动呼叫分析方法。在这方面,研究一直局限于开发数据量小或针对具体情况的分析方法。此外,缺乏关于什么类型的信息与不同呼叫分析任务相关的先验知识。为了解决这些问题,作为第一步,本文探讨了不同的特征表示方法,即基于HCTSA的手工特征Catch22,从人类语音训练的神经网络中提取的基于预训练的自监督学习(SSL)的特征,以及用于呼叫类型分类,呼叫者识别和呼叫者性别识别的端到端声学建模。通过对三种不同的绒猴呼叫数据集的调查,我们证明了基于SSL的特征表示和端到端声学建模往往会导致比Catch22功能更好的系统,用于呼叫类型和呼叫者分类。此外,我们还强调了信号带宽对所获得的任务性能的影响。
摘要:The acoustic analysis of marmoset (Callithrix jacchus) vocalizations is often used to understand the evolutionary origins of human language. Currently, the analysis is largely carried out in a manual or semi-manual manner. Thus, there is a need to develop automatic call analysis methods. In that direction, research has been limited to the development of analysis methods with small amounts of data or for specific scenarios. Furthermore, there is lack of prior knowledge about what type of information is relevant for different call analysis tasks. To address these issues, as a first step, this paper explores different feature representation methods, namely, HCTSA-based hand-crafted features Catch22, pre-trained self supervised learning (SSL) based features extracted from neural networks trained on human speech and end-to-end acoustic modeling for call-type classification, caller identification and caller sex identification. Through an investigation on three different marmoset call datasets, we demonstrate that SSL-based feature representations and end-to-end acoustic modeling tend to lead to better systems than Catch22 features for call-type and caller classification. Furthermore, we also highlight the impact of signal bandwidth on the obtained task performances.
【2】 StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models
标题: StableQuant:语音基础模型的层自适应训练后量化链接:https://arxiv.org/abs/2504.14915
备注:Accepted at ICASSP 2025
摘要:在本文中,我们提出了一种新的自适应后训练量化(PTQ)算法,广泛使用的语音基础模型(SFM)的StableQuant。虽然PTQ已经成功地用于压缩大型语言模型(LLM),因为它能够绕过额外的微调,但直接将这些技术应用于SFM可能不会产生最佳结果,因为SFM使用不同的网络架构进行特征提取。无论网络架构类型如何,StableQuant都能展示出最佳的量化性能,因为它通过分析尺度分布和整体性能来自适应地确定每层的量化范围。我们评估我们的算法上的两个SFM,HuBERT和wav2vec2.0,自动语音识别(ASR)的任务,并实现优越的性能相比,传统的PTQ方法。StableQuant成功地将SFM模型的大小减少到四分之一,并将推理速度提高一倍,同时将8位量化的字错误率(WER)性能下降限制在0.3%以下。
摘要:In this paper, we propose StableQuant, a novel adaptive post-training quantization (PTQ) algorithm for widely used speech foundation models (SFMs). While PTQ has been successfully employed for compressing large language models (LLMs) due to its ability to bypass additional fine-tuning, directly applying these techniques to SFMs may not yield optimal results, as SFMs utilize distinct network architecture for feature extraction. StableQuant demonstrates optimal quantization performance regardless of the network architecture type, as it adaptively determines the quantization range for each layer by analyzing both the scale distributions and overall performance. We evaluate our algorithm on two SFMs, HuBERT and wav2vec2.0, for an automatic speech recognition (ASR) task, and achieve superior performance compared to traditional PTQ methods. StableQuant successfully reduces the sizes of SFM models to a quarter and doubles the inference speed while limiting the word error rate (WER) performance drop to less than 0.3% with 8-bit quantization.
【3】 OmniAudio: Generating Spatial Audio from 360-Degree Video
标题: OmniAudio:从360度视频生成空间音频链接:https://arxiv.org/abs/2504.14906
备注:Work in Progress
摘要:传统的视频到音频生成技术主要集中在视场(FoV)视频和非空间音频上,通常缺少在3D环境中准确表示声源所必需的空间线索。为了解决这一限制,我们引入了一个新的任务,360 V2 SA,从360度视频中生成空间音频,特别是产生一阶高保真度立体声(FOA)音频-一种用于表示3D空间音频的标准格式,可以捕获声音方向性并实现逼真的3D音频再现。首先,我们创建了一个新的数据集Sphere 360,这是一个专门为这项任务定制的数据集,它是从真实世界的数据中挑选出来的。我们还设计了一个高效的半自动管道,用于收集和清理配对的视频-音频数据。为了从360度视频中生成空间音频,我们提出了一个新的框架OmniAudio,它利用了空间音频数据(FOA格式)和大规模非空间数据的自监督预训练。此外,OmniAudio具有双分支框架,利用全景和FoV视频输入从360度视频中捕获全面的本地和全局信息。实验结果表明,OmniAudio在Sphere 360上的客观和主观指标上都达到了最先进的性能。代码和数据集将在https://github.com/liuhuadai/OmniAudio上发布。演示页面可以在https://OmniAudio-360V2SA.github.io上找到。
摘要:Traditional video-to-audio generation techniques primarily focus on field-of-view (FoV) video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and FoV video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets will be released at https://github.com/liuhuadai/OmniAudio. The demo page is available at https://OmniAudio-360V2SA.github.io.
【4】 Quantitative Measures for Passive Sonar Texture Analysis
标题: 被动声纳纹理分析的量化措施链接:https://arxiv.org/abs/2504.14843
备注:5 pages, 2 figures
摘要:被动声纳信号包含复杂的特征,通常由环境噪声、船舶机械和传播效应引起。虽然卷积神经网络(CNN)在被动声纳分类任务中表现良好,但它们可能会难以应对数据中出现的统计变化。为了研究这一限制,合成的水下声学数据集生成的振幅和周期变化为中心。提出了两个度量的统计和结构纹理的背景下,被动声纳量化和验证这些特性。这些措施被应用到现实世界的被动声纳数据集,以评估信号中的纹理信息和相关的模型的性能。结果表明,CNN在统计纹理信号上表现不佳,但结合显式统计纹理建模会产生一致的改进。这些发现突出了量化的纹理信息的重要性被动声纳分类。
摘要:Passive sonar signals contain complex characteristics often arising from environmental noise, vessel machinery, and propagation effects. While convolutional neural networks (CNNs) perform well on passive sonar classification tasks, they can struggle with statistical variations that occur in the data. To investigate this limitation, synthetic underwater acoustic datasets are generated that centered on amplitude and period variations. Two metrics are proposed to quantify and validate these characteristics in the context of statistical and structural texture for passive sonar. These measures are applied to real-world passive sonar datasets to assess texture information in the signals and correlate the performances of the models. Results show that CNNs underperform on statistically textured signals, but incorporating explicit statistical texture modeling yields consistent improvements. These findings highlight the importance of quantifying texture information for passive sonar classification.
【5】 DNN based HRIRs Identification with a Continuously Rotating Speaker Array
标题: 基于DNN的连续旋转扬声器阵列HRIR识别链接:https://arxiv.org/abs/2504.14817
摘要:None
摘要:Conventional static measurement of head-related impulse responses (HRIRs) is time-consuming due to the need for repositioning a speaker array for each azimuth angle. Dynamic approaches using analytical models with a continuously rotating speaker array have been proposed, but their accuracy is significantly reduced at high rotational speeds. To address this limitation, we propose a DNN-based HRIRs identification using sequence-to-sequence learning. The proposed DNN model incorporates fully connected (FC) networks to effectively capture HRIR transitions and includes reset and update gates to identify HRIRs over a whole sequence. The model updates the HRIRs vector coefficients based on the gradient of the instantaneous square error (ISE). Additionally, we introduce a learnable normalization process based on the speaker excitation signals to stabilize the gradient scale of ISE across time. A training scheme, referred to as whole-sequence updating and optimization scheme, is also introduced to prevent overfitting. We evaluated the proposed method through simulations and experiments. Simulation results using the FABIAN database show that the proposed method outperforms previous analytic models, achieving over 7 dB improvement in normalized misalignment (NM) and maintaining log spectral distortion (LSD) below 2 dB at a rotational speed of 45{\deg}/s. Experimental results with a custom-built speaker array confirm that the proposed method successfully preserved accurate sound localization cues, consistent with those from static measurement. Source code is available at https://github.com/byko0810/DNN-based-HRIRs-identification
【6】 Predicting speech intelligibility in older adults using the Gammachirp Envelope Similarity Index, GESI
标题: 使用Gammachirp信封相似性指数(GESI)预测老年人的语音清晰度链接:https://arxiv.org/abs/2504.14437
备注:This manuscript was submitted to Speech Communication on 20 April 2025
摘要:我们提出了一个客观的可懂度测量(OIM),称为Gammachirp包络相似性指数(GESI),可以预测老年人的语音可懂度(SI)。GESI是一个自下而上的模型,基于从外围到中央听觉系统的心理声学知识,不需要训练数据。它使用伽马线性滤波器组(GCFB)、调制滤波器组和扩展余弦相似性度量来计算单个SI度量。它不仅考虑了听力图中表示的听力水平,而且还考虑了由时间调制传递函数(TMTF)捕获的时间处理特性。为了评估性能,SI实验进行了各种听力水平的老年人使用语音噪声与理想的语音增强熟悉度控制的话。将预测性能与HASPIw2进行比较,HASPIw2是为关键字SI预测而开发的。结果表明,GESI预测主观SI评分的准确性高于HASPIw2。将TMTF引入GESI算法的效果并不显著,这表明需要更多的研究来了解如何将时间响应特性引入OIM。
摘要:We propose an objective intelligibility measure (OIM), called the Gammachirp Envelope Similarity Index (GESI), that can predict speech intelligibility (SI) in older adults. GESI is a bottom-up model based on psychoacoustic knowledge from the peripheral to the central auditory system and requires no training data. It computes the single SI metric using the gammachirp filterbank (GCFB), the modulation filterbank, and the extended cosine similarity measure. It takes into account not only the hearing level represented in the audiogram, but also the temporal processing characteristics captured by the temporal modulation transfer function (TMTF). To evaluate performance, SI experiments were conducted with older adults of various hearing levels using speech-in-noise with ideal speech enhancement on familiarity-controlled words. The prediction performance was compared with HASPIw2, which was developed for keyword SI prediction. The results showed that GESI predicted the subjective SI scores more accurately than HASPIw2. The effect of introducing TMTF into the GESI algorithm was not significant, indicating that more research is needed to know how to introduce temporal response characteristics into the OIM.
【7】 Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
标题: 使用神经声学场和检索增强预训练进行数据增强链接:https://arxiv.org/abs/2504.14409
备注:Presented at ICASSP 2025 GenDA Workshop
摘要:本报告详细介绍了MERL的房间脉冲响应(RIR)估计系统,该系统已提交给ICASSP 2025的生成数据增强研讨会,用于增强RIR数据(任务1)和改进扬声器距离估计(任务2)。我们首先在外部大规模数据集上预训练由房间几何形状调节的神经声场,其中提供了成对的RIR和几何形状。然后通过使用登记数据使神经声场适应每个目标房间,其中我们根据可用性利用所提供的房间几何形状或从外部数据集检索的几何形状。最后,我们预测任务1中指定的每对源和接收器位置的RIR,并使用这些RIR来训练任务2中的说话人距离估计模型。
摘要:This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimation (Task 2). We first pre-train a neural acoustic field conditioned by room geometry on an external large-scale dataset in which pairs of RIRs and the geometries are provided. The neural acoustic field is then adapted to each target room by using the enrollment data, where we leverage either the provided room geometries or geometries retrieved from the external dataset, depending on availability. Lastly, we predict the RIRs for each pair of source and receiver locations specified by Task 1, and use these RIRs to train the speaker distance estimation model in Task 2.
【8】 The First VoicePrivacy Attacker Challenge
标题: 第一个语音隐私攻击者挑战链接:https://arxiv.org/abs/2504.14183
备注:Published in: ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
摘要:第一个语音隐私攻击者挑战赛是ICASSP 2025 SP大挑战赛,重点是针对提交给语音隐私2024挑战赛的一组语音匿名系统评估攻击者系统。提供了培训、开发和评估数据集以及基线攻击者。参与者以自动说话人验证系统的形式开发了他们的攻击系统,并提交了他们对开发和评估数据的评分。最好的攻击者系统相对于w.r.t.将相等错误率(EER)降低了25-44%。基线。
摘要:The First VoicePrivacy Attacker Challenge is an ICASSP 2025 SP Grand Challenge which focuses on evaluating attacker systems against a set of voice anonymization systems submitted to the VoicePrivacy 2024 Challenge. Training, development, and evaluation datasets were provided along with a baseline attacker. Participants developed their attacker systems in the form of automatic speaker verification systems and submitted their scores on the development and evaluation data. The best attacker systems reduced the equal error rate (EER) by 25-44% relative w.r.t. the baseline.
【9】 Testing LLMs' Capabilities in Annotating Translations Based on an Error Typology Designed for LSP Translation: First Experiments with ChatGPT
标题: 测试LLM基于为SPL翻译设计的错误类型注释翻译的能力:ChatGPT的首次实验链接:https://arxiv.org/abs/2504.15052
备注:Accepted for publication in the proceedings of MT Summit 2025
摘要:本研究调查了大型语言模型(LLM),特别是ChatGPT,在基于错误类型注释MT输出的能力。与以前主要关注一般语言的工作相比,我们探索了ChatGPT识别和分类专业翻译中错误的能力。通过测试两种不同的提示并基于自定义的错误类型,我们将ChatGPT注释与DeepL和ChatGPT本身产生的翻译的人类专家评估进行了比较。结果表明,对于DeepL生成的翻译,召回率和精确率都相当高。然而,错误分类的准确程度取决于提示的特定功能及其详细程度,ChatGPT在详细提示中表现得非常好。在评估自己的翻译时,ChatGPT的结果明显较差,显示了自我评估的局限性。这些结果突出了LLM用于翻译评估的潜力和局限性,特别是在专业领域。我们的实验为开源LLM的未来研究铺平了道路,这可以产生可比甚至更高质量的注释。在未来,我们的目标还包括在翻译培训的背景下测试这种自动化评估的实际有效性,特别是通过优化教师的人工评估过程,以及探索LLM注释对学生后期编辑和翻译学习的影响。
摘要:This study investigates the capabilities of large language models (LLMs), specifically ChatGPT, in annotating MT outputs based on an error typology. In contrast to previous work focusing mainly on general language, we explore ChatGPT's ability to identify and categorise errors in specialised translations. By testing two different prompts and based on a customised error typology, we compare ChatGPT annotations with human expert evaluations of translations produced by DeepL and ChatGPT itself. The results show that, for translations generated by DeepL, recall and precision are quite high. However, the degree of accuracy in error categorisation depends on the prompt's specific features and its level of detail, ChatGPT performing very well with a detailed prompt. When evaluating its own translations, ChatGPT achieves significantly poorer results, revealing limitations with self-assessment. These results highlight both the potential and the limitations of LLMs for translation evaluation, particularly in specialised domains. Our experiments pave the way for future research on open-source LLMs, which could produce annotations of comparable or even higher quality. In the future, we also aim to test the practical effectiveness of this automated evaluation in the context of translation training, particularly by optimising the process of human evaluation by teachers and by exploring the impact of annotations by LLMs on students' post-editing and translation learning.
【10】 DiffVox: A Differentiable Model for Capturing and Analysing Professional Effects Distributions
标题: 迪夫Vox:一种用于捕捉和分析专业效果分布的差异化模型链接:https://arxiv.org/abs/2504.14735
备注:Submitted to DAFx 2025
摘要:本研究介绍了一种新颖且可解释的模型DiffVox,用于匹配音乐制作中的声乐效果。DiffVox是“可微分声乐Fx”的缩写,它将参数均衡、动态范围控制、延迟和混响与有效的可微分实现集成在一起,以实现基于梯度的参数估计优化。从两个数据集中检索声乐曲目,包括来自MedleyDB的70首曲目和来自私人收藏的365首曲目。参数相关性的分析突出了效果和参数之间的强关系,例如高通和低架滤波器通常一起作用以形成低端,并且延迟时间与延迟信号的强度相关。主成分分析揭示了与麦克亚当斯音色维度的联系,其中最关键的成分调节了感知的宽敞度,而次要成分影响了光谱亮度。统计测试证实了参数分布的非高斯性质,突出了声音效果空间的复杂性。这些关于参数分布的初步研究结果为今后的声乐效果建模和自动混音研究奠定了基础。我们的源代码和数据集可以在https://github.com/SonyResearch/diffvox上访问。
摘要:This study introduces a novel and interpretable model, DiffVox, for matching vocal effects in music production. DiffVox, short for ``Differentiable Vocal Fx", integrates parametric equalisation, dynamic range control, delay, and reverb with efficient differentiable implementations to enable gradient-based optimisation for parameter estimation. Vocal presets are retrieved from two datasets, comprising 70 tracks from MedleyDB and 365 tracks from a private collection. Analysis of parameter correlations highlights strong relationships between effects and parameters, such as the high-pass and low-shelf filters often behaving together to shape the low end, and the delay time correlates with the intensity of the delayed signals. Principal component analysis reveals connections to McAdams' timbre dimensions, where the most crucial component modulates the perceived spaciousness while the secondary components influence spectral brightness. Statistical testing confirms the non-Gaussian nature of the parameter distribution, highlighting the complexity of the vocal effects space. These initial findings on the parameter distributions set the foundation for future research in vocal effects modelling and automatic mixing. Our source code and datasets are accessible at https://github.com/SonyResearch/diffvox.
【11】 Transformation of audio embeddings into interpretable, concept-based representations
标题: 将音频嵌入转换为可解释的、基于概念的表示链接:https://arxiv.org/abs/2504.14076
备注:Accepted to International Joint Conference on Neural Networks (IJCNN) 2025
摘要:音频神经网络的进步已经在下游音频任务上建立了最先进的结果。然而,这些模型的黑盒结构使得难以解释其内部音频表示中编码的信息。在这项工作中,我们通过利用CLAP(一种将音频和文本带入共享嵌入空间的对比学习模型)来探索从这些神经网络中提取的音频嵌入的语义可解释性。我们实现了一种事后方法来将CLAP嵌入转换为具有语义可解释性的基于概念的稀疏表示。定性和定量的评估表明,基于概念的表示优于或匹配的性能的原始音频嵌入下游任务,同时提供可解释性。此外,我们证明,微调的概念为基础的表示可以进一步提高其性能的下游任务。最后,我们发布了三个音频特定的词汇,基于概念的音频嵌入的可解释性。
摘要:Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio representations. In this work, we explore the semantic interpretability of audio embeddings extracted from these neural networks by leveraging CLAP, a contrastive learning model that brings audio and text into a shared embedding space. We implement a post-hoc method to transform CLAP embeddings into concept-based, sparse representations with semantic interpretability. Qualitative and quantitative evaluations show that the concept-based representations outperform or match the performance of original audio embeddings on downstream tasks while providing interpretability. Additionally, we demonstrate that fine-tuning the concept-based representations can further improve their performance on downstream tasks. Lastly, we publish three audio-specific vocabularies for concept-based interpretability of audio embeddings.
机器翻译由腾讯交互翻译提供,仅供参考
