微信公众号:arXiv_Daily
cs.SD语音
【1】Music Arena: Live Evaluation for Text-to-Music
标题:音乐竞技场:文本转音乐现场评估
链接:http://arxiv.org/pdf/2507.20900v1
摘要:我们提出了音乐竞技场,一个开放的平台,可扩展的人类偏好评估的文本到音乐(TTM)模型。通过听力研究征求人类偏好是TTM评估的黄金标准,但这些研究进行起来很昂贵,而且很难比较,因为不同系统的研究方案可能不同。此外,人类偏好可能有助于研究人员调整他们的TTM系统或改进自动评估指标,但目前还不存在开放和可更新的偏好来源。我们的目标是通过为TTM提供实时评估来填补这些空白。在Music Arena中,真实世界的用户输入他们选择的文本提示,并比较两个TTM系统的输出,他们的偏好被用来编制排行榜。虽然Music Arena遵循了其他AI领域的最新评估趋势,但我们还设计了针对音乐的关键功能:基于LLM的路由系统,用于导航TTM系统的异构类型签名,以及收集 详细 偏好,包括收听数据和自然语言反馈。我们还提出了一个滚动的数据发布政策,保证用户隐私,提供可再生的偏好数据来源,并提高平台的透明度。通过其标准化的评估协议、透明的数据访问策略和特定于音乐的功能,Music Arena不仅解决了TTM生态系统中的关键挑战,还展示了如何根据特定AI领域的独特特征对实时评估进行深思熟虑的调整。 Music Arena网站:https: music-arena.org
摘要:
We present Music Arena, an open platform for scalable human preference evaluation of text-to-music (TTM) models. Soliciting human preferences via listening studies is the gold standard for evaluation in TTM, but these studies are expensive to conduct and difficult to compare, as study protocols may differ across systems. Moreover, human preferences might help researchers align their TTM systems or improve automatic evaluation metrics, but an open and renewable source of preferences does not currently exist. We aim to fill these gaps by offering live evaluation for TTM. In Music Arena, real-world users input text prompts of their choosing and compare outputs from two TTM systems, and their preferences are used to compile a leaderboard. While Music Arena follows recent evaluation trends in other AI domains, we also design it with key features tailored to music: an LLM-based routing system to navigate the heterogeneous type signatures of TTM systems, and the collection of detailed preferences including listening data and natural language feedback. We also propose a rolling data release policy with user privacy guarantees, providing a renewable source of preference data and increasing platform transparency. Through its standardized evaluation protocol, transparent data access policies, and music-specific features, Music Arena not only addresses key challenges in the TTM ecosystem but also demonstrates how live evaluation can be thoughtfully adapted to unique characteristics of specific AI domains. Music Arena is available at: https: music-arena.org
【2】JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
标题:JAM:一款基于流的微型歌曲生成器,具有细粒度可控性和美观一致性
链接:http://arxiv.org/pdf/2507.20880v1
备注:his https URL
摘要:最近,扩散和流匹配模型彻底改变了自动文本到音频生成。这些模型越来越能够生成捕获语音和声学事件的高质量和忠实的音频输出。然而,在主要涉及音乐和歌曲的创造性音频生成方面仍有很大的改进空间。最近的开放式歌词到歌曲模型,如DiffRhythm、ACE-Step和LeVo,已经为娱乐用途的自动歌曲生成设定了可接受的标准。然而,这些模型缺乏音乐家在其工作流程中经常需要的细粒度单词级可控性。据我们所知,我们基于流匹配的JAM是第一次在歌曲生成中赋予单词级计时和持续时间控制的努力,允许细粒度的声乐控制。为了提高生成的歌曲的质量以更好地与人类偏好保持一致,我们通过直接偏好优化实现了美学对齐,该优化使用合成数据集迭代地优化模型,消除了手动数据注释的需要。此外,我们的目标是通过我们的公共评估数据集JAME标准化这些歌词到歌曲模型的评估。我们表明,JAM优于现有的模型在音乐的特定属性。
摘要:Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech and acoustic events. However, there is still much room for improvement in creative audio generation that primarily involves music and songs. Recent open lyrics-to-song models, such as, DiffRhythm, ACE-Step, and LeVo, have set an acceptable standard in automatic song generation for recreational use. However, these models lack fine-grained word-level controllability often desired by musicians in their workflows. To the best of our knowledge, our flow-matching-based JAM is the first effort toward endowing word-level timing and duration control in song generation, allowing fine-grained vocal control. To enhance the quality of generated songs to better align with human preferences, we implement aesthetic alignment through Direct Preference Optimization, which iteratively refines the model using a synthetic dataset, eliminating the need or manual data annotations. Furthermore, we aim to standardize the evaluation of such lyrics-to-song models through our public evaluation dataset JAME. We show that JAM outperforms the existing models in terms of the music-specific attributes.
【3】Learning Neural Vocoder from Range-Null Space Decomposition
标题:从范围-范围空间分解学习神经声码器
链接:http://arxiv.org/pdf/2507.20731v1
备注:10 pages, 7 figures, IJCAI2025
摘要:尽管近年来神经声码器发展迅速,但它们通常受到一些内在的挑战,如不透明的建模,参数性能权衡。在这项研究中,我们提出了一个创新的时频(T-F)域为基础的神经声码器,以解决上述挑战。具体而言,我们将经典的信号距离零分解(RND)理论与声码器任务相结合,将目标谱图的重构分解为距离空间与零空间的叠加,其中距离空间的重构是通过从原始梅尔尺度域到目标线性尺度域的线性域移位实现的,并且后者经由可学习网络被实例化以用于进一步的光谱细节生成。因此,我们提出了一种新的双路径框架,其中频谱是分层编码 解码,和交叉和窄带模块精心设计的高效的子带和顺序建模。在LJSpeech和LibriTTS基准上进行了全面的实验。定量和定性的结果表明,同时享有轻量级的网络参数,所提出的方法产生的现有的先进方法之间的最先进的性能。我们的代码和预训练的模型权重可以在https: github.com Andong-Li-speech RNDVoC上找到。
摘要:Despite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned challenges. To be specific, we bridge the connection between the classical signal range-null decomposition (RND) theory and vocoder task, and the reconstruction of target spectrogram can be decomposed into the superimposition between the range-space and null-space, where the former is enabled by a linear domain shift from the original mel-scale domain to the target linear-scale domain, and the latter is instantiated via a learnable network for further spectral detail generation. Accordingly, we propose a novel dual-path framework, where the spectrum is hierarchically encoded decoded, and the cross- and narrow-band modules are elaborately devised for efficient sub-band and sequential modeling. Comprehensive experiments are conducted on the LJSpeech and LibriTTS benchmarks. Quantitative and qualitative results show that while enjoying lightweight network parameters, the proposed approach yields state-of-the-art performance among existing advanced methods. Our code and the pretrained model weights are available at https: github.com Andong-Li-speech RNDVoC.
【4】Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
标题:具有多个时变条件的可控视频到音乐生成
链接:https://arxiv.org/abs/2507.20627
备注:Accepted by the 33rd ACM International Conference on Multimedia (ACMMM 2025). The project page is available at this https URL
摘要:音乐增强了视频叙事和情感,推动了对自动视频到音乐(V2M)生成的需求。然而,现有的V2M方法仅依赖于视觉特征或补充文本输入,以黑盒方式生成音乐,通常无法满足用户期望。为了解决这一挑战,我们提出了一种新的多条件引导的V2M生成框架,该框架结合了多个时变条件,以增强对音乐生成的控制。我们的方法使用两阶段的训练策略,使学习V2M的基本原理和视听时间同步,同时满足用户的多条件控制的需求。在第一阶段,我们引入了一个细粒度的特征选择模块和一个渐进的时间对齐注意机制,以确保灵活的特征对齐。在第二阶段,我们开发了一个动态条件融合模块和一个控制引导的解码器模块,以整合多个条件,准确地指导音乐创作过程。大量的实验表明,我们的方法在主观和客观评估方面都优于现有的V2M管道,显着增强了控制和与用户期望的一致性。
摘要:Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.
【5】Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
标题:音效链的订单感知分类的双曲嵌入
链接:http://arxiv.org/pdf/2507.20624v1
备注:7 pages, 3 figures, accepted for the 28th International Conference on Digital Audio Effects (DAFx25)
摘要:音频效果(AFX)是音乐制作中必不可少的工具,经常用于塑造音色和动态。链中AFX的顺序在确定最终声音中起着至关重要的作用,特别是当非线性(例如,失真)或时变(例如,合唱)处理器。尽管它的重要性,大多数AFX相关的研究主要集中在估计效果类型和它们的参数从湿信号。为了解决这个差距,我们制定AFX链识别的任务,联合估计AFX类型和他们的顺序从湿信号。我们提出了一种基于神经网络的方法,将湿信号嵌入到双曲空间中,并对其AFX链进行分类。双曲空间由于其指数扩展特性,可以比欧氏空间更有效地表示树结构数据。由于AFX链可以表示为树,AFX作为节点和边编码效果顺序,因此双曲空间非常适合建模有序AFX组合的指数增长和非交换性质,其中效果顺序的变化可能导致不同的最终声音。使用吉他声音的实验表明,与适当的曲率,所提出的方法优于其欧几里德对应。基于AFX类型和链长的进一步分析突出了所提出的方法在捕获AFX阶数方面的有效性。
摘要:Audio effects (AFXs) are essential tools in music production, frequently applied in chains to shape timbre and dynamics. The order of AFXs in a chain plays a crucial role in determining the final sound, particularly when non-linear (e.g., distortion) or time-variant (e.g., chorus) processors are involved. Despite its importance, most AFX-related studies have primarily focused on estimating effect types and their parameters from a wet signal. To address this gap, we formulate AFX chain recognition as the task of jointly estimating AFX types and their order from a wet signal. We propose a neural-network-based method that embeds wet signals into a hyperbolic space and classifies their AFX chains. Hyperbolic space can represent tree-structured data more efficiently than Euclidean space due to its exponential expansion property. Since AFX chains can be represented as trees, with AFXs as nodes and edges encoding effect order, hyperbolic space is well-suited for modeling the exponentially growing and non-commutative nature of ordered AFX combinations, where changes in effect order can result in different final sounds. Experiments using guitar sounds demonstrate that, with an appropriate curvature, the proposed method outperforms its Euclidean counterpart. Further analysis based on AFX type and chain length highlights the effectiveness of the proposed method in capturing AFX order.
【6】Sound Safeguarding for Acoustic Measurement Using Any Sounds: Tools and Applications
标题:使用任何声音进行声学测量的声音保护:工具和应用
链接:https://arxiv.org/abs/2507.20485
备注:2 pages, 2 figures, IEEE GCCE 2025 Demo session, Accepted
摘要:我们展示了基于“声音保护”方法开发的工具和应用程序,该方法使任何声音都能用于声学测量。我们开发了用于准备、交互式和实时测量以及报告生成的工具。在开发过程中,我们根据实际应用情况对该方法进行了扩展和改进。我们已经开源了这些工具,并鼓励潜在用户使用它们来改善他们的声学环境。
摘要:We demonstrate tools and applications developed based on the method of "sound safeguarding," which enables any sound to be used for acoustic measurements. We developed tools for preparation, interactive and real-time measurement, and report generation. We extended and modified the method during its development based on its application in various practical situations. We have open-sourced these tools and encourage prospective users to use them to improve their acoustic environments.
【7】Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection
标题:两种观点,一种真相:光谱和自监督特征融合用于稳健的语音深度伪造检测
链接:https://arxiv.org/abs/2507.20417
备注:ACCEPTED WASPAA 2025
摘要:合成语音的最新进展使音频deepfake变得越来越逼真,带来了重大的安全风险。现有的检测方法依赖于单一模态,无论是原始波形嵌入还是基于频谱的特征,都容易受到非欺骗干扰的影响,并且通常过拟合已知的伪造算法,导致对不可见攻击的泛化能力差。为了解决这些缺点,我们研究了混合融合框架,将基于自监督学习(SSL)的表示与手工制作的光谱描述符(MFCC,LFCC,CQCC)相结合。通过对齐和组合跨模态的互补信息,这些融合方法捕获单个特征方法通常忽略的细微伪影。我们探索了几种融合策略,包括简单的级联,交叉注意,相互交叉注意,和一个可学习的门控机制,以最佳方式混合SSL功能与细粒度的光谱线索。我们评估我们的方法在四个具有挑战性的公共基准和报告泛化性能。所有融合变体始终优于仅SSL基线,交叉注意策略实现了最佳泛化,等错误率(EER)相对降低38%。这些结果证实,波形和频谱视图的联合建模为音频深度伪造检测产生了鲁棒的域不可知表示。
摘要:Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features, are vulnerable to non spoof disturbances and often overfit to known forgery algorithms, resulting in poor generalization to unseen attacks. To address these shortcomings, we investigate hybrid fusion frameworks that integrate self supervised learning (SSL) based representations with handcrafted spectral descriptors (MFCC , LFCC, CQCC). By aligning and combining complementary information across modalities, these fusion approaches capture subtle artifacts that single feature approaches typically overlook. We explore several fusion strategies, including simple concatenation, cross attention, mutual cross attention, and a learnable gating mechanism, to optimally blend SSL features with fine grained spectral cues. We evaluate our approach on four challenging public benchmarks and report generalization performance. All fusion variants consistently outperform an SSL only baseline, with the cross attention strategy achieving the best generalization with a 38% relative reduction in equal error rate (EER). These results confirm that joint modeling of waveform and spectral views produces robust, domain agnostic representations for audio deepfake detection.
【8】Self-Improvement for Audio Large Language Model using Unlabeled Speech
标题:使用无标签语音的音频大语言模型自我改进
链接:https://arxiv.org/abs/2507.20169
备注:To appear in Interspeech 2025. 6 pages, 1 figure
摘要:最近的音频LLM迅速出现,在各种语音任务中表现出很强的泛化能力。然而,由于语音信号固有的复杂性,这些模型不可避免地遭受在特定的目标域的性能下降。为了解决这个问题,我们专注于在没有任何标记数据的情况下增强目标域中的音频LLM。我们提出了一种称为SI-SDA的自我改进方法,利用大模型解码中嵌入的信息来评估生成的伪标签的质量,然后基于强化学习优化进行域自适应。实验结果表明,我们的方法一致且显着地提高了音频LLM性能,在自动语音识别(ASR),口语问答(SQA)和语音到文本翻译(S2 TT)的多个公共数据集上,WER和BLEU的性能优于现有基线。此外,我们的方法具有很高的数据效率,强调了其在现实世界中部署的潜力。
摘要:Recent audio LLMs have emerged rapidly, demonstrating strong generalization across various speech tasks. However, given the inherent complexity of speech signals, these models inevitably suffer from performance degradation in specific target domains. To address this, we focus on enhancing audio LLMs in target domains without any labeled data. We propose a self-improvement method called SI-SDA, leveraging the information embedded in large-model decoding to evaluate the quality of generated pseudo labels and then perform domain adaptation based on reinforcement learning optimization. Experimental results show that our method consistently and significantly improves audio LLM performance, outperforming existing baselines in WER and BLEU across multiple public datasets of automatic speech recognition (ASR), spoken question-answering (SQA), and speech-to-text translation (S2TT). Furthermore, our approach exhibits high data efficiency, underscoring its potential for real-world deployment.
【9】Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
标题:不要模仿我的声音:Zero-Shot文本到语音的说话者身份的遗忘
链接:https://arxiv.org/abs/2507.20140
备注:Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), Vancouver, Canada. PMLR 267, 2025. Authors Jinju Kim and Taesoo Kim contributed equally
摘要:Zero-Shot Text-to-Speech(ZT-TTS)技术的快速发展使得能够从最少的音频线索进行高保真语音合成,这引起了重大的隐私和道德问题。尽管语音隐私受到威胁,但尚未探索选择性地从预训练的模型参数中删除复制不需要的个人语音的知识的研究。在本文中,我们解决了新的挑战,说话人身份unlearning的语音合成系统。为了实现这一目标,我们提出了第一个机器学习框架,特别是教师引导的学习(TGU),旨在确保模型忘记指定的扬声器身份,同时保留其为其他扬声器生成准确语音的能力。我们提出的方法结合了随机性,以防止一致的复制忘记扬声器的声音,确保未学习的身份仍然无法追踪。此外,我们提出了一个新的评估指标,扬声器零重训练遗忘(spk-ZRF)。这评估了模型忽略与被遗忘的说话者相关的提示的能力,有效地中和了它对这些声音的知识。在最先进的模型上进行的实验表明,TGU可以防止模型复制遗忘说话者的声音,同时保持其他说话者的高质量。该演示可在https://speechunlearn.github.io/上获得
摘要:The rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to selectively remove the knowledge to replicate unwanted individual voices from pre-trained model parameters has not been explored. In this paper, we address the new challenge of speaker identity unlearning for ZS-TTS systems. To meet this goal, we propose the first machine unlearning frameworks for ZS-TTS, especially Teacher-Guided Unlearning (TGU), designed to ensure the model forgets designated speaker identities while retaining its ability to generate accurate speech for other speakers. Our proposed methods incorporate randomness to prevent consistent replication of forget speakers' voices, assuring unlearned identities remain untraceable. Additionally, we propose a new evaluation metric, speaker-Zero Retrain Forgetting (spk-ZRF). This assesses the model's ability to disregard prompts associated with forgotten speakers, effectively neutralizing its knowledge of these voices. The experiments conducted on the state-of-the-art model demonstrate that TGU prevents the model from replicating forget speakers' voices while maintaining high quality for other speakers. The demo is available at https://speechunlearn.github.io/
【10】Diffusion-based Symbolic Music Generation with Structured State Space Models
标题:基于扩散的结构化状态空间模型符号音乐生成
链接:http://arxiv.org/pdf/2507.20128v1
备注:9 pages,3figures
摘要:扩散模型的最新进展显着改善了符号音乐的生成。然而,大多数方法依赖于具有自注意机制的基于变换器的架构,其受到二次计算复杂度的约束,限制了长序列的可扩展性。为了解决这个问题,我们提出了Mamba的符号音乐扩散(SMDIM),一种新的基于扩散的架构,集成了结构化状态空间模型(SSM)用于高效的全局上下文建模和Mamba前馈注意块(MFA)用于精确的局部细节保护。MFA Block结合了Mamba层的线性复杂性、前馈层的非线性细化和自注意机制的细粒度精度,实现了可扩展性和音乐表现力之间的平衡。SMDIM实现了接近线性的复杂性,使其对于长序列任务非常高效。在不同的数据集上进行评估,包括FolkDB,这是一个中国传统民间音乐的集合,代表了符号音乐生成中一个未充分探索的领域,SMDIM在生成质量和计算效率方面都优于最先进的模型。除了符号音乐,SMDIM的架构设计还展示了对广泛的长序列生成任务的适应性,为相干序列建模提供了可扩展和高效的解决方案。
摘要:Recent advancements in diffusion models have significantly improved symbolic music generation. However, most approaches rely on transformer-based architectures with self-attention mechanisms, which are constrained by quadratic computational complexity, limiting scalability for long sequences. To address this, we propose Symbolic Music Diffusion with Mamba (SMDIM), a novel diffusion-based architecture integrating Structured State Space Models (SSMs) for efficient global context modeling and the Mamba-FeedForward-Attention Block (MFA) for precise local detail preservation. The MFA Block combines the linear complexity of Mamba layers, the non-linear refinement of FeedForward layers, and the fine-grained precision of self-attention mechanisms, achieving a balance between scalability and musical expressiveness. SMDIM achieves near-linear complexity, making it highly efficient for long-sequence tasks. Evaluated on diverse datasets, including FolkDB, a collection of traditional Chinese folk music that represents an underexplored domain in symbolic music generation, SMDIM outperforms state-of-the-art models in both generation quality and computational efficiency. Beyond symbolic music, SMDIM's architectural design demonstrates adaptability to a broad range of long-sequence generation tasks, offering a scalable and efficient solution for coherent sequence modeling.
【11】Improving Deep Learning-based Respiratory Sound Analysis with Frequency Selection and Attention Mechanism
标题:利用频率选择和注意机制改进基于深度学习的呼吸音分析
链接:http://arxiv.org/pdf/2507.20128v1
备注:9 pages,3figures
摘要:扩散模型的最新进展显着改善了符号音乐的生成。然而,大多数方法依赖于具有自注意机制的基于变换器的架构,其受到二次计算复杂度的约束,限制了长序列的可扩展性。为了解决这个问题,我们提出了Mamba的符号音乐扩散(SMDIM),一种新的基于扩散的架构,集成了结构化状态空间模型(SSM)用于高效的全局上下文建模和Mamba前馈注意块(MFA)用于精确的局部细节保护。MFA Block结合了Mamba层的线性复杂性、前馈层的非线性细化和自注意机制的细粒度精度,实现了可扩展性和音乐表现力之间的平衡。SMDIM实现了接近线性的复杂性,使其对于长序列任务非常高效。在不同的数据集上进行评估,包括FolkDB,这是一个中国传统民间音乐的集合,代表了符号音乐生成中一个未充分探索的领域,SMDIM在生成质量和计算效率方面都优于最先进的模型。除了符号音乐,SMDIM的架构设计还展示了对广泛的长序列生成任务的适应性,为相干序列建模提供了可扩展和高效的解决方案。
摘要:Recent advancements in diffusion models have significantly improved symbolic music generation. However, most approaches rely on transformer-based architectures with self-attention mechanisms, which are constrained by quadratic computational complexity, limiting scalability for long sequences. To address this, we propose Symbolic Music Diffusion with Mamba (SMDIM), a novel diffusion-based architecture integrating Structured State Space Models (SSMs) for efficient global context modeling and the Mamba-FeedForward-Attention Block (MFA) for precise local detail preservation. The MFA Block combines the linear complexity of Mamba layers, the non-linear refinement of FeedForward layers, and the fine-grained precision of self-attention mechanisms, achieving a balance between scalability and musical expressiveness. SMDIM achieves near-linear complexity, making it highly efficient for long-sequence tasks. Evaluated on diverse datasets, including FolkDB, a collection of traditional Chinese folk music that represents an underexplored domain in symbolic music generation, SMDIM outperforms state-of-the-art models in both generation quality and computational efficiency. Beyond symbolic music, SMDIM's architectural design demonstrates adaptability to a broad range of long-sequence generation tasks, offering a scalable and efficient solution for coherent sequence modeling.
【12】Improving Audio Classification by Transitioning from Zero- to Few-Shot
标题:通过从零镜头过渡到Few-Shot改进音频分类
链接:http://arxiv.org/pdf/2507.20036v1
备注:Submitted to Interspeech 2025
摘要:现有技术的音频分类通常采用zero-shot方法,其涉及将音频嵌入与来自描述相应音频类的文本的嵌入进行比较。这些嵌入通常由通过对比学习训练的神经网络生成,以对齐音频和文本表示。识别音频类的最佳文本描述是具有挑战性的,特别是当该类包括各种各样的声音时。本文探讨了Few-Shot方法,旨在提高分类精度超过zero-shot的方法。具体而言,音频嵌入按类别分组并进行处理以替换固有的噪声文本嵌入。我们的研究结果表明,Few-Shot分类通常优于zero-shot基线。
摘要:State-of-the-art audio classification often employs a zero-shot approach, which involves comparing audio embeddings with embeddings from text describing the respective audio class. These embeddings are usually generated by neural networks trained through contrastive learning to align audio and text representations. Identifying the optimal text description for an audio class is challenging, particularly when the class comprises a wide variety of sounds. This paper examines few-shot methods designed to improve classification accuracy beyond the zero-shot approach. Specifically, audio embeddings are grouped by class and processed to replace the inherently noisy text embeddings. Our results demonstrate that few-shot classification typically outperforms the zero-shot baseline.
【13】Efficient Vocal-Conditioned Music Generation via Soft Alignment Attention and Latent Diffusion
标题:基于软对齐注意和潜在扩散的声乐条件音乐生成
链接:http://arxiv.org/pdf/2507.19991v1
备注:6 page, 3 figures
摘要:我们提出了一个轻量级的潜在扩散模型的声乐条件的音乐伴奏生成,解决现有的音乐AI系统的关键限制。我们的方法引入了一种新的软对齐注意机制,该机制基于扩散时间步长自适应地结合了局部和全局时间依赖性,从而能够有效地捕获多尺度音乐结构。该模型在预训练的变分自动编码器的压缩潜在空间中运行,与最先进的系统相比,参数减少了220倍,同时推理速度提高了52倍。实验评估表明,仅使用1500万个参数即可实现具有竞争力的性能,在制作质量和内容统一性方面优于OpenAI Juantic,同时保持合理的音乐连贯性。超轻量级架构支持在消费类硬件上进行实时部署,使AI辅助音乐创作可用于交互式应用程序和资源受限的环境。
摘要:
We present a lightweight latent diffusion model for vocal-conditioned musical accompaniment generation that addresses critical limitations in existing music AI systems. Our approach introduces a novel soft alignment attention mechanism that adaptively combines local and global temporal dependencies based on diffusion timesteps, enabling efficient capture of multi- scale musical structure. Operating in the compressed latent space of a pre-trained variational autoencoder, the model achieves a 220 times parameter reduction compared to state-of-the-art systems while delivering 52 times faster inference. Experimental evaluation demonstrates competitive performance with only 15M parame- ters, outperforming OpenAI Jukebox in production quality and content unity while maintaining reasonable musical coherence. The ultra-lightweight architecture enables real-time deployment on consumer hardware, making AI-assisted music creation ac- cessible for interactive applications and resource-constrained environments.
【14】ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
标题:ChoreoMuse:具有风格转移和节拍追随动作的强大音乐到舞蹈视频生成
链接:http://arxiv.org/pdf/2507.19836v1
作者:Xuanchen Wang, Heng Wang, Weidong Cai
备注:10 pages, 5 figures, accepted by the 33rd ACM International Conference on Multimedia (ACM MM 2025), demo page: this https URL
摘要:Modern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches often fail to produce high-quality dance videos that harmonize with both musical rhythm and user-defined choreography styles, limiting their applicability in real-world creative contexts. To address this gap, we introduce ChoreoMuse, a diffusion-based framework that uses SMPL format parameters and their variation version as intermediaries between music and video generation, thereby overcoming the usual constraints imposed by video resolution. Critically, ChoreoMuse supports style-controllable, high-fidelity dance video generation across diverse musical genres and individual dancer characteristics, including the flexibility to handle any reference individual at any resolution. Our method employs a novel music encoder MotionTune to capture motion cues from audio, ensuring that the generated choreography closely follows the beat and expressive qualities of the input music. To quantitatively evaluate how well the generated dances match both musical and choreographic styles, we introduce two new metrics that measure alignment with the intended stylistic cues. Extensive experiments confirm that ChoreoMuse achieves state-of-the-art performance across multiple dimensions, including video quality, beat alignment, dance diversity, and style adherence, demonstrating its potential as a robust solution for a wide range of creative applications. Video results can be found on our project page: https: choreomuse.github.io.
【15】SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
标题:SonicGauss:3D高斯表示的位置感知物理声音合成
链接:http://arxiv.org/pdf/2507.19835v1
作者:Chunshi Wang, Hongxing Li, Yawei Luo
备注:Accepted by ACMMM'25
摘要:While 3D Gaussian representations (3DGS) have proven effective for modeling the geometry and appearance of objects, their potential for capturing other physical attributes-such as sound-remains largely unexplored. In this paper, we present a novel framework dubbed SonicGauss for synthesizing impact sounds from 3DGS representations by leveraging their inherent geometric and material properties. Specifically, we integrate a diffusion-based sound synthesis model with a PointTransformer-based feature extractor to infer material characteristics and spatial-acoustic correlations directly from Gaussian ellipsoids. Our approach supports spatially varying sound responses conditioned on impact locations and generalizes across a wide range of object categories. Experiments on the ObjectFolder dataset and real-world recordings demonstrate that our method produces realistic, position-aware auditory feedback. The results highlight the framework's robustness and generalization ability, offering a promising step toward bridging 3D visual representations and interactive sound synthesis. Project page: https: chunshi.wang SonicGauss
【16】MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
标题:MCIF:多模式跨语言教学-遵循科学谈话的基准
链接:http://arxiv.org/pdf/2507.19634v1
作者:Sara Papi, Maike Züfle, Marco Gaido, Beatrice Savoldi, Danni Liu, Ioannis Douros, Luisa Bentivogli, Jan Niehues
备注:Work in progress
摘要:Recent advances in large language models have catalyzed the development of multimodal LLMs (MLLMs) that integrate text, speech, and vision within unified frameworks. As MLLMs evolve from narrow, monolingual, task-specific systems to general-purpose instruction-following models, a key frontier lies in evaluating their multilingual and multimodal capabilities over both long and short contexts. However, existing benchmarks fall short in evaluating these dimensions jointly: they are often limited to English, mostly focus on one single modality at a time, rely on short-form contexts, or lack human annotations--hindering comprehensive assessment of model performance across languages, modalities, and task complexity. To address these gaps, we introduce MCIF (Multimodal Crosslingual Instruction Following), the first multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities--speech, vision, and text--and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs' abilities to interpret instructions across languages and combine them with multimodal contextual information. MCIF is released under a CC-BY 4.0 license to encourage open research and progress in MLLMs development.
【17】Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
标题:低复杂度声学场景分类的联合特征和输出蒸馏
链接:https://arxiv.org/abs/2507.19557
备注:4 pages, submitted to DCASE2025 Challenge Task 1
摘要:本报告提出了一个双层次的知识蒸馏框架与多教师指导的低复杂度声学场景分类(ASC)在DCASE 2025任务1。我们提出了一个蒸馏策略,共同转移软logits和中间特征表示。具体来说,我们将PaSST和CP-ResNet模型作为教师模型进行了预训练。来自教师的Logits被平均以生成软目标,而一个CP-ResNet被选择用于特征级蒸馏。这使得紧凑的学生模型(CP-移动),以捕捉语义分布和结构信息,从教师的指导。在TAU Urban Acoustic Scenes 2022 Mobile数据集(开发集)上的实验表明,我们提交的系统达到了59.30%的准确率。
摘要:This report presents a dual-level knowledge distillation framework with multi-teacher guidance for low-complexity acoustic scene classification (ASC) in DCASE2025 Task 1. We propose a distillation strategy that jointly transfers both soft logits and intermediate feature representations. Specifically, we pre-trained PaSST and CP-ResNet models as teacher models. Logits from teachers are averaged to generate soft targets, while one CP-ResNet is selected for feature-level distillation. This enables the compact student model (CP-Mobile) to capture both semantic distribution and structural information from teacher guidance. Experiments on the TAU Urban Acoustic Scenes 2022 Mobile dataset (development set) demonstrate that our submitted systems achieve up to 59.30\% accuracy.
【18】MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection
标题:MIMIII-Agent:利用LLM并调用函数来对异常声音检测进行相对评估
链接:https://arxiv.org/abs/2507.20666
摘要:本文提出了一种生成特定于机器类型的异常的方法,以评估不同机器类型的无监督异常声音检测(UASD)系统的相对性能,即使在没有真实异常声音数据的情况下也是如此。传统的基于关键字的数据增强方法通常会产生不切实际的声音,因为它们依赖于手动定义的标签,限制了机器类型和异常模式多样化的可扩展性。高级音频生成模型,如MIMII-Gen,显示出希望,但通常依赖于异常训练数据,使得它们在各种异常示例不可用时效率较低。为了解决这些限制,我们提出了一种新的合成方法,利用大型语言模型(LLM)来解释故障的文本描述,并自动选择音频转换函数,将正常的机器声音转换为多样化和合理的异常声音。我们验证这种方法,通过评估UASD系统只训练正常的声音从五种机器类型,使用真实的和合成的异常数据。实验结果表明,在合成和真实异常之间,机器类型的相对检测难度趋势一致。这一发现支持了我们的假设,并强调了建议的LLM为基础的合成方法的UASD系统的相对评价的有效性。
摘要:This paper proposes a method for generating machine-type-specific anomalies to evaluate the relative performance of unsupervised anomalous sound detection (UASD) systems across different machine types, even in the absence of real anomaly sound data. Conventional keyword-based data augmentation methods often produce unrealistic sounds due to their reliance on manually defined labels, limiting scalability as machine types and anomaly patterns diversify. Advanced audio generative models, such as MIMII-Gen, show promise but typically depend on anomalous training data, making them less effective when diverse anomalous examples are unavailable. To address these limitations, we propose a novel synthesis approach leveraging large language models (LLMs) to interpret textual descriptions of faults and automatically select audio transformation functions, converting normal machine sounds into diverse and plausible anomalous sounds. We validate this approach by evaluating a UASD system trained only on normal sounds from five machine types, using both real and synthetic anomaly data. Experimental results reveal consistent trends in relative detection difficulty across machine types between synthetic and real anomalies. This finding supports our hypothesis and highlights the effectiveness of the proposed LLM-based synthesis approach for relative evaluation of UASD systems.
【19】Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots
标题:基于HRTF线索的仿人机器人双耳声音事件定位与检测
链接:https://arxiv.org/abs/2507.20530
备注:Submitted to IEEE/ACM TASLP
摘要:本文介绍了双耳声音事件定位和检测(BiSELD),一个任务,旨在联合检测和定位多个声音事件使用双耳音频,灵感来自人类的空间听觉机制。为了支持这一任务,我们提出了一个合成的基准数据集,称为双耳集,它模拟现实的听觉场景,使用测得的头部相关的传递函数(HRTF)和不同的声音事件。为了有效地解决BiSELD任务,我们提出了一种新的输入特征表示,称为双耳时频特征(BTFF),它编码双耳时间差(ITD),双耳电平差(ILD)和高频频谱线索(SC)从双耳信号。BTFF由八个通道组成,包括左和右梅尔频谱图、速度图、SC图和ITD/ILD图,旨在覆盖频带和空间轴上的不同空间线索。一个基于CRNN的模型,BiSELDnet,然后开发学习频谱时间模式和基于HRTF的定位线索从BTFF。在双耳集上的实验表明,每个BTFF子特征都增强了任务性能:V-map提高了检测,ITD-/ILD-map实现了准确的水平定位,SC-map捕获了垂直空间线索。最终的系统实现了0.110的SELD误差与87. 1% F-score和4.4{\deg}定位误差,证明了所提出的框架在模仿人类的听觉感知的有效性。
摘要:This paper introduces Binaural Sound Event Localization and Detection (BiSELD), a task that aims to jointly detect and localize multiple sound events using binaural audio, inspired by the spatial hearing mechanism of humans. To support this task, we present a synthetic benchmark dataset, called the Binaural Set, which simulates realistic auditory scenes using measured head-related transfer functions (HRTFs) and diverse sound events. To effectively address the BiSELD task, we propose a new input feature representation called the Binaural Time-Frequency Feature (BTFF), which encodes interaural time difference (ITD), interaural level difference (ILD), and high-frequency spectral cues (SC) from binaural signals. BTFF is composed of eight channels, including left and right mel-spectrograms, velocity-maps, SC-maps, and ITD-/ILD-maps, designed to cover different spatial cues across frequency bands and spatial axes. A CRNN-based model, BiSELDnet, is then developed to learn both spectro-temporal patterns and HRTF-based localization cues from BTFF. Experiments on the Binaural Set show that each BTFF sub-feature enhances task performance: V-map improves detection, ITD-/ILD-maps enable accurate horizontal localization, and SC-map captures vertical spatial cues. The final system achieves a SELD error of 0.110 with 87.1% F-score and 4.4{\deg} localization error, demonstrating the effectiveness of the proposed framework in mimicking human-like auditory perception.
【1】End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
标题:有噪多说话者场景下的端到端DoA引导语音提取
链接:https://arxiv.org/abs/2507.20926
备注:Accepted by INTERSPEECH 2025
摘要:目标说话人提取(TSE)在噪声和多说话人环境中对语音信号的增强起着至关重要的作用。本文提出了一种端到端的TSE模型,它结合了到达方向(DOA)和波束宽度嵌入,从以DOA为中心的指定空间区域中提取语音。我们的方法有效地捕捉空间和时间特征,在多个同时发言者的高度复杂的场景中实现强大的性能。实验结果表明,该模型不仅能在规定的波束宽度内显著增强目标语音,而且能有效抑制来自其他方向的干扰,产生清晰、孤立的目标语音。此外,该模型在下游自动语音识别(ASR)任务中实现了显着的改进,使其特别适合于现实世界的应用。
摘要:Target Speaker Extraction (TSE) plays a critical role in enhancing speech signals in noisy and multi-speaker environments. This paper presents an end-to-end TSE model that incorporates Direction of Arrival (DOA) and beamwidth embeddings to extract speech from a specified spatial region centered around the DOA. Our approach efficiently captures spatial and temporal features, enabling robust performance in highly complex scenarios with multiple simultaneous speakers. Experimental results demonstrate that the proposed model not only significantly enhances the target speech within the defined beamwidth but also effectively suppresses interference from other directions, producing a clear and isolated target voice. Furthermore, the model achieves remarkable improvements in downstream Automatic Speech Recognition (ASR) tasks, making it particularly suitable for real-world applications.
【2】MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection
标题:MIMIII-Agent:利用LLM并调用函数来对异常声音检测进行相对评估
链接:https://arxiv.org/abs/2507.20666
摘要:本文提出了一种生成特定于机器类型的异常的方法,以评估不同机器类型的无监督异常声音检测(UASD)系统的相对性能,即使在没有真实异常声音数据的情况下也是如此。传统的基于关键字的数据增强方法通常会产生不切实际的声音,因为它们依赖于手动定义的标签,限制了机器类型和异常模式多样化的可扩展性。高级音频生成模型,如MIMII-Gen,显示出希望,但通常依赖于异常训练数据,使得它们在各种异常示例不可用时效率较低。为了解决这些限制,我们提出了一种新的合成方法,利用大型语言模型(LLM)来解释故障的文本描述,并自动选择音频转换函数,将正常的机器声音转换为多样化和合理的异常声音。我们验证这种方法,通过评估UASD系统只训练正常的声音从五种机器类型,使用真实的和合成的异常数据。实验结果表明,在合成和真实异常之间,机器类型的相对检测难度趋势一致。这一发现支持了我们的假设,并强调了建议的LLM为基础的合成方法的UASD系统的相对评价的有效性。
摘要:This paper proposes a method for generating machine-type-specific anomalies to evaluate the relative performance of unsupervised anomalous sound detection (UASD) systems across different machine types, even in the absence of real anomaly sound data. Conventional keyword-based data augmentation methods often produce unrealistic sounds due to their reliance on manually defined labels, limiting scalability as machine types and anomaly patterns diversify. Advanced audio generative models, such as MIMII-Gen, show promise but typically depend on anomalous training data, making them less effective when diverse anomalous examples are unavailable. To address these limitations, we propose a novel synthesis approach leveraging large language models (LLMs) to interpret textual descriptions of faults and automatically select audio transformation functions, converting normal machine sounds into diverse and plausible anomalous sounds. We validate this approach by evaluating a UASD system trained only on normal sounds from five machine types, using both real and synthetic anomaly data. Experimental results reveal consistent trends in relative detection difficulty across machine types between synthetic and real anomalies. This finding supports our hypothesis and highlights the effectiveness of the proposed LLM-based synthesis approach for relative evaluation of UASD systems.
【3】Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots
标题:基于HRTF线索的仿人机器人双耳声音事件定位与检测
链接:https://arxiv.org/abs/2507.20530
备注:Submitted to IEEE/ACM TASLP
摘要:本文介绍了双耳声音事件定位和检测(BiSELD),一个任务,旨在联合检测和定位多个声音事件使用双耳音频,灵感来自人类的空间听觉机制。为了支持这一任务,我们提出了一个合成的基准数据集,称为双耳集,它模拟现实的听觉场景,使用测得的头部相关的传递函数(HRTF)和不同的声音事件。为了有效地解决BiSELD任务,我们提出了一种新的输入特征表示,称为双耳时频特征(BTFF),它编码双耳时间差(ITD),双耳电平差(ILD)和高频频谱线索(SC)从双耳信号。BTFF由八个通道组成,包括左和右梅尔频谱图、速度图、SC图和ITD/ILD图,旨在覆盖频带和空间轴上的不同空间线索。一个基于CRNN的模型,BiSELDnet,然后开发学习频谱时间模式和基于HRTF的定位线索从BTFF。在双耳集上的实验表明,每个BTFF子特征都增强了任务性能:V-map提高了检测,ITD-/ILD-map实现了准确的水平定位,SC-map捕获了垂直空间线索。最终的系统实现了0.110的SELD误差与87. 1% F-score和4.4{\deg}定位误差,证明了所提出的框架在模仿人类的听觉感知的有效性。
摘要:This paper introduces Binaural Sound Event Localization and Detection (BiSELD), a task that aims to jointly detect and localize multiple sound events using binaural audio, inspired by the spatial hearing mechanism of humans. To support this task, we present a synthetic benchmark dataset, called the Binaural Set, which simulates realistic auditory scenes using measured head-related transfer functions (HRTFs) and diverse sound events. To effectively address the BiSELD task, we propose a new input feature representation called the Binaural Time-Frequency Feature (BTFF), which encodes interaural time difference (ITD), interaural level difference (ILD), and high-frequency spectral cues (SC) from binaural signals. BTFF is composed of eight channels, including left and right mel-spectrograms, velocity-maps, SC-maps, and ITD-/ILD-maps, designed to cover different spatial cues across frequency bands and spatial axes. A CRNN-based model, BiSELDnet, is then developed to learn both spectro-temporal patterns and HRTF-based localization cues from BTFF. Experiments on the Binaural Set show that each BTFF sub-feature enhances task performance: V-map improves detection, ITD-/ILD-maps enable accurate horizontal localization, and SC-map captures vertical spatial cues. The final system achieves a SELD error of 0.110 with 87.1% F-score and 4.4{\deg} localization error, demonstrating the effectiveness of the proposed framework in mimicking human-like auditory perception.
【4】Binaural Localization Model for Speech in Noise
标题:噪音中语音的双耳定位模型
链接:https://arxiv.org/abs/2507.20027
摘要:双耳声源定位对于人类收听者的空间意识、交流和安全是重要的。本文提出了一种噪声环境下的端到端语音双耳定位模型。介绍了一种轻量级的卷积递归网络,它可以将声音定位在有噪混响双耳信号的正面方位平面中。该模型采用加性内耳噪声来表示典型听者的频率依赖性听力阈值。该模型的定位性能进行了比较与导向响应功率算法,并使用该模型作为双耳语音增强方法的耳间线索保存的措施进行了研究。进行了听力测试,以比较该模型的性能与人类在嘈杂的条件下的语音定位。
摘要:Binaural acoustic source localization is important to human listeners for spatial awareness, communication and safety. In this paper, an end-to-end binaural localization model for speech in noise is presented. A lightweight convolutional recurrent network that localizes sound in the frontal azimuthal plane for noisy reverberant binaural signals is introduced. The model incorporates additive internal ear noise to represent the frequency-dependent hearing threshold of a typical listener. The localization performance of the model is compared with the steered response power algorithm, and the use of the model as a measure of interaural cue preservation for binaural speech enhancement methods is studied. A listening test was performed to compare the performance of the model with human localization of speech in noisy conditions.
【5】Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
标题:使用复杂卷积回归网络的双耳语音增强
链接:https://arxiv.org/abs/2507.20023
摘要:从助听器到增强和虚拟现实设备,双耳语音增强算法已被确立为最先进的技术,以提高语音清晰度和听觉舒适度。在本文中,我们提出了一种端到端的双耳语音增强方法,该方法使用具有编码器-解码器架构的复杂递归卷积网络和放置在编码器和解码器之间的复杂LSTM递归块。一个损失函数,除了语音清晰度的改善和降噪的空间信息的保存。该网络在时频域中估计双耳听力设备的左耳和右耳通道的各个复比掩码。我们表明,与其他基线算法相比,所提出的方法显着提高了估计的语音清晰度并降低了噪声,同时在具有单个目标扬声器和各种类型各向同性噪声的声学情况下保留了双耳信号的空间信息。
摘要:From hearing aids to augmented and virtual reality devices, binaural speech enhancement algorithms have been established as state-of-the-art techniques to improve speech intelligibility and listening comfort. In this paper, we present an end-to-end binaural speech enhancement method using a complex recurrent convolutional network with an encoder-decoder architecture and a complex LSTM recurrent block placed between the encoder and decoder. A loss function that focuses on the preservation of spatial information in addition to speech intelligibility improvement and noise reduction is introduced. The network estimates individual complex ratio masks for the left and right-ear channels of a binaural hearing device in the time-frequency domain. We show that, compared to other baseline algorithms, the proposed method significantly improves the estimated speech intelligibility and reduces the noise while preserving the spatial information of the binaural signals in acoustic situations with a single target speaker and isotropic noise of various types.
【6】Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
标题:具有多个时变条件的可控视频到音乐生成
链接:https://arxiv.org/abs/2507.20627
备注:Accepted by the 33rd ACM International Conference on Multimedia (ACMMM 2025). The project page is available at this https URL
摘要:音乐增强了视频叙事和情感,推动了对自动视频到音乐(V2M)生成的需求。然而,现有的V2M方法仅依赖于视觉特征或补充文本输入,以黑盒方式生成音乐,通常无法满足用户期望。为了解决这一挑战,我们提出了一种新的多条件引导的V2M生成框架,该框架结合了多个时变条件,以增强对音乐生成的控制。我们的方法使用两阶段的训练策略,使学习V2M的基本原理和视听时间同步,同时满足用户的多条件控制的需求。在第一阶段,我们引入了一个细粒度的特征选择模块和一个渐进的时间对齐注意机制,以确保灵活的特征对齐。在第二阶段,我们开发了一个动态条件融合模块和一个控制引导的解码器模块,以整合多个条件,准确地指导音乐创作过程。大量的实验表明,我们的方法在主观和客观评估方面都优于现有的V2M管道,显着增强了控制和与用户期望的一致性。
摘要:Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.
【7】Sound Safeguarding for Acoustic Measurement Using Any Sounds: Tools and Applications
标题:使用任何声音进行声学测量的声音保护:工具和应用
链接:https://arxiv.org/abs/2507.20485
备注:2 pages, 2 figures, IEEE GCCE 2025 Demo session, Accepted
摘要:我们展示了基于“声音保护”方法开发的工具和应用程序,该方法使任何声音都能用于声学测量。我们开发了用于准备、交互式和实时测量以及报告生成的工具。在开发过程中,我们根据实际应用情况对该方法进行了扩展和改进。我们已经开源了这些工具,并鼓励潜在用户使用它们来改善他们的声学环境。
摘要:We demonstrate tools and applications developed based on the method of "sound safeguarding," which enables any sound to be used for acoustic measurements. We developed tools for preparation, interactive and real-time measurement, and report generation. We extended and modified the method during its development based on its application in various practical situations. We have open-sourced these tools and encourage prospective users to use them to improve their acoustic environments.
【8】Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection
标题:两种观点,一种真相:光谱和自监督特征融合用于稳健的语音深度伪造检测
链接:https://arxiv.org/abs/2507.20417
备注:ACCEPTED WASPAA 2025
摘要:合成语音的最新进展使音频deepfake变得越来越逼真,带来了重大的安全风险。现有的检测方法依赖于单一模态,无论是原始波形嵌入还是基于频谱的特征,都容易受到非欺骗干扰的影响,并且通常过拟合已知的伪造算法,导致对不可见攻击的泛化能力差。为了解决这些缺点,我们研究了混合融合框架,将基于自监督学习(SSL)的表示与手工制作的光谱描述符(MFCC,LFCC,CQCC)相结合。通过对齐和组合跨模态的互补信息,这些融合方法捕获单个特征方法通常忽略的细微伪影。我们探索了几种融合策略,包括简单的级联,交叉注意,相互交叉注意,和一个可学习的门控机制,以最佳方式混合SSL功能与细粒度的光谱线索。我们评估我们的方法在四个具有挑战性的公共基准和报告泛化性能。所有融合变体始终优于仅SSL基线,交叉注意策略实现了最佳泛化,等错误率(EER)相对降低38%。这些结果证实,波形和频谱视图的联合建模为音频深度伪造检测产生了鲁棒的域不可知表示。
摘要:Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features, are vulnerable to non spoof disturbances and often overfit to known forgery algorithms, resulting in poor generalization to unseen attacks. To address these shortcomings, we investigate hybrid fusion frameworks that integrate self supervised learning (SSL) based representations with handcrafted spectral descriptors (MFCC , LFCC, CQCC). By aligning and combining complementary information across modalities, these fusion approaches capture subtle artifacts that single feature approaches typically overlook. We explore several fusion strategies, including simple concatenation, cross attention, mutual cross attention, and a learnable gating mechanism, to optimally blend SSL features with fine grained spectral cues. We evaluate our approach on four challenging public benchmarks and report generalization performance. All fusion variants consistently outperform an SSL only baseline, with the cross attention strategy achieving the best generalization with a 38% relative reduction in equal error rate (EER). These results confirm that joint modeling of waveform and spectral views produces robust, domain agnostic representations for audio deepfake detection.
【9】Self-Improvement for Audio Large Language Model using Unlabeled Speech
标题:使用无标签语音的音频大语言模型自我改进
链接:https://arxiv.org/abs/2507.20169
备注:To appear in Interspeech 2025. 6 pages, 1 figure
摘要:最近的音频LLM迅速出现,在各种语音任务中表现出很强的泛化能力。然而,由于语音信号固有的复杂性,这些模型不可避免地遭受在特定的目标域的性能下降。为了解决这个问题,我们专注于在没有任何标记数据的情况下增强目标域中的音频LLM。我们提出了一种称为SI-SDA的自我改进方法,利用大模型解码中嵌入的信息来评估生成的伪标签的质量,然后基于强化学习优化进行域自适应。实验结果表明,我们的方法一致且显着地提高了音频LLM性能,在自动语音识别(ASR),口语问答(SQA)和语音到文本翻译(S2 TT)的多个公共数据集上,WER和BLEU的性能优于现有基线。此外,我们的方法具有很高的数据效率,强调了其在现实世界中部署的潜力。
摘要:Recent audio LLMs have emerged rapidly, demonstrating strong generalization across various speech tasks. However, given the inherent complexity of speech signals, these models inevitably suffer from performance degradation in specific target domains. To address this, we focus on enhancing audio LLMs in target domains without any labeled data. We propose a self-improvement method called SI-SDA, leveraging the information embedded in large-model decoding to evaluate the quality of generated pseudo labels and then perform domain adaptation based on reinforcement learning optimization. Experimental results show that our method consistently and significantly improves audio LLM performance, outperforming existing baselines in WER and BLEU across multiple public datasets of automatic speech recognition (ASR), spoken question-answering (SQA), and speech-to-text translation (S2TT). Furthermore, our approach exhibits high data efficiency, underscoring its potential for real-world deployment.
【10】Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
标题:不要模仿我的声音:Zero-Shot文本到语音的说话者身份的遗忘
链接:https://arxiv.org/abs/2507.20140
备注:Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), Vancouver, Canada. PMLR 267, 2025. Authors Jinju Kim and Taesoo Kim contributed equally
摘要:Zero-Shot Text-to-Speech(ZT-TTS)技术的快速发展使得能够从最少的音频线索进行高保真语音合成,这引起了重大的隐私和道德问题。尽管语音隐私受到威胁,但尚未探索选择性地从预训练的模型参数中删除复制不需要的个人语音的知识的研究。在本文中,我们解决了新的挑战,说话人身份unlearning的语音合成系统。为了实现这一目标,我们提出了第一个机器学习框架,特别是教师引导的学习(TGU),旨在确保模型忘记指定的扬声器身份,同时保留其为其他扬声器生成准确语音的能力。我们提出的方法结合了随机性,以防止一致的复制忘记扬声器的声音,确保未学习的身份仍然无法追踪。此外,我们提出了一个新的评估指标,扬声器零重训练遗忘(spk-ZRF)。这评估了模型忽略与被遗忘的说话者相关的提示的能力,有效地中和了它对这些声音的知识。在最先进的模型上进行的实验表明,TGU可以防止模型复制遗忘说话者的声音,同时保持其他说话者的高质量。该演示可在https://speechunlearn.github.io/上获得
摘要:The rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to selectively remove the knowledge to replicate unwanted individual voices from pre-trained model parameters has not been explored. In this paper, we address the new challenge of speaker identity unlearning for ZS-TTS systems. To meet this goal, we propose the first machine unlearning frameworks for ZS-TTS, especially Teacher-Guided Unlearning (TGU), designed to ensure the model forgets designated speaker identities while retaining its ability to generate accurate speech for other speakers. Our proposed methods incorporate randomness to prevent consistent replication of forget speakers' voices, assuring unlearned identities remain untraceable. Additionally, we propose a new evaluation metric, speaker-Zero Retrain Forgetting (spk-ZRF). This assesses the model's ability to disregard prompts associated with forgotten speakers, effectively neutralizing its knowledge of these voices. The experiments conducted on the state-of-the-art model demonstrate that TGU prevents the model from replicating forget speakers' voices while maintaining high quality for other speakers. The demo is available at https://speechunlearn.github.io/
【11】ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
标题:ProsodyLM:揭示语音语言模型中新兴的韵律处理能力
链接:https://arxiv.org/abs/2507.20091
摘要:语音语言模型是指具有语音处理和理解能力的语言模型。语音语言模型的一个关键的期望能力是捕捉内容和韵律之间复杂的相互依赖性的能力。训练语音语言模型的现有主流范式,将语音转换成离散的令牌,然后将它们送入LLM,是次优的学习韵律信息-我们发现,由此产生的LLM不表现出明显的新兴韵律处理能力,通过单独的预训练。为了克服这一点,我们提出了ProsodyLM,它引入了一个简单的标记化方案,适合学习韵律。每个语音话语首先被转录成文本,随后是一系列的词级韵律令牌。与传统的语音标记化方案相比,本文提出的标记化方案保留了更完整的韵律信息,并且更容易被基于文本的LLM理解。我们发现,ProsodyLM可以通过单独的预训练来学习令人惊讶的多样化新兴韵律处理能力,从利用生成的语音中的韵律细微差别,例如对比焦点,理解话语中的情感和压力,到在长上下文中保持韵律一致性。
摘要:Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody. The existing mainstream paradigm of training speech language models, which converts speech into discrete tokens before feeding them into LLMs, is sub-optimal in learning prosody information -- we find that the resulting LLMs do not exhibit obvious emerging prosody processing capabilities via pre-training alone. To overcome this, we propose ProsodyLM, which introduces a simple tokenization scheme amenable to learning prosody. Each speech utterance is first transcribed into text, followed by a sequence of word-level prosody tokens. Compared with conventional speech tokenization schemes, the proposed tokenization scheme retains more complete prosody information, and is more understandable to text-based LLMs. We find that ProsodyLM can learn surprisingly diverse emerging prosody processing capabilities through pre-training alone, ranging from harnessing the prosody nuances in generated speech, such as contrastive focus, understanding emotion and stress in an utterance, to maintaining prosody consistency in long contexts.
【12】Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
标题:低复杂度声学场景分类的联合特征和输出蒸馏
链接:https://arxiv.org/abs/2507.19557
备注:4 pages, submitted to DCASE2025 Challenge Task 1
摘要:本报告提出了一个双层次的知识蒸馏框架与多教师指导的低复杂度声学场景分类(ASC)在DCASE 2025任务1。我们提出了一个蒸馏策略,共同转移软logits和中间特征表示。具体来说,我们将PaSST和CP-ResNet模型作为教师模型进行了预训练。来自教师的Logits被平均以生成软目标,而一个CP-ResNet被选择用于特征级蒸馏。这使得紧凑的学生模型(CP-移动),以捕捉语义分布和结构信息,从教师的指导。在TAU Urban Acoustic Scenes 2022 Mobile数据集(开发集)上的实验表明,我们提交的系统达到了59.30%的准确率。
摘要:This report presents a dual-level knowledge distillation framework with multi-teacher guidance for low-complexity acoustic scene classification (ASC) in DCASE2025 Task 1. We propose a distillation strategy that jointly transfers both soft logits and intermediate feature representations. Specifically, we pre-trained PaSST and CP-ResNet models as teacher models. Logits from teachers are averaged to generate soft targets, while one CP-ResNet is selected for feature-level distillation. This enables the compact student model (CP-Mobile) to capture both semantic distribution and structural information from teacher guidance. Experiments on the TAU Urban Acoustic Scenes 2022 Mobile dataset (development set) demonstrate that our submitted systems achieve up to 59.30\% accuracy.
机器翻译由腾讯交互翻译提供,仅供参考
