今日论文合集:cs.SD语音18篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Voxtral
标题:沃克斯特拉尔
链接:https://arxiv.org/abs/2507.13264

作者: H. Liu, Andy Ehrenberg, Andy Lo, Clément Denoix, Corentin Barreau, Guillaume Lample, Jean-Malo Delignon, Khyathi Raghavi Chandu, Patrick von Platen, Pavankumar Reddy Muddireddy, Sanchit Gandhi, Soham Ghosh, Srijan Mishra, Thomas Foubert, Abhinav Rastogi, Adam Yang, Albert Q. Jiang, Alexandre Sablayrolles, Amélie Héliou, Amélie Martin, Anmol Agarwal, Antoine Roux, Arthur Darcet, Arthur Mensch, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Darius Dabert, Devendra Singh Chaplot, Devon Mizelle, Diego de las Casas, Elliot Chane-Sane, Emilien Fugier, Emma Bou Hanna, Gabrielle Berrada, Gauthier Delerce, Gauthier Guinet, Georgii Novikov, Guillaume Martin, Himanshu Jaju, Jan Ludziejewski, Jason Rute, Jean-Hadrien Chabran, Jessica Chudnovsky, Joachim Studnia, Joep Barmentlo, Jonas Amar, Josselin Somerville Roberts, Julien Denize, Karan Saxena, Karmesh Yadav, Kartik Khandelwal, Kush Jain, Lélio Renard Lavaud, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Marie Pellat, Mathilde Guillaumin, Mathis Felardos, Matthieu Dinot, Maxime Darrin, Maximilian Augustin, Mickaël Seznec, Neha Gupta, Nikhil Raghuraman, Olivier Duchenne, Patricia Wang, Patryk Saffer, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Rémi Delacourt, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Sandeep Subramanian, Saurabh Garg, Shashwat Dalal, Siddharth Gandhi, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Thibault Schueller, Thibaut Lavril, Thomas Robert, Thomas Wang, Timothée Lacroix, Tom Bewley, Valeriia Nemychnikova, Victor Paltz
备注:17 pages
摘要:我们提出Voxtral Mini和Voxtral Small,两个多模式音频聊天模型。Voxtral经过培训,能够理解语音和文本文档,在各种音频基准测试中实现最先进的性能,同时保持强大的文本功能。Voxtral Small的性能优于许多闭源模型,同时足够小,可以在本地运行。32K上下文窗口使模型能够处理长达40分钟的音频文件和长时间的多回合对话。我们还提供了三个基准评估语音理解模型的知识和琐事。这两个Voxtral模型都是在Apache 2.0许可下发布的。
摘要:We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio benchmarks, while preserving strong text capabilities. Voxtral Small outperforms a number of closed-source models, while being small enough to run locally. A 32K context window enables the model to handle audio files up to 40 minutes in duration and long multi-turn conversations. We also contribute three benchmarks for evaluating speech understanding models on knowledge and trivia. Both Voxtral models are released under Apache 2.0 license.


【2】SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks
标题:SHIELD:一种安全且高度增强的集成学习,用于针对对抗性攻击的稳健Deepfake检测
链接:https://arxiv.org/abs/2507.13170

作者:in, Awais Khan, Muhammad Umar Farooq, Khalid Malik
摘要:音频在说话人验证、支持语音的智能设备和音频会议等应用中起着至关重要的作用。然而,音频操纵,如deepfakes,通过传播错误信息带来了重大风险。我们的实证分析表明,现有的检测deepfake音频的方法往往容易受到反取证(AF)攻击,特别是那些使用生成对抗网络的攻击。在本文中,我们提出了一种新的协作学习方法SHIELD来防御生成式AF攻击。为了暴露AF签名,我们集成了一个辅助生成模型,称为防御(DF)生成模型,它通过结合输入和输出来促进协作学习。此外,我们设计了一个三元组模型来捕获真实和AF攻击音频与真实生成和攻击生成音频的相关性,使用辅助生成模型。所提出的SHIELD增强了对生成AF攻击的防御,并在各种生成模型中实现了鲁棒性能。对于三种不同的生成模型,所提出的AF将ASVspoof 2019的平均检测准确率从95.49%降至59.77%,将In-the-Wild的平均检测准确率从99.44%降至38.45%,将HalfTruth的平均检测准确率从98.41%降至51.18%。所提出的SHIELD机制对AF攻击具有鲁棒性,并且在ASVspoof 2019,In-the-Wild和HalfTruth数据集的匹配设置中分别实现了98.13%,98.58%和99.57%的平均准确率,以及98.78%,98.62%和98.85%的不匹配设置。
摘要:Audio plays a crucial role in applications like speaker verification, voice-enabled smart devices, and audio conferencing. However, audio manipulations, such as deepfakes, pose significant risks by enabling the spread of misinformation. Our empirical analysis reveals that existing methods for detecting deepfake audio are often vulnerable to anti-forensic (AF) attacks, particularly those attacked using generative adversarial networks. In this article, we propose a novel collaborative learning method called SHIELD to defend against generative AF attacks. To expose AF signatures, we integrate an auxiliary generative model, called the defense (DF) generative model, which facilitates collaborative learning by combining input and output. Furthermore, we design a triplet model to capture correlations for real and AF attacked audios with real-generated and attacked-generated audios using auxiliary generative models. The proposed SHIELD strengthens the defense against generative AF attacks and achieves robust performance across various generative models. The proposed AF significantly reduces the average detection accuracy from 95.49% to 59.77% for ASVspoof2019, from 99.44% to 38.45% for In-the-Wild, and from 98.41% to 51.18% for HalfTruth for three different generative models. The proposed SHIELD mechanism is robust against AF attacks and achieves an average accuracy of 98.13%, 98.58%, and 99.57% in match, and 98.78%, 98.62%, and 98.85% in mismatch settings for the ASVspoof2019, In-the-Wild, and HalfTruth datasets, respectively.


【3】NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

标题:非言语TTC:文本对齐非言语发声的公共英语数据库,具有情感注释,用于文本转语音

链接:http://arxiv.org/pdf/2507.13155v1

作者:risov, Egor Spirin, Daria Diatlova
摘要:目前的表达语音合成模型受到包含不同非言语发声(NV)的开源数据集有限的限制。在这项工作中,我们介绍了NonverbalTTS(NVTTS),这是一个17小时的开放获取数据集,注释了10种类型的NV(例如,笑声,咳嗽)和8个情绪类别。该数据集来自流行的来源,VoxCeleb和Exhibition,使用自动检测,然后进行人工验证。我们提出了一个综合的管道,集成了自动语音识别(ASR),NV标记,情感分类,和融合算法,以合并来自多个注释器的transmittance。在NVTTS数据集上微调开源文本到语音(TTS)模型,实现了与闭源系统(如CosyVoice2)的对等性,这是通过人工评估和自动度量(包括说话人相似性和NV保真度)来衡量的。通过发布NVTTS及其相应的注释指南,我们解决了表达性TTS研究中的一个关键瓶颈。该数据集可在https: huggingface.co datasets deepvk NonverbalTTS上获得。
摘要:

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour open-access dataset annotated with 10 types of NVs (e.g., laughter, coughs) and 8 emotional categories. The dataset is derived from popular sources, VoxCeleb and Expresso, using automated detection followed by human validation. We propose a comprehensive pipeline that integrates automatic speech recognition (ASR), NV tagging, emotion classification, and a fusion algorithm to merge transcriptions from multiple annotators. Fine-tuning open-source text-to-speech (TTS) models on the NVTTS dataset achieves parity with closed-source systems such as CosyVoice2, as measured by both human evaluation and automatic metrics, including speaker similarity and NV fidelity. By releasing NVTTS and its accompanying annotation guidelines, we address a key bottleneck in expressive TTS research. The dataset is available at https: huggingface.co datasets deepvk NonverbalTTS.


【4】Multi-Class-Token Transformer for Multitask Self-supervised Music Information Retrieval
标题:多任务自监督音乐信息检索的多类令牌Transformer
链接:http://arxiv.org/pdf/2507.12996v1

作者:ong, Vincent Lostanlen, Romain Hennequin, Mathieu Lagrange, Gabriel Meseguer-Brocal
摘要:对比学习和等变学习是自监督学习(SSL)音频内容分析的有效方法。然而,它们在音乐信息检索(MIR)中的应用面临着一个困境:前者在标记上更有效(例如,仪器识别)但对结构化预测不太有效(例如,音调估计);后者可以在其设计的特定任务上匹配监督方法,但它不能很好地推广到其他任务。在这篇文章中,我们采用了一种两全其美的方法,即同时在两种借口任务上训练深度神经网络。建议的新架构是一个Vision Transformer与1-D频谱图补丁(ViT-1D),配备了两个类令牌,这是专门为不同的自我监督的借口任务,但通过相同的模型进行优化:因此资格的自我监督多类令牌多任务(MT 2)。前一类令牌优化了交叉功率谱密度(CPSD),用于五分之一圈内的等变学习,而后者优化了归一化温度标度交叉熵(NT-Xent),用于对比学习。MT 2结合了两种借口任务的优势,并且始终优于使用对比或等变学习训练的单类令牌ViT-1D模型。平均两个类令牌进一步提高了几个任务的性能,突出了每个类令牌所学习的表示的互补性。此外,在最后一层的特征上使用相同的单线性层探测方法,MT 2在除节拍跟踪之外的所有任务上都优于MERT;由于其多任务处理能力,MT 2的参数减少了18倍。我们的SSL基准测试证明了我们的多类令牌多任务学习方法在MIR应用程序中的多功能性。
摘要:Contrastive learning and equivariant learning are effective methods for self-supervised learning (SSL) for audio content analysis. Yet, their application to music information retrieval (MIR) faces a dilemma: the former is more effective on tagging (e.g., instrument recognition) but less effective on structured prediction (e.g., tonality estimation); The latter can match supervised methods on the specific task it is designed for, but it does not generalize well to other tasks. In this article, we adopt a best-of-both-worlds approach by training a deep neural network on both kinds of pretext tasks at once. The proposed new architecture is a Vision Transformer with 1-D spectrogram patches (ViT-1D), equipped with two class tokens, which are specialized to different self-supervised pretext tasks but optimized through the same model: hence the qualification of self-supervised multi-class-token multitask (MT2). The former class token optimizes cross-power spectral density (CPSD) for equivariant learning over the circle of fifths, while the latter optimizes normalized temperature-scaled cross-entropy (NT-Xent) for contrastive learning. MT2 combines the strengths of both pretext tasks and outperforms consistently both single-class-token ViT-1D models trained with either contrastive or equivariant learning. Averaging the two class tokens further improves performance on several tasks, highlighting the complementary nature of the representations learned by each class token. Furthermore, using the same single-linear-layer probing method on the features of last layer, MT2 outperforms MERT on all tasks except for beat tracking; achieving this with 18x fewer parameters thanks to its multitasking capabilities. Our SSL benchmark demonstrates the versatility of our multi-class-token multitask learning approach for MIR applications.


【5】Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
标题:Enkidu:通用频率微扰,针对Voice Deepfakes的实时音频隐私保护
链接:http://arxiv.org/pdf/2507.12932v1

作者:Jiahao Chen, Chunyi Zhou, Yuwen Pu, Qingming Li, Tianyu Du, Shouling Ji
备注:Accepted by ACM MM 2025
摘要:语音Deepfake技术的快速发展引起了人们对用户音频隐私的严重担忧,因为攻击者越来越多地利用公开的语音数据来生成令人信服的虚假音频,用于身份盗窃,金融欺诈和错误信息活动等恶意目的。虽然现有的防御方法提供了部分保护,但它们面临着关键的限制,包括对不可见用户数据的适应性弱,对长音频的可扩展性差,严格依赖白盒知识,以及加密过程中的高计算和时间成本。为了应对这些挑战并抵御个性化的语音deepfake威胁,我们提出了Enkidu,这是一种新型的面向用户的隐私保护框架,它利用了通过黑盒知识和少量用户数据的Few-Shot训练生成的通用频率扰动。这些具有高度延展性的频域噪声补丁可以实现实时、轻量级的保护,在可变长度音频中具有强大的泛化能力,并对语音deepfake攻击具有强大的抵抗力,同时保持感知质量和语音清晰度。值得注意的是,与六种最先进的对策相比,Enkidu实现了超过50至200倍的处理内存效率(低至0.004千兆字节)和3至7000倍的运行时效率(实时系数低至0.004)。在六个主流的文本到语音模型和五个尖端的自动说话人验证模型上进行的广泛实验证明了Enkidu在防御普通和自适应语音deepfake攻击方面的有效性,可移植性和实用性。
摘要:The rapid advancement of voice deepfake technologies has raised serious concerns about user audio privacy, as attackers increasingly exploit publicly available voice data to generate convincing fake audio for malicious purposes such as identity theft, financial fraud, and misinformation campaigns. While existing defense methods offer partial protection, they face critical limitations, including weak adaptability to unseen user data, poor scalability to long audio, rigid reliance on white-box knowledge, and high computational and temporal costs during the encryption process. To address these challenges and defend against personalized voice deepfake threats, we propose Enkidu, a novel user-oriented privacy-preserving framework that leverages universal frequential perturbations generated through black-box knowledge and few-shot training on a small amount of user data. These highly malleable frequency-domain noise patches enable real-time, lightweight protection with strong generalization across variable-length audio and robust resistance to voice deepfake attacks, all while preserving perceptual quality and speech intelligibility. Notably, Enkidu achieves over 50 to 200 times processing memory efficiency (as low as 0.004 gigabytes) and 3 to 7000 times runtime efficiency (real-time coefficient as low as 0.004) compared to six state-of-the-art countermeasures. Extensive experiments across six mainstream text-to-speech models and five cutting-edge automated speaker verification models demonstrate the effectiveness, transferability, and practicality of Enkidu in defending against both vanilla and adaptive voice deepfake attacks.


【6】Best Practices and Considerations for Child Speech Corpus Collection and Curation in Educational, Clinical, and Forensic Scenarios

标题:在教育、临床和法医场景中收集和管理儿童语音语料库的最佳实践和考虑
链接:https://arxiv.org/abs/2507.12870

作者:en, Satwik Dutta, Ellen Grand
备注:5 pages, 0 figures, accepted at the 10th Workshop on Speech and Language Technology in Education (SLaTE 2025), a Satellite Workshop of the 2025 Interspeech Conference
摘要:孩子的语言能力会一直变化,直到他们成年。到7- 8岁时,他们的语音发展和语言结构迅速演变。他们的口语沟通技能和数据隐私的这种动态变化使得为儿童策划技术就绪的语音语料库具有挑战性。本研究旨在弥合这一差距,并为研究人员和从业人员提供基于预期目标开发此类语料库的最佳实践和考虑因素。虽然主要集中在教育目标,儿童语音数据的应用已经遍布包括临床和法医领域在内的各个领域。出于这一目标,我们描述了谁,什么,什么时候,在哪里收集数据的启发,以前的收集工作和我们的经验/知识。我们还提供了一个指南,以建立合作,信任和导航人类受试者的研究协议。本研究的结论是语料库质量检查,分类和注释的指导方针。
摘要:A child's spoken ability continues to change until their adult age. Until 7-8yrs, their speech sound development and language structure evolve rapidly. This dynamic shift in their spoken communication skills and data privacy make it challenging to curate technology-ready speech corpora for children. This study aims to bridge this gap and provide researchers and practitioners with the best practices and considerations for developing such a corpus based on an intended goal. Although primarily focused on educational goals, applications of child speech data have spread across fields including clinical and forensics fields. Motivated by this goal, we describe the WHO, WHAT, WHEN, and WHERE of data collection inspired by prior collection efforts and our experience/knowledge. We also provide a guide to establish collaboration, trust, and for navigating the human subjects research protocol. This study concludes with guidelines for corpus quality check, triage, and annotation.


【7】Autoregressive Speech Enhancement via Acoustic Tokens
标题:通过声学令牌的自回归语音增强
链接:https://arxiv.org/abs/2507.12825

作者:a Libera, Cem Subakan, Mirco Ravanelli
备注:5 pages, 2 figures
摘要:在语音处理管道中,提高真实世界录音的质量和清晰度至关重要。虽然监督回归是语音增强的主要方法,但音频标记化正在成为与其他模态平滑集成的有前途的替代方案。然而,利用离散表示进行语音增强的研究仍然有限。以前的工作主要集中在语义令牌,往往会丢弃关键的声学细节,如扬声器的身份。此外,这些研究通常采用非自回归模型,假设输出的条件独立性,忽略了自回归模型提供的潜在改进。为了解决这些差距,我们:1)对用于语音增强的声学令牌的性能进行全面研究,包括比特率和噪声强度的影响; 2)引入专门为此任务设计的基于换能器的自回归架构。在VoiceBank和Libri 1 Mix数据集上的实验表明,声学令牌在保留说话者身份方面优于语义令牌,并且我们的自回归方法可以进一步提高性能。尽管如此,我们观察到,离散表示仍然低于连续的,强调需要在这一领域进行进一步的研究。
摘要:In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising alternative for a smooth integration with other modalities. However, research on speech enhancement using discrete representations is still limited. Previous work has mainly focused on semantic tokens, which tend to discard key acoustic details such as speaker identity. Additionally, these studies typically employ non-autoregressive models, assuming conditional independence of outputs and overlooking the potential improvements offered by autoregressive modeling. To address these gaps we: 1) conduct a comprehensive study of the performance of acoustic tokens for speech enhancement, including the effect of bitrate and noise strength; 2) introduce a novel transducer-based autoregressive architecture specifically designed for this task. Experiments on VoiceBank and Libri1Mix datasets show that acoustic tokens outperform semantic tokens in terms of preserving speaker identity, and that our autoregressive approach can further improve performance. Nevertheless, we observe that discrete representations still fall short compared to continuous ones, highlighting the need for further research in this area.


【8】Large Language Models' Internal Perception of Symbolic Music
标题:大型语言模型对象征性音乐的内在感知
链接:https://arxiv.org/abs/2507.12808

作者:in, Kunitake Kaneko
摘要:大型语言模型(LLM)擅长建模自然语言中字符串之间的关系,并且在扩展到其他符号领域(如编码或数学)方面表现出了希望。然而,他们在多大程度上隐含的象征性音乐模型仍然未被探索。本文探讨了LLM如何通过从描述流派和风格组合的文本提示生成符号音乐数据来表示音乐概念,并通过识别和生成任务来评估其效用。我们生成一个LLM生成的MP3文件的数据集,而不依赖于明确的音乐训练。然后,我们完全在这个LLM生成的数据集上训练神经网络,并执行流派和风格分类以及旋律完成,将其性能与已建立的模型进行基准测试。我们的研究结果表明,LLM可以从文本中推断出基本的音乐结构和时间关系,突出了它们对音乐模式进行隐式编码的潜力,以及由于缺乏明确的音乐背景而导致的局限性,从而揭示了它们对符号音乐的生成能力。
摘要:Large language models (LLMs) excel at modeling relationships between strings in natural language and have shown promise in extending to other symbolic domains like coding or mathematics. However, the extent to which they implicitly model symbolic music remains underexplored. This paper investigates how LLMs represent musical concepts by generating symbolic music data from textual prompts describing combinations of genres and styles, and evaluating their utility through recognition and generation tasks. We produce a dataset of LLM-generated MIDI files without relying on explicit musical training. We then train neural networks entirely on this LLM-generated MIDI dataset and perform genre and style classification as well as melody completion, benchmarking their performance against established models. Our results demonstrate that LLMs can infer rudimentary musical structures and temporal relationships from text, highlighting both their potential to implicitly encode musical patterns and their limitations due to a lack of explicit musical context, shedding light on their generative capabilities for symbolic music.


【9】Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
标题:使用CNN-LSTM网络和基于MFCC的声学特征早期检测植物入侵的木材钻孔甲虫
链接:https://arxiv.org/abs/2507.12793

作者:n Sri Manukalpa, H. S. Bopage, W. A. M. Jayawardena, P. K. P. G. Panduwawala
备注:This is a preprint article
摘要:白蚁等结构性害虫对木结构建筑构成严重威胁,由于其隐蔽性和渐进性破坏,造成重大经济损失。传统的检测方法,如目视检查和化学处理,是侵入性的,劳动密集型的,并且对于早期侵染无效。为了弥补这一差距,本研究提出了一种基于非侵入性深度学习的声学分类框架,用于早期白蚁检测。我们的目标是开发一个强大的,可扩展的模型,区分白蚁产生的声音信号从背景噪声。我们介绍了一种混合卷积神经网络长短期记忆体系结构,捕捉白蚁活动的空间和时间特征。音频数据收集白蚁出没和干净的木材样品。我们提取Mel频率倒谱系数并训练CNN LSTM模型来分类信号。实验结果表明,高性能,94.5%的准确率,93.2%的精度,95.8%的召回率。比较分析表明,混合模型优于独立的CNN和LSTM架构,强调了其综合实力。值得注意的是,该模型产生的假阴性率很低,这对于及时干预至关重要。这项研究为早期白蚁检测提供了一种非侵入性的自动化解决方案,对改善害虫监测,最大限度地减少结构破坏以及房主和害虫控制专业人员更好地决策具有实际意义。未来的工作可能会集成物联网进行实时警报,并将检测扩展到其他结构性害虫。
摘要:Structural pests, such as termites, pose a serious threat to wooden buildings, resulting in significant economic losses due to their hidden and progressive damage. Traditional detection methods, such as visual inspections and chemical treatments, are invasive, labor intensive, and ineffective for early stage infestations. To bridge this gap, this study proposes a non invasive deep learning based acoustic classification framework for early termite detection. We aim to develop a robust, scalable model that distinguishes termite generated acoustic signals from background noise. We introduce a hybrid Convolutional Neural Network Long Short Term Memory architecture that captures both spatial and temporal features of termite activity. Audio data were collected from termite infested and clean wooden samples. We extracted Mel Frequency Cepstral Coefficients and trained the CNN LSTM model to classify the signals. Experimental results show high performance, with 94.5% accuracy, 93.2% precision, and 95.8% recall. Comparative analysis reveals that the hybrid model outperforms standalone CNN and LSTM architectures, underscoring its combined strength. Notably, the model yields low false-negative rates, which is essential for enabling timely intervention. This research contributes a non invasive, automated solution for early termite detection, with practical implications for improved pest monitoring, minimized structural damage, and better decision making by homeowners and pest control professionals. Future work may integrate IoT for real time alerts and extend detection to other structural pests.


【10】Sample-Constrained Black Box Optimization for Audio Personalization
标题:用于音频个性化的样本约束黑匣子优化
链接:https://arxiv.org/abs/2507.12773

作者: Rajagopalan, Yu-Lin Wei, Romit Roy Choudhury
备注:Published in AAAI 2024
摘要:我们考虑个性化音频以最大化用户体验的问题。简而言之,我们的目标是找到一个过滤器$h^*$,适用于任何音乐或语音,将最大限度地提高用户的满意度。这是一个黑盒优化问题,因为用户的满意度函数是未知的。在这个主题上已经做了大量的工作,其中关键思想是向用户播放音频样本,每个音频样本由不同的滤波器$h_i$整形,并向用户查询他们的满意度分数$f(h_i)$。一个家庭的"代理”功能,然后设计,以适应这些分数和优化方法逐渐完善这些功能,以达到过滤器$\hat{h}^*$,最大限度地提高满意度。在某些应用中,我们观察到第二种类型的查询是可能的,其中用户可以告诉我们最佳过滤器$h^*$的各个元素$h ^*[j]$。考虑一个与烹饪的类比,目标是烹饪一个最大化用户满意度的食谱。可以要求用户对各种烹饪食谱(例如,豆腐炒饭)或评分的个别成分(说,盐,糖,大米,鸡肉等)。给定$B$查询的预算,其中查询可以是任何一种类型,我们的目标是找到最大化用户满意度的配方。我们的建议建立在稀疏高斯过程回归(GPR)的基础上,并展示了混合方法如何优于任何一种类型的查询。我们的结果通过模拟和现实世界的实验,志愿者给音乐/语音音频的反馈,并能够实现高满意度进行验证。我们相信这种混合查询的想法在黑盒优化中打开了新的问题,解决方案可以使音频个性化之外的其他应用受益。
摘要:We consider the problem of personalizing audio to maximize user experience. Briefly, we aim to find a filter $h^*$, which applied to any music or speech, will maximize the user's satisfaction. This is a black-box optimization problem since the user's satisfaction function is unknown. Substantive work has been done on this topic where the key idea is to play audio samples to the user, each shaped by a different filter $h_i$, and query the user for their satisfaction scores $f(h_i)$. A family of ``surrogate" functions is then designed to fit these scores and the optimization method gradually refines these functions to arrive at the filter $\hat{h}^*$ that maximizes satisfaction. In certain applications, we observe that a second type of querying is possible where users can tell us the individual elements $h^*[j]$ of the optimal filter $h^*$. Consider an analogy from cooking where the goal is to cook a recipe that maximizes user satisfaction. A user can be asked to score various cooked recipes (e.g., tofu fried rice) or to score individual ingredients (say, salt, sugar, rice, chicken, etc.). Given a budget of $B$ queries, where a query can be of either type, our goal is to find the recipe that will maximize this user's satisfaction. Our proposal builds on Sparse Gaussian Process Regression (GPR) and shows how a hybrid approach can outperform any one type of querying. Our results are validated through simulations and real world experiments, where volunteers gave feedback on music/speech audio and were able to achieve high satisfaction levels. We believe this idea of hybrid querying opens new problems in black-box optimization and solutions can benefit other applications beyond audio personalization.


【11】Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
标题:合成视听伪造中真实音频恢复和篡改定位的跨模式水印
链接:https://arxiv.org/abs/2507.12723

作者:Kim, Sehwan Park, Sungmin Cha, Paul Hongsuck Seo
备注:5 pages, 2 figures, Interspeech 2025
摘要:语音克隆和唇同步模型的最新进展使合成视听语音(SAVF)成为可能,其中音频和视觉都被操纵以模仿目标说话者。这通过使虚假内容看起来真实来显著增加错误信息的风险。为了解决这个问题,现有方法检测或定位操纵,但不能恢复传达消息的语义内容的真实音频。这种限制降低了它们在打击视听误导方面的效力。在这项工作中,我们介绍了真实的音频恢复(AAR)和篡改定位音频(TLA)从SAVF的任务,并提出了一个跨模态水印框架嵌入真实的音频到视觉操作之前。这就实现了AAR、TLA和对错误信息的强大防御。大量实验证明,我们的方法在AAR和TLA中针对各种操作(包括语音克隆和嘴唇同步)具有强大的性能。
摘要:Recent advances in voice cloning and lip synchronization models have enabled Synthesized Audiovisual Forgeries (SAVFs), where both audio and visuals are manipulated to mimic a target speaker. This significantly increases the risk of misinformation by making fake content seem real. To address this issue, existing methods detect or localize manipulations but cannot recover the authentic audio that conveys the semantic content of the message. This limitation reduces their effectiveness in combating audiovisual misinformation. In this work, we introduce the task of Authentic Audio Recovery (AAR) and Tamper Localization in Audio (TLA) from SAVFs and propose a cross-modal watermarking framework to embed authentic audio into visuals before manipulation. This enables AAR, TLA, and a robust defense against misinformation. Extensive experiments demonstrate the strong performance of our method in AAR and TLA against various manipulations, including voice cloning and lip synchronization.


【12】AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
标题:AudioJudge:了解基于大型音频模型的语音评估中的有效方法
链接:https://arxiv.org/abs/2507.12705

作者:Manakul, Woody Haosheng Gan, Michael J. Ryan, Ali Sartaz Khan, Warit Sirichotedumrong, Kunat Pipatanakul, William Held, Diyi Yang
摘要:目前的语音评价受到两个关键的限制:需要和困难,设计专门的系统针对个人的音频特性,和自动评价方法和人类偏好之间的相关性差。这项工作提出了一个系统的研究大音频模型(LAM)作为一个法官,AudioJudge,调查它是否可以提供一个统一的评估框架,解决这两个挑战。我们系统地探索了AudioJudge在音频特征检测任务中的应用,包括发音、说话速率、说话人识别和语音质量,以及用于自动基准测试的系统级人类偏好模拟。我们研究了不同的提示工程策略,发现音频级联与上下文学习相结合,显着提高了音频特征检测和人类偏好模拟任务的性能。我们还引入了一个多方面的合奏AudioJudge,使通用的多方面的音频评估。该方法将语音评估分解为词汇内容,语音质量和语言特征的专业判断,在我们的系统排名基准上与人类偏好的相关性高达0.91斯皮尔曼。鲁棒性分析表明,虽然LAMs在声学噪声下保持强大的性能,但它们表现出显着的冗长性和位置偏差,需要仔细缓解。
摘要:Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor correlation between automatic evaluation methods and human preferences. This work presents a systematic study of Large Audio Model (LAM) as a Judge, AudioJudge, investigating whether it can provide a unified evaluation framework that addresses both challenges. We systematically explore AudioJudge across audio characteristic detection tasks, including pronunciation, speaking rate, speaker identification and speech quality, and system-level human preference simulation for automated benchmarking. We investigate different prompt engineering strategies, finding that audio concatenation combined with in-context learning significantly improves performance across both audio characteristic detection and human preference simulation tasks. We further introduce a multi-aspect ensemble AudioJudge to enable general-purpose multi-aspect audio evaluation. This method decomposes speech assessment into specialized judges for lexical content, speech quality, and paralinguistic features, achieving up to 0.91 Spearman correlation with human preferences on our system ranking benchmark. Robustness analysis reveals that while LAMs maintain strong performance under acoustic noise, they exhibit significant verbosity and positional biases that require careful mitigation.


【13】Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
标题:机器的特定任务音频编码:机器学习的潜在特征是该机器的代码
链接:https://arxiv.org/abs/2507.12701

作者: Kuznetsova, Inseon Jang, Wootaek Lim, Minje Kim
摘要:利用量化算法的神经音频编解码器已经显著影响了各种语音/音频任务。虽然高保真重建对于人类感知至关重要,但机器音频编码(ACoM)优先考虑有效的压缩和下游任务性能,而忽略感知的细微差别。这项工作介绍了一种有效的ACoM方法,可以压缩和压缩任何选定的中间特征表示已经训练的语音/音频下游模型。我们的方法采用任务特定的损失指导以及残余矢量量化(RVQ)损失,提供超低比特率(即,小于200 bps),下游模型性能的损失最小。由此产生的令牌化器可适应各种比特率和模型大小,以实现灵活的部署。在自动语音识别和音频分类上进行评估,我们的方法通过适当的正则化证明了其有效性和更广泛的任务和架构适用性的潜力。
摘要:Neural audio codecs, leveraging quantization algorithms, have significantly impacted various speech/audio tasks. While high-fidelity reconstruction is paramount for human perception, audio coding for machines (ACoM) prioritizes efficient compression and downstream task performance, disregarding perceptual nuances. This work introduces an efficient ACoM method that can compress and quantize any chosen intermediate feature representation of an already trained speech/audio downstream model. Our approach employs task-specific loss guidance alongside residual vector quantization (RVQ) losses, providing ultra-low bitrates (i.e., less than 200 bps) with a minimal loss of the downstream model performance. The resulting tokenizer is adaptable to various bitrates and model sizes for flexible deployment. Evaluated on automatic speech recognition and audio classification, our method demonstrates its efficacy and potential for broader task and architectural applicability through appropriate regularization.


【14】Keep the beat going: Automatic drum transcription with momentum
标题:保持节拍:有动力的自动鼓转录
链接:https://arxiv.org/abs/2507.12596

作者: Foster, Robert J. Webber
摘要:一个简单的,可解释的方法来执行自动鼓转录是通过使用部分固定的非负矩阵分解分解的幅度谱图的记录的音乐作品。有两种自然的方法来优化非负矩阵分解,包括乘法更新规则和动量投影梯度下降。这些方法在经验精度和理论收敛保证方面有所不同。本文总结了方法和它们的时间复杂性,并将方法应用于ENST鼓数据集和作者乐队的原始录音,评估了地面实况鼓注释的经验准确性。结果表明,带有动量的投影梯度下降算法在固定运行时间内具有更高的精度,并且满足更强的收敛保证。
摘要:A simple, interpretable way to perform automatic drum transcription is by factoring the magnitude spectrogram of a recorded musical piece using a partially fixed nonnegative matrix factorization. There are two natural ways to optimize the nonnegative matrix factorization, including a multiplicative update rule and projected gradient descent with momentum. The methods differ in their empirical accuracies and theoretical convergence guarantees. This paper summarizes the methods and their time complexities, and it applies the methods to the ENST-Drums data set and an original recording from the author's band, evaluating the empirical accuracy with respect to ground-truth drum annotations. The results indicate that projected gradient descent with momentum leads to higher accuracy for a fixed runtime, and it satisfies stronger convergence guarantees.


【15】Evaluation of Neural Surrogates for Physical Modelling Synthesis of Nonlinear Elastic Plates
标题:非线性弹性板物理模型综合的神经代理评价
链接:https://arxiv.org/abs/2507.12563

作者: La Vega Martin, Rodrigo Diaz Fernandez, Mark Sandler
摘要:物理建模合成旨在从振动结构的物理模拟生成音频。薄弹性板是鼓式膜的常见模型。传统的数值方法,如有限差分和有限元提供了高精度,但计算要求高,限制了它们在实时音频应用中的使用。本文提出了一种基于神经网络的方法来解决非线性弹性板的振动的比较分析。我们评估了几种最先进的模型,在短序列上训练,以自回归方式预测长序列。我们展示了这些模型的一些局限性,以及为什么不足以查看时域中的预测误差。我们讨论了实时音频合成的影响,并提出了未来的方向,以改善神经方法来模拟非线性振动。
摘要:Physical modelling synthesis aims to generate audio from physical simulations of vibrating structures. Thin elastic plates are a common model for drum membranes. Traditional numerical methods like finite differences and finite elements offer high accuracy but are computationally demanding, limiting their use in real-time audio applications. This paper presents a comparative analysis of neural network-based approaches for solving the vibration of nonlinear elastic plates. We evaluate several state-of-the-art models, trained on short sequences, for prediction of long sequences in an autoregressive fashion. We show some of the limitations of these models, and why is not enough to look at the prediction error in the time domain. We discuss the implications for real-time audio synthesis and propose future directions for improving neural approaches to model nonlinear vibration.


【16】AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
标题:AVFSNet:视听语音分离,实现灵活数量的扬声器,具有多规模和多任务学习
链接:https://arxiv.org/abs/2507.12972

作者:ang, Ying Wei
摘要:从含有灵活说话人数量的混合信号中分离目标语音是一项具有挑战性的任务。虽然现有的方法表现出很强的分离性能和噪声鲁棒性,他们主要假设先验知识的混合物中的扬声器计数。有限的研究解决未知的扬声器数量的情况下表现出显着的限制泛化能力在真实的声学环境。为了克服这些挑战,本文提出了AVFSNet -视听语音分离模型集成多尺度编码和并行架构-联合优化说话人计数和多说话人分离任务。该模型通过视觉信息的融合,在增强环境噪声适应性的同时,并行地独立分离每个说话人。全面的实验评估表明,AVFSNet在多个评估指标上实现了最先进的结果,并在不同的数据集上提供了出色的性能。
摘要:Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge of speaker counts in mixtures. The limited research addressing unknown speaker quantity scenarios exhibits significantly constrained generalization capabilities in real acoustic environments. To overcome these challenges, this paper proposes AVFSNet -- an audio-visual speech separation model integrating multi-scale encoding and parallel architecture -- jointly optimized for speaker counting and multi-speaker separation tasks. The model independently separates each speaker in parallel while enhancing environmental noise adaptability through visual information integration. Comprehensive experimental evaluations demonstrate that AVFSNet achieves state-of-the-art results across multiple evaluation metrics and delivers outstanding performance on diverse datasets.


【17】UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
标题:UniSLU:来自异类跨任务数据集的统一口语理解
链接:https://arxiv.org/abs/2507.12951

作者:heng, Shilin Zhou, Chen Gong, Zhenghua Li
备注:13 pages, 3 figures
摘要:口语理解(SLU)在以语音为中心的多媒体应用中起着至关重要的作用,使机器能够在会议、面试和客户服务交互等场景中理解口语。SLU包含多个任务,包括自动语音识别(ASR),口语命名实体识别(NER)和口语情感分析(SA)。然而,现有的方法通常依赖于单独的模型架构,如口语NER和SA,这增加了系统的复杂性,限制了跨任务的交互,并未能充分利用跨任务可用的异构数据集。为了解决这些限制,我们提出了UniSLU,这是一个统一的框架,可以在单个架构中联合建模多个SLU任务。具体来说,我们为不同的SLU任务提出了一个统一的表示,从而能够在多个任务中充分利用异构数据集。基于这种表示,我们提出了一种统一的生成方法,该方法联合建模ASR,口语NER和SA任务,增强任务交互,并实现与大型语言模型的无缝集成,以利用其强大的生成能力。在公共SLU数据集上的大量实验证明了我们方法的有效性,与几种基准方法相比,SLU性能优越,非常适合真实世界中基于语音的多媒体场景。我们将在github上发布所有代码和模型,以方便未来的研究。
摘要:Spoken Language Understanding (SLU) plays a crucial role in speech-centric multimedia applications, enabling machines to comprehend spoken language in scenarios such as meetings, interviews, and customer service interactions. SLU encompasses multiple tasks, including Automatic Speech Recognition (ASR), spoken Named Entity Recognition (NER), and spoken Sentiment Analysis (SA). However, existing methods often rely on separate model architectures for individual tasks such as spoken NER and SA, which increases system complexity, limits cross-task interaction, and fails to fully exploit heterogeneous datasets available across tasks. To address these limitations, we propose UniSLU, a unified framework that jointly models multiple SLU tasks within a single architecture. Specifically, we propose a unified representation for diverse SLU tasks, enabling full utilization of heterogeneous datasets across multiple tasks. Built upon this representation, we propose a unified generative method that jointly models ASR, spoken NER, and SA tasks, enhancing task interactions and enabling seamless integration with large language models to harness their powerful generative capabilities. Extensive experiments on public SLU datasets demonstrate the effectiveness of our approach, achieving superior SLU performance compared to several benchmark methods, making it well-suited for real-world speech-based multimedia scenarios. We will release all code and models at github to facilitate future research.


【18】DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
标题:迪夫节奏+:可控且灵活的全长歌曲生成,具有偏好优化
链接:https://arxiv.org/abs/2507.12890

作者:hen, Yuepeng Jiang, Guobin Ma, Chunbo Hao, Shuai Wang, Jixun Yao, Ziqian Ning, Meng Meng, Jian Luan, Lei Xie
摘要:歌曲作为音乐艺术的一种核心形式,体现了人类丰富的智慧和创造力。虽然生成建模的最新进展使长格式歌曲生成取得了显着进展,但当前的全长歌曲合成系统仍然面临着重大挑战,包括数据不平衡,可控性不足和音乐质量不一致。DiffRhythm是一个开创性的基于扩散的模型,通过生成具有富有表现力的人声和伴奏的全长歌曲来推进该领域。然而,它的性能受到不平衡的模型训练数据集和对音乐风格的有限控制的限制,导致明显的质量差异和有限的创作灵活性。为了解决这些限制,我们提出了DiffRhythm+,一个增强的基于扩散的框架,用于可控和灵活的全长歌曲生成。DiffRhythm+利用大幅扩展和平衡的训练数据集来缓解歌词重复和遗漏等问题,同时还培养了更丰富的音乐技巧和表现力。该框架引入了多模态风格调节策略,使用户能够通过描述性文本和参考音频精确地指定音乐风格,从而显着增强创意控制和多样性。我们进一步引入了与用户偏好相一致的直接性能优化,引导模型在评估指标中实现一致的首选输出。大量的实验表明,DiffRhythm+实现了显着的改善自然,安排的复杂性,和听众满意度比以前的系统。
摘要:Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current systems for full-length song synthesis still face major challenges, including data imbalance, insufficient controllability, and inconsistent musical quality. DiffRhythm, a pioneering diffusion-based model, advanced the field by generating full-length songs with expressive vocals and accompaniment. However, its performance was constrained by an unbalanced model training dataset and limited controllability over musical style, resulting in noticeable quality disparities and restricted creative flexibility. To address these limitations, we propose DiffRhythm+, an enhanced diffusion-based framework for controllable and flexible full-length song generation. DiffRhythm+ leverages a substantially expanded and balanced training dataset to mitigate issues such as repetition and omission of lyrics, while also fostering the emergence of richer musical skills and expressiveness. The framework introduces a multi-modal style conditioning strategy, enabling users to precisely specify musical styles through both descriptive text and reference audio, thereby significantly enhancing creative control and diversity. We further introduce direct performance optimization aligned with user preferences, guiding the model toward consistently preferred outputs across evaluation metrics. Extensive experiments demonstrate that DiffRhythm+ achieves significant improvements in naturalness, arrangement complexity, and listener satisfaction over previous systems.



eess.AS音频处理


【1】AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
标题:AVFSNet:视听语音分离,实现灵活数量的扬声器,具有多规模和多任务学习
链接:https://arxiv.org/abs/2507.12972

作者:ang, Ying Wei
摘要:从含有灵活说话人数量的混合信号中分离目标语音是一项具有挑战性的任务。虽然现有的方法表现出很强的分离性能和噪声鲁棒性,他们主要假设先验知识的混合物中的扬声器计数。有限的研究解决未知的扬声器数量的情况下表现出显着的限制泛化能力在真实的声学环境。为了克服这些挑战,本文提出了AVFSNet -视听语音分离模型集成多尺度编码和并行架构-联合优化说话人计数和多说话人分离任务。该模型通过视觉信息的融合,在增强环境噪声适应性的同时,并行地独立分离每个说话人。全面的实验评估表明,AVFSNet在多个评估指标上实现了最先进的结果,并在不同的数据集上提供了出色的性能。
摘要:Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge of speaker counts in mixtures. The limited research addressing unknown speaker quantity scenarios exhibits significantly constrained generalization capabilities in real acoustic environments. To overcome these challenges, this paper proposes AVFSNet -- an audio-visual speech separation model integrating multi-scale encoding and parallel architecture -- jointly optimized for speaker counting and multi-speaker separation tasks. The model independently separates each speaker in parallel while enhancing environmental noise adaptability through visual information integration. Comprehensive experimental evaluations demonstrate that AVFSNet achieves state-of-the-art results across multiple evaluation metrics and delivers outstanding performance on diverse datasets.


【2】UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
标题:UniSLU:来自异类跨任务数据集的统一口语理解
链接:https://arxiv.org/abs/2507.12951

作者:heng, Shilin Zhou, Chen Gong, Zhenghua Li
备注:13 pages, 3 figures
摘要:口语理解(SLU)在以语音为中心的多媒体应用中起着至关重要的作用,使机器能够在会议、面试和客户服务交互等场景中理解口语。SLU包含多个任务,包括自动语音识别(ASR),口语命名实体识别(NER)和口语情感分析(SA)。然而,现有的方法通常依赖于单独的模型架构,如口语NER和SA,这增加了系统的复杂性,限制了跨任务的交互,并未能充分利用跨任务可用的异构数据集。为了解决这些限制,我们提出了UniSLU,这是一个统一的框架,可以在单个架构中联合建模多个SLU任务。具体来说,我们为不同的SLU任务提出了一个统一的表示,从而能够在多个任务中充分利用异构数据集。基于这种表示,我们提出了一种统一的生成方法,该方法联合建模ASR,口语NER和SA任务,增强任务交互,并实现与大型语言模型的无缝集成,以利用其强大的生成能力。在公共SLU数据集上的大量实验证明了我们方法的有效性,与几种基准方法相比,SLU性能优越,非常适合真实世界中基于语音的多媒体场景。我们将在github上发布所有代码和模型,以方便未来的研究。
摘要:Spoken Language Understanding (SLU) plays a crucial role in speech-centric multimedia applications, enabling machines to comprehend spoken language in scenarios such as meetings, interviews, and customer service interactions. SLU encompasses multiple tasks, including Automatic Speech Recognition (ASR), spoken Named Entity Recognition (NER), and spoken Sentiment Analysis (SA). However, existing methods often rely on separate model architectures for individual tasks such as spoken NER and SA, which increases system complexity, limits cross-task interaction, and fails to fully exploit heterogeneous datasets available across tasks. To address these limitations, we propose UniSLU, a unified framework that jointly models multiple SLU tasks within a single architecture. Specifically, we propose a unified representation for diverse SLU tasks, enabling full utilization of heterogeneous datasets across multiple tasks. Built upon this representation, we propose a unified generative method that jointly models ASR, spoken NER, and SA tasks, enhancing task interactions and enabling seamless integration with large language models to harness their powerful generative capabilities. Extensive experiments on public SLU datasets demonstrate the effectiveness of our approach, achieving superior SLU performance compared to several benchmark methods, making it well-suited for real-world speech-based multimedia scenarios. We will release all code and models at github to facilitate future research.


【3】DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
标题:迪夫节奏+:可控且灵活的全长歌曲生成,具有偏好优化
链接:https://arxiv.org/abs/2507.12890

作者:hen, Yuepeng Jiang, Guobin Ma, Chunbo Hao, Shuai Wang, Jixun Yao, Ziqian Ning, Meng Meng, Jian Luan, Lei Xie
摘要:歌曲作为音乐艺术的一种核心形式,体现了人类丰富的智慧和创造力。虽然生成建模的最新进展使长格式歌曲生成取得了显着进展,但当前的全长歌曲合成系统仍然面临着重大挑战,包括数据不平衡,可控性不足和音乐质量不一致。DiffRhythm是一个开创性的基于扩散的模型,通过生成具有富有表现力的人声和伴奏的全长歌曲来推进该领域。然而,它的性能受到不平衡的模型训练数据集和对音乐风格的有限控制的限制,导致明显的质量差异和有限的创作灵活性。为了解决这些限制,我们提出了DiffRhythm+,一个增强的基于扩散的框架,用于可控和灵活的全长歌曲生成。DiffRhythm+利用大幅扩展和平衡的训练数据集来缓解歌词重复和遗漏等问题,同时还培养了更丰富的音乐技巧和表现力。该框架引入了多模态风格调节策略,使用户能够通过描述性文本和参考音频精确地指定音乐风格,从而显着增强创意控制和多样性。我们进一步引入了与用户偏好相一致的直接性能优化,引导模型在评估指标中实现一致的首选输出。大量的实验表明,DiffRhythm+实现了显着的改善自然,安排的复杂性,和听众满意度比以前的系统。
摘要:Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current systems for full-length song synthesis still face major challenges, including data imbalance, insufficient controllability, and inconsistent musical quality. DiffRhythm, a pioneering diffusion-based model, advanced the field by generating full-length songs with expressive vocals and accompaniment. However, its performance was constrained by an unbalanced model training dataset and limited controllability over musical style, resulting in noticeable quality disparities and restricted creative flexibility. To address these limitations, we propose DiffRhythm+, an enhanced diffusion-based framework for controllable and flexible full-length song generation. DiffRhythm+ leverages a substantially expanded and balanced training dataset to mitigate issues such as repetition and omission of lyrics, while also fostering the emergence of richer musical skills and expressiveness. The framework introduces a multi-modal style conditioning strategy, enabling users to precisely specify musical styles through both descriptive text and reference audio, thereby significantly enhancing creative control and diversity. We further introduce direct performance optimization aligned with user preferences, guiding the model toward consistently preferred outputs across evaluation metrics. Extensive experiments demonstrate that DiffRhythm+ achieves significant improvements in naturalness, arrangement complexity, and listener satisfaction over previous systems.


【4】Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
标题:通过协作多语言语音基础模型增强域内和域外语音造假检测
链接:https://arxiv.org/abs/2507.12595

作者:etia Phukan, Mohd Mujtaba Akhtar, Girish, Arun Balaji Buduru
摘要:在这项工作中,我们解决了EmoFake检测(EFD)问题。我们假设多语言语音基础模型(SFM)对EFD特别有效,因为它们在不同语言中进行了预训练,能够对音高,音调和强度的变化进行细致入微的理解。为了验证这一点,我们进行了全面的比较分析国家的最先进的(SOTA)的可持续森林管理。我们的研究结果表明,多语言的SFMs相同的语言(域内),以及跨语言(域外)的评价的优越性。为了我们的目的,我们还建议,THAMA融合的基础模型(FM)的相关研究的动机,结合FM已显示出更好的性能。THAMA利用Tucker分解和Hadamard乘积的互补结合进行有效融合。通过THAMA,与合作的多语言SFM协同,在域内和域外设置中实现了最高性能,优于单个FM,基线融合技术和先前的SOTA方法。
摘要:In this work, we address EmoFake Detection (EFD). We hypothesize that multilingual speech foundation models (SFMs) will be particularly effective for EFD due to their pre-training across diverse languages, enabling a nuanced understanding of variations in pitch, tone, and intensity. To validate this, we conduct a comprehensive comparative analysis of state-of-the-art (SOTA) SFMs. Our results shows the superiority of multilingual SFMs for same language (in-domain) as well as cross-lingual (out-domain) evaluation. To our end, we also propose, THAMA for fusion of foundation models (FMs) motivated by related research where combining FMs have shown improved performance. THAMA leverages the complementary conjunction of tucker decomposition and hadamard product for effective fusion. With THAMA, synergized with cooperative multilingual SFMs achieves topmost performance across in-domain and out-domain settings, outperforming individual FMs, baseline fusion techniques, and prior SOTA methods.


【5】Voxtral
标题:沃克斯特拉尔
链接:https://arxiv.org/abs/2507.13264

作者: H. Liu, Andy Ehrenberg, Andy Lo, Clément Denoix, Corentin Barreau, Guillaume Lample, Jean-Malo Delignon, Khyathi Raghavi Chandu, Patrick von Platen, Pavankumar Reddy Muddireddy, Sanchit Gandhi, Soham Ghosh, Srijan Mishra, Thomas Foubert, Abhinav Rastogi, Adam Yang, Albert Q. Jiang, Alexandre Sablayrolles, Amélie Héliou, Amélie Martin, Anmol Agarwal, Antoine Roux, Arthur Darcet, Arthur Mensch, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Darius Dabert, Devendra Singh Chaplot, Devon Mizelle, Diego de las Casas, Elliot Chane-Sane, Emilien Fugier, Emma Bou Hanna, Gabrielle Berrada, Gauthier Delerce, Gauthier Guinet, Georgii Novikov, Guillaume Martin, Himanshu Jaju, Jan Ludziejewski, Jason Rute, Jean-Hadrien Chabran, Jessica Chudnovsky, Joachim Studnia, Joep Barmentlo, Jonas Amar, Josselin Somerville Roberts, Julien Denize, Karan Saxena, Karmesh Yadav, Kartik Khandelwal, Kush Jain, Lélio Renard Lavaud, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Marie Pellat, Mathilde Guillaumin, Mathis Felardos, Matthieu Dinot, Maxime Darrin, Maximilian Augustin, Mickaël Seznec, Neha Gupta, Nikhil Raghuraman, Olivier Duchenne, Patricia Wang, Patryk Saffer, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Rémi Delacourt, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Sandeep Subramanian, Saurabh Garg, Shashwat Dalal, Siddharth Gandhi, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Thibault Schueller, Thibaut Lavril, Thomas Robert, Thomas Wang, Timothée Lacroix, Tom Bewley, Valeriia Nemychnikova, Victor Paltz
备注:17 pages
摘要:我们提出Voxtral Mini和Voxtral Small,两个多模式音频聊天模型。Voxtral经过培训,能够理解语音和文本文档,在各种音频基准测试中实现最先进的性能,同时保持强大的文本功能。Voxtral Small的性能优于许多闭源模型,同时足够小,可以在本地运行。32K上下文窗口使模型能够处理长达40分钟的音频文件和长时间的多回合对话。我们还提供了三个基准评估语音理解模型的知识和琐事。这两个Voxtral模型都是在Apache 2.0许可下发布的。
摘要:We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio benchmarks, while preserving strong text capabilities. Voxtral Small outperforms a number of closed-source models, while being small enough to run locally. A 32K context window enables the model to handle audio files up to 40 minutes in duration and long multi-turn conversations. We also contribute three benchmarks for evaluating speech understanding models on knowledge and trivia. Both Voxtral models are released under Apache 2.0 license.


【6】Automatically assessing oral narratives of Afrikaans and isiXhosa children
标题:自动评估南非荷兰语和西科萨语儿童的口头叙述
链接:https://arxiv.org/abs/2507.13205

作者:1), E. Sharratt (1), F. de Wet (1), C. Jacobs (1), A. Smith (1), H. Kamper (1) ((1) Stellenbosch University)
备注:Accepted to SLaTE 2025
摘要:在幼儿期发展叙述和理解技能对以后的识字至关重要。然而,大型学前班的教师很难准确识别需要干预的学生。我们提出了一个系统,自动评估在南非荷兰语和isiXhosa的学龄前儿童的口头叙述。该系统使用自动语音识别,然后使用机器学习评分模型来预测叙事和理解分数。对于评分预测成绩单,我们比较了线性模型和大型语言模型(LLM)。基于LLM的系统在大多数情况下优于线性模型,但线性系统尽管简单,但仍具有竞争力。基于LLM的系统相当于人类专家标记需要干预的儿童。我们为课堂上的自动口语评估奠定了基础,使教师能够专注于为儿童学习提供个性化支持。
摘要:Developing narrative and comprehension skills in early childhood is critical for later literacy. However, teachers in large preschool classrooms struggle to accurately identify students who require intervention. We present a system for automatically assessing oral narratives of preschool children in Afrikaans and isiXhosa. The system uses automatic speech recognition followed by a machine learning scoring model to predict narrative and comprehension scores. For scoring predicted transcripts, we compare a linear model to a large language model (LLM). The LLM-based system outperforms the linear model in most cases, but the linear system is competitive despite its simplicity. The LLM-based system is comparable to a human expert in flagging children who require intervention. We lay the foundation for automatic oral assessments in classrooms, giving teachers extra capacity to focus on personalised support for children's learning.


【7】SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks
标题:SHIELD:一种安全且高度增强的集成学习,用于针对对抗性攻击的稳健Deepfake检测
链接:https://arxiv.org/abs/2507.13170

作者:in, Awais Khan, Muhammad Umar Farooq, Khalid Malik
摘要:音频在说话人验证、支持语音的智能设备和音频会议等应用中起着至关重要的作用。然而,音频操纵,如deepfakes,通过传播错误信息带来了重大风险。我们的实证分析表明,现有的检测deepfake音频的方法往往容易受到反取证(AF)攻击,特别是那些使用生成对抗网络的攻击。在本文中,我们提出了一种新的协作学习方法SHIELD来防御生成式AF攻击。为了暴露AF签名,我们集成了一个辅助生成模型,称为防御(DF)生成模型,它通过结合输入和输出来促进协作学习。此外,我们设计了一个三元组模型来捕获真实和AF攻击音频与真实生成和攻击生成音频的相关性,使用辅助生成模型。所提出的SHIELD增强了对生成AF攻击的防御,并在各种生成模型中实现了鲁棒性能。对于三种不同的生成模型,所提出的AF将ASVspoof 2019的平均检测准确率从95.49%降至59.77%,将In-the-Wild的平均检测准确率从99.44%降至38.45%,将HalfTruth的平均检测准确率从98.41%降至51.18%。所提出的SHIELD机制对AF攻击具有鲁棒性,并且在ASVspoof 2019,In-the-Wild和HalfTruth数据集的匹配设置中分别实现了98.13%,98.58%和99.57%的平均准确率,以及98.78%,98.62%和98.85%的不匹配设置。
摘要:Audio plays a crucial role in applications like speaker verification, voice-enabled smart devices, and audio conferencing. However, audio manipulations, such as deepfakes, pose significant risks by enabling the spread of misinformation. Our empirical analysis reveals that existing methods for detecting deepfake audio are often vulnerable to anti-forensic (AF) attacks, particularly those attacked using generative adversarial networks. In this article, we propose a novel collaborative learning method called SHIELD to defend against generative AF attacks. To expose AF signatures, we integrate an auxiliary generative model, called the defense (DF) generative model, which facilitates collaborative learning by combining input and output. Furthermore, we design a triplet model to capture correlations for real and AF attacked audios with real-generated and attacked-generated audios using auxiliary generative models. The proposed SHIELD strengthens the defense against generative AF attacks and achieves robust performance across various generative models. The proposed AF significantly reduces the average detection accuracy from 95.49% to 59.77% for ASVspoof2019, from 99.44% to 38.45% for In-the-Wild, and from 98.41% to 51.18% for HalfTruth for three different generative models. The proposed SHIELD mechanism is robust against AF attacks and achieves an average accuracy of 98.13%, 98.58%, and 99.57% in match, and 98.78%, 98.62%, and 98.85% in mismatch settings for the ASVspoof2019, In-the-Wild, and HalfTruth datasets, respectively.


【8】Best Practices and Considerations for Child Speech Corpus Collection and Curation in Educational, Clinical, and Forensic Scenarios
标题:在教育、临床和法医场景中收集和管理儿童语音语料库的最佳实践和考虑
链接:https://arxiv.org/abs/2507.12870

作者:en, Satwik Dutta, Ellen Grand
备注:5 pages, 0 figures, accepted at the 10th Workshop on Speech and Language Technology in Education (SLaTE 2025), a Satellite Workshop of the 2025 Interspeech Conference
摘要:孩子的语言能力会一直变化,直到他们成年。到7- 8岁时,他们的语音发展和语言结构迅速演变。他们的口语沟通技能和数据隐私的这种动态变化使得为儿童策划技术就绪的语音语料库具有挑战性。本研究旨在弥合这一差距,并为研究人员和从业人员提供基于预期目标开发此类语料库的最佳实践和考虑因素。虽然主要集中在教育目标,儿童语音数据的应用已经遍布包括临床和法医领域在内的各个领域。出于这一目标,我们描述了谁,什么,什么时候,在哪里收集数据的启发,以前的收集工作和我们的经验/知识。我们还提供了一个指南,以建立合作,信任和导航人类受试者的研究协议。本研究的结论是语料库质量检查,分类和注释的指导方针。
摘要:A child's spoken ability continues to change until their adult age. Until 7-8yrs, their speech sound development and language structure evolve rapidly. This dynamic shift in their spoken communication skills and data privacy make it challenging to curate technology-ready speech corpora for children. This study aims to bridge this gap and provide researchers and practitioners with the best practices and considerations for developing such a corpus based on an intended goal. Although primarily focused on educational goals, applications of child speech data have spread across fields including clinical and forensics fields. Motivated by this goal, we describe the WHO, WHAT, WHEN, and WHERE of data collection inspired by prior collection efforts and our experience/knowledge. We also provide a guide to establish collaboration, trust, and for navigating the human subjects research protocol. This study concludes with guidelines for corpus quality check, triage, and annotation.


【9】Autoregressive Speech Enhancement via Acoustic Tokens
标题:通过声学令牌的自回归语音增强
链接:https://arxiv.org/abs/2507.12825

作者:a Libera, Cem Subakan, Mirco Ravanelli
备注:5 pages, 2 figures
摘要:在语音处理管道中,提高真实世界录音的质量和清晰度至关重要。虽然监督回归是语音增强的主要方法,但音频标记化正在成为与其他模态平滑集成的有前途的替代方案。然而,利用离散表示进行语音增强的研究仍然有限。以前的工作主要集中在语义令牌,往往会丢弃关键的声学细节,如扬声器的身份。此外,这些研究通常采用非自回归模型,假设输出的条件独立性,忽略了自回归模型提供的潜在改进。为了解决这些差距,我们:1)对用于语音增强的声学令牌的性能进行全面研究,包括比特率和噪声强度的影响; 2)引入专门为此任务设计的基于换能器的自回归架构。在VoiceBank和Libri 1 Mix数据集上的实验表明,声学令牌在保留说话者身份方面优于语义令牌,并且我们的自回归方法可以进一步提高性能。尽管如此,我们观察到,离散表示仍然低于连续的,强调需要在这一领域进行进一步的研究。
摘要:In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising alternative for a smooth integration with other modalities. However, research on speech enhancement using discrete representations is still limited. Previous work has mainly focused on semantic tokens, which tend to discard key acoustic details such as speaker identity. Additionally, these studies typically employ non-autoregressive models, assuming conditional independence of outputs and overlooking the potential improvements offered by autoregressive modeling. To address these gaps we: 1) conduct a comprehensive study of the performance of acoustic tokens for speech enhancement, including the effect of bitrate and noise strength; 2) introduce a novel transducer-based autoregressive architecture specifically designed for this task. Experiments on VoiceBank and Libri1Mix datasets show that acoustic tokens outperform semantic tokens in terms of preserving speaker identity, and that our autoregressive approach can further improve performance. Nevertheless, we observe that discrete representations still fall short compared to continuous ones, highlighting the need for further research in this area.


【10】Large Language Models' Internal Perception of Symbolic Music
标题:大型语言模型对象征性音乐的内在感知
链接:https://arxiv.org/abs/2507.12808

作者:in, Kunitake Kaneko
摘要:大型语言模型(LLM)擅长建模自然语言中字符串之间的关系,并且在扩展到其他符号领域(如编码或数学)方面表现出了希望。然而,他们在多大程度上隐含的象征性音乐模型仍然未被探索。本文探讨了LLM如何通过从描述流派和风格组合的文本提示生成符号音乐数据来表示音乐概念,并通过识别和生成任务来评估其效用。我们生成一个LLM生成的MP3文件的数据集,而不依赖于明确的音乐训练。然后,我们完全在这个LLM生成的数据集上训练神经网络,并执行流派和风格分类以及旋律完成,将其性能与已建立的模型进行基准测试。我们的研究结果表明,LLM可以从文本中推断出基本的音乐结构和时间关系,突出了它们对音乐模式进行隐式编码的潜力,以及由于缺乏明确的音乐背景而导致的局限性,从而揭示了它们对符号音乐的生成能力。
摘要:Large language models (LLMs) excel at modeling relationships between strings in natural language and have shown promise in extending to other symbolic domains like coding or mathematics. However, the extent to which they implicitly model symbolic music remains underexplored. This paper investigates how LLMs represent musical concepts by generating symbolic music data from textual prompts describing combinations of genres and styles, and evaluating their utility through recognition and generation tasks. We produce a dataset of LLM-generated MIDI files without relying on explicit musical training. We then train neural networks entirely on this LLM-generated MIDI dataset and perform genre and style classification as well as melody completion, benchmarking their performance against established models. Our results demonstrate that LLMs can infer rudimentary musical structures and temporal relationships from text, highlighting both their potential to implicitly encode musical patterns and their limitations due to a lack of explicit musical context, shedding light on their generative capabilities for symbolic music.


【11】Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
标题:使用CNN-LSTM网络和基于MFCC的声学特征早期检测植物入侵的木材钻孔甲虫
链接:https://arxiv.org/abs/2507.12793

作者:n Sri Manukalpa, H. S. Bopage, W. A. M. Jayawardena, P. K. P. G. Panduwawala
备注:This is a preprint article
摘要:白蚁等结构性害虫对木结构建筑构成严重威胁,由于其隐蔽性和渐进性破坏,造成重大经济损失。传统的检测方法,如目视检查和化学处理,是侵入性的,劳动密集型的,并且对于早期侵染无效。为了弥补这一差距,本研究提出了一种基于非侵入性深度学习的声学分类框架,用于早期白蚁检测。我们的目标是开发一个强大的,可扩展的模型,区分白蚁产生的声音信号从背景噪声。我们介绍了一种混合卷积神经网络长短期记忆体系结构,捕捉白蚁活动的空间和时间特征。音频数据收集白蚁出没和干净的木材样品。我们提取Mel频率倒谱系数并训练CNN LSTM模型来分类信号。实验结果表明,高性能,94.5%的准确率,93.2%的精度,95.8%的召回率。比较分析表明,混合模型优于独立的CNN和LSTM架构,强调了其综合实力。值得注意的是,该模型产生的假阴性率很低,这对于及时干预至关重要。这项研究为早期白蚁检测提供了一种非侵入性的自动化解决方案,对改善害虫监测,最大限度地减少结构破坏以及房主和害虫控制专业人员更好地决策具有实际意义。未来的工作可能会集成物联网进行实时警报,并将检测扩展到其他结构性害虫。
摘要:Structural pests, such as termites, pose a serious threat to wooden buildings, resulting in significant economic losses due to their hidden and progressive damage. Traditional detection methods, such as visual inspections and chemical treatments, are invasive, labor intensive, and ineffective for early stage infestations. To bridge this gap, this study proposes a non invasive deep learning based acoustic classification framework for early termite detection. We aim to develop a robust, scalable model that distinguishes termite generated acoustic signals from background noise. We introduce a hybrid Convolutional Neural Network Long Short Term Memory architecture that captures both spatial and temporal features of termite activity. Audio data were collected from termite infested and clean wooden samples. We extracted Mel Frequency Cepstral Coefficients and trained the CNN LSTM model to classify the signals. Experimental results show high performance, with 94.5% accuracy, 93.2% precision, and 95.8% recall. Comparative analysis reveals that the hybrid model outperforms standalone CNN and LSTM architectures, underscoring its combined strength. Notably, the model yields low false-negative rates, which is essential for enabling timely intervention. This research contributes a non invasive, automated solution for early termite detection, with practical implications for improved pest monitoring, minimized structural damage, and better decision making by homeowners and pest control professionals. Future work may integrate IoT for real time alerts and extend detection to other structural pests.


【12】Sample-Constrained Black Box Optimization for Audio Personalization
标题:用于音频个性化的样本约束黑匣子优化
链接:https://arxiv.org/abs/2507.12773

作者: Rajagopalan, Yu-Lin Wei, Romit Roy Choudhury
备注:Published in AAAI 2024
摘要:我们考虑个性化音频以最大化用户体验的问题。简而言之,我们的目标是找到一个过滤器$h^*$,适用于任何音乐或语音,将最大限度地提高用户的满意度。这是一个黑盒优化问题,因为用户的满意度函数是未知的。在这个主题上已经做了大量的工作,其中关键思想是向用户播放音频样本,每个音频样本由不同的滤波器$h_i$整形,并向用户查询他们的满意度分数$f(h_i)$。一个家庭的"代理”功能,然后设计,以适应这些分数和优化方法逐渐完善这些功能,以达到过滤器$\hat{h}^*$,最大限度地提高满意度。在某些应用中,我们观察到第二种类型的查询是可能的,其中用户可以告诉我们最佳过滤器$h^*$的各个元素$h ^*[j]$。考虑一个与烹饪的类比,目标是烹饪一个最大化用户满意度的食谱。可以要求用户对各种烹饪食谱(例如,豆腐炒饭)或评分的个别成分(说,盐,糖,大米,鸡肉等)。给定$B$查询的预算,其中查询可以是任何一种类型,我们的目标是找到最大化用户满意度的配方。我们的建议建立在稀疏高斯过程回归(GPR)的基础上,并展示了混合方法如何优于任何一种类型的查询。我们的结果通过模拟和现实世界的实验,志愿者给音乐/语音音频的反馈,并能够实现高满意度进行验证。我们相信这种混合查询的想法在黑盒优化中打开了新的问题,解决方案可以使音频个性化之外的其他应用受益。
摘要:We consider the problem of personalizing audio to maximize user experience. Briefly, we aim to find a filter $h^*$, which applied to any music or speech, will maximize the user's satisfaction. This is a black-box optimization problem since the user's satisfaction function is unknown. Substantive work has been done on this topic where the key idea is to play audio samples to the user, each shaped by a different filter $h_i$, and query the user for their satisfaction scores $f(h_i)$. A family of ``surrogate" functions is then designed to fit these scores and the optimization method gradually refines these functions to arrive at the filter $\hat{h}^*$ that maximizes satisfaction. In certain applications, we observe that a second type of querying is possible where users can tell us the individual elements $h^*[j]$ of the optimal filter $h^*$. Consider an analogy from cooking where the goal is to cook a recipe that maximizes user satisfaction. A user can be asked to score various cooked recipes (e.g., tofu fried rice) or to score individual ingredients (say, salt, sugar, rice, chicken, etc.). Given a budget of $B$ queries, where a query can be of either type, our goal is to find the recipe that will maximize this user's satisfaction. Our proposal builds on Sparse Gaussian Process Regression (GPR) and shows how a hybrid approach can outperform any one type of querying. Our results are validated through simulations and real world experiments, where volunteers gave feedback on music/speech audio and were able to achieve high satisfaction levels. We believe this idea of hybrid querying opens new problems in black-box optimization and solutions can benefit other applications beyond audio personalization.


【13】Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
标题:合成视听伪造中真实音频恢复和篡改定位的跨模式水印
链接:https://arxiv.org/abs/2507.12723

作者:Kim, Sehwan Park, Sungmin Cha, Paul Hongsuck Seo
备注:5 pages, 2 figures, Interspeech 2025
摘要:语音克隆和唇同步模型的最新进展使合成视听语音(SAVF)成为可能,其中音频和视觉都被操纵以模仿目标说话者。这通过使虚假内容看起来真实来显著增加错误信息的风险。为了解决这个问题,现有方法检测或定位操纵,但不能恢复传达消息的语义内容的真实音频。这种限制降低了它们在打击视听误导方面的效力。在这项工作中,我们介绍了真实的音频恢复(AAR)和篡改定位音频(TLA)从SAVF的任务,并提出了一个跨模态水印框架嵌入真实的音频到视觉操作之前。这就实现了AAR、TLA和对错误信息的强大防御。大量实验证明,我们的方法在AAR和TLA中针对各种操作(包括语音克隆和嘴唇同步)具有强大的性能。
摘要:Recent advances in voice cloning and lip synchronization models have enabled Synthesized Audiovisual Forgeries (SAVFs), where both audio and visuals are manipulated to mimic a target speaker. This significantly increases the risk of misinformation by making fake content seem real. To address this issue, existing methods detect or localize manipulations but cannot recover the authentic audio that conveys the semantic content of the message. This limitation reduces their effectiveness in combating audiovisual misinformation. In this work, we introduce the task of Authentic Audio Recovery (AAR) and Tamper Localization in Audio (TLA) from SAVFs and propose a cross-modal watermarking framework to embed authentic audio into visuals before manipulation. This enables AAR, TLA, and a robust defense against misinformation. Extensive experiments demonstrate the strong performance of our method in AAR and TLA against various manipulations, including voice cloning and lip synchronization.


【14】AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
标题:AudioJudge:了解基于大型音频模型的语音评估中的有效方法
链接:https://arxiv.org/abs/2507.12705

作者:Manakul, Woody Haosheng Gan, Michael J. Ryan, Ali Sartaz Khan, Warit Sirichotedumrong, Kunat Pipatanakul, William Held, Diyi Yang
摘要:目前的语音评价受到两个关键的限制:需要和困难,设计专门的系统针对个人的音频特性,和自动评价方法和人类偏好之间的相关性差。这项工作提出了一个系统的研究大音频模型(LAM)作为一个法官,AudioJudge,调查它是否可以提供一个统一的评估框架,解决这两个挑战。我们系统地探索了AudioJudge在音频特征检测任务中的应用,包括发音、说话速率、说话人识别和语音质量,以及用于自动基准测试的系统级人类偏好模拟。我们研究了不同的提示工程策略,发现音频级联与上下文学习相结合,显着提高了音频特征检测和人类偏好模拟任务的性能。我们还引入了一个多方面的合奏AudioJudge,使通用的多方面的音频评估。该方法将语音评估分解为词汇内容,语音质量和语言特征的专业判断,在我们的系统排名基准上与人类偏好的相关性高达0.91斯皮尔曼。鲁棒性分析表明,虽然LAMs在声学噪声下保持强大的性能,但它们表现出显着的冗长性和位置偏差,需要仔细缓解。
摘要:Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor correlation between automatic evaluation methods and human preferences. This work presents a systematic study of Large Audio Model (LAM) as a Judge, AudioJudge, investigating whether it can provide a unified evaluation framework that addresses both challenges. We systematically explore AudioJudge across audio characteristic detection tasks, including pronunciation, speaking rate, speaker identification and speech quality, and system-level human preference simulation for automated benchmarking. We investigate different prompt engineering strategies, finding that audio concatenation combined with in-context learning significantly improves performance across both audio characteristic detection and human preference simulation tasks. We further introduce a multi-aspect ensemble AudioJudge to enable general-purpose multi-aspect audio evaluation. This method decomposes speech assessment into specialized judges for lexical content, speech quality, and paralinguistic features, achieving up to 0.91 Spearman correlation with human preferences on our system ranking benchmark. Robustness analysis reveals that while LAMs maintain strong performance under acoustic noise, they exhibit significant verbosity and positional biases that require careful mitigation.


【15】Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
标题:机器的特定任务音频编码:机器学习的潜在特征是该机器的代码
链接:https://arxiv.org/abs/2507.12701

作者: Kuznetsova, Inseon Jang, Wootaek Lim, Minje Kim
摘要:利用量化算法的神经音频编解码器已经显著影响了各种语音/音频任务。虽然高保真重建对于人类感知至关重要,但机器音频编码(ACoM)优先考虑有效的压缩和下游任务性能,而忽略感知的细微差别。这项工作介绍了一种有效的ACoM方法,可以压缩和压缩任何选定的中间特征表示已经训练的语音/音频下游模型。我们的方法采用任务特定的损失指导以及残余矢量量化(RVQ)损失,提供超低比特率(即,小于200 bps),下游模型性能的损失最小。由此产生的令牌化器可适应各种比特率和模型大小,以实现灵活的部署。在自动语音识别和音频分类上进行评估,我们的方法通过适当的正则化证明了其有效性和更广泛的任务和架构适用性的潜力。
摘要:Neural audio codecs, leveraging quantization algorithms, have significantly impacted various speech/audio tasks. While high-fidelity reconstruction is paramount for human perception, audio coding for machines (ACoM) prioritizes efficient compression and downstream task performance, disregarding perceptual nuances. This work introduces an efficient ACoM method that can compress and quantize any chosen intermediate feature representation of an already trained speech/audio downstream model. Our approach employs task-specific loss guidance alongside residual vector quantization (RVQ) losses, providing ultra-low bitrates (i.e., less than 200 bps) with a minimal loss of the downstream model performance. The resulting tokenizer is adaptable to various bitrates and model sizes for flexible deployment. Evaluated on automatic speech recognition and audio classification, our method demonstrates its efficacy and potential for broader task and architectural applicability through appropriate regularization.


【16】Keep the beat going: Automatic drum transcription with momentum
标题:保持节拍:有动力的自动鼓转录
链接:https://arxiv.org/abs/2507.12596

作者: Foster, Robert J. Webber
摘要:一个简单的,可解释的方法来执行自动鼓转录是通过使用部分固定的非负矩阵分解分解的幅度谱图的记录的音乐作品。有两种自然的方法来优化非负矩阵分解,包括乘法更新规则和动量投影梯度下降。这些方法在经验精度和理论收敛保证方面有所不同。本文总结了方法和它们的时间复杂性,并将方法应用于ENST鼓数据集和作者乐队的原始录音,评估了地面实况鼓注释的经验准确性。结果表明,带有动量的投影梯度下降算法在固定运行时间内具有更高的精度,并且满足更强的收敛保证。
摘要:A simple, interpretable way to perform automatic drum transcription is by factoring the magnitude spectrogram of a recorded musical piece using a partially fixed nonnegative matrix factorization. There are two natural ways to optimize the nonnegative matrix factorization, including a multiplicative update rule and projected gradient descent with momentum. The methods differ in their empirical accuracies and theoretical convergence guarantees. This paper summarizes the methods and their time complexities, and it applies the methods to the ENST-Drums data set and an original recording from the author's band, evaluating the empirical accuracy with respect to ground-truth drum annotations. The results indicate that projected gradient descent with momentum leads to higher accuracy for a fixed runtime, and it satisfies stronger convergence guarantees.


【17】Evaluation of Neural Surrogates for Physical Modelling Synthesis of Nonlinear Elastic Plates
标题:非线性弹性板物理模型综合的神经代理评价
链接:https://arxiv.org/abs/2507.12563

作者: La Vega Martin, Rodrigo Diaz Fernandez, Mark Sandler
摘要:物理建模合成旨在从振动结构的物理模拟生成音频。薄弹性板是鼓式膜的常见模型。传统的数值方法,如有限差分和有限元提供了高精度,但计算要求高,限制了它们在实时音频应用中的使用。本文提出了一种基于神经网络的方法来解决非线性弹性板的振动的比较分析。我们评估了几种最先进的模型,在短序列上训练,以自回归方式预测长序列。我们展示了这些模型的一些局限性,以及为什么不足以查看时域中的预测误差。我们讨论了实时音频合成的影响,并提出了未来的方向,以改善神经方法来模拟非线性振动。
摘要:Physical modelling synthesis aims to generate audio from physical simulations of vibrating structures. Thin elastic plates are a common model for drum membranes. Traditional numerical methods like finite differences and finite elements offer high accuracy but are computationally demanding, limiting their use in real-time audio applications. This paper presents a comparative analysis of neural network-based approaches for solving the vibration of nonlinear elastic plates. We evaluate several state-of-the-art models, trained on short sequences, for prediction of long sequences in an autoregressive fashion. We show some of the limitations of these models, and why is not enough to look at the prediction error in the time domain. We discuss the implications for real-time audio synthesis and propose future directions for improving neural approaches to model nonlinear vibration.


机器翻译由腾讯交互翻译提供,仅供参考