微信公众号:arXiv_Daily
cs.SD语音
标题: 纳特:实时交互场景的神经声学传输
链接:https://arxiv.org/abs/2506.06190
摘要:以前的声学传输方法依赖于大量的数据预计算和存储,以实现实时交互和听觉反馈。然而,这些方法难以处理复杂的场景,特别是当对象位置、材料和大小的动态变化显著改变声音效果时。这些连续的变化导致波动的声学传递分布,使得用基本数据结构表示和实时高效渲染变得具有挑战性。为了解决这一挑战,我们提出了神经声学传输,一种新的方法,利用隐式神经表示编码预先计算的声学传输及其变化,允许在不同条件下实时预测声场。为了有效地生成神经声场所需的训练数据,我们开发了一种快速的基于蒙特卡罗的边界元法(BEM)近似的一般情况下光滑的诺依曼条件。此外,我们还为需要更高精度的场景实现了标准BEM的GPU加速版本。这些方法提供了必要的训练数据,使我们的神经网络能够准确地建模声辐射空间。我们证明了我们的方法的数值精度和运行时的效率(在几毫秒内为30秒的音频),通过全面的验证和比较,在不同的声学传输方案。我们的方法允许在动态变化的环境中有效和准确地建模声音行为,这可以有利于广泛的交互式应用,如虚拟现实,增强现实和先进的音频制作。
摘要:Previous acoustic transfer methods rely on extensive precomputation and storage of data to enable real-time interaction and auditory feedback. However, these methods struggle with complex scenes, especially when dynamic changes in object position, material, and size significantly alter sound effects. These continuous variations lead to fluctuating acoustic transfer distributions, making it challenging to represent with basic data structures and render efficiently in real time. To address this challenge, we present Neural Acoustic Transfer, a novel approach that utilizes an implicit neural representation to encode precomputed acoustic transfer and its variations, allowing for real-time prediction of sound fields under varying conditions. To efficiently generate the training data required for the neural acoustic field, we developed a fast Monte-Carlo-based boundary element method (BEM) approximation for general scenarios with smooth Neumann conditions. Additionally, we implemented a GPU-accelerated version of standard BEM for scenarios requiring higher precision. These methods provide the necessary training data, enabling our neural network to accurately model the sound radiation space. We demonstrate our method's numerical accuracy and runtime efficiency (within several milliseconds for 30s audio) through comprehensive validation and comparisons in diverse acoustic transfer scenarios. Our approach allows for efficient and accurate modeling of sound behavior in dynamically changing environments, which can benefit a wide range of interactive applications such as virtual reality, augmented reality, and advanced audio production.
【2】 Label-Context-Dependent Internal Language Model Estimation for CTC
链接:https://arxiv.org/abs/2506.06096
备注:accepted to Interspeech 2025
摘要:虽然连接主义时态分类(CTC)具有标签上下文无关性假设,但由于现代强大的编码器,它仍然可以隐式地学习上下文相关的内部语言模型(ILM)。在这项工作中,我们研究了隐含的上下文依赖建模的ILM的CTC。为此,我们提出了新的上下文相关的ILM估计方法的CTC知识蒸馏(KD)的基础上与理论的理由。此外,我们还介绍了KD的两种正则化方法。我们分别在Librispeech和TED-LIUM Release 2数据集上进行了域内和跨域评估实验。实验结果表明,上下文相关的ILM优于上下文无关的先验跨域评估,表明CTC学习上下文相关的ILM。所提出的标签级KD与平滑方法优于其他ILM估计方法,与浅融合相比,字错误率相对改善超过13%。
摘要:Although connectionist temporal classification (CTC) has the label context independence assumption, it can still implicitly learn a context-dependent internal language model (ILM) due to modern powerful encoders. In this work, we investigate the implicit context dependency modeled in the ILM of CTC. To this end, we propose novel context-dependent ILM estimation methods for CTC based on knowledge distillation (KD) with theoretical justifications. Furthermore, we introduce two regularization methods for KD. We conduct experiments on Librispeech and TED-LIUM Release 2 datasets for in-domain and cross-domain evaluation, respectively. Experimental results show that context-dependent ILMs outperform the context-independent priors in cross-domain evaluation, indicating that CTC learns a context-dependent ILM. The proposed label-level KD with smoothing method surpasses other ILM estimation approaches, with more than 13% relative improvement in word error rate compared to shallow fusion.
【3】 WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
链接:https://arxiv.org/abs/2506.05899
备注:3 pages
摘要:文本到音乐系统的平均意见得分(MOS)预测需要评估整体音乐质量和文本提示对齐。本文介绍了WhisQ,一种多模式架构,通过序列级共同关注和最佳运输规则化来解决这一双重评估挑战。WhisQ采用Whisper Base预训练模型进行时间音频编码,Qwen 3是一种0.6B小语言模型(SLM),用于文本编码,两者都保持了细粒度跨模态建模的序列结构。该架构具有专门的预测路径:OMQ是从池化的音频嵌入中预测的,而TA利用音频和文本之间的双向序列共同关注。Sinkhorn最优传输损失进一步加强共享嵌入空间中的语义对齐。在MusicEval Track-1数据集上,WhisQ在基线上实现了实质性的改进:OMQ的斯皮尔曼相关性提高了7%,TA提高了14%。消融研究表明,最佳的传输正则化提供了最大的性能增益(10% SRCC改善),显示了显式的跨模态对齐文本到音乐的评价的重要性。
摘要:Mean Opinion Score (MOS) prediction for text to music systems requires evaluating both overall musical quality and text prompt alignment. This paper introduces WhisQ, a multimodal architecture that addresses this dual-assessment challenge through sequence level co-attention and optimal transport regularization. WhisQ employs the Whisper Base pretrained model for temporal audio encoding and Qwen 3, a 0.6B Small Language Model (SLM), for text encoding, with both maintaining sequence structure for fine grained cross-modal modeling. The architecture features specialized prediction pathways: OMQ is predicted from pooled audio embeddings, while TA leverages bidirectional sequence co-attention between audio and text. Sinkhorn optimal transport loss further enforce semantic alignment in the shared embedding space. On the MusicEval Track-1 dataset, WhisQ achieves substantial improvements over the baseline: 7% improvement in Spearman correlation for OMQ and 14% for TA. Ablation studies reveal that optimal transport regularization provides the largest performance gain (10% SRCC improvement), demonstrating the importance of explicit cross-modal alignment for text-to-music evaluation.
【4】 WAKE: Watermarking Audio with Key Enrichment
链接:https://arxiv.org/abs/2506.05891
备注:Accepted by InterSpeech2025
摘要:随着深度学习在音频生成方面的进步,音频安全和版权保护方面的挑战凸显了对鲁棒音频水印的需求。最近基于神经网络的方法取得了进展,但仍然面临三个主要问题:防止未经授权的访问,多次嵌入后解码初始水印,以及嵌入不同长度的水印。为了解决这些问题,我们提出了WAKE,第一个密钥可控的音频水印框架。WAKE使用特定密钥嵌入水印,并使用相应的密钥恢复水印,通过使不正确的密钥解码成为可能来增强安全性。它还通过允许在多次嵌入后进行水印解码来解决水印问题,并支持可变长度的水印插入。WAKE在水印音频质量和水印检测精度方面都优于现有模型。代码、更多结果和演示页面:https://thuhcsi.github.io/WAKE。
摘要:As deep learning advances in audio generation, challenges in audio security and copyright protection highlight the need for robust audio watermarking. Recent neural network-based methods have made progress but still face three main issues: preventing unauthorized access, decoding initial watermarks after multiple embeddings, and embedding varying lengths of watermarks. To address these issues, we propose WAKE, the first key-controllable audio watermark framework. WAKE embeds watermarks using specific keys and recovers them with corresponding keys, enhancing security by making incorrect key decoding impossible. It also resolves the overwriting issue by allowing watermark decoding after multiple embeddings and supports variable-length watermark insertion. WAKE outperforms existing models in both watermarked audio quality and watermark detection accuracy. Code, more results, and demo page: https://thuhcsi.github.io/WAKE.
【5】 DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
链接:https://arxiv.org/abs/2506.05851
摘要:生成式人工智能发展迅速,允许创建非常逼真的操纵视频和音频。这一进展带来了重大的安全和道德威胁,因为恶意用户可以利用DeepFake技术传播错误信息。最近的DeepFake检测方法探索了多模态(音频-视频)威胁场景。特别是,现有数据集缺乏可重复性和关键问题-例如最近在广泛使用的FakeAVCeleb数据集中发现的沉默捷径。考虑到这一主题的重要性,我们的目标是更深入地了解影响音视频DeepFake检测基准测试的关键问题。我们通过三个核心基准支柱的镜头来检查这些挑战:数据集,检测方法和评估协议。为了解决这些问题,我们重点关注了最近的DeepSpeak v1数据集,并率先提出了一个评估协议,并使用SOTA模型对其进行了基准测试。我们引入SIMPLE Multimodal BASeline(SIMBA),这是一种具有竞争力但极简主义的方法,可以探索各种设计选择。我们还加深了对音频快捷方式问题的见解,并提出了一个有前途的缓解策略。最后,我们在广泛使用的FakeAVCeleb数据集上分析和增强了评估方案。我们的发现为音频视频DeepFake检测的复杂领域提供了一条前进的道路。
摘要:Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread misinformation. Recent DeepFake detection approaches explore the multimodal (audio-video) threat scenario. In particular, there is a lack of reproducibility and critical issues with existing datasets - such as the recently uncovered silence shortcut in the widely used FakeAVCeleb dataset. Considering the importance of this topic, we aim to gain a deeper understanding of the key issues affecting benchmarking in audio-video DeepFake detection. We examine these challenges through the lens of the three core benchmarking pillars: datasets, detection methods, and evaluation protocols. To address these issues, we spotlight the recent DeepSpeak v1 dataset and are the first to propose an evaluation protocol and benchmark it using SOTA models. We introduce SImple Multimodal BAseline (SIMBA), a competitive yet minimalistic approach that enables the exploration of diverse design choices. We also deepen insights into the issue of audio shortcuts and present a promising mitigation strategy. Finally, we analyze and enhance the evaluation scheme on the widely used FakeAVCeleb dataset. Our findings offer a way forward in the complex area of audio-video DeepFake detection.
【6】 Voice Impression Control in Zero-Shot TTS
链接:https://arxiv.org/abs/2506.05688
备注:5 pages,5 figures, Accepted to INTERSPEECH 2025
摘要:言语中的副/非语言信息是塑造听者印象的关键。虽然zero-shot文本到语音(TTS)已经实现了高说话者保真度,但是调制微妙的准/非语言信息以控制感知的语音特性,即,印象,仍然具有挑战性。因此,我们开发了一种zero-shot TTS中的语音印象控制方法,该方法利用低维向量来表示各种语音印象对的强度(例如,暗-亮)。客观和主观评价的结果都证明了我们的方法在印象控制的有效性。此外,通过大型语言模型生成该向量使得能够从所需印象的自然语言描述生成目标印象,从而消除对手动优化的需要。
摘要:Para-/non-linguistic information in speech is pivotal in shaping the listeners' impression. Although zero-shot text-to-speech (TTS) has achieved high speaker fidelity, modulating subtle para-/non-linguistic information to control perceived voice characteristics, i.e., impressions, remains challenging. We have therefore developed a voice impression control method in zero-shot TTS that utilizes a low-dimensional vector to represent the intensities of various voice impression pairs (e.g., dark-bright). The results of both objective and subjective evaluations have demonstrated our method's effectiveness in impression control. Furthermore, generating this vector via a large language model enables target-impression generation from a natural language description of the desired impression, thus eliminating the need for manual optimization.
【7】 Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
链接:https://arxiv.org/abs/2506.05593
备注:ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 11911-11915
摘要:近年来,端到端方法在解决说话人日记化的挑战方面取得了显着进展,其中涉及分割和识别多说话人录音中的说话人。一种这样的方法,编码器-解码器吸引器(EDA),已经被提出来处理可变的说话者计数,以及在训练期间更好地指导网络。在这项研究中,我们扩展吸引子范式超越直接扬声器建模,而是专注于通过多阶段的过程中,中间表示表示更详细的“扬声器属性”。此外,我们通过用conformer(一种卷积增强的Transformer)替换Transformers来增强架构,以建模局部依赖关系。实验表明,在CALLHOME数据集上的日记化性能有所改善。
摘要:In recent years, end-to-end approaches have made notable progress in addressing the challenge of speaker diarization, which involves segmenting and identifying speakers in multi-talker recordings. One such approach, Encoder-Decoder Attractors (EDA), has been proposed to handle variable speaker counts as well as better guide the network during training. In this study, we extend the attractor paradigm by moving beyond direct speaker modeling and instead focus on representing more detailed `speaker attributes' through a multi-stage process of intermediate representations. Additionally, we enhance the architecture by replacing transformers with conformers, a convolution-augmented transformer, to model local dependencies. Experiments demonstrate improved diarization performance on the CALLHOME dataset.
【8】 SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
链接:https://arxiv.org/abs/2506.05414
备注:Project website with demo videos: this https URL
摘要:动态视听环境中的3D空间推理是人类认知的基石,但现有的视听大语言模型(AV-LLM)和基准在很大程度上尚未探索,主要集中在静态或2D场景。我们介绍SAVVY-Bench,这是第一个在动态场景中使用同步空间音频进行3D空间推理的基准。SAVVY-Bench由数千个涉及静态和移动对象的关系组成,需要细粒度的时间基础,一致的3D定位和多模式注释。为了应对这一挑战,我们提出了SAVVY,这是一种新型的无训练推理管道,包括两个阶段:(i)自我中心空间轨迹估计,其利用AV-LLM以及其他视听方法来使用视觉和空间音频线索两者跟踪与查询相关的关键对象的轨迹,以及(ii)动态全局地图构建,其聚集多模态查询对象轨迹并将它们转换成统一的全局动态地图。使用构造的地图,通过将全局地图与查询的视点对齐的坐标变换来获得最终的QA答案。经验评估表明,SAVVY大大提高了最先进的AV-LLM的性能,为AV-LLM中的动态3D空间推理设定了新的标准和阶段。
摘要:3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3D spatial reasoning in dynamic scenes with synchronized spatial audio. SAVVY-Bench is comprised of thousands of relationships involving static and moving objects, and requires fine-grained temporal grounding, consistent 3D localization, and multi-modal annotation. To tackle this challenge, we propose SAVVY, a novel training-free reasoning pipeline that consists of two stages: (i) Egocentric Spatial Tracks Estimation, which leverages AV-LLMs as well as other audio-visual methods to track the trajectories of key objects related to the query using both visual and spatial audio cues, and (ii) Dynamic Global Map Construction, which aggregates multi-modal queried object trajectories and converts them into a unified global dynamic map. Using the constructed map, a final QA answer is obtained through a coordinate transformation that aligns the global map with the queried viewpoint. Empirical evaluation demonstrates that SAVVY substantially enhances performance of state-of-the-art AV-LLMs, setting a new standard and stage for approaching dynamic 3D spatial reasoning in AV-LLMs.
【1】 Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
链接:https://arxiv.org/abs/2506.06252
摘要:端到端自动语音识别(ASR)已经取得了显着的进步,但仍然与罕见和特定领域的实体斗争。本文介绍了一种简单而有效的基于上下文的ASR偏置技术,通过利用统一的多任务学习框架来提高识别精度。该方法包括两个关键组成部分:一个提示偏置模型,该模型被训练以确定何时关注提示中的实体,以及一个实体过滤机制,该机制有效地过滤掉不相关的实体。我们的方法显着提高了实体的ASR准确性,与内部域数据集上具有小实体列表和大实体列表的浅融合基线模型相比,实体字错误率分别降低了30.7%和18.0%。这种方法的主要优点在于其效率和简单性,而无需任何结构变化,使其重量轻,效率高。
摘要:End-to-End Automatic Speech Recognition (ASR) has advanced significantly yet still struggles with rare and domain-specific entities. This paper introduces a simple yet efficient prompt-based biasing technique for contextualized ASR, enhancing recognition accuracy by leverage a unified multitask learning framework. The approach comprises two key components: a prompt biasing model which is trained to determine when to focus on entities in prompt, and a entity filtering mechanism which efficiently filters out irrelevant entities. Our method significantly enhances ASR accuracy on entities, achieving a relative 30.7% and 18.0% reduction in Entity Word Error Rate compared to the baseline model with shallow fusion on in-house domain dataset with small and large entity lists, respectively. The primary advantage of this method lies in its efficiency and simplicity without any structure change, making it lightweight and highly efficient.
【2】 CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
链接:https://arxiv.org/abs/2506.06071
备注:8 pages
摘要:语音情感识别(SER)系统中的偏差通常源于说话者特征和情感标签之间的虚假相关性,导致对人口统计群体的不公平预测。许多现有的去偏方法需要模型特定的变化或人口统计注释,限制了它们的实际应用。我们提出了CO-VADA,一种以信心为导向的语音增强去偏方法,可以在不修改模型架构或依赖人口统计信息的情况下减轻偏差。CO-VADA识别反映训练数据中存在的偏差模式的训练样本,然后应用语音转换来改变不相关的属性并生成样本。这些增强的样本引入了与数据中的主导模式不同的说话者变化,引导模型更多地关注与情感相关的特征。我们的框架与各种SER模型和语音转换工具兼容,使其成为提高SER系统公平性的可扩展和实用的解决方案。
摘要:Bias in speech emotion recognition (SER) systems often stems from spurious correlations between speaker characteristics and emotional labels, leading to unfair predictions across demographic groups. Many existing debiasing methods require model-specific changes or demographic annotations, limiting their practical use. We present CO-VADA, a Confidence-Oriented Voice Augmentation Debiasing Approach that mitigates bias without modifying model architecture or relying on demographic information. CO-VADA identifies training samples that reflect bias patterns present in the training data and then applies voice conversion to alter irrelevant attributes and generate samples. These augmented samples introduce speaker variations that differ from dominant patterns in the data, guiding the model to focus more on emotion-relevant features. Our framework is compatible with various SER models and voice conversion tools, making it a scalable and practical solution for improving fairness in SER systems.
【3】 Audio-Aware Large Language Models as Judges for Speaking Styles
链接:https://arxiv.org/abs/2506.05984
摘要:音频感知大语言模型(ALLM)可以理解音频输入中的文本和非文本信息。在本文中,我们探索使用ALLM作为自动判断,以评估讲话风格的演讲。我们使用ALLM法官来评估两个任务:语音风格的指令以下和角色扮演的SLMs所产生的演讲。我们考虑的说话风格包括情感,音量,说话速度,单词强调,音高控制和非语言元素。我们使用四个口语模型(SLMs)来完成这两个任务,并使用人类和ALLM来判断SLMs的反应。我们比较了两个ALLM法官,GPT-4 O音频和双子座-2.5-亲,与人类的评价结果,并表明双子座和人类法官之间的协议是可比的人类评估员之间的协议。这些有希望的结果表明,ALLM可以作为一个法官来评价SLM。我们的研究结果还表明,目前的SLM,甚至GPT-40音频,仍然有改进的空间,在控制说话风格和产生自然的对话。
摘要:Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. We use ALLM judges to evaluate the speeches generated by SLMs on two tasks: voice style instruction following and role-playing. The speaking style we consider includes emotion, volume, speaking pace, word emphasis, pitch control, and non-verbal elements. We use four spoken language models (SLMs) to complete the two tasks and use humans and ALLMs to judge the SLMs' responses. We compare two ALLM judges, GPT-4o-audio and Gemini-2.5-pro, with human evaluation results and show that the agreement between Gemini and human judges is comparable to the agreement between human evaluators. These promising results show that ALLMs can be used as a judge to evaluate SLMs. Our results also reveal that current SLMs, even GPT-4o-audio, still have room for improvement in controlling the speaking style and generating natural dialogues.
【4】 TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
链接:https://arxiv.org/abs/2506.05802
备注:Accepted at Interspeech 2025
摘要:Deepfake检测在音频、文本和图像模式中获得了极大的关注,在区分真假方面具有很高的准确性。然而,识别确切的来源-例如Deepfake背后的系统或模型-仍然是一个研究较少的问题。在本文中,我们通过提出一种完全基于k最近邻(kNN)的免训练、绿色人工智能方法,在音频深度伪造模型归因或来源追踪方面迈出了重要一步。利用预训练的自监督学习(SSL)模型,我们证明了对来自同一生成器的样本进行分组是简单的-我们在五个deepfake数据集上获得了0.93的F1分数。该方法还展示了强大的域外(OOD)检测,有效地识别来自未知模型的样本,F1得分为0.84。 我们进一步分析这些结果在一个多维度的方法,并提供额外的见解。本工作中使用的所有代码和数据协议都可以在我们的开放存储库中找到:https://github.com/adrianastan/tada/。
摘要:Deepfake detection has gained significant attention across audio, text, and image modalities, with high accuracy in distinguishing real from fake. However, identifying the exact source--such as the system or model behind a deepfake--remains a less studied problem. In this paper, we take a significant step forward in audio deepfake model attribution or source tracing by proposing a training-free, green AI approach based entirely on k-Nearest Neighbors (kNN). Leveraging a pre-trained self-supervised learning (SSL) model, we show that grouping samples from the same generator is straightforward--we obtain an 0.93 F1-score across five deepfake datasets. The method also demonstrates strong out-of-domain (OOD) detection, effectively identifying samples from unseen models at an F1-score of 0.84. We further analyse these results in a multi-dimensional approach and provide additional insights. All code and data protocols used in this work are available in our open repository: https://github.com/adrianastan/tada/.
【5】 Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
链接:https://arxiv.org/abs/2506.05796
备注:Submitted to ASRU2025
摘要:多说话人自动语音识别(MS-ASR)在转录重叠语音方面面临着重大挑战,这是会议转录和会话分析等应用的关键任务。虽然序列化输出训练(SOT)风格的方法是常见的解决方案,但它们通常会丢弃绝对的时间信息,从而限制了它们在时间敏感场景中的实用性。利用大语言模型(LLM)的会话音频处理的最新进展,我们提出了一种新的diarization-aware多扬声器ASR系统,集成扬声器diarization与基于LLM的转录。我们的框架处理结构化的日记输入,以及框架级的扬声器和语义嵌入,使LLM生成段级transmittance。实验表明,该系统在多语言二元对话中具有鲁棒性,在复杂、高重叠的多人会议场景中表现出色。这项工作突出了LLM作为联合说话人感知分割和转录的统一后端的潜力。
摘要:Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training (SOT)-style methods serve as common solutions, they often discard absolute timing information, limiting their utility in time-sensitive scenarios. Leveraging recent advances in large language models (LLMs) for conversational audio processing, we propose a novel diarization-aware multi-speaker ASR system that integrates speaker diarization with LLM-based transcription. Our framework processes structured diarization inputs alongside frame-level speaker and semantic embeddings, enabling the LLM to generate segment-level transcriptions. Experiments demonstrate that the system achieves robust performance in multilingual dyadic conversations and excels in complex, high-overlap multi-speaker meeting scenarios. This work highlights the potential of LLMs as unified back-ends for joint speaker-aware segmentation and transcription.
【6】 Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
链接:https://arxiv.org/abs/2506.05706
摘要:将语音输入与大型语言模型(LLM)集成的一个挑战源于音频数据的连续性质与LLM的基于离散令牌的范例之间的差异。为了缓解这一差距,我们提出了一种方法,将矢量量化(VQ)集成到基于LLM的自动语音识别(ASR)。使用LLM嵌入表作为VQ码本,VQ模块将来自音频编码器的连续表示与离散LLM输入对齐,使得LLM能够对更好地反映语言结构的离散化音频表示进行操作。我们还通过更新码本并对码本嵌入执行加权求和来创建音频表示的软“离散化”。实证结果表明,我们提出的方法显着改善基于LLM的ASR基线,特别是在域外条件。这项工作突出了软离散化作为基于LLM的ASR中的模态桥梁的潜力。
摘要:One challenge of integrating speech input with large language models (LLMs) stems from the discrepancy between the continuous nature of audio data and the discrete token-based paradigm of LLMs. To mitigate this gap, we propose a method for integrating vector quantization (VQ) into LLM-based automatic speech recognition (ASR). Using the LLM embedding table as the VQ codebook, the VQ module aligns the continuous representations from the audio encoder with the discrete LLM inputs, enabling the LLM to operate on a discretized audio representation that better reflects the linguistic structure. We further create a soft "discretization" of the audio representation by updating the codebook and performing a weighted sum over the codebook embeddings. Empirical results demonstrate that our proposed method significantly improves upon the LLM-based ASR baseline, particularly in out-of-domain conditions. This work highlights the potential of soft discretization as a modality bridge in LLM-based ASR.
【7】 Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
链接:https://arxiv.org/abs/2506.05671
摘要:自动语音识别(ASR)的最新进展已经通过投影将语音编码器与大型语言模型(LLM)相结合,形成具有强大性能的语音LLM。然而,使它们适应新的领域仍然具有挑战性,特别是在配对的语音文本数据稀缺的低资源环境中。我们提出了一个纯文本微调策略的语音LLM使用不成对的目标域文本,而不需要额外的音频。为了保持语音文本对齐,我们在微调过程中引入了实时评估机制。这样可以在保持源域性能的同时实现有效的域自适应。LibriSpeech,SlideSpeech和Medical数据集上的实验表明,我们的方法具有竞争力的识别性能,与完整的音频文本微调相比,性能下降最小。它还提高了对新领域的泛化能力,而不会发生灾难性遗忘,突出了ASR低资源领域适应的纯文本微调潜力。
摘要:Recent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performance. However, adapting them to new domains remains challenging, especially in low-resource settings where paired speech-text data is scarce. We propose a text-only fine-tuning strategy for Speech LLMs using unpaired target-domain text without requiring additional audio. To preserve speech-text alignment, we introduce a real-time evaluation mechanism during fine-tuning. This enables effective domain adaptation while maintaining source-domain performance. Experiments on LibriSpeech, SlideSpeech, and Medical datasets show that our method achieves competitive recognition performance, with minimal degradation compared to full audio-text fine-tuning. It also improves generalization to new domains without catastrophic forgetting, highlighting the potential of text-only fine-tuning for low-resource domain adaptation of ASR.
【8】 NAT: Neural Acoustic Transfer for Interactive Scenes in Real Time
链接:https://arxiv.org/abs/2506.06190
摘要:以前的声学传输方法依赖于大量的数据预计算和存储,以实现实时交互和听觉反馈。然而,这些方法难以处理复杂的场景,特别是当对象位置、材料和大小的动态变化显著改变声音效果时。这些连续的变化导致波动的声学传递分布,使得用基本数据结构表示和实时高效渲染变得具有挑战性。为了解决这一挑战,我们提出了神经声学传输,一种新的方法,利用隐式神经表示编码预先计算的声学传输及其变化,允许在不同条件下实时预测声场。为了有效地生成神经声场所需的训练数据,我们开发了一种快速的基于蒙特卡罗的边界元法(BEM)近似的一般情况下光滑的诺依曼条件。此外,我们还为需要更高精度的场景实现了标准BEM的GPU加速版本。这些方法提供了必要的训练数据,使我们的神经网络能够准确地对声辐射空间进行建模。我们证明了我们的方法的数值精度和运行时的效率(在几毫秒内为30秒的音频),通过全面的验证和比较,在不同的声学传输方案。我们的方法允许在动态变化的环境中有效和准确地建模声音行为,这可以有利于广泛的交互式应用,如虚拟现实,增强现实和先进的音频制作。
摘要:Previous acoustic transfer methods rely on extensive precomputation and storage of data to enable real-time interaction and auditory feedback. However, these methods struggle with complex scenes, especially when dynamic changes in object position, material, and size significantly alter sound effects. These continuous variations lead to fluctuating acoustic transfer distributions, making it challenging to represent with basic data structures and render efficiently in real time. To address this challenge, we present Neural Acoustic Transfer, a novel approach that utilizes an implicit neural representation to encode precomputed acoustic transfer and its variations, allowing for real-time prediction of sound fields under varying conditions. To efficiently generate the training data required for the neural acoustic field, we developed a fast Monte-Carlo-based boundary element method (BEM) approximation for general scenarios with smooth Neumann conditions. Additionally, we implemented a GPU-accelerated version of standard BEM for scenarios requiring higher precision. These methods provide the necessary training data, enabling our neural network to accurately model the sound radiation space. We demonstrate our method's numerical accuracy and runtime efficiency (within several milliseconds for 30s audio) through comprehensive validation and comparisons in diverse acoustic transfer scenarios. Our approach allows for efficient and accurate modeling of sound behavior in dynamically changing environments, which can benefit a wide range of interactive applications such as virtual reality, augmented reality, and advanced audio production.
【9】 Label-Context-Dependent Internal Language Model Estimation for CTC
链接:https://arxiv.org/abs/2506.06096
备注:accepted to Interspeech 2025
摘要:虽然连接主义时态分类(CTC)具有标签上下文无关性假设,但由于现代强大的编码器,它仍然可以隐式地学习上下文相关的内部语言模型(ILM)。在这项工作中,我们研究了隐含的上下文依赖建模的ILM的CTC。为此,我们提出了新的上下文相关的ILM估计方法的CTC知识蒸馏(KD)的基础上与理论的理由。此外,我们还介绍了KD的两种正则化方法。我们分别在Librispeech和TED-LIUM Release 2数据集上进行了域内和跨域评估实验。实验结果表明,上下文相关的ILM优于上下文无关的先验跨域评估,表明CTC学习上下文相关的ILM。所提出的标签级KD与平滑方法优于其他ILM估计方法,与浅融合相比,字错误率相对改善超过13%。
摘要:Although connectionist temporal classification (CTC) has the label context independence assumption, it can still implicitly learn a context-dependent internal language model (ILM) due to modern powerful encoders. In this work, we investigate the implicit context dependency modeled in the ILM of CTC. To this end, we propose novel context-dependent ILM estimation methods for CTC based on knowledge distillation (KD) with theoretical justifications. Furthermore, we introduce two regularization methods for KD. We conduct experiments on Librispeech and TED-LIUM Release 2 datasets for in-domain and cross-domain evaluation, respectively. Experimental results show that context-dependent ILMs outperform the context-independent priors in cross-domain evaluation, indicating that CTC learns a context-dependent ILM. The proposed label-level KD with smoothing method surpasses other ILM estimation approaches, with more than 13% relative improvement in word error rate compared to shallow fusion.
【10】 WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
链接:https://arxiv.org/abs/2506.05899
备注:3 pages
摘要:文本到音乐系统的平均意见得分(MOS)预测需要评估整体音乐质量和文本提示对齐。本文介绍了WhisQ,一种多模式架构,通过序列级共同关注和最佳运输规则化来解决这一双重评估挑战。WhisQ采用Whisper Base预训练模型进行时间音频编码,Qwen 3是一种0.6B小语言模型(SLM),用于文本编码,两者都保持了细粒度跨模态建模的序列结构。该架构具有专门的预测路径:OMQ是从池化的音频嵌入中预测的,而TA利用音频和文本之间的双向序列共同关注。Sinkhorn最优传输损失进一步加强共享嵌入空间中的语义对齐。在MusicEval Track-1数据集上,WhisQ在基线上实现了实质性的改进:OMQ的斯皮尔曼相关性提高了7%,TA提高了14%。消融研究表明,最佳的传输正则化提供了最大的性能增益(10% SRCC改善),显示了显式的跨模态对齐文本到音乐的评价的重要性。
摘要:Mean Opinion Score (MOS) prediction for text to music systems requires evaluating both overall musical quality and text prompt alignment. This paper introduces WhisQ, a multimodal architecture that addresses this dual-assessment challenge through sequence level co-attention and optimal transport regularization. WhisQ employs the Whisper Base pretrained model for temporal audio encoding and Qwen 3, a 0.6B Small Language Model (SLM), for text encoding, with both maintaining sequence structure for fine grained cross-modal modeling. The architecture features specialized prediction pathways: OMQ is predicted from pooled audio embeddings, while TA leverages bidirectional sequence co-attention between audio and text. Sinkhorn optimal transport loss further enforce semantic alignment in the shared embedding space. On the MusicEval Track-1 dataset, WhisQ achieves substantial improvements over the baseline: 7% improvement in Spearman correlation for OMQ and 14% for TA. Ablation studies reveal that optimal transport regularization provides the largest performance gain (10% SRCC improvement), demonstrating the importance of explicit cross-modal alignment for text-to-music evaluation.
【11】 WAKE: Watermarking Audio with Key Enrichment
链接:https://arxiv.org/abs/2506.05891
备注:Accepted by InterSpeech2025
摘要:随着深度学习在音频生成方面的进步,音频安全和版权保护方面的挑战凸显了对鲁棒音频水印的需求。最近基于神经网络的方法取得了进展,但仍然面临三个主要问题:防止未经授权的访问,多次嵌入后解码初始水印,以及嵌入不同长度的水印。为了解决这些问题,我们提出了WAKE,第一个密钥可控的音频水印框架。WAKE使用特定密钥嵌入水印,并使用相应的密钥恢复水印,通过使不正确的密钥解码成为可能来增强安全性。它还通过允许在多次嵌入后进行水印解码来解决水印问题,并支持可变长度的水印插入。WAKE在水印音频质量和水印检测精度方面都优于现有模型。代码、更多结果和演示页面:https://thuhcsi.github.io/WAKE。
摘要:As deep learning advances in audio generation, challenges in audio security and copyright protection highlight the need for robust audio watermarking. Recent neural network-based methods have made progress but still face three main issues: preventing unauthorized access, decoding initial watermarks after multiple embeddings, and embedding varying lengths of watermarks. To address these issues, we propose WAKE, the first key-controllable audio watermark framework. WAKE embeds watermarks using specific keys and recovers them with corresponding keys, enhancing security by making incorrect key decoding impossible. It also resolves the overwriting issue by allowing watermark decoding after multiple embeddings and supports variable-length watermark insertion. WAKE outperforms existing models in both watermarked audio quality and watermark detection accuracy. Code, more results, and demo page: https://thuhcsi.github.io/WAKE.
【12】 DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
链接:https://arxiv.org/abs/2506.05851
摘要:生成式人工智能发展迅速,允许创建非常逼真的操纵视频和音频。这一进展带来了重大的安全和道德威胁,因为恶意用户可以利用DeepFake技术传播错误信息。最近的DeepFake检测方法探索了多模态(音频-视频)威胁场景。特别是,现有数据集缺乏可重复性和关键问题-例如最近在广泛使用的FakeAVCeleb数据集中发现的沉默捷径。考虑到这一主题的重要性,我们的目标是更深入地了解影响音视频DeepFake检测基准测试的关键问题。我们通过三个核心基准支柱的镜头来检查这些挑战:数据集,检测方法和评估协议。为了解决这些问题,我们重点关注了最近的DeepSpeak v1数据集,并率先提出了一个评估协议,并使用SOTA模型对其进行了基准测试。我们引入SIMPLE Multimodal BASeline(SIMBA),这是一种具有竞争力但极简主义的方法,可以探索各种设计选择。我们还加深了对音频快捷方式问题的见解,并提出了一个有前途的缓解策略。最后,我们在广泛使用的FakeAVCeleb数据集上分析和增强了评估方案。我们的研究结果为音视频DeepFake检测的复杂领域提供了一条前进的道路。
摘要:Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread misinformation. Recent DeepFake detection approaches explore the multimodal (audio-video) threat scenario. In particular, there is a lack of reproducibility and critical issues with existing datasets - such as the recently uncovered silence shortcut in the widely used FakeAVCeleb dataset. Considering the importance of this topic, we aim to gain a deeper understanding of the key issues affecting benchmarking in audio-video DeepFake detection. We examine these challenges through the lens of the three core benchmarking pillars: datasets, detection methods, and evaluation protocols. To address these issues, we spotlight the recent DeepSpeak v1 dataset and are the first to propose an evaluation protocol and benchmark it using SOTA models. We introduce SImple Multimodal BAseline (SIMBA), a competitive yet minimalistic approach that enables the exploration of diverse design choices. We also deepen insights into the issue of audio shortcuts and present a promising mitigation strategy. Finally, we analyze and enhance the evaluation scheme on the widely used FakeAVCeleb dataset. Our findings offer a way forward in the complex area of audio-video DeepFake detection.
【13】 Voice Impression Control in Zero-Shot TTS
链接:https://arxiv.org/abs/2506.05688
备注:5 pages,5 figures, Accepted to INTERSPEECH 2025
摘要:言语中的副/非语言信息是塑造听者印象的关键。虽然zero-shot文本到语音(TTS)已经实现了高说话者保真度,但是调制微妙的准/非语言信息以控制感知的语音特性,即,印象,仍然具有挑战性。因此,我们开发了一种zero-shot TTS中的语音印象控制方法,该方法利用低维向量来表示各种语音印象对的强度(例如,暗-亮)。客观和主观评价的结果都证明了我们的方法在印象控制的有效性。此外,通过大型语言模型生成该向量使得能够从所需印象的自然语言描述生成目标印象,从而消除对手动优化的需要。
摘要:Para-/non-linguistic information in speech is pivotal in shaping the listeners' impression. Although zero-shot text-to-speech (TTS) has achieved high speaker fidelity, modulating subtle para-/non-linguistic information to control perceived voice characteristics, i.e., impressions, remains challenging. We have therefore developed a voice impression control method in zero-shot TTS that utilizes a low-dimensional vector to represent the intensities of various voice impression pairs (e.g., dark-bright). The results of both objective and subjective evaluations have demonstrated our method's effectiveness in impression control. Furthermore, generating this vector via a large language model enables target-impression generation from a natural language description of the desired impression, thus eliminating the need for manual optimization.
【14】 Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
链接:https://arxiv.org/abs/2506.05593
备注:ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 11911-11915
摘要:近年来,端到端方法在解决说话人日记化的挑战方面取得了显着进展,其中涉及分割和识别多说话人录音中的说话人。一种这样的方法,编码器-解码器吸引器(EDA),已经被提出来处理可变的说话者计数,以及在训练期间更好地指导网络。在这项研究中,我们扩展吸引子范式超越直接扬声器建模,而是专注于通过多阶段的过程中,中间表示表示更详细的“扬声器属性”。此外,我们还通过用Conformer(一种卷积增强的Transformer)替换Transformers来增强架构,以对本地依赖关系进行建模。实验表明,在CALLHOME数据集上的日记化性能有所改善。
摘要:In recent years, end-to-end approaches have made notable progress in addressing the challenge of speaker diarization, which involves segmenting and identifying speakers in multi-talker recordings. One such approach, Encoder-Decoder Attractors (EDA), has been proposed to handle variable speaker counts as well as better guide the network during training. In this study, we extend the attractor paradigm by moving beyond direct speaker modeling and instead focus on representing more detailed `speaker attributes' through a multi-stage process of intermediate representations. Additionally, we enhance the architecture by replacing transformers with conformers, a convolution-augmented transformer, to model local dependencies. Experiments demonstrate improved diarization performance on the CALLHOME dataset.
【15】 SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
链接:https://arxiv.org/abs/2506.05414
备注:Project website with demo videos: this https URL
摘要:动态视听环境中的3D空间推理是人类认知的基石,但现有的视听大语言模型(AV-LLM)和基准在很大程度上尚未探索,主要集中在静态或2D场景。我们介绍SAVVY-Bench,这是第一个在动态场景中使用同步空间音频进行3D空间推理的基准。SAVVY-Bench由数千个涉及静态和移动对象的关系组成,需要细粒度的时间基础,一致的3D定位和多模式注释。为了应对这一挑战,我们提出了SAVVY,这是一种新型的无训练推理管道,包括两个阶段:(i)自我中心空间轨迹估计,其利用AV-LLM以及其他视听方法来使用视觉和空间音频线索两者跟踪与查询相关的关键对象的轨迹,以及(ii)动态全局地图构建,它聚合多模式查询对象轨迹并将其转换为统一的全局动态地图。使用构造的地图,通过将全局地图与查询的视点对齐的坐标变换来获得最终的QA答案。经验评估表明,SAVVY大大提高了最先进的AV-LLM的性能,为AV-LLM中的动态3D空间推理设定了新的标准和阶段。
摘要:3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3D spatial reasoning in dynamic scenes with synchronized spatial audio. SAVVY-Bench is comprised of thousands of relationships involving static and moving objects, and requires fine-grained temporal grounding, consistent 3D localization, and multi-modal annotation. To tackle this challenge, we propose SAVVY, a novel training-free reasoning pipeline that consists of two stages: (i) Egocentric Spatial Tracks Estimation, which leverages AV-LLMs as well as other audio-visual methods to track the trajectories of key objects related to the query using both visual and spatial audio cues, and (ii) Dynamic Global Map Construction, which aggregates multi-modal queried object trajectories and converts them into a unified global dynamic map. Using the constructed map, a final QA answer is obtained through a coordinate transformation that aligns the global map with the queried viewpoint. Empirical evaluation demonstrates that SAVVY substantially enhances performance of state-of-the-art AV-LLMs, setting a new standard and stage for approaching dynamic 3D spatial reasoning in AV-LLMs.
机器翻译由腾讯交互翻译提供,仅供参考
