微信公众号:arXiv_Daily
cs.SD语音
标题: DualTalk:用于3D说话头部对话的双扬声器交互
链接:https://arxiv.org/abs/2505.18096
备注:Accepted by CVPR 2025
摘要:在面对面的交谈中,人们需要在说话和倾听角色之间无缝切换。现有的3D说话头部生成模型只关注说话或倾听,忽略了交互式对话的自然动态,这导致不自然的交互和尴尬的过渡。为了解决这个问题,我们提出了一个新的任务-多轮双扬声器交互的3D说话头生成-这需要模型来处理和生成说话和倾听行为在连续的对话。为了解决这个问题,我们引入了DualTalk,一个新颖的统一框架,它集成了说话者和听众的动态行为,以模拟逼真和连贯的对话交互。该框架不仅在说话时合成逼真的说话头,而且在倾听时产生连续和生动的非语言反馈,有效地捕捉角色之间的相互作用。我们还创建了一个新的数据集,其中包含50小时的多轮对话,超过1,000个字符,参与者在说话和倾听角色之间不断切换。大量的实验表明,我们的方法显着提高的自然性和表现力的3D说话的头在双扬声器对话。我们建议观看补充视频:https://ziqiaopeng.github.io/dualtalk。
摘要:In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task -- multi-round dual-speaker interaction for 3D talking head generation -- which requires models to handle and generate both speaking and listening behaviors in continuous conversation. To solve this task, we introduce DualTalk, a novel unified framework that integrates the dynamic behaviors of speakers and listeners to simulate realistic and coherent dialogue interactions. This framework not only synthesizes lifelike talking heads when speaking but also generates continuous and vivid non-verbal feedback when listening, effectively capturing the interplay between the roles. We also create a new dataset featuring 50 hours of multi-round conversations with over 1,000 characters, where participants continuously switch between speaking and listening roles. Extensive experiments demonstrate that our method significantly enhances the naturalness and expressiveness of 3D talking heads in dual-speaker conversations. We recommend watching the supplementary video: https://ziqiaopeng.github.io/dualtalk.
【2】 Toward Optimal ANC: Establishing Mutual Information Lower Bound
标题: 迈向最佳非国大:建立互信息下限链接:https://arxiv.org/abs/2505.17877
摘要:主动噪声消除(ANC)算法旨在通过生成实时破坏性干扰原始噪声的抗噪声信号来抑制不必要的声学干扰。尽管最近基于深度学习的ANC算法已经设定了新的性能基准,但仍然缺乏严格评估其改进的理论限制。为了解决这个问题,我们推导出一个统一的下限消除性能由两个组件。第一个组成部分是信息理论:它将残余误差功率与抗噪声信号捕获的干扰熵的分数联系起来,从而量化信息处理能力所施加的限制。第二部分是基于支持的:它测量消除路径无法解决的频带中产生的不可约误差,反映了基本的物理约束。通过取这两项的最大值,我们的界限建立了任何ANC算法可达到的归一化均方误差(NMSE)的理论上限。我们在不同混响时间下的NOISEX数据集上经验性地验证了其紧密性,证明了在不同声学条件下的鲁棒性。
摘要:Active Noise Cancellation (ANC) algorithms aim to suppress unwanted acoustic disturbances by generating anti-noise signals that destructively interfere with the original noise in real time. Although recent deep learning-based ANC algorithms have set new performance benchmarks, there remains a shortage of theoretical limits to rigorously assess their improvements. To address this, we derive a unified lower bound on cancellation performance composed of two components. The first component is information-theoretic: it links residual error power to the fraction of disturbance entropy captured by the anti-noise signal, thereby quantifying limits imposed by information-processing capacity. The second component is support-based: it measures the irreducible error arising in frequency bands that the cancellation path cannot address, reflecting fundamental physical constraints. By taking the maximum of these two terms, our bound establishes a theoretical ceiling on the Normalized Mean Squared Error (NMSE) attainable by any ANC algorithm. We validate its tightness empirically on the NOISEX dataset under varying reverberation times, demonstrating robustness across diverse acoustic conditions.
【3】 CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
标题: CosyVoice 3:通过扩展和后训练实现野外语音生成链接:https://arxiv.org/abs/2505.17589
备注:Preprint, work in progress
摘要:在我们之前的工作中,我们介绍了一个可扩展的流语音合成模型,CosyVoice 2,它集成了一个大语言模型(LLM)和一个块感知流匹配(FM)模型,并实现了低延迟的双流语音合成和人类奇偶校验质量。尽管有这些进步,但CosyVoice 2在语言覆盖范围、领域多样性、数据量、文本格式和后期训练技术方面存在局限性。在本文中,我们提出了CosyVoice 3,一个改进的模型设计的zero-shot多语言语音合成在野外,超过其前身的内容一致性,说话人相似性,韵律自然。CosyVoice 3的主要功能包括:1)通过监督多任务训练开发的一种新型语音标记器,用于改善韵律自然度,包括自动语音识别,语音情感识别,语言识别,音频事件检测和说话人分析。2)提出了一种新的可微分后训练奖励模型,不仅适用于CosyVoice 3,也适用于其他基于LLM的语音合成模型。3)数据集大小缩放:训练数据从一万小时扩展到一百万小时,包括9种语言和18种汉语方言,涵盖各种领域和文本格式。4)模型尺寸缩放:模型参数从5亿增加到15亿,由于更大的模型容量,我们的多语言基准测试的性能得到了增强。这些进步对语音合成在野外的进展做出了重大贡献。我们鼓励读者在https://funaudiollm.github.io/cosyvoice3上收听演示。
摘要:In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.
【4】 JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
标题: JALMBench:音频语言模型中的越狱漏洞基准链接:https://arxiv.org/abs/2505.17568
摘要:音频语言模型(ALM)最近取得了重大进展。这些模型将音频模态直接集成到模型中,而不是将语音转换为文本并将文本输入到大型语言模型(LLM)中。虽然对LLM的越狱攻击已经得到了广泛的研究,但具有音频模式的ALM的安全性在很大程度上仍未得到探索。目前,缺乏对抗性音频数据集和专门用于评估和比较攻击和ALM的统一框架。在本文中,我们提出了JALMBench,\textit{第一}全面的基准来评估安全的ALM对越狱攻击。JALMBench包含一个数据集,包含2,200个文本样本和51,381个音频样本,超过268小时。支持12种主流ALM,4种文本传输和4种音频源攻击方式,5种防御方式。使用JALMBench,我们提供了攻击效率,主题敏感性,语音多样性和攻击表示的深入分析。此外,我们探讨缓解策略的攻击在提示级别和响应级别。
摘要:Audio Language Models (ALMs) have made significant progress recently. These models integrate the audio modality directly into the model, rather than converting speech into text and inputting text to Large Language Models (LLMs). While jailbreak attacks on LLMs have been extensively studied, the security of ALMs with audio modalities remains largely unexplored. Currently, there is a lack of an adversarial audio dataset and a unified framework specifically designed to evaluate and compare attacks and ALMs. In this paper, we present JALMBench, the \textit{first} comprehensive benchmark to assess the safety of ALMs against jailbreak attacks. JALMBench includes a dataset containing 2,200 text samples and 51,381 audio samples with over 268 hours. It supports 12 mainstream ALMs, 4 text-transferred and 4 audio-originated attack methods, and 5 defense methods. Using JALMBench, we provide an in-depth analysis of attack efficiency, topic sensitivity, voice diversity, and attack representations. Additionally, we explore mitigation strategies for the attacks at both the prompt level and the response level.
【5】 MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation
标题: MEGADance:适合具有流派意识的3D舞蹈一代的专家混合架构链接:https://arxiv.org/abs/2505.17543
备注:arXiv admin note: text overlap with arXiv:2505.14222
摘要:近年来,音乐驱动的3D舞蹈生成引起了越来越多的关注,在编舞,虚拟现实和创意内容创作中有着广阔的应用前景。先前的研究已经从音频信号中生成了有希望的逼真的舞蹈动作。然而,传统的方法没有充分利用体裁制约,往往把它作为辅助修饰语,而不是核心的语义驱动。这种疏忽损害了音乐动作的同步性,破坏了舞蹈类型的连续性,特别是在复杂的节奏过渡期间,从而导致视觉效果不令人满意。为了应对这一挑战,我们提出了MEGADance,一个新的架构,音乐驱动的3D舞蹈生成。通过将舞蹈的一致性解耦为舞蹈的一般性和体裁的特殊性,MEGADance展示了显著的舞蹈质量和强大的体裁可控性。它包括两个阶段:(1)高保真舞蹈量化阶段(HFDQ),其通过有限标量量化(FSQ)将舞蹈动作编码成潜在表示,并利用运动学-动态约束来重构它们,以及(2)体裁感知舞蹈生成阶段(GADG),其通过协同利用具有Mamba-Transformer混合骨干的专家混合(MoE)机制来将音乐映射成潜在表示。在FineDance和AIST++数据集上进行的大量实验证明了MEGADance在定性和定量方面的最新性能。代码将在接受后发布。
摘要:Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance movement from audio signals. However, traditional methods underutilize genre conditioning, often treating it as auxiliary modifiers rather than core semantic drivers. This oversight compromises music-motion synchronization and disrupts dance genre continuity, particularly during complex rhythmic transitions, thereby leading to visually unsatisfactory effects. To address the challenge, we propose MEGADance, a novel architecture for music-driven 3D dance generation. By decoupling choreographic consistency into dance generality and genre specificity, MEGADance demonstrates significant dance quality and strong genre controllability. It consists of two stages: (1) High-Fidelity Dance Quantization Stage (HFDQ), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) and reconstructs them with kinematic-dynamic constraints, and (2) Genre-Aware Dance Generation Stage (GADG), which maps music into the latent representation by synergistic utilization of Mixture-of-Experts (MoE) mechanism with Mamba-Transformer hybrid backbone. Extensive experiments on the FineDance and AIST++ dataset demonstrate the state-of-the-art performance of MEGADance both qualitatively and quantitatively. Code will be released upon acceptance.
【6】 Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
标题: 瑞典耳语者;利用大量语音数据库进行瑞典语音识别链接:https://arxiv.org/abs/2505.17538
备注:Submitted to Interspeech 2025
摘要:这项工作为瑞典语提供了一套微调的Whisper模型,该模型在这种中等资源语言的前所未有的大小和可变性的数据集上进行了训练。由于较小规模的语言在多语言训练数据集中往往代表性不足,因此可以通过微调现有的多语言模型来实现性能的大幅提高,如本文所示。这项工作报告了与OpenAI在瑞典评估的Whisper相比,模型大小的整体改进。最值得注意的是,在FLEURS、Common Voice和NST的评估中,我们报告称,与OpenAI的whisper-large-v3相比,我们的最佳性能模型的WER平均降低了47%。
摘要:This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual training datasets, substantial improvements in performance can be achieved by fine-tuning existing multilingual models, as shown in this work. This work reports an overall improvement across model sizes compared to OpenAI's Whisper evaluated on Swedish. Most notably, we report an average 47% reduction in WER comparing our best performing model to OpenAI's whisper-large-v3, in evaluations across FLEURS, Common Voice, and NST.
【7】 What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
标题: 读到的不是听到的:Deepfake语音检测中的语言敏感性链接:https://arxiv.org/abs/2505.17513
备注:15 pages, 2 fogures
摘要:文本转语音技术的最新进展使真实的语音生成成为可能,从而助长了基于音频的deepfake攻击,如欺诈和模仿。虽然音频反欺骗系统对于检测此类威胁至关重要,但先前的工作主要集中在声学水平的扰动上,而语言变化的影响在很大程度上未被探索。在本文中,我们通过引入转录水平的对抗性攻击来研究开源和商业反欺骗检测器的语言敏感性。我们的广泛评估表明,即使是轻微的语言干扰也会显着降低检测准确性:在几个开源检测器-语音对上,攻击成功率超过60%,特别是一个商业检测准确率从合成音频的100%下降到只有32%。通过全面的特征属性分析,我们发现语言复杂性和模型级音频嵌入相似性都对检测器脆弱性有很大贡献。我们通过复制Brad Pitt音频deepfake骗局的案例研究进一步展示了现实世界的风险,使用转录对抗攻击完全绕过商业检测器。这些结果强调了需要超越纯粹的声学防御,并考虑到语言的变化,在强大的反欺骗系统的设计。所有源代码都将公开。
摘要:Recent advances in text-to-speech technologies have enabled realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation. While audio anti-spoofing systems are critical for detecting such threats, prior work has predominantly focused on acoustic-level perturbations, leaving the impact of linguistic variation largely unexplored. In this paper, we investigate the linguistic sensitivity of both open-source and commercial anti-spoofing detectors by introducing transcript-level adversarial attacks. Our extensive evaluation reveals that even minor linguistic perturbations can significantly degrade detection accuracy: attack success rates surpass 60% on several open-source detector-voice pairs, and notably one commercial detection accuracy drops from 100% on synthetic audio to just 32%. Through a comprehensive feature attribution analysis, we identify that both linguistic complexity and model-level audio embedding similarity contribute strongly to detector vulnerability. We further demonstrate the real-world risk via a case study replicating the Brad Pitt audio deepfake scam, using transcript adversarial attacks to completely bypass commercial detectors. These results highlight the need to move beyond purely acoustic defenses and account for linguistic variation in the design of robust anti-spoofing systems. All source code will be publicly available.
【8】 Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
标题: 分析口语模型端到端训练中灾难性遗忘的缓解策略链接:https://arxiv.org/abs/2505.17496
备注:Accepted to Interspeech 2025
摘要:口语语言模型(SLM)的端到端训练通常涉及通过对ASR、TTS和口语问答(SQA)等各种任务进行多阶段训练,使预先训练的基于文本的大型语言模型(LLM)适应语音模态。尽管这种多阶段的持续学习使LLM具备了语音理解和生成能力,但各个阶段的任务和数据分布的巨大差异可能导致灾难性的遗忘,从而丢失先前获得的知识。本文研究了灾难性遗忘,并评估了三种缓解策略-模型合并,LoRA缩放因子折扣和经验重放,以平衡知识保留与新的学习。实验结果表明,经验回放法是最有效的,与其他方法相结合可以获得更大的收益。这些研究结果为开发更强大和有效的可持续土地管理培训管道提供了见解。
摘要:End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS and spoken question answering (SQA). Although this multi-stage continual learning equips LLMs with both speech understanding and generation capabilities, the substantial differences in task and data distributions across stages can lead to catastrophic forgetting, where previously acquired knowledge is lost. This paper investigates catastrophic forgetting and evaluates three mitigation strategies-model merging, discounting the LoRA scaling factor, and experience replay to balance knowledge retention with new learning. Results show that experience replay is the most effective, with further gains achieved by combining it with other methods. These findings provide insights for developing more robust and efficient SLM training pipelines.
【9】 Reverse-Speech-Finder: A Neural Network Backtracking Architecture for Generating Alzheimer's Disease Speech Samples and Improving Diagnosis Performance
标题: 反向语音识别:一种用于阿尔茨海默病语音样本生成和提高诊断性能的神经网络回溯结构链接:https://arxiv.org/abs/2505.17477
摘要:这项研究介绍了反向语音识别(RSF),这是一种开创性的神经网络回溯架构,旨在通过语音分析来增强阿尔茨海默病(AD)的诊断。RSF利用预先训练的大型语言模型的能力,识别并利用最可能的AD特定语音标记,解决了真实AD语音样本的稀缺性和现有模型中有限的可解释性的挑战。RSF的独特方法包括三个核心创新:首先,它利用了这样的观察,即最有可能预测AD的语音标记(定义为最可能的语音标记(MPM))必须具有激活那些具有最高预测AD概率的神经元(在神经网络中)的最高概率,定义为最可能的神经元(MPN)。其次,它在输入层利用语音令牌表示,允许从MPN回溯以识别AD的最可能的语音令牌(MPT)。最后,它开发了一种创新的回溯方法,从MPN向后跟踪到输入层,识别MPT和相应的MPM,并巧妙地发现新的语音标记AD检测。实验结果表明,RSF的优越性,比传统的方法,如SHAP和集成的修正,实现了3.5%的准确性和3.2%的F1分数提高。通过生成封装新标记的语音数据,RSF不仅减轻了真实数据稀缺的限制,而且还显著增强了AD诊断模型的鲁棒性和准确性。这些发现强调了RSF作为基于语音的AD检测的变革性工具的潜力,为AD相关的语言缺陷提供了新的见解,并为更有效的非侵入性早期干预策略铺平了道路。
摘要:This study introduces Reverse-Speech-Finder (RSF), a groundbreaking neural network backtracking architecture designed to enhance Alzheimer's Disease (AD) diagnosis through speech analysis. Leveraging the power of pre-trained large language models, RSF identifies and utilizes the most probable AD-specific speech markers, addressing both the scarcity of real AD speech samples and the challenge of limited interpretability in existing models. RSF's unique approach consists of three core innovations: Firstly, it exploits the observation that speech markers most probable of predicting AD, defined as the most probable speech-markers (MPMs), must have the highest probability of activating those neurons (in the neural network) with the highest probability of predicting AD, defined as the most probable neurons (MPNs). Secondly, it utilizes a speech token representation at the input layer, allowing backtracking from MPNs to identify the most probable speech-tokens (MPTs) of AD. Lastly, it develops an innovative backtracking method to track backwards from the MPNs to the input layer, identifying the MPTs and the corresponding MPMs, and ingeniously uncovering novel speech markers for AD detection. Experimental results demonstrate RSF's superiority over traditional methods such as SHAP and Integrated Gradients, achieving a 3.5% improvement in accuracy and a 3.2% boost in F1-score. By generating speech data that encapsulates novel markers, RSF not only mitigates the limitations of real data scarcity but also significantly enhances the robustness and accuracy of AD diagnostic models. These findings underscore RSF's potential as a transformative tool in speech-based AD detection, offering new insights into AD-related linguistic deficits and paving the way for more effective non-invasive early intervention strategies.
【10】 Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
标题: 探索分段和词汇大小对语音语言模型语音标记化的影响链接:https://arxiv.org/abs/2505.17446
备注:Accepted to Interspeech2025
摘要:语音标记化的目的是将语音信号转换为一系列离散表示,作为语音语言模型(SLM)的基础。虽然语音标记化有很多选择,但它们对SLM性能的影响仍然不清楚。本文研究了语音标记化的两个关键问题:分割宽度和离散单元的聚类大小。首先,我们将语音信号分割成固定/可变宽度和合并表示。然后,我们在多个集群大小中训练K均值模型。通过对零镜头(zero-shot)口语理解基准的评估,我们发现适度粗分割和较大的集群大小具有积极的作用。值得注意的是,在性能最好的模型中,最有效的模型实现了训练数据减少50%,训练运行时间减少70%。我们的分析强调了结合多个标记以增强细粒度口语理解的重要性。
摘要:The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, their effect on the performance of SLMs remains unclear. This paper investigates two key aspects of speech tokenization: the segmentation width and the cluster size of discrete units. First, we segment speech signals into fixed/variable widths and pooled representations. We then train K-means models in multiple cluster sizes. Through the evaluation on zero-shot spoken language understanding benchmarks, we find the positive effect of moderately coarse segmentation and bigger cluster size. Notably, among the best-performing models, the most efficient one achieves a 50% reduction in training data and a 70% decrease in training runtime. Our analysis highlights the importance of combining multiple tokens to enhance fine-grained spoken language understanding.
【11】 UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
标题: UniTTC:一个无需声学和语义信息脱钩的端到端TTC系统链接:https://arxiv.org/abs/2505.17426
摘要:诸如残差矢量量化(RVQ)和组矢量量化(GVQ)的多码本中性音频编解码器的出现显著地推进了基于大语言模型(LLM)的文本到语音(TTS)系统。这些编解码器在分离语义和声学信息同时有效利用语义先验方面至关重要。然而,由于语义和声学信息不能完全对齐,当应用于基于LLM的TTS时,这些方法的显著缺点是大型语言模型对全面音频信息的访问可能有限。为了解决这个问题,我们提出了DistilCodec和UniTTS,它们共同提供了以下优点:1)该方法可以将多码本音频编解码器提取为具有32,768个代码的单码本音频编解码器,同时实现接近100%的利用率。2)由于DistilCodec不采用语义对齐方案,因此大量高质量的未标记音频(例如具有声音效果的有声读物,歌曲等)可以在培训过程中纳入,进一步扩大数据多样性并扩大其适用性。3)利用DistilCodec的全面音频信息建模,我们将三个关键任务集成到UniTTS的预训练框架中:音频模态自回归,文本模态自回归和语音-文本跨模态自回归。这允许UniTTS接受交错的文本和语音/音频提示,同时基本上保留LLM的文本功能。4)UniTTS采用三个阶段的训练过程:预训练,监督微调(SFT)和对齐。源代码和模型检查点可在https://github.com/IDEA-Emdoor-Lab/UniTTS和https://github.com/IDEA-Emdoor-Lab/DistilCodec上公开获取。
摘要:The emergence of multi-codebook neutral audio codecs such as Residual Vector Quantization (RVQ) and Group Vector Quantization (GVQ) has significantly advanced Large-Language-Model (LLM) based Text-to-Speech (TTS) systems. These codecs are crucial in separating semantic and acoustic information while efficiently harnessing semantic priors. However, since semantic and acoustic information cannot be fully aligned, a significant drawback of these methods when applied to LLM-based TTS is that large language models may have limited access to comprehensive audio information. To address this limitation, we propose DistilCodec and UniTTS, which collectively offer the following advantages: 1) This method can distill a multi-codebook audio codec into a single-codebook audio codec with 32,768 codes while achieving a near 100\% utilization. 2) As DistilCodec does not employ a semantic alignment scheme, a large amount of high-quality unlabeled audio (such as audiobooks with sound effects, songs, etc.) can be incorporated during training, further expanding data diversity and broadening its applicability. 3) Leveraging the comprehensive audio information modeling of DistilCodec, we integrated three key tasks into UniTTS's pre-training framework: audio modality autoregression, text modality autoregression, and speech-text cross-modal autoregression. This allows UniTTS to accept interleaved text and speech/audio prompts while substantially preserving LLM's text capabilities. 4) UniTTS employs a three-stage training process: Pre-Training, Supervised Fine-Tuning (SFT), and Alignment. Source code and model checkpoints are publicly available at https://github.com/IDEA-Emdoor-Lab/UniTTS and https://github.com/IDEA-Emdoor-Lab/DistilCodec.
【12】 LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
标题: 基于LLM的具有合成数据和语音上下文的稀有词生成式错误纠正链接:https://arxiv.org/abs/2505.17410
备注:Accepted by INTERSPEECH 2025
摘要:基于大语言模型的生成式纠错(GER)是提高自动语音识别(ASR)性能的有效后处理方法。然而,由于训练数据有限,它经常遇到罕见或特定领域的单词。此外,现有的基于LLM的GER方法主要依赖于文本信息,忽略了语音线索,这导致过度校正。为了解决这些问题,我们提出了一种新的基于LLM的GER方法,目标是罕见的单词,并结合语音信息。首先,我们生成合成数据,以包含用于微调GER模型的稀有词。其次,我们整合ASR的N-最好的假设随着语音上下文,以减轻过度校正。实验结果表明,该方法不仅提高了稀有词的正确率,而且降低了英语和日语数据集的WER和CER。
摘要:Generative error correction (GER) with large language models (LLMs) has emerged as an effective post-processing approach to improve automatic speech recognition (ASR) performance. However, it often struggles with rare or domain-specific words due to limited training data. Furthermore, existing LLM-based GER approaches primarily rely on textual information, neglecting phonetic cues, which leads to over-correction. To address these issues, we propose a novel LLM-based GER approach that targets rare words and incorporates phonetic information. First, we generate synthetic data to contain rare words for fine-tuning the GER model. Second, we integrate ASR's N-best hypotheses along with phonetic context to mitigate over-correction. Experimental results show that our method not only improves the correction of rare words but also reduces the WER and CER across both English and Japanese datasets.
【13】 VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
标题: VoxRAG:在口语问题回答中迈向免转录RAG系统的一步链接:https://arxiv.org/abs/2505.17326
备注:Accepted to ACL 2025 Workshop MAGMaR
摘要:我们介绍VoxRAG,一个模块化的语音到语音检索增强生成系统,绕过转录检索语义相关的音频片段直接从口头查询。VoxRAG采用沉默感知分割,说话人日记化,CLAP音频嵌入和使用L2归一化余弦相似性的FAISS检索。我们构建了一个50个查询的测试集记录为口语输入的母语为英语的人。检索质量进行了评价,使用LLM-as-a-judge注释。对于非常相关的片段,余弦相似性达到0.34的Recall@10。对于一些相关的部分,Recall@10上升到0.60,nDCG@10上升到0.27,突出了强烈的主题对齐。答案质量在相关性、准确性、完整性和精确性方面以0- 2的尺度进行判断,平均得分分别为0.84、0.58、0.56和0.46。虽然精度和检索质量仍然是关键的限制,VoxRAG表明,转录免费语音到语音检索在RAG系统是可行的。
摘要:We introduce VoxRAG, a modular speech-to-speech retrieval-augmented generation system that bypasses transcription to retrieve semantically relevant audio segments directly from spoken queries. VoxRAG employs silence-aware segmentation, speaker diarization, CLAP audio embeddings, and FAISS retrieval using L2-normalized cosine similarity. We construct a 50-query test set recorded as spoken input by a native English speaker. Retrieval quality was evaluated using LLM-as-a-judge annotations. For very relevant segments, cosine similarity achieved a Recall@10 of 0.34. For somewhat relevant segments, Recall@10 rose to 0.60 and nDCG@10 to 0.27, highlighting strong topical alignment. Answer quality was judged on a 0--2 scale across relevance, accuracy, completeness, and precision, with mean scores of 0.84, 0.58, 0.56, and 0.46 respectively. While precision and retrieval quality remain key limitations, VoxRAG shows that transcription-free speech-to-speech retrieval is feasible in RAG systems.
【14】 Benchmarking Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
标题: 使用VITS和风格对表达性日语字符文本到语音进行基准测试-BERT-VITS 2链接:https://arxiv.org/abs/2505.17320
摘要:由于音高重音敏感性和风格的变化性,合成富有表现力的日语字符语音构成了独特的挑战。本文对两个开源的文本到语音转换模型--VITS和Style-BERT-VITS 2 JP Extra(SBV 2 JE)--在域内的字符驱动的日语语音上进行了基准测试。使用三个特定于字符的数据集,我们评估模型的自然度(平均意见和比较平均意见得分),可懂度(单词错误率)和说话人一致性。SBV 2 JE在自然度方面与人类地面实况相匹配(MOS 4.37 vs. 4.38),实现了较低的WER,并在CMOS中显示出轻微的偏好。SBV 2 JE通过音高重音控制和基于WavLM的语音增强,证明了它对语言学习和角色对话生成等应用程序的有效性,尽管计算要求更高。
摘要:Synthesizing expressive Japanese character speech poses unique challenges due to pitch-accent sensitivity and stylistic variability. This paper benchmarks two open-source text-to-speech models--VITS and Style-BERT-VITS2 JP Extra (SBV2JE)--on in-domain, character-driven Japanese speech. Using three character-specific datasets, we evaluate models across naturalness (mean opinion and comparative mean opinion score), intelligibility (word error rate), and speaker consistency. SBV2JE matches human ground truth in naturalness (MOS 4.37 vs. 4.38), achieves lower WER, and shows slight preference in CMOS. Enhanced by pitch-accent controls and a WavLM-based discriminator, SBV2JE proves effective for applications like language learning and character dialogue generation, despite higher computational demands.
【15】 Understanding the Algorithm Behind Audio Key Detection
标题: 了解音频密钥检测背后的算法链接:https://arxiv.org/abs/2505.17259
备注:Preprint. Describes an algorithmic approach to musical key detection implemented in Python. Includes conceptual explanation of audio feature extraction and key profile matching
摘要:音乐基调的确定是音乐理论和感知的一个基本方面,为旋律和和弦进行提供和声背景。自动化这个过程,称为自动键检测,是音乐信息检索(MIR)领域的一项重要任务。本文概述了一种通过数字信号处理技术分析录音的音调内容并与理论音调配置文件进行比较来估计录音音乐音调的算法方法。
摘要:The determination of musical key is a fundamental aspect of music theory and perception, providing a harmonic context for melodies and chord progressions. Automating this process, known as automatic key detection, is a significant task in the field of Music Information Retrieval (MIR). This article outlines an algorithmic methodology for estimating the musical key of an audio recording by analyzing its tonal content through digital signal processing techniques and comparison with theoretical key profiles.
【16】 Semantic-Aware Interpretable Multimodal Music Auto-Tagging
标题: 语义感知的可解释多模式音乐自动标记链接:https://arxiv.org/abs/2505.17233
摘要:音乐自动标记对于在广泛的数字图书馆中组织和发现音乐至关重要。虽然基础模型在这一领域取得了卓越的性能,但它们的输出往往缺乏可解释性,限制了研究人员和最终用户的信任和可用性。在这项工作中,我们提出了一个可解释的音乐自动标记框架,该框架利用了来自信号处理、深度学习、本体工程和自然语言处理的具有音乐意义的多模态特征。为了提高可解释性,我们聚类特征语义和采用期望最大化算法,分配不同的权重,每个组的基础上,其对标记过程的贡献。我们的方法实现了具有竞争力的标记性能,同时提供了对决策过程的更深入理解,为更透明和以用户为中心的音乐标记系统铺平了道路。
摘要:Music auto-tagging is essential for organizing and discovering music in extensive digital libraries. While foundation models achieve exceptional performance in this domain, their outputs often lack interpretability, limiting trust and usability for researchers and end-users alike. In this work, we present an interpretable framework for music auto-tagging that leverages groups of musically meaningful multimodal features, derived from signal processing, deep learning, ontology engineering, and natural language processing. To enhance interpretability, we cluster features semantically and employ an expectation maximization algorithm, assigning distinct weights to each group based on its contribution to the tagging process. Our method achieves competitive tagging performance while offering a deeper understanding of the decision-making process, paving the way for more transparent and user-centric music tagging systems.
【17】 Large Language Models Implicitly Learn to See and Hear Just By Reading
标题: 大型语言模型通过阅读就能学会看和听链接:https://arxiv.org/abs/2505.17091
备注:6 pages, 3 figures, 4 tables. Under Review WASPAA 2025
摘要:本文提出了一个有趣的发现:通过在文本标记上训练自回归LLM模型,文本模型内在地发展了内部理解图像和音频的能力,从而发展了仅通过阅读就能看到和听到的能力。流行的音频和视觉LLM模型微调文本LLM模型,以提供基于图像和音频嵌入的文本输出。另一方面,我们的架构接受图像、音频波形或令牌的补丁作为输入。它为我们提供了分类管道中典型的嵌入或类别标签。我们展示了文本权重在帮助数据集FSD-50 K和GTZAN的音频分类中的一般性。此外,我们还展示了CIFAR-10和Fashion-MNIST以及图像补丁上的图像分类。这推动了文本LLM学习强大的内部电路的概念,这些电路可以通过激活各种应用程序的必要连接来利用,而不是每次都从头开始训练模型。
摘要:This paper presents a fascinating find: By training an auto-regressive LLM model on text tokens, the text model inherently develops internally an ability to understand images and audio, thereby developing the ability to see and hear just by reading. Popular audio and visual LLM models fine-tune text LLM models to give text output conditioned on images and audio embeddings. On the other hand, our architecture takes in patches of images, audio waveforms or tokens as input. It gives us the embeddings or category labels typical of a classification pipeline. We show the generality of text weights in aiding audio classification for datasets FSD-50K and GTZAN. Further, we show this working for image classification on CIFAR-10 and Fashion-MNIST, as well on image patches. This pushes the notion of text-LLMs learning powerful internal circuits that can be utilized by activating necessary connections for various applications rather than training models from scratch every single time.
【18】 Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
标题: 帧速率对语音标记器的影响--以汉语和英语为例链接:https://arxiv.org/abs/2505.17076
备注:5 pages, 5 figures
摘要:语音标记器在最近的语音任务中起着至关重要的作用,通常充当语音信号和语言模型之间的桥梁。虽然低帧速率编解码器被广泛用作语音标记器,但帧速率对语音标记的影响仍然没有得到充分研究。在这项研究中,我们调查如何不同的帧速率影响语音标记通过检查普通话和英语,两种类型不同的语言。我们以不同的帧速率对语音进行编码,并在语音识别任务中评估由此产生的语义标记。我们的研究结果表明,帧速率的变化对每种语言的语音标记化的影响不同,突出了帧速率,语音密度和语言特定的声学特征之间的相互作用。研究结果为优化语音标记器的帧速率选择提供了见解,并对自动语音识别,文本到语音以及其他语音相关应用产生了影响。
摘要:The speech tokenizer plays a crucial role in recent speech tasks, generally serving as a bridge between speech signals and language models. While low-frame-rate codecs are widely employed as speech tokenizers, the impact of frame rates on speech tokens remains underexplored. In this study, we investigate how varying frame rates affect speech tokenization by examining Mandarin and English, two typologically distinct languages. We encode speech at different frame rates and evaluate the resulting semantic tokens in the speech recognition task. Our findings reveal that frame rate variations influence speech tokenization differently for each language, highlighting the interplay between frame rates, phonetic density, and language-specific acoustic features. The results provide insights into optimizing frame rate selection for speech tokenizers, with implications for automatic speech recognition, text-to-speech, and other speech-related applications.
【19】 Improving endpoint detection in end-to-end streaming ASR for conversational speech
标题: 改进对话语音的端到端流ASB中的端点检测链接:https://arxiv.org/abs/2505.17070
备注:Submitted to Interspeech 2024
摘要:ASR端点(EP)在人机对话中支持人类或人工代理的产品中提供良好的用户体验方面发挥着重要作用。基于传感器的ASR(T-ASR)是流媒体首选的端到端(E2 E)ASR建模技术。T-ASR的主要限制是ASR输出的延迟发射,这可能导致EP中的错误或延迟。不准确的EP将在说话时切断用户,返回不完整的转录,而EP中的延迟将增加感知延迟,降低用户体验。我们提出了通过解决延迟发射以及EP错误来改善EP的方法。为了解决延迟发射问题,我们在每个单词的末尾引入了一个单词结束标记,以及延迟惩罚。通过使用辅助网络获得可靠的帧级语音活动检测来解决EP延迟。我们将所提出的方法应用于Switchboard会话语音语料库,并对延迟惩罚方法进行评估。
摘要:ASR endpointing (EP) plays a major role in delivering a good user experience in products supporting human or artificial agents in human-human/machine conversations. Transducer-based ASR (T-ASR) is an end-to-end (E2E) ASR modelling technique preferred for streaming. A major limitation of T-ASR is delayed emission of ASR outputs, which could lead to errors or delays in EP. Inaccurate EP will cut the user off while speaking, returning incomplete transcript while delays in EP will increase the perceived latency, degrading the user experience. We propose methods to improve EP by addressing delayed emission along with EP mistakes. To address the delayed emission problem, we introduce an end-of-word token at the end of each word, along with a delay penalty. The EP delay is addressed by obtaining a reliable frame-level speech activity detection using an auxiliary network. We apply the proposed methods on Switchboard conversational speech corpus and evaluate it against a delay penalty method.
【20】 ReMi: A Random Recurrent Neural Network Approach to Music Production
标题: ReMi:音乐制作的随机回归神经网络方法链接:https://arxiv.org/abs/2505.17023
备注:Accepted for an Innovation Showcase Demo at International Computer Music Conference
摘要:生成型人工智能引起了人们对能源消耗、侵犯版权和创造性萎缩的担忧。我们证明了随机初始化的递归神经网络可以产生丰富且可配置的琶音和低频振荡。与旨在取代音乐家的端到端音乐生成相反,我们的方法扩展了他们的创造力,同时不需要数据和更少的计算能力。更多信息请访问:https://allendia.com/
摘要:Generative artificial intelligence raises concerns related to energy consumption, copyright infringement and creative atrophy. We show that randomly initialized recurrent neural networks can produce arpeggios and low-frequency oscillations that are rich and configurable. In contrast to end-to-end music generation that aims to replace musicians, our approach expands their creativity while requiring no data and much less computational power. More information can be found at: https://allendia.com/
【21】 Effects of auditory distance cues and reverberation on spatial perception and listening strategies
标题: 听觉距离线索和回响对空间感知和听力策略的影响链接:https://arxiv.org/abs/2505.18020
备注:13 pages, 6 figures
摘要:空间听觉是大脑利用听觉线索识别声音来源的能力,对日常听力至关重要。虽然简化的范式促进了对空间听觉的理解,但它们缺乏生态有效性,限制了它们在现实生活条件下的适用性。本研究旨在解决这一差距,通过调查听众的运动,混响和距离的影响定位精度在一个更生态有效的上下文中。参与者在无回声或混响条件下进行主动定位任务,没有具体的听力策略指示。结果表明,头部运动更频繁的混响环境中,提出了一种自适应策略,以减轻由于混响双耳线索的不确定性。虽然距离并不影响听力策略,但它影响了定位性能。我们的研究结果表明,听力行为适应取决于当前的声学条件,以支持有效的感知空间。
摘要:Spatial hearing, the brain's ability to use auditory cues to identify the origin of sounds, is crucial for everyday listening. While simplified paradigms have advanced the understanding of spatial hearing, their lack of ecological validity limits their applicability to real-life conditions. This study aims to address this gap by investigating the effects of listener movement, reverberation, and distance on localisation accuracy in a more ecologically valid context. Participants performed active localisation tasks with no specific instructions on listening strategy, in either anechoic or reverberant conditions. The results indicate that the head movements were more frequent in reverberant environments, suggesting an adaptive strategy to mitigate uncertainty in binaural cues due to reverberation. While distance did not affect the listening strategy, it influenced the localisation performance. Our outcomes suggest that listening behaviour is adapted depending on the current acoustic conditions to support an effective perception of the space.
【22】 Source Separation of Small Classical Ensembles: Challenges and Opportunities
标题: 小型古典合奏的来源分离:挑战与机遇链接:https://arxiv.org/abs/2505.17823
备注:5 pages, 4 figures, 2 tables, submitted to WASSPA 2025
摘要:使用非因果深度学习对西方流行音乐进行音乐(MSS)源分离可以非常有效。相比之下,古典音乐的MSS是一个未解决的问题。古典合奏比流行音乐更难分离,因为音乐中固有的更大变化;用于监督训练的地面真实录音的稀疏性;以及乐器之间的模糊性。华彩乐段项目一直在探索古典音乐的MSS。这样做是为了让音乐可以重新混合,以改善听力损失患者的听觉体验。为了实现这项工作,一个新的合成木管合奏数据库被创建,以克服EnsembleSet中的乐器不平衡。对于MSS,使用了一组ConvTasNet模型,每个模型都经过训练以提取弦乐器或木管乐器。之所以选择ConvTasNet,是因为它能够测试因果和非因果方法。非因果方法主导了MSS的工作,对录制的音乐很有用,但对于现场音乐或助听器处理,需要因果信号处理。MSS的性能进行了评估的两个小的数据集(巴赫10和URMP)的真实仪器记录,其中地面真理是可用的。因果和非因果系统的性能是相似的。将合成验证集(6.2 dB因果; 6.9非因果)的平均信号失真(SDR)与真实记录的评估集(0.3 dB因果,0.4 dB非因果)进行比较,表明合成数据与记录数据之间的失配是一个问题。未来的工作需要收集更多可用于训练的真实录音,或者提高合成录音的真实性和多样性,以减少不匹配。
摘要:Musical (MSS) source separation of western popular music using non-causal deep learning can be very effective. In contrast, MSS for classical music is an unsolved problem. Classical ensembles are harder to separate than popular music because of issues such as the inherent greater variation in the music; the sparsity of recordings with ground truth for supervised training; and greater ambiguity between instruments. The Cadenza project has been exploring MSS for classical music. This is being done so music can be remixed to improve listening experiences for people with hearing loss. To enable the work, a new database of synthesized woodwind ensembles was created to overcome instrumental imbalances in the EnsembleSet. For the MSS, a set of ConvTasNet models was used with each model being trained to extract a string or woodwind instrument. ConvTasNet was chosen because it enabled both causal and non-causal approaches to be tested. Non-causal approaches have dominated MSS work and are useful for recorded music, but for live music or processing on hearing aids, causal signal processing is needed. The MSS performance was evaluated on the two small datasets (Bach10 and URMP) of real instrument recordings where the ground-truth is available. The performances of the causal and non-causal systems were similar. Comparing the average Signal-to-Distortion (SDR) of the synthesized validation set (6.2 dB causal; 6.9 non-causal), to the real recorded evaluation set (0.3 dB causal, 0.4 dB non-causal), shows that mismatch between synthesized and recorded data is a problem. Future work needs to either gather more real recordings that can be used for training, or to improve the realism and diversity of the synthesized recordings to reduce the mismatch...
【23】 Audio-to-Audio Emotion Conversion With Pitch And Duration Style Transfer
标题: 具有音调和持续时间风格转换的音频到音频情感转换链接:https://arxiv.org/abs/2505.17655
备注:11 pages, 9 figures, 5 tables
摘要:给定一对源语音记录和参考语音记录,音频到音频(A2 A)风格转换涉及生成模仿参考语音的风格特征同时保留源语音的内容和说话者属性的输出语音。在本文中,我们提出了一种新的框架,称为A2 A Zero-shot情感风格转移(A2 A-ZEST),使参考情感属性的源,同时保留其扬声器和语音内容的转移。A2 A-ZEST框架由一个分析-合成管道组成,其中分析模块将语音分解为语义标记、说话者表示和情感嵌入。使用这些表示,学习音高轮廓估计器和持续时间预测器。此外,合成模块被设计为基于输入表示和导出因子来生成语音。这整个分析-合成的范例纯粹是以自我监督的方式训练的,具有自动编码损失。对于A2 A情感风格转换,在合成模块中使用从参考语音提取的情感嵌入以及来自源语音的其余表示来生成风格转换语音。在我们的实验中,我们对转换后的语音进行了内容/说话人保留(w.r.t.来源)以及情绪风格转移的有效性(w.r.t.参考)。该提案A2 A-ZEST被证明比其他先前的评估工作有所改进,从而在没有任何并行训练数据的情况下实现风格迁移。我们还说明了应用所提出的工作,情感识别任务中的数据增强。
摘要:Given a pair of source and reference speech recordings, audio-to-audio (A2A) style transfer involves the generation of an output speech that mimics the style characteristics of the reference while preserving the content and speaker attributes of the source. In this paper, we propose a novel framework, termed as A2A Zero-shot Emotion Style Transfer (A2A-ZEST), that enables the transfer of reference emotional attributes to the source while retaining its speaker and speech contents. The A2A-ZEST framework consists of an analysis-synthesis pipeline, where the analysis module decomposes speech into semantic tokens, speaker representations, and emotion embeddings. Using these representations, a pitch contour estimator and a duration predictor are learned. Further, a synthesis module is designed to generate speech based on the input representations and the derived factors. This entire paradigm of analysis-synthesis is trained purely in a self-supervised manner with an auto-encoding loss. For A2A emotion style transfer, the emotion embedding extracted from the reference speech along with the rest of the representations from the source speech are used in the synthesis module to generate the style translated speech. In our experiments, we evaluate the converted speech on content/speaker preservation (w.r.t. source) as well as on the effectiveness of the emotion style transfer (w.r.t. reference). The proposal, A2A-ZEST, is shown to improve over other prior works on these evaluations, thereby enabling style transfer without any parallel training data. We also illustrate the application of the proposed work for data augmentation in emotion recognition tasks.
【24】 Private kNN-VC: Interpretable Anonymization of Converted Speech
标题: 私人kNN-VC:转换语音的可解释分析链接:https://arxiv.org/abs/2505.17584
备注:Accepted by Interspeech 2025
摘要:说话人匿名化试图隐藏说话人的身份,同时保留他们的演讲的效用。所实现的隐私通常使用根据匿名语音训练的说话者识别模型进行评估。虽然这代表了一种强烈的攻击,但目前还不清楚语音的哪些方面被用来识别说话者。我们的研究旨在揭示这些方面。它从kNN-VC开始,这是一个强大的语音转换模型,作为匿名系统性能不佳,可能是因为韵律泄漏。为了验证这一假设,我们扩展了kNN-VC与两个可解释的组件,匿名的持续时间和变化的电话。这些组件增加隐私显着,证明所研究的韵律因素编码扬声器的身份,并利用隐私攻击。此外,我们表明,在目标选择算法的变化显着影响隐私攻击的结果。
摘要:Speaker anonymization seeks to conceal a speaker's identity while preserving the utility of their speech. The achieved privacy is commonly evaluated with a speaker recognition model trained on anonymized speech. Although this represents a strong attack, it is unclear which aspects of speech are exploited to identify the speakers. Our research sets out to unveil these aspects. It starts with kNN-VC, a powerful voice conversion model that performs poorly as an anonymization system, presumably because of prosody leakage. To test this hypothesis, we extend kNN-VC with two interpretable components that anonymize the duration and variation of phones. These components increase privacy significantly, proving that the studied prosodic factors encode speaker identity and are exploited by the privacy attack. Additionally, we show that changes in the target selection algorithm considerably influence the outcome of the privacy attack.
【25】 Speechless: Speech Instruction Training Without Speech for Low Resource Languages
标题: Speechless:低资源语言的无语音语音教学训练链接:https://arxiv.org/abs/2505.17417
备注:This paper was accepted by INTERSPEECH 2025
摘要:由大型语言模型(LLM)驱动的语音助理的快速增长突出了对语音指令数据的需求,以训练这些系统。尽管语音识别数据非常丰富,但语音指令数据却非常稀缺,这对于微调模型以理解和执行口头命令至关重要。生成高质量的合成语音需要良好的文本到语音(TTS)模型,这可能不适用于低资源语言。我们的新方法通过在语义表示级别停止合成来解决这一挑战,绕过了对TTS的需求。我们通过将合成语义表示与预训练的Whisper编码器对齐来实现这一点,使LLM能够在文本指令上进行微调,同时保持在推理过程中理解口头指令的能力。这种简化的培训过程是为低资源语言构建语音助手的一种很有前途的方法。
摘要:The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of speech recognition data, there is a notable scarcity of speech instruction data, which is essential for fine-tuning models to understand and execute spoken commands. Generating high-quality synthetic speech requires a good text-to-speech (TTS) model, which may not be available to low resource languages. Our novel approach addresses this challenge by halting synthesis at the semantic representation level, bypassing the need for TTS. We achieve this by aligning synthetic semantic representations with the pre-trained Whisper encoder, enabling an LLM to be fine-tuned on text instructions while maintaining the ability to understand spoken instructions during inference. This simplified training process is a promising approach to building voice assistant for low-resource languages.
【26】 From Weak Labels to Strong Results: Utilizing 5,000 Hours of Noisy Classroom Transcripts with Minimal Accurate Data
标题: 从弱标签到强结果:利用5,000小时喧闹的课堂成绩单和最低准确数据链接:https://arxiv.org/abs/2505.17088
摘要:语音识别的最新进展依赖于在大量标记数据上训练的模型。然而,课堂自动语音识别(ASR)面临着现实世界的挑战,即大量的弱成绩单只与少量准确的黄金标准数据配对。在这种低资源环境下,高转录成本使得重新转录不切实际。为解决这一问题,我们要求:当大量廉价的弱成绩单与有限的黄金标准数据共存时,如课堂语音数据的情况,什么是最好的方法?我们提出了弱监督预训练(WSP),这是一个两步的过程,首先以监督的方式对弱成绩单进行预训练,然后对准确的数据进行微调。我们的研究结果,基于合成和真正的弱成绩单,表明WSP优于替代方法,建立它作为一个有效的训练方法,在现实世界中的低资源ASR。
摘要:Recent progress in speech recognition has relied on models trained on vast amounts of labeled data. However, classroom Automatic Speech Recognition (ASR) faces the real-world challenge of abundant weak transcripts paired with only a small amount of accurate, gold-standard data. In such low-resource settings, high transcription costs make re-transcription impractical. To address this, we ask: what is the best approach when abundant inexpensive weak transcripts coexist with limited gold-standard data, as is the case for classroom speech data? We propose Weakly Supervised Pretraining (WSP), a two-step process where models are first pretrained on weak transcripts in a supervised manner, and then fine-tuned on accurate data. Our results, based on both synthetic and real weak transcripts, show that WSP outperforms alternative methods, establishing it as an effective training methodology for low-resource ASR in real-world scenarios.
标题: 听觉距离线索和回响对空间感知和听力策略的影响
链接:https://arxiv.org/abs/2505.18020
备注:13 pages, 6 figures
摘要:空间听觉是大脑利用听觉线索识别声音来源的能力,对日常听力至关重要。虽然简化的范式促进了对空间听觉的理解,但它们缺乏生态有效性,限制了它们在现实生活条件下的适用性。本研究旨在解决这一差距,通过调查听众的运动,混响和距离的影响定位精度在一个更生态有效的上下文中。参与者在无回声或混响条件下进行主动定位任务,没有具体的听力策略指示。结果表明,头部运动更频繁的混响环境中,提出了一种自适应策略,以减轻由于混响双耳线索的不确定性。虽然距离并不影响听力策略,但它影响了定位性能。我们的研究结果表明,听力行为适应取决于当前的声学条件,以支持有效的感知空间。
摘要:Spatial hearing, the brain's ability to use auditory cues to identify the origin of sounds, is crucial for everyday listening. While simplified paradigms have advanced the understanding of spatial hearing, their lack of ecological validity limits their applicability to real-life conditions. This study aims to address this gap by investigating the effects of listener movement, reverberation, and distance on localisation accuracy in a more ecologically valid context. Participants performed active localisation tasks with no specific instructions on listening strategy, in either anechoic or reverberant conditions. The results indicate that the head movements were more frequent in reverberant environments, suggesting an adaptive strategy to mitigate uncertainty in binaural cues due to reverberation. While distance did not affect the listening strategy, it influenced the localisation performance. Our outcomes suggest that listening behaviour is adapted depending on the current acoustic conditions to support an effective perception of the space.
【2】 Source Separation of Small Classical Ensembles: Challenges and Opportunities
标题: 小型古典合奏的来源分离:挑战与机遇链接:https://arxiv.org/abs/2505.17823
备注:5 pages, 4 figures, 2 tables, submitted to WASSPA 2025
摘要:使用非因果深度学习对西方流行音乐进行音乐(MSS)源分离可以非常有效。相比之下,古典音乐的MSS是一个未解决的问题。古典合奏比流行音乐更难分离,因为音乐中固有的更大变化;用于监督训练的地面真实录音的稀疏性;以及乐器之间的模糊性。华彩乐段项目一直在探索古典音乐的MSS。这样做是为了让音乐可以重新混合,以改善听力损失患者的听觉体验。为了实现这项工作,一个新的合成木管合奏数据库被创建,以克服EnsembleSet中的乐器不平衡。对于MSS,使用了一组ConvTasNet模型,每个模型都经过训练以提取弦乐器或木管乐器。之所以选择ConvTasNet,是因为它能够测试因果和非因果方法。非因果方法主导了MSS的工作,对录制的音乐很有用,但对于现场音乐或助听器处理,需要因果信号处理。MSS的性能进行了评估的两个小的数据集(巴赫10和URMP)的真实仪器记录,其中地面真理是可用的。因果和非因果系统的性能是相似的。将合成验证集(6.2 dB因果; 6.9非因果)的平均信号失真(SDR)与真实记录的评估集(0.3 dB因果,0.4 dB非因果)进行比较,表明合成数据与记录数据之间的失配是一个问题。未来的工作需要收集更多可用于训练的真实录音,或者提高合成录音的真实性和多样性,以减少不匹配。
摘要:Musical (MSS) source separation of western popular music using non-causal deep learning can be very effective. In contrast, MSS for classical music is an unsolved problem. Classical ensembles are harder to separate than popular music because of issues such as the inherent greater variation in the music; the sparsity of recordings with ground truth for supervised training; and greater ambiguity between instruments. The Cadenza project has been exploring MSS for classical music. This is being done so music can be remixed to improve listening experiences for people with hearing loss. To enable the work, a new database of synthesized woodwind ensembles was created to overcome instrumental imbalances in the EnsembleSet. For the MSS, a set of ConvTasNet models was used with each model being trained to extract a string or woodwind instrument. ConvTasNet was chosen because it enabled both causal and non-causal approaches to be tested. Non-causal approaches have dominated MSS work and are useful for recorded music, but for live music or processing on hearing aids, causal signal processing is needed. The MSS performance was evaluated on the two small datasets (Bach10 and URMP) of real instrument recordings where the ground-truth is available. The performances of the causal and non-causal systems were similar. Comparing the average Signal-to-Distortion (SDR) of the synthesized validation set (6.2 dB causal; 6.9 non-causal), to the real recorded evaluation set (0.3 dB causal, 0.4 dB non-causal), shows that mismatch between synthesized and recorded data is a problem. Future work needs to either gather more real recordings that can be used for training, or to improve the realism and diversity of the synthesized recordings to reduce the mismatch...
【3】 Audio-to-Audio Emotion Conversion With Pitch And Duration Style Transfer
标题: 具有音调和持续时间风格转换的音频到音频情感转换链接:https://arxiv.org/abs/2505.17655
备注:11 pages, 9 figures, 5 tables
摘要:给定一对源语音记录和参考语音记录,音频到音频(A2 A)风格转换涉及生成模仿参考语音的风格特征同时保留源语音的内容和说话者属性的输出语音。在本文中,我们提出了一种新的框架,称为A2 A Zero-shot情感风格转移(A2 A-ZEST),使参考情感属性的源,同时保留其扬声器和语音内容的转移。A2 A-ZEST框架由一个分析-合成管道组成,其中分析模块将语音分解为语义标记、说话者表示和情感嵌入。使用这些表示,学习音高轮廓估计器和持续时间预测器。此外,合成模块被设计为基于输入表示和导出因子来生成语音。这整个分析-合成的范例纯粹是以自我监督的方式训练的,具有自动编码损失。对于A2 A情感风格转换,在合成模块中使用从参考语音提取的情感嵌入以及来自源语音的其余表示来生成风格转换语音。在我们的实验中,我们对转换后的语音进行了内容/说话人保留(w.r.t.来源)以及情绪风格转移的有效性(w.r.t.参考)。该提案A2 A-ZEST被证明比其他先前的评估工作有所改进,从而在没有任何并行训练数据的情况下实现风格迁移。我们还说明了应用所提出的工作,情感识别任务中的数据增强。
摘要:Given a pair of source and reference speech recordings, audio-to-audio (A2A) style transfer involves the generation of an output speech that mimics the style characteristics of the reference while preserving the content and speaker attributes of the source. In this paper, we propose a novel framework, termed as A2A Zero-shot Emotion Style Transfer (A2A-ZEST), that enables the transfer of reference emotional attributes to the source while retaining its speaker and speech contents. The A2A-ZEST framework consists of an analysis-synthesis pipeline, where the analysis module decomposes speech into semantic tokens, speaker representations, and emotion embeddings. Using these representations, a pitch contour estimator and a duration predictor are learned. Further, a synthesis module is designed to generate speech based on the input representations and the derived factors. This entire paradigm of analysis-synthesis is trained purely in a self-supervised manner with an auto-encoding loss. For A2A emotion style transfer, the emotion embedding extracted from the reference speech along with the rest of the representations from the source speech are used in the synthesis module to generate the style translated speech. In our experiments, we evaluate the converted speech on content/speaker preservation (w.r.t. source) as well as on the effectiveness of the emotion style transfer (w.r.t. reference). The proposal, A2A-ZEST, is shown to improve over other prior works on these evaluations, thereby enabling style transfer without any parallel training data. We also illustrate the application of the proposed work for data augmentation in emotion recognition tasks.
【4】 Private kNN-VC: Interpretable Anonymization of Converted Speech
标题: 私人kNN-VC:转换语音的可解释分析链接:https://arxiv.org/abs/2505.17584
备注:Accepted by Interspeech 2025
摘要:说话人匿名化试图隐藏说话人的身份,同时保留他们的演讲的效用。所实现的隐私通常使用根据匿名语音训练的说话者识别模型进行评估。虽然这代表了一种强烈的攻击,但目前还不清楚语音的哪些方面被用来识别说话者。我们的研究旨在揭示这些方面。它从kNN-VC开始,这是一个强大的语音转换模型,作为匿名系统性能不佳,可能是因为韵律泄漏。为了验证这一假设,我们扩展了kNN-VC与两个可解释的组件,匿名的持续时间和变化的电话。这些组件增加隐私显着,证明所研究的韵律因素编码扬声器的身份,并利用隐私攻击。此外,我们表明,在目标选择算法的变化显着影响隐私攻击的结果。
摘要:Speaker anonymization seeks to conceal a speaker's identity while preserving the utility of their speech. The achieved privacy is commonly evaluated with a speaker recognition model trained on anonymized speech. Although this represents a strong attack, it is unclear which aspects of speech are exploited to identify the speakers. Our research sets out to unveil these aspects. It starts with kNN-VC, a powerful voice conversion model that performs poorly as an anonymization system, presumably because of prosody leakage. To test this hypothesis, we extend kNN-VC with two interpretable components that anonymize the duration and variation of phones. These components increase privacy significantly, proving that the studied prosodic factors encode speaker identity and are exploited by the privacy attack. Additionally, we show that changes in the target selection algorithm considerably influence the outcome of the privacy attack.
【5】 Speechless: Speech Instruction Training Without Speech for Low Resource Languages
标题: Speechless:低资源语言的无语音语音教学训练链接:https://arxiv.org/abs/2505.17417
备注:This paper was accepted by INTERSPEECH 2025
摘要:由大型语言模型(LLM)驱动的语音助理的快速增长突出了对语音指令数据的需求,以训练这些系统。尽管语音识别数据非常丰富,但语音指令数据却非常稀缺,这对于微调模型以理解和执行口头命令至关重要。生成高质量的合成语音需要良好的文本到语音(TTS)模型,这可能不适用于低资源语言。我们的新方法通过在语义表示级别停止合成来解决这一挑战,绕过了对TTS的需求。我们通过将合成语义表示与预训练的Whisper编码器对齐来实现这一点,使LLM能够在文本指令上进行微调,同时保持在推理过程中理解口头指令的能力。这种简化的培训过程是为低资源语言构建语音助手的一种很有前途的方法。
摘要:The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of speech recognition data, there is a notable scarcity of speech instruction data, which is essential for fine-tuning models to understand and execute spoken commands. Generating high-quality synthetic speech requires a good text-to-speech (TTS) model, which may not be available to low resource languages. Our novel approach addresses this challenge by halting synthesis at the semantic representation level, bypassing the need for TTS. We achieve this by aligning synthetic semantic representations with the pre-trained Whisper encoder, enabling an LLM to be fine-tuned on text instructions while maintaining the ability to understand spoken instructions during inference. This simplified training process is a promising approach to building voice assistant for low-resource languages.
【6】 Voicing Personas: Rewriting Persona Descriptions into Style Prompts for Controllable Text-to-Speech
标题: 配音角色:将角色扮演描述重写为可控文本到语音的风格预设链接:https://arxiv.org/abs/2505.17093
摘要:在本文中,我们提出了一个新的框架,以控制语音风格的基于文本的,可控的文本到语音系统,利用文本人物角色的语音风格提示。我们提出了两个人物角色重写策略,将通用人物角色描述转换为面向语音的提示,使细粒度的操纵韵律属性,如音高,情感和语速。实验结果表明,我们的方法提高了合成语音的自然度,清晰度和一致性。最后,我们分析了基于LLM的重写引入的隐性社会偏见,重点是性别。我们强调声音风格是角色驱动的AI对话系统的关键因素。
摘要:In this paper, we propose a novel framework to control voice style in prompt-based, controllable text-to-speech systems by leveraging textual personas as voice style prompts. We present two persona rewriting strategies to transform generic persona descriptions into speech-oriented prompts, enabling fine-grained manipulation of prosodic attributes such as pitch, emotion, and speaking rate. Experimental results demonstrate that our methods enhance the naturalness, clarity, and consistency of synthesized speech. Finally, we analyze implicit social biases introduced by LLM-based rewriting, with a focus on gender. We underscore voice style as a crucial factor for persona-driven AI dialogue systems.
【7】 From Weak Labels to Strong Results: Utilizing 5,000 Hours of Noisy Classroom Transcripts with Minimal Accurate Data
标题: 从弱标签到强结果:利用5,000小时喧闹的课堂成绩单和最低准确数据链接:https://arxiv.org/abs/2505.17088
摘要:语音识别的最新进展依赖于在大量标记数据上训练的模型。然而,课堂自动语音识别(ASR)面临着现实世界的挑战,即大量的弱成绩单只与少量准确的黄金标准数据配对。在这种低资源环境下,高转录成本使得重新转录不切实际。为解决这一问题,我们要求:当大量廉价的弱成绩单与有限的黄金标准数据共存时,如课堂语音数据的情况,什么是最好的方法?我们提出了弱监督预训练(WSP),这是一个两步的过程,首先以监督的方式对弱成绩单进行预训练,然后对准确的数据进行微调。我们的研究结果,基于合成和真正的弱成绩单,表明WSP优于替代方法,建立它作为一个有效的训练方法,在现实世界中的低资源ASR。
摘要:Recent progress in speech recognition has relied on models trained on vast amounts of labeled data. However, classroom Automatic Speech Recognition (ASR) faces the real-world challenge of abundant weak transcripts paired with only a small amount of accurate, gold-standard data. In such low-resource settings, high transcription costs make re-transcription impractical. To address this, we ask: what is the best approach when abundant inexpensive weak transcripts coexist with limited gold-standard data, as is the case for classroom speech data? We propose Weakly Supervised Pretraining (WSP), a two-step process where models are first pretrained on weak transcripts in a supervised manner, and then fine-tuned on accurate data. Our results, based on both synthetic and real weak transcripts, show that WSP outperforms alternative methods, establishing it as an effective training methodology for low-resource ASR in real-world scenarios.
【8】 DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
标题: DualTalk:用于3D说话头部对话的双扬声器交互链接:https://arxiv.org/abs/2505.18096
备注:Accepted by CVPR 2025
摘要:在面对面的交谈中,人们需要在说话和倾听角色之间无缝切换。现有的3D说话头部生成模型只关注说话或倾听,忽略了交互式对话的自然动态,这导致不自然的交互和尴尬的过渡。为了解决这个问题,我们提出了一个新的任务-多轮双扬声器交互的3D说话头生成-这需要模型来处理和生成说话和倾听行为在连续的对话。为了解决这个问题,我们引入了DualTalk,一个新颖的统一框架,它集成了说话者和听众的动态行为,以模拟逼真和连贯的对话交互。该框架不仅在说话时合成逼真的说话头,而且在倾听时产生连续和生动的非语言反馈,有效地捕捉角色之间的相互作用。我们还创建了一个新的数据集,其中包含50小时的多轮对话,超过1,000个字符,参与者在说话和倾听角色之间不断切换。大量的实验表明,我们的方法显着提高的自然性和表现力的3D说话的头在双扬声器对话。我们建议观看补充视频:https://ziqiaopeng.github.io/dualtalk。
摘要:In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task -- multi-round dual-speaker interaction for 3D talking head generation -- which requires models to handle and generate both speaking and listening behaviors in continuous conversation. To solve this task, we introduce DualTalk, a novel unified framework that integrates the dynamic behaviors of speakers and listeners to simulate realistic and coherent dialogue interactions. This framework not only synthesizes lifelike talking heads when speaking but also generates continuous and vivid non-verbal feedback when listening, effectively capturing the interplay between the roles. We also create a new dataset featuring 50 hours of multi-round conversations with over 1,000 characters, where participants continuously switch between speaking and listening roles. Extensive experiments demonstrate that our method significantly enhances the naturalness and expressiveness of 3D talking heads in dual-speaker conversations. We recommend watching the supplementary video: https://ziqiaopeng.github.io/dualtalk.
【9】 Toward Optimal ANC: Establishing Mutual Information Lower Bound
标题: 迈向最佳非国大:建立互信息下限链接:https://arxiv.org/abs/2505.17877
摘要:主动噪声消除(ANC)算法旨在通过生成实时破坏性干扰原始噪声的抗噪声信号来抑制不必要的声学干扰。尽管最近基于深度学习的ANC算法已经设定了新的性能基准,但仍然缺乏严格评估其改进的理论限制。为了解决这个问题,我们推导出一个统一的下限消除性能由两个组件。第一个组成部分是信息理论:它将残余误差功率与抗噪声信号捕获的干扰熵的分数联系起来,从而量化信息处理能力所施加的限制。第二部分是基于支持的:它测量消除路径无法解决的频带中产生的不可约误差,反映了基本的物理约束。通过取这两项的最大值,我们的界限建立了任何ANC算法可达到的归一化均方误差(NMSE)的理论上限。我们在不同混响时间下的NOISEX数据集上经验性地验证了其紧密性,证明了在不同声学条件下的鲁棒性。
摘要:Active Noise Cancellation (ANC) algorithms aim to suppress unwanted acoustic disturbances by generating anti-noise signals that destructively interfere with the original noise in real time. Although recent deep learning-based ANC algorithms have set new performance benchmarks, there remains a shortage of theoretical limits to rigorously assess their improvements. To address this, we derive a unified lower bound on cancellation performance composed of two components. The first component is information-theoretic: it links residual error power to the fraction of disturbance entropy captured by the anti-noise signal, thereby quantifying limits imposed by information-processing capacity. The second component is support-based: it measures the irreducible error arising in frequency bands that the cancellation path cannot address, reflecting fundamental physical constraints. By taking the maximum of these two terms, our bound establishes a theoretical ceiling on the Normalized Mean Squared Error (NMSE) attainable by any ANC algorithm. We validate its tightness empirically on the NOISEX dataset under varying reverberation times, demonstrating robustness across diverse acoustic conditions.
【10】 TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
标题: TEDI:用于分析和比较数据集文档的可信和道德数据集指标链接:https://arxiv.org/abs/2505.17841
摘要:数据集透明度是负责任的人工智能的关键推动因素,但对影响人工智能应用程序的可信和道德方面的多模态数据集属性的洞察仍然很少,并且难以在数据集之间进行比较。为了应对这一挑战,我们引入了值得信赖和道德的数据集指标(TEDI),以促进数据集文档的系统性实证分析。TEDI包含143个细粒度指标,表征多模式数据集及其收集过程的可信和道德属性。这些指标的框架是为了从数据集文件中提取可核实的信息。使用TEDI,我们手动注释和分析了100多个包含人类声音的多模态数据集。我们进一步注释了数据来源,大小和模态细节,以深入了解在数据集上形成可信和道德维度的因素。我们发现,只有少数数据集具有与同意,隐私和有害内容指标有关的记录属性和实践。这些和其他道德指标的处理程度因数据收集方法而异,通过众包和直接收集方法收集的数据集的文件更有可能提到这些指标。以道德指标为代价的刮取占主导地位的规模,但不是唯一可行的收集方法。我们的方法和经验见解有助于提高数据集在可信和道德维度上的透明度,并为将来从数据集文档中提取信息的繁琐任务自动化铺平道路。
摘要:Dataset transparency is a key enabler of responsible AI, but insights into multimodal dataset attributes that impact trustworthy and ethical aspects of AI applications remain scarce and are difficult to compare across datasets. To address this challenge, we introduce Trustworthy and Ethical Dataset Indicators (TEDI) that facilitate the systematic, empirical analysis of dataset documentation. TEDI encompasses 143 fine-grained indicators that characterize trustworthy and ethical attributes of multimodal datasets and their collection processes. The indicators are framed to extract verifiable information from dataset documentation. Using TEDI, we manually annotated and analyzed over 100 multimodal datasets that include human voices. We further annotated data sourcing, size, and modality details to gain insights into the factors that shape trustworthy and ethical dimensions across datasets. We find that only a select few datasets have documented attributes and practices pertaining to consent, privacy, and harmful content indicators. The extent to which these and other ethical indicators are addressed varies based on the data collection method, with documentation of datasets collected via crowdsourced and direct collection approaches being more likely to mention them. Scraping dominates scale at the cost of ethical indicators, but is not the only viable collection method. Our approach and empirical insights contribute to increasing dataset transparency along trustworthy and ethical dimensions and pave the way for automating the tedious task of extracting information from dataset documentation in future.
【11】 CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
标题: CosyVoice 3:通过扩展和后训练实现野外语音生成链接:https://arxiv.org/abs/2505.17589
备注:Preprint, work in progress
摘要:在我们之前的工作中,我们介绍了一个可扩展的流语音合成模型,CosyVoice 2,它集成了一个大语言模型(LLM)和一个块感知流匹配(FM)模型,并实现了低延迟的双流语音合成和人类奇偶校验质量。尽管有这些进步,但CosyVoice 2在语言覆盖范围、领域多样性、数据量、文本格式和后期训练技术方面存在局限性。在本文中,我们提出了CosyVoice 3,一个改进的模型设计的zero-shot多语言语音合成在野外,超过其前身的内容一致性,说话人相似性,韵律自然。CosyVoice 3的主要功能包括:1)通过监督多任务训练开发的新型语音标记器,可提高韵律自然度,包括自动语音识别、语音情感识别、语言识别、音频事件检测和说话者分析。2)提出了一种新的可微分后训练奖励模型,不仅适用于CosyVoice 3,也适用于其他基于LLM的语音合成模型。3)数据集大小缩放:训练数据从一万小时扩展到一百万小时,包括9种语言和18种汉语方言,涵盖各种领域和文本格式。4)模型尺寸缩放:模型参数从5亿增加到15亿,由于更大的模型容量,我们的多语言基准测试的性能得到了增强。这些进步对语音合成在野外的进展做出了重大贡献。我们鼓励读者在https://funaudiollm.github.io/cosyvoice3上收听演示。
摘要:In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.
【12】 JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
标题: JALMBench:音频语言模型中的越狱漏洞基准链接:https://arxiv.org/abs/2505.17568
摘要:音频语言模型(ALM)最近取得了重大进展。这些模型将音频模态直接集成到模型中,而不是将语音转换为文本并将文本输入到大型语言模型(LLM)中。虽然对LLM的越狱攻击已经得到了广泛的研究,但具有音频模式的ALM的安全性在很大程度上仍未得到探索。目前,缺乏对抗性音频数据集和专门用于评估和比较攻击和ALM的统一框架。在本文中,我们提出了JALMBench,\textit{第一}全面的基准来评估安全的ALM对越狱攻击。JALMBench包含一个数据集,包含2,200个文本样本和51,381个音频样本,超过268小时。支持12种主流ALM,4种文本传输和4种音频源攻击方式,5种防御方式。使用JALMBench,我们提供了攻击效率,主题敏感性,语音多样性和攻击表示的深入分析。此外,我们探讨缓解策略的攻击在提示级别和响应级别。
摘要:Audio Language Models (ALMs) have made significant progress recently. These models integrate the audio modality directly into the model, rather than converting speech into text and inputting text to Large Language Models (LLMs). While jailbreak attacks on LLMs have been extensively studied, the security of ALMs with audio modalities remains largely unexplored. Currently, there is a lack of an adversarial audio dataset and a unified framework specifically designed to evaluate and compare attacks and ALMs. In this paper, we present JALMBench, the \textit{first} comprehensive benchmark to assess the safety of ALMs against jailbreak attacks. JALMBench includes a dataset containing 2,200 text samples and 51,381 audio samples with over 268 hours. It supports 12 mainstream ALMs, 4 text-transferred and 4 audio-originated attack methods, and 5 defense methods. Using JALMBench, we provide an in-depth analysis of attack efficiency, topic sensitivity, voice diversity, and attack representations. Additionally, we explore mitigation strategies for the attacks at both the prompt level and the response level.
【13】 MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation
标题: MEGADance:适合具有流派意识的3D舞蹈一代的专家混合架构链接:https://arxiv.org/abs/2505.17543
备注:arXiv admin note: text overlap with arXiv:2505.14222
摘要:近年来,音乐驱动的3D舞蹈生成引起了越来越多的关注,在编舞,虚拟现实和创意内容创作中有着广阔的应用前景。先前的研究已经从音频信号中生成了有希望的逼真的舞蹈动作。然而,传统的方法没有充分利用体裁制约,往往把它作为辅助修饰语,而不是核心的语义驱动。这种疏忽损害了音乐动作的同步性,破坏了舞蹈类型的连续性,特别是在复杂的节奏过渡期间,从而导致视觉效果不令人满意。为了应对这一挑战,我们提出了MEGADance,一个新的架构,音乐驱动的3D舞蹈生成。通过将舞蹈的一致性解耦为舞蹈的一般性和体裁的特殊性,MEGADance展示了显著的舞蹈质量和强大的体裁可控性。它包括两个阶段:(1)高保真舞蹈量化阶段(HFDQ),其通过有限标量量化(FSQ)将舞蹈动作编码成潜在表示,并利用运动学-动态约束来重构它们,以及(2)体裁感知舞蹈生成阶段(GADG),其通过协同利用具有Mamba-Transformer混合骨干的专家混合(MoE)机制来将音乐映射成潜在表示。在FineDance和AIST++数据集上进行的大量实验证明了MEGADance在定性和定量方面的最新性能。代码将在接受后发布。
摘要:Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance movement from audio signals. However, traditional methods underutilize genre conditioning, often treating it as auxiliary modifiers rather than core semantic drivers. This oversight compromises music-motion synchronization and disrupts dance genre continuity, particularly during complex rhythmic transitions, thereby leading to visually unsatisfactory effects. To address the challenge, we propose MEGADance, a novel architecture for music-driven 3D dance generation. By decoupling choreographic consistency into dance generality and genre specificity, MEGADance demonstrates significant dance quality and strong genre controllability. It consists of two stages: (1) High-Fidelity Dance Quantization Stage (HFDQ), which encodes dance motions into a latent representation by Finite Scalar Quantization (FSQ) and reconstructs them with kinematic-dynamic constraints, and (2) Genre-Aware Dance Generation Stage (GADG), which maps music into the latent representation by synergistic utilization of Mixture-of-Experts (MoE) mechanism with Mamba-Transformer hybrid backbone. Extensive experiments on the FineDance and AIST++ dataset demonstrate the state-of-the-art performance of MEGADance both qualitatively and quantitatively. Code will be released upon acceptance.
【14】 Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
标题: 瑞典耳语者;利用大量语音数据库进行瑞典语音识别链接:https://arxiv.org/abs/2505.17538
备注:Submitted to Interspeech 2025
摘要:这项工作为瑞典语提供了一套微调的Whisper模型,该模型在这种中等资源语言的前所未有的大小和可变性的数据集上进行了训练。由于较小规模的语言在多语言训练数据集中的代表性往往不足,因此可以通过微调现有的多语言模型来实现性能的大幅提高,正如这项工作所示。这项工作报告了与OpenAI在瑞典评估的Whisper相比,模型大小的整体改进。最值得注意的是,在FLEURS、Common Voice和NST的评估中,我们报告称,与OpenAI的whisper-large-v3相比,我们的最佳性能模型的WER平均降低了47%。
摘要:This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual training datasets, substantial improvements in performance can be achieved by fine-tuning existing multilingual models, as shown in this work. This work reports an overall improvement across model sizes compared to OpenAI's Whisper evaluated on Swedish. Most notably, we report an average 47% reduction in WER comparing our best performing model to OpenAI's whisper-large-v3, in evaluations across FLEURS, Common Voice, and NST.
【15】 What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
标题: 读到的不是听到的:Deepfake语音检测中的语言敏感性链接:https://arxiv.org/abs/2505.17513
备注:15 pages, 2 fogures
摘要:文本转语音技术的最新进展使真实的语音生成成为可能,从而助长了基于音频的deepfake攻击,如欺诈和模仿。虽然音频反欺骗系统对于检测此类威胁至关重要,但先前的工作主要集中在声学水平的扰动上,而语言变化的影响在很大程度上未被探索。在本文中,我们通过引入转录水平的对抗性攻击来研究开源和商业反欺骗检测器的语言敏感性。我们的广泛评估表明,即使是轻微的语言干扰也会显着降低检测准确性:在几个开源检测器-语音对上,攻击成功率超过60%,特别是一个商业检测准确率从合成音频的100%下降到只有32%。通过全面的特征属性分析,我们发现语言复杂性和模型级音频嵌入相似性都对检测器脆弱性有很大贡献。我们通过复制Brad Pitt音频deepfake骗局的案例研究进一步展示了现实世界的风险,使用转录对抗攻击完全绕过商业检测器。这些结果强调了需要超越纯粹的声学防御,并考虑到语言的变化,在强大的反欺骗系统的设计。所有源代码都将公开。
摘要:Recent advances in text-to-speech technologies have enabled realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation. While audio anti-spoofing systems are critical for detecting such threats, prior work has predominantly focused on acoustic-level perturbations, leaving the impact of linguistic variation largely unexplored. In this paper, we investigate the linguistic sensitivity of both open-source and commercial anti-spoofing detectors by introducing transcript-level adversarial attacks. Our extensive evaluation reveals that even minor linguistic perturbations can significantly degrade detection accuracy: attack success rates surpass 60% on several open-source detector-voice pairs, and notably one commercial detection accuracy drops from 100% on synthetic audio to just 32%. Through a comprehensive feature attribution analysis, we identify that both linguistic complexity and model-level audio embedding similarity contribute strongly to detector vulnerability. We further demonstrate the real-world risk via a case study replicating the Brad Pitt audio deepfake scam, using transcript adversarial attacks to completely bypass commercial detectors. These results highlight the need to move beyond purely acoustic defenses and account for linguistic variation in the design of robust anti-spoofing systems. All source code will be publicly available.
【16】 Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
标题: 分析口语模型端到端训练中灾难性遗忘的缓解策略链接:https://arxiv.org/abs/2505.17496
备注:Accepted to Interspeech 2025
摘要:口语语言模型(SLM)的端到端训练通常涉及通过对ASR、TTS和口语问答(SQA)等各种任务进行多阶段训练,使预先训练的基于文本的大型语言模型(LLM)适应语音模态。尽管这种多阶段的持续学习使LLM具备了语音理解和生成能力,但各个阶段的任务和数据分布的巨大差异可能导致灾难性的遗忘,从而丢失先前获得的知识。本文研究了灾难性遗忘,并评估了三种缓解策略-模型合并,LoRA缩放因子折扣和经验重放,以平衡知识保留与新的学习。实验结果表明,经验回放法是最有效的,与其他方法相结合可以获得更大的收益。这些研究结果为开发更强大和有效的可持续土地管理培训管道提供了见解。
摘要:End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS and spoken question answering (SQA). Although this multi-stage continual learning equips LLMs with both speech understanding and generation capabilities, the substantial differences in task and data distributions across stages can lead to catastrophic forgetting, where previously acquired knowledge is lost. This paper investigates catastrophic forgetting and evaluates three mitigation strategies-model merging, discounting the LoRA scaling factor, and experience replay to balance knowledge retention with new learning. Results show that experience replay is the most effective, with further gains achieved by combining it with other methods. These findings provide insights for developing more robust and efficient SLM training pipelines.
【17】 Reverse-Speech-Finder: A Neural Network Backtracking Architecture for Generating Alzheimer's Disease Speech Samples and Improving Diagnosis Performance
标题: 反向语音识别:一种用于阿尔茨海默病语音样本生成和提高诊断性能的神经网络回溯结构链接:https://arxiv.org/abs/2505.17477
摘要:这项研究介绍了反向语音识别(RSF),这是一种开创性的神经网络回溯架构,旨在通过语音分析来增强阿尔茨海默病(AD)的诊断。RSF利用预先训练的大型语言模型的能力,识别并利用最可能的AD特定语音标记,解决了真实AD语音样本的稀缺性和现有模型中有限的可解释性的挑战。RSF的独特方法包括三个核心创新:首先,它利用了这样的观察,即最有可能预测AD的语音标记(定义为最可能的语音标记(MPM))必须具有激活那些具有最高预测AD概率的神经元(在神经网络中)的最高概率,定义为最可能的神经元(MPN)。其次,它在输入层利用语音令牌表示,允许从MPN回溯以识别AD的最可能的语音令牌(MPT)。最后,它开发了一种创新的回溯方法,从MPN向后跟踪到输入层,识别MPT和相应的MPM,并巧妙地发现新的语音标记AD检测。实验结果表明,RSF的优越性,比传统的方法,如SHAP和集成的修正,实现了3.5%的准确性和3.2%的F1分数提高。通过生成封装新标记的语音数据,RSF不仅减轻了真实数据稀缺的限制,而且还显著增强了AD诊断模型的鲁棒性和准确性。这些发现强调了RSF作为基于语音的AD检测的变革性工具的潜力,为AD相关的语言缺陷提供了新的见解,并为更有效的非侵入性早期干预策略铺平了道路。
摘要:This study introduces Reverse-Speech-Finder (RSF), a groundbreaking neural network backtracking architecture designed to enhance Alzheimer's Disease (AD) diagnosis through speech analysis. Leveraging the power of pre-trained large language models, RSF identifies and utilizes the most probable AD-specific speech markers, addressing both the scarcity of real AD speech samples and the challenge of limited interpretability in existing models. RSF's unique approach consists of three core innovations: Firstly, it exploits the observation that speech markers most probable of predicting AD, defined as the most probable speech-markers (MPMs), must have the highest probability of activating those neurons (in the neural network) with the highest probability of predicting AD, defined as the most probable neurons (MPNs). Secondly, it utilizes a speech token representation at the input layer, allowing backtracking from MPNs to identify the most probable speech-tokens (MPTs) of AD. Lastly, it develops an innovative backtracking method to track backwards from the MPNs to the input layer, identifying the MPTs and the corresponding MPMs, and ingeniously uncovering novel speech markers for AD detection. Experimental results demonstrate RSF's superiority over traditional methods such as SHAP and Integrated Gradients, achieving a 3.5% improvement in accuracy and a 3.2% boost in F1-score. By generating speech data that encapsulates novel markers, RSF not only mitigates the limitations of real data scarcity but also significantly enhances the robustness and accuracy of AD diagnostic models. These findings underscore RSF's potential as a transformative tool in speech-based AD detection, offering new insights into AD-related linguistic deficits and paving the way for more effective non-invasive early intervention strategies.
【18】 Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
标题: 探索分段和词汇大小对语音语言模型语音标记化的影响链接:https://arxiv.org/abs/2505.17446
备注:Accepted to Interspeech2025
摘要:语音标记化的目的是将语音信号转换为一系列离散表示,作为语音语言模型(SLM)的基础。虽然语音标记化有很多选择,但它们对SLM性能的影响仍然不清楚。本文研究了语音标记化的两个关键问题:分割宽度和离散单元的聚类大小。首先,我们将语音信号分割成固定/可变宽度和合并表示。然后,我们在多个集群大小中训练K均值模型。通过对零镜头(zero-shot)口语理解基准的评估,我们发现适度粗分割和较大的集群大小具有积极的作用。值得注意的是,在性能最好的模型中,最有效的模型实现了训练数据减少50%,训练运行时间减少70%。我们的分析强调了结合多个标记以增强细粒度口语理解的重要性。
摘要:The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, their effect on the performance of SLMs remains unclear. This paper investigates two key aspects of speech tokenization: the segmentation width and the cluster size of discrete units. First, we segment speech signals into fixed/variable widths and pooled representations. We then train K-means models in multiple cluster sizes. Through the evaluation on zero-shot spoken language understanding benchmarks, we find the positive effect of moderately coarse segmentation and bigger cluster size. Notably, among the best-performing models, the most efficient one achieves a 50% reduction in training data and a 70% decrease in training runtime. Our analysis highlights the importance of combining multiple tokens to enhance fine-grained spoken language understanding.
【19】 UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
标题: UniTTC:一个无需声学和语义信息脱钩的端到端TTC系统链接:https://arxiv.org/abs/2505.17426
摘要:诸如残差矢量量化(RVQ)和组矢量量化(GVQ)的多码本中性音频编解码器的出现显著地推进了基于大语言模型(LLM)的文本到语音(TTS)系统。这些编解码器在分离语义和声学信息同时有效利用语义先验方面至关重要。然而,由于语义和声学信息不能完全对齐,当应用于基于LLM的TTS时,这些方法的显著缺点是大型语言模型对全面音频信息的访问可能有限。为了解决这个问题,我们提出了DistilCodec和UniTTS,它们共同提供了以下优点:1)该方法可以将多码本音频编解码器提取为具有32,768个代码的单码本音频编解码器,同时实现接近100%的利用率。2)由于DistilCodec不采用语义对齐方案,因此大量高质量的未标记音频(例如具有声音效果的有声读物,歌曲等)可以在培训过程中纳入,进一步扩大数据多样性并扩大其适用性。3)利用DistilCodec的全面音频信息建模,我们将三个关键任务集成到UniTTS的预训练框架中:音频模态自回归,文本模态自回归和语音-文本跨模态自回归。这允许UniTTS接受交错的文本和语音/音频提示,同时基本上保留LLM的文本功能。4)UniTTS采用三个阶段的训练过程:预训练,监督微调(SFT)和对齐。源代码和模型检查点可在https://github.com/IDEA-Emdoor-Lab/UniTTS和https://github.com/IDEA-Emdoor-Lab/DistilCodec上公开获取。
摘要:The emergence of multi-codebook neutral audio codecs such as Residual Vector Quantization (RVQ) and Group Vector Quantization (GVQ) has significantly advanced Large-Language-Model (LLM) based Text-to-Speech (TTS) systems. These codecs are crucial in separating semantic and acoustic information while efficiently harnessing semantic priors. However, since semantic and acoustic information cannot be fully aligned, a significant drawback of these methods when applied to LLM-based TTS is that large language models may have limited access to comprehensive audio information. To address this limitation, we propose DistilCodec and UniTTS, which collectively offer the following advantages: 1) This method can distill a multi-codebook audio codec into a single-codebook audio codec with 32,768 codes while achieving a near 100\% utilization. 2) As DistilCodec does not employ a semantic alignment scheme, a large amount of high-quality unlabeled audio (such as audiobooks with sound effects, songs, etc.) can be incorporated during training, further expanding data diversity and broadening its applicability. 3) Leveraging the comprehensive audio information modeling of DistilCodec, we integrated three key tasks into UniTTS's pre-training framework: audio modality autoregression, text modality autoregression, and speech-text cross-modal autoregression. This allows UniTTS to accept interleaved text and speech/audio prompts while substantially preserving LLM's text capabilities. 4) UniTTS employs a three-stage training process: Pre-Training, Supervised Fine-Tuning (SFT), and Alignment. Source code and model checkpoints are publicly available at https://github.com/IDEA-Emdoor-Lab/UniTTS and https://github.com/IDEA-Emdoor-Lab/DistilCodec.
【20】 LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
标题: 基于LLM的具有合成数据和语音上下文的稀有词生成式错误纠正链接:https://arxiv.org/abs/2505.17410
备注:Accepted by INTERSPEECH 2025
摘要:基于大语言模型的生成式纠错(GER)是提高自动语音识别(ASR)性能的有效后处理方法。然而,由于训练数据有限,它经常遇到罕见或特定领域的单词。此外,现有的基于LLM的GER方法主要依赖于文本信息,忽略了语音线索,这导致过度校正。为了解决这些问题,我们提出了一种新的基于LLM的GER方法,目标是罕见的单词,并结合语音信息。首先,我们生成合成数据,以包含用于微调GER模型的稀有词。其次,我们整合ASR的N-最好的假设随着语音上下文,以减轻过度校正。实验结果表明,该方法不仅提高了稀有词的正确率,而且降低了英语和日语数据集的WER和CER。
摘要:Generative error correction (GER) with large language models (LLMs) has emerged as an effective post-processing approach to improve automatic speech recognition (ASR) performance. However, it often struggles with rare or domain-specific words due to limited training data. Furthermore, existing LLM-based GER approaches primarily rely on textual information, neglecting phonetic cues, which leads to over-correction. To address these issues, we propose a novel LLM-based GER approach that targets rare words and incorporates phonetic information. First, we generate synthetic data to contain rare words for fine-tuning the GER model. Second, we integrate ASR's N-best hypotheses along with phonetic context to mitigate over-correction. Experimental results show that our method not only improves the correction of rare words but also reduces the WER and CER across both English and Japanese datasets.
【21】 VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
标题: VoxRAG:在口语问题回答中迈向免转录RAG系统的一步链接:https://arxiv.org/abs/2505.17326
备注:Accepted to ACL 2025 Workshop MAGMaR
摘要:我们介绍VoxRAG,一个模块化的语音到语音检索增强生成系统,绕过转录检索语义相关的音频片段直接从口头查询。VoxRAG采用沉默感知分割,说话人日记化,CLAP音频嵌入和使用L2归一化余弦相似性的FAISS检索。我们构建了一个50个查询的测试集记录为口语输入的母语为英语的人。检索质量进行了评价,使用LLM-as-a-judge注释。对于非常相关的片段,余弦相似性达到0.34的Recall@10。对于一些相关的部分,Recall@10上升到0.60,nDCG@10上升到0.27,突出了强烈的主题对齐。答案质量在相关性、准确性、完整性和精确性方面以0- 2的尺度进行判断,平均得分分别为0.84、0.58、0.56和0.46。虽然精度和检索质量仍然是关键的限制,VoxRAG表明,转录免费语音到语音检索在RAG系统是可行的。
摘要:We introduce VoxRAG, a modular speech-to-speech retrieval-augmented generation system that bypasses transcription to retrieve semantically relevant audio segments directly from spoken queries. VoxRAG employs silence-aware segmentation, speaker diarization, CLAP audio embeddings, and FAISS retrieval using L2-normalized cosine similarity. We construct a 50-query test set recorded as spoken input by a native English speaker. Retrieval quality was evaluated using LLM-as-a-judge annotations. For very relevant segments, cosine similarity achieved a Recall@10 of 0.34. For somewhat relevant segments, Recall@10 rose to 0.60 and nDCG@10 to 0.27, highlighting strong topical alignment. Answer quality was judged on a 0--2 scale across relevance, accuracy, completeness, and precision, with mean scores of 0.84, 0.58, 0.56, and 0.46 respectively. While precision and retrieval quality remain key limitations, VoxRAG shows that transcription-free speech-to-speech retrieval is feasible in RAG systems.
【22】 Benchmarking Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
标题: 使用VITS和风格对表达性日语字符文本到语音进行基准测试-BERT-VITS 2链接:https://arxiv.org/abs/2505.17320
摘要:由于音高重音敏感性和风格的变化性,合成富有表现力的日语字符语音构成了独特的挑战。本文对两个开源的文本到语音转换模型--VITS和Style-BERT-VITS 2 JP Extra(SBV 2 JE)--在域内的字符驱动的日语语音上进行了基准测试。使用三个特定于字符的数据集,我们评估模型的自然度(平均意见和比较平均意见得分),可懂度(单词错误率)和说话人一致性。SBV 2 JE在自然度方面与人类地面实况相匹配(MOS 4.37 vs. 4.38),实现了较低的WER,并在CMOS中显示出轻微的偏好。SBV 2 JE通过音高重音控制和基于WavLM的语音增强,证明了它对语言学习和角色对话生成等应用程序的有效性,尽管计算要求更高。
摘要:Synthesizing expressive Japanese character speech poses unique challenges due to pitch-accent sensitivity and stylistic variability. This paper benchmarks two open-source text-to-speech models--VITS and Style-BERT-VITS2 JP Extra (SBV2JE)--on in-domain, character-driven Japanese speech. Using three character-specific datasets, we evaluate models across naturalness (mean opinion and comparative mean opinion score), intelligibility (word error rate), and speaker consistency. SBV2JE matches human ground truth in naturalness (MOS 4.37 vs. 4.38), achieves lower WER, and shows slight preference in CMOS. Enhanced by pitch-accent controls and a WavLM-based discriminator, SBV2JE proves effective for applications like language learning and character dialogue generation, despite higher computational demands.
【23】 Understanding the Algorithm Behind Audio Key Detection
标题: 了解音频密钥检测背后的算法链接:https://arxiv.org/abs/2505.17259
备注:Preprint. Describes an algorithmic approach to musical key detection implemented in Python. Includes conceptual explanation of audio feature extraction and key profile matching
摘要:音乐基调的确定是音乐理论和感知的一个基本方面,为旋律和和弦进行提供和声背景。自动化这个过程,称为自动键检测,是音乐信息检索(MIR)领域的一项重要任务。本文概述了一种算法方法,通过分析其音调内容,通过数字信号处理技术和理论的关键配置文件的比较,估计音乐的音频记录的关键。
摘要:The determination of musical key is a fundamental aspect of music theory and perception, providing a harmonic context for melodies and chord progressions. Automating this process, known as automatic key detection, is a significant task in the field of Music Information Retrieval (MIR). This article outlines an algorithmic methodology for estimating the musical key of an audio recording by analyzing its tonal content through digital signal processing techniques and comparison with theoretical key profiles.
【24】 Semantic-Aware Interpretable Multimodal Music Auto-Tagging
标题: 语义感知的可解释多模式音乐自动标记链接:https://arxiv.org/abs/2505.17233
摘要:音乐自动标记对于在广泛的数字图书馆中组织和发现音乐至关重要。虽然基础模型在这一领域取得了卓越的性能,但它们的输出往往缺乏可解释性,限制了研究人员和最终用户的信任和可用性。在这项工作中,我们提出了一个可解释的音乐自动标记框架,该框架利用了来自信号处理、深度学习、本体工程和自然语言处理的具有音乐意义的多模态特征。为了提高可解释性,我们聚类特征语义和采用期望最大化算法,分配不同的权重,每个组的基础上,其对标记过程的贡献。我们的方法实现了具有竞争力的标记性能,同时提供了对决策过程的更深入理解,为更透明和以用户为中心的音乐标记系统铺平了道路。
摘要:Music auto-tagging is essential for organizing and discovering music in extensive digital libraries. While foundation models achieve exceptional performance in this domain, their outputs often lack interpretability, limiting trust and usability for researchers and end-users alike. In this work, we present an interpretable framework for music auto-tagging that leverages groups of musically meaningful multimodal features, derived from signal processing, deep learning, ontology engineering, and natural language processing. To enhance interpretability, we cluster features semantically and employ an expectation maximization algorithm, assigning distinct weights to each group based on its contribution to the tagging process. Our method achieves competitive tagging performance while offering a deeper understanding of the decision-making process, paving the way for more transparent and user-centric music tagging systems.
【25】 Large Language Models Implicitly Learn to See and Hear Just By Reading
标题: 大型语言模型通过阅读就能学会看和听链接:https://arxiv.org/abs/2505.17091
备注:6 pages, 3 figures, 4 tables. Under Review WASPAA 2025
摘要:本文提出了一个有趣的发现:通过在文本标记上训练自回归LLM模型,文本模型内在地发展了内部理解图像和音频的能力,从而发展了仅通过阅读就能看到和听到的能力。流行的音频和视觉LLM模型微调文本LLM模型,以提供基于图像和音频嵌入的文本输出。另一方面,我们的架构接受图像、音频波形或令牌的补丁作为输入。它为我们提供了分类管道中典型的嵌入或类别标签。我们展示了文本权重在帮助数据集FSD-50K和GTZAN的音频分类中的一般性。此外,我们还展示了CIFAR-10和Fashion-MNIST以及图像补丁上的图像分类。这推动了文本LLM学习强大的内部电路的概念,这些电路可以通过激活各种应用程序的必要连接来利用,而不是每次都从头开始训练模型。
摘要:This paper presents a fascinating find: By training an auto-regressive LLM model on text tokens, the text model inherently develops internally an ability to understand images and audio, thereby developing the ability to see and hear just by reading. Popular audio and visual LLM models fine-tune text LLM models to give text output conditioned on images and audio embeddings. On the other hand, our architecture takes in patches of images, audio waveforms or tokens as input. It gives us the embeddings or category labels typical of a classification pipeline. We show the generality of text weights in aiding audio classification for datasets FSD-50K and GTZAN. Further, we show this working for image classification on CIFAR-10 and Fashion-MNIST, as well on image patches. This pushes the notion of text-LLMs learning powerful internal circuits that can be utilized by activating necessary connections for various applications rather than training models from scratch every single time.
【26】 Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
标题: 帧速率对语音标记器的影响--以汉语和英语为例链接:https://arxiv.org/abs/2505.17076
备注:5 pages, 5 figures
摘要:None
摘要:The speech tokenizer plays a crucial role in recent speech tasks, generally serving as a bridge between speech signals and language models. While low-frame-rate codecs are widely employed as speech tokenizers, the impact of frame rates on speech tokens remains underexplored. In this study, we investigate how varying frame rates affect speech tokenization by examining Mandarin and English, two typologically distinct languages. We encode speech at different frame rates and evaluate the resulting semantic tokens in the speech recognition task. Our findings reveal that frame rate variations influence speech tokenization differently for each language, highlighting the interplay between frame rates, phonetic density, and language-specific acoustic features. The results provide insights into optimizing frame rate selection for speech tokenizers, with implications for automatic speech recognition, text-to-speech, and other speech-related applications.
【27】 Improving endpoint detection in end-to-end streaming ASR for conversational speech
标题: 改进对话语音的端到端流ASB中的端点检测链接:https://arxiv.org/abs/2505.17070
备注:Submitted to Interspeech 2024
摘要:ASR端点(EP)在人机对话中支持人类或人工代理的产品中提供良好的用户体验方面发挥着重要作用。基于传感器的ASR(T-ASR)是流媒体首选的端到端(E2 E)ASR建模技术。T-ASR的主要限制是ASR输出的延迟发射,这可能导致EP中的错误或延迟。不准确的EP将在说话时切断用户,返回不完整的转录,而EP中的延迟将增加感知延迟,降低用户体验。我们提出了通过解决延迟发射以及EP错误来改善EP的方法。为了解决延迟发射问题,我们在每个单词的末尾引入了一个单词结束标记,以及延迟惩罚。通过使用辅助网络获得可靠的帧级语音活动检测来解决EP延迟。我们将所提出的方法应用于Switchboard会话语音语料库,并对延迟惩罚方法进行评估。
摘要:ASR endpointing (EP) plays a major role in delivering a good user experience in products supporting human or artificial agents in human-human/machine conversations. Transducer-based ASR (T-ASR) is an end-to-end (E2E) ASR modelling technique preferred for streaming. A major limitation of T-ASR is delayed emission of ASR outputs, which could lead to errors or delays in EP. Inaccurate EP will cut the user off while speaking, returning incomplete transcript while delays in EP will increase the perceived latency, degrading the user experience. We propose methods to improve EP by addressing delayed emission along with EP mistakes. To address the delayed emission problem, we introduce an end-of-word token at the end of each word, along with a delay penalty. The EP delay is addressed by obtaining a reliable frame-level speech activity detection using an auxiliary network. We apply the proposed methods on Switchboard conversational speech corpus and evaluate it against a delay penalty method.
【28】 ReMi: A Random Recurrent Neural Network Approach to Music Production
标题: ReMi:音乐制作的随机回归神经网络方法链接:https://arxiv.org/abs/2505.17023
备注:Accepted for an Innovation Showcase Demo at International Computer Music Conference
摘要:生成性人工智能引发了与能源消耗、版权侵犯和创造力萎缩相关的担忧。我们证明了随机初始化的递归神经网络可以产生丰富且可配置的琶音和低频振荡。与旨在取代音乐家的端到端音乐生成相反,我们的方法扩展了他们的创造力,同时不需要数据和更少的计算能力。更多信息请访问:https://allendia.com/
摘要:Generative artificial intelligence raises concerns related to energy consumption, copyright infringement and creative atrophy. We show that randomly initialized recurrent neural networks can produce arpeggios and low-frequency oscillations that are rich and configurable. In contrast to end-to-end music generation that aims to replace musicians, our approach expands their creativity while requiring no data and much less computational power. More information can be found at: https://allendia.com/
机器翻译由腾讯交互翻译提供,仅供参考
