今日论文合集:cs.SD语音19篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification
标题:Zero-Shot音频语言分类中利用预测熵自动提示加权
链接:https://arxiv.org/abs/2601.05011

作者:Karim El Khoury,Maxime Zanella,Tiffanie Godelaine,Christophe De Vleeschouwer,Benoit Macq
摘要:最近,音频语言模型通过利用自然语言监督来对没有标记的训练数据的音频事件进行分类,展示了强大的zero-shot能力。然而,它们的表现对文本提示的措辞高度敏感,微小的变化会导致准确性的大幅波动。以前的工作已经通过快速学习或快速集成来缓解这个问题。然而,这些策略要么需要注释的数据,要么无法考虑某些提示可能会对性能产生负面影响的事实。在这项工作中,我们提出了一种熵引导的提示加权方法,旨在找到提示贡献的鲁棒组合,以最大化预测置信度。为此,我们制定了一个量身定制的目标函数,最大限度地减少预测熵,以产生新的提示权重,利用低熵作为高置信度的代理。我们的方法可以应用于单个样本或一批音频样本,不需要额外的标签,产生的计算开销可以忽略不计。五个音频分类数据集,涵盖环境,城市和人声的实验,证明了一致的收益相比,经典的提示集成方法在zero-shot设置,精度提高5倍以上,在整个基准。
摘要:Audio-language models have recently demonstrated strong zero-shot capabilities by leveraging natural-language supervision to classify audio events without labeled training data. Yet, their performance is highly sensitive to the wording of text prompts, with small variations leading to large fluctuations in accuracy. Prior work has mitigated this issue through prompt learning or prompt ensembling. However, these strategies either require annotated data or fail to account for the fact that some prompts may negatively impact performance. In this work, we present an entropy-guided prompt weighting approach that aims to find a robust combination of prompt contributions to maximize prediction confidence. To this end, we formulate a tailored objective function that minimizes prediction entropy to yield new prompt weights, utilizing low-entropy as a proxy for high confidence. Our approach can be applied to individual samples or a batch of audio samples, requiring no additional labels and incurring negligible computational overhead. Experiments on five audio classification datasets covering environmental, urban, and vocal sounds, demonstrate consistent gains compared to classical prompt ensembling methods in a zero-shot setting, with accuracy improvements 5-times larger across the whole benchmark.


【2】A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction
标题:具有注入解释归因思维的类人互动统一口语模型
链接:https://arxiv.org/abs/2601.04960

作者:Qing Wang,Zehan Li,Yaodong Song,Hongjie Chen,Jian Kang,Jie Lian,Jie Li,Yongxiang Li,Xuelong Li
摘要:本文提出了一个统一的口语模型的情绪智力,增强了一种新的数据构建策略,称为注入式推理归因思维(IEAT)。IEAT将用户情绪状态及其潜在原因纳入模型的内部推理过程,使情绪感知推理能够内化,而不是被视为明确的监督。该模型采用两阶段渐进策略进行训练。第一阶段通过自升华进行语音-文本对齐和情感属性建模,而第二阶段进行端到端的跨模态联合优化,以确保文本和口语情感表达之间的一致性。在类人口语对话系统挑战赛(HumDial)情感智能基准测试上的实验表明,所提出的方法在基于LLM和人类评估的情感轨迹建模、情感推理和移情响应生成方面都取得了最佳性能。
摘要:This paper presents a unified spoken language model for emotional intelligence, enhanced by a novel data construction strategy termed Injected Emotional-Attribution Thinking (IEAT). IEAT incorporates user emotional states and their underlying causes into the model's internal reasoning process, enabling emotion-aware reasoning to be internalized rather than treated as explicit supervision. The model is trained with a two-stage progressive strategy. The first stage performs speech-text alignment and emotional attribute modeling via self-distillation, while the second stage conducts end-to-end cross-modal joint optimization to ensure consistency between textual and spoken emotional expressions. Experiments on the Human-like Spoken Dialogue Systems Challenge (HumDial) Emotional Intelligence benchmark demonstrate that the proposed approach achieves top-ranked performance across emotional trajectory modeling, emotional reasoning, and empathetic response generation under both LLM-based and human evaluations.


【3】ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
标题:ChronosAudio:用于评估音频大语言模型的全面长音频基准
链接:https://arxiv.org/abs/2601.04876

作者:Kaiwen Luo,Liang Lin,Yibo Zhang,Moayad Aloqaily,Dexian Wang,Zhenhong Zhou,Junwei Zhang,Kun Wang,Li Sun,Qingsong Wen
摘要:虽然音频大语言模型(ALLM)已经取得了实质性的进展,但其长音频理解能力仍未得到开发。对于一般的音频任务,已经提出了过多的基准,它们主要集中在短形式的剪辑上,在评估长时间的ALLM上没有达成共识。本文提出了ChronosAudio,这是第一个为ALLM中的长音频理解量身定制的多任务基准。它包括六个主要任务类别,包括36,000个测试实例,总计超过200小时的音频,分为短,中,长三类,以全面评估长度泛化。使用ChronosAudio对16个最先进的模型进行了广泛的实验,得出了三个关键的发现:1.偶然的长上下文崩溃:ALLM表现出严重的无法维持性能,从短到长的上下文转换触发了特定任务中超过90%的惊人性能下降。2.结构性注意力稀释:性能下降源于一个根本性的失败,在保持时间的局部性,注意力机制遭受显着的扩散在以后的序列。3.缓解的恢复性上限:目前的策略仅提供50%的恢复。这些发现揭示了长音频的重大挑战,强调了迫切需要实现强大的文档级音频推理的方法。
摘要:Although Audio Large Language Models (ALLMs) have witnessed substantial advancements, their long audio understanding capabilities remain unexplored. A plethora of benchmarks have been proposed for general audio tasks, they predominantly focus on short-form clips, leaving without a consensus on evaluating ALLMs over extended durations. This paper proposes ChronosAudio, the first multi-task benchmark tailored for long-audio understanding in ALLMs. It encompasses six major task categories and comprises 36,000 test instances totaling over 200 hours audio, stratified into short, middle, and long-form categories to comprehensively evaluate length generalization. Extensive experiments on 16 state-of-the-art models using ChronosAudio yield three critical findings: 1.Precipitous Long-Context Collapse: ALLMs exhibit a severe inability to sustain performance, with the transition from short to long contexts triggering a staggering performance degradation of over 90% in specific tasks. 2.Structural Attention Dilution: Performance degradation stems from a fundamental failure in maintaining temporal locality; attention mechanisms suffer from significant diffusion in later sequences. 3.Restorative Ceiling of Mitigation: Current strategies only offer 50% recovery. These findings reveal significant challenges in long-audio, underscoring the urgent need for approaches to achieve robust, document-level audio reasoning.


【4】Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling
标题:具有多层数据建模的语音对话半监督并行检测
链接:https://arxiv.org/abs/2601.04744

作者:Xingyuan Li,Mengyue Wu
摘要:从语音声学中检测医疗状况从根本上说是一个弱监督学习问题:一个单一的,通常是嘈杂的,会话级别的标签必须与一个长而复杂的音频记录中的细微差别联系起来。这一任务进一步受到严重的数据稀缺性和临床注释的主观性的阻碍。虽然半监督学习(SSL)提供了一条利用未标记数据的可行途径,但现有的音频方法往往无法解决患者语音中病理特征表达不一致的核心挑战。我们提出了一种新的,只有音频的SSL框架,明确模型,通过共同学习帧级,段级和会话级表示在未分割的临床对话的层次结构。我们的端到端方法动态地聚合这些多粒度特征,并生成高质量的伪标签,以有效地利用未标记的数据。大量的实验表明,该框架是模型不可知的,跨语言和条件的鲁棒性,以及高数据效率-例如,仅使用11个标记样本就实现了90%的全监督性能。这项工作提供了一个原则性的方法来学习从弱,远端监督在医疗语音分析。
摘要:Detecting medical conditions from speech acoustics is fundamentally a weakly-supervised learning problem: a single, often noisy, session-level label must be linked to nuanced patterns within a long, complex audio recording. This task is further hampered by severe data scarcity and the subjective nature of clinical annotations. While semi-supervised learning (SSL) offers a viable path to leverage unlabeled data, existing audio methods often fail to address the core challenge that pathological traits are not uniformly expressed in a patient's speech. We propose a novel, audio-only SSL framework that explicitly models this hierarchy by jointly learning from frame-level, segment-level, and session-level representations within unsegmented clinical dialogues. Our end-to-end approach dynamically aggregates these multi-granularity features and generates high-quality pseudo-labels to efficiently utilize unlabeled data. Extensive experiments show the framework is model-agnostic, robust across languages and conditions, and highly data-efficient-achieving, for instance, 90\% of fully-supervised performance using only 11 labeled samples. This work provides a principled approach to learning from weak, far-end supervision in medical speech analysis.


【5】LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence
标题:LAMB:基于LLM的音频字幕,通过Cauchy-Schwarz分歧弥合模式差距
链接:https://arxiv.org/abs/2601.04658

作者:Hyeongkeun Lee,Jongmin Choi,KiHyun Nam,Joon Son Chung
备注:5 pages, 2 figures;
摘要:自动音频字幕旨在描述输入音频的语义内容。最近的工作采用大型语言模型(LLM)作为文本解码器来利用其推理能力。然而,将音频特征投影到LLM嵌入空间中而不考虑跨模态对准的现有方法未能充分利用这些能力。为了解决这个问题,我们提出了LAMB,一个基于LLM的音频字幕框架,它弥合了音频嵌入和LLM文本嵌入空间之间的模态差距。LAMB集成了一个跨模态对齐器,可以最大限度地减少Cauchy-Schwarz分歧,同时最大限度地提高互信息,从而在全局和令牌级别上实现音频和文本之间的更紧密对齐。我们进一步设计了一个双流适配器,提取语义丰富的音频嵌入,从而提供更丰富的信息的跨模态对齐。最后,利用对齐的音频嵌入,提出的令牌指南直接计算LLM文本嵌入空间内的分数,以引导生成的字幕的输出logit。实验结果证实,我们的框架增强了LLM解码器的推理能力,在AudioCaps上实现了最先进的性能。
摘要:Automated Audio Captioning aims to describe the semantic content of input audio. Recent works have employed large language models (LLMs) as a text decoder to leverage their reasoning capabilities. However, prior approaches that project audio features into the LLM embedding space without considering cross-modal alignment fail to fully utilize these capabilities. To address this, we propose LAMB, an LLM-based audio captioning framework that bridges the modality gap between audio embeddings and the LLM text embedding space. LAMB incorporates a Cross-Modal Aligner that minimizes Cauchy-Schwarz divergence while maximizing mutual information, yielding tighter alignment between audio and text at both global and token levels. We further design a Two-Stream Adapter that extracts semantically enriched audio embeddings, thereby delivering richer information to the Cross-Modal Aligner. Finally, leveraging the aligned audio embeddings, a proposed Token Guide directly computes scores within the LLM text embedding space to steer the output logits of generated captions. Experimental results confirm that our framework strengthens the reasoning capabilities of the LLM decoder, achieving state-of-the-art performance on AudioCaps.


【6】FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions
标题:MIDI Voice:通过自然语言指令在Zero-ShotTTC中启用灵活的风格控制
链接:https://arxiv.org/abs/2601.04656

作者:Dekun Chen,Xueyao Zhang,Yuancheng Wang,Kenan Dai,Li Ma,Zhizheng Wu
摘要:本研究提出了一个文本到语音(TTS)合成系统,能够灵活的风格控制与zero-shot语音克隆。说话风格由自然语言指令控制,语音音色由语音参考以zero-shot方式提供。Voice是用LLM内核构建的,它以文本作为输入,还采用可选的自然语言指令和可选的语音参考来分别控制风格和音色。Vocabulary Voice配备了一种新颖的渐进式后期训练(PPT)方案,可逐步解锁准确和灵活的可控性。特别是,它首先采用直接偏好优化(DPO),使语音准确地遵循自然语言指令和语音参考同时进行。然后,它使用多目标组相对策略优化(GRPO)来解开风格指令,参考音色和文本内容。最后,它适应指令GRPO更先进的指令以下。实验结果表明,该算法优于同类算法,具有较强的解耦控制能力。人类评估进一步证实了其自然性,可控性和鲁棒性。音频样本可在https://flexi-voice.github.io上获得。
摘要:This study proposes FlexiVoice, a text-to-speech (TTS) synthesis system capable of flexible style control with zero-shot voice cloning. The speaking style is controlled by a natural-language instruction and the voice timbre is provided by a speech reference in zero-shot manner. FlexiVoice is built with an LLM core, which takes text as input, and also takes an optional natural language instruction and an optional speech reference to control style and timbre, respectively. FlexiVoice is equipped with a novel Progressive Post-Training (PPT) scheme that progressively unlocks accurate and flexible controllability. In particular, it first employs Direct Preference Optimization (DPO) to enable FlexiVoice to accurately follow both natural language instruction and speech reference simultaneously. It then uses a multi-objective Group Relative Policy Optimization (GRPO) to disentangle style instruction, reference timbre, and textual content. Finally, it adapts instruction GRPO for more advanced instruction following. Experimental results show that FlexiVoice surpasses competing baselines and demonstrates strong capability in decoupling control factors. Human evaluations further confirm its naturalness, controllability, and robustness. Audio samples are available at https://flexi-voice.github.io.


【7】Density Matrix RNN (DM-RNN): A Quantum Information Theoretic Framework for Modeling Musical Context and Polyphony
标题:密度矩阵RNN(DM-RNN):音乐语境和复调建模的量子信息理论框架
链接:https://arxiv.org/abs/2601.04592

作者:Joonwon Seo,Mariana Montiel
备注:Submitted to the 10th International Conference on Mathematics and Computation in Music (MCM 2026)
摘要:经典的递归神经网络(RNN)将音乐上下文总结为确定性的隐藏状态向量,造成了无法捕捉音乐中固有的模糊性的信息瓶颈。我们提出了密度矩阵RNN(DM-RNN),一种新的理论架构,利用密度矩阵。这使得模型能够保持音乐解释的统计合奏(混合状态),捕获经典概率和量子相干性。我们严格定义的时间动态使用量子通道(CPTP地图)。至关重要的是,我们详细介绍了基于Choi-Jamiolkowski同构的参数化策略,通过构造确保学习的动力学保持物理有效性(CPTP)。我们引入了一个分析框架,使用冯诺依曼熵来量化音乐的不确定性和量子互信息(QMI)来衡量声音之间的纠缠。DM-RNN为复杂、模糊的音乐结构建模提供了一个数学上严格的框架。
摘要:Classical Recurrent Neural Networks (RNNs) summarize musical context into a deterministic hidden state vector, imposing an information bottleneck that fails to capture the inherent ambiguity in music. We propose the Density Matrix RNN (DM-RNN), a novel theoretical architecture utilizing the Density Matrix. This allows the model to maintain a statistical ensemble of musical interpretations (a mixed state), capturing both classical probabilities and quantum coherences. We rigorously define the temporal dynamics using Quantum Channels (CPTP maps). Crucially, we detail a parameterization strategy based on the Choi-Jamiolkowski isomorphism, ensuring the learned dynamics remain physically valid (CPTP) by construction. We introduce an analytical framework using Von Neumann Entropy to quantify musical uncertainty and Quantum Mutual Information (QMI) to measure entanglement between voices. The DM-RNN provides a mathematically rigorous framework for modeling complex, ambiguous musical structures.


【8】When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
标题:当语气和词语不一致时:声学-语义冲突下的稳健语音情感识别
链接:https://arxiv.org/abs/2601.04564

作者:Dawei Huang,Yongjie Lv,Ruijie Xiong,Chunxiang Jin,Xiaojiang Peng
摘要:语音情感识别(SER)系统通常假设声音情感和词汇语义之间的一致性。然而,在现实世界的互动中,声音-语义冲突是常见的,但却被忽视了,其中语气传达的情感与口语的字面意思相矛盾。我们发现,最先进的SER模型,包括基于ASR的自监督学习(SSL)方法和音频语言模型(ALMs),由于语义偏见或纠缠的声学语义表示,在这种冲突下性能下降。为了解决这个问题,我们提出了融合声学-语义(FAS)框架,它明确地解开声学和语义的途径,并通过一个轻量级的,基于查询的注意力模块将它们连接起来。为了进行系统的评估,我们引入了声学-语义情感冲突(CASE),这是第一个在各种场景中由清晰和可解释的声学-语义冲突主导的数据集。大量的实验表明,FAS始终优于现有的方法在域和zero-shot设置。值得注意的是,在CASE基准测试中,传统的SER模型明显失败,而FAS设置了一个新的SOTA,准确率为59.38%。我们的代码和数据集可以在https://github.com/24DavidHuang/FAS上找到。
摘要:Speech Emotion Recognition (SER) systems often assume congruence between vocal emotion and lexical semantics. However, in real-world interactions, acoustic-semantic conflict is common yet overlooked, where the emotion conveyed by tone contradicts the literal meaning of spoken words. We show that state-of-the-art SER models, including ASR-based, self-supervised learning (SSL) approaches and Audio Language Models (ALMs), suffer performance degradation under such conflicts due to semantic bias or entangled acoustic-semantic representations. To address this, we propose the Fusion Acoustic-Semantic (FAS) framework, which explicitly disentangles acoustic and semantic pathways and bridges them through a lightweight, query-based attention module. To enable systematic evaluation, we introduce the Conflict in Acoustic-Semantic Emotion (CASE), the first dataset dominated by clear and interpretable acoustic-semantic conflicts in varied scenarios. Extensive experiments demonstrate that FAS consistently outperforms existing methods in both in-domain and zero-shot settings. Notably, on the CASE benchmark, conventional SER models fail dramatically, while FAS sets a new SOTA with 59.38% accuracy. Our code and datasets is available at https://github.com/24DavidHuang/FAS.


【9】WESR: Scaling and Evaluating Word-level Event-Speech Recognition
标题:WESR:缩放和评估单词级事件语音识别
链接:https://arxiv.org/abs/2601.04508

作者:Chenchen Yang,Kexin Huang,Liwei Fan,Qian Tu,Botian Jiang,Dong Zhang,Linqi Yin,Shimin Li,Zhaoye Fei,Qinyuan Cheng,Xipeng Qiu
备注:14 pages, 6 figures
摘要:言语不仅传达语言信息,还传达丰富的非言语声音事件,如笑和哭。虽然语义转录是很好的研究,非语言事件的精确定位仍然是一个关键的,但探索不足的挑战。目前的方法遭受有限的类别覆盖和模糊的时间粒度的任务定义不足。它们还缺乏标准化的评价框架,阻碍了下游应用程序的开发。为了弥合这一差距,我们首先开发了一个细化的分类21声乐活动,与一个新的分类到离散(独立)与连续(与语音混合)类型。基于改进的分类,我们引入了WESR-Bench,这是一个专家注释的评估集(900多个话语),具有一种新颖的位置感知协议,可以将ASR错误与事件检测分离,从而实现离散和连续事件的精确定位测量。我们还通过构建超过1,700小时的语料库建立了强大的基线,并训练专业模型,超越了开源音频语言模型和商业API,同时保持了ASR质量。我们预计,WESR将作为一个基础资源,为未来的研究建模丰富,真实世界的听觉场景。
摘要:Speech conveys not only linguistic information but also rich non-verbal vocal events such as laughing and crying. While semantic transcription is well-studied, the precise localization of non-verbal events remains a critical yet under-explored challenge. Current methods suffer from insufficient task definitions with limited category coverage and ambiguous temporal granularity. They also lack standardized evaluation frameworks, hindering the development of downstream applications. To bridge this gap, we first develop a refined taxonomy of 21 vocal events, with a new categorization into discrete (standalone) versus continuous (mixed with speech) types. Based on the refined taxonomy, we introduce WESR-Bench, an expert-annotated evaluation set (900+ utterances) with a novel position-aware protocol that disentangles ASR errors from event detection, enabling precise localization measurement for both discrete and continuous events. We also build a strong baseline by constructing a 1,700+ hour corpus, and train specialized models, surpassing both open-source audio-language models and commercial APIs while preserving ASR quality. We anticipate that WESR will serve as a foundational resource for future research in modeling rich, real-world auditory scenes.


【10】Summary of The Inaugural Music Source Restoration Challenge
标题:首届音乐来源恢复挑战赛总结
链接:https://arxiv.org/abs/2601.04343

作者:Yongyi Zang,Jiarui Hai,Wanying Ge,Qiuqiang Kong,Zheqi Dai,Helin Wang,Yuki Mitsufuji,Mark D. Plumbley
摘要:音乐源恢复(MSR)旨在从专业混合和降级的音频中恢复原始的未经处理的乐器,需要逆转制作效果和真实世界的降级。我们提出了首届MSR挑战赛,其特点是使用Multi-Mel-SNR,Zimtohrli和FAD-CLAP对工作室制作的混合物进行客观评估,同时对真实世界的降级录音进行主观评估。五支队伍参加了挑战。获奖系统实现了4.46 dB Multi-Mel-SNR和3.47 MOS-Overall,分别比第二名系统相对提高了91%和18%。每个词干的分析揭示了不同乐器的还原难度有很大的差异,所有团队的低音平均为4.59 dB,而打击乐平均只有0.29 dB。数据集、评估方案和基线可在https://msrchallenge.com/上获得。
摘要:Music Source Restoration (MSR) aims to recover original, unprocessed instrument stems from professionally mixed and degraded audio, requiring the reversal of both production effects and real-world degradations. We present the inaugural MSR Challenge, which features objective evaluation on studio-produced mixtures using Multi-Mel-SNR, Zimtohrli, and FAD-CLAP, alongside subjective evaluation on real-world degraded recordings. Five teams participated in the challenge. The winning system achieved 4.46 dB Multi-Mel-SNR and 3.47 MOS-Overall, corresponding to relative improvements of 91% and 18% over the second-place system, respectively. Per-stem analysis reveals substantial variation in restoration difficulty across instruments, with bass averaging 4.59 dB across all teams, while percussion averages only 0.29 dB. The dataset, evaluation protocols, and baselines are available at https://msrchallenge.com/.


【11】SmoothSync: Dual-Stream Diffusion Transformers for Jitter-Robust Beat-Synchronized Gesture Generation from Quantized Audio
标题:SmoothSync:用于从量化音频生成抖动鲁棒节拍同步手势的双流扩散变换器
链接:https://arxiv.org/abs/2601.04236

作者:Yujiao Jiang,Qingmin Liao,Zongqing Lu
摘要:协同语音手势生成是一个关键的研究领域,旨在合成语音同步的类人手势。现有的方法往往遭受的问题,如节奏不一致,运动抖动,脚滑动和有限的多采样多样性。在本文中,我们提出了SmoothSync,一个新的框架,利用量化的音频令牌在一个新的双流扩散Transformer(DiT)架构合成整体的姿态和提高采样变化。具体而言,我们(1)通过互补的Transformer流融合音频运动特征以实现更好的同步,(2)引入抖动抑制损失以提高时间平滑度,(3)实现概率音频量化以从相同的输入生成不同的手势序列。为了可靠地评估抖动下的节拍同步,我们引入了Smooth-BC,这是一种对运动噪声不太敏感的节拍一致性度量的鲁棒变体。在BEAT 2和SHOW数据集上的综合实验证明了SmoothSync的优越性,在BEAT 2上的表现优于最先进的方法,分别为-30.6% FGD,10.3% Smooth-BC和8.4% Diversity,同时分别减少了-62.9%和-17.1%的抖动和脚滑动。该代码将被释放,以方便未来的研究。
摘要:Co-speech gesture generation is a critical area of research aimed at synthesizing speech-synchronized human-like gestures. Existing methods often suffer from issues such as rhythmic inconsistency, motion jitter, foot sliding and limited multi-sampling diversity. In this paper, we present SmoothSync, a novel framework that leverages quantized audio tokens in a novel dual-stream Diffusion Transformer (DiT) architecture to synthesis holistic gestures and enhance sampling variation. Specifically, we (1) fuse audio-motion features via complementary transformer streams to achieve superior synchronization, (2) introduce a jitter-suppression loss to improve temporal smoothness, (3) implement probabilistic audio quantization to generate distinct gesture sequences from identical inputs. To reliably evaluate beat synchronization under jitter, we introduce Smooth-BC, a robust variant of the beat consistency metric less sensitive to motion noise. Comprehensive experiments on the BEAT2 and SHOW datasets demonstrate SmoothSync's superiority, outperforming state-of-the-art methods by -30.6% FGD, 10.3% Smooth-BC, and 8.4% Diversity on BEAT2, while reducing jitter and foot sliding by -62.9% and -17.1% respectively. The code will be released to facilitate future research.


【12】LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
标题:LEMAS:大型150 K小时大规模可扩展多语言音频套件,具有生成语音模型
链接:https://arxiv.org/abs/2601.04233

作者:Zhiyuan Zhao,Lijian Lin,Ye Zhu,Kai Xie,Yunfei Liu,Yu Li
备注:Demo page: https://lemas-project.github.io/LEMAS-Project
摘要:我们介绍了LEMAS数据集,据我们所知,这是目前最大的开源多语言语音语料库与单词级的时间戳。LEMAS-Dataset涵盖10种主要语言的150,000多个小时,通过高效的数据处理管道构建,确保高质量的数据和注释。为了验证LEMAS数据集在不同生成范式中的有效性,我们在该数据集上训练了两个具有不同架构和任务专业化的基准模型。LEMAS-TTS基于非自回归流匹配框架,利用数据集的大规模和语言多样性来实现鲁棒的zero-shot多语言合成。我们提出的口音对抗训练和CTC损失减轻了跨语言口音问题,增强了合成稳定性。作为补充,LEMAS-Edit采用自回归解码器专用架构,该架构将语音编辑制定为掩码令牌填充任务。通过利用精确的词级对齐来构造训练模板,并采用自适应解码策略,实现了无缝,平滑边界的语音编辑与自然过渡。实验结果表明,在LEMAS-Dataset上训练的模型提供了高质量的合成和编辑性能,证实了数据集的质量。我们设想,这种丰富的时间戳注释,细粒度的多语言语料库将推动未来的进步,基于语音生成系统。
摘要:We present the LEMAS-Dataset, which, to our knowledge, is currently the largest open-source multilingual speech corpus with word-level timestamps. Covering over 150,000 hours across 10 major languages, LEMAS-Dataset is constructed via a efficient data processing pipeline that ensures high-quality data and annotations. To validate the effectiveness of LEMAS-Dataset across diverse generative paradigms, we train two benchmark models with distinct architectures and task specializations on this dataset. LEMAS-TTS, built upon a non-autoregressive flow-matching framework, leverages the dataset's massive scale and linguistic diversity to achieve robust zero-shot multilingual synthesis. Our proposed accent-adversarial training and CTC loss mitigate cross-lingual accent issues, enhancing synthesis stability. Complementarily, LEMAS-Edit employs an autoregressive decoder-only architecture that formulates speech editing as a masked token infilling task. By exploiting precise word-level alignments to construct training masks and adopting adaptive decoding strategies, it achieves seamless, smooth-boundary speech editing with natural transitions. Experimental results demonstrate that models trained on LEMAS-Dataset deliver high-quality synthesis and editing performance, confirming the dataset's quality. We envision that this richly timestamp-annotated, fine-grained multilingual corpus will drive future advances in prompt-based speech generation systems.


【13】Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
标题:防御合成语音:实时检测RVC语音转换攻击
链接:https://arxiv.org/abs/2601.04227

作者:Prajwal Chinchmalatpure,Suyash Chinchmalatpure,Siddharth Chavan
摘要:生成音频技术现在可以实现高度逼真的语音克隆和实时语音转换,增加了电话和视频通话等通信渠道中的冒充、欺诈和错误信息的风险。这项研究调查了使用基于检索的语音转换(RVC)生成的AI生成语音的实时检测,该语音转换在DEEP-VOICE数据集上进行评估,该数据集包括来自多个知名扬声器的真实语音和语音转换语音样本。为了模拟现实条件,deepfake生成应用于孤立的声音成分,然后重新引入背景氛围以抑制琐碎的伪影并强调特定于转换的提示。我们将帧检测作为流分类任务,将音频划分为一秒段,提取时频和倒谱特征,并训练监督机器学习模型将每个段分类为真实或语音转换。所提出的系统能够实现低延迟推理,同时支持段级决策和调用级聚合。实验结果表明,短窗口声学特征可以可靠地捕获与RVC语音相关的判别模式,即使在嘈杂的背景。这些发现证明了实用、实时的deepfake语音检测的可行性,并强调了在真实的音频混合条件下进行评估以实现稳健部署的重要性。
摘要:Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study investigates real-time detection of AI-generated speech produced using Retrieval-based Voice Conversion (RVC), evaluated on the DEEP-VOICE dataset, which includes authentic and voice-converted speech samples from multiple well-known speakers. To simulate realistic conditions, deepfake generation is applied to isolated vocal components, followed by the reintroduction of background ambiance to suppress trivial artifacts and emphasize conversion-specific cues. We frame detection as a streaming classification task by dividing audio into one-second segments, extracting time-frequency and cepstral features, and training supervised machine learning models to classify each segment as real or voice-converted. The proposed system enables low-latency inference, supporting both segment-level decisions and call-level aggregation. Experimental results show that short-window acoustic features can reliably capture discriminative patterns associated with RVC speech, even in noisy backgrounds. These findings demonstrate the feasibility of practical, real-time deepfake speech detection and underscore the importance of evaluating under realistic audio mixing conditions for robust deployment.


【14】From Imitation to Innovation: The Divergent Paths of Techno in Germany and the USA
标题:从模仿到创新:德国和美国Techno的分歧之路
链接:https://arxiv.org/abs/2601.04222

作者:Tim Ziemer,Simon Linke
摘要:有许多关于早期豪斯音乐和电子音乐的纪录片。在这里,场景中的主角描述了影响音乐演变的关键元素和事件。在研究界,有一种共识,即必须对这种描述进行批判性的审查。然而,还没有人试图根据音频分析来证实这种说法。在这项研究中,来自德国和美国的9,000多首早期房屋和技术曲目使用录音室功能,机器学习和推理统计进行了分析。可以提出三点意见:(1)德国和美国的house/techno音乐是不同的,2。美国的风格更相似,3。几乎没有演变随着时间的推移相比,德国的房子/技术关于录音室的功能。这些研究结果与文献记载的说法一致,因此提供了一个基于音频的视角,解释为什么techno在德国成为一种大众现象,但在美国仍然是一种边缘现象。这样的观察可以帮助音乐行业估计新趋势是会经历突破还是会消失。
摘要:Many documentaries on early house and techno music exist. Here, protagonists from the scenes describe key elements and events that affected the evolution of the music. In the research community, there is consensus that such descriptions have to be examined critically. Yet, there have not been attempts to validate such statements on the basis of audio analyses. In this study, over 9,000 early house and techno tracks from Germany and the United States of America are analyzed using recording studio features, machine learning and inferential statistics. Three observations can be made: 1.) German and US house/techno music are distinct, 2.) US styles are much more alike, and 3.) scarcely evolved over time compared to German house/techno regarding the recording studio features. These findings are in agreement with documented statements and thus provide an audio-based perspective on why techno became a mass phenomenon in Germany but remained a fringe phenomenon in the USA. Observations like these can help the music industry estimate whether new trends will experience a breakthrough or disappear.


【15】Predictive Controlled Music
标题:预测控制音乐
链接:https://arxiv.org/abs/2601.04221

作者:Midhun T. Augustine
备注:10 pages, 4 figures
摘要:本文提出了一种新的算法作曲方法,称为预测控制音乐(PCM),它结合了模型预测控制(MPC)与音乐生成。PCM使用动态模型来预测和优化音乐生成过程,其中通过优化性能测量以类似于MPC问题的方式计算音符。基于前向神经网络的评估函数用于评估生成的乐谱,作为PCM优化问题的目标函数。此外,一个递归神经网络模型被用来捕捉音符中的变量之间的关系,然后使用这个模型来定义PCM中的约束。与MPC类似,所提出的PCM以滚动时域的方式计算音符,从而导致反馈控制预测。数值例子来说明PCM生成方法。
摘要:This paper presents a new approach to algorithmic composition, called predictive controlled music (PCM), which combines model predictive control (MPC) with music generation. PCM uses dynamic models to predict and optimize the music generation process, where musical notes are computed in a manner similar to an MPC problem by optimizing a performance measure. A feedforward neural network-based assessment function is used to evaluate the generated musical score, which serves as the objective function of the PCM optimization problem. Furthermore, a recurrent neural network model is employed to capture the relationships among the variables in the musical notes, and this model is then used to define the constraints in the PCM. Similar to MPC, the proposed PCM computes musical notes in a receding-horizon manner, leading to feedback controlled prediction. Numerical examples are presented to illustrate the PCM generation method.


【16】Listen to Rhythm, Choose Movements: Autoregressive Multimodal Dance Generation via Diffusion and Mamba with Decoupled Dance Dataset
标题:聆听节奏,选择动作:通过扩散和曼巴与脱钩舞蹈数据集的自回归多模式舞蹈生成
链接:https://arxiv.org/abs/2601.03323

作者:Oran Duan,Yinghua Shen,Yingzhu Lv,Luyang Jie,Yaxin Liu,Qiong Wu
备注:12 pages, 13 figures
摘要:生成模型和序列学习的发展极大地促进了舞蹈动作生成的研究,但目前的方法仍然存在语义控制粗糙和长序列连贯性差的问题。在这项工作中,我们提出了听节奏,选择动作(LRCM),多模态引导的扩散框架,支持不同的输入方式和自回归舞蹈运动生成。我们探索了一种舞蹈数据集的特征解耦范例,并将其推广到Motorica Dance数据集,分离运动捕捉数据,音频节奏以及专业注释的全局和局部文本描述。我们的扩散架构集成了一个音频潜在的构象和文本潜在的交叉构象,并采用了运动时间曼巴模块(MTMM),使平滑,长时间的自回归合成。实验结果表明,LRCM提供了强大的性能,在功能能力和定量指标,表现出显着的潜力,在多模态输入场景和扩展序列生成。我们将在接受后公开发布完整的代码库,数据集和预训练模型。
摘要:Advances in generative models and sequence learning have greatly promoted research in dance motion generation, yet current methods still suffer from coarse semantic control and poor coherence in long sequences. In this work, we present Listen to Rhythm, Choose Movements (LRCM), a multimodal-guided diffusion framework supporting both diverse input modalities and autoregressive dance motion generation. We explore a feature decoupling paradigm for dance datasets and generalize it to the Motorica Dance dataset, separating motion capture data, audio rhythm, and professionally annotated global and local text descriptions. Our diffusion architecture integrates an audio-latent Conformer and a text-latent Cross-Conformer, and incorporates a Motion Temporal Mamba Module (MTMM) to enable smooth, long-duration autoregressive synthesis. Experimental results indicate that LRCM delivers strong performance in both functional capability and quantitative metrics, demonstrating notable potential in multimodal input scenarios and extended sequence generation. We will release the full codebase, dataset, and pretrained models publicly upon acceptance.


【17】Gradient-based Optimisation of Modulation Effects
标题:基于对象的调制效果优化
链接:https://arxiv.org/abs/2601.04867

作者:Alistair Carson,Alec Wright,Stefan Bilbao
备注:Submitted to J. Audio Eng. Soc. Dec. 2025
摘要:调制效果,如相位器,凸缘和合唱效果大量使用与电吉他。近年来已经研究了模拟调制单元的基于机器学习的仿真,但是与规范数字实现相比,大多数方法要么被限制于一类效果,要么遭受高计算成本或延迟。在这里,我们建立在以前的工作,并提出了一个框架,用于建模镶边,合唱和相位器效果的基础上微分数字信号处理。该模型在时频域中进行训练,但在推理时在时域中操作,需要零延迟。我们研究了与基于梯度的优化这些影响相关的挑战,并表明,低频加权的损失函数,避免收敛到局部极小值时,学习延迟时间。我们表明,当对模拟效果单元进行训练时,模型的声音输出在某些情况下与参考声音在感知上无法区分,但对于具有长延迟时间和反馈的效果仍然存在挑战。
摘要:Modulation effects such as phasers, flangers and chorus effects are heavily used in conjunction with the electric guitar. Machine learning based emulation of analog modulation units has been investigated in recent years, but most methods have either been limited to one class of effect or suffer from a high computational cost or latency compared to canonical digital implementations. Here, we build on previous work and present a framework for modelling flanger, chorus and phaser effects based on differentiable digital signal processing. The model is trained in the time-frequency domain, but at inference operates in the time-domain, requiring zero latency. We investigate the challenges associated with gradient-based optimisation of such effects, and show that low-frequency weighting of loss functions avoids convergence to local minima when learning delay times. We show that when trained against analog effects units, sound output from the model is in some cases perceptually indistinguishable from the reference, but challenges still remain for effects with long delay times and feedback.


【18】LLMs-Integrated Automatic Hate Speech Recognition Using Controllable Text Generation Models
标题:使用可控文本生成模型的LLM集成自动仇恨语音识别
链接:https://arxiv.org/abs/2601.04654

作者:Ryutaro Oshima,Yuya Hosoda,Youji Iiguni
备注:In Proceedings of the 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2025)
摘要:提出了一种基于大语言模型的仇恨语音自动识别模型。所提出的方法将ASR模型的编码器与LLM的解码器集成在一起,从而能够同时执行转录和审查任务,以防止有害内容的暴露。LLM的指令调整以使用特定标记来屏蔽仇恨相关的单词需要注释的仇恨言论数据集,这是有限的。我们生成文本样本使用LLM与思想链(CoT)提示技术指导的文化背景和例子,然后将它们转换成语音样本使用文本到语音(TTS)系统。然而,其中一些包含与仇恨相关的词的非仇恨言论样本,这降低了审查性能。本文过滤文本分类模型正确标记为仇恨内容的样本。通过调整正确答案模型数量的阈值,我们可以控制生成的数据集中的仇恨水平,使我们能够通过课程学习以渐进的方式训练LLM。实验结果表明,该方法对仇恨相关词的掩蔽准确率达到了58.6%,超过了以往的基线.我们还确认,课程培训有助于提高转录和检查任务的效率。
摘要:This paper proposes an automatic speech recognition (ASR) model for hate speech using large language models (LLMs). The proposed method integrates the encoder of the ASR model with the decoder of the LLMs, enabling simultaneous transcription and censorship tasks to prevent the exposure of harmful content. Instruction tuning of the LLM to mask hate-related words with specific tokens requires an annotated hate speech dataset, which is limited. We generate text samples using an LLM with the Chain-of-Thought (CoT) prompting technique guided by cultural context and examples and then convert them into speech samples using a text-to-speech (TTS) system. However, some of them contain non-hate speech samples with hate-related words, which degrades the censorship performance. This paper filters the samples which text classification models correctly label as hate content. By adjusting the threshold for the number of correct answer models, we can control the level of hate in the generated dataset, allowing us to train the LLMs through curriculum learning in a gradual manner. Experimental results show that the proposed method achieves a masking accuracy of 58.6\% for hate-related words, surpassing previous baselines. We also confirm that the curriculum training contributes to the efficiency of both transcription and censorship tasks.


【19】Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
标题:基于流匹配的语音增强算法
链接:https://arxiv.org/abs/2601.04459

作者:Da-Hee Yang,Joon-Hyuk Chang
备注:Accepted for publication in IEEE Signal Processing Letters
摘要:噪声鲁棒的自动语音识别(ASR)通常通过在识别前在波形水平上应用语音增强(SE)来解决。然而,由于残余失真和与ASR编码器的潜在空间的不匹配,语音级增强并不总是转化为一致的识别改进。在这封信中,我们介绍了一种补充策略,称为潜在水平的增强,扭曲的表示在ASR推理过程中得到改善。具体来说,我们提出了一个即插即用的流匹配细化模块(FM-Refiner),该模块对预训练的基于CTC的ASR编码器的输出延迟进行操作。经过训练,可以将不完美的延迟(直接来自嘈杂的输入或来自增强但不完美的语音)映射到干净的对应物,FM-Refiner仅在推理时应用,而无需微调ASR参数。实验表明,FM-Refiner始终降低字错误率,无论是直接应用于噪声输入,当与传统的SE前端相结合。这些结果表明,通过流匹配的潜在级别的细化提供了一个轻量级的和有效的补充,现有的SE方法,强大的ASR。
摘要:Noise-robust automatic speech recognition (ASR) has been commonly addressed by applying speech enhancement (SE) at the waveform level before recognition. However, speech-level enhancement does not always translate into consistent recognition improvements due to residual distortions and mismatches with the latent space of the ASR encoder. In this letter, we introduce a complementary strategy termed latent-level enhancement, where distorted representations are refined during ASR inference. Specifically, we propose a plug-and-play Flow Matching Refinement module (FM-Refiner) that operates on the output latents of a pretrained CTC-based ASR encoder. Trained to map imperfect latents-either directly from noisy inputs or from enhanced-but-imperfect speech-toward their clean counterparts, the FM-Refiner is applied only at inference, without fine-tuning ASR parameters. Experiments show that FM-Refiner consistently reduces word error rate, both when directly applied to noisy inputs and when combined with conventional SE front-ends. These results demonstrate that latent-level refinement via flow matching provides a lightweight and effective complement to existing SE approaches for robust ASR.


eess.AS音频处理


【1】Gradient-based Optimisation of Modulation Effects
标题:基于对象的调制效果优化
链接:https://arxiv.org/abs/2601.04867

作者:Alistair Carson,Alec Wright,Stefan Bilbao
备注:Submitted to J. Audio Eng. Soc. Dec. 2025
摘要:调制效果,如相位器,凸缘和合唱效果大量使用与电吉他。近年来已经研究了模拟调制单元的基于机器学习的仿真,但是与规范数字实现相比,大多数方法要么被限制于一类效果,要么遭受高计算成本或延迟。在这里,我们建立在以前的工作,并提出了一个框架,用于建模镶边,合唱和相位器效果的基础上微分数字信号处理。该模型在时频域中进行训练,但在推理时在时域中操作,需要零延迟。我们研究了与基于梯度的优化这些影响相关的挑战,并表明,低频加权的损失函数,避免收敛到局部极小值时,学习延迟时间。我们表明,当对模拟效果单元进行训练时,模型的声音输出在某些情况下与参考声音在感知上无法区分,但对于具有长延迟时间和反馈的效果仍然存在挑战。
摘要:Modulation effects such as phasers, flangers and chorus effects are heavily used in conjunction with the electric guitar. Machine learning based emulation of analog modulation units has been investigated in recent years, but most methods have either been limited to one class of effect or suffer from a high computational cost or latency compared to canonical digital implementations. Here, we build on previous work and present a framework for modelling flanger, chorus and phaser effects based on differentiable digital signal processing. The model is trained in the time-frequency domain, but at inference operates in the time-domain, requiring zero latency. We investigate the challenges associated with gradient-based optimisation of such effects, and show that low-frequency weighting of loss functions avoids convergence to local minima when learning delay times. We show that when trained against analog effects units, sound output from the model is in some cases perceptually indistinguishable from the reference, but challenges still remain for effects with long delay times and feedback.


【2】LLMs-Integrated Automatic Hate Speech Recognition Using Controllable Text Generation Models
标题:使用可控文本生成模型的LLM集成自动仇恨语音识别
链接:https://arxiv.org/abs/2601.04654

作者:Ryutaro Oshima,Yuya Hosoda,Youji Iiguni
备注:In Proceedings of the 17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC 2025)
摘要:提出了一种基于大语言模型的仇恨语音自动识别模型。所提出的方法将ASR模型的编码器与LLM的解码器集成在一起,实现同时转录和审查任务,以防止有害内容的暴露。LLM的指令调整以使用特定标记来屏蔽仇恨相关的单词需要注释的仇恨言论数据集,这是有限的。我们生成文本样本使用LLM与思想链(CoT)提示技术指导的文化背景和例子,然后将它们转换成语音样本使用文本到语音(TTS)系统。然而,其中一些包含与仇恨相关的词的非仇恨言论样本,这降低了审查性能。本文过滤文本分类模型正确标记为仇恨内容的样本。通过调整正确答案模型数量的阈值,我们可以控制生成的数据集中的仇恨水平,使我们能够通过课程学习以渐进的方式训练LLM。实验结果表明,该方法对仇恨相关词的掩蔽准确率达到了58.6%,超过了以往的基线.我们还确认,课程培训有助于提高转录和检查任务的效率。
摘要:This paper proposes an automatic speech recognition (ASR) model for hate speech using large language models (LLMs). The proposed method integrates the encoder of the ASR model with the decoder of the LLMs, enabling simultaneous transcription and censorship tasks to prevent the exposure of harmful content. Instruction tuning of the LLM to mask hate-related words with specific tokens requires an annotated hate speech dataset, which is limited. We generate text samples using an LLM with the Chain-of-Thought (CoT) prompting technique guided by cultural context and examples and then convert them into speech samples using a text-to-speech (TTS) system. However, some of them contain non-hate speech samples with hate-related words, which degrades the censorship performance. This paper filters the samples which text classification models correctly label as hate content. By adjusting the threshold for the number of correct answer models, we can control the level of hate in the generated dataset, allowing us to train the LLMs through curriculum learning in a gradual manner. Experimental results show that the proposed method achieves a masking accuracy of 58.6\% for hate-related words, surpassing previous baselines. We also confirm that the curriculum training contributes to the efficiency of both transcription and censorship tasks.


【3】Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
标题:基于流匹配的语音增强算法
链接:https://arxiv.org/abs/2601.04459

作者:Da-Hee Yang,Joon-Hyuk Chang
备注:Accepted for publication in IEEE Signal Processing Letters
摘要:噪声鲁棒的自动语音识别(ASR)通常通过在识别前在波形水平上应用语音增强(SE)来解决。然而,由于残余失真和与ASR编码器的潜在空间的不匹配,语音级增强并不总是转化为一致的识别改进。在这封信中,我们介绍了一种补充策略,称为潜在水平的增强,扭曲的表示在ASR推理过程中得到改善。具体来说,我们提出了一个即插即用的流匹配细化模块(FM-Refiner),该模块对预训练的基于CTC的ASR编码器的输出延迟进行操作。经过训练,可以将不完美的延迟(直接来自嘈杂的输入或来自增强但不完美的语音)映射到干净的对应物,FM-Refiner仅在推理时应用,而无需微调ASR参数。实验表明,FM-Refiner始终降低字错误率,无论是直接应用于噪声输入,当与传统的SE前端相结合。这些结果表明,通过流匹配的潜在级别的细化提供了一个轻量级的和有效的补充,现有的SE方法,强大的ASR。
摘要:Noise-robust automatic speech recognition (ASR) has been commonly addressed by applying speech enhancement (SE) at the waveform level before recognition. However, speech-level enhancement does not always translate into consistent recognition improvements due to residual distortions and mismatches with the latent space of the ASR encoder. In this letter, we introduce a complementary strategy termed latent-level enhancement, where distorted representations are refined during ASR inference. Specifically, we propose a plug-and-play Flow Matching Refinement module (FM-Refiner) that operates on the output latents of a pretrained CTC-based ASR encoder. Trained to map imperfect latents-either directly from noisy inputs or from enhanced-but-imperfect speech-toward their clean counterparts, the FM-Refiner is applied only at inference, without fine-tuning ASR parameters. Experiments show that FM-Refiner consistently reduces word error rate, both when directly applied to noisy inputs and when combined with conventional SE front-ends. These results demonstrate that latent-level refinement via flow matching provides a lightweight and effective complement to existing SE approaches for robust ASR.


【4】Summary of The Inaugural Music Source Restoration Challenge
标题:首届音乐来源恢复挑战赛总结
链接:https://arxiv.org/abs/2601.04343

作者:Yongyi Zang,Jiarui Hai,Wanying Ge,Qiuqiang Kong,Zheqi Dai,Helin Wang,Yuki Mitsufuji,Mark D. Plumbley
摘要:音乐源恢复(MSR)旨在从专业混合和降级的音频中恢复原始的未经处理的乐器,需要反转制作效果和真实世界的降级。我们提出了首届MSR挑战赛,其特点是使用Multi-Mel-SNR,Zimtohrli和FAD-CLAP对工作室制作的混合物进行客观评估,同时对真实世界的降级录音进行主观评估。五支队伍参加了挑战。获奖系统实现了4.46 dB Multi-Mel-SNR和3.47 MOS-Overall,分别比第二名系统相对提高了91%和18%。每个词干的分析揭示了不同乐器的还原难度有很大的差异,所有团队的低音平均为4.59 dB,而打击乐平均只有0.29 dB。数据集、评估方案和基线可在https://msrchallenge.com/上获得。
摘要:Music Source Restoration (MSR) aims to recover original, unprocessed instrument stems from professionally mixed and degraded audio, requiring the reversal of both production effects and real-world degradations. We present the inaugural MSR Challenge, which features objective evaluation on studio-produced mixtures using Multi-Mel-SNR, Zimtohrli, and FAD-CLAP, alongside subjective evaluation on real-world degraded recordings. Five teams participated in the challenge. The winning system achieved 4.46 dB Multi-Mel-SNR and 3.47 MOS-Overall, corresponding to relative improvements of 91% and 18% over the second-place system, respectively. Per-stem analysis reveals substantial variation in restoration difficulty across instruments, with bass averaging 4.59 dB across all teams, while percussion averages only 0.29 dB. The dataset, evaluation protocols, and baselines are available at https://msrchallenge.com/.


【5】SmoothSync: Dual-Stream Diffusion Transformers for Jitter-Robust Beat-Synchronized Gesture Generation from Quantized Audio
标题:SmoothSync:用于从量化音频生成抖动鲁棒节拍同步手势的双流扩散变换器
链接:https://arxiv.org/abs/2601.04236

作者:Yujiao Jiang,Qingmin Liao,Zongqing Lu
摘要:协同语音手势生成是一个关键的研究领域,旨在合成语音同步的类人手势。现有的方法往往遭受的问题,如节奏不一致,运动抖动,脚滑动和有限的多采样多样性。在本文中,我们提出了SmoothSync,一个新的框架,利用量化的音频令牌在一个新的双流扩散Transformer(DiT)架构合成整体的姿态和提高采样变化。具体而言,我们(1)通过互补的Transformer流融合音频运动特征以实现更好的同步,(2)引入抖动抑制损失以提高时间平滑度,(3)实现概率音频量化以从相同的输入生成不同的手势序列。为了可靠地评估抖动下的节拍同步,我们引入了Smooth-BC,这是一种对运动噪声不太敏感的节拍一致性度量的鲁棒变体。在BEAT 2和SHOW数据集上的综合实验证明了SmoothSync的优越性,在BEAT 2上的表现优于最先进的方法,分别为-30.6% FGD,10.3% Smooth-BC和8.4% Diversity,同时分别减少了-62.9%和-17.1%的抖动和脚滑动。该代码将被释放,以方便未来的研究。
摘要:Co-speech gesture generation is a critical area of research aimed at synthesizing speech-synchronized human-like gestures. Existing methods often suffer from issues such as rhythmic inconsistency, motion jitter, foot sliding and limited multi-sampling diversity. In this paper, we present SmoothSync, a novel framework that leverages quantized audio tokens in a novel dual-stream Diffusion Transformer (DiT) architecture to synthesis holistic gestures and enhance sampling variation. Specifically, we (1) fuse audio-motion features via complementary transformer streams to achieve superior synchronization, (2) introduce a jitter-suppression loss to improve temporal smoothness, (3) implement probabilistic audio quantization to generate distinct gesture sequences from identical inputs. To reliably evaluate beat synchronization under jitter, we introduce Smooth-BC, a robust variant of the beat consistency metric less sensitive to motion noise. Comprehensive experiments on the BEAT2 and SHOW datasets demonstrate SmoothSync's superiority, outperforming state-of-the-art methods by -30.6% FGD, 10.3% Smooth-BC, and 8.4% Diversity on BEAT2, while reducing jitter and foot sliding by -62.9% and -17.1% respectively. The code will be released to facilitate future research.


【6】LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
标题:LEMAS:大型150 K小时大规模可扩展多语言音频套件,具有生成语音模型
链接:https://arxiv.org/abs/2601.04233

作者:Zhiyuan Zhao,Lijian Lin,Ye Zhu,Kai Xie,Yunfei Liu,Yu Li
备注:Demo page: https://lemas-project.github.io/LEMAS-Project
摘要:我们介绍了LEMAS数据集,据我们所知,这是目前最大的开源多语言语音语料库与单词级的时间戳。LEMAS-Dataset涵盖10种主要语言的150,000多个小时,通过高效的数据处理管道构建,确保高质量的数据和注释。为了验证LEMAS数据集在不同生成范式中的有效性,我们在该数据集上训练了两个具有不同架构和任务专业化的基准模型。LEMAS-TTS基于非自回归流匹配框架,利用数据集的大规模和语言多样性来实现鲁棒的zero-shot多语言合成。我们提出的口音对抗训练和CTC损失减轻了跨语言口音问题,增强了合成稳定性。作为补充,LEMAS-Edit采用自回归解码器专用架构,该架构将语音编辑制定为掩码令牌填充任务。通过利用精确的词级对齐来构造训练模板,并采用自适应解码策略,实现了无缝,平滑边界的语音编辑与自然过渡。实验结果表明,在LEMAS-Dataset上训练的模型提供了高质量的合成和编辑性能,证实了数据集的质量。我们设想,这种丰富的时间戳注释,细粒度的多语言语料库将推动未来的进步,基于语音生成系统。
摘要:We present the LEMAS-Dataset, which, to our knowledge, is currently the largest open-source multilingual speech corpus with word-level timestamps. Covering over 150,000 hours across 10 major languages, LEMAS-Dataset is constructed via a efficient data processing pipeline that ensures high-quality data and annotations. To validate the effectiveness of LEMAS-Dataset across diverse generative paradigms, we train two benchmark models with distinct architectures and task specializations on this dataset. LEMAS-TTS, built upon a non-autoregressive flow-matching framework, leverages the dataset's massive scale and linguistic diversity to achieve robust zero-shot multilingual synthesis. Our proposed accent-adversarial training and CTC loss mitigate cross-lingual accent issues, enhancing synthesis stability. Complementarily, LEMAS-Edit employs an autoregressive decoder-only architecture that formulates speech editing as a masked token infilling task. By exploiting precise word-level alignments to construct training masks and adopting adaptive decoding strategies, it achieves seamless, smooth-boundary speech editing with natural transitions. Experimental results demonstrate that models trained on LEMAS-Dataset deliver high-quality synthesis and editing performance, confirming the dataset's quality. We envision that this richly timestamp-annotated, fine-grained multilingual corpus will drive future advances in prompt-based speech generation systems.


【7】Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
标题:防御合成语音:实时检测RVC语音转换攻击
链接:https://arxiv.org/abs/2601.04227

作者:Prajwal Chinchmalatpure,Suyash Chinchmalatpure,Siddharth Chavan
摘要:生成音频技术现在可以实现高度逼真的语音克隆和实时语音转换,增加了电话和视频通话等通信渠道中的冒充、欺诈和错误信息的风险。这项研究调查了使用基于检索的语音转换(RVC)生成的AI生成语音的实时检测,该语音转换在DEEP-VOICE数据集上进行评估,该数据集包括来自多个知名扬声器的真实语音和语音转换语音样本。为了模拟现实条件,deepfake生成应用于孤立的声音成分,然后重新引入背景氛围以抑制琐碎的伪影并强调特定于转换的提示。我们将帧检测作为流分类任务,将音频划分为一秒段,提取时频和倒谱特征,并训练监督机器学习模型将每个段分类为真实或语音转换。所提出的系统能够实现低延迟推理,同时支持段级决策和调用级聚合。实验结果表明,短窗口声学特征可以可靠地捕获与RVC语音相关的判别模式,即使在嘈杂的背景。这些发现证明了实用、实时的deepfake语音检测的可行性,并强调了在真实的音频混合条件下进行评估以实现稳健部署的重要性。
摘要:Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study investigates real-time detection of AI-generated speech produced using Retrieval-based Voice Conversion (RVC), evaluated on the DEEP-VOICE dataset, which includes authentic and voice-converted speech samples from multiple well-known speakers. To simulate realistic conditions, deepfake generation is applied to isolated vocal components, followed by the reintroduction of background ambiance to suppress trivial artifacts and emphasize conversion-specific cues. We frame detection as a streaming classification task by dividing audio into one-second segments, extracting time-frequency and cepstral features, and training supervised machine learning models to classify each segment as real or voice-converted. The proposed system enables low-latency inference, supporting both segment-level decisions and call-level aggregation. Experimental results show that short-window acoustic features can reliably capture discriminative patterns associated with RVC speech, even in noisy backgrounds. These findings demonstrate the feasibility of practical, real-time deepfake speech detection and underscore the importance of evaluating under realistic audio mixing conditions for robust deployment.


【8】From Imitation to Innovation: The Divergent Paths of Techno in Germany and the USA
标题:从模仿到创新:德国和美国Techno的分歧之路
链接:https://arxiv.org/abs/2601.04222

作者:Tim Ziemer,Simon Linke
摘要:有许多关于早期豪斯音乐和电子音乐的纪录片。在这里,场景中的主角描述了影响音乐演变的关键元素和事件。在研究界,有一种共识,即必须对这种描述进行批判性的审查。然而,还没有人试图根据音频分析来证实这种说法。在这项研究中,来自德国和美国的9,000多首早期房屋和技术曲目使用录音室功能,机器学习和推理统计进行了分析。可以提出三点意见:(1)德国和美国的house/techno音乐是不同的,2。美国的风格更相似,3。几乎没有演变随着时间的推移相比,德国的房子/技术关于录音室的功能。这些研究结果与文献记载的说法一致,因此提供了一个基于音频的视角,解释为什么techno在德国成为一种大众现象,但在美国仍然是一种边缘现象。这样的观察可以帮助音乐行业估计新趋势是会经历突破还是会消失。
摘要:Many documentaries on early house and techno music exist. Here, protagonists from the scenes describe key elements and events that affected the evolution of the music. In the research community, there is consensus that such descriptions have to be examined critically. Yet, there have not been attempts to validate such statements on the basis of audio analyses. In this study, over 9,000 early house and techno tracks from Germany and the United States of America are analyzed using recording studio features, machine learning and inferential statistics. Three observations can be made: 1.) German and US house/techno music are distinct, 2.) US styles are much more alike, and 3.) scarcely evolved over time compared to German house/techno regarding the recording studio features. These findings are in agreement with documented statements and thus provide an audio-based perspective on why techno became a mass phenomenon in Germany but remained a fringe phenomenon in the USA. Observations like these can help the music industry estimate whether new trends will experience a breakthrough or disappear.


【9】Predictive Controlled Music
标题:预测控制音乐
链接:https://arxiv.org/abs/2601.04221

作者:Midhun T. Augustine
备注:10 pages, 4 figures
摘要:本文提出了一种新的算法作曲方法,称为预测控制音乐(PCM),它结合了模型预测控制(MPC)与音乐生成。PCM使用动态模型来预测和优化音乐生成过程,其中通过优化性能测量以类似于MPC问题的方式计算音符。基于前向神经网络的评估函数用于评估生成的乐谱,作为PCM优化问题的目标函数。此外,一个递归神经网络模型被用来捕捉音符中的变量之间的关系,然后使用这个模型来定义PCM中的约束。与MPC类似,所提出的PCM以滚动时域的方式计算音符,从而导致反馈控制预测。数值例子来说明PCM生成方法。
摘要:This paper presents a new approach to algorithmic composition, called predictive controlled music (PCM), which combines model predictive control (MPC) with music generation. PCM uses dynamic models to predict and optimize the music generation process, where musical notes are computed in a manner similar to an MPC problem by optimizing a performance measure. A feedforward neural network-based assessment function is used to evaluate the generated musical score, which serves as the objective function of the PCM optimization problem. Furthermore, a recurrent neural network model is employed to capture the relationships among the variables in the musical notes, and this model is then used to define the constraints in the PCM. Similar to MPC, the proposed PCM computes musical notes in a receding-horizon manner, leading to feedback controlled prediction. Numerical examples are presented to illustrate the PCM generation method.


机器翻译由腾讯交互翻译提供,仅供参考