今日论文合集:cs.SD语音15篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】DARC: Drum accompaniment generation with fine-grained rhythm control
标题:DARC:鼓伴奏生成与细粒度节奏控制
链接:https://arxiv.org/abs/2601.02357

作者:Trey Brosnan
摘要:在音乐创作中,快速原型对于探索和提炼想法至关重要,但当用户需要结构控制和风格灵活性时,现有的生成工具往往会出现不足。之前的茎到茎生成方法可以以其他音乐茎为条件,但对节奏的控制有限,而音色转移方法允许用户指定特定的节奏,但不能以音乐上下文为条件。我们介绍DARC,生成鼓伴奏模型,条件都从其他干音乐背景和明确的节奏提示,如beatboxing或轻敲轨道。使用参数有效的微调,我们增强了阶段,一个国家的最先进的鼓干发电机,与细粒度的节奏控制,同时保持音乐的上下文意识。
摘要:In music creation, rapid prototyping is essential for exploring and refining ideas, yet existing generative tools often fall short when users require both structural control and stylistic flexibility. Prior approaches in stem-to-stem generation can condition on other musical stems but offer limited control over rhythm, and timbre-transfer methods allow users to specify specific rhythms, but cannot condition on musical context. We introduce DARC, a generative drum accompaniment model that conditions both on musical context from other stems and explicit rhythm prompts such as beatboxing or tapping tracks. Using parameter-efficient fine-tuning, we augment STAGE, a state-of-the-art drum stem generator, with fine-grained rhythm control while maintaining musical context awareness.


【2】ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging
标题:ARCADE:一个城市规模的阿拉伯语方言标注语料库
链接:https://arxiv.org/abs/2601.02209

作者:Omer Nacar,Serry Sibaee,Adel Ammar,Yasser Alhabashi,Nadia Samer Sibai,Yara Farouk Ahmed,Ahmed Saud Alqusaiyer,Sulieman Mahmoud AlMahmoud,Abdulrhman Mamdoh Mukhaniq,Lubaba Raed,Sulaiman Mohammed Alatwah,Waad Nasser Alqahtani,Yousif Abdulmajeed Alnasser,Mohamed Aziz Khadraoui,Wadii Boulila
摘要:阿拉伯语的特点是有丰富的地区方言,在语音和词汇上有很大的差异,反映了其发言者的地理和文化多样性。尽管有许多多方言数据集可用,但将语音映射到细粒度的方言源(如城市)仍然未得到充分探索。我们提出了ARCADE(阿拉伯语广播语料库音频方言评估),第一个阿拉伯语语音数据集明确设计与城市级方言粒度。该语料库包括从阿拉伯世界的流媒体服务收集的阿拉伯语广播讲话。我们的数据管道从经过验证的无线电流中捕获30秒的片段,包括现代标准阿拉伯语(MSA)和各种方言语音。为了确保可靠性,每个片段都由一到三名阿拉伯语母语评论员进行注释,他们分配了丰富的元数据,包括情感,语音类型,方言类别和方言识别任务的有效性标志。由此产生的语料库包括6,907个注释和3,790个独特的音频片段,跨越19个国家的58个城市。这些细粒度的注释实现了强大的多任务学习,作为城市级方言标记的基准。我们详细介绍了数据收集方法,评估音频质量,并提供标签分布的全面分析。该数据集可在以下网站获得:https://huggingface.co/datasets/riotu-lab/ARCADE-full
摘要:The Arabic language is characterized by a rich tapestry of regional dialects that differ substantially in phonetics and lexicon, reflecting the geographic and cultural diversity of its speakers. Despite the availability of many multi-dialect datasets, mapping speech to fine-grained dialect sources, such as cities, remains underexplored. We present ARCADE (Arabic Radio Corpus for Audio Dialect Evaluation), the first Arabic speech dataset designed explicitly with city-level dialect granularity. The corpus comprises Arabic radio speech collected from streaming services across the Arab world. Our data pipeline captures 30-second segments from verified radio streams, encompassing both Modern Standard Arabic (MSA) and diverse dialectal speech. To ensure reliability, each clip was annotated by one to three native Arabic reviewers who assigned rich metadata, including emotion, speech type, dialect category, and a validity flag for dialect identification tasks. The resulting corpus comprises 6,907 annotations and 3,790 unique audio segments spanning 58 cities across 19 countries. These fine-grained annotations enable robust multi-task learning, serving as a benchmark for city-level dialect tagging. We detail the data collection methodology, assess audio quality, and provide a comprehensive analysis of label distributions. The dataset is available on: https://huggingface.co/datasets/riotu-lab/ARCADE-full


【3】A Mamba-Based Model for Automatic Chord Recognition
标题:基于Mamba的和弦自动识别模型
链接:https://arxiv.org/abs/2601.02101

作者:Chunyu Yuan,Johanna Devaney
备注:International Society of Music Information Retrieval, Late-Breaking Demo 2024
摘要:在这项工作中,我们提出了一个新的有效的解决方案,这是一个基于Mamba的模型命名为BMACE(双向Mamba为基础的网络,自动和弦估计),它利用双向Mamba层中的选择性结构化状态空间模型,有效地建模时间依赖性。我们的模型实现了与最先进的模型相媲美的高预测性能,具有所需参数少和计算资源少的优点
摘要:In this work, we propose a new efficient solution, which is a Mamba-based model named BMACE (Bidirectional Mamba-based network, for Automatic Chord Estimation), which utilizes selective structured state-space models in a bidirectional Mamba layer to effectively model temporal dependencies. Our model achieves high prediction performance comparable to state-of-the-art models, with the advantage of requiring fewer parameters and lower computational resources


【4】BeatlesFC: Harmonic function annotations of Isophonics' The Beatles dataset
标题:BeatlesFC:Isophonics的The Beatles数据集的调和函数注释
链接:https://arxiv.org/abs/2601.02099

作者:Ji Yeoung Sim,Rebecca Moranis,Johanna Devaney
备注:International Society for Music Information Retrieval, Late-Breaking Demo 2024
摘要:本文介绍了BeatlesFC,Isophonics的The Beatles数据集的一组调和函数注释。和声函数标注将和弦标签表征为稳定(主音)或不稳定(主要,属音)。它们在乐句的层次上运作,充当和弦标签和更高层次的正式结构之间的联系。
摘要:This paper presents BeatlesFC, a set of harmonic function annotations for Isophonics' The Beatles dataset. Harmonic function annotations characterize chord labels as stable (tonic) or unstable (predominant, dominant). They operate at the level of musical phrases, serving as a link between chord labels and higher-level formal structures.


【5】HyperCLOVA X 8B Omni
标题:HyperCLVA X 8B Omni
链接:https://arxiv.org/abs/2601.01792

作者:NAVER Cloud HyperCLOVA X Team
备注:Technical Report
摘要:在本报告中,我们介绍了HyperCLOVA X 8B Omni,这是HyperCLOVA X系列中第一款支持文本、音频和视觉作为输入和输出的全模态机型。通过将多模态理解和生成整合到单个模型中,而不是单独的特定模态管道中,HyperCLOVA X 8B Omni可作为8B级全向寻路点,实现实用的任意对任意全向助手。在高层次上,该模型通过一个共享的下一个令牌预测接口在一个交织的多模态序列上统一模态,而视觉和音频编码器则注入连续的嵌入,以实现细粒度的理解和基础。实证评估表明,在韩语和英语中,跨文本,音频和视觉的不同输入输出组合的小型模型的竞争力表现。我们预计HyperCLOVA X 8B Omni的开放重量版本将支持广泛的研究和部署场景。
摘要:In this report, we present HyperCLOVA X 8B Omni, the first any-to-any omnimodal model in the HyperCLOVA X family that supports text, audio, and vision as both inputs and outputs. By consolidating multimodal understanding and generation into a single model rather than separate modality-specific pipelines, HyperCLOVA X 8B Omni serves as an 8B-scale omni-pathfinding point toward practical any-to-any omni assistants. At a high level, the model unifies modalities through a shared next-token prediction interface over an interleaved multimodal sequence, while vision and audio encoders inject continuous embeddings for fine-grained understanding and grounding. Empirical evaluations demonstrate competitive performance against comparably sized models across diverse input-output combinations spanning text, audio, and vision, in both Korean and English. We anticipate that the open-weight release of HyperCLOVA X 8B Omni will support a wide range of research and deployment scenarios.


【6】MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
标题:MM-Sonate:具有零拍摄语音克隆的多模态可控音视频生成
链接:https://arxiv.org/abs/2601.01568

作者:Chunyu Qiang,Jun Wang,Xiaopeng Wang,Kang Yin,Yuxin Guo,Xijuan Zeng,Nan Li,Zihan Li,Yuzhe Liang,Ziyu Zhang,Teng Ma,Yushen Chen,Zhongliang Liu,Feng Deng,Chen Zhang,Pengfei Wan
摘要:联合音频视频生成的目的是合成同步的多感官内容,但目前的统一模型与细粒度的声学控制,特别是身份保留语音斗争。现有的方法或者由于级联生成而遭受时间不对准,或者缺乏在联合合成框架内执行zero-shot语音克隆的能力。在这项工作中,我们提出了MM-Sonate,一个多模式流匹配框架,它将可控的音频-视频联合生成与zero-shot语音克隆功能相结合。与依赖于粗糙的语义描述的先前作品不同,MM-Sonate利用统一的发音音素输入来执行严格的语言和时间对齐。为了实现zero-shot语音克隆,我们引入了一种音色注入机制,有效地从语言内容中提取说话人身份。此外,解决标准的无分类器的指导在多模态设置的局限性,我们提出了一个基于噪声的负调节策略,利用自然噪声先验显着提高声学保真度。经验评估表明,MM-Sonate在联合生成基准中建立了新的最先进的性能,在嘴唇同步和语音清晰度方面显著优于基线,同时实现了与专门的文本到语音系统相当的语音克隆保真度。
摘要:Joint audio-video generation aims to synthesize synchronized multisensory content, yet current unified models struggle with fine-grained acoustic control, particularly for identity-preserving speech. Existing approaches either suffer from temporal misalignment due to cascaded generation or lack the capability to perform zero-shot voice cloning within a joint synthesis framework. In this work, we present MM-Sonate, a multimodal flow-matching framework that unifies controllable audio-video joint generation with zero-shot voice cloning capabilities. Unlike prior works that rely on coarse semantic descriptions, MM-Sonate utilizes a unified instruction-phoneme input to enforce strict linguistic and temporal alignment. To enable zero-shot voice cloning, we introduce a timbre injection mechanism that effectively decouples speaker identity from linguistic content. Furthermore, addressing the limitations of standard classifier-free guidance in multimodal settings, we propose a noise-based negative conditioning strategy that utilizes natural noise priors to significantly enhance acoustic fidelity. Empirical evaluations demonstrate that MM-Sonate establishes new state-of-the-art performance in joint generation benchmarks, significantly outperforming baselines in lip synchronization and speech intelligibility, while achieving voice cloning fidelity comparable to specialized Text-to-Speech systems.


【7】MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization
标题:MOSS转录日记:通过扬声器日记进行准确转录
链接:https://arxiv.org/abs/2601.01554

作者:Donghua Yu,Zhengyuan Lin,Chen Yang,Yiyang Zhang,Zhaoye Fei,Hanfu Chen,Jingqi Chen,Ke Chen,Qinyuan Cheng,Liwei Fan,Yi Jiang,Jie Zhu,Muchen Li,Shimin Li,Wenxuan Wang,Yang Wang,Zhe Xu,Yitian Gong,Yuqian Zhang
摘要:Speaker-Attributed,Time-Stamped Transcription(SATS)旨在转录发言内容并精确确定每个发言者的时间,这对于会议转录特别有价值。现有的SATS系统很少采用端到端的公式,并进一步受到有限的上下文窗口,弱远程扬声器记忆,以及无法输出时间戳的限制。为了解决这些限制,我们提出了MOSS Transcribe Diarize,一个统一的多模态大型语言模型,它在端到端的范例中联合执行说话人属性,时间戳转录。经过大量真实野生数据的训练,并配备了一个128 k的上下文窗口,最多可输入90分钟,MOSS Transcribe Diarize扩展良好,泛化能力强。在综合评估中,它在多个公共和内部基准上的表现优于最先进的商业系统。
摘要:Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak long-range speaker memory, and the inability to output timestamps. To address these limitations, we present MOSS Transcribe Diarize, a unified multimodal large language model that jointly performs Speaker-Attributed, Time-Stamped Transcription in an end-to-end paradigm. Trained on extensive real wild data and equipped with a 128k context window for up to 90-minute inputs, MOSS Transcribe Diarize scales well and generalizes robustly. Across comprehensive evaluations, it outperforms state-of-the-art commercial systems on multiple public and in-house benchmarks.


【8】Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
标题:弥合差距:Speech-LLM和多语言会话ASR端到端架构的比较探索
链接:https://arxiv.org/abs/2601.01461

作者:Yuxiang Mei,Dongxing Xu,Jiaen Liang,Yanhua Long
备注:5 pages, 1 figure
摘要:INTERSPEECH 2025多语言会话语音语言模型(MLC-SLM)挑战赛通过大型语言模型(LLM)促进多语言会话ASR。我们之前的SHNU-mASR系统采用了一种具有竞争力的并行语音编码器架构,将Whisper和mHuBERT与LLM集成在一起。然而,它面临着两个挑战:简单的特征级联可能无法充分利用互补信息,以及基于LLM的ASR和端到端(E2 E)编码器-解码器ASR之间的性能差距仍未被探索。在这项工作中,我们提出了一个增强的基于LLM的ASR框架,该框架将微调的Whisper和mHuBERT编码器与LLM相结合,以丰富语音表示。我们首先使用LoRA评估E2 E Whisper模型,并在MLC-SLM ASR任务上进行全面微调,然后为并行语音编码器提出基于交叉注意的融合机制。在MLC-SLM挑战赛的官方评估集上,我们的系统达到了10.69%的CER/WER,与排名第一的Track 1系统不相上下,尽管与他们的大规模训练集相比,它只使用了1,500小时的基线训练数据。尽管如此,我们发现我们最终的基于LLM的ASR仍然与微调的E2 E Whisper模型的性能不匹配,为未来的Speech-LLM设计提供了宝贵的经验指导。我们的代码可在https://github.com/1535176727/MLC-SLM上公开获取。
摘要:The INTERSPEECH 2025 Challenge on Multilingual Conversational Speech Language Models (MLC-SLM) promotes multilingual conversational ASR with large language models (LLMs). Our previous SHNU-mASR system adopted a competitive parallel-speech-encoder architecture that integrated Whisper and mHuBERT with an LLM. However, it faced two challenges: simple feature concatenation may not fully exploit complementary information, and the performance gap between LLM-based ASR and end-to-end(E2E) encoder-decoder ASR remained unexplored. In this work, we present an enhanced LLM-based ASR framework that combines fine-tuned Whisper and mHuBERT encoders with an LLM to enrich speech representations. We first evaluate E2E Whisper models with LoRA and full fine-tuning on the MLC-SLM ASR task, and then propose cross-attention-based fusion mechanisms for the parallel-speech-encoder. On the official evaluation set of the MLC-SLM Challenge, our system achieves a CER/WER of 10.69%, ranking on par with the top-ranked Track 1 systems, even though it uses only 1,500 hours of baseline training data compared with their large-scale training sets. Nonetheless, we find that our final LLM-based ASR still does not match the performance of a fine-tuned E2E Whisper model, providing valuable empirical guidance for future Speech-LLM design. Our code is publicly available at https://github.com/1535176727/MLC-SLM.


【9】OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
标题:DV-DirectTTC:迈向开放词汇指导文本到语音
链接:https://arxiv.org/abs/2601.01459

作者:Yong Ren,Jiangyan Yi,Jianhua Tao,Haiyang Sun,Zhengqi Wen,Hao Gu,Le Xu,Ye Bai
摘要:指令文本到语音(InstructTTS)利用自然语言描述作为风格提示来指导语音合成。然而,现有的InstructTTS方法主要依赖于与音频相关的标签的直接组合或其不同的改写,使得难以处理灵活的、高级别的指令。这种严格的控制对于诸如希望用描述性指令来操纵生成的内容创建者之类的用户是不够的。为了解决这些限制,我们引入OV-InstructTTS,一个新的范例开放词汇的InstructTTS。我们提出了一个全面的解决方案,包括一个新策划的数据集,OV-Speech,和一个新的推理驱动的框架。OV-Speech数据集将语音与开放词汇指令配对,每个指令都通过将高级指令与声学特征连接起来的推理过程进行增强。推理驱动的框架在合成语音之前从开放词汇指令推断情感、声学和非语言信息。评估表明,这种推理驱动的方法显着提高了遵循的保真度和语音表达。我们相信这项工作可以激发下一个用户友好的InstructTTS系统具有更强的泛化能力和现实世界的适用性。数据集和演示在我们的项目页面上公开。
摘要:Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a direct combination of audio-related labels or their diverse rephrasings, making it difficult to handle flexible, high-level instructions. Such rigid control is insufficient for users such as content creators who wish to steer generation with descriptive instructions. To address these constraints, we introduce OV-InstructTTS, a new paradigm for open-vocabulary InstructTTS. We propose a comprehensive solution comprising a newly curated dataset, OV-Speech, and a novel reasoning-driven framework. The OV-Speech dataset pairs speech with open-vocabulary instructions, each augmented with a reasoning process that connects high-level instructions to acoustic features. The reasoning-driven framework infers emotional, acoustic, and paralinguistic information from open-vocabulary instructions before synthesizing speech. Evaluations show that this reasoning-driven approach significantly improves instruction-following fidelity and speech expressiveness. We believe this work can inspire the next user-friendly InstructTTS systems with stronger generalization and real-world applicability. The dataset and demos are publicly available on our project page.


【10】SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
标题:SAFE-QAAC:通过强化学习进行端到端慢思维音频文本欺诈检测
链接:https://arxiv.org/abs/2601.01392

作者:Peidong Wang,Zhiming Ma,Xin Dai,Yongkang Liu,Shi Feng,Xiaocui Yang,Wenxing Hu,Zhihao Wang,Mingjun Pan,Li Yuan,Daling Wang
摘要:现有的欺诈检测方法主要依赖于转录的文本,遭受ASR错误和丢失关键的声学线索,如音调和环境背景。这就限制了它们对抗复杂欺骗策略的有效性。为了应对这些挑战,我们首先提出了\textbf{SAFE-QAQ},这是一个端到端的综合框架,用于基于音频的慢思维欺诈检测。首先,SAFE-QAQ框架消除了转录错误对检测性能的影响。其次,我们提出了基于规则的慢思维奖励机制,系统地引导系统通过准确捕捉细粒度的音频细节,通过分层推理过程,识别欺诈指示模式。此外,我们的框架在实时呼叫期间引入了动态风险评估框架,从而能够早期发现和预防欺诈。在TeleAntiFraud-Bench上的实验表明,SAFE-QAQ在多个关键方面(包括准确性,推理效率和实时处理能力)比现有方法有了显着的改进。SAFE-QAQ目前每天部署并分析超过70,000个呼叫,可有效地自动执行复杂的欺诈检测,减少人工工作量和财务损失。代码:https://anonymous.4open.science/r/SAFE-QAQ。
摘要:Existing fraud detection methods predominantly rely on transcribed text, suffering from ASR errors and missing crucial acoustic cues like vocal tone and environmental context. This limits their effectiveness against complex deceptive strategies. To address these challenges, we first propose \textbf{SAFE-QAQ}, an end-to-end comprehensive framework for audio-based slow-thinking fraud detection. First, the SAFE-QAQ framework eliminates the impact of transcription errors on detection performance. Secondly, we propose rule-based slow-thinking reward mechanisms that systematically guide the system to identify fraud-indicative patterns by accurately capturing fine-grained audio details, through hierarchical reasoning processes. Besides, our framework introduces a dynamic risk assessment framework during live calls, enabling early detection and prevention of fraud. Experiments on the TeleAntiFraud-Bench demonstrate that SAFE-QAQ achieves dramatic improvements over existing methods in multiple key dimensions, including accuracy, inference efficiency, and real-time processing capabilities. Currently deployed and analyzing over 70,000 calls daily, SAFE-QAQ effectively automates complex fraud detection, reducing human workload and financial losses. Code: https://anonymous.4open.science/r/SAFE-QAQ.


【11】UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
标题:UltraEval-音频:音频基础模型综合评估的统一框架
链接:https://arxiv.org/abs/2601.01373

作者:Qundong Shi,Jie Zhou,Biyuan Lin,Junbo Cui,Guoyang Zeng,Yixuan Zhou,Ziyang Wang,Xin Liu,Zhen Luo,Yudong Wang,Zhiyuan Liu
备注:13 pages, 2 figures
摘要:自GPT-4 o出现以来,音频基础模型的发展迅速加速。然而,缺乏全面的评估已经成为该领域进一步发展的关键瓶颈,特别是在音频生成方面。当前的音频评估面临三大挑战:(1)音频评估缺乏统一的框架,数据集和代码分散在各个来源,阻碍了公平和有效的跨模型比较;(2)音频编解码器作为音频基础模型的关键组成部分,缺乏广泛接受的整体评估方法;(3)现有的语音基准严重依赖于英语,这使得客观地评估模型在中文上的性能具有挑战性。为了解决第一个问题,我们引入了UltraEval-Audio,这是一个针对音频基础模型的统一评估框架,专为音频理解和生成任务而设计。UltraEval-Audio采用模块化架构,支持10种语言和14个核心任务类别,同时无缝集成24个主流模型和36个权威基准。为了提高研究效率,该框架提供了一个命令评估功能,并配有实时公共排行榜。对于第二个挑战,UltraEval-Audio采用了一种新颖的音频编解码器综合评估方案,从三个关键维度评估性能:语义准确性、音色保真度和声学质量。为了解决第三个问题,我们提出了两个新的汉语基准,SpeechCMMLU和SpeechHSK,旨在评估汉语知识水平和语言流利度。我们希望UltraEval-Audio能够为学术界和工业界提供一个透明、高效和公平的音频模型比较平台。我们的代码、基准测试和排行榜可在https://github.com/OpenBMB/UltraEval-Audio上获得。
摘要:The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particularly in audio generation. Current audio evaluation faces three major challenges: (1) audio evaluation lacks a unified framework, with datasets and code scattered across various sources, hindering fair and efficient cross-model comparison;(2) audio codecs, as a key component of audio foundation models, lack a widely accepted and holistic evaluation methodology; (3) existing speech benchmarks are heavily reliant on English, making it challenging to objectively assess models' performance on Chinese. To address the first issue, we introduce UltraEval-Audio, a unified evaluation framework for audio foundation models, specifically designed for both audio understanding and generation tasks. UltraEval-Audio features a modular architecture, supporting 10 languages and 14 core task categories, while seamlessly integrating 24 mainstream models and 36 authoritative benchmarks. To enhance research efficiency, the framework provides a one-command evaluation feature, accompanied by real-time public leaderboards. For the second challenge, UltraEval-Audio adopts a novel comprehensive evaluation scheme for audio codecs, evaluating performance across three key dimensions: semantic accuracy, timbre fidelity, and acoustic quality. To address the third issue, we propose two new Chinese benchmarks, SpeechCMMLU and SpeechHSK, designed to assess Chinese knowledge proficiency and language fluency. We wish that UltraEval-Audio will provide both academia and industry with a transparent, efficient, and fair platform for comparison of audio models. Our code, benchmarks, and leaderboards are available at https://github.com/OpenBMB/UltraEval-Audio.


【12】Diffusion Timbre Transfer Via Mutual Information Guided Inpainting
标题:基于互信息引导的扩散音色修复
链接:https://arxiv.org/abs/2601.01294

作者:Ching Ho Lee,Javier Nistal,Stefan Lattner,Marco Pasini,George Fazekas
备注:6 pages, 2 figures, 3 tables
摘要:我们研究音色转移作为一个推断时间编辑问题的音乐音频。从一个强大的预先训练的潜在扩散模型开始,我们引入了一个轻量级的过程,不需要额外的训练:(i)一个维度上的噪声注入,目标是最能提供乐器身份信息的潜在通道,以及(ii)一个早期的箝位机制,在反向扩散过程中重新施加输入的旋律和节奏结构。该方法直接对音频潜伏期进行操作并且与文本/音频调节(例如,CLAP)。我们讨论了设计选择,分析了音色变化和结构保留之间的权衡,并表明简单的推理时间控制可以有意义地引导预训练模型用于风格转换用例。
摘要:We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a dimension-wise noise injection that targets latent channels most informative of instrument identity, and (ii) an early-step clamping mechanism that re-imposes the input's melodic and rhythmic structure during reverse diffusion. The method operates directly on audio latents and is compatible with text/audio conditioning (e.g., CLAP). We discuss design choices,analyze trade-offs between timbral change and structural preservation, and show that simple inference-time controls can meaningfully steer pre-trained models for style-transfer use cases.


【13】IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection
标题:IO-RAE:音频隐私保护的信息混淆可逆对抗示例
链接:https://arxiv.org/abs/2601.01239

作者:Jiajie Zhu,Xia Du,Xiaoyuan Liu,Jizhe Zhou,Qizhen Xu,Zheng Lin,Chi-Man Pun
备注:10 pages, 5 figures
摘要:人工智能的快速发展大大加速了语音识别技术的采用,导致其在各种应用中的广泛集成。然而,这种使用量的激增也凸显了一个关键问题:音频数据极易受到未经授权的暴露和分析,给企业和个人带来重大的隐私风险。本文介绍了一个信息混淆可逆对抗示例(IO-RAE)框架,这是一种开创性的方法,旨在使用可逆对抗示例来保护音频隐私。IO-RAE利用大型语言模型生成误导性但上下文一致的内容,有效防止人类和自动语音识别(ASR)系统进行未经授权的窃听。此外,我们还提出了累积信号攻击技术,该技术通过针对低频信号来减轻高频噪声并提高攻击效率。我们的方法确保了音频数据的保护,而不会降低其质量或我们的能力。实验评估表明,我们的方法的优越性,实现了96.5%的有针对性的误导率和显着的100%的非有针对性的误导率混淆目标关键字在多个ASR模型,包括商业黑盒系统从谷歌。此外,通过语音质量的感知评估得分来衡量,恢复的音频质量达到4.45,与高质量的原始录音相当。值得注意的是,ASR系统处理的恢复音频显示出0%的错误率,这表明几乎无损恢复。这些结果突出了我们的IO-RAE框架在保护敏感音频隐私方面的实用性和有效性。
摘要:The rapid advancements in artificial intelligence have significantly accelerated the adoption of speech recognition technology, leading to its widespread integration across various applications. However, this surge in usage also highlights a critical issue: audio data is highly vulnerable to unauthorized exposure and analysis, posing significant privacy risks for businesses and individuals. This paper introduces an Information-Obfuscation Reversible Adversarial Example (IO-RAE) framework, the pioneering method designed to safeguard audio privacy using reversible adversarial examples. IO-RAE leverages large language models to generate misleading yet contextually coherent content, effectively preventing unauthorized eavesdropping by humans and Automatic Speech Recognition (ASR) systems. Additionally, we propose the Cumulative Signal Attack technique, which mitigates high-frequency noise and enhances attack efficacy by targeting low-frequency signals. Our approach ensures the protection of audio data without degrading its quality or our ability. Experimental evaluations demonstrate the superiority of our method, achieving a targeted misguidance rate of 96.5% and a remarkable 100% untargeted misguidance rate in obfuscating target keywords across multiple ASR models, including a commercial black-box system from Google. Furthermore, the quality of the recovered audio, measured by the Perceptual Evaluation of Speech Quality score, reached 4.45, comparable to high-quality original recordings. Notably, the recovered audio processed by ASR systems exhibited an error rate of 0%, indicating nearly lossless recovery. These results highlight the practical applicability and effectiveness of our IO-RAE framework in protecting sensitive audio privacy.


【14】Index-ASR Technical Report
标题:指数-ASB技术报告
链接:https://arxiv.org/abs/2601.00890

作者:Zheshu Song,Lu Wang,Wei Deng,Zhuo Yang,Yong Wu,Bin Xia
备注:Index-ASR technical report
摘要:自动语音识别(ASR)近年来取得了显著的进展,主要是由于基于LLM的ASR范式的出现。尽管它们在各种开源基准测试中表现出色,但现有的基于LLM的ASR系统仍然受到两个关键限制。首先,它们容易产生幻觉错误,经常产生过长和重复的输出,这些输出在声学输入中没有很好的基础。其次,它们对灵活和细粒度的上下文定制提供了有限的支持。为了应对这些挑战,我们提出了Index-ASR,这是一个基于LLM的大规模ASR系统,旨在同时增强鲁棒性并支持可定制的热词识别。Index-ASR的核心思想在于将LLM与富含背景噪声和上下文信息的大规模训练数据相结合。实验结果表明,我们的Index-ASR在开源基准测试和内部测试集上都取得了很好的性能,突出了它对现实世界ASR应用的鲁棒性和实用性。
摘要:Automatic speech recognition (ASR) has witnessed remarkable progress in recent years, largely driven by the emergence of LLM-based ASR paradigm. Despite their strong performance on a variety of open-source benchmarks, existing LLM-based ASR systems still suffer from two critical limitations. First, they are prone to hallucination errors, often generating excessively long and repetitive outputs that are not well grounded in the acoustic input. Second, they provide limited support for flexible and fine-grained contextual customization. To address these challenges, we propose Index-ASR, a large-scale LLM-based ASR system designed to simultaneously enhance robustness and support customizable hotword recognition. The core idea of Index-ASR lies in the integration of LLM and large-scale training data enriched with background noise and contextual information. Experimental results show that our Index-ASR achieves strong performance on both open-source benchmarks and in-house test sets, highlighting its robustness and practicality for real-world ASR applications.


【15】Bayesian Negative Binomial Regression of Afrobeats Chart Persistence
标题:非洲节拍图表持续性的Bayesian负二项回归
链接:https://arxiv.org/abs/2601.01391

作者:Ian Jacob Cabansag,Paul Ntegeka
摘要:Afrobeats歌曲在流媒体平台上争夺关注,在这些平台上,图表的可见性可以影响收入和文化影响。本文使用2024年以来的每日尼日利亚Spotify Top 200数据,研究合作是否有助于歌曲在排行榜上保持更长时间。每个轨道是总结了多少天,它出现在前200名在一年中,其总的年度流在尼日利亚。应用贝叶斯负二项回归,以图表上的天数作为结果,以协作状态(单人与多艺术家)和日志总流作为预测因子。这种方法非常适合过度分散的计数数据,并允许在控制整体流行度的同时解释协作的效果。使用马尔可夫链蒙特卡罗进行后验推理,并使用率比,后验概率和预测检查评估结果。调查结果表明,在考虑到总流量后,合作曲目在排行榜上的天数往往略少于可比的独奏曲目。
摘要:Afrobeats songs compete for attention on streaming platforms, where chart visibility can influence both revenue and cultural impact. This paper examines whether collaborations help songs remain on the charts longer, using daily Nigeria Spotify Top 200 data from 2024. Each track is summarized by the number of days it appears in the Top 200 during the year and its total annual streams in Nigeria. A Bayesian negative binomial regression is applied, with days on chart as the outcome and collaboration status (solo versus multi-artist) and log total streams as predictors. This approach is well suited for overdispersed count data and allows the effect of collaboration to be interpreted while controlling for overall popularity. Posterior inference is conducted using Markov chain Monte Carlo, and results are assessed using rate ratios, posterior probabilities, and predictive checks. The findings indicate that, after accounting for total streams, collaboration tracks tend to spend slightly fewer days on the chart than comparable solo tracks.


eess.AS音频处理


【1】On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
标题:空间特征在基于基础模型的说话人Dialogue中的作用
链接:https://arxiv.org/abs/2601.02231

作者:Marc Deegen,Tobias Gburrek,Tobias Cord-Landwehr,Thilo von Neumann,Jiangyu Han,Lukáš Burget,Reinhold Haeb-Umbach
备注:Accepted at HSCMA 2026
摘要:说话人日志化的最新进展利用了大型预训练基础模型,如WavLM,以在多个数据集上实现最先进的性能。像DiariZen这样的系统利用了这些丰富的单声道表示,但仅限于单声道音频,从而阻止了在多声道录音中使用可用的空间线索。这项工作分析了将空间信息纳入一个国家的最先进的单通道日记系统的影响,通过评估几种策略,调节模型的多通道空间功能。会议风格的数据集上的实验表明,空间信息可以提高日志化性能,但整体改善小于预期的建议系统,这表明在所有WavLM层聚合的功能已经捕获了准确的说话人区分所需的信息,也在重叠的语音区域。这些研究结果提供了深入了解的潜力和局限性,使用空间线索,以提高基础模型为基础的日记。
摘要:Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZen leverage these rich single-channel representations, but are limited to single-channel audio, preventing the use of spatial cues available in multi-channel recordings. This work analyzes the impact of incorporating spatial information into a state-of-the-art single-channel diarization system by evaluating several strategies for conditioning the model on multi-channel spatial features. Experiments on meeting-style datasets indicate that spatial information can improve diarization performance, but the overall improvement is smaller than expected for the proposed system, suggesting that the features aggregated over all WavLM layers already capture much of the information needed for accurate speaker discrimination, also in overlapping speech regions. These findings provide insight into the potential and limitations of using spatial cues to enhance foundation model-based diarization.


【2】Towards Prosodically Informed Mizo TTS without Explicit Tone Markings
标题:迈向没有明确语气标记的韵律知情Mizo TTC
链接:https://arxiv.org/abs/2601.02073

作者:Abhijit Mohanta,Remruatpuii,Priyankoo Sarmah,Rohit Sinha,Wendy Lalhminghlui
摘要:本文报道了一个文本到语音(TTS)系统的米佐,一个低资源,音调,藏缅语主要在印度米佐拉姆邦发言的发展。TTS仅使用5.18小时的数据构建;然而,在主观和客观评价方面,输出被认为是感知上可接受和可理解的。使用Tacotron 2建立基线模型,然后使用相同的数据,使用VITS建立另一个TTS模型。在主观和客观评价中,VITS模型优于Tacotron 2模型。在音调合成方面,VITS模型显示出比Tacotron 2模型显著更低的音调误差。本文表明,一个非自回归,端到端的框架可以实现可接受的感知质量和可懂度的合成。
摘要:This paper reports on the development of a text-to-speech (TTS) system for Mizo, a low-resource, tonal, and Tibeto-Burman language spoken primarily in the Indian state of Mizoram. The TTS was built with only 5.18 hours of data; however, in terms of subjective and objective evaluations, the outputs were considered perceptually acceptable and intelligible. A baseline model using Tacotron2 was built, and then, with the same data, another TTS model was built with VITS. In both subjective and objective evaluations, the VITS model outperformed the Tacotron2 model. In terms of tone synthesis, the VITS model showed significantly lower tone errors than the Tacotron2 model. The paper demonstrates that a non-autoregressive, end-to-end framework can achieve synthesis of acceptable perceptual quality and intelligibility.


【3】MORE: Multi-Objective Adversarial Attacks on Speech Recognition
标题:更多:语音识别的多目标对抗攻击
链接:https://arxiv.org/abs/2601.01852

作者:Xiaoxue Gao,Zexin Li,Yiming Chen,Nancy F. Chen
备注:19 pages
摘要:大规模自动语音识别(ASR)模型(如Whisper)的出现大大扩展了它们在各种现实应用中的应用。因此,确保对即使是微小的输入扰动的鲁棒性对于在实时环境中保持可靠的性能至关重要。虽然之前的工作主要研究了对抗性攻击下的准确性下降,但在效率方面的鲁棒性在很大程度上仍未得到探索。这种狭隘的关注只提供了对ASR模型漏洞的部分理解。为了解决这一差距,我们进行了全面的研究,在多种攻击场景下的ASR鲁棒性。我们引入了MORE,一种多目标重复加倍鼓励攻击,通过分层的阶段性排斥锚定机制,共同降低识别精度和推理效率。具体来说,我们将多目标对抗优化重新表述为一个分层框架,依次实现双重目标。为了进一步提高有效性,我们提出了一种新的重复鼓励加倍目标(REDO),通过保持准确性下降和定期加倍预测序列长度来诱导重复文本生成。总体而言,MORE迫使ASR模型以更高的计算成本产生不正确的transmittance,由单个对抗性输入触发。实验表明,与现有基线相比,MORE始终产生显著更长的传输时间,同时保持较高的单词错误率,强调了其在多目标对抗攻击中的有效性。
摘要:The emergence of large-scale automatic speech recognition (ASR) models such as Whisper has greatly expanded their adoption across diverse real-world applications. Ensuring robustness against even minor input perturbations is therefore critical for maintaining reliable performance in real-time environments. While prior work has mainly examined accuracy degradation under adversarial attacks, robustness with respect to efficiency remains largely unexplored. This narrow focus provides only a partial understanding of ASR model vulnerabilities. To address this gap, we conduct a comprehensive study of ASR robustness under multiple attack scenarios. We introduce MORE, a multi-objective repetitive doubling encouragement attack, which jointly degrades recognition accuracy and inference efficiency through a hierarchical staged repulsion-anchoring mechanism. Specifically, we reformulate multi-objective adversarial optimization into a hierarchical framework that sequentially achieves the dual objectives. To further amplify effectiveness, we propose a novel repetitive encouragement doubling objective (REDO) that induces duplicative text generation by maintaining accuracy degradation and periodically doubling the predicted sequence length. Overall, MORE compels ASR models to produce incorrect transcriptions at a substantially higher computational cost, triggered by a single adversarial input. Experiments show that MORE consistently yields significantly longer transcriptions while maintaining high word error rates compared to existing baselines, underscoring its effectiveness in multi-objective adversarial attack.


【4】Bayesian Negative Binomial Regression of Afrobeats Chart Persistence
标题:非洲节拍图表持续性的Bayesian负二项回归
链接:https://arxiv.org/abs/2601.01391

作者:Ian Jacob Cabansag,Paul Ntegeka
摘要:Afrobeats歌曲在流媒体平台上争夺关注,在这些平台上,图表的可见性可以影响收入和文化影响。本文使用2024年以来的每日尼日利亚Spotify Top 200数据,研究合作是否有助于歌曲在排行榜上保持更长时间。每个轨道是总结了多少天,它出现在前200名在一年中,其总的年度流在尼日利亚。应用贝叶斯负二项回归,以图表上的天数作为结果,以协作状态(单人与多艺术家)和日志总流作为预测因子。这种方法非常适合过度分散的计数数据,并允许在控制整体流行度的同时解释协作的效果。使用马尔可夫链蒙特卡罗进行后验推理,并使用率比,后验概率和预测检查评估结果。调查结果表明,在考虑到总流量后,合作曲目在排行榜上的天数往往略少于可比的独奏曲目。
摘要:Afrobeats songs compete for attention on streaming platforms, where chart visibility can influence both revenue and cultural impact. This paper examines whether collaborations help songs remain on the charts longer, using daily Nigeria Spotify Top 200 data from 2024. Each track is summarized by the number of days it appears in the Top 200 during the year and its total annual streams in Nigeria. A Bayesian negative binomial regression is applied, with days on chart as the outcome and collaboration status (solo versus multi-artist) and log total streams as predictors. This approach is well suited for overdispersed count data and allows the effect of collaboration to be interpreted while controlling for overall popularity. Posterior inference is conducted using Markov chain Monte Carlo, and results are assessed using rate ratios, posterior probabilities, and predictive checks. The findings indicate that, after accounting for total streams, collaboration tracks tend to spend slightly fewer days on the chart than comparable solo tracks.


【5】Improving Code-Switching Speech Recognition with TTS Data Augmentation
标题:利用TTC数据增强改进代码转换语音识别
链接:https://arxiv.org/abs/2601.00935

作者:Yue Heng Yeo,Yuchen Hu,Shreyas Gopal,Yizhou Peng,Hexin Liu,Eng Siong Chng
备注:This paper was accepted by APSIPA 2025
摘要:由于缺乏真实的、高质量的标记语音数据,会话式语码转换语音的自动语音识别(ASR)仍然具有挑战性。本文探讨了多语言文本到语音(TTS)模型作为一种有效的数据增强技术,以解决这一不足。具体而言,我们在SEAME数据集上微调多语言CosyVoice2 TTS模型,以生成合成会话式汉英码切换语音,显着增加可用训练数据的数量和说话人多样性。我们的实验表明,用合成语音增强真实语音可以将DevMan上的混合错误率(MER)从12.1%降低到10.1%,并将DevSGE上的混合错误率从17.8%降低到16.0%,这表明了一致的性能增益。这些结果证实,多语言TTS是一个有效的和实用的工具,以提高在低资源会话代码转换的情况下,ASR的鲁棒性。
摘要:Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explores multilingual text-to-speech (TTS) models as an effective data augmentation technique to address this shortage. Specifically, we fine-tune the multilingual CosyVoice2 TTS model on the SEAME dataset to generate synthetic conversational Chinese-English code-switching speech, significantly increasing the quantity and speaker diversity of available training data. Our experiments demonstrate that augmenting real speech with synthetic speech reduces the mixed error rate (MER) from 12.1 percent to 10.1 percent on DevMan and from 17.8 percent to 16.0 percent on DevSGE, indicating consistent performance gains. These results confirm that multilingual TTS is an effective and practical tool for enhancing ASR robustness in low-resource conversational code-switching scenarios.


【6】Speak the Art: A Direct Speech to Image Generation Framework
标题:讲艺术:对图像生成框架的直接演讲
链接:https://arxiv.org/abs/2601.00827

作者:Mariam Saeed,Manar Amr,Farida Adel,Nada Hassan,Nour Walid,Eman Mohamed,Mohamed Hussein,Marwan Torki
摘要:直接语音到图像生成最近显示出有希望的结果。然而,与文本到图像的生成相比,仍然存在很大的差距。目前的方法使用两个阶段来解决这个任务:语音编码网络和图像生成对抗网络(GAN)。这些方法中的语音编码网络产生的嵌入不能捕获足够的语言信息来语义地表示输入语音。GANs存在不收敛、模式崩溃和梯度减弱等问题,分别导致模型参数不稳定、样本多样性有限和生成器学习无效。为了解决这些弱点,我们引入了一个名为\textbf{Speak the Art(STA)}的框架,该框架由语音编码网络和以语音嵌入为条件的VQ扩散网络组成。为了改进语音嵌入,语音编码网络在训练期间由大型预训练图像-文本模型监督。用扩散代替GANs可以实现更稳定的训练和生成多样化的图像。此外,我们调查的可行性,扩展我们的框架是多语言的。作为概念证明,我们用两种语言训练了我们的框架:英语和阿拉伯语。最后,我们表明,我们的结果超过了国家的最先进的模型由一个很大的保证金。
摘要:Direct speech-to-image generation has recently shown promising results. However, compared to text-to-image generation, there is still a large gap to enclose. Current approaches use two stages to tackle this task: speech encoding network and image generative adversarial network (GAN). The speech encoding networks in these approaches produce embeddings that do not capture sufficient linguistic information to semantically represent the input speech. GANs suffer from issues such as non-convergence, mode collapse, and diminished gradient, which result in unstable model parameters, limited sample diversity, and ineffective generator learning, respectively. To address these weaknesses, we introduce a framework called \textbf{Speak the Art (STA)} which consists of a speech encoding network and a VQ-Diffusion network conditioned on speech embeddings. To improve speech embeddings, the speech encoding network is supervised by a large pre-trained image-text model during training. Replacing GANs with diffusion leads to more stable training and the generation of diverse images. Additionally, we investigate the feasibility of extending our framework to be multilingual. As a proof of concept, we trained our framework with two languages: English and Arabic. Finally, we show that our results surpass state-of-the-art models by a large margin.


【7】DARC: Drum accompaniment generation with fine-grained rhythm control
标题:DARC:鼓伴奏生成与细粒度节奏控制
链接:https://arxiv.org/abs/2601.02357

作者:Trey Brosnan
摘要:在音乐创作中,快速原型对于探索和提炼想法至关重要,但当用户需要结构控制和风格灵活性时,现有的生成工具往往会出现不足。之前的茎到茎生成方法可以以其他音乐茎为条件,但对节奏的控制有限,而音色转移方法允许用户指定特定的节奏,但不能以音乐上下文为条件。我们介绍DARC,生成鼓伴奏模型,条件都从其他干音乐背景和明确的节奏提示,如beatboxing或轻敲轨道。使用参数有效的微调,我们增强了阶段,一个国家的最先进的鼓干发电机,与细粒度的节奏控制,同时保持音乐的上下文意识。
摘要:In music creation, rapid prototyping is essential for exploring and refining ideas, yet existing generative tools often fall short when users require both structural control and stylistic flexibility. Prior approaches in stem-to-stem generation can condition on other musical stems but offer limited control over rhythm, and timbre-transfer methods allow users to specify specific rhythms, but cannot condition on musical context. We introduce DARC, a generative drum accompaniment model that conditions both on musical context from other stems and explicit rhythm prompts such as beatboxing or tapping tracks. Using parameter-efficient fine-tuning, we augment STAGE, a state-of-the-art drum stem generator, with fine-grained rhythm control while maintaining musical context awareness.


【8】Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
标题:迈向多级别脚本分割:LoRA用于目录生成的微调
链接:https://arxiv.org/abs/2601.02128

作者:Steffen Freisinger,Philipp Seeberger,Thomas Ranzenberger,Tobias Bocklet,Korbinian Riedhammer
备注:Published in Proceedings of Interspeech 2025. Please cite the proceedings version (DOI: 10.21437/Interspeech.2025-2792)
摘要:将演讲稿分割成主题部分既有利于下游处理,也有利于依赖书面文本的用户。我们介绍了一种新的方法,在成绩单分层主题分割,生成多层次的内容表,捕捉主题和子主题的边界。我们在大型语言模型上比较了zero-shot提示和LoRA微调,同时还探索了高级语音暂停功能的集成。对英语会议录音和多语种演讲稿(葡萄牙语、德语)的评价显示,与既定的主题分割基线相比,有了显著改善。此外,我们采用了一个共同的评估措施,多层次分割,考虑到所有层次的一个指标。
摘要:Segmenting speech transcripts into thematic sections benefits both downstream processing and users who depend on written text for accessibility. We introduce a novel approach to hierarchical topic segmentation in transcripts, generating multi-level tables of contents that capture both topic and subtopic boundaries. We compare zero-shot prompting and LoRA fine-tuning on large language models, while also exploring the integration of high-level speech pause features. Evaluations on English meeting recordings and multilingual lecture transcripts (Portuguese, German) show significant improvements over established topic segmentation baselines. Additionally, we adapt a common evaluation measure for multi-level segmentation, taking into account all hierarchical levels within one metric.


【9】MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
标题:MM-Sonate:具有零拍摄语音克隆的多模态可控音视频生成
链接:https://arxiv.org/abs/2601.01568

作者:Chunyu Qiang,Jun Wang,Xiaopeng Wang,Kang Yin,Yuxin Guo,Xijuan Zeng,Nan Li,Zihan Li,Yuzhe Liang,Ziyu Zhang,Teng Ma,Yushen Chen,Zhongliang Liu,Feng Deng,Chen Zhang,Pengfei Wan
摘要:联合音频视频生成的目的是合成同步的多感官内容,但目前的统一模型与细粒度的声学控制,特别是身份保留语音斗争。现有的方法或者由于级联生成而遭受时间不对准,或者缺乏在联合合成框架内执行zero-shot语音克隆的能力。在这项工作中,我们提出了MM-Sonate,一个多模式流匹配框架,它将可控的音频-视频联合生成与zero-shot语音克隆功能相结合。与依赖于粗糙的语义描述的先前作品不同,MM-Sonate利用统一的发音音素输入来执行严格的语言和时间对齐。为了实现zero-shot语音克隆,我们引入了一种音色注入机制,有效地从语言内容中提取说话人身份。此外,解决标准的无分类器的指导在多模态设置的局限性,我们提出了一个基于噪声的负调节策略,利用自然噪声先验显着提高声学保真度。经验评估表明,MM-Sonate在联合生成基准中建立了新的最先进的性能,在嘴唇同步和语音清晰度方面显著优于基线,同时实现了与专门的文本到语音系统相当的语音克隆保真度。
摘要:Joint audio-video generation aims to synthesize synchronized multisensory content, yet current unified models struggle with fine-grained acoustic control, particularly for identity-preserving speech. Existing approaches either suffer from temporal misalignment due to cascaded generation or lack the capability to perform zero-shot voice cloning within a joint synthesis framework. In this work, we present MM-Sonate, a multimodal flow-matching framework that unifies controllable audio-video joint generation with zero-shot voice cloning capabilities. Unlike prior works that rely on coarse semantic descriptions, MM-Sonate utilizes a unified instruction-phoneme input to enforce strict linguistic and temporal alignment. To enable zero-shot voice cloning, we introduce a timbre injection mechanism that effectively decouples speaker identity from linguistic content. Furthermore, addressing the limitations of standard classifier-free guidance in multimodal settings, we propose a noise-based negative conditioning strategy that utilizes natural noise priors to significantly enhance acoustic fidelity. Empirical evaluations demonstrate that MM-Sonate establishes new state-of-the-art performance in joint generation benchmarks, significantly outperforming baselines in lip synchronization and speech intelligibility, while achieving voice cloning fidelity comparable to specialized Text-to-Speech systems.


【10】MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization
标题:MOSS转录日记:通过扬声器日记进行准确转录
链接:https://arxiv.org/abs/2601.01554

作者:Donghua Yu,Zhengyuan Lin,Chen Yang,Yiyang Zhang,Zhaoye Fei,Hanfu Chen,Jingqi Chen,Ke Chen,Qinyuan Cheng,Liwei Fan,Yi Jiang,Jie Zhu,Muchen Li,Shimin Li,Wenxuan Wang,Yang Wang,Zhe Xu,Yitian Gong,Yuqian Zhang
摘要:Speaker-Attributed,Time-Stamped Transcription(SATS)旨在转录发言内容并精确确定每个发言者的时间,这对于会议转录特别有价值。现有的SATS系统很少采用端到端的公式,并进一步受到有限的上下文窗口,弱远程扬声器记忆,以及无法输出时间戳的限制。为了解决这些限制,我们提出了MOSS Transcribe Diarize,一个统一的多模态大型语言模型,它在端到端的范例中联合执行说话人属性,时间戳转录。经过大量真实野生数据的训练,并配备了一个128 k的上下文窗口,最多可输入90分钟,MOSS Transcribe Diarize扩展良好,泛化能力强。在综合评估中,它在多个公共和内部基准上的表现优于最先进的商业系统。
摘要:Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meeting transcription. Existing SATS systems rarely adopt an end-to-end formulation and are further constrained by limited context windows, weak long-range speaker memory, and the inability to output timestamps. To address these limitations, we present MOSS Transcribe Diarize, a unified multimodal large language model that jointly performs Speaker-Attributed, Time-Stamped Transcription in an end-to-end paradigm. Trained on extensive real wild data and equipped with a 128k context window for up to 90-minute inputs, MOSS Transcribe Diarize scales well and generalizes robustly. Across comprehensive evaluations, it outperforms state-of-the-art commercial systems on multiple public and in-house benchmarks.


【11】Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
标题:弥合差距:Speech-LLM和多语言会话ASR端到端架构的比较探索
链接:https://arxiv.org/abs/2601.01461

作者:Yuxiang Mei,Dongxing Xu,Jiaen Liang,Yanhua Long
备注:5 pages, 1 figure
摘要:INTERSPEECH 2025多语言会话语音语言模型(MLC-SLM)挑战赛通过大型语言模型(LLM)促进多语言会话ASR。我们之前的SHNU-mASR系统采用了一种具有竞争力的并行语音编码器架构,将Whisper和mHuBERT与LLM集成在一起。然而,它面临着两个挑战:简单的特征级联可能无法充分利用互补信息,以及基于LLM的ASR和端到端(E2 E)编码器-解码器ASR之间的性能差距仍未被探索。在这项工作中,我们提出了一个增强的基于LLM的ASR框架,该框架将微调的Whisper和mHuBERT编码器与LLM相结合,以丰富语音表示。我们首先使用LoRA评估E2 E Whisper模型,并在MLC-SLM ASR任务上进行全面微调,然后为并行语音编码器提出基于交叉注意的融合机制。在MLC-SLM挑战赛的官方评估集上,我们的系统达到了10.69%的CER/WER,与排名第一的Track 1系统不相上下,尽管与他们的大规模训练集相比,它只使用了1,500小时的基线训练数据。尽管如此,我们发现,我们最终的基于LLM的ASR仍然不匹配微调E2 E耳语模型的性能,为未来的Speech-LLM设计提供了宝贵的经验指导。我们的代码可在https://github.com/1535176727/MLC-SLM上公开获取。
摘要:The INTERSPEECH 2025 Challenge on Multilingual Conversational Speech Language Models (MLC-SLM) promotes multilingual conversational ASR with large language models (LLMs). Our previous SHNU-mASR system adopted a competitive parallel-speech-encoder architecture that integrated Whisper and mHuBERT with an LLM. However, it faced two challenges: simple feature concatenation may not fully exploit complementary information, and the performance gap between LLM-based ASR and end-to-end(E2E) encoder-decoder ASR remained unexplored. In this work, we present an enhanced LLM-based ASR framework that combines fine-tuned Whisper and mHuBERT encoders with an LLM to enrich speech representations. We first evaluate E2E Whisper models with LoRA and full fine-tuning on the MLC-SLM ASR task, and then propose cross-attention-based fusion mechanisms for the parallel-speech-encoder. On the official evaluation set of the MLC-SLM Challenge, our system achieves a CER/WER of 10.69%, ranking on par with the top-ranked Track 1 systems, even though it uses only 1,500 hours of baseline training data compared with their large-scale training sets. Nonetheless, we find that our final LLM-based ASR still does not match the performance of a fine-tuned E2E Whisper model, providing valuable empirical guidance for future Speech-LLM design. Our code is publicly available at https://github.com/1535176727/MLC-SLM.


【12】OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
标题:DV-DirectTTC:迈向开放词汇指导文本到语音
链接:https://arxiv.org/abs/2601.01459

作者:Yong Ren,Jiangyan Yi,Jianhua Tao,Haiyang Sun,Zhengqi Wen,Hao Gu,Le Xu,Ye Bai
摘要:指令文本到语音(InstructTTS)利用自然语言描述作为风格提示来指导语音合成。然而,现有的InstructTTS方法主要依赖于与音频相关的标签的直接组合或其不同的改写,使得难以处理灵活的、高级别的指令。这种严格的控制对于诸如希望用描述性指令来操纵生成的内容创建者之类的用户是不够的。为了解决这些限制,我们引入OV-InstructTTS,一个新的范例开放词汇的InstructTTS。我们提出了一个全面的解决方案,包括一个新策划的数据集,OV-Speech,和一个新的推理驱动的框架。OV-Speech数据集将语音与开放词汇指令配对,每个指令都通过将高级指令与声学特征连接起来的推理过程进行增强。推理驱动的框架在合成语音之前从开放词汇指令推断情感、声学和非语言信息。评估表明,这种推理驱动的方法显着提高了遵循的保真度和语音表达。我们相信这项工作可以激发下一个用户友好的InstructTTS系统具有更强的泛化能力和现实世界的适用性。数据集和演示在我们的项目页面上公开。
摘要:Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a direct combination of audio-related labels or their diverse rephrasings, making it difficult to handle flexible, high-level instructions. Such rigid control is insufficient for users such as content creators who wish to steer generation with descriptive instructions. To address these constraints, we introduce OV-InstructTTS, a new paradigm for open-vocabulary InstructTTS. We propose a comprehensive solution comprising a newly curated dataset, OV-Speech, and a novel reasoning-driven framework. The OV-Speech dataset pairs speech with open-vocabulary instructions, each augmented with a reasoning process that connects high-level instructions to acoustic features. The reasoning-driven framework infers emotional, acoustic, and paralinguistic information from open-vocabulary instructions before synthesizing speech. Evaluations show that this reasoning-driven approach significantly improves instruction-following fidelity and speech expressiveness. We believe this work can inspire the next user-friendly InstructTTS systems with stronger generalization and real-world applicability. The dataset and demos are publicly available on our project page.


【13】SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
标题:SAFE-QAAC:通过强化学习进行端到端慢思维音频文本欺诈检测
链接:https://arxiv.org/abs/2601.01392

作者:Peidong Wang,Zhiming Ma,Xin Dai,Yongkang Liu,Shi Feng,Xiaocui Yang,Wenxing Hu,Zhihao Wang,Mingjun Pan,Li Yuan,Daling Wang
摘要:现有的欺诈检测方法主要依赖于转录的文本,遭受ASR错误和丢失关键的声学线索,如音调和环境背景。这就限制了它们对抗复杂欺骗策略的有效性。为了应对这些挑战,我们首先提出了\textbf{SAFE-QAQ},这是一个端到端的综合框架,用于基于音频的慢思维欺诈检测。首先,SAFE-QAQ框架消除了转录错误对检测性能的影响。其次,我们提出了基于规则的慢思维奖励机制,系统地引导系统通过准确捕捉细粒度的音频细节,通过分层推理过程,识别欺诈指示模式。此外,我们的框架在实时呼叫期间引入了动态风险评估框架,从而能够早期发现和预防欺诈。在TeleAntiFraud-Bench上的实验表明,SAFE-QAQ在多个关键方面(包括准确性,推理效率和实时处理能力)比现有方法有了显着的改进。SAFE-QAQ目前每天部署并分析超过70,000个呼叫,可有效地自动执行复杂的欺诈检测,减少人工工作量和财务损失。代码:https://anonymous.4open.science/r/SAFE-QAQ。
摘要:Existing fraud detection methods predominantly rely on transcribed text, suffering from ASR errors and missing crucial acoustic cues like vocal tone and environmental context. This limits their effectiveness against complex deceptive strategies. To address these challenges, we first propose \textbf{SAFE-QAQ}, an end-to-end comprehensive framework for audio-based slow-thinking fraud detection. First, the SAFE-QAQ framework eliminates the impact of transcription errors on detection performance. Secondly, we propose rule-based slow-thinking reward mechanisms that systematically guide the system to identify fraud-indicative patterns by accurately capturing fine-grained audio details, through hierarchical reasoning processes. Besides, our framework introduces a dynamic risk assessment framework during live calls, enabling early detection and prevention of fraud. Experiments on the TeleAntiFraud-Bench demonstrate that SAFE-QAQ achieves dramatic improvements over existing methods in multiple key dimensions, including accuracy, inference efficiency, and real-time processing capabilities. Currently deployed and analyzing over 70,000 calls daily, SAFE-QAQ effectively automates complex fraud detection, reducing human workload and financial losses. Code: https://anonymous.4open.science/r/SAFE-QAQ.


【14】UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
标题:UltraEval-音频:音频基础模型综合评估的统一框架
链接:https://arxiv.org/abs/2601.01373

作者:Qundong Shi,Jie Zhou,Biyuan Lin,Junbo Cui,Guoyang Zeng,Yixuan Zhou,Ziyang Wang,Xin Liu,Zhen Luo,Yudong Wang,Zhiyuan Liu
备注:13 pages, 2 figures
摘要:自GPT-4 o出现以来,音频基础模型的发展迅速加速。然而,缺乏全面的评估已经成为该领域进一步发展的关键瓶颈,特别是在音频生成方面。当前的音频评估面临三大挑战:(1)音频评估缺乏统一的框架,数据集和代码分散在各个来源,阻碍了公平和有效的跨模型比较;(2)音频编解码器作为音频基础模型的关键组成部分,缺乏广泛接受的整体评估方法;(3)现有的语音基准严重依赖于英语,这使得客观地评估模型在中文上的性能具有挑战性。为了解决第一个问题,我们引入了UltraEval-Audio,这是一个针对音频基础模型的统一评估框架,专为音频理解和生成任务而设计。UltraEval-Audio采用模块化架构,支持10种语言和14个核心任务类别,同时无缝集成24个主流模型和36个权威基准。为了提高研究效率,该框架提供了一个命令评估功能,并配有实时公共排行榜。对于第二个挑战,UltraEval-Audio采用了一种新颖的音频编解码器综合评估方案,从三个关键维度评估性能:语义准确性、音色保真度和声学质量。为了解决第三个问题,我们提出了两个新的汉语基准,SpeechCMMLU和SpeechHSK,旨在评估汉语知识水平和语言流利度。我们希望UltraEval-Audio能够为学术界和工业界提供一个透明、高效和公平的音频模型比较平台。我们的代码、基准测试和排行榜可在https://github.com/OpenBMB/UltraEval-Audio上获得。
摘要:The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particularly in audio generation. Current audio evaluation faces three major challenges: (1) audio evaluation lacks a unified framework, with datasets and code scattered across various sources, hindering fair and efficient cross-model comparison;(2) audio codecs, as a key component of audio foundation models, lack a widely accepted and holistic evaluation methodology; (3) existing speech benchmarks are heavily reliant on English, making it challenging to objectively assess models' performance on Chinese. To address the first issue, we introduce UltraEval-Audio, a unified evaluation framework for audio foundation models, specifically designed for both audio understanding and generation tasks. UltraEval-Audio features a modular architecture, supporting 10 languages and 14 core task categories, while seamlessly integrating 24 mainstream models and 36 authoritative benchmarks. To enhance research efficiency, the framework provides a one-command evaluation feature, accompanied by real-time public leaderboards. For the second challenge, UltraEval-Audio adopts a novel comprehensive evaluation scheme for audio codecs, evaluating performance across three key dimensions: semantic accuracy, timbre fidelity, and acoustic quality. To address the third issue, we propose two new Chinese benchmarks, SpeechCMMLU and SpeechHSK, designed to assess Chinese knowledge proficiency and language fluency. We wish that UltraEval-Audio will provide both academia and industry with a transparent, efficient, and fair platform for comparison of audio models. Our code, benchmarks, and leaderboards are available at https://github.com/OpenBMB/UltraEval-Audio.


【15】Diffusion Timbre Transfer Via Mutual Information Guided Inpainting
标题:基于互信息引导的扩散音色修复
链接:https://arxiv.org/abs/2601.01294

作者:Ching Ho Lee,Javier Nistal,Stefan Lattner,Marco Pasini,George Fazekas
备注:6 pages, 2 figures, 3 tables
摘要:我们研究音色转移作为一个推断时间编辑问题的音乐音频。从一个强大的预先训练的潜在扩散模型开始,我们引入了一个轻量级的过程,不需要额外的训练:(i)一个维度上的噪声注入,目标是最能提供乐器身份信息的潜在通道,以及(ii)一个早期的箝位机制,在反向扩散过程中重新施加输入的旋律和节奏结构。该方法直接对音频潜伏期进行操作并且与文本/音频调节(例如,CLAP)。我们讨论了设计选择,分析了音色变化和结构保留之间的权衡,并表明简单的推理时间控制可以有意义地引导预训练模型用于风格转换用例。
摘要:We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a dimension-wise noise injection that targets latent channels most informative of instrument identity, and (ii) an early-step clamping mechanism that re-imposes the input's melodic and rhythmic structure during reverse diffusion. The method operates directly on audio latents and is compatible with text/audio conditioning (e.g., CLAP). We discuss design choices,analyze trade-offs between timbral change and structural preservation, and show that simple inference-time controls can meaningfully steer pre-trained models for style-transfer use cases.


【16】IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection
标题:IO-RAE:音频隐私保护的信息混淆可逆对抗示例
链接:https://arxiv.org/abs/2601.01239

作者:Jiajie Zhu,Xia Du,Xiaoyuan Liu,Jizhe Zhou,Qizhen Xu,Zheng Lin,Chi-Man Pun
备注:10 pages, 5 figures
摘要:人工智能的快速发展大大加速了语音识别技术的采用,导致其在各种应用中的广泛集成。然而,这种使用量的激增也凸显了一个关键问题:音频数据极易受到未经授权的暴露和分析,给企业和个人带来重大的隐私风险。本文介绍了一个信息混淆可逆对抗示例(IO-RAE)框架,这是一种开创性的方法,旨在使用可逆对抗示例来保护音频隐私。IO-RAE利用大型语言模型生成误导性但上下文一致的内容,有效防止人类和自动语音识别(ASR)系统进行未经授权的窃听。此外,我们还提出了累积信号攻击技术,该技术通过针对低频信号来减轻高频噪声并提高攻击效率。我们的方法确保了音频数据的保护,而不会降低其质量或我们的能力。实验评估表明,我们的方法的优越性,实现了96.5%的有针对性的误导率和显着的100%的非有针对性的误导率混淆目标关键字在多个ASR模型,包括商业黑盒系统从谷歌。此外,通过语音质量的感知评估得分来衡量,恢复的音频质量达到4.45,与高质量的原始录音相当。值得注意的是,ASR系统处理的恢复音频显示出0%的错误率,这表明几乎无损恢复。这些结果突出了我们的IO-RAE框架在保护敏感音频隐私方面的实用性和有效性。
摘要:The rapid advancements in artificial intelligence have significantly accelerated the adoption of speech recognition technology, leading to its widespread integration across various applications. However, this surge in usage also highlights a critical issue: audio data is highly vulnerable to unauthorized exposure and analysis, posing significant privacy risks for businesses and individuals. This paper introduces an Information-Obfuscation Reversible Adversarial Example (IO-RAE) framework, the pioneering method designed to safeguard audio privacy using reversible adversarial examples. IO-RAE leverages large language models to generate misleading yet contextually coherent content, effectively preventing unauthorized eavesdropping by humans and Automatic Speech Recognition (ASR) systems. Additionally, we propose the Cumulative Signal Attack technique, which mitigates high-frequency noise and enhances attack efficacy by targeting low-frequency signals. Our approach ensures the protection of audio data without degrading its quality or our ability. Experimental evaluations demonstrate the superiority of our method, achieving a targeted misguidance rate of 96.5% and a remarkable 100% untargeted misguidance rate in obfuscating target keywords across multiple ASR models, including a commercial black-box system from Google. Furthermore, the quality of the recovered audio, measured by the Perceptual Evaluation of Speech Quality score, reached 4.45, comparable to high-quality original recordings. Notably, the recovered audio processed by ASR systems exhibited an error rate of 0%, indicating nearly lossless recovery. These results highlight the practical applicability and effectiveness of our IO-RAE framework in protecting sensitive audio privacy.


【17】Index-ASR Technical Report
标题:指数-ASB技术报告
链接:https://arxiv.org/abs/2601.00890

作者:Zheshu Song,Lu Wang,Wei Deng,Zhuo Yang,Yong Wu,Bin Xia
备注:Index-ASR technical report
摘要:自动语音识别(ASR)近年来取得了显著的进展,主要是由于基于LLM的ASR范式的出现。尽管它们在各种开源基准测试中表现出色,但现有的基于LLM的ASR系统仍然受到两个关键限制。首先,它们容易产生幻觉错误,经常产生过长和重复的输出,这些输出在声学输入中没有很好的基础。其次,它们对灵活和细粒度的上下文定制提供了有限的支持。为了应对这些挑战,我们提出了Index-ASR,这是一个基于LLM的大规模ASR系统,旨在同时增强鲁棒性并支持可定制的热词识别。Index-ASR的核心思想在于将LLM与富含背景噪声和上下文信息的大规模训练数据相结合。实验结果表明,我们的Index-ASR在开源基准测试和内部测试集上都取得了很好的性能,突出了它对现实世界ASR应用的鲁棒性和实用性。
摘要:Automatic speech recognition (ASR) has witnessed remarkable progress in recent years, largely driven by the emergence of LLM-based ASR paradigm. Despite their strong performance on a variety of open-source benchmarks, existing LLM-based ASR systems still suffer from two critical limitations. First, they are prone to hallucination errors, often generating excessively long and repetitive outputs that are not well grounded in the acoustic input. Second, they provide limited support for flexible and fine-grained contextual customization. To address these challenges, we propose Index-ASR, a large-scale LLM-based ASR system designed to simultaneously enhance robustness and support customizable hotword recognition. The core idea of Index-ASR lies in the integration of LLM and large-scale training data enriched with background noise and contextual information. Experimental results show that our Index-ASR achieves strong performance on both open-source benchmarks and in-house test sets, highlighting its robustness and practicality for real-world ASR applications.


机器翻译由腾讯交互翻译提供,仅供参考