今日论文合集:cs.SD语音4篇,eess.AS音频处理3篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】PromptSep: Generative Audio Separation via Multimodal Prompting
标题:格式Sep:通过多模式格式生成音频分离
链接:https://arxiv.org/abs/2511.04623

作者:Yutong Wen, Ke Chen, Prem Seetharaman, Oriol Nieto, Jiaqi Su, Rithesh Kumar, Minje Kim, Paris Smaragdis, Zeyu Jin, Justin Salamon
备注:Submitted to ICASSP 2026
摘要:语言查询音频源分离(LASS)的最新突破表明,生成模型可以实现比传统的基于掩蔽的方法更高的分离音频质量。然而,两个关键限制限制了它们的实际用途:(1)用户通常需要除分离之外的操作,例如声音去除;以及(2)仅依赖于文本提示对于指定声源可能是不直观的。在本文中,我们提出了一个更广泛的框架扩展LASS到通用的声音分离。AptSep利用经过精心设计的数据模拟增强的条件扩散模型来实现音频提取和声音去除。为了超越纯文本查询,我们通过将Sketch 2Sound作为数据增强策略,将声音模仿作为我们模型的附加和更直观的条件反射方式。多个基准测试的客观和主观评估都表明,MixtSep在声音去除和人声模仿引导的源分离方面达到了最先进的性能,同时在语言查询源分离方面保持了竞争力。
摘要:Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separation, such as sound removal; and (2) relying solely on text prompts can be unintuitive for specifying sound sources. In this paper, we propose PromptSep to extend LASS into a broader framework for general-purpose sound separation. PromptSep leverages a conditional diffusion model enhanced with elaborated data simulation to enable both audio extraction and sound removal. To move beyond text-only queries, we incorporate vocal imitation as an additional and more intuitive conditioning modality for our model, by incorporating Sketch2Sound as a data augmentation strategy. Both objective and subjective evaluations on multiple benchmarks demonstrate that PromptSep achieves state-of-the-art performance in sound removal and vocal-imitation-guided source separation, while maintaining competitive results on language-queried source separation.


【2】MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
标题:MusRec:通过Rectified Flow和Disffusion Transformers进行Zero-Shot文本到音乐编辑
链接:https://arxiv.org/abs/2511.04376

作者:Ali Boudaghi, Hadi Zare
摘要:音乐编辑已成为人工智能的一个重要和实用的领域,其应用范围从视频游戏和电影音乐制作到根据用户偏好个性化现有曲目。然而,现有的模型面临着很大的局限性,例如被限制为编辑由它们自己的模型生成的合成音乐,需要高度精确的提示,或者需要特定于任务的再训练,因此缺乏真正的zero-shot能力。利用整流和扩散Transformers的最新进展,我们推出了MusRec,这是第一个zero-shot文本到音乐编辑模型,能够高效地对真实世界的音乐执行各种编辑任务。实验结果表明,该方法在保留音乐内容、结构一致性和编辑保真度方面优于现有方法,为现实场景中的可控音乐编辑奠定了坚实的基础。
摘要:Music editing has emerged as an important and practical area of artificial intelligence, with applications ranging from video game and film music production to personalizing existing tracks according to user preferences. However, existing models face significant limitations, such as being restricted to editing synthesized music generated by their own models, requiring highly precise prompts, or necessitating task-specific retraining, thus lacking true zero-shot capability. Leveraging recent advances in rectified flow and diffusion transformers, we introduce MusRec, the first zero-shot text-to-music editing model capable of performing diverse editing tasks on real-world music efficiently and effectively. Experimental results demonstrate that our approach outperforms existing methods in preserving musical content, structural consistency, and editing fidelity, establishing a strong foundation for controllable music editing in real-world scenarios.


【3】CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
标题:CantoASB:针对低资源粤语的具有韵律意识的ASR-LALM合作
链接:https://arxiv.org/abs/2511.04139

作者:Dazhong Chen, Yi-Cheng Lin, Yuchen Huang, Ziwei Gong, Di Jiang, Zeying Xie, Yi R. (May)Fung
摘要:自动语音识别(ASR)对于语言的可访问性至关重要,但是由于有限的注释数据、六个词汇声调、连读变调和口音变化,低资源粤语仍然具有挑战性。现有的ASR模型,如Whisper,往往遭受高单词错误率。相比之下,大型音频语言模型(LALM)可以利用更广泛的上下文推理,但仍然需要明确的音调和韵律声学线索。我们介绍了CantoASR,这是一个协作的ASR-LALM纠错框架,它集成了用于声学特征提取的强制对齐,用于改进音调辨别的LoRA微调Whisper,以及用于韵律感知校正的调整Qwen音频。对自发广东话数据的评估显示,与Whisper-Large-V3相比,CER大幅增加。这些研究结果表明,整合声学线索与LALM推理提供了一个可扩展的策略,低资源音调和方言ASR。
摘要:Automatic speech recognition (ASR) is critical for language accessibility, yet low-resource Cantonese remains challenging due to limited annotated data, six lexical tones, tone sandhi, and accent variation. Existing ASR models, such as Whisper, often suffer from high word error rates. Large audio-language models (LALMs), in contrast, can leverage broader contextual reasoning but still require explicit tonal and prosodic acoustic cues. We introduce CantoASR, a collaborative ASR-LALM error correction framework that integrates forced alignment for acoustic feature extraction, a LoRA-finetuned Whisper for improved tone discrimination, and an instruction-tuned Qwen-Audio for prosody-aware correction. Evaluations on spontaneous Cantonese data show substantial CER gains over Whisper-Large-V3. These findings suggest that integrating acoustic cues with LALM reasoning provides a scalable strategy for low-resource tonal and dialectal ASR.


【4】MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
标题:MIDI-LLM:调整大型语言模型以实现文本到格式的音乐生成
链接:https://arxiv.org/abs/2511.03942

作者:Shih-Lun Wu, Yoon Kim, Cheng-Zhi Anna Huang
备注:To appear at NeurIPS 2025 Workshop on AI for Music
摘要:我们提出了MIDI-LLM,LLM用于从自由形式的文本提示生成多轨音乐。我们的方法扩展了文本LLM的词汇表,包括文本标记,并使用两个阶段的训练配方赋予文本的可扩展性。通过保留原始LLM的参数结构,我们可以直接利用vLLM库来加速推理。实验表明,MIDI-LLM实现了更高的质量,更好的文本控制,更快的推理相比,最近的Text 2 midi模型。在https://midi-llm-demo.vercel.app现场演示。
摘要:We present MIDI-LLM, an LLM for generating multitrack MIDI music from free-form text prompts. Our approach expands a text LLM's vocabulary to include MIDI tokens, and uses a two-stage training recipe to endow text-to-MIDI abilities. By preserving the original LLM's parameter structure, we can directly leverage the vLLM library for accelerated inference. Experiments show that MIDI-LLM achieves higher quality, better text control, and faster inference compared to the recent Text2midi model. Live demo at https://midi-llm-demo.vercel.app.


eess.AS音频处理


【1】CardioPHON: Quality assessment and self-supervised pretraining for screening of cardiac function based on phonocardiogram recordings
标题:CLARPHON:基于音素心电图记录筛查心功能的质量评估和自我监督预训练
链接:https://arxiv.org/abs/2511.04533

作者:Vladimir Despotovic, Peter Pocta, Andrej Zgank
备注:None
摘要:心血管疾病的远程监测在早期发现心脏功能异常、及时干预、改善预防护理和个性化患者治疗方面发挥着重要作用。心音异常可以通过计算机辅助决策支持系统自动检测,并用作检测心血管问题或监测治疗和干预效果的一线筛查工具。我们在本文中提出了一个集成的心音质量评估和分类工具,可用于筛选从心音图记录异常的心脏功能。该模型以自我监督的方式在六个中小型心音数据集的集合上进行预训练,能够自动删除低质量的录音,以确保心脏异常的细微声音不会被误诊,并为心音分类任务提供最先进的性能。结合了音频和社会人口统计特征的多模式模型表现出卓越的性能,在2022 George B的官方排行榜上获得最佳排名。Moody PhysioNet心音挑战,而单模态模型,即仅基于心音图记录,持有单模态方法中的第一位(总排名4),超过了利用多模态的模型。PHOSPON是心音记录领域第一个公开发布的预训练模型,有助于开发数据高效的人工智能模型,这些模型可以推广到心血管诊断中的各种下游任务。
摘要:Remote monitoring of cardiovascular diseases plays an essential role in early detection of abnormal cardiac function, enabling timely intervention, improved preventive care, and personalized patient treatment. Abnormalities in the heart sounds can be detected automatically via computer-assisted decision support systems, and used as the first-line screening tool for detection of cardiovascular problems, or for monitoring the effects of treatments and interventions. We propose in this paper CardioPHON, an integrated heart sound quality assessment and classification tool that can be used for screening of abnormal cardiac function from phonocardiogram recordings. The model is pretrained in a self-supervised fashion on a collection of six small- and mid-sized heart sound datasets, enables automatic removal of low quality recordings to ensure that subtle sounds of heart abnormalities are not misdiagnosed, and provides a state-of-the-art performance for the heart sound classification task. The multimodal model that combines audio and socio-demographic features demonstrated superior performance, achieving the best ranking on the official leaderboard of the 2022 George B. Moody PhysioNet heart sound challenge, whereas the unimodal model, that is based only on phonocardiogram recordings, holds the first position among the unimodal approaches (a total rank 4), surpassing the models utilizing multiple modalities. CardioPHON is the first publicly released pretrained model in the domain of heart sound recordings, facilitating the development of data-efficient artificial intelligence models that can generalize to various downstream tasks in cardiovascular diagnostics.


【2】PromptSep: Generative Audio Separation via Multimodal Prompting
标题:格式Sep:通过多模式格式生成音频分离
链接:https://arxiv.org/abs/2511.04623

作者:Yutong Wen, Ke Chen, Prem Seetharaman, Oriol Nieto, Jiaqi Su, Rithesh Kumar, Minje Kim, Paris Smaragdis, Zeyu Jin, Justin Salamon
备注:Submitted to ICASSP 2026
摘要:语言查询音频源分离(LASS)的最新突破表明,生成模型可以实现比传统的基于掩蔽的方法更高的分离音频质量。然而,两个关键限制限制了它们的实际用途:(1)用户通常需要除分离之外的操作,例如声音去除;以及(2)仅依赖于文本提示对于指定声源可能是不直观的。在本文中,我们提出了一个更广泛的框架扩展LASS到通用的声音分离。AptSep利用经过精心设计的数据模拟增强的条件扩散模型来实现音频提取和声音去除。为了超越纯文本查询,我们通过将Sketch 2Sound作为数据增强策略,将声音模仿作为我们模型的附加和更直观的条件反射方式。多个基准测试的客观和主观评估都表明,MixtSep在声音去除和人声模仿引导的源分离方面达到了最先进的性能,同时在语言查询源分离方面保持了竞争力。
摘要:Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separation, such as sound removal; and (2) relying solely on text prompts can be unintuitive for specifying sound sources. In this paper, we propose PromptSep to extend LASS into a broader framework for general-purpose sound separation. PromptSep leverages a conditional diffusion model enhanced with elaborated data simulation to enable both audio extraction and sound removal. To move beyond text-only queries, we incorporate vocal imitation as an additional and more intuitive conditioning modality for our model, by incorporating Sketch2Sound as a data augmentation strategy. Both objective and subjective evaluations on multiple benchmarks demonstrate that PromptSep achieves state-of-the-art performance in sound removal and vocal-imitation-guided source separation, while maintaining competitive results on language-queried source separation.


【3】MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
标题:MusRec:通过Rectified Flow和Disffusion Transformers进行Zero-Shot文本到音乐编辑
链接:https://arxiv.org/abs/2511.04376

作者:Ali Boudaghi, Hadi Zare
摘要:音乐编辑已成为人工智能的一个重要和实用的领域,其应用范围从视频游戏和电影音乐制作到根据用户偏好个性化现有曲目。然而,现有的模型面临着很大的局限性,例如被限制为编辑由它们自己的模型生成的合成音乐,需要高度精确的提示,或者需要特定于任务的再训练,因此缺乏真正的zero-shot能力。利用整流和扩散Transformers的最新进展,我们推出了MusRec,这是第一个zero-shot文本到音乐编辑模型,能够高效地对真实世界的音乐执行各种编辑任务。实验结果表明,我们的方法在保留音乐内容、结构一致性和编辑保真度方面优于现有方法,为现实场景中的可控音乐编辑奠定了坚实的基础。
摘要:Music editing has emerged as an important and practical area of artificial intelligence, with applications ranging from video game and film music production to personalizing existing tracks according to user preferences. However, existing models face significant limitations, such as being restricted to editing synthesized music generated by their own models, requiring highly precise prompts, or necessitating task-specific retraining, thus lacking true zero-shot capability. Leveraging recent advances in rectified flow and diffusion transformers, we introduce MusRec, the first zero-shot text-to-music editing model capable of performing diverse editing tasks on real-world music efficiently and effectively. Experimental results demonstrate that our approach outperforms existing methods in preserving musical content, structural consistency, and editing fidelity, establishing a strong foundation for controllable music editing in real-world scenarios.


机器翻译由腾讯交互翻译提供,仅供参考