微信公众号:arXiv_Daily
cs.SD语音
【1】Contrastive timbre representations for musical instrument and synthesizer retrieval
标题:用于乐器和合成器检索的对比音色表示
链接:https://arxiv.org/abs/2509.13285
摘要:从音频混合中有效地检索特定乐器音色仍然是数字音乐制作中的一个挑战。本文介绍了一种乐器检索的对比学习框架,使仪器数据库的直接查询,使用单一的模型,为单一和多乐器的声音。我们提出的技术,以产生现实的积极/消极的声音对虚拟乐器,如采样器和合成器,解决常见的音频数据增强方法的局限性。 第一个实验的重点是从3,884个乐器的数据集中检索乐器,使用单乐器音频作为输入。对比的方法是有竞争力的,与以前的作品的基础上分类预训练。第二个实验考虑了多乐器检索与混合的文书作为音频输入。在这种情况下,所提出的对比框架优于相关的工作,实现了81.7%的前1和95.7%的前5的准确度为三仪器的混合物。
摘要:Efficiently retrieving specific instrument timbres from audio mixtures remains a challenge in digital music production. This paper introduces a contrastive learning framework for musical instrument retrieval, enabling direct querying of instrument databases using a single model for both single- and multi-instrument sounds. We propose techniques to generate realistic positive/negative pairs of sounds for virtual musical instruments, such as samplers and synthesizers, addressing limitations in common audio data augmentation methods. The first experiment focuses on instrument retrieval from a dataset of 3,884 instruments, using single-instrument audio as input. Contrastive approaches are competitive with previous works based on classification pre-training. The second experiment considers multi-instrument retrieval with a mixture of instruments as audio input. In this case, the proposed contrastive framework outperforms related works, achieving 81.7\% top-1 and 95.7\% top-5 accuracies for three-instrument mixtures.
【2】Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
标题:大型音频语言模型能很好地理解音频吗?演讲、场景和活动了解LALM的基准
链接:https://arxiv.org/abs/2509.13148
备注:submitted to ICASSP 2026
摘要:最近,大型音频语言模型(LALM)发展迅速,通过跨模态集成在通用音频理解方面表现出强大的功效。为了评估LALM的音频理解性能,研究人员提出了不同的基准。然而,在现有的基准中,现实世界交互的关键方面没有得到充分的探索,即,音频信号通常包含语音和非语音分量,并且这些分量的能级在不同的场景中可能会有显着变化。此外,大多数基准测试不考虑同一音频片段内的语音、场景和事件的联合理解。在这项工作中,我们介绍了SSEU-Bench,这是第一个多功能的音频理解基准,它明确地解释了语音和非语音音频之间的能量差异,并为语音,场景和事件提供了独立和联合的理解设置。此外,我们表明,一些LALM往往表现不佳的某些任务,在联合理解设置。为了解决这个问题,我们引入了Chain-of-Thought,它通过将复杂的任务分解为更简单的推理步骤,有效地提高了LALM的联合音频理解性能
摘要:Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To evaluate the LALM's audio understanding performance, researchers have proposed different benchmarks. However, key aspects for real-world interactions are underexplored in existing benchmarks, i.e., audio signals typically contain both speech and non-speech components, and energy levels of these components can vary significantly across different scenarios. Moreover, most benchmarks do not consider the joint understanding of speech, scene, and events within the same audio clip. In this work, we introduce SSEU-Bench, the first versatile audio understanding benchmark that explicitly accounts for energy differences between speech and non-speech audio, with both independent and joint understanding settings for speech, scene, and events. Furthermore, we demonstrate that some LALMs tend to underperform on certain tasks in a joint understanding setting. To address this issue, we introduce Chain-of-Thought, which effectively improves the LALM's joint audio understanding performance by decomposing complex tasks into simpler reasoning steps
【3】UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
标题:UTI-LLM:基于多模式大语言模型的个性化关节言语治疗辅助系统
链接:https://arxiv.org/abs/2509.13145
摘要:言语治疗在训练由神经损伤如中风引起的言语障碍中起着关键作用。然而,传统的手动和计算机辅助系统在实时可访问性和发音运动反馈方面受到限制,限制了它们的实际效用。多模态大语言模型(MLLM)的最新进展在医疗保健领域表现出了巨大的潜力,特别是通过它们整合多模态数据进行自适应评估和治疗反馈的能力。然而,包括发音信息的获取和融合不足,发音器官运动轨迹的解析不足,以及缺乏高质量的特定领域数据集的挑战阻碍了MLLM在言语治疗中的应用。为了解决这些限制,我们提出了一个基于MLLM的语音康复辅助系统,协同利用超声舌头成像和语音信号,提供精确的,交互式的发音反馈。我们构建了一个高质量的特定领域的数据集,包括UTI语音对话对。该数据集有助于进行微调,以增强模型的临床适应性。在此数据集的基础上,我们的方法实现了超声视频和语音信号的时空融合训练策略,实现了细粒度的发音障碍分析,并最终生成可操作的反馈。
摘要:Speech therapy plays a critical role in training speech disorders caused by neurological impairments such as stroke. However, traditional manual and computer-assisted systems are limited in real-time accessibility and articulatory motion feedback, constraining their practical utility. Recent advances in multimodal large language models (MLLMs) have demonstrated significant potential in healthcare, particularly through their ability to integrate multimodal data for adaptive assessment and therapeutic feedback. Nevertheless, challenges including insufficient acquisition and fusion of articulatory information, inadequate parsing of articulatory organ motion trajectories, and the scarcity of high-quality domain-specific datasets hinder the application of MLLMs in speech therapy. To address these limitations, we propose an MLLM-based speech rehabilitation assistance system that synergistically leverages ultrasound tongue imaging and speech signals to deliver precise, interactive articulatory feedback. We construct a high-quality domain-specific dataset comprising UTI-speech dialogue pairs. This dataset facilitates fine-tuning to enhance the model's clinical adaptability. Building on this dataset, our methods achieves spatiotemporal fusion training strategy of ultrasound videos and speech signals, enabling fine-grained articulatory impairment analysis and ultimately generating actionable feedback.
【4】GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
标题:GRAD:针对多说话者ASB的全球本地感知动态专家混合
链接:https://arxiv.org/abs/2509.13093
摘要:端到端多说话人自动语音识别(MTASR)在准确转录重叠语音方面面临着重大挑战,特别是在高重叠条件下。为了应对这些挑战,我们提出了全局本地感知动态(GLAD)混合专家,动态融合说话人感知的全局信息和细粒度的本地功能,以指导专家选择。该机制通过利用全局上下文和局部声学线索来实现特定于说话者的路由。在LibriSpeechMix上的实验表明,GLAD优于现有的MTASR方法,特别是在具有挑战性的多说话者场景中。据我们所知,这是第一项将混合专家(MoE)应用于具有全局-局部融合策略的端到端MTASR的工作。我们的代码和训练数据集可以在https://github.com/NKU-HLT/GLAD上找到。
摘要:End-to-end multi-talker automatic speech recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, especially under high-overlap conditions. To address these challenges, we proposed Global-Local Aware Dynamic (GLAD) Mixture-of-Experts, which dynamically fuse speaker-aware global information and fine-grained local features to guide expert selection. This mechanism enables speaker-specific routing by leveraging both global context and local acoustic cues. Experiments on LibriSpeechMix show that GLAD outperforms existing MTASR approaches, particularly in challenging multi-talker scenarios. To our best knowledge, this is the first work to apply Mixture-of-Experts (MoE) to end-to-end MTASR with a global-local fusion strategy. Our code and train dataset can be found at https://github.com/NKU-HLT/GLAD.
【5】The CCF AATC 2025: Speech Restoration Challenge
标题:CPD AATC 2025:言语恢复挑战
链接:https://arxiv.org/abs/2509.12974
备注:Technical Report
摘要:真实世界的语音通信经常受到各种失真的阻碍,这些失真降低了质量和可懂度。虽然许多语音增强算法针对特定的退化,如噪声或混响,但它们通常在多个失真共存和相互作用的现实场景中不足。为了促进这一领域的研究,我们推出了语音恢复挑战赛,作为2025年中国计算机联合会(CCF)高级音频技术竞赛(AATC)的一部分。这一挑战的重点是恢复受三种退化类型复合影响的语音信号:(1)复杂的声学退化,包括非平稳噪声和混响;(2)信号链伪影,如MP3压缩;以及(3)由其他预处理增强模型引入的次级伪影。我们描述了挑战的背景,任务的设计,全面的数据集创建方法,以及详细的评估协议,评估目标性能和模型复杂性。主页:https://ccf-aatc.org.cn/。
摘要:Real-world speech communication is often hampered by a variety of distortions that degrade quality and intelligibility. While many speech enhancement algorithms target specific degradations like noise or reverberation, they often fall short in realistic scenarios where multiple distortions co-exist and interact. To spur research in this area, we introduce the Speech Restoration Challenge as part of the China Computer Federation (CCF) Advanced Audio Technology Competition (AATC) 2025. This challenge focuses on restoring speech signals affected by a composite of three degradation types: (1) complex acoustic degradations including non-stationary noise and reverberation; (2) signal-chain artifacts such as those from MP3 compression; and (3) secondary artifacts introduced by other pre-processing enhancement models. We describe the challenge's background, the design of the task, the comprehensive dataset creation methodology, and the detailed evaluation protocol, which assesses both objective performance and model complexity. Homepage: https://ccf-aatc.org.cn/.
【6】Improving Anomalous Sound Detection with Attribute-aware Representation from Domain-adaptive Pre-training
标题:利用域自适应预训练的属性感知表示改进异常声音检测
链接:https://arxiv.org/abs/2509.12845
备注:5 pages, 3 figures
摘要:异常声音检测(ASD)通常被制定为机器属性分类任务,这是一种只有正常数据可用于训练的常见场景所必需的策略。然而,机器属性标签的穷举收集是费力且不切实际的。为了解决缺少属性标签的挑战,本文提出了一种凝聚层次聚类方法,用于使用来自域自适应预训练模型的表示来分配伪属性标签,该表示有望捕获机器属性特征。然后,我们通过对机器属性分类进行监督微调,将模型自适应应用于这个预先训练好的模型,从而获得了新的最先进的性能。对声学场景和事件的检测和分类(DCASE)2025挑战数据集的评估表明,我们提出的方法产生了显着的性能提升,最终在挑战中超越了我们之前的顶级系统。
摘要:Anomalous Sound Detection (ASD) is often formulated as a machine attribute classification task, a strategy necessitated by the common scenario where only normal data is available for training. However, the exhaustive collection of machine attribute labels is laborious and impractical. To address the challenge of missing attribute labels, this paper proposes an agglomerative hierarchical clustering method for the assignment of pseudo-attribute labels using representations derived from a domain-adaptive pre-trained model, which are expected to capture machine attribute characteristics. We then apply model adaptation to this pre-trained model through supervised fine-tuning for machine attribute classification, resulting in a new state-of-the-art performance. Evaluation on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge dataset demonstrates that our proposed approach yields significant performance gains, ultimately outperforming our previous top-ranking system in the challenge.
【7】A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
标题:用于高噪语音语音克隆和准确唇同步合成的轻量级管道
链接:https://arxiv.org/abs/2509.12831
摘要:最近的发展,在语音克隆和说话的头部生成显示出令人印象深刻的能力,在合成自然的语音和逼真的嘴唇同步。当前的方法通常需要并在大规模数据集和计算密集型过程上进行训练,使用在嘈杂或低资源环境中不可行的干净工作室记录的输入。在本文中,我们介绍了一个新的模块化管道,包括Torrent文本语音。它是一个基于Transformer的潜在扩散模型,可以执行高保真zero shot的语音克隆,只需几个训练样本。我们使用一个轻量级的生成式对抗网络架构来实现鲁棒的实时嘴唇同步。该解决方案将有助于许多基本任务,这些任务涉及在嘈杂和不受约束的情况下减少对情感表达语音和嘴唇同步的大规模预训练生成的依赖。流水线的模块化结构允许未来的多模态和文本引导的语音调制的容易的扩展,并且它可以用于现实世界的系统中。
摘要:Recent developments in voice cloning and talking head generation demonstrate impressive capabilities in synthesizing natural speech and realistic lip synchronization. Current methods typically require and are trained on large scale datasets and computationally intensive processes using clean studio recorded inputs that is infeasible in noisy or low resource environments. In this paper, we introduce a new modular pipeline comprising Tortoise text to speech. It is a transformer based latent diffusion model that can perform high fidelity zero shot voice cloning given only a few training samples. We use a lightweight generative adversarial network architecture for robust real time lip synchronization. The solution will contribute to many essential tasks concerning less reliance on massive pre training generation of emotionally expressive speech and lip synchronization in noisy and unconstrained scenarios. The modular structure of the pipeline allows an easy extension for future multi modal and text guided voice modulation and it could be used in real world systems.
【8】Beyond Bars: Distribution of Edit Operations in Historical Prints
标题:超越酒吧:历史版画中编辑操作的分布
链接:https://arxiv.org/abs/2509.12786
摘要:在本文中,我们提出了一种方法进行比较语料库研究的音乐学,减少了耗时的数字化过程。而不是编码整个语料库的音乐来源,我们建议采样酒吧从这些来源。我们解决了选择代表性样本的挑战,并评估三种不同的采样方法。我们使用贝多芬的Bagatelles Op.33作为一个案例研究,以找到最好的方法,在寻找样本的代表性方面的差异。我们相信,这种方法提供了显着的价值,音乐学研究,使大规模的分析,从而统计健全的结果。此外,我们相信我们的工作是一个宝贵的一步,了解19世纪的编辑实践和丰富的历史音乐作品的学术编辑领域。
摘要:In this paper, we present a method for conducting comparative corpus studies in musicology that reduces the time-consuming digitization process. Instead of encoding whole corpora of musical sources, we suggest sampling bars from these sources. We address the challenge of selecting representative samples and evaluate three different sampling methods. We used Beethoven's Bagatelles Op. 33 as a case study to find the method that works best in finding samples representative with respect to differences. We believe that this approach offers significant value to musicological research by enabling large-scale analyses and thereby statistically sound results. Moreover, we believe our work to be a valuable step toward understanding nineteenth-century editorial practices and enriching the field of scholarly editing of historical musical works.
【9】Timbre-Adaptive Transcription: A Lightweight Architecture with Associative Memory for Dynamic Instrument Separation
标题:音色自适应转录:一种具有关联记忆的轻量级架构,用于动态乐器分离
链接:https://arxiv.org/abs/2509.12712
摘要:现有的多音色转录模型难以超越预先训练的乐器和严格的源计数约束。我们通过一个轻量级的深度聚类解决方案来解决这些限制,该解决方案具有以下特点:1)一个与音色无关的骨干,仅用可比模型的一半参数实现最先进的性能,以及2)一种新型的联想记忆机制,该机制模仿人类听觉认知,通过基于注意力的聚类来动态编码看不见的音色。我们的生物启发框架可以用最少的训练数据(12.5分钟)实现自适应复调分离,并由一种新的合成数据集方法提供支持,提供具有成本效益的高精度多音色生成。实验表明,音色不可知的转录模型优于现有的模型在公共基准测试,而分离模块表现出良好的音色歧视。这项工作为音色相关的音乐转录提供了一个有效的框架,并通过认知启发的架构探索了音色感知分离的新方向。
摘要:Existing multi-timbre transcription models struggle with generalization beyond pre-trained instruments and rigid source-count constraints. We address these limitations with a lightweight deep clustering solution featuring: 1) a timbre-agnostic backbone achieving state-of-the-art performance with only half the parameters of comparable models, and 2) a novel associative memory mechanism that mimics human auditory cognition to dynamically encode unseen timbres via attention-based clustering. Our biologically-inspired framework enables adaptive polyphonic separation with minimal training data (12.5 minutes), supported by a new synthetic dataset method offering cost-effective, high-precision multi-timbre generation. Experiments show the timbre-agnostic transcription model outperforms existing models on public benchmarks, while the separation module demonstrates promising timbre discrimination. This work provides an efficient framework for timbre-related music transcription and explores new directions for timbre-aware separation through cognitive-inspired architectures.
【10】Osu2MIR: Beat Tracking Dataset Derived From Osu! Data
标题:Osu 2 MIR:源自Osu的节拍跟踪数据集!数据
链接:https://arxiv.org/abs/2509.12667
备注:2 pages
摘要:在这项工作中,我们探索使用Osu!,基于社区的节奏游戏,作为节拍和强拍注释的替代来源。大须beatmaps是由一个庞大的、多样化的社区创建和完善的,跨越了代表性不足的流派,如动漫、Vocaloid和视频游戏音乐。我们引入了一个从Osu中提取注释的管道!beatmaps并将其划分为有意义的子集。通过手动分析,我们发现,具有单个时间点或间隔较远的多个时间点(间隔>=5秒)的节拍图提供了可靠的注释,而间隔较近的时间点(间隔<5秒)通常需要额外的策展。我们还观察到同一首歌的多个注释之间的高度一致性。这项研究证明了Osu的潜力!数据作为一个可扩展的,多样化的,社区驱动的资源MIR研究。我们发布了我们的管道和一个高质量的子集osu 2beat 2025,以支持进一步的探索:https://github.com/ziyunliu4444/osu2mir。
摘要:In this work, we explore the use of Osu!, a community-based rhythm game, as an alternative source of beat and downbeat annotations. Osu! beatmaps are created and refined by a large, diverse community and span underrepresented genres such as anime, Vocaloid, and video game music. We introduce a pipeline for extracting annotations from Osu! beatmaps and partition them into meaningful subsets. Through manual analysis, we find that beatmaps with a single timing point or widely spaced multiple timing points (>=5 seconds apart) provide reliable annotations, while closely spaced timing points (<5 seconds apart) often require additional curation. We also observe high consistency across multiple annotations of the same song. This study demonstrates the potential of Osu! data as a scalable, diverse, and community-driven resource for MIR research. We release our pipeline and a high-quality subset osu2beat2025 to support further exploration: https://github.com/ziyunliu4444/osu2mir.
【11】FunAudio-ASR Technical Report
标题:FunAudio-ASB技术报告
链接:https://arxiv.org/abs/2509.12508
摘要:近年来,自动语音识别(ASR)在三种互补范式的推动下取得了变革性的进步:数据缩放、模型大小缩放以及与大型语言模型(LLM)的深度集成。然而,LLM容易产生幻觉,这会显著降低现实世界ASR应用中的用户体验。在本文中,我们介绍了FunAudio-ASR,这是一个基于LLM的大规模ASR系统,它协同结合了海量数据、大模型容量、LLM集成和强化学习,以在各种复杂的语音识别场景中实现最先进的性能。此外,FunAudio-ASR针对实际部署进行了专门优化,增强了流媒体功能,噪声鲁棒性,代码切换,热词定制,并满足其他实际应用需求。实验结果表明,虽然大多数基于LLM的ASR系统在开源基准测试中表现出色,但它们在真实的行业评估集上往往表现不佳。得益于面向生产的优化,FunAudio-ASR在实际应用数据集上实现了SOTA性能,证明了其在实际环境中的有效性和鲁棒性。
摘要:In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present FunAudio-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, FunAudio-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, FunAudio-ASR achieves SOTA performance on real application datasets, demonstrating its effectiveness and robustness in practical settings.
【12】More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition
标题:相似多于不同:跨数据库语音情感识别的注释器建模
链接:https://arxiv.org/abs/2509.12295
备注:©20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:语音情感识别系统通常预测从多个注释者的评级生成的共识值。然而,这些模型预测任何一个人的注释的能力有限。或者,模型可以学习预测所有注释器的注释。使这些模型适应新的注释器是困难的,因为新的注释器必须单独提供足够的标记训练数据。我们建议通过使用在大型注释者群体上预训练的模型来利用注释者之间的相似性,以识别类似的,以前见过的注释者。给定一个新的、以前看不见的注释器和有限的注册数据,我们可以对类似的注释器进行预测,从而对目标数据集中看不见的数据进行现成的注释,从而提供一种成本极低的个性化机制。我们证明了我们的方法显着优于其他现成的方法,为轻量级的情感适应,实际的现实世界的部署铺平了道路。
摘要:Speech emotion recognition systems often predict a consensus value generated from the ratings of multiple annotators. However, these models have limited ability to predict the annotation of any one person. Alternatively, models can learn to predict the annotations of all annotators. Adapting such models to new annotators is difficult as new annotators must individually provide sufficient labeled training data. We propose to leverage inter-annotator similarity by using a model pre-trained on a large annotator population to identify a similar, previously seen annotator. Given a new, previously unseen, annotator and limited enrollment data, we can make predictions for a similar annotator, enabling off-the-shelf annotation of unseen data in target datasets, providing a mechanism for extremely low-cost personalization. We demonstrate our approach significantly outperforms other off-the-shelf approaches, paving the way for lightweight emotion adaptation, practical for real-world deployment.
【13】Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio questuin answering
标题:Omni-CLST:带引导的选择性思维链的音频问题回答的错误意识课程学习
链接:https://arxiv.org/abs/2509.12275
备注:5 pages, 1 figure, 2 tables
摘要:我们提出了Omni-CLST,一个错误意识的课程学习框架与指导选择性的思想链音频问题回答。该框架通过两个关键策略有效地利用了现有的高质量数据集:一个错误感知课程,按难度组织样本,以及一个引导思想的退出机制,将推理集中在具有挑战性的案例上。与GRPO训练相结合,这些策略使模型能够更有效地从信息样本中学习。在MMAU-mini和MMAR上的实验表明,Omni-CLST达到了有竞争力的准确性(在MMAU-mini上为73.80%),并建立了新的最新技术水平(在MMAR上为64.30%),凸显了其在多模态音频语言理解中的鲁棒性和泛化能力。
摘要:We propose Omni-CLST, an error-aware Curriculum Learning framework with guided Selective Chain-of-Thought for audio question answering. The framework efficiently leverages existing high-quality dataset through two key strategies: an error-aware curriculum that organizes samples by difficulty, and a guided thought dropout mechanism that focuses reasoning on challenging cases. Integrated with GRPO training, these strategies enable the model to learn more effectively from informative samples. Experiments on MMAU-mini and MMAR demonstrate that Omni-CLST achieves competitive accuracy (73.80% on MMAU-mini) and establishes a new state of the art (64.30% on MMAR), highlighting its robustness and generalization capability in multimodal audio-language understanding.
【14】A Traditional Approach to Symbolic Piano Continuation
标题:象征性钢琴延续的传统方法
链接:https://arxiv.org/abs/2509.12267
备注:3 pages, extended abstract, MIREX session at ISMIR 2025 LBD
摘要:我们为MIREX 2025象征音乐世代挑战赛提供了一种传统的象征钢琴音乐延续方法。虽然计算音乐生成最近专注于开发大型基础模型与复杂的架构修改,我们认为,更简单的方法仍然更有效的约束,单乐器的任务。因此,我们回到了一个简单的,未增强的下一个令牌预测目标,目标是通过使用更好的数据和更好的基本面来超越大型基础模型。我们在https://github.com/christianazinn/mirex2025上发布模型重量和代码。
摘要:We present a traditional approach to symbolic piano music continuation for the MIREX 2025 Symbolic Music Generation challenge. While computational music generation has recently focused on developing large foundation models with sophisticated architectural modifications, we argue that simpler approaches remain more effective for constrained, single-instrument tasks. We thus return to a simple, unaugmented next-token-prediction objective on tokenized raw MIDI, aiming to outperform large foundation models by using better data and better fundamentals. We release model weights and code at https://github.com/christianazinn/mirex2025.
【15】An Adaptive CMSA for Solving the Longest Filled Common Subsequence Problem with an Application in Audio Querying
标题:一种自适应CMSA算法及其在音频查询中的应用
链接:https://arxiv.org/abs/2509.12261
摘要:最长填充公共子序列(Longest Filled Common Subsequence,LFCS)问题是生物信息学中一个具有挑战性的NP难题,在基因突变预测和基因组数据重构等领域有着广泛的应用。现有的方法,包括精确的,元启发式和近似算法,主要是在小规模的实例,提供有限的见解,其可扩展性进行了评估。在这项工作中,我们引入了一个新的基准数据集,具有显着更大的实例,并证明现有的数据集缺乏有意义地评估算法性能所需的区分能力。为了有效地解决大型实例,我们利用自适应构造,合并,求解,适应(CMSA)框架,通过基于组件的构造迭代生成有前途的子问题,并使用先前迭代的反馈进行细化。使用外部黑盒求解器解决子问题。在标准和新引入的基准测试上的大量实验表明,所提出的自适应CMSA实现了最先进的性能,优于五种领先的方法。值得注意的是,在已知最佳解决方案的1,510个问题实例中,我们的方法解决了其中的1,486个-实现了超过99.9%的最佳解决方案质量,并展示了卓越的可扩展性。我们还提出了一个新的应用LFCS的歌曲识别从退化的音频摘录作为工程的贡献,使用现实世界的能量配置文件的实例,从流行音乐。最后,我们进行了一个实证的可解释性分析,以确定影响算法性能的关键特征组合,即,揭示了影响不同实例类型方法成功或失败的关键问题特征。
摘要:This paper addresses the Longest Filled Common Subsequence (LFCS) problem, a challenging NP-hard problem with applications in bioinformatics, including gene mutation prediction and genomic data reconstruction. Existing approaches, including exact, metaheuristic, and approximation algorithms, have primarily been evaluated on small-sized instances, which offer limited insights into their scalability. In this work, we introduce a new benchmark dataset with significantly larger instances and demonstrate that existing datasets lack the discriminative power needed to meaningfully assess algorithm performance at scale. To solve large instances efficiently, we utilize an adaptive Construct, Merge, Solve, Adapt (CMSA) framework that iteratively generates promising subproblems via component-based construction and refines them using feedback from prior iterations. Subproblems are solved using an external black-box solver. Extensive experiments on both standard and newly introduced benchmarks show that the proposed adaptive CMSA achieves state-of-the-art performance, outperforming five leading methods. Notably, on 1,510 problem instances with known optimal solutions, our approach solves 1,486 of them -- achieving over 99.9% optimal solution quality and demonstrating exceptional scalability. We additionally propose a novel application of LFCS for song identification from degraded audio excerpts as an engineering contribution, using real-world energy-profile instances from popular music. Finally, we conducted an empirical explainability analysis to identify critical feature combinations influencing algorithm performance, i.e., the key problem features contributing to success or failure of the approaches across different instance types are revealed.
【16】Multi-Modal Embedding-based Target Speaker Enhancement
标题:基于多模式嵌入的目标说话人增强
链接:https://arxiv.org/abs/2509.12583
摘要:目标说话人提取(TSE)是鸡尾酒会场景中的一个关键挑战。虽然利用多种模态,如语音,嘴唇,面部和表情嵌入,可以提高性能,但现实世界的应用程序往往会遭受间歇性的模态丢失。本文提出了一个全面的研究的相互作用和鲁棒性的各种多模态融合策略下不同程度的模态脱落。我们建立在一个国家的最先进的视听语音增强系统,并集成了四个不同的扬声器身份线索:唇嵌入同步的上下文信息,语音扬声器嵌入通过交叉注意提取声学一致性,静态面部嵌入扬声器身份,和一个新的动态表达嵌入帧明智的情感特征。我们系统地评估了这些模式的不同组合下两个关键的培训制度:零辍学和80%的方式辍学。大量的实验表明,虽然一个完整的多模态集成在理想的(零辍学)条件下实现最佳性能,其有效性显着降低时,测试时间发生辍学没有事先暴露在训练过程中。至关重要的是,我们表明,具有高(80%)模态退出率的训练大大增强了模型的鲁棒性,使系统即使在严重的测试时间缺失模态下也能保持卓越的性能。我们的研究结果强调,语音嵌入表现出一致的鲁棒性,而建议的表达嵌入提供了有价值的补充信息。这项工作强调了训练策略的重要性,考虑到现实世界的不完美,超越纯粹的性能最大化,以实现多模态语音增强系统的实际可靠性。
摘要:Target Speaker Extraction (TSE) is a critical challenge in cocktail party scenarios. While leveraging multiple modalities, such as voice, lip, face, and expression embeddings, can enhance performance, real-world applications often suffer from intermittent modality dropout. This paper presents a comprehensive study on the interactions and robustness of various multimodal fusion strategies under varying degrees of modality dropout. We build upon a state-of-the-art audio-visual speech enhancement system and integrate four distinct speaker identity cues: lip embeddings for synchronized contextual information, a voice speaker embedding extracted via cross-attention for acoustic consistency, a static face embedding for speaker identity, and a novel dynamic expression embedding for frame-wise emotional features. We systematically evaluate different combinations of these modalities under two key training regimes: zero dropout and 80% modality dropout. Extensive experiments demonstrate that while a full multimodal ensemble achieves optimal performance under ideal (zero dropout) conditions, its effectiveness diminishes significantly when test-time dropout occurs without prior exposure during training. Crucially, we show that training with a high (80%) modality dropout rate dramatically enhances model robustness, enabling the system to maintain superior performance even under severe test-time missing modalities. Our findings highlight that voice embeddings exhibit consistent robustness, while the proposed expression embedding provides valuable complementary information. This work underscores the importance of training strategies that account for real-world imperfection, moving beyond pure performance maximization to achieve practical reliability in multimodal speech enhancement systems.
【1】Importance-Weighted Domain Adaptation for Sound Source Tracking
标题:重要性加权域自适应用于声音源跟踪
链接:https://arxiv.org/abs/2509.13215
备注:Accepted paper: Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE 2025)
摘要:近年来,深度学习显著推进了声源定位(SSL)。然而,训练这样的模型需要大量的标记数据集,并且真实记录的注释成本很高,特别是如果源移动的话。虽然使用模拟房间脉冲响应(RIR)和噪声的合成数据提供了一种实用的替代方案,但在合成数据上训练的模型在真实环境中会受到域偏移的影响。无监督域自适应(UDA)可以通过对齐合成域和真实域来解决这个问题,而不依赖于后者的标签。然而,现有的几种UDA方法集中在静态SSL上,并且不考虑声源跟踪(SST)的问题,这提出了两个特定的域适应挑战。首先,可变长度的输入序列在域之间的特征维度中产生不匹配。其次,合成数据和真实数据的角度覆盖可能由于部分域重叠或由于批量大小约束而没有很好地对齐,我们称之为方向多样性失配。为了解决这些问题,我们提出了一种新的UDA方法,专门为SST的基础上两个关键功能。我们采用递归神经网络的最终隐藏状态作为固定维特征表示来处理可变长度序列。此外,我们使用重要性加权对抗训练,通过优先考虑与真实域相似的合成样本来解决方向多样性失配。实验结果表明,我们的方法成功地使合成训练的模型适应真实环境,提高了SST性能。
摘要:In recent years, deep learning has significantly advanced sound source localization (SSL). However, training such models requires large labeled datasets, and real recordings are costly to annotate in particular if sources move. While synthetic data using simulated room impulse responses (RIRs) and noise offers a practical alternative, models trained on synthetic data suffer from domain shift in real environments. Unsupervised domain adaptation (UDA) can address this by aligning synthetic and real domains without relying on labels from the latter. The few existing UDA approaches however focus on static SSL and do not account for the problem of sound source tracking (SST), which presents two specific domain adaptation challenges. First, variable-length input sequences create mismatches in feature dimensionality across domains. Second, the angular coverages of the synthetic and the real data may not be well aligned either due to partial domain overlap or due to batch size constraints, which we refer to as directional diversity mismatch. To address these, we propose a novel UDA approach tailored for SST based on two key features. We employ the final hidden state of a recurrent neural network as a fixed-dimensional feature representation to handle variable-length sequences. Further, we use importance-weighted adversarial training to tackle directional diversity mismatch by prioritizing synthetic samples similar to the real domain. Experimental results demonstrate that our approach successfully adapts synthetic-trained models to real environments, improving SST performance.
【2】Token-based Attractors and Cross-attention in Spoof Diarization
标题:恶搞日记中的基于代币的吸引器和交叉注意
链接:https://arxiv.org/abs/2509.13085
备注:Accepted to IEEE ASRU 2025
摘要:欺骗日志识别“什么时候欺骗”在一个给定的语音通过时间定位欺骗区域,并确定其操作技术。作为实现这一任务的第一步,先前的工作提出了一个用于本地化和欺骗类型聚类的两个分支模型,这为欺骗日志化奠定了基础。然而,其简单的结构限制了捕获复杂欺骗模式的能力,并且缺乏用于区分善意和各种欺骗类型的明确参考点。为了解决这些限制,我们的方法引入了可学习的令牌,其中每个令牌代表真实和欺骗语音的声学特征。这些吸引子与帧级嵌入相互作用,以提取有区别的表示,从而改善真实语音和生成语音之间的分离。在PartialSpoof数据集上的大量实验一致表明,我们的方法在善意检测和欺骗方法聚类方面优于现有方法。
摘要:Spoof diarization identifies ``what spoofed when" in a given speech by temporally locating spoofed regions and determining their manipulation techniques. As a first step toward this task, prior work proposed a two-branch model for localization and spoof type clustering, which laid the foundation for spoof diarization. However, its simple structure limits the ability to capture complex spoofing patterns and lacks explicit reference points for distinguishing between bona fide and various spoofing types. To address these limitations, our approach introduces learnable tokens where each token represents acoustic features of bona fide and spoofed speech. These attractors interact with frame-level embeddings to extract discriminative representations, improving separation between genuine and generated speech. Vast experiments on PartialSpoof dataset consistently demonstrate that our approach outperforms existing methods in bona fide detection and spoofing method clustering.
【3】MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
标题:MSR-Codec:一种用于具有信息解纠缠的高保真语音生成的低比特率多流剩余编解码器
链接:https://arxiv.org/abs/2509.13068
摘要:音频编解码器是现代语音生成系统的关键组件。本文介绍了一种低比特率,多尺度残差编解码器,编码语音成四个不同的流:语义,音色,韵律,和残留。这种架构实现了高保真语音重建在有竞争力的低比特率,同时表现出固有的信息解纠缠的能力。我们构建了一个两阶段的语言模型的文本到语音(TTS)合成使用这个编解码器,尽管其轻量级的设计和最小的数据要求,实现了最先进的字错误率(WER)和优越的扬声器相似性相比,几个较大的模型。此外,该编解码器的设计被证明是非常有效的语音转换,使扬声器的音色和韵律的独立操纵。
摘要:Audio codecs are a critical component of modern speech generation systems. This paper introduces a low-bitrate, multi-scale residual codec that encodes speech into four distinct streams: semantic, timbre, prosody, and residual. This architecture achieves high-fidelity speech reconstruction at competitive low bitrates while demonstrating an inherent ability for information disentanglement. We construct a two-stage language model for text-to-speech (TTS) synthesis using this codec, which, despite its lightweight design and minimal data requirements, achieves a state-of-the-art Word Error Rate (WER) and superior speaker similarity compared to several larger models. Furthermore, the codec's design proves highly effective for voice conversion, enabling independent manipulation of speaker timbre and prosody.
【4】Investigating the Potential of Multi-Stage Score Fusion in Spoofing-Aware Speaker Verification
标题:研究多阶段分数融合在欺骗意识说话者验证中的潜力
链接:https://arxiv.org/abs/2509.12668
备注:published in SIU2025
摘要:尽管自动说话人验证(ASV)有所改进,但针对欺骗攻击的漏洞仍然是一个主要问题。在这项研究中,我们调查的ASV和对策(CM)子系统集成到一个模块化的欺骗意识说话人验证(SASV)框架。与传统的单阶段分数级融合方法不同,我们探索了在多个阶段利用ASV和CM系统的多阶段方法的潜力。通过利用ECAPA-TDNN(ASV)和AASIST(CM)子系统,我们考虑支持向量机和逻辑回归分类器来实现SASV。在第二阶段,我们将它们的输出与原始分数相结合,以修改融合后端分类器。此外,我们还结合了来自RawGAT(CM)的另一个辅助评分,以进一步增强我们的SASV框架。我们的方法在SASV 2022挑战的评估数据集上产生了1.30%的等误差率(EER),比基线系统相对改善了24%。
摘要:Despite improvements in automatic speaker verification (ASV), vulnerability against spoofing attacks remains a major concern. In this study, we investigate the integration of ASV and countermeasure (CM) subsystems into a modular spoof-aware speaker verification (SASV) framework. Unlike conventional single-stage score-level fusion methods, we explore the potential of a multi-stage approach that utilizes the ASV and CM systems in multiple stages. By leveraging ECAPA-TDNN (ASV) and AASIST (CM) subsystems, we consider support vector machine and logistic regression classifiers to achieve SASV. In the second stage, we integrate their outputs with the original score to revise fusion back-end classifiers. Additionally, we incorporate another auxiliary score from RawGAT (CM) to further enhance our SASV framework. Our approach yields an equal error rate (EER) of 1.30% on the evaluation dataset of the SASV2022 challenge, representing a 24% relative improvement over the baseline system.
【5】Multi-Modal Embedding-based Target Speaker Enhancement
标题:基于多模式嵌入的目标说话人增强
链接:https://arxiv.org/abs/2509.12583
摘要:目标说话人提取(TSE)是鸡尾酒会场景中的一个关键挑战。虽然利用多种模态,如语音,嘴唇,面部和表情嵌入,可以提高性能,但现实世界的应用程序往往会遭受间歇性的模态丢失。本文提出了一个全面的研究的相互作用和鲁棒性的各种多模态融合策略下不同程度的模态脱落。我们建立在一个国家的最先进的视听语音增强系统,并集成了四个不同的扬声器身份线索:唇嵌入同步的上下文信息,语音扬声器嵌入通过交叉注意提取声学一致性,静态面部嵌入扬声器身份,和一个新的动态表达嵌入帧明智的情感特征。我们系统地评估了这些模式的不同组合下两个关键的培训制度:零辍学和80%的方式辍学。大量的实验表明,虽然一个完整的多模态集成在理想的(零辍学)条件下实现最佳性能,其有效性显着降低时,测试时间发生辍学没有事先暴露在训练过程中。至关重要的是,我们表明,具有高(80%)模态退出率的训练大大增强了模型的鲁棒性,使系统即使在严重的测试时间缺失模态下也能保持卓越的性能。我们的研究结果强调,语音嵌入表现出一致的鲁棒性,而建议的表达嵌入提供了有价值的补充信息。这项工作强调了训练策略的重要性,考虑到现实世界的不完美,超越纯粹的性能最大化,以实现多模态语音增强系统的实际可靠性。
摘要:Target Speaker Extraction (TSE) is a critical challenge in cocktail party scenarios. While leveraging multiple modalities, such as voice, lip, face, and expression embeddings, can enhance performance, real-world applications often suffer from intermittent modality dropout. This paper presents a comprehensive study on the interactions and robustness of various multimodal fusion strategies under varying degrees of modality dropout. We build upon a state-of-the-art audio-visual speech enhancement system and integrate four distinct speaker identity cues: lip embeddings for synchronized contextual information, a voice speaker embedding extracted via cross-attention for acoustic consistency, a static face embedding for speaker identity, and a novel dynamic expression embedding for frame-wise emotional features. We systematically evaluate different combinations of these modalities under two key training regimes: zero dropout and 80% modality dropout. Extensive experiments demonstrate that while a full multimodal ensemble achieves optimal performance under ideal (zero dropout) conditions, its effectiveness diminishes significantly when test-time dropout occurs without prior exposure during training. Crucially, we show that training with a high (80%) modality dropout rate dramatically enhances model robustness, enabling the system to maintain superior performance even under severe test-time missing modalities. Our findings highlight that voice embeddings exhibit consistent robustness, while the proposed expression embedding provides valuable complementary information. This work underscores the importance of training strategies that account for real-world imperfection, moving beyond pure performance maximization to achieve practical reliability in multimodal speech enhancement systems.
【6】The CCF AATC 2025: Speech Restoration Challenge
标题:CPD AATC 2025:言语恢复挑战
链接:https://arxiv.org/abs/2509.12974
备注:Technical Report
摘要:真实世界的语音通信经常受到各种失真的阻碍,这些失真降低了质量和可懂度。虽然许多语音增强算法针对特定的退化,如噪声或混响,但它们通常在多个失真共存和相互作用的现实场景中不足。为了促进这一领域的研究,我们推出了语音恢复挑战赛,作为2025年中国计算机联合会(CCF)高级音频技术竞赛(AATC)的一部分。这一挑战的重点是恢复受三种退化类型复合影响的语音信号:(1)复杂的声学退化,包括非平稳噪声和混响;(2)信号链伪影,如MP3压缩;以及(3)由其他预处理增强模型引入的次级伪影。我们描述了挑战的背景,任务的设计,全面的数据集创建方法,以及详细的评估协议,评估目标性能和模型复杂性。主页:https://ccf-aatc.org.cn/。
摘要:Real-world speech communication is often hampered by a variety of distortions that degrade quality and intelligibility. While many speech enhancement algorithms target specific degradations like noise or reverberation, they often fall short in realistic scenarios where multiple distortions co-exist and interact. To spur research in this area, we introduce the Speech Restoration Challenge as part of the China Computer Federation (CCF) Advanced Audio Technology Competition (AATC) 2025. This challenge focuses on restoring speech signals affected by a composite of three degradation types: (1) complex acoustic degradations including non-stationary noise and reverberation; (2) signal-chain artifacts such as those from MP3 compression; and (3) secondary artifacts introduced by other pre-processing enhancement models. We describe the challenge's background, the design of the task, the comprehensive dataset creation methodology, and the detailed evaluation protocol, which assesses both objective performance and model complexity. Homepage: https://ccf-aatc.org.cn/.
【7】PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
标题:PAC:发音感知上下文化基于大语言模型的自动语音识别
链接:https://arxiv.org/abs/2509.12647
备注:Submitted to ICASSP 2026
摘要:本文提出了一个发音感知的上下文(PAC)框架,以解决两个关键的挑战,在大语言模型(LLM)为基础的自动语音识别(ASR)系统:有效的发音建模和鲁棒的同音字歧视。这两个都是原始或长尾词识别所必需的。所提出的方法采用了两个阶段的学习范式。首先,我们介绍了一种发音引导的上下文学习方法。它采用了一个交错的字形音素上下文建模策略,结合字形只有干扰,鼓励模型利用音素线索准确识别。然后,我们提出了一种带有扰动标签采样的发音判别式强化学习方法,以进一步增强模型区分上下文同音异义词的能力。在公共英语Librispeech和普通话AISHELL-1数据集上的实验结果表明,PAC:(1)与预训练的基于LLM的ASR模型相比,相对单词错误率(WER)分别降低了30.2%和53.8%,(2)与强基线相比,长尾词的偏差WER分别降低了31.8%和60.5%。
摘要:This paper presents a Pronunciation-Aware Contextualized (PAC) framework to address two key challenges in Large Language Model (LLM)-based Automatic Speech Recognition (ASR) systems: effective pronunciation modeling and robust homophone discrimination. Both are essential for raw or long-tail word recognition. The proposed approach adopts a two-stage learning paradigm. First, we introduce a pronunciation-guided context learning method. It employs an interleaved grapheme-phoneme context modeling strategy that incorporates grapheme-only distractors, encouraging the model to leverage phonemic cues for accurate recognition. Then, we propose a pronunciation-discriminative reinforcement learning method with perturbed label sampling to further enhance the model\'s ability to distinguish contextualized homophones. Experimental results on the public English Librispeech and Mandarin AISHELL-1 datasets indicate that PAC: (1) reduces relative Word Error Rate (WER) by 30.2% and 53.8% compared to pre-trained LLM-based ASR models, and (2) achieves 31.8% and 60.5% relative reductions in biased WER for long-tail words compared to strong baselines, respectively.
【8】FunAudio-ASR Technical Report
标题:FunAudio-ASB技术报告
链接:https://arxiv.org/abs/2509.12508
摘要:近年来,自动语音识别(ASR)在三种互补范式的推动下取得了变革性的进步:数据缩放、模型大小缩放以及与大型语言模型(LLM)的深度集成。然而,LLM容易产生幻觉,这会显著降低现实世界ASR应用中的用户体验。在本文中,我们介绍了FunAudio-ASR,这是一个基于LLM的大规模ASR系统,它协同结合了海量数据、大模型容量、LLM集成和强化学习,以在各种复杂的语音识别场景中实现最先进的性能。此外,FunAudio-ASR针对实际部署进行了专门优化,增强了流媒体功能,噪声鲁棒性,代码切换,热词定制,并满足其他实际应用需求。实验结果表明,虽然大多数基于LLM的ASR系统在开源基准测试中表现出色,但它们在真实的行业评估集上往往表现不佳。得益于面向生产的优化,FunAudio-ASR在实际应用数据集上实现了SOTA性能,证明了其在实际环境中的有效性和鲁棒性。
摘要:In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present FunAudio-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, FunAudio-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, FunAudio-ASR achieves SOTA performance on real application datasets, demonstrating its effectiveness and robustness in practical settings.
【9】More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition
标题:相似多于不同:跨数据库语音情感识别的注释器建模
链接:https://arxiv.org/abs/2509.12295
备注:©20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:语音情感识别系统通常预测从多个注释者的评级生成的共识值。然而,这些模型预测任何一个人的注释的能力有限。或者,模型可以学习预测所有注释器的注释。使这些模型适应新的注释器是困难的,因为新的注释器必须单独提供足够的标记训练数据。我们建议通过使用在大型注释者群体上预训练的模型来利用注释者之间的相似性,以识别类似的,以前见过的注释者。给定一个新的、以前看不见的注释器和有限的注册数据,我们可以对类似的注释器进行预测,从而对目标数据集中看不见的数据进行现成的注释,从而提供一种成本极低的个性化机制。我们证明我们的方法显着优于其他现成的方法,为轻量级的情感适应铺平了道路,对于现实世界的部署是实用的。
摘要:Speech emotion recognition systems often predict a consensus value generated from the ratings of multiple annotators. However, these models have limited ability to predict the annotation of any one person. Alternatively, models can learn to predict the annotations of all annotators. Adapting such models to new annotators is difficult as new annotators must individually provide sufficient labeled training data. We propose to leverage inter-annotator similarity by using a model pre-trained on a large annotator population to identify a similar, previously seen annotator. Given a new, previously unseen, annotator and limited enrollment data, we can make predictions for a similar annotator, enabling off-the-shelf annotation of unseen data in target datasets, providing a mechanism for extremely low-cost personalization. We demonstrate our approach significantly outperforms other off-the-shelf approaches, paving the way for lightweight emotion adaptation, practical for real-world deployment.
【10】Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio questuin answering
标题:Omni-CLST:带引导的选择性思维链的音频问题回答的错误意识课程学习
链接:https://arxiv.org/abs/2509.12275
备注:5 pages, 1 figure, 2 tables
摘要:我们提出了Omni-CLST,一个错误意识的课程学习框架与指导选择性的思想链音频问题回答。该框架通过两个关键策略有效地利用了现有的高质量数据集:一个错误感知课程,按难度组织样本,以及一个引导思想的退出机制,将推理集中在具有挑战性的案例上。与GRPO训练相结合,这些策略使模型能够更有效地从信息样本中学习。在MMAU-mini和MMAR上的实验表明,Omni-CLST实现了具有竞争力的准确率(MMAU-mini上为73.80%),并建立了一个新的最先进的水平(MMAR上为64.30%),突出了其在多模态音频语言理解中的鲁棒性和泛化能力。
摘要:We propose Omni-CLST, an error-aware Curriculum Learning framework with guided Selective Chain-of-Thought for audio question answering. The framework efficiently leverages existing high-quality dataset through two key strategies: an error-aware curriculum that organizes samples by difficulty, and a guided thought dropout mechanism that focuses reasoning on challenging cases. Integrated with GRPO training, these strategies enable the model to learn more effectively from informative samples. Experiments on MMAU-mini and MMAR demonstrate that Omni-CLST achieves competitive accuracy (73.80% on MMAU-mini) and establishes a new state of the art (64.30% on MMAR), highlighting its robustness and generalization capability in multimodal audio-language understanding.
【11】A Traditional Approach to Symbolic Piano Continuation
标题:象征性钢琴延续的传统方法
链接:https://arxiv.org/abs/2509.12267
备注:3 pages, extended abstract, MIREX session at ISMIR 2025 LBD
摘要:我们为MIREX 2025象征音乐世代挑战赛提供了一种传统的象征钢琴音乐延续方法。虽然计算音乐生成最近专注于开发大型基础模型与复杂的架构修改,我们认为,更简单的方法仍然更有效的约束,单乐器的任务。因此,我们回到了一个简单的,未增强的下一个令牌预测目标,目标是通过使用更好的数据和更好的基本面来超越大型基础模型。我们在https://github.com/christianazinn/mirex2025上发布模型重量和代码。
摘要:We present a traditional approach to symbolic piano music continuation for the MIREX 2025 Symbolic Music Generation challenge. While computational music generation has recently focused on developing large foundation models with sophisticated architectural modifications, we argue that simpler approaches remain more effective for constrained, single-instrument tasks. We thus return to a simple, unaugmented next-token-prediction objective on tokenized raw MIDI, aiming to outperform large foundation models by using better data and better fundamentals. We release model weights and code at https://github.com/christianazinn/mirex2025.
【12】An Adaptive CMSA for Solving the Longest Filled Common Subsequence Problem with an Application in Audio Querying
标题:一种自适应CMSA算法及其在音频查询中的应用
链接:https://arxiv.org/abs/2509.12261
摘要:最长填充公共子序列(Longest Filled Common Subsequence,LFCS)问题是生物信息学中一个具有挑战性的NP难题,在基因突变预测和基因组数据重构等领域有着广泛的应用。现有的方法,包括精确的,元启发式和近似算法,主要是在小规模的实例,提供有限的见解,其可扩展性进行了评估。在这项工作中,我们引入了一个新的基准数据集,具有显着更大的实例,并证明现有的数据集缺乏有意义地评估算法性能所需的区分能力。为了有效地解决大型实例,我们利用自适应构造,合并,求解,适应(CMSA)框架,通过基于组件的构造迭代生成有前途的子问题,并使用先前迭代的反馈进行细化。使用外部黑盒求解器解决子问题。在标准和新引入的基准测试上的大量实验表明,所提出的自适应CMSA实现了最先进的性能,优于五种领先的方法。值得注意的是,在已知最佳解决方案的1,510个问题实例中,我们的方法解决了其中的1,486个-实现了超过99.9%的最佳解决方案质量,并展示了卓越的可扩展性。我们还提出了一个新的应用LFCS的歌曲识别从退化的音频摘录作为工程的贡献,使用现实世界的能量配置文件的实例,从流行音乐。最后,我们进行了一个实证的可解释性分析,以确定影响算法性能的关键特征组合,即,揭示了影响不同实例类型方法成功或失败的关键问题特征。
摘要:This paper addresses the Longest Filled Common Subsequence (LFCS) problem, a challenging NP-hard problem with applications in bioinformatics, including gene mutation prediction and genomic data reconstruction. Existing approaches, including exact, metaheuristic, and approximation algorithms, have primarily been evaluated on small-sized instances, which offer limited insights into their scalability. In this work, we introduce a new benchmark dataset with significantly larger instances and demonstrate that existing datasets lack the discriminative power needed to meaningfully assess algorithm performance at scale. To solve large instances efficiently, we utilize an adaptive Construct, Merge, Solve, Adapt (CMSA) framework that iteratively generates promising subproblems via component-based construction and refines them using feedback from prior iterations. Subproblems are solved using an external black-box solver. Extensive experiments on both standard and newly introduced benchmarks show that the proposed adaptive CMSA achieves state-of-the-art performance, outperforming five leading methods. Notably, on 1,510 problem instances with known optimal solutions, our approach solves 1,486 of them -- achieving over 99.9% optimal solution quality and demonstrating exceptional scalability. We additionally propose a novel application of LFCS for song identification from degraded audio excerpts as an engineering contribution, using real-world energy-profile instances from popular music. Finally, we conducted an empirical explainability analysis to identify critical feature combinations influencing algorithm performance, i.e., the key problem features contributing to success or failure of the approaches across different instance types are revealed.
机器翻译由腾讯交互翻译提供,仅供参考
