今日论文合集:cs.SD语音26篇,eess.AS音频处理14篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
标题:BridgeCode:自回归Zero-Shot文本到语音合成的双重语音表示范式
链接:https://arxiv.org/abs/2510.11646

作者:Jingyuan Xing, Mingru Yang, Zhipeng Li, Xiaofen Xing, Xiangmin Xu
摘要:自回归(AR)框架最近通过利用离散语音令牌和大语言模型技术在zero-shot文本到语音(TTS)方面取得了显著进展。尽管他们的成功,现有的基于AR的zero-shot TTS系统面临着两个关键的限制:(i)固有的速度-质量权衡,因为顺序令牌生成要么以表现力为代价降低帧速率,要么以效率为代价丰富令牌,以及(ii)面向文本的监督不匹配,因为交叉熵损失均匀地惩罚令牌错误,而不考虑相邻令牌之间的细粒度声学相似性。为了解决这些挑战,我们提出了BridgeTTS,一种新的AR-TTS框架建立在双重语音表示范式BridgeCode。BridgeTTS通过预测稀疏令牌来减少AR迭代,同时重建丰富的连续特征以实现高质量的合成。标记级和特征级目标的联合优化进一步增强了自然度和可理解性。实验表明,BridgeTTS实现竞争力的质量和扬声器相似性,同时显着加快合成。语音演示可在https://test1562.github.io/demo/上获得。
摘要:Autoregressive (AR) frameworks have recently achieved remarkable progress in zero-shot text-to-speech (TTS) by leveraging discrete speech tokens and large language model techniques. Despite their success, existing AR-based zero-shot TTS systems face two critical limitations: (i) an inherent speed-quality trade-off, as sequential token generation either reduces frame rates at the cost of expressiveness or enriches tokens at the cost of efficiency, and (ii) a text-oriented supervision mismatch, as cross-entropy loss penalizes token errors uniformly without considering the fine-grained acoustic similarity among adjacent tokens. To address these challenges, we propose BridgeTTS, a novel AR-TTS framework built upon the dual speech representation paradigm BridgeCode. BridgeTTS reduces AR iterations by predicting sparse tokens while reconstructing rich continuous features for high-quality synthesis. Joint optimization of token-level and feature-level objectives further enhances naturalness and intelligibility. Experiments demonstrate that BridgeTTS achieves competitive quality and speaker similarity while significantly accelerating synthesis. Speech demos are available at https://test1562.github.io/demo/.


【2】Automatic Music Sample Identification with Multi-Track Contrastive Learning
标题:多轨对比学习自动音乐样本识别
链接:https://arxiv.org/abs/2510.11507

作者:Alain Riou, Joan Serrà, Yuki Mitsufuji
摘要:采样,即重复使用现有音轨片段以创建新音乐内容的技术,是现代音乐制作中非常常见的做法。在本文中,我们解决了具有挑战性的任务,自动样本识别,也就是说,检测这样的采样内容和检索的材料从它的起源。为此,我们采用了一种自监督学习方法,该方法利用多轨道数据集来创建人工混合的正对,并设计了一种新的对比学习目标。我们表明,这种方法显着优于以前的国家的最先进的基线,这是强大的各种流派,以及规模时,增加参考数据库中的噪声歌曲的数量。此外,我们广泛分析了我们的培训管道的不同组成部分的贡献,并强调,特别是,需要高质量的分离干这项任务。
摘要:Sampling, the technique of reusing pieces of existing audio tracks to create new music content, is a very common practice in modern music production. In this paper, we tackle the challenging task of automatic sample identification, that is, detecting such sampled content and retrieving the material from which it originates. To do so, we adopt a self-supervised learning approach that leverages a multi-track dataset to create positive pairs of artificial mixes, and design a novel contrastive learning objective. We show that such method significantly outperforms previous state-of-the-art baselines, that is robust to various genres, and that scales well when increasing the number of noise songs in the reference database. In addition, we extensively analyze the contribution of the different components of our training pipeline and highlight, in particular, the need for high-quality separated stems for this task.


【3】Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
标题:音频大师:用工具增强推理增强大型音频语言模型
链接:https://arxiv.org/abs/2510.11454

作者:Kuan-Yi Lee, Tsung-En Lin, Hung-Yi Lee
备注:9pages
摘要:大型多模态模型(LMPs)的最新进展已经显示出强大的音频理解能力。然而,大多数系统仅依赖于端到端推理,限制了需要结构化知识或专门信号分析的任务的可解释性和准确性。在这项工作中,我们提出了Audio-Maestro -一个工具增强的音频推理框架,使音频语言模型能够自主调用外部工具,并将其时间戳输出集成到推理过程中。这种设计允许模型通过专门的工具来分析、转换和解释音频信号,而不是仅仅依赖于端到端的推理。实验表明,Audio-Maestro持续提高了一般音频推理性能:Gemini-2.5-flash在MMAU-Test上的平均准确率从67.4%上升到72.1%,DeSTA-2.5从58.3%上升到62.8%,GPT-4 o从60.8%上升到63.9%。据我们所知,Audio-Maestro是第一个将结构化工具输出集成到大型音频语言模型推理过程中的框架。
摘要:Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding. However, most systems rely solely on end-to-end reasoning, limiting interpretability and accuracy for tasks that require structured knowledge or specialized signal analysis. In this work, we present Audio-Maestro -- a tool-augmented audio reasoning framework that enables audio-language models to autonomously call external tools and integrate their timestamped outputs into the reasoning process. This design allows the model to analyze, transform, and interpret audio signals through specialized tools rather than relying solely on end-to-end inference. Experiments show that Audio-Maestro consistently improves general audio reasoning performance: Gemini-2.5-flash's average accuracy on MMAU-Test rises from 67.4% to 72.1%, DeSTA-2.5 from 58.3% to 62.8%, and GPT-4o from 60.8% to 63.9%. To our knowledge, Audio-Maestro is the first framework to integrate structured tool output into the large audio language model reasoning process.


【4】Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
标题:扩散-链接:弥合音频-文本情态差距的扩散概率模型
链接:https://arxiv.org/abs/2510.11330

作者:KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo, Joon Son Chung
备注:5 pages. Submitted to IEEE ICASSP 2026
摘要:对比音频语言预训练产生强大的联合表示,但持续的音频文本模态差距限制了多模态编码器与大型语言模型(LLM)耦合的好处。我们提出了扩散链接,一个基于扩散的模态桥接模块,生成映射到文本嵌入分布的音频嵌入。该模块在冻结的多模态编码器的输出嵌入处进行训练,并实现为具有三个残留MLP块的轻量级网络。为了评估扩散链接对多模态编码器-LLM耦合的影响,我们对自动音频字幕(AAC)进行了评估;据我们所知,这是基于扩散的模态桥接到AAC的第一个应用。我们报告两个结果。(1)模式差距分析:在相似性和几何标准上,Diffusion-Link在现有的基于扩散的方法中最大程度地减少了模态间隙,并且显示了音频嵌入向文本分布的集体迁移。(2)下游AAC:将Diffusion-Link附加到相同的多模态LLM基线,在没有外部知识的情况下,在zero-shot和完全监督字幕中实现了AudioCaps的最新技术水平,相对增益分别高达52.5%和7.5%。这些研究结果表明,关闭模态差距是多模态编码器和LLM之间有效耦合的关键,基于扩散的模态桥接提供了一个有前途的方向超越知识检索为中心的设计。代码将在接受后发布https://github.com/DevKiHyun/Diffusion-Link
摘要:Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Diffusion-Link, a diffusion-based modality-bridging module that generatively maps audio embeddings into the text-embedding distribution. The module is trained at the output embedding from the frozen multimodal encoder and implemented as a lightweight network with three residual MLP blocks. To assess the effect of Diffusion-Link on multimodal encoder-LLM coupling, we evaluate on Automatic Audio Captioning (AAC); to our knowledge, this is the first application of diffusion-based modality bridging to AAC. We report two results. (1) Modality-gap analysis: on similarity and geometric criteria, Diffusion-Link reduces the modality gap the most among prior diffusion-based methods and shows a collective migration of audio embeddings toward the text distribution. (2) Downstream AAC: attaching Diffusion-Link to the same multimodal LLM baseline achieves state-of-the-art on AudioCaps in both zero-shot and fully supervised captioning without external knowledge, with relative gains up to 52.5% and 7.5%, respectively. These findings show that closing the modality gap is pivotal for effective coupling between multimodal encoders and LLMs, and diffusion-based modality bridging offers a promising direction beyond knowledge-retrieval-centric designs. Code will be released upon acceptance https://github.com/DevKiHyun/Diffusion-Link


【5】Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker
标题:跨语言情感TTC的扰动自监督表示:情感和说话者的阶段化建模
链接:https://arxiv.org/abs/2510.11124

作者:Cheng Gong, Chunyu Qiang, Tianrui Wang, Yu Jiang, Yuheng Lu, Ruihao Jing, Xiaoxiao Miao, Xiaolei Zhang, Longbiao Wang, Jianwu Dang
备注:Submitted to Expert Systems with Applications,11 pages
摘要:跨语言情感文本到语音(TTS)的目标是产生一种语言的语音,捕捉另一种语言的说话者的情感,同时保持目标声音的音色。这种跨语言情感语音合成的过程提出了一个复杂的挑战,需要灵活控制情感,音色和语言。然而,情感和音色在语音信号中高度纠缠,使得细粒度控制具有挑战性。为了解决这个问题,我们提出了EMM-TTS,一个新的两阶段的跨语言情感语音合成框架的基础上扰动自监督学习(SSL)表示。在第一阶段,该模型显式和隐式编码的韵律线索,以捕捉情感表达,而第二阶段恢复的音色从扰动SSL表示。我们进一步研究了不同的扬声器扰动策略共振峰移位和扬声器匿名化对情感和音色的解纠缠的影响。为了加强说话人的保护和表达控制,我们引入了说话人一致性损失(SCL)和说话人情感自适应层归一化(SEALN)模块。此外,我们发现,结合明确的声学特征(例如,F0、能量和持续时间)以及预训练的潜在特征提高了语音克隆性能。全面的多指标评估,包括主观和客观的措施,表明EMM-TTS实现了卓越的自然性,情感的可转移性,跨语言的音色一致性。
摘要:Cross-lingual emotional text-to-speech (TTS) aims to produce speech in one language that captures the emotion of a speaker from another language while maintaining the target voice's timbre. This process of cross-lingual emotional speech synthesis presents a complex challenge, necessitating flexible control over emotion, timbre, and language. However, emotion and timbre are highly entangled in speech signals, making fine-grained control challenging. To address this issue, we propose EMM-TTS, a novel two-stage cross-lingual emotional speech synthesis framework based on perturbed self-supervised learning (SSL) representations. In the first stage, the model explicitly and implicitly encodes prosodic cues to capture emotional expressiveness, while the second stage restores the timbre from perturbed SSL representations. We further investigate the effect of different speaker perturbation strategies-formant shifting and speaker anonymization-on the disentanglement of emotion and timbre. To strengthen speaker preservation and expressive control, we introduce Speaker Consistency Loss (SCL) and Speaker-Emotion Adaptive Layer Normalization (SEALN) modules. Additionally, we find that incorporating explicit acoustic features (e.g., F0, energy, and duration) alongside pretrained latent features improves voice cloning performance. Comprehensive multi-metric evaluations, including both subjective and objective measures, demonstrate that EMM-TTS achieves superior naturalness, emotion transferability, and timbre consistency across languages.


【6】VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
标题:VCB Bench:基于音频的大型语言模型对话代理的评估基准
链接:https://arxiv.org/abs/2510.11098

作者:Jiliang Hu, Wenfu Wang, Zuchao Li, Chenxing Li, Yiyang Zhao, Hanzhao Li, Liqiang Zhang, Meng Yu, Dong Yu
备注:20 pages, 5 figures
摘要:大型音频语言模型(LALM)的最新进展极大地增强了多模态会话系统。然而,现有的基准仍然有限-它们主要以英语为中心,依赖于合成语音,并且缺乏跨多个维度的全面,有区别的评估。为了解决这些差距,我们提出了语音聊天机器人Bench(VCB Bench)-一个完全基于真实人类语音的高质量中文基准。VCB Bench从三个互补的角度评估LALM:指令遵循(包括文本命令之外的语音级别控制),知识理解(一般知识,推理和日常对话)和鲁棒性(在内容,环境和扬声器特性扰动下的稳定性)。在有代表性的LALM上的实验揭示了显著的性能差距,并突出了未来的改进方向。VCB Bench提供了一个可重复的细粒度评估框架,为推进中文语音会话模型提供了标准化的方法和实用的见解。
摘要:Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks remain limited -- they are mainly English-centric, rely on synthetic speech, and lack comprehensive, discriminative evaluation across multiple dimensions. To address these gaps, we present Voice Chat Bot Bench (VCB Bench) -- a high-quality Chinese benchmark built entirely on real human speech. VCB Bench evaluates LALMs from three complementary perspectives: instruction following (including speech-level control beyond text commands), knowledge understanding (general knowledge, reasoning, and daily dialogue), and robustness (stability under perturbations in content, environment, and speaker traits). Experiments on representative LALMs reveal notable performance gaps and highlight future directions for improvement. VCB Bench provides a reproducible and fine-grained evaluation framework, offering standardized methodology and practical insights for advancing Chinese voice conversational models.


【7】MSRBench: A Benchmarking Dataset for Music Source Restoration
标题:MSRBench:一个用于音乐源恢复的基准数据集
链接:https://arxiv.org/abs/2510.10995

作者:Yongyi Zang, Jiarui Hai, Wanying Ge, Qiuqiang Kong, Zheqi Dai, Helin Wang, Yuki Mitsufuji, Mark D. Plumbley
摘要:音乐源恢复(MSR)将源分离扩展到现实环境中,信号经历制作效果(均衡、压缩、混响)和真实世界的降级,目标是恢复原始未处理的源。现有的基准无法衡量恢复保真度:合成数据集使用未经处理的茎,但不切实际的混合物,而真正的生产数据集只提供已经处理过的茎,没有干净的参考。我们提出MSRBench,第一个明确设计用于MSR评估的基准。MSRBench包含八种乐器类别的原始茎混合物对,其中混合物由专业的混合工程师制作。这些原始处理的对能够直接评估分离精度和恢复保真度。除了受控的演播室条件外,混合物还增加了十二种真实世界的退化,包括模拟伪影、声学环境和有损编解码器。使用U-Net和BSRNN的基线实验分别实现了-37.8 dB和-23.4 dB的SI-SNR,感知质量(FAD CLAP)约为0.7-0.8,表明了改进的巨大空间和对特定架构的需求。
摘要:Music Source Restoration (MSR) extends source separation to realistic settings where signals undergo production effects (equalization, compression, reverb) and real-world degradations, with the goal of recovering the original unprocessed sources. Existing benchmarks cannot measure restoration fidelity: synthetic datasets use unprocessed stems but unrealistic mixtures, while real production datasets provide only already-processed stems without clean references. We present MSRBench, the first benchmark explicitly designed for MSR evaluation. MSRBench contains raw stem-mixture pairs across eight instrument classes, where mixtures are produced by professional mixing engineers. These raw-processed pairs enable direct evaluation of both separation accuracy and restoration fidelity. Beyond controlled studio conditions, the mixtures are augmented with twelve real-world degradations spanning analog artifacts, acoustic environments, and lossy codecs. Baseline experiments with U-Net and BSRNN achieve SI-SNR of -37.8 dB and -23.4 dB respectively, with perceptual quality (FAD CLAP) around 0.7-0.8, demonstrating substantial room for improvement and the need for restoration-specific architectures.


【8】Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
标题:通过嵌入有效排序统一通用音频表示神经缩放定律中的变量
链接:https://arxiv.org/abs/2510.10948

作者:Xuyao Deng, Yanjie Sun, Yong Dou, Kele Xu
摘要:标度律深刻地塑造了我们对计算机视觉和自然语言处理中模型性能的理解,但它们在一般音频表示学习中的应用仍有待探索。一个关键的挑战在于一般音频表示的多因素性质-表示质量受到诸如音频长度、嵌入维度、模型深度、模型架构、数据量等变量的共同影响,其中许多是难以分离或分析表达的。在这项工作中,我们提出了一个系统的研究一般音频表示的缩放律,利用嵌入有效秩(RankMe)作为一个统一的度量,封装不同的变量对表示质量的影响。RankMe能够对音频嵌入进行无标签的信息理论量化,使我们能够在广泛的超参数空间中检查缩放行为,包括模型大小,训练数据量,计算预算,架构配置等。我们的实证研究结果揭示了RankMe和表示质量之间一致的幂律关系,这表明嵌入有效秩充当用于评估和预测音频表示学习中的模型性能的可靠代理。这项工作不仅验证了经典的缩放原则,一般的音频域的适用性,但也提供了一个理论上的接地和经验的强大的框架,用于指导未来的模型缩放策略在音频基础模型。
摘要:Scaling laws have profoundly shaped our understanding of model performance in computer vision and natural language processing, yet their application to general audio representation learning remains underexplored. A key challenge lies in the multifactorial nature of general audio representation-representation quality is jointly influenced by variables such as audio length, embedding dimensionality, model depth, model architecture, data volume, etc., many of which are difficult to isolate or express analytically. In this work, we present a systematic study of scaling laws for general audio representations by utilizing embedding effective rank (RankMe) as a unifying metric that encapsulates the impact of diverse variables on representation quality. RankMe enables a label-free, information-theoretic quantification of audio embeddings, allowing us to examine scaling behaviors across a wide hyper-parameter space, including model size, training data volume, computational budget, architectural configurations, etc. Our empirical findings reveal a consistent power-law relationship between RankMe and representation quality, suggesting that embedding effective rank serves as a reliable proxy for assessing and predicting model performance in audio representation learning. This work not only validates the applicability of classical scaling principles to the general audio domain but also offers a theoretically grounded and empirically robust framework for guiding future model scaling strategies in audio foundation models.


【9】FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
标题:FAC-FACCodec:可控零冲击外国口音转换与分解语音编解码器
链接:https://arxiv.org/abs/2510.10785

作者:Yurii Halychanskyi, Cameron Churchwell, Yutong Wen, Volodymyr Kindratenko
备注:5 pages, 2 figures
摘要:以前的口音转换(AC)方法,包括外国口音转换(FAC),缺乏显式的控制程度的修改。由于口音修改可以改变感知的说话人身份,平衡转换强度和身份保护是至关重要的。我们提出了一个AC框架,提供了一个明确的,用户可控的参数口音修改。该方法针对发音,同时保留超音段线索,如语调和音素持续时间。结果表明,性能媲美最近的AC系统,更强的保存扬声器的身份,和独特的支持可控的口音转换。
摘要:Previous accent conversion (AC) methods, including foreign accent conversion (FAC), lack explicit control over the degree of modification. Because accent modification can alter the perceived speaker identity, balancing conversion strength and identity preservation is crucial. We present an AC framework that provides an explicit, user-controllable parameter for accent modification. The method targets pronunciation while preserving suprasegmental cues such as intonation and phoneme durations. Results show performance comparable to recent AC systems, stronger preservation of speaker identity, and unique support for controllable accent conversion.


【10】ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
标题:ParsVoice:一个用于文语合成的大规模多人波斯语语音语料库
链接:https://arxiv.org/abs/2510.10774

作者:Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery
摘要:波斯语尽管在全球有超过1亿人使用,但在高质量语音语料库中仍然严重不足,特别是对于文本到语音(TTS)合成应用程序。现有的波斯语语音数据集通常比英语语音数据集小,这对开发波斯语语音技术造成了关键限制。我们通过介绍ParsVoice来解决这个问题,ParsVoice是专门为TTS应用程序设计的最大的波斯语语音语料库。我们创建了一个自动化管道,将原始有声读物内容转换为TTS就绪的数据,其中包含基于BERT的句子完成检测器,用于精确音频文本对齐的二进制搜索边界优化方法以及为波斯语量身定制的多维质量评估框架等组件。该管道处理了2,000本有声读物,产生了3,526小时的干净语音,这些语音被进一步过滤成适合TTS的1,804小时高质量子集,其中包括470多个扬声器。ParsVoice是最大的高质量波斯语语音数据集,提供与主要英语语料库相当的扬声器多样性和音频质量。完整的数据集已经公开,以加速波斯语语音技术的发展,并作为其他低资源语言的模板。ParsVoice数据集可在ParsVoice(https://huggingface.co/Parsets/MohammadJRanjbar/ParsVoice)上公开获取。
摘要:Persian Language, despite being spoken by over 100 million people worldwide, remains severely underrepresented in high-quality speech corpora, particularly for text-to-speech (TTS) synthesis applications. Existing Persian speech datasets are typically smaller than their English counterparts, which creates a key limitation for developing Persian speech technologies. We address this gap by introducing ParsVoice, the largest Persian speech corpus designed specifically for TTS applications. We created an automated pipeline that transforms raw audiobook content into TTS-ready data, incorporating components such as a BERT-based sentence completion detector, a binary search boundary optimization method for precise audio-text alignment, and multi-dimensional quality assessment frameworks tailored to Persian. The pipeline processes 2,000 audiobooks, yielding 3,526 hours of clean speech, which was further filtered into a 1,804-hour high-quality subset suitable for TTS, featuring more than 470 speakers. ParsVoice is the largest high-quality Persian speech dataset, offering speaker diversity and audio quality comparable to major English corpora. The complete dataset has been made publicly available to accelerate the development of Persian speech technologies and to serve as a template for other low-resource languages. The ParsVoice dataset is publicly available at ParsVoice (https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice).


【11】Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting
标题:双数据扩展,实现稳健的两阶段用户定义关键字定位
链接:https://arxiv.org/abs/2510.10740

作者:Zhiqi Ai, Han Cheng, Yuxin Wang, Shiyi Mu, Shugong Xu, Yongjin Zhou
备注:5 pages, 3 figures
摘要:在本文中,我们提出了DS-KWS,一个两阶段的框架,强大的用户定义的关键字发现。它结合了CTC为基础的方法与流音素搜索模块来定位候选片段,其次是QbyT为基础的方法与音素匹配器模块验证在音素和话语水平。为了进一步提高性能,我们引入了双重数据缩放策略:(1)将ASR语料库从460小时扩展到1,460小时,以加强声学模型;(2)利用超过155 k个锚点类来训练音素匹配器,显著增强易混淆单词的区分。在LibriPhrase上的实验表明,DS-KWS算法在Hard子集上的EER和AUC分别达到6.13%和97.85%,明显优于已有的方法。在Hey-Snips上,它实现了与完整训练模型相当的zero-shot性能,每小时一次误报时的召回率达到99.13%。
摘要:In this paper, we propose DS-KWS, a two-stage framework for robust user-defined keyword spotting. It combines a CTC-based method with a streaming phoneme search module to locate candidate segments, followed by a QbyT-based method with a phoneme matcher module for verification at both the phoneme and utterance levels. To further improve performance, we introduce a dual data scaling strategy: (1) expanding the ASR corpus from 460 to 1,460 hours to strengthen the acoustic model; and (2) leveraging over 155k anchor classes to train the phoneme matcher, significantly enhancing the distinction of confusable words. Experiments on LibriPhrase show that DS-KWS significantly outperforms existing methods, achieving 6.13\% EER and 97.85\% AUC on the Hard subset. On Hey-Snips, it achieves zero-shot performance comparable to full-shot trained models, reaching 99.13\% recall at one false alarm per hour.


【12】Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR
标题:具有熟练度的自适应和数据增强,以实现稳健的L2 ASB
链接:https://arxiv.org/abs/2510.10738

作者:Ling Sun, Charlotte Zhu, Shuju Shi
备注:Submitted to ICASSP 2026
摘要:通用的ASR表现不佳的非典型扬声器,如L2学习者,加强偏见和限制使用的教育和可访问性。使用CEFR分级的发言和提高语料库,我们表明,天真的微调耳语降低平均WER,但同时扩大差距,不成比例地损害较低水平的学习者。为了解决这个问题,我们提出了两种策略:(i)熟练度感知的多任务学习,联合优化ASR与熟练度分类,以及(ii)有针对性的增强,将频谱图掩蔽应用于低熟练度语音以对抗不平衡。这些方法将WER降低了29.4%(相对),插入/缺失错误降低了58.6%(相对)。至关重要的是,尽管反映真实世界分布的数据集存在严重不平衡,但这两种策略始终缩小了熟练程度差距,促进了L2学习者的公平ASR。
摘要:General-purpose ASR underperforms for atypical speakers, such as L2 learners, reinforcing bias and limiting use in education and accessibility. Using the CEFR-graded Speak and Improve corpus, we show that naive fine-tuning of Whisper reduces average WER but simultaneously widens disparities and disproportionately harms lower-level learners. To address this, we propose two strategies: (i) proficiency-aware multitask learning, jointly optimizing ASR with proficiency classification, and (ii) targeted augmentation, applying spectrogram masking to low-proficiency speech to counter imbalance. These approaches reduce WER by up to 29.4 percent (relative) and insertion/deletion errors by as much as 58.6 percent (relative). Crucially, despite the severe imbalance of the dataset reflecting real-world distributions, both strategies consistently narrow proficiency gaps, advancing equitable ASR for L2 learners.


【13】SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
标题:SS-DPPN:可推广心脏音频表示的自我监督双路径基础模型
链接:https://arxiv.org/abs/2510.10719

作者:Ummy Maria Muna, Md Mehedi Hasan Shawon, Md Jobayer, Sumaiya Akter, Md Rakibul Hasan, Md. Golam Rabiul Alam
摘要:心音图的自动分析对于心血管疾病的早期诊断至关重要,但监督式深度学习往往受到缺乏专家注释数据的限制。在本文中,我们提出了自监督双路径原型网络(SS-DPPN),心脏音频表示和分类的基础模型从未标记的数据。该框架引入了一种基于双路径对比学习的架构,该架构使用新型混合损耗同时处理1D波形和2D频谱图。对于下游任务,使用了使用原型网络的度量学习方法,该方法增强了灵敏度并产生了校准良好且值得信赖的预测。SS-DPPN在四个心脏音频基准上实现了最先进的性能。该框架展示了卓越的数据效率,具有完全监督的模型,可将标记数据减少三倍。最后,学习的表示成功地推广到肺音分类和心率估计。我们的实验和研究结果验证SS-DPPN作为一个强大的,可靠的,可扩展的生理信号的基础模型。
摘要:The automated analysis of phonocardiograms is vital for the early diagnosis of cardiovascular disease, yet supervised deep learning is often constrained by the scarcity of expert-annotated data. In this paper, we propose the Self-Supervised Dual-Path Prototypical Network (SS-DPPN), a foundation model for cardiac audio representation and classification from unlabeled data. The framework introduces a dual-path contrastive learning based architecture that simultaneously processes 1D waveforms and 2D spectrograms using a novel hybrid loss. For the downstream task, a metric-learning approach using a Prototypical Network was used that enhances sensitivity and produces well-calibrated and trustworthy predictions. SS-DPPN achieves state-of-the-art performance on four cardiac audio benchmarks. The framework demonstrates exceptional data efficiency with a fully supervised model on three-fold reduction in labeled data. Finally, the learned representations generalize successfully across lung sound classification and heart rate estimation. Our experiments and findings validate SS-DPPN as a robust, reliable, and scalable foundation model for physiological signals.


【14】LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
标题:LSZone:一种用于实时车内多区域语音分离的轻量级空间信息建模架构
链接:https://arxiv.org/abs/2510.10687

作者:Jun Chen, Shichao Hu, Jiuxin Lin, Wenjie Li, Zihan Zhang, Xingchen Li, JinJiang Liu, Longshuai Xiao, Chao Weng, Lei Xie, Zhiyong Wu
备注:submitted to ICASSP 2026
摘要:车内多区域语音分离是从不同语音区域捕获语音,在人车交互中起着至关重要的作用。虽然以前的SpatialNet取得了显著的成绩,但其高计算成本仍然阻碍了车辆的实时应用。为此,本文提出了LSZone,一个轻量级的空间信息建模架构的实时车内多区域语音分离。我们设计了一个空间信息提取-压缩(SpaIEC)模块,它结合了Mel频谱图和耳间相位差(IPD),以减少计算负担,同时保持性能。此外,为了有效地建模空间信息,我们引入了一个非常轻量级的Conv-GRU交叉窄带处理(CNP)模块。实验结果表明,LSZone的复杂度为0.56G MAC,实时因子(RTF)为0.37,在复杂的噪声和多扬声器场景中提供了令人印象深刻的性能。
摘要:In-car multi-zone speech separation, which captures voices from different speech zones, plays a crucial role in human-vehicle interaction. Although previous SpatialNet has achieved notable results, its high computational cost still hinders real-time applications in vehicles. To this end, this paper proposes LSZone, a lightweight spatial information modeling architecture for real-time in-car multi-zone speech separation. We design a spatial information extraction-compression (SpaIEC) module that combines Mel spectrogram and Interaural Phase Difference (IPD) to reduce computational burden while maintaining performance. Additionally, to efficiently model spatial information, we introduce an extremely lightweight Conv-GRU crossband-narrowband processing (CNP) module. Experimental results demonstrate that LSZone, with a complexity of 0.56G MACs and a real-time factor (RTF) of 0.37, delivers impressive performance in complex noise and multi-speaker scenarios.


【15】A Machine Learning Approach for MIDI to Guitar Tablature Conversion
标题:一种机器学习方法用于琴弦转换
链接:https://arxiv.org/abs/2510.10619

作者:Maximos Kaliakatsos-Papakostas, Gregoris Bastas, Dimos Makris, Dorien Herremans, Vassilis Katsouros, Petros Maragos
备注:Proceedings of the 19th Sound and Music Computing Conference, June 5-12th, 2022, Saint-Étienne (France)
摘要:吉他指法谱转录包括推断字符串和烦恼的数字上,每个音符应发挥再现实际的音乐部分。这种分配应该导致整个轨道上可演奏的弦-品组合,一般来说,在连续组合之间保持简约的运动。在吉他演奏的历史中,不同的音乐风格都发展出了特定的和弦指法,这些指法促进了常见的惯用发音组合和它们之间的运动。本文提出了一种将吉他指法符号分配给给定的基于MIDI的音乐部分(可能由多个复调音轨组成)的方法,即不涉及关于吉他惯用表达特征的信息(例如弯曲等)。目前的策略是基于机器学习,需要一个关于手指在指板上可以伸展多少的基本假设;只检查标准的6弦吉他调音。所提出的方法还检查不打算由吉他演奏或不可能由吉他演奏的音乐作品(例如,可能是交响乐团部分)的转录,采用用于增强音乐信息和用人工数据训练/测试系统的基本方法。结果显示了系统在初始和增强数据集上训练时可以实现的有趣方面,表明使用增强数据的训练即使在简单的情况下(例如单声道)也可以提高性能。研究结果还指出了不足之处,并就可能的改进提出了有益的结论。
摘要:Guitar tablature transcription consists in deducing the string and the fret number on which each note should be played to reproduce the actual musical part. This assignment should lead to playable string-fret combinations throughout the entire track and, in general, preserve parsimonious motion between successive combinations. Throughout the history of guitar playing, specific chord fingerings have been developed across different musical styles that facilitate common idiomatic voicing combinations and motion between them. This paper presents a method for assigning guitar tablature notation to a given MIDI-based musical part (possibly consisting of multiple polyphonic tracks), i.e. no information about guitar-idiomatic expressional characteristics is involved (e.g. bending etc.) The current strategy is based on machine learning and requires a basic assumption about how much fingers can stretch on a fretboard; only standard 6-string guitar tuning is examined. The proposed method also examines the transcription of music pieces that was not meant to be played or could not possibly be played by a guitar (e.g. potentially a symphonic orchestra part), employing a rudimentary method for augmenting musical information and training/testing the system with artificial data. The results present interesting aspects about what the system can achieve when trained on the initial and augmented dataset, showing that the training with augmented data improves the performance even in simple, e.g. monophonic, cases. Results also indicate weaknesses and lead to useful conclusions about possible improvements.


【16】MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
标题:MARS-Sep:多模式对齐的增强声音分离
链接:https://arxiv.org/abs/2510.10509

作者:Zihan Zhang, Xize Cheng, Zhennan Jiang, Dongjie Fu, Jingyuan Chen, Zhou Zhao, Tao Jin
摘要:通用声音分离面临着一个基本的不一致:针对低电平信号度量优化的模型通常会产生语义污染的输出,无法抑制来自声学相似源的感知显著干扰。为了弥合这一差距,我们引入了MARS-Sep,这是一个强化学习框架,将分离重新定义为决策。MARS-Sep不是简单地回归地面真值掩码,而是学习一种因子化的Beta掩码策略,该策略通过熵正则化和组相对优势归一化的裁剪信任区域代理进行优化。具体地说,我们从冻结的旧政策中采样掩码,重建波形,并使用裁剪的重要性比率更新当前政策,从而产生更稳定和更有效的样本学习。来自音频-文本-视觉编码器的多模式奖励直接激励与查询提示的语义一致性。我们进一步提出了一个渐进的对齐方案来微调这个编码器,提高其跨模态的可辨别性和提高奖励的忠诚度。在多个基准测试上进行的大量实验表明,文本、音频和图像查询分离的效果一致,信号度量和语义质量也有显著改善。我们的代码可在https://anonymous.4open.science/r/MARS-Sep上获得。声音分离示例可在https://mars-sep.github.io/上获得。
摘要:Universal sound separation faces a fundamental misalignment: models optimized for low-level signal metrics often produce semantically contaminated outputs, failing to suppress perceptually salient interference from acoustically similar sources. To bridge this gap, we introduce MARS-Sep, a reinforcement learning framework that reformulates separation as decision making. Instead of simply regressing ground-truth masks, MARS-Sep learns a factorized Beta mask policy that is optimized by a clipped trust-region surrogate with entropy regularization and group-relative advantage normalization. Concretely, we sample masks from a frozen old policy, reconstruct waveforms, and update the current policy using clipped importance ratios-yielding substantially more stable and sample-efficient learning. Multimodal rewards, derived from an audio-text-vision encoder, directly incentivize semantic consistency with query prompts. We further propose a progressive alignment scheme to fine-tune this encoder, boosting its cross-modal discriminability and improving reward faithfulness. Extensive experiments on multiple benchmarks demonstrate consistent gains in Text-, Audio-, and Image-Queried separation, with notable improvements in signal metrics and semantic quality. Our code is available at https://anonymous.4open.science/r/MARS-Sep. Sound separation samples are available at https://mars-sep.github.io/.


【17】Knowledge-Decoupled Functionally Invariant Path with Synthetic Personal Data for Personalized ASR
标题:具有合成个人数据的知识脱钩功能不变路径用于个性化ASB
链接:https://arxiv.org/abs/2510.10401

作者:Yue Gu, Zhihao Du, Ying Shi, Jiqing Han, Yongjun He
备注:Accepted for publication in IEEE Signal Processing Letters, 2025
摘要:使用大规模合成个人数据微调通用ASR模型可以增强ASR模型的个性化,但它在适应合成个人数据而不忘记真实知识以及适应个人数据而不忘记通用知识方面带来了挑战。考虑到功能不变路径(FIP)框架使模型自适应,同时保留先验知识,在这封信中,我们引入FIP合成数据增强个性化ASR模型。然而,当应用FIP同时在所有三种类型的数据上训练模型时,该模型仍然难以平衡综合、个性化和通用知识的学习。为了解耦这个学习过程,并进一步解决上述两个挑战,我们集成了一个门控参数隔离策略到FIP,并提出了一个知识解耦的功能不变路径(KDFIP)框架,它存储在单独的模块通用和个性化的知识,并适用于FIP顺序。具体而言,KDFIP使个性化模块适应合成和真实的个人数据,使通用模块适应通用数据。这两个模块都是沿着个性化不变的路径更新的,它们的输出通过门控机制动态融合。通过增强合成数据,KDFIP在目标说话人上实现了29.38%的相对字符错误率降低,并保持了与未适应ASR基线相当的泛化性能。
摘要:Fine-tuning generic ASR models with large-scale synthetic personal data can enhance the personalization of ASR models, but it introduces challenges in adapting to synthetic personal data without forgetting real knowledge, and in adapting to personal data without forgetting generic knowledge. Considering that the functionally invariant path (FIP) framework enables model adaptation while preserving prior knowledge, in this letter, we introduce FIP into synthetic-data-augmented personalized ASR models. However, the model still struggles to balance the learning of synthetic, personalized, and generic knowledge when applying FIP to train the model on all three types of data simultaneously. To decouple this learning process and further address the above two challenges, we integrate a gated parameter-isolation strategy into FIP and propose a knowledge-decoupled functionally invariant path (KDFIP) framework, which stores generic and personalized knowledge in separate modules and applies FIP to them sequentially. Specifically, KDFIP adapts the personalized module to synthetic and real personal data and the generic module to generic data. Both modules are updated along personalization-invariant paths, and their outputs are dynamically fused through a gating mechanism. With augmented synthetic data, KDFIP achieves a 29.38% relative character error rate reduction on target speakers and maintains comparable generalization performance to the unadapted ASR baseline.


【18】MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
标题:MRSAaudio:具有细化注释的大规模多模式记录空间音频数据集
链接:https://arxiv.org/abs/2510.10396

作者:Wenxiang Guo, Changhao Pan, Zhiyuan Zhu, Xintong Hu, Yu Zhang, Li Tang, Rui Yang, Han Wang, Zongbao Zhang, Yuhan Wang, Yixuan Chen, Hankun Xu, Ke Xu, Pengfei Fan, Zhetao Chen, Yanhao Yu, Qiange Huang, Fei Wu, Zhou Zhao
备注:24 pages
摘要:人类依靠多感官整合来感知空间环境,其中听觉线索使声源定位在三维空间中。尽管空间音频在VR/AR等沉浸式技术中发挥着关键作用,但大多数现有的多模态数据集仅提供单声道音频,这限制了空间音频生成和理解的发展。为了应对这些挑战,我们引入了MRSAudio,这是一个大规模的多模态空间音频数据集,旨在推进空间音频理解和生成方面的研究。MRSAudio涵盖四个不同的组件:MRSLife,MRSSpeech,MRSMusic和MRSSing,涵盖各种真实场景。该数据集包括同步的双耳和立体混响音频、离心和自我中心视频、运动轨迹和细粒度注释,如转录、音素边界、歌词、乐谱和提示。为了展示MRSAudio的实用性和多功能性,我们建立了五个基本任务:音频空间化,空间文本到语音,空间歌唱声音合成,空间音乐生成和声音事件定位和检测。结果表明,MRSAudio能够实现高质量的空间建模,并支持广泛的空间音频研究。演示和数据集访问可在https://mrsaudio.github.io上获得。
摘要:Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audio, which limits the development of spatial audio generation and understanding. To address these challenges, we introduce MRSAudio, a large-scale multimodal spatial audio dataset designed to advance research in spatial audio understanding and generation. MRSAudio spans four distinct components: MRSLife, MRSSpeech, MRSMusic, and MRSSing, covering diverse real-world scenarios. The dataset includes synchronized binaural and ambisonic audio, exocentric and egocentric video, motion trajectories, and fine-grained annotations such as transcripts, phoneme boundaries, lyrics, scores, and prompts. To demonstrate the utility and versatility of MRSAudio, we establish five foundational tasks: audio spatialization, and spatial text to speech, spatial singing voice synthesis, spatial music generation and sound event localization and detection. Results show that MRSAudio enables high-quality spatial modeling and supports a broad range of spatial audio research. Demos and dataset access are available at https://mrsaudio.github.io.


【19】ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis
标题:ProGress:通过图形扩散和分层音乐分析的结构化音乐生成
链接:https://arxiv.org/abs/2510.10249

作者:Stephen Ni-Hahn, Chao Péter Yang, Mingchen Ma, Cynthia Rudin, Simon Mak, Yue Jiang
摘要:用于音乐生成的人工智能(AI)正在经历快速发展,最近的符号模型利用了复杂的深度学习和扩散模型算法。现有模型的一个缺点是它们缺乏结构凝聚力,特别是在和声旋律结构上。此外,这些现有的模型在本质上主要是“黑盒”的,并且在音乐上不可解释。本文通过一个新的生成音乐框架,结合申克分析(SchA)的概念与扩散建模框架,解决了这些限制。这个框架,我们称之为ProGress(延长增强的DiGress),采用了最先进的离散扩散深度模型(特别是Vignac等人的DiGress模型,2023年)的可解释和结构化的音乐生成。具体来说,我们的贡献包括1)新的适应DiGress模型的音乐生成,2)一个新的SchA启发的短语融合方法,和3)一个框架,允许用户控制生成过程的各个方面,以创建连贯的音乐作品。来自人体实验的结果表明优于现有的最先进的方法的性能。
摘要:Artificial Intelligence (AI) for music generation is undergoing rapid developments, with recent symbolic models leveraging sophisticated deep learning and diffusion model algorithms. One drawback with existing models is that they lack structural cohesion, particularly on harmonic-melodic structure. Furthermore, such existing models are largely "black-box" in nature and are not musically interpretable. This paper addresses these limitations via a novel generative music framework that incorporates concepts of Schenkerian analysis (SchA) in concert with a diffusion modeling framework. This framework, which we call ProGress (Prolongation-enhanced DiGress), adapts state-of-the-art deep models for discrete diffusion (in particular, the DiGress model of Vignac et al., 2023) for interpretable and structured music generation. Concretely, our contributions include 1) novel adaptations of the DiGress model for music generation, 2) a novel SchA-inspired phrase fusion methodology, and 3) a framework allowing users to control various aspects of the generation process to create coherent musical compositions. Results from human experiments suggest superior performance to existing state-of-the-art methods.


【20】Peransformer: Improving Low-informed Expressive Performance Rendering with Score-aware Discriminator
标题:Peransformer:使用分数感知识别器改善低信息表达性能渲染
链接:https://arxiv.org/abs/2510.10175

作者:Xian He, Wei Zeng, Ye Wang
备注:6 pages, 3 figures, accepted by APSIPA ASC 2025
摘要:高度信息化的表现力渲染(EPR)系统将具有丰富音乐注释的乐谱转换为类似人类的表现力渲染文件。虽然这些系统已经取得了可喜的成果,详细的乐谱的可用性是有限的,相比于MIDI文件,并不灵活地使用数字音频工作站(MIDI)。低信息EPR系统的最新进展提供了一个更容易获得的替代方案,直接利用得分派生的ESTA作为输入,但这些系统往往表现出次优的性能。与此同时,现有的作品进行评估与不同的自动度量和数据格式,阻碍直接客观的EPR系统之间的比较。在这项研究中,我们介绍Peransformer,变压器为基础的低知情的EPR系统,旨在弥合低知情和高知情的EPR系统之间的差距。我们的方法结合了一个分数感知的数据库,它利用了底层的分数派生的数据库文件,并在一个分数到性能配对,注意到注意对齐的数据库数据集上进行训练。实验结果表明,Peransformer实现国家的最先进的性能之间的低知情的系统,通过主观评价验证。此外,我们扩展了现有的EPR系统的自动评估指标,并引入了广义EPR指标(GEM),使EPR系统之间的比较更直接,准确和可靠。
摘要:Highly-informed Expressive Performance Rendering (EPR) systems transform music scores with rich musical annotations into human-like expressive performance MIDI files. While these systems have achieved promising results, the availability of detailed music scores is limited compared to MIDI files and are less flexible to work with using a digital audio workstation (DAW). Recent advancements in low-informed EPR systems offer a more accessible alternative by directly utilizing score-derived MIDI as input, but these systems often exhibit suboptimal performance. Meanwhile, existing works are evaluated with diverse automatic metrics and data formats, hindering direct objective comparisons between EPR systems. In this study, we introduce Peransformer, a transformer-based low-informed EPR system designed to bridge the gap between low-informed and highly-informed EPR systems. Our approach incorporates a score-aware discriminator that leverages the underlying score-derived MIDI files and is trained on a score-to-performance paired, note-to-note aligned MIDI dataset. Experimental results demonstrate that Peransformer achieves state-of-the-art performance among low-informed systems, as validated by subjective evaluations. Furthermore, we extend existing automatic evaluation metrics for EPR systems and introduce generalized EPR metrics (GEM), enabling more direct, accurate, and reliable comparisons across EPR systems.


【21】Chord Colourizer: A Near Real-Time System for Visualizing Musical Key
标题:Chord Colorurizer:一个近实时的音乐键可视化系统
链接:https://arxiv.org/abs/2510.10173

作者:Paul Haimes
备注:Author copy. This paper is in press for presentation at ADADA 2025. Please cite as: Haimes, P. (in press). Chord Colourizer: A near real-time system for visualizing musical key. In Proceedings of the 23rd International Conference of Asia Digital Art and Design Association (ADADA)
摘要:本文介绍了和弦着色器,一个近实时的系统,检测音频信号的音乐键,并通过一个新的图形用户界面(GUI)直观地表示它。该系统基于艾萨克·牛顿的原始色轮为音符分配颜色,保留了音高和色调之间的历史联系,并集成了Arduino控制的LED显示屏,使用3D打印的星形扩散器提供物理环境媒体表示。该方法采用恒定Q变换(CQT)色度特征进行和弦估计和可视化,然后进行基于阈值的滤波和音调增强以隔离根音、第三和第五音。为每个检测计算置信度得分以确保可靠性,并且仅可视化具有中等到非常强的确定性的和弦。图形界面动态更新颜色编码的键盘布局,而LED显示屏通过空间反馈提供相同的颜色信息。这种多模式系统增强了用户与和声内容的交互,为教育和艺术表演提供了创新的可能性。其局限性包括轻微的延迟和无法检测扩展和弦,未来的开发将旨在通过精细滤波,自适应阈值以及支持更复杂的和声(如七度和增强和弦)来解决。未来的工作还将探索与其他可视化风格的集成,以及音频分析库的比较,以提高检测速度和精度。计划还包括正式的用户测试,以评估感知,可用性和跨文化的解释颜色音高映射。
摘要:This paper introduces Chord Colourizer, a near real-time system that detects the musical key of an audio signal and visually represents it through a novel graphical user interface (GUI). The system assigns colours to musical notes based on Isaac Newton's original colour wheel, preserving historical links between pitch and hue, and also integrates an Arduino-controlled LED display using 3D-printed star-shaped diffusers to offer a physical ambient media representation. The method employs Constant-Q Transform (CQT) chroma features for chord estimation and visualization, followed by threshold-based filtering and tonal enhancement to isolate the root, third, and fifth. A confidence score is computed for each detection to ensure reliability, and only chords with moderate to very strong certainty are visualized. The graphical interface dynamically updates a colour-coded keyboard layout, while the LED display provides the same colour information via spatial feedback. This multi-modal system enhances user interaction with harmonic content, offering innovative possibilities for education and artistic performance. Limitations include slight latency and the inability to detect extended chords, which future development will aim to address through refined filtering, adaptive thresholds, and support for more complex harmonies such as sevenths and augmented chords. Future work will also explore integration with alternative visualization styles, and the comparison of audio analysis libraries to improve detection speed and precision. Plans also include formal user testing to evaluate perception, usability, and cross-cultural interpretations of colour-pitch mappings.


【22】Matchmaker: An Open-source Library for Real-time Piano Score Following and Systematic Evaluation
标题:Matchmaker:实时钢琴乐谱跟踪和系统评估的开源库
链接:https://arxiv.org/abs/2510.10087

作者:Jiyun Park, Carlos Cancino-Chacón, Suhit Chiruthapudi, Juhan Nam
备注:In Proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR), 2025
摘要:实时音乐对齐,也称为乐谱跟随,是一项具有悠久历史的基本MIR任务,对于许多交互式应用程序至关重要。尽管它的重要性,还没有一个统一的开放框架比较模型,主要是由于实时处理和语言或系统依赖的实现的固有的复杂性。此外,与现有MIR环境的兼容性低,使得难以使用近年来可用的大型数据集开发基准。虽然基于既定方法的新研究(例如,动态规划、概率模型),但大多数评价仅在同一系列内或在小的测试数据集上比较模型。本文介绍了Matchmaker,这是一个用于实时音乐对齐的开源Python库,易于使用,并与现代MIR库兼容。利用这一点,我们系统地比较沿两个方面的方法:音乐表示和对齐方法。我们在来自(n)ASAP、Batik和Vienna4x22数据集的大型钢琴独奏音乐测试集上评估了我们的方法,并使用了一套全面的指标来确保稳健的评估。我们的工作旨在为分数跟踪研究建立一个基准框架,同时提供一个实用的工具,开发人员可以很容易地集成到他们的应用程序。
摘要:Real-time music alignment, also known as score following, is a fundamental MIR task with a long history and is essential for many interactive applications. Despite its importance, there has not been a unified open framework for comparing models, largely due to the inherent complexity of real-time processing and the language- or system-dependent implementations. In addition, low compatibility with the existing MIR environment has made it difficult to develop benchmarks using large datasets available in recent years. While new studies based on established methods (e.g., dynamic programming, probabilistic models) have emerged, most evaluations compare models only within the same family or on small sets of test data. This paper introduces Matchmaker, an open-source Python library for real-time music alignment that is easy to use and compatible with modern MIR libraries. Using this, we systematically compare methods along two dimensions: music representations and alignment methods. We evaluated our approach on a large test set of solo piano music from the (n)ASAP, Batik, and Vienna4x22 datasets with a comprehensive set of metrics to ensure robust assessment. Our work aims to establish a benchmark framework for score-following research while providing a practical tool that developers can easily integrate into their applications.


【23】Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model
标题:用互信息正规生成模型改进语音情感识别
链接:https://arxiv.org/abs/2510.10078

作者:Chung-Soo Ahn, Rajib Rana, Sunil Sivadas, Carlos Busso, Jagath C. Rajapakse
摘要:尽管语音情感识别(SER)的研究已经取得了进展,但由于深度学习方法的存在,它仍然需要从大量质量标记的训练数据中获取输入。数据增强方法已被尝试来缓解这个问题,生成模型最近在其中显示出成功。我们提出了一个数据增强框架,它是由跨模态信息传输和互信息正则化的辅助。基于互信息的度量可以用作质量的指示符。此外,我们扩大了这种数据增强范围的多模态输入,由于相互信息确保模态之间的依赖性。我们的框架在三个基准数据集上进行了测试:IEMOCAP,MSP-IMPROV和MSP-Podcast。该实现被设计为生成输入特征,这些特征被馈送到最后一层用于情感分类。我们的框架提高了对现有作品的情感预测的性能。此外,我们发现我们的框架能够在没有任何跨模态信息的情况下生成新的输入。
摘要:Although speech emotion recognition (SER) research has been advanced, thanks to deep learning methods, it still suffers from obtaining inputs from large quality-labelled training data. Data augmentation methods have been attempted to mitigate this issue, generative models have shown success among them recently. We propose a data augmentation framework that is aided by cross-modal information transfer and mutual information regularization. Mutual information based metric can serve as an indicator for the quality. Furthermore, we expand this data augmentation scope to multimodal inputs, thanks to mutual information ensureing dependency between modalities. Our framework was tested on three benchmark datasets: IEMOCAP, MSP-IMPROV and MSP-Podcast. The implementation was designed to generate input features that are fed into last layer for emotion classification. Our framework improved the performance of emotion prediction against existing works. Also, we discovered that our framework is able to generate new inputs without any cross-modal information.


【24】MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
标题:MTP-S2 UT:通过多令牌预测提高语音翻译质量
链接:https://arxiv.org/abs/2510.10003

作者:Jianjin Wang, Runsong Zhao, Xiaoqian Liu, Yuan Ge, Ziqiang Xu, Tong Xiao, Shengxiang Gao, Zhengtao Yu, Jingbo Zhu
备注:Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:当前的直接语音到语音翻译方法主要采用语音标记作为中间表示。但是,单个语音标记在语义上并不密集,因此我们通常需要多个标记来表达一个完整的语义单元。为了解决这一限制,我们将多标记预测(MTP)丢失引入到语音到单元翻译(S2 UT)模型中,使模型能够在每个位置预测多个后续标记,从而捕获更完整的语义并提高每个位置的信息密度。初始MTP实现在最后一层应用损失,这改善了输出表示,但启动信息丰富太晚。我们假设将信息丰富过程推进到中间层可以更早更有效地增强隐藏表示。因此,我们提出了MTP-S2 UT损失,将MTP损失应用于计算CTC损失的隐藏表示。实验表明,所有MTP损失变体一致地提高了S2 UT翻译的质量,其中MTP-S2 UT实现了最佳性能。
摘要:Current direct speech-to-speech translation methods predominantly employ speech tokens as intermediate representations. However, a single speech token is not dense in semantics, so we generally need multiple tokens to express a complete semantic unit. To address this limitation, we introduce multi-token prediction (MTP) loss into speech-to-unit translation (S2UT) models, enabling models to predict multiple subsequent tokens at each position, thereby capturing more complete semantics and enhancing information density per position. Initial MTP implementations apply the loss at the final layer, which improves output representation but initiates information enrichment too late. We hypothesize that advancing the information enrichment process to intermediate layers can achieve earlier and more effective enhancement of hidden representation. Consequently, we propose MTP-S2UT loss, applying MTP loss to hidden representation where CTC loss is computed. Experiments demonstrate that all MTP loss variants consistently improve the quality of S2UT translation, with MTP-S2UT achieving the best performance.


【25】Universal Discrete-Domain Speech Enhancement
标题:通用离散域语音增强
链接:https://arxiv.org/abs/2510.09974

作者:Fei Liu, Yang Ai, Ye-Xin Lu, Rui-Chen Zheng, Hui-Peng Du, Zhen-Hua Ling
摘要:在现实世界中,语音信号不可避免地会受到各种类型的干扰,这使得语音增强(SE)成为鲁棒语音处理的关键任务。然而,大多数现有的SE方法只处理有限范围的失真,如加性噪声,混响,或频带限制,而在多个同时失真下的SE的研究仍然有限。为了解决这一问题,本文提出了一种通用离散域语音识别模型(Universal Discrete-domain SE Model,简称UDSE),与基于回归的语音识别模型不同,UDSE将语音识别重新定义为一个离散域的分类任务,相反,预测由预训练的神经语音编解码器的残差矢量量化器(RVQ)量化的干净的离散令牌。具体地说,UDSE首先从降级的语音中提取全局特征。在这些全局特征的指导下,每个VQ的干净令牌预测遵循RVQ的规则,其中每个VQ的预测依赖于前一个的结果。最后,对来自所有VQ的预测的干净令牌进行解码以重建干净的语音波形。在训练过程中,UDSE模型采用教师强制策略,并通过交叉熵损失进行优化。实验结果表明,该模型可以有效地增强语音退化的各种传统和非传统的失真,加性噪声、混响、频带限制、限幅、相位失真和压缩失真,以及它们的组合。这些结果表明,优越的通用性和实用性的UDSE相比,先进的回归为基础的SE方法。
摘要:In real-world scenarios, speech signals are inevitably corrupted by various types of interference, making speech enhancement (SE) a critical task for robust speech processing. However, most existing SE methods only handle a limited range of distortions, such as additive noise, reverberation, or band limitation, while the study of SE under multiple simultaneous distortions remains limited. This gap affects the generalization and practical usability of SE methods in real-world environments.To address this gap, this paper proposes a novel Universal Discrete-domain SE model called UDSE.Unlike regression-based SE models that directly predict clean speech waveform or continuous features, UDSE redefines SE as a discrete-domain classification task, instead predicting the clean discrete tokens quantized by the residual vector quantizer (RVQ) of a pre-trained neural speech codec.Specifically, UDSE first extracts global features from the degraded speech. Guided by these global features, the clean token prediction for each VQ follows the rules of RVQ, where the prediction of each VQ relies on the results of the preceding ones. Finally, the predicted clean tokens from all VQs are decoded to reconstruct the clean speech waveform. During training, the UDSE model employs a teacher-forcing strategy, and is optimized with cross-entropy loss. Experimental results confirm that the proposed UDSE model can effectively enhance speech degraded by various conventional and unconventional distortions, e.g., additive noise, reverberation, band limitation, clipping, phase distortion, and compression distortion, as well as their combinations. These results demonstrate the superior universality and practicality of UDSE compared to advanced regression-based SE methods.


【26】Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications
标题:利用复值CNN进行阶段感知深度学习用于音频信号应用
链接:https://arxiv.org/abs/2510.09926

作者:Naman Agrawal
摘要:本研究探讨了复值卷积神经网络(CVCNN)在音频信号处理中的设计和应用,重点是保留和利用实值网络中经常被忽视的相位信息。我们首先介绍CVCNN的基本理论概念,包括复卷积、池化层、基于Wirtinger的微分和各种复值激活函数。这些都是通过训练技术的关键调整来补充的,包括复杂的批量标准化和权重初始化方案,以确保训练动态的稳定性。经验评价分三个阶段进行。首先,CVCNN在标准图像数据集上进行基准测试,即使在合成复杂扰动下,它们也表现出与实值CNN竞争的性能。虽然我们的重点是音频信号处理,但我们首先在图像数据集上评估CVCNN,以建立基线性能并验证训练稳定性,然后再将其应用于音频任务。在第二个实验中,我们专注于使用梅尔频率倒谱系数(MFCC)的音频分类。在实值MFCC上训练的CVCNN的性能略优于真实CNN,而在输入工作流中保留阶段突出了在没有架构修改的情况下利用阶段的挑战。最后,第三个实验通过边缘加权引入GNN来建模相位信息,其中包含相位在二进制和多类流派分类中都产生了可测量的增益。这些结果强调了复值架构的表达能力,并确认相位在音频处理应用中是一个有意义和可利用的功能。虽然目前的方法显示出了希望,特别是对于像心形线这样的激活,但相位感知设计的未来进展对于利用神经网络中复杂表示的潜力至关重要。
摘要:This study explores the design and application of Complex-Valued Convolutional Neural Networks (CVCNNs) in audio signal processing, with a focus on preserving and utilizing phase information often neglected in real-valued networks. We begin by presenting the foundational theoretical concepts of CVCNNs, including complex convolutions, pooling layers, Wirtinger-based differentiation, and various complex-valued activation functions. These are complemented by critical adaptations of training techniques, including complex batch normalization and weight initialization schemes, to ensure stability in training dynamics. Empirical evaluations are conducted across three stages. First, CVCNNs are benchmarked on standard image datasets, where they demonstrate competitive performance with real-valued CNNs, even under synthetic complex perturbations. Although our focus is audio signal processing, we first evaluate CVCNNs on image datasets to establish baseline performance and validate training stability before applying them to audio tasks. In the second experiment, we focus on audio classification using Mel-Frequency Cepstral Coefficients (MFCCs). CVCNNs trained on real-valued MFCCs slightly outperform real CNNs, while preserving phase in input workflows highlights challenges in exploiting phase without architectural modifications. Finally, a third experiment introduces GNNs to model phase information via edge weighting, where the inclusion of phase yields measurable gains in both binary and multi-class genre classification. These results underscore the expressive capacity of complex-valued architectures and confirm phase as a meaningful and exploitable feature in audio processing applications. While current methods show promise, especially with activations like cardioid, future advances in phase-aware design will be essential to leverage the potential of complex representations in neural networks.


eess.AS音频处理


【1】ILD-VIT: A Unified Vision Transformer Architecture for Detection of Interstitial Lung Disease from Respiratory Sounds
标题:ILD-VIT:一种用于从呼吸音检测间质性肺疾病的统一视觉Transformer架构
链接:https://arxiv.org/abs/2510.11458

作者:Soubhagya Ranjan Hota, Arka Roy, Udit Satija
摘要:间质性肺病(ILD)是一组限制性慢性肺部疾病,通过引起肺内不可逆变化(如纤维化、实质瘢痕形成等)损害氧获取。ILD疾病通常通过各种临床模式(如肺量计、高分辨率肺成像技术、呼吸音(RS)等)进行诊断。在这封信中,我们开发了一种新的基于Vision Transformer(VIT)的深度学习框架,即ILD-VIT,以使用RS记录来检测ILD状况。所提出的框架包括三个主要阶段:预处理,梅尔频谱图提取和分类使用建议的VIT架构使用梅尔频谱图图像补丁。使用公开的BRACETS和KAUH数据库的实验结果表明,我们提出的ILD-VIT实现了准确性,灵敏度和特异性分别为84.86%,82.67%和86.91%,独立于受试者的盲测。所提出的框架在Raspberry-pi-4微控制器上的成功板载植入表明其在实际临床场景中作为ILD筛查的独立临床系统的潜力。
摘要:Interstitial lung disease (ILD) represents a group of restrictive chronic pulmonary diseases that impair oxygen acquisition by causing irreversible changes in the lungs such as fibrosis, scarring of parenchyma, etc. ILD conditions are often diagnosed by various clinical modalities such as spirometry, high-resolution lung imaging techniques, crackling respiratory sounds (RSs), etc. In this letter, we develop a novel vision transformer (VIT)-based deep learning framework namely, ILD-VIT, to detect the ILD condition using the RS recordings. The proposed framework comprises three major stages: pre-processing, mel spectrogram extraction, and classification using the proposed VIT architecture using the mel spectrogram image patches. Experimental results using the publicly available BRACETS and KAUH databases show that our proposed ILD-VIT achieves an accuracy, sensitivity, and specificity of 84.86%, 82.67%, and 86.91%, respectively, for subject-independent blind testing. The successful onboard implantation of the proposed framework on a Raspberry-pi-4 microcontroller indicates its potential as a standalone clinical system for ILD screening in a real clinical scenario.


【2】Dynamically Slimmable Speech Enhancement Network with Metric-Guided Training
标题:具有公制引导训练的动态可精简语音增强网络
链接:https://arxiv.org/abs/2510.11395

作者:Haixin Zhao, Kaixuan Yang, Nilesh Madhu
备注:Preprint version of a paper under review at ICASSP2026
摘要:为了进一步降低轻量级语音增强模型的复杂度,我们引入了一种基于门控的动态可精简网络(DSN)。DSN包括静态和动态组件。对于架构独立的适用性,我们引入了针对常用组件的不同动态结构,即分组递归神经网络单元,多头注意力,卷积和全连接层。策略模块根据输入信号质量以逐帧分辨率自适应地管理动态部分的使用,从而控制计算负载。我们进一步提出度量指导训练(MGT)明确指导政策模块评估输入语音质量。实验结果表明,DSN实现了可比的增强性能,在仪器指标的最先进的轻量级基线,而平均只使用73%的计算负载。动态组件使用率的评估表明,MGT-DSN可以适当地分配网络资源,根据输入信号失真的严重程度。
摘要:To further reduce the complexity of lightweight speech enhancement models, we introduce a gating-based Dynamically Slimmable Network (DSN). The DSN comprises static and dynamic components. For architecture-independent applicability, we introduce distinct dynamic structures targeting the commonly used components, namely, grouped recurrent neural network units, multi-head attention, convolutional, and fully connected layers. A policy module adaptively governs the use of dynamic parts at a frame-wise resolution according to the input signal quality, controlling computational load. We further propose Metric-Guided Training (MGT) to explicitly guide the policy module in assessing input speech quality. Experimental results demonstrate that the DSN achieves comparable enhancement performance in instrumental metrics to the state-of-the-art lightweight baseline, while using only 73% of its computational load on average. Evaluations of dynamic component usage ratios indicate that the MGT-DSN can appropriately allocate network resources according to the severity of input signal distortion.


【3】Phase Aware Ear-Conditioned Learning for Multi-Channel Binaural Speaker Separation
标题:基于相位感知耳条件学习的多通道双耳说话人分离
链接:https://arxiv.org/abs/2510.11366

作者:Ruben Johnson Robert Jeremiah, Peyman Goli, Steven van de Par
摘要:在混响环境中分离竞争语音需要在保持分离效率的同时保留空间线索的模型。我们提出了一个使用八个麦克风(PEASE-8)的相位感知耳调节扬声器分离网络,该网络消耗复杂的STFT并直接将原始STFT输入引入早期解码器层,绕过整个编码器路径以改善重建。该模型使用基于SI-SDR的目标针对直接路径耳目标进行端到端训练,在固定方位角中联合执行两个扬声器的分离和去混响,消除了对排列不变训练的需要。在消声、混响和噪声条件下的空间化双扬声器混音中,PEASE-8提供了强大的分离和可懂度。在混响环境中,它实现了12.37 dB SI-SDR,0.87 STOI和1.86 PESQ(T60 = 0.6 s),同时在消声条件下保持竞争力。
摘要:Separating competing speech in reverberant environments requires models that preserve spatial cues while maintaining separation efficiency. We present a Phase-aware Ear-conditioned speaker Separation network using eight microphones (PEASE-8) that consumes complex STFTs and directly introduces a raw-STFT input to the early decoder layer, bypassing the entire encoder pathway to improve reconstruction. The model is trained end-to-end with an SI-SDR-based objective against direct-path ear targets, jointly performing separation and dereverberation for two speakers in a fixed azimuth, eliminating the need for permutation invariant training. On spatialized two-speaker mixtures spanning anechoic, reverberant, and noisy conditions, PEASE-8 delivers strong separation and intelligibility. In reverberant environments, it achieves 12.37 dB SI-SDR, 0.87 STOI, and 1.86 PESQ at T60 = 0.6 s, while remaining competitive under anechoic conditions.


【4】Perceptual Compensation of Ambisonics Recordings for Reproduction in Room
标题:室内复制的立体声立体声录音的感知补偿
链接:https://arxiv.org/abs/2510.10883

作者:Ali Fallah, Shun Nakamura, Steven van de Par
备注:The manuscript was submitted to the JASA and is under review
摘要:高保真度立体声是一种用于准确地捕获和渲染声场的方法,假设回放房间的声学不会显著影响声场。然而,在实践中,回放房间的声学可能导致声音质量的显著劣化。我们提出了一种基于高保真度立体声的记录和渲染方法,该方法利用感知激励的方法来补偿回放房间的混响。在球面谐波(SH)域中记录的直达和混响声场分量被频谱和空间补偿以保留相关的听觉线索,包括直达声音的到达方向、直达和混响声音分量的频谱能量以及跨每个听觉频带的耳间相干性(IC)。与传统的高保真度立体声响复制相比,灵活数量的高保真度立体声响复制声道可以用于音频渲染。听力测试结果表明,所提出的方法提供了一个感知准确的渲染原始记录的声场,优于传统的高保真度立体声没有补偿,甚至理想的高保真度立体声渲染在模拟消声室。此外,坐在扬声器阵列的中心的听众的主观评价表明,该方法仍然是强大的头部旋转和微小的位移。
摘要:Ambisonics is a method for capturing and rendering a sound field accurately, assuming that the acoustics of the playback room does not significantly influence the sound field. However, in practice, the acoustics of the playback room may lead to a noticeable degradation in sound quality. We propose a recording and rendering method based on Ambisonics that utilizes a perceptually-motivated approach to compensate for the reverberation of the playback room. The recorded direct and reverberant sound field components in the spherical harmonics (SHs) domain are spectrally and spatially compensated to preserve the relevant auditory cues including the direction of arrival of the direct sound, the spectral energy of the direct and reverberant sound components, and the Interaural Coherence (IC) across each auditory band. In contrast to the conventional Ambisonics, a flexible number of Ambisonics channels can be used for audio rendering. Listening test results show that the proposed method provides a perceptually accurate rendering of the originally recorded sound field, outperforming both conventional Ambisonics without compensation and even ideal Ambisonics rendering in a simulated anechoic room. Additionally, subjective evaluations of listeners seated at the center of the loudspeaker array demonstrate that the method remains robust to head rotation and minor displacements.


【5】Automatic Music Sample Identification with Multi-Track Contrastive Learning
标题:多轨对比学习自动音乐样本识别
链接:https://arxiv.org/abs/2510.11507

作者:Alain Riou, Joan Serrà, Yuki Mitsufuji
摘要:采样,即重复使用现有音轨片段以创建新音乐内容的技术,是现代音乐制作中非常常见的做法。在本文中,我们解决了具有挑战性的任务,自动样本识别,也就是说,检测这样的采样内容和检索的材料从它的起源。为此,我们采用了一种自监督学习方法,该方法利用多轨道数据集来创建人工混合的正对,并设计了一种新的对比学习目标。我们表明,这种方法显着优于以前的国家的最先进的基线,这是强大的各种流派,以及规模时,增加参考数据库中的噪声歌曲的数量。此外,我们广泛分析了我们的培训管道的不同组成部分的贡献,并强调,特别是,需要高质量的分离干这项任务。
摘要:Sampling, the technique of reusing pieces of existing audio tracks to create new music content, is a very common practice in modern music production. In this paper, we tackle the challenging task of automatic sample identification, that is, detecting such sampled content and retrieving the material from which it originates. To do so, we adopt a self-supervised learning approach that leverages a multi-track dataset to create positive pairs of artificial mixes, and design a novel contrastive learning objective. We show that such method significantly outperforms previous state-of-the-art baselines, that is robust to various genres, and that scales well when increasing the number of noise songs in the reference database. In addition, we extensively analyze the contribution of the different components of our training pipeline and highlight, in particular, the need for high-quality separated stems for this task.


【6】Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
标题:扩散-链接:弥合音频-文本情态差距的扩散概率模型
链接:https://arxiv.org/abs/2510.11330

作者:KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo, Joon Son Chung
备注:5 pages. Submitted to IEEE ICASSP 2026
摘要:对比音频语言预训练产生强大的联合表示,但持续的音频文本模态差距限制了多模态编码器与大型语言模型(LLM)耦合的好处。我们提出了扩散链接,一个基于扩散的模态桥接模块,生成映射到文本嵌入分布的音频嵌入。该模块在冻结的多模态编码器的输出嵌入处进行训练,并实现为具有三个残留MLP块的轻量级网络。为了评估扩散链接对多模态编码器-LLM耦合的影响,我们对自动音频字幕(AAC)进行了评估;据我们所知,这是基于扩散的模态桥接到AAC的第一个应用。我们报告两个结果。(1)模式差距分析:在相似性和几何标准上,Diffusion-Link在现有的基于扩散的方法中最大程度地减少了模态间隙,并且显示了音频嵌入向文本分布的集体迁移。(2)下游AAC:将Diffusion-Link附加到相同的多模态LLM基线,在没有外部知识的情况下,在zero-shot和完全监督字幕中实现了AudioCaps的最新技术水平,相对增益分别高达52.5%和7.5%。这些研究结果表明,关闭模态差距是多模态编码器和LLM之间有效耦合的关键,基于扩散的模态桥接提供了一个有前途的方向超越知识检索为中心的设计。代码将在接受后发布https://github.com/DevKiHyun/Diffusion-Link
摘要:Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Diffusion-Link, a diffusion-based modality-bridging module that generatively maps audio embeddings into the text-embedding distribution. The module is trained at the output embedding from the frozen multimodal encoder and implemented as a lightweight network with three residual MLP blocks. To assess the effect of Diffusion-Link on multimodal encoder-LLM coupling, we evaluate on Automatic Audio Captioning (AAC); to our knowledge, this is the first application of diffusion-based modality bridging to AAC. We report two results. (1) Modality-gap analysis: on similarity and geometric criteria, Diffusion-Link reduces the modality gap the most among prior diffusion-based methods and shows a collective migration of audio embeddings toward the text distribution. (2) Downstream AAC: attaching Diffusion-Link to the same multimodal LLM baseline achieves state-of-the-art on AudioCaps in both zero-shot and fully supervised captioning without external knowledge, with relative gains up to 52.5% and 7.5%, respectively. These findings show that closing the modality gap is pivotal for effective coupling between multimodal encoders and LLMs, and diffusion-based modality bridging offers a promising direction beyond knowledge-retrieval-centric designs. Code will be released upon acceptance https://github.com/DevKiHyun/Diffusion-Link


【7】Efficient Edge Test-Time Adaptation via Latent Feature Coordinate Correction
标题:通过潜在特征坐标修正进行高效的边缘测试时间自适应
链接:https://arxiv.org/abs/2510.11068

作者:Xinyu Luo, Jie Liu, Kecheng Chen, Junyi Yang, Bo Ding, Arindam Basu, Haoliang Li
备注:Under review
摘要:由于计算资源有限和分布变化,边缘设备面临着重大挑战,这使得高效和适应性强的机器学习变得至关重要。现有的测试时间自适应(TTA)方法通常依赖于基于梯度的优化或批处理,由于其依赖于反向传播和高计算需求,这些方法本质上不适合资源受限的边缘场景。无故障的替代方案解决了这些问题,但往往受到有限的学习能力,缺乏灵活性,或强加架构约束。为了克服这些限制,我们提出了一种新的单实例TTA方法为边缘设备(TED),它采用了前向只协调优化的主要子空间的潜在使用协方差矩阵自适应进化策略(CMA-ES)。通过更新紧凑的低维向量,TED不仅增强了输出置信度,而且还将潜在表示更接近潜在主子空间内的源潜在分布。这是在没有反向传播的情况下实现的,保持模型参数冻结,并以最小的内存和计算开销实现高效的无遗忘自适应。在ImageNet和Google Speech Commands系列数据集上进行的图像分类和关键字定位任务的实验表明,TED实现了最先进的性能,同时将计算复杂度降低了63倍,为现实世界的边缘应用提供了实用且可扩展的解决方案。此外,我们还成功地在ZYNQ-7020平台上部署了TED,证明了其在实际部署中对于资源受限的边缘设备的可行性和有效性。
摘要:Edge devices face significant challenges due to limited computational resources and distribution shifts, making efficient and adaptable machine learning essential. Existing test-time adaptation (TTA) methods often rely on gradient-based optimization or batch processing, which are inherently unsuitable for resource-constrained edge scenarios due to their reliance on backpropagation and high computational demands. Gradient-free alternatives address these issues but often suffer from limited learning capacity, lack flexibility, or impose architectural constraints. To overcome these limitations, we propose a novel single-instance TTA method tailored for edge devices (TED), which employs forward-only coordinate optimization in the principal subspace of latent using the covariance matrix adaptation evolution strategy (CMA-ES). By updating a compact low-dimensional vector, TED not only enhances output confidence but also aligns the latent representation closer to the source latent distribution within the latent principal subspace. This is achieved without backpropagation, keeping the model parameters frozen, and enabling efficient, forgetting-free adaptation with minimal memory and computational overhead. Experiments on image classification and keyword spotting tasks across the ImageNet and Google Speech Commands series datasets demonstrate that TED achieves state-of-the-art performance while $\textit{reducing computational complexity by up to 63 times}$, offering a practical and scalable solution for real-world edge applications. Furthermore, we successfully $\textit{deployed TED on the ZYNQ-7020 platform}$, demonstrating its feasibility and effectiveness for resource-constrained edge devices in real-world deployments.


【8】Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
标题:通过嵌入有效排序统一通用音频表示神经缩放定律中的变量
链接:https://arxiv.org/abs/2510.10948

作者:Xuyao Deng, Yanjie Sun, Yong Dou, Kele Xu
摘要:标度律深刻地塑造了我们对计算机视觉和自然语言处理中模型性能的理解,但它们在一般音频表示学习中的应用仍有待探索。一个关键的挑战在于一般音频表示的多因素性质-表示质量受到诸如音频长度、嵌入维度、模型深度、模型架构、数据量等变量的共同影响,其中许多是难以分离或分析表达的。在这项工作中,我们提出了一个系统的研究一般音频表示的缩放律,利用嵌入有效秩(RankMe)作为一个统一的度量,封装不同的变量对表示质量的影响。RankMe能够对音频嵌入进行无标签的信息理论量化,使我们能够在广泛的超参数空间中检查缩放行为,包括模型大小,训练数据量,计算预算,架构配置等。我们的实证研究结果揭示了RankMe和表示质量之间一致的幂律关系,这表明嵌入有效秩充当用于评估和预测音频表示学习中的模型性能的可靠代理。这项工作不仅验证了经典的缩放原则,一般的音频域的适用性,但也提供了一个理论上的接地和经验的强大的框架,用于指导未来的模型缩放策略在音频基础模型。
摘要:Scaling laws have profoundly shaped our understanding of model performance in computer vision and natural language processing, yet their application to general audio representation learning remains underexplored. A key challenge lies in the multifactorial nature of general audio representation-representation quality is jointly influenced by variables such as audio length, embedding dimensionality, model depth, model architecture, data volume, etc., many of which are difficult to isolate or express analytically. In this work, we present a systematic study of scaling laws for general audio representations by utilizing embedding effective rank (RankMe) as a unifying metric that encapsulates the impact of diverse variables on representation quality. RankMe enables a label-free, information-theoretic quantification of audio embeddings, allowing us to examine scaling behaviors across a wide hyper-parameter space, including model size, training data volume, computational budget, architectural configurations, etc. Our empirical findings reveal a consistent power-law relationship between RankMe and representation quality, suggesting that embedding effective rank serves as a reliable proxy for assessing and predicting model performance in audio representation learning. This work not only validates the applicability of classical scaling principles to the general audio domain but also offers a theoretically grounded and empirically robust framework for guiding future model scaling strategies in audio foundation models.


【9】Delayed 1T to 2H Phase Transition Upon Electrochemical Delithiation of LiMoS2
标题:LiMoS2电化学脱锂的延迟1T到2H相变
链接:https://arxiv.org/abs/2510.10911

作者:Yerin Hong, Juhwan Lim, Jinhong Min, Nishkarsh Agarwal, Robert Hovden, Ageeth A. Bol, Yiyang Li
摘要:二硫化钼(MoS 2)是一种广泛研究的层状材料,用于电子,光学和催化应用。它可以在范德华层之间容纳锂离子,这引发了半导体2 H相和金属1 T相之间的相变。虽然锂插入触发相变到1 T相,但电化学锂去除时的相行为没有得到解决。在这项工作中,我们进行单片电化学(脱)锂的二硫化钼使用微电极阵列。通过电化学电压分析和相关的拉曼光谱,我们表明,电化学循环和脱锂二硫化钼片最初仍然在1 T相。然而,在几天的过程中,它转变回化学稳定的2 H相。该结果解决了脱锂后的相变途径,并展示了电化学合成亚稳1 T-MoS 2相的能力。
摘要:Molybdenum disulfide (MoS2) is a widely studied layered material for electronic, optical, and catalytic applications. It can host lithium ions between the van der Waals layers, which triggers a phase transition between the semiconducting 2H phase and metallic 1T phase. While lithium insertion triggers a phase transition to the 1T phase, the phase behavior upon electrochemical lithium removal is not resolved. In this work, we conduct single-flake electrochemical (de)lithiation of MoS2 using microelectrode arrays. Through both electrochemical voltage analysis and correlative Raman spectroscopy, we show that an electrochemically cycled and delithiated MoS2 flake initially remains in the 1T phase. However, over the course of several days, it transitions back into the thermodynamically stable 2H phase. This result resolves the phase transformation pathway upon delithiation and showcases the ability to electrochemically synthesize the metastable 1T-MoS2 phase.


【10】Bhasha-Rupantarika: Algorithm-Hardware Co-design approach for Multilingual Neural Machine Translation
标题:Bhasha-Rupantarika:多语言神经机器翻译的框架-硬件协同设计方法
链接:https://arxiv.org/abs/2510.10676

作者:Mukul Lokhande, Tanushree Dewangan, Mohd Sharik Mansoori, Tejas Chaudhari, Akarsh J., Damayanti Lokhande, Adam Teman, Santosh Kumar Vishvakarma
摘要:本文介绍了Bhasha-Rupantarika,一个通过算法硬件协同设计为资源有限的设置量身定制的轻量级高效的多语言翻译系统。该方法研究了子八位精度级别(FP 8,INT 8,INT 4和FP 4)的模型部署,实验结果表明模型大小(FP 4)减少了4.1倍,推理速度加快了4.2倍,这与66个令牌/s的吞吐量增加(提高了4.8倍)相关。这强调了超低精度量化对于使用FPGA加速器在物联网设备中实时部署的重要性,从而实现与预期相当的性能。我们的评估涵盖了印度和国际语言之间的双向翻译,展示了其在低资源语言环境中的适应性。FPGA部署表明LUT减少了1.96倍,FF减少了1.65倍,与OPU相比,吞吐量提高了2.2倍,与HPTA相比提高了4.6倍。总的来说,该评估提供了一个可行的解决方案,该解决方案基于量化感知翻译以及适用于可部署多语言AI系统的硬件效率。整个代码[https://github.com/mukullokhande99/Bhasha-Rupantarika/]和可重复性数据集是公开的,便于研究人员快速集成和进一步开发。
摘要:This paper introduces Bhasha-Rupantarika, a light and efficient multilingual translation system tailored through algorithm-hardware codesign for resource-limited settings. The method investigates model deployment at sub-octet precision levels (FP8, INT8, INT4, and FP4), with experimental results indicating a 4.1x reduction in model size (FP4) and a 4.2x speedup in inference speed, which correlates with an increased throughput of 66 tokens/s (improvement by 4.8x). This underscores the importance of ultra-low precision quantization for real-time deployment in IoT devices using FPGA accelerators, achieving performance on par with expectations. Our evaluation covers bidirectional translation between Indian and international languages, showcasing its adaptability in low-resource linguistic contexts. The FPGA deployment demonstrated a 1.96x reduction in LUTs and a 1.65x decrease in FFs, resulting in a 2.2x enhancement in throughput compared to OPU and a 4.6x enhancement compared to HPTA. Overall, the evaluation provides a viable solution based on quantisation-aware translation along with hardware efficiency suitable for deployable multilingual AI systems. The entire codes [https://github.com/mukullokhande99/Bhasha-Rupantarika/] and dataset for reproducibility are publicly available, facilitating rapid integration and further development by researchers.


【11】ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis
标题:ProGress:通过图形扩散和分层音乐分析的结构化音乐生成
链接:https://arxiv.org/abs/2510.10249

作者:Stephen Ni-Hahn, Chao Péter Yang, Mingchen Ma, Cynthia Rudin, Simon Mak, Yue Jiang
摘要:用于音乐生成的人工智能(AI)正在快速发展,最近的符号模型利用了复杂的深度学习和扩散模型算法。现有模型的一个缺点是它们缺乏结构凝聚力,特别是在和声旋律结构上。此外,这些现有的模型在本质上主要是“黑盒”的,并且在音乐上不可解释。本文通过一个新的生成音乐框架,结合申克分析(SchA)的概念与扩散建模框架,解决了这些限制。这个框架,我们称之为ProGress(延长增强的DiGress),采用了最先进的离散扩散深度模型(特别是Vignac等人的DiGress模型,2023年)的可解释和结构化的音乐生成。具体来说,我们的贡献包括1)新的适应DiGress模型的音乐生成,2)一个新的SchA启发的短语融合方法,和3)一个框架,允许用户控制生成过程的各个方面,以创建连贯的音乐作品。来自人体实验的结果表明优于现有的最先进的方法的性能。
摘要:Artificial Intelligence (AI) for music generation is undergoing rapid developments, with recent symbolic models leveraging sophisticated deep learning and diffusion model algorithms. One drawback with existing models is that they lack structural cohesion, particularly on harmonic-melodic structure. Furthermore, such existing models are largely "black-box" in nature and are not musically interpretable. This paper addresses these limitations via a novel generative music framework that incorporates concepts of Schenkerian analysis (SchA) in concert with a diffusion modeling framework. This framework, which we call ProGress (Prolongation-enhanced DiGress), adapts state-of-the-art deep models for discrete diffusion (in particular, the DiGress model of Vignac et al., 2023) for interpretable and structured music generation. Concretely, our contributions include 1) novel adaptations of the DiGress model for music generation, 2) a novel SchA-inspired phrase fusion methodology, and 3) a framework allowing users to control various aspects of the generation process to create coherent musical compositions. Results from human experiments suggest superior performance to existing state-of-the-art methods.


【12】Peransformer: Improving Low-informed Expressive Performance Rendering with Score-aware Discriminator
标题:Peransformer:使用分数感知识别器改善低信息表达性能渲染
链接:https://arxiv.org/abs/2510.10175

作者:Xian He, Wei Zeng, Ye Wang
备注:6 pages, 3 figures, accepted by APSIPA ASC 2025
摘要:高度信息化的表现力渲染(EPR)系统将具有丰富音乐注释的乐谱转换为类似人类的表现力渲染文件。虽然这些系统已经取得了可喜的成果,详细的乐谱的可用性是有限的,相比于MIDI文件,并不灵活地使用数字音频工作站(MIDI)。低信息EPR系统的最新进展提供了一个更容易获得的替代方案,直接利用得分派生的ESTA作为输入,但这些系统往往表现出次优的性能。与此同时,现有的作品进行评估与不同的自动度量和数据格式,阻碍直接客观的EPR系统之间的比较。在这项研究中,我们介绍Peransformer,变压器为基础的低知情的EPR系统,旨在弥合低知情和高知情的EPR系统之间的差距。我们的方法结合了一个分数感知的数据库,它利用了底层的分数派生的数据库文件,并在一个分数到性能配对,注意到注意对齐的数据库数据集上进行训练。实验结果表明,Peransformer实现国家的最先进的性能之间的低知情的系统,通过主观评价验证。此外,我们扩展了现有的EPR系统的自动评估指标,并引入了广义EPR指标(GEM),使EPR系统之间的比较更直接,准确和可靠。
摘要:Highly-informed Expressive Performance Rendering (EPR) systems transform music scores with rich musical annotations into human-like expressive performance MIDI files. While these systems have achieved promising results, the availability of detailed music scores is limited compared to MIDI files and are less flexible to work with using a digital audio workstation (DAW). Recent advancements in low-informed EPR systems offer a more accessible alternative by directly utilizing score-derived MIDI as input, but these systems often exhibit suboptimal performance. Meanwhile, existing works are evaluated with diverse automatic metrics and data formats, hindering direct objective comparisons between EPR systems. In this study, we introduce Peransformer, a transformer-based low-informed EPR system designed to bridge the gap between low-informed and highly-informed EPR systems. Our approach incorporates a score-aware discriminator that leverages the underlying score-derived MIDI files and is trained on a score-to-performance paired, note-to-note aligned MIDI dataset. Experimental results demonstrate that Peransformer achieves state-of-the-art performance among low-informed systems, as validated by subjective evaluations. Furthermore, we extend existing automatic evaluation metrics for EPR systems and introduce generalized EPR metrics (GEM), enabling more direct, accurate, and reliable comparisons across EPR systems.


【13】Chord Colourizer: A Near Real-Time System for Visualizing Musical Key
标题:Chord Colorurizer:一个近实时的音乐键可视化系统
链接:https://arxiv.org/abs/2510.10173

作者:Paul Haimes
备注:Author copy. This paper is in press for presentation at ADADA 2025. Please cite as: Haimes, P. (in press). Chord Colourizer: A near real-time system for visualizing musical key. In Proceedings of the 23rd International Conference of Asia Digital Art and Design Association (ADADA)
摘要:本文介绍了和弦着色器,一个近实时的系统,检测音频信号的音乐键,并通过一个新的图形用户界面(GUI)直观地表示它。该系统基于艾萨克·牛顿的原始色轮为音符分配颜色,保留了音高和色调之间的历史联系,并集成了Arduino控制的LED显示屏,使用3D打印的星形扩散器提供物理环境媒体表示。该方法采用恒定Q变换(CQT)色度特征进行和弦估计和可视化,然后进行基于阈值的滤波和音调增强以隔离根音、第三和第五音。为每个检测计算置信度得分以确保可靠性,并且仅可视化具有中等到非常强的确定性的和弦。图形界面动态更新颜色编码的键盘布局,而LED显示屏通过空间反馈提供相同的颜色信息。这种多模式系统增强了用户与和声内容的交互,为教育和艺术表演提供了创新的可能性。其局限性包括轻微的延迟和无法检测扩展和弦,未来的开发将旨在通过精细滤波,自适应阈值以及支持更复杂的和声(如七度和增强和弦)来解决。未来的工作还将探索与其他可视化风格的集成,以及音频分析库的比较,以提高检测速度和精度。计划还包括正式的用户测试,以评估感知,可用性和跨文化的解释颜色音高映射。
摘要:This paper introduces Chord Colourizer, a near real-time system that detects the musical key of an audio signal and visually represents it through a novel graphical user interface (GUI). The system assigns colours to musical notes based on Isaac Newton's original colour wheel, preserving historical links between pitch and hue, and also integrates an Arduino-controlled LED display using 3D-printed star-shaped diffusers to offer a physical ambient media representation. The method employs Constant-Q Transform (CQT) chroma features for chord estimation and visualization, followed by threshold-based filtering and tonal enhancement to isolate the root, third, and fifth. A confidence score is computed for each detection to ensure reliability, and only chords with moderate to very strong certainty are visualized. The graphical interface dynamically updates a colour-coded keyboard layout, while the LED display provides the same colour information via spatial feedback. This multi-modal system enhances user interaction with harmonic content, offering innovative possibilities for education and artistic performance. Limitations include slight latency and the inability to detect extended chords, which future development will aim to address through refined filtering, adaptive thresholds, and support for more complex harmonies such as sevenths and augmented chords. Future work will also explore integration with alternative visualization styles, and the comparison of audio analysis libraries to improve detection speed and precision. Plans also include formal user testing to evaluate perception, usability, and cross-cultural interpretations of colour-pitch mappings.


【14】MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
标题:MTP-S2 UT:通过多令牌预测提高语音翻译质量
链接:https://arxiv.org/abs/2510.10003

作者:Jianjin Wang, Runsong Zhao, Xiaoqian Liu, Yuan Ge, Ziqiang Xu, Tong Xiao, Shengxiang Gao, Zhengtao Yu, Jingbo Zhu
备注:Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
摘要:当前的直接语音到语音翻译方法主要采用语音标记作为中间表示。但是,单个语音标记在语义上并不密集,因此我们通常需要多个标记来表达一个完整的语义单元。为了解决这一限制,我们将多标记预测(MTP)丢失引入到语音到单元翻译(S2 UT)模型中,使模型能够在每个位置预测多个后续标记,从而捕获更完整的语义并提高每个位置的信息密度。初始MTP实现在最后一层应用损失,这改善了输出表示,但启动信息丰富太晚。我们假设将信息丰富过程推进到中间层可以更早更有效地增强隐藏表示。因此,我们提出了MTP-S2 UT损失,将MTP损失应用于计算CTC损失的隐藏表示。实验表明,所有MTP损失变体一致地提高了S2 UT翻译的质量,其中MTP-S2 UT实现了最佳性能。
摘要:Current direct speech-to-speech translation methods predominantly employ speech tokens as intermediate representations. However, a single speech token is not dense in semantics, so we generally need multiple tokens to express a complete semantic unit. To address this limitation, we introduce multi-token prediction (MTP) loss into speech-to-unit translation (S2UT) models, enabling models to predict multiple subsequent tokens at each position, thereby capturing more complete semantics and enhancing information density per position. Initial MTP implementations apply the loss at the final layer, which improves output representation but initiates information enrichment too late. We hypothesize that advancing the information enrichment process to intermediate layers can achieve earlier and more effective enhancement of hidden representation. Consequently, we propose MTP-S2UT loss, applying MTP loss to hidden representation where CTC loss is computed. Experiments demonstrate that all MTP loss variants consistently improve the quality of S2UT translation, with MTP-S2UT achieving the best performance.


机器翻译由腾讯交互翻译提供,仅供参考