今日论文合集:cs.SD语音13篇,eess.AS音频处理6篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
标题:用于带宽扩展的和声打击解开神经音频编解码器
链接:https://arxiv.org/pdf/2511.21580v1

作者:Benoît Giniès,Xiaoyu Bie,Olivier Fercoq,Gaël Richard
摘要:带宽扩展是音频处理中一个长期存在的问题,其任务是从音频信号的低通对应部分重构音频信号的高频分量。虽然传统方法随着信号处理的更广泛趋势而发展,但神经架构的最新进展显着提高了各种音频任务的性能,在这项工作中,我们通过将帧带宽扩展作为音频令牌预测问题来扩展这些进步。具体来说,我们训练基于变换器的语言模型上产生的离散表示的解纠缠神经音频编解码器,其中解纠缠是由输入信号的谐波打击分解,突出频谱结构特别相关的带宽扩展。我们的方法引入了一种新的编解码器设计,明确占下游令牌预测任务,使编解码器结构和Transformer建模之间的更有效的耦合。这种联合设计产生了高质量的原始信号重建,如客观指标和主观评价所测量的。这些结果突出了将编解码器解纠缠和表示学习与生成建模阶段对齐的重要性,并展示了全局表示感知设计用于推进带宽扩展的潜力。摘要:Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-pass counterpart, is a long-standing problem in audio processing. While traditional approaches have evolved alongside the broader trends in signal processing, recent advances in neural architectures have significantly improved performance across a wide range of audio tasks, In this work, we extend these advances by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.


【2】HarmonicAttack: An Adaptive Cross-Domain Audio Watermark Removal
标题:HarmonicAttack:自适应跨域音频水印删除
链接:https://arxiv.org/pdf/2511.21577v1

作者:Kexin Li,Xiao Hu,Ilya Grishchenko,David Lie
摘要:高质量的人工智能生成音频的可用性带来了安全挑战,例如错误信息活动和语音克隆欺诈。防止滥用人工智能生成的音频的一个关键防御措施是对其进行水印处理,以便很容易将其与真实音频区分开来。由于那些试图滥用AI生成的音频的人可能会因此试图删除音频水印,因此研究有效的水印删除技术对于能够客观地评估音频水印对删除的鲁棒性至关重要。以前的水印去除方案要么假设它们被设计为去除的水印的不切实际的知识,要么计算昂贵,潜在地在当前水印方案中产生虚假的置信度。 我们引入HarmonicAttack,一个有效的音频水印去除方法,只需要基本的能力,从目标方案生成水印,没有别的。这样,我们就能够训练一个通用的水印去除模型,该模型能够从任何带水印的音频样本中去除由目标方案生成的水印。HarmonicAttack采用双路径卷积自动编码器,在时域和频域中操作,以及GAN风格的训练,将水印从原始音频中分离出来。当对最先进的水印方案AudioSeal,WavMark和Silentcipher进行评估时,HarmonicAttack表现出比以前的水印去除方法更强的水印去除能力,具有接近实时的性能。此外,虽然HarmonicAttack需要训练,但我们发现它能够以最小的性能下降转移到分布外的样本。摘要:The availability of high-quality, AI-generated audio raises security challenges such as misinformation campaigns and voice-cloning fraud. A key defense against the misuse of AI-generated audio is by watermarking it, so that it can be easily distinguished from genuine audio. As those seeking to misuse AI-generated audio may thus seek to remove audio watermarks, studying effective watermark removal techniques is critical to being able to objectively evaluate the robustness of audio watermarks against removal. Previous watermark removal schemes either assume impractical knowledge of the watermarks they are designed to remove or are computationally expensive, potentially generating a false sense of confidence in current watermark schemes. We introduce HarmonicAttack, an efficient audio watermark removal method that only requires the basic ability to generate the watermarks from the targeted scheme and nothing else. With this, we are able to train a general watermark removal model that is able to remove the watermarks generated by the targeted scheme from any watermarked audio sample. HarmonicAttack employs a dual-path convolutional autoencoder that operates in both temporal and frequency domains, along with GAN-style training, to separate the watermark from the original audio. When evaluated against state-of-the-art watermark schemes AudioSeal, WavMark, and Silentcipher, HarmonicAttack demonstrates greater watermark removal ability than previous watermark removal methods with near real-time performance. Moreover, while HarmonicAttack requires training, we find that it is able to transfer to out-of-distribution samples with minimal degradation in performance.


【3】Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
标题:使用以音乐混合为条件的扩散模型生成分离的歌唱声音
链接:https://arxiv.org/pdf/2511.21342v1

作者:Genís Plaja-Roglans,Yun-Ning Hung,Xavier Serra,Igor Pereira

备注:Accepted for publication at WASPAA 2025

Journal-ref:2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)

摘要:分离音乐混合物中的各个元素是音乐分析和实践的重要过程。虽然这通常是使用优化的神经网络来解决的,以屏蔽或变换混合物的时频表示来提取目标源,但生成扩散模型的灵活性和泛化能力正在为这一复杂任务带来一类新的解决方案。在这项工作中,我们探索歌声分离从真正的音乐录音使用的扩散模型,该模型被训练生成的独唱声乐条件下相应的混合物。我们的方法改进了以前的生成系统,并在使用补充数据进行训练时,实现了与非生成基线相比具有竞争力的客观分数。扩散采样的迭代性质使用户能够控制质量-效率权衡,并在需要时改进输出。我们提出了一个消融研究的采样算法,突出用户可配置的参数的影响。摘要:Separating the individual elements in a musical mixture is an essential process for music analysis and practice. While this is generally addressed using neural networks optimized to mask or transform the time-frequency representation of a mixture to extract the target sources, the flexibility and generalization capabilities of generative diffusion models are giving rise to a novel class of solutions for this complicated task. In this work, we explore singing voice separation from real music recordings using a diffusion model which is trained to generate the solo vocals conditioned on the corresponding mixture. Our approach improves upon prior generative systems and achieves competitive objective scores against non-generative baselines when trained with supplementary data. The iterative nature of diffusion sampling enables the user to control the quality-efficiency trade-off, and also refine the output when needed. We present an ablation study of the sampling algorithm, highlighting the effects of the user-configurable parameters.


【4】SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
标题:SONAR:用于通用Deepfake检测的频谱对比音频残留
链接:https://arxiv.org/pdf/2511.21325v1

作者:Ido Nitzan HIdekel,Gal lifshitz,Khen Cohen,Dan Raviv
摘要:Deepfake(DF)音频检测器仍然难以推广到分布输入。一个核心原因是频谱偏差,即神经网络在高频(HF)细节之前学习低频结构的趋势,这既导致DF发生器留下HF伪影,又使普通检测器无法利用这些伪影。为了解决这一差距,我们提出了频谱-对比音频电阻(SONAR),一个频率引导的框架,明确地解开一个音频信号到互补的表示。XLSR编码器捕获主要的低频内容,而相同的克隆路径,前面是可学习的SRM,值约束高通滤波器,提取微弱的HF残差。频率交叉关注将长距离和短距离频率依赖性的两个视图重新结合起来,频率感知的Jensen-Shannon对比损失将真实的内容-噪声对拉在一起,同时将假嵌入分开,加速优化并锐化决策边界。在ASVspoof 2021和野外基准测试中进行评估,SONAR获得了最先进的性能,收敛速度比强基线快四倍。通过将微弱的高频残差提升为一流的学习信号,SONAR推出了一个完全数据驱动的频率引导对比框架,将潜在空间分为两个不相交的流形:真实音频的自然HF和合成音频的失真HF,从而锐化决策边界。由于该方案纯粹在表示级别上运行,因此它与架构无关,并且在未来的工作中,可以无缝集成到任何微妙的高频线索起决定性作用的模型或模态中。摘要:Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs. A central reason is spectral bias, the tendency of neural networks to learn low-frequency structure before high-frequency (HF) details, which both causes DF generators to leave HF artifacts and leaves those same artifacts under-exploited by common detectors. To address this gap, we propose Spectral-cONtrastive Audio Residuals (SONAR), a frequency-guided framework that explicitly disentangles an audio signal into complementary representations. An XLSR encoder captures the dominant low-frequency content, while the same cloned path, preceded by learnable SRM, value-constrained high-pass filters, distills faint HF residuals. Frequency cross-attention reunites the two views for long- and short-range frequency dependencies, and a frequency-aware Jensen-Shannon contrastive loss pulls real content-noise pairs together while pushing fake embeddings apart, accelerating optimization and sharpening decision boundaries. Evaluated on the ASVspoof 2021 and in-the-wild benchmarks, SONAR attains state-of-the-art performance and converges four times faster than strong baselines. By elevating faint high-frequency residuals to first-class learning signals, SONAR unveils a fully data-driven, frequency-guided contrastive framework that splits the latent space into two disjoint manifolds: natural-HF for genuine audio and distorted-HF for synthetic audio, thereby sharpening decision boundaries. Because the scheme operates purely at the representation level, it is architecture-agnostic and, in future work, can be seamlessly integrated into any model or modality where subtle high-frequency cues are decisive.


【5】Acoustic neural networks: Identifying design principles and exploring physical feasibility
标题:声学神经网络:识别设计原则并探索物理可行性
链接:https://arxiv.org/pdf/2511.21313v1

作者:Ivan Kalthoff,Marcel Rey,Raphael Wittkowski
备注:13 pages, 4 figures, 8 tables
摘要:基于波导的物理系统为超越传统电子学的节能模拟计算提供了一条有前途的途径。在这种情况下,声学神经网络代表了一种很有前途的方法,可以在电子设备效率低下或有限的环境中实现低功耗计算,但它们的系统设计在很大程度上仍未得到探索。在这里,我们介绍了一个设计和模拟声学神经网络的框架,它通过声波的传播来执行计算。使用数字孪生的方法,我们训练传统的神经网络架构下的物理激励的约束,包括非负信号和权重,偏置项的情况下,和非线性兼容的强度为基础的,非负的声学信号。我们的工作为声学神经网络提供了一个通用框架,该框架将可学习的网络组件直接连接到物理可测量的声学特性,从而能够系统地设计可实现的声学计算系统。我们证明,受约束的经常性和层次结构可以进行准确的语音分类,我们提出了SincHSRNN,一个混合模型,结合了可学习的声学带通滤波器与层次的时间处理。SincHSRNN在AudioMNIST数据集上实现了高达95%的准确度,同时保持与被动声学组件的兼容性。除了计算性能,学习的参数对应于可测量的材料和几何特性,如衰减和传输。我们的研究结果建立了物理上可实现的声学神经网络的一般设计原则,并概述了低功耗,基于波的神经计算的途径。摘要:Wave-guide-based physical systems provide a promising route toward energy-efficient analog computing beyond traditional electronics. Within this landscape, acoustic neural networks represent a promising approach for achieving low-power computation in environments where electronics are inefficient or limited, yet their systematic design has remained largely unexplored. Here we introduce a framework for designing and simulating acoustic neural networks, which perform computation through the propagation of sound waves. Using a digital-twin approach, we train conventional neural network architectures under physically motivated constraints including non-negative signals and weights, the absence of bias terms, and nonlinearities compatible with intensity-based, non-negative acoustic signals. Our work provides a general framework for acoustic neural networks that connects learnable network components directly to physically measurable acoustic properties, enabling the systematic design of realizable acoustic computing systems. We demonstrate that constrained recurrent and hierarchical architectures can perform accurate speech classification, and we propose the SincHSRNN, a hybrid model that combines learnable acoustic bandpass filters with hierarchical temporal processing. The SincHSRNN achieves up to 95% accuracy on the AudioMNIST dataset while remaining compatible with passive acoustic components. Beyond computational performance, the learned parameters correspond to measurable material and geometric properties such as attenuation and transmission. Our results establish general design principles for physically realizable acoustic neural networks and outline a pathway toward low-power, wave-based neural computing.


【6】Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
标题:大规模稳定和韵律单码本TTC LLM的多奖励GRPO
链接:https://arxiv.org/pdf/2511.21270v1

作者:Yicheng Zhong,Peiji Yang,Zhisheng Wang
备注:4 pages, 2 figures
摘要:大型语言模型(LLM)的最新进展已经改变了文本到语音(TTS)合成,激发了将语音表示为离散编解码器令牌序列的自回归框架。其中,单码本TTS LLM已经成为紧凑且可流式传输的架构,其联合建模语义和声学集成。然而,尽管它们的效率,这些模型往往表现出不稳定的韵律,扬声器漂移,自然度下降。为了解决这些问题,我们提出了一个多奖励组相对策略优化(GRPO)框架,直接优化单码本TTS LLM的令牌生成策略。除了标准的可懂度和说话人相似性目标,我们的设计集成了三个基于规则的奖励:持续时间一致性的长度惩罚,解码稳定性的熵正则化奖励,以及明确监督节奏的LLM注释的韵律对齐奖励。在这种韵律奖励中,外部推理LLM通过上下文学习预测多个合理的停顿结构,为GRPO训练提供与人类偏好一致的监督信号。为了评估普遍性,我们进一步在GRPO优化的AR主干上附加了一个流匹配(FM)解码器,并观察到一致的额外增益,这表明我们的强化优化增强了固有的AR策略。我们进一步进行了跨数据大小和模型规模的可扩展性分析,揭示了所提出的方法始终增强了韵律稳定性,说话人相似性和整体语音自然度在单码本TTS LLM。摘要:Recent advances in Large Language Models (LLMs) have transformed text-to-speech (TTS) synthesis, inspiring autoregressive frameworks that represent speech as sequences of discrete codec tokens. Among them, single-codebook TTS LLMs have emerged as compact and streamable architectures that jointly model semantic and acoustic integration. However, despite their efficiency, these models often exhibit unstable prosody, speaker drift, and degraded naturalness. To address these issues, we propose a multi-reward Group Relative Policy Optimization (GRPO) framework that directly optimizes the token generation policy of single-codebook TTS LLMs. Beyond standard intelligibility and speaker similarity objectives, our design integrates three rule-based rewards: a length penalty for duration consistency, an entropy regularization reward for decoding stability, and an LLM-annotated prosody alignment reward that explicitly supervises rhythm. In this prosody reward, an external reasoning LLM predicts multiple plausible pause structures via in-context learning, providing a human-preference-aligned supervisory signal for GRPO training. To assess universality, we further attach a flow-matching (FM) decoder on top of the GRPO-optimized AR backbone and observe consistent additional gains, indicating that our reinforcement optimization enhances the intrinsic AR policy. We further conduct a scalability analysis across data sizes and model scales, revealing that the proposed method consistently enhances prosodic stability, speaker similarity, and overall speech naturalness in single-codebook TTS LLMs.


【7】AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
标题:AV-Edit:通过视听语义关节控制进行多模式生成音效编辑
链接:https://arxiv.org/pdf/2511.21146v1

作者:Xinyue Guo,Xiaoran Yang,Lipan Zhang,Jianxuan Yang,Zhao Wang,Jian Luan
摘要:音效编辑-通过添加、删除或替换元素来修改音频-仍然受到现有方法的限制,这些方法仅依赖于低级信号处理或粗糙的文本提示,通常导致灵活性有限和次优音频质量。为了解决这个问题,我们提出了AV-Edit,一个生成的音效编辑框架,通过联合利用视觉,音频和文本语义,可以对视频中的现有音轨进行细粒度编辑。具体来说,所提出的方法采用了一个专门设计的对比视听掩蔽自动编码器(CAV-MAE-Edit)进行多模态预训练,学习对齐的跨模态表示。然后,这些表示被用于训练编辑的多模态扩散Transformer(MM-DiT),其能够通过基于相关性的特征选通训练策略去除视觉上不相关的声音并生成与视频内容一致的缺失的音频元素。此外,我们构建了一个专用的基于视频的声音编辑数据集作为评估基准。实验结果表明,AV-Edit能够根据视觉内容进行精确的修改,生成高质量的音频,在音效编辑领域具有最先进的性能,在音频生成领域具有较强的竞争力。摘要:Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit, a generative sound effect editing framework that enables fine-grained editing of existing audio tracks in videos by jointly leveraging visual, audio, and text semantics. Specifically, the proposed method employs a specially designed contrastive audio-visual masking autoencoder (CAV-MAE-Edit) for multimodal pre-training, learning aligned cross-modal representations. These representations are then used to train an editorial Multimodal Diffusion Transformer (MM-DiT) capable of removing visually irrelevant sounds and generating missing audio elements consistent with video content through a correlation-based feature gating training strategy. Furthermore, we construct a dedicated video-based sound editing dataset as an evaluation benchmark. Experiments demonstrate that the proposed AV-Edit generates high-quality audio with precise modifications based on visual content, achieving state-of-the-art performance in the field of sound effect editing and exhibiting strong competitiveness in the domain of audio generation.


【8】ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features
标题:使用语音特征的对齐增强转换器在低资源缅甸语中进行ASR纠错
链接:https://arxiv.org/pdf/2511.21088v1

作者:Ye Bhone Lin,Thura Aung,Ye Kyaw Thu,Thazin Myint Oo
备注:7 pages, 2 figures, 7 tables, Accepted to iSAI-NLP 2025
摘要:本文研究了用于低资源缅甸语自动语音识别(ASR)纠错的序列到序列Transformer模型,重点关注不同的特征集成策略,包括IPA和对齐信息。据我们所知,这是第一个专门针对缅甸语的ASR纠错研究。我们评估了五个ASR主干,并表明我们的ASR纠错(AEC)方法始终提高了基线输出的单词和字符级别的准确性。所提出的AEC模型,结合IPA和对齐功能,将ASR模型的平均WER从增强前的51.56降低到39.82(增强后从51.56降低到43.59),并将chrF++评分从0.5864提高到0.627,表明在没有AEC的情况下,与基线ASR输出相比,获得了一致的收益。我们的研究结果突出了AEC的鲁棒性和功能设计的重要性,以提高ASR输出在低资源设置。摘要:This paper investigates sequence-to-sequence Transformer models for automatic speech recognition (ASR) error correction in low-resource Burmese, focusing on different feature integration strategies including IPA and alignment information. To our knowledge, this is the first study addressing ASR error correction specifically for Burmese. We evaluate five ASR backbones and show that our ASR Error Correction (AEC) approaches consistently improve word- and character-level accuracy over baseline outputs. The proposed AEC model, combining IPA and alignment features, reduced the average WER of ASR models from 51.56 to 39.82 before augmentation (and 51.56 to 43.59 after augmentation) and improving chrF++ scores from 0.5864 to 0.627, demonstrating consistent gains over the baseline ASR outputs without AEC. Our results highlight the robustness of AEC and the importance of feature design for improving ASR outputs in low-resource settings.


【9】CartoonSing: Unifying Human and Nonhuman Timbres in Singing Generation
标题:卡通歌唱:歌唱一代中人类和非人类音色的统一
链接:https://arxiv.org/pdf/2511.21045v1

作者:Jionghao Han,Jiatong Shi,Zhuoyan Tao,Yuxun Tang,Yiwen Zhao,Gus Xia,Shinji Watanabe
摘要:歌唱声合成(SVS)和歌唱声转换(SVC)在生成自然发声的人类歌唱方面取得了显著的进展。然而,现有的系统仅限于人类音色,并且在人类范围之外合成语音的能力有限,这在诸如视频游戏、电影和虚拟角色的创造性应用中越来越多地被要求。我们介绍了非人类歌唱生成(NHSG),包括非人类歌唱声音合成(NHSVS)和非人类歌唱声音转换(NHSVC),作为一种新的机器学习任务,用于生成具有非人类音色特征的音乐连贯的歌唱。NHSG是特别具有挑战性的,由于非人类歌唱数据的稀缺性,缺乏符号对齐,以及人类和非人类声音之间的宽音色差距。为了解决这些挑战,我们提出了CartoonSing,一个统一的框架,集成了歌唱声音合成和转换,同时桥接人类和非人类歌唱一代。CartoonSing采用了两个阶段的管道:一个是用注释的人类歌唱训练的分数表示编码器,另一个是为人类和非人类音频重建波形的音色感知声码器。实验表明,CartoonSing成功地产生非人类的歌声,概括到新颖的音色,并扩展传统的SVS和SVC创造性的,非人类的歌唱生成。摘要:Singing voice synthesis (SVS) and singing voice conversion (SVC) have achieved remarkable progress in generating natural-sounding human singing. However, existing systems are restricted to human timbres and have limited ability to synthesize voices outside the human range, which are increasingly demanded in creative applications such as video games, movies, and virtual characters. We introduce Non-Human Singing Generation (NHSG), covering non-human singing voice synthesis (NHSVS) and non-human singing voice conversion (NHSVC), as a novel machine learning task for generating musically coherent singing with non-human timbral characteristics. NHSG is particularly challenging due to the scarcity of non-human singing data, the lack of symbolic alignment, and the wide timbral gap between human and non-human voices. To address these challenges, we propose CartoonSing, a unified framework that integrates singing voice synthesis and conversion while bridging human and non-human singing generation. CartoonSing employs a two-stage pipeline: a score representation encoder trained with annotated human singing and a timbre-aware vocoder that reconstructs waveforms for both human and non-human audio. Experiments demonstrate that CartoonSing successfully generates non-human singing voices, generalizes to novel timbres, and extends conventional SVS and SVC toward creative, non-human singing generation.


【10】SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications
标题:SingingDS:一个支持唱歌的口语对话系统,用于对话角色扮演应用
链接:https://arxiv.org/pdf/2511.20972v1

作者:Jionghao Han,Jiatong Shi,Masao Someki,Yuxun Tang,Lan Liu,Yiwen Zhao,Wenhao Feng,Shinji Watanabe
摘要:随着自动语音识别(ASR)、大语言模型(LLM)和文本到语音(TTS)技术的最新进展,口语对话系统(SDS)已经变得广泛可用。然而,大多数现有的SDS仅限于传统的口头响应。我们提出SingingSDS,级联SDS,通过唱歌而不是说话来响应,在基于角色的角色扮演和互动娱乐场景中培养更多情感,难忘和愉快的互动。SingingSDS采用模块化的ASR-LLM-SVS管道,并支持跨角色角色、ASR和LLM后端、SVS模型、旋律源和语音配置文件的各种配置,以满足延迟、质量和音乐风格方面的不同需求。SingingSDS是一个即插即用的Web演示,具有模块化的开源代码,支持自定义和扩展。演示:https: huggingface.co spaces espnet SingingSDS.代码:https: github.com SingingSDS SingingSDS.摘要:With recent advances in automatic speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS) technologies, spoken dialogue systems (SDS) have become widely accessible. However, most existing SDS are limited to conventional spoken responses. We present SingingSDS, a cascaded SDS that responds through singing rather than speaking, fostering more affective, memorable, and pleasurable interactions in character-based roleplay and interactive entertainment scenarios. SingingSDS employs a modular ASR-LLM-SVS pipeline and supports a wide range of configurations across character personas, ASR and LLM backends, SVS models, melody sources, and voice profiles, tailored to different needs in terms of latency, quality, and musical style. SingingSDS is available as a plug-and-play web demo, featuring modular, open-source code that supports customization and extension. Demo: https: huggingface.co spaces espnet SingingSDS. Code: https: github.com SingingSDS SingingSDS.


【11】Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
标题:乐谱理解基准:评估大型语言模型对完整乐谱的理解
链接:https://arxiv.org/pdf/2511.20697v1

作者:Congren Dai,Yue Yang,Krinos Li,Huichi Zhou,Shijie Liang,Zhang Bo,Enyang Liu,Ge Jin,Hongran An,Haosen Zhang,Peiyuan Jing,KinHei Lee,Zhenxuan Zhang,Xiaobing Li,Maosong Sun
摘要:理解完整的乐谱需要对符号结构进行推理,如音高,节奏,和声和形式。尽管大语言模型(LLM)和视觉语言模型(VLM)在自然语言和多模态任务中取得了快速进展,但它们理解乐谱的能力仍然没有得到充分的探索。我们介绍了乐谱理解基准(MSU-Bench),这是第一个大规模的,人工策划的基准,用于评估文本(ABC符号)和视觉(PDF)模式的乐谱水平音乐理解。MSU-Bench包含1,800个生成式问答(QA)对,来自巴赫,贝多芬,肖邦,德彪西等作品,分为四个渐进的理解水平:起始信息,记谱法和音符,和弦与和声,以及纹理与形式。通过对超过15个最先进(SOTA)模型的广泛zero-shot和微调评估,我们揭示了尖锐的模态差距,脆弱的水平成功率,以及维持多级正确性的困难。微调显著提高了两种模式的性能,同时保留了一般知识,将MSU-Bench建立为人工智能(AI),音乐学和多模态推理交叉领域未来研究的严格基础。摘要:Understanding complete musical scores requires reasoning over symbolic structures such as pitch, rhythm, harmony, and form. Despite the rapid progress of Large Language Models (LLMs) and Vision-Language Models (VLMs) in natural language and multimodal tasks, their ability to comprehend musical notation remains underexplored. We introduce Musical Score Understanding Benchmark (MSU-Bench), the first large-scale, human-curated benchmark for evaluating score-level musical understanding across both textual (ABC notation) and visual (PDF) modalities. MSU-Bench comprises 1,800 generative question-answer (QA) pairs drawn from works spanning Bach, Beethoven, Chopin, Debussy, and others, organised into four progressive levels of comprehension: Onset Information, Notation & Note, Chord & Harmony, and Texture & Form. Through extensive zero-shot and fine-tuned evaluations of over 15+ state-of-the-art (SOTA) models, we reveal sharp modality gaps, fragile level-wise success rates, and the difficulty of sustaining multilevel correctness. Fine-tuning markedly improves performance in both modalities while preserving general knowledge, establishing MSU-Bench as a rigorous foundation for future research at the intersection of Artificial Intelligence (AI), musicological, and multimodal reasoning.


【12】Seeing Beyond Sound: Visualization and Abstraction in Audio Data Representation
标题:超越声音:音频数据表示中的可视化和抽象
链接:https://arxiv.org/pdf/2511.20658v1

作者:Ashlae Blum'e
备注:23 pages, 3 figures
摘要:在音频信号处理中,使用视觉表示的复杂信息的解释通过其与人类感知系统的对准来增强模式识别。带有从其历史背景中继承的隐藏假设的软件工具有可能与现代工作流程不一致,因为设计起源变得模糊。我们认为,创建符合紧急需求的工具,提高了分析和创造性的产出,由于使用它们的亲和力增加。本文探讨了在可视化工具中添加维度和交互性以促进使用水母炸药软件进行音频信息研究的复杂工作流程的潜力。摘要:In audio signal processing, the interpretation of complex information using visual representation enhances pattern recognition through its alignment with human perceptual systems. Software tools that carry hidden assumptions inherited from their historical contexts risk misalignment with modern workflows as design origins become obscured. We argue that creating tools that align with emergent needs improves analytical and creative outputs due to an increased affinity for using them. This paper explores the potentials associated with adding dimensionality and interactivity into visualization tools to facilitate complex workflows in audio information research using the Jellyfish Dynamite software.


【13】The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval
标题:球体数据集:用于音乐源分离和信息检索的多轨弦录音
链接:https://arxiv.org/pdf/2511.21247v1

作者:Jaime Garcia-Martinez,David Diaz-Guerra,John Anderson,Ricardo Falcon-Perez,Pablo Cabañas-Molero,Tuomas Virtanen,Julio J. Carabias-Orti,Pedro Vera-Candeas
摘要:本文介绍了Spheres数据集,多轨管弦乐录音,旨在推进古典音乐领域音乐源分离和相关MIR任务的机器学习研究。该数据集由Colibrant Ensemble在The Spheres录音室录制的一个多小时的音乐作品组成,包括两部经典作品-柴可夫斯基的罗密欧与朱丽叶和莫扎特的第40号交响曲-以及每种乐器的半音阶和独奏节选。录音设置使用了23个麦克风,包括近点麦克风、主麦克风和环境麦克风,从而能够创建具有受控出血的逼真立体声混音,并为源分离模型的监督训练提供孤立的主干。此外,估计每个仪器位置的房间脉冲响应,提供有价值的声学表征的记录空间。我们提出的数据集结构,声学分析,和基线评估使用X-UMX为基础的模型,管弦乐家庭分离和麦克风debleeding。结果突出了复杂管弦乐场景中源分离的潜力和挑战,强调了数据集的基准测试和探索分离,本地化,去混响和沉浸式渲染古典音乐的新方法的价值。摘要:This paper introduces The Spheres dataset, multitrack orchestral recordings designed to advance machine learning research in music source separation and related MIR tasks within the classical music domain. The dataset is composed of over one hour recordings of musical pieces performed by the Colibrì Ensemble at The Spheres recording studio, capturing two canonical works - Tchaikovsky's Romeo and Juliet and Mozart's Symphony No. 40 - along with chromatic scales and solo excerpts for each instrument. The recording setup employed 23 microphones, including close spot, main, and ambient microphones, enabling the creation of realistic stereo mixes with controlled bleeding and providing isolated stems for supervised training of source separation models. In addition, room impulse responses were estimated for each instrument position, offering valuable acoustic characterization of the recording space. We present the dataset structure, acoustic analysis, and baseline evaluations using X-UMX based models for orchestral family separation and microphone debleeding. Results highlight both the potential and the challenges of source separation in complex orchestral scenarios, underscoring the dataset's value for benchmarking and for exploring new approaches to separation, localization, dereverberation, and immersive rendering of classical music.


eess.AS音频处理


【1】The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval
标题:球体数据集:用于音乐源分离和信息检索的多轨弦录音
链接:https://arxiv.org/pdf/2511.21247v1

作者:Jaime Garcia-Martinez,David Diaz-Guerra,John Anderson,Ricardo Falcon-Perez,Pablo Cabañas-Molero,Tuomas Virtanen,Julio J. Carabias-Orti,Pedro Vera-Candeas
摘要:本文介绍了Spheres数据集,多轨管弦乐录音,旨在推进古典音乐领域音乐源分离和相关MIR任务的机器学习研究。该数据集由Colibrant Ensemble在The Spheres录音室录制的一个多小时的音乐作品组成,包括两部经典作品-柴可夫斯基的罗密欧与朱丽叶和莫扎特的第40号交响曲-以及每种乐器的半音阶和独奏节选。录音设置使用了23个麦克风,包括近点麦克风、主麦克风和环境麦克风,从而能够创建具有受控出血的逼真立体声混音,并为源分离模型的监督训练提供孤立的主干。此外,估计每个仪器位置的房间脉冲响应,提供有价值的声学表征的记录空间。我们提出的数据集结构,声学分析,和基线评估使用X-UMX为基础的模型,管弦乐家庭分离和麦克风debleeding。结果突出了复杂管弦乐场景中源分离的潜力和挑战,强调了数据集的基准测试和探索分离,本地化,去混响和沉浸式渲染古典音乐的新方法的价值。摘要:This paper introduces The Spheres dataset, multitrack orchestral recordings designed to advance machine learning research in music source separation and related MIR tasks within the classical music domain. The dataset is composed of over one hour recordings of musical pieces performed by the Colibrì Ensemble at The Spheres recording studio, capturing two canonical works - Tchaikovsky's Romeo and Juliet and Mozart's Symphony No. 40 - along with chromatic scales and solo excerpts for each instrument. The recording setup employed 23 microphones, including close spot, main, and ambient microphones, enabling the creation of realistic stereo mixes with controlled bleeding and providing isolated stems for supervised training of source separation models. In addition, room impulse responses were estimated for each instrument position, offering valuable acoustic characterization of the recording space. We present the dataset structure, acoustic analysis, and baseline evaluations using X-UMX based models for orchestral family separation and microphone debleeding. Results highlight both the potential and the challenges of source separation in complex orchestral scenarios, underscoring the dataset's value for benchmarking and for exploring new approaches to separation, localization, dereverberation, and immersive rendering of classical music.


【2】Evaluation of an ITD-to-ILD Transformation as a Method to Restore the Spatial Benefit in Speech Intelligibility in Hearing Impaired Listeners
标题:评估ITD到ILD转换作为恢复听力障碍收听者言语可理解性空间效益的方法
链接:https://arxiv.org/pdf/2511.21222v1

作者:Timm-Jonas Bäumer,Johannes W. de Vries,Stephan Töpken,Richard C. Hendriks,Peyman Goli,Steven van de Par
备注:12 pages, 11 figues. Submitted to the special issue for the International Symposium on Hearing 2025 in Trends in Hearing
摘要:为了提高复杂日常情况下的语音清晰度,人类听觉系统部分依赖于耳间时间差(ITD)和耳间电平差(ILD)。然而,听力受损(HI)的听众往往表现出有限的敏感性ITD,导致语音清晰度性能下降。本研究旨在调查是否转换为ILD低频ITDs可以重新引入双耳的好处HI听众。我们进行了两个实验与HI听众。第一个实验使用不同频率的双耳相移正弦波来评估HI听者ITD敏感阈值。所有受试者在较高频率下ITD阈值增加,在较低频率下受试者之间ITD敏感性不同。在第二个实验中,通过操纵头相关传递函数(HRTF),测量了不同双耳配置下的语音接收延迟(SRT)。结果表明,尽管ITD灵敏度降低,但与ITD和ILD可用的未处理基线相比,去除ITD使SRT降低约1 dB。此外,用ILD代替低频ITD产生了横向目标扬声器的改进。在保留ITD的同时添加低频ILD,使得扬声器在各个方向上都有了显著的改善。这些发现表明,所提出的转换方法可以有效地恢复HI听众的双耳益处。这项研究的结果表明,使用这种转换技术来实现助听器和人工耳蜗植入,直接受益于HI听众。摘要:To improve speech intelligibility in complex everyday situations, the human auditory system partially relies on Interaural Time Differences (ITDs) and Interaural Level Differences (ILDs). However, hearing impaired (HI) listeners often exhibit limited sensitivity to ITDs, resulting in decreased speech intelligibility performance. This study aimed to investigate whether transforming low-frequency ITDs into ILDs could reintroduce a binaural benefit for HI listeners. We conducted two experiments with HI listeners. The first experiment used binaurally phase-shifted sinusoids at different frequencies to evaluate the HI listeners ITD sensitivity threshold. All subjects had an increased ITD threshold at higher frequencies, with different ITD sensitivities between the subjects in the lower frequencies. In the second experiment, Speech Reception Thresholds (SRTs) were measured in different binaural configurations by manipulating Head-Related Transfer Functions (HRTFs). The results showed that, despite the decreased ITD sensitivity, removing ITDs decreased SRTs by approximately 1 dB compared to the unprocessed baseline, where ITDs and ILDs are available. Furthermore, substituting low-frequency ITDs with ILDs yielded an improvement for a lateral target speaker. Adding the low-frequency ILDs while preserving the ITDs caused a significant improvement for speakers in all directions. These findings suggest that the proposed transformation method could be effective in restoring binaural benefits in HI listeners. The results of this study suggest the use of such transformation techniques to be implemented in hearing aids and cochlear implants, directly benefiting HI listeners.


【3】RosettaSpeech: Zero-Shot Speech-to-Speech Translation from Monolingual Data
标题:RosettaSpeech:从单语数据进行Zero-Shot语音翻译
链接:https://arxiv.org/pdf/2511.20974v1

作者:Zhisheng Zheng,Xiaohang Sun,Tuan Dinh,Abhishek Yanamandra,Abhinav Jain,Zhu Liu,Sunil Hadap,Vimal Bhat,Manoj Aggarwal,Gerard Medioni,David Harwath
备注:Work in progress
摘要:并行语音语料库的稀缺严重阻碍了语音到语音翻译(S2 ST),通常迫使依赖复杂的多阶段管道。本文介绍了RosettaSpeech,这是一种新颖且简化的zero-shot S2 ST框架,该框架在通过机器翻译监督增强的单语语音文本数据上进行训练。虽然我们的方法利用了基于文本的NMT模型中固有的语言知识,但它严格消除了对并行语音对的需求。我们的模型在训练过程中使用文本作为中间桥梁,但在推理时用作直接的端到端语音到语音模型。这种简化的方法在标准基准上实现了最先进的结果。例如,在CVSS-C测试集上,RosettaSpeech的表现优于领先的系统,德语到英语的ASR-BLEU得分为25.17,西班牙语到英语的ASR-BLEU得分为29.86,相对增益分别超过27%和14%。此外,我们证明了一个单一的模型可以提供强大的多对一的翻译性能(FR ES DE - EN)。我们还提供了训练数据缩放如何影响模型性能的基础分析。通过优先依赖大量的并行文本,而不是难以获取的并行语音,RosettaSpeech提供了一条可扩展的路径,为更广泛的语言创建高质量的、保留说话者的S2 ST。摘要:The scarcity of parallel speech corpora critically hampers speech-to-speech translation (S2ST), often forcing reliance on complex, multi-stage pipelines. This paper introduces RosettaSpeech, a novel and simplified framework for zero-shot S2ST that is trained on monolingual speech-text data augmented by machine translation supervision. While our method leverages the linguistic knowledge inherent in text-based NMT models, it strictly eliminates the need for parallel speech-to-speech pairs. Our model uniquely uses text as an intermediate bridge during training but functions as a direct, end-to-end speech-to-speech model at inference. This streamlined approach achieves state-of-the-art results on standard benchmarks. For instance, on the CVSS-C test set, RosettaSpeech outperforms leading systems, achieving an ASR-BLEU score of 25.17 for German-to-English and 29.86 for Spanish-to-English-relative gains of over 27% and 14%, respectively. Furthermore, we demonstrate that a single model can deliver strong many-to-one translation performance (FR ES DE - EN). We also provide a foundational analysis of how training data scaling impacts model performance. By prioritizing reliance on abundant parallel text rather than difficult-to-acquire parallel speech, RosettaSpeech offers a scalable path to creating high-quality, speaker-preserving S2ST for a much broader array of languages.


【4】Towards Audio Token Compression in Large Audio Language Models
标题:大型音频语言模型中的音频令牌压缩
链接:https://arxiv.org/pdf/2511.20973v1

作者:Saurabhchand Bhati,Samuel Thomas,Hilde Kuehne,Rogerio Feris,James Glass
摘要:大型音频语言模型(LALM)在从语音识别到一般音频理解的各种任务中表现出令人印象深刻的性能。然而,它们的可扩展性受到注意力的二次复杂度和音频信号的高令牌速率的限制。这些挑战使得难以将LALM扩展到长格式音频,并将其部署在资源受限的平台上,例如边缘设备。 在本文中,我们探讨了无监督分割,均匀平均池等技术,以减少由LALM的音频编码器生成但在它们被LLM解码器消耗之前的音频令牌的数量。为了减轻压缩表示引入的潜在性能下降,我们采用低秩适配器来微调模型。我们评估我们提出的模型上的两个任务,自动语音识别和语音到语音翻译任务,这是依赖于有效地揭示输入信号的基本词汇内容,并研究这些任务的影响下采样。实验结果表明,压缩的LALM可以实现更接近帧级LALM的性能,同时减少输入的音频令牌计数到LLM骨干之前的三倍。摘要:Large Audio Language Models (LALMs) demonstrate impressive performance across diverse tasks, ranging from speech recognition to general audio understanding. However, their scalability is limited by the quadratic complexity of attention and the high token rates of audio signals. These challenges make it difficult to extend LALMs to long-form audio and to deploy them on resource-constrained platforms such as edge devices. In this paper, we explore techniques such as unsupervised segmentation, uniform average pooling, etc., to reduce the number of audio tokens generated by the LALM's audio encoder but before they are consumed by the LLM decoder. To mitigate potential performance degradation introduced by the compressed representations, we employ low-rank adapters to finetune the model. We evaluate our proposed models on two tasks, automatic speech recognition and speech-to-speech translation tasks, that are dependent on effectively uncovering the underlying lexical content of the input signal and study the effect of downsampling on these tasks. Experimental results show that compressed LALMs can achieve performance closer to frame-level LALMs while reducing the input audio token count upto three times before the LLM backbone.


【5】Acoustic neural networks: Identifying design principles and exploring physical feasibility
标题:声学神经网络:识别设计原则并探索物理可行性
链接:https://arxiv.org/pdf/2511.21313v1

作者:Ivan Kalthoff,Marcel Rey,Raphael Wittkowski
备注:13 pages, 4 figures, 8 tables
摘要:基于波导的物理系统为超越传统电子学的节能模拟计算提供了一条有前途的途径。在这种情况下,声学神经网络代表了一种很有前途的方法,可以在电子设备效率低下或有限的环境中实现低功耗计算,但它们的系统设计在很大程度上仍未得到探索。在这里,我们介绍了一个设计和模拟声学神经网络的框架,它通过声波的传播来执行计算。使用数字孪生的方法,我们训练传统的神经网络架构下的物理激励的约束,包括非负信号和权重,偏置项的情况下,和非线性兼容的强度为基础的,非负的声学信号。我们的工作为声学神经网络提供了一个通用框架,该框架将可学习的网络组件直接连接到物理可测量的声学特性,从而能够系统地设计可实现的声学计算系统。我们证明,受约束的经常性和层次结构可以进行准确的语音分类,我们提出了SincHSRNN,一个混合模型,结合了可学习的声学带通滤波器与层次的时间处理。SincHSRNN在AudioMNIST数据集上实现了高达95%的准确度,同时保持与被动声学组件的兼容性。除了计算性能,学习的参数对应于可测量的材料和几何特性,如衰减和传输。我们的研究结果建立了物理上可实现的声学神经网络的一般设计原则,并概述了低功耗,基于波的神经计算的途径。摘要:Wave-guide-based physical systems provide a promising route toward energy-efficient analog computing beyond traditional electronics. Within this landscape, acoustic neural networks represent a promising approach for achieving low-power computation in environments where electronics are inefficient or limited, yet their systematic design has remained largely unexplored. Here we introduce a framework for designing and simulating acoustic neural networks, which perform computation through the propagation of sound waves. Using a digital-twin approach, we train conventional neural network architectures under physically motivated constraints including non-negative signals and weights, the absence of bias terms, and nonlinearities compatible with intensity-based, non-negative acoustic signals. Our work provides a general framework for acoustic neural networks that connects learnable network components directly to physically measurable acoustic properties, enabling the systematic design of realizable acoustic computing systems. We demonstrate that constrained recurrent and hierarchical architectures can perform accurate speech classification, and we propose the SincHSRNN, a hybrid model that combines learnable acoustic bandpass filters with hierarchical temporal processing. The SincHSRNN achieves up to 95% accuracy on the AudioMNIST dataset while remaining compatible with passive acoustic components. Beyond computational performance, the learned parameters correspond to measurable material and geometric properties such as attenuation and transmission. Our results establish general design principles for physically realizable acoustic neural networks and outline a pathway toward low-power, wave-based neural computing.


【6】Seeing Beyond Sound: Visualization and Abstraction in Audio Data Representation
标题:超越声音:音频数据表示中的可视化和抽象
链接:https://arxiv.org/pdf/2511.20658v1

作者:Ashlae Blum'e
备注:23 pages, 3 figures
摘要:在音频信号处理中,使用视觉表示的复杂信息的解释通过其与人类感知系统的对准来增强模式识别。带有从其历史背景中继承的隐藏假设的软件工具有可能与现代工作流程不一致,因为设计起源变得模糊。我们认为,创建符合紧急需求的工具,提高了分析和创造性的产出,由于使用它们的亲和力增加。本文探讨了在可视化工具中添加维度和交互性以促进使用水母炸药软件进行音频信息研究的复杂工作流程的潜力。摘要:In audio signal processing, the interpretation of complex information using visual representation enhances pattern recognition through its alignment with human perceptual systems. Software tools that carry hidden assumptions inherited from their historical contexts risk misalignment with modern workflows as design origins become obscured. We argue that creating tools that align with emergent needs improves analytical and creative outputs due to an increased affinity for using them. This paper explores the potentials associated with adding dimensionality and interactivity into visualization tools to facilitate complex workflows in audio information research using the Jellyfish Dynamite software.


机器翻译由腾讯交互翻译提供,仅供参考