今日论文合集:cs.SD语音8篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音



【1】 EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via  Reinforcement Learning
标题: EchoInk-R1:通过强化学习探索多模式LLM中的视听推理
链接:https://arxiv.org/abs/2505.04623
作者: Zhenghao Xing,  Xiaowei Hu,  Chi-Wing Fu,  Wenhai Wang,  Jifeng Dai,  Pheng-Ann Heng 
摘要:多模态大型语言模型(MLLM)具有跨文本、视觉和音频的高级感知能力,但它们通常难以进行结构化的跨模态推理,特别是在集成音频和视觉信号时。我们介绍EchoInk-R1,这是一个强化学习框架,可以增强MLLM中的这种推理。EchoInk-R1基于Qwen2.5-Omni-7 B基础,并使用组相对策略优化(GRPO)进行了优化,可以在同步的音频图像对上处理多项选择题。为了实现这一点,我们策划了AVQA-R1- 6 K,这是一个将此类音频图像输入与OmniInstruct-v1中的多项选择题配对的数据集。EchoInk-R1- 7 B在验证集上实现了85.77%的准确率,优于基础模型,后者的得分为80.53%,仅使用了562个强化学习步骤。除了准确性之外,EchoInk-R1还通过重新审视初始解释并在面对模糊的多模态输入时改进响应来展示反思性推理。这些结果表明,轻量级强化学习微调增强了MLLM中的跨模态推理。EchoInk-R1是第一个通过强化学习统一音频,视觉和文本形式的通用开放世界推理框架。代码和数据公开发布,以方便进一步的研究。
摘要:Multimodal large language models (MLLMs) have advanced perception across text, vision, and audio, yet they often struggle with structured cross-modal reasoning, particularly when integrating audio and visual signals. We introduce EchoInk-R1, a reinforcement learning framework that enhances such reasoning in MLLMs. Built upon the Qwen2.5-Omni-7B foundation and optimized with Group Relative Policy Optimization (GRPO), EchoInk-R1 tackles multiple-choice question answering over synchronized audio-image pairs. To enable this, we curate AVQA-R1-6K, a dataset pairing such audio-image inputs with multiple-choice questions derived from OmniInstruct-v1. EchoInk-R1-7B achieves 85.77% accuracy on the validation set, outperforming the base model, which scores 80.53%, using only 562 reinforcement learning steps. Beyond accuracy, EchoInk-R1 demonstrates reflective reasoning by revisiting initial interpretations and refining responses when facing ambiguous multimodal inputs. These results suggest that lightweight reinforcement learning fine-tuning enhances cross-modal reasoning in MLLMs. EchoInk-R1 is the first framework to unify audio, visual, and textual modalities for general open-world reasoning via reinforcement learning. Code and data are publicly released to facilitate further research.


【2】 Accelerating Audio Research with Robotic Dummy Heads

标题: 利用机器人假人头加速音频研究
链接:https://arxiv.org/abs/2505.04548
作者: Austin Lu,  Kanad Sarkar,  Yongjie Zhuang,  Leo Lin,  Ryan M Corey,  Andrew C Singer 
备注:WASPAA 2025
摘要:这项工作介绍了一个机器人假人头,融合了传统的听觉人体模型与机器人的流动性的声学现实主义。该设备能够像人一样移动,说话和倾听,并可用于自动化空间静止音频实验,从而加快音频研究的步伐。重要的是,由于其安静的电机,该设备还可以在动态实验中用作移动声源。这一功能使我们的工作与之前的机器人声学研究平台有所区别。通过各种实验和声学测量来验证机器人能够进行高质量的音频数据收集。这些实验还演示了机器人如何用于研究自适应双耳波束形成。设计文件作为开源提供,以刺激新的音频研究。
摘要:This work introduces a robotic dummy head that fuses the acoustic realism of conventional audiological mannequins with the mobility of robots. The proposed device is capable of moving, talking, and listening as people do, and can be used to automate spatially-stationary audio experiments, thus accelerating the pace of audio research. Critically, the device may also be used as a moving sound source in dynamic experiments, due to its quiet motor. This feature differentiates our work from previous robotic acoustic research platforms. Validation that the robot enables high quality audio data collection is provided through various experiments and acoustic measurements. These experiments also demonstrate how the robot might be used to study adaptive binaural beamforming. Design files are provided as open-source to stimulate novel audio research.


【3】 Aliasing Reduction in Neural Amp Modeling by Smoothing Activations

标题: 通过平滑激活减少神经元建模中的混叠
链接:https://arxiv.org/abs/2505.04082
作者: Ryota Sato,  Julius O. Smith III 
备注:Accepted to DAFx 2025
摘要:对模拟音频硬件(如老式吉他放大器)的高质量数字仿真的需求日益增长,导致了基于神经网络的黑盒建模的大量工作,WaveNet等深度学习架构显示出了良好的效果。然而,所有这些模型中的一个关键限制是在神经网络中使用非线性激活函数所产生的混叠伪影。在本文中,我们研究了新的和修改的激活函数,旨在减轻神经放大器模型内的混叠。为了支持这一点,我们引入了一个新的度量,混叠信号比(ASR),它定量评估混叠的水平,具有高精度。测量传统的错误信号比(ESR),我们进行了一系列的预先存在的和现代的激活功能与不同的拉伸因子的研究。我们的研究结果证实,具有更平滑曲线的激活函数倾向于实现更低的ASR值,这表明混叠明显减少。值得注意的是,这种混叠减少的改进是在没有显著增加ESR的情况下实现的,这表明在神经放大器模型中具有减少混叠的高建模精度的潜力。
摘要:The increasing demand for high-quality digital emulations of analog audio hardware such as vintage guitar amplifiers has led to numerous works in neural-network-based black-box modeling, with deep learning architectures like WaveNet showing promising results. However, a key limitation in all of these models is the aliasing artifacts that arise from the use of nonlinear activation functions in neural networks. In this paper, we investigate novel and modified activation functions aimed at mitigating aliasing within neural amplifier models. Supporting this, we introduce a novel metric, the Aliasing-to-Signal Ratio (ASR), which quantitatively assesses the level of aliasing with high accuracy. Measuring also the conventional Error-to-Signal Ratio (ESR), we conducted studies on a range of preexisting and modern activation functions with varying stretch factors. Our findings confirmed that activation functions with smoother curves tend to achieve lower ASR values, indicating a noticeable reduction in aliasing. Notably, this improvement in aliasing reduction was achievable without a substantial increase in ESR, demonstrating the potential for high modeling accuracy with reduced aliasing in neural amp models.


【4】 Score Distillation Sampling for Audio: Source Separation, Synthesis, and  Beyond

标题: 音频的分数蒸馏采样:源分离、合成等
链接:https://arxiv.org/abs/2505.04621
作者: Jessie Richter-Powell,  Antonio Torralba,  Jonathan Lorraine 
备注:See the project website at this https URL
摘要:我们介绍Audio-SDS,分数蒸馏采样(SDS)的推广文本条件音频扩散模型。虽然SDS最初是为使用图像扩散的文本到3D生成而设计的,但其将强大的生成先验提取为单独的参数表示的核心思想扩展到了音频领域。利用单个预训练模型,Audio-SDS可以实现广泛的任务,而无需专门的数据集。特别是,我们演示了如何Audio-SDS可以指导物理知情的影响声音模拟,校准FM合成参数,并执行指定的源分离。我们的研究结果说明了基于蒸馏的方法在不同模态中的多功能性,并为未来在音频任务中使用生成先验的工作奠定了坚实的基础。
摘要:We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of distilling a powerful generative prior into a separate parametric representation extends to the audio domain. Leveraging a single pretrained model, Audio-SDS enables a broad range of tasks without requiring specialized datasets. In particular, we demonstrate how Audio-SDS can guide physically informed impact sound simulations, calibrate FM-synthesis parameters, and perform prompt-specified source separation. Our findings illustrate the versatility of distillation-based methods across modalities and establish a robust foundation for future work using generative priors in audio tasks.


【5】 Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale  Data Restoration

标题: Miipher-2:用于百万小时规模数据恢复的通用语音恢复模型
链接:https://arxiv.org/abs/2505.04457
作者: Shigeki Karita,  Yuma Koizumi,  Heiga Zen,  Haruko Ishikawa,  Robin Scheibler,  Michiel Bacchiani 
摘要:训练数据清洗是基于生成模型的语音恢复的一个新的应用。本文介绍了Miipher-2,这是一种为百万小时规模数据设计的SR模型,用于训练大型语言模型等大型生成模型的数据清洗。解决的关键挑战包括对看不见的语言的泛化,没有显式条件的操作(例如,文本、扬声器ID)和计算效率。Miipher-2利用一个冻结的、预训练的通用语音模型(USM),支持300多种语言,作为一个强大的、无条件的特征提取器。为了优化效率并最大限度地减少内存,Miipher-2集成了并行适配器,用于从噪声输入中预测干净的USM特征,并采用WaneFit神经声码器进行波形合成。这些组件在3,000小时的多语言,录音室质量的录音中进行了训练,并增加了降级,而USM参数保持不变。实验结果表明,Miipher-2的优越或可比的性能,传统的SR模型在词的错误率,说话人相似性,客观和主观的声音质量分数在所有测试的语言。Miipher-2在消费级加速器上高效运行,实现了0.0078的实时系数,仅使用100个这样的加速器就可以在大约三天内处理100万小时的语音数据集。
摘要:Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, for training data cleaning for large-scale generative models like large language models. Key challenges addressed include generalization to unseen languages, operation without explicit conditioning (e.g., text, speaker ID), and computational efficiency. Miipher-2 utilizes a frozen, pre-trained Universal Speech Model (USM), supporting over 300 languages, as a robust, conditioning-free feature extractor. To optimize efficiency and minimize memory, Miipher-2 incorporates parallel adapters for predicting clean USM features from noisy inputs and employs the WaneFit neural vocoder for waveform synthesis. These components were trained on 3,000 hours of multi-lingual, studio-quality recordings with augmented degradations, while USM parameters remained fixed. Experimental results demonstrate Miipher-2's superior or comparable performance to conventional SR models in word-error-rate, speaker similarity, and both objective and subjective sound quality scores across all tested languages. Miipher-2 operates efficiently on consumer-grade accelerators, achieving a real-time factor of 0.0078, enabling the processing of a million-hour speech dataset in approximately three days using only 100 such accelerators.


【6】 Automatic Music Transcription using Convolutional Neural Networks and  Constant-Q transform

标题: 使用卷积神经网络和Constant-Q变换的自动音乐转录
链接:https://arxiv.org/abs/2505.04451
作者: Yohannis Telila,  Tommaso Cucinotta,  Davide Bacciu 
备注:6 pages
摘要:自动音乐转录(AMT)是分析音乐作品的音频记录并检测正在播放的音符的问题。AMT是一个具有挑战性的问题,特别是在复调音乐方面。AMT的目标是通过分析同时播放的包含多个音符的声音信号来产生音乐作品的乐谱表示。在这项工作中,我们设计了一个处理管道,可以将古典钢琴音频文件转换为.wav格式的乐谱表示。使用恒定Q变换从音频信号提取特征,并且将所得系数用作卷积神经网络(CNN)模型的输入。
摘要:Automatic music transcription (AMT) is the problem of analyzing an audio recording of a musical piece and detecting notes that are being played. AMT is a challenging problem, particularly when it comes to polyphonic music. The goal of AMT is to produce a score representation of a music piece, by analyzing a sound signal containing multiple notes played simultaneously. In this work, we design a processing pipeline that can transform classical piano audio files in .wav format into a music score representation. The features from the audio signals are extracted using the constant-Q transform, and the resulting coefficients are used as an input to the convolutional neural network (CNN) model.


【7】 ELGAR: Expressive Cello Performance Motion Generation for Audio  Rendition

标题: ELGAR:音频演绎的表现力大提琴表演动作生成
链接:https://arxiv.org/abs/2505.04203
作者: Zhiping Qiu,  Yitong Jin,  Yuan Wang,  Yi Shi,  Chongwu Wang,  Chao Tan,  Xiaobing Li,  Feng Yu,  Tao Yu,  Qionghai Dai 
备注:None
摘要:乐器演奏艺术是人类创造力和情感的生动体现。尽管如此,生成乐器演奏动作是一项极具挑战性的任务,因为它不仅需要捕捉复杂的动作,还需要重建演奏者-乐器交互的复杂动态。虽然现有的作品主要集中在部分身体运动建模,我们提出了表达细胞性能运动生成音频再现(ELGAR),一个国家的最先进的扩散为基础的框架,全身细粒度的乐器性能运动生成仅从音频。为了强调乐器演奏的交互性,我们引入了手交互接触损失(Hand Interactive Contact Loss,简称CARL)和弓交互接触损失(Bow Interactive Contact Loss,简称BICL),有效地保证了相互作用的真实性。此外,为了更好地评估所生成的运动是否与音乐音频的语义上下文相一致,我们专门为弦乐器演奏运动生成设计了新的指标,包括手指接触距离,弓弦距离和鞠躬得分。进行了广泛的评估和消融研究,以验证所提出的方法的疗效。此外,我们提出了一个运动生成数据集SPD-GEN,整理和规范化的MoCap数据集SPD。事实证明,ELGAR在生成具有复杂和快速交互的乐器演奏动作方面显示出巨大的潜力,这将促进动画,音乐教育,互动艺术创作等领域的进一步发展。
摘要:The art of instrument performance stands as a vivid manifestation of human creativity and emotion. Nonetheless, generating instrument performance motions is a highly challenging task, as it requires not only capturing intricate movements but also reconstructing the complex dynamics of the performer-instrument interaction. While existing works primarily focus on modeling partial body motions, we propose Expressive ceLlo performance motion Generation for Audio Rendition (ELGAR), a state-of-the-art diffusion-based framework for whole-body fine-grained instrument performance motion generation solely from audio. To emphasize the interactive nature of the instrument performance, we introduce Hand Interactive Contact Loss (HICL) and Bow Interactive Contact Loss (BICL), which effectively guarantee the authenticity of the interplay. Moreover, to better evaluate whether the generated motions align with the semantic context of the music audio, we design novel metrics specifically for string instrument performance motion generation, including finger-contact distance, bow-string distance, and bowing score. Extensive evaluations and ablation studies are conducted to validate the efficacy of the proposed methods. In addition, we put forward a motion generation dataset SPD-GEN, collated and normalized from the MoCap dataset SPD. As demonstrated, ELGAR has shown great potential in generating instrument performance motions with complicated and fast interactions, which will promote further development in areas such as animation, music education, interactive art creation, etc.


【8】 Advancing Zero-shot Text-to-Speech Intelligibility across Diverse  Domains via Preference Alignment

标题: 通过偏好对齐提高跨不同领域的Zero-Shot文本到语音的可理解性
链接:https://arxiv.org/abs/2505.04113
作者: Xueyao Zhang,  Yuancheng Wang,  Chaoren Wang,  Ziniu Li,  Zhuo Chen,  Zhizheng Wu 
摘要:现代zero-shot文本到语音(TTS)系统尽管使用了广泛的预训练,但通常在具有挑战性的场景中挣扎,例如绕口令、重复单词、代码切换和跨语言合成,从而导致可理解性问题。为了解决这些限制,本文利用偏好对齐技术,该技术能够有针对性地构建预训练分布外的数据以提高性能。我们引入了一个新的数据集,命名为可懂度偏好语音数据集(INTP),并扩展了直接偏好优化(DPO)框架,以适应不同的TTS架构。在INTP对齐之后,除了可懂度之外,我们还观察到跨不同领域的多个TTS模型的整体改善,包括自然度,相似性和音频质量。在此基础上,我们还验证了INTP对于更易于理解的模型(如CosyVoice 2和Ints)的弱到强的泛化能力。此外,我们展示了通过基于Ints的迭代对齐进行进一步改进的潜力。音频样本可在https://intalign.github.io/上获得。
摘要:Modern zero-shot text-to-speech (TTS) systems, despite using extensive pre-training, often struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis, leading to intelligibility issues. To address these limitations, this paper leverages preference alignment techniques, which enable targeted construction of out-of-pretraining-distribution data to enhance performance. We introduce a new dataset, named the Intelligibility Preference Speech Dataset (INTP), and extend the Direct Preference Optimization (DPO) framework to accommodate diverse TTS architectures. After INTP alignment, in addition to intelligibility, we observe overall improvements including naturalness, similarity, and audio quality for multiple TTS models across diverse domains. Based on that, we also verify the weak-to-strong generalization ability of INTP for more intelligible models such as CosyVoice 2 and Ints. Moreover, we showcase the potential for further improvements through iterative alignment based on Ints. Audio samples are available at https://intalign.github.io/.


eess.AS音频处理


【1】 EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via  Reinforcement Learning
标题: EchoInk-R1:通过强化学习探索多模式LLM中的视听推理
链接:https://arxiv.org/abs/2505.04623
作者: Zhenghao Xing,  Xiaowei Hu,  Chi-Wing Fu,  Wenhai Wang,  Jifeng Dai,  Pheng-Ann Heng 
摘要:多模态大型语言模型(MLLM)具有跨文本、视觉和音频的高级感知能力,但它们通常难以进行结构化的跨模态推理,特别是在集成音频和视觉信号时。我们介绍EchoInk-R1,这是一个强化学习框架,可以增强MLLM中的这种推理。EchoInk-R1基于Qwen2.5-Omni-7 B基础,并使用组相对策略优化(GRPO)进行了优化,可以在同步的音频图像对上处理多项选择题。为了实现这一点,我们策划了AVQA-R1- 6 K,这是一个将此类音频图像输入与OmniInstruct-v1中的多项选择题配对的数据集。EchoInk-R1- 7 B在验证集上实现了85.77%的准确率,优于基础模型,后者的得分为80.53%,仅使用了562个强化学习步骤。除了准确性之外,EchoInk-R1还通过重新审视初始解释并在面对模糊的多模态输入时改进响应来展示反思性推理。这些结果表明,轻量级强化学习微调增强了MLLM中的跨模态推理。EchoInk-R1是第一个通过强化学习统一音频,视觉和文本形式的通用开放世界推理框架。代码和数据公开发布,以方便进一步的研究。
摘要:Multimodal large language models (MLLMs) have advanced perception across text, vision, and audio, yet they often struggle with structured cross-modal reasoning, particularly when integrating audio and visual signals. We introduce EchoInk-R1, a reinforcement learning framework that enhances such reasoning in MLLMs. Built upon the Qwen2.5-Omni-7B foundation and optimized with Group Relative Policy Optimization (GRPO), EchoInk-R1 tackles multiple-choice question answering over synchronized audio-image pairs. To enable this, we curate AVQA-R1-6K, a dataset pairing such audio-image inputs with multiple-choice questions derived from OmniInstruct-v1. EchoInk-R1-7B achieves 85.77% accuracy on the validation set, outperforming the base model, which scores 80.53%, using only 562 reinforcement learning steps. Beyond accuracy, EchoInk-R1 demonstrates reflective reasoning by revisiting initial interpretations and refining responses when facing ambiguous multimodal inputs. These results suggest that lightweight reinforcement learning fine-tuning enhances cross-modal reasoning in MLLMs. EchoInk-R1 is the first framework to unify audio, visual, and textual modalities for general open-world reasoning via reinforcement learning. Code and data are publicly released to facilitate further research.


【2】 Accelerating Audio Research with Robotic Dummy Heads

标题: 利用机器人假人头加速音频研究
链接:https://arxiv.org/abs/2505.04548
作者: Austin Lu,  Kanad Sarkar,  Yongjie Zhuang,  Leo Lin,  Ryan M Corey,  Andrew C Singer 
备注:WASPAA 2025
摘要:这项工作介绍了一个机器人假人头,融合了传统的听觉人体模型与机器人的流动性的声学现实主义。该设备能够像人一样移动,说话和倾听,并可用于自动化空间静止音频实验,从而加快音频研究的步伐。重要的是,由于其安静的电机,该设备还可以在动态实验中用作移动声源。这一功能使我们的工作与之前的机器人声学研究平台有所区别。通过各种实验和声学测量来验证机器人能够进行高质量的音频数据收集。这些实验还演示了机器人如何用于研究自适应双耳波束形成。设计文件作为开源提供,以刺激新的音频研究。
摘要:This work introduces a robotic dummy head that fuses the acoustic realism of conventional audiological mannequins with the mobility of robots. The proposed device is capable of moving, talking, and listening as people do, and can be used to automate spatially-stationary audio experiments, thus accelerating the pace of audio research. Critically, the device may also be used as a moving sound source in dynamic experiments, due to its quiet motor. This feature differentiates our work from previous robotic acoustic research platforms. Validation that the robot enables high quality audio data collection is provided through various experiments and acoustic measurements. These experiments also demonstrate how the robot might be used to study adaptive binaural beamforming. Design files are provided as open-source to stimulate novel audio research.


【3】 Recognizing Ornaments in Vocal Indian Art Music with Active Annotation

标题: 用主动注释识别印度声乐艺术音乐中的装饰
链接:https://arxiv.org/abs/2505.04419
作者: Sumit Kumar,  Parampreet Singh,  Vipul Arora 
摘要:装饰、修饰或微音调变化对于许多音乐传统的旋律表达至关重要,为表演增加了深度、细微差别和情感影响。识别歌唱声音中的特征是MIR的关键,在音乐教学,歌手识别,流派分类和控制歌唱声音生成方面具有潜在的应用。然而,缺乏注释数据集和专门的建模方法仍然是这一研究领域取得进展的主要障碍。在这项工作中,我们介绍R\=aga装饰检测(ROD),一个新的数据集,包括印度古典音乐录音策划的专家音乐家。使用自定义的Human-in-the-Loop工具对数据集进行注释,用于标记为基于事件的标签的六个声乐装饰。使用这个数据集,我们开发了一个基于深度时间序列分析的装饰检测模型,在长音频记录的分块过程中保留装饰边界。我们在ROD数据集内使用不同的训练测试配置进行实验,并在印度古典音乐会录音的单独手动注释数据集上评估我们的方法。我们的实验结果支持我们提出的方法优于基线CRNN的性能。
摘要:Ornamentations, embellishments, or microtonal inflections are essential to melodic expression across many musical traditions, adding depth, nuance, and emotional impact to performances. Recognizing ornamentations in singing voices is key to MIR, with potential applications in music pedagogy, singer identification, genre classification, and controlled singing voice generation. However, the lack of annotated datasets and specialized modeling approaches remains a major obstacle for progress in this research area. In this work, we introduce R\=aga Ornamentation Detection (ROD), a novel dataset comprising Indian classical music recordings curated by expert musicians. The dataset is annotated using a custom Human-in-the-Loop tool for six vocal ornaments marked as event-based labels. Using this dataset, we develop an ornamentation detection model based on deep time-series analysis, preserving ornament boundaries during the chunking of long audio recordings. We conduct experiments using different train-test configurations within the ROD dataset and also evaluate our approach on a separate, manually annotated dataset of Indian classical concert recordings. Our experimental results support the superior performance of our proposed approach over the baseline CRNN.


【4】 Discrete Optimal Transport and Voice Conversion

标题: 离散最佳传输和语音转换
链接:https://arxiv.org/abs/2505.04382
作者: Anton Selitskiy,  Maitreya Kocharekar 
备注:4 pages, 6 figures, 1 table
摘要:在这项工作中,我们解决了语音转换(VC)的任务,使用基于矢量的接口。为了使扬声器之间的音频嵌入对齐,我们采用离散最优传输映射。我们的评估结果证明了这种方法的高质量和有效性。此外,我们表明,在音频生成中应用离散最优传输作为后处理步骤可能会导致合成音频作为真实的不正确分类。
摘要:In this work, we address the voice conversion (VC) task using a vector-based interface. To align audio embeddings between speakers, we employ discrete optimal transport mapping. Our evaluation results demonstrate the high quality and effectiveness of this method. Additionally, we show that applying discrete optimal transport as a post-processing step in audio generation can lead to the incorrect classification of synthetic audio as real.


【5】 Robust Speech Recognition with Schrödinger Bridge-Based Speech  Enhancement

标题: 基于薛定格桥的语音增强的鲁棒语音识别
链接:https://arxiv.org/abs/2505.04237
作者: Rauf Nasretdinov,  Roman Korostik,  Ante Jukić 
备注:5 pages. Published in ICASSP 2025
摘要:在这项工作中,我们研究应用生成语音增强,以提高噪声和混响条件下的ASR模型的鲁棒性。我们采用了最近提出的语音增强模型的基础上Schr\“odinger桥,这已被证明是表现良好的扩散为基础的方法相比。我们分析了模型缩放和不同采样方法对ASR性能的影响。此外,我们将所考虑的模型与预测和基于扩散的基线进行了比较,并分析了使用不同预训练ASR模型时的语音识别性能。所提出的方法显着降低了单词错误率,相对于未处理的语音信号降低了约40%,相对于类似大小的预测方法降低了约8%。
摘要:In this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schr\"odinger bridge, which has been shown to perform well compared to diffusion-based approaches. We analyze the impact of model scaling and different sampling methods on the ASR performance. Furthermore, we compare the considered model with predictive and diffusion-based baselines and analyze the speech recognition performance when using different pre-trained ASR models. The proposed approach significantly reduces the word error rate, reducing it by approximately 40% relative to the unprocessed speech signals and by approximately 8% relative to a similarly sized predictive approach.


【6】 Aliasing Reduction in Neural Amp Modeling by Smoothing Activations

标题: 通过平滑激活减少神经元建模中的混叠
链接:https://arxiv.org/abs/2505.04082
作者: Ryota Sato,  Julius O. Smith III 
备注:Accepted to DAFx 2025
摘要:对模拟音频硬件(如老式吉他放大器)的高质量数字仿真的需求日益增长,导致了基于神经网络的黑盒建模的大量工作,WaveNet等深度学习架构显示出了良好的效果。然而,所有这些模型中的一个关键限制是在神经网络中使用非线性激活函数所产生的混叠伪影。在本文中,我们研究了新的和修改的激活函数,旨在减轻神经放大器模型内的混叠。为了支持这一点,我们引入了一个新的度量,混叠信号比(ASR),它定量评估混叠的水平,具有高精度。测量传统的错误信号比(ESR),我们进行了一系列的预先存在的和现代的激活功能与不同的拉伸因子的研究。我们的研究结果证实,具有更平滑曲线的激活函数倾向于实现更低的ASR值,这表明混叠明显减少。值得注意的是,这种混叠减少的改进是在没有显著增加ESR的情况下实现的,这表明在神经放大器模型中具有减少混叠的高建模精度的潜力。
摘要:The increasing demand for high-quality digital emulations of analog audio hardware such as vintage guitar amplifiers has led to numerous works in neural-network-based black-box modeling, with deep learning architectures like WaveNet showing promising results. However, a key limitation in all of these models is the aliasing artifacts that arise from the use of nonlinear activation functions in neural networks. In this paper, we investigate novel and modified activation functions aimed at mitigating aliasing within neural amplifier models. Supporting this, we introduce a novel metric, the Aliasing-to-Signal Ratio (ASR), which quantitatively assesses the level of aliasing with high accuracy. Measuring also the conventional Error-to-Signal Ratio (ESR), we conducted studies on a range of preexisting and modern activation functions with varying stretch factors. Our findings confirmed that activation functions with smoother curves tend to achieve lower ASR values, indicating a noticeable reduction in aliasing. Notably, this improvement in aliasing reduction was achievable without a substantial increase in ESR, demonstrating the potential for high modeling accuracy with reduced aliasing in neural amp models.


【7】 Score Distillation Sampling for Audio: Source Separation, Synthesis, and  Beyond

标题: 音频的分数蒸馏采样:源分离、合成等
链接:https://arxiv.org/abs/2505.04621
作者: Jessie Richter-Powell,  Antonio Torralba,  Jonathan Lorraine 
备注:See the project website at this https URL
摘要:我们介绍Audio-SDS,分数蒸馏采样(SDS)的推广文本条件音频扩散模型。虽然SDS最初是为使用图像扩散的文本到3D生成而设计的,但其将强大的生成先验提取为单独的参数表示的核心思想扩展到了音频领域。利用单个预训练模型,Audio-SDS可以实现广泛的任务,而无需专门的数据集。特别是,我们演示了如何Audio-SDS可以指导物理知情的影响声音模拟,校准FM合成参数,并执行指定的源分离。我们的研究结果说明了基于蒸馏的方法在不同模态中的多功能性,并为未来在音频任务中使用生成先验的工作奠定了坚实的基础。
摘要:We introduce Audio-SDS, a generalization of Score Distillation Sampling (SDS) to text-conditioned audio diffusion models. While SDS was initially designed for text-to-3D generation using image diffusion, its core idea of distilling a powerful generative prior into a separate parametric representation extends to the audio domain. Leveraging a single pretrained model, Audio-SDS enables a broad range of tasks without requiring specialized datasets. In particular, we demonstrate how Audio-SDS can guide physically informed impact sound simulations, calibrate FM-synthesis parameters, and perform prompt-specified source separation. Our findings illustrate the versatility of distillation-based methods across modalities and establish a robust foundation for future work using generative priors in audio tasks.


【8】 Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale  Data Restoration

标题: Miipher-2:用于百万小时规模数据恢复的通用语音恢复模型
链接:https://arxiv.org/abs/2505.04457
作者: Shigeki Karita,  Yuma Koizumi,  Heiga Zen,  Haruko Ishikawa,  Robin Scheibler,  Michiel Bacchiani 
摘要:训练数据清洗是基于生成模型的语音恢复的一个新的应用。本文介绍了Miipher-2,这是一种为百万小时规模数据设计的SR模型,用于训练大型语言模型等大型生成模型的数据清洗。解决的关键挑战包括对看不见的语言的泛化,没有显式条件的操作(例如,文本、扬声器ID)和计算效率。Miipher-2利用一个冻结的、预训练的通用语音模型(USM),支持300多种语言,作为一个强大的、无条件的特征提取器。为了优化效率并最大限度地减少内存,Miipher-2集成了并行适配器,用于从噪声输入中预测干净的USM特征,并采用WaneFit神经声码器进行波形合成。这些组件在3,000小时的多语言,录音室质量的录音中进行了训练,并增加了降级,而USM参数保持不变。实验结果表明,Miipher-2的优越或可比的性能,传统的SR模型在词的错误率,说话人相似性,客观和主观的声音质量分数在所有测试的语言。Miipher-2在消费级加速器上高效运行,实现了0.0078的实时系数,仅使用100个这样的加速器就可以在大约三天内处理100万小时的语音数据集。
摘要:Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, for training data cleaning for large-scale generative models like large language models. Key challenges addressed include generalization to unseen languages, operation without explicit conditioning (e.g., text, speaker ID), and computational efficiency. Miipher-2 utilizes a frozen, pre-trained Universal Speech Model (USM), supporting over 300 languages, as a robust, conditioning-free feature extractor. To optimize efficiency and minimize memory, Miipher-2 incorporates parallel adapters for predicting clean USM features from noisy inputs and employs the WaneFit neural vocoder for waveform synthesis. These components were trained on 3,000 hours of multi-lingual, studio-quality recordings with augmented degradations, while USM parameters remained fixed. Experimental results demonstrate Miipher-2's superior or comparable performance to conventional SR models in word-error-rate, speaker similarity, and both objective and subjective sound quality scores across all tested languages. Miipher-2 operates efficiently on consumer-grade accelerators, achieving a real-time factor of 0.0078, enabling the processing of a million-hour speech dataset in approximately three days using only 100 such accelerators.


【9】 Automatic Music Transcription using Convolutional Neural Networks and  Constant-Q transform

标题: 使用卷积神经网络和Constant-Q变换的自动音乐转录
链接:https://arxiv.org/abs/2505.04451
作者: Yohannis Telila,  Tommaso Cucinotta,  Davide Bacciu 
备注:6 pages
摘要:自动音乐转录(AMT)是分析音乐作品的音频记录并检测正在播放的音符的问题。AMT是一个具有挑战性的问题,特别是在复调音乐方面。AMT的目标是通过分析同时播放的包含多个音符的声音信号来产生音乐作品的乐谱表示。在这项工作中,我们设计了一个处理管道,可以将古典钢琴音频文件转换为.wav格式的乐谱表示。使用恒定Q变换从音频信号提取特征,并且将所得系数用作卷积神经网络(CNN)模型的输入。
摘要:Automatic music transcription (AMT) is the problem of analyzing an audio recording of a musical piece and detecting notes that are being played. AMT is a challenging problem, particularly when it comes to polyphonic music. The goal of AMT is to produce a score representation of a music piece, by analyzing a sound signal containing multiple notes played simultaneously. In this work, we design a processing pipeline that can transform classical piano audio files in .wav format into a music score representation. The features from the audio signals are extracted using the constant-Q transform, and the resulting coefficients are used as an input to the convolutional neural network (CNN) model.


【10】 SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin  Transformer

标题: SwinLip:一种使用Swin Transformer进行唇读的高效视觉语音编码器
链接:https://arxiv.org/abs/2505.04394
作者: Young-Hu Park,  Rae-Hong Park,  Hyung-Min Park 
摘要:本文提出了一种高效的唇读视觉语音编码器。虽然最近的唇读研究已经基于ResNet架构,并取得了显着的成功,他们不足以有效地捕捉唇读功能,由于高计算复杂性建模的时空信息。此外,使用复杂的视觉模型不仅增加了唇读模型的复杂性,而且还在多模态研究的整个网络中引起延迟(例如,视听语音识别、语音增强和语音分离)。为了克服基于卷积神经网络(CNN)模型的局限性,我们将Swin Transformer的分层结构和窗口自注意应用于唇读。我们配置了一个新的轻量级规模的Swin Transformer适合处理唇读数据,并提出了SwinLip视觉语音编码器,它有效地降低了计算负荷,通过集成修改后的卷积增强的Transformer(构象)的时间嵌入与传统的空间嵌入的层次结构。通过大量的实验,我们已经验证了我们的SwinLip成功地提高了唇读网络的性能和推理速度,当应用于单词和句子识别的各种骨干时,减少了计算负载。特别是,我们的SwinLip在英语LRW和普通话LRW-1000数据集上都表现出了强大的性能,并在普通话LRW-1000数据集上实现了最先进的性能,与现有的最先进的模型相比,计算量更少。
摘要:This paper presents an efficient visual speech encoder for lip reading. While most recent lip reading studies have been based on the ResNet architecture and have achieved significant success, they are not sufficiently suitable for efficiently capturing lip reading features due to high computational complexity in modeling spatio-temporal information. Additionally, using a complex visual model not only increases the complexity of lip reading models but also induces delays in the overall network for multi-modal studies (e.g., audio-visual speech recognition, speech enhancement, and speech separation). To overcome the limitations of Convolutional Neural Network (CNN)-based models, we apply the hierarchical structure and window self-attention of the Swin Transformer to lip reading. We configure a new lightweight scale of the Swin Transformer suitable for processing lip reading data and present the SwinLip visual speech encoder, which efficiently reduces computational load by integrating modified Convolution-augmented Transformer (Conformer) temporal embeddings with conventional spatial embeddings in the hierarchical structure. Through extensive experiments, we have validated that our SwinLip successfully improves the performance and inference speed of the lip reading network when applied to various backbones for word and sentence recognition, reducing computational load. In particular, our SwinLip demonstrated robust performance in both English LRW and Mandarin LRW-1000 datasets and achieved state-of-the-art performance on the Mandarin LRW-1000 dataset with less computation compared to the existing state-of-the-art model.


【11】 ELGAR: Expressive Cello Performance Motion Generation for Audio  Rendition

标题: ELGAR:音频演绎的表现力大提琴表演动作生成
链接:https://arxiv.org/abs/2505.04203
作者: Zhiping Qiu,  Yitong Jin,  Yuan Wang,  Yi Shi,  Chongwu Wang,  Chao Tan,  Xiaobing Li,  Feng Yu,  Tao Yu,  Qionghai Dai 
备注:None
摘要:乐器演奏艺术是人类创造力和情感的生动体现。尽管如此,生成乐器演奏动作是一项极具挑战性的任务,因为它不仅需要捕捉复杂的动作,还需要重建演奏者-乐器交互的复杂动态。虽然现有的作品主要集中在部分身体运动建模,我们提出了表达细胞性能运动生成音频再现(ELGAR),一个国家的最先进的扩散为基础的框架,全身细粒度的乐器性能运动生成仅从音频。为了强调乐器演奏的交互性,我们引入了手交互接触损失(Hand Interactive Contact Loss,简称CARL)和弓交互接触损失(Bow Interactive Contact Loss,简称BICL),有效地保证了相互作用的真实性。此外,为了更好地评估所生成的运动是否与音乐音频的语义上下文相一致,我们专门为弦乐器演奏运动生成设计了新的指标,包括手指接触距离,弓弦距离和鞠躬得分。进行了广泛的评估和消融研究,以验证所提出的方法的疗效。此外,我们提出了一个运动生成数据集SPD-GEN,整理和规范化的MoCap数据集SPD。事实证明,ELGAR在生成具有复杂和快速交互的乐器演奏动作方面显示出巨大的潜力,这将促进动画,音乐教育,互动艺术创作等领域的进一步发展。
摘要:The art of instrument performance stands as a vivid manifestation of human creativity and emotion. Nonetheless, generating instrument performance motions is a highly challenging task, as it requires not only capturing intricate movements but also reconstructing the complex dynamics of the performer-instrument interaction. While existing works primarily focus on modeling partial body motions, we propose Expressive ceLlo performance motion Generation for Audio Rendition (ELGAR), a state-of-the-art diffusion-based framework for whole-body fine-grained instrument performance motion generation solely from audio. To emphasize the interactive nature of the instrument performance, we introduce Hand Interactive Contact Loss (HICL) and Bow Interactive Contact Loss (BICL), which effectively guarantee the authenticity of the interplay. Moreover, to better evaluate whether the generated motions align with the semantic context of the music audio, we design novel metrics specifically for string instrument performance motion generation, including finger-contact distance, bow-string distance, and bowing score. Extensive evaluations and ablation studies are conducted to validate the efficacy of the proposed methods. In addition, we put forward a motion generation dataset SPD-GEN, collated and normalized from the MoCap dataset SPD. As demonstrated, ELGAR has shown great potential in generating instrument performance motions with complicated and fast interactions, which will promote further development in areas such as animation, music education, interactive art creation, etc.


【12】 Advancing Zero-shot Text-to-Speech Intelligibility across Diverse  Domains via Preference Alignment

标题: 通过偏好对齐提高跨不同领域的Zero-Shot文本到语音的可理解性
链接:https://arxiv.org/abs/2505.04113
作者: Xueyao Zhang,  Yuancheng Wang,  Chaoren Wang,  Ziniu Li,  Zhuo Chen,  Zhizheng Wu 
摘要:现代zero-shot文本到语音(TTS)系统尽管使用了广泛的预训练,但通常在具有挑战性的场景中挣扎,例如绕口令、重复单词、代码切换和跨语言合成,从而导致可理解性问题。为了解决这些限制,本文利用偏好对齐技术,该技术能够有针对性地构建预训练分布外的数据以提高性能。我们引入了一个新的数据集,命名为可懂度偏好语音数据集(INTP),并扩展了直接偏好优化(DPO)框架,以适应不同的TTS架构。在INTP对齐之后,除了可懂度之外,我们还观察到跨不同领域的多个TTS模型的整体改善,包括自然度,相似性和音频质量。在此基础上,我们还验证了INTP对于更易于理解的模型(如CosyVoice 2和Ints)的弱到强的泛化能力。此外,我们展示了通过基于Ints的迭代对齐进行进一步改进的潜力。音频样本可在https://intalign.github.io/上获得。
摘要:Modern zero-shot text-to-speech (TTS) systems, despite using extensive pre-training, often struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis, leading to intelligibility issues. To address these limitations, this paper leverages preference alignment techniques, which enable targeted construction of out-of-pretraining-distribution data to enhance performance. We introduce a new dataset, named the Intelligibility Preference Speech Dataset (INTP), and extend the Direct Preference Optimization (DPO) framework to accommodate diverse TTS architectures. After INTP alignment, in addition to intelligibility, we observe overall improvements including naturalness, similarity, and audio quality for multiple TTS models across diverse domains. Based on that, we also verify the weak-to-strong generalization ability of INTP for more intelligible models such as CosyVoice 2 and Ints. Moreover, we showcase the potential for further improvements through iterative alignment based on Ints. Audio samples are available at https://intalign.github.io/.


【13】 LLAMAPIE: Proactive In-Ear Conversation Assistants

标题: LLAMAPIE:积极主动的入耳式对话助理
链接:https://arxiv.org/abs/2505.04066
作者: Tuochao Chen,  Nicholas Batchelder,  Alisa Liu,  Noah Smith,  Shyamnath Gollakota 
摘要:我们介绍LlamaPIE,这是第一个实时主动助理,旨在通过听觉设备提供的谨慎,简洁的指导来增强人类对话。与需要显式用户调用的传统语言模型不同,这个助手在后台运行,预测用户需求而不中断对话。我们解决了几个挑战,包括确定何时响应,制作简洁的响应以增强对话,利用用户的知识进行上下文感知帮助,以及实时的设备上处理。为了实现这一目标,我们构建了一个半合成对话数据集,并提出了一个双模型管道:一个决定何时响应的小模型和一个生成响应的大模型。我们在真实世界的数据集上评估了我们的方法,证明了它在提供有用的、不显眼的帮助方面的有效性。在Apple Silicon M2硬件上实施的对我们助手的用户研究显示,与没有帮助的基线和反应模型相比,我们对主动式助手有强烈的偏好,这突出了LlamaPie增强实时对话的潜力。
摘要:We introduce LlamaPIE, the first real-time proactive assistant designed to enhance human conversations through discreet, concise guidance delivered via hearable devices. Unlike traditional language models that require explicit user invocation, this assistant operates in the background, anticipating user needs without interrupting conversations. We address several challenges, including determining when to respond, crafting concise responses that enhance conversations, leveraging knowledge of the user for context-aware assistance, and real-time, on-device processing. To achieve this, we construct a semi-synthetic dialogue dataset and propose a two-model pipeline: a small model that decides when to respond and a larger model that generates the response. We evaluate our approach on real-world datasets, demonstrating its effectiveness in providing helpful, unobtrusive assistance. User studies with our assistant, implemented on Apple Silicon M2 hardware, show a strong preference for the proactive assistant over both a baseline with no assistance and a reactive model, highlighting the potential of LlamaPie to enhance live conversations.


机器翻译由腾讯交互翻译提供,仅供参考