微信公众号:arXiv_Daily
cs.SD语音
标题: 用有限的数据学习音乐音频表示
链接:https://arxiv.org/abs/2505.06042
备注:Presented at ICASSP 2025
摘要:大型音乐深度学习模型,包括那些专注于学习通用音乐音频表示的模型,通常被认为需要大量的训练数据才能实现高性能。如果是真的,这将在音频数据或注释稀缺的情况下带来挑战,例如代表性不足的音乐传统,非流行流派以及个性化音乐创作和收听。了解这些模型在有限数据场景中的行为对于开发解决这些问题的技术至关重要。 在这项工作中,我们研究了有限数据学习制度下的几种音乐音频表示模型的行为。我们考虑具有各种架构、训练范例和输入持续时间的音乐模型,并在5到8,000分钟长的数据收集上对它们进行训练。我们评估了各种音乐信息检索任务的学习表示,并分析其对噪声的鲁棒性。我们表明,在某些条件下,来自有限数据甚至随机模型的表示与来自大数据集模型的表示相比,表现更好,尽管手工制作的特征在某些任务中优于所有学习的表示。
摘要:Large deep-learning models for music, including those focused on learning general-purpose music audio representations, are often assumed to require substantial training data to achieve high performance. If true, this would pose challenges in scenarios where audio data or annotations are scarce, such as for underrepresented music traditions, non-popular genres, and personalized music creation and listening. Understanding how these models behave in limited-data scenarios could be crucial for developing techniques to tackle them. In this work, we investigate the behavior of several music audio representation models under limited-data learning regimes. We consider music models with various architectures, training paradigms, and input durations, and train them on data collections ranging from 5 to 8,000 minutes long. We evaluate the learned representations on various music information retrieval tasks and analyze their robustness to noise. We show that, under certain conditions, representations from limited-data and even random models perform comparably to ones from large-dataset models, though handcrafted features outperform all learned representations in some tasks.
【2】 Fast Differentiable Modal Simulation of Non-linear Strings, Membranes, and Plates
标题: 非线性弦、膜和板的快速微振模式模拟链接:https://arxiv.org/abs/2505.05940
备注:accepted to DAFx 2025
摘要:用于模拟弦、膜和板的振动的模态方法广泛用于声学和物理上知情的音频合成。然而,传统的实现,特别是对于非线性模型,如冯K\'arm\' an板,计算要求高,缺乏可微性,限制了逆建模和实时应用。我们引入了一个快速,可微分,GPU加速的模态框架,使用JAX库构建,提供高效的模拟,并实现基于梯度的逆向建模。基准测试表明,我们的方法显着优于基于CPU和GPU的实现,特别是对于许多模式的模拟。逆向建模实验表明,我们的方法可以恢复的物理参数,包括张力,刚度和几何形状,从合成和实验数据。虽然拟合的物理参数是更敏感的初始化相比,其他方法,它提供了更大的可解释性和更紧凑的参数化。该代码作为开源发布,以支持未来在可微物理建模和声音合成方面的研究和应用。
摘要:Modal methods for simulating vibrations of strings, membranes, and plates are widely used in acoustics and physically informed audio synthesis. However, traditional implementations, particularly for non-linear models like the von K\'arm\'an plate, are computationally demanding and lack differentiability, limiting inverse modelling and real-time applications. We introduce a fast, differentiable, GPU-accelerated modal framework built with the JAX library, providing efficient simulations and enabling gradient-based inverse modelling. Benchmarks show that our approach significantly outperforms CPU and GPU-based implementations, particularly for simulations with many modes. Inverse modelling experiments demonstrate that our approach can recover physical parameters, including tension, stiffness, and geometry, from both synthetic and experimental data. Although fitting physical parameters is more sensitive to initialisation compared to other methods, it provides greater interpretability and more compact parameterisation. The code is released as open source to support future research and applications in differentiable physical modelling and sound synthesis.
【3】 Toward a Sparse and Interpretable Audio Codec
标题: 迈向稀疏和可解释的音频编解码器链接:https://arxiv.org/abs/2505.05654
摘要:大多数广泛使用的现代音频编解码器,如Ogg Vorbis和MP3,以及最近的“神经”编解码器,如Meta的Encodec或Descript Audio Codec都是基于块编码的;音频被分成重叠的,固定大小的“帧”,然后被压缩。虽然它们通常可以产生出色的再现,并可用于文本到音频等下游任务,但它们不会产生直观的,可直接解释的表示。在这项工作中,我们介绍了一个概念验证的音频编码器,表示音频作为一个稀疏的事件集和它们的发生时间。基本的基于物理学的假设被用来模拟攻击和物理共振的乐器演奏和表演发生的房间,希望鼓励稀疏,简约,易于解释的表示。
摘要:Most widely-used modern audio codecs, such as Ogg Vorbis and MP3, as well as more recent "neural" codecs like Meta's Encodec or the Descript Audio Codec are based on block-coding; audio is divided into overlapping, fixed-size "frames" which are then compressed. While they often yield excellent reproductions and can be used for downstream tasks such as text-to-audio, they do not produce an intuitive, directly-interpretable representation. In this work, we introduce a proof-of-concept audio encoder that represents audio as a sparse set of events and their times-of-occurrence. Rudimentary physics-based assumptions are used to model attack and the physical resonance of both the instrument being played and the room in which a performance occurs, hopefully encouraging a sparse, parsimonious, and easy-to-interpret representation.
【4】 Unsupervised Blind Speech Separation with a Diffusion Prior
标题: 具有扩散先验的无监督盲语音分离链接:https://arxiv.org/abs/2505.05657
备注:Paper Accepted at ICML2025 Demo: this https URL Code: this https URL
摘要:盲语音分离(BSS)的目的是从麦克风阵列记录的音频混合中分离出多个语音源。该问题具有挑战性,因为它是一个盲逆问题,即,麦克风阵列几何形状、房间脉冲响应(RIR)和语音源都是未知的。我们提出ArrayDPS解决BSS问题的无监督,数组不可知,和生成的方式。其核心思想建立在扩散后验采样(DPS)的基础上,但与可能性易于处理的DPS不同,ArrayDPS必须通过制定单独的优化问题来近似可能性。优化的解决方案近似于室内声学和麦克风之间的相对传递函数。这些近似与扩散先验一起通过ArrayDPS采样过程进行迭代,并最终产生分离的语音源。我们只需要一个简单的单扬声器语音扩散模型作为先验,以及在麦克风处记录的混合物;不需要麦克风阵列信息。评估结果表明,ArrayDPS优于所有基线无监督方法,同时在SDR方面与有监督方法相当。音频演示请访问:https://arraydps.github.io/ArrayDPSDemo/。
摘要:Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geometry, the room impulse response (RIR), and the speech sources, are all unknown. We propose ArrayDPS to solve the BSS problem in an unsupervised, array-agnostic, and generative manner. The core idea builds on diffusion posterior sampling (DPS), but unlike DPS where the likelihood is tractable, ArrayDPS must approximate the likelihood by formulating a separate optimization problem. The solution to the optimization approximates room acoustics and the relative transfer functions between microphones. These approximations, along with the diffusion priors, iterate through the ArrayDPS sampling process and ultimately yield separated voice sources. We only need a simple single-speaker speech diffusion model as a prior along with the mixtures recorded at the microphones; no microphone array information is necessary. Evaluation results show that ArrayDPS outperforms all baseline unsupervised methods while being comparable to supervised methods in terms of SDR. Audio demos are provided at: https://arraydps.github.io/ArrayDPSDemo/.
标题: 具有扩散先验的无监督盲语音分离
链接:https://arxiv.org/abs/2505.05657
备注:Paper Accepted at ICML2025 Demo: this https URL Code: this https URL
摘要:盲语音分离(BSS)的目的是从麦克风阵列记录的音频混合中分离出多个语音源。该问题具有挑战性,因为它是一个盲逆问题,即,麦克风阵列几何形状、房间脉冲响应(RIR)和语音源都是未知的。我们提出ArrayDPS解决BSS问题的无监督,数组不可知,和生成的方式。其核心思想建立在扩散后验采样(DPS)的基础上,但与可能性易于处理的DPS不同,ArrayDPS必须通过制定单独的优化问题来近似可能性。优化的解决方案近似于室内声学和麦克风之间的相对传递函数。这些近似与扩散先验一起通过ArrayDPS采样过程进行迭代,并最终产生分离的语音源。我们只需要一个简单的单扬声器语音扩散模型作为先验,以及在麦克风处记录的混合物;不需要麦克风阵列信息。评估结果表明,ArrayDPS优于所有基线无监督方法,同时在SDR方面与有监督方法相当。音频演示请访问:https://arraydps.github.io/ArrayDPSDemo/。
摘要:Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geometry, the room impulse response (RIR), and the speech sources, are all unknown. We propose ArrayDPS to solve the BSS problem in an unsupervised, array-agnostic, and generative manner. The core idea builds on diffusion posterior sampling (DPS), but unlike DPS where the likelihood is tractable, ArrayDPS must approximate the likelihood by formulating a separate optimization problem. The solution to the optimization approximates room acoustics and the relative transfer functions between microphones. These approximations, along with the diffusion priors, iterate through the ArrayDPS sampling process and ultimately yield separated voice sources. We only need a simple single-speaker speech diffusion model as a prior along with the mixtures recorded at the microphones; no microphone array information is necessary. Evaluation results show that ArrayDPS outperforms all baseline unsupervised methods while being comparable to supervised methods in terms of SDR. Audio demos are provided at: https://arraydps.github.io/ArrayDPSDemo/.
【2】 Learning Music Audio Representations With Limited Data
标题: 用有限的数据学习音乐音频表示链接:https://arxiv.org/abs/2505.06042
备注:Presented at ICASSP 2025
摘要:大型音乐深度学习模型,包括那些专注于学习通用音乐音频表示的模型,通常被认为需要大量的训练数据才能实现高性能。如果是真的,这将在音频数据或注释稀缺的情况下带来挑战,例如代表性不足的音乐传统,非流行流派以及个性化音乐创作和收听。了解这些模型在有限数据场景中的行为对于开发解决这些问题的技术至关重要。 在这项工作中,我们研究了有限数据学习制度下的几种音乐音频表示模型的行为。我们考虑具有各种架构、训练范例和输入持续时间的音乐模型,并在5到8,000分钟长的数据收集上对它们进行训练。我们评估了各种音乐信息检索任务的学习表示,并分析其对噪声的鲁棒性。我们表明,在某些条件下,来自有限数据甚至随机模型的表示与来自大数据集模型的表示相比,表现更好,尽管手工制作的特征在某些任务中优于所有学习的表示。
摘要:Large deep-learning models for music, including those focused on learning general-purpose music audio representations, are often assumed to require substantial training data to achieve high performance. If true, this would pose challenges in scenarios where audio data or annotations are scarce, such as for underrepresented music traditions, non-popular genres, and personalized music creation and listening. Understanding how these models behave in limited-data scenarios could be crucial for developing techniques to tackle them. In this work, we investigate the behavior of several music audio representation models under limited-data learning regimes. We consider music models with various architectures, training paradigms, and input durations, and train them on data collections ranging from 5 to 8,000 minutes long. We evaluate the learned representations on various music information retrieval tasks and analyze their robustness to noise. We show that, under certain conditions, representations from limited-data and even random models perform comparably to ones from large-dataset models, though handcrafted features outperform all learned representations in some tasks.
【3】 Fast Differentiable Modal Simulation of Non-linear Strings, Membranes, and Plates
标题: 非线性弦、膜和板的快速微振模式模拟链接:https://arxiv.org/abs/2505.05940
备注:accepted to DAFx 2025
摘要:用于模拟弦、膜和板的振动的模态方法广泛用于声学和物理上知情的音频合成。然而,传统的实现,特别是对于非线性模型,如冯K\'arm\' an板,计算要求高,缺乏可微性,限制了逆建模和实时应用。我们引入了一个使用JAX库构建的快速、可微分、GPU加速的模型框架,提供高效的模拟并支持基于梯度的逆建模。基准测试表明,我们的方法显着优于基于CPU和GPU的实现,特别是对于许多模式的模拟。逆向建模实验表明,我们的方法可以恢复的物理参数,包括张力,刚度和几何形状,从合成和实验数据。虽然拟合的物理参数是更敏感的初始化相比,其他方法,它提供了更大的可解释性和更紧凑的参数化。该代码作为开源发布,以支持未来在可微物理建模和声音合成方面的研究和应用。
摘要:Modal methods for simulating vibrations of strings, membranes, and plates are widely used in acoustics and physically informed audio synthesis. However, traditional implementations, particularly for non-linear models like the von K\'arm\'an plate, are computationally demanding and lack differentiability, limiting inverse modelling and real-time applications. We introduce a fast, differentiable, GPU-accelerated modal framework built with the JAX library, providing efficient simulations and enabling gradient-based inverse modelling. Benchmarks show that our approach significantly outperforms CPU and GPU-based implementations, particularly for simulations with many modes. Inverse modelling experiments demonstrate that our approach can recover physical parameters, including tension, stiffness, and geometry, from both synthetic and experimental data. Although fitting physical parameters is more sensitive to initialisation compared to other methods, it provides greater interpretability and more compact parameterisation. The code is released as open source to support future research and applications in differentiable physical modelling and sound synthesis.
【4】 Toward a Sparse and Interpretable Audio Codec
标题: 迈向稀疏和可解释的音频编解码器链接:https://arxiv.org/abs/2505.05654
摘要:大多数广泛使用的现代音频编解码器,如Ogg Vorbis和MP3,以及最近的“神经”编解码器,如Meta的Encodec或Descript Audio Codec都是基于块编码的;音频被分成重叠的,固定大小的“帧”,然后被压缩。虽然它们通常可以产生出色的再现,并可用于文本到音频等下游任务,但它们不会产生直观的,可直接解释的表示。在这项工作中,我们介绍了一个概念验证的音频编码器,表示音频作为一个稀疏的事件集和它们的发生时间。基本的基于物理学的假设被用来模拟攻击和物理共振的乐器演奏和表演发生的房间,希望鼓励稀疏,简约,易于解释的表示。
摘要:Most widely-used modern audio codecs, such as Ogg Vorbis and MP3, as well as more recent "neural" codecs like Meta's Encodec or the Descript Audio Codec are based on block-coding; audio is divided into overlapping, fixed-size "frames" which are then compressed. While they often yield excellent reproductions and can be used for downstream tasks such as text-to-audio, they do not produce an intuitive, directly-interpretable representation. In this work, we introduce a proof-of-concept audio encoder that represents audio as a sparse set of events and their times-of-occurrence. Rudimentary physics-based assumptions are used to model attack and the physical resonance of both the instrument being played and the room in which a performance occurs, hopefully encouraging a sparse, parsimonious, and easy-to-interpret representation.
机器翻译由腾讯交互翻译提供,仅供参考
