本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
链接:https://arxiv.org/abs/2503.19677
备注:5 pages 8 figures
摘要:本文探讨了卷积神经网络CNN通过音频文件的Mel频谱图表示对语音中的情感进行分类的应用。传统的方法,如高斯混合模型和隐马尔可夫模型,已被证明不足以用于实际部署,促使人们转向深度学习技术。通过将音频数据转换为视觉格式,CNN模型可以自主学习识别复杂的模式,从而提高分类精度。开发的模型集成到一个用户友好的图形界面,促进实时预测和教育环境中的潜在应用。该研究旨在促进对语音情感识别中深度学习的理解,评估模型的可行性,并有助于将技术整合到学习环境中
摘要:This paper explores the application of Convolutional Neural Networks CNNs for classifying emotions in speech through Mel Spectrogram representations of audio files. Traditional methods such as Gaussian Mixture Models and Hidden Markov Models have proven insufficient for practical deployment, prompting a shift towards deep learning techniques. By transforming audio data into a visual format, the CNN model autonomously learns to identify intricate patterns, enhancing classification accuracy. The developed model is integrated into a user-friendly graphical interface, facilitating realtime predictions and potential applications in educational environments. The study aims to advance the understanding of deep learning in speech emotion recognition, assess the models feasibility, and contribute to the integration of technology in learning contexts
【2】 Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
标题: 高保真音乐一代的可分析音乐思想链预算
链接:https://arxiv.org/abs/2503.19611
备注:Preprint
摘要:自回归(AR)模型在生成高保真音乐方面表现出令人印象深刻的能力。然而,AR模型中的传统下一个令牌预测范式与音乐创作中的人类创作过程不一致,可能会损害生成样本的音乐性。为了克服这一限制,我们介绍了MusiCoT,一种新颖的思想链(CoT)提示技术为音乐生成量身定制。MusiCoT使AR模型能够在生成音频令牌之前首先概述整体音乐结构,从而增强所得作品的连贯性和创造性。通过利用对比语言音频预训练(CLAP)模型,我们建立了一个“音乐思想”链,使MusiCoT可扩展且独立于人类标记的数据,与传统的CoT方法相反。此外,MusiCoT允许对音乐结构进行深入分析,例如乐器编曲,并支持音乐参考-接受可变长度的音频输入作为可选的风格参考。这种创新的方法有效地解决了复制问题,将MusiCoT定位为音乐提示的重要实用方法。我们的实验结果表明,MusiCoT在客观和主观指标上始终实现卓越的性能,产生的音乐质量可与最先进的生成模型相媲美。 我们的样品可在https://MusiCoT.github.io/上获得。
摘要:Autoregressive (AR) models have demonstrated impressive capabilities in generating high-fidelity music. However, the conventional next-token prediction paradigm in AR models does not align with the human creative process in music composition, potentially compromising the musicality of generated samples. To overcome this limitation, we introduce MusiCoT, a novel chain-of-thought (CoT) prompting technique tailored for music generation. MusiCoT empowers the AR model to first outline an overall music structure before generating audio tokens, thereby enhancing the coherence and creativity of the resulting compositions. By leveraging the contrastive language-audio pretraining (CLAP) model, we establish a chain of "musical thoughts", making MusiCoT scalable and independent of human-labeled data, in contrast to conventional CoT methods. Moreover, MusiCoT allows for in-depth analysis of music structure, such as instrumental arrangements, and supports music referencing -- accepting variable-length audio inputs as optional style references. This innovative approach effectively addresses copying issues, positioning MusiCoT as a vital practical method for music prompting. Our experimental results indicate that MusiCoT consistently achieves superior performance across both objective and subjective metrics, producing music quality that rivals state-of-the-art generation models. Our samples are available at https://MusiCoT.github.io/.
【3】 QINCODEC: Neural Audio Compression with Implicit Neural Codebooks
标题: QINCodec:使用隐式神经代码簿的神经音频压缩
链接:https://arxiv.org/abs/2503.19597
摘要:神经音频编解码器,将波形压缩成离散令牌的神经网络,在音频生成模型的最新发展中起着至关重要的作用。最先进的编解码器依赖于自动编码器的端到端训练和量化瓶颈。然而,这种方法限制了量化方法的选择,因为它需要定义梯度如何通过量化器传播以及如何在线更新量化参数。在这项工作中,我们重新审视了联合训练的常见做法,并建议离线处理预先训练的自动编码器的潜在表示,然后对解码器进行可选的微调,以减轻量化带来的退化。该策略允许考虑任何现成的量化器,特别是具有隐式神经码本的最先进的可训练量化器,例如QINCO2。我们证明,与后者,我们提出的编解码器称为QINCODEC,是有竞争力的基线编解码器,同时显着更简单的训练。最后,我们的方法提供了一个通用框架,可以分摊自动编码器预训练的成本,并实现更灵活的编解码器设计。
摘要:Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a quantization bottleneck. However, this approach restricts the choice of the quantization methods as it requires to define how gradients propagate through the quantizer and how to update the quantization parameters online. In this work, we revisit the common practice of joint training and propose to quantize the latent representations of a pre-trained autoencoder offline, followed by an optional finetuning of the decoder to mitigate degradation from quantization. This strategy allows to consider any off-the-shelf quantizer, especially state-of-the-art trainable quantizers with implicit neural codebooks such as QINCO2. We demonstrate that with the latter, our proposed codec termed QINCODEC, is competitive with baseline codecs while being notably simpler to train. Finally, our approach provides a general framework that amortizes the cost of autoencoder pretraining, and enables more flexible codec design.
【4】 Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization
标题: 通过声学表示优化提高音频对抗示例的可移植性
链接:https://arxiv.org/abs/2503.19591
备注:Accepted to ICME 2025
摘要:随着自动语音识别(ASR)系统的广泛应用,其对抗性攻击的脆弱性得到了广泛的研究。然而,大多数现有的对抗性示例都是在特定的个体模型上生成的,导致缺乏可移植性。在现实场景中,攻击者通常无法访问有关目标模型的详细信息,使得基于查询的攻击不可行。为了应对这一挑战,我们提出了一种称为声学表示优化的技术,该技术将对抗性扰动与从语音表示模型中获得的低级别声学特征对齐。我们的方法不是依赖于特定于模型的高层抽象,而是利用在不同ASR架构中保持一致的基本声学表示。通过强制声学表示损失来引导对这些鲁棒的低级别表示的扰动,我们增强了对抗性示例的跨模型可转移性,而不会降低音频质量。我们的方法是即插即用的,可以与任何现有的攻击方法集成。我们在三个现代ASR模型上评估了我们的方法,实验结果表明,我们的方法显着提高了以前方法生成的对抗性示例的可移植性,同时保持了音频质量。
摘要:With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adversarial examples are generated on specific individual models, resulting in a lack of transferability. In real-world scenarios, attackers often cannot access detailed information about the target model, making query-based attacks unfeasible. To address this challenge, we propose a technique called Acoustic Representation Optimization that aligns adversarial perturbations with low-level acoustic characteristics derived from speech representation models. Rather than relying on model-specific, higher-layer abstractions, our approach leverages fundamental acoustic representations that remain consistent across diverse ASR architectures. By enforcing an acoustic representation loss to guide perturbations toward these robust, lower-level representations, we enhance the cross-model transferability of adversarial examples without degrading audio quality. Our method is plug-and-play and can be integrated with any existing attack methods. We evaluate our approach on three modern ASR models, and the experimental results demonstrate that our method significantly improves the transferability of adversarial examples generated by previous methods while preserving the audio quality.
【5】 Unifying EEG and Speech for Emotion Recognition: A Two-Step Joint Learning Framework for Handling Missing EEG Data During Inference
标题: 统一脑电和语音进行情感识别:用于处理推理期间缺失脑电数据的两步联合学习框架
链接:https://arxiv.org/abs/2503.18964
备注:10 pages, 5 figures
摘要:计算机接口正朝着使用多模态以实现更好的人机交互的方向发展。自动情感识别(AER)的使用可以使交互自然且有意义,从而增强用户体验。虽然言语是AER最直接和直观的形式,但它并不可靠,因为它可以被人类故意伪造。另一方面,像EEG这样的生理模式更可靠,不可能伪造。然而,由于需要专门的记录设置,EEG的使用对于现实场景使用是不可行的。在本文中,我们的主要目标之一是依靠EEG模态的可靠性,以促进语音模态上的鲁棒AER。我们的方法在训练过程中使用这两种模式,即使在没有更可靠的EEG模式的情况下,也能在推理时可靠地识别情绪。我们提出了一种两步联合多模态学习方法(JMML),该方法利用模态内和模态间特征来构建情感嵌入,从而丰富AER的性能。在第一步中,使用JEC-SSL,模态内学习独立于各个模态进行。随后是使用所提出的深度规范相关交叉模态自动编码器(E-DCC-CAE)的扩展变体的模态间学习。该方法通过将两种模态映射到一个公共表示空间来学习它们的联合属性,使得模态最大程度地相关。这些情感嵌入通过增强用于AER的ML分类器的性能来保持两种模态的属性。实验结果表明了该方法的有效性。据我们所知,这是第一次尝试将语音和EEG与联合多模态学习方法相结合,以获得可靠的AER。
摘要:Computer interfaces are advancing towards using multi-modalities to enable better human-computer interactions. The use of automatic emotion recognition (AER) can make the interactions natural and meaningful thereby enhancing the user experience. Though speech is the most direct and intuitive modality for AER, it is not reliable because it can be intentionally faked by humans. On the other hand, physiological modalities like EEG, are more reliable and impossible to fake. However, use of EEG is infeasible for realistic scenarios usage because of the need for specialized recording setup. In this paper, one of our primary aims is to ride on the reliability of the EEG modality to facilitate robust AER on the speech modality. Our approach uses both the modalities during training to reliably identify emotion at the time of inference, even in the absence of the more reliable EEG modality. We propose, a two-step joint multi-modal learning approach (JMML) that exploits both the intra- and inter- modal characteristics to construct emotion embeddings that enrich the performance of AER. In the first step, using JEC-SSL, intra-modal learning is done independently on the individual modalities. This is followed by an inter-modal learning using the proposed extended variant of deep canonically correlated cross-modal autoencoder (E-DCC-CAE). The approach learns the joint properties of both the modalities by mapping them into a common representation space, such that the modalities are maximally correlated. These emotion embeddings, hold properties of both the modalities there by enhancing the performance of ML classifier used for AER. Experimental results show the efficacy of the proposed approach. To best of our knowledge, this is the first attempt to combine speech and EEG with joint multi-modal learning approach for reliable AER.
【6】 Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach
标题: 跨音频域的音调轮廓探索:基于视觉的迁移学习方法
链接:https://arxiv.org/abs/2503.19161
摘要:本研究探讨音高轮廓作为一个统一的语义结构普遍存在于各种音频领域,包括音乐,语音,生物声学,和日常的声音。分析音高轮廓可以深入了解音高在音频信号感知处理中的普遍作用,并有助于更深入地了解人类和动物的听觉机制。传统的音高跟踪方法虽然针对音乐和语音进行了优化,但在处理其他音频域中发现的更宽的频率范围和更快的音高变化方面面临挑战。本研究介绍了一种基于视觉的方法来音高轮廓分析,消除了显式音高跟踪的需要。该方法使用卷积神经网络,预先训练用于自然图像中的对象检测,并使用合成生成的音高轮廓数据集进行微调,以从短音频片段的时频表示中提取关键轮廓参数。从四个音频域的一组不同的八个下游任务被选择为跨域音高轮廓分析提供一个具有挑战性的评估方案。结果表明,该方法始终优于传统的技术,基于音高跟踪的广泛的任务。这表明基于视觉的方法为跨不同音频域的音高轮廓特征的比较研究奠定了基础。
摘要:This study examines pitch contours as a unifying semantic construct prevalent across various audio domains including music, speech, bioacoustics, and everyday sounds. Analyzing pitch contours offers insights into the universal role of pitch in the perceptual processing of audio signals and contributes to a deeper understanding of auditory mechanisms in both humans and animals. Conventional pitch-tracking methods, while optimized for music and speech, face challenges in handling much broader frequency ranges and more rapid pitch variations found in other audio domains. This study introduces a vision-based approach to pitch contour analysis that eliminates the need for explicit pitch-tracking. The approach uses a convolutional neural network, pre-trained for object detection in natural images and fine-tuned with a dataset of synthetically generated pitch contours, to extract key contour parameters from the time-frequency representation of short audio segments. A diverse set of eight downstream tasks from four audio domains were selected to provide a challenging evaluation scenario for cross-domain pitch contour analysis. The results show that the proposed method consistently surpasses traditional techniques based on pitch-tracking on a wide range of tasks. This suggests that the vision-based approach establishes a foundation for comparative studies of pitch contour characteristics across diverse audio domains.
【1】 Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach
标题: 跨音频域的音调轮廓探索:基于视觉的迁移学习方法
链接:https://arxiv.org/abs/2503.19161
摘要:本研究探讨音高轮廓作为一个统一的语义结构普遍存在于各种音频领域,包括音乐,语音,生物声学,和日常的声音。分析音高轮廓可以深入了解音高在音频信号感知处理中的普遍作用,并有助于更深入地了解人类和动物的听觉机制。传统的音高跟踪方法虽然针对音乐和语音进行了优化,但在处理其他音频域中发现的更宽的频率范围和更快的音高变化方面面临挑战。本研究介绍了一种基于视觉的方法来音高轮廓分析,消除了显式音高跟踪的需要。该方法使用卷积神经网络,预先训练用于自然图像中的对象检测,并使用合成生成的音高轮廓数据集进行微调,以从短音频片段的时频表示中提取关键轮廓参数。从四个音频域的一组不同的八个下游任务被选择为跨域音高轮廓分析提供一个具有挑战性的评估方案。结果表明,该方法始终优于传统的技术,基于音高跟踪的广泛的任务。这表明基于视觉的方法为跨不同音频域的音高轮廓特征的比较研究奠定了基础。
摘要:This study examines pitch contours as a unifying semantic construct prevalent across various audio domains including music, speech, bioacoustics, and everyday sounds. Analyzing pitch contours offers insights into the universal role of pitch in the perceptual processing of audio signals and contributes to a deeper understanding of auditory mechanisms in both humans and animals. Conventional pitch-tracking methods, while optimized for music and speech, face challenges in handling much broader frequency ranges and more rapid pitch variations found in other audio domains. This study introduces a vision-based approach to pitch contour analysis that eliminates the need for explicit pitch-tracking. The approach uses a convolutional neural network, pre-trained for object detection in natural images and fine-tuned with a dataset of synthetically generated pitch contours, to extract key contour parameters from the time-frequency representation of short audio segments. A diverse set of eight downstream tasks from four audio domains were selected to provide a challenging evaluation scenario for cross-domain pitch contour analysis. The results show that the proposed method consistently surpasses traditional techniques based on pitch-tracking on a wide range of tasks. This suggests that the vision-based approach establishes a foundation for comparative studies of pitch contour characteristics across diverse audio domains.
【2】 Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
标题: 高保真音乐一代的可分析音乐思想链预算
链接:https://arxiv.org/abs/2503.19611
备注:Preprint
摘要:None
摘要:Autoregressive (AR) models have demonstrated impressive capabilities in generating high-fidelity music. However, the conventional next-token prediction paradigm in AR models does not align with the human creative process in music composition, potentially compromising the musicality of generated samples. To overcome this limitation, we introduce MusiCoT, a novel chain-of-thought (CoT) prompting technique tailored for music generation. MusiCoT empowers the AR model to first outline an overall music structure before generating audio tokens, thereby enhancing the coherence and creativity of the resulting compositions. By leveraging the contrastive language-audio pretraining (CLAP) model, we establish a chain of "musical thoughts", making MusiCoT scalable and independent of human-labeled data, in contrast to conventional CoT methods. Moreover, MusiCoT allows for in-depth analysis of music structure, such as instrumental arrangements, and supports music referencing -- accepting variable-length audio inputs as optional style references. This innovative approach effectively addresses copying issues, positioning MusiCoT as a vital practical method for music prompting. Our experimental results indicate that MusiCoT consistently achieves superior performance across both objective and subjective metrics, producing music quality that rivals state-of-the-art generation models. Our samples are available at https://MusiCoT.github.io/.
【3】 Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization
标题: 通过声学表示优化提高音频对抗示例的可移植性
链接:https://arxiv.org/abs/2503.19591
备注:Accepted to ICME 2025
摘要:随着自动语音识别(ASR)系统的广泛应用,其对抗性攻击的脆弱性得到了广泛的研究。然而,大多数现有的对抗性示例都是在特定的个体模型上生成的,导致缺乏可移植性。在现实场景中,攻击者通常无法访问有关目标模型的详细信息,使得基于查询的攻击不可行。为了应对这一挑战,我们提出了一种称为声学表示优化的技术,该技术将对抗性扰动与从语音表示模型中获得的低级别声学特征对齐。我们的方法不是依赖于特定于模型的高层抽象,而是利用在不同ASR架构中保持一致的基本声学表示。通过强制声学表示损失来引导对这些鲁棒的低级别表示的扰动,我们增强了对抗性示例的跨模型可转移性,而不会降低音频质量。我们的方法是即插即用的,可以与任何现有的攻击方法集成。我们在三种现代ASR模型上评估了我们的方法,实验结果表明,我们的方法显着提高了以前方法生成的对抗性示例的可移植性,同时保持了音频质量。
摘要:With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adversarial examples are generated on specific individual models, resulting in a lack of transferability. In real-world scenarios, attackers often cannot access detailed information about the target model, making query-based attacks unfeasible. To address this challenge, we propose a technique called Acoustic Representation Optimization that aligns adversarial perturbations with low-level acoustic characteristics derived from speech representation models. Rather than relying on model-specific, higher-layer abstractions, our approach leverages fundamental acoustic representations that remain consistent across diverse ASR architectures. By enforcing an acoustic representation loss to guide perturbations toward these robust, lower-level representations, we enhance the cross-model transferability of adversarial examples without degrading audio quality. Our method is plug-and-play and can be integrated with any existing attack methods. We evaluate our approach on three modern ASR models, and the experimental results demonstrate that our method significantly improves the transferability of adversarial examples generated by previous methods while preserving the audio quality.
