今日论文合集:cs.SD语音7篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and  Benchmark

标题:真实声场:视听室声学数据集和基准

链接:https://arxiv.org/abs/2403.18821

作者:Ziyang Chen,Israel D. Gebru,Christian Richardt,Anurag Kumar,William Laney,Andrew Owens,Alexander Richard

备注:Accepted to CVPR 2024. Project site: this https URL

摘要:我们提出了一个新的数据集,称为真实声场(RAF),从多个模态捕获真实的声学房间数据。该数据集包括与多视图图像配对的高质量和密集捕获的房间脉冲响应数据,以及房间中声音发射器和收听者的精确6DoF姿态跟踪数据。我们使用这个数据集来评估现有的新视图声学合成和脉冲响应生成方法,这些方法以前依赖于合成数据。在我们的评估中,我们根据多个标准彻底评估了现有的音频和视听模型,并提出了增强其在真实数据上的性能的设置。我们还进行了实验,以调查纳入视觉数据的影响(即,图像和深度)转换成神经声场模型。此外,我们还展示了一种简单的sim2real方法的有效性,其中模型使用模拟数据进行预训练,并使用稀疏的真实数据进行微调,从而显着改进了Few-Shot学习方法。RAF是第一个提供密集捕获的房间声学数据的数据集,使其成为研究音频和视听神经声场建模技术的研究人员的理想资源。演示和数据集可在我们的项目页面上获得:https://facebookresearch.github.io/real-acoustic-fields/

摘要:We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images, and precise 6DoF pose tracking data for sound emitters and listeners in the rooms. We used this dataset to evaluate existing methods for novel-view acoustic synthesis and impulse response generation which previously relied on synthetic data. In our evaluation, we thoroughly assessed existing audio and audio-visual models against multiple criteria and proposed settings to enhance their performance on real-world data. We also conducted experiments to investigate the impact of incorporating visual data (i.e., images and depth) into neural acoustic field models. Additionally, we demonstrated the effectiveness of a simple sim2real approach, where a model is pre-trained with simulated data and fine-tuned with sparse real-world data, resulting in significant improvements in the few-shot learning approach. RAF is the first dataset to provide densely captured room acoustic data, making it an ideal resource for researchers working on audio and audio-visual neural acoustic field modeling techniques. Demos and datasets are available on our project page: https://facebookresearch.github.io/real-acoustic-fields/


【2】 Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance  Accompaniment
标题:Duolando:Follower GPT与非政策强化学习的舞蹈伴奏
链接:https://arxiv.org/abs/2403.18811
作者:Li Siyao,Tianpei Gu,Zhitao Yang,Zhengyu Lin,Ziwei Liu,Henghui Ding,Lei Yang,Chen Change Loy
备注:ICLR 2024
摘要:我们介绍了一个新的任务领域内的3D舞蹈生成,称为舞蹈伴奏,这需要从一个舞蹈伙伴,“追随者”,与领舞者的动作和基本的音乐节奏同步的响应动作的生成。与现有的独舞或群舞生成任务不同,二重唱舞蹈场景需要两个参与者之间高度的互动,需要姿势和位置的微妙协调。为了支持这一任务,我们首先通过记录约117分钟的专业舞者表演来构建大规模和多样化的二重唱交互式舞蹈数据集DD 100。为了解决这个任务中固有的挑战,我们提出了一个基于GPT的模型,Duolando,它自回归预测随后的标记化运动的音乐,领导者和追随者的运动的协调信息的条件。为了进一步增强GPT在看不见的条件(音乐和领导者动作)下生成稳定结果的能力,我们设计了一种非策略强化学习策略,该策略允许模型在人类定义的奖励的指导下,从分布外采样中探索可行的轨迹。基于收集的数据集和提出的方法,我们建立了一个基准与几个精心设计的指标。
摘要:We introduce a novel task within the field of 3D dance generation, termed dance accompaniment, which necessitates the generation of responsive movements from a dance partner, the "follower", synchronized with the lead dancer's movements and the underlying musical rhythm. Unlike existing solo or group dance generation tasks, a duet dance scenario entails a heightened degree of interaction between the two participants, requiring delicate coordination in both pose and position. To support this task, we first build a large-scale and diverse duet interactive dance dataset, DD100, by recording about 117 minutes of professional dancers' performances. To address the challenges inherent in this task, we propose a GPT-based model, Duolando, which autoregressively predicts the subsequent tokenized motion conditioned on the coordinated information of the music, the leader's and the follower's movements. To further enhance the GPT's capabilities of generating stable results on unseen conditions (music and leader motions), we devise an off-policy reinforcement learning strategy that allows the model to explore viable trajectories from out-of-distribution samplings, guided by human-defined rewards. Based on the collected dataset and proposed method, we establish a benchmark with several carefully designed metrics.


【3】 Fusion approaches for emotion recognition from speech using acoustic and  text-based features
标题:基于声学和文本特征的语音情感识别融合方法
链接:https://arxiv.org/abs/2403.18635
作者:Leonardo Pepino,Pablo Riera,Luciana Ferrer,Agustin Gravano
备注:5 pages. Accepted in ICASSP 2020
摘要:在本文中,我们研究了不同的方法,从语音使用声学和基于文本的功能分类的情绪。我们建议使用BERT来获得上下文化的词嵌入,以表示语音transmitting中包含的信息,并表明这比使用手套嵌入具有更好的性能。我们还提出并比较了不同的策略来结合音频和文本模态,在IEMOCAP和MSP-PODCAST数据集上对其进行评估。我们发现,融合声学和基于文本的系统在两个数据集上都是有益的,尽管在评估的融合方法中只观察到细微的差异。最后,对于IEMOCAP,我们展示了用于定义交叉验证折叠的标准对结果的巨大影响。特别是,为该数据集创建折叠的标准方法导致对基于文本的系统的性能的高度乐观的估计,这表明以前的一些作品可能高估了结合transmittance的优势。
摘要:In this paper, we study different approaches for classifying emotions from speech using acoustic and text-based features. We propose to obtain contextualized word embeddings with BERT to represent the information contained in speech transcriptions and show that this results in better performance than using Glove embeddings. We also propose and compare different strategies to combine the audio and text modalities, evaluating them on IEMOCAP and MSP-PODCAST datasets. We find that fusing acoustic and text-based systems is beneficial on both datasets, though only subtle differences are observed across the evaluated fusion approaches. Finally, for IEMOCAP, we show the large effect that the criteria used to define the cross-validation folds have on results. In particular, the standard way of creating folds for this dataset results in a highly optimistic estimation of performance for the text-based system, suggesting that some previous works may overestimate the advantage of incorporating transcriptions.

【4】 ACES: Evaluating Automated Audio Captioning Models on the Semantics of  Sounds
标题:ACES:基于声音语义的自动音频字幕模型的评估
链接:https://arxiv.org/abs/2403.18572
作者:Gijs Wijngaard,Elia Formisano,Bruno L. Giordano,Michel Dumontier
摘要:自动音频字幕是一种多模态任务,旨在将音频内容转换为自然语言。音频字幕系统的评估通常基于应用于文本数据的定量度量。先前的研究采用来自机器翻译和图像字幕的度量来评估生成的音频字幕的质量。受听觉认知神经科学研究的启发,本文提出了一种新的评价方法--基于声音语义的音频字幕评价(ACES)。ACES考虑到人类听众如何从声音中解析语义信息,为自动音频字幕系统提供了一个新颖而全面的评估视角。ACES结合了语义相似性和语义实体标注。ACES在Clotho-Eval FENSE基准测试中的两个评估类别中优于类似的自动音频字幕指标。
摘要:Automated Audio Captioning is a multimodal task that aims to convert audio content into natural language. The assessment of audio captioning systems is typically based on quantitative metrics applied to text data. Previous studies have employed metrics derived from machine translation and image captioning to evaluate the quality of generated audio captions. Drawing inspiration from auditory cognitive neuroscience research, we introduce a novel metric approach -- Audio Captioning Evaluation on Semantics of Sound (ACES). ACES takes into account how human listeners parse semantic information from sounds, providing a novel and comprehensive evaluation perspective for automated audio captioning systems. ACES combines semantic similarities and semantic entity labeling. ACES outperforms similar automated audio captioning metrics on the Clotho-Eval FENSE benchmark in two evaluation categories.

【5】 A Diffusion-Based Generative Equalizer for Music Restoration
标题:一种基于扩散的音乐复原生成算法
链接:https://arxiv.org/abs/2403.18636
作者:Eloi Moliner,Maija Turunen,Filip Elvander,Vesa Välimäki备注:Submitted to DAFx24. Historical music restoration examples are available at: this http URL
摘要:本文提出了一种新的音频恢复方法,重点是增强低质量的音乐录音,特别是历史的。在以前的算法BABE(盲音频带宽扩展)的基础上,我们引入了BABE-2,它提出了一系列重大改进。这项研究拓宽了带宽扩展的概念,\n {生成均衡},一个新的任务,据我们所知,还没有明确解决在以前的研究。BABE-2是围绕优化算法构建的,该算法利用了扩散模型的先验知识,这些模型是使用一组精心策划的高质量音乐曲目进行训练或微调的。该算法同时执行两个关键任务:估计滤波器退化幅度响应和恢复音频的幻觉。所提出的方法是客观地评价历史钢琴录音,显示出显着的增强比以前的版本。这种方法在振兴著名声乐家恩里科·卡鲁索和内莉·梅尔巴的作品中也产生了同样令人印象深刻的结果。这项研究代表了历史音乐实际恢复的进步。
摘要:This paper presents a novel approach to audio restoration, focusing on the enhancement of low-quality music recordings, and in particular historical ones. Building upon a previous algorithm called BABE, or Blind Audio Bandwidth Extension, we introduce BABE-2, which presents a series of significant improvements. This research broadens the concept of bandwidth extension to \emph{generative equalization}, a novel task that, to the best of our knowledge, has not been explicitly addressed in previous studies. BABE-2 is built around an optimization algorithm utilizing priors from diffusion models, which are trained or fine-tuned using a curated set of high-quality music tracks. The algorithm simultaneously performs two critical tasks: estimation of the filter degradation magnitude response and hallucination of the restored audio. The proposed method is objectively evaluated on historical piano recordings, showing a marked enhancement over the prior version. The method yields similarly impressive results in rejuvenating the works of renowned vocalists Enrico Caruso and Nellie Melba. This research represents an advancement in the practical restoration of historical music.

【6】 Noise-Robust Keyword Spotting through Self-supervised Pretraining
标题:基于自监督预训练的噪声鲁棒关键词发现
链接:https://arxiv.org/abs/2403.18560
作者:Jacob Mørk,Holger Severin Bovbjerg,Gergely Kiss,Zheng-Hua Tan
摘要:语音助手现在已经广泛使用,为了激活它们,使用了关键字识别(KWS)算法。现代KWS系统主要使用监督学习方法进行训练,需要大量的标记数据才能实现良好的性能。通过自我监督学习(SSL)利用未标记的数据已被证明可以提高清洁条件下的准确性。本文探讨了SSL预训练(如Data2Vec)如何用于增强噪声条件下KWS模型的鲁棒性,这一点尚未得到充分研究。  使用不同的预训练方法对三种不同大小的模型进行预训练,然后针对KWS进行微调。然后测试这些模型,并与使用两种基线监督学习方法训练的模型进行比较,一种是使用干净数据的标准训练,另一种是多风格训练(MTR)。结果表明,在所有测试条件下,对干净数据进行预训练和微调优于对干净数据进行监督学习,并且在SNR高于5 dB的测试条件下优于监督MTR。这表明单独的预训练可以提高模型的鲁棒性。最后,我们发现,使用噪声数据进行预训练模型,特别是使用Data2Vec去噪方法,显着提高了KWS模型在噪声条件下的鲁棒性。
摘要:Voice assistants are now widely available, and to activate them a keyword spotting (KWS) algorithm is used. Modern KWS systems are mainly trained using supervised learning methods and require a large amount of labelled data to achieve a good performance. Leveraging unlabelled data through self-supervised learning (SSL) has been shown to increase the accuracy in clean conditions. This paper explores how SSL pretraining such as Data2Vec can be used to enhance the robustness of KWS models in noisy conditions, which is under-explored.  Models of three different sizes are pretrained using different pretraining approaches and then fine-tuned for KWS. These models are then tested and compared to models trained using two baseline supervised learning methods, one being standard training using clean data and the other one being multi-style training (MTR). The results show that pretraining and fine-tuning on clean data is superior to supervised learning on clean data across all testing conditions, and superior to supervised MTR for testing conditions of SNR above 5 dB. This indicates that pretraining alone can increase the model's robustness. Finally, it is found that using noisy data for pretraining models, especially with the Data2Vec-denoising approach, significantly enhances the robustness of KWS models in noisy conditions.


【7】 Dual-path Mamba: Short and Long-term Bidirectional Selective Structured  State Space Models for Speech Separation
标题:双路径Mamba:用于语音分离的短、长时双向选择性结构化状态空间模型
链接:https://arxiv.org/abs/2403.18257
作者:Xilin Jiang,Cong Han,Nima Mesgarani
摘要:Transformers已经成为用于各种语音建模任务(包括语音分离)的最成功的体系结构。然而,具有二次复杂度的Transformers中的自注意机制在计算和存储上是低效的。最近的模型结合了新的层和模块以及Transformers,以获得更好的性能,但也引入了额外的模型复杂性。在这项工作中,我们取代Transformers与曼巴,一个选择性的状态空间模型,语音分离。我们提出了双路径曼巴,它使用选择性状态空间模型的语音信号的短期和长期的前向和后向依赖。在WSJ 0 - 2 mix数据上的实验结果表明,我们的双路径Mamba模型与双路径Transformer模型Sepformer和QDPN的性能相当,前者的参数仅为前者的60%,后者的参数仅为前者的30%.我们的大型模型还达到了24.4 dB的最新SI-SNRi。
摘要:Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in computation and memory. Recent models incorporate new layers and modules along with transformers for better performance but also introduce extra model complexity. In this work, we replace transformers with Mamba, a selective state space model, for speech separation. We propose dual-path Mamba, which models short-term and long-term forward and backward dependency of speech signals using selective state spaces. Our experimental results on the WSJ0-2mix data show that our dual-path Mamba models match or outperform dual-path transformer models Sepformer with only 60% of its parameters, and the QDPN with only 30% of its parameters. Our large model also reaches a new state-of-the-art SI-SNRi of 24.4 dB.


eess.AS音频处理
【1】Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event  Detection
标题:注意领域差距:生物声事件检测的系统分析
链接:https://arxiv.org/abs/2403.18638
作者:Jinhua Liang,Ines Nolasco,Burooj Ghani,Huy Phan,Emmanouil Benetos,Dan Stowell
摘要:检测自然界中动物发声的存在对于研究动物种群及其行为至关重要。该领域的最新发展是引入了称为Few-Shot生物声学声音事件检测的任务,其目的是仅使用一小组音频样本来训练多功能动物声音检测器。以前在这一领域的努力利用不同的架构和数据增强技术,以提高模型的性能。然而,这些方法并没有完全弥合源和目标分布之间的域差距,限制了它们在现实世界中的应用场景。在这项工作中,我们引入了一个新的数据集,旨在增加的多样性和广度的类可用于Few-Shot生物声事件检测,我们以前的数据集的基础上建立。为了建立一个为DCASE 2024任务5挑战量身定制的强大基线系统,我们深入研究了一系列声学特征,并采用负硬采样作为我们的主要域适应策略。这种方法是根据挑战的指导方针选择的,该指导方针要求独立处理每个音频文件,避开了使用转导学习来确保合规性,同时旨在增强系统对域转移的适应性。实验结果表明,与普通原型网络相比,该基线系统具有更好的性能。研究结果还证实了每个域适应方法的有效性,通过消融网络内的不同组件。这突出了通过进一步减少域移位的影响来改进Few-Shot生物声学声音事件检测的潜力。
摘要:Detecting the presence of animal vocalisations in nature is essential to study animal populations and their behaviors. A recent development in the field is the introduction of the task known as few-shot bioacoustic sound event detection, which aims to train a versatile animal sound detector using only a small set of audio samples. Previous efforts in this area have utilized different architectures and data augmentation techniques to enhance model performance. However, these approaches have not fully bridged the domain gap between source and target distributions, limiting their applicability in real-world scenarios. In this work, we introduce an new dataset designed to augment the diversity and breadth of classes available for few-shot bioacoustic event detection, building on the foundations of our previous datasets. To establish a robust baseline system tailored for the DCASE 2024 Task 5 challenge, we delve into an array of acoustic features and adopt negative hard sampling as our primary domain adaptation strategy. This approach, chosen in alignment with the challenge's guidelines that necessitate the independent treatment of each audio file, sidesteps the use of transductive learning to ensure compliance while aiming to enhance the system's adaptability to domain shifts. Our experiments show that the proposed baseline system achieves a better performance compared with the vanilla prototypical network. The findings also confirm the effectiveness of each domain adaptation method by ablating different components within the networks. This highlights the potential to improve few-shot bioacoustic sound event detection by further reducing the impact of domain shift.

【2】 A Diffusion-Based Generative Equalizer for Music Restoration
标题:一种基于扩散的音乐复原生成算法
链接:https://arxiv.org/abs/2403.18636
作者:Eloi Moliner,Maija Turunen,Filip Elvander,Vesa Välimäki
备注:Submitted to DAFx24. Historical music restoration examples are available at: this http URL
摘要:本文提出了一种新的音频恢复方法,重点是增强低质量的音乐录音,特别是历史的。在以前的算法BABE(盲音频带宽扩展)的基础上,我们引入了BABE-2,它提出了一系列重大改进。这项研究拓宽了带宽扩展的概念,\n {生成均衡},一个新的任务,据我们所知,还没有明确解决在以前的研究。BABE-2是围绕优化算法构建的,该算法利用了扩散模型的先验知识,这些模型是使用一组精心策划的高质量音乐曲目进行训练或微调的。该算法同时执行两个关键任务:估计滤波器退化幅度响应和恢复音频的幻觉。所提出的方法是客观地评价历史钢琴录音,显示出显着的增强比以前的版本。这种方法在振兴著名声乐家恩里科·卡鲁索和内莉·梅尔巴的作品中也产生了同样令人印象深刻的结果。这项研究代表了历史音乐实际恢复的进步。
摘要:This paper presents a novel approach to audio restoration, focusing on the enhancement of low-quality music recordings, and in particular historical ones. Building upon a previous algorithm called BABE, or Blind Audio Bandwidth Extension, we introduce BABE-2, which presents a series of significant improvements. This research broadens the concept of bandwidth extension to \emph{generative equalization}, a novel task that, to the best of our knowledge, has not been explicitly addressed in previous studies. BABE-2 is built around an optimization algorithm utilizing priors from diffusion models, which are trained or fine-tuned using a curated set of high-quality music tracks. The algorithm simultaneously performs two critical tasks: estimation of the filter degradation magnitude response and hallucination of the restored audio. The proposed method is objectively evaluated on historical piano recordings, showing a marked enhancement over the prior version. The method yields similarly impressive results in rejuvenating the works of renowned vocalists Enrico Caruso and Nellie Melba. This research represents an advancement in the practical restoration of historical music.

【3】 Noise-Robust Keyword Spotting through Self-supervised Pretraining
标题:基于自监督预训练的噪声鲁棒关键词发现
链接:https://arxiv.org/abs/2403.18560
作者:Jacob Mørk,Holger Severin Bovbjerg,Gergely Kiss,Zheng-Hua Tan
摘要:语音助手现在已经广泛使用,为了激活它们,使用了关键字识别(KWS)算法。现代KWS系统主要使用监督学习方法进行训练,需要大量的标记数据才能实现良好的性能。通过自我监督学习(SSL)利用未标记的数据已被证明可以提高清洁条件下的准确性。本文探讨了SSL预训练(如Data2Vec)如何用于增强噪声条件下KWS模型的鲁棒性,这一点尚未得到充分研究。  使用不同的预训练方法对三种不同大小的模型进行预训练,然后针对KWS进行微调。然后测试这些模型,并与使用两种基线监督学习方法训练的模型进行比较,一种是使用干净数据的标准训练,另一种是多风格训练(MTR)。结果表明,在所有测试条件下,对干净数据进行预训练和微调优于对干净数据进行监督学习,并且在SNR高于5 dB的测试条件下优于监督MTR。这表明单独的预训练可以提高模型的鲁棒性。最后,我们发现,使用噪声数据进行预训练模型,特别是使用Data2Vec去噪方法,显着提高了KWS模型在噪声条件下的鲁棒性。
摘要:Voice assistants are now widely available, and to activate them a keyword spotting (KWS) algorithm is used. Modern KWS systems are mainly trained using supervised learning methods and require a large amount of labelled data to achieve a good performance. Leveraging unlabelled data through self-supervised learning (SSL) has been shown to increase the accuracy in clean conditions. This paper explores how SSL pretraining such as Data2Vec can be used to enhance the robustness of KWS models in noisy conditions, which is under-explored.  Models of three different sizes are pretrained using different pretraining approaches and then fine-tuned for KWS. These models are then tested and compared to models trained using two baseline supervised learning methods, one being standard training using clean data and the other one being multi-style training (MTR). The results show that pretraining and fine-tuning on clean data is superior to supervised learning on clean data across all testing conditions, and superior to supervised MTR for testing conditions of SNR above 5 dB. This indicates that pretraining alone can increase the model's robustness. Finally, it is found that using noisy data for pretraining models, especially with the Data2Vec-denoising approach, significantly enhances the robustness of KWS models in noisy conditions.


【4】 Dual-path Mamba: Short and Long-term Bidirectional Selective Structured  State Space Models for Speech Separation
标题:双路径Mamba:用于语音分离的短、长时双向选择性结构化状态空间模型
链接:https://arxiv.org/abs/2403.18257
作者:Xilin Jiang,Cong Han,Nima Mesgarani
摘要:Transformers已经成为用于各种语音建模任务(包括语音分离)的最成功的体系结构。然而,具有二次复杂度的Transformers中的自注意机制在计算和存储上是低效的。最近的模型结合了新的层和模块以及Transformers,以获得更好的性能,但也引入了额外的模型复杂性。在这项工作中,我们取代Transformers与曼巴,一个选择性的状态空间模型,语音分离。我们提出了双路径曼巴,它使用选择性状态空间模型的语音信号的短期和长期的前向和后向依赖。在WSJ 0 - 2 mix数据上的实验结果表明,我们的双路径Mamba模型与双路径Transformer模型Sepformer和QDPN的性能相当,前者的参数仅为前者的60%,后者的参数仅为前者的30%.我们的大型模型还达到了24.4 dB的最新SI-SNRi。
摘要:Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in computation and memory. Recent models incorporate new layers and modules along with transformers for better performance but also introduce extra model complexity. In this work, we replace transformers with Mamba, a selective state space model, for speech separation. We propose dual-path Mamba, which models short-term and long-term forward and backward dependency of speech signals using selective state spaces. Our experimental results on the WSJ0-2mix data show that our dual-path Mamba models match or outperform dual-path transformer models Sepformer with only 60% of its parameters, and the QDPN with only 30% of its parameters. Our large model also reaches a new state-of-the-art SI-SNRi of 24.4 dB.

【5】 Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and  Benchmark
标题:真实声场:视听室声学数据集和基准
链接:https://arxiv.org/abs/2403.18821
作者:Ziyang Chen,Israel D. Gebru,Christian Richardt,Anurag Kumar,William Laney,Andrew Owens,Alexander Richard
备注:Accepted to CVPR 2024. Project site: this https URL
摘要:我们提出了一个新的数据集,称为真实声场(RAF),从多个模态捕获真实的声学房间数据。该数据集包括与多视图图像配对的高质量和密集捕获的房间脉冲响应数据,以及房间中声音发射器和收听者的精确6DoF姿态跟踪数据。我们使用这个数据集来评估现有的新视图声学合成和脉冲响应生成方法,这些方法以前依赖于合成数据。在我们的评估中,我们根据多个标准彻底评估了现有的音频和视听模型,并提出了增强其在真实数据上的性能的设置。我们还进行了实验,以调查纳入视觉数据的影响(即,图像和深度)转换成神经声场模型。此外,我们还展示了一种简单的sim2real方法的有效性,其中模型使用模拟数据进行预训练,并使用稀疏的真实数据进行微调,从而显着改进了Few-Shot学习方法。RAF是第一个提供密集捕获的房间声学数据的数据集,使其成为研究音频和视听神经声场建模技术的研究人员的理想资源。演示和数据集可在我们的项目页面上获得:https://facebookresearch.github.io/real-acoustic-fields/
摘要:We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images, and precise 6DoF pose tracking data for sound emitters and listeners in the rooms. We used this dataset to evaluate existing methods for novel-view acoustic synthesis and impulse response generation which previously relied on synthetic data. In our evaluation, we thoroughly assessed existing audio and audio-visual models against multiple criteria and proposed settings to enhance their performance on real-world data. We also conducted experiments to investigate the impact of incorporating visual data (i.e., images and depth) into neural acoustic field models. Additionally, we demonstrated the effectiveness of a simple sim2real approach, where a model is pre-trained with simulated data and fine-tuned with sparse real-world data, resulting in significant improvements in the few-shot learning approach. RAF is the first dataset to provide densely captured room acoustic data, making it an ideal resource for researchers working on audio and audio-visual neural acoustic field modeling techniques. Demos and datasets are available on our project page: https://facebookresearch.github.io/real-acoustic-fields/


【6】 Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance  Accompaniment
标题:Duolando:Follower GPT与非政策强化学习的舞蹈伴奏
链接:https://arxiv.org/abs/2403.18811
作者:Li Siyao,Tianpei Gu,Zhitao Yang,Zhengyu Lin,Ziwei Liu,Henghui Ding,Lei Yang,Chen Change Loy
备注:ICLR 2024
摘要:我们介绍了一个新的任务领域内的3D舞蹈生成,称为舞蹈伴奏,这需要从一个舞蹈伙伴,“追随者”,与领舞者的动作和基本的音乐节奏同步的响应动作的生成。与现有的独舞或群舞生成任务不同,二重唱舞蹈场景需要两个参与者之间高度的互动,需要姿势和位置的微妙协调。为了支持这一任务,我们首先通过记录约117分钟的专业舞者表演来构建大规模和多样化的二重唱交互式舞蹈数据集DD 100。为了解决这个任务中固有的挑战,我们提出了一个基于GPT的模型,Duolando,它自回归预测随后的标记化运动的音乐,领导者和追随者的运动的协调信息的条件。为了进一步增强GPT在看不见的条件(音乐和领导者动作)下生成稳定结果的能力,我们设计了一种非策略强化学习策略,该策略允许模型在人类定义的奖励的指导下,从分布外采样中探索可行的轨迹。基于收集的数据集和提出的方法,我们建立了一个基准与几个精心设计的指标。
摘要:We introduce a novel task within the field of 3D dance generation, termed dance accompaniment, which necessitates the generation of responsive movements from a dance partner, the "follower", synchronized with the lead dancer's movements and the underlying musical rhythm. Unlike existing solo or group dance generation tasks, a duet dance scenario entails a heightened degree of interaction between the two participants, requiring delicate coordination in both pose and position. To support this task, we first build a large-scale and diverse duet interactive dance dataset, DD100, by recording about 117 minutes of professional dancers' performances. To address the challenges inherent in this task, we propose a GPT-based model, Duolando, which autoregressively predicts the subsequent tokenized motion conditioned on the coordinated information of the music, the leader's and the follower's movements. To further enhance the GPT's capabilities of generating stable results on unseen conditions (music and leader motions), we devise an off-policy reinforcement learning strategy that allows the model to explore viable trajectories from out-of-distribution samplings, guided by human-defined rewards. Based on the collected dataset and proposed method, we establish a benchmark with several carefully designed metrics.

【7】 Fusion approaches for emotion recognition from speech using acoustic and  text-based features标题:基于声学和文本特征的语音情感识别融合方法
链接:https://arxiv.org/abs/2403.18635
作者:Leonardo Pepino,Pablo Riera,Luciana Ferrer,Agustin Gravano
备注:5 pages. Accepted in ICASSP 2020
摘要:在本文中,我们研究了不同的方法,从语音使用声学和基于文本的功能分类的情绪。我们建议使用BERT来获得上下文化的词嵌入,以表示语音transmitting中包含的信息,并表明这比使用手套嵌入具有更好的性能。我们还提出并比较了不同的策略来结合音频和文本模态,在IEMOCAP和MSP-PODCAST数据集上对其进行评估。我们发现,融合声学和基于文本的系统在两个数据集上都是有益的,尽管在评估的融合方法中只观察到细微的差异。最后,对于IEMOCAP,我们展示了用于定义交叉验证折叠的标准对结果的巨大影响。特别是,为该数据集创建折叠的标准方法导致对基于文本的系统的性能的高度乐观的估计,这表明以前的一些作品可能高估了结合transmittance的优势。
摘要:In this paper, we study different approaches for classifying emotions from speech using acoustic and text-based features. We propose to obtain contextualized word embeddings with BERT to represent the information contained in speech transcriptions and show that this results in better performance than using Glove embeddings. We also propose and compare different strategies to combine the audio and text modalities, evaluating them on IEMOCAP and MSP-PODCAST datasets. We find that fusing acoustic and text-based systems is beneficial on both datasets, though only subtle differences are observed across the evaluated fusion approaches. Finally, for IEMOCAP, we show the large effect that the criteria used to define the cross-validation folds have on results. In particular, the standard way of creating folds for this dataset results in a highly optimistic estimation of performance for the text-based system, suggesting that some previous works may overestimate the advantage of incorporating transcriptions.

【8】 ACES: Evaluating Automated Audio Captioning Models on the Semantics of  Sounds
标题:ACES:基于声音语义的自动音频字幕模型的评估
链接:https://arxiv.org/abs/2403.18572
作者:Gijs Wijngaard,Elia Formisano,Bruno L. Giordano,Michel Dumontier
摘要:自动音频字幕是一种多模态任务,旨在将音频内容转换为自然语言。音频字幕系统的评估通常基于应用于文本数据的定量度量。先前的研究采用来自机器翻译和图像字幕的度量来评估生成的音频字幕的质量。受听觉认知神经科学研究的启发,本文提出了一种新的评价方法--基于声音语义的音频字幕评价(ACES)。ACES考虑到人类听众如何从声音中解析语义信息,为自动音频字幕系统提供了一个新颖而全面的评估视角。ACES结合了语义相似性和语义实体标注。ACES在Clotho-Eval FENSE基准测试中的两个评估类别中优于类似的自动音频字幕指标。
摘要:Automated Audio Captioning is a multimodal task that aims to convert audio content into natural language. The assessment of audio captioning systems is typically based on quantitative metrics applied to text data. Previous studies have employed metrics derived from machine translation and image captioning to evaluate the quality of generated audio captions. Drawing inspiration from auditory cognitive neuroscience research, we introduce a novel metric approach -- Audio Captioning Evaluation on Semantics of Sound (ACES). ACES takes into account how human listeners parse semantic information from sounds, providing a novel and comprehensive evaluation perspective for automated audio captioning systems. ACES combines semantic similarities and semantic entity labeling. ACES outperforms similar automated audio captioning metrics on the Clotho-Eval FENSE benchmark in two evaluation categories.


机器翻译由腾讯交互翻译提供,仅供参考