【1】Notochord: a Flexible Probabilistic Model for Real-Time MIDI Performance标题:NotoChord:一种灵活的实时MIDI性能概率模型链接:https://arxiv.org/abs/2403.12000作者:Victor Shepardson,Jack Armitage,Thor Magnusson备注:None摘要:基于深度学习的音乐数据概率模型正在产生越来越逼真的结果,并有望进入多种创意工作流程。然而,它们在性能设置中的研究很少,在性能设置中,用户操作的结果通常应该是即时的。为了实现这样的研究,我们设计了Notochord,这是一个针对结构化事件序列的深度概率模型,并在Lakh数据集上训练了它的一个实例。我们的概率公式允许在子事件级别进行可解释的干预,这使得一个模型能够作为各种交互式音乐功能的骨干,包括可操纵的生成,协调,机器即兴创作和基于可能性的界面。Notochord可以产生复音和多音轨的声音,并以低于10毫秒的延迟响应输入。培训代码、模型检查点和交互式示例作为开源软件提供。摘要:Deep learning-based probabilistic models of musical data are producing increasingly realistic results and promise to enter creative workflows of many kinds. Yet they have been little-studied in a performance setting, where the results of user actions typically ought to feel instantaneous. To enable such study, we designed Notochord, a deep probabilistic model for sequences of structured events, and trained an instance of it on the Lakh MIDI dataset. Our probabilistic formulation allows interpretable interventions at a sub-event level, which enables one model to act as a backbone for diverse interactive musical functions including steerable generation, harmonization, machine improvisation, and likelihood-based interfaces. Notochord can generate polyphonic and multi-track MIDI, and respond to inputs with latency below ten milliseconds. Training code, model checkpoints and interactive examples are provided as open source software. 【2】 Unimodal Multi-Task Fusion for Emotional Mimicry Prediciton标题:情绪模仿预测的单峰多任务融合链接:https://arxiv.org/abs/2403.11879作者:Tobias Hallmen,Fabian Deuser,Norbert Oswald,Elisabeth André摘要:在这项研究中,我们提出了一种方法的情绪模仿强度(EMI)估计任务的范围内的第六届研讨会和比赛的情感行为分析在野外。我们的方法利用Wav 2 Vec 2.0框架,在全面的播客数据集上进行预训练,以提取广泛的音频特征,包括语言和非语言元素。我们通过一种融合技术来增强特征表示,该技术将个体特征与全局均值向量相结合,将全局上下文洞察引入我们的分析中。此外,我们结合了来自Wav 2 Vec 2.0模型的预训练的效价-唤醒-优势(VAD)模块。我们的融合采用了长短期记忆(LSTM)架构,用于音频数据的有效时间分析。仅利用所提供的音频数据,我们的方法显示出显着的改进,超过了既定的基线。摘要:In this study, we propose a methodology for the Emotional Mimicry Intensity (EMI) Estimation task within the context of the 6th Workshop and Competition on Affective Behavior Analysis in-the-wild. Our approach leverages the Wav2Vec 2.0 framework, pre-trained on a comprehensive podcast dataset, to extract a broad range of audio features encompassing both linguistic and paralinguistic elements. We enhance feature representation through a fusion technique that integrates individual features with a global mean vector, introducing global contextual insights into our analysis. Additionally, we incorporate a pre-trained valence- arousal-dominance (VAD) module from the Wav2Vec 2.0 model. Our fusion employs a Long Short-Term Memory (LSTM) architecture for efficient temporal analysis of audio data. Utilizing only the provided audio data, our approach demonstrates significant improvements over the established baseline.
【3】 Sound Event Detection and Localization with Distance Estimation标题:基于距离估计的声音事件检测与定位链接:https://arxiv.org/abs/2403.11827作者:Daniel Aleksander Krause,Archontis Politis,Annamaria Mesaros备注:This paper has been submitted for the 32nd European Signal Processing Conference EUSIPCO 2024 in Lyon摘要:声音事件检测和定位(SELD)是一个识别声音事件及其相应的到达方向(DOA)的组合任务。虽然这项任务有许多应用,并已在近年来进行了广泛的研究,它不能提供有关声源位置的完整信息。在本文中,我们克服了这个问题,扩展到声音事件检测,定位与距离估计(3D SELD)的任务。我们研究了两种方法集成的SELD核心内的距离估计-一个多任务的方法,其中的问题是由一个单独的模型输出,和一个单任务的方法,通过扩展的多ACCDOA方法,包括距离信息。我们研究了STARSS 23:Sony-TAU Realistic Spatial Soundscapes 2023的Ambisonic和双耳版本的两种方法。此外,我们的研究涉及到距离估计部分的损失函数的实验。我们的研究结果表明,它是可能的,在声音事件检测和DOA估计的性能没有任何退化进行3D SELD。摘要:Sound Event Detection and Localization (SELD) is a combined task of identifying sound events and their corresponding direction-of-arrival (DOA). While this task has numerous applications and has been extensively researched in recent years, it fails to provide full information about the sound source position. In this paper, we overcome this problem by extending the task to Sound Event Detection, Localization with Distance Estimation (3D SELD). We study two ways of integrating distance estimation within the SELD core - a multi-task approach, in which the problem is tackled by a separate model output, and a single-task approach obtained by extending the multi-ACCDOA method to include distance information. We investigate both methods for the Ambisonic and binaural versions of STARSS23: Sony-TAU Realistic Spatial Soundscapes 2023. Moreover, our study involves experiments on the loss function related to the distance estimation part. Our results show that it is possible to perform 3D SELD without any degradation of performance in sound event detection and DOA estimation. 【4】 Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt标题:基于自然语言提示的可控歌唱声合成链接:https://arxiv.org/abs/2403.11780作者:Yongqi Wang,Ruofan Hu,Rongjie Huang,Zhiqing Hong,Ruiqi Li,Wenrui Liu,Fuming You,Tao Jin,Zhou Zhao备注:Accepted by NAACL 2024 (main conference)摘要:目前的歌唱声合成方法虽然取得了很好的音质和自然度,但缺乏对歌唱风格属性的控制能力。我们提出了SVS方法,第一个SVS方法,使属性控制歌手的性别,音域和音量与自然语言。我们采用了一个模型架构的基础上的解码器只有Transformer与多尺度层次结构,并设计了一个范围旋律解耦音高表示,使文本条件的音域控制,同时保持旋律的准确性。此外,我们探索了各种实验设置,包括不同类型的文本表示,文本编码器微调,并引入语音数据,以缓解数据稀缺,旨在促进进一步的研究。实验表明,该模型具有良好的控制能力和音频质量.音频样本可在http://prompt-singer.github.io上获得。摘要:Recent singing-voice-synthesis (SVS) methods have achieved remarkable audio quality and naturalness, yet they lack the capability to control the style attributes of the synthesized singing explicitly. We propose Prompt-Singer, the first SVS method that enables attribute controlling on singer gender, vocal range and volume with natural language. We adopt a model architecture based on a decoder-only transformer with a multi-scale hierarchy, and design a range-melody decoupled pitch representation that enables text-conditioned vocal range control while keeping melodic accuracy. Furthermore, we explore various experiment settings, including different types of text representations, text encoder fine-tuning, and introducing speech data to alleviate data scarcity, aiming to facilitate further research. Experiments show that our model achieves favorable controlling ability and audio quality. Audio samples are available at http://prompt-singer.github.io .
【5】 Towards the Development of a Real-Time Deepfake Audio Detection System in Communication Platforms标题:面向通信平台的实时深伪音频检测系统的开发链接:https://arxiv.org/abs/2403.11778作者:Jonat John Mathew,Rakin Ahsan,Sae Furukawa,Jagdish Gautham Krishna Kumar,Huzaifa Pallan,Agamjeet Singh Padda,Sara Adamski,Madhu Reddiboina,Arjun Pankajakshan摘要:Deepfake音频对通信平台构成了越来越大的威胁,需要实时检测音频流的完整性。与传统的非实时方法不同,这项研究评估了在实时通信平台中使用静态deepfake音频检测模型的可行性。可执行软件的开发是为了实现跨平台兼容性,从而实现实时执行。基于Resnet和LCNN架构的两个deepfake音频检测模型使用ASVspoof 2019数据集实现,与ASVspoof 2019挑战基线相比,实现了基准性能。该研究提出了增强这些模型的策略和框架,为通信平台中的实时深度伪造音频检测铺平了道路。这项工作有助于提高音频流安全性,确保在动态实时通信场景中具有强大的检测能力。摘要:Deepfake audio poses a rising threat in communication platforms, necessitating real-time detection for audio stream integrity. Unlike traditional non-real-time approaches, this study assesses the viability of employing static deepfake audio detection models in real-time communication platforms. An executable software is developed for cross-platform compatibility, enabling real-time execution. Two deepfake audio detection models based on Resnet and LCNN architectures are implemented using the ASVspoof 2019 dataset, achieving benchmark performances compared to ASVspoof 2019 challenge baselines. The study proposes strategies and frameworks for enhancing these models, paving the way for real-time deepfake audio detection in communication platforms. This work contributes to the advancement of audio stream security, ensuring robust detection capabilities in dynamic, real-time communication scenarios.
【6】 Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation标题:用于视听情感拟态强度估计的高效特征提取与后期融合策略链接:https://arxiv.org/abs/2403.11757作者:Jun Yu,Wangyuan Zhu,Jichao Zhu摘要:在本文中,我们提出了情绪模仿强度(EMI)估计挑战的解决方案,这是第六届情感行为分析野外(ABAW)比赛的一部分。EMI估计挑战任务旨在通过从一组预定义的情感类别(即,“钦佩”、“娱乐”、“决心”、“移情痛苦”、“兴奋”和“喜悦”)。摘要:In this paper, we present the solution to the Emotional Mimicry Intensity (EMI) Estimation challenge, which is part of 6th Affective Behavior Analysis in-the-wild (ABAW) Competition.The EMI Estimation challenge task aims to evaluate the emotional intensity of seed videos by assessing them from a set of predefined emotion categories (i.e., "Admiration," "Amusement," "Determination," "Empathic Pain," "Excitement," and "Joy").
【7】 Hallucination in Perceptual Metric-Driven Speech Enhancement Networks标题:感知度量驱动语音增强网络中的幻觉链接:https://arxiv.org/abs/2403.11732作者:George Close,Thomas Hain,Stefan Goetze备注:Submitted to EUSIPCO 2024摘要:在语音增强领域,人们对创建明确旨在提高处理后音频的感知质量的神经系统一直很感兴趣。与此相呼应的是非侵入式(即没有干净的参考)语音质量预测的主题,为此,神经网络被训练成直接从失真的音频中预测人类分配的质量标签。当组合时,这些领域允许创建强大的新语音增强系统,其可以通过将预先训练的语音质量预测器的推断作为语音增强系统的唯一损失函数来利用失真音频的大型真实世界数据集。本文的目的是确定一个潜在的陷阱,这种方法,即幻觉,这是由增强系统引入的“欺骗”的语音质量预测。摘要:Within the area of speech enhancement, there is an ongoing interest in the creation of neural systems which explicitly aim to improve the perceptual quality of the processed audio. In concert with this is the topic of non-intrusive (i.e. without clean reference) speech quality prediction, for which neural networks are trained to predict human-assigned quality labels directly from distorted audio. When combined, these areas allow for the creation of powerful new speech enhancement systems which can leverage large real-world datasets of distorted audio, by taking inference of a pre-trained speech quality predictor as the sole loss function of the speech enhancement system. This paper aims to identify a potential pitfall with this approach, namely hallucinations which are introduced by the enhancement system `tricking' the speech quality predictor. 【8】 Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models标题:文本条件音乐传播模型的广义多源推理链接:https://arxiv.org/abs/2403.11706作者:Emilian Postolache,Giorgio Mariani,Luca Cosmo,Emmanouil Benetos,Emanuele Rodolà备注:Accepted at ICASSP 2024摘要:多源扩散模型(MSDM)允许作曲音乐生成任务:生成一组相干源,创建重复,并执行源分离。尽管它们的通用性,它们需要估计的联合分布的来源,需要预先分离的音乐数据,这是很少可用的,并固定的数量和类型的来源在训练时间。本文将MSDM推广到以文本嵌入为条件的任意时域扩散模型。这些模型不需要单独的数据,因为它们是在混合物上训练的,可以参数化任意数量的源,并允许丰富的语义控制。我们提出了一个推理过程,使连贯的产生的来源和假设。此外,我们适应的狄拉克分离器的MSDM进行源分离。我们使用在Slakh 2100和MTG-Jamendo上训练的扩散模型进行实验,在宽松的数据设置中展示竞争性生成和分离结果。摘要:Multi-Source Diffusion Models (MSDM) allow for compositional musical generation tasks: generating a set of coherent sources, creating accompaniments, and performing source separation. Despite their versatility, they require estimating the joint distribution over the sources, necessitating pre-separated musical data, which is rarely available, and fixing the number and type of sources at training time. This paper generalizes MSDM to arbitrary time-domain diffusion models conditioned on text embeddings. These models do not require separated data as they are trained on mixtures, can parameterize an arbitrary number of sources, and allow for rich semantic control. We propose an inference procedure enabling the coherent generation of sources and accompaniments. Additionally, we adapt the Dirac separator of MSDM to perform source separation. We experiment with diffusion models trained on Slakh2100 and MTG-Jamendo, showcasing competitive generation and separation results in a relaxed data setting.
【9】 QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation标题:QEAN:视觉舞蹈生成的四元数增强注意力网络链接:https://arxiv.org/abs/2403.11626作者:Zhizhen Zhou,Yejing Huo,Guoheng Huang,An Zeng,Xuhang Chen,Lian Huang,Zinuo Li备注:Accepted by The Visual Computer Journal摘要:音乐生成舞蹈的研究是一个新颖而富有挑战性的图像生成课题。它的目的是输入一段音乐和种子运动,然后为后续的音乐生成自然的舞蹈动作。基于变换器的方法在与人类运动和音乐相关的时间序列预测任务中面临挑战,这是由于它们在捕获非线性关系和时间方面的斗争。这可能会导致关节变形,角色偏差,浮动和响应音乐产生的舞蹈动作不一致等问题。在本文中,我们提出了一个四元数增强注意力网络(QEAN)的视觉舞蹈合成从四元数的角度来看,它包括一个旋转位置嵌入(SPE)模块和四元数旋转注意力(QRA)模块。首先,SPE以旋转的方式将位置信息嵌入到自我注意中,从而更好地学习运动序列和音频序列的特征,并提高对音乐和舞蹈之间联系的理解。其次,QRA以一系列四元数的形式表示和融合了3D运动特征和音频特征,使模型能够在舞蹈生成的复杂时间周期条件下更好地学习音乐和舞蹈的时间协调。最后,我们在数据集AIST++上进行了实验,结果表明,我们的方法在生成准确,高质量的舞蹈动作方面取得了更好,更鲁棒的性能。我们的源代码和数据集可以分别从https://github.com/MarasyZZ/QEAN和https://google.github.io/aistplusplus_dataset获得。摘要:The study of music-generated dance is a novel and challenging Image generation task. It aims to input a piece of music and seed motions, then generate natural dance movements for the subsequent music. Transformer-based methods face challenges in time series prediction tasks related to human movements and music due to their struggle in capturing the nonlinear relationship and temporal aspects. This can lead to issues like joint deformation, role deviation, floating, and inconsistencies in dance movements generated in response to the music. In this paper, we propose a Quaternion-Enhanced Attention Network (QEAN) for visual dance synthesis from a quaternion perspective, which consists of a Spin Position Embedding (SPE) module and a Quaternion Rotary Attention (QRA) module. First, SPE embeds position information into self-attention in a rotational manner, leading to better learning of features of movement sequences and audio sequences, and improved understanding of the connection between music and dance. Second, QRA represents and fuses 3D motion features and audio features in the form of a series of quaternions, enabling the model to better learn the temporal coordination of music and dance under the complex temporal cycle conditions of dance generation. Finally, we conducted experiments on the dataset AIST++, and the results show that our approach achieves better and more robust performance in generating accurate, high-quality dance movements. Our source code and dataset can be available from https://github.com/MarasyZZ/QEAN and https://google.github.io/aistplusplus_dataset respectively.
【10】 Multitask frame-level learning for few-shot sound event detection标题:基于多任务帧级学习的Few-Shot声音事件检测链接:https://arxiv.org/abs/2403.11091作者:Liang Zou,Genwei Yan,Ruoyu Wang,Jun Du,Meng Lei,Tian Gao,Xin Fang备注:6 pages, 4 figures, conference摘要:本文研究了Few-Shot声音事件检测(SED),旨在利用有限的样本对声音事件进行自动识别和分类。然而,在Few-Shot SED中的流行方法主要依赖于片段级预测,其通常提供详细的细粒度预测,特别是对于持续时间短的事件。虽然已经提出了帧级预测策略来克服这些限制,但是这些策略通常面临由背景噪声引起的预测截断的困难。为了缓解这个问题,我们引入了一个创新的多任务帧级SED框架。此外,我们还引入了TimeFilterAug,一种用于数据增强的线性时序掩模,以提高模型对不同声学环境的鲁棒性和适应性。所提出的方法实现了63.8%的F分数,确保了2023年声学场景和事件挑战赛的检测和分类的Few-Shot生物声学事件检测类别中的第一名。摘要:This paper focuses on few-shot Sound Event Detection (SED), which aims to automatically recognize and classify sound events with limited samples. However, prevailing methods methods in few-shot SED predominantly rely on segment-level predictions, which often providing detailed, fine-grained predictions, particularly for events of brief duration. Although frame-level prediction strategies have been proposed to overcome these limitations, these strategies commonly face difficulties with prediction truncation caused by background noise. To alleviate this issue, we introduces an innovative multitask frame-level SED framework. In addition, we introduce TimeFilterAug, a linear timing mask for data augmentation, to increase the model's robustness and adaptability to diverse acoustic environments. The proposed method achieves a F-score of 63.8%, securing the 1st rank in the few-shot bioacoustic event detection category of the Detection and Classification of Acoustic Scenes and Events Challenge 2023. 【11】 Audio-Visual Segmentation via Unlabeled Frame Exploitation标题:基于无标记帧挖掘的视听分割链接:https://arxiv.org/abs/2403.11074作者:Jinxiang Liu,Yikun Liu,Fei Zhang,Chen Ju,Ya Zhang,Yanfeng Wang备注:Accepted by CVPR 2024摘要:音视频分割(AVS)的目的是分割视频帧中的声音对象。虽然已经取得了很大的进展,我们的实验表明,目前的方法达到边际性能增益内使用的未标记的帧,导致未充分利用的问题。为了充分探索未标记帧用于AVS的潜力,我们基于它们的时间特性将它们明确地分为两类,即,相邻帧(NF)和远距离帧(DF)。在时间上与标记帧相邻的NF通常包含丰富的运动信息,其有助于准确定位发声对象。与NF相反,DF与标记帧具有长的时间距离,其共享具有外观变化的语义相似的对象。考虑到它们的独特特性,我们提出了一个通用的框架,有效地利用它们来解决AVS。具体来说,对于NF,我们利用运动线索作为动态指导,以提高目标定位。此外,我们利用语义线索的DF处理他们作为有效的增强标记的帧,然后使用丰富的数据多样性的自我训练的方式。大量的实验结果表明,我们的方法的通用性和优越性,释放丰富的未标记帧的力量。摘要:Audio-visual segmentation (AVS) aims to segment the sounding objects in video frames. Although great progress has been witnessed, we experimentally reveal that current methods reach marginal performance gain within the use of the unlabeled frames, leading to the underutilization issue. To fully explore the potential of the unlabeled frames for AVS, we explicitly divide them into two categories based on their temporal characteristics, i.e., neighboring frame (NF) and distant frame (DF). NFs, temporally adjacent to the labeled frame, often contain rich motion information that assists in the accurate localization of sounding objects. Contrary to NFs, DFs have long temporal distances from the labeled frame, which share semantic-similar objects with appearance variations. Considering their unique characteristics, we propose a versatile framework that effectively leverages them to tackle AVS. Specifically, for NFs, we exploit the motion cues as the dynamic guidance to improve the objectness localization. Besides, we exploit the semantic cues in DFs by treating them as valid augmentations to the labeled frames, which are then used to enrich data diversity in a self-training manner. Extensive experimental results demonstrate the versatility and superiority of our method, unleashing the power of the abundant unlabeled frames.
【12】 Energy-Based Models with Applications to Speech and Language Processing标题:基于能量的模型及其在语音和语言处理中的应用链接:https://arxiv.org/abs/2403.10961作者:Zhijian Ou备注:None摘要:基于能量的模型(EBM)是一类重要的概率模型,也称为随机场和无向图模型。EBM是未归一化的,因此与其他流行的自归一化概率模型(如隐马尔可夫模型(HRM),自回归模型,生成对抗网络(GANs)和变分自动编码器(VAE))完全不同。在过去的几年里,EBM不仅吸引了核心机器学习社区的兴趣,而且由于重大的理论和算法进展,还吸引了语音,视觉,自然语言处理(NLP)等应用领域的兴趣。语音和语言的顺序性质也提出了特殊的挑战,并且需要与处理固定维度数据不同的处理(例如,图像)。因此,本专著的目的是系统地介绍基于能量的模型,包括语音和语言处理中的算法进展和应用。首先,介绍了EBM的基础知识,包括经典模型,最近的模型参数化的神经网络,采样方法,以及各种学习方法,从经典的学习算法到最先进的。然后,在三个不同的情况下,EBM的应用,即,分别用于边际分布、条件分布和联合分布的建模。1)序列数据的EBM及其在语言建模中的应用,主要关注序列本身的边缘分布:2)给定观察序列的目标序列的条件分布的EBM,其应用于语音识别、序列标记和文本生成; 3)用于对观测序列和目标序列两者的联合分布进行建模的EBM,以及它们在半监督学习和校准自然语言理解中的应用。摘要:Energy-Based Models (EBMs) are an important class of probabilistic models, also known as random fields and undirected graphical models. EBMs are un-normalized and thus radically different from other popular self-normalized probabilistic models such as hidden Markov models (HMMs), autoregressive models, generative adversarial nets (GANs) and variational auto-encoders (VAEs). Over the past years, EBMs have attracted increasing interest not only from the core machine learning community, but also from application domains such as speech, vision, natural language processing (NLP) and so on, due to significant theoretical and algorithmic progress. The sequential nature of speech and language also presents special challenges and needs a different treatment from processing fix-dimensional data (e.g., images). Therefore, the purpose of this monograph is to present a systematic introduction to energy-based models, including both algorithmic progress and applications in speech and language processing. First, the basics of EBMs are introduced, including classic models, recent models parameterized by neural networks, sampling methods, and various learning methods from the classic learning algorithms to the most advanced ones. Then, the application of EBMs in three different scenarios is presented, i.e., for modeling marginal, conditional and joint distributions, respectively. 1) EBMs for sequential data with applications in language modeling, where the main focus is on the marginal distribution of a sequence itself; 2) EBMs for modeling conditional distributions of target sequences given observation sequences, with applications in speech recognition, sequence labeling and text generation; 3) EBMs for modeling joint distributions of both sequences of observations and targets, and their applications in semi-supervised learning and calibrated natural language understanding.
【13】 Urban Sound Propagation: a Benchmark for 1-Step Generative Modeling of Complex Physical Systems标题:城市声传播:复杂物理系统一步生成性建模的基准链接:https://arxiv.org/abs/2403.10904作者:Martin Spitznagel,Janis Keuper摘要:复杂物理系统的数据驱动建模在仿真和机器学习社区中受到越来越多的关注。由于大多数物理模拟都是基于计算密集型的微分方程系统的迭代实现,因此用学习的一步推理模型(部分)替代具有在广泛的应用领域显着加速的潜力。在这种情况下,我们提出了一个新的基准评估1步生成学习模型的速度和物理正确性。我们的Urban Sound Propagation基准是基于物理上复杂且实际相关的,但直观上易于掌握的任务,即对城市环境中声源的波的2D传播进行建模。我们提供了一个具有10万个样本的数据集,其中每个样本由从OpenStreetmap绘制的真实2D建筑物地图对,参数化声源和给定场景的模拟地面真实声音传播组成。该数据集提供了四种不同的模拟任务,其反射、衍射和源方差的复杂性不断增加。常见的生成U-Net、GAN和扩散模型的第一基线评估表明,虽然这些模型能够在简单情况下很好地对声音传播进行建模,但由高阶方程表示的子系统的近似系统地失败。有关数据集、下载说明和源代码的信息,请访问我们的匿名网站:https://www.urban-sound-data.org。摘要:Data-driven modeling of complex physical systems is receiving a growing amount of attention in the simulation and machine learning communities. Since most physical simulations are based on compute-intensive, iterative implementations of differential equation systems, a (partial) replacement with learned, 1-step inference models has the potential for significant speedups in a wide range of application areas. In this context, we present a novel benchmark for the evaluation of 1-step generative learning models in terms of speed and physical correctness. Our Urban Sound Propagation benchmark is based on the physically complex and practically relevant, yet intuitively easy to grasp task of modeling the 2d propagation of waves from a sound source in an urban environment. We provide a dataset with 100k samples, where each sample consists of pairs of real 2d building maps drawn from OpenStreetmap, a parameterized sound source, and a simulated ground truth sound propagation for the given scene. The dataset provides four different simulation tasks with increasing complexity regarding reflection, diffraction and source variance. A first baseline evaluation of common generative U-Net, GAN and Diffusion models shows, that while these models are very well capable of modeling sound propagations in simple cases, the approximation of sub-systems represented by higher order equations systematically fails. Information about the dataset, download instructions and source codes are provided on our anonymous website: https://www.urban-sound-data.org. 【14】 Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference标题:语音驱动的个性化手势合成:利用自动模糊特征推理链接:https://arxiv.org/abs/2403.10805作者:Fan Zhang,Zhaohan Wang,Xin Lyu,Siyuan Zhao,Mengjian Li,Weidong Geng,Naye Ji,Hui Du,Fuxing Gao,Hao Wu,Shunman Li备注:12 pages,摘要:语音驱动的手势生成是虚拟人创作中的一个新兴领域。然而,一个重大的挑战在于准确地确定和处理大量的输入特征(如声学,语义,情感,个性,甚至微妙的未知特征)。传统的方法,依赖于各种显式的特征输入和复杂的多模态处理,限制了手势的表现力,并限制了它们的适用性。为了应对这些挑战,我们提出了Persona-Gestro,这是一种新型的端到端生成模型,旨在仅依赖原始语音音频生成高度个性化的3D全身手势。该模型结合了模糊特征提取器和非自回归自适应层归一化(AdaLN)Transformer扩散架构。模糊特征提取器利用模糊推理策略,自动推断隐含的、连续的模糊特征。这些模糊特征,表示为一个统一的潜在特征,被送入AdaLN Transformer。AdaLN Transformer引入了一种条件机制,该机制在所有令牌上应用统一函数,从而有效地对模糊特征和手势序列之间的相关性进行建模。该模块确保高水平的手势-语音同步,同时保持自然性。最后,我们采用扩散模型来训练和推断各种手势。对Trinity、ZEGGS和BEAT数据集进行的广泛的主观和客观评估证实了我们的模型优于当前最先进的方法。Persona-Gestro提高了系统的可用性和泛化能力,为语音驱动的手势合成设定了新的基准,并拓宽了虚拟人技术的视野。补充视频和代码可访问https://zf223669.github.io/Diffmotion-v2-website/摘要:Speech-driven gesture generation is an emerging field within virtual human creation. However, a significant challenge lies in accurately determining and processing the multitude of input features (such as acoustic, semantic, emotional, personality, and even subtle unknown features). Traditional approaches, reliant on various explicit feature inputs and complex multimodal processing, constrain the expressiveness of resulting gestures and limit their applicability. To address these challenges, we present Persona-Gestor, a novel end-to-end generative model designed to generate highly personalized 3D full-body gestures solely relying on raw speech audio. The model combines a fuzzy feature extractor and a non-autoregressive Adaptive Layer Normalization (AdaLN) transformer diffusion architecture. The fuzzy feature extractor harnesses a fuzzy inference strategy that automatically infers implicit, continuous fuzzy features. These fuzzy features, represented as a unified latent feature, are fed into the AdaLN transformer. The AdaLN transformer introduces a conditional mechanism that applies a uniform function across all tokens, thereby effectively modeling the correlation between the fuzzy features and the gesture sequence. This module ensures a high level of gesture-speech synchronization while preserving naturalness. Finally, we employ the diffusion model to train and infer various gestures. Extensive subjective and objective evaluations on the Trinity, ZEGGS, and BEAT datasets confirm our model's superior performance to the current state-of-the-art approaches. Persona-Gestor improves the system's usability and generalization capabilities, setting a new benchmark in speech-driven gesture synthesis and broadening the horizon for virtual human technology. Supplementary videos and code can be accessed at https://zf223669.github.io/Diffmotion-v2-website/ 【15】 CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing标题:Coplay:声学感知的音频不可知认知尺度链接:https://arxiv.org/abs/2403.10796作者:Yin Li,Rajalakshmi Nanadakumar摘要:通过利用智能设备上的扬声器和麦克风,声学感测在包括健康监测、手势接口和成像的各种应用中表现出巨大的潜力。然而,在声学传感的持续研究和开发中,一个问题经常被忽视:当同时用于传感和其他传统应用(如播放音乐)时,同一个扬声器可能会在两者中造成干扰,使其在现实世界中不切实际。强烈的超声波感应信号与音乐混合在一起,会使扬声器的混音器过载。为了应对过载信号的这个问题,当前的解决方案是限幅(clipping)或缩小(down-scaling),这两者都影响音乐回放质量以及感测范围和准确度。为了应对这一挑战,我们提出了CoPlay,这是一种基于深度学习的优化算法,可以认知地适应感知信号。它可以1)最大化由并发音乐留下的可用带宽内的感测信号幅度,以优化感测范围和准确度,以及2)最小化可能影响音乐回放的任何相应的频率失真。在这项工作中,我们设计了一个深度学习模型,并在常见类型的感知信号(正弦波或调频连续波FMCW)上进行测试,作为各种不可知并发音乐和语音的输入。首先,我们评估了模型的性能,以显示生成的信号的质量。然后,我们在现实世界中的下游声学传感任务进行了实地研究。一项针对12名用户的研究证明,使用我们的自适应信号进行呼吸监测和手势识别可以达到与无并发音乐场景相似的准确性,而剪切或缩小显示出更差的准确性。定性研究还表明,音乐播放质量没有下降,不像传统的裁剪或缩小的方法。摘要:Acoustic sensing manifests great potential in various applications that encompass health monitoring, gesture interface and imaging by leveraging the speakers and microphones on smart devices. However, in ongoing research and development in acoustic sensing, one problem is often overlooked: the same speaker, when used concurrently for sensing and other traditional applications (like playing music), could cause interference in both making it impractical to use in the real world. The strong ultrasonic sensing signals mixed with music would overload the speaker's mixer. To confront this issue of overloaded signals, current solutions are clipping or down-scaling, both of which affect the music playback quality and also sensing range and accuracy. To address this challenge, we propose CoPlay, a deep learning based optimization algorithm to cognitively adapt the sensing signal. It can 1) maximize the sensing signal magnitude within the available bandwidth left by the concurrent music to optimize sensing range and accuracy and 2) minimize any consequential frequency distortion that can affect music playback. In this work, we design a deep learning model and test it on common types of sensing signals (sine wave or Frequency Modulated Continuous Wave FMCW) as inputs with various agnostic concurrent music and speech. First, we evaluated the model performance to show the quality of the generated signals. Then we conducted field studies of downstream acoustic sensing tasks in the real world. A study with 12 users proved that respiration monitoring and gesture recognition using our adapted signal achieve similar accuracy as no-concurrent-music scenarios, while clipping or down-scaling manifests worse accuracy. A qualitative study also manifests that the music play quality is not degraded, unlike traditional clipping or down-scaling methods.
【16】 On-Device Domain Learning for Keyword Spotting on Low-Power Extreme Edge Embedded Systems标题:用于低功耗极边缘嵌入式系统关键词定位的设备上域学习链接:https://arxiv.org/abs/2403.10549作者:Cristian Cioflan,Lukas Cavigelli,Manuele Rusci,Miguel de Prado,Luca Benini备注:5 pages, 2 tables, 2 figures. Accepted at IEEE AICAS 2024摘要:当神经网络暴露在嘈杂的环境中时,关键字定位的准确性会降低。现场适应以前看不见的噪声对于恢复精度损失至关重要,需要在设备上学习以确保适应过程完全发生在边缘设备上。在这项工作中,我们提出了一个完全在设备上的域自适应系统,实现了高达14%的准确性增益超过已经强大的关键字定位模型。我们可以在不到10 kB的内存中实现设备学习,仅使用100个标记的话语,在适应复杂的语音噪声后恢复5%的准确率。我们证明,域自适应可以在超低功耗微控制器上实现,只需14秒,在始终开启,电池供电的设备上只需806 mJ。摘要:Keyword spotting accuracy degrades when neural networks are exposed to noisy environments. On-site adaptation to previously unseen noise is crucial to recovering accuracy loss, and on-device learning is required to ensure that the adaptation process happens entirely on the edge device. In this work, we propose a fully on-device domain adaptation system achieving up to 14% accuracy gains over already-robust keyword spotting models. We enable on-device learning with less than 10 kB of memory, using only 100 labeled utterances to recover 5% accuracy after adapting to the complex speech noise. We demonstrate that domain adaptation can be achieved on ultra-low-power microcontrollers with as little as 806 mJ in only 14 s on always-on, battery-operated devices.
【17】 Fine-Grained Engine Fault Sound Event Detection Using Multimodal Signals标题:基于多模态信号的发动机故障声事件检测链接:https://arxiv.org/abs/2403.11037作者:Dennis Fedorishin,Livio Forte III,Philip Schneider,Srirangaraj Setlur,Venu Govindaraju备注:Accepted to ICASSP 2024摘要:声音事件检测(SED)是音频研究的一个活跃领域,旨在检测声音的时间发生。在本文中,我们应用SED发动机故障检测,通过引入一个多模态SED框架,检测细粒度的发动机故障的汽车发动机使用音频和加速度计记录的振动。我们首先介绍的问题,发动机故障SED的数据集收集了大量的各种车辆与专业标记的发动机故障声音事件。接下来,我们提出了一个SED模型,用于在时间上检测车辆发动机内发生的十个细粒度发动机故障,并使用大规模弱标记发动机故障数据集进一步探索预训练策略。通过多次评估,我们表明我们提出的框架能够有效地检测发动机故障声音事件。最后,我们研究了每种模态的相互作用和特征,并表明融合音频和振动的特征可以提高整体发动机故障SED的能力。摘要:Sound event detection (SED) is an active area of audio research that aims to detect the temporal occurrence of sounds. In this paper, we apply SED to engine fault detection by introducing a multimodal SED framework that detects fine-grained engine faults of automobile engines using audio and accelerometer-recorded vibration. We first introduce the problem of engine fault SED on a dataset collected from a large variety of vehicles with expertly-labeled engine fault sound events. Next, we propose a SED model to temporally detect ten fine-grained engine faults that occur within vehicle engines and further explore a pretraining strategy using a large-scale weakly-labeled engine fault dataset. Through multiple evaluations, we show our proposed framework is able to effectively detect engine fault sound events. Finally, we investigate the interaction and characteristics of each modality and show that fusing features from audio and vibration improves overall engine fault SED capabilities.
【18】 Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval标题:基于音像时间一致性的音文交叉检索知识传递研究链接:https://arxiv.org/abs/2403.10756作者:Shunsuke Tsubaki,Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Keisuke Imoto备注:Submitted to EUSIPCO2024摘要:本研究的目的是改善知识转移的音频-图像的时间一致性的音频-文本交叉检索。为了解决配对的非语音音频-文本数据的有限可用性,已经研究了用于将从大量配对的音频-图像数据获取的知识转移到共享的音频-文本表示的学习方法,这表明了如何学习音频-图像共现的重要性。音频-图像学习中的常规方法将从对应视频流中随机选择的单个图像分配给整个音频剪辑,假设它们共同出现。然而,该方法可能无法准确地捕获目标音频和图像之间的时间一致性,因为单个图像只能表示场景的快照,尽管目标音频时刻变化。为了解决这个问题,我们提出了两种用于音频和图像匹配的方法,其有效地捕获时间信息:(i)最近匹配,其中基于与音频的相似性从多个时间帧中选择图像,以及(ii)多帧匹配,其中使用多个时间帧的音频和图像对。实验结果表明,方法(i)通过选择与音频信息最接近的图像并传递学习到的知识,提高了音频文本检索的性能。相反,方法(ii)提高了音频图像检索的性能,而在音频文本检索性能没有显着改善。这些结果表明,细化音频图像的时间协议可能有助于更好地知识转移到音频文本检索。摘要:The aim of this research is to refine knowledge transfer on audio-image temporal agreement for audio-text cross retrieval. To address the limited availability of paired non-speech audio-text data, learning methods for transferring the knowledge acquired from a large amount of paired audio-image data to shared audio-text representation have been investigated, suggesting the importance of how audio-image co-occurrence is learned. Conventional approaches in audio-image learning assign a single image randomly selected from the corresponding video stream to the entire audio clip, assuming their co-occurrence. However, this method may not accurately capture the temporal agreement between the target audio and image because a single image can only represent a snapshot of a scene, though the target audio changes from moment to moment. To address this problem, we propose two methods for audio and image matching that effectively capture the temporal information: (i) Nearest Match wherein an image is selected from multiple time frames based on similarity with audio, and (ii) Multiframe Match wherein audio and image pairs of multiple time frames are used. Experimental results show that method (i) improves the audio-text retrieval performance by selecting the nearest image that aligns with the audio information and transferring the learned knowledge. Conversely, method (ii) improves the performance of audio-image retrieval while not showing significant improvements in audio-text retrieval performance. These results indicate that refining audio-image temporal agreement may contribute to better knowledge transfer to audio-text retrieval. 【19】 PTSD-MDNN : Fusion tardive de réseaux de neurones profonds multimodaux pour la détection du trouble de stress post-traumatique标题:PTSD—MDNN:创伤后应激障碍多模式神经元的融合迟缓链接:https://arxiv.org/abs/2403.10565作者:Long Nguyen-Phuoc,Renald Gaboriau,Dimitri Delacroix,Laurent Navarro备注:in French language. GRETSI 2023摘要:为了给创伤后应激障碍(PTSD)的诊断提供一种更客观、更快速的方法,本文提出了一种融合两个单峰卷积神经网络的PTSD-MDNN,它具有较低的检测错误率。通过仅将视频和音频作为输入,该模型可用于远程咨询会话的配置,患者旅程的优化或人机交互。摘要:In order to provide a more objective and quicker way to diagnose post-traumatic stress disorder (PTSD), we present PTSD-MDNN which merges two unimodal convolutional neural networks and which gives low detection error rate. By taking only videos and audios as inputs, the model could be used in the configuration of teleconsultation sessions, in the optimization of patient journeys or for human-robot interaction. 【20】 Two-sided Acoustic Metascreen for Broadband and Individual Reflection and Transmission Control标题:用于宽带和单个反射和传输控制的双面声学元屏链接:https://arxiv.org/abs/2403.10548作者:Ao Chen,Xin Zhang摘要:声波调制在各种应用中起着关键作用,包括声场重建、无线通信和粒子操纵等。然而,当前的声学超材料和超表面设计通常集中于控制反射波或透射波,通常忽略声波的振幅和相位之间的耦合。为了填补这一空白,我们提出并通过实验验证了一种设计,该设计能够在4 kHz至8 kHz的频率范围内单独完全控制反射和透射声波,允许以宽带方式任意组合反射和透射声音的振幅和相位。此外,我们通过实现声学扩散,反射,聚焦和在三个不同频率下生成双面3D全息图来证明我们的方法对声音操作的重要性。这些发现为广泛工程声波开辟了另一条途径,在声学和相关领域有着广阔的应用前景。摘要:Acoustic wave modulation plays a pivotal role in various applications, including sound-field reconstruction, wireless communication, and particle manipulation, among others. However, current acoustic metamaterial and metasurface designs typically focus on controlling either reflection or transmission waves, often overlooking the coupling between amplitude and phase of acoustic waves. To fulfill this gap, we propose and experimentally validate a design enabling complete control of reflected and transmitted acoustic waves individually across a frequency range of 4 kHz to 8 kHz, allowing arbitrary combinations of amplitude and phase for reflected and transmitted sound in a broadband manner. Additionally, we demonstrate the significance of our approach for sound manipulation by achieving acoustic diffusion, reflection, focusing, and generating a two-sided 3D hologram at three distinct frequencies. These findings open an alternative avenue for extensively engineering sound waves, promising applications in acoustics and related fields.
eess.AS音频处理【1】 AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition标题:AdaMER—CTC:自适应最大熵正则化的连接时间分类算法链接:https://arxiv.org/abs/2403.11578作者:SooHwan Eom,Eunseop Yoon,Hee Suk Yoon,Chanwoo Kim,Mark Hasegawa-Johnson,Chang D. Yoo摘要:在自动语音识别(ASR)系统中,一个反复出现的障碍是生成的狭隘集中的输出分布。这种现象出现的一个副作用的连接时间分类(CTC),一个强大的序列学习工具,利用动态规划序列映射。虽然早期的努力试图将CTC损失与熵最大化正则化项相结合来缓解这个问题,但他们在训练期间对正则化采用了恒定的权重项,我们发现这可能不是最佳的。在这项工作中,我们引入了自适应最大熵正则化(AdaMER),这是一种可以在整个训练过程中调节熵正则化影响的技术。这种方法不仅改进了ASR模型训练,而且确保随着训练的进行,预测显示出所需的模型置信度。摘要:In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapping. While earlier efforts have tried to combine the CTC loss with an entropy maximization regularization term to mitigate this issue, they employed a constant weighting term on the regularization during the training, which we find may not be optimal. In this work, we introduce Adaptive Maximum Entropy Regularization (AdaMER), a technique that can modulate the impact of entropy regularization throughout the training process. This approach not only refines ASR model training but ensures that as training proceeds, predictions display the desired model confidence. 【2】 Discriminative Neighborhood Smoothing for Generative Anomalous Sound Detection标题:用于产生性异常声音检测的判别邻域平滑链接:https://arxiv.org/abs/2403.11508作者:Takuya Fujimura,Keisuke Imoto,Tomoki Toda备注:Submitted to EUSIPCO 2024摘要:我们提出了判别式邻域平滑的生成异常分数异常声音检测。虽然判别式方法通常比生成式方法具有更好的性能,但我们发现,由于训练数据和测试数据之间的差异,它有时会导致显着的性能下降,使其不如生成式方法鲁棒。我们提出的方法旨在通过结合生成和判别方法来弥补它们的缺点。生成的异常分数使用具有相似的判别特征的多个样本进行平滑,以在保持其鲁棒性的同时以集成方式提高生成方法的性能。实验结果表明,我们提出的方法大大提高了原始的生成方法,包括绝对提高了22%的AUC和鲁棒的作品,而一个判别式的方法遭受的差异。摘要:We propose discriminative neighborhood smoothing of generative anomaly scores for anomalous sound detection. While the discriminative approach is known to achieve better performance than generative approaches often, we have found that it sometimes causes significant performance degradation due to the discrepancy between the training and test data, making it less robust than the generative approach. Our proposed method aims to compensate for the disadvantages of generative and discriminative approaches by combining them. Generative anomaly scores are smoothed using multiple samples with similar discriminative features to improve the performance of the generative approach in an ensemble manner while keeping its robustness. Experimental results show that our proposed method greatly improves the original generative method, including absolute improvement of 22% in AUC and robustly works, while a discriminative method suffers from the discrepancy. 【3】 Fine-Grained Engine Fault Sound Event Detection Using Multimodal Signals标题:基于多模式信号的细粒度发动机故障声事件检测链接:https://arxiv.org/abs/2403.11037作者:Dennis Fedorishin,Livio Forte III,Philip Schneider,Srirangaraj Setlur,Venu Govindaraju备注:Accepted to ICASSP 2024摘要:声音事件检测(SED)是音频研究的一个活跃领域,旨在检测声音的时间发生。在本文中,我们应用SED发动机故障检测,通过引入一个多模态SED框架,检测细粒度的发动机故障的汽车发动机使用音频和加速度计记录的振动。我们首先介绍的问题,发动机故障SED的数据集收集了大量的各种车辆与专业标记的发动机故障声音事件。接下来,我们提出了一个SED模型,用于在时间上检测车辆发动机内发生的十个细粒度发动机故障,并使用大规模弱标记发动机故障数据集进一步探索预训练策略。通过多次评估,我们表明我们提出的框架能够有效地检测发动机故障声音事件。最后,我们研究了每种模态的相互作用和特征,并表明融合音频和振动的特征可以提高整体发动机故障SED的能力。摘要:Sound event detection (SED) is an active area of audio research that aims to detect the temporal occurrence of sounds. In this paper, we apply SED to engine fault detection by introducing a multimodal SED framework that detects fine-grained engine faults of automobile engines using audio and accelerometer-recorded vibration. We first introduce the problem of engine fault SED on a dataset collected from a large variety of vehicles with expertly-labeled engine fault sound events. Next, we propose a SED model to temporally detect ten fine-grained engine faults that occur within vehicle engines and further explore a pretraining strategy using a large-scale weakly-labeled engine fault dataset. Through multiple evaluations, we show our proposed framework is able to effectively detect engine fault sound events. Finally, we investigate the interaction and characteristics of each modality and show that fusing features from audio and vibration improves overall engine fault SED capabilities. 【4】 Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR标题:低资源ASR中改进格型搜索的最小增广语言模型初始译码链接:https://arxiv.org/abs/2403.10937作者:Savitha Murthy,Dinkar Sitaram备注:14 pages, 7 figures, Accepted in Sadhana Journal摘要:本文讨论了在低资源语言中,基线语言模型不足以生成包容性格的情况下,使用格重评分来提高语音识别准确率的问题。我们最低限度地增加基线语言模型与单词单字计数,存在于一个更大的文本语料库的目标语言,但在基线中不存在。用这种增强的基线语言模型解码后生成的格更全面。我们得到21.8%(泰卢固语)和41.8%(卡纳达语)的相对字错误减少与我们提出的方法。这种减少字错误率是可比的21.5%(泰卢固语)和45.9%(卡纳达语)通过解码与完整的维基百科文本增强语言模式获得的相对字错误减少,而我们的方法只消耗1/8的内存。我们证明了我们的方法与各种基于文本选择的语言模型增强是相当的,并且对于不同大小的数据集也是一致的。我们的方法适用于训练语音识别系统在低资源条件下,语音数据和计算资源不足,而有一个大的文本语料库,可在目标语言。我们的研究涉及解决基线的词汇表外单词的问题,而不是专注于解决命名实体的缺失。我们提出的方法是简单的,但计算成本较低。摘要:This paper addresses the problem of improving speech recognition accuracy with lattice rescoring in low-resource languages where the baseline language model is insufficient for generating inclusive lattices. We minimally augment the baseline language model with word unigram counts that are present in a larger text corpus of the target language but absent in the baseline. The lattices generated after decoding with such an augmented baseline language model are more comprehensive. We obtain 21.8% (Telugu) and 41.8% (Kannada) relative word error reduction with our proposed method. This reduction in word error rate is comparable to 21.5% (Telugu) and 45.9% (Kannada) relative word error reduction obtained by decoding with full Wikipedia text augmented language mode while our approach consumes only 1/8th the memory. We demonstrate that our method is comparable with various text selection-based language model augmentation and also consistent for data sets of different sizes. Our approach is applicable for training speech recognition systems under low resource conditions where speech data and compute resources are insufficient, while there is a large text corpus that is available in the target language. Our research involves addressing the issue of out-of-vocabulary words of the baseline in general and does not focus on resolving the absence of named entities. Our proposed method is simple and yet computationally less expensive.
【5】 Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval标题:音—文本交叉检索中音—图像时间一致性的知识传递精细化链接:https://arxiv.org/abs/2403.10756作者:Shunsuke Tsubaki,Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Keisuke Imoto备注:Submitted to EUSIPCO2024摘要:本研究的目的是改善知识转移的音频-图像的时间一致性的音频-文本交叉检索。为了解决配对的非语音音频-文本数据的有限可用性,已经研究了用于将从大量配对的音频-图像数据获取的知识转移到共享的音频-文本表示的学习方法,这表明了如何学习音频-图像共现的重要性。音频-图像学习中的常规方法将从对应视频流中随机选择的单个图像分配给整个音频剪辑,假设它们共同出现。然而,该方法可能无法准确地捕获目标音频和图像之间的时间一致性,因为单个图像只能表示场景的快照,尽管目标音频时刻变化。为了解决这个问题,我们提出了两种用于音频和图像匹配的方法,其有效地捕获时间信息:(i)最近匹配,其中基于与音频的相似性从多个时间帧中选择图像,以及(ii)多帧匹配,其中使用多个时间帧的音频和图像对。实验结果表明,方法(i)通过选择与音频信息最接近的图像并传递学习到的知识,提高了音频文本检索的性能。相反,方法(ii)提高了音频图像检索的性能,而在音频文本检索性能没有显着改善。这些结果表明,细化音频图像的时间协议可能有助于更好地知识转移到音频文本检索。摘要:The aim of this research is to refine knowledge transfer on audio-image temporal agreement for audio-text cross retrieval. To address the limited availability of paired non-speech audio-text data, learning methods for transferring the knowledge acquired from a large amount of paired audio-image data to shared audio-text representation have been investigated, suggesting the importance of how audio-image co-occurrence is learned. Conventional approaches in audio-image learning assign a single image randomly selected from the corresponding video stream to the entire audio clip, assuming their co-occurrence. However, this method may not accurately capture the temporal agreement between the target audio and image because a single image can only represent a snapshot of a scene, though the target audio changes from moment to moment. To address this problem, we propose two methods for audio and image matching that effectively capture the temporal information: (i) Nearest Match wherein an image is selected from multiple time frames based on similarity with audio, and (ii) Multiframe Match wherein audio and image pairs of multiple time frames are used. Experimental results show that method (i) improves the audio-text retrieval performance by selecting the nearest image that aligns with the audio information and transferring the learned knowledge. Conversely, method (ii) improves the performance of audio-image retrieval while not showing significant improvements in audio-text retrieval performance. These results indicate that refining audio-image temporal agreement may contribute to better knowledge transfer to audio-text retrieval.
【6】 PTSD-MDNN : Fusion tardive de réseaux de neurones profonds multimodaux pour la détection du trouble de stress post-traumatique标题:PTSD—MDNN:创伤后应激障碍多模式神经元的融合迟缓链接:https://arxiv.org/abs/2403.10565作者:Long Nguyen-Phuoc,Renald Gaboriau,Dimitri Delacroix,Laurent Navarro备注:in French language. GRETSI 2023摘要:为了给创伤后应激障碍(PTSD)的诊断提供一种更客观、更快速的方法,本文提出了一种融合两个单峰卷积神经网络的PTSD-MDNN,它具有较低的检测错误率。通过仅将视频和音频作为输入,该模型可用于远程咨询会话的配置,患者旅程的优化或人机交互。摘要:In order to provide a more objective and quicker way to diagnose post-traumatic stress disorder (PTSD), we present PTSD-MDNN which merges two unimodal convolutional neural networks and which gives low detection error rate. By taking only videos and audios as inputs, the model could be used in the configuration of teleconsultation sessions, in the optimization of patient journeys or for human-robot interaction.
【7】 Two-sided Acoustic Metascreen for Broadband and Individual Reflection and Transmission Control标题:用于宽带和单个反射和传输控制的双面声学元屏链接:https://arxiv.org/abs/2403.10548作者:Ao Chen,Xin Zhang摘要:声波调制在各种应用中起着关键作用,包括声场重建、无线通信和粒子操纵等。然而,当前的声学超材料和超表面设计通常集中于控制反射波或透射波,通常忽略声波的振幅和相位之间的耦合。为了填补这一空白,我们提出并通过实验验证了一种设计,该设计能够在4 kHz至8 kHz的频率范围内单独完全控制反射和透射声波,允许以宽带方式任意组合反射和透射声音的振幅和相位。此外,我们通过实现声学扩散,反射,聚焦和在三个不同频率下生成双面3D全息图来证明我们的方法对声音操作的重要性。这些发现为广泛工程声波开辟了另一条途径,在声学和相关领域有着广阔的应用前景。摘要:Acoustic wave modulation plays a pivotal role in various applications, including sound-field reconstruction, wireless communication, and particle manipulation, among others. However, current acoustic metamaterial and metasurface designs typically focus on controlling either reflection or transmission waves, often overlooking the coupling between amplitude and phase of acoustic waves. To fulfill this gap, we propose and experimentally validate a design enabling complete control of reflected and transmitted acoustic waves individually across a frequency range of 4 kHz to 8 kHz, allowing arbitrary combinations of amplitude and phase for reflected and transmitted sound in a broadband manner. Additionally, we demonstrate the significance of our approach for sound manipulation by achieving acoustic diffusion, reflection, focusing, and generating a two-sided 3D hologram at three distinct frequencies. These findings open an alternative avenue for extensively engineering sound waves, promising applications in acoustics and related fields.
【8】 Notochord: a Flexible Probabilistic Model for Real-Time MIDI Performance标题:Notochord:一种实时调度性能的灵活概率模型链接:https://arxiv.org/abs/2403.12000作者:Victor Shepardson,Jack Armitage,Thor Magnusson备注:None摘要:基于深度学习的音乐数据概率模型正在产生越来越逼真的结果,并有望进入多种创意工作流程。然而,它们在性能设置中的研究很少,在性能设置中,用户操作的结果通常应该是即时的。为了实现这样的研究,我们设计了Notochord,这是一个针对结构化事件序列的深度概率模型,并在Lakh数据集上训练了它的一个实例。我们的概率公式允许在子事件级别进行可解释的干预,这使得一个模型能够作为各种交互式音乐功能的骨干,包括可操纵的生成,协调,机器即兴创作和基于可能性的界面。Notochord可以产生复音和多音轨的声音,并以低于10毫秒的延迟响应输入。培训代码、模型检查点和交互式示例作为开源软件提供。摘要:Deep learning-based probabilistic models of musical data are producing increasingly realistic results and promise to enter creative workflows of many kinds. Yet they have been little-studied in a performance setting, where the results of user actions typically ought to feel instantaneous. To enable such study, we designed Notochord, a deep probabilistic model for sequences of structured events, and trained an instance of it on the Lakh MIDI dataset. Our probabilistic formulation allows interpretable interventions at a sub-event level, which enables one model to act as a backbone for diverse interactive musical functions including steerable generation, harmonization, machine improvisation, and likelihood-based interfaces. Notochord can generate polyphonic and multi-track MIDI, and respond to inputs with latency below ten milliseconds. Training code, model checkpoints and interactive examples are provided as open source software. 【9】 Unimodal Multi-Task Fusion for Emotional Mimicry Prediciton标题:情绪拟态预测的单峰多任务融合链接:https://arxiv.org/abs/2403.11879作者:Tobias Hallmen,Fabian Deuser,Norbert Oswald,Elisabeth André摘要:在这项研究中,我们提出了一种方法的情绪模仿强度(EMI)估计任务的范围内的第六届研讨会和比赛的情感行为分析在野外。我们的方法利用Wav 2 Vec 2.0框架,在全面的播客数据集上进行预训练,以提取广泛的音频特征,包括语言和非语言元素。我们通过一种融合技术来增强特征表示,该技术将个体特征与全局均值向量相结合,将全局上下文洞察引入我们的分析中。此外,我们结合了来自Wav 2 Vec 2.0模型的预训练的效价-唤醒-优势(VAD)模块。我们的融合采用了长短期记忆(LSTM)架构,用于音频数据的有效时间分析。仅利用所提供的音频数据,我们的方法显示出显着的改进,超过了既定的基线。摘要:In this study, we propose a methodology for the Emotional Mimicry Intensity (EMI) Estimation task within the context of the 6th Workshop and Competition on Affective Behavior Analysis in-the-wild. Our approach leverages the Wav2Vec 2.0 framework, pre-trained on a comprehensive podcast dataset, to extract a broad range of audio features encompassing both linguistic and paralinguistic elements. We enhance feature representation through a fusion technique that integrates individual features with a global mean vector, introducing global contextual insights into our analysis. Additionally, we incorporate a pre-trained valence- arousal-dominance (VAD) module from the Wav2Vec 2.0 model. Our fusion employs a Long Short-Term Memory (LSTM) architecture for efficient temporal analysis of audio data. Utilizing only the provided audio data, our approach demonstrates significant improvements over the established baseline.
【10】 Sound Event Detection and Localization with Distance Estimation标题:基于距离估计的声音事件检测与定位链接:https://arxiv.org/abs/2403.11827作者:Daniel Aleksander Krause,Archontis Politis,Annamaria Mesaros备注:This paper has been submitted for the 32nd European Signal Processing Conference EUSIPCO 2024 in Lyon摘要:声音事件检测和定位(SELD)是一个识别声音事件及其相应的到达方向(DOA)的组合任务。虽然这项任务有许多应用,并已在近年来进行了广泛的研究,它不能提供有关声源位置的完整信息。在本文中,我们克服了这个问题,扩展到声音事件检测,定位与距离估计(3D SELD)的任务。我们研究了两种方法集成的SELD核心内的距离估计-一个多任务的方法,其中的问题是由一个单独的模型输出,和一个单任务的方法,通过扩展的多ACCDOA方法,包括距离信息。我们研究了STARSS 23:Sony-TAU Realistic Spatial Soundscapes 2023的Ambisonic和双耳版本的两种方法。此外,我们的研究涉及到距离估计部分的损失函数的实验。我们的研究结果表明,它是可能的,在声音事件检测和DOA估计的性能没有任何退化进行3D SELD。摘要:Sound Event Detection and Localization (SELD) is a combined task of identifying sound events and their corresponding direction-of-arrival (DOA). While this task has numerous applications and has been extensively researched in recent years, it fails to provide full information about the sound source position. In this paper, we overcome this problem by extending the task to Sound Event Detection, Localization with Distance Estimation (3D SELD). We study two ways of integrating distance estimation within the SELD core - a multi-task approach, in which the problem is tackled by a separate model output, and a single-task approach obtained by extending the multi-ACCDOA method to include distance information. We investigate both methods for the Ambisonic and binaural versions of STARSS23: Sony-TAU Realistic Spatial Soundscapes 2023. Moreover, our study involves experiments on the loss function related to the distance estimation part. Our results show that it is possible to perform 3D SELD without any degradation of performance in sound event detection and DOA estimation. 【11】 Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt标题:提示-歌手:自然语言提示的可控演唱-声音-合成链接:https://arxiv.org/abs/2403.11780作者:Yongqi Wang,Ruofan Hu,Rongjie Huang,Zhiqing Hong,Ruiqi Li,Wenrui Liu,Fuming You,Tao Jin,Zhou Zhao备注:Accepted by NAACL 2024 (main conference)摘要:目前的歌唱声合成方法虽然取得了很好的音质和自然度,但缺乏对歌唱风格属性的控制能力。我们提出了SVS方法,第一个SVS方法,使属性控制歌手的性别,音域和音量与自然语言。我们采用了一个模型架构的基础上的解码器只有Transformer与多尺度层次结构,并设计了一个范围旋律解耦音高表示,使文本条件的音域控制,同时保持旋律的准确性。此外,我们探索了各种实验设置,包括不同类型的文本表示,文本编码器微调,并引入语音数据,以缓解数据稀缺,旨在促进进一步的研究。实验表明,该模型具有良好的控制能力和音频质量.音频样本可在http://prompt-singer.github.io上获得。摘要:Recent singing-voice-synthesis (SVS) methods have achieved remarkable audio quality and naturalness, yet they lack the capability to control the style attributes of the synthesized singing explicitly. We propose Prompt-Singer, the first SVS method that enables attribute controlling on singer gender, vocal range and volume with natural language. We adopt a model architecture based on a decoder-only transformer with a multi-scale hierarchy, and design a range-melody decoupled pitch representation that enables text-conditioned vocal range control while keeping melodic accuracy. Furthermore, we explore various experiment settings, including different types of text representations, text encoder fine-tuning, and introducing speech data to alleviate data scarcity, aiming to facilitate further research. Experiments show that our model achieves favorable controlling ability and audio quality. Audio samples are available at http://prompt-singer.github.io .
【12】 Towards the Development of a Real-Time Deepfake Audio Detection System in Communication Platforms标题:基于通信平台的实时深度假音频检测系统的开发链接:https://arxiv.org/abs/2403.11778作者:Jonat John Mathew,Rakin Ahsan,Sae Furukawa,Jagdish Gautham Krishna Kumar,Huzaifa Pallan,Agamjeet Singh Padda,Sara Adamski,Madhu Reddiboina,Arjun Pankajakshan摘要:Deepfake音频对通信平台构成了越来越大的威胁,需要实时检测音频流的完整性。与传统的非实时方法不同,这项研究评估了在实时通信平台中使用静态deepfake音频检测模型的可行性。可执行软件的开发是为了实现跨平台兼容性,从而实现实时执行。基于Resnet和LCNN架构的两个deepfake音频检测模型使用ASVspoof 2019数据集实现,与ASVspoof 2019挑战基线相比,实现了基准性能。该研究提出了增强这些模型的策略和框架,为通信平台中的实时深度伪造音频检测铺平了道路。这项工作有助于提高音频流安全性,确保在动态实时通信场景中具有强大的检测能力。摘要:Deepfake audio poses a rising threat in communication platforms, necessitating real-time detection for audio stream integrity. Unlike traditional non-real-time approaches, this study assesses the viability of employing static deepfake audio detection models in real-time communication platforms. An executable software is developed for cross-platform compatibility, enabling real-time execution. Two deepfake audio detection models based on Resnet and LCNN architectures are implemented using the ASVspoof 2019 dataset, achieving benchmark performances compared to ASVspoof 2019 challenge baselines. The study proposes strategies and frameworks for enhancing these models, paving the way for real-time deepfake audio detection in communication platforms. This work contributes to the advancement of audio stream security, ensuring robust detection capabilities in dynamic, real-time communication scenarios.
【13】 Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation标题:用于视听情感拟态强度估计的高效特征提取与后期融合策略链接:https://arxiv.org/abs/2403.11757作者:Jun Yu,Wangyuan Zhu,Jichao Zhu摘要:在本文中,我们提出了情绪模仿强度(EMI)估计挑战的解决方案,这是第六届情感行为分析野外(ABAW)比赛的一部分。EMI估计挑战任务旨在通过从一组预定义的情感类别(即,“钦佩”、“娱乐”、“决心”、“移情痛苦”、“兴奋”和“喜悦”)。摘要:In this paper, we present the solution to the Emotional Mimicry Intensity (EMI) Estimation challenge, which is part of 6th Affective Behavior Analysis in-the-wild (ABAW) Competition.The EMI Estimation challenge task aims to evaluate the emotional intensity of seed videos by assessing them from a set of predefined emotion categories (i.e., "Admiration," "Amusement," "Determination," "Empathic Pain," "Excitement," and "Joy").
【14】 Hallucination in Perceptual Metric-Driven Speech Enhancement Networks标题:感知度量驱动的语音增强网络中的幻觉链接:https://arxiv.org/abs/2403.11732作者:George Close,Thomas Hain,Stefan Goetze备注:Submitted to EUSIPCO 2024摘要:在语音增强领域,人们对创建明确旨在提高处理后音频的感知质量的神经系统一直很感兴趣。与此相呼应的是非侵入式(即没有干净的参考)语音质量预测的主题,为此,神经网络被训练成直接从失真的音频中预测人类分配的质量标签。当组合时,这些领域允许创建强大的新语音增强系统,其可以通过将预先训练的语音质量预测器的推断作为语音增强系统的唯一损失函数来利用失真音频的大型真实世界数据集。本文的目的是确定一个潜在的陷阱,这种方法,即幻觉,这是由增强系统引入的“欺骗”的语音质量预测。摘要:Within the area of speech enhancement, there is an ongoing interest in the creation of neural systems which explicitly aim to improve the perceptual quality of the processed audio. In concert with this is the topic of non-intrusive (i.e. without clean reference) speech quality prediction, for which neural networks are trained to predict human-assigned quality labels directly from distorted audio. When combined, these areas allow for the creation of powerful new speech enhancement systems which can leverage large real-world datasets of distorted audio, by taking inference of a pre-trained speech quality predictor as the sole loss function of the speech enhancement system. This paper aims to identify a potential pitfall with this approach, namely hallucinations which are introduced by the enhancement system `tricking' the speech quality predictor.
【15】 Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models标题:文本条件音乐传播模型的广义多源推理链接:https://arxiv.org/abs/2403.11706作者:Emilian Postolache,Giorgio Mariani,Luca Cosmo,Emmanouil Benetos,Emanuele Rodolà备注:Accepted at ICASSP 2024摘要:多源扩散模型(MSDM)允许作曲音乐生成任务:生成一组相干源,创建重复,并执行源分离。尽管它们的通用性,它们需要估计的联合分布的来源,需要预先分离的音乐数据,这是很少可用的,并固定的数量和类型的来源在训练时间。本文将MSDM推广到以文本嵌入为条件的任意时域扩散模型。这些模型不需要单独的数据,因为它们是在混合物上训练的,可以参数化任意数量的源,并允许丰富的语义控制。我们提出了一个推理过程,使连贯的产生的来源和假设。此外,我们适应的狄拉克分离器的MSDM进行源分离。我们使用在Slakh 2100和MTG-Jamendo上训练的扩散模型进行实验,在宽松的数据设置中展示竞争性生成和分离结果。摘要:Multi-Source Diffusion Models (MSDM) allow for compositional musical generation tasks: generating a set of coherent sources, creating accompaniments, and performing source separation. Despite their versatility, they require estimating the joint distribution over the sources, necessitating pre-separated musical data, which is rarely available, and fixing the number and type of sources at training time. This paper generalizes MSDM to arbitrary time-domain diffusion models conditioned on text embeddings. These models do not require separated data as they are trained on mixtures, can parameterize an arbitrary number of sources, and allow for rich semantic control. We propose an inference procedure enabling the coherent generation of sources and accompaniments. Additionally, we adapt the Dirac separator of MSDM to perform source separation. We experiment with diffusion models trained on Slakh2100 and MTG-Jamendo, showcasing competitive generation and separation results in a relaxed data setting.
【16】 QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation标题:QEAN:四元数增强的视觉舞蹈生成注意网络链接:https://arxiv.org/abs/2403.11626作者:Zhizhen Zhou,Yejing Huo,Guoheng Huang,An Zeng,Xuhang Chen,Lian Huang,Zinuo Li备注:Accepted by The Visual Computer Journal摘要:音乐生成舞蹈的研究是一个新颖而富有挑战性的图像生成课题。它的目的是输入一段音乐和种子运动,然后为后续的音乐生成自然的舞蹈动作。基于变换器的方法在与人类运动和音乐相关的时间序列预测任务中面临挑战,这是由于它们在捕获非线性关系和时间方面的斗争。这可能会导致关节变形,角色偏差,浮动和响应音乐产生的舞蹈动作不一致等问题。在本文中,我们提出了一个四元数增强注意力网络(QEAN)的视觉舞蹈合成从四元数的角度来看,它包括一个旋转位置嵌入(SPE)模块和四元数旋转注意力(QRA)模块。首先,SPE以旋转的方式将位置信息嵌入到自我注意中,从而更好地学习运动序列和音频序列的特征,并提高对音乐和舞蹈之间联系的理解。其次,QRA以一系列四元数的形式表示和融合了3D运动特征和音频特征,使模型能够在舞蹈生成的复杂时间周期条件下更好地学习音乐和舞蹈的时间协调。最后,我们在数据集AIST++上进行了实验,结果表明,我们的方法在生成准确,高质量的舞蹈动作方面取得了更好,更鲁棒的性能。我们的源代码和数据集可以分别从https://github.com/MarasyZZ/QEAN和https://google.github.io/aistplusplus_dataset获得。摘要:The study of music-generated dance is a novel and challenging Image generation task. It aims to input a piece of music and seed motions, then generate natural dance movements for the subsequent music. Transformer-based methods face challenges in time series prediction tasks related to human movements and music due to their struggle in capturing the nonlinear relationship and temporal aspects. This can lead to issues like joint deformation, role deviation, floating, and inconsistencies in dance movements generated in response to the music. In this paper, we propose a Quaternion-Enhanced Attention Network (QEAN) for visual dance synthesis from a quaternion perspective, which consists of a Spin Position Embedding (SPE) module and a Quaternion Rotary Attention (QRA) module. First, SPE embeds position information into self-attention in a rotational manner, leading to better learning of features of movement sequences and audio sequences, and improved understanding of the connection between music and dance. Second, QRA represents and fuses 3D motion features and audio features in the form of a series of quaternions, enabling the model to better learn the temporal coordination of music and dance under the complex temporal cycle conditions of dance generation. Finally, we conducted experiments on the dataset AIST++, and the results show that our approach achieves better and more robust performance in generating accurate, high-quality dance movements. Our source code and dataset can be available from https://github.com/MarasyZZ/QEAN and https://google.github.io/aistplusplus_dataset respectively. 【17】 Multitask frame-level learning for few-shot sound event detection标题:基于多任务帧级学习的Few-Shot声音事件检测链接:https://arxiv.org/abs/2403.11091作者:Liang Zou,Genwei Yan,Ruoyu Wang,Jun Du,Meng Lei,Tian Gao,Xin Fang备注:6 pages, 4 figures, conference摘要:本文研究了Few-Shot声音事件检测(SED),旨在利用有限的样本对声音事件进行自动识别和分类。然而,在Few-Shot SED中的流行方法主要依赖于片段级预测,其通常提供详细的细粒度预测,特别是对于持续时间短的事件。虽然已经提出了帧级预测策略来克服这些限制,但是这些策略通常面临由背景噪声引起的预测截断的困难。为了缓解这个问题,我们引入了一个创新的多任务帧级SED框架。此外,我们还引入了TimeFilterAug,一种用于数据增强的线性时序掩模,以提高模型对不同声学环境的鲁棒性和适应性。所提出的方法实现了63.8%的F分数,确保了2023年声学场景和事件挑战赛的检测和分类的Few-Shot生物声学事件检测类别中的第一名。摘要:This paper focuses on few-shot Sound Event Detection (SED), which aims to automatically recognize and classify sound events with limited samples. However, prevailing methods methods in few-shot SED predominantly rely on segment-level predictions, which often providing detailed, fine-grained predictions, particularly for events of brief duration. Although frame-level prediction strategies have been proposed to overcome these limitations, these strategies commonly face difficulties with prediction truncation caused by background noise. To alleviate this issue, we introduces an innovative multitask frame-level SED framework. In addition, we introduce TimeFilterAug, a linear timing mask for data augmentation, to increase the model's robustness and adaptability to diverse acoustic environments. The proposed method achieves a F-score of 63.8%, securing the 1st rank in the few-shot bioacoustic event detection category of the Detection and Classification of Acoustic Scenes and Events Challenge 2023. 【18】 Audio-Visual Segmentation via Unlabeled Frame Exploitation标题:基于无标记帧挖掘的视听分割链接:https://arxiv.org/abs/2403.11074作者:Jinxiang Liu,Yikun Liu,Fei Zhang,Chen Ju,Ya Zhang,Yanfeng Wang备注:Accepted by CVPR 2024摘要:音视频分割(AVS)的目的是分割视频帧中的声音对象。虽然已经取得了很大的进展,我们的实验表明,目前的方法达到边际性能增益内使用的未标记的帧,导致未充分利用的问题。为了充分探索未标记帧用于AVS的潜力,我们基于它们的时间特性将它们明确地分为两类,即,相邻帧(NF)和远距离帧(DF)。在时间上与标记帧相邻的NF通常包含丰富的运动信息,其有助于准确定位发声对象。与NF相反,DF与标记帧具有长的时间距离,其共享具有外观变化的语义相似的对象。考虑到它们的独特特性,我们提出了一个通用的框架,有效地利用它们来解决AVS。具体来说,对于NF,我们利用运动线索作为动态指导,以提高目标定位。此外,我们利用语义线索的DF处理他们作为有效的增强标记的帧,然后使用丰富的数据多样性的自我训练的方式。大量的实验结果表明,我们的方法的通用性和优越性,释放丰富的未标记帧的力量。摘要:Audio-visual segmentation (AVS) aims to segment the sounding objects in video frames. Although great progress has been witnessed, we experimentally reveal that current methods reach marginal performance gain within the use of the unlabeled frames, leading to the underutilization issue. To fully explore the potential of the unlabeled frames for AVS, we explicitly divide them into two categories based on their temporal characteristics, i.e., neighboring frame (NF) and distant frame (DF). NFs, temporally adjacent to the labeled frame, often contain rich motion information that assists in the accurate localization of sounding objects. Contrary to NFs, DFs have long temporal distances from the labeled frame, which share semantic-similar objects with appearance variations. Considering their unique characteristics, we propose a versatile framework that effectively leverages them to tackle AVS. Specifically, for NFs, we exploit the motion cues as the dynamic guidance to improve the objectness localization. Besides, we exploit the semantic cues in DFs by treating them as valid augmentations to the labeled frames, which are then used to enrich data diversity in a self-training manner. Extensive experimental results demonstrate the versatility and superiority of our method, unleashing the power of the abundant unlabeled frames.
【19】 Energy-Based Models with Applications to Speech and Language Processing标题:基于能量的模型及其在语音和语言处理中的应用链接:https://arxiv.org/abs/2403.10961作者:Zhijian Ou备注:None摘要:基于能量的模型(EBM)是一类重要的概率模型,也称为随机场和无向图模型。EBM是未归一化的,因此与其他流行的自归一化概率模型(如隐马尔可夫模型(HRM),自回归模型,生成对抗网络(GANs)和变分自动编码器(VAE))完全不同。在过去的几年里,EBM不仅吸引了核心机器学习社区的兴趣,而且由于重大的理论和算法进展,还吸引了语音,视觉,自然语言处理(NLP)等应用领域的兴趣。语音和语言的顺序性质也提出了特殊的挑战,并且需要与处理固定维度数据不同的处理(例如,图像)。因此,本专著的目的是系统地介绍基于能量的模型,包括语音和语言处理中的算法进展和应用。首先,介绍了EBM的基础知识,包括经典模型,最近的模型参数化的神经网络,采样方法,以及各种学习方法,从经典的学习算法到最先进的。然后,在三个不同的情况下,EBM的应用,即,分别用于边际分布、条件分布和联合分布的建模。1)序列数据的EBM及其在语言建模中的应用,主要关注序列本身的边缘分布:2)给定观察序列的目标序列的条件分布的EBM,其应用于语音识别、序列标记和文本生成; 3)用于对观测序列和目标序列两者的联合分布进行建模的EBM,以及它们在半监督学习和校准自然语言理解中的应用。摘要:Energy-Based Models (EBMs) are an important class of probabilistic models, also known as random fields and undirected graphical models. EBMs are un-normalized and thus radically different from other popular self-normalized probabilistic models such as hidden Markov models (HMMs), autoregressive models, generative adversarial nets (GANs) and variational auto-encoders (VAEs). Over the past years, EBMs have attracted increasing interest not only from the core machine learning community, but also from application domains such as speech, vision, natural language processing (NLP) and so on, due to significant theoretical and algorithmic progress. The sequential nature of speech and language also presents special challenges and needs a different treatment from processing fix-dimensional data (e.g., images). Therefore, the purpose of this monograph is to present a systematic introduction to energy-based models, including both algorithmic progress and applications in speech and language processing. First, the basics of EBMs are introduced, including classic models, recent models parameterized by neural networks, sampling methods, and various learning methods from the classic learning algorithms to the most advanced ones. Then, the application of EBMs in three different scenarios is presented, i.e., for modeling marginal, conditional and joint distributions, respectively. 1) EBMs for sequential data with applications in language modeling, where the main focus is on the marginal distribution of a sequence itself; 2) EBMs for modeling conditional distributions of target sequences given observation sequences, with applications in speech recognition, sequence labeling and text generation; 3) EBMs for modeling joint distributions of both sequences of observations and targets, and their applications in semi-supervised learning and calibrated natural language understanding. 【20】 Urban Sound Propagation: a Benchmark for 1-Step Generative Modeling of Complex Physical Systems标题:城市声传播:复杂物理系统一步生成性建模的基准链接:https://arxiv.org/abs/2403.10904作者:Martin Spitznagel,Janis Keuper摘要:复杂物理系统的数据驱动建模在仿真和机器学习社区中受到越来越多的关注。由于大多数物理模拟都是基于计算密集型的微分方程系统的迭代实现,因此用学习的一步推理模型(部分)替代具有在广泛的应用领域显着加速的潜力。在这种情况下,我们提出了一个新的基准评估1步生成学习模型的速度和物理正确性。我们的Urban Sound Propagation基准是基于物理上复杂且实际相关的,但直观上易于掌握的任务,即对城市环境中声源的波的2D传播进行建模。我们提供了一个具有10万个样本的数据集,其中每个样本由从OpenStreetmap绘制的真实2D建筑物地图对,参数化声源和给定场景的模拟地面真实声音传播组成。该数据集提供了四种不同的模拟任务,其反射、衍射和源方差的复杂性不断增加。常见的生成U-Net、GAN和扩散模型的第一基线评估表明,虽然这些模型能够在简单情况下很好地对声音传播进行建模,但由高阶方程表示的子系统的近似系统地失败。有关数据集、下载说明和源代码的信息,请访问我们的匿名网站:https://www.urban-sound-data.org。摘要:Data-driven modeling of complex physical systems is receiving a growing amount of attention in the simulation and machine learning communities. Since most physical simulations are based on compute-intensive, iterative implementations of differential equation systems, a (partial) replacement with learned, 1-step inference models has the potential for significant speedups in a wide range of application areas. In this context, we present a novel benchmark for the evaluation of 1-step generative learning models in terms of speed and physical correctness. Our Urban Sound Propagation benchmark is based on the physically complex and practically relevant, yet intuitively easy to grasp task of modeling the 2d propagation of waves from a sound source in an urban environment. We provide a dataset with 100k samples, where each sample consists of pairs of real 2d building maps drawn from OpenStreetmap, a parameterized sound source, and a simulated ground truth sound propagation for the given scene. The dataset provides four different simulation tasks with increasing complexity regarding reflection, diffraction and source variance. A first baseline evaluation of common generative U-Net, GAN and Diffusion models shows, that while these models are very well capable of modeling sound propagations in simple cases, the approximation of sub-systems represented by higher order equations systematically fails. Information about the dataset, download instructions and source codes are provided on our anonymous website: https://www.urban-sound-data.org. 【21】 Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference标题:语音驱动的个性化手势合成:利用自动模糊特征推理链接:https://arxiv.org/abs/2403.10805作者:Fan Zhang,Zhaohan Wang,Xin Lyu,Siyuan Zhao,Mengjian Li,Weidong Geng,Naye Ji,Hui Du,Fuxing Gao,Hao Wu,Shunman Li备注:12 pages,摘要:语音驱动的手势生成是虚拟人创作中的一个新兴领域。然而,一个重大的挑战在于准确地确定和处理大量的输入特征(如声学,语义,情感,个性,甚至微妙的未知特征)。传统的方法,依赖于各种显式的特征输入和复杂的多模态处理,限制了手势的表现力,并限制了它们的适用性。为了应对这些挑战,我们提出了Persona-Gestro,这是一种新型的端到端生成模型,旨在仅依赖原始语音音频生成高度个性化的3D全身手势。该模型结合了模糊特征提取器和非自回归自适应层归一化(AdaLN)Transformer扩散架构。模糊特征提取器利用模糊推理策略,自动推断隐含的、连续的模糊特征。这些模糊特征,表示为一个统一的潜在特征,被送入AdaLN Transformer。AdaLN Transformer引入了一种条件机制,该机制在所有令牌上应用统一函数,从而有效地对模糊特征和手势序列之间的相关性进行建模。该模块确保高水平的手势-语音同步,同时保持自然性。最后,我们采用扩散模型来训练和推断各种手势。对Trinity、ZEGGS和BEAT数据集进行的广泛的主观和客观评估证实了我们的模型优于当前最先进的方法。Persona-Gestro提高了系统的可用性和泛化能力,为语音驱动的手势合成设定了新的基准,并拓宽了虚拟人技术的视野。补充视频和代码可访问https://zf223669.github.io/Diffmotion-v2-website/摘要:Speech-driven gesture generation is an emerging field within virtual human creation. However, a significant challenge lies in accurately determining and processing the multitude of input features (such as acoustic, semantic, emotional, personality, and even subtle unknown features). Traditional approaches, reliant on various explicit feature inputs and complex multimodal processing, constrain the expressiveness of resulting gestures and limit their applicability. To address these challenges, we present Persona-Gestor, a novel end-to-end generative model designed to generate highly personalized 3D full-body gestures solely relying on raw speech audio. The model combines a fuzzy feature extractor and a non-autoregressive Adaptive Layer Normalization (AdaLN) transformer diffusion architecture. The fuzzy feature extractor harnesses a fuzzy inference strategy that automatically infers implicit, continuous fuzzy features. These fuzzy features, represented as a unified latent feature, are fed into the AdaLN transformer. The AdaLN transformer introduces a conditional mechanism that applies a uniform function across all tokens, thereby effectively modeling the correlation between the fuzzy features and the gesture sequence. This module ensures a high level of gesture-speech synchronization while preserving naturalness. Finally, we employ the diffusion model to train and infer various gestures. Extensive subjective and objective evaluations on the Trinity, ZEGGS, and BEAT datasets confirm our model's superior performance to the current state-of-the-art approaches. Persona-Gestor improves the system's usability and generalization capabilities, setting a new benchmark in speech-driven gesture synthesis and broadening the horizon for virtual human technology. Supplementary videos and code can be accessed at https://zf223669.github.io/Diffmotion-v2-website/
【22】 CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing标题:CoPlay:听觉不可知的认知尺度链接:https://arxiv.org/abs/2403.10796作者:Yin Li,Rajalakshmi Nanadakumar摘要:通过利用智能设备上的扬声器和麦克风,声学感测在包括健康监测、手势接口和成像的各种应用中表现出巨大的潜力。然而,在声学传感的持续研究和开发中,一个问题经常被忽视:当同时用于传感和其他传统应用(如播放音乐)时,同一个扬声器可能会在两者中造成干扰,使其在现实世界中不切实际。强烈的超声波感应信号与音乐混合在一起,会使扬声器的混音器过载。为了应对过载信号的这个问题,当前的解决方案是限幅(clipping)或缩小(down-scaling),这两者都影响音乐回放质量以及感测范围和准确度。为了应对这一挑战,我们提出了CoPlay,这是一种基于深度学习的优化算法,可以认知地适应感知信号。它可以1)最大化由并发音乐留下的可用带宽内的感测信号幅度,以优化感测范围和准确度,以及2)最小化可能影响音乐回放的任何相应的频率失真。在这项工作中,我们设计了一个深度学习模型,并在常见类型的感知信号(正弦波或调频连续波FMCW)上进行测试,作为各种不可知并发音乐和语音的输入。首先,我们评估了模型的性能,以显示生成的信号的质量。然后,我们在现实世界中的下游声学传感任务进行了实地研究。一项针对12名用户的研究证明,使用我们的自适应信号进行呼吸监测和手势识别可以达到与无并发音乐场景相似的准确性,而剪切或缩小显示出更差的准确性。定性研究还表明,音乐播放质量没有下降,不像传统的裁剪或缩小的方法。摘要:Acoustic sensing manifests great potential in various applications that encompass health monitoring, gesture interface and imaging by leveraging the speakers and microphones on smart devices. However, in ongoing research and development in acoustic sensing, one problem is often overlooked: the same speaker, when used concurrently for sensing and other traditional applications (like playing music), could cause interference in both making it impractical to use in the real world. The strong ultrasonic sensing signals mixed with music would overload the speaker's mixer. To confront this issue of overloaded signals, current solutions are clipping or down-scaling, both of which affect the music playback quality and also sensing range and accuracy. To address this challenge, we propose CoPlay, a deep learning based optimization algorithm to cognitively adapt the sensing signal. It can 1) maximize the sensing signal magnitude within the available bandwidth left by the concurrent music to optimize sensing range and accuracy and 2) minimize any consequential frequency distortion that can affect music playback. In this work, we design a deep learning model and test it on common types of sensing signals (sine wave or Frequency Modulated Continuous Wave FMCW) as inputs with various agnostic concurrent music and speech. First, we evaluated the model performance to show the quality of the generated signals. Then we conducted field studies of downstream acoustic sensing tasks in the real world. A study with 12 users proved that respiration monitoring and gesture recognition using our adapted signal achieve similar accuracy as no-concurrent-music scenarios, while clipping or down-scaling manifests worse accuracy. A qualitative study also manifests that the music play quality is not degraded, unlike traditional clipping or down-scaling methods.
【23】 On-Device Domain Learning for Keyword Spotting on Low-Power Extreme Edge Embedded Systems标题:低功耗Extreme Edge嵌入式系统中关键字检测的设备上领域学习链接:https://arxiv.org/abs/2403.10549作者:Cristian Cioflan,Lukas Cavigelli,Manuele Rusci,Miguel de Prado,Luca Benini备注:5 pages, 2 tables, 2 figures. Accepted at IEEE AICAS 2024摘要:当神经网络暴露在嘈杂的环境中时,关键字定位的准确性会降低。现场适应以前看不见的噪声对于恢复精度损失至关重要,需要在设备上学习以确保适应过程完全发生在边缘设备上。在这项工作中,我们提出了一个完全在设备上的域自适应系统,实现了高达14%的准确性增益超过已经强大的关键字定位模型。我们可以在不到10 kB的内存中实现设备学习,仅使用100个标记的话语,在适应复杂的语音噪声后恢复5%的准确率。我们证明,域自适应可以在超低功耗微控制器上实现,只需14秒,在始终开启,电池供电的设备上只需806 mJ。摘要:Keyword spotting accuracy degrades when neural networks are exposed to noisy environments. On-site adaptation to previously unseen noise is crucial to recovering accuracy loss, and on-device learning is required to ensure that the adaptation process happens entirely on the edge device. In this work, we propose a fully on-device domain adaptation system achieving up to 14% accuracy gains over already-robust keyword spotting models. We enable on-device learning with less than 10 kB of memory, using only 100 labeled utterances to recover 5% accuracy after adapting to the complex speech noise. We demonstrate that domain adaptation can be achieved on ultra-low-power microcontrollers with as little as 806 mJ in only 14 s on always-on, battery-operated devices.