今日论文合集:cs.SD语音9篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation  Guided by the Characteristic Dance Primitives

标题:Lodge:一种以特征舞蹈基元为导向的长舞世代由粗到精的传播网络

链接:https://arxiv.org/abs/2403.10518

作者:Ronghui Li,YuXiang Zhang,Yachao Zhang,Hongwen Zhang,Jie Guo,Yan Zhang,Yebin Liu,Xiu Li摘要:我们提出洛奇,一个网络能够产生非常长的舞蹈序列的条件下,给定的音乐。我们设计洛奇作为一个两阶段的粗到细的扩散架构,并提出了具有显着的表现力的特征舞蹈原语作为两个扩散模型之间的中间表示。第一阶段是全球扩散,重点是理解粗层次的音乐-舞蹈相关性和生产特征的舞蹈原语。第二阶段是局部扩散,在舞蹈原语和编舞规则的指导下,局部扩散生成详细的动作序列。此外,我们提出了一个脚优化块,以优化脚和地面之间的接触,增强运动的物理现实主义。我们的方法可以快速生成非常长的舞蹈序列,在全局编舞模式和局部运动质量和表现力之间取得平衡。大量的实验验证了我们的方法的有效性。

摘要:We propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations between two diffusion models. The first stage is global diffusion, which focuses on comprehending the coarse-level music-dance correlation and production characteristic dance primitives. In contrast, the second-stage is the local diffusion, which parallelly generates detailed motion sequences under the guidance of the dance primitives and choreographic rules. In addition, we propose a Foot Refine Block to optimize the contact between the feet and the ground, enhancing the physical realism of the motion. Our approach can parallelly generate dance sequences of extremely long length, striking a balance between global choreographic patterns and local motion quality and expressiveness. Extensive experiments validate the efficacy of our method.


【2】 MusicHiFi: Fast High-Fidelity Stereo Vocoding
标题:MusicHiFi:快速高保真立体声编码
链接:https://arxiv.org/abs/2403.10493
作者:Ge Zhu,Juan-Pablo Caceres,Zhiyao Duan,Nicholas J. Bryan
摘要:基于扩散的音频和音乐生成模型通常通过构造音频的图像表示(例如,MEL频谱图),然后使用相位重构模型或声码器将其转换成音频。然而,典型的声码器以较低的分辨率产生单声道音频(例如,16-24 kHz),这限制了它们的有效性。我们提出了MusicHiFi -一种高效的高保真立体声声码器。我们的方法采用了三个生成对抗网络(GANs)的级联,将低分辨率梅尔频谱图转换为音频,通过带宽扩展上采样为高分辨率音频,并上混为立体声音频。与以前的工作相比,我们提出了1)一个统一的基于GAN的生成器和FPGA架构以及级联每个阶段的训练过程,2)一个新的快速,接近下采样兼容的带宽扩展模块,以及3)一个新的快速下混兼容的单声道到立体声上混器,确保在输出中保留单声道内容。我们使用客观和主观的听力测试来评估我们的方法,并发现我们的方法与过去的工作相比,可以产生相当或更好的音频质量,更好的空间化控制和更快的推理速度。健全的例子是在https://MusicHiFi.github.io/web/。
摘要:Diffusion-based audio and music generation models commonly generate music by constructing an image representation of audio (e.g., a mel-spectrogram) and then converting it to audio using a phase reconstruction model or vocoder. Typical vocoders, however, produce monophonic audio at lower resolutions (e.g., 16-24 kHz), which limits their effectiveness. We propose MusicHiFi -- an efficient high-fidelity stereophonic vocoder. Our method employs a cascade of three generative adversarial networks (GANs) that convert low-resolution mel-spectrograms to audio, upsamples to high-resolution audio via bandwidth expansion, and upmixes to stereophonic audio. Compared to previous work, we propose 1) a unified GAN-based generator and discriminator architecture and training procedure for each stage of our cascade, 2) a new fast, near downsampling-compatible bandwidth extension module, and 3) a new fast downmix-compatible mono-to-stereo upmixer that ensures the preservation of monophonic content in the output. We evaluate our approach using both objective and subjective listening tests and find our approach yields comparable or better audio quality, better spatialization control, and significantly faster inference speed compared to past work. Sound examples are at https://MusicHiFi.github.io/web/.

【3】 Joint Multimodal Transformer for Dimensional Emotional Recognition in  the Wild
标题:用于野外空间情绪识别的联合多模式转换器
链接:https://arxiv.org/abs/2403.10488
作者:Paul Waligora,Osama Zeeshan,Haseeb Aslam,Soufiane Belharbi,Alessandro Lameiras Koerich,Marco Pedersoli,Simon Bacon,Eric Granger
备注:5 pages, 1 figure
摘要:视频中的视听情感识别(ER)在单峰性能上具有巨大的潜力。它有效地利用了视觉和听觉模态之间的模态间和模态内依赖性。本文提出了一种新的视听情感识别系统,该系统采用基于键的交叉注意的联合多模态Transformer架构。该框架旨在利用视频中音频和视觉线索(面部表情和声音模式)的互补性,与仅依赖单一模态相比,具有更好的性能。所提出的模型利用单独的骨干捕捉内模态的时间依赖性在每个模态(音频和视频)。随后,联合多模态Transformer架构集成了各个模态嵌入,使模型能够有效地捕获模态间(音频和视觉之间)和模态内(每个模态内)的关系。在具有挑战性的Affwild2数据集上进行的广泛评估表明,所提出的模型在ER任务中显着优于基线和最先进的方法。
摘要:Audiovisual emotion recognition (ER) in videos has immense potential over unimodal performance. It effectively leverages the inter- and intra-modal dependencies between visual and auditory modalities. This work proposes a novel audio-visual emotion recognition system utilizing a joint multimodal transformer architecture with key-based cross-attention. This framework aims to exploit the complementary nature of audio and visual cues (facial expressions and vocal patterns) in videos, leading to superior performance compared to solely relying on a single modality. The proposed model leverages separate backbones for capturing intra-modal temporal dependencies within each modality (audio and visual). Subsequently, a joint multimodal transformer architecture integrates the individual modality embeddings, enabling the model to effectively capture inter-modal (between audio and visual) and intra-modal (within each modality) relationships. Extensive evaluations on the challenging Affwild2 dataset demonstrate that the proposed model significantly outperforms baseline and state-of-the-art methods in ER tasks.

【4】 BirdSet: A Multi-Task Benchmark for Classification in Avian Bioacoustics
标题:BirdSet:一个用于鸟类生物声学分类的多任务基准
链接:https://arxiv.org/abs/2403.10380
作者:Lukas Rauch,Raphael Schwinger,Moritz Wirth,René Heinrich,Jonas Lange,Stefan Kahl,Bernhard Sick,Sven Tomforde,Christoph Scholz
备注:Work in progress, to be submitted @DMLR next month
摘要:深度学习(DL)模型已经成为鸟类生物声学诊断环境健康和生物多样性的强大工具。然而,研究中的不一致性构成了阻碍这一领域取得进展的显著挑战。可靠的DL模型需要灵活地分析各种物种和环境中的鸟类叫声,以充分利用生物声学在经济高效的被动声学监测方案中的潜力。跨研究的数据碎片化和不透明性使一般模型性能的综合评价复杂化。为了克服这些挑战,我们提出了BirdSet基准,一个统一的框架,巩固研究工作的整体方法分类鸟类生物声学中的鸟类发声。BirdSet将开源的鸟类记录整合到一个精心策划的数据集中。这种统一的方法提供了对模型性能的深入理解,并确定了不同任务中的潜在缺点。通过建立当前模型的基线结果,BirdSet旨在促进可比性,指导后续数据收集,并增加新来者对鸟类生物声学的可访问性。
摘要:Deep learning (DL) models have emerged as a powerful tool in avian bioacoustics to diagnose environmental health and biodiversity. However, inconsistencies in research pose notable challenges hindering progress in this domain. Reliable DL models need to analyze bird calls flexibly across various species and environments to fully harness the potential of bioacoustics in a cost-effective passive acoustic monitoring scenario. Data fragmentation and opacity across studies complicate a comprehensive evaluation of general model performance. To overcome these challenges, we present the BirdSet benchmark, a unified framework consolidating research efforts with a holistic approach for classifying bird vocalizations in avian bioacoustics. BirdSet harmonizes open-source bird recordings into a curated dataset collection. This unified approach provides an in-depth understanding of model performance and identifies potential shortcomings across different tasks. By establishing baseline results of current models, BirdSet aims to facilitate comparability, guide subsequent data collection, and increase accessibility for newcomers to avian bioacoustics.

【5】 Multiscale Matching Driven by Cross-Modal Similarity Consistency for  Audio-Text Retrieval
标题:基于跨模态相似一致性驱动的多尺度匹配音频文本检索
链接:https://arxiv.org/abs/2403.10146
作者:Qian Wang,Jia-Chen Gu,Zhen-Hua Ling
备注:5 pages, accepted to ICASSP2024
摘要:音频文本检索(ATR),即从给定的音频片段(A2 T)中检索相关的字幕,反之亦然(T2 A),近年来引起了广泛的研究关注。现有的方法通常将来自每个模态的信息聚合到单个向量中进行匹配,但是这牺牲了局部细节并且难以捕获模态内和模态之间的复杂关系。此外,目前的ATR数据集缺乏全面的对齐信息,简单的二进制对比学习标签忽略了样本之间细粒度语义差异的测量。为了应对这些挑战,我们提出了一个新的ATR框架,全面捕捉多模态信息的匹配关系,从不同的角度和更细的粒度。具体来说,引入了一种细粒度的对齐方法,通过从局部到全局的多尺度过程来捕获细致的跨模态关系,从而实现更面向细节的匹配。此外,我们开创了跨模态相似性一致性的应用,利用模态内相似性关系作为软监督来促进更复杂的对齐。大量的实验验证了我们的方法的有效性,在AudioCaps数据集上至少有3.9%(T2 A)/ 6.9%(A2 T)R@1的显着裕度,在Clotho数据集上至少有2.9%(T2 A)/ 5.4%(A2 T)R@1的显着裕度。
摘要:Audio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single vector for matching, but this sacrifices local details and can hardly capture intricate relationships within and between modalities. Furthermore, current ATR datasets lack comprehensive alignment information, and simple binary contrastive learning labels overlook the measurement of fine-grained semantic differences between samples. To counter these challenges, we present a novel ATR framework that comprehensively captures the matching relationships of multimodal information from different perspectives and finer granularities. Specifically, a fine-grained alignment method is introduced, achieving a more detail-oriented matching through a multiscale process from local to global levels to capture meticulous cross-modal relationships. In addition, we pioneer the application of cross-modal similarity consistency, leveraging intra-modal similarity relationships as soft supervision to boost more intricate alignment. Extensive experiments validate the effectiveness of our approach, outperforming previous methods by significant margins of at least 3.9% (T2A) / 6.9% (A2T) R@1 on the AudioCaps dataset and 2.9% (T2A) / 5.4% (A2T) R@1 on the Clotho dataset.

【6】 MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate  Instrument Leakage
标题:MR—MT3:记忆保留多轨音乐转录以减轻仪器泄漏
链接:https://arxiv.org/abs/2403.10024
作者:Hao Hao Tan,Kin Wai Cheuk,Taemin Cho,Wei-Hsiang Liao,Yuki Mitsufuji
摘要:本文介绍了MT3模型的增强,这是一种基于SOTA令牌的多乐器自动音乐转录(AMT)模型。尽管有SOTA性能,但MT3存在仪器泄漏问题,其中跨不同仪器的传输是分散的。为了缓解这一问题,我们提出了MR-MT3,并提出了增强功能,包括内存保留机制,事先令牌采样和令牌洗牌。这些方法在Slakh 2100数据集上进行了评估,证明了改善的发作F1评分和减少的仪器泄漏。除了传统的多仪器转录F1分数外,还引入了仪器泄漏率和仪器检测F1分数等新指标,以更全面地评估转录质量。该研究还通过在ComMU和NSynth等单乐器单声道数据集上评估MT3来探索域过拟合问题。研究结果以及源代码将被共享,以促进未来旨在改进基于令牌的多仪器AMT模型的工作。
摘要:This paper presents enhancements to the MT3 model, a state-of-the-art (SOTA) token-based multi-instrument automatic music transcription (AMT) model. Despite SOTA performance, MT3 has the issue of instrument leakage, where transcriptions are fragmented across different instruments. To mitigate this, we propose MR-MT3, with enhancements including a memory retention mechanism, prior token sampling, and token shuffling are proposed. These methods are evaluated on the Slakh2100 dataset, demonstrating improved onset F1 scores and reduced instrument leakage. In addition to the conventional multi-instrument transcription F1 score, new metrics such as the instrument leakage ratio and the instrument detection F1 score are introduced for a more comprehensive assessment of transcription quality. The study also explores the issue of domain overfitting by evaluating MT3 on single-instrument monophonic datasets such as ComMU and NSynth. The findings, along with the source code, are shared to facilitate future work aimed at refining token-based multi-instrument AMT models.

【7】 SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification  of Spoken Numbers in Different Languages
标题:口语-100:用于不同语言口语数字分类的跨语言基准数据集
链接:https://arxiv.org/abs/2403.09753
作者:René Groh,Nina Goes,Andreas M. Kist
备注:Accepted as a full paper by the tinyML Research Symposium 2024
摘要:基准测试在评估和增强紧凑型深度学习模型的性能方面发挥着关键作用,这些模型是为在资源受限的设备(如微控制器)上执行而设计的。我们的研究引入了一个全新的、完全人工生成的、专为语音识别量身定制的基准测试数据集,这是微型深度学习领域的一个核心挑战。SpokeN-100由32个不同的说话者用四种不同的语言(英语、普通话、德语和法语)说出的从0到99的数字组成,产生了12,800个音频样本。我们确定听觉特征,并使用UMAP(均匀流形近似和投影降维)作为降维方法,以显示数据集的多样性和丰富性。为了突出数据集的用例,我们引入了两个基准任务:给定音频样本,对(i)使用的语言和/或(ii)口语数字进行分类。我们优化了最先进的深度神经网络,并执行了进化神经架构搜索,以找到针对32位ARM Cortex-M4 nRF 52840微控制器优化的微型架构。我们的结果代表了SpokeN-100实现的第一个基准数据。
摘要:Benchmarking plays a pivotal role in assessing and enhancing the performance of compact deep learning models designed for execution on resource-constrained devices, such as microcontrollers. Our study introduces a novel, entirely artificially generated benchmarking dataset tailored for speech recognition, representing a core challenge in the field of tiny deep learning. SpokeN-100 consists of spoken numbers from 0 to 99 spoken by 32 different speakers in four different languages, namely English, Mandarin, German and French, resulting in 12,800 audio samples. We determine auditory features and use UMAP (Uniform Manifold Approximation and Projection for Dimension Reduction) as a dimensionality reduction method to show the diversity and richness of the dataset. To highlight the use case of the dataset, we introduce two benchmark tasks: given an audio sample, classify (i) the used language and/or (ii) the spoken number. We optimized state-of-the-art deep neural networks and performed an evolutionary neural architecture search to find tiny architectures optimized for the 32-bit ARM Cortex-M4 nRF52840 microcontroller. Our results represent the first benchmark data achieved for SpokeN-100.


【8】 Multi-Source Localization and Data Association for Time-Difference of  Arrival Measurements
标题:到达时间差测量的多源定位和数据关联
链接:https://arxiv.org/abs/2403.10329
作者:Gabrielle Flood,Filip Elvander
摘要:在这项工作中,我们考虑的问题,定位多个信号源的到达时间差(TDOA)测量的基础上。在盲环境中,源信号是未知的,定位任务是具有挑战性的,由于数据关联问题。也就是说,不知道哪个TDOA测量值对应于相同的源。在这里,我们建议通过一个最佳的运输配方进行联合定位和数据关联。该方法通过寻找TDOA测量的最佳分组并将其与候选源位置相关联来操作。为了允许在三维空间中的计算上可行的定位,使用基于接收器对的最小集合的最小迭代求解器来构造候选位置的有效集合。在数值模拟中,我们证明了所提出的方法是强大的测量噪声和TDOA检测误差。此外,它表明,所提出的方法提供的数据关联允许统计上有效的估计的源位置。
摘要:In this work, we consider the problem of localizing multiple signal sources based on time-difference of arrival (TDOA) measurements. In the blind setting, in which the source signals are not known, the localization task is challenging due to the data association problem. That is, it is not known which of the TDOA measurements correspond to the same source. Herein, we propose to perform joint localization and data association by means of an optimal transport formulation. The method operates by finding optimal groupings of TDOA measurements and associating these with candidate source locations. To allow for computationally feasible localization in three-dimensional space, an efficient set of candidate locations is constructed using a minimal multilateration solver based on minimal sets of receiver pairs. In numerical simulations, we demonstrate that the proposed method is robust both to measurement noise and TDOA detection errors. Furthermore, it is shown that the data association provided by the proposed method allows for statistically efficient estimates of the source locations.


【9】 Audiosockets: A Python socket package for Real-Time Audio Processing
标题:Audiosockets:用于实时音频处理的Python套接字包
链接:https://arxiv.org/abs/2403.09789
作者:Nicolas Shu,David V. Anderson
备注:4 pages, 2 figures
摘要:Python中有许多包允许对音频数据进行实时处理。不幸的是,由于语言的同步性质,缺乏一个框架,它允许分布式并行处理的数据,而不需要一个大的编程开销,其中的数据采集不会被阻塞的后续处理操作。这项工作改进了用于音频数据收集的软件包,具有轻量级后端和简单的接口,允许通过基于套接字的结构进行分布式处理。这是为了在Python中进行实时音频机器学习和数据处理,并在同一数据上快速部署多个并行操作,使用户能够花费更少的时间进行调试,并将更多的时间用于开发。
摘要:There are many packages in Python which allow one to perform real-time processing on audio data. Unfortunately, due to the synchronous nature of the language, there lacks a framework which allows for distributed parallel processing of the data without requiring a large programming overhead and in which the data acquisition is not blocked by subsequent processing operations. This work improves on packages used for audio data collection with a light-weight backend and a simple interface that allows for distributed processing through a socket-based structure. This is intended for real-time audio machine learning and data processing in Python with a quick deployment of multiple parallel operations on the same data, allowing users to spend less time debugging and more time developing.

eess.AS音频处理
【1】 How to train your ears: Auditory-model emulation for large-dynamic-range  inputs and mild-to-severe hearing losses
标题:如何训练你的耳朵:大动态范围输入和轻度到重度听力损失的听觉模型仿真
链接:https://arxiv.org/abs/2403.10428
作者:Peter Leer,Jesper Jensen,Zheng-Hua Tan,Jan Østergaard,Lars Bramsløw
备注:Accepted by IEEE/ACM Transactions on Audio, Speech and Language Processing. This version is the authors' version and may vary from the final publication in details
摘要:高级听觉模型在设计用于听力损失补偿或语音增强的信号处理算法时非常有用。这样的听觉模型提供了丰富和详细的描述的听觉通路,并可能允许个性化的信号处理策略,基于生理测量。然而,这些听觉模型通常在计算上要求很高,需要大量时间来计算。为了解决这个问题,以前的研究已经探索了使用深度神经网络来模拟听觉模型并减少推理时间。虽然这些深度神经网络在计算时间方面提供了令人印象深刻的效率增益,但它们可能会受到仿真性能不均匀的影响,这是模型频率通道和输入声压级的函数,使它们不适合许多任务。在这项研究中,我们证明了现有最先进方法中使用的传统机器学习优化目标是这种限制的主要来源。具体而言,优化目标未能考虑听觉模型的频率和电平依赖性,这是由听觉模型模拟的大输入动态范围和不同类型的听力损失引起的。为了克服这一限制,我们提出了一个新的优化目标,明确嵌入的频率和水平依赖的听觉模型。我们的研究结果表明,这种新的优化目标显着提高了深度神经网络在相关输入声级和模型频率通道上的仿真性能,而不会增加推理过程中的计算负载。解决这些限制对于推进听觉模型在信号处理任务中的应用至关重要,确保其在不同场景中的有效性。
摘要:Advanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and might allow for individualization of signal-processing strategies, based on physiological measurements. However, these auditory models are often computationally demanding, requiring significant time to compute. To address this issue, previous studies have explored the use of deep neural networks to emulate auditory models and reduce inference time. While these deep neural networks offer impressive efficiency gains in terms of computational time, they may suffer from uneven emulation performance as a function of auditory-model frequency-channels and input sound pressure level, making them unsuitable for many tasks. In this study, we demonstrate that the conventional machine-learning optimization objective used in existing state-of-the-art methods is the primary source of this limitation. Specifically, the optimization objective fails to account for the frequency- and level-dependencies of the auditory model, caused by a large input dynamic range and different types of hearing losses emulated by the auditory model. To overcome this limitation, we propose a new optimization objective that explicitly embeds the frequency- and level-dependencies of the auditory model. Our results show that this new optimization objective significantly improves the emulation performance of deep neural networks across relevant input sound levels and auditory-model frequency channels, without increasing the computational load during inference. Addressing these limitations is essential for advancing the application of auditory models in signal-processing tasks, ensuring their efficacy in diverse scenarios.

【2】 Neural Networks Hear You Loud And Clear: Hearing Loss Compensation Using  Deep Neural Networks
标题:神经网络听得清清楚楚:使用深度神经网络进行听力损失补偿
链接:https://arxiv.org/abs/2403.10420
作者:Peter Leer,Jesper Jensen,Laurel Carney,Zheng-Hua Tan,Jan Østergaard,Lars Bramsløw
摘要:本文研究了深度神经网络(DNN)在听力损失补偿中的应用。听力损失是影响全球数百万人的普遍问题,传统助听器在提供满意的补偿方面存在局限性。DNN在各种听觉任务中表现出卓越的性能,包括语音识别,说话人识别和音乐分类。在这项研究中,我们提出了一种基于DNN的听力损失补偿方法,该方法是在听力受损和听力正常的基于DNN的听觉模型对语音信号的输出上进行训练的。首先,我们介绍了一个使用DNN仿真听觉模型的框架,重点是听觉通路中的神经元模型。我们提出了一种基于DNN的方法的线性化,我们使用它来分析基于DNN的听力损失补偿。此外,我们开发了一个简单的方法来选择用于补偿策略的听觉模型的声学中心频率。最后,我们评估了基于DNN的听力损失补偿策略,使用听力受损的听众的听力测试。结果表明,所提出的方法是可行的听力损失补偿策略。我们提出的方法被证明可以提供语音清晰度的增加,并被发现在感知语音质量方面优于传统的方法。
摘要:This article investigates the use of deep neural networks (DNNs) for hearing-loss compensation. Hearing loss is a prevalent issue affecting millions of people worldwide, and conventional hearing aids have limitations in providing satisfactory compensation. DNNs have shown remarkable performance in various auditory tasks, including speech recognition, speaker identification, and music classification. In this study, we propose a DNN-based approach for hearing-loss compensation, which is trained on the outputs of hearing-impaired and normal-hearing DNN-based auditory models in response to speech signals. First, we introduce a framework for emulating auditory models using DNNs, focusing on an auditory-nerve model in the auditory pathway. We propose a linearization of the DNN-based approach, which we use to analyze the DNN-based hearing-loss compensation. Additionally we develop a simple approach to choose the acoustic center frequencies of the auditory model used for the compensation strategy. Finally, we evaluate the DNN-based hearing-loss compensation strategies using listening tests with hearing impaired listeners. The results demonstrate that the proposed approach results in feasible hearing-loss compensation strategies. Our proposed approach was shown to provide an increase in speech intelligibility and was found to outperform a conventional approach in terms of perceived speech quality.


【3】 Multi-Source Localization and Data Association for Time-Difference of  Arrival Measurements
标题:到达时差测量的多源定位与数据关联
链接:https://arxiv.org/abs/2403.10329
作者:Gabrielle Flood,Filip Elvander
摘要:在这项工作中,我们考虑的问题,定位多个信号源的到达时间差(TDOA)测量的基础上。在盲环境中,源信号是未知的,定位任务是具有挑战性的,由于数据关联问题。也就是说,不知道哪个TDOA测量值对应于相同的源。在这里,我们建议通过一个最佳的运输配方进行联合定位和数据关联。该方法通过寻找TDOA测量的最佳分组并将其与候选源位置相关联来操作。为了允许在三维空间中的计算上可行的定位,使用基于接收器对的最小集合的最小迭代求解器来构造候选位置的有效集合。在数值模拟中,我们证明了所提出的方法是强大的测量噪声和TDOA检测误差。此外,它表明,所提出的方法提供的数据关联允许统计上有效的估计的源位置。
摘要:In this work, we consider the problem of localizing multiple signal sources based on time-difference of arrival (TDOA) measurements. In the blind setting, in which the source signals are not known, the localization task is challenging due to the data association problem. That is, it is not known which of the TDOA measurements correspond to the same source. Herein, we propose to perform joint localization and data association by means of an optimal transport formulation. The method operates by finding optimal groupings of TDOA measurements and associating these with candidate source locations. To allow for computationally feasible localization in three-dimensional space, an efficient set of candidate locations is constructed using a minimal multilateration solver based on minimal sets of receiver pairs. In numerical simulations, we demonstrate that the proposed method is robust both to measurement noise and TDOA detection errors. Furthermore, it is shown that the data association provided by the proposed method allows for statistically efficient estimates of the source locations.

【4】 SuperME: Supervised and Mixture-to-Mixture Co-Learning for Speech  Enhancement and Robust ASR
标题:SuperME:用于语音增强和鲁棒ASR的监督和混合到混合协同学习
链接:https://arxiv.org/abs/2403.10271
作者:Zhong-Qiu Wang
备注:in submission
摘要:目前,神经语音增强的主要方法是基于使用模拟训练数据的监督学习。然而,经过训练的模型通常对真实记录的数据表现出有限的泛化能力。为了解决这个问题,我们直接在真实的目标域数据上研究训练模型,并提出了两种算法,混合到混合(M2M)训练和一种协同学习算法,该算法在监督算法的帮助下改进了M2M。当成对的近距离说话和远场混合可用于训练时,M2M通过训练深度神经网络(DNN)来实现语音增强,以产生语音和噪声估计,使得它们可以被线性过滤以重建近距离说话和远场混合。通过这种方式,DNN可以直接在真实混合物上进行训练,并且可以利用近距离混合物作为弱监督来增强远场混合物。为了改进M2M,我们将其与监督方法相结合来共同训练DNN,其中小批量的真实近距离说话和远场混合对以及小批量的模拟混合和干净语音对交替地馈送到DNN,并且损失函数分别是(a)真实近场和远场混合上的混合重构损失和(b)在模拟的干净语音和噪声上的规则增强损失。我们发现,通过这种方式,DNN可以从真实和模拟数据中学习,以更好地泛化到真实数据。我们将此算法命名为SuperME、$\underline{super}$vised和$\underline{m}$ixture-to-mixtur$\underline{e}$ co-learning。CHiME-4数据集上的评估结果表明了其有效性和潜力。
摘要:The current dominant approach for neural speech enhancement is based on supervised learning by using simulated training data. The trained models, however, often exhibit limited generalizability to real-recorded data. To address this, we investigate training models directly on real target-domain data, and propose two algorithms, mixture-to-mixture (M2M) training and a co-learning algorithm that improves M2M with the help of supervised algorithms. When paired close-talk and far-field mixtures are available for training, M2M realizes speech enhancement by training a deep neural network (DNN) to produce speech and noise estimates in a way such that they can be linearly filtered to reconstruct the close-talk and far-field mixtures. This way, the DNN can be trained directly on real mixtures, and can leverage close-talk mixtures as a weak supervision to enhance far-field mixtures. To improve M2M, we combine it with supervised approaches to co-train the DNN, where mini-batches of real close-talk and far-field mixture pairs and mini-batches of simulated mixture and clean speech pairs are alternately fed to the DNN, and the loss functions are respectively (a) the mixture reconstruction loss on the real close-talk and far-field mixtures and (b) the regular enhancement loss on the simulated clean speech and noise. We find that, this way, the DNN can learn from real and simulated data to achieve better generalization to real data. We name this algorithm SuperME, $\underline{super}$vised and $\underline{m}$ixture-to-mixtur$\underline{e}$ co-learning. Evaluation results on the CHiME-4 dataset show its effectiveness and potential.


【5】 Audiosockets: A Python socket package for Real-Time Audio Processing
标题:Audiosockets:用于实时音频处理的Python套接字包
链接:https://arxiv.org/abs/2403.09789
作者:Nicolas Shu,David V. Anderson
备注:4 pages, 2 figures
摘要:Python中有许多包允许对音频数据进行实时处理。不幸的是,由于语言的同步性质,缺乏一个框架,它允许分布式并行处理的数据,而不需要一个大的编程开销,其中的数据采集不会被阻塞的后续处理操作。这项工作改进了用于音频数据收集的软件包,具有轻量级后端和简单的接口,允许通过基于套接字的结构进行分布式处理。这是为了在Python中进行实时音频机器学习和数据处理,并在同一数据上快速部署多个并行操作,使用户能够花费更少的时间进行调试,并将更多的时间用于开发。
摘要:There are many packages in Python which allow one to perform real-time processing on audio data. Unfortunately, due to the synchronous nature of the language, there lacks a framework which allows for distributed parallel processing of the data without requiring a large programming overhead and in which the data acquisition is not blocked by subsequent processing operations. This work improves on packages used for audio data collection with a light-weight backend and a simple interface that allows for distributed processing through a socket-based structure. This is intended for real-time audio machine learning and data processing in Python with a quick deployment of multiple parallel operations on the same data, allowing users to spend less time debugging and more time developing.

【6】 Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation  Guided by the Characteristic Dance Primitives
标题:洛奇:以特色舞蹈原始为导向的长舞世代由粗到细的传播网络
链接:https://arxiv.org/abs/2403.10518
作者:Ronghui Li,YuXiang Zhang,Yachao Zhang,Hongwen Zhang,Jie Guo,Yan Zhang,Yebin Liu,Xiu Li摘要:我们提出洛奇,一个网络能够产生非常长的舞蹈序列的条件下,给定的音乐。我们设计洛奇作为一个两阶段的粗到细的扩散架构,并提出了具有显着的表现力的特征舞蹈原语作为两个扩散模型之间的中间表示。第一阶段是全球扩散,重点是理解粗层次的音乐-舞蹈相关性和生产特征的舞蹈原语。第二阶段是局部扩散,在舞蹈原语和编舞规则的指导下,局部扩散生成详细的动作序列。此外,我们提出了一个脚优化块,以优化脚和地面之间的接触,增强运动的物理现实主义。我们的方法可以快速生成非常长的舞蹈序列,在全局编舞模式和局部运动质量和表现力之间取得平衡。大量的实验验证了我们的方法的有效性。
摘要:We propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations between two diffusion models. The first stage is global diffusion, which focuses on comprehending the coarse-level music-dance correlation and production characteristic dance primitives. In contrast, the second-stage is the local diffusion, which parallelly generates detailed motion sequences under the guidance of the dance primitives and choreographic rules. In addition, we propose a Foot Refine Block to optimize the contact between the feet and the ground, enhancing the physical realism of the motion. Our approach can parallelly generate dance sequences of extremely long length, striking a balance between global choreographic patterns and local motion quality and expressiveness. Extensive experiments validate the efficacy of our method.

【7】 MusicHiFi: Fast High-Fidelity Stereo Vocoding
标题:MusicHiFi:快速高保真立体声编码
链接:https://arxiv.org/abs/2403.10493
作者:Ge Zhu,Juan-Pablo Caceres,Zhiyao Duan,Nicholas J. Bryan
摘要:基于扩散的音频和音乐生成模型通常通过构造音频的图像表示(例如,MEL频谱图),然后使用相位重构模型或声码器将其转换成音频。然而,典型的声码器以较低的分辨率产生单声道音频(例如,16-24 kHz),这限制了它们的有效性。我们提出了MusicHiFi -一种高效的高保真立体声声码器。我们的方法采用了三个生成对抗网络(GANs)的级联,将低分辨率梅尔频谱图转换为音频,通过带宽扩展上采样为高分辨率音频,并上混为立体声音频。与以前的工作相比,我们提出了1)一个统一的基于GAN的生成器和FPGA架构以及级联每个阶段的训练过程,2)一个新的快速,接近下采样兼容的带宽扩展模块,以及3)一个新的快速下混兼容的单声道到立体声上混器,确保在输出中保留单声道内容。我们使用客观和主观的听力测试来评估我们的方法,并发现我们的方法与过去的工作相比,可以产生相当或更好的音频质量,更好的空间化控制和更快的推理速度。健全的例子是在https://MusicHiFi.github.io/web/。
摘要:Diffusion-based audio and music generation models commonly generate music by constructing an image representation of audio (e.g., a mel-spectrogram) and then converting it to audio using a phase reconstruction model or vocoder. Typical vocoders, however, produce monophonic audio at lower resolutions (e.g., 16-24 kHz), which limits their effectiveness. We propose MusicHiFi -- an efficient high-fidelity stereophonic vocoder. Our method employs a cascade of three generative adversarial networks (GANs) that convert low-resolution mel-spectrograms to audio, upsamples to high-resolution audio via bandwidth expansion, and upmixes to stereophonic audio. Compared to previous work, we propose 1) a unified GAN-based generator and discriminator architecture and training procedure for each stage of our cascade, 2) a new fast, near downsampling-compatible bandwidth extension module, and 3) a new fast downmix-compatible mono-to-stereo upmixer that ensures the preservation of monophonic content in the output. We evaluate our approach using both objective and subjective listening tests and find our approach yields comparable or better audio quality, better spatialization control, and significantly faster inference speed compared to past work. Sound examples are at https://MusicHiFi.github.io/web/.

【8】 Joint Multimodal Transformer for Dimensional Emotional Recognition in  the Wild
标题:用于野外维度情感识别的联合多模Transformer
链接:https://arxiv.org/abs/2403.10488
作者:Paul Waligora,Osama Zeeshan,Haseeb Aslam,Soufiane Belharbi,Alessandro Lameiras Koerich,Marco Pedersoli,Simon Bacon,Eric Granger
备注:5 pages, 1 figure
摘要:视频中的视听情感识别(ER)在单峰性能上具有巨大的潜力。它有效地利用了视觉和听觉模态之间的模态间和模态内依赖性。本文提出了一种新的视听情感识别系统,该系统采用基于键的交叉注意的联合多模态Transformer架构。该框架旨在利用视频中音频和视觉线索(面部表情和声音模式)的互补性,与仅依赖单一模态相比,具有更好的性能。所提出的模型利用单独的骨干捕捉内模态的时间依赖性在每个模态(音频和视频)。随后,联合多模态Transformer架构集成了各个模态嵌入,使模型能够有效地捕获模态间(音频和视觉之间)和模态内(每个模态内)的关系。在具有挑战性的Affwild2数据集上进行的广泛评估表明,所提出的模型在ER任务中显着优于基线和最先进的方法。
摘要:Audiovisual emotion recognition (ER) in videos has immense potential over unimodal performance. It effectively leverages the inter- and intra-modal dependencies between visual and auditory modalities. This work proposes a novel audio-visual emotion recognition system utilizing a joint multimodal transformer architecture with key-based cross-attention. This framework aims to exploit the complementary nature of audio and visual cues (facial expressions and vocal patterns) in videos, leading to superior performance compared to solely relying on a single modality. The proposed model leverages separate backbones for capturing intra-modal temporal dependencies within each modality (audio and visual). Subsequently, a joint multimodal transformer architecture integrates the individual modality embeddings, enabling the model to effectively capture inter-modal (between audio and visual) and intra-modal (within each modality) relationships. Extensive evaluations on the challenging Affwild2 dataset demonstrate that the proposed model significantly outperforms baseline and state-of-the-art methods in ER tasks.

【9】 BirdSet: A Multi-Task Benchmark for Classification in Avian Bioacoustics
标题:BirdSet:一个用于鸟类生物声学分类的多任务基准
链接:https://arxiv.org/abs/2403.10380
作者:Lukas Rauch,Raphael Schwinger,Moritz Wirth,René Heinrich,Jonas Lange,Stefan Kahl,Bernhard Sick,Sven Tomforde,Christoph Scholz
备注:Work in progress, to be submitted @DMLR next month
摘要:深度学习(DL)模型已经成为鸟类生物声学诊断环境健康和生物多样性的强大工具。然而,研究中的不一致性构成了阻碍这一领域取得进展的显著挑战。可靠的DL模型需要灵活地分析各种物种和环境中的鸟类叫声,以充分利用生物声学在经济高效的被动声学监测方案中的潜力。跨研究的数据碎片化和不透明性使一般模型性能的综合评价复杂化。为了克服这些挑战,我们提出了BirdSet基准,一个统一的框架,巩固研究工作的整体方法分类鸟类生物声学中的鸟类发声。BirdSet将开源的鸟类记录整合到一个精心策划的数据集中。这种统一的方法提供了对模型性能的深入理解,并确定了不同任务中的潜在缺点。通过建立当前模型的基线结果,BirdSet旨在促进可比性,指导后续数据收集,并增加新来者对鸟类生物声学的可访问性。
摘要:Deep learning (DL) models have emerged as a powerful tool in avian bioacoustics to diagnose environmental health and biodiversity. However, inconsistencies in research pose notable challenges hindering progress in this domain. Reliable DL models need to analyze bird calls flexibly across various species and environments to fully harness the potential of bioacoustics in a cost-effective passive acoustic monitoring scenario. Data fragmentation and opacity across studies complicate a comprehensive evaluation of general model performance. To overcome these challenges, we present the BirdSet benchmark, a unified framework consolidating research efforts with a holistic approach for classifying bird vocalizations in avian bioacoustics. BirdSet harmonizes open-source bird recordings into a curated dataset collection. This unified approach provides an in-depth understanding of model performance and identifies potential shortcomings across different tasks. By establishing baseline results of current models, BirdSet aims to facilitate comparability, guide subsequent data collection, and increase accessibility for newcomers to avian bioacoustics.

【10】 Multiscale Matching Driven by Cross-Modal Similarity Consistency for  Audio-Text Retrieval
标题:基于跨模态相似一致性驱动的多尺度匹配音频文本检索
链接:https://arxiv.org/abs/2403.10146
作者:Qian Wang,Jia-Chen Gu,Zhen-Hua Ling
备注:5 pages, accepted to ICASSP2024
摘要:音频文本检索(ATR),即从给定的音频片段(A2 T)中检索相关的字幕,反之亦然(T2 A),近年来引起了广泛的研究关注。现有的方法通常将来自每个模态的信息聚合到单个向量中进行匹配,但是这牺牲了局部细节并且难以捕获模态内和模态之间的复杂关系。此外,目前的ATR数据集缺乏全面的对齐信息,简单的二进制对比学习标签忽略了样本之间细粒度语义差异的测量。为了应对这些挑战,我们提出了一个新的ATR框架,全面捕捉多模态信息的匹配关系,从不同的角度和更细的粒度。具体来说,引入了一种细粒度的对齐方法,通过从局部到全局的多尺度过程来捕获细致的跨模态关系,从而实现更面向细节的匹配。此外,我们开创了跨模态相似性一致性的应用,利用模态内相似性关系作为软监督来促进更复杂的对齐。大量的实验验证了我们的方法的有效性,在AudioCaps数据集上至少有3.9%(T2 A)/ 6.9%(A2 T)R@1的显着裕度,在Clotho数据集上至少有2.9%(T2 A)/ 5.4%(A2 T)R@1的显着裕度。
摘要:Audio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single vector for matching, but this sacrifices local details and can hardly capture intricate relationships within and between modalities. Furthermore, current ATR datasets lack comprehensive alignment information, and simple binary contrastive learning labels overlook the measurement of fine-grained semantic differences between samples. To counter these challenges, we present a novel ATR framework that comprehensively captures the matching relationships of multimodal information from different perspectives and finer granularities. Specifically, a fine-grained alignment method is introduced, achieving a more detail-oriented matching through a multiscale process from local to global levels to capture meticulous cross-modal relationships. In addition, we pioneer the application of cross-modal similarity consistency, leveraging intra-modal similarity relationships as soft supervision to boost more intricate alignment. Extensive experiments validate the effectiveness of our approach, outperforming previous methods by significant margins of at least 3.9% (T2A) / 6.9% (A2T) R@1 on the AudioCaps dataset and 2.9% (T2A) / 5.4% (A2T) R@1 on the Clotho dataset.


【11】 MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate  Instrument Leakage
标题:MR—MT3:记忆保留多轨音乐转录以减轻仪器泄漏
链接:https://arxiv.org/abs/2403.10024
作者:Hao Hao Tan,Kin Wai Cheuk,Taemin Cho,Wei-Hsiang Liao,Yuki Mitsufuji
摘要:本文介绍了MT3模型的增强,这是一种基于SOTA令牌的多乐器自动音乐转录(AMT)模型。尽管有SOTA性能,但MT3存在仪器泄漏问题,其中跨不同仪器的传输是分散的。为了缓解这一问题,我们提出了MR-MT3,并提出了增强功能,包括内存保留机制,事先令牌采样和令牌洗牌。这些方法在Slakh 2100数据集上进行了评估,证明了改善的发作F1评分和减少的仪器泄漏。除了传统的多仪器转录F1分数外,还引入了仪器泄漏率和仪器检测F1分数等新指标,以更全面地评估转录质量。该研究还通过在ComMU和NSynth等单乐器单声道数据集上评估MT3来探索域过拟合问题。研究结果以及源代码将被共享,以促进未来旨在改进基于令牌的多仪器AMT模型的工作。
摘要:This paper presents enhancements to the MT3 model, a state-of-the-art (SOTA) token-based multi-instrument automatic music transcription (AMT) model. Despite SOTA performance, MT3 has the issue of instrument leakage, where transcriptions are fragmented across different instruments. To mitigate this, we propose MR-MT3, with enhancements including a memory retention mechanism, prior token sampling, and token shuffling are proposed. These methods are evaluated on the Slakh2100 dataset, demonstrating improved onset F1 scores and reduced instrument leakage. In addition to the conventional multi-instrument transcription F1 score, new metrics such as the instrument leakage ratio and the instrument detection F1 score are introduced for a more comprehensive assessment of transcription quality. The study also explores the issue of domain overfitting by evaluating MT3 on single-instrument monophonic datasets such as ComMU and NSynth. The findings, along with the source code, are shared to facilitate future work aimed at refining token-based multi-instrument AMT models.


【12】 SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification  of Spoken Numbers in Different Languages
标题:SpokeN—100:一个用于不同语言口语数分类的跨语言基准数据集
链接:https://arxiv.org/abs/2403.09753
作者:René Groh,Nina Goes,Andreas M. Kist
备注:Accepted as a full paper by the tinyML Research Symposium 2024
摘要:基准测试在评估和增强紧凑型深度学习模型的性能方面发挥着关键作用,这些模型是为在资源受限的设备(如微控制器)上执行而设计的。我们的研究引入了一个全新的、完全人工生成的、专为语音识别量身定制的基准测试数据集,这是微型深度学习领域的一个核心挑战。SpokeN-100由32个不同的说话者用四种不同的语言(英语、普通话、德语和法语)说出的从0到99的数字组成,产生了12,800个音频样本。我们确定听觉特征,并使用UMAP(均匀流形近似和投影降维)作为降维方法,以显示数据集的多样性和丰富性。为了突出数据集的用例,我们引入了两个基准任务:给定音频样本,对(i)使用的语言和/或(ii)口语数字进行分类。我们优化了最先进的深度神经网络,并执行了进化神经架构搜索,以找到针对32位ARM Cortex-M4 nRF 52840微控制器优化的微型架构。我们的结果代表了SpokeN-100实现的第一个基准数据。
摘要:Benchmarking plays a pivotal role in assessing and enhancing the performance of compact deep learning models designed for execution on resource-constrained devices, such as microcontrollers. Our study introduces a novel, entirely artificially generated benchmarking dataset tailored for speech recognition, representing a core challenge in the field of tiny deep learning. SpokeN-100 consists of spoken numbers from 0 to 99 spoken by 32 different speakers in four different languages, namely English, Mandarin, German and French, resulting in 12,800 audio samples. We determine auditory features and use UMAP (Uniform Manifold Approximation and Projection for Dimension Reduction) as a dimensionality reduction method to show the diversity and richness of the dataset. To highlight the use case of the dataset, we introduce two benchmark tasks: given an audio sample, classify (i) the used language and/or (ii) the spoken number. We optimized state-of-the-art deep neural networks and performed an evolutionary neural architecture search to find tiny architectures optimized for the 32-bit ARM Cortex-M4 nRF52840 microcontroller. Our results represent the first benchmark data achieved for SpokeN-100.
机器翻译由腾讯交互翻译提供,仅供参考