今日论文合集:cs.SD语音9篇,eess.AS音频处理9篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】DeepSRGM -- Sequence Classification and Ranking in Indian Classical  Music with Deep Learning
标题:DeepSRGM--基于深度学习的印度古典音乐序列分类与排名
链接:https://arxiv.org/abs/2402.10168
作者:Sathwik Tejaswi Madhusudhan,Girish Chowdhary
摘要:印度古典音乐(ICM)的一个重要方面是Raga,它是作曲和即兴创作的旋律框架。Raga识别是ICM中一个重要的音乐信息检索任务,因为它可以帮助许多下游应用,从音乐推荐到组织庞大的音乐收藏。在这项工作中,我们提出了一种基于深度学习的Raga识别方法。我们的方法采用有效的预处理和学习时间序列的音乐数据使用长短期记忆的递归神经网络(LSTM-RNN)。我们在从原始音频采样的较小序列上对网络进行训练和测试,而最终的推理是在整个音频上执行的。我们的方法在Comp Music Carnatic数据集及其10个Raga子集上的推理过程中分别实现了88.1%和97%的准确率,使其成为Raga识别任务的最新技术。我们的方法还使序列排名,这有助于我们检索旋律模式从给定的音乐数据库,密切相关的查询序列。
摘要:A vital aspect of Indian Classical Music (ICM) is Raga, which serves as a melodic framework for compositions and improvisations alike. Raga Recognition is an important music information retrieval task in ICM as it can aid numerous downstream applications ranging from music recommendations to organizing huge music collections. In this work, we propose a deep learning based approach to Raga recognition. Our approach employs efficient pre possessing and learns temporal sequences in music data using Long Short Term Memory based Recurrent Neural Networks (LSTM-RNN). We train and test the network on smaller sequences sampled from the original audio while the final inference is performed on the audio as a whole. Our method achieves an accuracy of 88.1% and 97 % during inference on the Comp Music Carnatic dataset and its 10 Raga subset respectively making it the state-of-the-art for the Raga recognition task. Our approach also enables sequence ranking which aids us in retrieving melodic patterns from a given music data base that are closely related to the presented query sequence.

【2】 Tuning In: Analysis of Audio Classifier Performance in Clinical Settings  with Limited Data
标题:调谐:有限数据下临床环境下音频分类器的性能分析
链接:https://arxiv.org/abs/2402.10100
作者:Hamza Mahdi,Eptehal Nashnoush,Rami Saab,Arjun Balachandar,Rishit Dagli,Lucas X. Perri,Houman Khosravani
摘要:这项研究评估了在临床环境中用于音频分类的深度学习模型,这些模型具有反映真实世界前瞻性数据收集的小数据集的约束。我们分析CNN,包括DenseNet和ConvNeXt,以及Transformer模型,如ViT,SWIN和AST,并将它们与预先训练的音频模型,如YAMNet和VGGish进行比较。我们的方法强调了在对特定临床数据进行微调之前对大型数据集进行预训练的好处。我们前瞻性地收集了两个来自中风患者的首个患者音频数据集。我们研究了各种预处理技术,发现RGB和灰度谱图变换根据它们从预训练中学习到的先验知识对模型性能产生不同的影响。我们的研究结果表明,CNN可以在小数据集上下文中匹配或超过Transformer模型,DenseNet-Contrastive和AST模型表现出显着的性能。这项研究强调了通过模型选择,预训练和预处理在声音分类中的增量边际增益的重要性;这为依赖于音频分类的临床诊断提供了有价值的见解。
摘要:This study assesses deep learning models for audio classification in a clinical setting with the constraint of small datasets reflecting real-world prospective data collection. We analyze CNNs, including DenseNet and ConvNeXt, alongside transformer models like ViT, SWIN, and AST, and compare them against pre-trained audio models such as YAMNet and VGGish. Our method highlights the benefits of pre-training on large datasets before fine-tuning on specific clinical data. We prospectively collected two first-of-their-kind patient audio datasets from stroke patients. We investigated various preprocessing techniques, finding that RGB and grayscale spectrogram transformations affect model performance differently based on the priors they learn from pre-training. Our findings indicate CNNs can match or exceed transformer models in small dataset contexts, with DenseNet-Contrastive and AST models showing notable performance. This study highlights the significance of incremental marginal gains through model selection, pre-training, and preprocessing in sound classification; this offers valuable insights for clinical diagnostics that rely on audio classification.

【3】 Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
标题:基于DDPM倒置的Zero-Shot无监督文本音频编辑
链接:https://arxiv.org/abs/2402.10009
作者:Hila Manor,Tomer Michaeli
备注:Examples and code available in this https URL
摘要:使用大的预训练模型以zero-shot方式编辑信号最近在图像领域中取得了快速发展。然而,这一浪潮尚未到达音频领域。在本文中,我们探讨了两个zero-shot编辑技术的音频信号,使用预先训练的扩散模型的DDPM反演。第一个是从图像域中采用的,允许基于文本的编辑。第二,是一种新的方法,发现语义有意义的编辑方向没有监督。当应用于音乐信号时,这种方法暴露了一系列音乐上有趣的修改,从控制特定乐器的参与到旋律的即兴创作。示例可以在我们的示例页面https://hilamanor.github.io/AudioEditing/上找到,代码可以在https://github.com/hilamanor/AudioEditing/上找到。
摘要:Editing signals using large pre-trained models, in a zero-shot manner, has recently seen rapid advancements in the image domain. However, this wave has yet to reach the audio domain. In this paper, we explore two zero-shot editing techniques for audio signals, which use DDPM inversion on pre-trained diffusion models. The first, adopted from the image domain, allows text-based editing. The second, is a novel approach for discovering semantically meaningful editing directions without supervision. When applied to music signals, this method exposes a range of musically interesting modifications, from controlling the participation of specific instruments to improvisations on the melody. Samples can be found on our examples page in https://hilamanor.github.io/AudioEditing/ and code can be found in https://github.com/hilamanor/AudioEditing/ .

【4】 ML-ASPA: A Contemplation of Machine Learning-based Acoustic Signal  Processing Analysis for Sounds, & Strains Emerging Technology
标题:ML-ASPA:一种基于机器学习的声学信号处理分析新兴技术
链接:https://arxiv.org/abs/2402.10005
作者:Ratul Ali,Aktarul Islam,Md. Shohel Rana,Saila Nasrin,Sohel Afzal Shajol,Professor Dr. A. H. M. Saifullah Sadi
备注:7 pages, 5 figures, Article
摘要:声学数据是推进跨生物学、通信、海洋和地球科学等不同学科的科学和工程理解的基本基石。这项调查仔细探索了声学领域的最新进展和变革潜力,特别关注机器学习(ML)和深度学习。ML包含大量的统计技术,对于自主识别和利用数据中的模式是必不可少的。与传统的声学和信号处理相比,ML采用数据驱动的方法,在给定大量训练数据的情况下,揭示了特征与所需标签或动作之间以及特征本身之间的复杂关系。将ML应用于大量训练数据集有助于发现阐明复杂声学现象(如人类语音和混响)的模型。ML在声学中的动态演变产生了令人信服的结果,并为未来带来了巨大的希望。电子听诊器和模拟记录和数据记录设备的出现已经将声学信号处理概念的应用扩展到肠鸣音的分析。本文批判性地回顾了现有的文献中关于肠音分析的声学信号处理,概述了基本方法和适用的机器学习原理。它记录了信号处理技术的历史进展,这些技术促进了从肠鸣音中提取有价值的信息,强调了降噪,分割,信号增强,特征提取,声音定位和机器学习技术的进步。
摘要:Acoustic data serves as a fundamental cornerstone in advancing scientific and engineering understanding across diverse disciplines, spanning biology, communications, and ocean and Earth science. This inquiry meticulously explores recent advancements and transformative potential within the domain of acoustics, specifically focusing on machine learning (ML) and deep learning. ML, comprising an extensive array of statistical techniques, proves indispensable for autonomously discerning and leveraging patterns within data. In contrast to traditional acoustics and signal processing, ML adopts a data-driven approach, unveiling intricate relationships between features and desired labels or actions, as well as among features themselves, given ample training data. The application of ML to expansive sets of training data facilitates the discovery of models elucidating complex acoustic phenomena such as human speech and reverberation. The dynamic evolution of ML in acoustics yields compelling results and holds substantial promise for the future. The advent of electronic stethoscopes and analogous recording and data logging devices has expanded the application of acoustic signal processing concepts to the analysis of bowel sounds. This paper critically reviews existing literature on acoustic signal processing for bowel sound analysis, outlining fundamental approaches and applicable machine learning principles. It chronicles historical progress in signal processing techniques that have facilitated the extraction of valuable information from bowel sounds, emphasizing advancements in noise reduction, segmentation, signal enhancement, feature extraction, sound localization, and machine learning techniques...

【5】 MuChin: A Chinese Colloquial Description Benchmark for Evaluating  Language Models in the Field of Music
标题:MuChin:一个用于评估音乐领域语言模型的汉语口语描述基准
链接:https://arxiv.org/abs/2402.09871
作者:Zihao Wang,Shuyu Li,Tao Zhang,Qi Wang,Pengfei Yu,Jinyang Luo,Yan Liu,Ming Xi,Kejun Zhang
摘要:快速发展的多模态大语言模型(LLM)迫切需要新的基准来统一评估其在理解和文本描述音乐方面的性能。然而,由于音乐信息检索(MIR)算法和人类理解之间的语义差距,专业人士和公众之间的差异,以及注释的低精度,现有的音乐描述数据集不能作为基准。为此,我们提出了MuChin,这是第一个用汉语口语描述的开源音乐描述基准,旨在评估多模态LLM在理解和描述音乐方面的性能。我们建立了彩虫音乐注释平台(CaiMAP),采用创新的多人多阶段保证方法,招募业余和专业人士,以确保注释的准确性和与流行语义的一致性。利用这种方法,我们建立了一个具有多维,高精度音乐注释的数据集,CaiMD,并精心挑选了1,000个高质量的条目作为MuChin的测试集。基于MuChin,我们分析了专业人士和业余爱好者在音乐描述方面的差异,并实证证明了注释数据用于微调LLM的有效性。最后,我们聘请MuChin评估现有的音乐理解模型提供口语化音乐描述的能力。所有与基准相关的数据和评分代码都是开源的。
摘要:The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between Music Information Retrieval (MIR) algorithms and human understanding, discrepancies between professionals and the public, and low precision of annotations, existing music description datasets cannot serve as benchmarks. To this end, we present MuChin, the first open-source music description benchmark in Chinese colloquial language, designed to evaluate the performance of multimodal LLMs in understanding and describing music. We established the Caichong Music Annotation Platform (CaiMAP) that employs an innovative multi-person, multi-stage assurance method, and recruited both amateurs and professionals to ensure the precision of annotations and alignment with popular semantics. Utilizing this method, we built a dataset with multi-dimensional, high-precision music annotations, the Caichong Music Dataset (CaiMD), and carefully selected 1,000 high-quality entries to serve as the test set for MuChin. Based on MuChin, we analyzed the discrepancies between professionals and amateurs in terms of music description, and empirically demonstrated the effectiveness of annotated data for fine-tuning LLMs. Ultimately, we employed MuChin to evaluate existing music understanding models on their ability to provide colloquial descriptions of music. All data related to the benchmark and the code for scoring have been open-sourced.

【6】 A cross-talk robust multichannel VAD model for multiparty agent  interactions trained using synthetic re-recordings
标题:使用合成重录训练多方代理交互的串扰稳健多通道VAD模型
链接:https://arxiv.org/abs/2402.09797
作者:Hyewon Han,Naveen Kumar
备注:Accepted for presentation at the Hands-free Speech Communication and Microphone Arrays (HSCMA 2024)
摘要:在这项工作中,我们提出了一种新的串扰拒绝框架的多通道多说话者设置的现场多方互动节目。我们的远场音频设置要求在现场互动期间免提,并在同一空间内包括四个带有定向麦克风的相邻扬声器。这样的设置通常会在通道之间引入严重的串扰,导致自动语音识别(ASR)和自然语言理解(NLU)性能降低。为了解决这个问题,我们提出了语音活动检测(VAD)模型的所有说话者使用多通道信息,然后用于过滤音频的下游任务。我们采用了一种合成训练数据生成方法,通过回放和重新记录这种情况下,模拟具有挑战性的语音重叠条件。我们训练我们的模型上的合成数据,并证明我们的方法优于单通道VAD模型和基于能量的多通道VAD算法在各种声学环境。除了VAD结果,我们还提供了多方ASR评估结果,以突出使用我们的VAD模型通过显着减少插入错误来过滤下游任务中的音频的影响。
摘要:In this work, we propose a novel cross-talk rejection framework for a multi-channel multi-talker setup for a live multiparty interactive show. Our far-field audio setup is required to be hands-free during live interaction and comprises four adjacent talkers with directional microphones in the same space. Such setups often introduce heavy cross-talk between channels, resulting in reduced automatic speech recognition (ASR) and natural language understanding (NLU) performance. To address this problem, we propose voice activity detection (VAD) model for all talkers using multichannel information, which is then used to filter audio for downstream tasks. We adopt a synthetic training data generation approach through playback and re-recording for such scenarios, simulating challenging speech overlap conditions. We train our models on this synthetic data and demonstrate that our approach outperforms single-channel VAD models and energy-based multi-channel VAD algorithm in various acoustic environments. In addition to VAD results, we also present multiparty ASR evaluation results to highlight the impact of using our VAD model for filtering audio in downstream tasks by significantly reducing the insertion error.


【7】 Domain Adaptation for Contrastive Audio-Language Models
标题:对比听觉语言模型的领域适应
链接:https://arxiv.org/abs/2402.09585
作者:Soham Deshmukh,Rita Singh,Bhiksha Raj
摘要:音频语言模型(ALM)的目标是通过在测试时提供zero-shot能力成为通用音频模型。通过为每个域使用合适的文本提示,ALM的zero-shot性能得到了提高。文本提示通常是通过特定过程手工制作的,这会导致ALM泛化和分发外性能的下降。现有的提高领域性能的方法,如Few-Shot学习或微调,需要访问带注释的数据和迭代训练。因此,我们提出了一个测试时域自适应方法的ALM,不需要访问注释。我们的方法通过在测试音频的增强视图中执行一致性来学习域向量。我们广泛地评估了我们的方法在12个跨域的下游任务。仅举一个例子,我们的域自适应方法导致平均zero-shot性能提高3.2%(最大8.4%)。自适应后,模型仍然保留了自适应层模型的泛化特性。
摘要:Audio-Language Models (ALM) aim to be general-purpose audio models by providing zero-shot capabilities at test time. The zero-shot performance of ALM improves by using suitable text prompts for each domain. The text prompts are usually hand-crafted through an ad-hoc process and lead to a drop in ALM generalization and out-of-distribution performance. Existing approaches to improve domain performance, like few-shot learning or fine-tuning, require access to annotated data and iterations of training. Therefore, we propose a test-time domain adaptation method for ALMs that does not require access to annotations. Our method learns a domain vector by enforcing consistency across augmented views of the testing audio. We extensively evaluate our approach on 12 downstream tasks across domains. With just one example, our domain adaptation method leads to 3.2% (max 8.4%) average zero-shot performance improvement. After adaptation, the model still retains the generalization property of ALMs.

【8】 Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation  and Editing via Content-based Controls
标题:安排,修补和优化:通过基于内容的控件可操纵的长期音乐音频生成和编辑
链接:https://arxiv.org/abs/2402.09508
作者:Liwei Lin,Gus Xia,Yixiao Zhang,Junyan Jiang
摘要:可控的音乐生成在人类-AI音乐共同创作中起着至关重要的作用。虽然大型语言模型(LLM)在生成高质量音乐方面表现出了希望,但它们对自回归生成的关注限制了它们在音乐编辑任务中的实用性。为了弥补这一差距,我们引入了一种新的参数有效的微调(PEFT)方法。这种方法使自回归语言模型能够无缝地解决音乐修复任务。此外,我们的PEFT方法集成了帧级的基于内容的控制,促进轨道条件的音乐细化和评分条件的音乐安排。我们应用这种方法来微调MusicGen,一个领先的自回归音乐生成模型。我们的实验在多个音乐编辑任务中展示了有希望的结果,为未来的AI驱动的音乐编辑工具提供了更灵活的控制。一个演示页面\footnote{\url{https://kikyo-16.github.io/AIR/}.}展示我们的工作和源代码\footnote{\url{https://github.com/Kikyo-16/airgen}.}可以在网上找到。
摘要:Controllable music generation plays a vital role in human-AI music co-creation. While Large Language Models (LLMs) have shown promise in generating high-quality music, their focus on autoregressive generation limits their utility in music editing tasks. To bridge this gap, we introduce a novel Parameter-Efficient Fine-Tuning (PEFT) method. This approach enables autoregressive language models to seamlessly address music inpainting tasks. Additionally, our PEFT method integrates frame-level content-based controls, facilitating track-conditioned music refinement and score-conditioned music arrangement. We apply this method to fine-tune MusicGen, a leading autoregressive music generation model. Our experiments demonstrate promising results across multiple music editing tasks, offering more flexible controls for future AI-driven music editing tools. A demo page\footnote{\url{https://kikyo-16.github.io/AIR/}.} showcasing our work and source codes\footnote{\url{https://github.com/Kikyo-16/airgen}.} are available online.

【9】 Diffusion Models for Audio Restoration
标题:用于音频恢复的扩散模型
链接:https://arxiv.org/abs/2402.09821
作者:Jean-Marie Lemercier,Julius Richter,Simon Welker,Eloi Moliner,Vesa Välimäki,Timo Gerkmann
备注:Full paper invited to the IEEE Signal Processing Magazine Special Issue "Model-based and Data-Driven Audio Signal Processing"
摘要:随着音频播放设备和快速数据传输的发展,娱乐和通信对高音质的需求正在上升。在追求更好的音质的过程中,来自录音端或由不完美的传输管道引起的失真和干扰带来了挑战。为了解决这个问题,音频恢复方法旨在从损坏的输入数据中恢复干净的声音信号。我们在这里提出了基于扩散模型的音频恢复算法,重点是语音增强和音乐恢复任务。传统的方法,通常是基于手工制作的规则和统计分析,塑造了我们对音频信号的理解。在过去的几十年里,数据驱动的方法已经发生了显著的转变,这些方法利用了深度神经网络(DNN)的建模能力。深度生成模型,以及其中的扩散模型,已经成为学习复杂数据分布的强大技术。然而,仅仅依赖基于DNN的学习方法会降低可解释性,特别是在使用端到端模型时。尽管如此,与基于统计模型的框架相比,数据驱动的方法具有更大的灵活性,因为基于统计模型的框架的性能取决于可能难以保证的分布和统计假设。在这里,我们的目标是表明,扩散模型可以结合两个世界的最好的,并提供机会,设计音频恢复算法具有良好的可解释性和显着的性能方面的声音质量。
摘要:With the development of audio playback devices and fast data transmission, the demand for high sound quality is rising, for both entertainment and communications. In this quest for better sound quality, challenges emerge from distortions and interferences originating at the recording side or caused by an imperfect transmission pipeline. To address this problem, audio restoration methods aim to recover clean sound signals from the corrupted input data. We present here audio restoration algorithms based on diffusion models, with a focus on speech enhancement and music restoration tasks. Traditional approaches, often grounded in handcrafted rules and statistical heuristics, have shaped our understanding of audio signals. In the past decades, there has been a notable shift towards data-driven methods that exploit the modeling capabilities of deep neural networks (DNNs). Deep generative models, and among them diffusion models, have emerged as powerful techniques for learning complex data distributions. However, relying solely on DNN-based learning approaches carries the risk of reducing interpretability, particularly when employing end-to-end models. Nonetheless, data-driven approaches allow more flexibility in comparison to statistical model-based frameworks whose performance depends on distributional and statistical assumptions that can be difficult to guarantee. Here, we aim to show that diffusion models can combine the best of both worlds and offer the opportunity to design audio restoration algorithms with a good degree of interpretability and a remarkable performance in terms of sound quality.


eess.AS音频处理
【1】 Diffusion Models for Audio Restoration
标题:用于音频恢复的扩散模型
链接:https://arxiv.org/abs/2402.09821
作者:Jean-Marie Lemercier,Julius Richter,Simon Welker,Eloi Moliner,Vesa Välimäki,Timo Gerkmann
备注:Full paper invited to the IEEE Signal Processing Magazine Special Issue "Model-based and Data-Driven Audio Signal Processing"
摘要:随着音频播放设备和快速数据传输的发展,娱乐和通信对高音质的需求正在上升。在追求更好的音质的过程中,来自录音端或由不完美的传输管道引起的失真和干扰带来了挑战。为了解决这个问题,音频恢复方法旨在从损坏的输入数据中恢复干净的声音信号。我们在这里提出了基于扩散模型的音频恢复算法,重点是语音增强和音乐恢复任务。传统的方法,通常是基于手工制作的规则和统计分析,塑造了我们对音频信号的理解。在过去的几十年里,数据驱动的方法已经发生了显著的转变,这些方法利用了深度神经网络(DNN)的建模能力。深度生成模型,以及其中的扩散模型,已经成为学习复杂数据分布的强大技术。然而,仅仅依赖基于DNN的学习方法会降低可解释性,特别是在使用端到端模型时。尽管如此,与基于统计模型的框架相比,数据驱动的方法具有更大的灵活性,因为基于统计模型的框架的性能取决于可能难以保证的分布和统计假设。在这里,我们的目标是表明,扩散模型可以结合两个世界的最好的,并提供机会,设计音频恢复算法具有良好的可解释性和显着的性能方面的声音质量。
摘要:With the development of audio playback devices and fast data transmission, the demand for high sound quality is rising, for both entertainment and communications. In this quest for better sound quality, challenges emerge from distortions and interferences originating at the recording side or caused by an imperfect transmission pipeline. To address this problem, audio restoration methods aim to recover clean sound signals from the corrupted input data. We present here audio restoration algorithms based on diffusion models, with a focus on speech enhancement and music restoration tasks. Traditional approaches, often grounded in handcrafted rules and statistical heuristics, have shaped our understanding of audio signals. In the past decades, there has been a notable shift towards data-driven methods that exploit the modeling capabilities of deep neural networks (DNNs). Deep generative models, and among them diffusion models, have emerged as powerful techniques for learning complex data distributions. However, relying solely on DNN-based learning approaches carries the risk of reducing interpretability, particularly when employing end-to-end models. Nonetheless, data-driven approaches allow more flexibility in comparison to statistical model-based frameworks whose performance depends on distributional and statistical assumptions that can be difficult to guarantee. Here, we aim to show that diffusion models can combine the best of both worlds and offer the opportunity to design audio restoration algorithms with a good degree of interpretability and a remarkable performance in terms of sound quality.


【2】 Implementation of the Multichannel Filtered Reference Least Mean Square  (McFxLMS) Algorithm with an Arbitrary Number of Channels by Using MATLAB
标题:任意通道数多通道滤波参考最小均方(McFxLMS)算法的MATLAB实现
链接:https://arxiv.org/abs/2402.09449
作者:Boxiang Wang
摘要:多通道滤波参考最小均方(McFxLMS)算法广泛应用于自适应多通道有源噪声控制(MCANC)应用中。作为一个关键的和高计算效率的自适应关键算法,它也通常作为一个基准的比较研究的新算法提出的同行和研究人员。然而,到目前为止,很少有FxLMS算法的开源代码,特别是对于大计数通道。因此,本工作为McFxLMS算法提供了一个MATLAB代码,可用于任意通道数的系统。代码可以在GitHub和Mathworks上找到。
摘要:Multichannel filtered reference least mean square (McFxLMS) algorithms are widely utilized in adaptive multichannel active noise control (MCANC) applications. As a critical and high-computationally efficient adaptive critical algorithm, it also typically works as a benchmark for comparative studies of the new algorithms proposed by peers and researchers. However, up to now, there are few open-source codes for the FxLMS algorithm, especially for large-count channels. Therefore, this work provides a MATLAB code for the McFxLMS algorithm, which can be used for the arbitrary number of channels system. The code is available on GitHub and Mathworks.


【3】 DeepSRGM -- Sequence Classification and Ranking in Indian Classical  Music with Deep Learning
标题:DeepSRGM--基于深度学习的印度古典音乐序列分类与排名
链接:https://arxiv.org/abs/2402.10168
作者:Sathwik Tejaswi Madhusudhan,Girish Chowdhary
摘要:印度古典音乐(ICM)的一个重要方面是Raga,它是作曲和即兴创作的旋律框架。Raga识别是ICM中一个重要的音乐信息检索任务,因为它可以帮助许多下游应用,从音乐推荐到组织庞大的音乐收藏。在这项工作中,我们提出了一种基于深度学习的Raga识别方法。我们的方法采用有效的预处理和学习时间序列的音乐数据使用长短期记忆的递归神经网络(LSTM-RNN)。我们在从原始音频采样的较小序列上对网络进行训练和测试,而最终的推理是在整个音频上执行的。我们的方法在Comp Music Carnatic数据集及其10个Raga子集上的推理过程中分别实现了88.1%和97%的准确率,使其成为Raga识别任务的最新技术。我们的方法还使序列排名,这有助于我们检索旋律模式从给定的音乐数据库,密切相关的查询序列。
摘要:A vital aspect of Indian Classical Music (ICM) is Raga, which serves as a melodic framework for compositions and improvisations alike. Raga Recognition is an important music information retrieval task in ICM as it can aid numerous downstream applications ranging from music recommendations to organizing huge music collections. In this work, we propose a deep learning based approach to Raga recognition. Our approach employs efficient pre possessing and learns temporal sequences in music data using Long Short Term Memory based Recurrent Neural Networks (LSTM-RNN). We train and test the network on smaller sequences sampled from the original audio while the final inference is performed on the audio as a whole. Our method achieves an accuracy of 88.1% and 97 % during inference on the Comp Music Carnatic dataset and its 10 Raga subset respectively making it the state-of-the-art for the Raga recognition task. Our approach also enables sequence ranking which aids us in retrieving melodic patterns from a given music data base that are closely related to the presented query sequence.

【4】 Tuning In: Analysis of Audio Classifier Performance in Clinical Settings  with Limited Data
标题:调谐:有限数据下临床环境下音频分类器的性能分析
链接:https://arxiv.org/abs/2402.10100
作者:Hamza Mahdi,Eptehal Nashnoush,Rami Saab,Arjun Balachandar,Rishit Dagli,Lucas X. Perri,Houman Khosravani
摘要:这项研究评估了在临床环境中用于音频分类的深度学习模型,这些模型具有反映真实世界前瞻性数据收集的小数据集的约束。我们分析CNN,包括DenseNet和ConvNeXt,以及Transformer模型,如ViT,SWIN和AST,并将它们与预先训练的音频模型,如YAMNet和VGGish进行比较。我们的方法强调了在对特定临床数据进行微调之前对大型数据集进行预训练的好处。我们前瞻性地收集了两个来自中风患者的首个患者音频数据集。我们研究了各种预处理技术,发现RGB和灰度谱图变换根据它们从预训练中学习到的先验知识对模型性能产生不同的影响。我们的研究结果表明,CNN可以在小数据集上下文中匹配或超过Transformer模型,DenseNet-Contrastive和AST模型表现出显着的性能。这项研究强调了通过模型选择,预训练和预处理在声音分类中的增量边际增益的重要性;这为依赖于音频分类的临床诊断提供了有价值的见解。
摘要:This study assesses deep learning models for audio classification in a clinical setting with the constraint of small datasets reflecting real-world prospective data collection. We analyze CNNs, including DenseNet and ConvNeXt, alongside transformer models like ViT, SWIN, and AST, and compare them against pre-trained audio models such as YAMNet and VGGish. Our method highlights the benefits of pre-training on large datasets before fine-tuning on specific clinical data. We prospectively collected two first-of-their-kind patient audio datasets from stroke patients. We investigated various preprocessing techniques, finding that RGB and grayscale spectrogram transformations affect model performance differently based on the priors they learn from pre-training. Our findings indicate CNNs can match or exceed transformer models in small dataset contexts, with DenseNet-Contrastive and AST models showing notable performance. This study highlights the significance of incremental marginal gains through model selection, pre-training, and preprocessing in sound classification; this offers valuable insights for clinical diagnostics that rely on audio classification.

【5】 Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
标题:基于DDPM倒置的Zero-Shot无监督文本音频编辑
链接:https://arxiv.org/abs/2402.10009
作者:Hila Manor,Tomer Michaeli
备注:Examples and code available in this https URL
摘要:使用大的预训练模型以zero-shot方式编辑信号最近在图像领域中取得了快速发展。然而,这一浪潮尚未到达音频领域。在本文中,我们探讨了两个zero-shot编辑技术的音频信号,使用预先训练的扩散模型的DDPM反演。第一个是从图像域中采用的,允许基于文本的编辑。第二,是一种新的方法,发现语义有意义的编辑方向没有监督。当应用于音乐信号时,这种方法暴露了一系列音乐上有趣的修改,从控制特定乐器的参与到旋律的即兴创作。示例可以在我们的示例页面https://hilamanor.github.io/AudioEditing/上找到,代码可以在https://github.com/hilamanor/AudioEditing/上找到。
摘要:Editing signals using large pre-trained models, in a zero-shot manner, has recently seen rapid advancements in the image domain. However, this wave has yet to reach the audio domain. In this paper, we explore two zero-shot editing techniques for audio signals, which use DDPM inversion on pre-trained diffusion models. The first, adopted from the image domain, allows text-based editing. The second, is a novel approach for discovering semantically meaningful editing directions without supervision. When applied to music signals, this method exposes a range of musically interesting modifications, from controlling the participation of specific instruments to improvisations on the melody. Samples can be found on our examples page in https://hilamanor.github.io/AudioEditing/ and code can be found in https://github.com/hilamanor/AudioEditing/ .

【6】 ML-ASPA: A Contemplation of Machine Learning-based Acoustic Signal  Processing Analysis for Sounds, & Strains Emerging Technology
标题:ML-ASPA:一种基于机器学习的声学信号处理分析新兴技术
链接:https://arxiv.org/abs/2402.10005
作者:Ratul Ali,Aktarul Islam,Md. Shohel Rana,Saila Nasrin,Sohel Afzal Shajol,Professor Dr. A. H. M. Saifullah Sadi
备注:7 pages, 5 figures, Article
摘要:声学数据是推进跨生物学、通信、海洋和地球科学等不同学科的科学和工程理解的基本基石。这项调查仔细探索了声学领域的最新进展和变革潜力,特别关注机器学习(ML)和深度学习。ML包含大量的统计技术,对于自主识别和利用数据中的模式是必不可少的。与传统的声学和信号处理相比,ML采用数据驱动的方法,在给定大量训练数据的情况下,揭示了特征与所需标签或动作之间以及特征本身之间的复杂关系。将ML应用于大量训练数据集有助于发现阐明复杂声学现象(如人类语音和混响)的模型。ML在声学中的动态演变产生了令人信服的结果,并为未来带来了巨大的希望。电子听诊器和模拟记录和数据记录设备的出现已经将声学信号处理概念的应用扩展到肠鸣音的分析。本文批判性地回顾了现有的文献中关于肠音分析的声学信号处理,概述了基本方法和适用的机器学习原理。它记录了信号处理技术的历史进展,这些技术促进了从肠鸣音中提取有价值的信息,强调了降噪,分割,信号增强,特征提取,声音定位和机器学习技术的进步。
摘要:Acoustic data serves as a fundamental cornerstone in advancing scientific and engineering understanding across diverse disciplines, spanning biology, communications, and ocean and Earth science. This inquiry meticulously explores recent advancements and transformative potential within the domain of acoustics, specifically focusing on machine learning (ML) and deep learning. ML, comprising an extensive array of statistical techniques, proves indispensable for autonomously discerning and leveraging patterns within data. In contrast to traditional acoustics and signal processing, ML adopts a data-driven approach, unveiling intricate relationships between features and desired labels or actions, as well as among features themselves, given ample training data. The application of ML to expansive sets of training data facilitates the discovery of models elucidating complex acoustic phenomena such as human speech and reverberation. The dynamic evolution of ML in acoustics yields compelling results and holds substantial promise for the future. The advent of electronic stethoscopes and analogous recording and data logging devices has expanded the application of acoustic signal processing concepts to the analysis of bowel sounds. This paper critically reviews existing literature on acoustic signal processing for bowel sound analysis, outlining fundamental approaches and applicable machine learning principles. It chronicles historical progress in signal processing techniques that have facilitated the extraction of valuable information from bowel sounds, emphasizing advancements in noise reduction, segmentation, signal enhancement, feature extraction, sound localization, and machine learning techniques...


【7】 MuChin: A Chinese Colloquial Description Benchmark for Evaluating  Language Models in the Field of Music
标题:MuChin:一个用于评估音乐领域语言模型的汉语口语描述基准
链接:https://arxiv.org/abs/2402.09871
作者:Zihao Wang,Shuyu Li,Tao Zhang,Qi Wang,Pengfei Yu,Jinyang Luo,Yan Liu,Ming Xi,Kejun Zhang摘要:快速发展的多模态大语言模型(LLM)迫切需要新的基准来统一评估其在理解和文本描述音乐方面的性能。然而,由于音乐信息检索(MIR)算法和人类理解之间的语义差距,专业人士和公众之间的差异,以及注释的低精度,现有的音乐描述数据集不能作为基准。为此,我们提出了MuChin,这是第一个用汉语口语描述的开源音乐描述基准,旨在评估多模态LLM在理解和描述音乐方面的性能。我们建立了彩虫音乐注释平台(CaiMAP),采用创新的多人多阶段保证方法,招募业余和专业人士,以确保注释的准确性和与流行语义的一致性。利用这种方法,我们建立了一个具有多维,高精度音乐注释的数据集,CaiMD,并精心挑选了1,000个高质量的条目作为MuChin的测试集。基于MuChin,我们分析了专业人士和业余爱好者在音乐描述方面的差异,并实证证明了注释数据用于微调LLM的有效性。最后,我们聘请MuChin评估现有的音乐理解模型提供口语化音乐描述的能力。所有与基准相关的数据和评分代码都是开源的。
摘要:The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between Music Information Retrieval (MIR) algorithms and human understanding, discrepancies between professionals and the public, and low precision of annotations, existing music description datasets cannot serve as benchmarks. To this end, we present MuChin, the first open-source music description benchmark in Chinese colloquial language, designed to evaluate the performance of multimodal LLMs in understanding and describing music. We established the Caichong Music Annotation Platform (CaiMAP) that employs an innovative multi-person, multi-stage assurance method, and recruited both amateurs and professionals to ensure the precision of annotations and alignment with popular semantics. Utilizing this method, we built a dataset with multi-dimensional, high-precision music annotations, the Caichong Music Dataset (CaiMD), and carefully selected 1,000 high-quality entries to serve as the test set for MuChin. Based on MuChin, we analyzed the discrepancies between professionals and amateurs in terms of music description, and empirically demonstrated the effectiveness of annotated data for fine-tuning LLMs. Ultimately, we employed MuChin to evaluate existing music understanding models on their ability to provide colloquial descriptions of music. All data related to the benchmark and the code for scoring have been open-sourced.

【8】 A cross-talk robust multichannel VAD model for multiparty agent  interactions trained using synthetic re-recordings
标题:使用合成重新记录训练的多方代理交互的串扰鲁棒多通道VAD模型
链接:https://arxiv.org/abs/2402.09797
作者:Hyewon Han,Naveen Kumar
备注:Accepted for presentation at the Hands-free Speech Communication and Microphone Arrays (HSCMA 2024)
摘要:在这项工作中,我们提出了一种新的串扰拒绝框架的多通道多说话者设置的现场多方互动节目。我们的远场音频设置要求在现场互动期间免提,并在同一空间内包括四个带有定向麦克风的相邻扬声器。这样的设置通常会在通道之间引入严重的串扰,导致自动语音识别(ASR)和自然语言理解(NLU)性能降低。为了解决这个问题,我们提出了语音活动检测(VAD)模型的所有说话者使用多通道信息,然后用于过滤音频的下游任务。我们采用了一种合成训练数据生成方法,通过回放和重新记录这种情况下,模拟具有挑战性的语音重叠条件。我们训练我们的模型上的合成数据,并证明我们的方法优于单通道VAD模型和基于能量的多通道VAD算法在各种声学环境。除了VAD结果,我们还提供了多方ASR评估结果,以突出使用我们的VAD模型通过显着减少插入错误来过滤下游任务中的音频的影响。
摘要:In this work, we propose a novel cross-talk rejection framework for a multi-channel multi-talker setup for a live multiparty interactive show. Our far-field audio setup is required to be hands-free during live interaction and comprises four adjacent talkers with directional microphones in the same space. Such setups often introduce heavy cross-talk between channels, resulting in reduced automatic speech recognition (ASR) and natural language understanding (NLU) performance. To address this problem, we propose voice activity detection (VAD) model for all talkers using multichannel information, which is then used to filter audio for downstream tasks. We adopt a synthetic training data generation approach through playback and re-recording for such scenarios, simulating challenging speech overlap conditions. We train our models on this synthetic data and demonstrate that our approach outperforms single-channel VAD models and energy-based multi-channel VAD algorithm in various acoustic environments. In addition to VAD results, we also present multiparty ASR evaluation results to highlight the impact of using our VAD model for filtering audio in downstream tasks by significantly reducing the insertion error.


【9】 Domain Adaptation for Contrastive Audio-Language Models
标题:对比性音频语言模型的领域自适应
链接:https://arxiv.org/abs/2402.09585
作者:Soham Deshmukh,Rita Singh,Bhiksha Raj
摘要:音频语言模型(ALM)的目标是通过在测试时提供zero-shot能力成为通用音频模型。通过为每个域使用合适的文本提示,ALM的zero-shot性能得到了提高。文本提示通常是通过特定过程手工制作的,这会导致ALM泛化和分发外性能的下降。现有的提高领域性能的方法,如Few-Shot学习或微调,需要访问带注释的数据和迭代训练。因此,我们提出了一个测试时域自适应方法的ALM,不需要访问注释。我们的方法通过在测试音频的增强视图中执行一致性来学习域向量。我们广泛地评估了我们的方法在12个跨域的下游任务。仅举一个例子,我们的域自适应方法导致平均zero-shot性能提高3.2%(最大8.4%)。自适应后,模型仍然保留了自适应层模型的泛化特性。
摘要:Audio-Language Models (ALM) aim to be general-purpose audio models by providing zero-shot capabilities at test time. The zero-shot performance of ALM improves by using suitable text prompts for each domain. The text prompts are usually hand-crafted through an ad-hoc process and lead to a drop in ALM generalization and out-of-distribution performance. Existing approaches to improve domain performance, like few-shot learning or fine-tuning, require access to annotated data and iterations of training. Therefore, we propose a test-time domain adaptation method for ALMs that does not require access to annotations. Our method learns a domain vector by enforcing consistency across augmented views of the testing audio. We extensively evaluate our approach on 12 downstream tasks across domains. With just one example, our domain adaptation method leads to 3.2% (max 8.4%) average zero-shot performance improvement. After adaptation, the model still retains the generalization property of ALMs.

机器翻译由腾讯交互翻译提供,仅供参考