今日论文合集:cs.SD语音24篇,eess.AS音频处理31篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Exploring the Capability of Mamba in Speech Applications
标题: 探索曼巴在语音应用中的能力
作者:Koichi Miyazaki,Yoshiki Masuyama,Masato Murata
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文探讨了Mamba的能力,最近提出的基于状态空间模型(SSM)的架构,作为一个有竞争力的替代基于transformer的模型。在语音领域,设计良好的基于Transformer的模型,如Conformer和E-Branchformer,已经成为事实上的标准。广泛的评估已经证明了这些基于transformer的模型在广泛的语音任务中的有效性。相比之下,SSM的评估仅限于少数任务,如自动语音识别(ASR)和语音合成。在本文中,我们比较了Mamba与最先进的Transformer变体在各种语音应用中的应用,包括ASR、文本到语音、口语理解和语音摘要。实验评估表明,Mamba实现了与基于transformer的模型相当或更好的性能,并证明了其在长形式语音处理中的效率。摘要:This paper explores the capability of Mamba, a recently proposed architecture based on state space models (SSMs), as a competitive alternative to Transformer-based models. In the speech domain, well-designed Transformer-based models, such as the Conformer and E-Branchformer, have become the de facto standards. Extensive evaluations have demonstrated the effectiveness of these Transformer-based models across a wide range of speech tasks. In contrast, the evaluation of SSMs has been limited to a few tasks, such as automatic speech recognition (ASR) and speech synthesis. In this paper, we compared Mamba with state-of-the-art Transformer variants for various speech applications, including ASR, text-to-speech, spoken language understanding, and speech summarization. Experimental evaluations revealed that Mamba achieves comparable or better performance than Transformer-based models, and demonstrated its efficiency in long-form speech processing.

【2】 Towards Zero-Shot Text-To-Speech for Arabic Dialects
标题: 阿拉伯方言的Zero-Shot文本语音转换
作者:Khai Duy Doan,Abdul Waheed,Muhammad Abdul-Mageed
链接:点击下载PDF文件
摘要:Zero-shot多说话人文本到语音转换(ZERO SHOT MULTI-SPEAKETTTS)系统已经在英语方面取得了进展,然而,由于资源不足,它仍然落后。我们首先通过调整现有的大量数据集来满足语音合成的需求,从而解决阿拉伯语(一种拥有超过4.5亿母语者的语言)的这一差距。此外,我们采用了一套阿拉伯语方言识别模型,探讨在多方言设置的影响,预定义的方言标签,以改善语音-TTS模型。随后,我们微调了XTTS footnote{https: docs.coqui.ai en latest models xtts.html} footnote{https: medium.com machine-learns xtts-v2-new-version-of-the-open-source-text-to-speech-model-af 73914 db 81 f} footnote{https: medium.com @erogol xtts-v1-techincal-notes-eb83ff05bdc}模型,这是一个开源架构。然后,我们在一个包含31个看不见的说话者和一个内部方言数据集的数据集上评估我们的模型。我们的自动和人工评估结果显示出令人信服的性能,同时能够生成方言语音。我们的研究突出了阿拉伯语研究这一新兴领域的重大改进潜力。摘要:Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native speakers, by first adapting a sizeable existing dataset to suit the needs of speech synthesis. Additionally, we employ a set of Arabic dialect identification models to explore the impact of pre-defined dialect labels on improving the ZS-TTS model in a multi-dialect setting. Subsequently, we fine-tune the XTTS footnote{https: docs.coqui.ai en latest models xtts.html} footnote{https: medium.com machine-learns xtts-v2-new-version-of-the-open-source-text-to-speech-model-af73914db81f} footnote{https: medium.com @erogol xtts-v1-techincal-notes-eb83ff05bdc} model, an open-source architecture. We then evaluate our models on a dataset comprising 31 unseen speakers and an in-house dialectal dataset. Our automated and human evaluation results show convincing performance while capable of generating dialectal speech. Our study highlights significant potential for improvements in this emerging area of research in Arabic.

【3】 SNR-Progressive Model with Harmonic Compensation for Low-SNR Speech Enhancement
标题: 用于低SNR语音增强的带调和补偿的SNR递进模型
作者:Zhongshu Hou,Qinwen Hu,Zhanzhong Cao,Ming Tang,Jing Lu
链接:点击下载PDF文件
摘要:尽管在过去十年中取得了重大进展,但基于深度神经网络(DNN)的语音增强(SE)仍然面临着在低信噪比(SNR)条件下恢复语音质量显着下降的挑战。在这封信中,我们提出了一个信噪比渐进语音增强模型与谐波补偿低信噪比SE。从中间输出获得可靠的音高估计,其具有比粗略估计保留更多语音分量同时拥有比输入噪声语音显著更高的SNR的益处。为了更好地恢复谐波,引入了有效的谐波补偿机制。大量的实验证明了我们提出的模型的优点。基于所提出的骨干模型的多模态语音提取系统在ICASSP 2024 MISP挑战赛中排名第一:https: mispchallenge.github.io mispchallenge2023 index.html。摘要:Despite significant progress made in the last decade, deep neural network (DNN) based speech enhancement (SE) still faces the challenge of notable degradation in the quality of recovered speech under low signal-to-noise ratio (SNR) conditions. In this letter, we propose an SNR-progressive speech enhancement model with harmonic compensation for low-SNR SE. Reliable pitch estimation is obtained from the intermediate output, which has the benefit of retaining more speech components than the coarse estimate while possessing a significant higher SNR than the input noisy speech. An effective harmonic compensation mechanism is introduced for better harmonic recovery. Extensive ex-periments demonstrate the advantage of our proposed model. A multi-modal speech extraction system based on the proposed backbone model ranks first in the ICASSP 2024 MISP Challenge: https: mispchallenge.github.io mispchallenge2023 index.html.

【4】 Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
标题: 聆听并移动:提高不可知声音到视频生成中的GAN一致性
作者:Rafael Redondo
Journal-ref:Abstract version published in the ICCV 2023 workshop "AV4D: Visual Learning of Sounds in Spaces"
链接:点击下载PDF文件
摘要:深度生成模型已经证明了创建逼真视听内容的能力,有时由不同性质的域驱动。然而,在视频生成中,平滑的时间动态是一个具有挑战性的问题。这项工作的重点是通用的声音到视频的生成,并提出了三个主要特征,以提高生成对抗模型中的图像质量和时间相干性:三重声音路由方案,用于扩展声音分析的多尺度残差和扩张递归网络,以及用于视频预测的新型递归和定向卷积层。所提出的每个特征在质量和一致性方面都改进了SoTA中通常使用的基线神经架构,其中视频预测层提供了额外的时间细化。摘要:Deep generative models have demonstrated the ability to create realistic audiovisual content, sometimes driven by domains of different nature. However, smooth temporal dynamics in video generation is a challenging problem. This work focuses on generic sound-to-video generation and proposes three main features to enhance both image quality and temporal coherency in generative adversarial models: a triple sound routing scheme, a multi-scale residual and dilated recurrent network for extended sound analysis, and a novel recurrent and directional convolutional layer for video prediction. Each of the proposed features improves, in both quality and coherency, the baseline neural architecture typically used in the SoTA, with the video prediction layer providing an extra temporal refinement.

【5】 Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking
标题: 迈向开放式呼吸声学基础模型:预训练和基准
作者:Yuwei Zhang,Tong Xia,Jing Han,Yu Wu,Georgios Rizos,Yang Liu,Mohammed Mosuily,Jagmohan Chauhan,Cecilia Mascolo
链接:点击下载PDF文件
摘要:呼吸音频(如咳嗽和呼吸声)具有广泛的医疗保健应用的预测能力,但目前尚未得到充分开发。这些应用程序的主要问题来自于难以收集用于模型开发的大型标记特定于任务的数据。用未标记数据预训练的可推广的呼吸声学基础模型将提供吸引人的优势,并可能打破这一僵局。然而,考虑到医疗保健应用程序的安全关键性,确保任何建议的基础模型解决方案的开放性和可复制性也是至关重要的。为此,我们引入OPERA,OPEn呼吸声学基础模型预训练和基准测试系统,作为满足这一需求的第一种方法。我们策划了大规模的呼吸音频数据集(约136K样本,440小时),预训练了三个开创性的基础模型,并建立了一个由19个下游呼吸健康任务组成的基准进行评估。我们的预训练模型表现出卓越的性能(与现有的声学模型相比,在19个任务中有16个任务使用普通音频进行了预训练)和可推广性(对于看不见的数据集和新的呼吸音频模式)。这突出了呼吸声学基础模型的巨大前景,并鼓励更多的研究使用OPERA作为一个开放的资源,以加速对健康的呼吸音频的研究。该系统可从https: github.com evelyn0414 OPERA访问。摘要:Respiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Generalizable respiratory acoustic foundation models pretrained with unlabeled data would offer appealing advantages and possibly unlock this impasse. However, given the safety-critical nature of healthcare applications, it is pivotal to also ensure openness and replicability for any proposed foundation model solution. To this end, we introduce OPERA, an OPEn Respiratory Acoustic foundation model pretraining and benchmarking system, as the first approach answering this need. We curate large-scale respiratory audio datasets (~136K samples, 440 hours), pretrain three pioneering foundation models, and build a benchmark consisting of 19 downstream respiratory health tasks for evaluation. Our pretrained models demonstrate superior performance (against existing acoustic models pretrained with general audio on 16 out of 19 tasks) and generalizability (to unseen datasets and new respiratory audio modalities). This highlights the great promise of respiratory acoustic foundation models and encourages more studies using OPERA as an open resource to accelerate research on respiratory audio for health. The system is accessible from https: github.com evelyn0414 OPERA.

【6】 Speech Representation Analysis based on Inter- and Intra-Model Similarities
标题: 基于模型间和模型内相似性的语音表示分析
作者:Yassine El Kheir,Ahmed Ali,Shammur Absar Chowdhury
备注:5 pages, Accepted to appear in ICASSP XAI-SA Workshop
链接:点击下载PDF文件
摘要:自监督模型已经彻底改变了语音处理,在有限的资源下在各种任务中实现了新的性能水平。然而,这些模型的内部运作仍然不透明。在本文中,我们的目标是分析这些基础模型的编码上下文表示的基础上,他们之间和模型内的相似性,独立于任何外部注释和特定于任务的约束。我们研究了不同的SSL模型,这些模型的训练模式不同-对比模型(Wav2Vec2.0)和预测模型(HuBERT);以及模型大小(基础和大型)。我们在不同层次的信息定位 分布上探索这些模型,包括(i)单个神经元;(ii)层表示;(iii)注意力权重和(iv)将表征与其微调的对应物进行比较。我们的结果强调,这些模型收敛到相似的表征子空间,但不收敛到相似的神经元局部概念 脚注{概念表示知识的连贯片段,例如 一个包含某些对象作为元素的类,其中对象具有某些属性。我们公开了代码,以促进进一步的研究,我们公开发布了我们的代码。摘要:Self-supervised models have revolutionized speech processing, achieving new levels of performance in a wide variety of tasks with limited resources. However, the inner workings of these models are still opaque. In this paper, we aim to analyze the encoded contextual representation of these foundation models based on their inter- and intra-model similarity, independent of any external annotation and task-specific constraint. We examine different SSL models varying their training paradigm -- Contrastive (Wav2Vec2.0) and Predictive models (HuBERT); and model sizes (base and large). We explore these models on different levels of localization distributivity of information including (i) individual neurons; (ii) layer representation; (iii) attention weights and (iv) compare the representations with their finetuned counterparts.Our results highlight that these models converge to similar representation subspaces but not to similar neuron-localized concepts footnote{A concept represents a coherent fragment of knowledge, such as a class containing certain objects as elements, where the objects have certain properties. We made the code publicly available for facilitating further research, we publicly released our code.

【7】 AudioBench: A Universal Benchmark for Audio Large Language Models
标题: AudioBench:音频大型语言模型的通用基准
作者:Bin Wang,Xunlong Zou,Geyu Lin,Shuo Sun,Zhuohan Liu,Wenyu Zhang,Zhengyuan Liu,AiTi Aw,Nancy F. Chen
备注:20 pages; Preprint; Code: this https URL
链接:点击下载PDF文件
摘要:我们介绍AudioBench,一个新的基准测试,旨在评估音频大语言模型(AudioLLM)。AudioBench包含8个不同的任务和26个精心挑选或新策划的数据集,专注于语音理解,语音解释和音频场景理解。尽管大型语言模型(包括多模式版本)的发展迅速,但在全面评估其能力的综合基准方面存在着重大差距。AudioBench通过提供相关的数据集和评估指标来解决这一差距。在我们的研究中,我们评估了四种模型在各个方面的能力,发现没有一种模型在所有任务中都表现出色。我们概述了AudioLLM的研究前景,并预计我们的开源代码,数据和排行榜将为未来的模型开发提供一个强大的测试平台。摘要:We introduce AudioBench, a new benchmark designed to evaluate audio large language models (AudioLLMs). AudioBench encompasses 8 distinct tasks and 26 carefully selected or newly curated datasets, focusing on speech understanding, voice interpretation, and audio scene understanding. Despite the rapid advancement of large language models, including multimodal versions, a significant gap exists in comprehensive benchmarks for thoroughly evaluating their capabilities. AudioBench addresses this gap by providing relevant datasets and evaluation metrics. In our study, we evaluated the capabilities of four models across various aspects and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-source code, data, and leaderboard will offer a robust testbed for future model developments.

【8】 Predicting Individual Depression Symptoms from Acoustic Features During Speech
标题: 根据言语过程中的声学特征预测个人抑郁症状
作者:Sebastian Rodriguez,Sri Harsha Dumpala,Katerina Dikaios,Sheri Rempel,Rudolf Uher,Sageev Oore
链接:点击下载PDF文件
摘要:目前的自动抑郁症检测系统直接提供预测,而不依赖于临床抑郁症评定量表中所示的抑郁症的个体症状 项目。相比之下,临床医生在临床环境中评估抑郁症评定量表中的每个项目,从而隐含地为抑郁症诊断提供更详细的理由。在这项工作中,我们做了第一步,使用语音的声学特征来预测个别项目的抑郁症评级量表,然后获得最终的抑郁症预测。为此,我们使用卷积(CNN)和循环(长短期记忆(LSTM))神经网络。我们考虑不同的方法来学习语音的时间背景。此外,我们分析了两个变量的投票方案,个别项目的预测和抑郁症的检测。我们还包括一个动画可视化,显示随着时间的推移,随着讲话的进展项目预测的例子。摘要:Current automatic depression detection systems provide predictions directly without relying on the individual symptoms items of depression as denoted in the clinical depression rating scales. In contrast, clinicians assess each item in the depression rating scale in a clinical setting, thus implicitly providing a more detailed rationale for a depression diagnosis. In this work, we make a first step towards using the acoustic features of speech to predict individual items of the depression rating scale before obtaining the final depression prediction. For this, we use convolutional (CNN) and recurrent (long short-term memory (LSTM)) neural networks. We consider different approaches to learning the temporal context of speech. Further, we analyze two variants of voting schemes for individual item prediction and depression detection. We also include an animated visualization that shows an example of item prediction over time as the speech progresses.

【9】 Real-time Speech Summarization for Medical Conversations
标题: 医疗对话的实时语音总结
作者:Khai Le-Duc,Khai-Nguyen Nguyen,Long Vo-Dang,Truong-Son Hy
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:在医患对话中,识别医学相关信息至关重要,因此需要进行对话摘要。在这项工作中,我们提出了第一个可部署的实时语音摘要系统在工业中的实际应用,它产生一个本地摘要后,每N个语音话语在一个对话和一个全球性的总结后,结束对话。我们的系统可以从商业角度增强用户体验,同时从技术角度降低计算成本。其次,我们介绍了VietMed-Sum,据我们所知,它是第一个用于医疗对话的语音摘要数据集。第三,我们是第一个利用LLM和人类注释器协同创建医疗会话摘要的黄金标准和综合摘要的公司。最后,我们提出了最先进的模型VietMed-Sum的基线结果。所有代码、数据(英语翻译和越南语)和模型均可在线获取:https: github.com leduckhai MultiMed摘要:In doctor-patient conversations, identifying medically relevant information is crucial, posing the need for conversation summarization. In this work, we propose the first deployable real-time speech summarization system for real-world applications in industry, which generates a local summary after every N speech utterances within a conversation and a global summary after the end of a conversation. Our system could enhance user experience from a business standpoint, while also reducing computational costs from a technical perspective. Secondly, we present VietMed-Sum which, to our knowledge, is the first speech summarization dataset for medical conversations. Thirdly, we are the first to utilize LLM and human annotators collaboratively to create gold standard and synthetic summaries for medical conversation summarization. Finally, we present baseline results of state-of-the-art models on VietMed-Sum. All code, data (English-translated and Vietnamese) and models are available online: https: github.com leduckhai MultiMed

【10】 The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
标题: 音乐大师或音乐大师,大型语言模型的大规模音乐评估基准
作者:Jiajia Li,Lu Yang,Mingni Tang,Cong Chen,Zuchao Li,Ping Wang,Hai Zhao
备注:Accepted to ACL-Findings 2024
链接:点击下载PDF文件
摘要:Benchmark在评估大型语言模型(LLM)的进步方面发挥着关键作用。虽然已经提出了许多基准来评估LLM的能力,但值得注意的是缺乏一个专门的基准来评估他们的音乐能力。为了解决这一差距,我们提出了ZIQI-Eval,这是一个全面的大规模音乐基准,专门用于评估LLM的音乐相关能力。ZIQI-Eval包含广泛的问题,涵盖10个主要类别和56个子类别,导致超过14,000个精心策划的数据条目。通过利用ZIQI-Eval,我们对16个LLM进行了综合评估,以评估和分析LLM在音乐领域的表现。结果表明,所有的LLM表现不佳的ZIQI-Eval基准,这表明显着的空间,改善他们的音乐能力。通过ZIQI-Eval,我们的目标是提供一个标准化和强大的评估框架,以促进对LLM音乐相关能力的全面评估。该数据集可在GitHub footnote{https: github.com zcli-charlie ZIQI-Eval}和HuggingFace footnote{https: huggingface.co MyTH-Lab ZIQI-Eval}获得。摘要:Benchmark plays a pivotal role in assessing the advancements of large language models (LLMs). While numerous benchmarks have been proposed to evaluate LLMs' capabilities, there is a notable absence of a dedicated benchmark for assessing their musical abilities. To address this gap, we present ZIQI-Eval, a comprehensive and large-scale music benchmark specifically designed to evaluate the music-related capabilities of LLMs. ZIQI-Eval encompasses a wide range of questions, covering 10 major categories and 56 subcategories, resulting in over 14,000 meticulously curated data entries. By leveraging ZIQI-Eval, we conduct a comprehensive evaluation over 16 LLMs to evaluate and analyze LLMs' performance in the domain of music. Results indicate that all LLMs perform poorly on the ZIQI-Eval benchmark, suggesting significant room for improvement in their musical capabilities. With ZIQI-Eval, we aim to provide a standardized and robust evaluation framework that facilitates a comprehensive assessment of LLMs' music-related abilities. The dataset is available at GitHub footnote{https: github.com zcli-charlie ZIQI-Eval} and HuggingFace footnote{https: huggingface.co datasets MYTH-Lab ZIQI-Eval}.

【11】 AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
标题: 基于人工智能的无人机在灾难环境中协助人类救援:挑战和机遇
作者:Narek Papyan,Michel Kulhandjian,Hovannes Kulhandjian,Levon Hakob Aslanyan
Journal-ref:Pattern Recognit. Image Anal. 34 (2024)
链接:点击下载PDF文件
摘要:在这项调查中,我们专注于利用基于无人机的系统来检测个人,特别是通过识别人类的尖叫声和其他求救信号。这项研究在灾后情景中具有重要意义,包括地震,飓风,军事冲突,野火等事件。这些无人机能够在受灾地区上空盘旋,这可能对救援队直接进入造成挑战。无人驾驶飞行器(UAV),通常被称为无人机,经常被部署用于灾害情况下的搜索和救援任务。通常情况下,无人机会捕捉空中图像,以评估结构性损坏并确定灾害的程度。他们还采用热成像技术来检测身体热量信号,这可以帮助定位个人。在某些情况下,大型无人机被用来向被困在孤立的受灾地区的人们运送必需品。在我们的讨论中,我们深入研究了与通过航空声学定位人类相关的独特挑战。听觉系统必须区分人类的哭声和自然发生的声音,如动物的叫声和风声。此外,它应该能够识别与信号相关的不同模式,如喊叫,鼓掌或人们试图向救援队发出信号的其他方式。为了应对这一挑战,一种解决方案涉及利用人工智能(AI)来分析声音频率并识别常见的音频签名。基于深度学习的网络,如卷积神经网络(CNN),可以使用这些特征进行训练,以过滤掉无人机电机和其他环境因素产生的噪声。此外,采用基于麦克风阵列信号的波达方向(DOA)等信号处理技术可以提高跟踪人类噪声源的精度。摘要:In this survey we are focusing on utilizing drone-based systems for the detection of individuals, particularly by identifying human screams and other distress signals. This study has significant relevance in post-disaster scenarios, including events such as earthquakes, hurricanes, military conflicts, wildfires, and more. These drones are capable of hovering over disaster-stricken areas that may be challenging for rescue teams to access directly. Unmanned aerial vehicles (UAVs), commonly referred to as drones, are frequently deployed for search-and-rescue missions during disaster situations. Typically, drones capture aerial images to assess structural damage and identify the extent of the disaster. They also employ thermal imaging technology to detect body heat signatures, which can help locate individuals. In some cases, larger drones are used to deliver essential supplies to people stranded in isolated disaster-stricken areas. In our discussions, we delve into the unique challenges associated with locating humans through aerial acoustics. The auditory system must distinguish between human cries and sounds that occur naturally, such as animal calls and wind. Additionally, it should be capable of recognizing distinct patterns related to signals like shouting, clapping, or other ways in which people attempt to signal rescue teams. To tackle this challenge, one solution involves harnessing artificial intelligence (AI) to analyze sound frequencies and identify common audio signatures. Deep learning-based networks, such as convolutional neural networks (CNNs), can be trained using these signatures to filter out noise generated by drone motors and other environmental factors. Furthermore, employing signal processing techniques like the direction of arrival (DOA) based on microphone array signals can enhance the precision of tracking the source of human noises.

【12】 Multimodal Segmentation for Vocal Tract Modeling
标题: 用于人声建模的多模式分割
作者:Rishi Jain,Bohan Yu,Peter Wu,Tejas Prabhune,Gopala Anumanchipalli
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:声道的精确建模对于构建可解释的语音处理和语言学的发音表示是必要的。然而,声道建模是具有挑战性的,因为许多内部发音被外部运动捕捉技术遮挡。实时磁共振成像(RT-MRI)允许测量语音期间内部发音器官的精确运动,但是由于耗时且计算昂贵的标记方法,MRI的注释数据集的大小受到限制。我们首先提出了一个深度标记策略的RT-MRI视频使用的视觉分割方法。然后,我们引入了一个多模态算法,使用音频,以改善分割的发音器官。总之,我们为MRI视频分割中的声道建模设定了一个新的基准,并使用它来发布75扬声器RT-MRI数据集的标签,将声道的标记公共RT-MRI数据量增加了9倍以上。代码和数据集标签可以在 url{rishiraij.github.io multimodal-mri-avatar }找到。摘要:Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articulators are occluded from external motion capture technologies. Real-time magnetic resonance imaging (RT-MRI) allows measuring precise movements of internal articulators during speech, but annotated datasets of MRI are limited in size due to time-consuming and computationally expensive labeling methods. We first present a deep labeling strategy for the RT-MRI video using a vision-only segmentation approach. We then introduce a multimodal algorithm using audio to improve segmentation of vocal articulators. Together, we set a new benchmark for vocal tract modeling in MRI video segmentation and use this to release labels for a 75-speaker RT-MRI dataset, increasing the amount of labeled public RT-MRI data of the vocal tract by over a factor of 9. The code and dataset labels can be found at url{rishiraij.github.io multimodal-mri-avatar }.

【13】 Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
标题: 使用GAN和集成的未对齐清洁数据改进无监督清洁到渲染的吉他音调转换
作者:Yu-Hua Chen,Woosung Choi,Wei-Hsiang Liao,Marco Martínez-Ramírez,Kin Wai Cheuk,Yuki Mitsufuji,Jyh-Shing Roger Jang,Yi-Hsuan Yang
备注:Accepted to DAFx 2024
链接:点击下载PDF文件
摘要:近年来,人们对将深度学习方法应用于吉他放大器或效果踏板建模的兴趣越来越大。现有方法主要基于监督方法,需要未处理和渲染音频的时间对齐数据对。然而,由于创建数据对涉及复杂的过程,这种方法的扩展性不佳。Wright等人最近的一项工作探索了利用未配对数据进行训练的潜力,使用基于生成对抗网络(GAN)的框架。本文通过在GAN中使用更高级的鉴别器,并使用更多的未配对数据进行训练,扩展了他们的工作。具体来说,从神经声码器的最新进展中汲取灵感,我们在基于GAN的吉他放大器模型中采用了两组鉴别器,一组基于多尺度鉴别器(MSD),另一组基于多周期鉴别器(MPD)。此外,我们尝试将未处理的音频信号添加到训练数据中,这些信号没有目标音调的相应渲染音频,以查看GAN模型从未配对数据中受益多少。我们的实验表明,提出的两个扩展有助于建模的低增益和高增益吉他放大器。摘要:Recent years have seen increasing interest in applying deep learning methods to the modeling of guitar amplifiers or effect pedals. Existing methods are mainly based on the supervised approach, requiring temporally-aligned data pairs of unprocessed and rendered audio. However, this approach does not scale well, due to the complicated process involved in creating the data pairs. A very recent work done by Wright et al. has explored the potential of leveraging unpaired data for training, using a generative adversarial network (GAN)-based framework. This paper extends their work by using more advanced discriminators in the GAN, and using more unpaired data for training. Specifically, drawing inspiration from recent advancements in neural vocoders, we employ in our GAN-based model for guitar amplifier modeling two sets of discriminators, one based on multi-scale discriminator (MSD) and the other multi-period discriminator (MPD). Moreover, we experiment with adding unprocessed audio signals that do not have the corresponding rendered audio of a target tone to the training data, to see how much the GAN model benefits from the unpaired data. Our experiments show that the proposed two extensions contribute to the modeling of both low-gain and high-gain guitar amplifiers.

【14】 Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
标题: 用于平衡多方面发音评估的声学特征混合
作者:Heejin Do,Wonjun Lee,Gary Geunbae Lee
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:在自动发音评估中,最近的重点逐渐在于评估多个方面以提供丰富的反馈。然而,获取非母语学习者语音的多方面分数标记数据提出了挑战,而且,它往往会导致分数不平衡分布。在本文中,我们提出了两个声学特征混合策略,线性和非线性插值与批平均功能,以解决数据稀缺和分数标签的不平衡。首先,使用良好的发音作为声学特征,我们量身定制的混合设计,以适应发音评估。此外,我们通过将语音识别结果与原始答案音素进行比较来集成细粒度错误率特征,从而给出发音错误的直接提示。声学特征的有效混合显着增强了speechocean 762数据集的整体评分性能,并且详细的分析强调了我们预测看不见的失真的潜力。摘要:In automated pronunciation assessment, recent emphasis progressively lies on evaluating multiple aspects to provide enriched feedback. However, acquiring multi-aspect-score labeled data for non-native language learners' speech poses challenges; moreover, it often leads to score-imbalanced distributions. In this paper, we propose two Acoustic Feature Mixup strategies, linearly and non-linearly interpolating with the in-batch averaged feature, to address data scarcity and score-label imbalances. Primarily using goodness-of-pronunciation as an acoustic feature, we tailor mixup designs to suit pronunciation assessment. Further, we integrate fine-grained error-rate features by comparing speech recognition results with the original answer phonemes, giving direct hints for mispronunciation. Effective mixing of the acoustic features notably enhances overall scoring performances on the speechocean762 dataset, and detailed analysis highlights our potential to predict unseen distortions.

【15】 PI-Whisper: An Adaptive and Incremental ASR Framework for Diverse and Evolving Speaker Characteristics
标题: PI-Whisper:一种适应性和增量式的ASB框架,用于多样化和不断发展的说话者特征
作者:Amir Nassereldine,Dancheng Liu,Chenhui Xu,Jinjun Xiong
备注:11 pages, 3 figures
链接:点击下载PDF文件
摘要:随着基于边缘的自动语音识别(ASR)技术在智能和个性化助理的开发中变得越来越普遍,这些资源受限的ASR模型必须解决三个重要挑战,即,适应性、递增性和包容性。我们提出了一种新的ASR框架,PI耳语,在这项工作中,并显示它可以通过实时识别不同的扬声器的特性,如何自适应地提高ASR的识别能力,这样的适应可以逐步进行,而无需重复的再训练,以及它如何可以提高不同的扬声器组的公平性和公正性。更令人印象深刻的是,我们提出的PI Whisper框架实现了所有这些良好的属性,同时仍然实现了最先进的准确性,字错误率(WER)降低了13.7%,并且相对于计算资源具有线性可扩展性。摘要:As edge-based automatic speech recognition (ASR) technologies become increasingly prevalent for the development of intelligent and personalized assistants, three important challenges must be addressed for these resource-constrained ASR models, i.e., adaptivity, incrementality, and inclusivity. We propose a novel ASR framework, PI-Whisper, in this work and show how it can improve an ASR's recognition capabilities adaptively by identifying different speakers' characteristics in real-time, how such an adaption can be performed incrementally without repetitive retraining, and how it can improve the equity and fairness for diverse speaker groups. More impressively, our proposed PI-Whisper framework attains all of these nice properties while still achieving state-of-the-art accuracy with up to 13.7% reduction of the word error rate (WER) with linear scalability with respect to computing resources.

【16】 Generating Music with Structure Using Self-Similarity as Attention
标题: 利用自相似性作为注意力生成具有结构的音乐
作者:Sophia Hager,Kathleen Hablutzel,Katherine Kinnaird
链接:点击下载PDF文件
摘要:尽管在深度学习和生成式人工智能方面有了创新,但创建长期结构以及音乐作品中常见的重复结构层仍然是音乐生成中的一个公开挑战。我们提出了一个注意力层,它使用一种新的方法,将用户提供的自相似性矩阵应用于先前的时间步长,并在我们的相似性激励神经生成器(SING)系统中进行了演示,SING系统是一个具有两层的深度学习自主音乐生成系统。第一个是普通的长短期记忆层,第二个是建议的注意力层。在生成过程中,这种注意力机制将来自模板作品的建议结构强加于生成的音乐上。我们训练SING的MAESTRO数据集上使用一种新的变量的方法,并比较其性能相同的模型没有注意力机制。我们提出的注意力机制的加入显着提高了网络复制特定结构的能力,并且它在看不见的测试集上的表现比没有注意力机制的模型更好。摘要:Despite the innovations in deep learning and generative AI, creating long term structure as well as the layers of repeated structure common in musical works remains an open challenge in music generation. We propose an attention layer that uses a novel approach applying user-supplied self-similarity matrices to previous time steps, and demonstrate it in our Similarity Incentivized Neural Generator (SING) system, a deep learning autonomous music generation system with two layers. The first is a vanilla Long Short Term Memory layer, and the second is the proposed attention layer. During generation, this attention mechanism imposes a suggested structure from a template piece on the generated music. We train SING on the MAESTRO dataset using a novel variable batching method, and compare its performance to the same model without the attention mechanism. The addition of our proposed attention mechanism significantly improves the network's ability to replicate specific structures, and it performs better on an unseen test set than a model without the attention mechanism.

【17】 System Description for the Displace Speaker Diarization Challenge 2023
标题: 2023年Displace Speaker Diaration Challenge的系统描述
作者:Ali Aliyev
链接:点击下载PDF文件
摘要:本文介绍了我们的解决方案,对话环境中的说话者和语言的日记化挑战(位移2023)。我们使用VAD的组合来寻找语音片段,基于Resnet架构的CNN用于从这些片段中提取特征,以及谱聚类用于特征聚类。即使它没有使用印地语进行训练,所描述的算法也实现了以下指标:DER 27。1%,27。4%,分别用于数据集的开发和第1阶段评估部分。摘要:This paper describes our solution for the Diarization of Speaker and Language in Conversational Environments Challenge (Displace 2023). We used a combination of VAD for finding segfments with speech, Resnet architecture based CNN for feature extraction from these segments, and spectral clustering for features clustering. Even though it was not trained with using Hindi, the described algorithm achieves the following metrics: DER 27. 1% and DER 27. 4%, on the development and phase-1 evaluation parts of the dataset, respectively.

【18】 Improving Text-To-Audio Models with Synthetic Captions
标题: 使用合成字幕改进文本到音频模型
作者:Zhifeng Kong,Sang-gil Lee,Deepanway Ghosal,Navonil Majumder,Ambuj Mehrish,Rafael Valle,Soujanya Poria,Bryan Catanzaro
链接:点击下载PDF文件
摘要:为文本到音频模型获得高质量的训练数据,特别是字幕,是一个公开的挑战。虽然先前的方法已经利用了纯文本语言模型来增强和改进字幕,但这些方法在音频和字幕之间的规模和连贯性方面存在局限性。在这项工作中,我们提出了一个音频字幕管道,使用 textit{音频语言模型}来合成准确和多样化的字幕音频规模。我们利用这个管道为AudioSet生成一个名为 texttt{AF-AudioSet}的合成字幕数据集,然后评估这些合成字幕的预训练文本到音频模型的好处。通过对AudioCaps和MusicCaps的系统评估,我们发现利用我们的管道和合成字幕可以显著提高音频生成质量,实现新的 textit{最先进的}。摘要:It is an open challenge to obtain high quality training data, especially captions, for text-to-audio models. Although prior methods have leveraged textit{text-only language models} to augment and improve captions, such methods have limitations related to scale and coherence between audio and captions. In this work, we propose an audio captioning pipeline that uses an textit{audio language model} to synthesize accurate and diverse captions for audio at scale. We leverage this pipeline to produce a dataset of synthetic captions for AudioSet, named texttt{AF-AudioSet}, and then evaluate the benefit of pre-training text-to-audio models on these synthetic captions. Through systematic evaluations on AudioCaps and MusicCaps, we find leveraging our pipeline and synthetic captions leads to significant improvements on audio generation quality, achieving a new textit{state-of-the-art}.

【19】 One-Class Learning with Adaptive Centroid Shift for Audio Deepfake Detection
标题: 用于音频深度伪造检测的自适应中心漂移的一类学习
作者:Hyun Myung Kim,Kangwook Jang,Hoirin Kim
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,随着语音合成系统不断取得显著进步,在看不见的系统中表现良好的鲁棒深度伪造检测系统的重要性也在增长。在本文中,我们提出了一种新的自适应质心移位(ACS)的方法,更新的质心表示不断移动的加权平均值的善意表示。我们的方法只使用真正的样本来定义它们的质心,这可以为一类学习产生一个专门的质心。将我们的ACS与单类学习集成在一起,将真正的表示收集到一个集群中,形成分离良好的嵌入,对看不见的欺骗攻击具有鲁棒性。我们提出的方法在ASVspoof 2021 deepfake数据集上实现了2.19%的等错误率(EER),优于所有现有系统。此外,t-SNE可视化表明,我们的方法有效地映射到一个单一的集群的bonafide嵌入,并成功地解开bonafide和欺骗类。摘要:As speech synthesis systems continue to make remarkable advances in recent years, the importance of robust deepfake detection systems that perform well in unseen systems has grown. In this paper, we propose a novel adaptive centroid shift (ACS) method that updates the centroid representation by continually shifting as the weighted average of bonafide representations. Our approach uses only bonafide samples to define their centroid, which can yield a specialized centroid for one-class learning. Integrating our ACS with one-class learning gathers bonafide representations into a single cluster, forming well-separated embeddings robust to unseen spoofing attacks. Our proposed method achieves an equal error rate (EER) of 2.19% on the ASVspoof 2021 deepfake dataset, outperforming all existing systems. Furthermore, the t-SNE visualization illustrates that our method effectively maps the bonafide embeddings into a single cluster and successfully disentangles the bonafide and spoof classes.

【20】 Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework
标题: 使用神经分析和合成框架进行端到端神经歌手扩张的歌曲数据清理
作者:Hokuto Munakata,Ryo Terashima,Yusuke Fujita
备注:INTERSPEECH 2024 accepted
链接:点击下载PDF文件
摘要:我们提出了一种数据清洗方法,利用神经分析和合成(NANSY++)框架训练一个端到端的神经日记模型(EEND)的歌手日记。我们提出的模型将歌曲数据与合唱,这是通常包含在流行音乐和不适合生成模拟数据集的独唱数据。这种净化基于NANSY++,NANSY ++是一种经过训练以重建输入非重叠音频信号的框架。我们利用预训练的NANSY++将合唱转换为干净,不重叠的音频。这种净化过程减轻了将合唱误标记为独唱的情况,并有助于有效地训练EEND模型,即使大多数可用的歌曲数据包含合唱部分。我们通过实验评估了使用我们提出的方法使用注释的流行二重唱歌曲训练的数据集的EEND模型。结果,我们提出的方法改善了14.8点的日志化错误率。摘要:We propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization. Our proposed model converts song data with choral singing which is commonly contained in popular music and unsuitable for generating a simulated dataset to the solo singing data. This cleansing is based on NANSY++, which is a framework trained to reconstruct an input non-overlapped audio signal. We exploit the pre-trained NANSY++ to convert choral singing into clean, non-overlapped audio. This cleansing process mitigates the mislabeling of choral singing to solo singing and helps the effective training of EEND models even when the majority of available song data contains choral singing sections. We experimentally evaluated the EEND model trained with a dataset using our proposed method using annotated popular duet songs. As a result, our proposed method improved 14.8 points in diarization error rate.

【21】 Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
标题: 具有中间偏置损失的上下文化端到端自动语音识别
作者:Muhammad Shakeel,Yui Sudo,Yifan Peng,Shinji Watanabe
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语境化的端到端自动语音识别一直是一个活跃的研究领域,最近的努力集中在基于最终损失目标的语境短语的内隐学习。然而,这些方法忽略了有用的上下文知识编码在中间层。我们假设采用显式偏置损失作为编码器中间层中的辅助任务可以更好地将文本令牌或音频帧与所需目标对齐。我们提出的中间偏置损失为网络带来了更多的正则化和上下文化。我们的方法优于LibriSpeech语料库上的传统上下文偏置基线,实现了22.5%的偏置词错误率(B-WER)的相对改善,与偏置列表大小为100的非上下文基线相比高达44%。此外,采用RNN传感器驱动的联合解码进一步降低了无偏字错误率(U-WER),从而使网络更加鲁棒。摘要:Contextualized end-to-end automatic speech recognition has been an active research area, with recent efforts focusing on the implicit learning of contextual phrases based on the final loss objective. However, these approaches ignore the useful contextual knowledge encoded in the intermediate layers. We hypothesize that employing explicit biasing loss as an auxiliary task in the encoder intermediate layers may better align text tokens or audio frames with the desired objectives. Our proposed intermediate biasing loss brings more regularization and contextualization to the network. Our method outperforms a conventional contextual biasing baseline on the LibriSpeech corpus, achieving a relative improvement of 22.5% in biased word error rate (B-WER) and up to 44% compared to the non-contextual baseline with a biasing list size of 100. Moreover, employing RNN-transducer-driven joint decoding further reduces the unbiased word error rate (U-WER), resulting in a more robust network.

【22】 Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
标题: 融合音频和元数据嵌入改进了基于数据的音频检索
作者:Paul Primus,Gerhard Widmer
备注:EUSIPCO 2024
链接:点击下载PDF文件
摘要:将原始音频信号与文本描述进行匹配需要理解音频的内容和描述的语义,然后在两种模态之间建立联系。本文研究了一个混合检索系统,利用音频元数据作为一个额外的线索,以了解音频信号的内容之前,匹配它们与文本查询。我们试验了经常附加到音频记录的元数据,例如关键字和自然语言描述,并研究了后期和中期融合策略,以合并音频和元数据。我们的混合方法与关键字元数据和后期融合提高了检索性能超过基于内容的基线由2.36和3.69页。分别在ClothoV2和AudioCaps基准上的mAP@10。摘要:Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid retrieval system that utilizes audio metadata as an additional clue to understand the content of audio signals before matching them with textual queries. We experimented with metadata often attached to audio recordings, such as keywords and natural-language descriptions, and we investigated late and mid-level fusion strategies to merge audio and metadata. Our hybrid approach with keyword metadata and late fusion improved the retrieval performance over a content-based baseline by 2.36 and 3.69 pp. mAP@10 on the ClothoV2 and AudioCaps benchmarks, respectively.

【23】 Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes
标题: 具有粗预测池和声音事件边界盒的自训练和集成频率相关网络
作者:Hyeonuk Nam,Deokki Min,Seungdeok Choi,Inhan Choi,Yong-Hwa Park
备注:DCASE 2024 Challenge Task 4 technical report
链接:点击下载PDF文件
摘要:为了解决声音事件检测(SED)任务,我们提出了频率相关网络(FreDNets),它严重利用频率相关的方法。我们应用频率扭曲和FilterAugment,这是频率相关的数据增强方法。该模型结构由3个分支组成:音频教师-学生Transformer(ATST)分支、BEAT分支和CNN分支,包括部分扩频动态卷积(PDFD)或具有时帧频率方式SE(tfwSE)的挤压和激励(SE)。为了训练具有粗略时间分辨率的MAESTRO标签,我们对MAESTRO数据集的预测应用最大池化。利用最佳集成模型,通过自训练从DESED弱集、DESED未标记集和AudioSet中获取伪标记。AudioSet标签被过滤以集中于高置信度伪标签,并且AudioSet伪标签仅用于在DESED标签上进行训练。我们使用基于变化检测的声音事件边界框(cSEBBs)作为自训练和提交模型的集成模型的后处理。摘要:To tackle sound event detection (SED) task, we propose frequency dependent networks (FreDNets), which heavily leverage frequency-dependent methods. We apply frequency warping and FilterAugment, which are frequency-dependent data augmentation methods. The model architecture consists of 3 branches: audio teacher-student transformer (ATST) branch, BEATs branch and CNN branch including either partial dilated frequency dynamic convolution (PDFD) or squeeze-and-Excitation (SE) with time-frame frequency-wise SE (tfwSE). To train MAESTRO labels with coarse temporal resolution, we apply max pooling on prediction for the MAESTRO dataset. Using best ensemble model, we apply self training to obtain pseudo label from DESED weak set, DESED unlabeled set and AudioSet. AudioSet labels are filtered to focus on high-confidence pseudo labels and AudioSet pseudo labels are used to train on DESED labels only. We used change-detection-based sound event bounding boxes (cSEBBs) as post processing for ensemble models on self training and submission models.

【24】 R&B -- Rhythm and Brain: Cross-subject Decoding of Music from Human Brain Activity
标题: R & B --节奏与大脑:从人脑活动中解读音乐的跨学科
作者:Matteo Ferrante,Matteo Ciferri,Nicola Toschi
备注:The first two authors contributed equally to this work
链接:点击下载PDF文件
摘要:音乐是一种普遍现象,深刻影响着不同文化的人类体验。这项研究调查了音乐是否可以从人类大脑活动中解码,在其感知过程中使用功能性MRI(fMRI)进行测量。利用最近在广泛的数据集和预先训练的计算模型方面的进展,我们构建了神经数据和音乐刺激的潜在表征之间的映射。我们的方法集成了功能和解剖对齐技术,以促进跨学科的解码,解决了fMRI数据的低时间分辨率和信噪比(SNR)所带来的挑战。从GTZan fMRI数据集开始,其中5名参与者在记录大脑活动的同时听取了来自10种不同流派的540种音乐刺激,我们使用CLAP(对比听觉-音频预训练)模型来提取音乐刺激的潜在表征,并开发了体素编码模型来识别对这些刺激做出反应的大脑区域。通过对预测和实际大脑活动之间的关联应用阈值,我们确定了特定的感兴趣区域(ROI),这些区域可以被解释为音乐处理中的关键角色。我们的解码管道主要基于检索,采用线性映射将大脑活动投射到相应的CLAP特征。这使我们能够预测和检索与fMRI数据来源最相似的音乐刺激。我们的结果证明了最先进的识别精度,我们的方法显着优于现有的方法。我们的研究结果表明,基于神经的音乐检索系统可以实现个性化推荐和治疗应用。未来的工作可以使用更高的时间分辨率神经成像和生成模型来提高解码准确性,并探索音乐感知和情感的神经基础。摘要:Music is a universal phenomenon that profoundly influences human experiences across cultures. This study investigates whether music can be decoded from human brain activity measured with functional MRI (fMRI) during its perception. Leveraging recent advancements in extensive datasets and pre-trained computational models, we construct mappings between neural data and latent representations of musical stimuli. Our approach integrates functional and anatomical alignment techniques to facilitate cross-subject decoding, addressing the challenges posed by the low temporal resolution and signal-to-noise ratio (SNR) in fMRI data. Starting from the GTZan fMRI dataset, where five participants listened to 540 musical stimuli from 10 different genres while their brain activity was recorded, we used the CLAP (Contrastive Language-Audio Pretraining) model to extract latent representations of the musical stimuli and developed voxel-wise encoding models to identify brain regions responsive to these stimuli. By applying a threshold to the association between predicted and actual brain activity, we identified specific regions of interest (ROIs) which can be interpreted as key players in music processing. Our decoding pipeline, primarily retrieval-based, employs a linear map to project brain activity to the corresponding CLAP features. This enables us to predict and retrieve the musical stimuli most similar to those that originated the fMRI data. Our results demonstrate state-of-the-art identification accuracy, with our methods significantly outperforming existing approaches. Our findings suggest that neural-based music retrieval systems could enable personalized recommendations and therapeutic applications. Future work could use higher temporal resolution neuroimaging and generative models to improve decoding accuracy and explore the neural underpinnings of music perception and emotion.


eess.AS音频处理
【1】 One-Class Learning with Adaptive Centroid Shift for Audio Deepfake Detection
标题: 用于音频深度伪造检测的自适应中心漂移的一类学习
作者:Hyun Myung Kim,Kangwook Jang,Hoirin Kim
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,随着语音合成系统不断取得显著进步,在看不见的系统中表现良好的鲁棒深度伪造检测系统的重要性也在增长。在本文中,我们提出了一种新的自适应质心移位(ACS)的方法,更新的质心表示不断移动的加权平均值的善意表示。我们的方法只使用真正的样本来定义它们的质心,这可以为一类学习产生一个专门的质心。将我们的ACS与单类学习集成在一起,将真正的表示收集到一个集群中,形成分离良好的嵌入,对看不见的欺骗攻击具有鲁棒性。我们提出的方法在ASVspoof 2021 deepfake数据集上实现了2.19%的等错误率(EER),优于所有现有系统。此外,t-SNE可视化表明,我们的方法有效地映射到一个单一的集群的bonafide嵌入,并成功地解开bonafide和欺骗类。摘要:As speech synthesis systems continue to make remarkable advances in recent years, the importance of robust deepfake detection systems that perform well in unseen systems has grown. In this paper, we propose a novel adaptive centroid shift (ACS) method that updates the centroid representation by continually shifting as the weighted average of bonafide representations. Our approach uses only bonafide samples to define their centroid, which can yield a specialized centroid for one-class learning. Integrating our ACS with one-class learning gathers bonafide representations into a single cluster, forming well-separated embeddings robust to unseen spoofing attacks. Our proposed method achieves an equal error rate (EER) of 2.19% on the ASVspoof 2021 deepfake dataset, outperforming all existing systems. Furthermore, the t-SNE visualization illustrates that our method effectively maps the bonafide embeddings into a single cluster and successfully disentangles the bonafide and spoof classes.

【2】 RefXVC: Cross-Lingual Voice Conversion with Enhanced Reference Leveraging
标题: RefXVC:具有增强参考利用的跨语言语音转换
作者:Mingyang Zhang,Yi Zhou,Yi Ren,Chen Zhang,Xiang Yin,Haizhou Li
备注:Manuscript under review by TASLP
链接:点击下载PDF文件
摘要:本文提出了RefXVC,跨语言语音转换(XVC)的方法,利用参考信息,以提高转换性能。以前的XVC作品通常采用平均的说话人嵌入来调节说话人身份,这并没有考虑到不同发音所发生的语音音色的变化。为了解决这个问题,我们的方法使用全局和本地扬声器嵌入来捕获语音转换过程中的音色变化。此外,我们观察到不同语言中音色和发音之间的联系,并通过将音色编码器和发音匹配网络纳入我们的模型来利用这一点。此外,我们发现音调的变化在一个句子中没有得到充分的反映,因此,我们使用了多个参考来更好地捕捉说话者的声音范围。该方法在语音质量和说话人相似性方面优于现有系统,突出了利用参考信息进行跨语言语音转换的有效性。转换后的语音样本可以在以下网站上找到: url{http:refxvc.dn3point.com}摘要:This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally take an average speaker embedding to condition the speaker identity, which does not account for the changing timbre of speech that occurs with different pronunciations. To address this, our method uses both global and local speaker embeddings to capture the timbre changes during speech conversion. Additionally, we observed a connection between timbre and pronunciation in different languages and utilized this by incorporating a timbre encoder and a pronunciation matching network into our model. Furthermore, we found that the variation in tones is not adequately reflected in a sentence, and therefore, we used multiple references to better capture the range of a speaker's voice. The proposed method outperformed existing systems in terms of both speech quality and speaker similarity, highlighting the effectiveness of leveraging reference information in cross-lingual voice conversion. The converted speech samples can be found on the website: url{http: refxvc.dn3point.com}

【3】 Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework
标题: 使用神经分析和合成框架进行端到端神经歌手扩张的歌曲数据清理
作者:Hokuto Munakata,Ryo Terashima,Yusuke Fujita
备注:INTERSPEECH 2024 accepted
链接:点击下载PDF文件
摘要:我们提出了一种数据清洗方法,利用神经分析和合成(NANSY++)框架训练一个端到端的神经日记模型(EEND)的歌手日记。我们提出的模型将歌曲数据与合唱,这是通常包含在流行音乐和不适合生成模拟数据集的独唱数据。这种净化基于NANSY++,NANSY ++是一种经过训练以重建输入非重叠音频信号的框架。我们利用预训练的NANSY++将合唱转换为干净,不重叠的音频。这种净化过程减轻了将合唱误标记为独唱的情况,并有助于有效地训练EEND模型,即使大多数可用的歌曲数据包含合唱部分。我们通过实验评估了使用我们提出的方法使用注释的流行二重唱歌曲训练的数据集的EEND模型。结果,我们提出的方法改善了14.8点的日志化错误率。摘要:We propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization. Our proposed model converts song data with choral singing which is commonly contained in popular music and unsuitable for generating a simulated dataset to the solo singing data. This cleansing is based on NANSY++, which is a framework trained to reconstruct an input non-overlapped audio signal. We exploit the pre-trained NANSY++ to convert choral singing into clean, non-overlapped audio. This cleansing process mitigates the mislabeling of choral singing to solo singing and helps the effective training of EEND models even when the majority of available song data contains choral singing sections. We experimentally evaluated the EEND model trained with a dataset using our proposed method using annotated popular duet songs. As a result, our proposed method improved 14.8 points in diarization error rate.

【4】 DreamVoice: Text-Guided Voice Conversion
标题: DreamVoice:文本引导语音转换
作者:Jiarui Hai,Karan Thakkar,Helin Wang,Zengyi Qin,Mounya Elhilali
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:生成语音技术正在迅速发展,为更个性化和更具包容性的体验提供了机会。传统的单次语音转换(VC)需要在推理过程中进行目标录音,这限制了生成所需语音音色的易用性。文本引导生成提供了一个直观的解决方案,可以根据用户的需求将语音转换为所需的“DreamVoices”。本文介绍了VC技术的两个主要贡献:(1)DreamVoiceDB,一个来自VCTK和LibriTTS的900名说话者的语音音色注释的强大数据集。(2)两种文本引导VC方法:DreamVC,一种基于端到端扩散的文本引导VC模型;和DreamVG,一种通用的文本到语音生成插件,可以与任何一次性VC模型相结合。实验结果表明,我们提出的方法在DreamVoiceDB数据集上训练,生成的语音音色与文本提示准确对齐,实现了高质量的VC。摘要:Generative voice technologies are rapidly evolving, offering opportunities for more personalized and inclusive experiences. Traditional one-shot voice conversion (VC) requires a target recording during inference, limiting ease of usage in generating desired voice timbres. Text-guided generation offers an intuitive solution to convert voices to desired "DreamVoices" according to the users' needs. Our paper presents two major contributions to VC technology: (1) DreamVoiceDB, a robust dataset of voice timbre annotations for 900 speakers from VCTK and LibriTTS. (2) Two text-guided VC methods: DreamVC, an end-to-end diffusion-based text-guided VC model; and DreamVG, a versatile text-to-voice generation plugin that can be combined with any one-shot VC models. The experimental results demonstrate that our proposed methods trained on the DreamVoiceDB dataset generate voice timbres accurately aligned with the text prompt and achieve high-quality VC.

【5】 Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
标题: 利用基础模型和语音增强在现实世界操作条件下从语音中检测帕金森病
作者:Moreno La Quatra,Maria Francesca Turco,Torbjørn Svendsen,Giampiero Salvi,Juan Rafael Orozco-Arroyave,Sabato Marco Siniscalchi
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:这项工作涉及设计一个强大的帕金森氏症(PD)疾病检测器从语音在现实世界的操作条件下使用(i)基础模型,和(ii)语音增强(SE)的方法。为此,我们首先在标准PC-GITA(s-PC-GITA)干净数据上微调几个基于基础的模型。我们的研究结果表明,优越的性能,以前提出的模型。其次,我们评估了PD模型对在真实操作条件下收集的扩展PC-GITA(e-PC-GITA)记录的泛化能力,并观察到从理想条件到真实条件的性能严重下降。第三,我们调整训练和测试条件,在e-PC-GITA上应用现成的SE技术,并且仅在基于基础的模型中观察到性能的显着提高。最后,结合在s-PC-GITA上训练的两个最好的基于基础的模型,即WavLM Base和Hubert Base,在增强的e-PC-GITA上产生了最佳性能。摘要:This work is concerned with devising a robust Parkinson's (PD) disease detector from speech in real-world operating conditions using (i) foundational models, and (ii) speech enhancement (SE) methods. To this end, we first fine-tune several foundational-based models on the standard PC-GITA (s-PC-GITA) clean data. Our results demonstrate superior performance to previously proposed models. Second, we assess the generalization capability of the PD models on the extended PC-GITA (e-PC-GITA) recordings, collected in real-world operative conditions, and observe a severe drop in performance moving from ideal to real-world conditions. Third, we align training and testing conditions applaying off-the-shelf SE techniques on e-PC-GITA, and a significant boost in performance is observed only for the foundational-based models. Finally, combining the two best foundational-based models trained on s-PC-GITA, namely WavLM Base and Hubert Base, yielded top performance on the enhanced e-PC-GITA.

【6】 Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
标题: 具有中间偏置损失的上下文化端到端自动语音识别
作者:Muhammad Shakeel,Yui Sudo,Yifan Peng,Shinji Watanabe
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语境化的端到端自动语音识别一直是一个活跃的研究领域,最近的努力集中在基于最终损失目标的语境短语的内隐学习。然而,这些方法忽略了有用的上下文知识编码在中间层。我们假设采用显式偏置损失作为编码器中间层中的辅助任务可以更好地将文本令牌或音频帧与所需目标对齐。我们提出的中间偏置损失为网络带来了更多的正则化和上下文化。我们的方法优于LibriSpeech语料库上的传统上下文偏置基线,实现了22.5%的偏置词错误率(B-WER)的相对改善,与偏置列表大小为100的非上下文基线相比高达44%。此外,采用RNN传感器驱动的联合解码进一步降低了无偏字错误率(U-WER),从而使网络更加鲁棒。摘要:Contextualized end-to-end automatic speech recognition has been an active research area, with recent efforts focusing on the implicit learning of contextual phrases based on the final loss objective. However, these approaches ignore the useful contextual knowledge encoded in the intermediate layers. We hypothesize that employing explicit biasing loss as an auxiliary task in the encoder intermediate layers may better align text tokens or audio frames with the desired objectives. Our proposed intermediate biasing loss brings more regularization and contextualization to the network. Our method outperforms a conventional contextual biasing baseline on the LibriSpeech corpus, achieving a relative improvement of 22.5% in biased word error rate (B-WER) and up to 44% compared to the non-contextual baseline with a biasing list size of 100. Moreover, employing RNN-transducer-driven joint decoding further reduces the unbiased word error rate (U-WER), resulting in a more robust network.

【7】 Decoder-only Architecture for Streaming End-to-end Speech Recognition
标题: 流媒体端到端语音识别的纯解码器架构
作者:Emiru Tsunoo,Hayato Futami,Yosuke Kashiwagi,Siddhant Arora,Shinji Watanabe
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:仅解码器语言模型(LM)已成功用于语音处理任务,包括自动语音识别(ASR)。LM具有足够的表现力,并且高效地执行。该效率是ASR的流应用的合适特性。在这项工作中,我们建议使用一个解码器只架构的块流ASR。在我们的方法中,语音特征压缩使用CTC输出和上下文嵌入使用分块语音子网络,并依次提供提示的解码器。解码器在每个块处及时估计输出令牌。为此,我们还提出了一种新的训练方案,使用随机长度的前缀提示,使模型鲁棒截断提示块处理所造成的。实验比较表明,我们提出的仅解码器流ASR在LibriSpeech测试中实现了8%的相对字错误率降低,同时速度是基线模型的两倍。摘要:Decoder-only language models (LMs) have been successfully adopted for speech-processing tasks including automatic speech recognition (ASR). The LMs have ample expressiveness and perform efficiently. This efficiency is a suitable characteristic for streaming applications of ASR. In this work, we propose to use a decoder-only architecture for blockwise streaming ASR. In our approach, speech features are compressed using CTC output and context embedding using blockwise speech subnetwork, and are sequentially provided as prompts to the decoder. The decoder estimates the output tokens promptly at each block. To this end, we also propose a novel training scheme using random-length prefix prompts to make the model robust to the truncated prompts caused by blockwise processing. An experimental comparison shows that our proposed decoder-only streaming ASR achieves 8% relative word error rate reduction in the LibriSpeech test-other set while being twice as fast as the baseline model.

【8】 Text-Queried Target Sound Event Localization
标题: 文本查询目标声音事件定位
作者:Jinzheng Zhao,Xinyuan Qian,Yong Xu,Haohe Liu,Yin Cao,Davide Berghi,Wenwu Wang
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
摘要:声音事件定位和检测(SELD)的目的是确定声音类别的外观,以及它们的到达方向(DOA)。然而,目前的SELD系统只能预测特定类别的活动,例如,DCASE挑战中的13个类别。在本文中,我们提出了文本查询的目标声音事件定位(SEL),一个新的范例,允许用户输入的文本来描述的声音事件,和SEL模型可以预测相关的声音事件的位置。该任务为人机交互提供了一种更友好的方式。我们提供了一个基准研究所提出的任务和模拟房间脉冲响应(RIR)和真正的RIR创建的数据集上进行实验,以验证所提出的方法的有效性。我们希望我们的基准将激发的兴趣和额外的研究文本查询声源定位。摘要:Sound event localization and detection (SELD) aims to determine the appearance of sound classes, together with their Direction of Arrival (DOA). However, current SELD systems can only predict the activities of specific classes, for example, 13 classes in DCASE challenges. In this paper, we propose text-queried target sound event localization (SEL), a new paradigm that allows the user to input the text to describe the sound event, and the SEL model can predict the location of the related sound event. The proposed task presents a more user-friendly way for human-computer interaction. We provide a benchmark study for the proposed task and perform experiments on datasets created by simulated room impulse response (RIR) and real RIR to validate the effectiveness of the proposed methods. We hope that our benchmark will inspire the interest and additional research for text-queried sound source localization.

【9】 Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
标题: 融合音频和元数据嵌入改进了基于数据的音频检索
作者:Paul Primus,Gerhard Widmer
备注:EUSIPCO 2024
链接:点击下载PDF文件
摘要:将原始音频信号与文本描述进行匹配需要理解音频的内容和描述的语义,然后在两种模态之间建立联系。本文研究了一个混合检索系统,利用音频元数据作为一个额外的线索,以了解音频信号的内容之前,匹配它们与文本查询。我们试验了经常附加到音频记录的元数据,例如关键字和自然语言描述,并研究了后期和中期融合策略,以合并音频和元数据。我们的混合方法与关键字元数据和后期融合提高了检索性能超过基于内容的基线由2.36和3.69页。分别在ClothoV2和AudioCaps基准上的mAP@10。摘要:Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid retrieval system that utilizes audio metadata as an additional clue to understand the content of audio signals before matching them with textual queries. We experimented with metadata often attached to audio recordings, such as keywords and natural-language descriptions, and we investigated late and mid-level fusion strategies to merge audio and metadata. Our hybrid approach with keyword metadata and late fusion improved the retrieval performance over a content-based baseline by 2.36 and 3.69 pp. mAP@10 on the ClothoV2 and AudioCaps benchmarks, respectively.

【10】 TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
标题: TacoLM:配备GaTed注意力的编解码器语言模型是高效的Zero-Shot文本到语音合成器
作者:Yakun Song,Zhuo Chen,Xiaofei Wang,Ziyang Ma,Guanrou Yang,Xie Chen
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:神经编解码器语言模型(LM)在零触发(zero-shot)文本到语音(TTS)合成方面表现出强大的能力。然而,编解码器LM由于其自回归性质和文本与音频之间的隐式对齐而经常受到推理速度和稳定性的限制。在这项工作中,为了应对这些挑战,我们引入了一种新的神经编解码器LM,即TacoLM。具体来说,TacoLM引入了门控注意力机制,以提高训练和推理效率,并减少模型大小。同时,每个解码器层都包含一个额外的门控交叉注意层,提高了合成语音的效率和内容准确性。在Librisepeech语料库的评估中,与VALL-E相比,提出的TacoLM实现了更好的单词错误率、说话者相似度和平均意见得分,参数减少了90%,速度提高了5.2倍。演示和代码可在https: ereboas.github.io TacoLM 上获取。摘要:Neural codec language model (LM) has demonstrated strong capability in zero-shot text-to-speech (TTS) synthesis. However, the codec LM often suffers from limitations in inference speed and stability, due to its auto-regressive nature and implicit alignment between text and audio. In this work, to handle these challenges, we introduce a new variant of neural codec LM, namely TacoLM. Specifically, TacoLM introduces a gated attention mechanism to improve the training and inference efficiency and reduce the model size. Meanwhile, an additional gated cross-attention layer is included for each decoder layer, which improves the efficiency and content accuracy of the synthesized speech. In the evaluation of the Librispeech corpus, the proposed TacoLM achieves a better word error rate, speaker similarity, and mean opinion score, with 90% fewer parameters and 5.2 times speed up, compared with VALL-E. Demo and code is available at https: ereboas.github.io TacoLM .

【11】 Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes
标题: 具有粗预测池和声音事件边界盒的自训练和集成频率相关网络
作者:Hyeonuk Nam,Deokki Min,Seungdeok Choi,Inhan Choi,Yong-Hwa Park
备注:DCASE 2024 Challenge Task 4 technical report
链接:点击下载PDF文件
摘要:为了解决声音事件检测(SED)任务,我们提出了频率相关网络(FreDNets),它严重利用频率相关的方法。我们应用频率扭曲和FilterAugment,这是频率相关的数据增强方法。该模型结构由3个分支组成:音频教师-学生Transformer(ATST)分支、BEAT分支和CNN分支,包括部分扩频动态卷积(PDFD)或具有时帧频率方式SE(tfwSE)的挤压和激励(SE)。为了训练具有粗略时间分辨率的MAESTRO标签,我们对MAESTRO数据集的预测应用最大池化。利用最佳集成模型,通过自训练从DESED弱集、DESED未标记集和AudioSet中获取伪标记。AudioSet标签被过滤以集中于高置信度伪标签,并且AudioSet伪标签仅用于在DESED标签上进行训练。我们使用基于变化检测的声音事件边界框(cSEBBs)作为自训练和提交模型的集成模型的后处理。摘要:To tackle sound event detection (SED) task, we propose frequency dependent networks (FreDNets), which heavily leverage frequency-dependent methods. We apply frequency warping and FilterAugment, which are frequency-dependent data augmentation methods. The model architecture consists of 3 branches: audio teacher-student transformer (ATST) branch, BEATs branch and CNN branch including either partial dilated frequency dynamic convolution (PDFD) or squeeze-and-Excitation (SE) with time-frame frequency-wise SE (tfwSE). To train MAESTRO labels with coarse temporal resolution, we apply max pooling on prediction for the MAESTRO dataset. Using best ensemble model, we apply self training to obtain pseudo label from DESED weak set, DESED unlabeled set and AudioSet. AudioSet labels are filtered to focus on high-confidence pseudo labels and AudioSet pseudo labels are used to train on DESED labels only. We used change-detection-based sound event bounding boxes (cSEBBs) as post processing for ensemble models on self training and submission models.

【12】 Exploring the Capability of Mamba in Speech Applications
标题: 探索曼巴在语音应用中的能力
作者:Koichi Miyazaki,Yoshiki Masuyama,Masato Murata
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文探讨了Mamba的能力,最近提出的基于状态空间模型(SSM)的架构,作为一个有竞争力的替代基于transformer的模型。在语音领域,设计良好的基于Transformer的模型,如Conformer和E-Branchformer,已经成为事实上的标准。广泛的评估已经证明了这些基于transformer的模型在广泛的语音任务中的有效性。相比之下,SSM的评估仅限于少数任务,如自动语音识别(ASR)和语音合成。在本文中,我们比较了Mamba与最先进的Transformer变体在各种语音应用中的应用,包括ASR、文本到语音、口语理解和语音摘要。实验评估表明,Mamba实现了与基于transformer的模型相当或更好的性能,并证明了其在长形式语音处理中的效率。摘要:This paper explores the capability of Mamba, a recently proposed architecture based on state space models (SSMs), as a competitive alternative to Transformer-based models. In the speech domain, well-designed Transformer-based models, such as the Conformer and E-Branchformer, have become the de facto standards. Extensive evaluations have demonstrated the effectiveness of these Transformer-based models across a wide range of speech tasks. In contrast, the evaluation of SSMs has been limited to a few tasks, such as automatic speech recognition (ASR) and speech synthesis. In this paper, we compared Mamba with state-of-the-art Transformer variants for various speech applications, including ASR, text-to-speech, spoken language understanding, and speech summarization. Experimental evaluations revealed that Mamba achieves comparable or better performance than Transformer-based models, and demonstrated its efficiency in long-form speech processing.

【13】 Towards Zero-Shot Text-To-Speech for Arabic Dialects
标题: 阿拉伯方言的Zero-Shot文本语音转换
作者:Khai Duy Doan,Abdul Waheed,Muhammad Abdul-Mageed
链接:点击下载PDF文件
摘要:Zero-shot多说话人文本到语音转换(ZERO SHOT MULTI-SPEAKETTTS)系统已经在英语方面取得了进展,然而,由于资源不足,它仍然落后。我们首先通过调整现有的大量数据集来满足语音合成的需求,从而解决阿拉伯语(一种拥有超过4.5亿母语者的语言)的这一差距。此外,我们采用了一套阿拉伯语方言识别模型,探讨在多方言设置的影响,预定义的方言标签,以改善语音-TTS模型。随后,我们微调了XTTS footnote{https: docs.coqui.ai en latest models xtts.html} footnote{https: medium.com machine-learns xtts-v2-new-version-of-the-open-source-text-to-speech-model-af 73914 db 81 f} footnote{https: medium.com @erogol xtts-v1-techincal-notes-eb83ff05bdc}模型,这是一个开源架构。然后,我们在一个包含31个看不见的说话者和一个内部方言数据集的数据集上评估我们的模型。我们的自动和人工评估结果显示出令人信服的性能,同时能够生成方言语音。我们的研究突出了阿拉伯语研究这一新兴领域的重大改进潜力。摘要:Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native speakers, by first adapting a sizeable existing dataset to suit the needs of speech synthesis. Additionally, we employ a set of Arabic dialect identification models to explore the impact of pre-defined dialect labels on improving the ZS-TTS model in a multi-dialect setting. Subsequently, we fine-tune the XTTS footnote{https: docs.coqui.ai en latest models xtts.html} footnote{https: medium.com machine-learns xtts-v2-new-version-of-the-open-source-text-to-speech-model-af73914db81f} footnote{https: medium.com @erogol xtts-v1-techincal-notes-eb83ff05bdc} model, an open-source architecture. We then evaluate our models on a dataset comprising 31 unseen speakers and an in-house dialectal dataset. Our automated and human evaluation results show convincing performance while capable of generating dialectal speech. Our study highlights significant potential for improvements in this emerging area of research in Arabic.

【14】 SNR-Progressive Model with Harmonic Compensation for Low-SNR Speech Enhancement
标题: 用于低SNR语音增强的带调和补偿的SNR递进模型
作者:Zhongshu Hou,Qinwen Hu,Zhanzhong Cao,Ming Tang,Jing Lu
链接:点击下载PDF文件
摘要:尽管在过去十年中取得了重大进展,但基于深度神经网络(DNN)的语音增强(SE)仍然面临着在低信噪比(SNR)条件下恢复语音质量显着下降的挑战。在这封信中,我们提出了一个信噪比渐进语音增强模型与谐波补偿低信噪比SE。从中间输出获得可靠的音高估计,其具有比粗略估计保留更多语音分量同时拥有比输入噪声语音显著更高的SNR的益处。为了更好地恢复谐波,引入了有效的谐波补偿机制。大量的实验证明了我们提出的模型的优点。基于所提出的骨干模型的多模态语音提取系统在ICASSP 2024 MISP挑战赛中排名第一:https: mispchallenge.github.io mispchallenge2023 index.html。摘要:Despite significant progress made in the last decade, deep neural network (DNN) based speech enhancement (SE) still faces the challenge of notable degradation in the quality of recovered speech under low signal-to-noise ratio (SNR) conditions. In this letter, we propose an SNR-progressive speech enhancement model with harmonic compensation for low-SNR SE. Reliable pitch estimation is obtained from the intermediate output, which has the benefit of retaining more speech components than the coarse estimate while possessing a significant higher SNR than the input noisy speech. An effective harmonic compensation mechanism is introduced for better harmonic recovery. Extensive ex-periments demonstrate the advantage of our proposed model. A multi-modal speech extraction system based on the proposed backbone model ranks first in the ICASSP 2024 MISP Challenge: https: mispchallenge.github.io mispchallenge2023 index.html.

【15】 Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
标题: 聆听并移动:提高不可知声音到视频生成中的GAN一致性
作者:Rafael Redondo
Journal-ref:Abstract version published in the ICCV 2023 workshop "AV4D: Visual Learning of Sounds in Spaces"
链接:点击下载PDF文件
摘要:深度生成模型已经证明了创建逼真视听内容的能力,有时由不同性质的域驱动。然而,在视频生成中,平滑的时间动态是一个具有挑战性的问题。这项工作的重点是通用的声音到视频的生成,并提出了三个主要特征,以提高生成对抗模型中的图像质量和时间相干性:三重声音路由方案,用于扩展声音分析的多尺度残差和扩张递归网络,以及用于视频预测的新型递归和定向卷积层。所提出的每个特征在质量和一致性方面都改进了SoTA中通常使用的基线神经架构,其中视频预测层提供了额外的时间细化。摘要:Deep generative models have demonstrated the ability to create realistic audiovisual content, sometimes driven by domains of different nature. However, smooth temporal dynamics in video generation is a challenging problem. This work focuses on generic sound-to-video generation and proposes three main features to enhance both image quality and temporal coherency in generative adversarial models: a triple sound routing scheme, a multi-scale residual and dilated recurrent network for extended sound analysis, and a novel recurrent and directional convolutional layer for video prediction. Each of the proposed features improves, in both quality and coherency, the baseline neural architecture typically used in the SoTA, with the video prediction layer providing an extra temporal refinement.

【16】 Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking
标题: 迈向开放式呼吸声学基础模型:预训练和基准
作者:Yuwei Zhang,Tong Xia,Jing Han,Yu Wu,Georgios Rizos,Yang Liu,Mohammed Mosuily,Jagmohan Chauhan,Cecilia Mascolo
链接:点击下载PDF文件
摘要:None摘要:Respiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Generalizable respiratory acoustic foundation models pretrained with unlabeled data would offer appealing advantages and possibly unlock this impasse. However, given the safety-critical nature of healthcare applications, it is pivotal to also ensure openness and replicability for any proposed foundation model solution. To this end, we introduce OPERA, an OPEn Respiratory Acoustic foundation model pretraining and benchmarking system, as the first approach answering this need. We curate large-scale respiratory audio datasets (~136K samples, 440 hours), pretrain three pioneering foundation models, and build a benchmark consisting of 19 downstream respiratory health tasks for evaluation. Our pretrained models demonstrate superior performance (against existing acoustic models pretrained with general audio on 16 out of 19 tasks) and generalizability (to unseen datasets and new respiratory audio modalities). This highlights the great promise of respiratory acoustic foundation models and encourages more studies using OPERA as an open resource to accelerate research on respiratory audio for health. The system is accessible from https: github.com evelyn0414 OPERA.

【17】 Speech Representation Analysis based on Inter- and Intra-Model Similarities
标题: 基于模型间和模型内相似性的语音表示分析
作者:Yassine El Kheir,Ahmed Ali,Shammur Absar Chowdhury
备注:5 pages, Accepted to appear in ICASSP XAI-SA Workshop
链接:点击下载PDF文件
摘要:自监督模型已经彻底改变了语音处理,在有限的资源下在各种任务中实现了新的性能水平。然而,这些模型的内部运作仍然不透明。在本文中,我们的目标是分析这些基础模型的编码上下文表示的基础上,他们之间和模型内的相似性,独立于任何外部注释和特定于任务的约束。我们研究了不同的SSL模型,这些模型的训练模式不同-对比模型(Wav2Vec2.0)和预测模型(HuBERT);以及模型大小(基础和大型)。我们在不同层次的信息定位 分布上探索这些模型,包括(i)单个神经元;(ii)层表示;(iii)注意力权重和(iv)将表征与其微调的对应物进行比较。我们的结果强调,这些模型收敛到相似的表征子空间,但不收敛到相似的神经元局部概念 脚注{概念表示知识的连贯片段,例如 一个包含某些对象作为元素的类,其中对象具有某些属性。我们公开了代码,以促进进一步的研究,我们公开发布了我们的代码。摘要:Self-supervised models have revolutionized speech processing, achieving new levels of performance in a wide variety of tasks with limited resources. However, the inner workings of these models are still opaque. In this paper, we aim to analyze the encoded contextual representation of these foundation models based on their inter- and intra-model similarity, independent of any external annotation and task-specific constraint. We examine different SSL models varying their training paradigm -- Contrastive (Wav2Vec2.0) and Predictive models (HuBERT); and model sizes (base and large). We explore these models on different levels of localization distributivity of information including (i) individual neurons; (ii) layer representation; (iii) attention weights and (iv) compare the representations with their finetuned counterparts.Our results highlight that these models converge to similar representation subspaces but not to similar neuron-localized concepts footnote{A concept represents a coherent fragment of knowledge, such as a class containing certain objects as elements, where the objects have certain properties. We made the code publicly available for facilitating further research, we publicly released our code.

【18】 AudioBench: A Universal Benchmark for Audio Large Language Models
标题: AudioBench:音频大型语言模型的通用基准
作者:Bin Wang,Xunlong Zou,Geyu Lin,Shuo Sun,Zhuohan Liu,Wenyu Zhang,Zhengyuan Liu,AiTi Aw,Nancy F. Chen
备注:20 pages; Preprint; Code: this https URL
链接:点击下载PDF文件
摘要:我们介绍AudioBench,一个新的基准测试,旨在评估音频大语言模型(AudioLLM)。AudioBench包含8个不同的任务和26个精心挑选或新策划的数据集,专注于语音理解,语音解释和音频场景理解。尽管大型语言模型(包括多模式版本)的发展迅速,但在全面评估其能力的综合基准方面存在着重大差距。AudioBench通过提供相关的数据集和评估指标来解决这一差距。在我们的研究中,我们评估了四种模型在各个方面的能力,发现没有一种模型在所有任务中都表现出色。我们概述了AudioLLM的研究前景,并预计我们的开源代码,数据和排行榜将为未来的模型开发提供一个强大的测试平台。摘要:We introduce AudioBench, a new benchmark designed to evaluate audio large language models (AudioLLMs). AudioBench encompasses 8 distinct tasks and 26 carefully selected or newly curated datasets, focusing on speech understanding, voice interpretation, and audio scene understanding. Despite the rapid advancement of large language models, including multimodal versions, a significant gap exists in comprehensive benchmarks for thoroughly evaluating their capabilities. AudioBench addresses this gap by providing relevant datasets and evaluation metrics. In our study, we evaluated the capabilities of four models across various aspects and found that no single model excels consistently across all tasks. We outline the research outlook for AudioLLMs and anticipate that our open-source code, data, and leaderboard will offer a robust testbed for future model developments.

【19】 Predicting Individual Depression Symptoms from Acoustic Features During Speech
标题: 根据言语过程中的声学特征预测个人抑郁症状
作者:Sebastian Rodriguez,Sri Harsha Dumpala,Katerina Dikaios,Sheri Rempel,Rudolf Uher,Sageev Oore
链接:点击下载PDF文件
摘要:目前的自动抑郁症检测系统直接提供预测,而不依赖于临床抑郁症评定量表中所示的抑郁症的个体症状 项目。相比之下,临床医生在临床环境中评估抑郁症评定量表中的每个项目,从而隐含地为抑郁症诊断提供更详细的理由。在这项工作中,我们做了第一步,使用语音的声学特征来预测个别项目的抑郁症评级量表,然后获得最终的抑郁症预测。为此,我们使用卷积(CNN)和循环(长短期记忆(LSTM))神经网络。我们考虑不同的方法来学习语音的时间背景。此外,我们分析了两个变量的投票方案,个别项目的预测和抑郁症的检测。我们还包括一个动画可视化,显示随着时间的推移,随着讲话的进展项目预测的例子。摘要:Current automatic depression detection systems provide predictions directly without relying on the individual symptoms items of depression as denoted in the clinical depression rating scales. In contrast, clinicians assess each item in the depression rating scale in a clinical setting, thus implicitly providing a more detailed rationale for a depression diagnosis. In this work, we make a first step towards using the acoustic features of speech to predict individual items of the depression rating scale before obtaining the final depression prediction. For this, we use convolutional (CNN) and recurrent (long short-term memory (LSTM)) neural networks. We consider different approaches to learning the temporal context of speech. Further, we analyze two variants of voting schemes for individual item prediction and depression detection. We also include an animated visualization that shows an example of item prediction over time as the speech progresses.

【20】 Real-time Speech Summarization for Medical Conversations
标题: 医疗对话的实时语音总结
作者:Khai Le-Duc,Khai-Nguyen Nguyen,Long Vo-Dang,Truong-Son Hy
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:在医患对话中,识别医学相关信息至关重要,因此需要进行对话摘要。在这项工作中,我们提出了第一个可部署的实时语音摘要系统在工业中的实际应用,它产生一个本地摘要后,每N个语音话语在一个对话和一个全球性的总结后,结束对话。我们的系统可以从商业角度增强用户体验,同时从技术角度降低计算成本。其次,我们介绍了VietMed-Sum,据我们所知,它是第一个用于医疗对话的语音摘要数据集。第三,我们是第一个利用LLM和人类注释器协同创建医疗会话摘要的黄金标准和综合摘要的公司。最后,我们提出了最先进的模型VietMed-Sum的基线结果。所有代码、数据(英语翻译和越南语)和模型均可在线获取:https: github.com leduckhai MultiMed摘要:In doctor-patient conversations, identifying medically relevant information is crucial, posing the need for conversation summarization. In this work, we propose the first deployable real-time speech summarization system for real-world applications in industry, which generates a local summary after every N speech utterances within a conversation and a global summary after the end of a conversation. Our system could enhance user experience from a business standpoint, while also reducing computational costs from a technical perspective. Secondly, we present VietMed-Sum which, to our knowledge, is the first speech summarization dataset for medical conversations. Thirdly, we are the first to utilize LLM and human annotators collaboratively to create gold standard and synthetic summaries for medical conversation summarization. Finally, we present baseline results of state-of-the-art models on VietMed-Sum. All code, data (English-translated and Vietnamese) and models are available online: https: github.com leduckhai MultiMed

【21】 The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
标题: 音乐大师或音乐大师,大型语言模型的大规模音乐评估基准
作者:Jiajia Li,Lu Yang,Mingni Tang,Cong Chen,Zuchao Li,Ping Wang,Hai Zhao
备注:Accepted to ACL-Findings 2024
链接:点击下载PDF文件
摘要:Benchmark在评估大型语言模型(LLM)的进步方面发挥着关键作用。虽然已经提出了许多基准来评估LLM的能力,但值得注意的是缺乏一个专门的基准来评估他们的音乐能力。为了解决这一差距,我们提出了ZIQI-Eval,这是一个全面的大规模音乐基准,专门用于评估LLM的音乐相关能力。ZIQI-Eval包含广泛的问题,涵盖10个主要类别和56个子类别,导致超过14,000个精心策划的数据条目。通过利用ZIQI-Eval,我们对16个LLM进行了综合评估,以评估和分析LLM在音乐领域的表现。结果表明,所有的LLM表现不佳的ZIQI-Eval基准,这表明显着的空间,改善他们的音乐能力。通过ZIQI-Eval,我们的目标是提供一个标准化和强大的评估框架,以促进对LLM音乐相关能力的全面评估。该数据集可在GitHub footnote{https: github.com zcli-charlie ZIQI-Eval}和HuggingFace footnote{https: huggingface.co MyTH-Lab ZIQI-Eval}获得。摘要:Benchmark plays a pivotal role in assessing the advancements of large language models (LLMs). While numerous benchmarks have been proposed to evaluate LLMs' capabilities, there is a notable absence of a dedicated benchmark for assessing their musical abilities. To address this gap, we present ZIQI-Eval, a comprehensive and large-scale music benchmark specifically designed to evaluate the music-related capabilities of LLMs. ZIQI-Eval encompasses a wide range of questions, covering 10 major categories and 56 subcategories, resulting in over 14,000 meticulously curated data entries. By leveraging ZIQI-Eval, we conduct a comprehensive evaluation over 16 LLMs to evaluate and analyze LLMs' performance in the domain of music. Results indicate that all LLMs perform poorly on the ZIQI-Eval benchmark, suggesting significant room for improvement in their musical capabilities. With ZIQI-Eval, we aim to provide a standardized and robust evaluation framework that facilitates a comprehensive assessment of LLMs' music-related abilities. The dataset is available at GitHub footnote{https: github.com zcli-charlie ZIQI-Eval} and HuggingFace footnote{https: huggingface.co datasets MYTH-Lab ZIQI-Eval}.

【22】 AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
标题: 基于人工智能的无人机在灾难环境中协助人类救援:挑战和机遇
作者:Narek Papyan,Michel Kulhandjian,Hovannes Kulhandjian,Levon Hakob Aslanyan
Journal-ref:Pattern Recognit. Image Anal. 34 (2024)
链接:点击下载PDF文件
摘要:在这项调查中,我们专注于利用基于无人机的系统来检测个人,特别是通过识别人类的尖叫声和其他求救信号。这项研究在灾后情景中具有重要意义,包括地震,飓风,军事冲突,野火等事件。这些无人机能够在受灾地区上空盘旋,这可能对救援队直接进入造成挑战。无人驾驶飞行器(UAV),通常被称为无人机,经常被部署用于灾害情况下的搜索和救援任务。通常情况下,无人机会捕捉空中图像,以评估结构性损坏并确定灾害的程度。他们还采用热成像技术来检测身体热量信号,这可以帮助定位个人。在某些情况下,大型无人机被用来向被困在孤立的受灾地区的人们运送必需品。在我们的讨论中,我们深入研究了与通过航空声学定位人类相关的独特挑战。听觉系统必须区分人类的哭声和自然发生的声音,如动物的叫声和风声。此外,它应该能够识别与信号相关的不同模式,如喊叫,鼓掌或人们试图向救援队发出信号的其他方式。为了应对这一挑战,一种解决方案涉及利用人工智能(AI)来分析声音频率并识别常见的音频签名。基于深度学习的网络,如卷积神经网络(CNN),可以使用这些特征进行训练,以过滤掉无人机电机和其他环境因素产生的噪声。此外,采用基于麦克风阵列信号的波达方向(DOA)等信号处理技术可以提高跟踪人类噪声源的精度。摘要:In this survey we are focusing on utilizing drone-based systems for the detection of individuals, particularly by identifying human screams and other distress signals. This study has significant relevance in post-disaster scenarios, including events such as earthquakes, hurricanes, military conflicts, wildfires, and more. These drones are capable of hovering over disaster-stricken areas that may be challenging for rescue teams to access directly. Unmanned aerial vehicles (UAVs), commonly referred to as drones, are frequently deployed for search-and-rescue missions during disaster situations. Typically, drones capture aerial images to assess structural damage and identify the extent of the disaster. They also employ thermal imaging technology to detect body heat signatures, which can help locate individuals. In some cases, larger drones are used to deliver essential supplies to people stranded in isolated disaster-stricken areas. In our discussions, we delve into the unique challenges associated with locating humans through aerial acoustics. The auditory system must distinguish between human cries and sounds that occur naturally, such as animal calls and wind. Additionally, it should be capable of recognizing distinct patterns related to signals like shouting, clapping, or other ways in which people attempt to signal rescue teams. To tackle this challenge, one solution involves harnessing artificial intelligence (AI) to analyze sound frequencies and identify common audio signatures. Deep learning-based networks, such as convolutional neural networks (CNNs), can be trained using these signatures to filter out noise generated by drone motors and other environmental factors. Furthermore, employing signal processing techniques like the direction of arrival (DOA) based on microphone array signals can enhance the precision of tracking the source of human noises.

【23】 Revisiting Interpolation Augmentation for Speech-to-Text Generation
标题: 重新审视语音到文本生成的内插增强
作者:Chen Xu,Jie Wang,Xiaoqian Liu,Qianqian Dong,Chunliang Zhang,Tong Xiao,Jingbo Zhu,Dapeng Man,Wu Yang
备注:ACL 2024 Findings
链接:点击下载PDF文件
摘要:语音到文本(S2T)生成系统在低资源场景中经常面临挑战,主要是由于缺乏广泛的标记数据集。一种新兴的解决方案是通过内插输入和标签来构建虚拟训练样本,这显著增强了系统在其他领域的泛化能力。尽管它的潜力,这项技术在S2T任务中的应用仍然没有得到充分的探索。在本文中,我们深入到插值增强的效用,引导几个关键问题。我们的研究结果表明,在插值增强中采用适当的策略可以显着提高不同任务,架构和数据规模的性能,为资源受限环境中更强大的S2T系统提供了一条有前途的途径。摘要:Speech-to-text (S2T) generation systems frequently face challenges in low-resource scenarios, primarily due to the lack of extensive labeled datasets. One emerging solution is constructing virtual training samples by interpolating inputs and labels, which has notably enhanced system generalization in other domains. Despite its potential, this technique's application in S2T tasks has remained under-explored. In this paper, we delve into the utility of interpolation augmentation, guided by several pivotal questions. Our findings reveal that employing an appropriate strategy in interpolation augmentation significantly enhances performance across diverse tasks, architectures, and data scales, offering a promising avenue for more robust S2T systems in resource-constrained settings.

【24】 Multimodal Segmentation for Vocal Tract Modeling
标题: 用于人声建模的多模式分割
作者:Rishi Jain,Bohan Yu,Peter Wu,Tejas Prabhune,Gopala Anumanchipalli
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:声道的精确建模对于构建可解释的语音处理和语言学的发音表示是必要的。然而,声道建模是具有挑战性的,因为许多内部发音被外部运动捕捉技术遮挡。实时磁共振成像(RT-MRI)允许测量语音期间内部发音器官的精确运动,但是由于耗时且计算昂贵的标记方法,MRI的注释数据集的大小受到限制。我们首先提出了一个深度标记策略的RT-MRI视频使用的视觉分割方法。然后,我们引入了一个多模态算法,使用音频,以改善分割的发音器官。总之,我们为MRI视频分割中的声道建模设定了一个新的基准,并使用它来发布75扬声器RT-MRI数据集的标签,将声道的标记公共RT-MRI数据量增加了9倍以上。代码和数据集标签可以在 url{rishiraij.github.io multimodal-mri-avatar }找到。摘要:Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articulators are occluded from external motion capture technologies. Real-time magnetic resonance imaging (RT-MRI) allows measuring precise movements of internal articulators during speech, but annotated datasets of MRI are limited in size due to time-consuming and computationally expensive labeling methods. We first present a deep labeling strategy for the RT-MRI video using a vision-only segmentation approach. We then introduce a multimodal algorithm using audio to improve segmentation of vocal articulators. Together, we set a new benchmark for vocal tract modeling in MRI video segmentation and use this to release labels for a 75-speaker RT-MRI dataset, increasing the amount of labeled public RT-MRI data of the vocal tract by over a factor of 9. The code and dataset labels can be found at url{rishiraij.github.io multimodal-mri-avatar }.

【25】 Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
标题: 使用GAN和集成的未对齐清洁数据改进无监督清洁到渲染的吉他音调转换
作者:Yu-Hua Chen,Woosung Choi,Wei-Hsiang Liao,Marco Martínez-Ramírez,Kin Wai Cheuk,Yuki Mitsufuji,Jyh-Shing Roger Jang,Yi-Hsuan Yang
备注:Accepted to DAFx 2024
链接:点击下载PDF文件
摘要:近年来,人们越来越关注将深度学习方法应用于吉他放大器或效果踏板的建模。现有的方法主要是基于监督的方法,需要未处理和渲染的音频的时间对齐的数据对。然而,由于创建数据对所涉及的复杂过程,这种方法不能很好地扩展。Wright等人最近的一项工作探索了利用未配对数据进行训练的潜力,使用基于生成对抗网络(GAN)的框架。本文通过在GAN中使用更高级的鉴别器,并使用更多的未配对数据进行训练,扩展了他们的工作。具体来说,从神经声码器的最新进展中汲取灵感,我们在基于GAN的吉他放大器模型中采用了两组鉴别器,一组基于多尺度鉴别器(MSD),另一组基于多周期鉴别器(MPD)。此外,我们尝试将未处理的音频信号添加到训练数据中,这些信号没有目标音调的相应渲染音频,以查看GAN模型从未配对数据中受益多少。我们的实验表明,提出的两个扩展有助于建模的低增益和高增益吉他放大器。摘要:Recent years have seen increasing interest in applying deep learning methods to the modeling of guitar amplifiers or effect pedals. Existing methods are mainly based on the supervised approach, requiring temporally-aligned data pairs of unprocessed and rendered audio. However, this approach does not scale well, due to the complicated process involved in creating the data pairs. A very recent work done by Wright et al. has explored the potential of leveraging unpaired data for training, using a generative adversarial network (GAN)-based framework. This paper extends their work by using more advanced discriminators in the GAN, and using more unpaired data for training. Specifically, drawing inspiration from recent advancements in neural vocoders, we employ in our GAN-based model for guitar amplifier modeling two sets of discriminators, one based on multi-scale discriminator (MSD) and the other multi-period discriminator (MPD). Moreover, we experiment with adding unprocessed audio signals that do not have the corresponding rendered audio of a target tone to the training data, to see how much the GAN model benefits from the unpaired data. Our experiments show that the proposed two extensions contribute to the modeling of both low-gain and high-gain guitar amplifiers.

【26】 Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
标题: 用于平衡多方面发音评估的声学特征混合
作者:Heejin Do,Wonjun Lee,Gary Geunbae Lee
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:在自动发音评估中,最近的重点逐渐在于评估多个方面以提供丰富的反馈。然而,获取非母语学习者语音的多方面分数标记数据提出了挑战,而且,它往往会导致分数不平衡分布。在本文中,我们提出了两个声学特征混合策略,线性和非线性插值与批平均功能,以解决数据稀缺和分数标签的不平衡。首先,使用良好的发音作为声学特征,我们量身定制的混合设计,以适应发音评估。此外,我们通过将语音识别结果与原始答案音素进行比较来集成细粒度错误率特征,从而给出发音错误的直接提示。声学特征的有效混合显着增强了speechocean 762数据集的整体评分性能,并且详细的分析强调了我们预测看不见的失真的潜力。摘要:In automated pronunciation assessment, recent emphasis progressively lies on evaluating multiple aspects to provide enriched feedback. However, acquiring multi-aspect-score labeled data for non-native language learners' speech poses challenges; moreover, it often leads to score-imbalanced distributions. In this paper, we propose two Acoustic Feature Mixup strategies, linearly and non-linearly interpolating with the in-batch averaged feature, to address data scarcity and score-label imbalances. Primarily using goodness-of-pronunciation as an acoustic feature, we tailor mixup designs to suit pronunciation assessment. Further, we integrate fine-grained error-rate features by comparing speech recognition results with the original answer phonemes, giving direct hints for mispronunciation. Effective mixing of the acoustic features notably enhances overall scoring performances on the speechocean762 dataset, and detailed analysis highlights our potential to predict unseen distortions.

【27】 PI-Whisper: An Adaptive and Incremental ASR Framework for Diverse and Evolving Speaker Characteristics
标题: PI-Whisper:一种适应性和增量式的ASB框架,用于多样化和不断发展的说话者特征
作者:Amir Nassereldine,Dancheng Liu,Chenhui Xu,Jinjun Xiong
备注:11 pages, 3 figures
链接:点击下载PDF文件
摘要:随着基于边缘的自动语音识别(ASR)技术在智能和个性化助理的开发中变得越来越普遍,这些资源受限的ASR模型必须解决三个重要挑战,即,适应性、递增性和包容性。我们提出了一种新的ASR框架,PI耳语,在这项工作中,并显示它可以通过实时识别不同的扬声器的特性,如何自适应地提高ASR的识别能力,这样的适应可以逐步进行,而无需重复的再训练,以及它如何可以提高不同的扬声器组的公平性和公正性。更令人印象深刻的是,我们提出的PI Whisper框架实现了所有这些良好的属性,同时仍然实现了最先进的准确性,字错误率(WER)降低了13.7%,并且相对于计算资源具有线性可扩展性。摘要:As edge-based automatic speech recognition (ASR) technologies become increasingly prevalent for the development of intelligent and personalized assistants, three important challenges must be addressed for these resource-constrained ASR models, i.e., adaptivity, incrementality, and inclusivity. We propose a novel ASR framework, PI-Whisper, in this work and show how it can improve an ASR's recognition capabilities adaptively by identifying different speakers' characteristics in real-time, how such an adaption can be performed incrementally without repetitive retraining, and how it can improve the equity and fairness for diverse speaker groups. More impressively, our proposed PI-Whisper framework attains all of these nice properties while still achieving state-of-the-art accuracy with up to 13.7% reduction of the word error rate (WER) with linear scalability with respect to computing resources.

【28】 Generating Music with Structure Using Self-Similarity as Attention
标题: 利用自相似性作为注意力生成具有结构的音乐
作者:Sophia Hager,Kathleen Hablutzel,Katherine Kinnaird
链接:点击下载PDF文件
摘要:尽管在深度学习和生成式人工智能方面有了创新,但创建长期结构以及音乐作品中常见的重复结构层仍然是音乐生成中的一个公开挑战。我们提出了一个注意力层,它使用一种新的方法,将用户提供的自相似性矩阵应用于先前的时间步长,并在我们的相似性激励神经生成器(SING)系统中进行了演示,SING系统是一个具有两层的深度学习自主音乐生成系统。第一个是普通的长短期记忆层,第二个是建议的注意力层。在生成过程中,这种注意力机制将来自模板作品的建议结构强加于生成的音乐上。我们训练SING的MAESTRO数据集上使用一种新的变量的方法,并比较其性能相同的模型没有注意力机制。我们提出的注意力机制的加入显着提高了网络复制特定结构的能力,并且它在看不见的测试集上的表现比没有注意力机制的模型更好。摘要:Despite the innovations in deep learning and generative AI, creating long term structure as well as the layers of repeated structure common in musical works remains an open challenge in music generation. We propose an attention layer that uses a novel approach applying user-supplied self-similarity matrices to previous time steps, and demonstrate it in our Similarity Incentivized Neural Generator (SING) system, a deep learning autonomous music generation system with two layers. The first is a vanilla Long Short Term Memory layer, and the second is the proposed attention layer. During generation, this attention mechanism imposes a suggested structure from a template piece on the generated music. We train SING on the MAESTRO dataset using a novel variable batching method, and compare its performance to the same model without the attention mechanism. The addition of our proposed attention mechanism significantly improves the network's ability to replicate specific structures, and it performs better on an unseen test set than a model without the attention mechanism.

【29】 R&B -- Rhythm and Brain: Cross-subject Decoding of Music from Human Brain Activity
标题: R & B --节奏与大脑:从人脑活动中解读音乐的跨学科
作者:Matteo Ferrante,Matteo Ciferri,Nicola Toschi
备注:The first two authors contributed equally to this work
链接:点击下载PDF文件
摘要:音乐是一种普遍现象,深刻影响着不同文化的人类体验。这项研究调查了音乐是否可以从人类大脑活动中解码,在其感知过程中使用功能性MRI(fMRI)进行测量。利用最近在广泛的数据集和预先训练的计算模型方面的进展,我们构建了神经数据和音乐刺激的潜在表征之间的映射。我们的方法集成了功能和解剖对齐技术,以促进跨学科的解码,解决了fMRI数据的低时间分辨率和信噪比(SNR)所带来的挑战。从GTZan fMRI数据集开始,其中5名参与者在记录大脑活动的同时听取了来自10种不同流派的540种音乐刺激,我们使用CLAP(对比听觉-音频预训练)模型来提取音乐刺激的潜在表征,并开发了体素编码模型来识别对这些刺激做出反应的大脑区域。通过对预测和实际大脑活动之间的关联应用阈值,我们确定了特定的感兴趣区域(ROI),这些区域可以被解释为音乐处理中的关键角色。我们的解码管道主要基于检索,采用线性映射将大脑活动投射到相应的CLAP特征。这使我们能够预测和检索与fMRI数据来源最相似的音乐刺激。我们的结果证明了最先进的识别精度,我们的方法显着优于现有的方法。我们的研究结果表明,基于神经的音乐检索系统可以实现个性化推荐和治疗应用。未来的工作可以使用更高的时间分辨率神经成像和生成模型来提高解码准确性,并探索音乐感知和情感的神经基础。摘要:Music is a universal phenomenon that profoundly influences human experiences across cultures. This study investigates whether music can be decoded from human brain activity measured with functional MRI (fMRI) during its perception. Leveraging recent advancements in extensive datasets and pre-trained computational models, we construct mappings between neural data and latent representations of musical stimuli. Our approach integrates functional and anatomical alignment techniques to facilitate cross-subject decoding, addressing the challenges posed by the low temporal resolution and signal-to-noise ratio (SNR) in fMRI data. Starting from the GTZan fMRI dataset, where five participants listened to 540 musical stimuli from 10 different genres while their brain activity was recorded, we used the CLAP (Contrastive Language-Audio Pretraining) model to extract latent representations of the musical stimuli and developed voxel-wise encoding models to identify brain regions responsive to these stimuli. By applying a threshold to the association between predicted and actual brain activity, we identified specific regions of interest (ROIs) which can be interpreted as key players in music processing. Our decoding pipeline, primarily retrieval-based, employs a linear map to project brain activity to the corresponding CLAP features. This enables us to predict and retrieve the musical stimuli most similar to those that originated the fMRI data. Our results demonstrate state-of-the-art identification accuracy, with our methods significantly outperforming existing approaches. Our findings suggest that neural-based music retrieval systems could enable personalized recommendations and therapeutic applications. Future work could use higher temporal resolution neuroimaging and generative models to improve decoding accuracy and explore the neural underpinnings of music perception and emotion.

【30】 System Description for the Displace Speaker Diarization Challenge 2023
标题: 2023年Displace Speaker Diaration Challenge的系统描述
作者:Ali Aliyev
链接:点击下载PDF文件
摘要:本文介绍了我们的解决方案,对话环境中的说话者和语言的日记化挑战(位移2023)。我们使用VAD的组合来寻找语音片段,基于Resnet架构的CNN用于从这些片段中提取特征,以及谱聚类用于特征聚类。即使它没有使用印地语进行训练,所描述的算法也实现了以下指标:DER 27。1%,27。4%,分别用于数据集的开发和第1阶段评估部分。摘要:This paper describes our solution for the Diarization of Speaker and Language in Conversational Environments Challenge (Displace 2023). We used a combination of VAD for finding segfments with speech, Resnet architecture based CNN for feature extraction from these segments, and spectral clustering for features clustering. Even though it was not trained with using Hindi, the described algorithm achieves the following metrics: DER 27. 1% and DER 27. 4%, on the development and phase-1 evaluation parts of the dataset, respectively.

【31】 Improving Text-To-Audio Models with Synthetic Captions
标题: 使用合成字幕改进文本到音频模型
作者:Zhifeng Kong,Sang-gil Lee,Deepanway Ghosal,Navonil Majumder,Ambuj Mehrish,Rafael Valle,Soujanya Poria,Bryan Catanzaro
链接:点击下载PDF文件
摘要:为文本到音频模型获得高质量的训练数据,特别是字幕,是一个公开的挑战。虽然先前的方法已经利用了纯文本语言模型来增强和改进字幕,但这些方法在音频和字幕之间的规模和连贯性方面存在局限性。在这项工作中,我们提出了一个音频字幕管道,使用 textit{音频语言模型}来合成准确和多样化的字幕音频规模。我们利用这个管道为AudioSet生成一个名为 texttt{AF-AudioSet}的合成字幕数据集,然后评估这些合成字幕的预训练文本到音频模型的好处。通过对AudioCaps和MusicCaps的系统评估,我们发现利用我们的管道和合成字幕可以显著提高音频生成质量,实现新的 textit{最先进的}。摘要:It is an open challenge to obtain high quality training data, especially captions, for text-to-audio models. Although prior methods have leveraged textit{text-only language models} to augment and improve captions, such methods have limitations related to scale and coherence between audio and captions. In this work, we propose an audio captioning pipeline that uses an textit{audio language model} to synthesize accurate and diverse captions for audio at scale. We leverage this pipeline to produce a dataset of synthetic captions for AudioSet, named texttt{AF-AudioSet}, and then evaluate the benefit of pre-training text-to-audio models on these synthetic captions. Through systematic evaluations on AudioCaps and MusicCaps, we find leveraging our pipeline and synthetic captions leads to significant improvements on audio generation quality, achieving a new textit{state-of-the-art}.


机器翻译,仅供参考