今日论文合集:cs.SD语音39篇,eess.AS音频处理55篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Multimodal Contextualized Semantic Parsing from Speech
标题: 语音的多模式上下文化语义解析
作者:Jordan Voas,Raymond Mooney,David Harwath
备注:10 Pages, 3 figures, ACL 2024 Main
链接:点击下载PDF文件
摘要:我们介绍了上下文环境中的语义解析(SPICE),旨在提高人工代理的上下文意识的任务,通过整合多模态输入与先前的上下文。SPICE超越了传统的语义解析,提供了一个结构化的,可解释的框架,用于动态更新代理的知识与新信息,反映了人类沟通的复杂性。我们开发的VG-SPICE数据集,精心制作的挑战代理与视觉场景图建设从口语对话交流,突出语音和视觉数据集成。我们还提出了视听对话场景解析器(AViD-SP)开发的VG-SPICE上使用。这些创新旨在改善多模式信息处理和整合。VG-SPICE数据集和AViD-SP模型都是公开的。摘要:We introduce Semantic Parsing in Contextual Environments (SPICE), a task designed to enhance artificial agents' contextual awareness by integrating multimodal inputs with prior contexts. SPICE goes beyond traditional semantic parsing by offering a structured, interpretable framework for dynamically updating an agent's knowledge with new information, mirroring the complexity of human communication. We develop the VG-SPICE dataset, crafted to challenge agents with visual scene graph construction from spoken conversational exchanges, highlighting speech and visual data integration. We also present the Audio-Vision Dialogue Scene Parser (AViD-SP) developed for use on VG-SPICE. These innovations aim to improve multimodal information processing and integration. Both the VG-SPICE dataset and the AViD-SP model are publicly available.

【2】 Controlling Emotion in Text-to-Speech with Natural Language Prompts
标题: 使用自然语言脚本控制文本到语音中的情感
作者:Thomas Bott,Florian Lux,Ngoc Thang Vu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,由于对自然语言的直观使用,提示已迅速成为引导生成式机器学习模型输出的标准方式之一。在这项工作中,我们提出了一个以嵌入为条件的系统,该嵌入源自作为提示的情感丰富的文本。因此,扬声器和提示嵌入的联合表示被集成在基于变换器的架构内的多个点处。我们的方法在合并的情感语音和文本数据集上进行训练,并在每次训练迭代中改变提示,以提高模型的泛化能力。客观和主观的评价结果表明,有条件的合成系统的能力,准确地转移到语音提示中存在的情绪。同时,说话人身份的精确易处理性以及整体高语音质量和可懂度得以保持。摘要:In recent years, prompting has quickly become one of the standard ways of steering the outputs of generative machine learning models, due to its intuitive use of natural language. In this work, we propose a system conditioned on embeddings derived from an emotionally rich text that serves as prompt. Thereby, a joint representation of speaker and prompt embeddings is integrated at several points within a transformer-based architecture. Our approach is trained on merged emotional speech and text datasets and varies prompts in each training iteration to increase the generalization capabilities of the model. Objective and subjective evaluation results demonstrate the ability of the conditioned synthesis system to accurately transfer the emotions present in a prompt to speech. At the same time, precise tractability of speaker identities as well as overall high speech quality and intelligibility are maintained.

【3】 Meta Learning Text-to-Speech Synthesis in over 7000 Languages
标题: Meta学习7000多种语言的文本到语音合成
作者:Florian Lux,Sarina Meyer,Lyonel Behringer,Frank Zalkow,Phat Do,Matt Coler,Emanuël A. P. Habets,Ngoc Thang Vu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:在这项工作中,我们承担了一个具有挑战性的任务,建立一个单一的文本到语音合成系统,能够生成语音超过7000种语言,其中许多缺乏足够的数据,传统的TTS开发。通过利用大规模多语言预训练和Meta学习的新集成来近似语言表示,我们的方法可以在没有任何可用数据的情况下实现语言的zero-shot语音合成。我们通过客观的措施和人类的评价,在不同的语言景观验证我们的系统的性能。通过公开发布我们的代码和模型,我们的目标是赋予社区有限的语言资源,并促进语音技术领域的进一步创新。摘要:In this work, we take on the challenging task of building a single text-to-speech synthesis system that is capable of generating speech in over 7000 languages, many of which lack sufficient data for traditional TTS development. By leveraging a novel integration of massively multilingual pretraining and meta learning to approximate language representations, our approach enables zero-shot speech synthesis in languages without any available data. We validate our system's performance through objective measures and human evaluation across a diverse linguistic landscape. By releasing our code and models publicly, we aim to empower communities with limited linguistic resources and foster further innovation in the field of speech technology.

【4】 MOSA: Music Motion with Semantic Annotation Dataset for Cross-Modal Music Processing
标题: MOSA:用于跨模式音乐处理的具有语义注释数据集的音乐运动
作者:Yu-Fen Huang,Nikki Moran,Simon Coleman,Jon Kelly,Shun-Hwa Wei,Po-Yin Chen,Yun-Hsin Huang,Tsung-Ping Chen,Yu-Chia Kuo,Yu-Chi Wei,Chih-Hsuan Li,Da-Yu Huang,Hsuan-Kai Kao,Ting-Wei Lin,Li Su
备注:IEEEACM Transactions on Audio, Speech, and Language Processing, 2024. 14 pages, 7 figures. Dataset is available on: this https URL and this https URL
链接:点击下载PDF文件
摘要:在跨模态音乐处理中,视觉,听觉和语义内容之间的翻译开辟了新的可能性和挑战。这种变革性计划的构建取决于具有全面数据基础设施的基准语料库。特别是,组装一个大规模的跨模态数据集提出了重大挑战。在本文中,我们提出了MOSA(音乐运动与语义注释)数据集,其中包含高质量的3-D运动捕捉数据,对齐的音频记录,并注意到由23个专业音乐家的742个专业音乐表演的音高,节拍,乐句,动态,清晰度和和声的语义注释,包括超过30小时和570 K个音符的数据。据我们所知,这是迄今为止最大的具有音符级注释的跨模态音乐数据集。为了演示MOSA数据集的使用,我们提出了几个创新的跨模态音乐信息检索(MIR)和音乐内容生成任务,包括从音频、视频和运动数据中检测节拍、强拍、短语和表达内容,以及从给定的音乐音频中生成音乐家的身体动作。该数据集和代码与本出版物一起提供(https: github.com yufenhuang MOSA-Music-mOtion-and-Semantic-Annotation-dataset)。摘要:In cross-modal music processing, translation between visual, auditory, and semantic content opens up new possibilities as well as challenges. The construction of such a transformative scheme depends upon a benchmark corpus with a comprehensive data infrastructure. In particular, the assembly of a large-scale cross-modal dataset presents major challenges. In this paper, we present the MOSA (Music mOtion with Semantic Annotation) dataset, which contains high quality 3-D motion capture data, aligned audio recordings, and note-by-note semantic annotations of pitch, beat, phrase, dynamic, articulation, and harmony for 742 professional music performances by 23 professional musicians, comprising more than 30 hours and 570 K notes of data. To our knowledge, this is the largest cross-modal music dataset with note-level annotations to date. To demonstrate the usage of the MOSA dataset, we present several innovative cross-modal music information retrieval (MIR) and musical content generation tasks, including the detection of beats, downbeats, phrase, and expressive contents from audio, video and motion data, and the generation of musicians' body motion from given music audio. The dataset and codes are available alongside this publication (https: github.com yufenhuang MOSA-Music-mOtion-and-Semantic-Annotation-dataset).

【5】 mHuBERT-147: A Compact Multilingual HuBERT Model
标题: mHuBERT-147:紧凑的多语言HuBERT模型
作者:Marcely Zanon Boito,Vivek Iyer,Nikolaos Lagos,Laurent Besacier,Ioan Calapodescu
备注:Extended version of the Interspeech 2024 paper of same name
链接:点击下载PDF文件
摘要:mHuBERT-147是第一个通用的大规模多语言HuBERT语音表示模型,基于9万小时的干净,开放许可证的数据进行训练。为了扩展多迭代HuBERT方法,我们使用基于faiss的聚类,实现了比原始方法快5.2倍的标签分配。我们还应用了一种新的多语言上采样策略,利用语言和数据集的多样性。经过3次训练迭代,并且只有95 M个参数,mHuBERT-147的性能优于在更多数据上训练的大型模型。我们在ML-SUPERB 10分钟 1小时排行榜上分别排名第二和第一,所有LID任务的SOTA得分。在ASR LID任务中,我们的模型始终超过XLS-R(300 M参数; 436 K小时),并与更大的MMS(1B参数; 491 K小时)相比表现出强大的竞争力。我们的研究结果表明,mHuBERT-147是一个有前途的多语言语音处理任务的模型,提供了一个前所未有的高性能和参数效率之间的平衡。摘要:We present mHuBERT-147, the first general-purpose massively multilingual HuBERT speech representation model trained on 90K hours of clean, open-license data. To scale up the multi-iteration HuBERT approach, we use faiss-based clustering, achieving 5.2x faster label assignment over the original method. We also apply a new multilingual batching up-sampling strategy, leveraging both language and dataset diversity. After 3 training iterations and with only 95M parameters, mHuBERT-147 outperforms larger models trained on substantially more data. We rank second and first on the ML-SUPERB 10min 1h leaderboards respectively, with SOTA scores for all LID tasks. Across ASR LID tasks, our model consistently surpasses XLS-R (300M params; 436K hours) and demonstrates strong competitiveness against the much larger MMS (1B params; 491K hours). Our findings suggest that mHuBERT-147 is a promising model for multilingual speech processing tasks, offering an unprecedented balance between high performance and parameter efficiency.

【6】 Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
标题: 使用数据驱动和基于知识的特征从语音预测心脏活动
作者:Gasser Elbanna,Zohreh Mostaani,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:准确预测心脏活动和其他生物信号对于诊断和监测至关重要。鉴于语音是多个生理系统的产物,大量工作研究了心脏活动的声学相关性。最近,与传统的声学方法相比,自监督模型在语音相关的任务中表现出色。然而,数据驱动表示在预测心脏活动方面的鲁棒性仍然未被探索。在这项研究中,我们证明了自我监督的语音模型在预测心脏活动参数方面优于声学特征。我们还强调了个体变异性对模型泛化能力的影响。这些发现强调了数据驱动表示在此类任务中的价值,以及需要更多基于语音的生理数据来减轻与说话者相关的挑战。摘要:Accurately predicting heart activity and other biological signals is crucial for diagnosis and monitoring. Given that speech is an outcome of multiple physiological systems, a significant body of work studied the acoustic correlates of heart activity. Recently, self-supervised models have excelled in speech-related tasks compared to traditional acoustic methods. However, the robustness of data-driven representations in predicting heart activity remained unexplored. In this study, we demonstrate that self-supervised speech models outperform acoustic features in predicting heart activity parameters. We also emphasize the impact of individual variability on model generalizability. These findings underscore the value of data-driven representations in such tasks and the need for more speech-based physiological data to mitigate speaker-related challenges.

【7】 Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
标题: 基于音频的运行步数估计--窗口和神经网络基线
作者:Philipp Wagner,Andreas Triantafyllopoulos,Alexander Gebhard,Björn Schuller
备注:Accepted at EUSIPCO 2024
链接:点击下载PDF文件
摘要:近几十年来,跑步已成为越来越受欢迎的消遣活动,因为它的可访问性,易于实践和预期的健康益处。然而,对于不同经验水平的跑步者来说,跑步相关伤害的风险是很大的。几种常见的受伤形式是由于过度使用--超过推荐的跑步时间和强度。最近,基于音频的跟踪已经成为监测跑步行为和表现的另一种方式,以前的研究主要集中在预测跑步者疲劳。在这项工作中,我们研究了基于音频的步数估计在户外运行,实现了平均绝对误差为1.098,基于窗口的步数差异和皮尔逊相关系数为0.479时,预测的步骤数在5秒的音频窗口。因此,我们的工作展示了基于音频的监测估计重要的生理变量的可行性,并为进一步利用音频传感器更彻底地表征跑步者行为奠定了基础。摘要:In recent decades, running has become an increasingly popular pastime activity due to its accessibility, ease of practice, and anticipated health benefits. However, the risk of running-related injuries is substantial for runners of different experience levels. Several common forms of injuries result from overuse -- extending beyond the recommended running time and intensity. Recently, audio-based tracking has emerged as yet another modality for monitoring running behaviour and performance, with previous studies largely concentrating on predicting runner fatigue. In this work, we investigate audio-based step count estimation during outdoor running, achieving a mean absolute error of 1.098 in window-based step-count differences and a Pearson correlation coefficient of 0.479 when predicting the number of steps in a 5-second window of audio. Our work thus showcases the feasibility of audio-based monitoring for estimating important physiological variables and lays the foundations for further utilising audio sensors for a more thorough characterisation of runner behaviour.

【8】 An automatic analysis of ultrasound vocalisations for the prediction of interaction context in captive Egyptian fruit bats
标题: 超声发声的自动分析用于预测圈养埃及果蝠的互动环境
作者:Andreas Triantafyllopoulos,Alexander Gebhard,Manuel Milling,Simon Rampp,Björn Schuller
备注:Accepted at EUSIPCO 2024
链接:点击下载PDF文件
摘要:计算生物声学的先前工作主要集中在特定栖息地中动物存在的检测上。然而,动物的声音包含的信息比仅仅存在的信息丰富得多;除此之外,它们还包含了这些动物与其物种其他成员的互动。在自然主义环境中研究这些相互作用几乎是不可能的,因为通常缺乏基本事实。相反,圈养动物的使用提供了一种可行的替代途径。然而,大多数以前的作品遵循传统的,基于地理学的方法来分析相互作用。在目前的工作中,我们超越了这个标准框架,试图使用深度神经网络来预测捕获的埃及伊蚊之间相互作用的潜在背景。我们达到了超过30%的未加权平均召回率-超过机会水平的三倍-并显示出与我们的统计分析不同的错误模式。因此,这项工作代表了从声音中自动分析动物状态的重要一步。摘要:Prior work in computational bioacoustics has mostly focused on the detection of animal presence in a particular habitat. However, animal sounds contain much richer information than mere presence; among others, they encapsulate the interactions of those animals with other members of their species. Studying these interactions is almost impossible in a naturalistic setting, as the ground truth is often lacking. The use of animals in captivity instead offers a viable alternative pathway. However, most prior works follow a traditional, statistics-based approach to analysing interactions. In the present work, we go beyond this standard framework by attempting to predict the underlying context in interactions between captive emph{Rousettus Aegyptiacus} using deep neural networks. We reach an unweighted average recall of over 30 % -- more than thrice the chance level -- and show error patterns that differ from our statistical analysis. This work thus represents an important step towards the automatic analysis of states in animals from sound.

【9】 Unsupervised Improved MVDR Beamforming for Sound Enhancement
标题: 用于声音增强的无监督改进的MVDR束形成
作者:Jacob Kealey,John Hershey,François Grondin
链接:点击下载PDF文件
摘要:神经网络最近已经成为声音分离的主要方法。它们的良好性能依赖于孤立记录的大型数据集。对于语音和音乐,隔离的单通道数据是容易获得的;然而,在多通道情况下,以及大多数其他声音类别中,情况并非如此。多通道方法有可能优于单通道方法,因为它们可以利用空间和光谱特征,但缺乏训练数据仍然是一个挑战。我们提出了无监督改进的最小变差无失真响应(UIMVDR),它使多通道分离能够通过无监督训练和波束成形来利用野外单通道数据。结果表明,UIMVDR的推广以及监督模型相比,提高分离性能,特别是在有限的监督数据的情况下。通过使用在线数据,它还减少了为多渠道方法收集数据所需的工作。摘要:Neural networks have recently become the dominant approach to sound separation. Their good performance relies on large datasets of isolated recordings. For speech and music, isolated single channel data are readily available; however the same does not hold in the multi-channel case, and with most other sound classes. Multi-channel methods have the potential to outperform single channel approaches as they can exploit both spatial and spectral features, but the lack of training data remains a challenge. We propose unsupervised improved minimum variation distortionless response (UIMVDR), which enables multi-channel separation to leverage in-the-wild single-channel data through unsupervised training and beamforming. Results show that UIMVDR generalizes well and improves separation performance compared to supervised models, particularly in cases with limited supervised data. By using data available online, it also reduces the effort required to gather data for multi-channel approaches.

【10】 Zero-Shot Audio Captioning Using Soft and Hard Prompts
标题: 使用软和硬字幕的Zero-Shot音频字幕
作者:Yiming Zhang,Xuenan Xu,Ruoyi Du,Haohe Liu,Yuan Dong,Zheng-Hua Tan,Wenwu Wang,Zhanyu Ma
备注:Submitted to IEEEACM Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
摘要:在传统的音频字幕方法中,模型通常使用包含音频文本对的人工注释数据集以完全监督的方式进行训练,然后在来自相同数据集的测试集上进行评估。这种方法有两个局限性。首先,这些方法通常需要大量数据,并且需要耗时且昂贵的人工注释来获得音频文本对。第二,这些模型在跨域场景中经常遭受性能降级,即,当输入音频来自与训练集不同的域时,然而,这很少受到关注。本文提出了一种基于对比语言-音频预训练(CLAP)模型的音频字幕生成方法。我们提出的方法只需要文本数据进行训练,使模型能够从跨模态语义空间中的文本特征生成文本。在推理阶段,模型通过利用CLAP的音频-文本对齐,从音频特征生成给定音频的描述性文本。我们设计了两种策略来减轻文本和音频嵌入之间的差异:基于混合增强的软提示和基于检索的声学感知硬提示。这些方法的目的是提高我们提出的模型的泛化性能,促进模型生成字幕更强大,更准确。在AudioCaps和Clotho基准测试上的大量实验表明了该方法的有效性,在域内场景中的性能优于其他zero-shot音频字幕方法,在跨域场景中的性能优于其他方法,从而突出了该方法的泛化能力。摘要:In traditional audio captioning methods, a model is usually trained in a fully supervised manner using a human-annotated dataset containing audio-text pairs and then evaluated on the test sets from the same dataset. Such methods have two limitations. First, these methods are often data-hungry and require time-consuming and expensive human annotations to obtain audio-text pairs. Second, these models often suffer from performance degradation in cross-domain scenarios, i.e., when the input audio comes from a different domain than the training set, which, however, has received little attention. We propose an effective audio captioning method based on the contrastive language-audio pre-training (CLAP) model to address these issues. Our proposed method requires only textual data for training, enabling the model to generate text from the textual feature in the cross-modal semantic space.In the inference stage, the model generates the descriptive text for the given audio from the audio feature by leveraging the audio-text alignment from CLAP.We devise two strategies to mitigate the discrepancy between text and audio embeddings: a mixed-augmentation-based soft prompt and a retrieval-based acoustic-aware hard prompt. These approaches are designed to enhance the generalization performance of our proposed model, facilitating the model to generate captions more robustly and accurately. Extensive experiments on AudioCaps and Clotho benchmarks show the effectiveness of our proposed method, which outperforms other zero-shot audio captioning approaches for in-domain scenarios and outperforms the compared methods for cross-domain scenarios, underscoring the generalization ability of our method.

【11】 Quantifying the effect of speech pathology on automatic and human speaker verification
标题: 量化言语病理学对自动和人类说话人验证的影响
作者:Bence Mark Halpern,Thomas Tienkamp,Wen-Chin Huang,Lester Phillip Violeta,Teja Rebernik,Sebastiaan de Visscher,Max Witjes,Martijn Wieling,Defne Abur,Tomoki Toda
备注:5 pages, 2 figures, 2 tables. Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本研究探讨了如何手术干预语音病理(特别是口腔癌手术的结果)影响的自动说话人确认(ASV)系统的性能。使用两个最近收集的荷兰数据集与并行术前和术后的音频从同一扬声器,NKI-OC-VC和SPOKE,我们评估在何种程度上语音病理影响ASV的性能,以及是否客观 主观措施的语音严重程度与性能。最后,我们进行了一个感性的研究,比较ASV和人类听众的判断。我们的研究结果表明,病理性言语会对ASV的表现产生负面影响,并且言语的严重程度与ASV的表现呈负相关。说话人相似性和严重性的感知和客观得分有适度的一致性,但是,我们不能清楚地建立在感知研究中,是否同样的现象也存在于人类的感知。摘要:This study investigates how surgical intervention for speech pathology (specifically, as a result of oral cancer surgery) impacts the performance of an automatic speaker verification (ASV) system. Using two recently collected Dutch datasets with parallel pre and post-surgery audio from the same speaker, NKI-OC-VC and SPOKE, we assess the extent to which speech pathology influences ASV performance, and whether objective subjective measures of speech severity are correlated with the performance. Finally, we carry out a perceptual study to compare judgements of ASV and human listeners. Our findings reveal that pathological speech negatively affects ASV performance, and the severity of the speech is negatively correlated with the performance. There is a moderate agreement in perceptual and objective scores of speaker similarity and severity, however, we could not clearly establish in the perceptual study, whether the same phenomenon also exists in human perception.

【12】 Thunder : Unified Regression-Diffusion Speech Enhancement with a Single Reverse Step using Brownian Bridge
标题: Thunder:使用布朗桥进行单反向步骤的统一回归扩散语音增强
作者:Thanapat Trachu,Chawan Piansaddhayanon,Ekapol Chuangsuwanich
备注:5 pages, 3 figures, 4 tables, This paper will be submitted in the interspeech conference
链接:点击下载PDF文件
摘要:基于扩散的语音增强已经显示出有希望的结果,但可能遭受较慢的推理时间。利用由基于回归的模型生成的增强音频来初始化扩散过程可以用于减少所需的计算步骤。然而,这些方法通常需要回归模型,进一步增加了系统的复杂性。我们提出了雷霆,一个统一的回归扩散模型,利用布朗桥过程,可以让该模型在这两种模式下的行为。通过将扩散时间步长设置为接近1,可以进入回归模式。然而,由于梯度不稳定性,标准的基于分数的扩散建模在此设置中表现不佳。为了缓解这个问题,我们修改了扩散模型来预测干净的语音,而不是分数函数,以更紧凑的模型大小和更少的反向步骤实现有竞争力的性能。摘要:Diffusion-based speech enhancement has shown promising results, but can suffer from a slower inference time. Initializing the diffusion process with the enhanced audio generated by a regression-based model can be used to reduce the computational steps required. However, these approaches often necessitate a regression model, further increasing the system's complexity. We propose Thunder, a unified regression-diffusion model that utilizes the Brownian bridge process which can allow the model to act in both modes. The regression mode can be accessed by setting the diffusion time step closed to 1. However, the standard score-based diffusion modeling does not perform well in this setup due to gradient instability. To mitigate this problem, we modify the diffusion model to predict the clean speech instead of the score function, achieving competitive performance with a more compact model size and fewer reverse steps.

【13】 StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
标题: StreamAtt:通过基于注意力的音频历史选择直接流媒体语音到文本翻译
作者:Sara Papi,Marco Gaido,Matteo Negri,Luisa Bentivogli
备注:Accepted at ACL 2024 main conference
链接:点击下载PDF文件
摘要:流式语音到文本翻译(StreamST)是在增量接收音频流的同时自动翻译语音的任务。与处理预分段语音的同步ST(SimulST)不同,StreamST面临着处理连续和无限音频流的挑战。这需要关于保留先前历史的什么的附加决定,由于延迟和计算约束,完全保留先前历史是不切实际的。尽管现实世界对实时ST的需求,但对流式翻译的研究仍然有限,现有的作品仅关注SimulST。为了填补这一空白,我们引入StreamAtt,第一个StreamST政策,并提出StreamLAAL,第一个StreamST延迟度量,旨在与现有的指标SimulST相媲美。在MuST-C v1.0的所有8种语言中进行的广泛实验显示了StreamAtt与原始流基线和相关的最先进的SimulST策略相比的有效性,为StreamST研究提供了第一步。摘要:Streaming speech-to-text translation (StreamST) is the task of automatically translating speech while incrementally receiving an audio stream. Unlike simultaneous ST (SimulST), which deals with pre-segmented speech, StreamST faces the challenges of handling continuous and unbounded audio streams. This requires additional decisions about what to retain of the previous history, which is impractical to keep entirely due to latency and computational constraints. Despite the real-world demand for real-time ST, research on streaming translation remains limited, with existing works solely focusing on SimulST. To fill this gap, we introduce StreamAtt, the first StreamST policy, and propose StreamLAAL, the first StreamST latency metric designed to be comparable with existing metrics for SimulST. Extensive experiments across all 8 languages of MuST-C v1.0 show the effectiveness of StreamAtt compared to a naive streaming baseline and the related state-of-the-art SimulST policy, providing a first step in StreamST research.

【14】 RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
标题: RawBMamba:用于音频深度伪造检测的端到端双向状态空间模型
作者:Yujie Chen,Jiangyan Yi,Jun Xue,Chenglong Wang,Xiaohui Zhang,Shunbo Dong,Siding Zeng,Jianhua Tao,Lv Zhao,Cunhang Fan
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:用于区分真实音频和假音频的假伪像可以存在于短距离段和长距离段两者中。因此,结合局部和全局特征信息可以有效地区分真假音频。本文提出了一种名为RawBMamba的端到端双向状态空间模型,用于捕获用于音频深度伪造检测的短距离和长距离区分信息。具体来说,我们使用sinc Layer和多个卷积层来捕获短程特征,然后设计一个双向Mamba来解决Mamba的单向建模问题,并进一步捕获长程特征信息。此外,我们开发了一个双向融合模块来整合嵌入,增强音频上下文表示,并结合短距离和长距离的信息。实验结果表明,在ASVspoof 2021 LA数据集上,RawBMamba算法比Rawformer算法性能提高了34.1%,在其他数据集上也表现出了较好的性能.摘要:Fake artefacts for discriminating between bonafide and fake audio can exist in both short- and long-range segments. Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio. This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short- and long-range discriminative information for audio deepfake detection. Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information. Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining short- and long-range information. The results show that our proposed RawBMamba achieves a 34.1 % improvement over Rawformer on ASVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.

【15】 Contrastive Learning from Synthetic Audio Doppelgangers
标题: 合成音频分身的对比学习
作者:Manuel Cherep,Nikhil Singh
备注:17 pages, 6 figures
链接:点击下载PDF文件
摘要:学习鲁棒的音频表示目前需要大量的真实世界录音数据集。通过对这些记录进行人工转换,模型可以学习识别相似性,尽管通过对比学习等技术存在细微变化。然而,这些转换只是真实世界声音中真实多样性的近似值,这些声音是由复杂的物理过程相互作用产生的,从声带振动到乐器的共振。我们提出了一个解决方案,利用合成音频的数据规模和转换的限制。通过随机扰动声音合成器的参数,我们生成音频doppelg “angers合成的积极对与因果操纵的音色,音高和时间包络的变化。这些变化,难以实现通过现有的音频转换,提供了丰富的对比信息的来源。尽管转移到随机生成的合成数据,我们的方法产生强有力的表示,与标准音频分类基准的真实数据竞争。值得注意的是,我们的方法是轻量级的,不需要数据存储,只有一个超参数,我们广泛分析。我们提供这种方法作为现有的音频对比学习策略的补充,使用合成的声音来减少从业者的数据负担。摘要:Learning robust audio representations currently demands extensive datasets of real-world sound recordings. By applying artificial transformations to these recordings, models can learn to recognize similarities despite subtle variations through techniques like contrastive learning. However, these transformations are only approximations of the true diversity found in real-world sounds, which are generated by complex interactions of physical processes, from vocal cord vibrations to the resonance of musical instruments. We propose a solution to both the data scale and transformation limitations, leveraging synthetic audio. By randomly perturbing the parameters of a sound synthesizer, we generate audio doppelg "angers-synthetic positive pairs with causally manipulated variations in timbre, pitch, and temporal envelopes. These variations, difficult to achieve through transformations of existing audio, provide a rich source of contrastive information. Despite the shift to randomly generated synthetic data, our method produces strong representations, competitive with real data on standard audio classification benchmarks. Notably, our approach is lightweight, requires no data storage, and has only a single hyperparameter, which we extensively analyze. We offer this method as a complement to existing strategies for contrastive learning in audio, using synthesized sounds to reduce the data burden on practitioners.

【16】 Zero-Shot End-To-End Spoken Question Answering In Medical Domain
标题: 医疗领域的Zero-Shot端到端口语问答
作者:Yanis Labrak,Adel Moumen,Richard Dufour,Mickael Rouvier
Journal-ref:InterSpeech 2024
链接:点击下载PDF文件
摘要:在快速发展的口语问答(SQA)领域,大型语言模型(LLM)的集成已经成为一种变革性的发展。传统的方法通常需要使用单独的模型来进行问题音频转录和答案选择,从而导致显著的资源利用和错误累积。为了应对这些挑战,我们探索了医疗领域SQA的端到端(E2E)方法的有效性。我们的研究介绍了一种新的zero-shot SQA方法,相比传统的级联系统。通过对8个医疗任务和48小时合成音频的新开放基准进行的全面评估,我们证明了我们的方法所需的资源比1.3B参数LLM与1.55B参数ASR模型的组合少14.7倍,同时平均精度提高了0.5%。这些发现强调了在资源受限的情况下,E2E方法用于SQA的潜力。摘要:In the rapidly evolving landscape of spoken question-answering (SQA), the integration of large language models (LLMs) has emerged as a transformative development. Conventional approaches often entail the use of separate models for question audio transcription and answer selection, resulting in significant resource utilization and error accumulation. To tackle these challenges, we explore the effectiveness of end-to-end (E2E) methodologies for SQA in the medical domain. Our study introduces a novel zero-shot SQA approach, compared to traditional cascade systems. Through a comprehensive evaluation conducted on a new open benchmark of 8 medical tasks and 48 hours of synthetic audio, we demonstrate that our approach requires up to 14.7 times fewer resources than a combined 1.3B parameters LLM with a 1.55B parameters ASR model while improving average accuracy by 0.5 %. These findings underscore the potential of E2E methodologies for SQA in resource-constrained contexts.

【17】 Source -Free Domain Adaptation for Speaker Verification in Data-Scarce Languages and Noisy Channels
标题: 源-免费域自适应,用于数据稀缺语言和有噪通道中的说话人验证
作者:Shlomo Salo Elia,Aviad Malachi,Vered Aharonson,Gadi Pinkas
链接:点击下载PDF文件
摘要:领域自适应常常受到极小的目标数据集和不可访问的源数据的阻碍。这些情况在语音验证中普遍存在,其中隐私政策和 或具有稀缺语音资源的语言限制了足够数据的可用性。本文探讨了数据稀缺语言中说话人确认的无源域自适应技术。研究了源语和目标语之间的语言和通道不匹配。在不同大小的标记目标数据中评估和比较了微调方法。针对未标记目标数据集,研究了一种新的迭代聚类学习算法。摘要:Domain adaptation is often hampered by exceedingly small target datasets and inaccessible source data. These conditions are prevalent in speech verification, where privacy policies and or languages with scarce speech resources limit the availability of sufficient data. This paper explored techniques of sourcefree domain adaptation unto a limited target speech dataset for speaker verificationin data-scarce languages. Both language and channel mis-match between source and target were investigated. Fine-tuning methods were evaluated and compared across different sizes of labeled target data. A novel iterative cluster-learn algorithm was studied for unlabeled target datasets.

【18】 Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
标题: 预算真的会提示吗?探索Whisper的快速理解能力
作者:Chih-Kai Yang,Kuan-Po Huang,Hung-yi Lee
备注:In progress
链接:点击下载PDF文件
摘要:本研究探讨Whisper,一个高性能的语音识别模型,和提示信息之间的相互作用。我们的研究结果出乎意料地表明,耳语可能无法完全掌握预期的文本提示。此外,我们发现,即使在文本提示中更严格地遵守主题信息,也不能保证性能提高。值得注意的是,英语提示在两种语言的数据集上通常优于普通话提示,这可能是由于这些语言的训练数据分布的差异。相反,我们发现,耳语表现出意识的误导性信息的语言标记,有效地忽略不正确的语言标记,并专注于正确的。总之,这项工作提出了关于耳语的快速理解能力的问题,并鼓励进一步的研究。摘要:This research explores the interaction between Whisper, a high-performing speech recognition model, and information in prompts. Our results unexpectedly show that Whisper may not fully grasp textual prompts as anticipated. Additionally, we find that performance improvement is not guaranteed even with stronger adherence to the topic information in textual prompts. It is also noted that English prompts generally outperform Mandarin ones on datasets of both languages, likely due to differences in training data distributions for these languages. Conversely, we discover that Whisper exhibits awareness of misleading information in language tokens by effectively ignoring incorrect language tokens and focusing on the correct ones. In summary, this work raises questions about Whisper's prompt understanding capability and encourages further studies.

【19】 Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
标题: 优化多口吃语音分类:利用Whisper的编码器在自动评估中有效简化参数
作者:Huma Ameer,Seemab Latif,Rabia Latif
链接:点击下载PDF文件
摘要:口吃语音的自动分类对及时评估语音语言病理学家提供帮助具有重要意义。尽管在该领域取得了显着的进步,但在言语中发生多个不流利的情况需要注意。我们已经采取了一种渐进的方法来填补这一空白,更有效地分类多口吃的语音。这个问题已经通过首先从SEP-28 k音频片段中策划多口吃不流利的数据集来解决。其次,采用Whisper,一个国家的最先进的语音识别模型已利用其编码器和多标签分类的问题。第三,使用6个编码器层Whisper并试验各种层冻结策略,确定了模型的计算效率配置。所提出的配置在外部测试数据集(即Fluency-Bank)上相应地实现了0.88、0.85和0.87的微观、宏观和加权F1分数。此外,通过层冻结策略,我们能够通过微调单个编码器层来实现上述结果,从而将模型的可训练参数从2027万减少到329万。这项研究揭示了最后一个编码器层在识别口吃语音中的不流利性方面的贡献。因此,它导致了一种计算效率高的方法,使模型更适合各种方言和语言。摘要:The automated classification of stuttered speech has significant implications for timely assessments providing assistance to speech language pathologists. Despite notable advancements in the field, the cases in which multiple disfluencies occur in speech require attention. We have taken a progressive approach to fill this gap by classifying multi-stuttered speech more efficiently. The problem has been addressed by firstly curating a dataset of multi-stuttered disfluencies from SEP-28k audio clips. Secondly, employing Whisper, a state-of-the-art speech recognition model has been leveraged by using its encoder and taking the problem as multi-label classification. Thirdly, using a 6 encoder layer Whisper and experimenting with various layer freezing strategies, a computationally efficient configuration of the model was identified. The proposed configuration achieved micro, macro, and weighted F1- scores of 0.88, 0.85, and 0.87, correspondingly on an external test dataset i.e. Fluency-Bank. In addition, through layer freezing strategies, we were able to achieve the aforementioned results by fine-tuning a single encoder layer, consequently, reducing the model's trainable parameters from 20.27 million to 3.29 million. This research study unveils the contribution of the last encoder layer in the identification of disfluencies in stuttered speech. Consequently, it has led to a computationally efficient approach which makes the model more adaptable for various dialects and languages.

【20】 SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
标题: SPA-SRC:用于歌唱声音转换的自我监督音调增强
作者:Bingsong Bai,Fengping Wang,Yingming Gao,Ya Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:与传统方法相比,基于扩散的歌唱声转换(SVC)模型显示出更好的合成质量。然而,在跨域SVC场景中,源语音域和目标语音域之间的音高存在显著差异,这些模型往往会生成声音嘶哑的音频,这对实现高质量的语音输出提出了挑战。因此,在本文中,我们提出了一种自监督音高增强方法歌唱语音转换(SPA-SVC),它可以提高语音质量在SVC任务,而不需要额外的数据或增加模型参数。我们创新性地引入了一个周期性的音高变换训练策略和结构相似性指数(SSIM)损失到我们的SVC模型,有效地提高了它的性能。在公共歌唱数据集M4 Singer上的实验结果表明,我们提出的方法显着提高了模型的性能,在一般的SVC场景,特别是在跨域SVC场景。摘要:Diffusion-based singing voice conversion (SVC) models have shown better synthesis quality compared to traditional methods. However, in cross-domain SVC scenarios, where there is a significant disparity in pitch between the source and target voice domains, the models tend to generate audios with hoarseness, posing challenges in achieving high-quality vocal outputs. Therefore, in this paper, we propose a Self-supervised Pitch Augmentation method for Singing Voice Conversion (SPA-SVC), which can enhance the voice quality in SVC tasks without requiring additional data or increasing model parameters. We innovatively introduce a cycle pitch shifting training strategy and Structural Similarity Index (SSIM) loss into our SVC model, effectively enhancing its performance. Experimental results on the public singing datasets M4Singer indicate that our proposed method significantly improves model performance in both general SVC scenarios and particularly in cross-domain SVC scenarios.

【21】 Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
标题: 利用分层韵律建模实现表达性Zero-Shot语音合成
作者:Yuepeng Jiang,Tao Li,Fengyu Yang,Lei Xie,Meng Meng,Yujun Wang
备注:5 pages, 2 figures, accepted by Interspeech2024
链接:点击下载PDF文件
摘要:近年来zero-shot语音合成的研究在说话人相似度方面取得了很大的进展。然而,目前的努力集中在音色的泛化,而不是韵律建模,这导致有限的自然性和表现力。为了解决这个问题,我们引入了一种新的语音合成模型,在大规模数据集上训练,包括音色和层次韵律建模。由于音色是一个与表现力密切相关的全局属性,我们采用一个全局向量对说话人音色进行建模,同时指导韵律建模。此外,由于韵律包含全局一致性和局部变化,我们引入了扩散模型作为音高预测器,并采用韵律适配器对韵律进行分层建模,进一步提高了合成语音的韵律质量。实验结果表明,我们的模型不仅保持可比的音色质量的基线,但也表现出更好的自然度和表现力。摘要:Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody modeling, which results in limited naturalness and expressiveness. To address this, we introduce a novel speech synthesis model trained on large-scale datasets, including both timbre and hierarchical prosody modeling. As timbre is a global attribute closely linked to expressiveness, we adopt a global vector to model speaker timbre while guiding prosody modeling. Besides, given that prosody contains both global consistency and local variations, we introduce a diffusion model as the pitch predictor and employ a prosody adaptor to model prosody hierarchically, further enhancing the prosody quality of the synthesized speech. Experimental results show that our model not only maintains comparable timbre quality to the baseline but also exhibits better naturalness and expressiveness.

【22】 Heart Sound Segmentation Using Deep Learning Techniques
标题: 使用深度学习技术的人声分割
作者:Manas Madine
链接:点击下载PDF文件
摘要:心脏病仍然是全球死亡的主要原因。听诊,即倾听心音的过程,可以通过使用心音图(PCG)信号的计算机辅助分析来增强。本文提出了一种新的方法,心音分割和分类为S1(LUB)和S2(DUB)的声音。我们采用基于FFT的过滤,动态规划的事件检测,和一个暹罗网络的强大的分类。与现有方法相比,我们的方法在PASCAL心音数据集上表现出优越的性能。摘要:Heart disease remains a leading cause of mortality worldwide. Auscultation, the process of listening to heart sounds, can be enhanced through computer-aided analysis using Phonocardiogram (PCG) signals. This paper presents a novel approach for heart sound segmentation and classification into S1 (LUB) and S2 (DUB) sounds. We employ FFT-based filtering, dynamic programming for event detection, and a Siamese network for robust classification. Our method demonstrates superior performance on the PASCAL heart sound dataset compared to existing approaches.

【23】 Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
作者:Mark Hamilton,Andrew Zisserman,John R. Hershey,William T. Freeman
备注:Computer Vision and Pattern Recognition 2024
链接:点击下载PDF文件
摘要:我们提出了DenseAV,一种新颖的双编码器接地架构,仅通过观看视频来学习高分辨率,语义有意义和视听对齐的功能。我们表明,DenseAV可以发现单词的“意义”和声音的“位置”,而无需明确的本地化监督。此外,它可以自动发现和区分这两种类型的关联,而无需监督。我们表明,DenseAV的定位能力来自一个新的多头特征聚合算子,该算子直接比较密集的图像和音频表示进行对比学习。相比之下,许多其他学习“全球”音频和视频表示的系统不能本地化单词和声音。最后,我们贡献了两个新的数据集,通过语音和声音提示的语义分割来提高AV表示的评估。在这些和其他数据集上,我们显示DenseAV在语音和声音提示的语义分割方面显着优于现有技术。DenseAV在使用不到一半的参数进行跨模态检索方面优于以前的最先进的ImageBind。项目页面: href{https: aka.ms densav}{https: aka.ms densav}摘要:We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching videos. We show that DenseAV can discover the meaning'' of words and the location'' of sounds without explicit localization supervision. Furthermore, it automatically discovers and distinguishes between these two types of associations without supervision. We show that DenseAV's localization abilities arise from a new multi-head feature aggregation operator that directly compares dense image and audio representations for contrastive learning. In contrast, many other systems that learn global'' audio and video representations cannot localize words and sound. Finally, we contribute two new datasets to improve the evaluation of AV representations through speech and sound prompted semantic segmentation. On these and other datasets we show DenseAV dramatically outperforms the prior art on speech and sound prompted semantic segmentation. DenseAV outperforms the previous state-of-the-art, ImageBind, on cross-modal retrieval using fewer than half of the parameters. Project Page: href{https: aka.ms denseav}{https: aka.ms denseav}

【24】 Exploring the Benefits of Tokenization of Discrete Acoustic Units
标题: 探索离散声学单元代币化的好处
作者:Avihu Dekel,Raul Fernandez
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:将基本词汇表的单元合并为更大的可变速率单元的标记化算法已经成为自然语言处理任务中的标准。然而,当词汇表由音素或离散声学单位(DAU)组成时,这个想法大多被忽视了,由于离散语言建模技术的成功,基于音频的表示正在发挥越来越重要的作用。在本文中,我们展示了语音单元和DAU的标记化在三个预测任务上的优势:字素到音素,字素到DAU,以及使用DAU语言建模的无监督语音生成。我们证明了令牌化在所有三个任务中的性能以及训练和推理速度方面都有显着的改进。我们还提供理论见解,为观察到的卓越性能提供一些解释。摘要:Tokenization algorithms that merge the units of a base vocabulary into larger, variable-rate units have become standard in natural language processing tasks. This idea, however, has been mostly overlooked when the vocabulary consists of phonemes or Discrete Acoustic Units (DAUs), an audio-based representation that is playing an increasingly important role due to the success of discrete language-modeling techniques. In this paper, we showcase the advantages of tokenization of phonetic units and of DAUs on three prediction tasks: grapheme-to-phoneme, grapheme-to-DAUs, and unsupervised speech generation using DAU language modeling. We demonstrate that tokenization yields significant improvements in terms of performance, as well as training and inference speed, across all three tasks. We also offer theoretical insights to provide some explanation for the superior performance observed.

【25】 Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation
标题: 嗯你说什么?使用心理物理反向相关揭示第一语言和第二语言单词感知中的远端和近端上下文效应
作者:Paige Tuttösí,H. Henny Yeung,Yue Wang,Fenqi Wang,Guillaume Denis,Jean-Julien Aucouturier,Angelica Lim
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:声学语境效应,即周围音高、速率或音色的变化影响声音的感知,在言语感知中有很好的记录,但它们如何与语言背景相互作用仍不清楚。使用反向相关的方法,我们系统地改变了音高和语速在短语周围的不同对元音的第二语言(L2)的英语( i - I )和法语( u - y )的扬声器,从而重建,在数据驱动的方式,韵律配置文件,偏见他们的看法。测试英语和法语的发言者(n=25),我们发现,元音感知实际上是由周围的音高和语音速率的冲突影响:一致的近端效应0.2秒前的目标和远端对比效应高达1秒前,并发现,L1和L2扬声器表现出惊人的相似的韵律配置文件的感知。我们提供了一种新的方法来调查声环境的影响,跨刺激,时间尺度,和声学域。摘要:Acoustic context effects, where surrounding changes in pitch, rate or timbre influence the perception of a sound, are well documented in speech perception, but how they interact with language background remains unclear. Using a reverse-correlation approach, we systematically varied the pitch and speech rate in phrases around different pairs of vowels for second language (L2) speakers of English ( i - I ) and French ( u - y ), thus reconstructing, in a data-driven manner, the prosodic profiles that bias their perception. Testing English and French speakers (n=25), we showed that vowel perception is in fact influenced by conflicting effects from the surrounding pitch and speech rate: a congruent proximal effect 0.2s pre-target and a distal contrastive effect up to 1s before; and found that L1 and L2 speakers exhibited strikingly similar prosodic profiles in perception. We provide a novel method to investigate acoustic context effects across stimuli, timescales, and acoustic domain.

【26】 DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
标题: DAISY:语音表示模型的数据自适应自我监督早期退出
作者:Tzu-Quan Lin,Hung-yi Lee,Hao Tang
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自监督语音模型已被证明对各种任务都很有用,但它们的大尺寸限制了在计算能力和内存较低的设备中的使用。在这项工作中,我们探索早期退出,通过提前退出网络的转发过程来减少延迟的方法。大多数提前退出方法需要为每个任务提供单独的提前退出模型,有些甚至需要对整个预训练模型进行微调。我们引入了数据自适应自监督早期退出(DAISY),这是一种基于自监督损失决定何时退出的方法,无需多轮训练和微调。DAISY在MiniSUPERB基准测试中与HuBERT的性能相当,但推理时间要快得多。我们对DAISY自适应性的分析表明,该模型在干净数据上提前退出(使用较少的层),而在有噪声的数据上延迟退出(使用更多的层),根据每个样本的噪声水平动态调整推理的计算成本。摘要:Self-supervised speech models have shown to be useful for various tasks, but their large size limits the use in devices with low computing power and memory. In this work, we explore early exit, an approach for reducing latency by exiting the forward process of a network early. Most approaches of early exit need a separate early exit model for each task, with some even requiring fine-tuning of the entire pretrained model. We introduce Data Adaptive Self-Supervised Early Exit (DAISY), an approach that decides when to exit based on the self-supervised loss, eliminating the need for multiple round of training and fine-tuning. DAISY matches the performance of HuBERT on the MiniSUPERB benchmark, but with much faster inference times. Our analysis on the adaptivity of DAISY shows that the model exits early (using fewer layers) on clean data while exits late (using more layers) on noisy data, dynamically adjusting the computational cost of inference based on the noise level of each sample.

【27】 VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
标题: WAL-E 2:神经编解码器语言模型是人类对等Zero-Shot文本到语音合成器
作者:Sanyuan Chen,Shujie Liu,Long Zhou,Yanqing Liu,Xu Tan,Jinyu Li,Sheng Zhao,Yao Qian,Furu Wei
链接:点击下载PDF文件
摘要:本文介绍了VALL-E 2,神经编解码器语言模型的最新进展,标志着zero-shot文本到语音合成(TTS)的一个里程碑,首次实现人类平等。在其前身VALL-E的基础上,新迭代引入了两项重大增强:重复感知采样通过考虑解码历史中的令牌重复来改进原始核采样过程。它不仅稳定了解码,还避免了无限循环问题。分组代码建模将编解码器代码组织成组,有效缩短序列长度,不仅提高了推理速度,还解决了长序列建模的挑战。我们在LibriSpeech和VCTK数据集上的实验表明,VALL-E 2在语音鲁棒性、自然度和说话人相似性方面优于以前的系统。这是第一个在这些基准上实现人类平等的国家。此外,VALL-E 2始终能够合成高质量的语音,即使是由于其复杂性或重复短语而传统上具有挑战性的句子。这项工作的优势可能有助于有价值的努力,例如为失语症患者或肌萎缩侧索硬化症患者生成语音。VALL-E 2的演示将发布到https: aka.ms valle2。摘要:This paper introduces VALL-E 2, the latest advancement in neural codec language models that marks a milestone in zero-shot text-to-speech synthesis (TTS), achieving human parity for the first time. Based on its predecessor, VALL-E, the new iteration introduces two significant enhancements: Repetition Aware Sampling refines the original nucleus sampling process by accounting for token repetition in the decoding history. It not only stabilizes the decoding but also circumvents the infinite loop issue. Grouped Code Modeling organizes codec codes into groups to effectively shorten the sequence length, which not only boosts inference speed but also addresses the challenges of long sequence modeling. Our experiments on the LibriSpeech and VCTK datasets show that VALL-E 2 surpasses previous systems in speech robustness, naturalness, and speaker similarity. It is the first of its kind to reach human parity on these benchmarks. Moreover, VALL-E 2 consistently synthesizes high-quality speech, even for sentences that are traditionally challenging due to their complexity or repetitive phrases. The advantages of this work could contribute to valuable endeavors, such as generating speech for individuals with aphasia or people with amyotrophic lateral sclerosis. Demos of VALL-E 2 will be posted to https: aka.ms valle2.

【28】 Sample Rate Independent Recurrent Neural Networks for Audio Effects Processing
标题: 用于音效处理的独立采样率的回归神经网络
作者:Alistair Carson,Alec Wright,Jatin Chowdhury,Vesa Välimäki,Stefan Bilbao
备注:Accepted for publication in Proc. DAFx24, Guildford, UK, September 2024
链接:点击下载PDF文件
摘要:近年来,对吉他放大器和效果踏板进行建模的机器学习方法得到了广泛的研究,并已成为一些消费产品的标准做法。特别是,递归神经网络(RNN)是对非线性设备(如真空管放大器和失真电路)建模的热门选择。这种模型的一个局限性是,它们是在特定采样率下对音频进行训练的,因此在以另一种速率操作时会给出不可靠的结果。在这里,我们研究了几种修改RNN结构的方法,使它们近似独立于采样率,重点是过采样。在整数过采样的情况下,我们证明了先前提出的基于延迟的方法提供高保真的采样率转换,同时还减少了混叠。对于非整数采样率调整,我们提出了两种新的方法,并表明,其中之一,基于三次拉格朗日插值的延迟线,提供了显着的改进现有的方法。据我们所知,这项工作提供了第一次深入研究这个问题。摘要:In recent years, machine learning approaches to modelling guitar amplifiers and effects pedals have been widely investigated and have become standard practice in some consumer products. In particular, recurrent neural networks (RNNs) are a popular choice for modelling non-linear devices such as vacuum tube amplifiers and distortion circuitry. One limitation of such models is that they are trained on audio at a specific sample rate and therefore give unreliable results when operating at another rate. Here, we investigate several methods of modifying RNN structures to make them approximately sample rate independent, with a focus on oversampling. In the case of integer oversampling, we demonstrate that a previously proposed delay-based approach provides high fidelity sample rate conversion whilst additionally reducing aliasing. For non-integer sample rate adjustment, we propose two novel methods and show that one of these, based on cubic Lagrange interpolation of a delay-line, provides a significant improvement over existing methods. To our knowledge, this work provides the first in-depth study into this problem.

【29】 Label-Looping: Highly Efficient Decoding for Transducers
标题: 标签循环:传感器的高效解码
作者:Vladimir Bataev,Hainan Xu,Daniel Galvez,Vitaly Lavrukhin,Boris Ginsburg
链接:点击下载PDF文件
摘要:本文介绍了一种用于传感器推理的高效贪婪译码算法。我们提出了一种新的数据结构,使用CUDA张量表示部分假设在一批支持并行化的假设操作。在解码过程中,我们的算法通过采用嵌套循环设计来最大限度地提高GPU并行度,其中内部循环消耗所有空白预测,而非空白预测在外部循环中处理。我们的算法是通用的,可以与传统的传感器和令牌和持续时间传感器。实验表明,在批量大小为32的情况下,标签循环算法比传统的批量解码算法可以带来2.0倍的加速比,并且可以与其他编译器或GPU调用相关技术相结合,带来更大的加速比。我们将开源我们的实现,以造福研究界。摘要:This paper introduces a highly efficient greedy decoding algorithm for Transducer inference. We propose a novel data structure using CUDA tensors to represent partial hypotheses in a batch that supports parallelized hypothesis manipulations. During decoding, our algorithm maximizes GPU parallelism by adopting a nested-loop design, where the inner loop consumes all blank predictions, while non-blank predictions are handled in the outer loop. Our algorithm is general-purpose and can work with both conventional Transducers and Token-and-Duration Transducers. Experiments show that the label-looping algorithm can bring a speedup up to 2.0X compared to conventional batched decoding algorithms when using batch size 32, and can be combined with other compiler or GPU call-related techniques to bring more speedup. We will open-source our implementation to benefit the research community.

【30】 EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
标题: EARS:用于语音增强和去回响的无回声全带语音数据集
作者:Julius Richter,Yi-Chiao Wu,Steven Krenn,Simon Welker,Bunlong Lay,Shinji Watanabe,Alexander Richard,Timo Gerkmann
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:我们发布了EARS(表达性无回声语音记录)数据集,这是一个高质量的语音数据集,包括来自不同背景的107个扬声器,总计100小时的干净无回声语音数据。该数据集涵盖了大量不同的说话风格,包括情感语音,不同的阅读风格,非语言声音和会话自由形式的语音。我们对数据集上的语音增强和去混响的各种方法进行基准测试,并通过一组工具度量来评估它们的性能。此外,我们进行了听力测试,20名参与者的语音增强任务,其中生成方法是首选。我们引入了一个盲测试集,允许自动在线评估上传的数据。数据集下载链接和自动评估服务器可以在网上找到。摘要:We release the EARS (Expressive Anechoic Recordings of Speech) dataset, a high-quality speech dataset comprising 107 speakers from diverse backgrounds, totaling in 100 hours of clean, anechoic speech data. The dataset covers a large range of different speaking styles, including emotional speech, different reading styles, non-verbal sounds, and conversational freeform speech. We benchmark various methods for speech enhancement and dereverberation on the dataset and evaluate their performance through a set of instrumental metrics. In addition, we conduct a listening test with 20 participants for the speech enhancement task, where a generative method is preferred. We introduce a blind test set that allows for automatic online evaluation of uploaded data. Dataset download links and automatic evaluation server can be found online.

【31】 JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
标题: JenGAN:基于GAN的语音合成中的堆叠位移过滤器
作者:Hyunjae Cho,Junhyeok Lee,Wonbin Jung
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:基于非自回归GAN的神经声码器由于其快速推理速度和高感知质量而被广泛使用。然而,它们经常遭受可听伪像,例如在其生成的结果中的音调伪像。因此,我们提出了JenGAN,一种新的训练策略,涉及堆叠移位低通滤波器,以确保移位等变属性。此方法有助于防止混淆并减少伪影,同时保留推理期间使用的模型结构。在我们的实验评估中,JenGAN始终如一地增强了声码器模型的性能,在大多数评估指标中获得了显著的优异分数。摘要:Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we propose JenGAN, a new training strategy that involves stacking shifted low-pass filters to ensure the shift-equivariant property. This method helps prevent aliasing and reduce artifacts while preserving the model structure used during inference. In our experimental evaluation, JenGAN consistently enhances the performance of vocoder models, yielding significantly superior scores across the majority of evaluation metrics.

【32】 Soundscape Captioning using Sound Affective Quality Network and Large Language Model
标题: 使用声音情感质量网络和大型语言模型的声景字幕
作者:Yuanbo Hou,Qiaoqiao Ren,Andrew Mitchell,Wenwu Wang,Jian Kang,Tony Belpaeme,Dick Botteldooren
备注:Code: this https URL
链接:点击下载PDF文件
摘要:我们生活在一个丰富多样的声学世界中,个人或社区将其作为声景体验。计算听觉场景分析,通过检测和分类事件来解开声学场景,专注于声音的客观属性,例如它们的类别和时间特征,忽略了声音对人的影响,并且未能探索声音与它们在上下文中唤起的情感之间的关系。为了填补这一空白,并自动化声景分析,传统上依赖于劳动密集型的主观评级和调查,我们提出了声景字幕(SoundSCap)任务。SoundSCap通过捕获声学场景、事件信息和相应的人类情感品质来生成上下文感知的声景描述。为此,我们提出了一个自动音景字幕(SoundSCaper)组成的声学模型,SoundAQnet,和一个通用的大语言模型(LLM)。SoundAQnet同时对声学场景、事件和感知情感质量的多尺度信息进行建模,而LLM通过将SoundAQnet捕获的信息解析为通用语言来生成音景字幕。音景字幕的质量由16位音频 音景专家组成的评审团进行评估。SoundScaper生成的字幕的平均得分(满分5分)比两个音景专家生成的字幕的得分分别低0.21和0.25,在评估集和具有不同长度和声学属性的模型未知混合外部数据集上,但差异在统计学上不显著。总体而言,SoundScaper生成的字幕显示出有前途的性能相比,由音景专家注释的字幕。模型的代码,LLM脚本,人类评估数据和说明,以及专家评估统计数据都是公开的。摘要:We live in a rich and varied acoustic world, which is experienced by individuals or communities as a soundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring the effect of sounds on people and failing to explore the relationship between sounds and the emotions they evoke within a context. To fill this gap and to automate soundscape analysis, which traditionally relies on labour-intensive subjective ratings and surveys, we propose the soundscape captioning (SoundSCap) task. SoundSCap generates context-aware soundscape descriptions by capturing the acoustic scene, event information, and the corresponding human affective qualities. To this end, we propose an automatic soundscape captioner (SoundSCaper) composed of an acoustic model, SoundAQnet, and a general large language model (LLM). SoundAQnet simultaneously models multi-scale information about acoustic scenes, events, and perceived affective qualities, while LLM generates soundscape captions by parsing the information captured by SoundAQnet to a common language. The soundscape caption's quality is assessed by a jury of 16 audio soundscape experts. The average score (out of 5) of SoundSCaper-generated captions is lower than the score of captions generated by two soundscape experts by 0.21 and 0.25, respectively, on the evaluation set and the model-unknown mixed external dataset with varying lengths and acoustic properties, but the differences are not statistically significant. Overall, SoundSCaper-generated captions show promising performance compared to captions annotated by soundscape experts. The models' code, LLM scripts, human assessment data and instructions, and expert evaluation statistics are all publicly available.

【33】 Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
标题: 用于文本到语音合成的自回归扩散Transformer
作者:Zhijun Liu,Shuai Wang,Sho Inoue,Qibing Bai,Haizhou Li
链接:点击下载PDF文件
摘要:音频语言模型最近已经成为各种音频生成任务的一种有前途的方法,依赖于音频标记器将波形编码成离散符号序列。音频标记化通常在编码比特率和重建精度之间提出必要的折衷。当处理低比特率音频代码时,语言模型被限制为仅处理嵌入在音频中的信息的子集,这反过来限制了它们的生成能力。为了规避这些问题,我们建议在连续空间$ mathbb R^d$中将音频编码为矢量序列,并使用仅解码器扩散Transformer(ARDiT)自回归生成这些序列。我们的研究结果表明,ARDiT在zero-shot文本到语音转换方面表现出色,并且表现出与最先进的模型相比甚至超过的性能。高比特率的连续语音表示可以实现几乎完美的重建,使我们的模型能够实现近乎完美的语音编辑。我们的实验表明,采用积分Kullback-Leibler(IKL)发散蒸馏在每个自回归步骤显着提高感知质量的样本。同时,它将扩散模型的迭代采样过程压缩为一个步骤。此外,ARDiT可以被训练成在一个步骤中预测多个连续向量,从而显著减少采样期间的延迟。令人印象深刻的是,我们的模型之一可以产生$170$ ms的$24$ kHz的语音每评估步骤,性能下降最小。音频样本可在http: ardit-tts.github.io 上获得。摘要:Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete symbols. Audio tokenization often poses a necessary compromise between code bitrate and reconstruction accuracy. When dealing with low-bitrate audio codes, language models are constrained to process only a subset of the information embedded in the audio, which in turn restricts their generative capabilities. To circumvent these issues, we propose encoding audio as vector sequences in continuous space $ mathbb R^d$ and autoregressively generating these sequences using a decoder-only diffusion transformer (ARDiT). Our findings indicate that ARDiT excels in zero-shot text-to-speech and exhibits performance that compares to or even surpasses that of state-of-the-art models. High-bitrate continuous speech representation enables almost flawless reconstruction, allowing our model to achieve nearly perfect speech editing. Our experiments reveal that employing Integral Kullback-Leibler (IKL) divergence for distillation at each autoregressive step significantly boosts the perceived quality of the samples. Simultaneously, it condenses the iterative sampling process of the diffusion model into a single step. Furthermore, ARDiT can be trained to predict several continuous vectors in one step, significantly reducing latency during sampling. Impressively, one of our models can generate $170$ ms of $24$ kHz speech per evaluation step with minimal degradation in performance. Audio samples are available at http: ardit-tts.github.io .

【34】 Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
标题: 您应该在TTC中使用概率持续时间模型吗?可能! 尤其是对于自发的言论
作者:Shivam Mehta,Harm Lameris,Rajiv Punmiya,Jonas Beskow,Éva Székely,Gustav Eje Henter
备注:5 pages, 2 figures. Final version, accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在TTS中将输入符号转换为输出音频需要对语音声音的持续时间进行建模。领先的非自回归(NAR)TTS模型将持续时间建模视为回归问题。然后,每次都以相同的时间说出相同的话语,这与人类说话不同。已经提出了持续时间的概率模型,但有各种证据表明它们的好处。然而,先前的研究一般只考虑朗读的言语,而忽略了自发言语,尽管后者是一种更常见和更多变的说话模式。我们比较了传统的确定性持续时间建模的效果,从一个强大的概率模型采样的持续时间基于条件流匹配(OT-CFM),在三个不同的NAR TTS方法:基于回归,深度生成,和端到端。在四个不同的语料库中,随机持续时间建模提高了概率NAR TTS方法,特别是对于自发语音。请访问https: shivammehta25.github.io prob_dur 获取音频和资源。摘要:Converting input symbols to output audio in TTS requires modelling the durations of speech sounds. Leading non-autoregressive (NAR) TTS models treat duration modelling as a regression problem. The same utterance is then spoken with identical timings every time, unlike when a human speaks. Probabilistic models of duration have been proposed, but there is mixed evidence of their benefits. However, prior studies generally only consider speech read aloud, and ignore spontaneous speech, despite the latter being both a more common and a more variable mode of speaking. We compare the effect of conventional deterministic duration modelling to durations sampled from a powerful probabilistic model based on conditional flow matching (OT-CFM), in three different NAR TTS approaches: regression-based, deep generative, and end-to-end. Across four different corpora, stochastic duration modelling improves probabilistic NAR TTS approaches, especially for spontaneous speech. Please see https: shivammehta25.github.io prob_dur for audio and resources.

【35】 Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
标题: 通过自适应神经网络量化实现轻量级说话人验证
作者:Bei Liu,Haoyu Wang,Yanmin Qian
备注:submitted to IEEEACM Transactions on Audio Speech and Language Processing (Under Review)
链接:点击下载PDF文件
摘要:现代说话人验证(SV)系统通常需要昂贵的存储和计算资源,从而阻碍了它们在移动设备上的部署。在本文中,我们探索自适应神经网络量化轻量级说话人确认。首先,我们提出了一种新的自适应均匀精度量化方法,使量化质心的动态生成定制的每个网络层的基础上的k均值聚类。通过将其应用于预训练的SV系统,我们获得了一系列具有不同位宽的量化变体。为了提高低比特量化模型的性能,进一步引入了混合精度量化算法以及多级微调(MSFT)策略。与均匀精度量化不同,混合精度方法允许将不同的位宽分配给不同的网络层。当确定比特组合时,采用MSFT以特定顺序逐步地对网络进行优化和微调。最后,我们设计了两种不同的二进制量化方案,以减轻1位量化模型的性能下降:静态和自适应量化器。在VoxCeleb上的实验表明,ResNets和DF-ResNets都实现了无损4位均匀精度量化,产生了约8的有希望的压缩比。此外,与均匀精度方法相比,混合精度量化不仅获得了类似模型大小的额外性能改进,而且还提供了针对任何期望模型大小生成比特组合的灵活性。此外,我们建议的1位量化方案显着提高二值化模型的性能。最后,与现有的轻量级SV系统的全面比较表明,我们提出的模型优于所有以前的方法在各种模型尺寸范围内的大幅度。摘要:Modern speaker verification (SV) systems typically demand expensive storage and computing resources, thereby hindering their deployment on mobile devices. In this paper, we explore adaptive neural network quantization for lightweight speaker verification. Firstly, we propose a novel adaptive uniform precision quantization method which enables the dynamic generation of quantization centroids customized for each network layer based on k-means clustering. By applying it to the pre-trained SV systems, we obtain a series of quantized variants with different bit widths. To enhance the performance of low-bit quantized models, a mixed precision quantization algorithm along with a multi-stage fine-tuning (MSFT) strategy is further introduced. Unlike uniform precision quantization, mixed precision approach allows for the assignment of varying bit widths to different network layers. When bit combination is determined, MSFT is employed to progressively quantize and fine-tune network in a specific order. Finally, we design two distinct binary quantization schemes to mitigate performance degradation of 1-bit quantized models: the static and adaptive quantizers. Experiments on VoxCeleb demonstrate that lossless 4-bit uniform precision quantization is achieved on both ResNets and DF-ResNets, yielding a promising compression ratio of around 8. Moreover, compared to uniform precision approach, mixed precision quantization not only obtains additional performance improvements with a similar model size but also offers the flexibility to generate bit combination for any desirable model size. In addition, our suggested 1-bit quantization schemes remarkably boost the performance of binarized models. Finally, a thorough comparison with existing lightweight SV systems reveals that our proposed models outperform all previous methods by a large margin across various model size ranges.

【36】 Diversifying and Expanding Frequency-Adaptive Convolution Kernels for Sound Event Detection
标题: 多样化和扩展用于声音事件检测的频率自适应卷积核
作者:Hyeonuk Nam,Seong-Hu Kim,Deokki Min,Junhyeok Lee,Yong-Hwa Park
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:频率动态卷积(FDY conv)在声音事件检测(SED)中显示了最先进的性能,其使用通过基核的频率变化组合获得的频率自适应核。然而,FDY conv缺乏使频率自适应内核多样化的明确手段,从而潜在地限制了性能。此外,基核的大小是有限的,而时频模式跨越更大的谱-时间范围。因此,我们提出了扩展频率动态卷积(DFD conv),它通过向基础内核引入不同的扩展大小来多样化和扩展频率自适应内核。实验表明,沿频率维改变膨胀尺度的优点,并对注意力权重方差分析证明了膨胀基核的有效多样化。通过采用基于交集的F1分数的类中值滤波器,DFD-CRNN在复调音检测分数(PSDS)方面优于FDY-CRNN 3.12%。摘要:Frequency dynamic convolution (FDY conv) has shown the state-of-the-art performance in sound event detection (SED) using frequency-adaptive kernels obtained by frequency-varying combination of basis kernels. However, FDY conv lacks an explicit mean to diversify frequency-adaptive kernels, potentially limiting the performance. In addition, size of basis kernels is limited while time-frequency patterns span larger spectro-temporal range. Therefore, we propose dilated frequency dynamic convolution (DFD conv) which diversifies and expands frequency-adaptive kernels by introducing different dilation sizes to basis kernels. Experiments showed advantages of varying dilation sizes along frequency dimension, and analysis on attention weight variance proved dilated basis kernels are effectively diversified. By adapting class-wise median filter with intersection-based F1 score, proposed DFD-CRNN outperforms FDY-CRNN by 3.12% in terms of polyphonic sound detection score (PSDS).

【37】 LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
标题: LDM-SRC:基于潜在扩散模型的Zero-Shot任意歌唱声音转换,具有歌手引导
作者:Shihao Chen,Yu Gu,Jie Zhang,Na Li,Rilin Chen,Liping Chen,Lirong Dai
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:任意到任意歌声转换(SVC)是一种有趣的音频编辑技术,旨在将一个歌手的歌声转换为另一个歌手的歌声,只需几秒钟的演唱数据。然而,在转换过程中,音色泄漏的问题是不可避免的:转换后的歌声听起来仍然像原始歌手的声音。为了解决这个问题,我们提出了一个潜在的扩散模型SVC(LDM-SVC)在这项工作中,试图执行SVC在潜在的空间使用LDM。我们使用基于VITS框架的开源So-VITS-SVC项目预训练变分自动编码器结构,然后将其用于LDM训练。此外,我们提出了一种基于无分类器指导的歌唱者指导训练方法,以进一步抑制原歌唱者的音色。实验结果表明,该方法优于以往的作品在主观和客观评价的音色相似性。摘要:Any-to-any singing voice conversion (SVC) is an interesting audio editing technique, aiming to convert the singing voice of one singer into that of another, given only a few seconds of singing data. However, during the conversion process, the issue of timbre leakage is inevitable: the converted singing voice still sounds like the original singer's voice. To tackle this, we propose a latent diffusion model for SVC (LDM-SVC) in this work, which attempts to perform SVC in the latent space using an LDM. We pretrain a variational autoencoder structure using the noted open-source So-VITS-SVC project based on the VITS framework, which is then used for the LDM training. Besides, we propose a singer guidance training method based on classifier-free guidance to further suppress the timbre of the original singer. Experimental results show the superiority of the proposed method over previous works in both subjective and objective evaluations of timbre similarity.

【38】 Signal processing algorithm effective for sound quality of hearing loss simulators
标题: 对听力损失模拟器音质有效的信号处理算法
作者:Toshio Irino,Shintaro Doan,Minami Ishikawa
备注:This paper has been accepted for publication in Interspeech 2024
链接:点击下载PDF文件
摘要:听力损失(HL)模拟器,它允许正常听力(NH)的听众体验HL,已被用于语音清晰度实验,但由于可感知的失真,而不是在音质实验。如果它们产生较少的失真,它们可能有助于NH听众评估助听器等的声音质量。我们进行了感知音质实验,比较剑桥版本的HL模拟器(CamHLS)和和歌山版本的HL模拟器(WHIS),其中有两个算法的滤波器组分析合成(FBAS)和直接时变滤波器(DTVF)。实验结果表明,即使在非线性过程中,DTVF的WHIS也比CamHLS和FBAS的WHIS产生更少的语音失真。这一优势主要是由于DTVF算法的使用,该算法可以应用于具有滤波器组分析的各种信号合成应用。摘要:Hearing loss (HL) simulators, which allow normal hearing (NH) listeners to experience HL, have been used in speech intelligibility experiments, but not in sound quality experiments due to perceptible distortion. If they produced less distortion, they might be useful for NH listeners to evaluate the sound quality of, for example, hearing aids. We conducted perceptual sound quality experiments to compare the Cambridge version of HL simulator (CamHLS) and the Wakayama version of the HL simulator (WHIS), which has the two algorithms of filterbank analysis synthesis (FBAS) and direct time-varying filter (DTVF). The experimental results showed that WHIS with DTVF produces less perceptible distortion in speech sounds than CamHLS and WHIS with FBAS, even when the nonlinear process is working. This advantage is mainly due to the use of the DTVF algorithm, which could be applied to various signal synthesis applications with filterbank analysis.

【39】 XANE: eXplainable Acoustic Neural Embeddings
标题: XANE:可扩展的声学神经嵌入
作者:Sri Harsha Dumpala,Dushyant Sharma,Chandramouli Shama Sastri,Stanislav Kruchinin,James Fosburgh,Patrick A. Naylor
链接:点击下载PDF文件
摘要:我们提出了一种新的方法来提取神经嵌入模型的语音信号的背景声学。所提取的嵌入用于以非侵入性方式估计与信号的背景声学特性相关的特定参数,这允许嵌入根据这些参数来解释。我们通过对看不见的测试数据进行聚类实验来说明这些嵌入的价值,并表明所提出的嵌入在三个不同的任务中实现了95.2%的平均F1得分,显著优于基于WavLM的信号嵌入。我们还表明,所提出的方法可以解释的嵌入估计14个声学参数表征的背景声学,包括混响和噪声水平,重叠语音检测,编解码器类型检测和噪声类型检测具有高精度和实时因素17倍低于外部基线方法。摘要:We present a novel method for extracting neural embeddings that model the background acoustics of a speech signal. The extracted embeddings are used to estimate specific parameters related to the background acoustic properties of the signal in a non-intrusive manner, which allows the embeddings to be explainable in terms of those parameters. We illustrate the value of these embeddings by performing clustering experiments on unseen test data and show that the proposed embeddings achieve a mean F1 score of 95.2 % for three different tasks, outperforming significantly the WavLM based signal embeddings. We also show that the proposed method can explain the embeddings by estimating 14 acoustic parameters characterizing the background acoustics, including reverberation and noise levels, overlapped speech detection, CODEC type detection and noise type detection with high accuracy and a real-time factor 17 times lower than an external baseline method.


eess.AS音频处理
【1】 Sample Rate Independent Recurrent Neural Networks for Audio Effects Processing
标题: 用于音效处理的独立采样率的回归神经网络
作者:Alistair Carson,Alec Wright,Jatin Chowdhury,Vesa Välimäki,Stefan Bilbao
备注:Accepted for publication in Proc. DAFx24, Guildford, UK, September 2024
链接:点击下载PDF文件
摘要:近年来,对吉他放大器和效果踏板进行建模的机器学习方法得到了广泛的研究,并已成为一些消费产品的标准做法。特别是,递归神经网络(RNN)是对非线性设备(如真空管放大器和失真电路)建模的热门选择。这种模型的一个局限性是,它们是在特定采样率下对音频进行训练的,因此在以另一种速率操作时会给出不可靠的结果。在这里,我们研究了几种修改RNN结构的方法,使它们近似独立于采样率,重点是过采样。在整数过采样的情况下,我们证明了先前提出的基于延迟的方法提供高保真的采样率转换,同时还减少了混叠。对于非整数采样率调整,我们提出了两种新的方法,并表明,其中之一,基于三次拉格朗日插值的延迟线,提供了显着的改进现有的方法。据我们所知,这项工作提供了第一次深入研究这个问题。摘要:In recent years, machine learning approaches to modelling guitar amplifiers and effects pedals have been widely investigated and have become standard practice in some consumer products. In particular, recurrent neural networks (RNNs) are a popular choice for modelling non-linear devices such as vacuum tube amplifiers and distortion circuitry. One limitation of such models is that they are trained on audio at a specific sample rate and therefore give unreliable results when operating at another rate. Here, we investigate several methods of modifying RNN structures to make them approximately sample rate independent, with a focus on oversampling. In the case of integer oversampling, we demonstrate that a previously proposed delay-based approach provides high fidelity sample rate conversion whilst additionally reducing aliasing. For non-integer sample rate adjustment, we propose two novel methods and show that one of these, based on cubic Lagrange interpolation of a delay-line, provides a significant improvement over existing methods. To our knowledge, this work provides the first in-depth study into this problem.

【2】 Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
标题: 通过有效的微调学习语音生成的细粒度可控性
作者:Chung-Ming Chien,Andros Tjandra,Apoorv Vyas,Matt Le,Bowen Shi,Wei-Ning Hsu
备注:Accepted by InterSpeech 2024
链接:点击下载PDF文件
摘要:随着生成模型的规模不断增长,预训练模型的有效重用和适应已成为关键考虑因素。在这项工作中,我们提出了Voicebox Adapter,一种新的方法,它使用交叉注意模块将细粒度条件集成到预先训练的Voicebox语音生成模型中。为了确保新添加的模块与预先训练的模块的顺利集成,我们探索了各种有效的微调方法。我们的实验表明,LoRA与偏置调谐配置产生最佳的性能,提高可控性,而不影响语音质量。通过三个细粒度条件生成任务,我们证明了Voicebox Adapter的有效性和资源效率。后续实验进一步强调了Voicebox Adapter在不同数据设置中的鲁棒性。摘要:As the scale of generative models continues to grow, efficient reuse and adaptation of pre-trained models have become crucial considerations. In this work, we propose Voicebox Adapter, a novel approach that integrates fine-grained conditions into a pre-trained Voicebox speech generation model using a cross-attention module. To ensure a smooth integration of newly added modules with pre-trained ones, we explore various efficient fine-tuning approaches. Our experiment shows that the LoRA with bias-tuning configuration yields the best performance, enhancing controllability without compromising speech quality. Across three fine-grained conditional generation tasks, we demonstrate the effectiveness and resource efficiency of Voicebox Adapter. Follow-up experiments further highlight the robustness of Voicebox Adapter across diverse data setups.

【3】 Label-Looping: Highly Efficient Decoding for Transducers
标题: 标签循环:传感器的高效解码
作者:Vladimir Bataev,Hainan Xu,Daniel Galvez,Vitaly Lavrukhin,Boris Ginsburg
链接:点击下载PDF文件
摘要:本文介绍了一种用于传感器推理的高效贪婪译码算法。我们提出了一种新的数据结构,使用CUDA张量表示部分假设在一批支持并行化的假设操作。在解码过程中,我们的算法通过采用嵌套循环设计来最大限度地提高GPU并行度,其中内部循环消耗所有空白预测,而非空白预测在外部循环中处理。我们的算法是通用的,可以与传统的传感器和令牌和持续时间传感器。实验表明,在批量大小为32的情况下,标签循环算法比传统的批量解码算法可以带来2.0倍的加速比,并且可以与其他编译器或GPU调用相关技术相结合,带来更大的加速比。我们将开源我们的实现,以造福研究界。摘要:This paper introduces a highly efficient greedy decoding algorithm for Transducer inference. We propose a novel data structure using CUDA tensors to represent partial hypotheses in a batch that supports parallelized hypothesis manipulations. During decoding, our algorithm maximizes GPU parallelism by adopting a nested-loop design, where the inner loop consumes all blank predictions, while non-blank predictions are handled in the outer loop. Our algorithm is general-purpose and can work with both conventional Transducers and Token-and-Duration Transducers. Experiments show that the label-looping algorithm can bring a speedup up to 2.0X compared to conventional batched decoding algorithms when using batch size 32, and can be combined with other compiler or GPU call-related techniques to bring more speedup. We will open-source our implementation to benefit the research community.

【4】 EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
标题: EARS:用于语音增强和去回响的无回声全带语音数据集
作者:Julius Richter,Yi-Chiao Wu,Steven Krenn,Simon Welker,Bunlong Lay,Shinji Watanabe,Alexander Richard,Timo Gerkmann
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:我们发布了EARS(表达性无回声语音记录)数据集,这是一个高质量的语音数据集,包括来自不同背景的107个扬声器,总计100小时的干净无回声语音数据。该数据集涵盖了大量不同的说话风格,包括情感语音,不同的阅读风格,非语言声音和会话自由形式的语音。我们对数据集上的语音增强和去混响的各种方法进行基准测试,并通过一组工具度量来评估它们的性能。此外,我们进行了听力测试,20名参与者的语音增强任务,其中生成方法是首选。我们引入了一个盲测试集,允许自动在线评估上传的数据。数据集下载链接和自动评估服务器可以在网上找到。摘要:We release the EARS (Expressive Anechoic Recordings of Speech) dataset, a high-quality speech dataset comprising 107 speakers from diverse backgrounds, totaling in 100 hours of clean, anechoic speech data. The dataset covers a large range of different speaking styles, including emotional speech, different reading styles, non-verbal sounds, and conversational freeform speech. We benchmark various methods for speech enhancement and dereverberation on the dataset and evaluate their performance through a set of instrumental metrics. In addition, we conduct a listening test with 20 participants for the speech enhancement task, where a generative method is preferred. We introduce a blind test set that allows for automatic online evaluation of uploaded data. Dataset download links and automatic evaluation server can be found online.

【5】 The Effect of Training Dataset Size on Discriminative and Diffusion-Based Speech Enhancement Systems
标题: 训练数据集大小对区分性和基于扩散的语音增强系统的影响
作者:Philippe Gonzalez,Zheng-Hua Tan,Jan Østergaard,Jesper Jensen,Tommy Sonne Alstrøm,Tobias May
链接:点击下载PDF文件
摘要:基于深度神经网络的语音增强系统的性能通常随着训练数据集的大小而增加。然而,研究训练数据集大小对语音增强性能的影响的研究没有考虑最近的方法,例如基于扩散的生成模型。扩散模型通常使用大量数据集进行训练以用于图像生成任务,但这是否也需要用于语音增强尚不清楚。此外,研究训练数据集大小的影响的研究没有控制数据多样性。因此,不清楚性能的提高是由于数据集大小的增加还是多样性的增加。因此,我们系统地研究了训练数据集大小对流行的最先进的基于判别和扩散的语音增强系统的性能的影响。我们通过使用一组固定的语音话语,噪声段和双耳房间脉冲响应来生成不同大小的数据集,从而控制数据的多样性。我们发现,基于扩散的系统并不像判别系统那样受益于增加训练数据集的大小。它们相对于具有10小时或更少数据集的判别系统表现最好,但它们优于具有100小时或更长数据集的判别系统。摘要:The performance of deep neural network-based speech enhancement systems typically increases with the training dataset size. However, studies that investigated the effect of training dataset size on speech enhancement performance did not consider recent approaches, such as diffusion-based generative models. Diffusion models are typically trained with massive datasets for image generation tasks, but whether this is also required for speech enhancement is unknown. Moreover, studies that investigated the effect of training dataset size did not control for the data diversity. It is thus unclear whether the performance improvement was due to the increased dataset size or diversity. Therefore, we systematically investigate the effect of training dataset size on the performance of popular state-of-the-art discriminative and diffusion-based speech enhancement systems. We control for the data diversity by using a fixed set of speech utterances, noise segments and binaural room impulse responses to generate datasets of different sizes. We find that the diffusion-based systems do not benefit from increasing the training dataset size as much as the discriminative systems. They perform the best relative to the discriminative systems with datasets of 10 h or less, but they are outperformed by the discriminative systems with datasets of 100 h or more.

【6】 JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
标题: JenGAN:基于GAN的语音合成中的堆叠位移过滤器
作者:Hyunjae Cho,Junhyeok Lee,Wonbin Jung
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:基于非自回归GAN的神经声码器由于其快速推理速度和高感知质量而被广泛使用。然而,它们经常遭受可听伪像,例如在其生成的结果中的音调伪像。因此,我们提出了JenGAN,一种新的训练策略,涉及堆叠移位低通滤波器,以确保移位等变属性。此方法有助于防止混淆并减少伪影,同时保留推理期间使用的模型结构。在我们的实验评估中,JenGAN始终如一地增强了声码器模型的性能,在大多数评估指标中获得了显著的优异分数。摘要:Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we propose JenGAN, a new training strategy that involves stacking shifted low-pass filters to ensure the shift-equivariant property. This method helps prevent aliasing and reduce artifacts while preserving the model structure used during inference. In our experimental evaluation, JenGAN consistently enhances the performance of vocoder models, yielding significantly superior scores across the majority of evaluation metrics.

【7】 Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech Separation
标题: 分离和重建:语音分离的非对称编解码器
作者:Ui-Hyeop Shin,Sangyoun Lee,Taehan Kim,Hyung-Min Park
备注:Project Page this https URL
链接:点击下载PDF文件
摘要:自时域语音分离成功以来,通过扩展特征序列的长度和通道以增加计算量,已经做出了进一步的改进。在语音分离的研究中,当特征在时间上扩展为长序列时,通常采用双路径模型将其分割成块。特别是,它是共同的过程中分离的功能对应于每个扬声器位于网络的最后一级。然而,主动扩展特征序列以包括说话者的数量作为额外维度是更有利和直观的。在本文中,我们提出了一种非对称的策略,其中的编码器和解码器进行分区,以执行不同的处理分离任务。编码器分析特征,并且编码器的输出被分成要分离的扬声器的数量。分离后的序列在跨话者处理的基础上,通过权值共享解码器进行重构,形成连体网络。通过在解码器中使用Siamese网络,不使用说话人信息,网络直接学习使用分离目标来区分特征。利用公共分割层,用于跳过连接的中间编码器特征也被分割用于基于U-Net结构的重构解码器。此外,我们设计了全局和局部Transformer块来直接处理长序列,而不是将特征分割成块作为双路径。实验结果表明,该分离重构框架是有效的,全局和局部Transformer的结合可以充分替代双路径结构中组块间和组块内处理的作用。最后,所提出的模型,包括这两个实现了最先进的性能与更少的计算比以前在各种基准数据集。摘要:Since the success of a time-domain speech separation, further improvements have been made by expanding the length and channel of a feature sequence to increase the amount of computation. When temporally expanded to a long sequence, the feature is segmented into chunks as a dual-path model in most studies of speech separation. In particular, it is common for the process of separating features corresponding to each speaker to be located in the final stage of the network. However, it is more advantageous and intuitive to proactively expand the feature sequence to include the number of speakers as an extra dimension. In this paper, we present an asymmetric strategy in which the encoder and decoder are partitioned to perform distinct processing in separation tasks. The encoder analyzes features, and the output of the encoder is split into the number of speakers to be separated. The separated sequences are then reconstructed by the weight-shared decoder, as Siamese network, in addition to cross-speaker processing. By using the Siamese network in the decoder, without using speaker information, the network directly learns to discriminate the features using a separation objective. With a common split layer, intermediate encoder features for skip connections are also split for the reconstruction decoder based on the U-Net structure. In addition, instead of segmenting the feature into chunks as dual-path, we design global and local Transformer blocks to directly process long sequences. The experimental results demonstrated that this separation-and-reconstruction framework is effective and that the combination of proposed global and local Transformer can sufficiently replace the role of inter- and intra-chunk processing in dual-path structure. Finally, the presented model including both of these achieved state-of-the-art performance with less computation than before in various benchmark datasets.

【8】 Prompting Large Language Models with Audio for General-Purpose Speech Summarization
标题: 使用音频嵌入大型语言模型以实现通用语音摘要
作者:Wonjune Kang,Deb Roy
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了一个框架,利用大型语言模型(LLM)的处理和推理能力的语音摘要。我们提出了一个端到端的系统,它结合了一个语音调谐LLM与音频编码器,将语音转换成令牌表示的LLM可以解释。使用具有配对的语音-文本数据的数据集,整个系统被训练为对具有相同语义信息的提示生成一致的响应,而不管输入模态如何。由此产生的框架允许LLM以与文本相同的方式处理语音输入,通过简单地提示LLM来实现语音摘要。与以前的方法不同,我们的方法能够从任何任意领域总结口语内容,并且可以通过改变LLM提示策略来生成不同风格的摘要。实验表明,我们的方法优于级联基线的语音识别,其次是LLM文本处理。摘要:In this work, we introduce a framework for speech summarization that leverages the processing and reasoning capabilities of large language models (LLMs). We propose an end-to-end system that combines an instruction-tuned LLM with an audio encoder that converts speech into token representations that the LLM can interpret. Using a dataset with paired speech-text data, the overall system is trained to generate consistent responses to prompts with the same semantic information regardless of the input modality. The resulting framework allows the LLM to process speech inputs in the same way as text, enabling speech summarization by simply prompting the LLM. Unlike prior approaches, our method is able to summarize spoken content from any arbitrary domain, and it can produce summaries in different styles by varying the LLM prompting strategy. Experiments demonstrate that our approach outperforms a cascade baseline of speech recognition followed by LLM text processing.

【9】 MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
标题: MakeSinger:一种半监督训练方法,通过无分类器扩散引导实现数据高效的歌唱声音合成
作者:Semin Kim,Myeonghun Jeong,Hyeonseung Lee,Minchan Kim,Byoung Jin Choi,Nam Soo Kim
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了MakeSinger,一种通过无分类器扩散指导的歌唱声音合成(SVS)的半监督训练方法。SVS中的挑战在于收集对齐的文本、音高和音频数据集的昂贵过程。MakeSinger能够从任何语音和歌声数据中训练基于扩散的SVS模型,而不管其标签如何,从而提高了具有大量未标记数据的生成语音的质量。在推理中,我们的新的双重指导机制通过估计掩蔽输入的分数在反向扩散步骤上提供文本和音高指导。实验结果表明,以半监督方式训练的模型在发音,音高准确性和整体质量方面优于仅在标记数据上训练的其他基线。此外,我们证明,通过在训练中添加文本到语音(TTS)数据,该模型可以合成TTS扬声器的歌声,即使没有他们的歌声。摘要:In this paper, we propose MakeSinger, a semi-supervised training method for singing voice synthesis (SVS) via classifier-free diffusion guidance. The challenge in SVS lies in the costly process of gathering aligned sets of text, pitch, and audio data. MakeSinger enables the training of the diffusion-based SVS model from any speech and singing voice data regardless of its labeling, thereby enhancing the quality of generated voices with large amount of unlabeled data. At inference, our novel dual guiding mechanism gives text and pitch guidance on the reverse diffusion step by estimating the score of masked input. Experimental results show that the model trained in a semi-supervised manner outperforms other baselines trained only on the labeled data in terms of pronunciation, pitch accuracy and overall quality. Furthermore, we demonstrate that by adding Text-to-Speech (TTS) data in training, the model can synthesize the singing voices of TTS speakers even without their singing voices.

【10】 BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
标题: BS-PLCNet 2:具有模型内知识蒸馏的两级带分裂分组丢失隐藏网络
作者:Zihan Zhang,Xianjun Xia,Chuanzeng Huang,Yijian Xiao,Lei Xie
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:语音数据包丢失是实时语音通信中不可避免的问题。最近提出了一种针对全频带信号的频带分裂分组丢失隐藏网络(BS-PLCNet)。尽管BS-PLCNet在ICASSP 2024 PLC挑战赛中表现出色,但它是一个大型模型,具有8.95G FLOPS的高计算复杂度。本文提出了其更新版本BS-PLCNet 2,以降低计算复杂度并进一步提高性能。具体来说,为了补偿丢失的未来信息,在宽带模块中,我们设计了一个双路径编码器结构(具有非因果路径和因果路径),并利用模型内知识蒸馏策略从非因果教师到临时学生提取未来信息。此外,我们引入了一个轻量级的后处理模块后,数据包丢失恢复恢复语音失真和去除残留噪声的音频信号。BS-PLCNet 2仅保留了40%的原始参数,比ICASSP 2024 PLC挑战盲集提高了0.18 PLCMOS,在该数据集上实现了最先进的性能。摘要:Audio packet loss is an inevitable problem in real-time speech communication. A band-split packet loss concealment network (BS-PLCNet) targeting full-band signals was recently proposed. Although it performs superiorly in the ICASSP 2024 PLC Challenge, BS-PLCNet is a large model with high computational complexity of 8.95G FLOPS. This paper presents its updated version, BS-PLCNet 2, to reduce computational complexity and improve performance further. Specifically, to compensate for the missing future information, in the wide-band module, we design a dual-path encoder structure (with non-causal and causal path) and leverage an intra-model knowledge distillation strategy to distill the future information from the non-causal teacher to the casual student. Moreover, we introduce a lightweight post-processing module after packet loss restoration to recover speech distortions and remove residual noise in the audio signal. With only 40% of original parameters in BS-PLCNet, BS-PLCNet 2 brings 0.18 PLCMOS improvement on the ICASSP 2024 PLC challenge blind set, achieving state-of-the-art performance on this dataset.

【11】 Accent Conversion with Articulatory Representations
标题: 带有关节表达的口音转换
作者:Yashish M. Siriwardena,Nathan Swedlow,Audrey Howard,Evan Gitterman,Dan Darcy,Carol Espy-Wilson,Andrea Fanelli
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:将非本族口音语音转换为本族(美国)英语具有广泛的应用,例如提高非本族语音的可懂度。先前在这一领域的工作已经使用语音后验图作为目标语音表示来训练声学模型,该声学模型然后用于提取输入语音的紧凑表示以用于口音转换。在这项工作中,我们介绍了使用一个有效的发音语音表示,从声学发音语音反转系统中提取的想法,以提高口音转换中使用的声学模型。结合发音表征的想法源于他们在演讲中很好地描述口音的能力。为了将发音表征与传统的语音后验图相结合,提出了一种基于多任务学习的声学模型。客观和主观评价表明,发音表征的使用可以提高口音转换的有效性。摘要:Conversion of non-native accented speech to native (American) English has a wide range of applications such as improving intelligibility of non-native speech. Previous work on this domain has used phonetic posteriograms as the target speech representation to train an acoustic model which is then used to extract a compact representation of input speech for accent conversion. In this work, we introduce the idea of using an effective articulatory speech representation, extracted from an acoustic-to-articulatory speech inversion system, to improve the acoustic model used in accent conversion. The idea to incorporate articulatory representations originates from their ability to well characterize accents in speech. To incorporate articulatory representations with conventional phonetic posteriograms, a multi-task learning based acoustic model is proposed. Objective and subjective evaluations show that the use of articulatory representations can improve the effectiveness of accent conversion.

【12】 Soundscape Captioning using Sound Affective Quality Network and Large Language Model
标题: 使用声音情感质量网络和大型语言模型的声景字幕
作者:Yuanbo Hou,Qiaoqiao Ren,Andrew Mitchell,Wenwu Wang,Jian Kang,Tony Belpaeme,Dick Botteldooren
备注:Code: this https URL
链接:点击下载PDF文件
摘要:我们生活在一个丰富多样的声学世界中,个人或社区将其作为声景体验。计算听觉场景分析,通过检测和分类事件来解开声学场景,专注于声音的客观属性,例如它们的类别和时间特征,忽略了声音对人的影响,并且未能探索声音与它们在上下文中唤起的情感之间的关系。为了填补这一空白,并自动化声景分析,传统上依赖于劳动密集型的主观评级和调查,我们提出了声景字幕(SoundSCap)任务。SoundSCap通过捕获声学场景、事件信息和相应的人类情感品质来生成上下文感知的声景描述。为此,我们提出了一个自动音景字幕(SoundSCaper)组成的声学模型,SoundAQnet,和一个通用的大语言模型(LLM)。SoundAQnet同时对声学场景、事件和感知情感质量的多尺度信息进行建模,而LLM通过将SoundAQnet捕获的信息解析为通用语言来生成音景字幕。音景字幕的质量由16位音频 音景专家组成的评审团进行评估。SoundScaper生成的字幕的平均得分(满分5分)比两个音景专家生成的字幕的得分分别低0.21和0.25,在评估集和具有不同长度和声学属性的模型未知混合外部数据集上,但差异在统计学上不显著。总体而言,SoundScaper生成的字幕显示出有前途的性能相比,由音景专家注释的字幕。模型的代码,LLM脚本,人类评估数据和说明,以及专家评估统计数据都是公开的。摘要:We live in a rich and varied acoustic world, which is experienced by individuals or communities as a soundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring the effect of sounds on people and failing to explore the relationship between sounds and the emotions they evoke within a context. To fill this gap and to automate soundscape analysis, which traditionally relies on labour-intensive subjective ratings and surveys, we propose the soundscape captioning (SoundSCap) task. SoundSCap generates context-aware soundscape descriptions by capturing the acoustic scene, event information, and the corresponding human affective qualities. To this end, we propose an automatic soundscape captioner (SoundSCaper) composed of an acoustic model, SoundAQnet, and a general large language model (LLM). SoundAQnet simultaneously models multi-scale information about acoustic scenes, events, and perceived affective qualities, while LLM generates soundscape captions by parsing the information captured by SoundAQnet to a common language. The soundscape caption's quality is assessed by a jury of 16 audio soundscape experts. The average score (out of 5) of SoundSCaper-generated captions is lower than the score of captions generated by two soundscape experts by 0.21 and 0.25, respectively, on the evaluation set and the model-unknown mixed external dataset with varying lengths and acoustic properties, but the differences are not statistically significant. Overall, SoundSCaper-generated captions show promising performance compared to captions annotated by soundscape experts. The models' code, LLM scripts, human assessment data and instructions, and expert evaluation statistics are all publicly available.

【13】 MaLa-ASR: Multimedia-Assisted LLM-Based ASR
标题: MaLa-ASB:多媒体辅助的基于LLM的ASB
作者:Guanrou Yang,Ziyang Ma,Fan Yu,Zhifu Gao,Shiliang Zhang,Xie Chen
链接:点击下载PDF文件
摘要:随着越来越多的信息丰富的数据,如视频变得可用,利用多模态辅助信息来增强音频任务引起了广泛的研究兴趣。最近对基于LLM的音频模型的研究激增,为解决音频任务提供了新的视角。鉴于LLM可以灵活地摄取多个输入,我们提出了MaLa-ASR,一个基于LLM的ASR模型,可以集成从演示幻灯片中提取的文本关键字,以提高对会议内容的识别。MaLa-ASR在SlideSpeech语料库的L95和S95子集上产生的平均WER分别为9.4%和11.7%,比SlideSpeech中报告的基线模型显著相对WER下降了27.9%和44.7%。MaLa-ASR强调了LLM在语音任务中的强大性能和方便整合辅助信息的能力。通过在输入提示中添加关键词,偏误率(B-WER)相对降低了46.0%和44.2%,建立了一个新的SOTA。摘要:As more and more information-rich data like video become available, utilizing multi-modal auxiliary information to enhance audio tasks has sparked widespread research interest. The recent surge in research on LLM-based audio models provides fresh perspectives for tackling audio tasks. Given that LLM can flexibly ingest multiple inputs, we propose MaLa-ASR, an LLM-based ASR model that can integrate textual keywords extracted from presentation slides to improve recognition of conference content. MaLa-ASR yields average WERs of 9.4% and 11.7% on the L95 and S95 subsets of the SlideSpeech corpus, representing a significant relative WER drop of 27.9% and 44.7% over the baseline model reported in SlideSpeech. MaLa-ASR underscores LLM's strong performance in speech tasks and the capability to integrate auxiliary information conveniently. By adding keywords to the input prompt, the biased word error rate (B-WER) reduces relatively by 46.0% and 44.2%, establishing a new SOTA on this dataset.

【14】 WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
标题: WenetSpeech 4TTC:用于大型语音生成模型基准的12,800小时普通话TTC数据库
作者:Linhan Ma,Dake Guo,Kun Song,Yuepeng Jiang,Shuai Wang,Liumeng Xue,Weiming Xu,Huan Zhao,Binbin Zhang,Lei Xie
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:随着大规模文本到语音(TTS)模型的发展和训练数据的扩大,最先进的TTS系统已经取得了令人印象深刻的性能。在本文中,我们提出了WenetSpeech 4 TTS,一个多域的普通话语料库来自开源的WenetSpeech数据集。针对文本到语音的任务,我们通过调整段边界,提高音频质量,并消除每个段内的扬声器混合来改进WenetSpeech。在更准确的转录过程和基于质量的数据过滤过程之后,获得的WenetSpeech 4 TTS语料库包含12,800 $小时的配对音频文本数据。此外,我们还创建了不同大小的子集,按片段质量分数进行分类,以允许TTS模型训练和微调。VALL-E和NaturalSpeech 2系统在这些子集上进行了训练和微调,以验证WenetSpeech 4 TTS的可用性,建立了TTS系统公平比较的基准基线。语料库和相应的基准可以在huggingface上公开获得。摘要:With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus derived from the open-sourced WenetSpeech dataset. Tailored for the text-to-speech tasks, we refined WenetSpeech by adjusting segment boundaries, enhancing the audio quality, and eliminating speaker mixing within each segment. Following a more accurate transcription process and quality-based data filtering process, the obtained WenetSpeech4TTS corpus contains $12,800$ hours of paired audio-text data. Furthermore, we have created subsets of varying sizes, categorized by segment quality scores to allow for TTS model training and fine-tuning. VALL-E and NaturalSpeech 2 systems are trained and fine-tuned on these subsets to validate the usability of WenetSpeech4TTS, establishing baselines on benchmark for fair comparison of TTS systems. The corpus and corresponding benchmarks are publicly available on huggingface.

【15】 An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
标题: 基于流匹配的零激发TTC的噪音鲁棒性研究
作者:Xiaofei Wang,Sefik Emre Eskimez,Manthan Thakker,Hemin Yang,Zirun Zhu,Min Tang,Yufei Xia,Jinzhu Li,Sheng Zhao,Jinyu Li,Naoyuki Kanda
备注:Accepted to INTERSPEECH2024
链接:点击下载PDF文件
摘要:最近,能够从短的音频提示合成任何说话者的语音的zero-shot文本到语音(TTS)系统取得了迅速的进步。然而,当音频提示包含噪声时,生成的语音的质量显著恶化,并且已经进行了有限的研究来解决这个问题。在本文中,我们探讨了各种策略,以提高从嘈杂的音频提示的背景下,流匹配的zero-shot TTS的音频质量。我们的研究包括全面的训练策略:无监督预训练与掩蔽语音去噪,多说话人检测和基于DNSMOS的数据过滤预训练数据,以及微调与随机噪声混合。我们的实验结果表明,显着改善的可懂度,扬声器的相似性,和整体音频质量相比,应用语音增强的音频提示的方法。摘要:Recently, zero-shot text-to-speech (TTS) systems, capable of synthesizing any speaker's voice from a short audio prompt, have made rapid advancements. However, the quality of the generated speech significantly deteriorates when the audio prompt contains noise, and limited research has been conducted to address this issue. In this paper, we explored various strategies to enhance the quality of audio generated from noisy audio prompts within the context of flow-matching-based zero-shot TTS. Our investigation includes comprehensive training strategies: unsupervised pre-training with masked speech denoising, multi-speaker detection and DNSMOS-based data filtering on the pre-training data, and fine-tuning with random noise mixing. The results of our experiments demonstrate significant improvements in intelligibility, speaker similarity, and overall audio quality compared to the approach of applying speech enhancement to the audio prompt.

【16】 Text-aware and Context-aware Expressive Audiobook Speech Synthesis
标题: 文本感知和上下文感知表达式有声读物语音合成
作者:Dake Guo,Xinfa Zhu,Liumeng Xue,Yongmao Zhang Li,Wenjie Tian,Lei Xie
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:文语转换技术的最新进展极大地提高了合成语音的表达能力。然而,一个主要的挑战仍然是在生成的语音,捕捉专业的叙述者在有声读物中表现出的不同风格,而不依赖于手动标记的数据或参考speech.To解决这个问题,我们提出了一个文本感知和上下文感知(TACA)风格建模方法表达有声读物语音合成。我们首先建立了一个文本感知的风格空间,通过对比学习的监督下的讲话风格涵盖不同的风格。同时,我们采用了一个上下文编码器,将跨句子的信息和风格嵌入从文本中获得。最后,我们将上下文编码器引入到两种典型的文语转换系统模型中,包括基于VITS的文语转换系统和基于语言模型的文语转换系统。实验结果表明,该方法可以有效地捕捉不同的风格和连贯的韵律,从而提高了有声读物语音合成的自然度和表现力。摘要:Recent advances in text-to-speech have significantly improved the expressiveness of synthetic speech. However, a major challenge remains in generating speech that captures the diverse styles exhibited by professional narrators in audiobooks without relying on manually labeled data or reference speech.To address this problem, we propose a text-aware and context-aware(TACA) style modeling approach for expressive audiobook speech synthesis. We first establish a text-aware style space to cover diverse styles via contrastive learning with the supervision of the speech style. Meanwhile, we adopt a context encoder to incorporate cross-sentence information and the style embedding obtained from text. Finally, we introduce the context encoder to two typical TTS models, including VITS-based TTS and language model-based TTS. Experimental results demonstrate that our proposed approach can effectively capture diverse styles and coherent prosody, and consequently improves naturalness and expressiveness in audiobook speech synthesis.

【17】 Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
标题: 用于文本到语音合成的自回归扩散Transformer
作者:Zhijun Liu,Shuai Wang,Sho Inoue,Qibing Bai,Haizhou Li
链接:点击下载PDF文件
摘要:音频语言模型最近已经成为各种音频生成任务的一种有前途的方法,依赖于音频标记器将波形编码成离散符号序列。音频标记化通常在编码比特率和重建精度之间提出必要的折衷。当处理低比特率音频代码时,语言模型被限制为仅处理嵌入在音频中的信息的子集,这反过来限制了它们的生成能力。为了规避这些问题,我们建议在连续空间$ mathbb R^d$中将音频编码为矢量序列,并使用仅解码器扩散Transformer(ARDiT)自回归生成这些序列。我们的研究结果表明,ARDiT在zero-shot文本到语音转换方面表现出色,并且表现出与最先进的模型相比甚至超过的性能。高比特率的连续语音表示可以实现几乎完美的重建,使我们的模型能够实现近乎完美的语音编辑。我们的实验表明,采用积分Kullback-Leibler(IKL)发散蒸馏在每个自回归步骤显着提高感知质量的样本。同时,它将扩散模型的迭代采样过程压缩为一个步骤。此外,ARDiT可以被训练成在一个步骤中预测多个连续向量,从而显著减少采样期间的延迟。令人印象深刻的是,我们的模型之一可以产生$170$ ms的$24$ kHz的语音每评估步骤,性能下降最小。音频样本可在http: ardit-tts.github.io 上获得。摘要:Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete symbols. Audio tokenization often poses a necessary compromise between code bitrate and reconstruction accuracy. When dealing with low-bitrate audio codes, language models are constrained to process only a subset of the information embedded in the audio, which in turn restricts their generative capabilities. To circumvent these issues, we propose encoding audio as vector sequences in continuous space $ mathbb R^d$ and autoregressively generating these sequences using a decoder-only diffusion transformer (ARDiT). Our findings indicate that ARDiT excels in zero-shot text-to-speech and exhibits performance that compares to or even surpasses that of state-of-the-art models. High-bitrate continuous speech representation enables almost flawless reconstruction, allowing our model to achieve nearly perfect speech editing. Our experiments reveal that employing Integral Kullback-Leibler (IKL) divergence for distillation at each autoregressive step significantly boosts the perceived quality of the samples. Simultaneously, it condenses the iterative sampling process of the diffusion model into a single step. Furthermore, ARDiT can be trained to predict several continuous vectors in one step, significantly reducing latency during sampling. Impressively, one of our models can generate $170$ ms of $24$ kHz speech per evaluation step with minimal degradation in performance. Audio samples are available at http: ardit-tts.github.io .

【18】 Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
标题: 您应该在TTC中使用概率持续时间模型吗?可能! 尤其是对于自发的言论
作者:Shivam Mehta,Harm Lameris,Rajiv Punmiya,Jonas Beskow,Éva Székely,Gustav Eje Henter
备注:5 pages, 2 figures. Final version, accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在TTS中将输入符号转换为输出音频需要对语音声音的持续时间进行建模。领先的非自回归(NAR)TTS模型将持续时间建模视为回归问题。然后,每次都以相同的时间说出相同的话语,这与人类说话不同。已经提出了持续时间的概率模型,但有各种证据表明它们的好处。然而,先前的研究一般只考虑朗读的言语,而忽略了自发言语,尽管后者是一种更常见和更多变的说话模式。我们比较了传统的确定性持续时间建模的效果,从一个强大的概率模型采样的持续时间基于条件流匹配(OT-CFM),在三个不同的NAR TTS方法:基于回归,深度生成,和端到端。在四个不同的语料库中,随机持续时间建模提高了概率NAR TTS方法,特别是对于自发语音。请访问https: shivammehta25.github.io prob_dur 获取音频和资源。摘要:Converting input symbols to output audio in TTS requires modelling the durations of speech sounds. Leading non-autoregressive (NAR) TTS models treat duration modelling as a regression problem. The same utterance is then spoken with identical timings every time, unlike when a human speaks. Probabilistic models of duration have been proposed, but there is mixed evidence of their benefits. However, prior studies generally only consider speech read aloud, and ignore spontaneous speech, despite the latter being both a more common and a more variable mode of speaking. We compare the effect of conventional deterministic duration modelling to durations sampled from a powerful probabilistic model based on conditional flow matching (OT-CFM), in three different NAR TTS approaches: regression-based, deep generative, and end-to-end. Across four different corpora, stochastic duration modelling improves probabilistic NAR TTS approaches, especially for spontaneous speech. Please see https: shivammehta25.github.io prob_dur for audio and resources.

【19】 Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
标题: 通过自适应神经网络量化实现轻量级说话人验证
作者:Bei Liu,Haoyu Wang,Yanmin Qian
备注:submitted to IEEEACM Transactions on Audio Speech and Language Processing (Under Review)
链接:点击下载PDF文件
摘要:现代说话人验证(SV)系统通常需要昂贵的存储和计算资源,从而阻碍了它们在移动设备上的部署。在本文中,我们探索自适应神经网络量化轻量级说话人确认。首先,我们提出了一种新的自适应均匀精度量化方法,使量化质心的动态生成定制的每个网络层的基础上的k均值聚类。通过将其应用于预训练的SV系统,我们获得了一系列具有不同位宽的量化变体。为了提高低比特量化模型的性能,进一步引入了混合精度量化算法以及多级微调(MSFT)策略。与均匀精度量化不同,混合精度方法允许将不同的位宽分配给不同的网络层。当确定比特组合时,采用MSFT以特定顺序逐步地对网络进行优化和微调。最后,我们设计了两种不同的二进制量化方案,以减轻1位量化模型的性能下降:静态和自适应量化器。在VoxCeleb上的实验表明,ResNets和DF-ResNets都实现了无损4位均匀精度量化,产生了约8的有希望的压缩比。此外,与均匀精度方法相比,混合精度量化不仅获得了类似模型大小的额外性能改进,而且还提供了针对任何期望模型大小生成比特组合的灵活性。此外,我们建议的1位量化方案显着提高二值化模型的性能。最后,与现有的轻量级SV系统的全面比较表明,我们提出的模型优于所有以前的方法在各种模型尺寸范围内的大幅度。摘要:Modern speaker verification (SV) systems typically demand expensive storage and computing resources, thereby hindering their deployment on mobile devices. In this paper, we explore adaptive neural network quantization for lightweight speaker verification. Firstly, we propose a novel adaptive uniform precision quantization method which enables the dynamic generation of quantization centroids customized for each network layer based on k-means clustering. By applying it to the pre-trained SV systems, we obtain a series of quantized variants with different bit widths. To enhance the performance of low-bit quantized models, a mixed precision quantization algorithm along with a multi-stage fine-tuning (MSFT) strategy is further introduced. Unlike uniform precision quantization, mixed precision approach allows for the assignment of varying bit widths to different network layers. When bit combination is determined, MSFT is employed to progressively quantize and fine-tune network in a specific order. Finally, we design two distinct binary quantization schemes to mitigate performance degradation of 1-bit quantized models: the static and adaptive quantizers. Experiments on VoxCeleb demonstrate that lossless 4-bit uniform precision quantization is achieved on both ResNets and DF-ResNets, yielding a promising compression ratio of around 8. Moreover, compared to uniform precision approach, mixed precision quantization not only obtains additional performance improvements with a similar model size but also offers the flexibility to generate bit combination for any desirable model size. In addition, our suggested 1-bit quantization schemes remarkably boost the performance of binarized models. Finally, a thorough comparison with existing lightweight SV systems reveals that our proposed models outperform all previous methods by a large margin across various model size ranges.

【20】 Diversifying and Expanding Frequency-Adaptive Convolution Kernels for Sound Event Detection
标题: 多样化和扩展用于声音事件检测的频率自适应卷积核
作者:Hyeonuk Nam,Seong-Hu Kim,Deokki Min,Junhyeok Lee,Yong-Hwa Park
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:频率动态卷积(FDY conv)在声音事件检测(SED)中显示了最先进的性能,其使用通过基核的频率变化组合获得的频率自适应核。然而,FDY conv缺乏使频率自适应内核多样化的明确手段,从而潜在地限制了性能。此外,基核的大小是有限的,而时频模式跨越更大的谱-时间范围。因此,我们提出了扩展频率动态卷积(DFD conv),它通过向基础内核引入不同的扩展大小来多样化和扩展频率自适应内核。实验表明,沿频率维改变膨胀尺度的优点,并对注意力权重方差分析证明了膨胀基核的有效多样化。通过采用基于交集的F1分数的类中值滤波器,DFD-CRNN在复调音检测分数(PSDS)方面优于FDY-CRNN 3.12%。摘要:Frequency dynamic convolution (FDY conv) has shown the state-of-the-art performance in sound event detection (SED) using frequency-adaptive kernels obtained by frequency-varying combination of basis kernels. However, FDY conv lacks an explicit mean to diversify frequency-adaptive kernels, potentially limiting the performance. In addition, size of basis kernels is limited while time-frequency patterns span larger spectro-temporal range. Therefore, we propose dilated frequency dynamic convolution (DFD conv) which diversifies and expands frequency-adaptive kernels by introducing different dilation sizes to basis kernels. Experiments showed advantages of varying dilation sizes along frequency dimension, and analysis on attention weight variance proved dilated basis kernels are effectively diversified. By adapting class-wise median filter with intersection-based F1 score, proposed DFD-CRNN outperforms FDY-CRNN by 3.12% in terms of polyphonic sound detection score (PSDS).

【21】 To what extent can ASV systems naturally defend against spoofing attacks?
标题: ASV系统可以在多大程度上自然防御欺骗攻击?
作者:Jee-weon Jung,Xin Wang,Nicholas Evans,Shinji Watanabe,Hye-jin Shim,Hemlata Tak,Sidhhant Arora,Junichi Yamagishi,Joon Son Chung
备注:5 pages, 3 figures, 3 tables, Interspeech 2024
链接:点击下载PDF文件
摘要:目前的自动说话人确认(ASV)任务涉及两种类型的试验:目标和非目标的二元决策。然而,语音生成技术的新兴进步对ASV系统的可靠性构成了重大威胁。这项研究调查了ASV是否毫不费力地获得了对欺骗攻击的鲁棒性(即,zero-shot能力),通过系统地探索各种ASV系统和欺骗攻击,从传统的尖端技术。通过对8个不同的ASV系统和29个欺骗攻击系统进行广泛的分析,我们证明了ASV的演变本质上包含了对欺骗攻击的防御机制。然而,我们的研究结果也强调,欺骗攻击的进步远远超过ASV系统,因此有必要进一步研究欺骗强大的ASV方法。摘要:The current automatic speaker verification (ASV) task involves making binary decisions on two types of trials: target and non-target. However, emerging advancements in speech generation technology pose significant threats to the reliability of ASV systems. This study investigates whether ASV effortlessly acquires robustness against spoofing attacks (i.e., zero-shot capability) by systematically exploring diverse ASV systems and spoofing attacks, ranging from traditional to cutting-edge techniques. Through extensive analyses conducted on eight distinct ASV systems and 29 spoofing attack systems, we demonstrate that the evolution of ASV inherently incorporates defense mechanisms against spoofing attacks. Nevertheless, our findings also underscore that the advancement of spoofing attacks far outpaces that of ASV systems, hence necessitating further research on spoofing-robust ASV methodologies.

【22】 LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
标题: LDM-SRC:基于潜在扩散模型的Zero-Shot任意歌唱声音转换,具有歌手引导
作者:Shihao Chen,Yu Gu,Jie Zhang,Na Li,Rilin Chen,Liping Chen,Lirong Dai
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:任意到任意歌声转换(SVC)是一种有趣的音频编辑技术,旨在将一个歌手的歌声转换为另一个歌手的歌声,只需几秒钟的演唱数据。然而,在转换过程中,音色泄漏的问题是不可避免的:转换后的歌声听起来仍然像原始歌手的声音。为了解决这个问题,我们提出了一个潜在的扩散模型SVC(LDM-SVC)在这项工作中,试图执行SVC在潜在的空间使用LDM。我们使用基于VITS框架的开源So-VITS-SVC项目预训练变分自动编码器结构,然后将其用于LDM训练。此外,我们提出了一种基于无分类器指导的歌唱者指导训练方法,以进一步抑制原歌唱者的音色。实验结果表明,该方法优于以往的作品在主观和客观评价的音色相似性。摘要:Any-to-any singing voice conversion (SVC) is an interesting audio editing technique, aiming to convert the singing voice of one singer into that of another, given only a few seconds of singing data. However, during the conversion process, the issue of timbre leakage is inevitable: the converted singing voice still sounds like the original singer's voice. To tackle this, we propose a latent diffusion model for SVC (LDM-SVC) in this work, which attempts to perform SVC in the latent space using an LDM. We pretrain a variational autoencoder structure using the noted open-source So-VITS-SVC project based on the VITS framework, which is then used for the LDM training. Besides, we propose a singer guidance training method based on classifier-free guidance to further suppress the timbre of the original singer. Experimental results show the superiority of the proposed method over previous works in both subjective and objective evaluations of timbre similarity.

【23】 Relational Proxy Loss for Audio-Text based Keyword Spotting
标题: 基于音频文本的关键词发现的关系代理丢失
作者:Youngmoon Jung,Seungjin Lee,Joon-Young Yang,Jaeyoung Roh,Chang Woo Han,Hoon-Young Cho
备注:5 pages, 2 figures, Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,人们越来越关注用户的便利性,这导致人们对基于文本的关键字注册系统的兴趣增加。由于系统在注册阶段使用文本输入,在实际使用过程中使用音频输入,因此我们将此任务称为基于音频文本的KWS。为了实现这一任务,声学和文本编码器通常都使用深度度量学习损失函数进行训练,例如基于三元组和代理的损失。本研究旨在通过利用声学嵌入和文本嵌入中的结构关系来改进现有方法。与以前的研究,只比较声学和文本嵌入在点对点的基础上,我们的方法侧重于嵌入空间内的关系结构,通过引入关系代理损失(RPL)的概念。通过将RPL,我们证明了提高性能的华尔街日报(WSJ)语料库。摘要:In recent years, there has been an increasing focus on user convenience, leading to increased interest in text-based keyword enrollment systems for keyword spotting (KWS). Since the system utilizes text input during the enrollment phase and audio input during actual usage, we call this task audio-text based KWS. To enable this task, both acoustic and text encoders are typically trained using deep metric learning loss functions, such as triplet- and proxy-based losses. This study aims to improve existing methods by leveraging the structural relations within acoustic embeddings and within text embeddings. Unlike previous studies that only compare acoustic and text embeddings on a point-to-point basis, our approach focuses on the relational structures within the embedding space by introducing the concept of Relational Proxy Loss (RPL). By incorporating RPL, we demonstrated improved performance on the Wall Street Journal (WSJ) corpus.

【24】 Spectral Codecs: Spectrogram-Based Audio Codecs for High Quality Speech Synthesis
标题: 频谱编解码器:用于高质量语音合成的基于频谱的音频编解码器
作者:Ryan Langman,Ante Jukić,Kunal Dhawan,Nithin Rao Koluguri,Boris Ginsburg
链接:点击下载PDF文件
摘要:从历史上看,机器学习中的大多数语音模型都使用梅尔频谱图作为语音表示。最近,由神经音频编解码器产生的离散音频令牌已经成为语音合成任务(诸如文本到语音(TTS))的流行的替代语音表示。然而,由这样的编解码器产生的数据分布对于一些TTS模型来说太复杂而无法预测,因此需要大的自回归模型来获得合理的质量。典型的音频编解码器压缩和重构时域音频信号。我们提出了一个频谱编码器,压缩梅尔频谱图和重建的时域音频信号。客观的音频质量指标的研究表明,我们的频谱编解码器具有相当的感知质量等效的音频编解码器。此外,使用所提出的频谱编解码器训练的非自回归TTS模型生成的音频质量明显高于使用梅尔频谱图或音频编解码器训练的音频质量。摘要:Historically, most speech models in machine-learning have used the mel-spectrogram as a speech representation. Recently, discrete audio tokens produced by neural audio codecs have become a popular alternate speech representation for speech synthesis tasks such as text-to-speech (TTS). However, the data distribution produced by such codecs is too complex for some TTS models to predict, hence requiring large autoregressive models to get reasonable quality. Typical audio codecs compress and reconstruct the time-domain audio signal. We propose a spectral codec which compresses the mel-spectrogram and reconstructs the time-domain audio signal. A study of objective audio quality metrics suggests that our spectral codec has comparable perceptual quality to equivalent audio codecs. Furthermore, non-autoregressive TTS models trained with the proposed spectral codec generate audio with significantly higher quality than when trained with mel-spectrograms or audio codecs.

【25】 Signal processing algorithm effective for sound quality of hearing loss simulators
标题: 对听力损失模拟器音质有效的信号处理算法
作者:Toshio Irino,Shintaro Doan,Minami Ishikawa
备注:This paper has been accepted for publication in Interspeech 2024
链接:点击下载PDF文件
摘要:听力损失(HL)模拟器,它允许正常听力(NH)的听众体验HL,已被用于语音清晰度实验,但由于可感知的失真,而不是在音质实验。如果它们产生较少的失真,它们可能有助于NH听众评估助听器等的声音质量。我们进行了感知音质实验,比较剑桥版本的HL模拟器(CamHLS)和和歌山版本的HL模拟器(WHIS),其中有两个算法的滤波器组分析合成(FBAS)和直接时变滤波器(DTVF)。实验结果表明,即使在非线性过程中,DTVF的WHIS也比CamHLS和FBAS的WHIS产生更少的语音失真。这一优势主要是由于DTVF算法的使用,该算法可以应用于具有滤波器组分析的各种信号合成应用。摘要:Hearing loss (HL) simulators, which allow normal hearing (NH) listeners to experience HL, have been used in speech intelligibility experiments, but not in sound quality experiments due to perceptible distortion. If they produced less distortion, they might be useful for NH listeners to evaluate the sound quality of, for example, hearing aids. We conducted perceptual sound quality experiments to compare the Cambridge version of HL simulator (CamHLS) and the Wakayama version of the HL simulator (WHIS), which has the two algorithms of filterbank analysis synthesis (FBAS) and direct time-varying filter (DTVF). The experimental results showed that WHIS with DTVF produces less perceptible distortion in speech sounds than CamHLS and WHIS with FBAS, even when the nonlinear process is working. This advantage is mainly due to the use of the DTVF algorithm, which could be applied to various signal synthesis applications with filterbank analysis.

【26】 A model of early word acquisition based on realistic-scale audiovisual naming events
标题: 基于现实规模视听命名事件的早期词汇习得模型
作者:Khazar Khorrami,Okko Räsänen
备注:22 pages, 4 figures, journal article, submitted for review
链接:点击下载PDF文件
摘要:婴儿逐渐学会将连续语音解析成单词,并将名称与物体联系起来,但早期单词感知技能发展背后的机制仍然未知。我们研究了在何种程度上早期的话可以通过统计学习,从视听感官输入的学习。我们在一个现实的环境中模拟了12个月大的婴儿的单词学习,使用一个模型,该模型仅从未注释的原始语音和像素级视觉输入的统计数据中学习。至关重要的是,物体命名事件的数量是经过精心设计的,以匹配可比年龄的婴儿。结果表明,该模型有效地学习识别单词并将其与相应的视觉对象相关联,其词汇增长率与婴儿中观察到的词汇增长率相当。研究结果支持的可行性,一般的统计学习的早期单词的感知,展示了学习如何运作,而无需假设任何先前的语言能力。摘要:Infants gradually learn to parse continuous speech into words and connect names with objects, yet the mechanisms behind development of early word perception skills remain unknown. We studied the extent to which early words can be acquired through statistical learning from regularities in audiovisual sensory input. We simulated word learning in infants up to 12 months of age in a realistic setting, using a model that solely learns from statistical regularities in unannotated raw speech and pixel-level visual input. Crucially, the quantity of object naming events was carefully designed to match that accessible to infants of comparable ages. Results show that the model effectively learns to recognize words and associate them with corresponding visual objects, with a vocabulary growth rate comparable to that observed in infants. The findings support the viability of general statistical learning for early word perception, demonstrating how learning can operate without assuming any prior linguistic capabilities.

【27】 XANE: eXplainable Acoustic Neural Embeddings
标题: XANE:可扩展的声学神经嵌入
作者:Sri Harsha Dumpala,Dushyant Sharma,Chandramouli Shama Sastri,Stanislav Kruchinin,James Fosburgh,Patrick A. Naylor
链接:点击下载PDF文件
摘要:我们提出了一种新的方法来提取神经嵌入模型的语音信号的背景声学。所提取的嵌入用于以非侵入性方式估计与信号的背景声学特性相关的特定参数,这允许嵌入根据这些参数来解释。我们通过对看不见的测试数据进行聚类实验来说明这些嵌入的价值,并表明所提出的嵌入在三个不同的任务中实现了95.2%的平均F1得分,显著优于基于WavLM的信号嵌入。我们还表明,所提出的方法可以解释的嵌入估计14个声学参数表征的背景声学,包括混响和噪声水平,重叠语音检测,编解码器类型检测和噪声类型检测具有高精度和实时因素17倍低于外部基线方法。摘要:We present a novel method for extracting neural embeddings that model the background acoustics of a speech signal. The extracted embeddings are used to estimate specific parameters related to the background acoustic properties of the signal in a non-intrusive manner, which allows the embeddings to be explainable in terms of those parameters. We illustrate the value of these embeddings by performing clustering experiments on unseen test data and show that the proposed embeddings achieve a mean F1 score of 95.2 % for three different tasks, outperforming significantly the WavLM based signal embeddings. We also show that the proposed method can explain the embeddings by estimating 14 acoustic parameters characterizing the background acoustics, including reverberation and noise levels, overlapped speech detection, CODEC type detection and noise type detection with high accuracy and a real-time factor 17 times lower than an external baseline method.

【28】 Multimodal Contextualized Semantic Parsing from Speech
标题: 语音的多模式上下文化语义解析
作者:Jordan Voas,Raymond Mooney,David Harwath
备注:10 Pages, 3 figures, ACL 2024 Main
链接:点击下载PDF文件
摘要:我们介绍了上下文环境中的语义解析(SPICE),旨在提高人工代理的上下文意识的任务,通过整合多模态输入与先前的上下文。SPICE超越了传统的语义解析,提供了一个结构化的,可解释的框架,用于动态更新代理的知识与新信息,反映了人类沟通的复杂性。我们开发的VG-SPICE数据集,精心制作的挑战代理与视觉场景图建设从口语对话交流,突出语音和视觉数据集成。我们还提出了视听对话场景解析器(AViD-SP)开发的VG-SPICE上使用。这些创新旨在改善多模式信息处理和整合。VG-SPICE数据集和AViD-SP模型都是公开的。摘要:We introduce Semantic Parsing in Contextual Environments (SPICE), a task designed to enhance artificial agents' contextual awareness by integrating multimodal inputs with prior contexts. SPICE goes beyond traditional semantic parsing by offering a structured, interpretable framework for dynamically updating an agent's knowledge with new information, mirroring the complexity of human communication. We develop the VG-SPICE dataset, crafted to challenge agents with visual scene graph construction from spoken conversational exchanges, highlighting speech and visual data integration. We also present the Audio-Vision Dialogue Scene Parser (AViD-SP) developed for use on VG-SPICE. These innovations aim to improve multimodal information processing and integration. Both the VG-SPICE dataset and the AViD-SP model are publicly available.

【29】 Controlling Emotion in Text-to-Speech with Natural Language Prompts
标题: 使用自然语言脚本控制文本到语音中的情感
作者:Thomas Bott,Florian Lux,Ngoc Thang Vu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,由于对自然语言的直观使用,提示已迅速成为引导生成式机器学习模型输出的标准方式之一。在这项工作中,我们提出了一个以嵌入为条件的系统,该嵌入源自作为提示的情感丰富的文本。因此,扬声器和提示嵌入的联合表示被集成在基于变换器的架构内的多个点处。我们的方法在合并的情感语音和文本数据集上进行训练,并在每次训练迭代中改变提示,以提高模型的泛化能力。客观和主观的评价结果表明,有条件的合成系统的能力,准确地转移到语音提示中存在的情绪。同时,说话人身份的精确易处理性以及整体高语音质量和可懂度得以保持。摘要:In recent years, prompting has quickly become one of the standard ways of steering the outputs of generative machine learning models, due to its intuitive use of natural language. In this work, we propose a system conditioned on embeddings derived from an emotionally rich text that serves as prompt. Thereby, a joint representation of speaker and prompt embeddings is integrated at several points within a transformer-based architecture. Our approach is trained on merged emotional speech and text datasets and varies prompts in each training iteration to increase the generalization capabilities of the model. Objective and subjective evaluation results demonstrate the ability of the conditioned synthesis system to accurately transfer the emotions present in a prompt to speech. At the same time, precise tractability of speaker identities as well as overall high speech quality and intelligibility are maintained.

【30】 Meta Learning Text-to-Speech Synthesis in over 7000 Languages
标题: Meta学习7000多种语言的文本到语音合成
作者:Florian Lux,Sarina Meyer,Lyonel Behringer,Frank Zalkow,Phat Do,Matt Coler,Emanuël A. P. Habets,Ngoc Thang Vu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:在这项工作中,我们承担了一个具有挑战性的任务,建立一个单一的文本到语音合成系统,能够生成语音超过7000种语言,其中许多缺乏足够的数据,传统的TTS开发。通过利用大规模多语言预训练和Meta学习的新集成来近似语言表示,我们的方法可以在没有任何可用数据的情况下实现语言的zero-shot语音合成。我们通过客观的措施和人类的评价,在不同的语言景观验证我们的系统的性能。通过公开发布我们的代码和模型,我们的目标是赋予社区有限的语言资源,并促进语音技术领域的进一步创新。摘要:In this work, we take on the challenging task of building a single text-to-speech synthesis system that is capable of generating speech in over 7000 languages, many of which lack sufficient data for traditional TTS development. By leveraging a novel integration of massively multilingual pretraining and meta learning to approximate language representations, our approach enables zero-shot speech synthesis in languages without any available data. We validate our system's performance through objective measures and human evaluation across a diverse linguistic landscape. By releasing our code and models publicly, we aim to empower communities with limited linguistic resources and foster further innovation in the field of speech technology.

【31】 MOSA: Music Motion with Semantic Annotation Dataset for Cross-Modal Music Processing
标题: MOSA:用于跨模式音乐处理的具有语义注释数据集的音乐运动
作者:Yu-Fen Huang,Nikki Moran,Simon Coleman,Jon Kelly,Shun-Hwa Wei,Po-Yin Chen,Yun-Hsin Huang,Tsung-Ping Chen,Yu-Chia Kuo,Yu-Chi Wei,Chih-Hsuan Li,Da-Yu Huang,Hsuan-Kai Kao,Ting-Wei Lin,Li Su
备注:IEEEACM Transactions on Audio, Speech, and Language Processing, 2024. 14 pages, 7 figures. Dataset is available on: this https URL and this https URL
链接:点击下载PDF文件
摘要:在跨模态音乐处理中,视觉,听觉和语义内容之间的翻译开辟了新的可能性和挑战。这种变革性计划的构建取决于具有全面数据基础设施的基准语料库。特别是,组装一个大规模的跨模态数据集提出了重大挑战。在本文中,我们提出了MOSA(音乐运动与语义注释)数据集,其中包含高质量的3-D运动捕捉数据,对齐的音频记录,并注意到由23个专业音乐家的742个专业音乐表演的音高,节拍,乐句,动态,清晰度和和声的语义注释,包括超过30小时和570 K个音符的数据。据我们所知,这是迄今为止最大的具有音符级注释的跨模态音乐数据集。为了演示MOSA数据集的使用,我们提出了几个创新的跨模态音乐信息检索(MIR)和音乐内容生成任务,包括从音频、视频和运动数据中检测节拍、强拍、短语和表达内容,以及从给定的音乐音频中生成音乐家的身体动作。该数据集和代码与本出版物一起提供(https: github.com yufenhuang MOSA-Music-mOtion-and-Semantic-Annotation-dataset)。摘要:In cross-modal music processing, translation between visual, auditory, and semantic content opens up new possibilities as well as challenges. The construction of such a transformative scheme depends upon a benchmark corpus with a comprehensive data infrastructure. In particular, the assembly of a large-scale cross-modal dataset presents major challenges. In this paper, we present the MOSA (Music mOtion with Semantic Annotation) dataset, which contains high quality 3-D motion capture data, aligned audio recordings, and note-by-note semantic annotations of pitch, beat, phrase, dynamic, articulation, and harmony for 742 professional music performances by 23 professional musicians, comprising more than 30 hours and 570 K notes of data. To our knowledge, this is the largest cross-modal music dataset with note-level annotations to date. To demonstrate the usage of the MOSA dataset, we present several innovative cross-modal music information retrieval (MIR) and musical content generation tasks, including the detection of beats, downbeats, phrase, and expressive contents from audio, video and motion data, and the generation of musicians' body motion from given music audio. The dataset and codes are available alongside this publication (https: github.com yufenhuang MOSA-Music-mOtion-and-Semantic-Annotation-dataset).

【32】 mHuBERT-147: A Compact Multilingual HuBERT Model
标题: mHuBERT-147:紧凑的多语言HuBERT模型
作者:Marcely Zanon Boito,Vivek Iyer,Nikolaos Lagos,Laurent Besacier,Ioan Calapodescu
备注:Extended version of the Interspeech 2024 paper of same name
链接:点击下载PDF文件
摘要:mHuBERT-147是第一个通用的大规模多语言HuBERT语音表示模型,基于9万小时的干净,开放许可证的数据进行训练。为了扩展多迭代HuBERT方法,我们使用基于faiss的聚类,实现了比原始方法快5.2倍的标签分配。我们还应用了一种新的多语言上采样策略,利用语言和数据集的多样性。经过3次训练迭代,并且只有95 M个参数,mHuBERT-147的性能优于在更多数据上训练的大型模型。我们在ML-SUPERB 10分钟 1小时排行榜上分别排名第二和第一,所有LID任务的SOTA得分。在ASR LID任务中,我们的模型始终超过XLS-R(300 M参数; 436 K小时),并与更大的MMS(1B参数; 491 K小时)相比表现出强大的竞争力。我们的研究结果表明,mHuBERT-147是一个有前途的多语言语音处理任务的模型,提供了一个前所未有的高性能和参数效率之间的平衡。摘要:We present mHuBERT-147, the first general-purpose massively multilingual HuBERT speech representation model trained on 90K hours of clean, open-license data. To scale up the multi-iteration HuBERT approach, we use faiss-based clustering, achieving 5.2x faster label assignment over the original method. We also apply a new multilingual batching up-sampling strategy, leveraging both language and dataset diversity. After 3 training iterations and with only 95M parameters, mHuBERT-147 outperforms larger models trained on substantially more data. We rank second and first on the ML-SUPERB 10min 1h leaderboards respectively, with SOTA scores for all LID tasks. Across ASR LID tasks, our model consistently surpasses XLS-R (300M params; 436K hours) and demonstrates strong competitiveness against the much larger MMS (1B params; 491K hours). Our findings suggest that mHuBERT-147 is a promising model for multilingual speech processing tasks, offering an unprecedented balance between high performance and parameter efficiency.

【33】 Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
标题: 使用数据驱动和基于知识的特征从语音预测心脏活动
作者:Gasser Elbanna,Zohreh Mostaani,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:准确预测心脏活动和其他生物信号对于诊断和监测至关重要。鉴于语音是多个生理系统的产物,大量工作研究了心脏活动的声学相关性。最近,与传统的声学方法相比,自监督模型在语音相关的任务中表现出色。然而,数据驱动表示在预测心脏活动方面的鲁棒性仍然未被探索。在这项研究中,我们证明了自我监督的语音模型在预测心脏活动参数方面优于声学特征。我们还强调了个体变异性对模型泛化能力的影响。这些发现强调了数据驱动表示在此类任务中的价值,以及需要更多基于语音的生理数据来减轻与说话者相关的挑战。摘要:Accurately predicting heart activity and other biological signals is crucial for diagnosis and monitoring. Given that speech is an outcome of multiple physiological systems, a significant body of work studied the acoustic correlates of heart activity. Recently, self-supervised models have excelled in speech-related tasks compared to traditional acoustic methods. However, the robustness of data-driven representations in predicting heart activity remained unexplored. In this study, we demonstrate that self-supervised speech models outperform acoustic features in predicting heart activity parameters. We also emphasize the impact of individual variability on model generalizability. These findings underscore the value of data-driven representations in such tasks and the need for more speech-based physiological data to mitigate speaker-related challenges.

【34】 Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
标题: 基于音频的运行步数估计--窗口和神经网络基线
作者:Philipp Wagner,Andreas Triantafyllopoulos,Alexander Gebhard,Björn Schuller
备注:Accepted at EUSIPCO 2024
链接:点击下载PDF文件
摘要:近几十年来,跑步已成为越来越受欢迎的消遣活动,因为它的可访问性,易于实践和预期的健康益处。然而,对于不同经验水平的跑步者来说,跑步相关伤害的风险是很大的。几种常见的受伤形式是由于过度使用--超过推荐的跑步时间和强度。最近,基于音频的跟踪已经成为监测跑步行为和表现的另一种方式,以前的研究主要集中在预测跑步者疲劳。在这项工作中,我们研究了基于音频的步数估计在户外运行,实现了平均绝对误差为1.098,基于窗口的步数差异和皮尔逊相关系数为0.479时,预测的步骤数在5秒的音频窗口。因此,我们的工作展示了基于音频的监测估计重要的生理变量的可行性,并为进一步利用音频传感器更彻底地表征跑步者行为奠定了基础。摘要:In recent decades, running has become an increasingly popular pastime activity due to its accessibility, ease of practice, and anticipated health benefits. However, the risk of running-related injuries is substantial for runners of different experience levels. Several common forms of injuries result from overuse -- extending beyond the recommended running time and intensity. Recently, audio-based tracking has emerged as yet another modality for monitoring running behaviour and performance, with previous studies largely concentrating on predicting runner fatigue. In this work, we investigate audio-based step count estimation during outdoor running, achieving a mean absolute error of 1.098 in window-based step-count differences and a Pearson correlation coefficient of 0.479 when predicting the number of steps in a 5-second window of audio. Our work thus showcases the feasibility of audio-based monitoring for estimating important physiological variables and lays the foundations for further utilising audio sensors for a more thorough characterisation of runner behaviour.

【35】 An automatic analysis of ultrasound vocalisations for the prediction of interaction context in captive Egyptian fruit bats
标题: 超声发声的自动分析用于预测圈养埃及果蝠的互动环境
作者:Andreas Triantafyllopoulos,Alexander Gebhard,Manuel Milling,Simon Rampp,Björn Schuller
备注:Accepted at EUSIPCO 2024
链接:点击下载PDF文件
摘要:计算生物声学的先前工作主要集中在特定栖息地中动物存在的检测上。然而,动物的声音包含的信息比仅仅存在的信息丰富得多;除此之外,它们还包含了这些动物与其物种其他成员的互动。在自然主义环境中研究这些相互作用几乎是不可能的,因为通常缺乏基本事实。相反,圈养动物的使用提供了一种可行的替代途径。然而,大多数以前的作品遵循传统的,基于地理学的方法来分析相互作用。在目前的工作中,我们超越了这个标准框架,试图使用深度神经网络来预测捕获的埃及伊蚊之间相互作用的潜在背景。我们达到了超过30%的未加权平均召回率-超过机会水平的三倍-并显示出与我们的统计分析不同的错误模式。因此,这项工作代表了从声音中自动分析动物状态的重要一步。摘要:Prior work in computational bioacoustics has mostly focused on the detection of animal presence in a particular habitat. However, animal sounds contain much richer information than mere presence; among others, they encapsulate the interactions of those animals with other members of their species. Studying these interactions is almost impossible in a naturalistic setting, as the ground truth is often lacking. The use of animals in captivity instead offers a viable alternative pathway. However, most prior works follow a traditional, statistics-based approach to analysing interactions. In the present work, we go beyond this standard framework by attempting to predict the underlying context in interactions between captive emph{Rousettus Aegyptiacus} using deep neural networks. We reach an unweighted average recall of over 30 % -- more than thrice the chance level -- and show error patterns that differ from our statistical analysis. This work thus represents an important step towards the automatic analysis of states in animals from sound.

【36】 A Parameter-efficient Language Extension Framework for Multilingual ASR
标题: 多语言ASB的参数高效语言扩展框架
作者:Wei Liu,Jingyong Hou,Dong Yang,Muyong Cao,Tan Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:用多语言语音识别模型(MASR)覆盖所有语言是非常困难的。在现有的MASR之上执行语言扩展是一个理想的选择。在这项研究中,MASR持续学习问题被概率分解为语言身份预测(LP)和跨语言适应(XLA)的子问题。在此基础上,我们提出了一个基于体系结构的语言扩展框架,可以从根本上解决灾难性遗忘,被称为PELE。PELE被设计成参数高效的,渐进地合并附加模块以适应新语言。具体而言,不同的参数有效的微调(PEFT)模块和它们的变体被探索作为执行XLA的潜在候选者。实验进行了5个新的语言与广泛的低资源的数据大小。表现最好的PEFT候选者可以在所有语言中获得令人满意的表现,并在五种语言中的三种语言中表现出优于持续联合学习设置的优势。值得注意的是,PEFT方法专注于权重参数或输入特征被揭示为性能有限,与在层之间插入轻量级模块(例如适配器)相比,显示出明显较差的扩展能力。摘要:Covering all languages with a multilingual speech recognition model (MASR) is very difficult. Performing language extension on top of an existing MASR is a desirable choice. In this study, the MASR continual learning problem is probabilistically decomposed into language identity prediction (LP) and cross-lingual adaptation (XLA) sub-problems. Based on this, we propose an architecture-based framework for language extension that can fundamentally solve catastrophic forgetting, debudded as PELE. PELE is designed to be parameter-efficient, incrementally incorporating an add-on module to adapt to a new language. Specifically, different parameter-efficient fine-tuning (PEFT) modules and their variants are explored as potential candidates to perform XLA. Experiments are carried out on 5 new languages with a wide range of low-resourced data sizes. The best-performing PEFT candidate can achieve satisfactory performance across all languages and demonstrates superiority in three of five languages over the continual joint learning setting. Notably, PEFT methods focusing on weight parameters or input features are revealed to be limited in performance, showing significantly inferior extension capabilities compared to inserting a lightweight module in between layers such as an Adapter.

【37】 Unsupervised Improved MVDR Beamforming for Sound Enhancement
标题: 用于声音增强的无监督改进的MVDR束形成
作者:Jacob Kealey,John Hershey,François Grondin
链接:点击下载PDF文件
摘要:神经网络最近已经成为声音分离的主要方法。它们的良好性能依赖于孤立记录的大型数据集。对于语音和音乐,隔离的单通道数据是容易获得的;然而,在多通道情况下,以及大多数其他声音类别中,情况并非如此。多通道方法有可能优于单通道方法,因为它们可以利用空间和光谱特征,但缺乏训练数据仍然是一个挑战。我们提出了无监督改进的最小变差无失真响应(UIMVDR),它使多通道分离能够通过无监督训练和波束成形来利用野外单通道数据。结果表明,UIMVDR的推广以及监督模型相比,提高分离性能,特别是在有限的监督数据的情况下。通过使用在线数据,它还减少了为多渠道方法收集数据所需的工作。摘要:Neural networks have recently become the dominant approach to sound separation. Their good performance relies on large datasets of isolated recordings. For speech and music, isolated single channel data are readily available; however the same does not hold in the multi-channel case, and with most other sound classes. Multi-channel methods have the potential to outperform single channel approaches as they can exploit both spatial and spectral features, but the lack of training data remains a challenge. We propose unsupervised improved minimum variation distortionless response (UIMVDR), which enables multi-channel separation to leverage in-the-wild single-channel data through unsupervised training and beamforming. Results show that UIMVDR generalizes well and improves separation performance compared to supervised models, particularly in cases with limited supervised data. By using data available online, it also reduces the effort required to gather data for multi-channel approaches.

【38】 Zero-Shot Audio Captioning Using Soft and Hard Prompts
标题: 使用软和硬字幕的Zero-Shot音频字幕
作者:Yiming Zhang,Xuenan Xu,Ruoyi Du,Haohe Liu,Yuan Dong,Zheng-Hua Tan,Wenwu Wang,Zhanyu Ma
备注:Submitted to IEEEACM Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
摘要:在传统的音频字幕方法中,模型通常使用包含音频文本对的人工注释数据集以完全监督的方式进行训练,然后在来自相同数据集的测试集上进行评估。这种方法有两个局限性。首先,这些方法通常需要大量数据,并且需要耗时且昂贵的人工注释来获得音频文本对。第二,这些模型在跨域场景中经常遭受性能降级,即,当输入音频来自与训练集不同的域时,然而,这很少受到关注。本文提出了一种基于对比语言-音频预训练(CLAP)模型的音频字幕生成方法。我们提出的方法只需要文本数据进行训练,使模型能够从跨模态语义空间中的文本特征生成文本。在推理阶段,模型通过利用CLAP的音频-文本对齐,从音频特征生成给定音频的描述性文本。我们设计了两种策略来减轻文本和音频嵌入之间的差异:基于混合增强的软提示和基于检索的声学感知硬提示。这些方法的目的是提高我们提出的模型的泛化性能,促进模型生成字幕更强大,更准确。在AudioCaps和Clotho基准测试上的大量实验表明了该方法的有效性,在域内场景中的性能优于其他zero-shot音频字幕方法,在跨域场景中的性能优于其他方法,从而突出了该方法的泛化能力。摘要:In traditional audio captioning methods, a model is usually trained in a fully supervised manner using a human-annotated dataset containing audio-text pairs and then evaluated on the test sets from the same dataset. Such methods have two limitations. First, these methods are often data-hungry and require time-consuming and expensive human annotations to obtain audio-text pairs. Second, these models often suffer from performance degradation in cross-domain scenarios, i.e., when the input audio comes from a different domain than the training set, which, however, has received little attention. We propose an effective audio captioning method based on the contrastive language-audio pre-training (CLAP) model to address these issues. Our proposed method requires only textual data for training, enabling the model to generate text from the textual feature in the cross-modal semantic space.In the inference stage, the model generates the descriptive text for the given audio from the audio feature by leveraging the audio-text alignment from CLAP.We devise two strategies to mitigate the discrepancy between text and audio embeddings: a mixed-augmentation-based soft prompt and a retrieval-based acoustic-aware hard prompt. These approaches are designed to enhance the generalization performance of our proposed model, facilitating the model to generate captions more robustly and accurately. Extensive experiments on AudioCaps and Clotho benchmarks show the effectiveness of our proposed method, which outperforms other zero-shot audio captioning approaches for in-domain scenarios and outperforms the compared methods for cross-domain scenarios, underscoring the generalization ability of our method.

【39】 Quantifying the effect of speech pathology on automatic and human speaker verification
标题: 量化言语病理学对自动和人类说话人验证的影响
作者:Bence Mark Halpern,Thomas Tienkamp,Wen-Chin Huang,Lester Phillip Violeta,Teja Rebernik,Sebastiaan de Visscher,Max Witjes,Martijn Wieling,Defne Abur,Tomoki Toda
备注:5 pages, 2 figures, 2 tables. Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本研究探讨了如何手术干预语音病理(特别是口腔癌手术的结果)影响的自动说话人确认(ASV)系统的性能。使用两个最近收集的荷兰数据集与并行术前和术后的音频从同一扬声器,NKI-OC-VC和SPOKE,我们评估在何种程度上语音病理影响ASV的性能,以及是否客观 主观措施的语音严重程度与性能。最后,我们进行了一个感性的研究,比较ASV和人类听众的判断。我们的研究结果表明,病理性言语会对ASV的表现产生负面影响,并且言语的严重程度与ASV的表现呈负相关。说话人相似性和严重性的感知和客观得分有适度的一致性,但是,我们不能清楚地建立在感知研究中,是否同样的现象也存在于人类的感知。摘要:This study investigates how surgical intervention for speech pathology (specifically, as a result of oral cancer surgery) impacts the performance of an automatic speaker verification (ASV) system. Using two recently collected Dutch datasets with parallel pre and post-surgery audio from the same speaker, NKI-OC-VC and SPOKE, we assess the extent to which speech pathology influences ASV performance, and whether objective subjective measures of speech severity are correlated with the performance. Finally, we carry out a perceptual study to compare judgements of ASV and human listeners. Our findings reveal that pathological speech negatively affects ASV performance, and the severity of the speech is negatively correlated with the performance. There is a moderate agreement in perceptual and objective scores of speaker similarity and severity, however, we could not clearly establish in the perceptual study, whether the same phenomenon also exists in human perception.

【40】 Thunder : Unified Regression-Diffusion Speech Enhancement with a Single Reverse Step using Brownian Bridge
标题: Thunder:使用布朗桥进行单反向步骤的统一回归扩散语音增强
作者:Thanapat Trachu,Chawan Piansaddhayanon,Ekapol Chuangsuwanich
备注:5 pages, 3 figures, 4 tables, This paper will be submitted in the interspeech conference
链接:点击下载PDF文件
摘要:基于扩散的语音增强已经显示出有希望的结果,但可能遭受较慢的推理时间。利用由基于回归的模型生成的增强音频来初始化扩散过程可以用于减少所需的计算步骤。然而,这些方法通常需要回归模型,进一步增加了系统的复杂性。我们提出了雷霆,一个统一的回归扩散模型,利用布朗桥过程,可以让该模型在这两种模式下的行为。通过将扩散时间步长设置为接近1,可以进入回归模式。然而,由于梯度不稳定性,标准的基于分数的扩散建模在此设置中表现不佳。为了缓解这个问题,我们修改了扩散模型来预测干净的语音,而不是分数函数,以更紧凑的模型大小和更少的反向步骤实现有竞争力的性能。摘要:Diffusion-based speech enhancement has shown promising results, but can suffer from a slower inference time. Initializing the diffusion process with the enhanced audio generated by a regression-based model can be used to reduce the computational steps required. However, these approaches often necessitate a regression model, further increasing the system's complexity. We propose Thunder, a unified regression-diffusion model that utilizes the Brownian bridge process which can allow the model to act in both modes. The regression mode can be accessed by setting the diffusion time step closed to 1. However, the standard score-based diffusion modeling does not perform well in this setup due to gradient instability. To mitigate this problem, we modify the diffusion model to predict the clean speech instead of the score function, achieving competitive performance with a more compact model size and fewer reverse steps.

【41】 StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
标题: StreamAtt:通过基于注意力的音频历史选择直接流媒体语音到文本翻译
作者:Sara Papi,Marco Gaido,Matteo Negri,Luisa Bentivogli
备注:Accepted at ACL 2024 main conference
链接:点击下载PDF文件
摘要:流式语音到文本翻译(StreamST)是在增量接收音频流的同时自动翻译语音的任务。与处理预分段语音的同步ST(SimulST)不同,StreamST面临着处理连续和无限音频流的挑战。这需要关于保留先前历史的什么的附加决定,由于延迟和计算约束,完全保留先前历史是不切实际的。尽管现实世界对实时ST的需求,但对流式翻译的研究仍然有限,现有的作品仅关注SimulST。为了填补这一空白,我们引入StreamAtt,第一个StreamST政策,并提出StreamLAAL,第一个StreamST延迟度量,旨在与现有的指标SimulST相媲美。在MuST-C v1.0的所有8种语言中进行的广泛实验显示了StreamAtt与原始流基线和相关的最先进的SimulST策略相比的有效性,为StreamST研究提供了第一步。摘要:Streaming speech-to-text translation (StreamST) is the task of automatically translating speech while incrementally receiving an audio stream. Unlike simultaneous ST (SimulST), which deals with pre-segmented speech, StreamST faces the challenges of handling continuous and unbounded audio streams. This requires additional decisions about what to retain of the previous history, which is impractical to keep entirely due to latency and computational constraints. Despite the real-world demand for real-time ST, research on streaming translation remains limited, with existing works solely focusing on SimulST. To fill this gap, we introduce StreamAtt, the first StreamST policy, and propose StreamLAAL, the first StreamST latency metric designed to be comparable with existing metrics for SimulST. Extensive experiments across all 8 languages of MuST-C v1.0 show the effectiveness of StreamAtt compared to a naive streaming baseline and the related state-of-the-art SimulST policy, providing a first step in StreamST research.

【42】 RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
标题: RawBMamba:用于音频深度伪造检测的端到端双向状态空间模型
作者:Yujie Chen,Jiangyan Yi,Jun Xue,Chenglong Wang,Xiaohui Zhang,Shunbo Dong,Siding Zeng,Jianhua Tao,Lv Zhao,Cunhang Fan
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:用于区分真实音频和假音频的假伪像可以存在于短距离段和长距离段两者中。因此,结合局部和全局特征信息可以有效地区分真假音频。本文提出了一种名为RawBMamba的端到端双向状态空间模型,用于捕获用于音频深度伪造检测的短距离和长距离区分信息。具体来说,我们使用sinc Layer和多个卷积层来捕获短程特征,然后设计一个双向Mamba来解决Mamba的单向建模问题,并进一步捕获长程特征信息。此外,我们开发了一个双向融合模块来整合嵌入,增强音频上下文表示,并结合短距离和长距离的信息。实验结果表明,在ASVspoof 2021 LA数据集上,RawBMamba算法比Rawformer算法性能提高了34.1%,在其他数据集上也表现出了较好的性能.摘要:Fake artefacts for discriminating between bonafide and fake audio can exist in both short- and long-range segments. Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio. This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short- and long-range discriminative information for audio deepfake detection. Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information. Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining short- and long-range information. The results show that our proposed RawBMamba achieves a 34.1 % improvement over Rawformer on ASVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.

【43】 Contrastive Learning from Synthetic Audio Doppelgangers
标题: 合成音频分身的对比学习
作者:Manuel Cherep,Nikhil Singh
备注:17 pages, 6 figures
链接:点击下载PDF文件
摘要:学习鲁棒的音频表示目前需要大量的真实世界录音数据集。通过对这些记录进行人工转换,模型可以学习识别相似性,尽管通过对比学习等技术存在细微变化。然而,这些转换只是真实世界声音中真实多样性的近似值,这些声音是由复杂的物理过程相互作用产生的,从声带振动到乐器的共振。我们提出了一个解决方案,利用合成音频的数据规模和转换的限制。通过随机扰动声音合成器的参数,我们生成音频doppelg “angers合成的积极对与因果操纵的音色,音高和时间包络的变化。这些变化,难以实现通过现有的音频转换,提供了丰富的对比信息的来源。尽管转移到随机生成的合成数据,我们的方法产生强有力的表示,与标准音频分类基准的真实数据竞争。值得注意的是,我们的方法是轻量级的,不需要数据存储,只有一个超参数,我们广泛分析。我们提供这种方法作为现有的音频对比学习策略的补充,使用合成的声音来减少从业者的数据负担。摘要:Learning robust audio representations currently demands extensive datasets of real-world sound recordings. By applying artificial transformations to these recordings, models can learn to recognize similarities despite subtle variations through techniques like contrastive learning. However, these transformations are only approximations of the true diversity found in real-world sounds, which are generated by complex interactions of physical processes, from vocal cord vibrations to the resonance of musical instruments. We propose a solution to both the data scale and transformation limitations, leveraging synthetic audio. By randomly perturbing the parameters of a sound synthesizer, we generate audio doppelg "angers-synthetic positive pairs with causally manipulated variations in timbre, pitch, and temporal envelopes. These variations, difficult to achieve through transformations of existing audio, provide a rich source of contrastive information. Despite the shift to randomly generated synthetic data, our method produces strong representations, competitive with real data on standard audio classification benchmarks. Notably, our approach is lightweight, requires no data storage, and has only a single hyperparameter, which we extensively analyze. We offer this method as a complement to existing strategies for contrastive learning in audio, using synthesized sounds to reduce the data burden on practitioners.

【44】 Zero-Shot End-To-End Spoken Question Answering In Medical Domain
标题: 医疗领域的Zero-Shot端到端口语问答
作者:Yanis Labrak,Adel Moumen,Richard Dufour,Mickael Rouvier
Journal-ref:InterSpeech 2024
链接:点击下载PDF文件
摘要:在快速发展的口语问答(SQA)领域,大型语言模型(LLM)的集成已经成为一种变革性的发展。传统的方法通常需要使用单独的模型来进行问题音频转录和答案选择,从而导致显著的资源利用和错误累积。为了应对这些挑战,我们探索了医疗领域SQA的端到端(E2E)方法的有效性。我们的研究介绍了一种新的zero-shot SQA方法,相比传统的级联系统。通过对8个医疗任务和48小时合成音频的新开放基准进行的全面评估,我们证明了我们的方法所需的资源比1.3B参数LLM与1.55B参数ASR模型的组合少14.7倍,同时平均精度提高了0.5%。这些发现强调了在资源受限的情况下,E2E方法用于SQA的潜力。摘要:In the rapidly evolving landscape of spoken question-answering (SQA), the integration of large language models (LLMs) has emerged as a transformative development. Conventional approaches often entail the use of separate models for question audio transcription and answer selection, resulting in significant resource utilization and error accumulation. To tackle these challenges, we explore the effectiveness of end-to-end (E2E) methodologies for SQA in the medical domain. Our study introduces a novel zero-shot SQA approach, compared to traditional cascade systems. Through a comprehensive evaluation conducted on a new open benchmark of 8 medical tasks and 48 hours of synthetic audio, we demonstrate that our approach requires up to 14.7 times fewer resources than a combined 1.3B parameters LLM with a 1.55B parameters ASR model while improving average accuracy by 0.5 %. These findings underscore the potential of E2E methodologies for SQA in resource-constrained contexts.

【45】 Source -Free Domain Adaptation for Speaker Verification in Data-Scarce Languages and Noisy Channels
标题: 源-免费域自适应,用于数据稀缺语言和有噪通道中的说话人验证
作者:Shlomo Salo Elia,Aviad Malachi,Vered Aharonson,Gadi Pinkas
链接:点击下载PDF文件
摘要:领域自适应常常受到极小的目标数据集和不可访问的源数据的阻碍。这些情况在语音验证中普遍存在,其中隐私政策和 或具有稀缺语音资源的语言限制了足够数据的可用性。本文探讨了数据稀缺语言中说话人确认的无源域自适应技术。研究了源语和目标语之间的语言和通道不匹配。在不同大小的标记目标数据中评估和比较了微调方法。针对未标记目标数据集,研究了一种新的迭代聚类学习算法。摘要:Domain adaptation is often hampered by exceedingly small target datasets and inaccessible source data. These conditions are prevalent in speech verification, where privacy policies and or languages with scarce speech resources limit the availability of sufficient data. This paper explored techniques of sourcefree domain adaptation unto a limited target speech dataset for speaker verificationin data-scarce languages. Both language and channel mis-match between source and target were investigated. Fine-tuning methods were evaluated and compared across different sizes of labeled target data. A novel iterative cluster-learn algorithm was studied for unlabeled target datasets.

【46】 Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
标题: 预算真的会提示吗?探索Whisper的快速理解能力
作者:Chih-Kai Yang,Kuan-Po Huang,Hung-yi Lee
备注:In progress
链接:点击下载PDF文件
摘要:本研究探讨Whisper,一个高性能的语音识别模型,和提示信息之间的相互作用。我们的研究结果出乎意料地表明,耳语可能无法完全掌握预期的文本提示。此外,我们发现,即使在文本提示中更严格地遵守主题信息,也不能保证性能提高。值得注意的是,英语提示在两种语言的数据集上通常优于普通话提示,这可能是由于这些语言的训练数据分布的差异。相反,我们发现,耳语表现出意识的误导性信息的语言标记,有效地忽略不正确的语言标记,并专注于正确的。总之,这项工作提出了关于耳语的快速理解能力的问题,并鼓励进一步的研究。摘要:This research explores the interaction between Whisper, a high-performing speech recognition model, and information in prompts. Our results unexpectedly show that Whisper may not fully grasp textual prompts as anticipated. Additionally, we find that performance improvement is not guaranteed even with stronger adherence to the topic information in textual prompts. It is also noted that English prompts generally outperform Mandarin ones on datasets of both languages, likely due to differences in training data distributions for these languages. Conversely, we discover that Whisper exhibits awareness of misleading information in language tokens by effectively ignoring incorrect language tokens and focusing on the correct ones. In summary, this work raises questions about Whisper's prompt understanding capability and encourages further studies.

【47】 Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
标题: 优化多口吃语音分类:利用Whisper的编码器在自动评估中有效简化参数
作者:Huma Ameer,Seemab Latif,Rabia Latif
链接:点击下载PDF文件
摘要:口吃语音的自动分类对及时评估语音语言病理学家提供帮助具有重要意义。尽管在该领域取得了显着的进步,但在言语中发生多个不流利的情况需要注意。我们已经采取了一种渐进的方法来填补这一空白,更有效地分类多口吃的语音。这个问题已经通过首先从SEP-28 k音频片段中策划多口吃不流利的数据集来解决。其次,采用Whisper,一个国家的最先进的语音识别模型已利用其编码器和多标签分类的问题。第三,使用6个编码器层Whisper并试验各种层冻结策略,确定了模型的计算效率配置。所提出的配置在外部测试数据集(即Fluency-Bank)上相应地实现了0.88、0.85和0.87的微观、宏观和加权F1分数。此外,通过层冻结策略,我们能够通过微调单个编码器层来实现上述结果,从而将模型的可训练参数从2027万减少到329万。这项研究揭示了最后一个编码器层在识别口吃语音中的不流利性方面的贡献。因此,它导致了一种计算效率高的方法,使模型更适合各种方言和语言。摘要:The automated classification of stuttered speech has significant implications for timely assessments providing assistance to speech language pathologists. Despite notable advancements in the field, the cases in which multiple disfluencies occur in speech require attention. We have taken a progressive approach to fill this gap by classifying multi-stuttered speech more efficiently. The problem has been addressed by firstly curating a dataset of multi-stuttered disfluencies from SEP-28k audio clips. Secondly, employing Whisper, a state-of-the-art speech recognition model has been leveraged by using its encoder and taking the problem as multi-label classification. Thirdly, using a 6 encoder layer Whisper and experimenting with various layer freezing strategies, a computationally efficient configuration of the model was identified. The proposed configuration achieved micro, macro, and weighted F1- scores of 0.88, 0.85, and 0.87, correspondingly on an external test dataset i.e. Fluency-Bank. In addition, through layer freezing strategies, we were able to achieve the aforementioned results by fine-tuning a single encoder layer, consequently, reducing the model's trainable parameters from 20.27 million to 3.29 million. This research study unveils the contribution of the last encoder layer in the identification of disfluencies in stuttered speech. Consequently, it has led to a computationally efficient approach which makes the model more adaptable for various dialects and languages.

【48】 SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
标题: SPA-SRC:用于歌唱声音转换的自我监督音调增强
作者:Bingsong Bai,Fengping Wang,Yingming Gao,Ya Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:与传统方法相比,基于扩散的歌唱声转换(SVC)模型显示出更好的合成质量。然而,在跨域SVC场景中,源语音域和目标语音域之间的音高存在显著差异,这些模型往往会生成声音嘶哑的音频,这对实现高质量的语音输出提出了挑战。因此,在本文中,我们提出了一种自监督音高增强方法歌唱语音转换(SPA-SVC),它可以提高语音质量在SVC任务,而不需要额外的数据或增加模型参数。我们创新性地引入了一个周期性的音高变换训练策略和结构相似性指数(SSIM)损失到我们的SVC模型,有效地提高了它的性能。在公共歌唱数据集M4 Singer上的实验结果表明,我们提出的方法显着提高了模型的性能,在一般的SVC场景,特别是在跨域SVC场景。摘要:Diffusion-based singing voice conversion (SVC) models have shown better synthesis quality compared to traditional methods. However, in cross-domain SVC scenarios, where there is a significant disparity in pitch between the source and target voice domains, the models tend to generate audios with hoarseness, posing challenges in achieving high-quality vocal outputs. Therefore, in this paper, we propose a Self-supervised Pitch Augmentation method for Singing Voice Conversion (SPA-SVC), which can enhance the voice quality in SVC tasks without requiring additional data or increasing model parameters. We innovatively introduce a cycle pitch shifting training strategy and Structural Similarity Index (SSIM) loss into our SVC model, effectively enhancing its performance. Experimental results on the public singing datasets M4Singer indicate that our proposed method significantly improves model performance in both general SVC scenarios and particularly in cross-domain SVC scenarios.

【49】 Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
标题: 利用分层韵律建模实现表达性Zero-Shot语音合成
作者:Yuepeng Jiang,Tao Li,Fengyu Yang,Lei Xie,Meng Meng,Yujun Wang
备注:5 pages, 2 figures, accepted by Interspeech2024
链接:点击下载PDF文件
摘要:近年来zero-shot语音合成的研究在说话人相似度方面取得了很大的进展。然而,目前的努力集中在音色的泛化,而不是韵律建模,这导致有限的自然性和表现力。为了解决这个问题,我们引入了一种新的语音合成模型,在大规模数据集上训练,包括音色和层次韵律建模。由于音色是一个与表现力密切相关的全局属性,我们采用一个全局向量对说话人音色进行建模,同时指导韵律建模。此外,由于韵律包含全局一致性和局部变化,我们引入了扩散模型作为音高预测器,并采用韵律适配器对韵律进行分层建模,进一步提高了合成语音的韵律质量。实验结果表明,我们的模型不仅保持可比的音色质量的基线,但也表现出更好的自然度和表现力。摘要:Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody modeling, which results in limited naturalness and expressiveness. To address this, we introduce a novel speech synthesis model trained on large-scale datasets, including both timbre and hierarchical prosody modeling. As timbre is a global attribute closely linked to expressiveness, we adopt a global vector to model speaker timbre while guiding prosody modeling. Besides, given that prosody contains both global consistency and local variations, we introduce a diffusion model as the pitch predictor and employ a prosody adaptor to model prosody hierarchically, further enhancing the prosody quality of the synthesized speech. Experimental results show that our model not only maintains comparable timbre quality to the baseline but also exhibits better naturalness and expressiveness.

【50】 Heart Sound Segmentation Using Deep Learning Techniques
标题: 使用深度学习技术的人声分割
作者:Manas Madine
链接:点击下载PDF文件
摘要:心脏病仍然是全球死亡的主要原因。听诊,即倾听心音的过程,可以通过使用心音图(PCG)信号的计算机辅助分析来增强。本文提出了一种新的方法,心音分割和分类为S1(LUB)和S2(DUB)的声音。我们采用基于FFT的过滤,动态规划的事件检测,和一个暹罗网络的强大的分类。与现有方法相比,我们的方法在PASCAL心音数据集上表现出优越的性能。摘要:Heart disease remains a leading cause of mortality worldwide. Auscultation, the process of listening to heart sounds, can be enhanced through computer-aided analysis using Phonocardiogram (PCG) signals. This paper presents a novel approach for heart sound segmentation and classification into S1 (LUB) and S2 (DUB) sounds. We employ FFT-based filtering, dynamic programming for event detection, and a Siamese network for robust classification. Our method demonstrates superior performance on the PASCAL heart sound dataset compared to existing approaches.

【51】 Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
作者:Mark Hamilton,Andrew Zisserman,John R. Hershey,William T. Freeman
备注:Computer Vision and Pattern Recognition 2024
链接:点击下载PDF文件
摘要:我们提出了DenseAV,一种新颖的双编码器接地架构,仅通过观看视频来学习高分辨率,语义有意义和视听对齐的功能。我们表明,DenseAV可以发现单词的“意义”和声音的“位置”,而无需明确的本地化监督。此外,它可以自动发现和区分这两种类型的关联,而无需监督。我们表明,DenseAV的定位能力来自一个新的多头特征聚合算子,该算子直接比较密集的图像和音频表示进行对比学习。相比之下,许多其他学习“全球”音频和视频表示的系统不能本地化单词和声音。最后,我们贡献了两个新的数据集,通过语音和声音提示的语义分割来提高AV表示的评估。在这些和其他数据集上,我们显示DenseAV在语音和声音提示的语义分割方面显着优于现有技术。DenseAV在使用不到一半的参数进行跨模态检索方面优于以前的最先进的ImageBind。项目页面: href{https: aka.ms densav}{https: aka.ms densav}摘要:We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching videos. We show that DenseAV can discover the meaning'' of words and the location'' of sounds without explicit localization supervision. Furthermore, it automatically discovers and distinguishes between these two types of associations without supervision. We show that DenseAV's localization abilities arise from a new multi-head feature aggregation operator that directly compares dense image and audio representations for contrastive learning. In contrast, many other systems that learn global'' audio and video representations cannot localize words and sound. Finally, we contribute two new datasets to improve the evaluation of AV representations through speech and sound prompted semantic segmentation. On these and other datasets we show DenseAV dramatically outperforms the prior art on speech and sound prompted semantic segmentation. DenseAV outperforms the previous state-of-the-art, ImageBind, on cross-modal retrieval using fewer than half of the parameters. Project Page: href{https: aka.ms denseav}{https: aka.ms denseav}

【52】 Exploring the Benefits of Tokenization of Discrete Acoustic Units
标题: 探索离散声学单元代币化的好处
作者:Avihu Dekel,Raul Fernandez
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:将基本词汇表的单元合并为更大的可变速率单元的标记化算法已经成为自然语言处理任务中的标准。然而,当词汇表由音素或离散声学单位(DAU)组成时,这个想法大多被忽视了,由于离散语言建模技术的成功,基于音频的表示正在发挥越来越重要的作用。在本文中,我们展示了语音单元和DAU的标记化在三个预测任务上的优势:字素到音素,字素到DAU,以及使用DAU语言建模的无监督语音生成。我们证明了令牌化在所有三个任务中的性能以及训练和推理速度方面都有显着的改进。我们还提供理论见解,为观察到的优异性能提供一些解释。摘要:Tokenization algorithms that merge the units of a base vocabulary into larger, variable-rate units have become standard in natural language processing tasks. This idea, however, has been mostly overlooked when the vocabulary consists of phonemes or Discrete Acoustic Units (DAUs), an audio-based representation that is playing an increasingly important role due to the success of discrete language-modeling techniques. In this paper, we showcase the advantages of tokenization of phonetic units and of DAUs on three prediction tasks: grapheme-to-phoneme, grapheme-to-DAUs, and unsupervised speech generation using DAU language modeling. We demonstrate that tokenization yields significant improvements in terms of performance, as well as training and inference speed, across all three tasks. We also offer theoretical insights to provide some explanation for the superior performance observed.

【53】 Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation
标题: 嗯你说什么?使用心理物理反向相关揭示第一语言和第二语言单词感知中的远端和近端上下文效应
作者:Paige Tuttösí,H. Henny Yeung,Yue Wang,Fenqi Wang,Guillaume Denis,Jean-Julien Aucouturier,Angelica Lim
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:声学语境效应,即周围音高、速率或音色的变化影响声音的感知,在言语感知中有很好的记录,但它们如何与语言背景相互作用仍不清楚。使用反向相关的方法,我们系统地改变了音高和语速在短语周围的不同对元音的第二语言(L2)的英语( i - I )和法语( u - y )的扬声器,从而重建,在数据驱动的方式,韵律配置文件,偏见他们的看法。测试英语和法语的发言者(n=25),我们发现,元音感知实际上是由周围的音高和语音速率的冲突影响:一致的近端效应0.2秒前的目标和远端对比效应高达1秒前,并发现,L1和L2扬声器表现出惊人的相似的韵律配置文件的感知。我们提供了一种新的方法来调查声环境的影响,跨刺激,时间尺度,和声学域。摘要:Acoustic context effects, where surrounding changes in pitch, rate or timbre influence the perception of a sound, are well documented in speech perception, but how they interact with language background remains unclear. Using a reverse-correlation approach, we systematically varied the pitch and speech rate in phrases around different pairs of vowels for second language (L2) speakers of English ( i - I ) and French ( u - y ), thus reconstructing, in a data-driven manner, the prosodic profiles that bias their perception. Testing English and French speakers (n=25), we showed that vowel perception is in fact influenced by conflicting effects from the surrounding pitch and speech rate: a congruent proximal effect 0.2s pre-target and a distal contrastive effect up to 1s before; and found that L1 and L2 speakers exhibited strikingly similar prosodic profiles in perception. We provide a novel method to investigate acoustic context effects across stimuli, timescales, and acoustic domain.

【54】 DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
标题: DAISY:语音表示模型的数据自适应自我监督早期退出
作者:Tzu-Quan Lin,Hung-yi Lee,Hao Tang
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自监督语音模型已被证明对各种任务都很有用,但它们的大尺寸限制了在计算能力和内存较低的设备中的使用。在这项工作中,我们探索早期退出,通过提前退出网络的转发过程来减少延迟的方法。大多数提前退出方法需要为每个任务提供单独的提前退出模型,有些甚至需要对整个预训练模型进行微调。我们引入了数据自适应自监督早期退出(DAISY),这是一种基于自监督损失决定何时退出的方法,无需多轮训练和微调。DAISY在MiniSUPERB基准测试中与HuBERT的性能相当,但推理时间要快得多。我们对DAISY自适应性的分析表明,该模型在干净数据上提前退出(使用较少的层),而在有噪声的数据上延迟退出(使用更多的层),根据每个样本的噪声水平动态调整推理的计算成本。摘要:Self-supervised speech models have shown to be useful for various tasks, but their large size limits the use in devices with low computing power and memory. In this work, we explore early exit, an approach for reducing latency by exiting the forward process of a network early. Most approaches of early exit need a separate early exit model for each task, with some even requiring fine-tuning of the entire pretrained model. We introduce Data Adaptive Self-Supervised Early Exit (DAISY), an approach that decides when to exit based on the self-supervised loss, eliminating the need for multiple round of training and fine-tuning. DAISY matches the performance of HuBERT on the MiniSUPERB benchmark, but with much faster inference times. Our analysis on the adaptivity of DAISY shows that the model exits early (using fewer layers) on clean data while exits late (using more layers) on noisy data, dynamically adjusting the computational cost of inference based on the noise level of each sample.

【55】 VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
标题: WAL-E 2:神经编解码器语言模型是人类对等Zero-Shot文本到语音合成器
作者:Sanyuan Chen,Shujie Liu,Long Zhou,Yanqing Liu,Xu Tan,Jinyu Li,Sheng Zhao,Yao Qian,Furu Wei
链接:点击下载PDF文件
摘要:本文介绍了VALL-E 2,神经编解码器语言模型的最新进展,标志着zero-shot文本到语音合成(TTS)的一个里程碑,首次实现人类平等。在其前身VALL-E的基础上,新的迭代引入了两个显著的增强功能:重复感知采样通过考虑解码历史中的令牌重复来改进原始核采样过程。它不仅稳定了解码,而且还避免了无限循环问题。分组代码建模将编解码器代码组织成组,以有效地缩短序列长度,这不仅提高了推理速度,而且解决了长序列建模的挑战。我们在LibriSpeech和VCTK数据集上的实验表明,VALL-E 2在语音鲁棒性、自然度和说话人相似性方面优于以前的系统。这是第一个在这些基准上实现人类平等的国家。此外,VALL-E 2始终能够合成高质量的语音,即使是由于其复杂性或重复短语而传统上具有挑战性的句子。这项工作的优势可能有助于有价值的努力,例如为失语症患者或肌萎缩侧索硬化症患者生成语音。VALL-E 2的演示将发布到https: aka.ms valle2。摘要:This paper introduces VALL-E 2, the latest advancement in neural codec language models that marks a milestone in zero-shot text-to-speech synthesis (TTS), achieving human parity for the first time. Based on its predecessor, VALL-E, the new iteration introduces two significant enhancements: Repetition Aware Sampling refines the original nucleus sampling process by accounting for token repetition in the decoding history. It not only stabilizes the decoding but also circumvents the infinite loop issue. Grouped Code Modeling organizes codec codes into groups to effectively shorten the sequence length, which not only boosts inference speed but also addresses the challenges of long sequence modeling. Our experiments on the LibriSpeech and VCTK datasets show that VALL-E 2 surpasses previous systems in speech robustness, naturalness, and speaker similarity. It is the first of its kind to reach human parity on these benchmarks. Moreover, VALL-E 2 consistently synthesizes high-quality speech, even for sentences that are traditionally challenging due to their complexity or repetitive phrases. The advantages of this work could contribute to valuable endeavors, such as generating speech for individuals with aphasia or people with amyotrophic lateral sclerosis. Demos of VALL-E 2 will be posted to https: aka.ms valle2.


机器翻译,仅供参考