今日论文合集:cs.SD语音33篇,eess.AS音频处理41篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Hearing Anything Anywhere
标题: 任何地方都能听到任何声音
作者:Mason Wang,Ryosuke Sawata,Samuel Clarke,Ruohan Gao,Shangzhe Wu,Jiajun Wu
备注:CVPR 2024. The first two authors contributed equally. Project page: this https URL
链接:点击下载PDF文件
摘要:近年来,3D计算机视觉和计算机图形学取得了巨大的进步,新兴的工具可以为许多混合现实(XR)应用虚拟化真实世界的3D环境。然而,除了沉浸式视觉体验,沉浸式听觉体验对我们对环境的整体感知同样重要。在本文中,我们的目标是重建的空间声学特性的任意环境给定的只有一个稀疏的(大约12)房间脉冲响应(RIR)记录和平面重建的场景,一个设置,是很容易实现的普通用户。为此,我们介绍了DiffRIR,一个可区分的RIR渲染框架与可解释的参数模型的显着的声学特征的场景,包括声源的方向性和表面反射率。这使我们能够通过空间与任何源音频合成新颖的听觉体验。为了评估我们的方法,我们收集了四个不同的真实环境中的RIR录音和音乐数据集。我们表明,我们的模型在渲染单声道和双耳RIR和音乐方面优于最先进的基线,并学习了物理上可解释的参数,这些参数表征了场景中声源和表面的声学特性。摘要:Recent years have seen immense progress in 3D computer vision and computer graphics, with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However, alongside immersive visual experiences, immersive auditory experiences are equally vital to our holistic perception of an environment. In this paper, we aim to reconstruct the spatial acoustic characteristics of an arbitrary environment given only a sparse set of (roughly 12) room impulse response (RIR) recordings and a planar reconstruction of the scene, a setup that is easily achievable by ordinary users. To this end, we introduce DiffRIR, a differentiable RIR rendering framework with interpretable parametric models of salient acoustic features of the scene, including sound source directivity and surface reflectivity. This allows us to synthesize novel auditory experiences through the space with any source audio. To evaluate our method, we collect a dataset of RIR recordings and music in four diverse, real environments. We show that our model outperforms state-ofthe-art baselines on rendering monaural and binaural RIRs and music at unseen locations, and learns physically interpretable parameters characterizing acoustic properties of the sound source and surfaces in the scene.

【2】 RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
标题: RaD-Net 2:因果两阶段修复和去噪语音增强网络,具有知识提炼和复杂的轴向自我注意力
作者:Mingshuai Liu,Zhuangqi Chen,Xiaopeng Yan,Yuanjun Lv,Xianjun Xia,Chuanzeng Huang,Yijian Xiao,Lei Xie
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在实时语音通信系统中,语音信号经常会受到多种失真的影响。最近,在ICASSP 2024语音信号改善(SSI)挑战赛中提出了一种两阶段修复和去噪网络(RaD-Net),具有出色的语音质量改善。然而,未能利用未来的信息和限制卷积层的感受野限制了系统的性能。为了缓解这些问题,我们将RaD-Net扩展到其升级版本RaD-Net 2。具体而言,在第一阶段引入了基于因果关系的知识蒸馏,以因果关系的方式使用未来信息。我们使用非因果修复网络作为教师,以提高因果修复网络的性能。此外,在第二阶段,复轴自注意力应用于去噪网络的复特征编码器 解码器。在ICASSP 2024 SSI Challenge盲测试集上的实验结果表明,与RaD-Net相比,RaD-Net 2带来了0.10 OVRL DNSMOS的改进。摘要:In real-time speech communication systems, speech signals are often degraded by multiple distortions. Recently, a two-stage Repair-and-Denoising network (RaD-Net) was proposed with superior speech quality improvement in the ICASSP 2024 Speech Signal Improvement (SSI) Challenge. However, failure to use future information and constraint receptive field of convolution layers limit the system's performance. To mitigate these problems, we extend RaD-Net to its upgraded version, RaD-Net 2. Specifically, a causality-based knowledge distillation is introduced in the first stage to use future information in a causal way. We use the non-causal repairing network as the teacher to improve the performance of the causal repairing network. In addition, in the second stage, complex axial self-attention is applied in the denoising network's complex feature encoder decoder. Experimental results on the ICASSP 2024 SSI Challenge blind test set show that RaD-Net 2 brings 0.10 OVRL DNSMOS improvement compared to RaD-Net.

【3】 A pilot protocol and cohort for the investigation of non-pathological variability in speech
标题: 研究言语非病理变异性的试点方案和队列
作者:Nicholas Cummins,Lauren L. White,Zahia Rahman,Catriona Lucas,Tian Pan,Ewan Carr,Faith Matcham,Johnny Downs,Richard J. Dobson,Judith Dineley
备注:29 pages. Pre peer review
链接:点击下载PDF文件
摘要:背景基于语音的生物标志物有潜力作为一种手段,定期,客观评估症状的严重程度,远程和临床结合先进的分析模型。然而,言语的复杂性和与健康相关的微妙变化意味着研究结果高度依赖于方法和队列选择。这些往往没有充分报道的研究调查基于语音的健康评估目的开发和应用一个示范协议,以生成一个试点数据集的健康语音与详细的元数据的评估因素的语音记录分析管道,包括设备的选择,语音诱导任务和非病理变异。方法在专题文献综述的基础上,我们制定了我们的收集协议和选择的范例语音特征。我们的协议包括三种不同的语音类型的启发。为了专注于远程应用,我们还选择使用三种不同的麦克风类型来收集语音。我们开发了一个管道来提取一组14个范例语音特征。结果28名受试者每天采集3次语音,8-11周后重复采集; 25名健康受试者每周采集3次语音。收集的参与者特征包括参与者的性别、年龄、母语状态和语音使用习惯。一个初步的14个语音特征,涵盖时间,韵律,语音质量,清晰度和频谱矩特征提取,提供了一个资源的规范值。结论言语录音的收集、加工和分析涉及多种方法学因素。迫切需要一致的报告和更大程度的研究方案协调,以帮助将语音处理转化为临床研究和实践。摘要:Background Speech-based biomarkers have potential as a means for regular, objective assessment of symptom severity, remotely and in-clinic in combination with advanced analytical models. However, the complex nature of speech and the often subtle changes associated with health mean that findings are highly dependent on methodological and cohort choices. These are often not reported adequately in studies investigating speech-based health assessment Objective To develop and apply an exemplar protocol to generate a pilot dataset of healthy speech with detailed metadata for the assessment of factors in the speech recording-analysis pipeline, including device choice, speech elicitation task and non-pathological variability. Methods We developed our collection protocol and choice of exemplar speech features based on a thematic literature review. Our protocol includes the elicitation of three different speech types. With a focus towards remote applications, we also choose to collect speech with three different microphone types. We developed a pipeline to extract a set of 14 exemplar speech features. Results We collected speech from 28 individuals three times in one day, repeated at the same times 8-11 weeks later, and from 25 healthy individuals three times in one week. Participant characteristics collected included sex, age, native language status and voice use habits of the participant. A preliminary set of 14 speech features covering timing, prosody, voice quality, articulation and spectral moment characteristics were extracted that provide a resource of normative values. Conclusions There are multiple methodological factors involved in the collection, processing and analysis of speech recordings. Consistent reporting and greater harmonisation of study protocols are urgently required to aid the translation of speech processing into clinical research and practice.

【4】 Multimodal Belief Prediction
标题: 多峰信念预测
作者:John Murzaku,Adil Soubki,Owen Rambow
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
摘要:识别说话者对信念的承诺程度是一项艰巨的任务;人类不仅在上下文中解释单词的含义,而且还理解来自语调和音频信号的其他方面的线索。NLP社区中的许多论文和语料库都使用纯文本方法来处理信念预测任务。我们是第一个框架和目前的多模态信念预测任务的结果。我们使用CB韵律语料库(CBP),包含对齐的文本和音频与扬声器信念注释。我们首先使用声学韵律特征和传统的机器学习方法报告基线和重要特征。然后,我们分别在BERT和Whisper上呈现用于CBP语料库微调的文本和音频基线。最后,我们提出了我们的多模态架构,微调BERT和耳语,并使用多种融合方法,单独改善这两种方式。摘要:Recognizing a speaker's level of commitment to a belief is a difficult task; humans do not only interpret the meaning of the words in context, but also understand cues from intonation and other aspects of the audio signal. Many papers and corpora in the NLP community have approached the belief prediction task using text-only approaches. We are the first to frame and present results on the multimodal belief prediction task. We use the CB-Prosody corpus (CBP), containing aligned text and audio with speaker belief annotations. We first report baselines and significant features using acoustic-prosodic features and traditional machine learning methods. We then present text and audio baselines for the CBP corpus fine-tuning on BERT and Whisper respectively. Finally, we present our multimodal architecture which fine-tunes on BERT and Whisper and uses multiple fusion methods, improving on both modalities alone.

【5】 Graph-based multi-Feature fusion method for speech emotion recognition
标题: 基于图的语音情感识别多特征融合方法
作者:Xueyu Liu,Jie Lin,Wang Chao
备注:25 pages,4 figures
链接:点击下载PDF文件
摘要:不同的语音特征可以提供反映人类情感状态的互补线索,因此探索合适的方法进行多语音特征融合对于跨语料语音情感识别至关重要。虽然大多数以前的方法只提取一个单一的语音特征的情感识别,现有的融合方法,如级联,并行连接,拼接忽略了异构模式之间的相互作用的功能和功能,导致现有系统的性能。在本文中,我们提出了一种新的基于图的融合方法来明确建模每对语音特征之间的关系。具体来说,我们提出了一种多维边缘特征学习策略,称为基于图的多特征融合方法用于语音情感识别。它将每个语音特征表示为一个节点,并学习多维边缘特征,以显式描述情感识别背景下每个特征-特征对之间的关系。这样,学习到的多维边缘特征就可以对来自顶点和边缘维度的语音特征级信息进行编码。我们的方法由三个模块组成:音频特征生成(AFG)模块、音频特征多维边缘特征(AMEF)模块和语音情感识别(SER)模块。所提出的方法在SEWA数据集上取得了令人满意的结果。此外,与AVEC 2019研讨会和挑战赛中的基线相比,该方法表现出更高的性能。我们使用来自两种文化的数据作为我们的训练和验证集:在SEWA数据集上包含德语和匈牙利语的两种文化,德语的CCC分数在唤醒方面提高了17.28%,在喜欢方面提高了7.93%。我们的方法的结果表明,替代融合技术,包括那些采用一维基于边缘的特征融合方法的13%的改善。摘要:Exploring proper way to conduct multi-speech feature fusion for cross-corpus speech emotion recognition is crucial as different speech features could provide complementary cues reflecting human emotion status. While most previous approaches only extract a single speech feature for emotion recognition, existing fusion methods such as concatenation, parallel connection, and splicing ignore heterogeneous patterns in the interaction between features and features, resulting in performance of existing systems. In this paper, we propose a novel graph-based fusion method to explicitly model the relationships between every pair of speech features. Specifically, we propose a multi-dimensional edge features learning strategy called Graph-based multi-Feature fusion method for speech emotion recognition. It represents each speech feature as a node and learns multi-dimensional edge features to explicitly describe the relationship between each feature-feature pair in the context of emotion recognition. This way, the learned multi-dimensional edge features encode speech feature-level information from both the vertex and edge dimensions. Our Approach consists of three modules: an Audio Feature Generation(AFG)module, an Audio-Feature Multi-dimensional Edge Feature(AMEF) module and a Speech Emotion Recognition (SER) module. The proposed methodology yielded satisfactory outcomes on the SEWA dataset. Furthermore, the method demonstrated enhanced performance compared to the baseline in the AVEC 2019 Workshop and Challenge. We used data from two cultures as our training and validation sets: two cultures containing German and Hungarian on the SEWA dataset, the CCC scores for German are improved by 17.28% for arousal and 7.93% for liking. The outcomes of our methodology demonstrate a 13% improvement over alternative fusion techniques, including those employing one dimensional edge-based feature fusion approach.

【6】 A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
标题: 说话人识别中说话人增强的综合研究
作者:Zhenyu Zhou,Shibiao Xu,Shi Yin,Lantian Li,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:数据增强(DA)在深度说话人识别的成功中发挥了关键作用。目前的DA技术主要集中在说话人保留增强,这不会改变语音的说话人特征,也不会创建新的说话人。最近的研究揭示了说话人增强的潜力,它可以生成新的说话人来丰富训练数据集。在这项研究中,我们深入到两个扬声器增强方法:速度扰动(SP)和声道长度扰动(VTLP)。尽管这两种方法的经验利用,其有效性缺乏全面的调查。我们使用两个公共数据集VoxCeleb和CN-Celeb进行的研究表明,SP和VTLP都擅长生成新的说话人,从而显著提高了说话人识别的性能。此外,他们表现出不同的属性的扰动因素和数据复杂性的敏感性,暗示在其融合的潜在好处。我们的研究强调了扬声器增强的巨大潜力,强调了深入探索和分析的重要性。摘要:Data augmentation (DA) has played a pivotal role in the success of deep speaker recognition. Current DA techniques primarily focus on speaker-preserving augmentation, which does not change the speaker trait of the speech and does not create new speakers. Recent research has shed light on the potential of speaker augmentation, which generates new speakers to enrich the training dataset. In this study, we delve into two speaker augmentation approaches: speed perturbation (SP) and vocal tract length perturbation (VTLP). Despite the empirical utilization of both methods, a comprehensive investigation into their efficacy is lacking. Our study, conducted using two public datasets, VoxCeleb and CN-Celeb, revealed that both SP and VTLP are proficient at generating new speakers, leading to significant performance improvements in speaker recognition. Furthermore, they exhibit distinct properties in sensitivity to perturbation factors and data complexity, hinting at the potential benefits of their fusion. Our research underscores the substantial potential of speaker augmentation, highlighting the importance of in-depth exploration and analysis.

【7】 CTC-based Non-autoregressive Textless Speech-to-Speech Translation
标题: 基于ATC的非自回归无文本语音翻译
作者:Qingkai Fang,Zhengrui Ma,Yan Zhou,Min Zhang,Yang Feng
备注:ACL 2024 Findings
链接:点击下载PDF文件
摘要:直接语音到语音翻译(S2 ST)已经取得了令人印象深刻的翻译质量,但它往往面临的挑战,由于相当长的语音序列的解码速度慢。最近,一些研究已经转向非自回归(NAR)模型来加速解码,但翻译质量通常明显落后于自回归(AR)模型。在本文中,我们研究了S2 ST中基于CTC的NAR模型的性能,因为这些模型在机器翻译中表现出了令人印象深刻的结果。实验结果表明,通过结合预训练,知识蒸馏,和先进的NAR训练技术,如扫视训练和非单调潜在对齐,基于CTC的NAR模型实现翻译质量相媲美的AR模型,同时保持高达26.81$ times$解码加速比。摘要:Direct speech-to-speech translation (S2ST) has achieved impressive translation quality, but it often faces the challenge of slow decoding due to the considerable length of speech sequences. Recently, some research has turned to non-autoregressive (NAR) models to expedite decoding, yet the translation quality typically lags behind autoregressive (AR) models significantly. In this paper, we investigate the performance of CTC-based NAR models in S2ST, as these models have shown impressive results in machine translation. Experimental results demonstrate that by combining pretraining, knowledge distillation, and advanced NAR training techniques such as glancing training and non-monotonic latent alignments, CTC-based NAR models achieve translation quality comparable to the AR model, while preserving up to 26.81$ times$ decoding speedup.

【8】 Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
标题: 我们能否在没有并行语音数据的情况下实现高质量的直接语音到语音翻译?
作者:Qingkai Fang,Shaolei Zhang,Zhengrui Ma,Min Zhang,Yang Feng
备注:ACL 2024 main conference. Project Page: this https URL
链接:点击下载PDF文件
摘要:最近提出的两遍直接语音到语音翻译(S2 ST)模型将任务分解为语音到文本翻译(S2 TT)和文本到语音(TTS)的端到端模型,产生有希望的结果。然而,这些模型的训练仍然依赖于并行语音数据,这对收集极具挑战性。相比之下,S2 TT和TTS积累了大量的数据和预训练模型,这些数据和模型在S2 ST模型的开发中没有得到充分利用。受此启发,在本文中,我们首先介绍了一种名为ComSpeech的复合S2 ST模型,它可以将任何预训练的S2 TT和TTS模型无缝集成到直接S2 ST模型中。此外,为了消除对并行语音数据的依赖,我们提出了一种新的训练方法ComSpeech-ZS,它只利用S2 TT和TTS数据。它通过对比学习对齐潜在空间中的表示,使从TTS数据学习的语音合成能力以zero-shot方式推广到S2 ST。在CVSS数据集上的实验结果表明,当并行语音数据可用时,ComSpeech在翻译质量和解码速度方面都超过了以前的两遍模型,如UnitY和Translatotron 2。当没有并行语音数据时,ComSpeech-ZS仅落后于 name 0.7 ASR-BLEU,并且优于级联模型。摘要:Recently proposed two-pass direct speech-to-speech translation (S2ST) models decompose the task into speech-to-text translation (S2TT) and text-to-speech (TTS) within an end-to-end model, yielding promising results. However, the training of these models still relies on parallel speech data, which is extremely challenging to collect. In contrast, S2TT and TTS have accumulated a large amount of data and pretrained models, which have not been fully utilized in the development of S2ST models. Inspired by this, in this paper, we first introduce a composite S2ST model named ComSpeech, which can seamlessly integrate any pretrained S2TT and TTS models into a direct S2ST model. Furthermore, to eliminate the reliance on parallel speech data, we propose a novel training method ComSpeech-ZS that solely utilizes S2TT and TTS data. It aligns representations in the latent space through contrastive learning, enabling the speech synthesis capability learned from the TTS data to generalize to S2ST in a zero-shot manner. Experimental results on the CVSS dataset show that when the parallel speech data is available, ComSpeech surpasses previous two-pass models like UnitY and Translatotron 2 in both translation quality and decoding speed. When there is no parallel speech data, ComSpeech-ZS lags behind name by only 0.7 ASR-BLEU and outperforms the cascaded models.

【9】 Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
标题: 使用记录质量和环境的潜在变量通过条件去噪训练进行噪音稳健的语音转换
作者:Takuto Igarashi,Yuki Saito,Kentaro Seki,Shinnosuke Takamichi,Ryuichi Yamamoto,Kentaro Tachibana,Hiroshi Saruwatari
备注:5 pages, accepted for INTERSPEECH 2024, audio samples: this http URL
链接:点击下载PDF文件
摘要:我们提出了噪声鲁棒的语音转换(VC),它考虑到录音质量和环境的噪声源语音。传统的去噪训练通过学习噪声到干净的VC过程来提高VC模型的噪声鲁棒性。然而,当源语音的噪声在训练期间不可见时,转换语音的自然度是有限的。为此,我们提出的训练条件的VC模型上的两个潜在变量代表的录音质量和环境的源语音。这些潜在变量来自于深度神经网络,这些神经网络在记录质量评估和声学场景分类方面进行了预训练,并以话语或帧方式进行计算。因此,训练的VC模型可以在训练期间显式地学习关于语音退化的信息。客观和主观评价表明,我们的训练提高了转换后的语音质量相比,传统的训练。摘要:We propose noise-robust voice conversion (VC) which takes into account the recording quality and environment of noisy source speech. Conventional denoising training improves the noise robustness of a VC model by learning noisy-to-clean VC process. However, the naturalness of the converted speech is limited when the noise of the source speech is unseen during the training. To this end, our proposed training conditions a VC model on two latent variables representing the recording quality and environment of the source speech. These latent variables are derived from deep neural networks pre-trained on recording quality assessment and acoustic scene classification and calculated in an utterance-wise or frame-wise manner. As a result, the trained VC model can explicitly learn information about speech degradation during the training. Objective and subjective evaluations show that our training improves the quality of the converted speech compared to the conventional training.

【10】 AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
标题: AS-70:用于自动语音识别和口吃事件检测的普通话口吃语音数据集
作者:Rong Gong,Hongfei Xue,Lezhi Wang,Xin Xu,Qisheng Li,Lei Xie,Hui Bu,Shaomei Wu,Jiaming Zhou,Yong Qin,Binbin Zhang,Jun Du,Jia Bin,Ming Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在过去的二十年里,语音技术的快速发展已经使自动语音识别(ASR)等任务的性能达到了人类水平。然而,当应用于非典型语音(如口吃)时,这些模型的有效性会降低。本文介绍了AS-70,第一个公开可用的普通话口吃语音数据集,这是同类数据集中最大的数据集。包括会话和语音命令阅读语音,AS-70包括逐字手动转录,使其适用于各种语音相关的任务。此外,基线系统的建立,实验结果的ASR和口吃事件检测(SED)的任务。通过将该数据集纳入模型微调,最先进的ASR模型得到了显著改进,例如,耳语和休伯特,观察,提高他们的包容性,在解决口吃的讲话。摘要:The rapid advancements in speech technologies over the past two decades have led to human-level performance in tasks like automatic speech recognition (ASR) for fluent speech. However, the efficacy of these models diminishes when applied to atypical speech, such as stuttering. This paper introduces AS-70, the first publicly available Mandarin stuttered speech dataset, which stands out as the largest dataset in its category. Encompassing conversational and voice command reading speech, AS-70 includes verbatim manual transcription, rendering it suitable for various speech-related tasks. Furthermore, baseline systems are established, and experimental results are presented for ASR and stuttering event detection (SED) tasks. By incorporating this dataset into the model fine-tuning, significant improvements in the state-of-the-art ASR models, e.g., Whisper and Hubert, are observed, enhancing their inclusivity in addressing stuttered speech.

【11】 SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
标题: SRC 4VC:智能手机录制的语音转换基准数据库
作者:Yuki Saito,Takuto Igarashi,Kentaro Seki,Shinnosuke Takamichi,Ryuichi Yamamoto,Kentaro Tachibana,Hiroshi Saruwatari
备注:Accepted for INTERSPEECH 2024, corpus project page: this https URL
链接:点击下载PDF文件
摘要:我们提出了SRC4VC,一个新的语料库,包含100名日本人在智能手机上记录的11小时的语音。虽然高质量的多说话人语料库可以促进语音转换(VC)技术的发展,但当低质量的语音记录作为输入时,它们并不总是适合测试VC。为此,我们首先要求100名众包工作者使用智能手机记录他们的声音样本。然后,我们用说话者的录音质量分数和话语感知情感标签来注释记录的样本。我们还在任何对任何VC上对SRC4VC进行了基准测试,其中我们在高质量语音上训练了多扬声器VC模型,并使用SRC4VC扬声器的语音样本作为VC中的源。结果表明,训练数据和评价数据之间的记录质量不匹配会显著降低VC的性能,通过对低质量源语音样本进行语音增强可以改善VC的性能。摘要:We present SRC4VC, a new corpus containing 11 hours of speech recorded on smartphones by 100 Japanese speakers. Although high-quality multi-speaker corpora can advance voice conversion (VC) technologies, they are not always suitable for testing VC when low-quality speech recording is given as the input. To this end, we first asked 100 crowdworkers to record their voice samples using smartphones. Then, we annotated the recorded samples with speaker-wise recording-quality scores and utterance-wise perceived emotion labels. We also benchmark SRC4VC on any-to-any VC, in which we trained a multi-speaker VC model on high-quality speech and used the SRC4VC speakers' voice samples as the source in VC. The results show that the recording quality mismatch between the training and evaluation data significantly degrades the VC performance, which can be improved by applying speech enhancement to the low-quality source speech samples.

【12】 ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
标题: ParaCLAP --迈向计算非语言任务的通用语言音频模型
作者:Xin Jing,Andreas Triantafyllopoulos,Björn Schuller
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:对比语言-音频预训练(CLAP)最近出现了一种使音频分析更具普遍性的方法。具体来说,CLAP风格的模型能够“回答”一组不同的语言查询,将音频模型的功能扩展到一组封闭的标签之外。然而,CLAP依赖于大量的(音频,查询)对进行预训练。虽然这样的集合可用于一般的音频任务,如字幕或声音事件检测,但没有用于计算语言学(CP)任务的匹配音频和文本查询的数据集。因此,社区依赖于为通用音频训练的通用CLAP模型,但成功率有限。在本研究中,我们探讨了ParaCLAP的训练考虑因素,这是一种适合CP的CLAP风格模型,包括一种用于创建音频语言查询的新过程。我们证明了其有效性的一组计算语言学的任务,它被证明是超越性能的开源国家的最先进的模型。摘要:Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable. Specifically, CLAP-style models are able to answer' a diverse set of language queries, extending the capabilities of audio models beyond a closed set of labels. However, CLAP relies on a large set of (audio, query) pairs for pretraining. While such sets are available for general audio tasks, like captioning or sound event detection, there are no datasets with matched audio and text queries for computational paralinguistic (CP) tasks. As a result, the community relies on generic CLAP models trained for general audio with limited success. In the present study, we explore training considerations for ParaCLAP, a CLAP-style model suited to CP, including a novel process for creating audio-language queries. We demonstrate its effectiveness on a set of computational paralinguistic tasks, where it is shown to surpass the performance of open-source state-of-the-art models.

【13】 EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
标题: 收件箱:多语言多文本语音情感识别工具包和基准
作者:Ziyang Ma,Mingjie Chen,Hezhao Zhang,Zhisheng Zheng,Wenxi Chen,Xiquan Li,Jiaxin Ye,Xie Chen,Thomas Hain
备注:Accepted by INTERSPEECH 2024. GitHub Repository: this https URL
链接:点击下载PDF文件
摘要:语音情感识别是人机交互的重要组成部分,受到工业界和学术界的广泛关注。然而,目前的SER研究领域长期存在以下问题:1)很少有合理的和通用的数据集分裂,使不同的模型和方法的比较困难。2)没有一个通用的基准涵盖众多的语料库和语言供研究人员参考,这使得复制成为一种负担。在本文中,我们提出了一个开箱即用的多语言多语料库语音情感识别工具包,以及一个基准的语料库内和跨语料库设置。对于语料库内设置,我们仔细设计了不同数据集的数据划分。对于跨语料库设置,我们采用了一个基础SER模型,emotion2vec,以减轻注释错误,并获得一个测试集,是完全平衡的扬声器和情绪分布。基于WEBBox,我们给出了10个预训练语音模型在14种语言的32个情感数据集上的语料库内SER结果,以及在4个完全平衡测试集上的跨语料库SER结果。据我们所知,这是跨语言范围和数量尺度的最大SER基准。我们希望我们的工具包和基准可以促进社会对SER的研究。摘要:Speech emotion recognition (SER) is an important part of human-computer interaction, receiving extensive attention from both industry and academia. However, the current research field of SER has long suffered from the following problems: 1) There are few reasonable and universal splits of the datasets, making comparing different models and methods difficult. 2) No commonly used benchmark covers numerous corpus and languages for researchers to refer to, making reproduction a burden. In this paper, we propose EmoBox, an out-of-the-box multilingual multi-corpus speech emotion recognition toolkit, along with a benchmark for both intra-corpus and cross-corpus settings. For intra-corpus settings, we carefully designed the data partitioning for different datasets. For cross-corpus settings, we employ a foundation SER model, emotion2vec, to mitigate annotation errors and obtain a test set that is fully balanced in speakers and emotions distributions. Based on EmoBox, we present the intra-corpus SER results of 10 pre-trained speech models on 32 emotion datasets with 14 languages, and the cross-corpus SER results on 4 datasets with the fully balanced test sets. To the best of our knowledge, this is the largest SER benchmark, across language scopes and quantity scales. We hope that our toolkit and benchmark can facilitate the research of SER in the community.

【14】 ICGAN: An implicit conditioning method for interpretable feature control of neural audio synthesis
标题: ICGAN:一种用于神经音频合成可解释特征控制的隐式条件反射方法
作者:Yunyi Liu,Craig Jin
链接:点击下载PDF文件
摘要:神经音频合成方法可以通过利用深度生成模型来实现高保真和逼真的声音生成。这样的模型通常依赖于外部标签,其通常是离散的作为调节信息以实现引导声音生成。然而,如果没有适当的描述性标签,仍然很难控制声音的细微变化,特别是在有限的数据集下。本文提出了一种使用生成对抗网络进行神经音频合成的隐式条件反射方法,该方法允许对合成声音的声学特征进行可解释的控制。我们的技术创建了一个连续的调节空间,使音色操作,而不依赖于显式标签。我们进一步引入了一个评估指标,探索可控性,并证明我们的方法是有效的,使一定程度的控制变化的不同合成的声音效果的域内和跨域的声音。摘要:Neural audio synthesis methods can achieve high-fidelity and realistic sound generation by utilizing deep generative models. Such models typically rely on external labels which are often discrete as conditioning information to achieve guided sound generation. However, it remains difficult to control the subtle changes in sounds without appropriate and descriptive labels, especially given a limited dataset. This paper proposes an implicit conditioning method for neural audio synthesis using generative adversarial networks that allows for interpretable control of the acoustic features of synthesized sounds. Our technique creates a continuous conditioning space that enables timbre manipulation without relying on explicit labels. We further introduce an evaluation metric to explore controllability and demonstrate that our approach is effective in enabling a degree of controlled variation of different synthesized sound effects for in-domain and cross-domain sounds.

【15】 Bridging Language Gaps in Audio-Text Retrieval
标题: 弥合音频文本检索中的语言差距
作者:Zhiyong Yan,Heinrich Dinkel,Yongqing Wang,Jizhong Liu,Junbo Zhang,Yujun Wang,Bin Wang
备注:interspeech2024
链接:点击下载PDF文件
摘要:音频文本检索是一项具有挑战性的任务,需要在数据库中搜索音频片段或文本标题。鉴于现实世界数据中大量的非英语内容,现有的研究主要集中在英语描述上,这对此类模型的适用性构成了限制。为了解决这些语言差异,我们提出了一种语言增强(LE),使用多语言文本编码器(SONAR)来编码具有特定语言信息的文本数据。此外,我们通过应用一致性集成蒸馏(CED)优化音频编码器,增强对可变长度音频文本检索的支持。我们的方法在英语音频文本检索方面表现出色,在AudioCaps和Clotho等常用数据集上展示了最先进的(SOTA)性能。同时,该方法在检索其他七种语言的内容时表现出熟练程度,只有10%的额外语言增强训练数据,产生了有希望的结果。源代码可在https: github.com zyyan4 ml-clap上公开获取。摘要:Audio-text retrieval is a challenging task, requiring the search for an audio clip or a text caption within a database. The predominant focus of existing research on English descriptions poses a limitation on the applicability of such models, given the abundance of non-English content in real-world data. To address these linguistic disparities, we propose a language enhancement (LE), using a multilingual text encoder (SONAR) to encode the text data with language-specific information. Additionally, we optimize the audio encoder through the application of consistent ensemble distillation (CED), enhancing support for variable-length audio-text retrieval. Our methodology excels in English audio-text retrieval, demonstrating state-of-the-art (SOTA) performance on commonly used datasets such as AudioCaps and Clotho. Simultaneously, the approach exhibits proficiency in retrieving content in seven other languages with only 10% of additional language-enhanced training data, yielding promising results. The source code is publicly available https: github.com zyyan4 ml-clap.

【16】 Scaling up masked audio encoder learning for general audio classification
标题: 扩大掩蔽音频编码器学习以实现一般音频分类
作者:Heinrich Dinkel,Zhiyong Yan,Yongqing Wang,Junbo Zhang,Yujun Wang,Bin Wang
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:尽管音频分类取得了进展,但语音和其他声音域(如环境声音和音乐)之间仍存在泛化差距。为语音任务训练的模型通常在环境或音乐音频任务上表现不佳,反之亦然。虽然自监督(SSL)音频表示提供了一种替代方案,但对基于SSL的通用音频分类的模型和数据集大小进行缩放的探索有限。我们介绍了一个简单的SSL音频编码器,基于高效的掩码自动编码器框架。在272,356小时的多样化音频上训练了12亿个参数,Dasheng在HEAR基准测试中获得了显着的性能提升。它在CREMA-D,LibriCount,Speech Commands,VoxLingua上的表现优于以前的作品,并且在音乐和环境分类方面表现出色。最近邻分类实验表明,大声特征本身包含丰富的语音、音乐和环境信息。代码可从https: github.com richermans dasheng 获得。摘要:Despite progress in audio classification, a generalization gap remains between speech and other sound domains, such as environmental sounds and music. Models trained for speech tasks often fail to perform well on environmental or musical audio tasks, and vice versa. While self-supervised (SSL) audio representations offer an alternative, there has been limited exploration of scaling both model and dataset sizes for SSL-based general audio classification. We introduce Dasheng, a simple SSL audio encoder, based on the efficient masked autoencoder framework. Trained with 1.2 billion parameters on 272,356 hours of diverse audio, Dasheng obtains significant performance gains on the HEAR benchmark. It outperforms previous works on CREMA-D, LibriCount, Speech Commands, VoxLingua, and competes well in music and environment classification. Dasheng features inherently contain rich speech, music, and environmental information, as shown in nearest-neighbor classification experiments. Code is available https: github.com richermans dasheng .

【17】 AudioMarkBench: Benchmarking Robustness of Audio Watermarking
标题: AudioMarkBench:音频水印的稳健性基准
作者:Hongbin Liu,Moyang Guo,Zhengyuan Jiang,Lun Wang,Neil Zhenqiang Gong
链接:点击下载PDF文件
摘要:在文本到语音模型的进步的推动下,合成语音的日益真实性引起了对模仿和虚假信息的道德关注。音频水印技术通过在人工智能生成的音频中嵌入人类不可感知的水印,提供了一种很有前途的解决方案。然而,音频水印对共同 敌对扰动的鲁棒性仍然研究不足。我们提出AudioMarkBench,第一个系统的基准评估音频水印对水印去除和水印伪造的鲁棒性。AudioMarkBench包括一个从Common-Voice创建的跨语言,生物性别和年龄的新数据集,3种最先进的水印方法和15种扰动类型。我们基准这些方法对扰动的鲁棒性在无盒,黑盒和白盒设置。我们的研究结果突出了当前水印技术的漏洞,并强调需要更强大和公平的音频水印解决方案。我们的数据集和代码可在 url{https: github.com moyangkuo AudioMarkBench}上公开获取。摘要:The increasing realism of synthetic speech, driven by advancements in text-to-speech models, raises ethical concerns regarding impersonation and disinformation. Audio watermarking offers a promising solution via embedding human-imperceptible watermarks into AI-generated audios. However, the robustness of audio watermarking against common adversarial perturbations remains understudied. We present AudioMarkBench, the first systematic benchmark for evaluating the robustness of audio watermarking against watermark removal and watermark forgery. AudioMarkBench includes a new dataset created from Common-Voice across languages, biological sexes, and ages, 3 state-of-the-art watermarking methods, and 15 types of perturbations. We benchmark the robustness of these methods against the perturbations in no-box, black-box, and white-box settings. Our findings highlight the vulnerabilities of current watermarking techniques and emphasize the need for more robust and fair audio watermarking solutions. Our dataset and code are publicly available at url{https: github.com moyangkuo AudioMarkBench}.

【18】 Missingness-resilient Video-enhanced Multimodal Disfluency Detection
标题: 具有丢失弹性的视频增强的多模式不流利检测
作者:Payal Mohapatra,Shamika Likhite,Subrata Biswas,Bashima Islam,Qi Zhu
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:大多数现有的语音不流利检测技术仅依赖于声学数据。在这项工作中,我们提出了一种实用的多模式不流利检测方法,该方法利用可用的视频数据和音频。我们策划了一个视听数据集,并提出了一种新型融合技术,该技术具有统一的权重共享模式不可知编码器来学习时间和语义上下文。我们的弹性设计适应现实世界的场景,其中视频模态在推理过程中有时可能会丢失。我们还提出了替代融合策略时,这两种方式都保证是完整的。在五个不流利检测任务的实验中,我们的统一多模态方法显著优于仅音频的单峰方法,平均绝对改善10%(即,当视频和音频模态始终可用时,增加10个百分点),即使在一半的样本中缺少视频模态,也增加7%。摘要:Most existing speech disfluency detection techniques only rely upon acoustic data. In this work, we present a practical multimodal disfluency detection approach that leverages available video data together with audio. We curate an audiovisual dataset and propose a novel fusion technique with unified weight-sharing modality-agnostic encoders to learn the temporal and semantic context. Our resilient design accommodates real-world scenarios where the video modality may sometimes be missing during inference. We also present alternative fusion strategies when both modalities are assured to be complete. In experiments across five disfluency-detection tasks, our unified multimodal approach significantly outperforms Audio-only unimodal methods, yielding an average absolute improvement of 10% (i.e., 10 percentage point increase) when both video and audio modalities are always available, and 7% even when video modality is missing in half of the samples.

【19】 A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation
标题: 端到端同时语音对任意翻译的非自回归生成框架
作者:Zhengrui Ma,Qingkai Fang,Shaolei Zhang,Shoutao Guo,Yang Feng,Min Zhang
备注:ACL 2024; Codes and demos are at this https URL
链接:点击下载PDF文件
摘要:同声翻译模型在促进沟通方面发挥着至关重要的作用。然而,现有的研究主要集中在文本到文本或语音到文本模型上,需要额外的级联组件来实现语音到语音翻译。这些流水线方法会受到误差传播的影响,并在每个级联组件中累积延迟,从而导致说话者和听众之间的同步降低。为了克服这些挑战,我们提出了一种新型的同步语音翻译非自回归生成框架(NAST-S2 X),它将语音到文本和语音到语音任务集成到统一的端到端框架中。我们开发了一个非自回归解码器,能够同时生成多个文本或声学单元令牌接收固定长度的语音块。解码器可以生成空白或重复的令牌,并采用CTC解码来动态调整其延迟。实验结果表明,NAST-S2 X在语音到文本和语音到语音任务中的性能优于最先进的模型。它在不到3秒的延迟内实现了高质量的同声传译,并在离线生成时提供了28倍的解码加速。摘要:Simultaneous translation models play a crucial role in facilitating communication. However, existing research primarily focuses on text-to-text or speech-to-text models, necessitating additional cascade components to achieve speech-to-speech translation. These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener. To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2X), which integrates speech-to-text and speech-to-speech tasks into a unified end-to-end framework. We develop a non-autoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks. The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency. Experimental results show that NAST-S2X outperforms state-of-the-art models in both speech-to-text and speech-to-speech tasks. It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28 times decoding speedup in offline generation.

【20】 BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification
标题: BTS:元数据辅助呼吸音分类的文本和声音模式的桥梁
作者:June-Woo Kim,Miika Toikkanen,Yera Choi,Seoung-Eun Moon,Ho-Young Jung
备注:Accepted INTERSPEECH 2024
链接:点击下载PDF文件
摘要:呼吸音分类(RSC)是具有挑战性的,由于不同的声学签名,主要受患者人口统计和记录环境的影响。为了解决这个问题,我们引入了一个文本音频多模态模型,利用呼吸音的元数据,这提供了有用的补充信息RSC。具体来说,我们使用来自声音样本元数据的自由文本描述来微调预训练的文本-音频多模态模型,其中包括患者的性别和年龄,记录设备的类型以及患者身体上的记录位置。我们的方法在ICBHI数据集上实现了最先进的性能,超过了之前的最佳结果1.17%。该结果验证了利用元数据和呼吸声样本在增强RSC性能方面的有效性。此外,我们研究了在元数据部分不可用的情况下的模型性能,这可能发生在现实世界的临床环境中。摘要:Respiratory sound classification (RSC) is challenging due to varied acoustic signatures, primarily influenced by patient demographics and recording environments. To address this issue, we introduce a text-audio multimodal model that utilizes metadata of respiratory sounds, which provides useful complementary information for RSC. Specifically, we fine-tune a pretrained text-audio multimodal model using free-text descriptions derived from the sound samples' metadata which includes the gender and age of patients, type of recording devices, and recording location on the patient's body. Our method achieves state-of-the-art performance on the ICBHI dataset, surpassing the previous best result by a notable margin of 1.17%. This result validates the effectiveness of leveraging metadata and respiratory sound samples in enhancing RSC performance. Additionally, we investigate the model performance in the case where metadata is partially unavailable, which may occur in real-world clinical setting.

【21】 SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound
标题: SEE-2-SOUND:Zero-Shot空间环境到空间声音
作者:Rishit Dagli,Shivesh Prakash,Robert Wu,Houman Khosravani
备注:Project Page: this https URL
链接:点击下载PDF文件
摘要:生成组合的视觉和听觉感官体验对于沉浸式内容的消费至关重要。神经生成模型的最新进展使得能够跨多种模式创建高分辨率内容,例如图像,文本,语音和视频。尽管取得了这些成功,但在生成补充生成的视觉内容的高质量空间音频方面仍然存在重大差距。此外,当前的音频生成模型在生成自然音频或语音或音乐方面表现出色,但在集成沉浸式体验所需的空间音频线索方面却有所欠缺。在这项工作中,我们介绍了SEE-2-SOUND,一种zero-shot方法,将任务分解为(1)识别感兴趣的视觉区域;(2)在3D空间中定位这些元素;(3)为每个元素生成单声道音频;(4)将它们集成到空间音频中。使用我们的框架,我们展示了令人信服的结果,从互联网上生成高质量的视频,图像和动态图像的空间音频,以及学习方法生成的媒体。摘要:Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across multiple modalities such as images, text, speech, and videos. Despite these successes, there remains a significant gap in the generation of high-quality spatial audio that complements generated visual content. Furthermore, current audio generation models excel in either generating natural audio or speech or music but fall short in integrating spatial audio cues necessary for immersive experiences. In this work, we introduce SEE-2-SOUND, a zero-shot approach that decomposes the task into (1) identifying visual regions of interest; (2) locating these elements in 3D space; (3) generating mono-audio for each; and (4) integrating them into spatial audio. Using our framework, we demonstrate compelling results for generating spatial audio for high-quality videos, images, and dynamic images from the internet, as well as media generated by learned approaches.

【22】 A Human-in-the-Loop Approach to Improving Cross-Text Prosody Transfer
标题: 改善跨文本韵律迁移的人在环方法
作者:Himanshu Maurya,Atli Sigurgeirsson
备注:4 pages (+1 references), 4 figures, to be presented at Interspeech 2024
链接:点击下载PDF文件
摘要:文语转换(TTS)韵律转换模型可以对同一文本生成不同的韵律再现。这些模型用与目标话语相同的参考来训练。但是,当参考话语与目标文本不同时,如在跨文本韵律迁移中,这些模型很难将韵律从文本中分离出来,导致感知自然度降低。为了解决这个问题,我们提出了一个人在环(HitL)的方法。HitL使用者在保持整体参照韵律效果的前提下,调整韵律的显著关联,使韵律更适合目标语篇。人类调整的翻译保持参考韵律,同时被评为更适合的目标文本的时间为57.8%。我们的分析表明,有限的用户努力足以实现这些改进,并且潜在参考空间中的接近度对于跨文本条件不是可靠的韵律相似性度量。摘要:Text-To-Speech (TTS) prosody transfer models can generate varied prosodic renditions, for the same text, by conditioning on a reference utterance. These models are trained with a reference that is identical to the target utterance. But when the reference utterance differs from the target text, as in cross-text prosody transfer, these models struggle to separate prosody from text, resulting in reduced perceived naturalness. To address this, we propose a Human-in-the-Loop (HitL) approach. HitL users adjust salient correlates of prosody to make the prosody more appropriate for the target text, while maintaining the overall reference prosodic effect. Human adjusted renditions maintain the reference prosody while being rated as more appropriate for the target text $57.8 %$ of the time. Our analysis suggests that limited user effort suffices for these improvements, and that closeness in the latent reference space is not a reliable prosodic similarity metric for the cross-text condition.

【23】 MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
标题: MM-KWS:用于多语言用户定义关键词定位的多模式预算
作者:Zhiqi Ai,Zhiyong Chen,Shugong Xu
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了MM-KWS,一种新的方法,用户定义的关键字定位利用多模态登记的文本和语音模板。与以前的方法,只专注于文本或语音特征,MM-KWS提取音素,文本和语音嵌入这两种方式。然后将这些嵌入与查询语音嵌入进行比较以检测目标关键字。为了确保MM-KWS在不同语言中的适用性,我们使用了一个包含多个多语言预训练模型的特征提取器。随后,我们验证了其有效性的普通话和英语任务。此外,我们还集成了先进的数据增强工具,用于硬案例挖掘,以增强MM-KWS在区分易混淆单词方面的能力。在LibriPhrase和WenetPhrase数据集上的实验结果表明,MM-KWS的性能明显优于现有方法。摘要:In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS extracts phoneme, text, and speech embeddings from both modalities. These embeddings are then compared with the query speech embedding to detect the target keywords. To ensure the applicability of MM-KWS across diverse languages, we utilize a feature extractor incorporating several multilingual pre-trained models. Subsequently, we validate its effectiveness on Mandarin and English tasks. In addition, we have integrated advanced data augmentation tools for hard case mining to enhance MM-KWS in distinguishing confusable words. Experimental results on the LibriPhrase and WenetPhrase datasets demonstrate that MM-KWS outperforms prior methods significantly.

【24】 Description and Discussion on DCASE 2024 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
标题: DUSE 2024挑战任务2的描述和讨论:机器状态监控的第一次无监督异常声音检测
作者:Tomoya Nishida,Noboru Harada,Daisuke Niizumi,Davide Albertini,Roberto Sannino,Simone Pradolini,Filippo Augusti,Keisuke Imoto,Kota Dohi,Harsh Purohit,Takashi Endo,Yohei Kawaguchi
备注:anomaly detection, acoustic condition monitoring, domain shift, first-shot problem, DCASE Challenge. arXiv admin note: text overlap with arXiv:2305.07828
链接:点击下载PDF文件
摘要:我们介绍了声学场景和事件的检测和分类(DCASE)2024挑战任务2的任务描述:用于机器状态监测的首次无监督异常声音检测(ASD)。从去年的DCASE 2023挑战任务2开始,我们将该任务组织为域泛化所需设置下的第一个问题。第一枪问题的主要目标是使ASD系统能够快速部署到新类型的机器上,而不需要特定于机器的超参数调整。这种问题设置是通过以下方式实现的:(1)为每种机器类型仅提供一个部分,以及(2)为开发和评估数据集提供完全不同的机器类型。对于DCASE 2024挑战任务2,全新机器类型的数据被新收集并作为评估数据集提供。此外,隐藏了几种机器类型的属性信息(如机器操作条件),以模拟此类信息不可用的情况。我们将在挑战提交截止日期后添加挑战结果和提交分析。摘要:We present the task description of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 Challenge Task 2: First-shot unsupervised anomalous sound detection (ASD) for machine condition monitoring. Continuing from last year's DCASE 2023 Challenge Task 2, we organize the task as a first-shot problem under domain generalization required settings. The main goal of the first-shot problem is to enable rapid deployment of ASD systems for new kinds of machines without the need for machine-specific hyperparameter tunings. This problem setting was realized by (1) giving only one section for each machine type and (2) having completely different machine types for the development and evaluation datasets. For the DCASE 2024 Challenge Task 2, data of completely new machine types were newly collected and provided as the evaluation dataset. In addition, attribute information such as the machine operation conditions were concealed for several machine types to mimic situations where such information are unavailable. We will add challenge results and analysis of the submissions after the challenge submission deadline.

【25】 CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
标题: CodecFake:增强针对基于编解码器的语音合成系统Deepfake Audios的反欺骗模型
作者:Haibin Wu,Yuan Tseng,Hung-yi Lee
备注:Accepted to Interspeech 2024, project page: this https URL
链接:点击下载PDF文件
摘要:目前最先进的(SOTA)基于编解码器的音频合成系统可以模仿任何人的声音,只需3秒的样本,从特定的看不见的扬声器。不幸的是,恶意攻击者可能会利用这些技术,导致滥用和安全问题。已经开发了反欺骗模型来检测虚假语音。然而,当前的SOTA反欺骗模型是否可以有效地对抗来自基于编解码器的语音合成系统的deepfake音频的公开问题仍然没有答案。在本文中,我们策划了一个广泛的当代SOTA编解码器模型,利用它们来重新创建合成语音。这一努力导致了CodecFake的创建,这是第一个基于编解码器的deepfake音频数据集。此外,我们还验证了在常用数据集上训练的反欺骗模型无法从当前基于编解码器的语音生成系统中检测合成语音。拟议的CodecFake数据集使这些模型能够有效应对这一挑战。摘要:Current state-of-the-art (SOTA) codec-based audio synthesis systems can mimic anyone's voice with just a 3-second sample from that specific unseen speaker. Unfortunately, malicious attackers may exploit these technologies, causing misuse and security issues. Anti-spoofing models have been developed to detect fake speech. However, the open question of whether current SOTA anti-spoofing models can effectively counter deepfake audios from codec-based speech synthesis systems remains unanswered. In this paper, we curate an extensive collection of contemporary SOTA codec models, employing them to re-create synthesized speech. This endeavor leads to the creation of CodecFake, the first codec-based deepfake audio dataset. Additionally, we verify that anti-spoofing models trained on commonly used datasets cannot detect synthesized speech from current codec-based speech generation systems. The proposed CodecFake dataset empowers these models to counter this challenge effectively.

【26】 Translating speech with just images
标题: 仅用图像翻译语音
作者:Dan Oneata,Herman Kamper
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:基于视觉的语音模型将语音与图像联系起来。我们通过现有的图像字幕系统将图像链接到文本来扩展这种连接,从而获得将语音音频直接映射到文本的能力。这种方法可以用于语音翻译,只有图像,通过具有不同的语言从生成的字幕的音频。我们在一个真正的低资源语言Yor ub 'a上研究了这样一个系统,并提出了一个Yor ub ' a到英语的语音翻译模型,该模型利用预先训练的组件,以便能够在低资源环境中学习。为了限制过拟合,我们发现使用一种解码方案来产生用于训练的不同图像标题是至关重要的。结果表明,预测的翻译捕获的主要语义的口语音频,虽然在一个更简单,更短的形式。摘要:Visually grounded speech models link speech to images. We extend this connection by linking images to text via an existing image captioning system, and as a result gain the ability to map speech audio directly to text. This approach can be used for speech translation with just images by having the audio in a different language from the generated captions. We investigate such a system on a real low-resource language, Yor ub 'a, and propose a Yor ub 'a-to-English speech translation model that leverages pretrained components in order to be able to learn in the low-resource regime. To limit overfitting, we find that it is essential to use a decoding scheme that produces diverse image captions for training. Results show that the predicted translations capture the main semantics of the spoken audio, albeit in a simpler and shorter form.

【27】 Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
标题: 具有基于CTE的Word Spotter的用于TLC和Transducer ASB模型的快速上下文偏置
作者:Andrei Andrusenko,Aleksandr Laptev,Vladimir Bataev,Vitaly Lavrukhin,Boris Ginsburg
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:生僻词和新词的准确识别一直是语境化自动语音识别(ASR)系统亟待解决的问题。大多数上下文偏置方法涉及修改ASR模型或波束搜索解码算法,使模型重用复杂化并减慢推理。这项工作提出了一种新的方法来快速上下文偏置与CTC和传感器(RNN-T)ASR模型的CTC为基础的词点(CTC-WS)。所提出的方法匹配CTC对数概率对一个紧凑的上下文图,以检测潜在的上下文偏置的候选人。然后,有效的候选者在相应的帧间隔中替换它们的贪婪识别对应物。Hybrid Transducer-CTC模型支持针对Transducer模型的CTC-WS应用。结果表明,与基线方法相比,上下文偏置识别的速度显着加快,同时提高了F分数和WER。所提出的方法在NVIDIA NeMo工具包中公开提供。摘要:Accurate recognition of rare and new words remains a pressing problem for contextualized Automatic Speech Recognition (ASR) systems. Most context-biasing methods involve modification of the ASR model or the beam-search decoding algorithm, complicating model reuse and slowing down inference. This work presents a new approach to fast context-biasing with CTC-based Word Spotter (CTC-WS) for CTC and Transducer (RNN-T) ASR models. The proposed method matches CTC log-probabilities against a compact context graph to detect potential context-biasing candidates. The valid candidates then replace their greedy recognition counterparts in corresponding frame intervals. A Hybrid Transducer-CTC model enables the CTC-WS application for the Transducer model. The results demonstrate a significant acceleration of the context-biasing recognition with a simultaneous improvement in F-score and WER compared to baseline methods. The proposed method is publicly available in the NVIDIA NeMo toolkit.

【28】 The Reasonable Effectiveness of Speaker Embeddings for Violence Detection
标题: 说话者嵌入暴力检测的合理有效性
作者:Sarthak Jain,Orchid Chetia Phukan,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 24 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:在本文中,我们专注于音频暴力检测(AVD)。AVD是必要的,有几个原因,特别是在维护安全,防止伤害,并确保各种环境中的安全性。这就需要精确的AVD系统。与音频处理中的许多相关应用一样,提高性能的最常见方法是利用自监督(SSL)预训练模型(PTM)。然而,由于这些SSL模型是具有数百万个参数的非常大的模型,这可能会阻碍实际部署,特别是在计算约束环境中。为了解决这个问题,我们建议使用说话人识别模型,这是更小的SSL模型相比。使用SVM和随机森林作为分类器的说话人识别模型嵌入的实验表明,说话人识别模型嵌入与最先进的(SOTA)SSL模型相比表现最好,并达到SOTA结果。摘要:In this paper, we focus on audio violence detection (AVD). AVD is necessary for several reasons, especially in the context of maintaining safety, preventing harm, and ensuring security in various environments. This calls for accurate AVD systems. Like many related applications in audio processing, the most common approach for improving the performance, would be by leveraging self-supervised (SSL) pre-trained models (PTMs). However, as these SSL models are very large models with million of parameters and this can hinder real-world deployment especially in compute-constraint environment. To resolve this, we propose the usage of speaker recognition models which are much smaller compared to the SSL models. Experimentation with speaker recognition model embeddings with SVM & Random Forest as classifiers, we show that speaker recognition model embeddings perform the best in comparison to state-of-the-art (SOTA) SSL models and achieve SOTA results.

【29】 PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
标题: PERSONA:情感识别、性别识别和年龄估计的应用
作者:Devyani Koshal,Orchid Chetia Phukan,Sarthak Jain,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 2024 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:情感识别(ER)、性别识别(GR)和年龄估计(AE)构成了不依赖于口语内容而主要依赖于语音特征(如音高和音调)的非语言学任务。虽然以前的研究在为每个任务单独开发模型方面取得了重大进展,但相对较少强调同时学习这些任务,尽管它们具有内在的相互关联性。因此,在这个演示中,我们提出了PERSONA,一个用于预测ER,GR和AE的应用程序,在后端使用单个模型。值得注意的一点是,我们通过比较研究表明,说话人识别预训练模型(PTM)的表示比最先进的(SOTA)自监督(SSL)PTM更适合这种多任务学习格式。我们的方法避免了为每个任务部署单独模型的需要,并且可以在训练和部署阶段节省资源和时间。摘要:Emotion Recognition (ER), Gender Recognition (GR), and Age Estimation (AE) constitute paralinguistic tasks that rely not on the spoken content but primarily on speech characteristics such as pitch and tone. While previous research has made significant strides in developing models for each task individually, there has been comparatively less emphasis on concurrently learning these tasks, despite their inherent interconnectedness. As such in this demonstration, we present PERSONA, an application for predicting ER, GR, and AE with a single model in the backend. One notable point is we show that representations from speaker recognition pre-trained model (PTM) is better suited for such a multi-task learning format than the state-of-the-art (SOTA) self-supervised (SSL) PTM by carrying out a comparative study. Our methodology obviates the need for deploying separate models for each task and can potentially conserve resources and time during the training and deployment phases.

【30】 ComFeAT: Combination of Neural and Spectral Features for Improved Depression Detection
标题: ComFeAT:结合神经和光谱特征来改进抑郁检测
作者:Orchid Chetia Phukan,Sarthak Jain,Shubham Singh,Muskaan Singh,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 2024 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:在这项工作中,我们专注于通过语音分析检测抑郁症。以前的研究已经广泛地探索了从主要为语言任务训练的预训练模型(PTM)中提取的特征。尽管这些功能已经在基于语音的抑郁症检测方面取得了足够的进展,但它们在现实世界中的性能却有所下降。为了解决这个问题,在本文中,我们介绍了ComFeAT,这是一个应用程序,它采用了一个CNN模型,该模型是在从PTM提取的特征组合上训练的,也就是说,神经特征和频谱特征来增强抑郁检测。谱特征对域变化是鲁棒的,但是,它们在性能上不如神经特征,令人惊讶的是,将它们组合起来显示出互补的行为,并且单独改善了神经和谱特征。所提出的方法也改善了以前的国家的最先进的(SOTA)的作品E-DAIC基准。摘要:In this work, we focus on the detection of depression through speech analysis. Previous research has widely explored features extracted from pre-trained models (PTMs) primarily trained for paralinguistic tasks. Although these features have led to sufficient advances in speech-based depression detection, their performance declines in real-world settings. To address this, in this paper, we introduce ComFeAT, an application that employs a CNN model trained on a combination of features extracted from PTMs, a.k.a. neural features and spectral features to enhance depression detection. Spectral features are robust to domain variations, but, they are not as good as neural features in performance, suprisingly, combining them shows complementary behavior and improves over both neural and spectral features individually. The proposed method also improves over previous state-of-the-art (SOTA) works on E-DAIC benchmark.

【31】 ASTRA: Aligning Speech and Text Representations for Asr without Sampling
标题: ASTRA:无需采样即可对齐Asr的语音和文本表示
作者:Neeraj Gaur,Rohan Agrawal,Gary Wang,Parisa Haghani,Andrew Rosenberg,Bhuvana Ramabhadran
备注:To be published in Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了一种通过文本注入改进自动语音识别(ASR)的新方法ASTRA,与现有技术不同的是,ASTRA不需要通过采样来匹配语音和文本模态之间的序列长度。相反,它利用了CTC RNNT模型中学到的固有对齐。这种方法提供了以下两个优点,即,避免了可能由上采样引起的语音和文本特征之间的潜在不对准,并且消除了对模型的需要,以准确地预测子词令牌的持续时间。这种新颖的形式(长度)匹配作为加权RNNT目标匹配的性能最先进的基于持续时间的方法在FLEURS基准,同时开辟了其他途径的语音处理研究。摘要:This paper introduces ASTRA, a novel method for improving Automatic Speech Recognition (ASR) through text injection.Unlike prevailing techniques, ASTRA eliminates the need for sampling to match sequence lengths between speech and text modalities. Instead, it leverages the inherent alignments learned within CTC RNNT models. This approach offers the following two advantages, namely, avoiding potential misalignment between speech and text features that could arise from upsampling and eliminating the need for models to accurately predict duration of sub-word tokens. This novel formulation of modality (length) matching as a weighted RNNT objective matches the performance of the state-of-the-art duration-based methods on the FLEURS benchmark, while opening up other avenues of research in speech processing.

【32】 Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
标题: 具有强度知识的感知语音自我监督表示学习
作者:Rui Liu,Zening Ma
备注:Accepted by InterSpeech2024
链接:点击下载PDF文件
摘要:语音自监督学习(SSL)在各种下游任务中表现出相当大的效率。然而,目前的自我监督模型往往忽略了情感相关的先验信息的纳入,从而忽视了潜在的增强情感任务的理解,通过情感先验知识的讲话。在本文中,我们提出了一个情感感知的语音表示学习强度知识。具体来说,我们提取帧级的情感强度使用一个已建立的语音情感理解模型。随后,我们提出了一种新型的情感掩蔽策略(EMS),将情感强度纳入掩蔽过程。我们选择了两个基于Transformer和CNN的代表性模型,即MockingJay和非自回归预测编码(NPC),并在IEMOCAP数据集上进行了实验。实验表明,从我们提出的方法得到的表示优于原始模型在SER任务。摘要:Speech Self-Supervised Learning (SSL) has demonstrated considerable efficacy in various downstream tasks. Nevertheless, prevailing self-supervised models often overlook the incorporation of emotion-related prior information, thereby neglecting the potential enhancement of emotion task comprehension through emotion prior knowledge in speech. In this paper, we propose an emotion-aware speech representation learning with intensity knowledge. Specifically, we extract frame-level emotion intensities using an established speech-emotion understanding model. Subsequently, we propose a novel emotional masking strategy (EMS) to incorporate emotion intensities into the masking process. We selected two representative models based on Transformer and CNN, namely MockingJay and Non-autoregressive Predictive Coding (NPC), and conducted experiments on IEMOCAP dataset. Experiments have demonstrated that the representations derived from our proposed method outperform the original model in SER task.

【33】 Sparse Binarization for Fast Keyword Spotting
标题: 用于快速关键词发现的稀疏二进制化
作者:Jonathan Svirsky,Uri Shaham,Ofir Lindenbaum
链接:点击下载PDF文件
摘要:随着语音激活设备和应用程序的日益普及,关键字识别(KWS)模型使用户能够免提与技术进行交互,从而提高各种环境中的便利性和可访问性。在边缘设备(如智能手机和嵌入式系统)上部署KWS模型,可为实时应用程序、隐私和带宽效率带来显著好处。然而,这些设备通常具有有限的计算能力和存储器。这需要优化神经网络模型的效率,而不会显着影响其准确性。为了解决这些挑战,我们提出了一种新的关键字定位模型的基础上稀疏输入表示,其次是线性分类器。该模型比之前最先进的边缘设备兼容模型快四倍,精度更高。我们表明,我们的方法在噪声环境中也更鲁棒,同时速度更快。我们的代码可在https: github.com jsvir sparknet上获得。摘要:With the increasing prevalence of voice-activated devices and applications, keyword spotting (KWS) models enable users to interact with technology hands-free, enhancing convenience and accessibility in various contexts. Deploying KWS models on edge devices, such as smartphones and embedded systems, offers significant benefits for real-time applications, privacy, and bandwidth efficiency. However, these devices often possess limited computational power and memory. This necessitates optimizing neural network models for efficiency without significantly compromising their accuracy. To address these challenges, we propose a novel keyword-spotting model based on sparse input representation followed by a linear classifier. The model is four times faster than the previous state-of-the-art edge device-compatible model with better accuracy. We show that our method is also more robust in noisy environments while being fast. Our code is available at: https: github.com jsvir sparknet.


eess.AS音频处理
【1】 Noise-robust Speech Separation with Fast Generative Correction
标题: 具有快速生成纠正的抗噪语音分离
作者:Helin Wang,Jesus Villalba,Laureano Moro-Velazquez,Jiarui Hai,Thomas Thebaud,Najim Dehak
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音分离是从混合音频信号中分离出多个语音源的任务,在嘈杂的环境中仍然具有挑战性。在本文中,我们提出了一种生成校正方法,以提高判别分离器的输出。通过利用基于扩散模型的生成校正器,我们通过去除噪声和感知上不自然的失真来改进单通道混合语音的分离过程。此外,我们使用预测损失来优化生成模型,以将扩散模型的反向过程简化为单个步骤,并通过反向过程来纠正任何相关的错误。我们的方法在域内Libri 2 Mix噪声数据集和具有各种噪声的域外WSJ上实现了最先进的性能,相对于SepFormer,将SI-SNR提高了22-35%,表现出鲁棒性和强大的泛化能力。摘要:Speech separation, the task of isolating multiple speech sources from a mixed audio signal, remains challenging in noisy environments. In this paper, we propose a generative correction method to enhance the output of a discriminative separator. By leveraging a generative corrector based on a diffusion model, we refine the separation process for single-channel mixture speech by removing noises and perceptually unnatural distortions. Furthermore, we optimize the generative model using a predictive loss to streamline the diffusion model's reverse process into a single step and rectify any associated errors by the reverse process. Our method achieves state-of-the-art performance on the in-domain Libri2Mix noisy dataset, and out-of-domain WSJ with a variety of noises, improving SI-SNR by 22-35% relative to SepFormer, demonstrating robustness and strong generalization capabilities.

【2】 Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
标题: 单编解码器:单码本语音编解码器迈向高性能语音生成
作者:Hanzhao Li,Liumeng Xue,Haohan Guo,Xinfa Zhu,Yuanjun Lv,Lei Xie,Yunlin Chen,Hao Yin,Zhifei Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:多码本语音编解码器使大语言模型(LLM)在TTS中的应用成为可能,但由于多序列预测,效率和鲁棒性受到瓶颈。为了避免这个障碍,我们提出了单编解码器,单码本单序列编解码器,它采用了一个解纠缠的VQ-VAE解耦语音到一个时不变的嵌入和语音丰富的离散序列。此外,编码器通过以下来增强:1)利用BLSTM模块的上下文建模以利用时间信息,2)混合采样模块以减轻来自上采样和下采样的失真,以及3)重新采样模块以鼓励离散单元携带更多的语音信息。与多码本编解码器相比,EnCodec和TiCodec,Single-Codec以仅304 bps的较低带宽展示了更高的重建质量。LLM-TTS实验进一步验证了单码的有效性,表现出更好的自然度和可懂度。摘要:The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid this obstacle, we propose Single-Codec, a single-codebook single-sequence codec, which employs a disentangled VQ-VAE to decouple speech into a time-invariant embedding and a phonetically-rich discrete sequence. Furthermore, the encoder is enhanced with 1) contextual modeling with a BLSTM module to exploit the temporal information, 2) a hybrid sampling module to alleviate distortion from upsampling and downsampling, and 3) a resampling module to encourage discrete units to carry more phonetic information. Compared with multi-codebook codecs, e.g., EnCodec and TiCodec, Single-Codec demonstrates higher reconstruction quality with a lower bandwidth of only 304bps. The effectiveness of Single-Code is further validated by LLM-TTS experiments, showing improved naturalness and intelligibility.

【3】 Clever Hans Effect Found in Automatic Detection of Alzheimer's Disease through Speech
标题: 通过言语自动检测阿尔茨海默病发现聪明的汉斯效应
作者:Yin-Long Liu,Rui Feng,Jia-Hong Yuan,Zhen-Hua Ling
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:我们发现一个潜在的偏见存在于从皮特语料库的图片描述任务产生的录音,最大的公开访问的数据库阿尔茨海默氏病(AD)检测研究。即使只利用这些音频记录的无声片段,我们也可以实现近100%的AD检测准确率。然而,对其他数据集和预处理的Pitt记录采用相同的方法会导致AD检测准确度的典型水平(约80%)。这些结果表明,聪明的汉斯效果在AD检测皮特语料库。我们的研究结果强调了对用于训练深度学习模型的数据集的固有偏差保持警惕的至关重要性,并强调了更好地了解模型性能的必要性。摘要:We uncover an underlying bias present in the audio recordings produced from the picture description task of the Pitt corpus, the largest publicly accessible database for Alzheimer's Disease (AD) detection research. Even by solely utilizing the silent segments of these audio recordings, we achieve nearly 100% accuracy in AD detection. However, employing the same methods to other datasets and preprocessed Pitt recordings results in typical levels (approximately 80%) of AD detection accuracy. These results demonstrate a Clever Hans effect in AD detection on the Pitt corpus. Our findings emphasize the crucial importance of maintaining vigilance regarding inherent biases in datasets utilized for training deep learning models, and highlight the necessity for a better understanding of the models' performance.

【4】 MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
标题: MM-KWS:用于多语言用户定义关键词定位的多模式预算
作者:Zhiqi Ai,Zhiyong Chen,Shugong Xu
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了MM-KWS,一种新的方法,用户定义的关键字定位利用多模态登记的文本和语音模板。与以前的方法,只专注于文本或语音特征,MM-KWS提取音素,文本和语音嵌入这两种方式。然后将这些嵌入与查询语音嵌入进行比较以检测目标关键字。为了确保MM-KWS在不同语言中的适用性,我们使用了一个包含多个多语言预训练模型的特征提取器。随后,我们验证了其有效性的普通话和英语任务。此外,我们还集成了先进的数据增强工具,用于硬案例挖掘,以增强MM-KWS在区分易混淆单词方面的能力。在LibriPhrase和WenetPhrase数据集上的实验结果表明,MM-KWS的性能明显优于现有方法。摘要:In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS extracts phoneme, text, and speech embeddings from both modalities. These embeddings are then compared with the query speech embedding to detect the target keywords. To ensure the applicability of MM-KWS across diverse languages, we utilize a feature extractor incorporating several multilingual pre-trained models. Subsequently, we validate its effectiveness on Mandarin and English tasks. In addition, we have integrated advanced data augmentation tools for hard case mining to enhance MM-KWS in distinguishing confusable words. Experimental results on the LibriPhrase and WenetPhrase datasets demonstrate that MM-KWS outperforms prior methods significantly.

【5】 Description and Discussion on DCASE 2024 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
标题: DUSE 2024挑战任务2的描述和讨论:机器状态监控的第一次无监督异常声音检测
作者:Tomoya Nishida,Noboru Harada,Daisuke Niizumi,Davide Albertini,Roberto Sannino,Simone Pradolini,Filippo Augusti,Keisuke Imoto,Kota Dohi,Harsh Purohit,Takashi Endo,Yohei Kawaguchi
备注:anomaly detection, acoustic condition monitoring, domain shift, first-shot problem, DCASE Challenge. arXiv admin note: text overlap with arXiv:2305.07828
链接:点击下载PDF文件
摘要:我们介绍了声学场景和事件的检测和分类(DCASE)2024挑战任务2的任务描述:用于机器状态监测的首次无监督异常声音检测(ASD)。从去年的DCASE 2023挑战任务2开始,我们将该任务组织为域泛化所需设置下的第一个问题。第一枪问题的主要目标是使ASD系统能够快速部署到新类型的机器上,而不需要特定于机器的超参数调整。这种问题设置是通过以下方式实现的:(1)为每种机器类型仅提供一个部分,以及(2)为开发和评估数据集提供完全不同的机器类型。对于DCASE 2024挑战任务2,全新机器类型的数据被新收集并作为评估数据集提供。此外,隐藏了几种机器类型的属性信息(如机器操作条件),以模拟此类信息不可用的情况。我们将在挑战提交截止日期后添加挑战结果和提交分析。摘要:We present the task description of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 Challenge Task 2: First-shot unsupervised anomalous sound detection (ASD) for machine condition monitoring. Continuing from last year's DCASE 2023 Challenge Task 2, we organize the task as a first-shot problem under domain generalization required settings. The main goal of the first-shot problem is to enable rapid deployment of ASD systems for new kinds of machines without the need for machine-specific hyperparameter tunings. This problem setting was realized by (1) giving only one section for each machine type and (2) having completely different machine types for the development and evaluation datasets. For the DCASE 2024 Challenge Task 2, data of completely new machine types were newly collected and provided as the evaluation dataset. In addition, attribute information such as the machine operation conditions were concealed for several machine types to mimic situations where such information are unavailable. We will add challenge results and analysis of the submissions after the challenge submission deadline.

【6】 CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
标题: CodecFake:增强针对基于编解码器的语音合成系统Deepfake Audios的反欺骗模型
作者:Haibin Wu,Yuan Tseng,Hung-yi Lee
备注:Accepted to Interspeech 2024, project page: this https URL
链接:点击下载PDF文件
摘要:目前最先进的(SOTA)基于编解码器的音频合成系统可以模仿任何人的声音,只需3秒的样本,从特定的看不见的扬声器。不幸的是,恶意攻击者可能会利用这些技术,导致滥用和安全问题。已经开发了反欺骗模型来检测虚假语音。然而,当前的SOTA反欺骗模型是否可以有效地对抗来自基于编解码器的语音合成系统的deepfake音频的公开问题仍然没有答案。在本文中,我们策划了一个广泛的当代SOTA编解码器模型,利用它们来重新创建合成语音。这一努力导致了CodecFake的创建,这是第一个基于编解码器的deepfake音频数据集。此外,我们还验证了在常用数据集上训练的反欺骗模型无法从当前基于编解码器的语音生成系统中检测合成语音。拟议的CodecFake数据集使这些模型能够有效应对这一挑战。摘要:Current state-of-the-art (SOTA) codec-based audio synthesis systems can mimic anyone's voice with just a 3-second sample from that specific unseen speaker. Unfortunately, malicious attackers may exploit these technologies, causing misuse and security issues. Anti-spoofing models have been developed to detect fake speech. However, the open question of whether current SOTA anti-spoofing models can effectively counter deepfake audios from codec-based speech synthesis systems remains unanswered. In this paper, we curate an extensive collection of contemporary SOTA codec models, employing them to re-create synthesized speech. This endeavor leads to the creation of CodecFake, the first codec-based deepfake audio dataset. Additionally, we verify that anti-spoofing models trained on commonly used datasets cannot detect synthesized speech from current codec-based speech generation systems. The proposed CodecFake dataset empowers these models to counter this challenge effectively.

【7】 Target Speech Diarization with Multimodal Prompts
标题: 基于多模式子集的目标语音规模化
作者:Yidi Jiang,Ruijie Tao,Zhengyang Chen,Yanmin Qian,Haizhou Li
备注:13 pages, 7 figures
链接:点击下载PDF文件
摘要:传统的说话人日志化试图根据说话人的特征来检测“谁在什么时候说话”。扩展到目标语音日志,根据语音的语义特征检测“目标事件何时发生”。我们提出了一种新的多模态目标语音日记(MM-TSD)框架,它容纳了多样化和多模态的提示,以灵活和用户友好的方式指定目标事件,包括语义语言描述,预注册的语音,预注册的人脸图像,和音频语言逻辑提示。我们进一步提出了一个语音-人脸对齐模块,将人类的语音和人脸表示投射到一个共享空间中。我们开发了一个基于VoxCeleb 2的多模态数据集,用于MM-TSD训练和评估。此外,我们对每一类提示进行比较分析和消融研究,以验证拟议框架中每个组件的有效性。此外,我们的框架在执行各种信号处理任务,包括扬声器日记和重叠语音检测,使用特定于任务的提示,表现出多功能性。MM-TSD作为一个统一的系统,与专门的模型相比,实现了强大和可比的性能。此外,MM-TSD显示了处理真实世界数据集的复杂对话的能力。摘要:Traditional speaker diarization seeks to detect who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect when target event occurs'' according to the semantic characteristics of speech. We propose a novel Multimodal Target Speech Diarization (MM-TSD) framework, which accommodates diverse and multi-modal prompts to specify target events in a flexible and user-friendly manner, including semantic language description, pre-enrolled speech, pre-registered face image, and audio-language logical prompts. We further propose a voice-face aligner module to project human voice and face representation into a shared space. We develop a multi-modal dataset based on VoxCeleb2 for MM-TSD training and evaluation. Additionally, we conduct comparative analysis and ablation studies for each category of prompts to validate the efficacy of each component in the proposed framework. Furthermore, our framework demonstrates versatility in performing various signal processing tasks, including speaker diarization and overlap speech detection, using task-specific prompts. MM-TSD achieves robust and comparable performance as a unified system compared to specialized models. Moreover, MM-TSD shows capability to handle complex conversations for real-world dataset.

【8】 Translating speech with just images
标题: 仅用图像翻译语音
作者:Dan Oneata,Herman Kamper
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:基于视觉的语音模型将语音与图像联系起来。我们通过现有的图像字幕系统将图像链接到文本来扩展这种连接,从而获得将语音音频直接映射到文本的能力。这种方法可以用于语音翻译,只有图像,通过具有不同的语言从生成的字幕的音频。我们在一个真正的低资源语言Yor ub 'a上研究了这样一个系统,并提出了一个Yor ub ' a到英语的语音翻译模型,该模型利用预先训练的组件,以便能够在低资源环境中学习。为了限制过拟合,我们发现使用一种解码方案来产生用于训练的不同图像标题是至关重要的。结果表明,预测的翻译捕获的主要语义的口语音频,虽然在一个更简单,更短的形式。摘要:Visually grounded speech models link speech to images. We extend this connection by linking images to text via an existing image captioning system, and as a result gain the ability to map speech audio directly to text. This approach can be used for speech translation with just images by having the audio in a different language from the generated captions. We investigate such a system on a real low-resource language, Yor ub 'a, and propose a Yor ub 'a-to-English speech translation model that leverages pretrained components in order to be able to learn in the low-resource regime. To limit overfitting, we find that it is essential to use a decoding scheme that produces diverse image captions for training. Results show that the predicted translations capture the main semantics of the spoken audio, albeit in a simpler and shorter form.

【9】 MR-RawNet: Speaker verification system with multiple temporal resolutions for variable duration utterances using raw waveforms
标题: MR-RawNet:具有多个时间分辨率的说话者验证系统,使用原始波来实现可变持续时间的发声
作者:Seung-bin Kim,Chan-yeong Lim,Jungwoo Heo,Ju-ho Kim,Hyun-seo Shin,Kyo-Won Koo,Ha-Jin Yu
备注:5 pages, accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在说话人验证系统中,短话语的利用提出了一个持续的挑战,导致性能下降,主要是由于语音信息不足,以表征说话人。为了克服这一障碍,我们提出了一种新的结构,MR-RawNet,旨在提高说话人验证系统的鲁棒性对可变持续时间的话语使用原始波形。MR-RawNet通过多分辨率特征提取器从原始波形中提取时频表示,该提取器可同时优化调整时间和频谱分辨率。此外,我们应用了一个多分辨率的注意力块,专注于不同的和广泛的时间背景,确保对话语长度的变化的鲁棒性。在VoxCeleb 1数据集上进行的实验结果表明,与其他基于原始波形的系统相比,MR-RawNet在处理可变持续时间的话语方面表现出优越的性能。摘要:In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To overcome this obstacle, we propose a novel structure, MR-RawNet, designed to enhance the robustness of speaker verification systems against variable duration utterances using raw waveforms. The MR-RawNet extracts time-frequency representations from raw waveforms via a multi-resolution feature extractor that optimally adjusts both temporal and spectral resolutions simultaneously. Furthermore, we apply a multi-resolution attention block that focuses on diverse and extensive temporal contexts, ensuring robustness against changes in utterance length. The experimental results, conducted on VoxCeleb1 dataset, demonstrate that the MR-RawNet exhibits superior performance in handling utterances of variable duration compared to other raw waveform-based systems.

【10】 Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
标题: 具有基于CTE的Word Spotter的用于TLC和Transducer ASB模型的快速上下文偏置
作者:Andrei Andrusenko,Aleksandr Laptev,Vladimir Bataev,Vitaly Lavrukhin,Boris Ginsburg
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:生僻词和新词的准确识别一直是语境化自动语音识别(ASR)系统亟待解决的问题。大多数上下文偏置方法涉及修改ASR模型或波束搜索解码算法,使模型重用复杂化并减慢推理。这项工作提出了一种新的方法来快速上下文偏置与CTC和传感器(RNN-T)ASR模型的CTC为基础的词点(CTC-WS)。所提出的方法匹配CTC对数概率对一个紧凑的上下文图,以检测潜在的上下文偏置的候选人。然后,有效的候选者在相应的帧间隔中替换它们的贪婪识别对应物。Hybrid Transducer-CTC模型支持针对Transducer模型的CTC-WS应用。结果表明,与基线方法相比,上下文偏置识别的速度显着加快,同时提高了F分数和WER。所提出的方法在NVIDIA NeMo工具包中公开提供。摘要:Accurate recognition of rare and new words remains a pressing problem for contextualized Automatic Speech Recognition (ASR) systems. Most context-biasing methods involve modification of the ASR model or the beam-search decoding algorithm, complicating model reuse and slowing down inference. This work presents a new approach to fast context-biasing with CTC-based Word Spotter (CTC-WS) for CTC and Transducer (RNN-T) ASR models. The proposed method matches CTC log-probabilities against a compact context graph to detect potential context-biasing candidates. The valid candidates then replace their greedy recognition counterparts in corresponding frame intervals. A Hybrid Transducer-CTC model enables the CTC-WS application for the Transducer model. The results demonstrate a significant acceleration of the context-biasing recognition with a simultaneous improvement in F-score and WER compared to baseline methods. The proposed method is publicly available in the NVIDIA NeMo toolkit.

【11】 Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech
标题: 使用特定领域语音克隆语音的口语库增强
作者:Mateusz Czyżnikiewicz,Łukasz Bondaruk,Jakub Kubiak,Adam Wiącek,Łukasz Degórski,Marek Kubis,Paweł Skórzewski
链接:点击下载PDF文件
摘要:在本文中,我们研究的影响,增强口语语料库与特定领域的合成样本的目的是训练语音识别系统。使用传统的神经TTS系统和具有语音克隆能力的zero-shot系统,我们生成语音数量不同的语音语料库。我们将使用这两种方法生成的不同数量的合成数据训练的语音识别模型与仅在语音记录上训练的基线模型进行比较。我们表明,虽然语音克隆数据集的质量较低,但其增加的多声音性使其比使用传统神经TTS系统合成的只有几个声音的数据集更有效。此外,我们的实验表明,使用低可变性合成语音迅速导致饱和的ASR的质量,而高可变性语音提供改善,即使当用于训练的数据总量增加30%。摘要:In this paper we study the impact of augmenting spoken language corpora with domain-specific synthetic samples for the purpose of training a speech recognition system. Using both a conventional neural TTS system and a zero-shot one with voice cloning ability we generate speech corpora that vary in the number of voices. We compare speech recognition models trained with addition of different amounts of synthetic data generated using these two methods with a baseline model trained solely on voice recordings. We show that while the quality of voice-cloned dataset is lower, its increased multivoiceity makes it much more effective than the one with only a few voices synthesized with the use of a conventional neural TTS system. Furthermore, our experiments indicate that using low variability synthetic speech quickly leads to saturation in the quality of the ASR whereas high variability speech provides improvement even when increasing total amount of data used for training by 30%.

【12】 The Reasonable Effectiveness of Speaker Embeddings for Violence Detection
标题: 说话者嵌入暴力检测的合理有效性
作者:Sarthak Jain,Orchid Chetia Phukan,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 24 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:在本文中,我们专注于音频暴力检测(AVD)。AVD是必要的,有几个原因,特别是在维护安全,防止伤害,并确保各种环境中的安全性。这就需要精确的AVD系统。与音频处理中的许多相关应用一样,提高性能的最常见方法是利用自监督(SSL)预训练模型(PTM)。然而,由于这些SSL模型是具有数百万个参数的非常大的模型,这可能会阻碍实际部署,特别是在计算约束环境中。为了解决这个问题,我们建议使用说话人识别模型,这是更小的SSL模型相比。使用SVM和随机森林作为分类器的说话人识别模型嵌入的实验表明,说话人识别模型嵌入与最先进的(SOTA)SSL模型相比表现最好,并达到SOTA结果。摘要:In this paper, we focus on audio violence detection (AVD). AVD is necessary for several reasons, especially in the context of maintaining safety, preventing harm, and ensuring security in various environments. This calls for accurate AVD systems. Like many related applications in audio processing, the most common approach for improving the performance, would be by leveraging self-supervised (SSL) pre-trained models (PTMs). However, as these SSL models are very large models with million of parameters and this can hinder real-world deployment especially in compute-constraint environment. To resolve this, we propose the usage of speaker recognition models which are much smaller compared to the SSL models. Experimentation with speaker recognition model embeddings with SVM & Random Forest as classifiers, we show that speaker recognition model embeddings perform the best in comparison to state-of-the-art (SOTA) SSL models and achieve SOTA results.

【13】 PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
标题: PERSONA:情感识别、性别识别和年龄估计的应用
作者:Devyani Koshal,Orchid Chetia Phukan,Sarthak Jain,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 2024 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:情感识别(ER)、性别识别(GR)和年龄估计(AE)构成了不依赖于口语内容而主要依赖于语音特征(如音高和音调)的非语言学任务。虽然以前的研究在为每个任务单独开发模型方面取得了重大进展,但相对较少强调同时学习这些任务,尽管它们具有内在的相互关联性。因此,在这个演示中,我们提出了PERSONA,一个用于预测ER,GR和AE的应用程序,在后端使用单个模型。值得注意的一点是,我们通过比较研究表明,说话人识别预训练模型(PTM)的表示比最先进的(SOTA)自监督(SSL)PTM更适合这种多任务学习格式。我们的方法避免了为每个任务部署单独模型的需要,并且可以在训练和部署阶段节省资源和时间。摘要:Emotion Recognition (ER), Gender Recognition (GR), and Age Estimation (AE) constitute paralinguistic tasks that rely not on the spoken content but primarily on speech characteristics such as pitch and tone. While previous research has made significant strides in developing models for each task individually, there has been comparatively less emphasis on concurrently learning these tasks, despite their inherent interconnectedness. As such in this demonstration, we present PERSONA, an application for predicting ER, GR, and AE with a single model in the backend. One notable point is we show that representations from speaker recognition pre-trained model (PTM) is better suited for such a multi-task learning format than the state-of-the-art (SOTA) self-supervised (SSL) PTM by carrying out a comparative study. Our methodology obviates the need for deploying separate models for each task and can potentially conserve resources and time during the training and deployment phases.

【14】 ComFeAT: Combination of Neural and Spectral Features for Improved Depression Detection
标题: ComFeAT:结合神经和光谱特征来改进抑郁检测
作者:Orchid Chetia Phukan,Sarthak Jain,Shubham Singh,Muskaan Singh,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to INTERSPEECH 2024 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:在这项工作中,我们专注于通过语音分析检测抑郁症。以前的研究已经广泛地探索了从主要为语言任务训练的预训练模型(PTM)中提取的特征。尽管这些功能已经在基于语音的抑郁症检测方面取得了足够的进展,但它们在现实世界中的性能却有所下降。为了解决这个问题,在本文中,我们介绍了ComFeAT,这是一个应用程序,它采用了一个CNN模型,该模型是在从PTM提取的特征组合上训练的,也就是说,神经特征和频谱特征来增强抑郁检测。谱特征对域变化是鲁棒的,但是,它们在性能上不如神经特征,令人惊讶的是,将它们组合起来显示出互补的行为,并且单独改善了神经和谱特征。所提出的方法也改善了以前的国家的最先进的(SOTA)的作品E-DAIC基准。摘要:In this work, we focus on the detection of depression through speech analysis. Previous research has widely explored features extracted from pre-trained models (PTMs) primarily trained for paralinguistic tasks. Although these features have led to sufficient advances in speech-based depression detection, their performance declines in real-world settings. To address this, in this paper, we introduce ComFeAT, an application that employs a CNN model trained on a combination of features extracted from PTMs, a.k.a. neural features and spectral features to enhance depression detection. Spectral features are robust to domain variations, but, they are not as good as neural features in performance, suprisingly, combining them shows complementary behavior and improves over both neural and spectral features individually. The proposed method also improves over previous state-of-the-art (SOTA) works on E-DAIC benchmark.

【15】 ASTRA: Aligning Speech and Text Representations for Asr without Sampling
标题: ASTRA:无需采样即可对齐Asr的语音和文本表示
作者:Neeraj Gaur,Rohan Agrawal,Gary Wang,Parisa Haghani,Andrew Rosenberg,Bhuvana Ramabhadran
备注:To be published in Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了一种通过文本注入改进自动语音识别(ASR)的新方法ASTRA,与现有技术不同的是,ASTRA不需要通过采样来匹配语音和文本模态之间的序列长度。相反,它利用了CTC RNNT模型中学到的固有对齐。这种方法提供了以下两个优点,即,避免了可能由上采样引起的语音和文本特征之间的潜在不对准,并且消除了对模型的需要,以准确地预测子词令牌的持续时间。这种新颖的形式(长度)匹配作为加权RNNT目标匹配的性能最先进的基于持续时间的方法在FLEURS基准,同时开辟了其他途径的语音处理研究。摘要:This paper introduces ASTRA, a novel method for improving Automatic Speech Recognition (ASR) through text injection.Unlike prevailing techniques, ASTRA eliminates the need for sampling to match sequence lengths between speech and text modalities. Instead, it leverages the inherent alignments learned within CTC RNNT models. This approach offers the following two advantages, namely, avoiding potential misalignment between speech and text features that could arise from upsampling and eliminating the need for models to accurately predict duration of sub-word tokens. This novel formulation of modality (length) matching as a weighted RNNT objective matches the performance of the state-of-the-art duration-based methods on the FLEURS benchmark, while opening up other avenues of research in speech processing.

【16】 Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
标题: 具有强度知识的感知语音自我监督表示学习
作者:Rui Liu,Zening Ma
备注:Accepted by InterSpeech2024
链接:点击下载PDF文件
摘要:语音自监督学习(SSL)在各种下游任务中表现出相当大的效率。然而,目前的自我监督模型往往忽略了情感相关的先验信息的纳入,从而忽视了潜在的增强情感任务的理解,通过情感先验知识的讲话。在本文中,我们提出了一个情感感知的语音表示学习强度知识。具体来说,我们提取帧级的情感强度使用一个已建立的语音情感理解模型。随后,我们提出了一种新的情感掩蔽策略(EMS),将情感强度的掩蔽过程。我们选择了两个基于Transformer和CNN的代表性模型,即MockingJay和非自回归预测编码(NPC),并在IEMOCAP数据集上进行了实验。实验表明,从我们提出的方法得到的表示优于原始模型在SER任务。摘要:Speech Self-Supervised Learning (SSL) has demonstrated considerable efficacy in various downstream tasks. Nevertheless, prevailing self-supervised models often overlook the incorporation of emotion-related prior information, thereby neglecting the potential enhancement of emotion task comprehension through emotion prior knowledge in speech. In this paper, we propose an emotion-aware speech representation learning with intensity knowledge. Specifically, we extract frame-level emotion intensities using an established speech-emotion understanding model. Subsequently, we propose a novel emotional masking strategy (EMS) to incorporate emotion intensities into the masking process. We selected two representative models based on Transformer and CNN, namely MockingJay and Non-autoregressive Predictive Coding (NPC), and conducted experiments on IEMOCAP dataset. Experiments have demonstrated that the representations derived from our proposed method outperform the original model in SER task.

【17】 Sparse Binarization for Fast Keyword Spotting
标题: 用于快速关键词发现的稀疏二进制化
作者:Jonathan Svirsky,Uri Shaham,Ofir Lindenbaum
链接:点击下载PDF文件
摘要:随着语音激活设备和应用程序的日益普及,关键字识别(KWS)模型使用户能够免提与技术进行交互,从而提高各种环境中的便利性和可访问性。在边缘设备(如智能手机和嵌入式系统)上部署KWS模型,可为实时应用程序、隐私和带宽效率带来显著好处。然而,这些设备通常具有有限的计算能力和存储器。这需要优化神经网络模型的效率,而不会显着影响其准确性。为了解决这些挑战,我们提出了一种新的关键字定位模型的基础上稀疏输入表示,其次是线性分类器。该模型比之前最先进的边缘设备兼容模型快四倍,精度更高。我们表明,我们的方法在噪声环境中也更鲁棒,同时速度更快。我们的代码可在https: github.com jsvir sparknet上获得。摘要:With the increasing prevalence of voice-activated devices and applications, keyword spotting (KWS) models enable users to interact with technology hands-free, enhancing convenience and accessibility in various contexts. Deploying KWS models on edge devices, such as smartphones and embedded systems, offers significant benefits for real-time applications, privacy, and bandwidth efficiency. However, these devices often possess limited computational power and memory. This necessitates optimizing neural network models for efficiency without significantly compromising their accuracy. To address these challenges, we propose a novel keyword-spotting model based on sparse input representation followed by a linear classifier. The model is four times faster than the previous state-of-the-art edge device-compatible model with better accuracy. We show that our method is also more robust in noisy environments while being fast. Our code is available at: https: github.com jsvir sparknet.

【18】 LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
标题: LoRA-Whisper:参数高效且可扩展的多语言ASB
作者:Zheshu Song,Jianheng Zhuo,Yifan Yang,Ziyang Ma,Shixiong Zhang,Xie Chen
备注:5 pages, 2 figures, conference
链接:点击下载PDF文件
摘要:近年来,在端到端(E2 E)模型的出现和多语言数据集的扩展的推动下,多语言自动语音识别(ASR)取得了重大进展。尽管如此,多语言ASR仍然存在两个主要挑战:语言干扰和在不降低现有语言性能的情况下纳入新语言。本文提出了LoRA-Whisper算法,将LoRA矩阵与Whisper算法相结合,实现多语言语音识别,有效缓解语言干扰。此外,通过利用LoRA和语言之间的相似性,我们可以在新语言上实现更好的性能,同时保持原始语言的一致性能。在八种语言的真实任务上的实验表明,我们提出的LoRA-Whisper在多语言ASR和语言扩展方面分别比基线系统产生了18.5%和23.0%的相对增益。摘要:Recent years have witnessed significant progress in multilingual automatic speech recognition (ASR), driven by the emergence of end-to-end (E2E) models and the scaling of multilingual datasets. Despite that, two main challenges persist in multilingual ASR: language interference and the incorporation of new languages without degrading the performance of the existing ones. This paper proposes LoRA-Whisper, which incorporates LoRA matrix into Whisper for multilingual ASR, effectively mitigating language interference. Furthermore, by leveraging LoRA and the similarities between languages, we can achieve better performance on new languages while upholding consistent performance on original ones. Experiments on a real-world task across eight languages demonstrate that our proposed LoRA-Whisper yields a relative gain of 18.5% and 23.0% over the baseline system for multilingual ASR and language expansion respectively.

【19】 Hearing Anything Anywhere
标题: 任何地方都能听到任何声音
作者:Mason Wang,Ryosuke Sawata,Samuel Clarke,Ruohan Gao,Shangzhe Wu,Jiajun Wu
备注:CVPR 2024. The first two authors contributed equally. Project page: this https URL
链接:点击下载PDF文件
摘要:近年来,3D计算机视觉和计算机图形学取得了巨大的进步,新兴的工具可以为许多混合现实(XR)应用虚拟化真实世界的3D环境。然而,除了沉浸式视觉体验,沉浸式听觉体验对我们对环境的整体感知同样重要。在本文中,我们的目标是重建的空间声学特性的任意环境给定的只有一个稀疏的(大约12)房间脉冲响应(RIR)记录和平面重建的场景,一个设置,是很容易实现的普通用户。为此,我们介绍了DiffRIR,一个可区分的RIR渲染框架与可解释的参数模型的显着的声学特征的场景,包括声源的方向性和表面反射率。这使我们能够通过空间与任何源音频合成新颖的听觉体验。为了评估我们的方法,我们收集了四个不同的真实环境中的RIR录音和音乐数据集。我们表明,我们的模型在渲染单声道和双耳RIR和音乐方面优于最先进的基线,并学习了物理上可解释的参数,这些参数表征了场景中声源和表面的声学特性。摘要:Recent years have seen immense progress in 3D computer vision and computer graphics, with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However, alongside immersive visual experiences, immersive auditory experiences are equally vital to our holistic perception of an environment. In this paper, we aim to reconstruct the spatial acoustic characteristics of an arbitrary environment given only a sparse set of (roughly 12) room impulse response (RIR) recordings and a planar reconstruction of the scene, a setup that is easily achievable by ordinary users. To this end, we introduce DiffRIR, a differentiable RIR rendering framework with interpretable parametric models of salient acoustic features of the scene, including sound source directivity and surface reflectivity. This allows us to synthesize novel auditory experiences through the space with any source audio. To evaluate our method, we collect a dataset of RIR recordings and music in four diverse, real environments. We show that our model outperforms state-ofthe-art baselines on rendering monaural and binaural RIRs and music at unseen locations, and learns physically interpretable parameters characterizing acoustic properties of the sound source and surfaces in the scene.

【20】 RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
标题: RaD-Net 2:因果两阶段修复和去噪语音增强网络,具有知识提炼和复杂的轴向自我注意力
作者:Mingshuai Liu,Zhuangqi Chen,Xiaopeng Yan,Yuanjun Lv,Xianjun Xia,Chuanzeng Huang,Yijian Xiao,Lei Xie
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在实时语音通信系统中,语音信号经常会受到多种失真的影响。最近,在ICASSP 2024语音信号改善(SSI)挑战赛中提出了一种两阶段修复和去噪网络(RaD-Net),具有出色的语音质量改善。然而,未能利用未来的信息和限制卷积层的感受野限制了系统的性能。为了缓解这些问题,我们将RaD-Net扩展到其升级版本RaD-Net 2。具体而言,在第一阶段引入了基于因果关系的知识蒸馏,以因果关系的方式使用未来信息。我们使用非因果修复网络作为教师,以提高因果修复网络的性能。此外,在第二阶段,复轴自注意力应用于去噪网络的复特征编码器 解码器。在ICASSP 2024 SSI Challenge盲测试集上的实验结果表明,与RaD-Net相比,RaD-Net 2带来了0.10 OVRL DNSMOS的改进。摘要:In real-time speech communication systems, speech signals are often degraded by multiple distortions. Recently, a two-stage Repair-and-Denoising network (RaD-Net) was proposed with superior speech quality improvement in the ICASSP 2024 Speech Signal Improvement (SSI) Challenge. However, failure to use future information and constraint receptive field of convolution layers limit the system's performance. To mitigate these problems, we extend RaD-Net to its upgraded version, RaD-Net 2. Specifically, a causality-based knowledge distillation is introduced in the first stage to use future information in a causal way. We use the non-causal repairing network as the teacher to improve the performance of the causal repairing network. In addition, in the second stage, complex axial self-attention is applied in the denoising network's complex feature encoder decoder. Experimental results on the ICASSP 2024 SSI Challenge blind test set show that RaD-Net 2 brings 0.10 OVRL DNSMOS improvement compared to RaD-Net.

【21】 A pilot protocol and cohort for the investigation of non-pathological variability in speech
标题: 研究言语非病理变异性的试点方案和队列
作者:Nicholas Cummins,Lauren L. White,Zahia Rahman,Catriona Lucas,Tian Pan,Ewan Carr,Faith Matcham,Johnny Downs,Richard J. Dobson,Judith Dineley
备注:29 pages. Pre peer review
链接:点击下载PDF文件
摘要:背景基于语音的生物标志物有潜力作为一种手段,定期,客观评估症状的严重程度,远程和临床结合先进的分析模型。然而,言语的复杂性和与健康相关的微妙变化意味着研究结果高度依赖于方法和队列选择。这些往往没有充分报道的研究调查基于语音的健康评估目的开发和应用一个示范协议,以生成一个试点数据集的健康语音与详细的元数据的评估因素的语音记录分析管道,包括设备的选择,语音诱导任务和非病理变异。方法在专题文献综述的基础上,我们制定了我们的收集协议和选择的范例语音特征。我们的协议包括三种不同的语音类型的启发。为了专注于远程应用,我们还选择使用三种不同的麦克风类型来收集语音。我们开发了一个管道来提取一组14个范例语音特征。结果28名受试者每天采集3次语音,8-11周后重复采集; 25名健康受试者每周采集3次语音。收集的参与者特征包括参与者的性别、年龄、母语状态和语音使用习惯。一个初步的14个语音特征,涵盖时间,韵律,语音质量,清晰度和频谱矩特征提取,提供了一个资源的规范值。结论言语录音的收集、加工和分析涉及多种方法学因素。迫切需要一致的报告和更大程度的研究方案协调,以帮助将语音处理转化为临床研究和实践。摘要:Background Speech-based biomarkers have potential as a means for regular, objective assessment of symptom severity, remotely and in-clinic in combination with advanced analytical models. However, the complex nature of speech and the often subtle changes associated with health mean that findings are highly dependent on methodological and cohort choices. These are often not reported adequately in studies investigating speech-based health assessment Objective To develop and apply an exemplar protocol to generate a pilot dataset of healthy speech with detailed metadata for the assessment of factors in the speech recording-analysis pipeline, including device choice, speech elicitation task and non-pathological variability. Methods We developed our collection protocol and choice of exemplar speech features based on a thematic literature review. Our protocol includes the elicitation of three different speech types. With a focus towards remote applications, we also choose to collect speech with three different microphone types. We developed a pipeline to extract a set of 14 exemplar speech features. Results We collected speech from 28 individuals three times in one day, repeated at the same times 8-11 weeks later, and from 25 healthy individuals three times in one week. Participant characteristics collected included sex, age, native language status and voice use habits of the participant. A preliminary set of 14 speech features covering timing, prosody, voice quality, articulation and spectral moment characteristics were extracted that provide a resource of normative values. Conclusions There are multiple methodological factors involved in the collection, processing and analysis of speech recordings. Consistent reporting and greater harmonisation of study protocols are urgently required to aid the translation of speech processing into clinical research and practice.

【22】 Multimodal Belief Prediction
标题: 多峰信念预测
作者:John Murzaku,Adil Soubki,Owen Rambow
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
摘要:识别说话者对信念的承诺程度是一项艰巨的任务;人类不仅在上下文中解释单词的含义,而且还理解来自语调和音频信号的其他方面的线索。NLP社区中的许多论文和语料库都使用纯文本方法来处理信念预测任务。我们是第一个框架和目前的多模态信念预测任务的结果。我们使用CB韵律语料库(CBP),包含对齐的文本和音频与扬声器信念注释。我们首先使用声学韵律特征和传统的机器学习方法报告基线和重要特征。然后,我们分别在BERT和Whisper上呈现用于CBP语料库微调的文本和音频基线。最后,我们提出了我们的多模态架构,微调BERT和耳语,并使用多种融合方法,单独改善这两种方式。摘要:Recognizing a speaker's level of commitment to a belief is a difficult task; humans do not only interpret the meaning of the words in context, but also understand cues from intonation and other aspects of the audio signal. Many papers and corpora in the NLP community have approached the belief prediction task using text-only approaches. We are the first to frame and present results on the multimodal belief prediction task. We use the CB-Prosody corpus (CBP), containing aligned text and audio with speaker belief annotations. We first report baselines and significant features using acoustic-prosodic features and traditional machine learning methods. We then present text and audio baselines for the CBP corpus fine-tuning on BERT and Whisper respectively. Finally, we present our multimodal architecture which fine-tunes on BERT and Whisper and uses multiple fusion methods, improving on both modalities alone.

【23】 Graph-based multi-Feature fusion method for speech emotion recognition
标题: 基于图的语音情感识别多特征融合方法
作者:Xueyu Liu,Jie Lin,Wang Chao
备注:25 pages,4 figures
链接:点击下载PDF文件
摘要:不同的语音特征可以提供反映人类情感状态的互补线索,因此探索合适的方法进行多语音特征融合对于跨语料语音情感识别至关重要。虽然大多数以前的方法只提取一个单一的语音特征的情感识别,现有的融合方法,如级联,并行连接,拼接忽略了异构模式之间的相互作用的功能和功能,导致现有系统的性能。在本文中,我们提出了一种新的基于图的融合方法来明确建模每对语音特征之间的关系。具体来说,我们提出了一种多维边缘特征学习策略,称为基于图的多特征融合方法用于语音情感识别。它将每个语音特征表示为一个节点,并学习多维边缘特征,以明确描述情感识别上下文中每个特征-特征对之间的关系。这样,学习的多维边缘特征从顶点和边缘维度编码语音特征级信息。我们的方法包括三个模块:音频特征生成(AFG)模块,音频特征多维边缘特征(AMEF)模块和语音情感识别(SER)模块。拟议的方法在自营职业妇女协会数据集上取得了令人满意的结果。此外,与AVEC 2019研讨会和挑战赛中的基线相比,该方法表现出更高的性能。我们使用来自两种文化的数据作为我们的训练和验证集:在SEWA数据集上包含德语和匈牙利语的两种文化,德语的CCC分数在唤醒方面提高了17.28%,在喜欢方面提高了7.93%。我们的方法的结果表明,替代融合技术,包括那些采用一维基于边缘的特征融合方法的13%的改善。摘要:Exploring proper way to conduct multi-speech feature fusion for cross-corpus speech emotion recognition is crucial as different speech features could provide complementary cues reflecting human emotion status. While most previous approaches only extract a single speech feature for emotion recognition, existing fusion methods such as concatenation, parallel connection, and splicing ignore heterogeneous patterns in the interaction between features and features, resulting in performance of existing systems. In this paper, we propose a novel graph-based fusion method to explicitly model the relationships between every pair of speech features. Specifically, we propose a multi-dimensional edge features learning strategy called Graph-based multi-Feature fusion method for speech emotion recognition. It represents each speech feature as a node and learns multi-dimensional edge features to explicitly describe the relationship between each feature-feature pair in the context of emotion recognition. This way, the learned multi-dimensional edge features encode speech feature-level information from both the vertex and edge dimensions. Our Approach consists of three modules: an Audio Feature Generation(AFG)module, an Audio-Feature Multi-dimensional Edge Feature(AMEF) module and a Speech Emotion Recognition (SER) module. The proposed methodology yielded satisfactory outcomes on the SEWA dataset. Furthermore, the method demonstrated enhanced performance compared to the baseline in the AVEC 2019 Workshop and Challenge. We used data from two cultures as our training and validation sets: two cultures containing German and Hungarian on the SEWA dataset, the CCC scores for German are improved by 17.28% for arousal and 7.93% for liking. The outcomes of our methodology demonstrate a 13% improvement over alternative fusion techniques, including those employing one dimensional edge-based feature fusion approach.

【24】 A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
标题: 说话人识别中说话人增强的综合研究
作者:Zhenyu Zhou,Shibiao Xu,Shi Yin,Lantian Li,Dong Wang
备注:to be published in INTERSPEECH 2024
链接:点击下载PDF文件
摘要:数据增强(DA)在深度说话人识别的成功中发挥了关键作用。目前的DA技术主要集中在说话人保留增强,这不会改变语音的说话人特征,也不会创建新的说话人。最近的研究揭示了说话人增强的潜力,它可以生成新的说话人来丰富训练数据集。在这项研究中,我们深入到两个扬声器增强方法:速度扰动(SP)和声道长度扰动(VTLP)。尽管这两种方法的经验利用,其有效性缺乏全面的调查。我们使用两个公共数据集VoxCeleb和CN-Celeb进行的研究表明,SP和VTLP都擅长生成新的说话人,从而显著提高了说话人识别的性能。此外,他们表现出不同的属性的扰动因素和数据复杂性的敏感性,暗示在其融合的潜在好处。我们的研究强调了扬声器增强的巨大潜力,强调了深入探索和分析的重要性。摘要:Data augmentation (DA) has played a pivotal role in the success of deep speaker recognition. Current DA techniques primarily focus on speaker-preserving augmentation, which does not change the speaker trait of the speech and does not create new speakers. Recent research has shed light on the potential of speaker augmentation, which generates new speakers to enrich the training dataset. In this study, we delve into two speaker augmentation approaches: speed perturbation (SP) and vocal tract length perturbation (VTLP). Despite the empirical utilization of both methods, a comprehensive investigation into their efficacy is lacking. Our study, conducted using two public datasets, VoxCeleb and CN-Celeb, revealed that both SP and VTLP are proficient at generating new speakers, leading to significant performance improvements in speaker recognition. Furthermore, they exhibit distinct properties in sensitivity to perturbation factors and data complexity, hinting at the potential benefits of their fusion. Our research underscores the substantial potential of speaker augmentation, highlighting the importance of in-depth exploration and analysis.

【25】 CTC-based Non-autoregressive Textless Speech-to-Speech Translation
标题: 基于ATC的非自回归无文本语音翻译
作者:Qingkai Fang,Zhengrui Ma,Yan Zhou,Min Zhang,Yang Feng
备注:ACL 2024 Findings
链接:点击下载PDF文件
摘要:直接语音到语音翻译(S2 ST)已经取得了令人印象深刻的翻译质量,但它往往面临的挑战,由于相当长的语音序列的解码速度慢。最近,一些研究已经转向非自回归(NAR)模型来加速解码,但翻译质量通常明显落后于自回归(AR)模型。在本文中,我们研究了S2 ST中基于CTC的NAR模型的性能,因为这些模型在机器翻译中表现出了令人印象深刻的结果。实验结果表明,通过结合预训练,知识蒸馏,和先进的NAR训练技术,如扫视训练和非单调潜在对齐,基于CTC的NAR模型实现翻译质量相媲美的AR模型,同时保持高达26.81$ times$解码加速比。摘要:Direct speech-to-speech translation (S2ST) has achieved impressive translation quality, but it often faces the challenge of slow decoding due to the considerable length of speech sequences. Recently, some research has turned to non-autoregressive (NAR) models to expedite decoding, yet the translation quality typically lags behind autoregressive (AR) models significantly. In this paper, we investigate the performance of CTC-based NAR models in S2ST, as these models have shown impressive results in machine translation. Experimental results demonstrate that by combining pretraining, knowledge distillation, and advanced NAR training techniques such as glancing training and non-monotonic latent alignments, CTC-based NAR models achieve translation quality comparable to the AR model, while preserving up to 26.81$ times$ decoding speedup.

【26】 Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
标题: 我们能否在没有并行语音数据的情况下实现高质量的直接语音到语音翻译?
作者:Qingkai Fang,Shaolei Zhang,Zhengrui Ma,Min Zhang,Yang Feng
备注:ACL 2024 main conference. Project Page: this https URL
链接:点击下载PDF文件
摘要:最近提出的两遍直接语音到语音翻译(S2 ST)模型将任务分解为语音到文本翻译(S2 TT)和文本到语音(TTS)的端到端模型,产生有希望的结果。然而,这些模型的训练仍然依赖于并行语音数据,这对收集极具挑战性。相比之下,S2 TT和TTS积累了大量的数据和预训练模型,这些数据和模型在S2 ST模型的开发中没有得到充分利用。受此启发,在本文中,我们首先介绍了一种名为ComSpeech的复合S2 ST模型,它可以将任何预训练的S2 TT和TTS模型无缝集成到直接S2 ST模型中。此外,为了消除对并行语音数据的依赖,我们提出了一种新的训练方法ComSpeech-ZS,它只利用S2 TT和TTS数据。它通过对比学习对齐潜在空间中的表示,使从TTS数据学习的语音合成能力以zero-shot方式推广到S2 ST。在CVSS数据集上的实验结果表明,当并行语音数据可用时,ComSpeech在翻译质量和解码速度方面都超过了以前的两遍模型,如UnitY和Translatotron 2。当没有并行语音数据时,ComSpeech-ZS仅落后于 name 0.7 ASR-BLEU,并且优于级联模型。摘要:Recently proposed two-pass direct speech-to-speech translation (S2ST) models decompose the task into speech-to-text translation (S2TT) and text-to-speech (TTS) within an end-to-end model, yielding promising results. However, the training of these models still relies on parallel speech data, which is extremely challenging to collect. In contrast, S2TT and TTS have accumulated a large amount of data and pretrained models, which have not been fully utilized in the development of S2ST models. Inspired by this, in this paper, we first introduce a composite S2ST model named ComSpeech, which can seamlessly integrate any pretrained S2TT and TTS models into a direct S2ST model. Furthermore, to eliminate the reliance on parallel speech data, we propose a novel training method ComSpeech-ZS that solely utilizes S2TT and TTS data. It aligns representations in the latent space through contrastive learning, enabling the speech synthesis capability learned from the TTS data to generalize to S2ST in a zero-shot manner. Experimental results on the CVSS dataset show that when the parallel speech data is available, ComSpeech surpasses previous two-pass models like UnitY and Translatotron 2 in both translation quality and decoding speed. When there is no parallel speech data, ComSpeech-ZS lags behind name by only 0.7 ASR-BLEU and outperforms the cascaded models.

【27】 Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
标题: 使用记录质量和环境的潜在变量通过条件去噪训练进行噪音稳健的语音转换
作者:Takuto Igarashi,Yuki Saito,Kentaro Seki,Shinnosuke Takamichi,Ryuichi Yamamoto,Kentaro Tachibana,Hiroshi Saruwatari
备注:5 pages, accepted for INTERSPEECH 2024, audio samples: this http URL
链接:点击下载PDF文件
摘要:我们提出了噪声鲁棒的语音转换(VC),它考虑到录音质量和环境的噪声源语音。传统的去噪训练通过学习噪声到干净的VC过程来提高VC模型的噪声鲁棒性。然而,当源语音的噪声在训练期间不可见时,转换语音的自然度是有限的。为此,我们提出的训练条件的VC模型上的两个潜在变量代表的录音质量和环境的源语音。这些潜在变量来自于深度神经网络,这些神经网络在记录质量评估和声学场景分类方面进行了预训练,并以话语或帧方式进行计算。因此,训练的VC模型可以在训练期间显式地学习关于语音退化的信息。客观和主观评价表明,我们的训练提高了转换后的语音质量相比,传统的训练。摘要:We propose noise-robust voice conversion (VC) which takes into account the recording quality and environment of noisy source speech. Conventional denoising training improves the noise robustness of a VC model by learning noisy-to-clean VC process. However, the naturalness of the converted speech is limited when the noise of the source speech is unseen during the training. To this end, our proposed training conditions a VC model on two latent variables representing the recording quality and environment of the source speech. These latent variables are derived from deep neural networks pre-trained on recording quality assessment and acoustic scene classification and calculated in an utterance-wise or frame-wise manner. As a result, the trained VC model can explicitly learn information about speech degradation during the training. Objective and subjective evaluations show that our training improves the quality of the converted speech compared to the conventional training.

【28】 AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
标题: AS-70:用于自动语音识别和口吃事件检测的普通话口吃语音数据集
作者:Rong Gong,Hongfei Xue,Lezhi Wang,Xin Xu,Qisheng Li,Lei Xie,Hui Bu,Shaomei Wu,Jiaming Zhou,Yong Qin,Binbin Zhang,Jun Du,Jia Bin,Ming Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在过去的二十年里,语音技术的快速发展已经使自动语音识别(ASR)等任务的性能达到了人类水平。然而,当应用于非典型语音(如口吃)时,这些模型的有效性会降低。本文介绍了AS-70,第一个公开可用的普通话口吃语音数据集,这是同类数据集中最大的数据集。包括会话和语音命令阅读语音,AS-70包括逐字手动转录,使其适用于各种语音相关的任务。此外,基线系统的建立,实验结果的ASR和口吃事件检测(SED)的任务。通过将该数据集纳入模型微调,最先进的ASR模型得到了显著改进,例如,耳语和休伯特,观察,提高他们的包容性,在解决口吃的讲话。摘要:The rapid advancements in speech technologies over the past two decades have led to human-level performance in tasks like automatic speech recognition (ASR) for fluent speech. However, the efficacy of these models diminishes when applied to atypical speech, such as stuttering. This paper introduces AS-70, the first publicly available Mandarin stuttered speech dataset, which stands out as the largest dataset in its category. Encompassing conversational and voice command reading speech, AS-70 includes verbatim manual transcription, rendering it suitable for various speech-related tasks. Furthermore, baseline systems are established, and experimental results are presented for ASR and stuttering event detection (SED) tasks. By incorporating this dataset into the model fine-tuning, significant improvements in the state-of-the-art ASR models, e.g., Whisper and Hubert, are observed, enhancing their inclusivity in addressing stuttered speech.

【29】 SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
标题: SRC 4VC:智能手机录制的语音转换基准数据库
作者:Yuki Saito,Takuto Igarashi,Kentaro Seki,Shinnosuke Takamichi,Ryuichi Yamamoto,Kentaro Tachibana,Hiroshi Saruwatari
备注:Accepted for INTERSPEECH 2024, corpus project page: this https URL
链接:点击下载PDF文件
摘要:我们提出了SRC4VC,一个新的语料库,包含100名日本人在智能手机上记录的11小时的语音。虽然高质量的多说话人语料库可以促进语音转换(VC)技术的发展,但当低质量的语音记录作为输入时,它们并不总是适合测试VC。为此,我们首先要求100名众包工作者使用智能手机记录他们的声音样本。然后,我们用说话者的录音质量分数和话语感知情感标签来注释记录的样本。我们还在任何对任何VC上对SRC4VC进行了基准测试,其中我们在高质量语音上训练了多扬声器VC模型,并使用SRC4VC扬声器的语音样本作为VC中的源。结果表明,训练数据和评价数据之间的记录质量不匹配会显著降低VC的性能,通过对低质量源语音样本进行语音增强可以改善VC的性能。摘要:We present SRC4VC, a new corpus containing 11 hours of speech recorded on smartphones by 100 Japanese speakers. Although high-quality multi-speaker corpora can advance voice conversion (VC) technologies, they are not always suitable for testing VC when low-quality speech recording is given as the input. To this end, we first asked 100 crowdworkers to record their voice samples using smartphones. Then, we annotated the recorded samples with speaker-wise recording-quality scores and utterance-wise perceived emotion labels. We also benchmark SRC4VC on any-to-any VC, in which we trained a multi-speaker VC model on high-quality speech and used the SRC4VC speakers' voice samples as the source in VC. The results show that the recording quality mismatch between the training and evaluation data significantly degrades the VC performance, which can be improved by applying speech enhancement to the low-quality source speech samples.

【30】 ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
标题: ParaCLAP --迈向计算非语言任务的通用语言音频模型
作者:Xin Jing,Andreas Triantafyllopoulos,Björn Schuller
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:对比语言-音频预训练(CLAP)最近出现了一种使音频分析更具普遍性的方法。具体来说,CLAP风格的模型能够“回答”一组不同的语言查询,将音频模型的功能扩展到一组封闭的标签之外。然而,CLAP依赖于大量的(音频,查询)对进行预训练。虽然这样的集合可用于一般的音频任务,如字幕或声音事件检测,但没有用于计算语言学(CP)任务的匹配音频和文本查询的数据集。因此,社区依赖于为通用音频训练的通用CLAP模型,但成功率有限。在本研究中,我们探讨了ParaCLAP的训练考虑因素,这是一种适合CP的CLAP风格模型,包括一种用于创建音频语言查询的新过程。我们证明了其有效性的一组计算语言学的任务,它被证明是超越性能的开源国家的最先进的模型。摘要:Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable. Specifically, CLAP-style models are able to answer' a diverse set of language queries, extending the capabilities of audio models beyond a closed set of labels. However, CLAP relies on a large set of (audio, query) pairs for pretraining. While such sets are available for general audio tasks, like captioning or sound event detection, there are no datasets with matched audio and text queries for computational paralinguistic (CP) tasks. As a result, the community relies on generic CLAP models trained for general audio with limited success. In the present study, we explore training considerations for ParaCLAP, a CLAP-style model suited to CP, including a novel process for creating audio-language queries. We demonstrate its effectiveness on a set of computational paralinguistic tasks, where it is shown to surpass the performance of open-source state-of-the-art models.

【31】 EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
标题: 收件箱:多语言多文本语音情感识别工具包和基准
作者:Ziyang Ma,Mingjie Chen,Hezhao Zhang,Zhisheng Zheng,Wenxi Chen,Xiquan Li,Jiaxin Ye,Xie Chen,Thomas Hain
备注:Accepted by INTERSPEECH 2024. GitHub Repository: this https URL
链接:点击下载PDF文件
摘要:语音情感识别是人机交互的重要组成部分,受到工业界和学术界的广泛关注。然而,目前的SER研究领域长期存在以下问题:1)很少有合理的和通用的数据集分裂,使不同的模型和方法的比较困难。2)没有一个通用的基准涵盖众多的语料库和语言供研究人员参考,这使得复制成为一种负担。在本文中,我们提出了一个开箱即用的多语言多语料库语音情感识别工具包,以及一个基准的语料库内和跨语料库设置。对于语料库内设置,我们仔细设计了不同数据集的数据划分。对于跨语料库设置,我们采用了一个基础SER模型,emotion2vec,以减轻注释错误,并获得一个测试集,是完全平衡的扬声器和情绪分布。基于WEBBox,我们给出了10个预训练语音模型在14种语言的32个情感数据集上的语料库内SER结果,以及在4个完全平衡测试集上的跨语料库SER结果。据我们所知,这是跨语言范围和数量尺度的最大SER基准。我们希望我们的工具包和基准可以促进社会对SER的研究。摘要:Speech emotion recognition (SER) is an important part of human-computer interaction, receiving extensive attention from both industry and academia. However, the current research field of SER has long suffered from the following problems: 1) There are few reasonable and universal splits of the datasets, making comparing different models and methods difficult. 2) No commonly used benchmark covers numerous corpus and languages for researchers to refer to, making reproduction a burden. In this paper, we propose EmoBox, an out-of-the-box multilingual multi-corpus speech emotion recognition toolkit, along with a benchmark for both intra-corpus and cross-corpus settings. For intra-corpus settings, we carefully designed the data partitioning for different datasets. For cross-corpus settings, we employ a foundation SER model, emotion2vec, to mitigate annotation errors and obtain a test set that is fully balanced in speakers and emotions distributions. Based on EmoBox, we present the intra-corpus SER results of 10 pre-trained speech models on 32 emotion datasets with 14 languages, and the cross-corpus SER results on 4 datasets with the fully balanced test sets. To the best of our knowledge, this is the largest SER benchmark, across language scopes and quantity scales. We hope that our toolkit and benchmark can facilitate the research of SER in the community.

【32】 ICGAN: An implicit conditioning method for interpretable feature control of neural audio synthesis
标题: ICGAN:一种用于神经音频合成可解释特征控制的隐式条件反射方法
作者:Yunyi Liu,Craig Jin
链接:点击下载PDF文件
摘要:神经音频合成方法可以通过利用深度生成模型来实现高保真和逼真的声音生成。这样的模型通常依赖于外部标签,其通常是离散的作为调节信息以实现引导声音生成。然而,如果没有适当的描述性标签,仍然很难控制声音的细微变化,特别是在有限的数据集下。本文提出了一种使用生成对抗网络进行神经音频合成的隐式条件反射方法,该方法允许对合成声音的声学特征进行可解释的控制。我们的技术创建了一个连续的调节空间,使音色操作,而不依赖于显式标签。我们进一步引入了一个评估指标,探索可控性,并证明我们的方法是有效的,使一定程度的控制变化的不同合成的声音效果的域内和跨域的声音。摘要:Neural audio synthesis methods can achieve high-fidelity and realistic sound generation by utilizing deep generative models. Such models typically rely on external labels which are often discrete as conditioning information to achieve guided sound generation. However, it remains difficult to control the subtle changes in sounds without appropriate and descriptive labels, especially given a limited dataset. This paper proposes an implicit conditioning method for neural audio synthesis using generative adversarial networks that allows for interpretable control of the acoustic features of synthesized sounds. Our technique creates a continuous conditioning space that enables timbre manipulation without relying on explicit labels. We further introduce an evaluation metric to explore controllability and demonstrate that our approach is effective in enabling a degree of controlled variation of different synthesized sound effects for in-domain and cross-domain sounds.

【33】 Bridging Language Gaps in Audio-Text Retrieval
标题: 弥合音频文本检索中的语言差距
作者:Zhiyong Yan,Heinrich Dinkel,Yongqing Wang,Jizhong Liu,Junbo Zhang,Yujun Wang,Bin Wang
备注:interspeech2024
链接:点击下载PDF文件
摘要:音频文本检索是一项具有挑战性的任务,需要在数据库中搜索音频片段或文本标题。鉴于现实世界数据中大量的非英语内容,现有的研究主要集中在英语描述上,这对此类模型的适用性构成了限制。为了解决这些语言差异,我们提出了一种语言增强(LE),使用多语言文本编码器(SONAR)来编码具有特定语言信息的文本数据。此外,我们通过应用一致性集成蒸馏(CED)优化音频编码器,增强对可变长度音频文本检索的支持。我们的方法在英语音频文本检索方面表现出色,在AudioCaps和Clotho等常用数据集上展示了最先进的(SOTA)性能。同时,该方法在检索其他七种语言的内容时表现出熟练程度,只有10%的额外语言增强训练数据,产生了有希望的结果。源代码可在https: github.com zyyan4 ml-clap上公开获取。摘要:Audio-text retrieval is a challenging task, requiring the search for an audio clip or a text caption within a database. The predominant focus of existing research on English descriptions poses a limitation on the applicability of such models, given the abundance of non-English content in real-world data. To address these linguistic disparities, we propose a language enhancement (LE), using a multilingual text encoder (SONAR) to encode the text data with language-specific information. Additionally, we optimize the audio encoder through the application of consistent ensemble distillation (CED), enhancing support for variable-length audio-text retrieval. Our methodology excels in English audio-text retrieval, demonstrating state-of-the-art (SOTA) performance on commonly used datasets such as AudioCaps and Clotho. Simultaneously, the approach exhibits proficiency in retrieving content in seven other languages with only 10% of additional language-enhanced training data, yielding promising results. The source code is publicly available https: github.com zyyan4 ml-clap.

【34】 Scaling up masked audio encoder learning for general audio classification
标题: 扩大掩蔽音频编码器学习以实现一般音频分类
作者:Heinrich Dinkel,Zhiyong Yan,Yongqing Wang,Junbo Zhang,Yujun Wang,Bin Wang
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:尽管音频分类取得了进展,但语音和其他声音域(如环境声音和音乐)之间仍存在泛化差距。为语音任务训练的模型通常在环境或音乐音频任务上表现不佳,反之亦然。虽然自监督(SSL)音频表示提供了一种替代方案,但对基于SSL的通用音频分类的模型和数据集大小进行缩放的探索有限。我们介绍了一个简单的SSL音频编码器,基于高效的掩码自动编码器框架。在272,356小时的多样化音频中训练了12亿个参数,Dasheng在HEAR基准测试中获得了显着的性能提升。它在CREMA-D,LibriCount,Speech Commands,VoxLingua上的表现优于以前的作品,并在音乐和环境分类方面表现出色。最近邻分类实验表明,大声特征本身包含丰富的语音、音乐和环境信息。代码可从https: github.com richermans dasheng 获得。摘要:Despite progress in audio classification, a generalization gap remains between speech and other sound domains, such as environmental sounds and music. Models trained for speech tasks often fail to perform well on environmental or musical audio tasks, and vice versa. While self-supervised (SSL) audio representations offer an alternative, there has been limited exploration of scaling both model and dataset sizes for SSL-based general audio classification. We introduce Dasheng, a simple SSL audio encoder, based on the efficient masked autoencoder framework. Trained with 1.2 billion parameters on 272,356 hours of diverse audio, Dasheng obtains significant performance gains on the HEAR benchmark. It outperforms previous works on CREMA-D, LibriCount, Speech Commands, VoxLingua, and competes well in music and environment classification. Dasheng features inherently contain rich speech, music, and environmental information, as shown in nearest-neighbor classification experiments. Code is available https: github.com richermans dasheng .

【35】 AudioMarkBench: Benchmarking Robustness of Audio Watermarking
标题: AudioMarkBench:音频水印的稳健性基准
作者:Hongbin Liu,Moyang Guo,Zhengyuan Jiang,Lun Wang,Neil Zhenqiang Gong
链接:点击下载PDF文件
摘要:在文本到语音模型的进步的推动下,合成语音的日益真实性引起了对模仿和虚假信息的道德关注。音频水印技术通过在人工智能生成的音频中嵌入人类不可感知的水印,提供了一种很有前途的解决方案。然而,音频水印对共同 敌对扰动的鲁棒性仍然研究不足。我们提出AudioMarkBench,第一个系统的基准评估音频水印对水印去除和水印伪造的鲁棒性。AudioMarkBench包括一个从Common-Voice创建的跨语言,生物性别和年龄的新数据集,3种最先进的水印方法和15种扰动类型。我们基准这些方法对扰动的鲁棒性在无盒,黑盒和白盒设置。我们的研究结果突出了当前水印技术的漏洞,并强调需要更强大和公平的音频水印解决方案。我们的数据集和代码可在 url{https: github.com moyangkuo AudioMarkBench}上公开获取。摘要:The increasing realism of synthetic speech, driven by advancements in text-to-speech models, raises ethical concerns regarding impersonation and disinformation. Audio watermarking offers a promising solution via embedding human-imperceptible watermarks into AI-generated audios. However, the robustness of audio watermarking against common adversarial perturbations remains understudied. We present AudioMarkBench, the first systematic benchmark for evaluating the robustness of audio watermarking against watermark removal and watermark forgery. AudioMarkBench includes a new dataset created from Common-Voice across languages, biological sexes, and ages, 3 state-of-the-art watermarking methods, and 15 types of perturbations. We benchmark the robustness of these methods against the perturbations in no-box, black-box, and white-box settings. Our findings highlight the vulnerabilities of current watermarking techniques and emphasize the need for more robust and fair audio watermarking solutions. Our dataset and code are publicly available at url{https: github.com moyangkuo AudioMarkBench}.

【36】 Missingness-resilient Video-enhanced Multimodal Disfluency Detection
标题: 具有丢失弹性的视频增强的多模式不流利检测
作者:Payal Mohapatra,Shamika Likhite,Subrata Biswas,Bashima Islam,Qi Zhu
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:大多数现有的语音不流利检测技术仅依赖于声学数据。在这项工作中,我们提出了一个实用的多模态不流利检测方法,利用可用的视频数据与音频。我们策划了一个视听数据集,并提出了一种新的融合技术与统一的权重共享模态不可知的编码器学习的时间和语义上下文。我们的弹性设计适应现实世界的情况下,视频模态有时可能会在推理过程中丢失。我们还提出了替代融合策略时,这两种方式都保证是完整的。在五个不流利检测任务的实验中,我们的统一多模态方法显著优于仅音频的单峰方法,平均绝对改善10%(即,当视频和音频模态始终可用时,增加10个百分点),即使在一半的样本中缺少视频模态,也增加7%。摘要:Most existing speech disfluency detection techniques only rely upon acoustic data. In this work, we present a practical multimodal disfluency detection approach that leverages available video data together with audio. We curate an audiovisual dataset and propose a novel fusion technique with unified weight-sharing modality-agnostic encoders to learn the temporal and semantic context. Our resilient design accommodates real-world scenarios where the video modality may sometimes be missing during inference. We also present alternative fusion strategies when both modalities are assured to be complete. In experiments across five disfluency-detection tasks, our unified multimodal approach significantly outperforms Audio-only unimodal methods, yielding an average absolute improvement of 10% (i.e., 10 percentage point increase) when both video and audio modalities are always available, and 7% even when video modality is missing in half of the samples.

【37】 A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Any Translation
标题: 端到端同时语音对任意翻译的非自回归生成框架
作者:Zhengrui Ma,Qingkai Fang,Shaolei Zhang,Shoutao Guo,Yang Feng,Min Zhang
备注:ACL 2024; Codes and demos are at this https URL
链接:点击下载PDF文件
摘要:同声翻译模型在促进沟通方面发挥着至关重要的作用。然而,现有的研究主要集中在文本到文本或语音到文本模型上,需要额外的级联组件来实现语音到语音翻译。这些流水线方法会受到误差传播的影响,并在每个级联组件中累积延迟,从而导致说话者和听众之间的同步降低。为了克服这些挑战,我们提出了一种新型的同步语音翻译非自回归生成框架(NAST-S2 X),它将语音到文本和语音到语音任务集成到统一的端到端框架中。我们开发了一个非自回归解码器,能够同时生成多个文本或声学单元令牌接收固定长度的语音块。解码器可以生成空白或重复的令牌,并采用CTC解码来动态调整其延迟。实验结果表明,NAST-S2 X在语音到文本和语音到语音任务中的性能优于最先进的模型。它在不到3秒的延迟内实现了高质量的同声传译,并在离线生成时提供了28倍的解码加速。摘要:Simultaneous translation models play a crucial role in facilitating communication. However, existing research primarily focuses on text-to-text or speech-to-text models, necessitating additional cascade components to achieve speech-to-speech translation. These pipeline methods suffer from error propagation and accumulate delays in each cascade component, resulting in reduced synchronization between the speaker and listener. To overcome these challenges, we propose a novel non-autoregressive generation framework for simultaneous speech translation (NAST-S2X), which integrates speech-to-text and speech-to-speech tasks into a unified end-to-end framework. We develop a non-autoregressive decoder capable of concurrently generating multiple text or acoustic unit tokens upon receiving fixed-length speech chunks. The decoder can generate blank or repeated tokens and employ CTC decoding to dynamically adjust its latency. Experimental results show that NAST-S2X outperforms state-of-the-art models in both speech-to-text and speech-to-speech tasks. It achieves high-quality simultaneous interpretation within a delay of less than 3 seconds and provides a 28 times decoding speedup in offline generation.

【38】 BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification
标题: BTS:元数据辅助呼吸音分类的文本和声音模式的桥梁
作者:June-Woo Kim,Miika Toikkanen,Yera Choi,Seoung-Eun Moon,Ho-Young Jung
备注:Accepted INTERSPEECH 2024
链接:点击下载PDF文件
摘要:呼吸音分类(RSC)是具有挑战性的,由于不同的声学签名,主要受患者人口统计和记录环境的影响。为了解决这个问题,我们引入了一个文本音频多模态模型,利用呼吸音的元数据,这提供了有用的补充信息RSC。具体来说,我们使用来自声音样本元数据的自由文本描述来微调预训练的文本-音频多模态模型,其中包括患者的性别和年龄,记录设备的类型以及患者身体上的记录位置。我们的方法在ICBHI数据集上实现了最先进的性能,超过了之前的最佳结果1.17%。该结果验证了利用元数据和呼吸声样本在增强RSC性能方面的有效性。此外,我们研究了在元数据部分不可用的情况下的模型性能,这可能发生在现实世界的临床环境中。摘要:Respiratory sound classification (RSC) is challenging due to varied acoustic signatures, primarily influenced by patient demographics and recording environments. To address this issue, we introduce a text-audio multimodal model that utilizes metadata of respiratory sounds, which provides useful complementary information for RSC. Specifically, we fine-tune a pretrained text-audio multimodal model using free-text descriptions derived from the sound samples' metadata which includes the gender and age of patients, type of recording devices, and recording location on the patient's body. Our method achieves state-of-the-art performance on the ICBHI dataset, surpassing the previous best result by a notable margin of 1.17%. This result validates the effectiveness of leveraging metadata and respiratory sound samples in enhancing RSC performance. Additionally, we investigate the model performance in the case where metadata is partially unavailable, which may occur in real-world clinical setting.

【39】 SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound
标题: SEE-2-SOUND:Zero-Shot空间环境到空间声音
作者:Rishit Dagli,Shivesh Prakash,Robert Wu,Houman Khosravani
备注:Project Page: this https URL
链接:点击下载PDF文件
摘要:生成组合的视觉和听觉感官体验对于沉浸式内容的消费至关重要。神经生成模型的最新进展使得能够跨多种模式创建高分辨率内容,例如图像,文本,语音和视频。尽管取得了这些成功,但在生成补充生成的视觉内容的高质量空间音频方面仍然存在重大差距。此外,当前的音频生成模型在生成自然音频或语音或音乐方面表现出色,但在集成沉浸式体验所需的空间音频线索方面却有所欠缺。在这项工作中,我们介绍了SEE-2-SOUND,一种zero-shot方法,将任务分解为(1)识别感兴趣的视觉区域;(2)在3D空间中定位这些元素;(3)为每个元素生成单声道音频;(4)将它们集成到空间音频中。使用我们的框架,我们展示了令人信服的结果,从互联网上生成高质量的视频,图像和动态图像的空间音频,以及学习方法生成的媒体。摘要:Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across multiple modalities such as images, text, speech, and videos. Despite these successes, there remains a significant gap in the generation of high-quality spatial audio that complements generated visual content. Furthermore, current audio generation models excel in either generating natural audio or speech or music but fall short in integrating spatial audio cues necessary for immersive experiences. In this work, we introduce SEE-2-SOUND, a zero-shot approach that decomposes the task into (1) identifying visual regions of interest; (2) locating these elements in 3D space; (3) generating mono-audio for each; and (4) integrating them into spatial audio. Using our framework, we demonstrate compelling results for generating spatial audio for high-quality videos, images, and dynamic images from the internet, as well as media generated by learned approaches.

【40】 A Human-in-the-Loop Approach to Improving Cross-Text Prosody Transfer
标题: 改善跨文本韵律迁移的人在环方法
作者:Himanshu Maurya,Atli Sigurgeirsson
备注:4 pages (+1 references), 4 figures, to be presented at Interspeech 2024
链接:点击下载PDF文件
摘要:文语转换(TTS)韵律转换模型可以对同一文本生成不同的韵律再现。这些模型用与目标话语相同的参考来训练。但是,当参考话语与目标文本不同时,如在跨文本韵律迁移中,这些模型很难将韵律从文本中分离出来,导致感知自然度降低。为了解决这个问题,我们提出了一个人在环(HitL)的方法。HitL使用者在保持整体参照韵律效果的前提下,调整韵律的显著关联,使韵律更适合目标语篇。人类调整的翻译保持参考韵律,同时被评为更适合的目标文本的时间为57.8%。我们的分析表明,有限的用户努力足以实现这些改进,并且潜在参考空间中的接近度对于跨文本条件不是可靠的韵律相似性度量。摘要:Text-To-Speech (TTS) prosody transfer models can generate varied prosodic renditions, for the same text, by conditioning on a reference utterance. These models are trained with a reference that is identical to the target utterance. But when the reference utterance differs from the target text, as in cross-text prosody transfer, these models struggle to separate prosody from text, resulting in reduced perceived naturalness. To address this, we propose a Human-in-the-Loop (HitL) approach. HitL users adjust salient correlates of prosody to make the prosody more appropriate for the target text, while maintaining the overall reference prosodic effect. Human adjusted renditions maintain the reference prosody while being rated as more appropriate for the target text $57.8 %$ of the time. Our analysis suggests that limited user effort suffices for these improvements, and that closeness in the latent reference space is not a reliable prosodic similarity metric for the cross-text condition.

【41】 Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
标题: 具有预训练大语言模型的离散多模式转换器用于混合监督语音处理
作者:Viet Anh Trinh,Rosy Southwell,Yiwen Guan,Xinlu He,Zhiyong Wang,Jacob Whitehill
链接:点击下载PDF文件
摘要:最近关于离散语音标记化的工作为可以跨模态无缝执行多个任务的模型铺平了道路,例如,语音识别、文本到语音、语音到语音翻译。此外,从大量文本语料库中预训练的大型语言模型(LLM)包含丰富的语言信息,可以提高各种任务的准确性。在本文中,我们提出了一个仅解码器的离散多模态语言模型(DMLM),它可以灵活地应用于多个任务(ASR,T2S,S2TT等)。和模式(文本、语音、视觉)。我们探讨了离散多模态模型的几个关键方面,包括损失函数,权重初始化,混合训练监督和码本。我们的研究结果表明,DMLM在多个任务和数据集上从监督和无监督训练的组合中受益匪浅。此外,对于ASR,它受益于从预训练的LLM初始化DMLM,以及从Whisper激活导出的码本。摘要:Recent work on discrete speech tokenization has paved the way for models that can seamlessly perform multiple tasks across modalities, e.g., speech recognition, text to speech, speech to speech translation. Moreover, large language models (LLMs) pretrained from vast text corpora contain rich linguistic information that can improve accuracy in a variety of tasks. In this paper, we present a decoder-only Discrete Multimodal Language Model (DMLM), which can be flexibly applied to multiple tasks (ASR, T2S, S2TT, etc.) and modalities (text, speech, vision). We explore several critical aspects of discrete multi-modal models, including the loss function, weight initialization, mixed training supervision, and codebook. Our results show that DMLM benefits significantly, across multiple tasks and datasets, from a combination of supervised and unsupervised training. Moreover, for ASR, it benefits from initializing DMLM from a pretrained LLM, and from a codebook derived from Whisper activations.


机器翻译,仅供参考