今日论文合集:cs.SD语音34篇,eess.AS音频处理38篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
标题: DiscreteSLU:一个具有自监督离散语音单元的大型语言模型,用于口语理解
作者:Suwon Shon,Kwangyoun Kim,Yi-Te Hsu,Prashant Sridhar,Shinji Watanabe,Karen Livescu
链接:点击下载PDF文件
摘要:预训练的基于文本的大型语言模型(LLM)与语音输入的集成为不同的语音任务提供了指令遵循能力。这种集成需要使用语音编码器、语音适配器和接受过不同任务培训的LLM。我们建议使用离散语音单元(DSU),而不是连续值语音编码器输出,并使用语音适配器将其转换到LLM令牌嵌入空间。我们使用自监督语音编码器生成DSU,然后进行k均值聚类。所提出的模型表现出强大的性能语音输入从可见 不可见的域和解释以下的能力,在口语问答。我们还探讨了从自监督语音编码器的不同层中提取的各种类型的DSU,以及Mel频率倒谱系数(MFCC)。我们的研究结果表明,ASR任务和数据集是不是至关重要的提示调整口语问答任务。摘要:The integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks. This integration requires the use of a speech encoder, a speech adapter, and an LLM, trained on diverse tasks. We propose the use of discrete speech units (DSU), rather than continuous-valued speech encoder outputs, that are converted to the LLM token embedding space using the speech adapter. We generate DSU using a self-supervised speech encoder followed by k-means clustering. The proposed model shows robust performance on speech inputs from seen unseen domains and instruction-following capability in spoken question answering. We also explore various types of DSU extracted from different layers of the self-supervised speech encoder, as well as Mel frequency Cepstral Coefficients (MFCC). Our findings suggest that the ASR task and datasets are not crucial in instruction-tuning for spoken question answering tasks.

【2】 PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance
标题: PianoMotion10 M:钢琴演奏中手部运动生成的数据集和基准
作者:Qijun Gan,Song Wang,Shengtao Wu,Jianke Zhu
备注:Codes and Dataset: this https URL
链接:点击下载PDF文件
摘要:近年来,人工智能技术在教育领域受到越来越多的关注,但设计有效的乐器教学系统仍然是一个悬而未决的问题。虽然按键可以直接从乐谱中得到,但按键之间的过渡动作在钢琴演奏中需要更广泛的指导。在这项工作中,我们构建了一个钢琴手运动生成基准,以指导钢琴演奏的手运动和指法。为此,我们收集了一个带注释的数据集PianoMotion10M,其中包含116小时的钢琴演奏视频,这些视频来自鸟瞰图,其中包含1000万个带注释的手部姿势。我们还引入了一个强大的基线模型,通过位置预测器和位置引导的手势生成器从钢琴音频生成手部动作。此外,设计了一系列的评估指标来评估基线模型的性能,包括运动相似性,平滑性,左右手的位置准确性,以及运动分布的整体保真度。尽管钢琴按键与乐谱或音频已经可以访问,PianoMotion10M旨在提供指导钢琴指法的教学目的。数据集和源代码可以在https: agnjason.github.io PianoMotion-page上访问。摘要:Recently, artificial intelligence techniques for education have been received increasing attentions, while it still remains an open problem to design the effective music instrument instructing systems. Although key presses can be directly derived from sheet music, the transitional movements among key presses require more extensive guidance in piano performance. In this work, we construct a piano-hand motion generation benchmark to guide hand movements and fingerings for piano playing. To this end, we collect an annotated dataset, PianoMotion10M, consisting of 116 hours of piano playing videos from a bird's-eye view with 10 million annotated hand poses. We also introduce a powerful baseline model that generates hand motions from piano audios through a position predictor and a position-guided gesture generator. Furthermore, a series of evaluation metrics are designed to assess the performance of the baseline model, including motion similarity, smoothness, positional accuracy of left and right hands, and overall fidelity of movement distribution. Despite that piano key presses with respect to music scores or audios are already accessible, PianoMotion10M aims to provide guidance on piano fingering for instruction purposes. The dataset and source code can be accessed at https: agnjason.github.io PianoMotion-page.

【3】 On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
标题: 论异类数据源对语音转文本基础模型的影响
作者:Jinchuan Tian,Yifan Peng,William Chen,Kwanghee Choi,Karen Livescu,Shinji Watanabe
链接:点击下载PDF文件
摘要:开放式耳语风格语音模型(OWSM)系列旨在实现构建高级语音到文本(S2T)基础模型的完全透明性。为此,OWSM模型在25个公共语音数据集上进行了训练,这些数据集在多种方面都是异构的。在这项研究中,我们通过引入OWSM v3.2来推进OWSM系列,该系列通过调查和解决这种数据异质性的影响来改进先前的模型。我们的研究从每个数据集的详细分析开始,从中我们得出两个关键策略:使用代理任务进行数据过滤以提高数据质量,以及使用开放式大型语言模型(LLM)合并标点符号和真大小写。在所有其他配置保持不变的情况下,OWSM v3.2比OWSM v3.1基准提高了性能,同时使用的训练数据减少了15%。摘要:The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are heterogeneous in multiple ways. In this study, we advance the OWSM series by introducing OWSM v3.2, which improves on prior models by investigating and addressing the impacts of this data heterogeneity. Our study begins with a detailed analysis of each dataset, from which we derive two key strategies: data filtering with proxy task to enhance data quality, and the incorporation of punctuation and true-casing using an open large language model (LLM). With all other configurations staying the same, OWSM v3.2 improves performance over the OWSM v3.1 baseline while using 15% less training data.

【4】 Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
标题: SYS 2 Sound:从以自我为中心的视频中产生具有环境意识的动作声音
作者:Changan Chen,Puyuan Peng,Ami Baid,Zihui Xue,Wei-Ning Hsu,David Harwarth,Kristen Grauman
备注:Project page: this https URL
链接:点击下载PDF文件
摘要:为人类交互生成逼真的音频对于许多应用都很重要,例如为电影或虚拟现实游戏创建声音效果。现有的方法隐含地假设在训练过程中视频和音频之间完全对应,但许多声音发生在屏幕外,与视觉效果的对应性很弱甚至没有对应性-导致在测试时出现不受控制的环境声音或幻觉。我们提出了一种新的环境感知音频生成模型,AV-LDM。我们设计了一种新的音频调节机制,以学习在野外训练视频中将前景动作声音与环境背景声音分开。给定一个新颖的无声视频,我们的模型使用检索增强生成来创建在语义和时间上都与视觉内容相匹配的音频。我们训练和评估我们的模型在两个在野外自我中心的视频数据集Ego 4D和EPIC-KITCHENS。我们的模型优于一系列现有的方法,允许可控生成的环境声音,甚至显示出推广到计算机图形游戏剪辑的承诺。总的来说,我们的工作是第一个将视频到音频的生成忠实地集中在所观察到的视觉内容上,尽管训练来自具有自然背景声音的未经策划的剪辑。摘要:Generating realistic audio for human interactions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio during training, yet many sounds happen off-screen and have weak to no correspondence with the visuals -- resulting in uncontrolled ambient sounds or hallucinations at test time. We propose a novel ambient-aware audio generation model, AV-LDM. We devise a novel audio-conditioning mechanism to learn to disentangle foreground action sounds from the ambient background sounds in in-the-wild training videos. Given a novel silent video, our model uses retrieval-augmented generation to create audio that matches the visual content both semantically and temporally. We train and evaluate our model on two in-the-wild egocentric video datasets Ego4D and EPIC-KITCHENS. Our model outperforms an array of existing methods, allows controllable generation of the ambient sound, and even shows promise for generalizing to computer graphics game clips. Overall, our work is the first to focus video-to-audio generation faithfully on the observed visual content despite training from uncurated clips with natural background sounds.

【5】 Vision Transformer Segmentation for Visual Bird Sound Denoising
标题: 视觉Transformer分割用于视觉鸟声去噪
作者:Sahil Kumar,Jialu Li,Youshan Zhang
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:由于持续的残留噪声,音频去噪,特别是在鸟鸣的背景下,仍然是一项具有挑战性的任务。传统和深度学习方法经常与人为或低频噪声作斗争。在这项工作中,我们提出了ViTVS,一种新的方法,利用权力的Vision Transformer(ViT)架构。ViTVS巧妙地结合了分割技术,从复杂的信号混合中分离出干净的音频。我们的主要贡献包括ViTVS的开发,引入了全面的,长距离的和多尺度的表示。这些贡献直接解决了传统方法固有的局限性。大量的实验表明,ViTVS优于最先进的方法,将其定位为现实世界的鸟声去噪应用的基准解决方案。源代码可在https: github.com aiai-4 ViVTS上获得。摘要:Audio denoising, especially in the context of bird sounds, remains a challenging task due to persistent residual noise. Traditional and deep learning methods often struggle with artificial or low-frequency noise. In this work, we propose ViTVS, a novel approach that leverages the power of the vision transformer (ViT) architecture. ViTVS adeptly combines segmentation techniques to disentangle clean audio from complex signal mixtures. Our key contributions encompass the development of ViTVS, introducing comprehensive, long-range, and multi-scale representations. These contributions directly tackle the limitations inherent in conventional approaches. Extensive experiments demonstrate that ViTVS outperforms state-of-the-art methods, positioning it as a benchmark solution for real-world bird sound denoising applications. Source code is available at: https: github.com aiai-4 ViVTS.

【6】 Complex Image-Generative Diffusion Transformer for Audio Denoising
标题: 用于音频去噪的复杂图像生成扩散Transformer
作者:Junhui Li,Pu Wang,Jialu Li,Youshan Zhang
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:音频去噪技术在深度神经网络领域引起了广泛的关注。最近,音频去噪问题已被转换为图像生成任务,基于深度学习的方法已被应用于解决这个问题。但其表现仍然有限,留有进一步提升的空间。为了提高音频去噪性能,本文介绍了一种复图像生成扩散Transformer,从复傅立叶域捕获更多的信息。通过将Transformer与扩散模型相结合,提出了一种新型的扩散Transformer.我们提出的模型证明了Transformer的可扩展性,并使用注意扩散扩展稀疏注意的感受野。我们的工作是第一个利用扩散Transformers来处理音频去噪的图像生成任务。在两个基准数据集上的大量实验表明,我们提出的模型优于最先进的方法。摘要:The audio denoising technique has captured widespread attention in the deep neural network field. Recently, the audio denoising problem has been converted into an image generation task, and deep learning-based approaches have been applied to tackle this problem. However, its performance is still limited, leaving room for further improvement. In order to enhance audio denoising performance, this paper introduces a complex image-generative diffusion transformer that captures more information from the complex Fourier domain. We explore a novel diffusion transformer by integrating the transformer with a diffusion model. Our proposed model demonstrates the scalability of the transformer and expands the receptive field of sparse attention using attention diffusion. Our work is among the first to utilize diffusion transformers to deal with the image generation task for audio denoising. Extensive experiments on two benchmark datasets demonstrate that our proposed model outperforms state-of-the-art methods.

【7】 Towards Multilingual Audio-Visual Question Answering
标题: 走向多语言视听问答
作者:Orchid Chetia Phukan,Priyabrata Mallick,Swarup Ranjan Behera,Aalekhya Satya Narayani,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们努力扩展视听问题问答(AVQA)的多语言设置。现有的AVQA研究主要围绕英语进行,复制它以解决其他语言的AVQA需要大量的资源分配。作为一个可扩展的解决方案,我们利用机器翻译,并提出了两个多语言的AVQA数据集,从现有的基准AVQA数据集创建的八种语言。这可以防止人工手动收集问题和答案的额外人工注释工作。为此,我们建议,MERA框架,利用国家的最先进的(SOTA)的视频,音频和文本基础模型AVQA在多种语言。我们引入了一套模型,即MERA-L,MERA-C,MERA-T,具有不同的模型架构,以基准测试所提出的数据集。我们相信我们的工作将开辟新的研究方向,并作为未来多语言AVQA工作的参考基准。摘要:In this paper, we work towards extending Audio-Visual Question Answering (AVQA) to multilingual settings. Existing AVQA research has predominantly revolved around English and replicating it for addressing AVQA in other languages requires a substantial allocation of resources. As a scalable solution, we leverage machine translation and present two multilingual AVQA datasets for eight languages created from existing benchmark AVQA datasets. This prevents extra human annotation efforts of collecting questions and answers manually. To this end, we propose, MERA framework, by leveraging state-of-the-art (SOTA) video, audio, and textual foundation models for AVQA in multiple languages. We introduce a suite of models namely MERA-L, MERA-C, MERA-T with varied model architectures to benchmark the proposed datasets. We believe our work will open new research directions and act as a reference benchmark for future works in multilingual AVQA.

【8】 Diffusion Gaussian Mixture Audio Denoise
标题: 扩散高斯混合音频降噪
作者:Pu Wang,Junhui Li,Jialu Li,Liangdong Guo,Youshan Zhang
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最近的扩散模型在音频去噪任务中取得了很好的性能。逆过程的独特性质可以恢复干净的信号。然而,现实世界中的噪声分布并不符合单一的高斯分布,甚至是未知的。高斯噪声条件下的采样限制了其应用场景。为了克服这些挑战,我们提出了一个DiffGMM模型,一个基于扩散和高斯混合模型的去噪模型。我们采用逆过程来估计高斯混合模型的参数。给定一个带噪的音频信号,我们首先应用一维U-Net来提取特征,并训练线性层来估计高斯混合模型的参数,然后我们近似真实的噪声分布。从估计的噪声中连续地减去噪声信号以输出干净的音频信号。大量的实验结果表明,所提出的DiffGMM模型达到了最先进的性能。摘要:Recent diffusion models have achieved promising performances in audio-denoising tasks. The unique property of the reverse process could recover clean signals. However, the distribution of real-world noises does not comply with a single Gaussian distribution and is even unknown. The sampling of Gaussian noise conditions limits its application scenarios. To overcome these challenges, we propose a DiffGMM model, a denoising model based on the diffusion and Gaussian mixture models. We employ the reverse process to estimate parameters for the Gaussian mixture model. Given a noisy audio signal, we first apply a 1D-U-Net to extract features and train linear layers to estimate parameters for the Gaussian mixture model, and we approximate the real noise distributions. The noisy signal is continuously subtracted from the estimated noise to output clean audio signals. Extensive experimental results demonstrate that the proposed DiffGMM model achieves state-of-the-art performance.

【9】 LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
标题: 激光:通过调整语音的自我监督表示来学习以改善内容相关任务
作者:Amit Meghanani,Thomas Hain
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:基于自监督学习(SSL)的语音模型被广泛用于全栈语音处理。然而,已经观察到,使用未标记的语音用于内容相关的任务来改进基于SSL的语音表示是具有挑战性的并且在计算上昂贵。最近的尝试已经解决这个问题,具有成本效益的自我监督微调(SSFT)的方法。继续在这个方向上,一个具有成本效益的SSFT方法命名为“激光:通过对齐自监督表示学习”。LASER基于具有时间正则化项的软DTW对准损失。实验进行了HuBERT和WavLM模型和评估SUPERB基准两个内容相关的任务:自动语音识别(ASR)和音素识别(PR)。对于ASR和PR任务,HuBERT的相对改进分别为3.7%和8.2%,WavLM的相对改进分别为4.1%和11.7%,在单个GPU上仅进行了< 3小时的微调。摘要:Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-related tasks is challenging and computationally expensive. Recent attempts have been made to address this issue with cost-effective self-supervised fine-tuning (SSFT) approaches. Continuing in this direction, a cost-effective SSFT method named "LASER: Learning by Aligning Self-supervised Representations" is presented. LASER is based on the soft-DTW alignment loss with temporal regularisation term. Experiments are conducted with HuBERT and WavLM models and evaluated on the SUPERB benchmark for two content-related tasks: automatic speech recognition (ASR) and phoneme recognition (PR). A relative improvement of 3.7% and 8.2% for HuBERT, and 4.1% and 11.7% for WavLM are observed, for the ASR and PR tasks respectively, with only < 3 hours of fine-tuning on a single GPU.

【10】 Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
标题: 探索多语言看不见的说话者情感识别:在多任务学习中利用共同注意线索
作者:Arnav Goel,Medha Hira,Anubha Gupta
备注:5 pages, Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:现代深度学习技术的出现带来了语音情感识别(SER)领域的进步。然而,该领域中流行的大多数系统未能推广到在训练期间未看到的扬声器。这项研究的重点是处理多语言SER的挑战,特别是看不见的扬声器。我们介绍CAMuLeNet,一种新的架构,利用基于共同注意力的融合和多任务学习来解决这个问题。此外,我们使用10倍leave-speaker-out交叉验证对Whisper,HuBERT,Wav2Vec2.0和WavLM的预训练编码器进行了基准测试,这些测试针对五个现有的多语言基准数据集:IEMOCAP,RAVDESS,CREMA-D,CARDB和CaFE,并在印地语(BhavVani)上发布了SER的新数据集。CAMuLeNet显示,在我们的交叉验证策略确定的未见过的扬声器的所有基准测试中,平均提高了约8%。摘要:Advent of modern deep learning techniques has given rise to advancements in the field of Speech Emotion Recognition (SER). However, most systems prevalent in the field fail to generalize to speakers not seen during training. This study focuses on handling challenges of multilingual SER, specifically on unseen speakers. We introduce CAMuLeNet, a novel architecture leveraging co-attention based fusion and multitask learning to address this problem. Additionally, we benchmark pretrained encoders of Whisper, HuBERT, Wav2Vec2.0, and WavLM using 10-fold leave-speaker-out cross-validation on five existing multilingual benchmark datasets: IEMOCAP, RAVDESS, CREMA-D, EmoDB and CaFE and, release a novel dataset for SER on the Hindi language (BhavVani). CAMuLeNet shows an average improvement of approximately 8% over all benchmarks on unseen speakers determined by our cross-validation strategy.

【11】 AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
标题: AV-GS:新视图声学合成的学习材料和几何感知先验
作者:Swapnil Bhosale,Haosen Yang,Diptesh Kanojia,Jiankang Deng,Xiatian Zhu
链接:点击下载PDF文件
摘要:新颖视图声学合成(NVAS)旨在在给定由3D场景处的声源发出的单声道音频的情况下在任何目标视点处渲染双耳音频。现有的方法已经提出了基于NeRF的隐式模型,以利用视觉线索作为合成双耳音频的条件。然而,除了源于繁重NeRF渲染的低效率之外,这些方法都具有表征整个场景环境的有限能力,例如房间几何形状,材料属性以及收听者和声源之间的空间关系。为了解决这些问题,我们提出了一种新的视听高斯飞溅(AV-GS)模型。为了获得音频合成的材料感知和几何感知条件,我们学习了一个显式的基于点的场景表示,在本地初始化的高斯点上具有音频指导参数,同时考虑到来自听众和声源的空间关系。为了使视觉场景模型音频自适应,我们提出了一种点致密化和修剪策略来最佳地分布高斯点,其中每个点在声音传播中的贡献(例如,无纹理的壁表面需要更多的点,因为它们影响声音路径转移)。大量的实验验证了我们的AV-GS优于现实世界RWAS和基于模拟的SoundSpaces数据集上的现有替代品。摘要:Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a condition for synthesizing binaural audio. However, in addition to low efficiency originating from heavy NeRF rendering, these methods all have a limited ability of characterizing the entire scene environment such as room geometry, material properties, and the spatial relation between the listener and sound source. To address these issues, we propose a novel Audio-Visual Gaussian Splatting (AV-GS) model. To obtain a material-aware and geometry-aware condition for audio synthesis, we learn an explicit point-based scene representation with an audio-guidance parameter on locally initialized Gaussian points, taking into account the space relation from the listener and sound source. To make the visual scene model audio adaptive, we propose a point densification and pruning strategy to optimally distribute the Gaussian points, with the per-point contribution in sound propagation (e.g., more points needed for texture-less wall surfaces as they affect sound path diversion). Extensive experiments validate the superiority of our AV-GS over existing alternatives on the real-world RWAS and simulation-based SoundSpaces datasets.

【12】 Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
标题: 用于噪音和回响多说话人自动语音识别的语音分离模型的免转录微调
作者:William Ravenscroft,George Close,Stefan Goetze,Thomas Hain,Mohammad Soleymanpour,Anurag Chowdhury,Mark C. Fuhs
备注:5 pages, 3 Figures, 3 Tables, Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:重叠说话者的自动语音识别(ASR)的一种解决方案是分离语音,然后对分离的信号执行ASR。通常,分离器会产生伪影,这通常会降低ASR性能。解决这个问题通常需要参考传输来联合训练分离和ASR网络。这对于在参考转录信息并不总是可用的真实世界域内音频上进行训练通常是不可行的。本文提出了一种只使用音频信号进行联合训练的免转录方法。所提出的方法使用预训练的ASR编码器的嵌入差异作为损失,并对称为引导PIT(GPIT)的排列不变训练(PIT)进行修改。该方法实现了6.4%的改善,字错误率(WER)的措施超过了信号电平的损失,也表现出增强改善感知措施,如短期客观可懂度(STOI)。摘要:One solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals. Commonly, the separator produces artefacts which often degrade ASR performance. Addressing this issue typically requires reference transcriptions to jointly train the separation and ASR networks. This is often not viable for training on real-world in-domain audio where reference transcript information is not always available. This paper proposes a transcription-free method for joint training using only audio signals. The proposed method uses embedding differences of pre-trained ASR encoders as a loss with a proposed modification to permutation invariant training (PIT) called guided PIT (GPIT). The method achieves a 6.4% improvement in word error rate (WER) measures over a signal-level loss and also shows enhancement improvements in perceptual measures such as short-time objective intelligibility (STOI).

【13】 SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
标题: SingOMD:从语音模型构建面向歌唱的多分辨率离散表示
作者:Yuxun Tang,Yuning Wu,Jiatong Shi,Qin Jin
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:离散表示在语音生成任务中已经显示出优势,其中离散令牌是通过将来自自监督学习(SSL)预训练模型的隐藏特征离散化来导出的。然而,语音SSL模型直接应用于歌唱生成遇到语音和歌唱之间的领域差距。此外,歌唱生成需要比典型语音更精细的表示。为了解决这些挑战,我们引入SingOMD,一种新的方法来提取面向唱歌的多分辨率离散表示语音SSL模型。具体来说,我们首先通过重新合成任务适应语音SSL的功能,并结合基于重新合成的多分辨率模块,以更好地服务于歌唱生成。这些适应多分辨率功能,然后通过聚类离散化。大量的实验表明,这些表示在歌唱声码器和歌唱声音合成的鲁棒性,效率和有效性。摘要:Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre-trained models. However, the direct application of speech SSL models to singing generation encounters domain gaps between speech and singing. Furthermore, singing generation necessitates a more refined representation than typical speech. To address these challenges, we introduce SingOMD, a novel method to extract singing-oriented multi-resolution discrete representations from speech SSL models. Specifically, we first adapt the features from speech SSL through a resynthesis task and incorporate multi-resolution modules based on resampling to better serve singing generation. These adapted multi-resolution features are then discretized via clustering. Extensive experiments demonstrate the robustness, efficiency, and effectiveness of these representations in singing vocoders and singing voice synthesis.

【14】 AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers
标题: AdaPTin:《Transformer》中产品双胞胎的低成本自适应压缩
作者:Emil Biju,Anirudh Sriram,Mert Pilanci
备注:12 pages, 3 figures, submitted to NeurIPS 2024
链接:点击下载PDF文件
摘要:虽然大型的基于变换的模型在说话人无关的语音识别中表现出了卓越的性能,但它们的大尺寸和计算要求使得它们在资源受限的环境中使用时昂贵或不切实际。在这项工作中,我们提出了一种低秩自适应压缩技术,称为AdaPTwin,共同压缩产品依赖的权重矩阵对的Transformer注意层。我们的方法可以优先考虑压缩模型的性能在一个特定的扬声器,同时保持推广到新的扬声器和声学条件。值得注意的是,我们的技术只需要8小时的语音数据进行微调,可以在20分钟内完成,与其他压缩方法相比,它具有很高的成本效益。我们通过将Whisper和Distil-Whisper模型压缩高达45%来证明我们的方法的有效性,同时单词错误率增加不到2%。摘要:While large transformer-based models have exhibited remarkable performance in speaker-independent speech recognition, their large size and computational requirements make them expensive or impractical to use in resource-constrained settings. In this work, we propose a low-rank adaptive compression technique called AdaPTwin that jointly compresses product-dependent pairs of weight matrices in the transformer attention layer. Our approach can prioritize the compressed model's performance on a specific speaker while maintaining generalizability to new speakers and acoustic conditions. Notably, our technique requires only 8 hours of speech data for fine-tuning, which can be accomplished in under 20 minutes, making it highly cost-effective compared to other compression methods. We demonstrate the efficacy of our approach by compressing the Whisper and Distil-Whisper models by up to 45% while incurring less than a 2% increase in word error rate.

【15】 A Single-Step Non-Autoregressive Automatic Speech Recognition Architecture with High Accuracy and Inference Speed
标题: 一种高准确度和推理速度的分步非自回归自动语音识别架构
作者:Ziyang Zhuang,Chenfeng Miao,Kun Zou,Shuai Gong,Ming Fang,Tao Wei,Zijian Li,Wei Hu,Shaojun Wang,Jing Xiao
链接:点击下载PDF文件
摘要:非自回归(NAR)自动语音识别(ASR)模型能够独立地、同时地预测语音标记,具有很高的推理速度。然而,与自回归(AR)模型相比,NAR模型的准确性仍然存在差距。为了进一步缩小NAR和AR模型之间的差距,我们提出了一种具有高精度和推理速度的单步NAR ASR架构,称为EfficientASR。它使用基于索引映射向量(IMV)的对齐生成器来在训练期间生成对齐,并使用对齐预测器来学习对齐以进行推理。它可以在交叉熵损失与对齐损失相结合的情况下进行端到端(E2 E)训练。与最先进的(SOTA)模型相比,所提出的EfficientASR在AISHELL-1和AISHELL-2基准测试中取得了有竞争力的结果。具体来说,它在AISHELL-1 dev test数据集上实现了4.26% 4.62%的字符错误率(CER),比SOTA AR Conformer的推理速度提高了约30倍。摘要:Non-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. To further narrow the gap between the NAR and AR models, we propose a single-step NAR ASR architecture with high accuracy and inference speed, called EfficientASR. It uses an Index Mapping Vector (IMV) based alignment generator to generate alignments during training, and an alignment predictor to learn the alignments for inference. It can be trained end-to-end (E2E) with cross-entropy loss combined with alignment loss. The proposed EfficientASR achieves competitive results on the AISHELL-1 and AISHELL-2 benchmarks compared to the state-of-the-art (SOTA) models. Specifically, it achieves character error rates (CER) of 4.26% 4.62% on the AISHELL-1 dev test dataset, which outperforms the SOTA AR Conformer with about 30x inference speedup.

【16】 Interpretable Temporal Class Activation Representation for Audio Spoofing Detection
标题: 用于音频欺骗检测的可解释时间类激活表示
作者:Menglu Li,Xiao-Ping Zhang
备注:10 pages, 5 figures, Accepted to Interspeech2024
链接:点击下载PDF文件
摘要:解释音频欺骗检测模型做出的决定对于培养对检测结果的信任至关重要。然而,目前对检测模型可解释性的研究仅限于将XAI工具应用于后训练模型。在本文中,我们利用wav 2 vec 2.0模型和专注的话语级功能直接集成到模型的架构的可解释性,从而提高决策过程的透明度。具体来说,我们提出了一个类激活表示本地化的歧视性帧有助于检测。此外,我们还证明了基于欺骗类型的多标签训练,而不是像善意和欺骗那样的二进制标签,使模型能够学习不同攻击的不同特征,从而显着提高检测性能。我们的模型实现了最先进的结果,ASVspoof 2019-LA集的EER为0.51%,最小t-DCF为0.0165。摘要:Explaining the decisions made by audio spoofing detection models is crucial for fostering trust in detection outcomes. However, current research on the interpretability of detection models is limited to applying XAI tools to post-trained models. In this paper, we utilize the wav2vec 2.0 model and attentive utterance-level features to integrate interpretability directly into the model's architecture, thereby enhancing transparency of the decision-making process. Specifically, we propose a class activation representation to localize the discriminative frames contributing to detection. Furthermore, we demonstrate that multi-label training based on spoofing types, rather than binary labels as bonafide and spoofed, enables the model to learn distinct characteristics of different attacks, significantly improving detection performance. Our model achieves state-of-the-art results, with an EER of 0.51% and a min t-DCF of 0.0165 on the ASVspoof2019-LA set.

【17】 Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
标题: 通过为预训练的多说话者文本转语音系统分配语音输入即兴生成说话者
作者:Zhengyang Chen,Xuechen Liu,Erica Cooper,Junichi Yamagishi,Yanmin Qian
备注:Accepted for presentation at Interspeech 2024 (with more analysis in the final Appendix part)
链接:点击下载PDF文件
摘要:本文提出了一种语音合成系统,允许用户指定和控制的扬声器的声学特性的提示描述合成语音的扬声器的特点。与以前的方法不同,我们的方法利用听众的印象来构建提示,这更容易收集和更自然地与日常描述的扬声器特质。我们采用低秩自适应(LoRA)技术来快速定制预训练的语言模型,以满足我们的需求,便于从提示文本中提取说话者相关特征。此外,与其他语音驱动的文本到语音(TTS)系统不同,我们将语音到说话人模块从多说话人TTS系统中分离出来,增强了系统的灵活性和与各种预训练的多说话人TTS系统的兼容性。此外,对于说话人-说话人特征模块,我们还比较了判别方法和基于流匹配的生成方法,我们发现,结合这两种方法可以帮助系统同时捕获说话人相关的信息,从提示更好地生成语音具有更高的保真度。摘要:This paper proposes a speech synthesis system that allows users to specify and control the acoustic characteristics of a speaker by means of prompts describing the speaker's traits of synthesized speech. Unlike previous approaches, our method utilizes listener impressions to construct prompts, which are easier to collect and align more naturally with everyday descriptions of speaker traits. We adopt the Low-rank Adaptation (LoRA) technique to swiftly tailor a pre-trained language model to our needs, facilitating the extraction of speaker-related traits from the prompt text. Besides, different from other prompt-driven text-to-speech (TTS) systems, we separate the prompt-to-speaker module from the multi-speaker TTS system, enhancing system flexibility and compatibility with various pre-trained multi-speaker TTS systems. Moreover, for the prompt-to-speaker characteristic module, we also compared the discriminative method and flow-matching based generative method and we found that combining both methods can help the system simultaneously capture speaker-related information from prompts better and generate speech with higher fidelity.

【18】 Are we there yet? A brief survey of Music Emotion Prediction Datasets, Models and Outstanding Challenges
标题: 我们到了吗?音乐情感预测数据集、模型和突出挑战的简要调查
作者:Jaeyong Kang,Dorien Herremans
链接:点击下载PDF文件
摘要:在过去的几年里,音乐的深度学习模型取得了巨大的进步。但是,如今机器学习模型在捕捉情感方面有多好,研究人员面临着哪些挑战?在本文中,我们提供了一个全面的概述,现有的音乐情感数据集,并讨论了评估标准以及在该领域的竞争。我们还简要概述了多年来建立的各种类型的音乐情感预测模型,为该领域的各种方法提供了见解。通过这次考试,我们强调坚持准确捕捉音乐中的情感的挑战。认识到这一领域的动态性质,我们用一个附带的GitHub存储库补充了我们的发现。该存储库包含音乐情感数据集和最新预测模型的全面列表。摘要:Deep learning models for music have advanced drastically in the last few years. But how good are machine learning models at capturing emotion these days and what challenges are researchers facing? In this paper, we provide a comprehensive overview of the available music-emotion datasets and discuss evaluation standards as well as competitions in the field. We also provide a brief overview of various types of music emotion prediction models that have been built over the years, offering insights into the diverse approaches within the field. Through this examination, we highlight the challenges that persist in accurately capturing emotion in music. Recognizing the dynamic nature of this field, we have complemented our findings with an accompanying GitHub repository. This repository contains a comprehensive list of music emotion datasets and recent predictive models.

【19】 Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
标题: 来自生成基础模型的合成音频可以帮助音频识别和语音建模吗?
作者:Tiantian Feng,Dimitrios Dimitriadis,Shrikanth Narayanan
备注:Accepted to 2024 INTERSPEECH
链接:点击下载PDF文件
摘要:基础模型的最新进展使音频生成模型能够产生与音乐,事件和人类行为相关的高保真声音。尽管现代音频生成模型取得了成功,但评估音频生成质量的传统方法在很大程度上依赖于距离度量,如Frechet音频距离。相比之下,我们的目标是通过检查使用它们作为训练数据的有效性来评估音频生成的质量。具体来说,我们进行研究,探索使用合成音频的音频识别。此外,我们研究合成音频是否可以作为语音相关建模中的数据增强资源。我们的综合实验证明了使用合成音频进行音频识别和语音相关建模的潜力。我们的代码可以在https: github.com usc-sail SynthAudio上找到。摘要:Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional approach to assessing the quality of the audio generation relies heavily on distance metrics like Frechet Audio Distance. In contrast, we aim to evaluate the quality of audio generation by examining the effectiveness of using them as training data. Specifically, we conduct studies to explore the use of synthetic audio for audio recognition. Moreover, we investigate whether synthetic audio can serve as a resource for data augmentation in speech-related modeling. Our comprehensive experiments demonstrate the potential of using synthetic audio for audio recognition and speech-related modeling. Our code is available at https: github.com usc-sail SynthAudio.

【20】 MFF-EINV2: Multi-scale Feature Fusion across Spectral-Spatial-Temporal Domains for Sound Event Localization and Detection
标题: MFF-EINV 2:跨频谱-空间-时间域的多尺度特征融合,用于声音事件定位和检测
作者:Da Mu,Zhicheng Zhang,Haobo Yue
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:声音事件定位和检测(SELD)涉及使用多声道声音记录来检测和定位声音事件。先前提出的事件无关网络V2(EINV 2)在SELD上取得了出色的性能。然而,它仍然面临着有效地提取跨光谱,空间和时间域的特征的挑战。本文提出了一个三阶段的网络结构命名为多尺度特征融合(MFF)模块,以充分提取跨光谱,空间和时间域的多尺度特征。MFF模块采用并行子网络结构生成多尺度光谱和空间特征。TF-卷积模块用于提供多尺度时间特征。我们将MFF合并到EINV 2中,并将所提出的方法称为MFF-EINV 2。在2022年和2023年DCASE挑战任务3数据集上的实验结果表明了我们的MFF-EINV 2的有效性,与已发表的方法相比,它实现了最先进的(SOTA)性能。摘要:Sound Event Localization and Detection (SELD) involves detecting and localizing sound events using multichannel sound recordings. Previously proposed Event-Independent Network V2 (EINV2) has achieved outstanding performance on SELD. However, it still faces challenges in effectively extracting features across spectral, spatial, and temporal domains. This paper proposes a three-stage network structure named Multi-scale Feature Fusion (MFF) module to fully extract multi-scale features across spectral, spatial, and temporal domains. The MFF module utilizes parallel subnetworks architecture to generate multi-scale spectral and spatial features. The TF-Convolution Module is employed to provide multi-scale temporal features. We incorporated MFF into EINV2 and term the proposed method as MFF-EINV2. Experimental results in 2022 and 2023 DCASE challenge task3 datasets show the effectiveness of our MFF-EINV2, which achieves state-of-the-art (SOTA) performance compared to published methods.

【21】 VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
标题: Visinger 2+:通过自我监督学习表示增强的端到端歌唱声音合成
作者:Yifeng Yu,Jiatong Shi,Yuning Wu,Shinji Watanabe
备注:4 pages, 2 figures
链接:点击下载PDF文件
摘要:随着深度学习技术的出现,歌唱语音合成(SVS)取得了重大进展。然而,SVS中的一个重大挑战是标记的歌唱声音数据的稀缺性,这限制了监督学习方法的有效性。为了应对这一挑战,本文介绍了一种新的方法,通过利用来自预训练的自监督学习模型的未标记数据来提高SVS的质量。在现有VISinger2框架的基础上,本研究将额外的光谱特征信息集成到系统中,以提高其性能。该集成旨在利用来自预训练模型的丰富声学特征,从而丰富合成并产生更自然和更具表现力的歌声。在不同语料库中的实验结果表明,该方法在提高合成歌声的客观和主观指标的整体质量的有效性。摘要:Singing Voice Synthesis (SVS) has witnessed significant advancements with the advent of deep learning techniques. However, a significant challenge in SVS is the scarcity of labeled singing voice data, which limits the effectiveness of supervised learning methods. In response to this challenge, this paper introduces a novel approach to enhance the quality of SVS by leveraging unlabeled data from pre-trained self-supervised learning models. Building upon the existing VISinger2 framework, this study integrates additional spectral feature information into the system to enhance its performance. The integration aims to harness the rich acoustic features from the pre-trained models, thereby enriching the synthesis and yielding a more natural and expressive singing voice. Experimental results in various corpora demonstrate the efficacy of this approach in improving the overall quality of synthesized singing voices in both objective and subjective metrics.

【22】 TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
标题: TSE-PI:具有音调信息的回响环境下的目标声音提取
作者:Yiwen Wang,Xihong Wu
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:目标声音提取(TSE)根据提供的线索从混合信号中分离出目标声音。然而,现有模型的性能显着降低混响条件下。受听觉场景分析(ASA)的启发,本文提出了一种具有基音信息的TSE模型TSE-PI。条件音高提取是通过具有声音类别标签的逐行线性调制层来实现的。一个修改后的波形模型结合音高信息,采用一个可学习的伽玛通滤波器组的卷积编码器的地方,用于目标声音提取。包含音高信息的目的是提高模型的性能。在FSD 50 K数据集上的实验结果表明,当结合音高信息和伽马通滤波器组时,混响环境下的目标声音提取提高了2.4 dB。摘要:Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis (ASA), this work proposes a TSE model provided with pitch information named TSE-PI. Conditional pitch extraction is achieved through the Feature-wise Linearly Modulated layer with the sound-class label. A modified Waveformer model combined with pitch information, employing a learnable Gammatone filterbank in place of the convolutional encoder, is used for target sound extraction. The inclusion of pitch information is aimed at improving the model's performance. The experimental results on the FSD50K dataset illustrate 2.4 dB improvements of target sound extraction under reverberant environments when incorporating pitch information and Gammatone filterbank.

【23】 ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
标题: ML-SURB 2.0:跨建模约束、语言和数据集的多语言语音模型基准测试
作者:Jiatong Shi,Shih-Heng Wang,William Chen,Martijn Bartelds,Vanya Bannihatti Kumar,Jinchuan Tian,Xuankai Chang,Dan Jurafsky,Karen Livescu,Hung-yi Lee,Shinji Watanabe
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:ML-SUPERB评估语言识别和自动语音识别(ASR)任务的自监督学习(SSL)模型。该基准测试将模型视为特征提取器,并使用单个浅下游模型,该模型可以针对下游任务进行微调。然而,现实世界的用例可能需要不同的配置。本文介绍了ML-SUPERB~2.0,这是一个新的基准,用于跨下游模型,微调设置和有效的模型自适应方法评估预训练的SSL和监督语音模型。我们发现ML-SUPERB的设置性能有所改善。然而,性能取决于下游模型设计。此外,我们发现语言和数据集之间存在很大的性能差异,这表明需要更有针对性的方法来提高多语言ASR性能。摘要:ML-SUPERB evaluates self-supervised learning (SSL) models on the tasks of language identification and automatic speech recognition (ASR). This benchmark treats the models as feature extractors and uses a single shallow downstream model, which can be fine-tuned for a downstream task. However, real-world use cases may require different configurations. This paper presents ML-SUPERB~2.0, which is a new benchmark for evaluating pre-trained SSL and supervised speech models across downstream models, fine-tuning setups, and efficient model adaptation approaches. We find performance improvements over the setup of ML-SUPERB. However, performance depends on the downstream model design. Also, we find large performance differences between languages and datasets, suggesting the need for more targeted approaches to improve multilingual ASR performance.

【24】 Emotion Manipulation Through Music -- A Deep Learning Interactive Visual Approach
标题: 通过音乐操纵情绪--深度学习交互视觉方法
作者:Adel N. Abdalla,Jared Osborne,Razvan Andonie
链接:点击下载PDF文件
摘要:音乐能唤起许多人的感情。我们介绍了一种使用AI工具操纵歌曲情感内容的新方法。我们的目标是达到所需的情感,同时保持原始旋律尽可能完整。为此,我们创建了一个交互式管道,能够将输入的歌曲转换为完全相反的情感,并通过Russel的Circumplex模型将此结果可视化。我们的方法是音乐语义操纵的概念验证,这是一个旨在修改现有音乐情感内容的新领域。我们设计了一个深度学习模型,能够评估我们对键、SoundFont乐器和其他音乐功能的修改的准确性。我们模型的准确性与4Q Emotion数据集上的最新技术水平一致。随着进一步的完善,这项研究可能有助于按需定制音乐的生成,现有作品的自动混音,以及为情感发展调整的音乐播放列表。摘要:Music evokes emotion in many people. We introduce a novel way to manipulate the emotional content of a song using AI tools. Our goal is to achieve the desired emotion while leaving the original melody as intact as possible. For this, we create an interactive pipeline capable of shifting an input song into a diametrically opposed emotion and visualize this result through Russel's Circumplex model. Our approach is a proof-of-concept for Semantic Manipulation of Music, a novel field aimed at modifying the emotional content of existing music. We design a deep learning model able to assess the accuracy of our modifications to key, SoundFont instrumentation, and other musical features. The accuracy of our model is in-line with the current state of the art techniques on the 4Q Emotion dataset. With further refinement, this research may contribute to on-demand custom music generation, the automated remixing of existing work, and music playlists tuned for emotional progression.

【25】 Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
标题: 通过文本到发音障碍语音合成增强训练数据用于发音障碍自动语音识别
作者:Wing-Zin Leung,Mattias Cross,Anton Ragni,Stefan Goetze
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,自动语音识别(ASR)的研究取得了令人印象深刻的成绩,并在增强和替代通信(AAC)和家庭环境系统中为构音障碍(PwD)患者提供接入方面具有巨大的潜力。然而,构音障碍性ASR(DASR)的进展受到构音障碍性言语的高度变异性和构音障碍训练数据的有限公共可用性的限制。本文证明了使用文本到构音障碍语音(TTDS)合成来微调大型ASR模型的数据增强对DASR是有效的。具体来说,基于扩散的文本到语音(TTS)模型可以产生类似于构音障碍语音的语音样本,这些语音样本可以用作微调ASR基础模型的额外训练数据,在这种情况下是Whisper。结果表明,改进的合成指标和ASR性能的建议多扬声器扩散为基础的TTDS数据增强ASR微调相比,目前的DASR基线。摘要:Automatic speech recognition (ASR) research has achieved impressive performance in recent years and has significant potential for enabling access for people with dysarthria (PwD) in augmentative and alternative communication (AAC) and home environment systems. However, progress in dysarthric ASR (DASR) has been limited by high variability in dysarthric speech and limited public availability of dysarthric training data. This paper demonstrates that data augmentation using text-to-dysarthic-speech (TTDS) synthesis for finetuning large ASR models is effective for DASR. Specifically, diffusion-based text-to-speech (TTS) models can produce speech samples similar to dysarthric speech that can be used as additional training data for fine-tuning ASR foundation models, in this case Whisper. Results show improved synthesis metrics and ASR performance for the proposed multi-speaker diffusion-based TTDS data augmentation for ASR fine-tuning compared to current DASR baselines.

【26】 Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
标题: 探索多语言广播和机构语音自动转录的口语识别策略
作者:Martina Valente,Fabio Brugnara,Giovanni Morrone,Enrico Zovato,Leonardo Badino
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文讨论了口语识别(SLI)和语音识别的多语种广播和机构的讲话,真正的应用场景,很少在SLI文献中得到解决。观察到在这些领域的语言变化大多与说话人的变化,我们提出了一个级联系统,包括说话人日记和语言识别,并比较它与更传统的语言识别和语言日记系统。实验结果表明,该系统可以实现较低的语言分类和语言日志化错误率(相对语言日志错误减少高达10%,相对语言混淆减少60%),并导致多语言测试集的WER较低(相对WER减少8%以上),而同时不会对单语音频上的语音识别产生负面影响(相对于语音识别,绝对WER增加在0.1%和0.7%之间。单语ASR)。摘要:This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Observing that in these domains language changes are mostly associated with speaker changes, we propose a cascaded system consisting of speaker diarization and language identification and compare it with more traditional language identification and language diarization systems. Results show that the proposed system often achieves lower language classification and language diarization error rates (up to 10% relative language diarization error reduction and 60% relative language confusion reduction) and leads to lower WERs on multilingual test sets (more than 8% relative WER reduction), while at the same time does not negatively affect speech recognition on monolingual audio (with an absolute WER increase between 0.1% and 0.7% w.r.t. monolingual ASR).

【27】 FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
标题: FlowAVSE:具有条件流匹配的高效视听语音增强
作者:Chaeyoung Jung,Suyeon Lee,Ji-Hoon Kim,Joon Son Chung
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:这项工作提出了一种有效的方法来提高质量的损坏的语音信号,利用声学和视觉线索。虽然现有的基于扩散的方法已经证明了显着的质量,其适用性受到缓慢的推理速度和计算复杂性的限制。为了解决这个问题,我们提出了FlowAVSE,它提高了推理速度,减少了可学习参数的数量,而不会降低输出质量。特别是,我们采用了一个条件流匹配算法,使高质量的语音在一个单一的采样步骤的生成。此外,我们通过优化基于扩散的系统的底层U-网架构来提高效率。我们的实验表明,FlowAVSE实现了22倍的推理速度,并减少了一半的模型大小,同时保持输出质量。演示页面可在https: cyongong.github.io FlowAVSE.github.io 上找到摘要:This work proposes an efficient method to enhance the quality of corrupted speech signals by leveraging both acoustic and visual cues. While existing diffusion-based approaches have demonstrated remarkable quality, their applicability is limited by slow inference speeds and computational complexity. To address this issue, we present FlowAVSE which enhances the inference speed and reduces the number of learnable parameters without degrading the output quality. In particular, we employ a conditional flow matching algorithm that enables the generation of high-quality speech in a single sampling step. Moreover, we increase efficiency by optimizing the underlying U-net architecture of diffusion-based systems. Our experiments demonstrate that FlowAVSE achieves 22 times faster inference speed and reduces the model size by half while maintaining the output quality. The demo page is available at: https: cyongong.github.io FlowAVSE.github.io

【28】 ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
标题: ToneUnit:音调语言语音合成的语音离散化方法
作者:Dehua Tao,Daxin Tan,Yu Ting Yeung,Xiao Chen,Tan Lee
链接:点击下载PDF文件
摘要:将语音表示为离散单元在支持下游口语处理任务方面具有许多好处。然而,这种方法在汉语普通话等有声调语言的语音合成中的研究较少。我们在汉语语音合成上的初步实验揭示了“声调偏移”的问题,即合成的语音话语包含正确的基本音节,但不正确的声调。为了解决这个问题,我们提出了ToneUnit框架,该框架利用带有音调标签的注释数据作为CTC监督来学习汉语普通话语音的音调感知离散语音单元。我们的研究结果表明,通过TonUnit获得的离散单元解决了汉语合成语音中的“变调”问题,并在英语合成中取得了良好的效果。此外,实验结果表明,有限标量量化增强了ToneUnit的有效性。值得注意的是,ToneUnit即使使用最少的注释数据也可以有效地工作。摘要:Representing speech as discretized units has numerous benefits in supporting downstream spoken language processing tasks. However, the approach has been less explored in speech synthesis of tonal languages like Mandarin Chinese. Our preliminary experiments on Chinese speech synthesis reveal the issue of "tone shift", where a synthesized speech utterance contains correct base syllables but incorrect tones. To address the issue, we propose the ToneUnit framework, which leverages annotated data with tone labels as CTC supervision to learn tone-aware discrete speech units for Mandarin Chinese speech. Our findings indicate that the discrete units acquired through the TonUnit resolve the "tone shift" issue in synthesized Chinese speech and yield favorable results in English synthesis. Moreover, the experimental results suggest that finite scalar quantization enhances the effectiveness of ToneUnit. Notably, ToneUnit can work effectively even with minimal annotated data.

【29】 Cascaded noise reduction and acoustic echo cancellation based on an extended noise reduction
标题: 基于扩展降噪的级联降噪和声回声消除
作者:Arnout Roebben,Toon van Waterschoot,Marc Moonen
备注:Accepted for publication in EUSIPCO 2024
链接:点击下载PDF文件
摘要:在许多语音记录应用中,所记录的期望语音被噪声和声学回声破坏,使得需要组合降噪(NR)和声学回声消除(AEC)。常见的级联设计对应于AEC滤波器之前的NR个滤波器。这些NR滤波器旨在降低近端房间噪声(以及可能部分回声),并且仅对麦克风进行操作,因此需要AEC滤波器对回声路径和NR滤波器进行建模。然而,在本文中,我们提出了一种设计扩展NR(NRext)过滤器前AEC过滤器的回声路径是添加剂的地图的假设下,从而保留了加法运算。在这里,NRext滤波器旨在减少回声中的近端房间噪声和远端房间噪声分量,并且在麦克风和扬声器上操作。我们表明,成功的AEC过滤器显着成为独立的NRext过滤器,这样的AEC过滤器只需要模拟的回声路径,提高AEC性能。此外,NRext滤波器中的自由度随着扬声器的数量而缩放,这不是NR滤波器的情况,从而导致改善的NR性能。摘要:In many speech recording applications, the recorded desired speech is corrupted by both noise and acoustic echo, such that combined noise reduction (NR) and acoustic echo cancellation (AEC) is called for. A common cascaded design corresponds to NR filters preceding AEC filters. These NR filters aim at reducing the near-end room noise (and possibly partially the echo) and operate on the microphones only, consequently requiring the AEC filters to model both the echo paths and the NR filters. In this paper, however, we propose a design with extended NR (NRext) filters preceding AEC filters under the assumption of the echo paths being additive maps, thus preserving the addition operation. Here, the NRext filters aim at reducing both the near-end room noise and the far-end room noise component in the echo, and operate on both the microphones and loudspeakers. We show that the succeeding AEC filters remarkably become independent of the NRext filters, such that the AEC filters are only required to model the echo paths, improving the AEC performance. Further, the degrees of freedom in the NRext filters scale with the number of loudspeakers, which is not the case for the NR filters, resulting in an improved NR performance.

【30】 Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
标题: 使用超声波麦克风阵列和CNN预测NC车床操作中的刀具磨损
作者:Jan Steckel,Arne Aerts,Erik Verreycken,Dennis Laurijssen,Walter Daems
链接:点击下载PDF文件
摘要:本文介绍了一种新的方法来预测刀具磨损数控车削操作,结合超声波麦克风阵列和卷积神经网络(CNN)。使用波束形成技术增强OkHz和60kHz之间的高频声发射以提高信噪比。然后通过CNN分析处理后的声学数据,预测切削刀具的剩余使用寿命(RUL)。通过对350个用单一硬质合金刀片加工的工件的数据进行训练,该模型可以准确地预测硬质合金刀片的RUL。我们的研究结果表明,通过将先进的超声波传感器与深度学习相结合,可以在CNC加工中实现精确的预测性维护任务。摘要:This paper introduces a novel method for predicting tool wear in CNC turning operations, combining ultrasonic microphone arrays and convolutional neural networks (CNNs). High-frequency acoustic emissions between 0 kHz and 60 kHz are enhanced using beamforming techniques to improve the signal- to-noise ratio. The processed acoustic data is then analyzed by a CNN, which predicts the Remaining Useful Life (RUL) of cutting tools. Trained on data from 350 workpieces machined with a single carbide insert, the model can accurately predict the RUL of the carbide insert. Our results demonstrate the potential gained by integrating advanced ultrasonic sensors with deep learning for accurate predictive maintenance tasks in CNC machining.

【31】 On Improving Error Resilience of Neural End-to-End Speech Coders
标题: 提高神经端到端语音编码器的抗错误能力
作者:Kishan Gupta,Nicola Pia,Srikanth Korse,Andreas Brendel,Guillaume Fuchs,Markus Multrus
链接:点击下载PDF文件
摘要:丢包隐藏(PLC)和前向纠错(FEC)等容错工具对于维护IP语音(VoIP)等应用的可靠语音通信至关重要,因为在这些应用中,数据包经常延迟和丢失。近年来,端到端神经语音编解码器由于其以低比特率传输语音信号的能力而显着增加,但很少考虑其在实际系统中的错误恢复能力。最近推出的神经端到端语音编解码器(NESC)可以在低比特率下再现高质量的自然语音。我们通过增加一个低复杂度的网络来预测潜在空间中的码本索引,从而扩展其对分组丢失的鲁棒性。此外,我们提出了一种方法来添加一个带内FEC在一个额外的比特率为0.8 kbps。主观和客观的评估表明,所提出的方法的有效性,并表明耦合PLC和FEC提供了显着的鲁棒性对数据包丢失。摘要:Error resilient tools like Packet Loss Concealment (PLC) and Forward Error Correction (FEC) are essential to maintain a reliable speech communication for applications like Voice over Internet Protocol (VoIP), where packets are frequently delayed and lost. In recent times, end-to-end neural speech codecs have seen a significant rise, due to their ability to transmit speech signal at low bitrates but few considerations were made about their error resilience in a real system. Recently introduced Neural End-to-End Speech Codec (NESC) can reproduce high quality natural speech at low bitrates. We extend its robustness to packet losses by adding a low complexity network to predict the codebook indices in latent space. Furthermore, we propose a method to add an in-band FEC at an additional bitrate of 0.8 kbps. Both subjective and objective assessment indicate the effectiveness of proposed methods, and demonstrate that coupling PLC and FEC provide significant robustness against packet losses.

【32】 DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
标题: DubWise:多模式基于LLM的文本转语音配音中的视频引导语音持续时间控制
作者:Neha Sahipjohn,Ashishkumar Gudmalwar,Nirmesh Shah,Pankaj Wasnik,Rajiv Ratn Shah
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:配音后的视听对齐是一个具有挑战性的研究问题。为此,我们提出了一种新的方法,基于DubWise多模态大语言模型(LLM)的文本到语音(TTS),它可以控制合成语音的语音持续时间,使其与参考视频中给出的扬声器嘴唇运动保持一致,即使口语文本不同或使用不同的语言。为了实现这一点,我们建议在预先训练的基于GPT的TTS中利用跨模态注意力技术。我们结合语言令牌从文本,扬声器身份令牌通过语音克隆网络,和视频令牌通过一个建议的持续时间控制器网络。我们证明了我们的系统在Lip 2 Wav-Chemistry和LRS 2数据集上的有效性。此外,与用于相同语言但不同文本的SOTA相比,所提出的方法实现了改进的唇同步和自然度(即,非平行)和不同的语言,不同的文本(即,跨语言)场景。摘要:Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech (TTS), which can control the speech duration of synthesized speech in such a way that it aligns well with the speakers lip movements given in the reference video even when the spoken text is different or in a different language. To accomplish this, we propose to utilize cross-modal attention techniques in a pre-trained GPT-based TTS. We combine linguistic tokens from text, speaker identity tokens via a voice cloning network, and video tokens via a proposed duration controller network. We demonstrate the effectiveness of our system on the Lip2Wav-Chemistry and LRS2 datasets. Also, the proposed method achieves improved lip sync and naturalness compared to the SOTAs for the same language but different text (i.e., non-parallel) and the different language, different text (i.e., cross-lingual) scenarios.

【33】 Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
标题: 从脑电信号实现完全端到端的听听语音解码
作者:Jihwan Lee,Aditya Kommineni,Tiantian Feng,Kleanthis Avramidis,Xuan Shi,Sudarsana Kadiri,Shrikanth Narayanan
备注:accepted to Interspeech2024
链接:点击下载PDF文件
摘要:从EEG信号中解码语音是一项具有挑战性的任务,其中对大脑活动进行建模以估计声学刺激的显著特征。我们提出了FESDE,一个新的框架,从脑电信号完全端到端的语音解码。我们的方法的目的是直接重建听到的语音波形给定的EEG信号,其中没有中间的声学特征处理步骤是必需的。所提出的方法包括EEG模块和语音模块以及连接器。EEG模块学习更好地表示EEG信号,而语音模块从模型表示生成语音波形。连接器学习桥接EEG和语音的潜在空间的分布。建议的框架是简单和有效的,通过允许单步推理,并优于以前的作品的客观指标。细粒度的音素分析进行揭示语音解码的模型特性。源代码可以在这里找到:github.com lee-jhwn fesde。摘要:Speech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli. We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals. Our approach aims to directly reconstruct listened speech waveforms given EEG signals, where no intermediate acoustic feature processing step is required. The proposed method consists of an EEG module and a speech module along with a connector. The EEG module learns to better represent EEG signals, while the speech module generates speech waveforms from model representations. The connector learns to bridge the distributions of the latent spaces of EEG and speech. The proposed framework is both simple and efficient, by allowing single-step inference, and outperforms prior works on objective metrics. A fine-grained phoneme analysis is conducted to unveil model characteristics of speech decoding. The source code is available here: github.com lee-jhwn fesde.

【34】 DB3V: A Dialect Dominated Dataset of Bird Vocalisation for Cross-corpus Bird Species Recognition
标题: DB 3V:一个基于方言的鸟类发声数据集,用于跨库鸟类物种识别
作者:Xin Jing,Luyang Zhang,Jiangjian Xie,Alexander Gebhard,Alice Baird,Bjoern Schuller
备注:accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在鸟类学中,已知鸟类物种有多种多样的方言,这一点得到了广泛的承认,即鸟类物种在不同地区的叫声中表现出不同的方言。因此,仅通过叫声识别鸟类的计算方法面临着重大挑战。有越来越多的兴趣,了解鸟类物种识别方法的有效性,物种特异性方言的影响。尽管通过扩大方言数据集可能会缓解这种情况,但目前缺乏公开的测试数据阻碍了强有力的基准测试工作。本文介绍了方言主导的鸟类发声数据集,第一个跨语料库数据集,重点是方言的鸟类发声。DB3V包括超过25小时的音频记录,来自10种鸟类,分布在美国毗连的三个不同地区。除了呈现数据集之外,我们还进行分析并建立跨语料库鸟类识别的基线模型。数据和代码可在网上公开获取:https: zenodo.org records 11544734摘要:In ornithology, bird species are known to have variedit's widely acknowledged that bird species display diverse dialects in their calls across different regions. Consequently, computational methods to identify bird species onsolely through their calls face critsignificalnt challenges. There is growing interest in understanding the impact of species-specific dialects on the effectiveness of bird species recognition methods. Despite potential mitigation through the expansion of dialect datasets, the absence of publicly available testing data currently impedes robust benchmarking efforts. This paper presents the Dialect Dominated Dataset of Bird Vocalisation, the first cross-corpus dataset that focuses on dialects in bird vocalisations. The DB3V comprises more than 25 hours of audio recordings from 10 bird species distributed across three distinct regions in the contiguous United States (CONUS). In addition to presenting the dataset, we conduct analyses and establish baseline models for cross-corpus bird recognition. The data and code are publicly available online: https: zenodo.org records 11544734


eess.AS音频处理

【1】 Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
标题: 探索多语言广播和机构语音自动转录的口语识别策略
作者:Martina Valente,Fabio Brugnara,Giovanni Morrone,Enrico Zovato,Leonardo Badino
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文讨论了口语识别(SLI)和语音识别的多语种广播和机构的讲话,真正的应用场景,很少在SLI文献中得到解决。观察到在这些领域的语言变化大多与说话人的变化,我们提出了一个级联系统,包括说话人日记和语言识别,并比较它与更传统的语言识别和语言日记系统。实验结果表明,该系统可以实现较低的语言分类和语言日志化错误率(相对语言日志错误减少高达10%,相对语言混淆减少60%),并导致多语言测试集的WER较低(相对WER减少8%以上),而同时不会对单语音频上的语音识别产生负面影响(相对于语音识别,绝对WER增加在0.1%和0.7%之间。单语ASR)。摘要:This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Observing that in these domains language changes are mostly associated with speaker changes, we propose a cascaded system consisting of speaker diarization and language identification and compare it with more traditional language identification and language diarization systems. Results show that the proposed system often achieves lower language classification and language diarization error rates (up to 10% relative language diarization error reduction and 60% relative language confusion reduction) and leads to lower WERs on multilingual test sets (more than 8% relative WER reduction), while at the same time does not negatively affect speech recognition on monolingual audio (with an absolute WER increase between 0.1% and 0.7% w.r.t. monolingual ASR).

【2】 FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
标题: FlowAVSE:具有条件流匹配的高效视听语音增强
作者:Chaeyoung Jung,Suyeon Lee,Ji-Hoon Kim,Joon Son Chung
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:这项工作提出了一种有效的方法来提高质量的损坏的语音信号,利用声学和视觉线索。虽然现有的基于扩散的方法已经证明了显着的质量,其适用性受到缓慢的推理速度和计算复杂性的限制。为了解决这个问题,我们提出了FlowAVSE,它提高了推理速度,减少了可学习参数的数量,而不会降低输出质量。特别是,我们采用了一个条件流匹配算法,使高质量的语音在一个单一的采样步骤的生成。此外,我们通过优化基于扩散的系统的底层U-网架构来提高效率。我们的实验表明,FlowAVSE实现了22倍的推理速度,并减少了一半的模型大小,同时保持输出质量。演示页面可在https: cyongong.github.io FlowAVSE.github.io 上找到摘要:This work proposes an efficient method to enhance the quality of corrupted speech signals by leveraging both acoustic and visual cues. While existing diffusion-based approaches have demonstrated remarkable quality, their applicability is limited by slow inference speeds and computational complexity. To address this issue, we present FlowAVSE which enhances the inference speed and reduces the number of learnable parameters without degrading the output quality. In particular, we employ a conditional flow matching algorithm that enables the generation of high-quality speech in a single sampling step. Moreover, we increase efficiency by optimizing the underlying U-net architecture of diffusion-based systems. Our experiments demonstrate that FlowAVSE achieves 22 times faster inference speed and reduces the model size by half while maintaining the output quality. The demo page is available at: https: cyongong.github.io FlowAVSE.github.io

【3】 End-to-end Streaming model for Low-Latency Speech Anonymization
标题: 低延迟语音模拟的端到端流媒体模型
作者:Waris Quamer,Ricardo Gutierrez-Osuna
链接:点击下载PDF文件
摘要:说话人匿名化的目的是隐藏说话人身份的线索,同时保留语言内容。目前基于机器学习的方法需要大量的计算资源,阻碍了实时流媒体应用。为了解决这些问题,我们提出了一个流模型,实现扬声器匿名低延迟。该系统以端到端自动编码器的方式进行训练,使用轻量级内容编码器提取类似于HuBERT的信息,预训练的扬声器编码器提取扬声器身份,以及方差编码器注入音高和能量信息。这三个解纠缠的表示被馈送到重新合成语音信号的解码器。我们从我们的系统的两个实现,一个完整的模型,实现了230毫秒的延迟,和一个精简版(0.1倍大小),进一步减少延迟到66毫秒,同时保持最先进的性能在自然,可理解性和隐私保护的评估结果。摘要:Speaker anonymization aims to conceal cues to speaker identity while preserving linguistic content. Current machine learning based approaches require substantial computational resources, hindering real-time streaming applications. To address these concerns, we propose a streaming model that achieves speaker anonymization with low latency. The system is trained in an end-to-end autoencoder fashion using a lightweight content encoder that extracts HuBERT-like information, a pretrained speaker encoder that extract speaker identity, and a variance encoder that injects pitch and energy information. These three disentangled representations are fed to a decoder that resynthesizes the speech signal. We present evaluation results from two implementations of our system, a full model that achieves a latency of 230ms, and a lite version (0.1x in size) that further reduces latency to 66ms while maintaining state-of-the-art performance in naturalness, intelligibility, and privacy preservation.

【4】 ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
标题: ToneUnit:音调语言语音合成的语音离散化方法
作者:Dehua Tao,Daxin Tan,Yu Ting Yeung,Xiao Chen,Tan Lee
链接:点击下载PDF文件
摘要:将语音表示为离散单元在支持下游口语处理任务方面具有许多好处。然而,这种方法在汉语普通话等有声调语言的语音合成中的研究较少。我们在汉语语音合成上的初步实验揭示了“声调偏移”的问题,即合成的语音话语包含正确的基本音节,但不正确的声调。为了解决这个问题,我们提出了ToneUnit框架,该框架利用带有音调标签的注释数据作为CTC监督来学习汉语普通话语音的音调感知离散语音单元。我们的研究结果表明,通过TonUnit获得的离散单元解决了汉语合成语音中的“变调”问题,并在英语合成中取得了良好的效果。此外,实验结果表明,有限标量量化增强了ToneUnit的有效性。值得注意的是,ToneUnit即使使用最少的注释数据也可以有效地工作。摘要:Representing speech as discretized units has numerous benefits in supporting downstream spoken language processing tasks. However, the approach has been less explored in speech synthesis of tonal languages like Mandarin Chinese. Our preliminary experiments on Chinese speech synthesis reveal the issue of "tone shift", where a synthesized speech utterance contains correct base syllables but incorrect tones. To address the issue, we propose the ToneUnit framework, which leverages annotated data with tone labels as CTC supervision to learn tone-aware discrete speech units for Mandarin Chinese speech. Our findings indicate that the discrete units acquired through the TonUnit resolve the "tone shift" issue in synthesized Chinese speech and yield favorable results in English synthesis. Moreover, the experimental results suggest that finite scalar quantization enhances the effectiveness of ToneUnit. Notably, ToneUnit can work effectively even with minimal annotated data.

【5】 Cascaded noise reduction and acoustic echo cancellation based on an extended noise reduction
标题: 基于扩展降噪的级联降噪和声回声消除
作者:Arnout Roebben,Toon van Waterschoot,Marc Moonen
备注:Accepted for publication in EUSIPCO 2024
链接:点击下载PDF文件
摘要:在许多语音记录应用中,所记录的期望语音被噪声和声学回声破坏,使得需要组合降噪(NR)和声学回声消除(AEC)。常见的级联设计对应于AEC滤波器之前的NR个滤波器。这些NR滤波器旨在降低近端房间噪声(以及可能部分回声),并且仅对麦克风进行操作,因此需要AEC滤波器对回声路径和NR滤波器进行建模。然而,在本文中,我们提出了一种设计扩展NR(NRext)过滤器前AEC过滤器的回声路径是添加剂的地图的假设下,从而保留了加法运算。在这里,NRext滤波器旨在减少回声中的近端房间噪声和远端房间噪声分量,并且在麦克风和扬声器上操作。我们表明,成功的AEC过滤器显着成为独立的NRext过滤器,这样的AEC过滤器只需要模拟的回声路径,提高AEC性能。此外,NRext滤波器中的自由度随着扬声器的数量而缩放,这不是NR滤波器的情况,从而导致改善的NR性能。摘要:In many speech recording applications, the recorded desired speech is corrupted by both noise and acoustic echo, such that combined noise reduction (NR) and acoustic echo cancellation (AEC) is called for. A common cascaded design corresponds to NR filters preceding AEC filters. These NR filters aim at reducing the near-end room noise (and possibly partially the echo) and operate on the microphones only, consequently requiring the AEC filters to model both the echo paths and the NR filters. In this paper, however, we propose a design with extended NR (NRext) filters preceding AEC filters under the assumption of the echo paths being additive maps, thus preserving the addition operation. Here, the NRext filters aim at reducing both the near-end room noise and the far-end room noise component in the echo, and operate on both the microphones and loudspeakers. We show that the succeeding AEC filters remarkably become independent of the NRext filters, such that the AEC filters are only required to model the echo paths, improving the AEC performance. Further, the degrees of freedom in the NRext filters scale with the number of loudspeakers, which is not the case for the NR filters, resulting in an improved NR performance.

【6】 Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
标题: 使用超声波麦克风阵列和CNN预测NC车床操作中的刀具磨损
作者:Jan Steckel,Arne Aerts,Erik Verreycken,Dennis Laurijssen,Walter Daems
链接:点击下载PDF文件
摘要:本文介绍了一种新的方法来预测刀具磨损数控车削操作,结合超声波麦克风阵列和卷积神经网络(CNN)。使用波束形成技术增强OkHz和60kHz之间的高频声发射以提高信噪比。然后通过CNN分析处理后的声学数据,预测切削刀具的剩余使用寿命(RUL)。通过对350个用单一硬质合金刀片加工的工件的数据进行训练,该模型可以准确地预测硬质合金刀片的RUL。我们的研究结果表明,通过将先进的超声波传感器与深度学习相结合,可以在CNC加工中实现精确的预测性维护任务。摘要:This paper introduces a novel method for predicting tool wear in CNC turning operations, combining ultrasonic microphone arrays and convolutional neural networks (CNNs). High-frequency acoustic emissions between 0 kHz and 60 kHz are enhanced using beamforming techniques to improve the signal- to-noise ratio. The processed acoustic data is then analyzed by a CNN, which predicts the Remaining Useful Life (RUL) of cutting tools. Trained on data from 350 workpieces machined with a single carbide insert, the model can accurately predict the RUL of the carbide insert. Our results demonstrate the potential gained by integrating advanced ultrasonic sensors with deep learning for accurate predictive maintenance tasks in CNC machining.

【7】 On Improving Error Resilience of Neural End-to-End Speech Coders
标题: 提高神经端到端语音编码器的抗错误能力
作者:Kishan Gupta,Nicola Pia,Srikanth Korse,Andreas Brendel,Guillaume Fuchs,Markus Multrus
链接:点击下载PDF文件
摘要:丢包隐藏(PLC)和前向纠错(FEC)等容错工具对于维护IP语音(VoIP)等应用的可靠语音通信至关重要,因为在这些应用中,数据包经常延迟和丢失。近年来,端到端神经语音编解码器由于其以低比特率传输语音信号的能力而显着增加,但很少考虑其在实际系统中的错误恢复能力。最近推出的神经端到端语音编解码器(NESC)可以在低比特率下再现高质量的自然语音。我们通过增加一个低复杂度的网络来预测潜在空间中的码本索引,从而扩展其对分组丢失的鲁棒性。此外,我们提出了一种方法来添加一个带内FEC在一个额外的比特率为0.8 kbps。主观和客观的评估表明,所提出的方法的有效性,并表明耦合PLC和FEC提供了显着的鲁棒性对数据包丢失。摘要:Error resilient tools like Packet Loss Concealment (PLC) and Forward Error Correction (FEC) are essential to maintain a reliable speech communication for applications like Voice over Internet Protocol (VoIP), where packets are frequently delayed and lost. In recent times, end-to-end neural speech codecs have seen a significant rise, due to their ability to transmit speech signal at low bitrates but few considerations were made about their error resilience in a real system. Recently introduced Neural End-to-End Speech Codec (NESC) can reproduce high quality natural speech at low bitrates. We extend its robustness to packet losses by adding a low complexity network to predict the codebook indices in latent space. Furthermore, we propose a method to add an in-band FEC at an additional bitrate of 0.8 kbps. Both subjective and objective assessment indicate the effectiveness of proposed methods, and demonstrate that coupling PLC and FEC provide significant robustness against packet losses.

【8】 DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
标题: DisfluencySpeech --具有副语言的单说话者对话语音数据集
作者:Kyra Wang,Dorien Herremans
备注:4 pages, 1 figure, submitted to IEEE TENCON 2024
链接:点击下载PDF文件
摘要:笑、叹息、口吃和其他形式的语言并不直接为言语提供词汇意义,但它们提供了关键的命题语境,有助于语义和语用过程,如讽刺。因此,重要的是人工社会代理都理解,并能够生成语音与语义重要的语言。大多数语音数据集不包括转录的非词汇语音声音和不流利,而那些包括的通常是多说话者数据集,其中每个说话者提供相对较少的音频。这使得训练包括这样的语言成分的会话文本到语音(TTS)合成模型具有挑战性。 因此,我们提出了DisfluencySpeech,一个工作室质量的标记英语语音数据集,具有可编程语言。一个单一的扬声器再现近10个小时的表达话语从交换机-1电话语音语料库(交换机),模拟现实的非正式对话。为了帮助开发一个能够从没有这些组件的文本中预测合成语言的TTS模型,我们提供了三种不同的转录本,它们处于不同的信息去除水平(去除非语音事件,去除非句子元素和去除错误开始),以及在每个级别上训练的基准TTS模型。摘要:Laughing, sighing, stuttering, and other forms of paralanguage do not contribute any direct lexical meaning to speech, but they provide crucial propositional context that aids semantic and pragmatic processes such as irony. It is thus important for artificial social agents to both understand and be able to generate speech with semantically-important paralanguage. Most speech datasets do not include transcribed non-lexical speech sounds and disfluencies, while those that do are typically multi-speaker datasets where each speaker provides relatively little audio. This makes it challenging to train conversational Text-to-Speech (TTS) synthesis models that include such paralinguistic components. We thus present DisfluencySpeech, a studio-quality labeled English speech dataset with paralanguage. A single speaker recreates nearly 10 hours of expressive utterances from the Switchboard-1 Telephone Speech Corpus (Switchboard), simulating realistic informal conversations. To aid the development of a TTS model that is able to predictively synthesise paralanguage from text without such components, we provide three different transcripts at different levels of information removal (removal of non-speech events, removal of non-sentence elements, and removal of false starts), as well as benchmark TTS models trained on each of these levels.

【9】 DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
标题: DubWise:多模式基于LLM的文本转语音配音中的视频引导语音持续时间控制
作者:Neha Sahipjohn,Ashishkumar Gudmalwar,Nirmesh Shah,Pankaj Wasnik,Rajiv Ratn Shah
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:配音后的视听对齐是一个具有挑战性的研究问题。为此,我们提出了一种新的方法,基于DubWise多模态大语言模型(LLM)的文本到语音(TTS),它可以控制合成语音的语音持续时间,使其与参考视频中给出的扬声器嘴唇运动保持一致,即使口语文本不同或使用不同的语言。为了实现这一点,我们建议在预先训练的基于GPT的TTS中利用跨模态注意力技术。我们结合语言令牌从文本,扬声器身份令牌通过语音克隆网络,和视频令牌通过一个建议的持续时间控制器网络。我们证明了我们的系统在Lip 2 Wav-Chemistry和LRS 2数据集上的有效性。此外,与用于相同语言但不同文本的SOTA相比,所提出的方法实现了改进的唇同步和自然度(即,非平行)和不同的语言,不同的文本(即,跨语言)场景。摘要:Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech (TTS), which can control the speech duration of synthesized speech in such a way that it aligns well with the speakers lip movements given in the reference video even when the spoken text is different or in a different language. To accomplish this, we propose to utilize cross-modal attention techniques in a pre-trained GPT-based TTS. We combine linguistic tokens from text, speaker identity tokens via a voice cloning network, and video tokens via a proposed duration controller network. We demonstrate the effectiveness of our system on the Lip2Wav-Chemistry and LRS2 datasets. Also, the proposed method achieves improved lip sync and naturalness compared to the SOTAs for the same language but different text (i.e., non-parallel) and the different language, different text (i.e., cross-lingual) scenarios.

【10】 Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
标题: 从脑电信号实现完全端到端的听听语音解码
作者:Jihwan Lee,Aditya Kommineni,Tiantian Feng,Kleanthis Avramidis,Xuan Shi,Sudarsana Kadiri,Shrikanth Narayanan
备注:accepted to Interspeech2024
链接:点击下载PDF文件
摘要:从EEG信号中解码语音是一项具有挑战性的任务,其中对大脑活动进行建模以估计声学刺激的显著特征。我们提出了FESDE,一个新的框架,从脑电信号完全端到端的语音解码。我们的方法的目的是直接重建听到的语音波形给定的EEG信号,其中没有中间的声学特征处理步骤是必需的。所提出的方法包括EEG模块和语音模块以及连接器。EEG模块学习更好地表示EEG信号,而语音模块从模型表示生成语音波形。连接器学习桥接EEG和语音的潜在空间的分布。建议的框架是简单和有效的,通过允许单步推理,并优于以前的作品的客观指标。细粒度的音素分析进行揭示语音解码的模型特性。源代码可以在这里找到:github.com lee-jhwn fesde。摘要:Speech decoding from EEG signals is a challenging task, where brain activity is modeled to estimate salient characteristics of acoustic stimuli. We propose FESDE, a novel framework for Fully-End-to-end Speech Decoding from EEG signals. Our approach aims to directly reconstruct listened speech waveforms given EEG signals, where no intermediate acoustic feature processing step is required. The proposed method consists of an EEG module and a speech module along with a connector. The EEG module learns to better represent EEG signals, while the speech module generates speech waveforms from model representations. The connector learns to bridge the distributions of the latent spaces of EEG and speech. The proposed framework is both simple and efficient, by allowing single-step inference, and outperforms prior works on objective metrics. A fine-grained phoneme analysis is conducted to unveil model characteristics of speech decoding. The source code is available here: github.com lee-jhwn fesde.

【11】 DB3V: A Dialect Dominated Dataset of Bird Vocalisation for Cross-corpus Bird Species Recognition
标题: DB 3V:一个基于方言的鸟类发声数据集,用于跨库鸟类物种识别
作者:Xin Jing,Luyang Zhang,Jiangjian Xie,Alexander Gebhard,Alice Baird,Bjoern Schuller
备注:accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在鸟类学中,已知鸟类物种有多种多样的方言,这一点得到了广泛的承认,即鸟类物种在不同地区的叫声中表现出不同的方言。因此,仅通过叫声识别鸟类的计算方法面临着重大挑战。有越来越多的兴趣,了解鸟类物种识别方法的有效性,物种特异性方言的影响。尽管通过扩大方言数据集可能会缓解这种情况,但目前缺乏公开的测试数据阻碍了强有力的基准测试工作。本文介绍了方言主导的鸟类发声数据集,第一个跨语料库数据集,重点是方言的鸟类发声。DB3V包括超过25小时的音频记录,来自10种鸟类,分布在美国毗连的三个不同地区。除了呈现数据集之外,我们还进行分析并建立跨语料库鸟类识别的基线模型。数据和代码可在网上公开获取:https: zenodo.org records 11544734摘要:In ornithology, bird species are known to have variedit's widely acknowledged that bird species display diverse dialects in their calls across different regions. Consequently, computational methods to identify bird species onsolely through their calls face critsignificalnt challenges. There is growing interest in understanding the impact of species-specific dialects on the effectiveness of bird species recognition methods. Despite potential mitigation through the expansion of dialect datasets, the absence of publicly available testing data currently impedes robust benchmarking efforts. This paper presents the Dialect Dominated Dataset of Bird Vocalisation, the first cross-corpus dataset that focuses on dialects in bird vocalisations. The DB3V comprises more than 25 hours of audio recordings from 10 bird species distributed across three distinct regions in the contiguous United States (CONUS). In addition to presenting the dataset, we conduct analyses and establish baseline models for cross-corpus bird recognition. The data and code are publicly available online: https: zenodo.org records 11544734

【12】 DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
标题: DiscreteSLU:一个具有自监督离散语音单元的大型语言模型,用于口语理解
作者:Suwon Shon,Kwangyoun Kim,Yi-Te Hsu,Prashant Sridhar,Shinji Watanabe,Karen Livescu
链接:点击下载PDF文件
摘要:将预先训练的基于文本的大型语言模型(LLM)与语音输入相结合,实现了针对不同语音任务的自动跟随能力。这种集成需要使用语音编码器,语音适配器和LLM,在不同的任务上进行训练。我们建议使用离散语音单元(DSU),而不是连续值的语音编码器的输出,这是转换到LLM令牌嵌入空间使用语音适配器。我们生成DSU使用自监督语音编码器,其次是k-均值聚类。所提出的模型表现出强大的性能语音输入从可见 不可见的域和解释以下的能力,在口语问答。我们还探讨了从自监督语音编码器的不同层中提取的各种类型的DSU,以及Mel频率倒谱系数(MFCC)。我们的研究结果表明,ASR任务和数据集是不是至关重要的提示调整口语问答任务。摘要:The integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks. This integration requires the use of a speech encoder, a speech adapter, and an LLM, trained on diverse tasks. We propose the use of discrete speech units (DSU), rather than continuous-valued speech encoder outputs, that are converted to the LLM token embedding space using the speech adapter. We generate DSU using a self-supervised speech encoder followed by k-means clustering. The proposed model shows robust performance on speech inputs from seen unseen domains and instruction-following capability in spoken question answering. We also explore various types of DSU extracted from different layers of the self-supervised speech encoder, as well as Mel frequency Cepstral Coefficients (MFCC). Our findings suggest that the ASR task and datasets are not crucial in instruction-tuning for spoken question answering tasks.

【13】 PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance
标题: PianoMotion10 M:钢琴演奏中手部运动生成的数据集和基准
作者:Qijun Gan,Song Wang,Shengtao Wu,Jianke Zhu
备注:Codes and Dataset: this https URL
链接:点击下载PDF文件
摘要:近年来,人工智能技术在教育领域受到越来越多的关注,但设计有效的乐器教学系统仍然是一个悬而未决的问题。虽然按键可以直接从乐谱中得到,但按键之间的过渡动作在钢琴演奏中需要更广泛的指导。在这项工作中,我们构建了一个钢琴手运动生成基准,以指导钢琴演奏的手运动和指法。为此,我们收集了一个带注释的数据集PianoMotion10M,其中包含116小时的钢琴演奏视频,这些视频来自鸟瞰图,其中包含1000万个带注释的手部姿势。我们还引入了一个强大的基线模型,通过位置预测器和位置引导的手势生成器从钢琴音频生成手部动作。此外,设计了一系列的评估指标来评估基线模型的性能,包括运动相似性,平滑性,左右手的位置准确性,以及运动分布的整体保真度。尽管钢琴按键与乐谱或音频已经可以访问,PianoMotion10M旨在提供指导钢琴指法的教学目的。数据集和源代码可以在https: agnjason.github.io PianoMotion-page上访问。摘要:Recently, artificial intelligence techniques for education have been received increasing attentions, while it still remains an open problem to design the effective music instrument instructing systems. Although key presses can be directly derived from sheet music, the transitional movements among key presses require more extensive guidance in piano performance. In this work, we construct a piano-hand motion generation benchmark to guide hand movements and fingerings for piano playing. To this end, we collect an annotated dataset, PianoMotion10M, consisting of 116 hours of piano playing videos from a bird's-eye view with 10 million annotated hand poses. We also introduce a powerful baseline model that generates hand motions from piano audios through a position predictor and a position-guided gesture generator. Furthermore, a series of evaluation metrics are designed to assess the performance of the baseline model, including motion similarity, smoothness, positional accuracy of left and right hands, and overall fidelity of movement distribution. Despite that piano key presses with respect to music scores or audios are already accessible, PianoMotion10M aims to provide guidance on piano fingering for instruction purposes. The dataset and source code can be accessed at https: agnjason.github.io PianoMotion-page.

【14】 On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
标题: 论异类数据源对语音转文本基础模型的影响
作者:Jinchuan Tian,Yifan Peng,William Chen,Kwanghee Choi,Karen Livescu,Shinji Watanabe
链接:点击下载PDF文件
摘要:开放式耳语风格语音模型(OWSM)系列旨在实现构建高级语音到文本(S2T)基础模型的完全透明性。为此,OWSM模型在25个公共语音数据集上进行了训练,这些数据集在多种方面都是异构的。在这项研究中,我们通过引入OWSM v3.2来推进OWSM系列,该系列通过调查和解决这种数据异质性的影响来改进先前的模型。我们的研究从每个数据集的详细分析开始,从中我们得出两个关键策略:使用代理任务进行数据过滤以提高数据质量,以及使用开放式大型语言模型(LLM)合并标点符号和真大小写。在所有其他配置保持不变的情况下,OWSM v3.2比OWSM v3.1基准提高了性能,同时使用的训练数据减少了15%。摘要:The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are heterogeneous in multiple ways. In this study, we advance the OWSM series by introducing OWSM v3.2, which improves on prior models by investigating and addressing the impacts of this data heterogeneity. Our study begins with a detailed analysis of each dataset, from which we derive two key strategies: data filtering with proxy task to enhance data quality, and the incorporation of punctuation and true-casing using an open large language model (LLM). With all other configurations staying the same, OWSM v3.2 improves performance over the OWSM v3.1 baseline while using 15% less training data.

【15】 Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
标题: SYS 2 Sound:从以自我为中心的视频中产生具有环境意识的动作声音
作者:Changan Chen,Puyuan Peng,Ami Baid,Zihui Xue,Wei-Ning Hsu,David Harwarth,Kristen Grauman
备注:Project page: this https URL
链接:点击下载PDF文件
摘要:为人类交互生成逼真的音频对于许多应用都很重要,例如为电影或虚拟现实游戏创建声音效果。现有的方法隐含地假设在训练过程中视频和音频之间完全对应,但许多声音发生在屏幕外,与视觉效果的对应性很弱甚至没有对应性-导致在测试时出现不受控制的环境声音或幻觉。我们提出了一种新的环境感知音频生成模型,AV-LDM。我们设计了一种新的音频调节机制,以学习在野外训练视频中将前景动作声音与环境背景声音分开。给定一个新颖的无声视频,我们的模型使用检索增强生成来创建在语义和时间上都与视觉内容相匹配的音频。我们训练和评估我们的模型在两个在野外自我中心的视频数据集Ego4D和EPIC-KITCHENS。我们的模型优于一系列现有的方法,允许可控生成的环境声音,甚至显示出推广到计算机图形游戏剪辑的承诺。总的来说,我们的工作是第一个将视频到音频的生成忠实地集中在所观察到的视觉内容上,尽管训练来自具有自然背景声音的未经策划的剪辑。摘要:Generating realistic audio for human interactions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio during training, yet many sounds happen off-screen and have weak to no correspondence with the visuals -- resulting in uncontrolled ambient sounds or hallucinations at test time. We propose a novel ambient-aware audio generation model, AV-LDM. We devise a novel audio-conditioning mechanism to learn to disentangle foreground action sounds from the ambient background sounds in in-the-wild training videos. Given a novel silent video, our model uses retrieval-augmented generation to create audio that matches the visual content both semantically and temporally. We train and evaluate our model on two in-the-wild egocentric video datasets Ego4D and EPIC-KITCHENS. Our model outperforms an array of existing methods, allows controllable generation of the ambient sound, and even shows promise for generalizing to computer graphics game clips. Overall, our work is the first to focus video-to-audio generation faithfully on the observed visual content despite training from uncurated clips with natural background sounds.

【16】 Vision Transformer Segmentation for Visual Bird Sound Denoising
标题: 视觉Transformer分割用于视觉鸟声去噪
作者:Sahil Kumar,Jialu Li,Youshan Zhang
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:由于持续的残留噪声,音频去噪,特别是在鸟鸣的背景下,仍然是一项具有挑战性的任务。传统和深度学习方法经常与人为或低频噪声作斗争。在这项工作中,我们提出了ViTVS,一种新的方法,利用权力的Vision Transformer(ViT)架构。ViTVS巧妙地结合了分割技术,从复杂的信号混合中分离出干净的音频。我们的主要贡献包括ViTVS的开发,引入了全面的,长距离的和多尺度的表示。这些贡献直接解决了传统方法固有的局限性。大量的实验表明,ViTVS优于最先进的方法,将其定位为现实世界的鸟声去噪应用的基准解决方案。源代码可在https: github.com aiai-4 ViVTS上获得。摘要:Audio denoising, especially in the context of bird sounds, remains a challenging task due to persistent residual noise. Traditional and deep learning methods often struggle with artificial or low-frequency noise. In this work, we propose ViTVS, a novel approach that leverages the power of the vision transformer (ViT) architecture. ViTVS adeptly combines segmentation techniques to disentangle clean audio from complex signal mixtures. Our key contributions encompass the development of ViTVS, introducing comprehensive, long-range, and multi-scale representations. These contributions directly tackle the limitations inherent in conventional approaches. Extensive experiments demonstrate that ViTVS outperforms state-of-the-art methods, positioning it as a benchmark solution for real-world bird sound denoising applications. Source code is available at: https: github.com aiai-4 ViVTS.

【17】 Complex Image-Generative Diffusion Transformer for Audio Denoising
标题: 用于音频去噪的复杂图像生成扩散Transformer
作者:Junhui Li,Pu Wang,Jialu Li,Youshan Zhang
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:音频去噪技术在深度神经网络领域引起了广泛的关注。最近,音频去噪问题已被转换为图像生成任务,基于深度学习的方法已被应用于解决这个问题。但其表现仍然有限,留有进一步提升的空间。为了提高音频去噪性能,本文介绍了一种复图像生成扩散Transformer,从复傅立叶域捕获更多的信息。通过将Transformer与扩散模型相结合,提出了一种新型的扩散Transformer.我们提出的模型证明了Transformer的可扩展性,并使用注意扩散扩展稀疏注意的感受野。我们的工作是第一个利用扩散Transformers来处理音频去噪的图像生成任务。在两个基准数据集上的大量实验表明,我们提出的模型优于最先进的方法。摘要:The audio denoising technique has captured widespread attention in the deep neural network field. Recently, the audio denoising problem has been converted into an image generation task, and deep learning-based approaches have been applied to tackle this problem. However, its performance is still limited, leaving room for further improvement. In order to enhance audio denoising performance, this paper introduces a complex image-generative diffusion transformer that captures more information from the complex Fourier domain. We explore a novel diffusion transformer by integrating the transformer with a diffusion model. Our proposed model demonstrates the scalability of the transformer and expands the receptive field of sparse attention using attention diffusion. Our work is among the first to utilize diffusion transformers to deal with the image generation task for audio denoising. Extensive experiments on two benchmark datasets demonstrate that our proposed model outperforms state-of-the-art methods.

【18】 Towards Multilingual Audio-Visual Question Answering
标题: 走向多语言视听问答
作者:Orchid Chetia Phukan,Priyabrata Mallick,Swarup Ranjan Behera,Aalekhya Satya Narayani,Arun Balaji Buduru,Rajesh Sharma
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们努力扩展视听问题问答(AVQA)的多语言设置。现有的AVQA研究主要围绕英语进行,复制它以解决其他语言的AVQA需要大量的资源分配。作为一个可扩展的解决方案,我们利用机器翻译,并提出了两个多语言的AVQA数据集,从现有的基准AVQA数据集创建的八种语言。这可以防止人工手动收集问题和答案的额外人工注释工作。为此,我们建议,MERA框架,利用国家的最先进的(SOTA)的视频,音频和文本基础模型AVQA在多种语言。我们引入了一套模型,即MERA-L,MERA-C,MERA-T,具有不同的模型架构,以基准测试所提出的数据集。我们相信我们的工作将开辟新的研究方向,并作为未来多语言AVQA工作的参考基准。摘要:In this paper, we work towards extending Audio-Visual Question Answering (AVQA) to multilingual settings. Existing AVQA research has predominantly revolved around English and replicating it for addressing AVQA in other languages requires a substantial allocation of resources. As a scalable solution, we leverage machine translation and present two multilingual AVQA datasets for eight languages created from existing benchmark AVQA datasets. This prevents extra human annotation efforts of collecting questions and answers manually. To this end, we propose, MERA framework, by leveraging state-of-the-art (SOTA) video, audio, and textual foundation models for AVQA in multiple languages. We introduce a suite of models namely MERA-L, MERA-C, MERA-T with varied model architectures to benchmark the proposed datasets. We believe our work will open new research directions and act as a reference benchmark for future works in multilingual AVQA.

【19】 Diffusion Gaussian Mixture Audio Denoise
标题: 扩散高斯混合音频降噪
作者:Pu Wang,Junhui Li,Jialu Li,Liangdong Guo,Youshan Zhang
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:最近的扩散模型在音频去噪任务中取得了很好的性能。逆过程的独特性质可以恢复干净的信号。然而,现实世界中的噪声分布并不符合单一的高斯分布,甚至是未知的。高斯噪声条件下的采样限制了其应用场景。为了克服这些挑战,我们提出了一个DiffGMM模型,一个基于扩散和高斯混合模型的去噪模型。我们采用逆过程来估计高斯混合模型的参数。给定一个带噪的音频信号,我们首先应用一维U-Net来提取特征,并训练线性层来估计高斯混合模型的参数,然后我们近似真实的噪声分布。从估计的噪声中连续地减去噪声信号以输出干净的音频信号。大量的实验结果表明,所提出的DiffGMM模型达到了最先进的性能。摘要:Recent diffusion models have achieved promising performances in audio-denoising tasks. The unique property of the reverse process could recover clean signals. However, the distribution of real-world noises does not comply with a single Gaussian distribution and is even unknown. The sampling of Gaussian noise conditions limits its application scenarios. To overcome these challenges, we propose a DiffGMM model, a denoising model based on the diffusion and Gaussian mixture models. We employ the reverse process to estimate parameters for the Gaussian mixture model. Given a noisy audio signal, we first apply a 1D-U-Net to extract features and train linear layers to estimate parameters for the Gaussian mixture model, and we approximate the real noise distributions. The noisy signal is continuously subtracted from the estimated noise to output clean audio signals. Extensive experimental results demonstrate that the proposed DiffGMM model achieves state-of-the-art performance.

【20】 LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
标题: 激光:通过调整语音的自我监督表示来学习以改善内容相关任务
作者:Amit Meghanani,Thomas Hain
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:基于自监督学习(SSL)的语音模型被广泛用于全栈语音处理。然而,已经观察到,使用未标记的语音用于内容相关的任务来改进基于SSL的语音表示是具有挑战性的并且在计算上昂贵。最近的尝试已经解决这个问题,具有成本效益的自我监督微调(SSFT)的方法。继续在这个方向上,一个具有成本效益的SSFT方法命名为“激光:通过对齐自监督表示学习”。LASER基于具有时间正则化项的软DTW对准损失。实验进行了HuBERT和WavLM模型和评估SUPERB基准两个内容相关的任务:自动语音识别(ASR)和音素识别(PR)。对于ASR和PR任务,HuBERT的相对改进分别为3.7%和8.2%,WavLM的相对改进分别为4.1%和11.7%,在单个GPU上仅进行了< 3小时的微调。摘要:Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-related tasks is challenging and computationally expensive. Recent attempts have been made to address this issue with cost-effective self-supervised fine-tuning (SSFT) approaches. Continuing in this direction, a cost-effective SSFT method named "LASER: Learning by Aligning Self-supervised Representations" is presented. LASER is based on the soft-DTW alignment loss with temporal regularisation term. Experiments are conducted with HuBERT and WavLM models and evaluated on the SUPERB benchmark for two content-related tasks: automatic speech recognition (ASR) and phoneme recognition (PR). A relative improvement of 3.7% and 8.2% for HuBERT, and 4.1% and 11.7% for WavLM are observed, for the ASR and PR tasks respectively, with only < 3 hours of fine-tuning on a single GPU.

【21】 Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
标题: 探索多语言看不见的说话者情感识别:在多任务学习中利用共同注意线索
作者:Arnav Goel,Medha Hira,Anubha Gupta
备注:5 pages, Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:现代深度学习技术的出现带来了语音情感识别(SER)领域的进步。然而,该领域中流行的大多数系统未能推广到在训练期间未看到的扬声器。这项研究的重点是处理多语言SER的挑战,特别是看不见的扬声器。我们介绍CAMuLeNet,一种新的架构,利用基于共同注意力的融合和多任务学习来解决这个问题。此外,我们使用10倍leave-speaker-out交叉验证对Whisper,HuBERT,Wav2Vec2.0和WavLM的预训练编码器进行了基准测试,这些测试针对五个现有的多语言基准数据集:IEMOCAP,RAVDESS,CREMA-D,CARDB和CaFE,并在印地语(BhavVani)上发布了SER的新数据集。CAMuLeNet显示,在我们的交叉验证策略确定的未见过的扬声器的所有基准测试中,平均提高了约8%。摘要:Advent of modern deep learning techniques has given rise to advancements in the field of Speech Emotion Recognition (SER). However, most systems prevalent in the field fail to generalize to speakers not seen during training. This study focuses on handling challenges of multilingual SER, specifically on unseen speakers. We introduce CAMuLeNet, a novel architecture leveraging co-attention based fusion and multitask learning to address this problem. Additionally, we benchmark pretrained encoders of Whisper, HuBERT, Wav2Vec2.0, and WavLM using 10-fold leave-speaker-out cross-validation on five existing multilingual benchmark datasets: IEMOCAP, RAVDESS, CREMA-D, EmoDB and CaFE and, release a novel dataset for SER on the Hindi language (BhavVani). CAMuLeNet shows an average improvement of approximately 8% over all benchmarks on unseen speakers determined by our cross-validation strategy.

【22】 AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
标题: AV-GS:新视图声学合成的学习材料和几何感知先验
作者:Swapnil Bhosale,Haosen Yang,Diptesh Kanojia,Jiankang Deng,Xiatian Zhu
链接:点击下载PDF文件
摘要:新颖视图声学合成(NVAS)旨在在给定由3D场景处的声源发出的单声道音频的情况下在任何目标视点处渲染双耳音频。现有的方法已经提出了基于NeRF的隐式模型,以利用视觉线索作为合成双耳音频的条件。然而,除了源于繁重NeRF渲染的低效率之外,这些方法都具有表征整个场景环境的有限能力,例如房间几何形状,材料属性以及收听者和声源之间的空间关系。为了解决这些问题,我们提出了一种新的视听高斯飞溅(AV-GS)模型。为了获得音频合成的材料感知和几何感知条件,我们学习了一个显式的基于点的场景表示,在本地初始化的高斯点上具有音频指导参数,同时考虑到来自听众和声源的空间关系。为了使视觉场景模型音频自适应,我们提出了一种点致密化和修剪策略来最佳地分布高斯点,其中每个点在声音传播中的贡献(例如,无纹理的壁表面需要更多的点,因为它们影响声音路径转移)。大量的实验验证了我们的AV-GS优于现实世界RWAS和基于模拟的SoundSpaces数据集上的现有替代品。摘要:Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a condition for synthesizing binaural audio. However, in addition to low efficiency originating from heavy NeRF rendering, these methods all have a limited ability of characterizing the entire scene environment such as room geometry, material properties, and the spatial relation between the listener and sound source. To address these issues, we propose a novel Audio-Visual Gaussian Splatting (AV-GS) model. To obtain a material-aware and geometry-aware condition for audio synthesis, we learn an explicit point-based scene representation with an audio-guidance parameter on locally initialized Gaussian points, taking into account the space relation from the listener and sound source. To make the visual scene model audio adaptive, we propose a point densification and pruning strategy to optimally distribute the Gaussian points, with the per-point contribution in sound propagation (e.g., more points needed for texture-less wall surfaces as they affect sound path diversion). Extensive experiments validate the superiority of our AV-GS over existing alternatives on the real-world RWAS and simulation-based SoundSpaces datasets.

【23】 Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
标题: 用于噪音和回响多说话人自动语音识别的语音分离模型的免转录微调
作者:William Ravenscroft,George Close,Stefan Goetze,Thomas Hain,Mohammad Soleymanpour,Anurag Chowdhury,Mark C. Fuhs
备注:5 pages, 3 Figures, 3 Tables, Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:重叠说话者的自动语音识别(ASR)的一种解决方案是分离语音,然后对分离的信号执行ASR。通常,分离器会产生伪影,这通常会降低ASR性能。解决这个问题通常需要参考传输来联合训练分离和ASR网络。这对于在参考转录信息并不总是可用的真实世界域内音频上进行训练通常是不可行的。本文提出了一种只使用音频信号进行联合训练的免转录方法。所提出的方法使用预训练的ASR编码器的嵌入差异作为损失,并对称为引导PIT(GPIT)的排列不变训练(PIT)进行修改。该方法实现了6.4%的改善,字错误率(WER)的措施超过了信号电平的损失,也表现出增强改善感知措施,如短期客观可懂度(STOI)。摘要:One solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals. Commonly, the separator produces artefacts which often degrade ASR performance. Addressing this issue typically requires reference transcriptions to jointly train the separation and ASR networks. This is often not viable for training on real-world in-domain audio where reference transcript information is not always available. This paper proposes a transcription-free method for joint training using only audio signals. The proposed method uses embedding differences of pre-trained ASR encoders as a loss with a proposed modification to permutation invariant training (PIT) called guided PIT (GPIT). The method achieves a 6.4% improvement in word error rate (WER) measures over a signal-level loss and also shows enhancement improvements in perceptual measures such as short-time objective intelligibility (STOI).

【24】 An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
标题: 低资源场景下TTC系统语言适应的初步研究
作者:Cheng Gong,Erica Cooper,Xin Wang,Chunyu Qiang,Mengzhe Geng,Dan Wells,Longbiao Wang,Jianwu Dang,Marc Tessier,Aidan Pine,Korin Richmond,Junichi Yamagishi
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:来自大规模多语言模型的自监督学习(SSL)表示为低资源语言语音任务提供了一个有前途的解决方案。尽管取得了进步,但TTS系统中的语言适应仍然是一个悬而未决的问题。本文探讨的语言适应能力的ZMM-TTS,最近基于SSL的多语言TTS系统在我们以前的工作中提出。我们对12种语言进行了实验,使用有限的数据与各种微调配置。我们证明了预训练和目标语言之间的语音相似性,以及语言类别,影响目标语言的适应性能。此外,我们发现微调数据集的大小和扬声器的数量影响适应性。令人惊讶的是,我们还观察到,与仅音频数据相比,使用配对数据进行微调并不总是最佳的。除了语音清晰度,我们的分析还包括说话人相似性,语言识别和预测MOS。摘要:Self-supervised learning (SSL) representations from massively multilingual models offer a promising solution for low-resource language speech tasks. Despite advancements, language adaptation in TTS systems remains an open problem. This paper explores the language adaptation capability of ZMM-TTS, a recent SSL-based multilingual TTS system proposed in our previous work. We conducted experiments on 12 languages using limited data with various fine-tuning configurations. We demonstrate that the similarity in phonetics between the pre-training and target languages, as well as the language category, affects the target language's adaptation performance. Additionally, we find that the fine-tuning dataset size and number of speakers influence adaptability. Surprisingly, we also observed that using paired data for fine-tuning is not always optimal compared to audio-only data. Beyond speech intelligibility, our analysis covers speaker similarity, language identification, and predicted MOS.

【25】 SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
标题: SingOMD:从语音模型构建面向歌唱的多分辨率离散表示
作者:Yuxun Tang,Yuning Wu,Jiatong Shi,Qin Jin
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:离散表示在语音生成任务中已经显示出优势,其中离散令牌是通过将来自自监督学习(SSL)预训练模型的隐藏特征离散化来导出的。然而,语音SSL模型直接应用于歌唱生成遇到语音和歌唱之间的领域差距。此外,歌唱生成需要比典型语音更精细的表示。为了解决这些挑战,我们引入SingOMD,一种新的方法来提取面向唱歌的多分辨率离散表示语音SSL模型。具体来说,我们首先通过重新合成任务适应语音SSL的功能,并结合基于重新合成的多分辨率模块,以更好地服务于歌唱生成。这些适应多分辨率功能,然后通过聚类离散化。大量的实验表明,这些表示在歌唱声码器和歌唱声音合成的鲁棒性,效率和有效性。摘要:Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre-trained models. However, the direct application of speech SSL models to singing generation encounters domain gaps between speech and singing. Furthermore, singing generation necessitates a more refined representation than typical speech. To address these challenges, we introduce SingOMD, a novel method to extract singing-oriented multi-resolution discrete representations from speech SSL models. Specifically, we first adapt the features from speech SSL through a resynthesis task and incorporate multi-resolution modules based on resampling to better serve singing generation. These adapted multi-resolution features are then discretized via clustering. Extensive experiments demonstrate the robustness, efficiency, and effectiveness of these representations in singing vocoders and singing voice synthesis.

【26】 AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers
标题: AdaPTin:《Transformer》中产品双胞胎的低成本自适应压缩
作者:Emil Biju,Anirudh Sriram,Mert Pilanci
备注:12 pages, 3 figures, submitted to NeurIPS 2024
链接:点击下载PDF文件
摘要:虽然大型的基于变换的模型在说话人无关的语音识别中表现出了卓越的性能,但它们的大尺寸和计算要求使得它们在资源受限的环境中使用时昂贵或不切实际。在这项工作中,我们提出了一种低秩自适应压缩技术,称为AdaPTwin,共同压缩产品依赖的权重矩阵对的Transformer注意层。我们的方法可以优先考虑压缩模型的性能在一个特定的扬声器,同时保持推广到新的扬声器和声学条件。值得注意的是,我们的技术只需要8小时的语音数据进行微调,可以在20分钟内完成,与其他压缩方法相比,它具有很高的成本效益。我们通过将Whisper和Distil-Whisper模型压缩高达45%来证明我们的方法的有效性,同时单词错误率增加不到2%。摘要:While large transformer-based models have exhibited remarkable performance in speaker-independent speech recognition, their large size and computational requirements make them expensive or impractical to use in resource-constrained settings. In this work, we propose a low-rank adaptive compression technique called AdaPTwin that jointly compresses product-dependent pairs of weight matrices in the transformer attention layer. Our approach can prioritize the compressed model's performance on a specific speaker while maintaining generalizability to new speakers and acoustic conditions. Notably, our technique requires only 8 hours of speech data for fine-tuning, which can be accomplished in under 20 minutes, making it highly cost-effective compared to other compression methods. We demonstrate the efficacy of our approach by compressing the Whisper and Distil-Whisper models by up to 45% while incurring less than a 2% increase in word error rate.

【27】 A Single-Step Non-Autoregressive Automatic Speech Recognition Architecture with High Accuracy and Inference Speed
标题: 一种高准确度和推理速度的分步非自回归自动语音识别架构
作者:Ziyang Zhuang,Chenfeng Miao,Kun Zou,Shuai Gong,Ming Fang,Tao Wei,Zijian Li,Wei Hu,Shaojun Wang,Jing Xiao
链接:点击下载PDF文件
摘要:非自回归(NAR)自动语音识别(ASR)模型独立、同时预测令牌,带来高推理速度。然而,与自回归(AR)模型相比,NAR模型的准确性仍然存在差距。为了进一步缩小NAR和AR模型之间的差距,我们提出了一种具有高准确性和推理速度的单步NAR ASR架构,称为EfficientASR。它使用基于索引映射向量(IMV)的对齐生成器来在训练期间生成对齐,并使用对齐预测器来学习对齐以进行推理。它可以在交叉熵损失与对齐损失相结合的情况下进行端到端(E2 E)训练。与最先进的(SOTA)模型相比,所提出的EfficientASR在AISHELL-1和AISHELL-2基准测试中取得了有竞争力的结果。具体来说,它在AISHELL-1 dev test数据集上实现了4.26% 4.62%的字符错误率(CER),比SOTA AR Conformer的推理速度提高了约30倍。摘要:Non-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. To further narrow the gap between the NAR and AR models, we propose a single-step NAR ASR architecture with high accuracy and inference speed, called EfficientASR. It uses an Index Mapping Vector (IMV) based alignment generator to generate alignments during training, and an alignment predictor to learn the alignments for inference. It can be trained end-to-end (E2E) with cross-entropy loss combined with alignment loss. The proposed EfficientASR achieves competitive results on the AISHELL-1 and AISHELL-2 benchmarks compared to the state-of-the-art (SOTA) models. Specifically, it achieves character error rates (CER) of 4.26% 4.62% on the AISHELL-1 dev test dataset, which outperforms the SOTA AR Conformer with about 30x inference speedup.

【28】 Interpretable Temporal Class Activation Representation for Audio Spoofing Detection
标题: 用于音频欺骗检测的可解释时间类激活表示
作者:Menglu Li,Xiao-Ping Zhang
备注:10 pages, 5 figures, Accepted to Interspeech2024
链接:点击下载PDF文件
摘要:解释音频欺骗检测模型做出的决定对于培养对检测结果的信任至关重要。然而,目前对检测模型可解释性的研究仅限于将XAI工具应用于后训练模型。在本文中,我们利用wav 2 vec 2.0模型和专注的话语级功能直接集成到模型的架构的可解释性,从而提高决策过程的透明度。具体来说,我们提出了一个类激活表示本地化的歧视性帧有助于检测。此外,我们还证明了基于欺骗类型的多标签训练,而不是像善意和欺骗那样的二进制标签,使模型能够学习不同攻击的不同特征,从而显着提高检测性能。我们的模型实现了最先进的结果,ASVspoof 2019-LA集的EER为0.51%,最小t-DCF为0.0165。摘要:Explaining the decisions made by audio spoofing detection models is crucial for fostering trust in detection outcomes. However, current research on the interpretability of detection models is limited to applying XAI tools to post-trained models. In this paper, we utilize the wav2vec 2.0 model and attentive utterance-level features to integrate interpretability directly into the model's architecture, thereby enhancing transparency of the decision-making process. Specifically, we propose a class activation representation to localize the discriminative frames contributing to detection. Furthermore, we demonstrate that multi-label training based on spoofing types, rather than binary labels as bonafide and spoofed, enables the model to learn distinct characteristics of different attacks, significantly improving detection performance. Our model achieves state-of-the-art results, with an EER of 0.51% and a min t-DCF of 0.0165 on the ASVspoof2019-LA set.

【29】 Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
标题: 通过为预训练的多说话者文本转语音系统分配语音输入即兴生成说话者
作者:Zhengyang Chen,Xuechen Liu,Erica Cooper,Junichi Yamagishi,Yanmin Qian
备注:Accepted for presentation at Interspeech 2024 (with more analysis in the final Appendix part)
链接:点击下载PDF文件
摘要:本文提出了一种语音合成系统,允许用户指定和控制的扬声器的声学特性的提示描述合成语音的扬声器的特点。与以前的方法不同,我们的方法利用听众的印象来构建提示,这更容易收集和更自然地与日常描述的扬声器特质。我们采用低秩自适应(LoRA)技术来快速定制预训练的语言模型,以满足我们的需求,便于从提示文本中提取说话者相关特征。此外,与其他语音驱动的文本到语音(TTS)系统不同,我们将语音到说话人模块从多说话人TTS系统中分离出来,增强了系统的灵活性和与各种预训练的多说话人TTS系统的兼容性。此外,对于说话人-说话人特征模块,我们还比较了判别方法和基于流匹配的生成方法,我们发现,结合这两种方法可以帮助系统同时捕获说话人相关的信息,从提示更好地生成语音具有更高的保真度。摘要:This paper proposes a speech synthesis system that allows users to specify and control the acoustic characteristics of a speaker by means of prompts describing the speaker's traits of synthesized speech. Unlike previous approaches, our method utilizes listener impressions to construct prompts, which are easier to collect and align more naturally with everyday descriptions of speaker traits. We adopt the Low-rank Adaptation (LoRA) technique to swiftly tailor a pre-trained language model to our needs, facilitating the extraction of speaker-related traits from the prompt text. Besides, different from other prompt-driven text-to-speech (TTS) systems, we separate the prompt-to-speaker module from the multi-speaker TTS system, enhancing system flexibility and compatibility with various pre-trained multi-speaker TTS systems. Moreover, for the prompt-to-speaker characteristic module, we also compared the discriminative method and flow-matching based generative method and we found that combining both methods can help the system simultaneously capture speaker-related information from prompts better and generate speech with higher fidelity.

【30】 Are we there yet? A brief survey of Music Emotion Prediction Datasets, Models and Outstanding Challenges
标题: 我们到了吗?音乐情感预测数据集、模型和突出挑战的简要调查
作者:Jaeyong Kang,Dorien Herremans
链接:点击下载PDF文件
摘要:在过去的几年里,音乐的深度学习模型取得了巨大的进步。但是,如今机器学习模型在捕捉情感方面有多好,研究人员面临着哪些挑战?在本文中,我们提供了一个全面的概述,现有的音乐情感数据集,并讨论了评估标准以及在该领域的竞争。我们还简要概述了多年来建立的各种类型的音乐情感预测模型,为该领域的各种方法提供了见解。通过这次考试,我们强调坚持准确捕捉音乐中的情感的挑战。认识到这一领域的动态性质,我们用一个附带的GitHub存储库补充了我们的发现。该存储库包含音乐情感数据集和最新预测模型的全面列表。摘要:Deep learning models for music have advanced drastically in the last few years. But how good are machine learning models at capturing emotion these days and what challenges are researchers facing? In this paper, we provide a comprehensive overview of the available music-emotion datasets and discuss evaluation standards as well as competitions in the field. We also provide a brief overview of various types of music emotion prediction models that have been built over the years, offering insights into the diverse approaches within the field. Through this examination, we highlight the challenges that persist in accurately capturing emotion in music. Recognizing the dynamic nature of this field, we have complemented our findings with an accompanying GitHub repository. This repository contains a comprehensive list of music emotion datasets and recent predictive models.

【31】 Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
标题: 来自生成基础模型的合成音频可以帮助音频识别和语音建模吗?
作者:Tiantian Feng,Dimitrios Dimitriadis,Shrikanth Narayanan
备注:Accepted to 2024 INTERSPEECH
链接:点击下载PDF文件
摘要:基础模型的最新进展使音频生成模型能够产生与音乐,事件和人类行为相关的高保真声音。尽管现代音频生成模型取得了成功,但评估音频生成质量的传统方法在很大程度上依赖于距离度量,如Frechet音频距离。相比之下,我们的目标是通过检查使用它们作为训练数据的有效性来评估音频生成的质量。具体来说,我们进行研究,探索使用合成音频的音频识别。此外,我们研究合成音频是否可以作为语音相关建模中的数据增强资源。我们的综合实验证明了使用合成音频进行音频识别和语音相关建模的潜力。我们的代码可以在https: github.com usc-sail SynthAudio上找到。摘要:Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional approach to assessing the quality of the audio generation relies heavily on distance metrics like Frechet Audio Distance. In contrast, we aim to evaluate the quality of audio generation by examining the effectiveness of using them as training data. Specifically, we conduct studies to explore the use of synthetic audio for audio recognition. Moreover, we investigate whether synthetic audio can serve as a resource for data augmentation in speech-related modeling. Our comprehensive experiments demonstrate the potential of using synthetic audio for audio recognition and speech-related modeling. Our code is available at https: github.com usc-sail SynthAudio.

【32】 MFF-EINV2: Multi-scale Feature Fusion across Spectral-Spatial-Temporal Domains for Sound Event Localization and Detection
标题: MFF-EINV 2:跨频谱-空间-时间域的多尺度特征融合,用于声音事件定位和检测
作者:Da Mu,Zhicheng Zhang,Haobo Yue
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:声音事件定位和检测(SELD)涉及使用多声道声音记录来检测和定位声音事件。先前提出的事件无关网络V2(EINV 2)在SELD上取得了出色的性能。然而,它仍然面临着有效地提取跨光谱,空间和时间域的特征的挑战。本文提出了一个三阶段的网络结构命名为多尺度特征融合(MFF)模块,以充分提取跨光谱,空间和时间域的多尺度特征。MFF模块采用并行子网络结构生成多尺度光谱和空间特征。TF-卷积模块用于提供多尺度时间特征。我们将MFF合并到EINV 2中,并将所提出的方法称为MFF-EINV 2。在2022年和2023年DCASE挑战任务3数据集上的实验结果表明了我们的MFF-EINV 2的有效性,与已发表的方法相比,它实现了最先进的(SOTA)性能。摘要:Sound Event Localization and Detection (SELD) involves detecting and localizing sound events using multichannel sound recordings. Previously proposed Event-Independent Network V2 (EINV2) has achieved outstanding performance on SELD. However, it still faces challenges in effectively extracting features across spectral, spatial, and temporal domains. This paper proposes a three-stage network structure named Multi-scale Feature Fusion (MFF) module to fully extract multi-scale features across spectral, spatial, and temporal domains. The MFF module utilizes parallel subnetworks architecture to generate multi-scale spectral and spatial features. The TF-Convolution Module is employed to provide multi-scale temporal features. We incorporated MFF into EINV2 and term the proposed method as MFF-EINV2. Experimental results in 2022 and 2023 DCASE challenge task3 datasets show the effectiveness of our MFF-EINV2, which achieves state-of-the-art (SOTA) performance compared to published methods.

【33】 VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
标题: Visinger 2+:通过自我监督学习表示增强的端到端歌唱声音合成
作者:Yifeng Yu,Jiatong Shi,Yuning Wu,Shinji Watanabe
备注:4 pages, 2 figures
链接:点击下载PDF文件
摘要:随着深度学习技术的出现,歌唱语音合成(SVS)取得了重大进展。然而,SVS中的一个重大挑战是标记的歌唱声音数据的稀缺性,这限制了监督学习方法的有效性。为了应对这一挑战,本文介绍了一种新的方法,通过利用来自预训练的自监督学习模型的未标记数据来提高SVS的质量。在现有VISinger2框架的基础上,本研究将额外的光谱特征信息集成到系统中,以提高其性能。该集成旨在利用来自预训练模型的丰富声学特征,从而丰富合成并产生更自然和更具表现力的歌声。在不同语料库中的实验结果表明,该方法在提高合成歌声的客观和主观指标的整体质量的有效性。摘要:Singing Voice Synthesis (SVS) has witnessed significant advancements with the advent of deep learning techniques. However, a significant challenge in SVS is the scarcity of labeled singing voice data, which limits the effectiveness of supervised learning methods. In response to this challenge, this paper introduces a novel approach to enhance the quality of SVS by leveraging unlabeled data from pre-trained self-supervised learning models. Building upon the existing VISinger2 framework, this study integrates additional spectral feature information into the system to enhance its performance. The integration aims to harness the rich acoustic features from the pre-trained models, thereby enriching the synthesis and yielding a more natural and expressive singing voice. Experimental results in various corpora demonstrate the efficacy of this approach in improving the overall quality of synthesized singing voices in both objective and subjective metrics.

【34】 TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
标题: TSE-PI:具有音调信息的回响环境下的目标声音提取
作者:Yiwen Wang,Xihong Wu
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:目标声音提取(TSE)根据提供的线索从混合信号中分离出目标声音。然而,现有模型的性能显着降低混响条件下。受听觉场景分析(ASA)的启发,本文提出了一种具有基音信息的TSE模型TSE-PI。条件音高提取是通过具有声音类别标签的逐行线性调制层来实现的。一个修改后的波形模型结合音高信息,采用一个可学习的伽玛通滤波器组的卷积编码器的地方,用于目标声音提取。包含音高信息的目的是提高模型的性能。在FSD 50 K数据集上的实验结果表明,当结合音高信息和伽马通滤波器组时,混响环境下的目标声音提取提高了2.4 dB。摘要:Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis (ASA), this work proposes a TSE model provided with pitch information named TSE-PI. Conditional pitch extraction is achieved through the Feature-wise Linearly Modulated layer with the sound-class label. A modified Waveformer model combined with pitch information, employing a learnable Gammatone filterbank in place of the convolutional encoder, is used for target sound extraction. The inclusion of pitch information is aimed at improving the model's performance. The experimental results on the FSD50K dataset illustrate 2.4 dB improvements of target sound extraction under reverberant environments when incorporating pitch information and Gammatone filterbank.

【35】 ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
标题: ML-SURB 2.0:跨建模约束、语言和数据集的多语言语音模型基准测试
作者:Jiatong Shi,Shih-Heng Wang,William Chen,Martijn Bartelds,Vanya Bannihatti Kumar,Jinchuan Tian,Xuankai Chang,Dan Jurafsky,Karen Livescu,Hung-yi Lee,Shinji Watanabe
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:ML-SUPERB评估语言识别和自动语音识别(ASR)任务的自监督学习(SSL)模型。该基准测试将模型视为特征提取器,并使用单个浅下游模型,该模型可以针对下游任务进行微调。然而,现实世界的用例可能需要不同的配置。本文介绍了ML-SUPERB~2.0,这是一个新的基准,用于跨下游模型,微调设置和有效的模型自适应方法评估预训练的SSL和监督语音模型。我们发现ML-SUPERB的设置性能有所改善。然而,性能取决于下游模型设计。此外,我们发现语言和数据集之间存在很大的性能差异,这表明需要更有针对性的方法来提高多语言ASR性能。摘要:ML-SUPERB evaluates self-supervised learning (SSL) models on the tasks of language identification and automatic speech recognition (ASR). This benchmark treats the models as feature extractors and uses a single shallow downstream model, which can be fine-tuned for a downstream task. However, real-world use cases may require different configurations. This paper presents ML-SUPERB~2.0, which is a new benchmark for evaluating pre-trained SSL and supervised speech models across downstream models, fine-tuning setups, and efficient model adaptation approaches. We find performance improvements over the setup of ML-SUPERB. However, performance depends on the downstream model design. Also, we find large performance differences between languages and datasets, suggesting the need for more targeted approaches to improve multilingual ASR performance.

【36】 Emotion Manipulation Through Music -- A Deep Learning Interactive Visual Approach
标题: 通过音乐操纵情绪--深度学习交互视觉方法
作者:Adel N. Abdalla,Jared Osborne,Razvan Andonie
链接:点击下载PDF文件
摘要:音乐能唤起许多人的感情。我们介绍了一种使用AI工具操纵歌曲情感内容的新方法。我们的目标是达到所需的情感,同时保持原始旋律尽可能完整。为此,我们创建了一个交互式管道,能够将输入的歌曲转换为完全相反的情感,并通过Russel的Circumplex模型将此结果可视化。我们的方法是音乐语义操纵的概念验证,这是一个旨在修改现有音乐情感内容的新领域。我们设计了一个深度学习模型,能够评估我们对键、SoundFont乐器和其他音乐功能的修改的准确性。我们模型的准确性与4Q Emotion数据集上的最新技术水平一致。随着进一步的完善,这项研究可能有助于按需定制音乐的生成,现有作品的自动混音,以及为情感发展调整的音乐播放列表。摘要:Music evokes emotion in many people. We introduce a novel way to manipulate the emotional content of a song using AI tools. Our goal is to achieve the desired emotion while leaving the original melody as intact as possible. For this, we create an interactive pipeline capable of shifting an input song into a diametrically opposed emotion and visualize this result through Russel's Circumplex model. Our approach is a proof-of-concept for Semantic Manipulation of Music, a novel field aimed at modifying the emotional content of existing music. We design a deep learning model able to assess the accuracy of our modifications to key, SoundFont instrumentation, and other musical features. The accuracy of our model is in-line with the current state of the art techniques on the 4Q Emotion dataset. With further refinement, this research may contribute to on-demand custom music generation, the automated remixing of existing work, and music playlists tuned for emotional progression.

【37】 Self-Supervised Speech Representations are More Phonetic than Semantic
标题: 自我监督的言语表达更多的是语音而不是语义
作者:Kwanghee Choi,Ankita Pasad,Tomohiko Nakamura,Satoru Fukayama,Karen Livescu,Shinji Watanabe
备注:Accepted to Interspeech 2024. Source code at this https URL
链接:点击下载PDF文件
摘要:自监督语音模型(Self-supervised Speech Models,S3 M)已成为语音应用的有效支撑。各种分析表明,S3 Ms编码语言属性。在这项工作中,我们寻求一个更细粒度的分析编码在S3 Ms的单词级的语言属性。具体来说,我们策划了一个新的近同音字(语音相似)和同义词(语义相似)词对数据集,并测量S3 M词表示对之间的相似性。我们的研究表明,S3 M表示一致,显着表现出更多的语音比语义相似性。此外,我们质疑是否广泛使用的意图分类数据集,如流利的语音命令和Snips Smartlights,足以衡量语义能力。我们的简单基线,只使用身份这个词,超越了基于S3 M的模型。这证实了我们的发现,并表明这些数据集上的高分并不一定保证语义内容的存在。摘要:Self-supervised speech models (S3Ms) have become an effective backbone for speech applications. Various analyses suggest that S3Ms encode linguistic properties. In this work, we seek a more fine-grained analysis of the word-level linguistic properties encoded in S3Ms. Specifically, we curate a novel dataset of near homophone (phonetically similar) and synonym (semantically similar) word pairs and measure the similarities between S3M word representation pairs. Our study reveals that S3M representations consistently and significantly exhibit more phonetic than semantic similarity. Further, we question whether widely used intent classification datasets such as Fluent Speech Commands and Snips Smartlights are adequate for measuring semantic abilities. Our simple baseline, using only the word identity, surpasses S3M-based models. This corroborates our findings and suggests that high scores on these datasets do not necessarily guarantee the presence of semantic content.

【38】 Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
标题: 通过文本到发音障碍语音合成增强训练数据用于发音障碍自动语音识别
作者:Wing-Zin Leung,Mattias Cross,Anton Ragni,Stefan Goetze
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,自动语音识别(ASR)的研究取得了令人印象深刻的成绩,并在增强和替代通信(AAC)和家庭环境系统中为构音障碍(PwD)患者提供接入方面具有巨大的潜力。然而,构音障碍性ASR(DASR)的进展受到构音障碍性言语的高度变异性和构音障碍训练数据的有限公共可用性的限制。本文证明了使用文本到构音障碍语音(TTDS)合成来微调大型ASR模型的数据增强对DASR是有效的。具体来说,基于扩散的文本到语音(TTS)模型可以产生类似于构音障碍语音的语音样本,这些语音样本可以用作微调ASR基础模型的额外训练数据,在这种情况下是Whisper。结果表明,改进的合成指标和ASR性能的建议多扬声器扩散为基础的TTDS数据增强ASR微调相比,目前的DASR基线。摘要:Automatic speech recognition (ASR) research has achieved impressive performance in recent years and has significant potential for enabling access for people with dysarthria (PwD) in augmentative and alternative communication (AAC) and home environment systems. However, progress in dysarthric ASR (DASR) has been limited by high variability in dysarthric speech and limited public availability of dysarthric training data. This paper demonstrates that data augmentation using text-to-dysarthic-speech (TTDS) synthesis for finetuning large ASR models is effective for DASR. Specifically, diffusion-based text-to-speech (TTS) models can produce speech samples similar to dysarthric speech that can be used as additional training data for fine-tuning ASR foundation models, in this case Whisper. Results show improved synthesis metrics and ASR performance for the proposed multi-speaker diffusion-based TTDS data augmentation for ASR fine-tuning compared to current DASR baselines.


机器翻译,仅供参考