【1】Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio标题:利用预先训练的自动编码器进行音乐音频的可解释原型学习链接:https://arxiv.org/abs/2402.09318作者:Pablo Alonso-Jiménez,Leonardo Pepino,Roser Batlle-Roca,Pablo Zinemanas,Dmitry Bogdanov,Xavier Serra,Martín Rocamora摘要:我们提出了PECMAE,一个基于原型学习的音乐音频分类的可解释模型。我们的模型基于以前的方法APNet,它联合学习自动编码器和原型网络。相反,我们建议将两个训练过程解耦。这使我们能够利用现有的自监督自动编码器在更大的数据(EnCodecMAE)上进行预训练,提供具有更好泛化能力的表示。APNet允许原型重构为波形,以实现依赖于最近的训练数据样本的可解释性。相比之下,我们探索使用扩散解码器,允许重建没有这种依赖性。我们在乐器分类(Medley-Solos-DB)和流派识别(GTZAN和一个更大的内部数据集)的数据集上评估了我们的方法,后者是一个更具挑战性的任务,以前没有用原型网络解决。我们发现基于原型的模型保留了自动编码器嵌入所实现的大部分性能,而原型的发音有利于理解分类器的行为。摘要:We present PECMAE, an interpretable model for music audio classification based on prototype learning. Our model is based on a previous method, APNet, which jointly learns an autoencoder and a prototypical network. Instead, we propose to decouple both training processes. This enables us to leverage existing self-supervised autoencoders pre-trained on much larger data (EnCodecMAE), providing representations with better generalization. APNet allows prototypes' reconstruction to waveforms for interpretability relying on the nearest training data samples. In contrast, we explore using a diffusion decoder that allows reconstruction without such dependency. We evaluate our method on datasets for music instrument classification (Medley-Solos-DB) and genre recognition (GTZAN and a larger in-house dataset), the latter being a more challenging task not addressed with prototypical networks before. We find that the prototype-based models preserve most of the performance achieved with the autoencoder embeddings, while the sonification of prototypes benefits understanding the behavior of the classifier. 【2】 An Embarrassingly Simple Approach for LLM with Strong ASR Capacity标题:对于拥有强大ASR能力的LLM来说,一种令人尴尬的简单方法链接:https://arxiv.org/abs/2402.08846作者:Ziyang Ma,Guanrou Yang,Yifan Yang,Zhifu Gao,Jiaming Wang,Zhihao Du,Fan Yu,Qian Chen,Siqi Zheng,Shiliang Zhang,Xie Chen备注:Working in progress and will open-source soon摘要:在本文中,我们专注于解决语音处理领域中最重要的任务之一,即,自动语音识别(ASR),具有语音基础编码器和大型语言模型(LLM)。最近的作品有复杂的设计,如压缩输出时间的语音编码器,解决模态对准的投影仪,并利用参数有效的微调的LLM。我们发现,微妙的设计是没有必要的,而一个非常简单的组成现成的语音编码器,LLM,和唯一的可训练线性投影仪是胜任的ASR任务。更具体地说,我们对LLM和语音编码器的各种组合进行了基准测试和探索,从而得到了基于LLM的最佳ASR系统,我们称之为SLAM-ASR。所提出的SLAM-ASR提供了一个干净的设置和很少的特定于任务的设计,其中只有线性投影仪被训练。据我们所知,SLAM-ASR在基于LLM的ASR模型中在Libripeech基准测试中取得了最佳性能,甚至优于在大量配对数据上训练的最新基于LLM的音频通用模型。最后,我们探讨了基于LLM的ASR在模态对齐过程中的能力涌现。我们希望我们的研究可以促进研究扩展LLM与跨模态的能力,并揭示了基于LLM的ASR社区。摘要:In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Recent works have complex designs such as compressing the output temporally for the speech encoder, tackling modal alignment for the projector, and utilizing parameter-efficient fine-tuning for the LLM. We found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for the ASR task. To be more specific, we benchmark and explore various combinations of LLMs and speech encoders, leading to the optimal LLM-based ASR system, which we call SLAM-ASR. The proposed SLAM-ASR provides a clean setup and little task-specific design, where only the linear projector is trained. To the best of our knowledge, SLAM-ASR achieves the best performance on the Librispeech benchmark among LLM-based ASR models and even outperforms the latest LLM-based audio-universal model trained on massive pair data. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community.
【3】 Syllable based DNN-HMM Cantonese Speech to Text System标题:基于音节的DNN-HMM粤语语音文本转换系统链接:https://arxiv.org/abs/2402.08788作者:Timothy Wong,Claire Li,Sam Lam,Billy Chiu,Qin Lu,Minglei Li,Dan Xiong,Roy Shing Yu,Vincent T. Y. Ng备注:7 pages, 3 figures, LREC 2016摘要:本文介绍了一个基于音节声学模型的粤语语音到文本(STT)转换系统。这是建立一个STT系统的努力的一部分,以帮助那些在写作技能上有认知缺陷,但通过演讲表达思想没有问题的阅读障碍学生。在粤语语音识别中,声学模型的基本单位可以是传统的IF音节,也可以是ONC音节,ONC音节将韵母进一步分为核心和结尾,以反映粤语音节内的变化。通过使用Kaldi工具包,我们的系统在GPU的帮助下使用随机梯度下降优化模型进行训练,用于混合深度神经网络和隐马尔可夫模型(DNN-HMM),具有和不具有基于I向量的说话人自适应训练技术。在所有情况下,使用具有说话人自适应训练的相同高斯混合模型(GMM-SAT)到DNN的输入特征。实验结果表明,基于ONC的音节声学建模与基于I向量的DNN-HMM取得了最好的性能,字错误率(WER)为9.66%,实时因子(RTF)为1.38812。摘要:This paper reports our work on building up a Cantonese Speech-to-Text (STT) system with a syllable based acoustic model. This is a part of an effort in building a STT system to aid dyslexic students who have cognitive deficiency in writing skills but have no problem expressing their ideas through speech. For Cantonese speech recognition, the basic unit of acoustic models can either be the conventional Initial-Final (IF) syllables, or the Onset-Nucleus-Coda (ONC) syllables where finals are further split into nucleus and coda to reflect the intra-syllable variations in Cantonese. By using the Kaldi toolkit, our system is trained using the stochastic gradient descent optimization model with the aid of GPUs for the hybrid Deep Neural Network and Hidden Markov Model (DNN-HMM) with and without I-vector based speaker adaptive training technique. The input features of the same Gaussian Mixture Model with speaker adaptive training (GMM-SAT) to DNN are used in all cases. Experiments show that the ONC-based syllable acoustic modeling with I-vector based DNN-HMM achieves the best performance with the word error rate (WER) of 9.66% and the real time factor (RTF) of 1.38812.
【4】 If Turing played piano with an artificial partner标题:如果图灵和一个人造舞伴一起弹钢琴链接:https://arxiv.org/abs/2402.08690作者:Dobromir Dotov,Dante Camarena,Zack Harris,Joanna Spyra,Pietro Gagliano,Laurel Trainor摘要:音乐是一种固有的社会活动,它允许人们分享经验并感受彼此之间的联系。在设计一种与另一个人玩耍时表现出类似社会体验的人造伙伴方面,进展甚微。实现生成模型(如大型语言模型)的神经网络架构适用于生成乐谱。然而,社会性地演奏音乐不仅仅是演奏乐谱;它必须补充其他音乐家的想法并正确地保持时间。我们解决了一个问题,即一个经过训练的生成模型是否可以产生令人信服的社会体验,而不一定是为了同步和连续而优化的。该网络是一种在大型数字乐谱语料库上训练的变分自动编码器,适用于与人类合作伙伴进行定时呼叫和响应任务。参与者与人类或人工伙伴以各种配置演奏钢琴,并对演奏质量和自我-他人整合的第一人称体验进行评级。总的来说,人工伴侣有希望,但评级低于人类伴侣。具有最简单设计和最高相似性参数的人工伙伴在某些指标上与人类伙伴的评分没有差异,这表明交互而不是生成复杂性在实现社交AI方面很重要。摘要:Music is an inherently social activity that allows people to share experiences and feel connected with one another. There has been little progress in designing artificial partners exhibiting a similar social experience as playing with another person. Neural network architectures that implement generative models, such as large language models, are suited for producing musical scores. Playing music socially, however, involves more than playing a score; it must complement the other musicians' ideas and keep time correctly. We addressed the question of whether a convincing social experience is made possible by a generative model trained to produce musical scores, not necessarily optimized for synchronization and continuation. The network, a variational autoencoder trained on a large corpus of digital scores, was adapted for a timed call-and-response task with a human partner. Participants played piano with a human or artificial partner-in various configurations-and rated the performance quality and first-person experience of self-other integration. Overall, the artificial partners held promise but were rated lower than human partners. The artificial partner with simplest design and highest similarity parameter was not rated differently from the human partners on some measures, suggesting that interactive rather than generative sophistication is important in enabling social AI. 【5】 MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech标题:MobileSpeech:一种快速高保真的移动Zero-Shot文语转换框架链接:https://arxiv.org/abs/2402.09378作者:Shengpeng Ji,Ziyue Jiang,Hanting Wang,Jialong Zuo,Zhou Zhao摘要:Zero-shot文本到语音(TTS)由于其强大的语音克隆功能而受到广泛关注,只需要几秒钟看不见的扬声器语音提示。然而,所有以前的工作都是针对基于云的系统开发的。以自回归模型为例,虽然这些方法实现了高保真的语音克隆,但它们在推理速度,模型大小和鲁棒性方面存在不足。因此,我们提出了MobileSpeech,这是一个快速,轻量级的,和鲁棒的zero-shot文本到语音系统的移动设备的基础上的第一次。具体而言:1)利用离散编解码器,设计了一个并行语音掩码解码器模块SMD,该模块在生成过程中融合了来自语音编解码器的层次信息和不同编解码器层的权重机制。此外,为了弥合文本和语音之间的差距,我们引入了一个高层次的概率掩码,模拟语音生成过程中信息流从少到多的进展。2)对于说话人提示,我们从提示语音中提取细粒度的提示持续时间,并在SMD中通过交叉注意将文本,提示语音合并。我们展示了MobileSpeech在不同级别的多语言数据集上的有效性,在生成速度和语音质量方面取得了最先进的结果。MobileSpeech在单个A100 GPU上实现了0.09的RTF,我们已经成功地在移动设备上部署了MobileSpeech。音频示例可在\url{https://mobilespeech.github.io/}上获得。摘要:Zero-shot text-to-speech (TTS) has gained significant attention due to its powerful voice cloning capabilities, requiring only a few seconds of unseen speaker voice prompts. However, all previous work has been developed for cloud-based systems. Taking autoregressive models as an example, although these approaches achieve high-fidelity voice cloning, they fall short in terms of inference speed, model size, and robustness. Therefore, we propose MobileSpeech, which is a fast, lightweight, and robust zero-shot text-to-speech system based on mobile devices for the first time. Specifically: 1) leveraging discrete codec, we design a parallel speech mask decoder module called SMD, which incorporates hierarchical information from the speech codec and weight mechanisms across different codec layers during the generation process. Moreover, to bridge the gap between text and speech, we introduce a high-level probabilistic mask that simulates the progression of information flow from less to more during speech generation. 2) For speaker prompts, we extract fine-grained prompt duration from the prompt speech and incorporate text, prompt speech by cross attention in SMD. We demonstrate the effectiveness of MobileSpeech on multilingual datasets at different levels, achieving state-of-the-art results in terms of generating speed and speech quality. MobileSpeech achieves RTF of 0.09 on a single A100 GPU and we have successfully deployed MobileSpeech on mobile devices. Audio samples are available at \url{https://mobilespeech.github.io/} .
【6】 Listening to Multi-talker Conversations: Modular and End-to-end Perspectives标题:聆听多人对话:模块化和端到端视角链接:https://arxiv.org/abs/2402.08932作者:Desh Raj备注:Ph.D. dissertation摘要:自30多年前第一个语音识别系统问世以来,语音技术的改进已经使智能助理和自动化客户支持等应用成为可能。然而,未来的对话智能需要识别自由流动的多方对话,这是一个关键和具有挑战性的组成部分,仍然没有解决。在这篇论文中,我们专注于这个问题的说话人属性的多说话人语音识别,并提出了两个观点,导致其概率公式。 在模块化的角度来看,我们建立了一个管道的子任务,包括说话人日记,目标说话人提取,语音识别。我们的第一个贡献是一种方法来执行的约束优化问题,通过重新制定的谱聚类的感知日志。我们还描述了一个算法,合奏日记输出,无论是结合感知系统或后期融合执行多通道日记。一旦扬声器段被识别,我们鲁棒地提取单扬声器话语的混合使用GPU加速实现的引导源分离,这使我们能够使用现成的ASR系统来获得扬声器属性的成绩单。 由于模块化的方法遭受错误传播,我们提出了一个替代的“端到端”的角度来看这个问题。为此,我们描述了流式解混和识别传感器(SURT)。我们展示了如何通过仔细设计网络架构、目标函数和混合仿真技术来有效地训练SURT模型。最后,我们增加了一个辅助扬声器分支,使联合预测的扬声器标签同步的语音令牌。我们证明,在合成混合物上进行训练并使用真实数据进行调整有助于这些模型很好地传输真实会议会话的流式转录。摘要:Since the first speech recognition systems were built more than 30 years ago, improvement in voice technology has enabled applications such as smart assistants and automated customer support. However, conversation intelligence of the future requires recognizing free-flowing multi-party conversations, which is a crucial and challenging component that still remains unsolved. In this dissertation, we focus on this problem of speaker-attributed multi-talker speech recognition, and propose two perspectives which result from its probabilistic formulation. In the modular perspective, we build a pipeline of sub-tasks involving speaker diarization, target speaker extraction, and speech recognition. Our first contribution is a method to perform overlap-aware diarization by reformulating spectral clustering as a constrained optimization problem. We also describe an algorithm to ensemble diarization outputs, either to combine overlap-aware systems or to perform multi-channel diarization by late fusion. Once speaker segments are identified, we robustly extract single-speaker utterances from the mixture using a GPU-accelerated implementation of guided source separation, which allows us to use an off-the-shelf ASR system to obtain speaker-attributed transcripts. Since the modular approach suffers from error propagation, we propose an alternate "end-to-end" perspective on the problem. For this, we describe the Streaming Unmixing and Recognition Transducer (SURT). We show how to train SURT models efficiently by carefully designing the network architecture, objective functions, and mixture simulation techniques. Finally, we add an auxiliary speaker branch to enable joint prediction of speaker labels synchronized with the speech tokens. We demonstrate that training on synthetic mixtures and adapting with real data helps these models transfer well for streaming transcription of real meeting sessions.
【7】 Sound Field Reconstruction Using a Compact Acoustics-informed Neural Network标题:基于紧凑型声学信息神经网络的声场重建链接:https://arxiv.org/abs/2402.08904作者:Fei Ma,Sipei Zhao,Ian S. Burnett摘要:声场重建(SFR)增强了由麦克风阵列捕获的声场的信息。使用基函数分解的常规SFR方法是直接的且计算高效的,但是可能需要比测量声场所需的更多的麦克风。最近的研究表明,纯数据驱动和基于学习的方法在一些SFR任务中很有前途,但它们通常计算量很大,并且可能无法重建物理上有效的声场。本文提出了一种紧凑的声学通知的神经网络(AINN)方法SFR,利用亥姆霍兹方程来正则化神经网络。与纯粹依赖于测量声压的纯数据驱动方法相反,亥姆霍兹方程的集成提高了神经网络对测量过程中变化的鲁棒性,并促使生成物理上有效的重建。AINN被设计为紧凑的,并且能够基于沿边界测量的声压来预测感兴趣的空间区域内的声压和声压梯度。通过对不同环境下测量的声学传递函数进行数值实验,验证了AINN方法相对于传统的圆柱谐波分解和奇异值分解方法的优越性。摘要:Sound field reconstruction (SFR) augments the information of a sound field captured by a microphone array. Conventional SFR methods using basis function decomposition are straightforward and computationally efficient, but may require more microphones than needed to measure the sound field. Recent studies show that pure data-driven and learning-based methods are promising in some SFR tasks, but they are usually computationally heavy and may fail to reconstruct a physically valid sound field. This paper proposes a compact acoustics-informed neural network (AINN) method for SFR, whereby the Helmholtz equation is exploited to regularize the neural network. As opposed to pure data-driven approaches that solely rely on measured sound pressures, the integration of the Helmholtz equation improves robustness of the neural network against variations during the measurement processes and prompts the generation of physically valid reconstructions. The AINN is designed to be compact, and is able to predict not only the sound pressures but also sound pressure gradients within a spatial region of interest based on measured sound pressures along the boundary. Numerical experiments with acoustic transfer functions measured in different environments demonstrate the superiority of the AINN method over the traditional cylinder harmonic decomposition and the singular value decomposition methods. 【8】 UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models标题:UniEnc-CASSNAT:一种用于语音SSL模型的仅编码非自回归ASR链接:https://arxiv.org/abs/2402.08898作者:Ruchao Fan,Natarajan Balaji Shanka,Abeer Alwan备注:Published in IEEE Signal Processing Letters摘要:非自回归自动语音识别(NASR)模型由于其并行性和快速推理而受到关注。基于编码器的NASR,例如连接时间分类(CTC),可以从语音基础模型(SFM)初始化,但不考虑中间令牌之间的任何依赖性。基于编码器-解码器的NASR,如基于CTC分解的单步非自回归Transformer(CASS-NAT),可以减轻依赖性问题,但不能有效地集成SFM。受最近成功的基于Transformer编码器的语音-文本联合预训练的启发,本文提出了一种新的基于编码器的NASR(UniEnc-CASSNAT),它结合了CTC和CASS-NAT的优点。该编码器通过两次前向传递同时扮演CASS-NAT编码器和解码器的角色。编码器的第一遍接受语音信号作为输入,而语音信号和令牌级声学嵌入的级联被用作第二遍的输入。在Librispeech 100 h、MyST和Aishell 1数据集上进行了检查,所提出的UniEnc-CASSNAT实现了最先进的NASR结果,并且比CASS-NAT更好或相当,只有一个编码器,因此模型参数更少。我们的代码是公开的。摘要:Non-autoregressive automatic speech recognition (NASR) models have gained attention due to their parallelism and fast inference. The encoder-based NASR, e.g. connectionist temporal classification (CTC), can be initialized from the speech foundation models (SFM) but does not account for any dependencies among intermediate tokens. The encoder-decoder-based NASR, like CTC alignment-based single-step non-autoregressive transformer (CASS-NAT), can mitigate the dependency problem but is not able to efficiently integrate SFM. Inspired by the success of recent work of speech-text joint pre-training with a shared transformer encoder, we propose a new encoder-based NASR, UniEnc-CASSNAT, to combine the advantages of CTC and CASS-NAT. UniEnc-CASSNAT consists of only an encoder as the major module, which can be the SFM. The encoder plays the role of both the CASS-NAT encoder and decoder by two forward passes. The first pass of the encoder accepts the speech signal as input, while the concatenation of the speech signal and the token-level acoustic embedding is used as the input for the second pass. Examined on the Librispeech 100h, MyST, and Aishell1 datasets, the proposed UniEnc-CASSNAT achieves state-of-the-art NASR results and is better or comparable to CASS-NAT with only an encoder and hence, fewer model parameters. Our codes are publicly available. eess.AS音频处理【1】 MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech标题:MobileSpeech:一种快速高保真的移动Zero-Shot文语转换框架链接:https://arxiv.org/abs/2402.09378作者:Shengpeng Ji,Ziyue Jiang,Hanting Wang,Jialong Zuo,Zhou Zhao摘要:Zero-shot文本到语音(TTS)由于其强大的语音克隆功能而受到广泛关注,只需要几秒钟看不见的扬声器语音提示。然而,所有以前的工作都是针对基于云的系统开发的。以自回归模型为例,虽然这些方法实现了高保真的语音克隆,但它们在推理速度,模型大小和鲁棒性方面存在不足。因此,我们提出了MobileSpeech,这是一个快速,轻量级的,和鲁棒的zero-shot文本到语音系统的移动设备的基础上的第一次。具体而言:1)利用离散编解码器,设计了一个并行语音掩码解码器模块SMD,该模块在生成过程中融合了来自语音编解码器的层次信息和不同编解码器层的权重机制。此外,为了弥合文本和语音之间的差距,我们引入了一个高层次的概率掩码,模拟语音生成过程中信息流从少到多的进展。2)对于说话人提示,我们从提示语音中提取细粒度的提示持续时间,并在SMD中通过交叉注意将文本,提示语音合并。我们展示了MobileSpeech在不同级别的多语言数据集上的有效性,在生成速度和语音质量方面取得了最先进的结果。MobileSpeech在单个A100 GPU上实现了0.09的RTF,我们已经成功地在移动设备上部署了MobileSpeech。音频示例可在\url{https://mobilespeech.github.io/}上获得。摘要:Zero-shot text-to-speech (TTS) has gained significant attention due to its powerful voice cloning capabilities, requiring only a few seconds of unseen speaker voice prompts. However, all previous work has been developed for cloud-based systems. Taking autoregressive models as an example, although these approaches achieve high-fidelity voice cloning, they fall short in terms of inference speed, model size, and robustness. Therefore, we propose MobileSpeech, which is a fast, lightweight, and robust zero-shot text-to-speech system based on mobile devices for the first time. Specifically: 1) leveraging discrete codec, we design a parallel speech mask decoder module called SMD, which incorporates hierarchical information from the speech codec and weight mechanisms across different codec layers during the generation process. Moreover, to bridge the gap between text and speech, we introduce a high-level probabilistic mask that simulates the progression of information flow from less to more during speech generation. 2) For speaker prompts, we extract fine-grained prompt duration from the prompt speech and incorporate text, prompt speech by cross attention in SMD. We demonstrate the effectiveness of MobileSpeech on multilingual datasets at different levels, achieving state-of-the-art results in terms of generating speed and speech quality. MobileSpeech achieves RTF of 0.09 on a single A100 GPU and we have successfully deployed MobileSpeech on mobile devices. Audio samples are available at \url{https://mobilespeech.github.io/} . 【2】 Mixture to Mixture: Leveraging Close-talk Mixtures as Weak-supervision for Speech Separation标题:混合到混合:利用近距离混合作为语音分离的弱监督链接:https://arxiv.org/abs/2402.09313作者:Zhong-Qiu Wang备注:in submission摘要:我们提出了混合到混合(M2M)训练,这是一种弱监督神经语音分离算法,它利用近距离说话混合作为训练判别模型的弱监督来分离远场混合。我们的想法是,对于一个目标说话人,其近距离说话混合有一个更高的信噪比(SNR)的目标说话人比任何远场混合,因此可以用来设计一个弱监督分离。为了实现这一点,在每个训练步骤中,我们将远场混合输入到深度神经网络(DNN)中,为每个扬声器生成中间估计,并且对于每个考虑的近距离通话和远场麦克风,我们对DNN估计进行线性滤波并优化损失,以便所有扬声器的滤波估计可以总计为每个考虑的麦克风捕获的混合。在模拟混响条件下的2扬声器分离任务的评估结果表明,M2M可以有效地利用近距离说话混合作为分离远场混合的弱监督。摘要:We propose mixture to mixture (M2M) training, a weakly-supervised neural speech separation algorithm that leverages close-talk mixtures as a weak supervision for training discriminative models to separate far-field mixtures. Our idea is that, for a target speaker, its close-talk mixture has a much higher signal-to-noise ratio (SNR) of the target speaker than any far-field mixtures, and hence could be utilized to design a weak supervision for separation. To realize this, at each training step we feed a far-field mixture to a deep neural network (DNN) to produce an intermediate estimate for each speaker, and, for each of considered close-talk and far-field microphones, we linearly filter the DNN estimates and optimize a loss so that the filtered estimates of all the speakers can sum up to the mixture captured by each of the considered microphones. Evaluation results on a 2-speaker separation task in simulated reverberant conditions show that M2M can effectively leverage close-talk mixtures as a weak supervision for separating far-field mixtures.
【3】 Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality标题:L3DAS23视听扩展现实挑战赛综述链接:https://arxiv.org/abs/2402.09245作者:Christian Marinoni,Riccardo Fosco Gramaccioni,Changan Chen,Aurelio Uncini,Danilo Comminiello备注:Accepted to 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2023)摘要:ICASSP 2023的L3 DAS 23信号处理大挑战赛的主要目标是促进和支持3D音频信号处理机器学习的合作研究,特别强调延展实境应用中的3D语音增强和3D声音事件定位和检测。作为我们最新竞赛的一部分,我们提供了一个全新的数据集,它保持了L3 DAS 21和L3 DAS 22数据集的相同一般特征,但具有来自多个混响模拟环境的一阶高保真度立体声录音。此外,我们开始探索一个视听方案,通过提供这些环境的图像,通过不同的麦克风位置和方向感知。我们还为这两个任务提出了更新的基线模型,这些模型现在可以支持音频-图像耦合作为输入,并支持API来复制我们的结果。最后,我们介绍了参与者的结果。有关挑战的更多详情,请访问https://www.l3das.com/icassp2023。摘要:The primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing, with a specific emphasis on 3D speech enhancement and 3D Sound Event Localization and Detection in Extended Reality applications. As part of our latest competition, we provide a brand-new dataset, which maintains the same general characteristics of the L3DAS21 and L3DAS22 datasets, but with first-order Ambisonics recordings from multiple reverberant simulated environments. Moreover, we start exploring an audio-visual scenario by providing images of these environments, as perceived by the different microphone positions and orientations. We also propose updated baseline models for both tasks that can now support audio-image couples as input and a supporting API to replicate our results. Finally, we present the results of the participants. Further details about the challenge are available at https://www.l3das.com/icassp2023. 【4】 Listening to Multi-talker Conversations: Modular and End-to-end Perspectives标题:倾听多人对话:模块化和端到端视角链接:https://arxiv.org/abs/2402.08932作者:Desh Raj备注:Ph.D. dissertation摘要:自30多年前第一个语音识别系统问世以来,语音技术的改进已经使智能助理和自动化客户支持等应用成为可能。然而,未来的对话智能需要识别自由流动的多方对话,这是一个关键和具有挑战性的组成部分,仍然没有解决。在这篇论文中,我们专注于这个问题的说话人属性的多说话人语音识别,并提出了两个观点,导致其概率公式。 在模块化的角度来看,我们建立了一个管道的子任务,包括说话人日记,目标说话人提取,语音识别。我们的第一个贡献是一种方法来执行的约束优化问题,通过重新制定的谱聚类的感知日志。我们还描述了一个算法,合奏日记输出,无论是结合感知系统或后期融合执行多通道日记。一旦扬声器段被识别,我们鲁棒地提取单扬声器话语的混合使用GPU加速实现的引导源分离,这使我们能够使用现成的ASR系统来获得扬声器属性的成绩单。 由于模块化的方法遭受错误传播,我们提出了一个替代的“端到端”的角度来看这个问题。为此,我们描述了流式解混和识别传感器(SURT)。我们展示了如何通过仔细设计网络架构、目标函数和混合仿真技术来有效地训练SURT模型。最后,我们增加了一个辅助扬声器分支,使联合预测的扬声器标签同步的语音令牌。我们证明,在合成混合物上进行训练并使用真实数据进行调整有助于这些模型很好地传输真实会议会话的流式转录。摘要:Since the first speech recognition systems were built more than 30 years ago, improvement in voice technology has enabled applications such as smart assistants and automated customer support. However, conversation intelligence of the future requires recognizing free-flowing multi-party conversations, which is a crucial and challenging component that still remains unsolved. In this dissertation, we focus on this problem of speaker-attributed multi-talker speech recognition, and propose two perspectives which result from its probabilistic formulation. In the modular perspective, we build a pipeline of sub-tasks involving speaker diarization, target speaker extraction, and speech recognition. Our first contribution is a method to perform overlap-aware diarization by reformulating spectral clustering as a constrained optimization problem. We also describe an algorithm to ensemble diarization outputs, either to combine overlap-aware systems or to perform multi-channel diarization by late fusion. Once speaker segments are identified, we robustly extract single-speaker utterances from the mixture using a GPU-accelerated implementation of guided source separation, which allows us to use an off-the-shelf ASR system to obtain speaker-attributed transcripts. Since the modular approach suffers from error propagation, we propose an alternate "end-to-end" perspective on the problem. For this, we describe the Streaming Unmixing and Recognition Transducer (SURT). We show how to train SURT models efficiently by carefully designing the network architecture, objective functions, and mixture simulation techniques. Finally, we add an auxiliary speaker branch to enable joint prediction of speaker labels synchronized with the speech tokens. We demonstrate that training on synthetic mixtures and adapting with real data helps these models transfer well for streaming transcription of real meeting sessions. 【5】 Sound Field Reconstruction Using a Compact Acoustics-informed Neural Network标题:基于紧凑型声学信息神经网络的声场重建链接:https://arxiv.org/abs/2402.08904作者:Fei Ma,Sipei Zhao,Ian S. Burnett摘要:声场重建(SFR)增强了由麦克风阵列捕获的声场的信息。使用基函数分解的常规SFR方法是直接的且计算高效的,但是可能需要比测量声场所需的更多的麦克风。最近的研究表明,纯数据驱动和基于学习的方法在一些SFR任务中很有前途,但它们通常计算量很大,并且可能无法重建物理上有效的声场。本文提出了一种紧凑的声学通知的神经网络(AINN)方法SFR,利用亥姆霍兹方程来正则化神经网络。与纯粹依赖于测量声压的纯数据驱动方法相反,亥姆霍兹方程的集成提高了神经网络对测量过程中变化的鲁棒性,并促使生成物理上有效的重建。AINN被设计为紧凑的,并且能够基于沿边界测量的声压来预测感兴趣的空间区域内的声压和声压梯度。通过对不同环境下测量的声学传递函数进行数值实验,验证了AINN方法相对于传统的圆柱谐波分解和奇异值分解方法的优越性。摘要:Sound field reconstruction (SFR) augments the information of a sound field captured by a microphone array. Conventional SFR methods using basis function decomposition are straightforward and computationally efficient, but may require more microphones than needed to measure the sound field. Recent studies show that pure data-driven and learning-based methods are promising in some SFR tasks, but they are usually computationally heavy and may fail to reconstruct a physically valid sound field. This paper proposes a compact acoustics-informed neural network (AINN) method for SFR, whereby the Helmholtz equation is exploited to regularize the neural network. As opposed to pure data-driven approaches that solely rely on measured sound pressures, the integration of the Helmholtz equation improves robustness of the neural network against variations during the measurement processes and prompts the generation of physically valid reconstructions. The AINN is designed to be compact, and is able to predict not only the sound pressures but also sound pressure gradients within a spatial region of interest based on measured sound pressures along the boundary. Numerical experiments with acoustic transfer functions measured in different environments demonstrate the superiority of the AINN method over the traditional cylinder harmonic decomposition and the singular value decomposition methods.
【6】 UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models标题:UniEnc-CASSNAT:一种用于语音SSL模型的仅编码非自回归ASR链接:https://arxiv.org/abs/2402.08898作者:Ruchao Fan,Natarajan Balaji Shanka,Abeer Alwan备注:Published in IEEE Signal Processing Letters摘要:非自回归自动语音识别(NASR)模型由于其并行性和快速推理而受到关注。基于编码器的NASR,例如连接时间分类(CTC),可以从语音基础模型(SFM)初始化,但不考虑中间令牌之间的任何依赖性。基于编码器-解码器的NASR,如基于CTC分解的单步非自回归Transformer(CASS-NAT),可以减轻依赖性问题,但不能有效地集成SFM。受最近成功的基于Transformer编码器的语音-文本联合预训练的启发,本文提出了一种新的基于编码器的NASR(UniEnc-CASSNAT),它结合了CTC和CASS-NAT的优点。该编码器通过两次前向传递同时扮演CASS-NAT编码器和解码器的角色。编码器的第一遍接受语音信号作为输入,而语音信号和令牌级声学嵌入的级联被用作第二遍的输入。在Librispeech 100 h、MyST和Aishell 1数据集上进行了检查,所提出的UniEnc-CASSNAT实现了最先进的NASR结果,并且比CASS-NAT更好或相当,只有一个编码器,因此模型参数更少。我们的代码是公开的。摘要:Non-autoregressive automatic speech recognition (NASR) models have gained attention due to their parallelism and fast inference. The encoder-based NASR, e.g. connectionist temporal classification (CTC), can be initialized from the speech foundation models (SFM) but does not account for any dependencies among intermediate tokens. The encoder-decoder-based NASR, like CTC alignment-based single-step non-autoregressive transformer (CASS-NAT), can mitigate the dependency problem but is not able to efficiently integrate SFM. Inspired by the success of recent work of speech-text joint pre-training with a shared transformer encoder, we propose a new encoder-based NASR, UniEnc-CASSNAT, to combine the advantages of CTC and CASS-NAT. UniEnc-CASSNAT consists of only an encoder as the major module, which can be the SFM. The encoder plays the role of both the CASS-NAT encoder and decoder by two forward passes. The first pass of the encoder accepts the speech signal as input, while the concatenation of the speech signal and the token-level acoustic embedding is used as the input for the second pass. Examined on the Librispeech 100h, MyST, and Aishell1 datasets, the proposed UniEnc-CASSNAT achieves state-of-the-art NASR results and is better or comparable to CASS-NAT with only an encoder and hence, fewer model parameters. Our codes are publicly available. 【7】 Leveraging cough sounds to optimize chest x-ray usage in low-resource settings标题:利用咳嗽声音在低资源环境中优化胸部X光使用链接:https://arxiv.org/abs/2402.08789作者:Alexander Philip,Sanya Chawla,Lola Jover,George P. Kafentzis,Joe Brew,Vishakh Saraf,Shibu Vijayan,Peter Small,Carlos Chaccour摘要:胸部X光检查是呼吸道疾病分诊、诊断和管理的常用工具。在资源有限的环境中,优化这一资源可以为医疗保健系统和患者节省宝贵的成本,并改善咨询时间。我们使用了来自印度比哈尔邦普尔尼亚基督教医疗中心和医院(CMCH)的137名胸部X线检查患者的前瞻性收集数据。每名患者在等待X线摄影时至少咳嗽五次。使用声学AI方法分析收集的咳嗽声。对每个患者咳嗽声的时间和频谱特征进行交叉验证。使用标准统计方法总结特征。三个模型的开发,测试和比较,在他们的能力,以预测一个异常的结果,在胸部X射线。所有这三种方法产生的模型,可以在一定程度上区分正常和异常的逻辑回归表现最好的接收器工作特征曲线下的面积范围从0.7到0.78。尽管存在局限性和样本量相对较小,但这项研究表明,AI算法可以使用咳嗽声来预测哪些接受胸部X光检查的人会有正常或异常的结果。鉴于低收入和中等收入国家有限的卫生保健资源的潜在优化,这些结果要求扩大这项研究。摘要:Chest X-ray is a commonly used tool during triage, diagnosis and management of respiratory diseases. In resource-constricted settings, optimizing this resource can lead to valuable cost savings for the health care system and the patients as well as to and improvement in consult time. We used prospectively-collected data from 137 patients referred for chest X-ray at the Christian Medical Center and Hospital (CMCH) in Purnia, Bihar, India. Each patient provided at least five coughs while awaiting radiography. Collected cough sounds were analyzed using acoustic AI methods. Cross-validation was done on temporal and spectral features on the cough sounds of each patient. Features were summarized using standard statistical approaches. Three models were developed, tested and compared in their capacity to predict an abnormal result in the chest X-ray. All three methods yielded models that could discriminate to some extent between normal and abnormal with the logistic regression performing best with an area under the receiver operating characteristic curves ranging from 0.7 to 0.78. Despite limitations and its relatively small sample size, this study shows that AI-enabled algorithms can use cough sounds to predict which individuals presenting for chest radiographic examination will have a normal or abnormal results. These results call for expanding this research given the potential optimization of limited health care resources in low- and middle-income countries.
【8】 Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio标题:利用预先训练的自动编码器进行音乐音频的可解释原型学习链接:https://arxiv.org/abs/2402.09318作者:Pablo Alonso-Jiménez,Leonardo Pepino,Roser Batlle-Roca,Pablo Zinemanas,Dmitry Bogdanov,Xavier Serra,Martín Rocamora摘要:我们提出了PECMAE,一个基于原型学习的音乐音频分类的可解释模型。我们的模型基于以前的方法APNet,它联合学习自动编码器和原型网络。相反,我们建议将两个训练过程解耦。这使我们能够利用现有的自监督自动编码器在更大的数据(EnCodecMAE)上进行预训练,提供具有更好泛化能力的表示。APNet允许原型重构为波形,以实现依赖于最近的训练数据样本的可解释性。相比之下,我们探索使用扩散解码器,允许重建没有这种依赖性。我们在乐器分类(Medley-Solos-DB)和流派识别(GTZAN和一个更大的内部数据集)的数据集上评估了我们的方法,后者是一个更具挑战性的任务,以前没有用原型网络解决。我们发现基于原型的模型保留了自动编码器嵌入所实现的大部分性能,而原型的发音有利于理解分类器的行为。摘要:We present PECMAE, an interpretable model for music audio classification based on prototype learning. Our model is based on a previous method, APNet, which jointly learns an autoencoder and a prototypical network. Instead, we propose to decouple both training processes. This enables us to leverage existing self-supervised autoencoders pre-trained on much larger data (EnCodecMAE), providing representations with better generalization. APNet allows prototypes' reconstruction to waveforms for interpretability relying on the nearest training data samples. In contrast, we explore using a diffusion decoder that allows reconstruction without such dependency. We evaluate our method on datasets for music instrument classification (Medley-Solos-DB) and genre recognition (GTZAN and a larger in-house dataset), the latter being a more challenging task not addressed with prototypical networks before. We find that the prototype-based models preserve most of the performance achieved with the autoencoder embeddings, while the sonification of prototypes benefits understanding the behavior of the classifier.
【9】 An Embarrassingly Simple Approach for LLM with Strong ASR Capacity标题:对于拥有强大ASR能力的LLM来说,一种令人尴尬的简单方法链接:https://arxiv.org/abs/2402.08846作者:Ziyang Ma,Guanrou Yang,Yifan Yang,Zhifu Gao,Jiaming Wang,Zhihao Du,Fan Yu,Qian Chen,Siqi Zheng,Shiliang Zhang,Xie Chen备注:Working in progress and will open-source soon摘要:在本文中,我们专注于解决语音处理领域中最重要的任务之一,即,自动语音识别(ASR),具有语音基础编码器和大型语言模型(LLM)。最近的作品有复杂的设计,如压缩输出时间的语音编码器,解决模态对准的投影仪,并利用参数有效的微调的LLM。我们发现,微妙的设计是没有必要的,而一个非常简单的组成现成的语音编码器,LLM,和唯一的可训练线性投影仪是胜任的ASR任务。更具体地说,我们对LLM和语音编码器的各种组合进行了基准测试和探索,从而得到了基于LLM的最佳ASR系统,我们称之为SLAM-ASR。所提出的SLAM-ASR提供了一个干净的设置和很少的特定于任务的设计,其中只有线性投影仪被训练。据我们所知,SLAM-ASR在基于LLM的ASR模型中在Libripeech基准测试中取得了最佳性能,甚至优于在大量配对数据上训练的最新基于LLM的音频通用模型。最后,我们探讨了基于LLM的ASR在模态对齐过程中的能力涌现。我们希望我们的研究可以促进研究扩展LLM与跨模态的能力,并揭示了基于LLM的ASR社区。摘要:In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Recent works have complex designs such as compressing the output temporally for the speech encoder, tackling modal alignment for the projector, and utilizing parameter-efficient fine-tuning for the LLM. We found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for the ASR task. To be more specific, we benchmark and explore various combinations of LLMs and speech encoders, leading to the optimal LLM-based ASR system, which we call SLAM-ASR. The proposed SLAM-ASR provides a clean setup and little task-specific design, where only the linear projector is trained. To the best of our knowledge, SLAM-ASR achieves the best performance on the Librispeech benchmark among LLM-based ASR models and even outperforms the latest LLM-based audio-universal model trained on massive pair data. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community.
【10】 Syllable based DNN-HMM Cantonese Speech to Text System标题:基于音节的DNN-HMM粤语语音文本转换系统链接:https://arxiv.org/abs/2402.08788作者:Timothy Wong,Claire Li,Sam Lam,Billy Chiu,Qin Lu,Minglei Li,Dan Xiong,Roy Shing Yu,Vincent T. Y. Ng备注:7 pages, 3 figures, LREC 2016摘要:本文介绍了一个基于音节声学模型的粤语语音到文本(STT)转换系统。这是建立一个STT系统的努力的一部分,以帮助那些在写作技能上有认知缺陷,但通过演讲表达思想没有问题的阅读障碍学生。在粤语语音识别中,声学模型的基本单位可以是传统的IF音节,也可以是ONC音节,ONC音节将韵母进一步分为核心和结尾,以反映粤语音节内的变化。通过使用Kaldi工具包,我们的系统在GPU的帮助下使用随机梯度下降优化模型进行训练,用于混合深度神经网络和隐马尔可夫模型(DNN-HMM),具有和不具有基于I向量的说话人自适应训练技术。在所有情况下,使用具有说话人自适应训练的相同高斯混合模型(GMM-SAT)到DNN的输入特征。实验结果表明,基于ONC的音节声学建模与基于I向量的DNN-HMM取得了最好的性能,字错误率(WER)为9.66%,实时因子(RTF)为1.38812。摘要:This paper reports our work on building up a Cantonese Speech-to-Text (STT) system with a syllable based acoustic model. This is a part of an effort in building a STT system to aid dyslexic students who have cognitive deficiency in writing skills but have no problem expressing their ideas through speech. For Cantonese speech recognition, the basic unit of acoustic models can either be the conventional Initial-Final (IF) syllables, or the Onset-Nucleus-Coda (ONC) syllables where finals are further split into nucleus and coda to reflect the intra-syllable variations in Cantonese. By using the Kaldi toolkit, our system is trained using the stochastic gradient descent optimization model with the aid of GPUs for the hybrid Deep Neural Network and Hidden Markov Model (DNN-HMM) with and without I-vector based speaker adaptive training technique. The input features of the same Gaussian Mixture Model with speaker adaptive training (GMM-SAT) to DNN are used in all cases. Experiments show that the ONC-based syllable acoustic modeling with I-vector based DNN-HMM achieves the best performance with the word error rate (WER) of 9.66% and the real time factor (RTF) of 1.38812.