今天跟大家分享一篇语音相关的论文合集:cs.SD语音15篇,eess.AS音频处理17篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily
cs.SD语音

【1】 Paraformer: Fast and Accurate Parallel Transformer for  Non-autoregressive End-to-End Speech Recognition

标题:Paraformer:快速准确的非自回归端到端语音识别并行转换器

链接:https://arxiv.org/abs/2206.08317

作者:Zhifu Gao,Shiliang Zhang,Ian McLoughlin,Zhijie Yan
机构:Speech Lab, Alibaba Group, China, ICT Cluster, Singapore Institute of Technology, Singapore
备注:5 pages, 3 figures, accepted by InterSpeech2022
摘要:Transformer最近在ASR领域占据主导地位。虽然能够产生良好的性能,但它们涉及一个自回归(AR)解码器来逐个生成令牌,这在计算上效率很低。为了加快推理速度,设计了非自回归(NAR)方法,例如单步NAR,以实现并行生成。然而,由于输出标记内的独立性假设,单步NAR的性能不如AR模型,尤其是在大规模语料库中。改进单步NAR面临两个挑战:一是准确预测输出令牌数并提取隐藏变量;其次,增强输出标记之间相互依赖性的建模。为了应对这两个挑战,我们提出了一种快速准确的并联Transformer,称为并联Transformer。这利用了一个连续集成和基于fire的预测器来预测令牌的数量并生成隐藏变量。然后,GLM采样器生成语义嵌入,以增强NAR解码器对上下文相互依赖性建模的能力。最后,我们设计了一种生成负样本的策略,用于最小字错误率训练,以进一步提高性能。使用公共AISHELL-1、AISHELL-2基准和工业级20000小时任务进行的实验表明,所提出的并行器可以达到与最先进的ARTransformer相当的性能,加速比超过10倍。
摘要:Transformers have recently dominated the ASR field. Although able to yield good performance, they involve an autoregressive (AR) decoder to generate tokens one by one, which is computationally inefficient. To speed up inference, non-autoregressive (NAR) methods, e.g. single-step NAR, were designed, to enable parallel generation. However, due to an independence assumption within the output tokens, performance of single-step NAR is inferior to that of AR models, especially with a large-scale corpus. There are two challenges to improving single-step NAR: Firstly to accurately predict the number of output tokens and extract hidden variables; secondly, to enhance modeling of interdependence between output tokens. To tackle both challenges, we propose a fast and accurate parallel transformer, termed Paraformer. This utilizes a continuous integrate-and-fire based predictor to predict the number of tokens and generate hidden variables. A glancing language model (GLM) sampler then generates semantic embeddings to enhance the NAR decoder's ability to model context interdependence. Finally, we design a strategy to generate negative samples for minimum word error rate training to further improve performance. Experiments using the public AISHELL-1, AISHELL-2 benchmark, and an industrial-level 20,000 hour task demonstrate that the proposed Paraformer can attain comparable performance to the state-of-the-art AR transformer, with more than 10x speedup.


【2】 SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning

标题:SoundSpaces 2.0:视听学习模拟平台

链接:https://arxiv.org/abs/2206.08312

作者:Changan Chen,Carl Schissler,Sanchit Garg,Philip Kobernik,Alexander Clegg,Paul Calamia,Dhruv Batra,Philip W Robinson,Kristen Grauman
机构:Philip Robinson, UT Austin,  Reality Labs at Meta, Georgia Tech, Meta AI
备注:Website: this https URL
摘要:我们将介绍SoundSpaces 2.0,这是一个用于3D环境中基于动态几何体的音频渲染平台。给定真实世界环境的3D网格,Soundspace可以为从任意麦克风位置捕获的任意声音生成高度逼真的声学效果。它与现有的3D视觉资产一起,支持一系列视听研究任务,例如视听导航、映射、源定位和分离以及声学匹配。与现有资源相比,SoundSpaces 2.0具有以下优势:允许连续的空间采样、对新环境的泛化,以及可配置的麦克风和材料属性。据我们所知,这是第一次基于几何的声学模拟,它提供了高保真度和真实感,同时速度也足够快,可以用于具体学习。我们展示了模拟器的特性,并根据真实世界的音频测量对其性能进行了基准测试。此外,通过两个下游任务,包括嵌入式导航和远场自动语音识别,突出了后者的sim2real性能。SoundSpaces 2.0已公开提供,以促进对既能看到又能听到的感知系统进行更广泛的研究。
摘要:We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D visual assets, it supports an array of audio-visual research tasks, such as audio-visual navigation, mapping, source localization and separation, and acoustic matching. Compared to existing resources, SoundSpaces 2.0 has the advantages of allowing continuous spatial sampling, generalization to novel environments, and configurable microphone and material properties. To our best knowledge, this is the first geometry-based acoustic simulation that offers high fidelity and realism while also being fast enough to use for embodied learning. We showcase the simulator's properties and benchmark its performance against real-world audio measurements. In addition, through two downstream tasks covering embodied navigation and far-field automatic speech recognition, highlighting sim2real performance for the latter. SoundSpaces 2.0 is publicly available to facilitate wider research for perceptual systems that can both see and hear.


【3】 GoodBye WaveNet -- A Language Model for Raw Audio with Context of 1/2  Million Samples

标题:再见WaveNet--一个具有50万个样本的原始音频语言模型

链接:https://arxiv.org/abs/2206.08297

作者:Prateek Verma
机构:Stanford University
备注:12 pages, 1 figure. Technical Report at Stanford University. Ongoing Work
摘要:建模音频信号的长期依赖性是一个特别具有挑战性的问题,因为即使是很小的时间尺度也会产生大约十万个样本。随着最近Transformer的出现,神经架构变得善于在更长的时间尺度上建模依赖关系,但它们在缩放时受到二次约束。我们提出了一种生成式自回归架构,它可以在相当大的背景下(超过500000个样本)对音频波形进行建模。我们的工作适应于通过学习CNN前端的潜在表示来学习时间依赖性,然后使用Transformer编码器(经过充分训练的端到端)学习这些表示的依赖性:从而允许学习它认为适合下一个样本的表示。与之前比较不同时间尺度以显示改进的工作不同,我们使用标准数据集,使用相同数量的参数/上下文来显示改进。与Wavenet、SaSHMI和Sample RNN等其他方法相比,我们在用于长期结构建模的标准数据集上实现了最先进的性能。这项工作为该领域提供了非常令人兴奋的方向,因为上下文建模的改进可以用更多的数据进行缩放,并且通过使用数十亿/万亿的参数可能会获得更好的结果。
摘要:Modeling long-term dependencies for audio signals is a particularly challenging problem, as even small-time scales yield on the order of a hundred thousand samples. With the recent advent of Transformers, neural architectures became good at modeling dependencies over longer time scales, but they suffered from quadratic constraints to scale them. We propose a generative auto-regressive architecture that can model audio waveforms over quite a large context, greater than 500,000 samples. Our work is adapted to learn time dependencies by learning a latent representation by a CNN front-end, and then learning dependencies over these representations using Transformer encoders, fully trained end-to-end: thereby allowing to learn representations as it deems fit for the next sample. Unlike previous works that compared different time scales to show improvement, we use a standard dataset, with the same number of parameters/context to show improvements. We achieve a state-of-the-art performance as compared to other approaches such as Wavenet, SaSHMI, and Sample-RNN on a standard dataset for modeling long-term structure. This work gives very exciting direction for the field, given improvements in context modeling that can be scaled with more data, as well as potentially better results by using billions/trillions of parameters.


【4】 Event-related data conditioning for acoustic event classification

标题:用于声事件分类的事件相关数据条件

链接:https://arxiv.org/abs/2206.08233

作者:Yuanbo Hou,Dick Botteldooren
机构:WAVES, Ghent University, Belgium
备注:Accepted by INTERSPEECH 2022
摘要:基于不同注意机制的模型最近在与声学事件分类(AEC)相关的任务中大放异彩。其中,自我注意通常用于纯音频任务,以帮助模型识别不同的声学事件。自我注意依赖于时间框架之间的相似性,并使用整个片段的全局信息来突出框架内的特定特征。在现实生活中,与声学事件相关的信息会随时间衰减,这意味着事件周围某些帧内的信息比可能与事件无关的远距离全局信息更值得关注。本文表明,自我注意可能会过度增强某些音频表征片段,并平滑事件表征和背景噪声之间的边界。因此,本文提出了一种用于AEC的事件相关数据调节(EDC)。EDC直接处理光谱图。EDC的思想是基于声学特征自适应地选择与帧相关的注意范围,并收集事件相关的局部信息来表示帧。实验表明:1)与基于谱图的数据增强方法和可训练的特征加权和自我注意方法相比,EDC在原始尺寸模式和增强模式下均优于它们;2) EDC有效地收集与事件相关的本地信息,增强事件和背景之间的边界,提高AEC的性能。
摘要:Models based on diverse attention mechanisms have recently shined in tasks related to acoustic event classification (AEC). Among them, self-attention is often used in audio-only tasks to help the model recognize different acoustic events. Self-attention relies on the similarity between time frames, and uses global information from the whole segment to highlight specific features within a frame. In real life, information related to acoustic events will attenuate over time, which means the information within some frames around the event deserves more attention than distant time global information that may be unrelated to the event. This paper shows that self-attention may over-enhance certain segments of audio representations, and smooth out the boundaries between events representations and background noises. Hence, this paper proposes an event-related data conditioning (EDC) for AEC. EDC directly works on spectrograms. The idea of EDC is to adaptively select the frame-related attention range based on acoustic features, and gather the event-related local information to represent the frame. Experiments show that: 1) compared with spectrogram-based data augmentation methods and trainable feature weighting and self-attention, EDC outperforms them in both the original-size mode and the augmented mode; 2) EDC effectively gathers event-related local information and enhances boundaries between events and backgrounds, improving the performance of AEC.


【5】 Censer: Curriculum Semi-supervised Learning for Speech Recognition Based  on Self-supervised Pre-training

标题:基于自监督预训练的语音识别课程半监督学习

链接:https://arxiv.org/abs/2206.08189

作者:Bowen Zhang,Songjun Cao,Xiaoming Zhang,Yike Zhang,Long Ma,Takahiro Shinozaki
机构:Tokyo Institute of Technology, Tokyo, Japan, Tencent Cloud Xiaowei, Beijing, China
摘要:最近的研究表明,自我监督的预训练和自我训练(伪标记)所提供的益处是互补的。然而,在预训练框架下的半监督微调策略仍然没有得到足够的研究。此外,现代半监督语音识别算法要么不加区分地处理未标记的数据,要么用置信阈值过滤掉噪声样本。不同的未标记数据之间的差异常常被忽略。本文提出了一种基于自监督预训练的半监督语音识别算法Censer,以最大限度地提高未标记数据的利用率。Censer的预训练阶段采用wav2vec2.0,微调阶段采用改进的slimIPL半监督学习算法,该算法根据未标记数据的伪标签质量逐步利用未标记数据。我们还结合了一个时间伪标签池和一个指数移动平均值来控制伪标签的更新频率并避免模型发散。在Libri-Light和LibriSpeech数据集上的实验结果表明,与现有方法相比,我们提出的方法在更加统一的同时取得了更好的性能。
摘要:Recent studies have shown that the benefits provided by self-supervised pre-training and self-training (pseudo-labeling) are complementary. Semi-supervised fine-tuning strategies under the pre-training framework, however, remain insufficiently studied. Besides, modern semi-supervised speech recognition algorithms either treat unlabeled data indiscriminately or filter out noisy samples with a confidence threshold. The dissimilarities among different unlabeled data are often ignored. In this paper, we propose Censer, a semi-supervised speech recognition algorithm based on self-supervised pre-training to maximize the utilization of unlabeled data. The pre-training stage of Censer adopts wav2vec2.0 and the fine-tuning stage employs an improved semi-supervised learning algorithm from slimIPL, which leverages unlabeled data progressively according to their pseudo labels' qualities. We also incorporate a temporal pseudo label pool and an exponential moving average to control the pseudo labels' update frequency and to avoid model divergence. Experimental results on Libri-Light and LibriSpeech datasets manifest our proposed method achieves better performance compared to existing approaches while being more unified.


【6】 Adversarial Privacy Protection on Speech Enhancement

标题:语音增强中的对抗性隐私保护

链接:https://arxiv.org/abs/2206.08170

作者:Mingyu Dong,Diqun Yan,Rangding Wang
机构:College of Information Science and Engineering, Ningbo University
备注:5 pages, 6 figures
摘要:语音很容易在不知不觉中泄露,例如在不同情况下被手机录制。通过语音增强技术可以恶意提取语音中的私有内容。语音增强技术随着深度神经网络(DNN)的发展迅速,但对抗性的例子可能会导致DNN失败。在这项工作中,我们提出了一种对抗性的方法来降低语音增强系统的性能。实验结果表明,生成的对抗性示例可以擦除原始示例中的大部分内容信息,或者通过语音增强将其替换为目标语音内容。增强的原始示例和增强的对抗性示例识别结果之间的单词错误率(WER)可以达到89.0%。增强对抗示例和目标示例之间的目标攻击功率低至33.75%。对抗性扰动可以使原始示例的变化率超过1.4430。这项工作可以防止恶意的语音提取。
摘要:Speech is easily leaked imperceptibly, such as being recorded by mobile phones in different situations. Private content in speech may be maliciously extracted through speech enhancement technology. Speech enhancement technology has developed rapidly along with deep neural networks (DNNs), but adversarial examples can cause DNNs to fail. In this work, we propose an adversarial method to degrade speech enhancement systems. Experimental results show that generated adversarial examples can erase most content information in original examples or replace it with target speech content through speech enhancement. The word error rate (WER) between an enhanced original example and enhanced adversarial example recognition result can reach 89.0%. WER of target attack between enhanced adversarial example and target example is low to 33.75% . Adversarial perturbation can bring the rate of change to the original example to more than 1.4430. This work can prevent the malicious extraction of speech.


【7】 Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis  Using Linguistic and Prosodic Contexts of Dialogue History

标题:利用对话历史的语言和韵律语境进行端到端移情对话语音合成的声学建模

链接:https://arxiv.org/abs/2206.08039

作者:Yuto Nishimura,Yuki Saito,Shinnosuke Takamichi,Kentaro Tachibana,Hiroshi Saruwatari
机构:The University of Tokyo, Japan,LINE Corp., Japan.
备注:5 pages, 3 figures, Accepted for INTERSPEECH2022
摘要:我们提出了一个端到端移情对话语音合成(DSS)模型,该模型考虑了对话历史的语言和韵律背景。移情是人类在对话中主动尝试进入对话者内部,而移情DSS是在口语对话系统中实现这一行为的技术。我们的模型以语言和韵律特征的历史为条件,以预测合适的对话语境。因此,它可以看作是传统的基于语言特征的对话历史建模的扩展。为了有效地训练移情DSS模型,我们研究了1)用大型语音语料库预训练的自我监督学习模型,2)使用对话语境嵌入预测的当前话语韵律嵌入的风格引导训练,3)结合文本和语音模式的跨模态注意,4)句子嵌入,实现细粒度的韵律建模,而不是话语建模。评价结果表明:1)在移情决策支持系统中,单纯考虑对话历史的韵律语境并不能提高语音质量;2)引入风格引导训练和句子嵌入建模可以获得比传统方法更高的语音质量。
摘要:We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history. Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and empathetic DSS is a technology to implement this act in spoken dialogue systems. Our model is conditioned by the history of linguistic and prosody features for predicting appropriate dialogue context. As such, it can be regarded as an extension of the conventional linguistic-feature-based dialogue history modeling. To train the empathetic DSS model effectively, we investigate 1) a self-supervised learning model pretrained with large speech corpora, 2) a style-guided training using a prosody embedding of the current utterance to be predicted by the dialogue context embedding, 3) a cross-modal attention to combine text and speech modalities, and 4) a sentence-wise embedding to achieve fine-grained prosody modeling rather than utterance-wise modeling. The evaluation results demonstrate that 1) simply considering prosodic contexts of the dialogue history does not improve the quality of speech in empathetic DSS and 2) introducing style-guided training and sentence-wise embedding modeling achieves higher speech quality than that by the conventional method.


【8】 DCASE 2022: Comparative Analysis Of CNNs For Acoustic Scene  Classification Under Low-Complexity Considerations

标题:DCASE 2022:低复杂度条件下声场分类的CNN比较分析

链接:https://arxiv.org/abs/2206.08007

作者:Josep Zaragoza-Paredes,Javier Naranjo-Alcazar,Valery Naranjo,Pedro Zuccarello
机构:Pedro Zuccarello 2 1 Universitat Politecnica de Valencia
摘要:声学场景分类是一个自动监听问题,其目的是根据音频数据将音频记录分配给预定义的场景。多年来(以及在DCASE的过去版本中),这个问题通常通过称为集成的技术(使用多个机器学习模型在推理阶段组合预测)来解决。虽然这些解决方案可以在准确性方面显示性能,但在计算能力方面可能非常昂贵,因此无法在物联网设备中部署它们。由于这一研究领域的漂移,这项任务在模型复杂性方面有两个局限性。应该注意的是,失配设备也会增加复杂性(提供的音频由不同的信息源记录)。本技术报告对两种不同的网络架构进行了比较研究:传统的CNN和Conv混音器。虽然两个网络都超过了比赛要求的基线,但传统的有线电视新闻网表现出了更高的性能,超过了基线8个百分点。基于Conv-mixer体系结构的解决方案性能较差,尽管它们是更轻的解决方案。
摘要:Acoustic scene classification is an automatic listening problem that aims to assign an audio recording to a pre-defined scene based on its audio data. Over the years (and in past editions of the DCASE) this problem has often been solved with techniques known as ensembles (use of several machine learning models to combine their predictions in the inference phase). While these solutions can show performance in terms of accuracy, they can be very expensive in terms of computational capacity, making it impossible to deploy them in IoT devices. Due to the drift in this field of study, this task has two limitations in terms of model complexity. It should be noted that there is also the added complexity of mismatching devices (the audios provided are recorded by different sources of information). This technical report makes a comparative study of two different network architectures: conventional CNN and Conv-mixer. Although both networks exceed the baseline required by the competition, the conventional CNN shows a higher performance, exceeding the baseline by 8 percentage points. Solutions based on Conv-mixer architectures show worse performance although they are much lighter solutions.


【9】 Automatic Prosody Annotation with Pre-Trained Text-Speech Model

标题:基于预训练文本-语音模型的韵律自动标注

链接:https://arxiv.org/abs/2206.07956

作者:Ziqian Dai,Jianwei Yu,Yan Wang,Nuo Chen,Yanyao Bian,Guangzhi Li,Deng Cai,Dong Yu
机构:Tencent AI Lab,Peking University
备注:accepted by INTERSPEECH2022
摘要:韵律边界在文语合成(TTS)的自然性和可读性方面起着重要作用。然而,韵律边界标签的获取依赖于手动注释,这既昂贵又耗时。在这篇文章中,我们提出通过一个带有预训练音频编码器的神经文本语音模型,从文本音频数据中自动提取韵律边界标签。该模型分别对文本和语音数据进行预训练,并以三元组格式对TTS数据进行联合微调:{语音,文本,韵律}。自动评价和人工评价的实验结果表明:1)提出的文本-语音韵律标注框架显著优于纯文本基线;2) 自动韵律边界标注的质量与人类标注相当;3) 使用模型注释边界训练的TTS系统略优于使用手动边界的系统。
摘要:Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming. In this paper, we propose to automatically extract prosodic boundary labels from text-audio data via a neural text-speech model with pre-trained audio encoders. This model is pre-trained on text and speech data separately and jointly fine-tuned on TTS data in a triplet format: {speech, text, prosody}. The experimental results on both automatic evaluation and human evaluation demonstrate that: 1) the proposed text-speech prosody annotation framework significantly outperforms text-only baselines; 2) the quality of automatic prosodic boundary annotations is comparable to human annotations; 3) TTS systems trained with model-annotated boundaries are slightly better than systems that use manual ones.


【10】 Accelerating Inference and Language Model Fusion of Recurrent Neural  Network Transducers via End-to-End 4-bit Quantization

标题:端到端4位量化加速递归神经网络传感器的推理和语言模型融合

链接:https://arxiv.org/abs/2206.07882

作者:Andrea Fasoli,Chia-Yu Chen,Mauricio Serrano,Swagath Venkataramani,George Saon,Xiaodong Cui,Brian Kingsbury,Kailash Gopalakrishnan
机构:IBM Research, USA
备注:5 pages, 2 figures, 1 table. Paper accepted to Interspeech 2022
摘要:我们报告了积极的量化策略,大大加快了递归神经网络传感器(RNN-T)的推理。我们对权重和激活使用4位整数表示,并应用量化感知训练(QAT)对完整模型(声学编码器和语言模型)进行重新训练,以达到接近iso的精度。我们表明,根据网络的局部特性定制的量化方案对于在限制QAT计算开销的同时获得良好性能至关重要。密度比语言模型融合在RNN-T工作负载上显示出显著的精度提高,但它严重增加了推理的计算成本。我们表明,我们的量化策略能够使用大波束宽度进行假设搜索,同时实现流兼容的运行时,与全精度模型相比,全模型压缩比为7.6$\倍。通过硬件模拟,我们估计端到端量化RNN-T(包括LM融合)从FP16到INT4的加速度为3.4$\倍,导致实时因子(RTF)为0.06。在NIST Hub5 2000、Hub5 2001和RT-03测试集上,我们保留了与LM fusion相关的大部分收益,将平均WER提高了$>$1.5%。
摘要:We report on aggressive quantization strategies that greatly accelerate inference of Recurrent Neural Network Transducers (RNN-T). We use a 4 bit integer representation for both weights and activations and apply Quantization Aware Training (QAT) to retrain the full model (acoustic encoder and language model) and achieve near-iso-accuracy. We show that customized quantization schemes that are tailored to the local properties of the network are essential to achieve good performance while limiting the computational overhead of QAT.  Density ratio Language Model fusion has shown remarkable accuracy gains on RNN-T workloads but it severely increases the computational cost of inference. We show that our quantization strategies enable using large beam widths for hypothesis search while achieving streaming-compatible runtimes and a full model compression ratio of 7.6$\times$ compared to the full precision model.  Via hardware simulations, we estimate a 3.4$\times$ acceleration from FP16 to INT4 for the end-to-end quantized RNN-T inclusive of LM fusion, resulting in a Real Time Factor (RTF) of 0.06. On the NIST Hub5 2000, Hub5 2001, and RT-03 test sets, we retain most of the gains associated with LM fusion, improving the average WER by $>$1.5%.


【11】 EPG2S: Speech Generation and Speech Enhancement based on  Electropalatography and Audio Signals using Multimodal Learning

标题:EPG2S:基于腭电和音频信号的多模式学习语音生成与增强

链接:https://arxiv.org/abs/2206.07860

作者:Li-Chin Chen,Po-Hsun Chen,Richard Tzong-Han Tsai,Yu Tsao
备注:Accepted By IEEE Signal Processing Letter
摘要:基于发音运动的语音生成和增强有助于在缺乏言语交流的情况下进行交流,例如在丧失说话能力的患者中。尽管为此提出了各种技术,但腭电图(EPG)作为一种记录语言过程中舌头和硬腭之间接触的监测技术,尚未得到充分的探索。在此,我们提出了一种新的多模式EPG-to-speech(EPG2S)系统,该系统利用EPG和语音信号进行语音生成和增强。研究了基于EPG和噪声语音信号多重组合的不同融合策略,并对该方法的可行性进行了研究。实验结果表明,EPG2S仅基于EPG信号就可以获得理想的语音生成效果。此外,观察到添加噪声语音信号以提高质量和可懂度。此外,观察到EPG2S可实现仅基于音频信号的高质量语音增强,添加EPG信号可进一步提高性能。延迟融合策略被认为是同步语音生成和增强的最有效方法。
摘要:Speech generation and enhancement based on articulatory movements facilitate communication when the scope of verbal communication is absent, e.g., in patients who have lost the ability to speak. Although various techniques have been proposed to this end, electropalatography (EPG), which is a monitoring technique that records contact between the tongue and hard palate during speech, has not been adequately explored. Herein, we propose a novel multimodal EPG-to-speech (EPG2S) system that utilizes EPG and speech signals for speech generation and enhancement. Different fusion strategies based on multiple combinations of EPG and noisy speech signals are examined, and the viability of the proposed method is investigated. Experimental results indicate that EPG2S achieves desirable speech generation outcomes based solely on EPG signals. Further, the addition of noisy speech signals is observed to improve quality and intelligibility. Additionally, EPG2S is observed to achieve high-quality speech enhancement based solely on audio signals, with the addition of EPG signals further improving the performance. The late fusion strategy is deemed to be the most effective approach for simultaneous speech generation and enhancement.


【12】 Strategies to Improve Robustness of Target Speech Extraction to  Enrollment Variations

标题:提高目标语音提取对注册变化的稳健性的策略

链接:https://arxiv.org/abs/2206.08174

作者:Hiroshi Sato,Tsubasa Ochiai,Marc Delcroix,Keisuke Kinoshita,Takafumi Moriya,Naoki Makishima,Mana Ihori,Tomohiro Tanaka,Ryo Masumura
机构:NTT Corporation, Japan
备注:5 pages, 2 figures, 3 tables Submitted to Interspeech 2022
摘要:目标语音提取是一种使用预先记录的注册话语从混合信号中提取目标说话人语音的技术,该注册话语表征了目标说话人的语音特征。目标语音提取的一个主要困难在于处理“说话人内部”特征的可变性,即目标语音和注册话语之间的特征不匹配。虽然大多数传统的方法侧重于在给定一组注册语句的情况下提高{\it平均性能},但在这里,我们建议保证{\it最差性能},我们认为这具有非常重要的实际意义。在这项工作中,我们提出了一种称为最差注册源失真比(SDR)的评估指标,以定量衡量对注册变化的鲁棒性。我们还介绍了一种新的训练方案,该方案旨在直接优化最坏情况下的性能,方法是将重点放在难以注册的情况下进行训练,其中提取效果不佳。此外,我们还研究了辅助说话人识别丢失(SI丢失)作为提高注册鲁棒性的另一种方法的有效性。实验验证表明,最差注册目标训练和SI损失训练都可以通过提高说话人的辨别能力来提高对注册变化的鲁棒性。
摘要:Target speech extraction is a technique to extract the target speaker's voice from mixture signals using a pre-recorded enrollment utterance that characterize the voice characteristics of the target speaker. One major difficulty of target speech extraction lies in handling variability in ``intra-speaker'' characteristics, i.e., characteristics mismatch between target speech and an enrollment utterance. While most conventional approaches focus on improving {\it average performance} given a set of enrollment utterances, here we propose to guarantee the {\it worst performance}, which we believe is of great practical importance. In this work, we propose an evaluation metric called worst-enrollment source-to-distortion ratio (SDR) to quantitatively measure the robustness towards enrollment variations. We also introduce a novel training scheme that aims at directly optimizing the worst-case performance by focusing on training with difficult enrollment cases where extraction does not perform well. In addition, we investigate the effectiveness of auxiliary speaker identification loss (SI-loss) as another way to improve robustness over enrollments. Experimental validation reveals the effectiveness of both worst-enrollment target training and SI-loss training to improve robustness against enrollment variations, by increasing speaker discriminability.


【13】 Nonwords Pronunciation Classification in Language Development Tests for  Preschool Children

标题:学龄前儿童语言发展测验中的非单词发音分类

链接:https://arxiv.org/abs/2206.08058

作者:Ilja Baumann,Dominik Wagner,Sebastian Bayerl,Tobias Bocklet
机构:Technische Hochschule N¨urnberg Georg Simon Ohm, Germany, Intel Labs
备注:Accepted to Interspeech 2022
摘要:这项工作旨在自动评估儿童的语言发展是否与年龄相适应。经过验证的语音和语言测试用于测试听觉记忆。在这项工作中,任务是确定所说的非单词是否正确说出。我们比较了不同的建模特定语言结构的方法:低层特征(FFT)、说话人嵌入(ECAPA-TDNN)、字素嵌入(wav2vec 2.0)和senones形式的语音嵌入(ASR声学模型)。每种方法都为类似VGG的5层CNN分类器提供输入。我们还研究了每个非单词的顺应性。使用来自不同幼儿园的非单词口语录音对提议的系统进行评估。ECAPA-TDNN和低层FFT特征没有明确建模语音信息;wav2vec2.0是在字形标签上训练的,我们的ASR声学模型功能包含(子)语音信息。我们发现,语音建模的粒度越大,识别率越高。使用VTLN对ASR声学模型特征进行训练的最佳系统的精确度为89.4%,ROC(接收机工作特性)曲线(AUC)下的面积为0.923。与FFT基线相比,这相当于精确度提高了20.2%,AUC为0.309。
摘要:This work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine whether spoken nonwords have been uttered correctly. We compare different approaches that are motivated to model specific language structures: Low-level features (FFT), speaker embeddings (ECAPA-TDNN), grapheme-motivated embeddings (wav2vec 2.0), and phonetic embeddings in form of senones (ASR acoustic model). Each of the approaches provides input for VGG-like 5-layer CNN classifiers. We also examine the adaptation per nonword. The evaluation of the proposed systems was performed using recordings from different kindergartens of spoken nonwords. ECAPA-TDNN and low-level FFT features do not explicitly model phonetic information; wav2vec2.0 is trained on grapheme labels, our ASR acoustic model features contain (sub-)phonetic information. We found that the more granular the phonetic modeling is, the higher are the achieved recognition rates. The best system trained on ASR acoustic model features with VTLN achieved an accuracy of 89.4% and an area under the ROC (Receiver Operating Characteristic) curve (AUC) of 0.923. This corresponds to an improvement in accuracy of 20.2% and AUC of 0.309 relative compared to the FFT-baseline.


【14】 DRAFT: A Novel Framework to Reduce Domain Shifting in Self-supervised  Learning and Its Application to Children's ASR

标题:一种减少自我监督学习领域转移的新框架及其在儿童ASR中的应用

链接:https://arxiv.org/abs/2206.07931

作者:Ruchao Fan,Abeer Alwan
机构:Dept. of Electrical and Computer Engineering, University of California, Los Angeles, USA
备注:Accepted to Interspeech 2022
摘要:在预训练阶段使用无注释语音数据的自监督学习(SSL)在低资源自动语音识别(ASR)任务中取得了成功。然而,通过SSL训练的模型偏向于预训练数据,而预训练数据通常与微调任务中使用的数据不同,导致领域转移问题,从而导致有限的知识转移。我们提出了一个新的框架,域负责自适应和微调(草案),通过一个额外的自适应阶段来减少预训练语音模型中的域转移。在草稿中,将剩余适配器(RAs)插入到预训练模型中,以学习与域相关的信息,其SSL丢失与预训练阶段相同。在自适应阶段,仅更新RA参数。DRAFT与所使用的SSL方法的类型无关,并使用三种广泛使用的方法进行评估:APC、Wav2vec2.0和HuBERT。在两个子ASR任务(OGI和MyST数据库)上,使用用未注释成人语音数据(Librispeech)训练的SSL模型,与未经调整的预训练模型相比,相对WER提高了19.7%。其他实验检验了两个数据集之间交叉知识转移的潜力,结果很有希望,表明拟议框架草案的使用范围更广。
摘要:Self-supervised learning (SSL) in the pretraining stage using un-annotated speech data has been successful in low-resource automatic speech recognition (ASR) tasks. However, models trained through SSL are biased to the pretraining data which is usually different from the data used in finetuning tasks, causing a domain shifting problem, and thus resulting in limited knowledge transfer. We propose a novel framework, domain responsible adaptation and finetuning (DRAFT), to reduce domain shifting in pretrained speech models through an additional adaptation stage. In DRAFT, residual adapters (RAs) are inserted in the pretrained model to learn domain-related information with the same SSL loss as the pretraining stage. Only RA parameters are updated during the adaptation stage. DRAFT is agnostic to the type of SSL method used and is evaluated with three widely used approaches: APC, Wav2vec2.0, and HuBERT. On two child ASR tasks (OGI and MyST databases), using SSL models trained with un-annotated adult speech data (Librispeech), relative WER improvements of up to 19.7% are observed when compared to the pretrained models without adaptation. Additional experiments examined the potential of cross knowledge transfer between the two datasets and the results are promising, showing a broader usage of the proposed DRAFT framework.


【15】 To Dereverb Or Not to Dereverb? Perceptual Studies On Real-Time  Dereverberation Targets

标题:去德韦伯还是不去德韦伯?实时去混响目标的感知研究

链接:https://arxiv.org/abs/2206.07917

作者:Jean-Marc Valin,Ritwik Giri,Shrikant Venkataramani,Umut Isik,Arvindh Krishnaswamy
机构:Amazon Web Services, Palo Alto, CA, USA
备注:5 pages
摘要:在现实生活中,房间效应(也称为房间混响)和当前背景噪声会降低语音质量。近年来,基于深度学习的语音增强方法显示出了巨大的潜力,并超越了传统的去噪和去冗余方法。这些最先进的去噪算法显著改善了人类听者感知到的语音质量,这一点也已得到公认。但去冗余对主观(感知)语音质量的作用,以及去冗余引入的附加伪影是否弊大于利,仍不清楚。在本文中,我们试图通过对一个最先进的语音增强系统进行综合主观评估来回答这些问题,该系统针对不同的去冗余目标选择进行评估。
摘要:In real life, room effect, also known as room reverberation, and the present background noise degrade the quality of speech. Recently, deep learning-based speech enhancement approaches have shown a lot of promise and surpassed traditional denoising and dereverberation methods. It is also well established that these state-of-the-art denoising algorithms significantly improve the quality of speech as perceived by human listeners. But the role of dereverberation on subjective (perceived) speech quality, and whether the additional artifacts introduced by dereverberation cause more harm than good are still unclear. In this paper, we attempt to answer these questions by evaluating a state of the art speech enhancement system in a comprehensive subjective evaluation study for different choices of dereverberation targets.


eess.AS音频处理

【1】 Strategies to Improve Robustness of Target Speech Extraction to  Enrollment Variations

标题:提高目标语音提取对注册变化的稳健性的策略

链接:https://arxiv.org/abs/2206.08174

作者:Hiroshi Sato,Tsubasa Ochiai,Marc Delcroix,Keisuke Kinoshita,Takafumi Moriya,Naoki Makishima,Mana Ihori,Tomohiro Tanaka,Ryo Masumura
机构:NTT Corporation, Japan
备注:5 pages, 2 figures, 3 tables Submitted to Interspeech 2022
摘要:目标语音提取是一种使用预先记录的注册话语从混合信号中提取目标说话人语音的技术,该注册话语表征了目标说话人的语音特征。目标语音提取的一个主要困难在于处理“说话人内部”特征的可变性,即目标语音和注册话语之间的特征不匹配。虽然大多数传统的方法侧重于在给定一组注册语句的情况下提高{\it平均性能},但在这里,我们建议保证{\it最差性能},我们认为这具有非常重要的实际意义。在这项工作中,我们提出了一种称为最差注册源失真比(SDR)的评估指标,以定量衡量对注册变化的鲁棒性。我们还介绍了一种新的训练方案,该方案旨在直接优化最坏情况下的性能,方法是将重点放在难以注册的情况下进行训练,其中提取效果不佳。此外,我们还研究了辅助说话人识别丢失(SI丢失)作为提高注册鲁棒性的另一种方法的有效性。实验验证表明,最差注册目标训练和SI损失训练都可以通过提高说话人的辨别能力来提高对注册变化的鲁棒性。
摘要:Target speech extraction is a technique to extract the target speaker's voice from mixture signals using a pre-recorded enrollment utterance that characterize the voice characteristics of the target speaker. One major difficulty of target speech extraction lies in handling variability in ``intra-speaker'' characteristics, i.e., characteristics mismatch between target speech and an enrollment utterance. While most conventional approaches focus on improving {\it average performance} given a set of enrollment utterances, here we propose to guarantee the {\it worst performance}, which we believe is of great practical importance. In this work, we propose an evaluation metric called worst-enrollment source-to-distortion ratio (SDR) to quantitatively measure the robustness towards enrollment variations. We also introduce a novel training scheme that aims at directly optimizing the worst-case performance by focusing on training with difficult enrollment cases where extraction does not perform well. In addition, we investigate the effectiveness of auxiliary speaker identification loss (SI-loss) as another way to improve robustness over enrollments. Experimental validation reveals the effectiveness of both worst-enrollment target training and SI-loss training to improve robustness against enrollment variations, by increasing speaker discriminability.


【2】 Nonwords Pronunciation Classification in Language Development Tests for  Preschool Children

标题:学龄前儿童语言发展测验中的非单词发音分类

链接:https://arxiv.org/abs/2206.08058

作者:Ilja Baumann,Dominik Wagner,Sebastian Bayerl,Tobias Bocklet
机构:Technische Hochschule N¨urnberg Georg Simon Ohm, Germany, Intel Labs
备注:Accepted to Interspeech 2022
摘要:这项工作旨在自动评估儿童的语言发展是否与年龄相适应。经过验证的语音和语言测试用于测试听觉记忆。在这项工作中,任务是确定所说的非单词是否正确说出。我们比较了不同的建模特定语言结构的方法:低层特征(FFT)、说话人嵌入(ECAPA-TDNN)、字素嵌入(wav2vec 2.0)和senones形式的语音嵌入(ASR声学模型)。每种方法都为类似VGG的5层CNN分类器提供输入。我们还研究了每个非单词的顺应性。使用来自不同幼儿园的非单词口语录音对提议的系统进行评估。ECAPA-TDNN和低层FFT特征没有明确建模语音信息;wav2vec2.0是在字形标签上训练的,我们的ASR声学模型功能包含(子)语音信息。我们发现,语音建模的粒度越大,识别率越高。使用VTLN对ASR声学模型特征进行训练的最佳系统的精确度为89.4%,ROC(接收机工作特性)曲线(AUC)下的面积为0.923。与FFT基线相比,这相当于精确度提高了20.2%,AUC为0.309。
摘要:This work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine whether spoken nonwords have been uttered correctly. We compare different approaches that are motivated to model specific language structures: Low-level features (FFT), speaker embeddings (ECAPA-TDNN), grapheme-motivated embeddings (wav2vec 2.0), and phonetic embeddings in form of senones (ASR acoustic model). Each of the approaches provides input for VGG-like 5-layer CNN classifiers. We also examine the adaptation per nonword. The evaluation of the proposed systems was performed using recordings from different kindergartens of spoken nonwords. ECAPA-TDNN and low-level FFT features do not explicitly model phonetic information; wav2vec2.0 is trained on grapheme labels, our ASR acoustic model features contain (sub-)phonetic information. We found that the more granular the phonetic modeling is, the higher are the achieved recognition rates. The best system trained on ASR acoustic model features with VTLN achieved an accuracy of 89.4% and an area under the ROC (Receiver Operating Characteristic) curve (AUC) of 0.923. This corresponds to an improvement in accuracy of 20.2% and AUC of 0.309 relative compared to the FFT-baseline.


【3】 A CTC Triggered Siamese Network with Spatial-Temporal Dropout for Speech  Recognition

标题:一种用于语音识别的时空丢弃CTC触发暹罗网络

链接:https://arxiv.org/abs/2206.08031

作者:Yingying Gao,Junlan Feng,Tianrui Wang,Chao Deng,Shilei Zhang
机构:China Mobile Research Institute, Institute of Information Science, Beijing Jiaotong University
摘要:暹罗网络在无监督视觉表征学习中显示出了有效的效果。这些模型旨在通过最大化两个增广的相似度来学习一个输入的两个增广的不变表示。本文提出了一种有效的暹罗网络来提高端到端自动语音识别(ASR)的鲁棒性。我们引入时空辍学来支持对暹罗ASR框架更猛烈的干扰。此外,我们还放松了相似性正则化,以最大化连接主义时间分类(CTC)尖峰出现的帧上分布的相似性,而不是所有帧上分布的相似性。在AISHELL-1和Librispeech两个基准上对所提出的体系结构的效率进行了评估,分别降低了7.13%和6.59%的相对字符错误率(CER)和字错误率(WER)。分析表明,我们提出的方法为训练模型带来了更好的一致性,并明显增大了CTC峰值。
摘要:Siamese networks have shown effective results in unsupervised visual representation learning. These models are designed to learn an invariant representation of two augmentations for one input by maximizing their similarity. In this paper, we propose an effective Siamese network to improve the robustness of End-to-End automatic speech recognition (ASR). We introduce spatial-temporal dropout to support a more violent disturbance for Siamese-ASR framework. Besides, we also relax the similarity regularization to maximize the similarities of distributions on the frames that connectionist temporal classification (CTC) spikes occur rather than on all of them. The efficiency of the proposed architecture is evaluated on two benchmarks, AISHELL-1 and Librispeech, resulting in 7.13% and 6.59% relative character error rate (CER) and word error rate (WER) reductions respectively. Analysis shows that our proposed approach brings a better uniformity for the trained model and enlarges the CTC spikes obviously.


【4】 DRAFT: A Novel Framework to Reduce Domain Shifting in Self-supervised  Learning and Its Application to Children's ASR

标题:一种减少自我监督学习领域转移的新框架及其在儿童ASR中的应用

链接:https://arxiv.org/abs/2206.07931

作者:Ruchao Fan,Abeer Alwan
机构:Dept. of Electrical and Computer Engineering, University of California, Los Angeles, USA
备注:Accepted to Interspeech 2022
摘要:在预训练阶段使用无注释语音数据的自监督学习(SSL)在低资源自动语音识别(ASR)任务中取得了成功。然而,通过SSL训练的模型偏向于预训练数据,而预训练数据通常与微调任务中使用的数据不同,导致领域转移问题,从而导致有限的知识转移。我们提出了一个新的框架,域负责自适应和微调(草案),通过一个额外的自适应阶段来减少预训练语音模型中的域转移。在草稿中,将剩余适配器(RAs)插入到预训练模型中,以学习与域相关的信息,其SSL丢失与预训练阶段相同。在自适应阶段,仅更新RA参数。DRAFT与所使用的SSL方法的类型无关,并使用三种广泛使用的方法进行评估:APC、Wav2vec2.0和HuBERT。在两个子ASR任务(OGI和MyST数据库)上,使用用未注释成人语音数据(Librispeech)训练的SSL模型,与未经调整的预训练模型相比,相对WER提高了19.7%。其他实验检验了两个数据集之间交叉知识转移的潜力,结果很有希望,表明拟议框架草案的使用范围更广。
摘要:Self-supervised learning (SSL) in the pretraining stage using un-annotated speech data has been successful in low-resource automatic speech recognition (ASR) tasks. However, models trained through SSL are biased to the pretraining data which is usually different from the data used in finetuning tasks, causing a domain shifting problem, and thus resulting in limited knowledge transfer. We propose a novel framework, domain responsible adaptation and finetuning (DRAFT), to reduce domain shifting in pretrained speech models through an additional adaptation stage. In DRAFT, residual adapters (RAs) are inserted in the pretrained model to learn domain-related information with the same SSL loss as the pretraining stage. Only RA parameters are updated during the adaptation stage. DRAFT is agnostic to the type of SSL method used and is evaluated with three widely used approaches: APC, Wav2vec2.0, and HuBERT. On two child ASR tasks (OGI and MyST databases), using SSL models trained with un-annotated adult speech data (Librispeech), relative WER improvements of up to 19.7% are observed when compared to the pretrained models without adaptation. Additional experiments examined the potential of cross knowledge transfer between the two datasets and the results are promising, showing a broader usage of the proposed DRAFT framework.


【5】 To Dereverb Or Not to Dereverb? Perceptual Studies On Real-Time  Dereverberation Targets

标题:去德韦伯还是不去德韦伯?实时去混响目标的感知研究

链接:https://arxiv.org/abs/2206.07917

作者:Jean-Marc Valin,Ritwik Giri,Shrikant Venkataramani,Umut Isik,Arvindh Krishnaswamy
机构:Amazon Web Services, Palo Alto, CA, USA
备注:5 pages
摘要:在现实生活中,房间效应(也称为房间混响)和当前背景噪声会降低语音质量。近年来,基于深度学习的语音增强方法显示出了巨大的潜力,并超越了传统的去噪和去冗余方法。这些最先进的去噪算法显著改善了人类听者感知到的语音质量,这一点也已得到公认。但去冗余对主观(感知)语音质量的作用,以及去冗余引入的附加伪影是否弊大于利,仍不清楚。在本文中,我们试图通过对一个最先进的语音增强系统进行综合主观评估来回答这些问题,该系统针对不同的去冗余目标选择进行评估。
摘要:In real life, room effect, also known as room reverberation, and the present background noise degrade the quality of speech. Recently, deep learning-based speech enhancement approaches have shown a lot of promise and surpassed traditional denoising and dereverberation methods. It is also well established that these state-of-the-art denoising algorithms significantly improve the quality of speech as perceived by human listeners. But the role of dereverberation on subjective (perceived) speech quality, and whether the additional artifacts introduced by dereverberation cause more harm than good are still unclear. In this paper, we attempt to answer these questions by evaluating a state of the art speech enhancement system in a comprehensive subjective evaluation study for different choices of dereverberation targets.


【6】 The Scattering Transform Network with Generalized Morse Wavelets and Its  Application to Music Genre Classification

标题:基于广义Morse小波的散射变换网络及其在音乐体裁分类中的应用

链接:https://arxiv.org/abs/2206.07857

作者:Wai Ho Chak,Naoki Saito,David Weber
机构:Graduate Group in Applied Mathematics, University of California, Davis, USA.
摘要:我们建议使用广义莫尔斯小波(GMW)代替散射变换网络(STN)中常用的Morlet(或Gabor)小波,我们称之为GMW-STN,用于信号分类问题。GMW形成了一个真正解析小波的参数化族,而Morlet小波只是近似解析的。STN中底层小波滤波器的分析性对于非平稳振荡信号(如音乐信号)尤其重要,因为它通过提供输入信号的多尺度振幅和相位(以及频率)信息来提高STN表示的可解释性。我们使用所谓的GTZAN数据库证明了GMW-STN在音乐流派分类方面优于传统STN。此外,我们通过将GMW-STN的层数增加到典型的两层STN的三层,展示了GMW-STN的性能改进。}
摘要:We propose to use the Generalized Morse Wavelets (GMWs) instead of commonly-used Morlet (or Gabor) wavelets in the Scattering Transform Network (STN), which we call the GMW-STN, for signal classification problems. The GMWs form a parameterized family of truly analytic wavelets while the Morlet wavelets are only approximately analytic. The analyticity of underlying wavelet filters in the STN is particularly important for nonstationary oscillatory signals such as music signals because it improves interpretability of the STN representations by providing multiscale amplitude and phase (and consequently frequency) information of input signals. We demonstrate the superiority of the GMW-STN over the conventional STN in music genre classification using the so-called GTZAN database. Moreover, we show the performance improvement of the GMW-STN by increasing its number of layers to three over the typical two-layer STN.}


【7】 Paraformer: Fast and Accurate Parallel Transformer for  Non-autoregressive End-to-End Speech Recognition

标题:Paraformer:快速准确的非自回归端到端语音识别并行转换器

链接:https://arxiv.org/abs/2206.08317

作者:Zhifu Gao,Shiliang Zhang,Ian McLoughlin,Zhijie Yan
机构:Speech Lab, Alibaba Group, China, ICT Cluster, Singapore Institute of Technology, Singapore
备注:5 pages, 3 figures, accepted by InterSpeech2022
摘要:Transformer最近在ASR领域占据主导地位。虽然能够产生良好的性能,但它们涉及一个自回归(AR)解码器来逐个生成令牌,这在计算上效率很低。为了加快推理速度,设计了非自回归(NAR)方法,例如单步NAR,以实现并行生成。然而,由于输出标记内的独立性假设,单步NAR的性能不如AR模型,尤其是在大规模语料库中。改进单步NAR面临两个挑战:一是准确预测输出令牌数并提取隐藏变量;其次,增强输出标记之间相互依赖性的建模。为了应对这两个挑战,我们提出了一种快速准确的并联Transformer,称为并联Transformer。这利用了一个连续集成和基于fire的预测器来预测令牌的数量并生成隐藏变量。然后,GLM采样器生成语义嵌入,以增强NAR解码器对上下文相互依赖性建模的能力。最后,我们设计了一种生成负样本的策略,用于最小字错误率训练,以进一步提高性能。使用公共AISHELL-1、AISHELL-2基准和工业级20000小时任务进行的实验表明,所提出的并行器可以达到与最先进的ARTransformer相当的性能,加速比超过10倍。
摘要:Transformers have recently dominated the ASR field. Although able to yield good performance, they involve an autoregressive (AR) decoder to generate tokens one by one, which is computationally inefficient. To speed up inference, non-autoregressive (NAR) methods, e.g. single-step NAR, were designed, to enable parallel generation. However, due to an independence assumption within the output tokens, performance of single-step NAR is inferior to that of AR models, especially with a large-scale corpus. There are two challenges to improving single-step NAR: Firstly to accurately predict the number of output tokens and extract hidden variables; secondly, to enhance modeling of interdependence between output tokens. To tackle both challenges, we propose a fast and accurate parallel transformer, termed Paraformer. This utilizes a continuous integrate-and-fire based predictor to predict the number of tokens and generate hidden variables. A glancing language model (GLM) sampler then generates semantic embeddings to enhance the NAR decoder's ability to model context interdependence. Finally, we design a strategy to generate negative samples for minimum word error rate training to further improve performance. Experiments using the public AISHELL-1, AISHELL-2 benchmark, and an industrial-level 20,000 hour task demonstrate that the proposed Paraformer can attain comparable performance to the state-of-the-art AR transformer, with more than 10x speedup.


【8】 SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning

标题:SoundSpaces 2.0:视听学习模拟平台

链接:https://arxiv.org/abs/2206.08312

作者:Changan Chen,Carl Schissler,Sanchit Garg,Philip Kobernik,Alexander Clegg,Paul Calamia,Dhruv Batra,Philip W Robinson,Kristen Grauman
机构:Philip Robinson, UT Austin,  Reality Labs at Meta, Georgia Tech, Meta AI
备注:Website: this https URL
摘要:我们将介绍SoundSpaces 2.0,这是一个用于3D环境中基于动态几何体的音频渲染平台。给定真实世界环境的3D网格,Soundspace可以为从任意麦克风位置捕获的任意声音生成高度逼真的声学效果。它与现有的3D视觉资产一起,支持一系列视听研究任务,例如视听导航、映射、源定位和分离以及声学匹配。与现有资源相比,SoundSpaces 2.0具有以下优势:允许连续的空间采样、对新环境的泛化,以及可配置的麦克风和材料属性。据我们所知,这是第一次基于几何的声学模拟,它提供了高保真度和真实感,同时速度也足够快,可以用于具体学习。我们展示了模拟器的特性,并根据真实世界的音频测量对其性能进行了基准测试。此外,通过两个下游任务,包括嵌入式导航和远场自动语音识别,突出了后者的sim2real性能。SoundSpaces 2.0已公开提供,以促进对既能看到又能听到的感知系统进行更广泛的研究。
摘要:We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D visual assets, it supports an array of audio-visual research tasks, such as audio-visual navigation, mapping, source localization and separation, and acoustic matching. Compared to existing resources, SoundSpaces 2.0 has the advantages of allowing continuous spatial sampling, generalization to novel environments, and configurable microphone and material properties. To our best knowledge, this is the first geometry-based acoustic simulation that offers high fidelity and realism while also being fast enough to use for embodied learning. We showcase the simulator's properties and benchmark its performance against real-world audio measurements. In addition, through two downstream tasks covering embodied navigation and far-field automatic speech recognition, highlighting sim2real performance for the latter. SoundSpaces 2.0 is publicly available to facilitate wider research for perceptual systems that can both see and hear.


【9】 GoodBye WaveNet -- A Language Model for Raw Audio with Context of 1/2  Million Samples

标题:再见WaveNet--一个具有50万个样本的原始音频语言模型

链接:https://arxiv.org/abs/2206.08297

作者:Prateek Verma机
构:Stanford University
备注:12 pages, 1 figure. Technical Report at Stanford University. Ongoing Work
摘要:建模音频信号的长期依赖性是一个特别具有挑战性的问题,因为即使是很小的时间尺度也会产生大约十万个样本。随着最近Transformer的出现,神经架构变得善于在更长的时间尺度上建模依赖关系,但它们在缩放时受到二次约束。我们提出了一种生成式自回归架构,它可以在相当大的背景下(超过500000个样本)对音频波形进行建模。我们的工作适应于通过学习CNN前端的潜在表示来学习时间依赖性,然后使用Transformer编码器(经过充分训练的端到端)学习这些表示的依赖性:从而允许学习它认为适合下一个样本的表示。与之前比较不同时间尺度以显示改进的工作不同,我们使用标准数据集,使用相同数量的参数/上下文来显示改进。与Wavenet、SaSHMI和Sample RNN等其他方法相比,我们在用于长期结构建模的标准数据集上实现了最先进的性能。这项工作为该领域提供了非常令人兴奋的方向,因为上下文建模的改进可以用更多的数据进行缩放,并且通过使用数十亿/万亿的参数可能会获得更好的结果。
摘要:Modeling long-term dependencies for audio signals is a particularly challenging problem, as even small-time scales yield on the order of a hundred thousand samples. With the recent advent of Transformers, neural architectures became good at modeling dependencies over longer time scales, but they suffered from quadratic constraints to scale them. We propose a generative auto-regressive architecture that can model audio waveforms over quite a large context, greater than 500,000 samples. Our work is adapted to learn time dependencies by learning a latent representation by a CNN front-end, and then learning dependencies over these representations using Transformer encoders, fully trained end-to-end: thereby allowing to learn representations as it deems fit for the next sample. Unlike previous works that compared different time scales to show improvement, we use a standard dataset, with the same number of parameters/context to show improvements. We achieve a state-of-the-art performance as compared to other approaches such as Wavenet, SaSHMI, and Sample-RNN on a standard dataset for modeling long-term structure. This work gives very exciting direction for the field, given improvements in context modeling that can be scaled with more data, as well as potentially better results by using billions/trillions of parameters.


【10】 Event-related data conditioning for acoustic event classification

标题:用于声事件分类的事件相关数据条件

链接:https://arxiv.org/abs/2206.08233

作者:Yuanbo Hou,Dick Botteldooren
机构:WAVES, Ghent University, Belgium
备注:Accepted by INTERSPEECH 2022
摘要:基于不同注意机制的模型最近在与声学事件分类(AEC)相关的任务中大放异彩。其中,自我注意通常用于纯音频任务,以帮助模型识别不同的声学事件。自我注意依赖于时间框架之间的相似性,并使用整个片段的全局信息来突出框架内的特定特征。在现实生活中,与声学事件相关的信息会随时间衰减,这意味着事件周围某些帧内的信息比可能与事件无关的远距离全局信息更值得关注。本文表明,自我注意可能会过度增强某些音频表征片段,并平滑事件表征和背景噪声之间的边界。因此,本文提出了一种用于AEC的事件相关数据调节(EDC)。EDC直接处理光谱图。EDC的思想是基于声学特征自适应地选择与帧相关的注意范围,并收集事件相关的局部信息来表示帧。实验表明:1)与基于谱图的数据增强方法和可训练的特征加权和自我注意方法相比,EDC在原始尺寸模式和增强模式下均优于它们;2) EDC有效地收集与事件相关的本地信息,增强事件和背景之间的边界,提高AEC的性能。
摘要:Models based on diverse attention mechanisms have recently shined in tasks related to acoustic event classification (AEC). Among them, self-attention is often used in audio-only tasks to help the model recognize different acoustic events. Self-attention relies on the similarity between time frames, and uses global information from the whole segment to highlight specific features within a frame. In real life, information related to acoustic events will attenuate over time, which means the information within some frames around the event deserves more attention than distant time global information that may be unrelated to the event. This paper shows that self-attention may over-enhance certain segments of audio representations, and smooth out the boundaries between events representations and background noises. Hence, this paper proposes an event-related data conditioning (EDC) for AEC. EDC directly works on spectrograms. The idea of EDC is to adaptively select the frame-related attention range based on acoustic features, and gather the event-related local information to represent the frame. Experiments show that: 1) compared with spectrogram-based data augmentation methods and trainable feature weighting and self-attention, EDC outperforms them in both the original-size mode and the augmented mode; 2) EDC effectively gathers event-related local information and enhances boundaries between events and backgrounds, improving the performance of AEC.


【11】 Censer: Curriculum Semi-supervised Learning for Speech Recognition Based  on Self-supervised Pre-training

标题:基于自监督预训练的语音识别课程半监督学习

链接:https://arxiv.org/abs/2206.08189

作者:Bowen Zhang,Songjun Cao,Xiaoming Zhang,Yike Zhang,Long Ma,Takahiro Shinozaki
机构:Tokyo Institute of Technology, Tokyo, Japan, Tencent Cloud Xiaowei, Beijing, China
摘要:最近的研究表明,自我监督的预训练和自我训练(伪标记)所提供的益处是互补的。然而,在预训练框架下的半监督微调策略仍然没有得到足够的研究。此外,现代半监督语音识别算法要么不加区分地处理未标记的数据,要么用置信阈值过滤掉噪声样本。不同的未标记数据之间的差异常常被忽略。本文提出了一种基于自监督预训练的半监督语音识别算法Censer,以最大限度地提高未标记数据的利用率。Censer的预训练阶段采用wav2vec2.0,微调阶段采用改进的slimIPL半监督学习算法,该算法根据未标记数据的伪标签质量逐步利用未标记数据。我们还结合了一个时间伪标签池和一个指数移动平均值来控制伪标签的更新频率并避免模型发散。在Libri-Light和LibriSpeech数据集上的实验结果表明,与现有方法相比,我们提出的方法在更加统一的同时取得了更好的性能。
摘要:Recent studies have shown that the benefits provided by self-supervised pre-training and self-training (pseudo-labeling) are complementary. Semi-supervised fine-tuning strategies under the pre-training framework, however, remain insufficiently studied. Besides, modern semi-supervised speech recognition algorithms either treat unlabeled data indiscriminately or filter out noisy samples with a confidence threshold. The dissimilarities among different unlabeled data are often ignored. In this paper, we propose Censer, a semi-supervised speech recognition algorithm based on self-supervised pre-training to maximize the utilization of unlabeled data. The pre-training stage of Censer adopts wav2vec2.0 and the fine-tuning stage employs an improved semi-supervised learning algorithm from slimIPL, which leverages unlabeled data progressively according to their pseudo labels' qualities. We also incorporate a temporal pseudo label pool and an exponential moving average to control the pseudo labels' update frequency and to avoid model divergence. Experimental results on Libri-Light and LibriSpeech datasets manifest our proposed method achieves better performance compared to existing approaches while being more unified.


【12】 Adversarial Privacy Protection on Speech Enhancement

标题:语音增强中的对抗性隐私保护

链接:https://arxiv.org/abs/2206.08170

作者:Mingyu Dong,Diqun Yan,Rangding Wang
机构:College of Information Science and Engineering, Ningbo University
备注:5 pages, 6 figures
摘要:语音很容易在不知不觉中泄露,例如在不同情况下被手机录制。通过语音增强技术可以恶意提取语音中的私有内容。语音增强技术随着深度神经网络(DNN)的发展迅速,但对抗性的例子可能会导致DNN失败。在这项工作中,我们提出了一种对抗性的方法来降低语音增强系统的性能。实验结果表明,生成的对抗性示例可以擦除原始示例中的大部分内容信息,或者通过语音增强将其替换为目标语音内容。增强的原始示例和增强的对抗性示例识别结果之间的单词错误率(WER)可以达到89.0%。增强对抗示例和目标示例之间的目标攻击功率低至33.75%。对抗性扰动可以使原始示例的变化率超过1.4430。这项工作可以防止恶意的语音提取。
摘要:Speech is easily leaked imperceptibly, such as being recorded by mobile phones in different situations. Private content in speech may be maliciously extracted through speech enhancement technology. Speech enhancement technology has developed rapidly along with deep neural networks (DNNs), but adversarial examples can cause DNNs to fail. In this work, we propose an adversarial method to degrade speech enhancement systems. Experimental results show that generated adversarial examples can erase most content information in original examples or replace it with target speech content through speech enhancement. The word error rate (WER) between an enhanced original example and enhanced adversarial example recognition result can reach 89.0%. WER of target attack between enhanced adversarial example and target example is low to 33.75% . Adversarial perturbation can bring the rate of change to the original example to more than 1.4430. This work can prevent the malicious extraction of speech.


【13】 Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis  Using Linguistic and Prosodic Contexts of Dialogue History

标题:利用对话历史的语言和韵律语境进行端到端移情对话语音合成的声学建模

链接:https://arxiv.org/abs/2206.08039

作者:Yuto Nishimura,Yuki Saito,Shinnosuke Takamichi,Kentaro Tachibana,Hiroshi Saruwatari
机构:The University of Tokyo, Japan,LINE Corp., Japan.
备注:5 pages, 3 figures, Accepted for INTERSPEECH2022
摘要:我们提出了一个端到端移情对话语音合成(DSS)模型,该模型考虑了对话历史的语言和韵律背景。移情是人类在对话中主动尝试进入对话者内部,而移情DSS是在口语对话系统中实现这一行为的技术。我们的模型以语言和韵律特征的历史为条件,以预测合适的对话语境。因此,它可以看作是传统的基于语言特征的对话历史建模的扩展。为了有效地训练移情DSS模型,我们研究了1)用大型语音语料库预训练的自我监督学习模型,2)使用对话语境嵌入预测的当前话语韵律嵌入的风格引导训练,3)结合文本和语音模式的跨模态注意,4)句子嵌入,实现细粒度的韵律建模,而不是话语建模。评价结果表明:1)在移情决策支持系统中,单纯考虑对话历史的韵律语境并不能提高语音质量;2)引入风格引导训练和句子嵌入建模可以获得比传统方法更高的语音质量。
摘要:We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history. Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and empathetic DSS is a technology to implement this act in spoken dialogue systems. Our model is conditioned by the history of linguistic and prosody features for predicting appropriate dialogue context. As such, it can be regarded as an extension of the conventional linguistic-feature-based dialogue history modeling. To train the empathetic DSS model effectively, we investigate 1) a self-supervised learning model pretrained with large speech corpora, 2) a style-guided training using a prosody embedding of the current utterance to be predicted by the dialogue context embedding, 3) a cross-modal attention to combine text and speech modalities, and 4) a sentence-wise embedding to achieve fine-grained prosody modeling rather than utterance-wise modeling. The evaluation results demonstrate that 1) simply considering prosodic contexts of the dialogue history does not improve the quality of speech in empathetic DSS and 2) introducing style-guided training and sentence-wise embedding modeling achieves higher speech quality than that by the conventional method.


【14】 DCASE 2022: Comparative Analysis Of CNNs For Acoustic Scene  Classification Under Low-Complexity Considerations

标题:DCASE 2022:低复杂度条件下声场分类的CNN比较分析

链接:https://arxiv.org/abs/2206.08007

作者:Josep Zaragoza-Paredes,Javier Naranjo-Alcazar,Valery Naranjo,Pedro Zuccarello
机构:Pedro Zuccarello 2 1 Universitat Politecnica de Valencia
摘要:声学场景分类是一个自动监听问题,其目的是根据音频数据将音频记录分配给预定义的场景。多年来(以及在DCASE的过去版本中),这个问题通常通过称为集成的技术(使用多个机器学习模型在推理阶段组合预测)来解决。虽然这些解决方案可以在准确性方面显示性能,但在计算能力方面可能非常昂贵,因此无法在物联网设备中部署它们。由于这一研究领域的漂移,这项任务在模型复杂性方面有两个局限性。应该注意的是,失配设备也会增加复杂性(提供的音频由不同的信息源记录)。本技术报告对两种不同的网络架构进行了比较研究:传统的CNN和Conv混音器。虽然两个网络都超过了比赛要求的基线,但传统的有线电视新闻网表现出了更高的性能,超过了基线8个百分点。基于Conv-mixer体系结构的解决方案性能较差,尽管它们是更轻的解决方案。
摘要:Acoustic scene classification is an automatic listening problem that aims to assign an audio recording to a pre-defined scene based on its audio data. Over the years (and in past editions of the DCASE) this problem has often been solved with techniques known as ensembles (use of several machine learning models to combine their predictions in the inference phase). While these solutions can show performance in terms of accuracy, they can be very expensive in terms of computational capacity, making it impossible to deploy them in IoT devices. Due to the drift in this field of study, this task has two limitations in terms of model complexity. It should be noted that there is also the added complexity of mismatching devices (the audios provided are recorded by different sources of information). This technical report makes a comparative study of two different network architectures: conventional CNN and Conv-mixer. Although both networks exceed the baseline required by the competition, the conventional CNN shows a higher performance, exceeding the baseline by 8 percentage points. Solutions based on Conv-mixer architectures show worse performance although they are much lighter solutions.


【15】 Automatic Prosody Annotation with Pre-Trained Text-Speech Model

标题:基于预训练文本-语音模型的韵律自动标注

链接:https://arxiv.org/abs/2206.07956

作者:Ziqian Dai,Jianwei Yu,Yan Wang,Nuo Chen,Yanyao Bian,Guangzhi Li,Deng Cai,Dong Yu
机构:Tencent AI Lab,Peking University
备注:accepted by INTERSPEECH2022
摘要:韵律边界在文语合成(TTS)的自然性和可读性方面起着重要作用。然而,韵律边界标签的获取依赖于手动注释,这既昂贵又耗时。在这篇文章中,我们提出通过一个带有预训练音频编码器的神经文本语音模型,从文本音频数据中自动提取韵律边界标签。该模型分别对文本和语音数据进行预训练,并以三元组格式对TTS数据进行联合微调:{语音,文本,韵律}。自动评价和人工评价的实验结果表明:1)提出的文本-语音韵律标注框架显著优于纯文本基线;2) 自动韵律边界标注的质量与人类标注相当;3) 使用模型注释边界训练的TTS系统略优于使用手动边界的系统。
摘要:Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming. In this paper, we propose to automatically extract prosodic boundary labels from text-audio data via a neural text-speech model with pre-trained audio encoders. This model is pre-trained on text and speech data separately and jointly fine-tuned on TTS data in a triplet format: {speech, text, prosody}. The experimental results on both automatic evaluation and human evaluation demonstrate that: 1) the proposed text-speech prosody annotation framework significantly outperforms text-only baselines; 2) the quality of automatic prosodic boundary annotations is comparable to human annotations; 3) TTS systems trained with model-annotated boundaries are slightly better than systems that use manual ones.


【16】 Accelerating Inference and Language Model Fusion of Recurrent Neural  Network Transducers via End-to-End 4-bit Quantization

标题:端到端4位量化加速递归神经网络传感器的推理和语言模型融合

链接:https://arxiv.org/abs/2206.07882

作者:Andrea Fasoli,Chia-Yu Chen,Mauricio Serrano,Swagath Venkataramani,George Saon,Xiaodong Cui,Brian Kingsbury,Kailash Gopalakrishnan
机构:IBM Research, USA
备注:5 pages, 2 figures, 1 table. Paper accepted to Interspeech 2022
摘要:我们报告了积极的量化策略,大大加快了递归神经网络传感器(RNN-T)的推理。我们对权重和激活使用4位整数表示,并应用量化感知训练(QAT)对完整模型(声学编码器和语言模型)进行重新训练,以达到接近iso的精度。我们表明,根据网络的局部特性定制的量化方案对于在限制QAT计算开销的同时获得良好性能至关重要。密度比语言模型融合在RNN-T工作负载上显示出显著的精度提高,但它严重增加了推理的计算成本。我们表明,我们的量化策略能够使用大波束宽度进行假设搜索,同时实现流兼容的运行时,与全精度模型相比,全模型压缩比为7.6$\倍。通过硬件模拟,我们估计端到端量化RNN-T(包括LM融合)从FP16到INT4的加速度为3.4$\倍,导致实时因子(RTF)为0.06。在NIST Hub5 2000、Hub5 2001和RT-03测试集上,我们保留了与LM fusion相关的大部分收益,将平均WER提高了$>$1.5%。
摘要:We report on aggressive quantization strategies that greatly accelerate inference of Recurrent Neural Network Transducers (RNN-T). We use a 4 bit integer representation for both weights and activations and apply Quantization Aware Training (QAT) to retrain the full model (acoustic encoder and language model) and achieve near-iso-accuracy. We show that customized quantization schemes that are tailored to the local properties of the network are essential to achieve good performance while limiting the computational overhead of QAT.  Density ratio Language Model fusion has shown remarkable accuracy gains on RNN-T workloads but it severely increases the computational cost of inference. We show that our quantization strategies enable using large beam widths for hypothesis search while achieving streaming-compatible runtimes and a full model compression ratio of 7.6$\times$ compared to the full precision model.  Via hardware simulations, we estimate a 3.4$\times$ acceleration from FP16 to INT4 for the end-to-end quantized RNN-T inclusive of LM fusion, resulting in a Real Time Factor (RTF) of 0.06. On the NIST Hub5 2000, Hub5 2001, and RT-03 test sets, we retain most of the gains associated with LM fusion, improving the average WER by $>$1.5%.


【17】 EPG2S: Speech Generation and Speech Enhancement based on  Electropalatography and Audio Signals using Multimodal Learning

标题:EPG2S:基于腭电和音频信号的多模式学习语音生成与增强

链接:https://arxiv.org/abs/2206.07860

作者:Li-Chin Chen,Po-Hsun Chen,Richard Tzong-Han Tsai,Yu Tsao
备注:Accepted By IEEE Signal Processing Letter
摘要:基于发音运动的语音生成和增强有助于在缺乏言语交流的情况下进行交流,例如在丧失说话能力的患者中。尽管为此提出了各种技术,但腭电图(EPG)作为一种记录语言过程中舌头和硬腭之间接触的监测技术,尚未得到充分的探索。在此,我们提出了一种新的多模式EPG-to-speech(EPG2S)系统,该系统利用EPG和语音信号进行语音生成和增强。研究了基于EPG和噪声语音信号多重组合的不同融合策略,并对该方法的可行性进行了研究。实验结果表明,EPG2S仅基于EPG信号就可以获得理想的语音生成效果。此外,观察到添加噪声语音信号以提高质量和可懂度。此外,观察到EPG2S可实现仅基于音频信号的高质量语音增强,添加EPG信号可进一步提高性能。延迟融合策略被认为是同步语音生成和增强的最有效方法。
摘要:Speech generation and enhancement based on articulatory movements facilitate communication when the scope of verbal communication is absent, e.g., in patients who have lost the ability to speak. Although various techniques have been proposed to this end, electropalatography (EPG), which is a monitoring technique that records contact between the tongue and hard palate during speech, has not been adequately explored. Herein, we propose a novel multimodal EPG-to-speech (EPG2S) system that utilizes EPG and speech signals for speech generation and enhancement. Different fusion strategies based on multiple combinations of EPG and noisy speech signals are examined, and the viability of the proposed method is investigated. Experimental results indicate that EPG2S achieves desirable speech generation outcomes based solely on EPG signals. Further, the addition of noisy speech signals is observed to improve quality and intelligibility. Additionally, EPG2S is observed to achieve high-quality speech enhancement based solely on audio signals, with the addition of EPG signals further improving the performance. The late fusion strategy is deemed to be the most effective approach for simultaneous speech generation and enhancement.


机器翻译,仅供参考