今天跟大家分享一篇语音相关的论文合集:cs.SD语音17篇,eess.AS音频处理17篇。本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
【1】 Improving generalizability of distilled self-supervised speech processing models under distorted settings
标题:提高失真环境下提取的自监督语音处理模型的泛化能力
链接:https://arxiv.org/abs/2210.07978
作者:Kuan-Po Huang,Yu-Kuan Fu,Tsu-Yuan Hsu,Fabian Ritter Gutierrez,Fan-Lin Wang,Liang-Hsuan Tseng,Yu Zhang,Hung-yi Lee机构:National Taiwan University, ASUS Intelligent Cloud Services, Nanyang Technological University, Google Brain备注:Accepted by IEEE SLT2022摘要:自监督学习(SSL)语音预训练模型在各种语音处理任务中表现良好。已开发出SSL模型的精简版本,以满足设备上语音应用程序的需求。尽管与原始SSL模型具有相似的性能,但在失真环境中,经过提炼的对应模型的性能下降甚至比原始版本更严重。本文提出在SSL模型的知识提取过程中引入交叉失真映射和领域对抗训练,以缓解领域不匹配问题带来的性能差距。结果表明,在域内和域外失真设置下,对于不同的下游任务,在保持有效的模型大小的同时,性能得到了一致的改善。摘要:Self-supervised learned (SSL) speech pre-trained models perform well across various speech processing tasks. Distilled versions of SSL models have been developed to match the needs of on-device speech applications. Though having similar performance as original SSL models, distilled counterparts suffer from performance degradation even more than their original versions in distorted environments. This paper proposes to apply Cross-Distortion Mapping and Domain Adversarial Training to SSL models during knowledge distillation to alleviate the performance gap caused by the domain mismatch problem. Results show consistent performance improvements under both in- and out-of-domain distorted setups for different downstream tasks while keeping efficient model size.
【2】 Bringing NURC/SP to Digital Life: the Role of Open-source Automatic Speech Recognition Models
标题:将NURC/SP带入数字生活:开源自动语音识别模型的作用
链接:https://arxiv.org/abs/2210.07852
作者:Lucas Rafael Stefanel Gris,Arnaldo Candido Junior,Vinícius G. dos Santos,Bruno A. Papa Dias,Marli Quadros Leite,Flaviane Romani Fernandes Svartman,Sandra Aluísio机构:Federal University of Goi´as, Brazil, S˜ao Paulo State University, Brazil, University of S˜ao Paulo, Brazil, lucas.gris(at)ufg.discente.br, arnaldo.candido(at)unesp.br, sandra(at)icmc.usp.br摘要:1969年开始的NURC项目研究了巴西五个首都的城市文化语言规范,负责为每个首都汇编一个大型语料库。数字化的NURC/SP包括在圣保罗首都拍摄的334小时记录中的375个查询。虽然47项询问有录音誊本,但录音誊本之间没有对齐,328项询问没有录音誊本。本文对三个用葡萄牙语自发语音训练的自动语音识别模型和一个用准备语音训练的模型进行了评价和误差分析。通过评估,我们可以使用WER和CER指标,在NURC/SP的手动对齐样本中选择最佳模型,以自动转录284小时。摘要:The NURC Project that started in 1969 to study the cultured linguistic urban norm spoken in five Brazilian capitals, was responsible for compiling a large corpus for each capital. The digitized NURC/SP comprises 375 inquiries in 334 hours of recordings taken in S\~ao Paulo capital. Although 47 inquiries have transcripts, there was no alignment between the audio-transcription, and 328 inquiries were not transcribed. This article presents an evaluation and error analysis of three automatic speech recognition models trained with spontaneous speech in Portuguese and one model trained with prepared speech. The evaluation allowed us to choose the best model, using WER and CER metrics, in a manually aligned sample of NURC/SP, to automatically transcribe 284 hours.
【3】 Contrastive Audio-Visual Masked Autoencoder
标题:对比式视听屏蔽式自动编码器
链接:https://arxiv.org/abs/2210.07839
作者:Yuan Gong,Andrew Rouditchenko,Alexander H. Liu,David Harwath,Leonid Karlinsky,Hilde Kuehne,James Glass机构:MIT CSAIL; ,UT Austin; ,MIT-IBM Watson AI Lab; ,Goethe University Frankfurt摘要:本文首先将最新的掩蔽自动编码器(MAE)模型从单模态扩展到视听多模态。在此基础上,结合对比学习和掩蔽数据建模这两种主要的自监督学习框架,提出了对比视听掩蔽自动编码器(CAV-MAE),用于学习联合协调的视听表示。实验结果表明,对比性视听对应学习目标不仅能使模型完成视听检索任务,而且能帮助模型学习到更好的联合表征。结果表明,在VGGSound上,我们的完全自监督预训练CAV-MAE达到了65.9%的新SOTA准确率,并且在视听事件分类任务中与之前在AudioSet上的最佳监督预训练模型相当。摘要:In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by combining contrastive learning and masked data modeling, two major self-supervised learning frameworks, to learn a joint and coordinated audio-visual representation. Our experiments show that the contrastive audio-visual correspondence learning objective not only enables the model to perform audio-visual retrieval tasks, but also helps the model learn a better joint representation. As a result, our fully self-supervised pretrained CAV-MAE achieves a new SOTA accuracy of 65.9% on VGGSound, and is comparable with the previous best supervised pretrained model on AudioSet in the audio-visual event classification task.
【4】 Intel Labs at Ego4D Challenge 2022: A Better Baseline for Audio-Visual Diarization
标题:英特尔实验室参加2022年Ego4D挑战赛:视听失真的更好基准
链接:https://arxiv.org/abs/2210.07764
备注:Validation report for the Ego4D challenge at ECCV 2022摘要:本报告描述了我们在2022年Ego4D挑战赛中完成视听诊断(AVD)任务的方法。具体而言,我们在官方基准基础上进行了多项技术改进。首先,通过修改模型的训练方案,提高了摄像头佩戴者语音活动的检测性能。第二,我们发现现成的语音活动检测模型在仅应用于相机佩戴者的语音活动时可以有效地去除假阳性。最后,我们证明了更好的活动说话人检测导致更好的AVD结果。我们的最终方法在Ego4D的测试集上获得了65.9%的DER,显著优于所有基线。我们的作品在2022年Ego4D挑战赛中获得第一名。摘要:This report describes our approach for the Audio-Visual Diarization (AVD) task of the Ego4D Challenge 2022. Specifically, we present multiple technical improvements over the official baselines. First, we improve the detection performance of the camera wearer's voice activity by modifying the training scheme of its model. Second, we discover that an off-the-shelf voice activity detection model can effectively remove false positives when it is applied solely to the camera wearer's voice activities. Lastly, we show that better active speaker detection leads to a better AVD outcome. Our final method obtains 65.9% DER on the test set of Ego4D, which significantly outperforms all the baselines. Our submission achieved 1st place in the Ego4D Challenge 2022.
【5】 Accelerating RNN-based Speech Enhancement on a Multi-Core MCU with Mixed FP16-INT8 Post-Training Quantization
标题:基于混合FP16-INT8后训练量化的多核MCU加速RNN语音增强
链接:https://arxiv.org/abs/2210.07692
作者:Manuele Rusci,Marco Fariselli,Martin Croome,Francesco Paci,Eric Flamand机构:Universita’ di Bologna, Bologna, ITA, Greenwaves Technologies, Grenoble, FRA备注:Accepted at the ITEM Workshop 2022 (located at ECML-PKDD2022)摘要:提出了一种基于递归神经网络(RNN)的语音增强(SE)算法的优化设计方法,并将其部署在具有1+8通用RISC-V核的先进微控制器单元(MCU)上。为了实现低延迟执行,我们提出了一种优化的软件流水线交错并行计算LSTM或GRU循环块,具有矢量化的8位整数(INT 8)和16位浮点(FP 16)计算单元,以及手动管理的模型参数内存传输。为了保证相对于全精度模型的最小精度下降,我们提出了一种新的FP 16-INT 8混合精度训练后量化(PTQ)方案,该方案将递归层压缩到8位,而其余层的位精度保持为FP 16。实验在Valentini数据集上训练的多个基于LSTM和GRU的SE模型上进行,具有多达1.24M个参数。由于所提出的方法,相对于无损FP 16基线,我们将计算速度提高了4倍。与使PESQ分数平均降低0.3的均匀8位量化不同,混合精度PTQ方案导致仅0.06的低降低,同时实现1.4- 1.7倍的存储器节省。得益于这种压缩,我们通过在有限的片内非易失性存储器上安装大型模型,降低了外部存储器的功耗成本;通过将电源电压从0.8V降至0.65V,MCU的功耗最多可降低2.5倍,同时仍能满足实时性要求。与部署在单核MCU上的最先进SE解决方案相比,我们的设计能效高出10倍,这些解决方案利用了更小的模型和量化感知训练。摘要:This paper presents an optimized methodology to design and deploy Speech Enhancement (SE) algorithms based on Recurrent Neural Networks (RNNs) on a state-of-the-art MicroController Unit (MCU), with 1+8 general-purpose RISC-V cores. To achieve low-latency execution, we propose an optimized software pipeline interleaving parallel computation of LSTM or GRU recurrent blocks, featuring vectorized 8-bit integer (INT8) and 16-bit floating-point (FP16) compute units, with manually-managed memory transfers of model parameters. To ensure minimal accuracy degradation with respect to the full-precision models, we propose a novel FP16-INT8 Mixed-Precision Post-Training Quantization (PTQ) scheme that compresses the recurrent layers to 8-bit while the bit precision of remaining layers is kept to FP16. Experiments are conducted on multiple LSTM and GRU based SE models trained on the Valentini dataset, featuring up to 1.24M parameters. Thanks to the proposed approaches, we speed-up the computation by up to 4x with respect to the lossless FP16 baselines. Differently from a uniform 8-bit quantization that degrades the PESQ score by 0.3 on average, the Mixed-Precision PTQ scheme leads to a low-degradation of only 0.06, while achieving a 1.4-1.7x memory saving. Thanks to this compression, we cut the power cost of the external memory by fitting the large models on the limited on-chip non-volatile memory and we gain a MCU power saving of up to 2.5x by reducing the supply voltage from 0.8V to 0.65V while still matching the real-time constraints. Our design results 10x more energy efficient than state-of-the-art SE solutions deployed on single-core MCUs that make use of smaller models and quantization-aware training.
【6】 Full-Stack Bioacoustics: Field Kit to AI to Action (Workshop report)
标题:全套生物声学:从现场套件到人工智能到行动(研讨会报告)
链接:https://arxiv.org/abs/2210.07685
作者:Dan Stowell,Caitlin Black,Florencia Noriega,Sarab S. Sethi机构:•, Sarab Sethi, University of Cambridge, This report contains an overview of the workshop aims and structure, as well as, reports from the six groups., Scientific case备注:Workshop report: Lorentz Center, Leiden, the Netherlands, 1-5 August 2022摘要:声学数据(声音记录)是探测、计数和区分野生动物的重要证据来源。由于信号处理和机器学习、记录设备以及数据处理和存储能力的巨大进步,“生物声学”领域在过去十年中得到了发展。许多研究论文描述了Raspberry Pi或类似设备用于声学监测的用途,并且其他研究论文描述了通过机器学习对动物声音的自动分类。但对于大多数生态学家、动物学家、自然资源保护主义者来说,这些拼图并没有拼在一起:域被分割。在这次洛伦兹研讨会上,我们将汇集生物声学监测和机器学习领域的开放硬件和开源软件的主要代表,以及生态学家和其他领域的研究人员,以弥合这一差距。我们在分享技能的同时,也为“生物声学AI”的未来发展建立了愿景。 本报告概述了讲习班的目标和结构,以及六个小组的报告。摘要:Acoustic data (sound recordings) are a vital source of evidence for detecting, counting, and distinguishing wildlife. This domain of "bioacoustics" has grown in the past decade due to the massive advances in signal processing and machine learning, recording devices, and the capacity of data processing and storage. Numerous research papers describe the use of Raspberry Pi or similar devices for acoustic monitoring, and other research papers describe automatic classification of animal sounds by machine learning. But for most ecologists, zoologists, conservationists, the pieces of the puzzle do not come together: the domain is fragmented. In this Lorentz workshop we bridge this gap by bringing together leading exponents of open hardware and open-source software for bioacoustic monitoring and machine learning, as well as ecologists and other field researchers. We share skills while also building a vision for the future development of "bioacoustic AI". This report contains an overview of the workshop aims and structure, as well as reports from the six groups.
【7】 Training speech emotion classifier without categorical annotations
标题:训练不带类别标注的语音情感分类器
链接:https://arxiv.org/abs/2210.07642
作者:Meysam Shamsi,Marie Tahon机构:LIUM, Le Mans University, Avenue Olivier Messiaen, Le Mans, France摘要:情感表征有两种范式:范畴标注和连续空间维度描述。因此,情绪识别任务可以被视为分类或回归。本研究的主要目的是研究这两种表示法之间的关系,并提出一种只使用维度标注的分类管道。提出的方法包含一个回归器模型,该模型被训练来预测给定语音音频的维度表示中的连续值的向量。该模型的输出可以使用映射算法解释为情感类别。我们研究了三种特征提取器、三种神经网络结构和三种映射算法在两个不同语料库上的性能。我们的研究显示了通过回归方法进行分类的优点和局限性。摘要:There are two paradigms of emotion representation, categorical labeling and dimensional description in continuous space. Therefore, the emotion recognition task can be treated as a classification or regression. The main aim of this study is to investigate the relation between these two representations and propose a classification pipeline that uses only dimensional annotation. The proposed approach contains a regressor model which is trained to predict a vector of continuous values in dimensional representation for given speech audio. The output of this model can be interpreted as an emotional category using a mapping algorithm. We investigated the performances of a combination of three feature extractors, three neural network architectures, and three mapping algorithms on two different corpora. Our study shows the advantages and limitations of the classification via regression approach.
【8】 Empirical Study Incorporating Linguistic Knowledge on Filled Pauses for Personalized Spontaneous Speech Synthesis
标题:个性化自然语音合成中填充停顿与语言学知识结合的实证研究
链接:https://arxiv.org/abs/2210.07559
作者:Yuta Matsunaga,Takaaki Saeki,Shinnosuke Takamichi,Hiroshi Saruwatari机构:∗ Graduate School of Information Science and Technology, The University of Tokyo, Japan.备注:Accepted to APSIPA ASC 2022摘要:我们提出了一个全面的实证研究个性化自发语音合成的基础上的语言知识。随着用于阅读式语音 合成的声音克隆的出现 , 需要一种用于类人和自发语音合成的新的声音克隆范例。因此,我们将重点放在个人化的自发语音合成上 ,它可以同时复制个人的声音音色和语音 不 流畅 。具体来说 , 我们研究 的 是填充停顿,它是言语不流畅的主要来源,在心理学和语言学中,它在言语生成和交际中起着重要作用 。为了比较个性化填充停顿插入和非个性化填充 停顿 预测 方法,提出了一种基于多说话人语料库训练 的 非个性化外部填充停顿 预测器的语音合成方法 。研究结果表明,填充停顿的位置 - 词纠缠 ,在合成语音的评价中,为了自然性而精确预测位置的必要 性和为了个性而精确预测单词的必要性 。摘要:We present a comprehensive empirical study for personalized spontaneous speech synthesis on the basis of linguistic knowledge. With the advent of voice cloning for reading-style speech synthesis, a new voice cloning paradigm for human-like and spontaneous speech synthesis is required. We, therefore, focus on personalized spontaneous speech synthesis that can clone both the individual's voice timbre and speech disfluency. Specifically, we deal with filled pauses, a major source of speech disfluency, which is known to play an important role in speech generation and communication in psychology and linguistics. To comparatively evaluate personalized filled pause insertion and non-personalized filled pause prediction methods, we developed a speech synthesis method with a non-personalized external filled pause predictor trained with a multi-speaker corpus. The results clarify the position-word entanglement of filled pauses, i.e., the necessity of precisely predicting positions for naturalness and the necessity of precisely predicting words for individuality on the evaluation of synthesized speech.
【9】 Transformer-Based Speech Synthesizer Attribution in an Open Set Scenario
标题:开集场景下基于变换的语音合成器属性
链接:https://arxiv.org/abs/2210.07546
作者:Emily R. Bartusiak,Edward J. Delp机构:Video and Image Processing Lab, School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN摘要:语音合成方法可以创建听起来逼真的语音,其可用于欺诈、欺骗和误导活动。检测合成语音的取证方法对于防止这种攻击是重要的。取证归属方法提供关于合成语音信号的性质的甚至更多信息,语音合成器),用于创建语音信号。由于越来越多的真实声音的语音合成器,我们提出了一种语音归属的方法,推广到新的合成器没有看到在训练。为了做到这一点,我们研究了语音合成器在封闭集和开放集两种情况下的归因。换句话说,我们认为一些语音合成器是“已知的”合成器(即,而其它的合成器是“未知”合成器(即,开集的一部分)。我们将语音信号表示为谱图,并在闭集上训练我们提出的方法,称为紧凑属性Transformer(CAT),用于多类分类。然后,我们将我们的分析扩展到开集,以将合成的语音信号归属于已知和未知合成器。我们在训练过的CAT的潜在空间上利用t分布随机邻居嵌入(tSNE)来区分每个未知合成器。此外,我们还探索了poly-1丢失配方以改善归因结果。我们提出的方法在闭集和开集两种情况下都成功地将合成的语音信号归属到它们各自的语音合成器。摘要:Speech synthesis methods can create realistic-sounding speech, which may be used for fraud, spoofing, and misinformation campaigns. Forensic methods that detect synthesized speech are important for protection against such attacks. Forensic attribution methods provide even more information about the nature of synthesized speech signals because they identify the specific speech synthesis method (i.e., speech synthesizer) used to create a speech signal. Due to the increasing number of realistic-sounding speech synthesizers, we propose a speech attribution method that generalizes to new synthesizers not seen during training. To do so, we investigate speech synthesizer attribution in both a closed set scenario and an open set scenario. In other words, we consider some speech synthesizers to be "known" synthesizers (i.e., part of the closed set) and others to be "unknown" synthesizers (i.e., part of the open set). We represent speech signals as spectrograms and train our proposed method, known as compact attribution transformer (CAT), on the closed set for multi-class classification. Then, we extend our analysis to the open set to attribute synthesized speech signals to both known and unknown synthesizers. We utilize a t-distributed stochastic neighbor embedding (tSNE) on the latent space of the trained CAT to differentiate between each unknown synthesizer. Additionally, we explore poly-1 loss formulations to improve attribution results. Our proposed approach successfully attributes synthesized speech signals to their respective speech synthesizers in both closed and open set scenarios.
【10】 Hierarchical Diffusion Models for Singing Voice Neural Vocoder
标题:歌唱语音神经声码器的分层扩散模型
链接:https://arxiv.org/abs/2210.07508
作者:Naoya Takahashi,Mayank Kumar,Singh,Yuki Mitsufuji机构:Sony Group Corporation, Japan摘要:深度生成模型的最新进展提高了语音域神经声码器的质量。然而,由于音高、响度和发音方面的音乐表现形式的多样性,产生高质量的歌唱声音仍然具有挑战性。在这项工作中,我们提出了一个层次扩散模型的歌唱声神经声码器。所提出的方法由以不同采样率操作的多个扩散模型组成;最低采样率下的模型集中于生成准确的低频分量,例如音调,而其它模型基于较低采样率下的数据和声学特征逐渐生成较高采样率下的波形。实验结果表明,该方法能够为多个演唱者生成高质量的演唱声,在计算量相近的情况下,其性能优于现有的神经声码器。摘要:Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, it remains challenging to generate high-quality singing voice due to a wider variety of musical expressions in pitch, loudness, and pronunciations. In this work, we propose a hierarchical diffusion model for singing voice neural vocoders. The proposed method consists of multiple diffusion models operating in different sampling rates; the model at the lowest sampling rate focuses on generating accurate low frequency components such as pitch, and other models progressively generate the waveform at the higher sampling rates based on the data at the lower sampling rate and acoustic features. Experimental results show that the proposed method produces high-quality singing voice for multiple singers, outperforming state-of-the-art neural vocoders with a similar range of computational costs.
【11】 Bayes risk CTC: Controllable CTC alignment in Sequence-to-Sequence tasks
标题:贝叶斯风险CTC:序列到序列任务中的可控CTC比对
链接:https://arxiv.org/abs/2210.07499
作者:Jinchuan Tian,Brian Yan,Jianwei Yu,Chao Weng,Dong Yu,Shinji Watanabe机构:Tencent AI LAB, Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA , USA摘要:序列到序列(seq 2seq)任务将输入序列转录为靶序列。连接主义时态分类(CTC)准则被广泛应用于多个seq 2seq任务。除了预测目标序列之外,CTC的副产品是预测比对,比对是最可能的输入长序列,其指定输入单元和目标单元之间的硬比对关系。由于在CTC公式中同等考虑了多个潜在的比对序列(称为路径),因此选择哪条路径最有可能成为预测的比对总是不确定的。此外,通常观察到,由vanilla CTC预测的比对与其参考相比将漂移,并且很少提供实际功能。因此,这项工作的动机是使CTC对准预测可控,从而为CTC配备额外的功能。在此基础上,提出了贝叶斯风险CTC(BRCTC)准则,该准则采用一个可定制的贝叶斯风险函数来增强预测序列的期望特性. BRCTC是一个通用的框架,通过风险函数对路径采用一些可定制的偏好,以便将后验概率集中到路径的特定子集中。在应用中,我们探索了一种特殊的偏好,它产生的模型具有下采样能力和降低的推理成本。通过使用BRCTC和另一个早期发射偏好,我们获得了在线模型的改进的性能-延迟权衡。实验结果表明,BRCTC算法在不降低系统性能的前提下,将离线模型的推理代价降低了47%,并将在线系统的整体延迟降低到了一个不可见的水平.摘要:Sequence-to-Sequence (seq2seq) tasks transcribe the input sequence to a target sequence. The Connectionist Temporal Classification (CTC) criterion is widely used in multiple seq2seq tasks. Besides predicting the target sequence, a side product of CTC is to predict the alignment, which is the most probable input-long sequence that specifies a hard aligning relationship between the input and target units. As there are multiple potential aligning sequences (called paths) that are equally considered in CTC formulation, the choice of which path will be most probable and become the predicted alignment is always uncertain. In addition, it is usually observed that the alignment predicted by vanilla CTC will drift compared with its reference and rarely provides practical functionalities. Thus, the motivation of this work is to make the CTC alignment prediction controllable and thus equip CTC with extra functionalities. The Bayes risk CTC (BRCTC) criterion is then proposed in this work, in which a customizable Bayes risk function is adopted to enforce the desired characteristics of the predicted alignment. With the risk function, the BRCTC is a general framework to adopt some customizable preference over the paths in order to concentrate the posterior into a particular subset of the paths. In applications, we explore one particular preference which yields models with the down-sampling ability and reduced inference costs. By using BRCTC with another preference for early emissions, we obtain an improved performance-latency trade-off for online models. Experimentally, the proposed BRCTC reduces the inference cost of offline models by up to 47% without performance degradation and cuts down the overall latency of online systems to an unseen level.
【12】 JOIST: A Joint Speech and Text Streaming Model For ASR
标题:Joist:一种面向ASR的语音和文本联合流媒体模型
链接:https://arxiv.org/abs/2210.07353
作者:Tara N. Sainath,Rohit Prabhavalkar,Ankur Bapna,Yu Zhang,Zhouyuan Huo,Zhehuai Chen,Bo Li,Weiran Wang,Trevor Strohman摘要:本文提出了一种新的编码算法JOIST,该算法可以训练一个具有语音-文本成对输入和仅文本非成对输入的流、级联、端到端(E2 E)编码器模型。与以往的研究不同,我们探索了两种模式的联合训练,而不是预先训练和微调。此外,我们使用一个流E2 E模型来探索JOIST,数据量增加了一个数量级,这与以前的工作相比也是新颖的。通过一系列的烧蚀研究,我们探索了不同类型的文本建模,包括如何对文本序列的长度进行建模以及合适的文本子词单元表示。我们发现,与未使用文本训练的模型相比,JOIST的最佳文本表示在各种搜索和稀有词测试集上将WER提高了4-14%。此外,我们定量地表明,JOIST保持了流功能,这对良好的用户级体验很重要。摘要:We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous works, we explore joint training with both modalities, rather than pre-training and fine-tuning. In addition, we explore JOIST using a streaming E2E model with an order of magnitude more data, which are also novelties compared to previous works. Through a series of ablation studies, we explore different types of text modeling, including how to model the length of the text sequence and the appropriate text sub-word unit representation. We find that best text representation for JOIST improves WER across a variety of search and rare-word test sets by 4-14% relative, compared to a model not trained with text. In addition, we quantitatively show that JOIST maintains streaming capabilities, which is important for good user-level experience.
【13】 HuBERT-TR: Reviving Turkish Automatic Speech Recognition with Self-supervised Speech Representation Learning标题:Hubert-tr:用自监督语音表征学习恢复土耳其语自动语音识别链接:https://arxiv.org/abs/2210.07323
作者:Ali Safaya,Engin Erzin机构:KUIS AI Center, Computer Engineering Department, Koc¸ University摘要:虽然土耳其语被列为低资源语言之一,但有关土耳其语自动语音识别(ASR)的文献相对较早。本文提出了一种基于HuBERT的土耳其语语音表示模型HuBERT-TR。HuBERT-TR在几个土耳其ASR数据集上获得了最先进的结果。我们使用从在线资源中收集的大规模数据研究了土耳其语的预训练HuBERT。我们使用从YouTube上收集的超过6,500小时的语音数据对HuBERT-TR进行预训练,这些数据在质量和类型方面具有广泛的可变性。我们表明,多语言设置中的预训练模型劣于特定语言模型,其中我们的土耳其语模型HuBERT-TR base的性能优于其10倍大的多语言对应物XLS-R-1B。此外,我们通过将我们的模型缩放到1B个参数来研究缩放对ASR性能的影响。我们的最佳模型在Turkish Broadcast News数据集上产生了4.97%的最先进的单词错误率。有关型号的信息,请访问huggingface.co/asafaya。摘要:While the Turkish language is listed among low-resource languages, literature on Turkish automatic speech recognition (ASR) is relatively old. In this paper, we present HuBERT-TR, a speech representation model for Turkish based on HuBERT. HuBERT-TR achieves state-of-the-art results on several Turkish ASR datasets. We investigate pre-training HuBERT for Turkish with large-scale data curated from online resources. We pre-train HuBERT-TR using over 6,500 hours of speech data curated from YouTube that includes extensive variability in terms of quality and genre. We show that pre-trained models within a multi-lingual setup are inferior to language-specific models, where our Turkish model HuBERT-TR base performs better than its x10 times larger multi-lingual counterpart XLS-R-1B. Moreover, we study the effect of scaling on ASR performance by scaling our models up to 1B parameters. Our best model yields a state-of-the-art word error rate of 4.97% on the Turkish Broadcast News dataset. Models are available at huggingface.co/asafaya .
【14】 Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system标题:描述和分析DCASE关于基线系统的任务4 2022中引入的新内容链接:https://arxiv.org/abs/2210.07856
作者:Francesca Ronchini,Samuele Cornell,Romain Serizel,Nicolas Turpault,Eduardo Fonseca,Daniel P. W. Ellis机构:Universite de Lorraine, CNRS, Inria, Loria, Nancy, France, Department of Information Engineering, Italy, Google Research, United States摘要:声学场景和事件的检测和分类挑战任务4的目的是使用异构数据集评估家庭环境中声音事件检测系统。系统需要能够正确地检测记录的音频剪辑中存在的声音事件,以及及时地定位事件。今年的任务是DCASE 2021 Task 4的后续,有一些重要的新功能。本文的目标是描述和激励这些新添加的内容,并报告它们对基线系统的影响的分析。我们推出了三个主要的新功能:外部数据集的使用,包括最近发布的来自Audioset的强注释剪辑,利用预先训练的模型的可能性,以及新的能耗度量,以提高对训练声音事件检测器的生态影响的认识。基线系统的结果表明,利用AudioSet上的开源预训练可以显著改善事件分类的结果,但不能改善事件分割的结果。摘要:The aim of the Detection and Classification of Acoustic Scenes and Events Challenge Task 4 is to evaluate systems for the detection of sound events in domestic environments using an heterogeneous dataset. The systems need to be able to correctly detect the sound events present in a recorded audio clip, as well as localize the events in time. This year's task is a follow-up of DCASE 2021 Task 4, with some important novelties. The goal of this paper is to describe and motivate these new additions, and report an analysis of their impact on the baseline system. We introduced three main novelties: the use of external datasets, including recently released strongly annotated clips from Audioset, the possibility of leveraging pre-trained models, and a new energy consumption metric to raise awareness about the ecological impact of training sound events detectors. The results on the baseline system show that leveraging open-source pretrained on AudioSet improves the results significantly in terms of event classification but not in terms of event segmentation.
【15】 Learning to Jointly Transcribe and Subtitle for End-to-End Spontaneous Speech Recognition标题:学习联合转录和字幕进行端到端自发语音识别链接:https://arxiv.org/abs/2210.07771
作者:Jakob Poncelet,Hugo Van hamme机构:KU Leuven, Department Electrical Engineering ESAT-PSI, Kasteelpark Arenberg , Bus , B-, Leuven, Belgium备注:Accepted at SLT 2022摘要:电视字幕是许多类型演讲的转录的丰富来源,从新闻报道中的朗读到脱口秀和肥皂剧中的会话和自发演讲。然而,字幕不是语音的逐字(即,精确)转录,因此它们不能直接用于改进自动语音识别(ASR)模型。我们提出一个多任务双解码器Transformer模型,联合执行ASR和自动字幕。ASR解码器(可能是预先训练的)预测逐字输出,而字幕解码器生成字幕,同时共享编码器。这两个解码器可以是独立的或连接的。该模型被训练来联合执行两个任务,并且能够有效地使用字幕数据。我们显示了通过结合附加字幕解码器对常规ASR以及自发和会话ASR的改进。该方法不需要预处理(对齐、过滤、伪标记等)。字幕。摘要:TV subtitles are a rich source of transcriptions of many types of speech, ranging from read speech in news reports to conversational and spontaneous speech in talk shows and soaps. However, subtitles are not verbatim (i.e. exact) transcriptions of speech, so they cannot be used directly to improve an Automatic Speech Recognition (ASR) model. We propose a multitask dual-decoder Transformer model that jointly performs ASR and automatic subtitling. The ASR decoder (possibly pre-trained) predicts the verbatim output and the subtitle decoder generates a subtitle, while sharing the encoder. The two decoders can be independent or connected. The model is trained to perform both tasks jointly, and is able to effectively use subtitle data. We show improvements on regular ASR and on spontaneous and conversational ASR by incorporating the additional subtitle decoder. The method does not require preprocessing (aligning, filtering, pseudo-labeling, ...) of the subtitles.
【16】 LeVoice ASR Systems for the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge标题:面向ISCSLP 2022智能驾驶舱语音识别挑战赛的LeVoice ASR系统链接:https://arxiv.org/abs/2210.07749
作者:Yan Jia,Mi Hong,Jingyu Hou,Kailong Ren,Sifan Ma,Jin Wang,Fangzhen Peng,Yinglin Ji,Lin Yang,Junjie Wang机构:Lenovo Research, Beijing, China摘要:本文介绍了LeVoice自动语音识别系统在2022年智能座舱语音识别挑战赛中的track 2。Track 2是对模型大小的范围没有限制的语音识别任务。我们的主要观点包括基于深度学习的语音增强、基于文本到语音的语音生成、通过各种技术增强训练数据以及语音识别模型融合。我们比较并融合了混合架构和两种端到端架构。对于端到端建模,我们使用了基于连接主义时间分类/基于注意力的编码器-解码器架构和递归神经网络传感器/基于注意力的编码器-解码器架构的模型。这些模型的性能通过附加语言模型来评估,以改善单词错误率。在挑战测试集数据上,我们的系统达到了10.2%的字符错误率,在挑战赛提交的系统中排名第三。摘要:This paper describes LeVoice automatic speech recognition systems to track2 of intelligent cockpit speech recognition challenge 2022. Track2 is a speech recognition task without limits on the scope of model size. Our main points include deep learning based speech enhancement, text-to-speech based speech generation, training data augmentation via various techniques and speech recognition model fusion. We compared and fused the hybrid architecture and two kinds of end-to-end architecture. For end-to-end modeling, we used models based on connectionist temporal classification/attention-based encoder-decoder architecture and recurrent neural network transducer/attention-based encoder-decoder architecture. The performance of these models is evaluated with an additional language model to improve word error rates. As a result, our system achieved 10.2\% character error rate on the challenge test set data and ranked third place among the submitted systems in the challenge.
【17】 TransFusion: Transcribing Speech with Multinomial Diffusion标题:输血:用多项式扩散法转录语音链接:https://arxiv.org/abs/2210.07677
作者:Matthew Baas,Kevin Eloff,Herman Kamper机构:MediaLab, Department of Electronic & Electrical Engineering, Stellenbosch University, South Africa备注:12 pages, 4 figures, 1 table. Accepted at SACAIR 2022摘要:扩散模型已经在图像合成领域中显示出异常的缩放特性,并且最初的尝试已经显示出将扩散应用于无条件文本合成的类似益处。去噪扩散模型试图迭代地细化采样的噪声信号,直到它类似于相干信号(诸如图像或书面句子)。在这项工作中,我们的目标是看看扩散模型的好处是否也可以实现语音识别。为此,我们提出了一种新的方法来执行语音识别使用扩散模型的条件下预先训练的语音特征。具体而言,我们建议采用TransFusion:转录扩散模型,其迭代地将随机字符序列去噪为与条件发音的转录相对应的连贯文本。我们在LibriSpeech语音识别基准测试中证明了与现有高性能对比模型相当的性能。据我们所知,我们是第一个将去噪扩散应用于语音识别的人。我们还提出了有效采样和解码多项式扩散模型的新技术。这些是必需的,因为传统的声学模型采样方法不可能使用我们的新离散扩散方法。代码和经过培训的模型可用:https://github.com/RF5/transfusion-asr摘要:Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional text synthesis. Denoising diffusion models attempt to iteratively refine a sampled noise signal until it resembles a coherent signal (such as an image or written sentence). In this work we aim to see whether the benefits of diffusion models can also be realized for speech recognition. To this end, we propose a new way to perform speech recognition using a diffusion model conditioned on pretrained speech features. Specifically, we propose TransFusion: a transcribing diffusion model which iteratively denoises a random character sequence into coherent text corresponding to the transcript of a conditioning utterance. We demonstrate comparable performance to existing high-performing contrastive models on the LibriSpeech speech recognition benchmark. To the best of our knowledge, we are the first to apply denoising diffusion to speech recognition. We also propose new techniques for effectively sampling and decoding multinomial diffusion models. These are required because traditional methods of sampling from acoustic models are not possible with our new discrete diffusion approach. Code and trained models are available: https://github.com/RF5/transfusion-asr
【1】 Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system
标题:描述和分析DCASE关于基线系统的任务4 2022中引入的新内容
链接:https://arxiv.org/abs/2210.07856
* 与cs.SD语音【14】为同一篇
作者:Francesca Ronchini,Samuele Cornell,Romain Serizel,Nicolas Turpault,Eduardo Fonseca,Daniel P. W. Ellis机构:Universite de Lorraine, CNRS, Inria, Loria, Nancy, France, Department of Information Engineering, Italy, Google Research, United States摘要:声学场景和事件的检测和分类挑战任务4的目的是使用异构数据集评估家庭环境中声音事件检测系统。系统需要能够正确地检测记录的音频剪辑中存在的声音事件,以及及时地定位事件。今年的任务是DCASE 2021 Task 4的后续,有一些重要的新功能。本文的目标是描述和激励这些新添加的内容,并报告它们对基线系统的影响的分析。我们推出了三个主要的新功能:外部数据集的使用,包括最近发布的来自Audioset的强注释剪辑,利用预先训练的模型的可能性,以及新的能耗度量,以提高对训练声音事件检测器的生态影响的认识。基线系统的结果表明,利用AudioSet上的开源预训练可以显著改善事件分类的结果,但不能改善事件分割的结果。摘要:The aim of the Detection and Classification of Acoustic Scenes and Events Challenge Task 4 is to evaluate systems for the detection of sound events in domestic environments using an heterogeneous dataset. The systems need to be able to correctly detect the sound events present in a recorded audio clip, as well as localize the events in time. This year's task is a follow-up of DCASE 2021 Task 4, with some important novelties. The goal of this paper is to describe and motivate these new additions, and report an analysis of their impact on the baseline system. We introduced three main novelties: the use of external datasets, including recently released strongly annotated clips from Audioset, the possibility of leveraging pre-trained models, and a new energy consumption metric to raise awareness about the ecological impact of training sound events detectors. The results on the baseline system show that leveraging open-source pretrained on AudioSet improves the results significantly in terms of event classification but not in terms of event segmentation.
【2】 Learning to Jointly Transcribe and Subtitle for End-to-End Spontaneous Speech Recognition
标题:学习联合转录和字幕进行端到端自发语音识别
链接:https://arxiv.org/abs/2210.07771
* 与cs.SD语音【15】为同一篇
作者:Jakob Poncelet,Hugo Van hamme机构:KU Leuven, Department Electrical Engineering ESAT-PSI, Kasteelpark Arenberg , Bus , B-, Leuven, Belgium摘要:电视字幕是许多类型演讲的转录的丰富来源,从新闻报道中的朗读到脱口秀和肥皂剧中的会话和自发演讲。然而,字幕不是语音的逐字(即,精确)转录,因此它们不能直接用于改进自动语音识别(ASR)模型。我们提出一个多任务双解码器Transformer模型,联合执行ASR和自动字幕。ASR解码器(可能是预先训练的)预测逐字输出,而字幕解码器生成字幕,同时共享编码器。这两个解码器可以是独立的或连接的。该模型被训练来联合执行两个任务,并且能够有效地使用字幕数据。我们显示了通过结合附加字幕解码器对常规ASR以及自发和会话ASR的改进。该方法不需要预处理(对齐、过滤、伪标记等)。字幕。摘要:TV subtitles are a rich source of transcriptions of many types of speech, ranging from read speech in news reports to conversational and spontaneous speech in talk shows and soaps. However, subtitles are not verbatim (i.e. exact) transcriptions of speech, so they cannot be used directly to improve an Automatic Speech Recognition (ASR) model. We propose a multitask dual-decoder Transformer model that jointly performs ASR and automatic subtitling. The ASR decoder (possibly pre-trained) predicts the verbatim output and the subtitle decoder generates a subtitle, while sharing the encoder. The two decoders can be independent or connected. The model is trained to perform both tasks jointly, and is able to effectively use subtitle data. We show improvements on regular ASR and on spontaneous and conversational ASR by incorporating the additional subtitle decoder. The method does not require preprocessing (aligning, filtering, pseudo-labeling, ...) of the subtitles.
【3】 LeVoice ASR Systems for the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge
标题:面向ISCSLP 2022智能驾驶舱语音识别挑战赛的LeVoice ASR系统
链接:https://arxiv.org/abs/2210.07749
* 与cs.SD语音【16】为同一篇
作者:Yan Jia,Mi Hong,Jingyu Hou,Kailong Ren,Sifan Ma,Jin Wang,Fangzhen Peng,Yinglin Ji,Lin Yang,Junjie Wang机构:Lenovo Research, Beijing, China摘要:本文介绍了LeVoice自动语音识别系统在2022年智能座舱语音识别挑战赛中的track 2。Track 2是对模型大小的范围没有限制的语音识别任务。我们的主要观点包括基于深度学习的语音增强、基于文本到语音的语音生成、通过各种技术增强训练数据以及语音识别模型融合。我们比较并融合了混合架构和两种端到端架构。对于端到端建模,我们使用了基于连接主义时间分类/基于注意力的编码器-解码器架构和递归神经网络传感器/基于注意力的编码器-解码器架构的模型。这些模型的性能通过附加语言模型来评估,以改善单词错误率。在挑战测试集数据上,我们的系统达到了10.2%的字符错误率,在挑战赛提交的系统中排名第三。摘要:This paper describes LeVoice automatic speech recognition systems to track2 of intelligent cockpit speech recognition challenge 2022. Track2 is a speech recognition task without limits on the scope of model size. Our main points include deep learning based speech enhancement, text-to-speech based speech generation, training data augmentation via various techniques and speech recognition model fusion. We compared and fused the hybrid architecture and two kinds of end-to-end architecture. For end-to-end modeling, we used models based on connectionist temporal classification/attention-based encoder-decoder architecture and recurrent neural network transducer/attention-based encoder-decoder architecture. The performance of these models is evaluated with an additional language model to improve word error rates. As a result, our system achieved 10.2\% character error rate on the challenge test set data and ranked third place among the submitted systems in the challenge.
【4】 TransFusion: Transcribing Speech with Multinomial Diffusion
标题:输血:用多项式扩散法转录语音
链接:https://arxiv.org/abs/2210.07677
* 与cs.SD语音【17】为同一篇
作者:Matthew Baas,Kevin Eloff,Herman Kamper机构:MediaLab, Department of Electronic & Electrical Engineering, Stellenbosch University, South Africa备注:12 pages, 4 figures, 1 table. Accepted at SACAIR 2022摘要:扩散模型已经在图像合成领域中显示出异常的缩放特性,并且最初的尝试已经显示出将扩散应用于无条件文本合成的类似益处。去噪扩散模型试图迭代地细化采样的噪声信号,直到它类似于相干信号(诸如图像或书面句子)。在这项工作中,我们的目标是看看扩散模型的好处是否也可以实现语音识别。为此,我们提出了一种新的方法来执行语音识别使用扩散模型的条件下预先训练的语音特征。具体而言,我们建议采用TransFusion:转录扩散模型,其迭代地将随机字符序列去噪为与条件发音的转录相对应的连贯文本。我们在LibriSpeech语音识别基准测试中证明了与现有高性能对比模型相当的性能。据我们所知,我们是第一个将去噪扩散应用于语音识别的人。我们还提出了有效采样和解码多项式扩散模型的新技术。这些是必需的,因为传统的声学模型采样方法不可能使用我们的新离散扩散方法。代码和经过培训的模型可用:https://github.com/RF5/transfusion-asr摘要:Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional text synthesis. Denoising diffusion models attempt to iteratively refine a sampled noise signal until it resembles a coherent signal (such as an image or written sentence). In this work we aim to see whether the benefits of diffusion models can also be realized for speech recognition. To this end, we propose a new way to perform speech recognition using a diffusion model conditioned on pretrained speech features. Specifically, we propose TransFusion: a transcribing diffusion model which iteratively denoises a random character sequence into coherent text corresponding to the transcript of a conditioning utterance. We demonstrate comparable performance to existing high-performing contrastive models on the LibriSpeech speech recognition benchmark. To the best of our knowledge, we are the first to apply denoising diffusion to speech recognition. We also propose new techniques for effectively sampling and decoding multinomial diffusion models. These are required because traditional methods of sampling from acoustic models are not possible with our new discrete diffusion approach. Code and trained models are available: https://github.com/RF5/transfusion-asr
【5】 Improving generalizability of distilled self-supervised speech processing models under distorted settings
标题:提高失真环境下提取的自监督语音处理模型的泛化能力
链接:https://arxiv.org/abs/2210.07978
* 与cs.SD语音【1】为同一篇
作者:Kuan-Po Huang,Yu-Kuan Fu,Tsu-Yuan Hsu,Fabian Ritter Gutierrez,Fan-Lin Wang,Liang-Hsuan Tseng,Yu Zhang,Hung-yi Lee机构:National Taiwan University, ASUS Intelligent Cloud Services, Nanyang Technological University, Google Brain备注:Accepted by IEEE SLT2022摘要:自监督学习(SSL)语音预训练模型在各种语音处理任务中表现良好。已开发出SSL模型的精简版本,以满足设备上语音应用程序的需求。尽管与原始SSL模型具有相似的性能,但在失真环境中,经过提炼的对应模型的性能下降甚至比原始版本更严重。本文提出在SSL模型的知识提取过程中引入交叉失真映射和领域对抗训练,以缓解领域不匹配问题带来的性能差距。结果表明,在域内和域外失真设置下,对于不同的下游任务,在保持有效的模型大小的同时,性能得到了一致的改善。摘要:Self-supervised learned (SSL) speech pre-trained models perform well across various speech processing tasks. Distilled versions of SSL models have been developed to match the needs of on-device speech applications. Though having similar performance as original SSL models, distilled counterparts suffer from performance degradation even more than their original versions in distorted environments. This paper proposes to apply Cross-Distortion Mapping and Domain Adversarial Training to SSL models during knowledge distillation to alleviate the performance gap caused by the domain mismatch problem. Results show consistent performance improvements under both in- and out-of-domain distorted setups for different downstream tasks while keeping efficient model size.
【6】 Bringing NURC/SP to Digital Life: the Role of Open-source Automatic Speech Recognition Models
标题:将NURC/SP带入数字生活:开源自动语音识别模型的作用
链接:https://arxiv.org/abs/2210.07852
* 与cs.SD语音【2】为同一篇
作者:Lucas Rafael Stefanel Gris,Arnaldo Candido Junior,Vinícius G. dos Santos,Bruno A. Papa Dias,Marli Quadros Leite,Flaviane Romani Fernandes Svartman,Sandra Aluísio机构:Federal University of Goi´as, Brazil, S˜ao Paulo State University, Brazil, University of S˜ao Paulo, Brazil, lucas.gris(at)ufg.discente.br, arnaldo.candido(at)unesp.br, sandra(at)icmc.usp.br摘要:1969年开始的NURC项目研究了巴西五个首都的城市文化语言规范,负责为每个首都汇编一个大型语料库。数字化的NURC/SP包括在圣保罗首都拍摄的334小时记录中的375个查询。虽然47项询问有录音誊本,但录音誊本之间没有对齐,328项询问没有录音誊本。本文对三个用葡萄牙语自发语音训练的自动语音识别模型和一个用准备语音训练的模型进行了评价和误差分析。通过评估,我们可以使用WER和CER指标,在NURC/SP的手动对齐样本中选择最佳模型,以自动转录284小时。摘要:The NURC Project that started in 1969 to study the cultured linguistic urban norm spoken in five Brazilian capitals, was responsible for compiling a large corpus for each capital. The digitized NURC/SP comprises 375 inquiries in 334 hours of recordings taken in S\~ao Paulo capital. Although 47 inquiries have transcripts, there was no alignment between the audio-transcription, and 328 inquiries were not transcribed. This article presents an evaluation and error analysis of three automatic speech recognition models trained with spontaneous speech in Portuguese and one model trained with prepared speech. The evaluation allowed us to choose the best model, using WER and CER metrics, in a manually aligned sample of NURC/SP, to automatically transcribe 284 hours.
【7】 Contrastive Audio-Visual Masked Autoencoder
标题:对比式视听屏蔽式自动编码器
链接:https://arxiv.org/abs/2210.07839
* 与cs.SD语音【3】为同一篇
作者:Yuan Gong,Andrew Rouditchenko,Alexander H. Liu,David Harwath,Leonid Karlinsky,Hilde Kuehne,James Glass机构:MIT CSAIL; ,UT Austin; ,MIT-IBM Watson AI Lab; ,Goethe University Frankfurt摘要:本文首先将最新的掩蔽自动编码器(MAE)模型从单模态扩展到视听多模态。在此基础上,结合对比学习和掩蔽数据建模这两种主要的自监督学习框架,提出了对比视听掩蔽自动编码器(CAV-MAE),用于学习联合协调的视听表示。实验结果表明,对比性视听对应学习目标不仅能使模型完成视听检索任务,而且能帮助模型学习到更好的联合表征。结果表明,在VGGSound上,我们的完全自监督预训练CAV-MAE达到了65.9%的新SOTA准确率,并且在视听事件分类任务中与之前在AudioSet上的最佳监督预训练模型相当。摘要:In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by combining contrastive learning and masked data modeling, two major self-supervised learning frameworks, to learn a joint and coordinated audio-visual representation. Our experiments show that the contrastive audio-visual correspondence learning objective not only enables the model to perform audio-visual retrieval tasks, but also helps the model learn a better joint representation. As a result, our fully self-supervised pretrained CAV-MAE achieves a new SOTA accuracy of 65.9% on VGGSound, and is comparable with the previous best supervised pretrained model on AudioSet in the audio-visual event classification task.
【8】 Intel Labs at Ego4D Challenge 2022: A Better Baseline for Audio-Visual Diarization
标题:英特尔实验室参加2022年Ego4D挑战赛:视听失真的更好基准
链接:https://arxiv.org/abs/2210.07764
* 与cs.SD语音【4】为同一篇
备注:Validation report for the Ego4D challenge at ECCV 2022摘要:本报告描述了我们在2022年Ego4D挑战赛中完成视听诊断(AVD)任务的方法。具体而言,我们在官方基准基础上进行了多项技术改进。首先,通过修改模型的训练方案,提高了摄像头佩戴者语音活动的检测性能。第二,我们发现现成的语音活动检测模型在仅应用于相机佩戴者的语音活动时可以有效地去除假阳性。最后,我们证明了更好的活动说话人检测导致更好的AVD结果。我们的最终方法在Ego4D的测试集上获得了65.9%的DER,显著优于所有基线。我们的作品在2022年Ego4D挑战赛中获得第一名。摘要:This report describes our approach for the Audio-Visual Diarization (AVD) task of the Ego4D Challenge 2022. Specifically, we present multiple technical improvements over the official baselines. First, we improve the detection performance of the camera wearer's voice activity by modifying the training scheme of its model. Second, we discover that an off-the-shelf voice activity detection model can effectively remove false positives when it is applied solely to the camera wearer's voice activities. Lastly, we show that better active speaker detection leads to a better AVD outcome. Our final method obtains 65.9% DER on the test set of Ego4D, which significantly outperforms all the baselines. Our submission achieved 1st place in the Ego4D Challenge 2022.
【9】 Accelerating RNN-based Speech Enhancement on a Multi-Core MCU with Mixed FP16-INT8 Post-Training Quantization
标题:基于混合FP16-INT8后训练量化的多核MCU加速RNN语音增强
链接:https://arxiv.org/abs/2210.07692
* 与cs.SD语音【5】为同一篇
作者:Manuele Rusci,Marco Fariselli,Martin Croome,Francesco Paci,Eric Flamand机构:Universita’ di Bologna, Bologna, ITA, Greenwaves Technologies, Grenoble, FRA备注:Accepted at the ITEM Workshop 2022 (located at ECML-PKDD2022)摘要:提出了一种基于递归神经网络(RNN)的语音增强(SE)算法的优化设计方法,并将其部署在具有1+8通用RISC-V核的先进微控制器单元(MCU)上。为了实现低延迟执行,我们提出了一种优化的软件流水线交错并行计算LSTM或GRU循环块,具有矢量化的8位整数(INT 8)和16位浮点(FP 16)计算单元,以及手动管理的模型参数内存传输。为了保证相对于全精度模型的最小精度下降,我们提出了一种新的FP 16-INT 8混合精度训练后量化(PTQ)方案,该方案将递归层压缩到8位,而其余层的位精度保持为FP 16。实验在Valentini数据集上训练的多个基于LSTM和GRU的SE模型上进行,具有多达1.24M个参数。由于所提出的方法,相对于无损FP 16基线,我们将计算速度提高了4倍。与使PESQ分数平均降低0.3的均匀8位量化不同,混合精度PTQ方案导致仅0.06的低降低,同时实现1.4- 1.7倍的存储器节省。得益于这种压缩,我们通过在有限的片内非易失性存储器上安装大型模型,降低了外部存储器的功耗成本;通过将电源电压从0.8V降至0.65V,MCU的功耗最多可降低2.5倍,同时仍能满足实时性要求。与部署在单核MCU上的最先进SE解决方案相比,我们的设计能效高出10倍,这些解决方案利用了更小的模型和量化感知训练。摘要:This paper presents an optimized methodology to design and deploy Speech Enhancement (SE) algorithms based on Recurrent Neural Networks (RNNs) on a state-of-the-art MicroController Unit (MCU), with 1+8 general-purpose RISC-V cores. To achieve low-latency execution, we propose an optimized software pipeline interleaving parallel computation of LSTM or GRU recurrent blocks, featuring vectorized 8-bit integer (INT8) and 16-bit floating-point (FP16) compute units, with manually-managed memory transfers of model parameters. To ensure minimal accuracy degradation with respect to the full-precision models, we propose a novel FP16-INT8 Mixed-Precision Post-Training Quantization (PTQ) scheme that compresses the recurrent layers to 8-bit while the bit precision of remaining layers is kept to FP16. Experiments are conducted on multiple LSTM and GRU based SE models trained on the Valentini dataset, featuring up to 1.24M parameters. Thanks to the proposed approaches, we speed-up the computation by up to 4x with respect to the lossless FP16 baselines. Differently from a uniform 8-bit quantization that degrades the PESQ score by 0.3 on average, the Mixed-Precision PTQ scheme leads to a low-degradation of only 0.06, while achieving a 1.4-1.7x memory saving. Thanks to this compression, we cut the power cost of the external memory by fitting the large models on the limited on-chip non-volatile memory and we gain a MCU power saving of up to 2.5x by reducing the supply voltage from 0.8V to 0.65V while still matching the real-time constraints. Our design results 10x more energy efficient than state-of-the-art SE solutions deployed on single-core MCUs that make use of smaller models and quantization-aware training.
【10】 Full-Stack Bioacoustics: Field Kit to AI to Action (Workshop report)
标题:全套生物声学:从现场套件到人工智能到行动(研讨会报告)
链接:https://arxiv.org/abs/2210.07685
* 与cs.SD语音【6】为同一篇
作者:Dan Stowell,Caitlin Black,Florencia Noriega,Sarab S. Sethi机构:•, Sarab Sethi, University of Cambridge, This report contains an overview of the workshop aims and structure, as well as, reports from the six groups., Scientific case备注:Workshop report: Lorentz Center, Leiden, the Netherlands, 1-5 August 2022摘要:声学数据(声音记录)是探测、计数和区分野生动物的重要证据来源。由于信号处理和机器学习、记录设备以及数据处理和存储能力的巨大进步,“生物声学”领域在过去十年中得到了发展。许多研究论文描述了Raspberry Pi或类似设备用于声学监测的用途,并且其他研究论文描述了通过机器学习对动物声音的自动分类。但对于大多数生态学家、动物学家、自然资源保护主义者来说,这些拼图并没有拼在一起:域被分割。在这次洛伦兹研讨会上,我们将汇集生物声学监测和机器学习领域的开放硬件和开源软件的主要代表,以及生态学家和其他领域的研究人员,以弥合这一差距。我们在分享技能的同时,也为“生物声学AI”的未来发展建立了愿景。 本报告概述了讲习班的目标和结构,以及六个小组的报告。摘要:Acoustic data (sound recordings) are a vital source of evidence for detecting, counting, and distinguishing wildlife. This domain of "bioacoustics" has grown in the past decade due to the massive advances in signal processing and machine learning, recording devices, and the capacity of data processing and storage. Numerous research papers describe the use of Raspberry Pi or similar devices for acoustic monitoring, and other research papers describe automatic classification of animal sounds by machine learning. But for most ecologists, zoologists, conservationists, the pieces of the puzzle do not come together: the domain is fragmented. In this Lorentz workshop we bridge this gap by bringing together leading exponents of open hardware and open-source software for bioacoustic monitoring and machine learning, as well as ecologists and other field researchers. We share skills while also building a vision for the future development of "bioacoustic AI". This report contains an overview of the workshop aims and structure, as well as reports from the six groups.
【11】 Training speech emotion classifier without categorical annotations
标题:训练不带类别标注的语音情感分类器
链接:https://arxiv.org/abs/2210.07642
* 与cs.SD语音【7】为同一篇
作者:Meysam Shamsi,Marie Tahon机构:LIUM, Le Mans University, Avenue Olivier Messiaen, Le Mans, France摘要:情感表征有两种范式:范畴标注和连续空间维度描述。因此,情绪识别任务可以被视为分类或回归。本研究的主要目的是研究这两种表示法之间的关系,并提出一种只使用维度标注的分类管道。提出的方法包含一个回归器模型,该模型被训练来预测给定语音音频的维度表示中的连续值的向量。该模型的输出可以使用映射算法解释为情感类别。我们研究了三种特征提取器、三种神经网络结构和三种映射算法在两个不同语料库上的性能。我们的研究显示了通过回归方法进行分类的优点和局限性。摘要:There are two paradigms of emotion representation, categorical labeling and dimensional description in continuous space. Therefore, the emotion recognition task can be treated as a classification or regression. The main aim of this study is to investigate the relation between these two representations and propose a classification pipeline that uses only dimensional annotation. The proposed approach contains a regressor model which is trained to predict a vector of continuous values in dimensional representation for given speech audio. The output of this model can be interpreted as an emotional category using a mapping algorithm. We investigated the performances of a combination of three feature extractors, three neural network architectures, and three mapping algorithms on two different corpora. Our study shows the advantages and limitations of the classification via regression approach.
【12】 Empirical Study Incorporating Linguistic Knowledge on Filled Pauses for Personalized Spontaneous Speech Synthesis
标题:个性化自然语音合成中填充停顿与语言学知识结合的实证研究
链接:https://arxiv.org/abs/2210.07559
* 与cs.SD语音【8】为同一篇
作者:Yuta Matsunaga,Takaaki Saeki,Shinnosuke Takamichi,Hiroshi Saruwatari机构:∗ Graduate School of Information Science and Technology, The University of Tokyo, Japan.备注:Accepted to APSIPA ASC 2022摘要:我们提出了一个全面的实证研究个性化自发语音合成的基础上的语言知识。随着用于阅读式语音合成的声音克隆的出现,需要一种用于类人和自发语音合成的新的声音克隆范例。因此,我们将重点放在个人化的自发语音合成上,它可以同时复制个人的声音音色和语音不流畅。具体来说,我们研究的是填充停顿,它是言语不流畅的主要来源,在心理学和语言学中,它在言语生成和交际中起着重要作用。为了比较个性化填充停顿插入和非个性化填充停顿预测方法,提出了一种基于多说话人语料库训练的非个性化外部填充停顿预测器的语音合成方法。研究结果表明,填充停顿的位置-词纠缠,在合成语音的评价中,为了自然性而精确预测位置的必要性和为了个性而精确预测单词的必要性。摘要:We present a comprehensive empirical study for personalized spontaneous speech synthesis on the basis of linguistic knowledge. With the advent of voice cloning for reading-style speech synthesis, a new voice cloning paradigm for human-like and spontaneous speech synthesis is required. We, therefore, focus on personalized spontaneous speech synthesis that can clone both the individual's voice timbre and speech disfluency. Specifically, we deal with filled pauses, a major source of speech disfluency, which is known to play an important role in speech generation and communication in psychology and linguistics. To comparatively evaluate personalized filled pause insertion and non-personalized filled pause prediction methods, we developed a speech synthesis method with a non-personalized external filled pause predictor trained with a multi-speaker corpus. The results clarify the position-word entanglement of filled pauses, i.e., the necessity of precisely predicting positions for naturalness and the necessity of precisely predicting words for individuality on the evaluation of synthesized speech.
【13】 Transformer-Based Speech Synthesizer Attribution in an Open Set Scenario
标题:开集场景下基于变换的语音合成器属性
链接:https://arxiv.org/abs/2210.07546
* 与cs.SD语音【9】为同一篇
作者:Emily R. Bartusiak,Edward J. Delp机构:Video and Image Processing Lab, School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN摘要:语音合成方法可以创建听起来逼真的语音,其可用于欺诈、欺骗和误导活动。检测合成语音的取证方法对于防止这种攻击是重要的。取证归属方法提供关于合成语音信号的性质的甚至更多信息,语音合成器),用于创建语音信号。由于越来越多的真实声音的语音合成器,我们提出了一种语音归属的方法,推广到新的合成器没有看到在训练。为了做到这一点,我们研究了语音合成器在封闭集和开放集两种情况下的归因。换句话说,我们认为一些语音合成器是“已知的”合成器(即,而其它的合成器是“未知”合成器(即,开集的一部分)。我们将语音信号表示为谱图,并在闭集上训练我们提出的方法,称为紧凑属性Transformer(CAT),用于多类分类。然后,我们将我们的分析扩展到开集,以将合成的语音信号归属于已知和未知合成器。我们在训练过的CAT的潜在空间上利用t分布随机邻居嵌入(tSNE)来区分每个未知合成器。此外,我们还探索了poly-1丢失配方以改善归因结果。我们提出的方法在闭集和开集两种情况下都成功地将合成的语音信号归属到它们各自的语音合成器。摘要:Speech synthesis methods can create realistic-sounding speech, which may be used for fraud, spoofing, and misinformation campaigns. Forensic methods that detect synthesized speech are important for protection against such attacks. Forensic attribution methods provide even more information about the nature of synthesized speech signals because they identify the specific speech synthesis method (i.e., speech synthesizer) used to create a speech signal. Due to the increasing number of realistic-sounding speech synthesizers, we propose a speech attribution method that generalizes to new synthesizers not seen during training. To do so, we investigate speech synthesizer attribution in both a closed set scenario and an open set scenario. In other words, we consider some speech synthesizers to be "known" synthesizers (i.e., part of the closed set) and others to be "unknown" synthesizers (i.e., part of the open set). We represent speech signals as spectrograms and train our proposed method, known as compact attribution transformer (CAT), on the closed set for multi-class classification. Then, we extend our analysis to the open set to attribute synthesized speech signals to both known and unknown synthesizers. We utilize a t-distributed stochastic neighbor embedding (tSNE) on the latent space of the trained CAT to differentiate between each unknown synthesizer. Additionally, we explore poly-1 loss formulations to improve attribution results. Our proposed approach successfully attributes synthesized speech signals to their respective speech synthesizers in both closed and open set scenarios.
【14】 Hierarchical Diffusion Models for Singing Voice Neural Vocoder
标题:歌唱语音神经声码器的分层扩散模型
链接:https://arxiv.org/abs/2210.07508
* 与cs.SD语音【10】为同一篇
作者:Naoya Takahashi,Mayank Kumar,Singh,Yuki Mitsufuji机构:Sony Group Corporation, Japan摘要:深度生成模型的最新进展提高了语音域神经声码器的质量。然而,由于音高、响度和发音方面的音乐表现形式的多样性,产生高质量的歌唱声音仍然具有挑战性。在这项工作中,我们提出了一个层次扩散模型的歌唱声神经声码器。所提出的方法由以不同采样率操作的多个扩散模型组成;最低采样率下的模型集中于生成准确的低频分量,例如音调,而其它模型基于较低采样率下的数据和声学特征逐渐生成较高采样率下的波形。实验结果表明,该方法能够为多个演唱者生成高质量的演唱声,在计算量相近的情况下,其性能优于现有的神经声码器。摘要:Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, it remains challenging to generate high-quality singing voice due to a wider variety of musical expressions in pitch, loudness, and pronunciations. In this work, we propose a hierarchical diffusion model for singing voice neural vocoders. The proposed method consists of multiple diffusion models operating in different sampling rates; the model at the lowest sampling rate focuses on generating accurate low frequency components such as pitch, and other models progressively generate the waveform at the higher sampling rates based on the data at the lower sampling rate and acoustic features. Experimental results show that the proposed method produces high-quality singing voice for multiple singers, outperforming state-of-the-art neural vocoders with a similar range of computational costs.
【15】 Bayes risk CTC: Controllable CTC alignment in Sequence-to-Sequence tasks
标题:贝叶斯风险CTC:序列到序列任务中的可控CTC比对
链接:https://arxiv.org/abs/2210.07499
* 与cs.SD语音【11】为同一篇
作者:Jinchuan Tian,Brian Yan,Jianwei Yu,Chao Weng,Dong Yu,Shinji Watanabe机构:Tencent AI LAB, Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA , USA摘要:序列到序列(seq 2seq)任务将输入序列转录为靶序列。连接主义时态分类(CTC)准则被广泛应用于多个seq 2seq任务。除了预测目标序列之外,CTC的副产品是预测比对,比对是最可能的输入长序列,其指定输入单元和目标单元之间的硬比对关系。由于在CTC公式中同等考虑了多个潜在的比对序列(称为路径),因此选择哪条路径最有可能成为预测的比对总是不确定的。此外,通常观察到,由vanilla CTC预测的比对与其参考相比将漂移,并且很少提供实际功能。因此,这项工作的动机是使CTC对准预测可控,从而为CTC配备额外的功能。在此基础上,提出了贝叶斯风险CTC(BRCTC)准则,该准则采用一个可定制的贝叶斯风险函数来增强预测序列的期望特性. BRCTC是一个通用的框架,通过风险函数对路径采用一些可定制的偏好,以便将后验概率集中到路径的特定子集中。在应用中,我们探索了一种特殊的偏好,它产生的模型具有下采样能力和降低的推理成本。通过使用BRCTC和另一个早期发射偏好,我们获得了在线模型的改进的性能-延迟权衡。实验结果表明,BRCTC算法在不降低系统性能的前提下,将离线模型的推理代价降低了47%,并将在线系统的整体延迟降低到了一个不可见的水平.摘要:Sequence-to-Sequence (seq2seq) tasks transcribe the input sequence to a target sequence. The Connectionist Temporal Classification (CTC) criterion is widely used in multiple seq2seq tasks. Besides predicting the target sequence, a side product of CTC is to predict the alignment, which is the most probable input-long sequence that specifies a hard aligning relationship between the input and target units. As there are multiple potential aligning sequences (called paths) that are equally considered in CTC formulation, the choice of which path will be most probable and become the predicted alignment is always uncertain. In addition, it is usually observed that the alignment predicted by vanilla CTC will drift compared with its reference and rarely provides practical functionalities. Thus, the motivation of this work is to make the CTC alignment prediction controllable and thus equip CTC with extra functionalities. The Bayes risk CTC (BRCTC) criterion is then proposed in this work, in which a customizable Bayes risk function is adopted to enforce the desired characteristics of the predicted alignment. With the risk function, the BRCTC is a general framework to adopt some customizable preference over the paths in order to concentrate the posterior into a particular subset of the paths. In applications, we explore one particular preference which yields models with the down-sampling ability and reduced inference costs. By using BRCTC with another preference for early emissions, we obtain an improved performance-latency trade-off for online models. Experimentally, the proposed BRCTC reduces the inference cost of offline models by up to 47% without performance degradation and cuts down the overall latency of online systems to an unseen level.
【16】 JOIST: A Joint Speech and Text Streaming Model For ASR
标题:Joist:一种面向ASR的语音和文本联合流媒体模型
链接:https://arxiv.org/abs/2210.07353
* 与cs.SD语音【12】为同一篇
作者:Tara N. Sainath,Rohit Prabhavalkar,Ankur Bapna,Yu Zhang,Zhouyuan Huo,Zhehuai Chen,Bo Li,Weiran Wang,Trevor Strohman摘要:本文提出了一种新的编码算法JOIST,该算法可以训练一个具有语音-文本成对输入和仅文本非成对输入的流、级联、端到端(E2 E)编码器模型。与以往的研究不同,我们探索了两种模式的联合训练,而不是预先训练和微调。此外,我们使用一个流E2 E模型来探索JOIST,数据量增加了一个数量级,这与以前的工作相比也是新颖的。通过一系列的烧蚀研究,我们探索了不同类型的文本建模,包括如何对文本序列的长度进行建模以及合适的文本子词单元表示。我们发现,与未使用文本训练的模型相比,JOIST的最佳文本表示在各种搜索和稀有词测试集上将WER提高了4-14%。此外,我们定量地表明,JOIST保持了流功能,这对良好的用户级体验很重要。摘要:We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous works, we explore joint training with both modalities, rather than pre-training and fine-tuning. In addition, we explore JOIST using a streaming E2E model with an order of magnitude more data, which are also novelties compared to previous works. Through a series of ablation studies, we explore different types of text modeling, including how to model the length of the text sequence and the appropriate text sub-word unit representation. We find that best text representation for JOIST improves WER across a variety of search and rare-word test sets by 4-14% relative, compared to a model not trained with text. In addition, we quantitatively show that JOIST maintains streaming capabilities, which is important for good user-level experience.
【17】 HuBERT-TR: Reviving Turkish Automatic Speech Recognition with Self-supervised Speech Representation Learning
标题:Hubert-tr:用自监督语音表征学习恢复土耳其语自动语音识别
链接:https://arxiv.org/abs/2210.07323
* 与cs.SD语音【13】为同一篇
作者:Ali Safaya,Engin Erzin机构:KUIS AI Center, Computer Engineering Department, Koc¸ University摘要:虽然土耳其语被列为低资源语言之一,但有关土耳其语自动语音识别(ASR)的文献相对较早。本文提出了一种基于HuBERT的土耳其语语音表示模型HuBERT-TR。HuBERT-TR在几个土耳其ASR数据集上获得了最先进的结果。我们使用从在线资源中收集的大规模数据研究了土耳其语的预训练HuBERT。我们使用从YouTube上收集的超过6,500小时的语音数据对HuBERT-TR进行预训练,这些数据在质量和类型方面具有广泛的可变性。我们表明,多语言设置中的预训练模型劣于特定语言模型,其中我们的土耳其语模型HuBERT-TR base的性能优于其10倍大的多语言对应物XLS-R-1B。此外,我们通过将我们的模型缩放到1B个参数来研究缩放对ASR性能的影响。我们的最佳模型在Turkish Broadcast News数据集上产生了4.97%的最先进的单词错误率。有关型号的信息,请访问huggingface.co/asafaya。摘要:While the Turkish language is listed among low-resource languages, literature on Turkish automatic speech recognition (ASR) is relatively old. In this paper, we present HuBERT-TR, a speech representation model for Turkish based on HuBERT. HuBERT-TR achieves state-of-the-art results on several Turkish ASR datasets. We investigate pre-training HuBERT for Turkish with large-scale data curated from online resources. We pre-train HuBERT-TR using over 6,500 hours of speech data curated from YouTube that includes extensive variability in terms of quality and genre. We show that pre-trained models within a multi-lingual setup are inferior to language-specific models, where our Turkish model HuBERT-TR base performs better than its x10 times larger multi-lingual counterpart XLS-R-1B. Moreover, we study the effect of scaling on ASR performance by scaling our models up to 1B parameters. Our best model yields a state-of-the-art word error rate of 4.97% on the Turkish Broadcast News dataset. Models are available at huggingface.co/asafaya .
机器翻译,仅供参考