今日论文合集:cs.SD语音13篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily


cs.SD语音

【1】 Sub-8-bit quantization for on-device speech recognition: a  regularization-free approach

标题:用于设备上语音识别的亚8位量化:一种无需正则化的方法

链接:https://arxiv.org/abs/2210.09188

作者:Kai Zhen,Martin Radfar,Hieu Duy Nguyen,Grant P. Strimel,Nathan Susanj,Athanasios Mouchtaris机构:Amazon Alexa AI
备注:Accepted for publication at IEEE SLT'22
摘要:对于设备上自动语音识别(ASR),量化感知训练(QAT)是普遍存在的,以实现模型预测性能和效率之间的折衷。在现有的QAT方法中,一个主要的缺点是量化质心必须预先确定和固定。为了克服这一限制,我们引入了一种无正则化的“软到硬”压缩机制,该机制在μ定律约束空间中具有自调整质心,从而产生一种更简单但更通用的量化方案,称为通用量化器(GQ)。我们使用递归神经网络转换器(RNN-T)和Conformer架构,在LibriSpeech和去识别远场数据集上将GQ应用于ASR任务。在不降低精度的情况下,GQ可以将RNN-T和Conformer压缩到8位以下,对于一些RNN-T层,可以压缩到1位,以实现快速和准确的推断。通过物理设备基准测试,我们发现与8位QAT相比,内存占用空间节省了30.73%,用户感知延迟减少了31.75%。
摘要:For on-device automatic speech recognition (ASR), quantization aware training (QAT) is ubiquitous to achieve the trade-off between model predictive performance and efficiency. Among existing QAT methods, one major drawback is that the quantization centroids have to be predetermined and fixed. To overcome this limitation, we introduce a regularization-free, "soft-to-hard" compression mechanism with self-adjustable centroids in a mu-Law constrained space, resulting in a simpler yet more versatile quantization scheme, called General Quantizer (GQ). We apply GQ to ASR tasks using Recurrent Neural Network Transducer (RNN-T) and Conformer architectures on both LibriSpeech and de-identified far-field datasets. Without accuracy degradation, GQ can compress both RNN-T and Conformer into sub-8-bit, and for some RNN-T layers, to 1-bit for fast and accurate inference. We observe a 30.73% memory footprint saving and 31.75% user-perceived latency reduction compared to 8-bit QAT via physical device benchmarking.


【2】 Visual onoma-to-wave: environmental sound synthesis from visual  onomatopoeias and sound-source images

标题:视觉拟声学:从视觉拟声词和声源图像合成环境声音

链接:https://arxiv.org/abs/2210.09173

作者:Hien Ohnaka,Shinnosuke Takamichi,Keisuke Imoto,Yuki Okamoto,Kazuki Fujii,Hiroshi Saruwatari
机构:National Institute of Technology, Tokuyama College, Japan. ,The University of Tokyo, Japan., Doshisha University, Japan. ,Ritsumeikan University, Japan.
备注:Submitted to ICASSP 2023
摘要:我们提出一种从视觉表示的拟声词和声源合成环境声音的方法。拟声词是模仿声音结构的词,即,声音的文本表示。从这一角度出发,拟声-波技术被提出来从所需的拟声文本中合成环境声音。拟声诗还有另一种表现形式:漫画、广告和虚拟现实中声音的视觉-文本表示。视觉拟声词(视觉文本拟声词)包含文本中不存在的丰富信息,如长-短持续时间的图像,因此使用这种表示法有望合成出多样化的声音。因此,我们提出视觉拟声转波的方法,以从视觉拟声合成环境声音。该方法可以将视觉文本和声源图像的视觉概念转换为合成声音。我们还提出了一种数据扩充方法,重点关注拟声词的重复,以提高我们的方法的性能。实验结果表明,该方法能够从视觉文本和声源图像中合成出多种环境声音。
摘要:We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental sounds from the desired onomatopoeia texts. Onomatopoeias have another representation: visual-text representations of sounds in comics, advertisements, and virtual reality. A visual onomatopoeia (visual text of onomatopoeia) contains rich information that is not present in the text, such as a long-short duration of the image, so the use of this representation is expected to synthesize diverse sounds. Therefore, we propose visual onoma-to-wave for environmental sound synthesis from visual onomatopoeia. The method can transfer visual concepts of the visual text and sound-source image to the synthesized sound. We also propose a data augmentation method focusing on the repetition of onomatopoeias to enhance the performance of our method. An experimental evaluation shows that the methods can synthesize diverse environmental sounds from visual text and sound-source images.


【3】 Language-agnostic Code-Switching in End-To-End Speech Recognition

标题:端到端语音识别中的语言不可知码转换

链接:https://arxiv.org/abs/2210.08992

作者:Enes Yavuz Ugan,Christian Huber,Juan Hussain,Alexander Waibel
机构:Interactive Systems Lab, Karlsruhe Institute of Technology, Karlsruhe, Germany, Carnegie Mellon University, Pittsburgh PA, USA
备注:6 pages
摘要:语码转换是指不同语言中的词和短语交替使用的现象。虽然当今的神经端到端(E2E)模型在自动语音识别(ASR)任务上提供了最先进的性能,但众所周知,这些系统是非常数据密集型的。然而,只有少数转录和对齐的CS语音可用。为了克服这一问题并训练能够转录CS语音的多语言系统,我们提出了一种简单而有效的数据扩充,其中音频和不同源语言的对应标签被级联。通过使用该训练数据,我们的E2E模型改进了转录CS语音,并且还改进了多语言模型的性能。实验结果表明,这种增强技术可以使模型对训练过程中未发现的句间语言转换的预测性能提高5.03WER.
摘要:Code-Switching (CS) is referred to the phenomenon of alternately using words and phrases from different languages. While today's neural end-to-end (E2E) models deliver state-of-the-art performances on the task of automatic speech recognition (ASR) it is commonly known that these systems are very data-intensive. However, there is only a few transcribed and aligned CS speech available. To overcome this problem and train multilingual systems which can transcribe CS speech, we propose a simple yet effective data augmentation in which audio and corresponding labels of different source languages are concatenated. By using this training data, our E2E model improves on transcribing CS speech and improves performance over the multilingual model, as well. The results show that this augmentation technique can even improve the model's performance on inter-sentential language switches not seen during training by 5,03\% WER.


【4】 How to Leverage DNN-based speech enhancement for multi-channel speaker  verification?

标题:如何利用基于DNN的语音增强进行多声道说话人验证?

链接:https://arxiv.org/abs/2210.08834

作者:Sandipana Dowerah,Romain Serizel,Denis Jouvet,Mohammad Mohammadamini,Driss Matrouf
机构:Université de Lorraine, CNRS, Inria-Loria., Nancy, France,  Avignon Université, Laboratoire Informatique d'Avignon, Avignon, France
备注:None
摘要:由于环境噪声和室内混响的影响,说话人确认(SV)在远场场景中的性能并不理想。本文提出了一个用于远场说话人确认的多通道语音增强基准。一种是基于深度神经网络的方法,另一种是深度神经网络与信号处理相结合的方法。我们将DNN架构与信号处理技术相结合来进行各种实验。将我们的方法与现有的最新方法进行了比较。我们检验了入组在预处理中的重要性,这在以前的研究中很大程度上被忽视了。实验结果表明,只要对注册文件进行类似于测试数据的预处理,并且测试和注册发生在相似的信噪比范围内,预处理可以提高SV的性能。在VOiCES数据集的生成和所有噪声条件上获得了相当大的改善。
摘要:Speaker verification (SV) suffers from unsatisfactory performance in far-field scenarios due to environmental noise andthe adverse impact of room reverberation. This work presents a benchmark of multichannel speech enhancement for far-fieldspeaker verification. One approach is a deep neural network-based, and the other is a combination of deep neural network andsignal processing. We integrated a DNN architecture with signal processing techniques to carry out various experiments. Ourapproach is compared to the existing state-of-the-art approaches. We examine the importance of enrollment in pre-processing,which has been largely overlooked in previous studies. Experimental evaluation shows that pre-processing can improve the SVperformance as long as the enrollment files are processed similarly to the test data and that test and enrollment occur within similarSNR ranges. Considerable improvement is obtained on the generated and all the noise conditions of the VOiCES dataset.


【5】 SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of  Self-Supervised Speech Representation Learning

标题:Superb@SLT 2022:自我监督语音表征学习的泛化和效率挑战

链接:https://arxiv.org/abs/2210.08634

作者:Tzu-hsun Feng,Annie Dong,Ching-Feng Yeh,Shu-wen Yang,Tzu-Quan Lin,Jiatong Shi,Kai-Wei Chang,Zili Huang,Haibin Wu,Xuankai Chang,Shinji Watanabe,Abdelrahman Mohamed,Shang-Wen Li,Hung-yi Lee
机构:National Taiwan University, Taiwan,Meta, USA, Carnegie Mellon University, USA,Johns Hopkins University, USA
备注:Accepted by 2022 SLT Workshop
摘要:我们在SLT 2022上提出了SUPERB挑战,其目的是学习自监督语音表示以获得更好的性能、泛化和效率。该挑战建立在SUPERB基准之上,并实施了度量标准来衡量自监督学习(SSL)表示的计算需求,并评估其在不同SUPERB任务中的推广性和性能。SUPERB性能指标评测全面涵盖了流行的语音处理任务,从语音和说话人识别到音频生成和语义理解。随着SSL技术在语音领域的广泛应用,并取得了令人瞩目的成果,我们设想通过激发任务性能之外的更实用的技术设计来提升SSL技术的影响力。本文总结了14个已提交模型的结果。最后,我们讨论了这些论文的主要发现和SSL研究的未来方向。
摘要:We present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency. The challenge builds upon the SUPERB benchmark and implements metrics to measure the computation requirements of self-supervised learning (SSL) representation and to evaluate its generalizability and performance across the diverse SUPERB tasks. The SUPERB benchmark provides comprehensive coverage of popular speech processing tasks, from speech and speaker recognition to audio generation and semantic understanding. As SSL has gained interest in the speech community and showed promising outcomes, we envision the challenge to uplevel the impact of SSL techniques by motivating more practical designs of techniques beyond task performance. We summarize the results of 14 submitted models in this paper. We also discuss the main findings from those submissions and the future directions of SSL research.


【6】 Robust, General, and Low Complexity Acoustic Scene Classification  Systems and An Effective Visualization for Presenting a Sound Scene Context

标题:稳健、通用、低复杂度的声音场景分类系统和用于呈现声音场景环境的有效可视化

链接:https://arxiv.org/abs/2210.08610

作者:Lam Pham,Dusan Salovic,Anahid Jalali,Alexander Schindler,Khoa Tran,Canh Vu,Phu X. Nguyen机构:Nguyen is with University ofFPT,  Tran is with Da Nang University
摘要:在本文中,我们提出了一个全面的分析声学场景分类(ASC),任务是识别一个音频记录的场景从其声学签名。特别地,我们首先提出了一个基于开始的低足迹ASC模型,称为ASC基线。然后将提出的ASC基线与MobileNetV 1、MobileNetV 2、VGG 16、VGG 19、ResNet 50 V2、ResNet 152 V2、DenseNet 121、DenseNet 201和Xception的基准和高复杂度网络架构进行了比较。其次,我们提出了一种新的深度神经网络结构,利用残差起始结构和多个核来改进ASC基线。给出了一种新的剩余起始(NRI)模型,进一步评估了模型复杂度和模型精度性能之间的权衡.最后,我们评估了在声音场景记录中发生的声音事件是否有助于提高ASC的准确性,并指出如何通过结合声音场景和声音事件信息来很好地呈现声音场景上下文。我们对各种ASC数据集进行了广泛的实验,包括拥挤场景、IEEE AASP声学场景和事件检测和分类挑战(DCASE)2018年任务1A和1B、2019年任务1A和1B、2020年任务1A、2021年任务1A、2022年任务1。几种不同ASC挑战的实验结果突出了两个主要成就:第一个是提出鲁棒的、通用的和低复杂度的ASC系统,其适合于在广泛的边缘设备和移动设备上的实际应用;二是提出了一种有效的可视化方法,用于全面地呈现声音场景上下文。
摘要:In this paper, we present a comprehensive analysis of Acoustic Scene Classification (ASC), the task of identifying the scene of an audio recording from its acoustic signature. In particular, we firstly propose an inception-based and low footprint ASC model, referred to as the ASC baseline. The proposed ASC baseline is then compared with benchmark and high-complexity network architectures of MobileNetV1, MobileNetV2, VGG16, VGG19, ResNet50V2, ResNet152V2, DenseNet121, DenseNet201, and Xception. Next, we improve the ASC baseline by proposing a novel deep neural network architecture which leverages residual-inception architectures and multiple kernels. Given the novel residual-inception (NRI) model, we further evaluate the trade off between the model complexity and the model accuracy performance. Finally, we evaluate whether sound events occurring in a sound scene recording can help to improve ASC accuracy, then indicate how a sound scene context is well presented by combining both sound scene and sound event information. We conduct extensive experiments on various ASC datasets, including Crowded Scenes, IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 Task 1A and 1B, 2019 Task 1A and 1B, 2020 Task 1A, 2021 Task 1A, 2022 Task 1. The experimental results on several different ASC challenges highlight two main achievements; the first is to propose robust, general, and low complexity ASC systems which are suitable for real-life applications on a wide range of edge devices and mobiles; the second is to propose an effective visualization method for comprehensively presenting a sound scene context.


【7】 A Policy-based Approach to the SpecAugment Method for Low Resource E2E  ASR

标题:一种基于策略的低资源E2E ASR规范调整方法

链接:https://arxiv.org/abs/2210.08520

作者:Rui Li,Guodong Ma,Dexin Zhao,Ranran Zeng,Xiaoyu Li,Hao Huang
机构:∗ School of Information Science and Engineering, Xinjiang University, Urumqi, China, † China Telecom Beijing Research Institute, Beijing, China, ‡ Xinjiang Provincial Key Laboratory of Multi-lingual Information Technology, Urumqi, China, § Equal contributions
备注:Accepted to APSIPA ASC 2022
摘要:SpecAugment对于HMM和基于E2 E的自动语音识别(ASR)系统都是一种非常有效的数据增强方法。特别是,它还适用于低资源方案。然而,SpecAugment在固定的增强策略中屏蔽了时域或频域的频谱,这可能会给低资源ASR带来相对较小的数据分集。本文提出了一种基于策略的SpecAugment(Policy-SpecAugment)方法来缓解上述问题。其思想是采用增广选择策略和增广参数改变策略来求解固定路径。这些策略是基于验证集的丢失来学习的,验证集的丢失被应用于对应的增强策略。它的目的是鼓励模型学习更多的多样性数据,这是模型相对需要的。在实验中,我们评估了我们的方法在低资源场景下的有效性,即,100小时自由演讲任务。根据结果和分析,我们可以看到,使用我们的建议可以明显缓解上述问题。实验结果表明,与现有的规范增强算法相比,Policy-SpecAugment算法在测试/开发-清洁集上的相对WER降低超过10%,在测试/开发-其他集上的相对WER降低超过5%,在所有测试集上的绝对WER降低超过1%。
摘要:SpecAugment is a very effective data augmentation method for both HMM and E2E-based automatic speech recognition (ASR) systems. Especially, it also works in low-resource scenarios. However, SpecAugment masks the spectrum of time or the frequency domain in a fixed augmentation policy, which may bring relatively less data diversity to the low-resource ASR. In this paper, we propose a policy-based SpecAugment (Policy-SpecAugment) method to alleviate the above problem. The idea is to use the augmentation-select policy and the augmentation-parameter changing policy to solve the fixed way. These policies are learned based on the loss of validation set, which is applied to the corresponding augmentation policies. It aims to encourage the model to learn more diverse data, which the model relatively requires. In experiments, we evaluate the effectiveness of our approach in low-resource scenarios, i.e., the 100 hours librispeech task. According to the results and analysis, we can see that the above issue can be obviously alleviated using our proposal. In addition, the experimental results show that, compared with the state-of-the-art SpecAugment, the proposed Policy-SpecAugment has a relative WER reduction of more than 10% on the Test/Dev-clean set, more than 5% on the Test/Dev-other set, and an absolute WER reduction of more than 1% on all test sets.


【8】 Learning Invariant Representation and Risk Minimized for Unsupervised  Accent Domain Adaptation

标题:无监督重音域自适应的学习不变表征和风险最小化

链接:https://arxiv.org/abs/2210.08182

作者:Chendong Zhao,Jianzong Wang,Xiaoyang Qu,Haoqian Wang,Jing Xiao
机构:† Ping An Technology (Shenzhen) Co., Ltd., China, ⋆ The Shenzhen International Graduate School, Tsinghua University, China
备注:Accepted to SLT 2022
摘要:针对语音音频的无监督表示学习在语音识别任务中获得了令人印象深刻的性能,特别是当注释语音受限时。然而,无监督范式需要仔细设计,并且对于这些表示获得什么性质知之甚少。不能保证模型学习有价值信息的有意义表示以用于识别。此外,学习的表示对其他域的适应能力仍然需要估计。在这项工作中,我们探索学习领域不变的表示,通过直接映射的语音表示,以其相应的高级语言信息。实验结果表明,学习后的潜台词不仅能捕捉到每个音素的发音特征,而且能提高适应能力,在重音基准测试中的表现大大优于基线。
摘要:Unsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the unsupervised paradigm needs to be carefully designed and little is known about what properties these representations acquire. There is no guarantee that the model learns meaningful representations for valuable information for recognition. Moreover, the adaptation ability of the learned representations to other domains still needs to be estimated. In this work, we explore learning domain-invariant representations via a direct mapping of speech representations to their corresponding high-level linguistic informations. Results prove that the learned latents not only capture the articulatory feature of each phoneme but also enhance the adaptation ability, outperforming the baseline largely on accented benchmarks.


【9】 spatial-dccrn: dccrn equipped with frame-level angle feature and hybrid  filtering for multi-channel speech enhancement

标题:基于空域角度特征和混合滤波的多声道语音增强

链接:https://arxiv.org/abs/2210.08802

作者:Shubo Lv,Yihui Fu,Yukai Jv,Lei Xie,Weixin Zhu,Wei Rao,Yannan Wang
机构:Northwestern Polytechnical University, Xi’an, China, Tencent Ethereal Audio Lab, Tencent Corporation, Shenzhen, China
摘要:近年来,多通道语音增强由于利用空间信息来区分目标语音和干扰信号而引起了人们的广泛关注。为了充分利用空间信息和基于神经网络的掩蔽估计,提出了一种多通道去噪神经网络-- Spatial DCCRN。首先,将S-DCCRN算法扩展到多信道场景,实现子信道级联和全信道处理策略,分别对不同的信道进行建模。此外,该模型不仅采用多通道频谱或级联第一通道幅度和IPD作为模型的输入,还采用了角度特征提取模块(AFE)提取帧级的角度特征嵌入,使模型能够更直观地感知空间信息。最后,针对噪声和语音同时存在于同一个时频(TF)面元时,残留噪声现象会更加严重的问题,设计了一种掩蔽映射滤波方法来代替传统的滤波求和运算,实现粗去噪、去混响和残留噪声抑制的级联。在L3 DAS 22 Challenge数据集上,提议的模型Spatial-DCCRN已经超过了EaBNet、FasNet以及几个竞争模型。不仅在3D场景中,在多通道ConferencingSpeech 2021 Challenge数据集上,Spatial-DCCRN在多个评估指标上均大幅优于最先进的(SOTA)模型MIMO-UNet。消融研究也证明了不同贡献的有效性。
摘要:Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking estimation, we propose a multi-channel denoising neural network -- Spatial DCCRN. Firstly, we extend S-DCCRN to multi-channel scenario, aiming at performing cascaded sub-channel and full-channel processing strategy, which can model different channels separately. Moreover, instead of only adopting multi-channel spectrum or concatenating first-channel's magnitude and IPD as the model's inputs, we apply an angle feature extraction module (AFE) to extract frame-level angle feature embeddings, which can help the model to apparently perceive spatial information. Finally, since the phenomenon of residual noise will be more serious when the noise and speech exist in the same time frequency (TF) bin, we particularly design a masking and mapping filtering method to substitute the traditional filter-and-sum operation, with the purpose of cascading coarsely denoising, dereverberation and residual noise suppression. The proposed model, Spatial-DCCRN, has surpassed EaBNet, FasNet as well as several competitive models on the L3DAS22 Challenge dataset. Not only the 3D scenario, Spatial-DCCRN outperforms state-of-the-art (SOTA) model MIMO-UNet by a large margin in multiple evaluation metrics on the multi-channel ConferencingSpeech2021 Challenge dataset. Ablation studies also demonstrate the effectiveness of different contributions.


【10】 Acoustic-aware Non-autoregressive Spell Correction with Mask Sample  Decoding

标题:基于模板样本译码的声学非自回归拼写纠正

链接:https://arxiv.org/abs/2210.08665

作者:Ruchao Fan,Guoli Ye,Yashesh Gaur,Jinyu Li
机构:University of California, Los Angeles,Microsoft Corporation
摘要:屏蔽语言模型(MLM)已被广泛应用于任务理解,如BERT。最近,MLM也被用于生成任务。在语音识别中最常用的一种方法是使用Mask-CTC进行非自回归语音识别。本文进一步探讨了使用MLM作为变换器-转换器(transformer-transducer,TT)的非自回归拼写校正(non-autoregressive spell correction,SC)模型(MLM-SC)的可能性,初步实验表明MLM-SC对Librispeech数据没有改善。问题可能是模型单元(词块)的选择和英语数据的TT置信度得分的不准确性。为了解决这个问题,我们提出了一种掩码样本解码(MS-decode)方法,其中掩码标记可以选择被掩码或不被掩码以补偿不准确性。结果,我们分别将Librispeech测试-其他数据上的流式TT的WER从7.6%降低到6.5%,并将Aishell测试数据上的CER从7.3%降低到6.1%。
摘要:Masked language model (MLM) has been widely used for understanding tasks, e.g. BERT. Recently, MLM has also been used for generation tasks. The most popular one in speech is using Mask-CTC for non-autoregressive speech recognition. In this paper, we take one step further, and explore the possibility of using MLM as a non-autoregressive spell correction (SC) model for transformer-transducer (TT), denoted as MLM-SC. Our initial experiments show that MLM-SC provides no improvements on Librispeech data. The problem might be the choice of modeling units (word pieces) and the inaccuracy of the TT confidence scores for English data. To solve the problem, we propose a mask sample decoding (MS-decode) method where the masked tokens can have the choice of being masked or not to compensate for the inaccuracy. As a result, we reduce the WER of a streaming TT from 7.6% to 6.5% on the Librispeech test-other data and the CER from 7.3% to 6.1% on the Aishell test data, respectively.


【11】 Attention-Based Audio Embeddings for Query-by-Example

标题:基于注意力的逐例查询音频嵌入

链接:https://arxiv.org/abs/2210.08624

作者:Anup Singh,Kris Demuynck,Vipul Arora
机构:IDLab, Department of Electronics and Information Systems, imec - Ghent University, Belgium,  Department of Electrical Engineering, Indian Institute of Technology Kanpur, India
摘要:一个理想的音频检索系统有效地和鲁棒地从大量的数据库中识别短的查询片段。然而,公知的音频指纹识别系统的性能在高信号失真水平上不足。提出了一种基于对比学习框架的音频检索系统,该系统能够生成噪声和混响鲁棒的音频指纹。使用这些指纹,该方法执行全面搜索以识别查询音频并精确地估计其在参考音频中的时间戳。我们的框架涉及训练CNN,以最大化从干净音频提取的嵌入对与其相应的失真和时移版本之间的相似性。我们采用通道式频谱-时间注意机制,通过赋予信号中显著的频谱-时间碎片更多的权重来更好地区分音频。实验结果表明,与现有的同类系统相比,该系统在计算量和内存利用率上都有很大的提高,尤其是在较高的失真水平下,其计算精度更高,并且可以扩展到更大的数据库。
摘要:An ideal audio retrieval system efficiently and robustly recognizes a short query snippet from an extensive database. However, the performance of well-known audio fingerprinting systems falls short at high signal distortion levels. This paper presents an audio retrieval system that generates noise and reverberation robust audio fingerprints using the contrastive learning framework. Using these fingerprints, the method performs a comprehensive search to identify the query audio and precisely estimate its timestamp in the reference audio. Our framework involves training a CNN to maximize the similarity between pairs of embeddings extracted from clean audio and its corresponding distorted and time-shifted version. We employ a channel-wise spectral-temporal attention mechanism to better discriminate the audio by giving more weight to the salient spectral-temporal patches in the signal. Experimental results indicate that our system is efficient in computation and memory usage while being more accurate, particularly at higher distortion levels, than competing state-of-the-art systems and scalable to a larger database.


【12】 CTCBERT: Advancing Hidden-unit BERT with CTC Objectives

标题:CTCBERT:以CTC目标推进隐藏单元BERT

链接:https://arxiv.org/abs/2210.08603

作者:Ruchao Fan,Yiming Wang,Yashesh Gaur,Jinyu Li
机构:University of California, Los Angeles, Microsoft Corporation
摘要:本文提出了一种简单有效的隐单元BERT(HuBERT)改进算法CTCBERT. HuBERT应用帧级交叉熵(CE)损失,这类似于大多数声学模型训练。然而,CTCBERT在去除每个掩蔽区域中的重复ID之后,使用连接主义时间分类(CTC)目标来执行模型训练。这个想法源于这样的观察,即当使用聚集的或对齐的ID时,在对齐中可能存在显著的错误。CTC隐式地学习对齐,这表明当存在对齐时,使用CTC的学习可以更加灵活。我们检查了来自HuBERT Iter 1、HuBERT Iter 2和PBERT的ID上的CTCBERT。与CE培训相比,CTC培训带来了一致的改进。此外,当在微调期间加载空白相关参数时,观察到轻微的改进。在Librispeech 960- 100 h环境下,CTCBERT相对于HuBERT和PERT在测试-其他数据上的WER改善为2%-11%。
摘要:In this work, we present a simple but effective method, CTCBERT, for advancing hidden-unit BERT (HuBERT). HuBERT applies a frame-level cross-entropy (CE) loss, which is similar to most acoustic model training. However, CTCBERT performs the model training with the Connectionist Temporal Classification (CTC) objective after removing duplicated IDs in each masked region. The idea stems from the observation that there can be significant errors in alignments when using clustered or aligned IDs. CTC learns alignments implicitly, indicating that learning with CTC can be more flexible when misalignment exists. We examine CTCBERT on IDs from HuBERT Iter1, HuBERT Iter2, and PBERT. The CTC training brings consistent improvements compared to the CE training. Furthermore, when loading blank-related parameters during finetuning, slight improvements are observed. Evaluated on the Librispeech 960-100h setting, the relative WER improvements of CTCBERT are 2%-11% over HuBERT and PERT on test-other data.


【13】 End-to-end Two-dimensional Sound Source Localization With Ad-hoc  Microphone Arrays

标题:基于自组织麦克风阵列的端到端二维声源定位

链接:https://arxiv.org/abs/2210.08484

作者:Yijun Gong,Shupei Liu,Xiao-Lei Zhang
机构:School of Marine Science and Technology, Northwestern Polytechnical University, China
备注:6 pages, 4 figures, coference
摘要:传统的声源定位方法大多是基于多个麦克风组成的单个麦克风阵列。它们通常被表述为到达方向的估计问题。本文提出了一种基于深度学习的自组织麦克风阵列端到端声源定位方法,其中自组织麦克风阵列是一组随机分布的麦克风阵列,它们之间相互协作.它可以产生扬声器的二维位置,每个节点只有一个麦克风。具体地,我们将目标室内空间划分为多个局部区域。我们用一个独热码对每个局部区域进行编码,这样,节点和说话人的位置就可以用独热码来表示。因此,声源定位问题被公式化为这样的分类任务:给定麦克风节点及其语音记录的一位热码,识别说话者的一位热码。针对分类问题,设计了一种端到端的时空深度模型。该算法采用时空关注度结构,在结构中间插入融合层,在模型训练和测试过程中能够处理任意数量的麦克风节点。实验结果表明,该方法在强混响和噪声环境下具有良好的性能。
摘要:Conventional sound source localization methods are mostly based on a single microphone array that consists of multiple microphones. They are usually formulated as the estimation of the direction of arrival problem. In this paper, we propose a deep-learning-based end-to-end sound source localization method with ad-hoc microphone arrays, where an ad-hoc microphone array is a set of randomly distributed microphone arrays that collaborate with each other. It can produce two-dimensional locations of speakers with only a single microphone per node. Specifically, we divide a targeted indoor space into multiple local areas. We encode each local area by a one-hot code, therefore, the node and speaker locations can be represented by the one-hot codes. Accordingly, the sound source localization problem is formulated as such a classification task of recognizing the one-hot code of the speaker given the one hot codes of the microphone nodes and their speech recordings. An end-to-end spatial-temporal deep model is designed for the classification problem. It utilizes a spatial-temporal attention architecture with a fusion layer inserted in the middle of the architecture, which is able to handle arbitrarily different numbers of microphone nodes during the model training and test. Experimental results show that the proposed method yields good performance in highly reverberant and noisy environments.


eess.AS音频处理

【1】 spatial-dccrn: dccrn equipped with frame-level angle feature and hybrid  filtering for multi-channel speech enhancement

标题:基于空域角度特征和混合滤波的多声道语音增强

链接:https://arxiv.org/abs/2210.08802

* 与cs.SD语音【9】为同一篇

作者:Shubo Lv,Yihui Fu,Yukai Jv,Lei Xie,Weixin Zhu,Wei Rao,Yannan Wang
机构:Northwestern Polytechnical University, Xi’an, China, Tencent Ethereal Audio Lab, Tencent Corporation, Shenzhen, China
摘要:近年来,多通道语音增强由于利用空间信息来区分目标语音和干扰信号而引起了人们的广泛关注。为了充分利用空间信息和基于神经网络的掩蔽估计,提出了一种多通道去噪神经网络-- Spatial DCCRN。首先,将S-DCCRN算法扩展到多信道场景,实现子信道级联和全信道处理策略,分别对不同的信道进行建模。此外,该模型不仅采用多通道频谱或级联第一通道幅度和IPD作为模型的输入,还采用了角度特征提取模块(AFE)提取帧级的角度特征嵌入,使模型能够更直观地感知空间信息。最后,针对噪声和语音同时存在于同一个时频(TF)面元时,残留噪声现象会更加严重的问题,设计了一种掩蔽映射滤波方法来代替传统的滤波求和运算,实现粗去噪、去混响和残留噪声抑制的级联。在L3 DAS 22 Challenge数据集上,提议的模型Spatial-DCCRN已经超过了EaBNet、FasNet以及几个竞争模型。不仅在3D场景中,在多通道ConferencingSpeech 2021 Challenge数据集上,Spatial-DCCRN在多个评估指标上均大幅优于最先进的(SOTA)模型MIMO-UNet。消融研究也证明了不同贡献的有效性。
摘要:Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking estimation, we propose a multi-channel denoising neural network -- Spatial DCCRN. Firstly, we extend S-DCCRN to multi-channel scenario, aiming at performing cascaded sub-channel and full-channel processing strategy, which can model different channels separately. Moreover, instead of only adopting multi-channel spectrum or concatenating first-channel's magnitude and IPD as the model's inputs, we apply an angle feature extraction module (AFE) to extract frame-level angle feature embeddings, which can help the model to apparently perceive spatial information. Finally, since the phenomenon of residual noise will be more serious when the noise and speech exist in the same time frequency (TF) bin, we particularly design a masking and mapping filtering method to substitute the traditional filter-and-sum operation, with the purpose of cascading coarsely denoising, dereverberation and residual noise suppression. The proposed model, Spatial-DCCRN, has surpassed EaBNet, FasNet as well as several competitive models on the L3DAS22 Challenge dataset. Not only the 3D scenario, Spatial-DCCRN outperforms state-of-the-art (SOTA) model MIMO-UNet by a large margin in multiple evaluation metrics on the multi-channel ConferencingSpeech2021 Challenge dataset. Ablation studies also demonstrate the effectiveness of different contributions.


【2】 Acoustic-aware Non-autoregressive Spell Correction with Mask Sample  Decoding

标题:基于模板样本译码的声学非自回归拼写纠正

链接:https://arxiv.org/abs/2210.08665

* 与cs.SD语音【10】为同一篇

作者:Ruchao Fan,Guoli Ye,Yashesh Gaur,Jinyu Li
机构:University of California, Los Angeles,Microsoft Corporation
摘要:屏蔽语言模型(MLM)已被广泛应用于任务理解,如BERT。最近,MLM也被用于生成任务。在语音识别中最常用的一种方法是使用Mask-CTC进行非自回归语音识别。本文进一步探讨了使用MLM作为变换器-转换器(transformer-transducer,TT)的非自回归拼写校正(non-autoregressive spell correction,SC)模型(MLM-SC)的可能性,初步实验表明MLM-SC对Librispeech数据没有改善。问题可能是模型单元(词块)的选择和英语数据的TT置信度得分的不准确性。为了解决这个问题,我们提出了一种掩码样本解码(MS-decode)方法,其中掩码标记可以选择被掩码或不被掩码以补偿不准确性。结果,我们分别将Librispeech测试-其他数据上的流式TT的WER从7.6%降低到6.5%,并将Aishell测试数据上的CER从7.3%降低到6.1%。
摘要:Masked language model (MLM) has been widely used for understanding tasks, e.g. BERT. Recently, MLM has also been used for generation tasks. The most popular one in speech is using Mask-CTC for non-autoregressive speech recognition. In this paper, we take one step further, and explore the possibility of using MLM as a non-autoregressive spell correction (SC) model for transformer-transducer (TT), denoted as MLM-SC. Our initial experiments show that MLM-SC provides no improvements on Librispeech data. The problem might be the choice of modeling units (word pieces) and the inaccuracy of the TT confidence scores for English data. To solve the problem, we propose a mask sample decoding (MS-decode) method where the masked tokens can have the choice of being masked or not to compensate for the inaccuracy. As a result, we reduce the WER of a streaming TT from 7.6% to 6.5% on the Librispeech test-other data and the CER from 7.3% to 6.1% on the Aishell test data, respectively.


【3】 Attention-Based Audio Embeddings for Query-by-Example

标题:基于注意力的逐例查询音频嵌入

链接:https://arxiv.org/abs/2210.08624

* 与cs.SD语音【11】为同一篇

作者:Anup Singh,Kris Demuynck,Vipul Arora
机构:IDLab, Department of Electronics and Information Systems, imec - Ghent University, Belgium,  Department of Electrical Engineering, Indian Institute of Technology Kanpur, India
摘要:一个理想的音频检索系统有效地和鲁棒地从大量的数据库中识别短的查询片段。然而,公知的音频指纹识别系统的性能在高信号失真水平上不足。提出了一种基于对比学习框架的音频检索系统,该系统能够生成噪声和混响鲁棒的音频指纹。使用这些指纹,该方法执行全面搜索以识别查询音频并精确地估计其在参考音频中的时间戳。我们的框架涉及训练CNN,以最大化从干净音频提取的嵌入对与其相应的失真和时移版本之间的相似性。我们采用通道式频谱-时间注意机制,通过赋予信号中显著的频谱-时间碎片更多的权重来更好地区分音频。实验结果表明,与现有的同类系统相比,该系统在计算量和内存利用率上都有很大的提高,尤其是在较高的失真水平下,其计算精度更高,并且可以扩展到更大的数据库。
摘要:An ideal audio retrieval system efficiently and robustly recognizes a short query snippet from an extensive database. However, the performance of well-known audio fingerprinting systems falls short at high signal distortion levels. This paper presents an audio retrieval system that generates noise and reverberation robust audio fingerprints using the contrastive learning framework. Using these fingerprints, the method performs a comprehensive search to identify the query audio and precisely estimate its timestamp in the reference audio. Our framework involves training a CNN to maximize the similarity between pairs of embeddings extracted from clean audio and its corresponding distorted and time-shifted version. We employ a channel-wise spectral-temporal attention mechanism to better discriminate the audio by giving more weight to the salient spectral-temporal patches in the signal. Experimental results indicate that our system is efficient in computation and memory usage while being more accurate, particularly at higher distortion levels, than competing state-of-the-art systems and scalable to a larger database.


【4】 CTCBERT: Advancing Hidden-unit BERT with CTC Objectives

标题:CTCBERT:以CTC目标推进隐藏单元BERT

链接:https://arxiv.org/abs/2210.08603

* 与cs.SD语音【12】为同一篇

作者:Ruchao Fan,Yiming Wang,Yashesh Gaur,Jinyu Li
机构:University of California, Los Angeles, Microsoft Corporation
摘要:本文提出了一种简单有效的隐单元BERT(HuBERT)改进算法CTCBERT. HuBERT应用帧级交叉熵(CE)损失,这类似于大多数声学模型训练。然而,CTCBERT在去除每个掩蔽区域中的重复ID之后,使用连接主义时间分类(CTC)目标来执行模型训练。这个想法源于这样的观察,即当使用聚集的或对齐的ID时,在对齐中可能存在显著的错误。CTC隐式地学习对齐,这表明当存在对齐时,使用CTC的学习可以更加灵活。我们检查了来自HuBERT Iter 1、HuBERT Iter 2和PBERT的ID上的CTCBERT。与CE培训相比,CTC培训带来了一致的改进。此外,当在微调期间加载空白相关参数时,观察到轻微的改进。在Librispeech 960- 100 h环境下,CTCBERT相对于HuBERT和PERT在测试-其他数据上的WER改善为2%-11%。
摘要:In this work, we present a simple but effective method, CTCBERT, for advancing hidden-unit BERT (HuBERT). HuBERT applies a frame-level cross-entropy (CE) loss, which is similar to most acoustic model training. However, CTCBERT performs the model training with the Connectionist Temporal Classification (CTC) objective after removing duplicated IDs in each masked region. The idea stems from the observation that there can be significant errors in alignments when using clustered or aligned IDs. CTC learns alignments implicitly, indicating that learning with CTC can be more flexible when misalignment exists. We examine CTCBERT on IDs from HuBERT Iter1, HuBERT Iter2, and PBERT. The CTC training brings consistent improvements compared to the CE training. Furthermore, when loading blank-related parameters during finetuning, slight improvements are observed. Evaluated on the Librispeech 960-100h setting, the relative WER improvements of CTCBERT are 2%-11% over HuBERT and PERT on test-other data.


【5】 End-to-end Two-dimensional Sound Source Localization With Ad-hoc  Microphone Arrays

标题:基于自组织麦克风阵列的端到端二维声源定位

链接:https://arxiv.org/abs/2210.08484

* 与cs.SD语音【13】为同一篇

作者:Yijun Gong,Shupei Liu,Xiao-Lei Zhang
机构:School of Marine Science and Technology, Northwestern Polytechnical University, China
备注:6 pages, 4 figures, coference
摘要:传统的声源定位方法大多是基于多个麦克风组成的单个麦克风阵列。它们通常被表述为到达方向的估计问题。本文提出了一种基于深度学习的自组织麦克风阵列端到端声源定位方法,其中自组织麦克风阵列是一组随机分布的麦克风阵列,它们之间相互协作.它可以产生扬声器的二维位置,每个节点只有一个麦克风。具体地,我们将目标室内空间划分为多个局部区域。我们用一个独热码对每个局部区域进行编码,这样,节点和说话人的位置就可以用独热码来表示。因此,声源定位问题被公式化为这样的分类任务:给定麦克风节点及其语音记录的一位热码,识别说话者的一位热码。针对分类问题,设计了一种端到端的时空深度模型。该算法采用时空关注度结构,在结构中间插入融合层,在模型训练和测试过程中能够处理任意数量的麦克风节点。实验结果表明,该方法在强混响和噪声环境下具有良好的性能。
摘要:Conventional sound source localization methods are mostly based on a single microphone array that consists of multiple microphones. They are usually formulated as the estimation of the direction of arrival problem. In this paper, we propose a deep-learning-based end-to-end sound source localization method with ad-hoc microphone arrays, where an ad-hoc microphone array is a set of randomly distributed microphone arrays that collaborate with each other. It can produce two-dimensional locations of speakers with only a single microphone per node. Specifically, we divide a targeted indoor space into multiple local areas. We encode each local area by a one-hot code, therefore, the node and speaker locations can be represented by the one-hot codes. Accordingly, the sound source localization problem is formulated as such a classification task of recognizing the one-hot code of the speaker given the one hot codes of the microphone nodes and their speech recordings. An end-to-end spatial-temporal deep model is designed for the classification problem. It utilizes a spatial-temporal attention architecture with a fusion layer inserted in the middle of the architecture, which is able to handle arbitrarily different numbers of microphone nodes during the model training and test. Experimental results show that the proposed method yields good performance in highly reverberant and noisy environments.


【6】 Sub-8-bit quantization for on-device speech recognition: a  regularization-free approach

标题:用于设备上语音识别的亚8位量化:一种无需正则化的方法

链接:https://arxiv.org/abs/2210.09188

* 与cs.SD语音【1】为同一篇

作者:Kai Zhen,Martin Radfar,Hieu Duy Nguyen,Grant P. Strimel,Nathan Susanj,Athanasios Mouchtaris机构:Amazon Alexa AI
备注:Accepted for publication at IEEE SLT'22
摘要:对于设备上自动语音识别(ASR),量化感知训练(QAT)是普遍存在的,以实现模型预测性能和效率之间的折衷。在现有的QAT方法中,一个主要的缺点是量化质心必须预先确定和固定。为了克服这一限制,我们引入了一种无正则化的“软到硬”压缩机制,该机制在μ定律约束空间中具有自调整质心,从而产生一种更简单但更通用的量化方案,称为通用量化器(GQ)。我们使用递归神经网络转换器(RNN-T)和Conformer架构,在LibriSpeech和去识别远场数据集上将GQ应用于ASR任务。在不降低精度的情况下,GQ可以将RNN-T和Conformer压缩到8位以下,对于一些RNN-T层,可以压缩到1位,以实现快速和准确的推断。通过物理设备基准测试,我们发现与8位QAT相比,内存占用空间节省了30.73%,用户感知延迟减少了31.75%。
摘要:For on-device automatic speech recognition (ASR), quantization aware training (QAT) is ubiquitous to achieve the trade-off between model predictive performance and efficiency. Among existing QAT methods, one major drawback is that the quantization centroids have to be predetermined and fixed. To overcome this limitation, we introduce a regularization-free, "soft-to-hard" compression mechanism with self-adjustable centroids in a mu-Law constrained space, resulting in a simpler yet more versatile quantization scheme, called General Quantizer (GQ). We apply GQ to ASR tasks using Recurrent Neural Network Transducer (RNN-T) and Conformer architectures on both LibriSpeech and de-identified far-field datasets. Without accuracy degradation, GQ can compress both RNN-T and Conformer into sub-8-bit, and for some RNN-T layers, to 1-bit for fast and accurate inference. We observe a 30.73% memory footprint saving and 31.75% user-perceived latency reduction compared to 8-bit QAT via physical device benchmarking.


【7】 Visual onoma-to-wave: environmental sound synthesis from visual  onomatopoeias and sound-source images

标题:视觉拟声学:从视觉拟声词和声源图像合成环境声音

链接:https://arxiv.org/abs/2210.09173

* 与cs.SD语音【2】为同一篇

作者:Hien Ohnaka,Shinnosuke Takamichi,Keisuke Imoto,Yuki Okamoto,Kazuki Fujii,Hiroshi Saruwatari
机构:National Institute of Technology, Tokuyama College, Japan. ,The University of Tokyo, Japan., Doshisha University, Japan. ,Ritsumeikan University, Japan.
备注:Submitted to ICASSP 2023
摘要:我们提出一种从视觉表示的拟声词和声源合成环境声音的方法。拟声词是模仿声音结构的词,即,声音的文本表示。从这一角度出发,拟声-波技术被提出来从所需的拟声文本中合成环境声音。拟声诗还有另一种表现形式:漫画、广告和虚拟现实中声音的视觉-文本表示。视觉拟声词(视觉文本拟声词)包含文本中不存在的丰富信息,如长-短持续时间的图像,因此使用这种表示法有望合成出多样化的声音。因此,我们提出视觉拟声转波的方法,以从视觉拟声合成环境声音。该方法可以将视觉文本和声源图像的视觉概念转换为合成声音。我们还提出了一种数据扩充方法,重点关注拟声词的重复,以提高我们的方法的性能。实验结果表明,该方法能够从视觉文本和声源图像中合成出多种环境声音。
摘要:We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental sounds from the desired onomatopoeia texts. Onomatopoeias have another representation: visual-text representations of sounds in comics, advertisements, and virtual reality. A visual onomatopoeia (visual text of onomatopoeia) contains rich information that is not present in the text, such as a long-short duration of the image, so the use of this representation is expected to synthesize diverse sounds. Therefore, we propose visual onoma-to-wave for environmental sound synthesis from visual onomatopoeia. The method can transfer visual concepts of the visual text and sound-source image to the synthesized sound. We also propose a data augmentation method focusing on the repetition of onomatopoeias to enhance the performance of our method. An experimental evaluation shows that the methods can synthesize diverse environmental sounds from visual text and sound-source images.


【8】 Language-agnostic Code-Switching in End-To-End Speech Recognition

标题:端到端语音识别中的语言不可知码转换

链接:https://arxiv.org/abs/2210.08992

* 与cs.SD语音【3】为同一篇

作者:Enes Yavuz Ugan,Christian Huber,Juan Hussain,Alexander Waibel
机构:Interactive Systems Lab, Karlsruhe Institute of Technology, Karlsruhe, Germany, Carnegie Mellon University, Pittsburgh PA, USA
备注:6 pages
摘要:语码转换是指不同语言中的词和短语交替使用的现象。虽然当今的神经端到端(E2E)模型在自动语音识别(ASR)任务上提供了最先进的性能,但众所周知,这些系统是非常数据密集型的。然而,只有少数转录和对齐的CS语音可用。为了克服这一问题并训练能够转录CS语音的多语言系统,我们提出了一种简单而有效的数据扩充,其中音频和不同源语言的对应标签被级联。通过使用该训练数据,我们的E2E模型改进了转录CS语音,并且还改进了多语言模型的性能。实验结果表明,这种增强技术可以使模型对训练过程中未发现的句间语言转换的预测性能提高5.03WER.
摘要:Code-Switching (CS) is referred to the phenomenon of alternately using words and phrases from different languages. While today's neural end-to-end (E2E) models deliver state-of-the-art performances on the task of automatic speech recognition (ASR) it is commonly known that these systems are very data-intensive. However, there is only a few transcribed and aligned CS speech available. To overcome this problem and train multilingual systems which can transcribe CS speech, we propose a simple yet effective data augmentation in which audio and corresponding labels of different source languages are concatenated. By using this training data, our E2E model improves on transcribing CS speech and improves performance over the multilingual model, as well. The results show that this augmentation technique can even improve the model's performance on inter-sentential language switches not seen during training by 5,03\% WER.


【9】 How to Leverage DNN-based speech enhancement for multi-channel speaker  verification?

标题:如何利用基于DNN的语音增强进行多声道说话人验证?

链接:https://arxiv.org/abs/2210.08834

* 与cs.SD语音【4】为同一篇

作者:Sandipana Dowerah,Romain Serizel,Denis Jouvet,Mohammad Mohammadamini,Driss Matrouf
机构:Université de Lorraine, CNRS, Inria-Loria., Nancy, France,  Avignon Université, Laboratoire Informatique d'Avignon, Avignon, France
备注:None
摘要:由于环境噪声和室内混响的影响,说话人确认(SV)在远场场景中的性能并不理想。本文提出了一个用于远场说话人确认的多通道语音增强基准。一种是基于深度神经网络的方法,另一种是深度神经网络与信号处理相结合的方法。我们将DNN架构与信号处理技术相结合来进行各种实验。将我们的方法与现有的最新方法进行了比较。我们检验了入组在预处理中的重要性,这在以前的研究中很大程度上被忽视了。实验结果表明,只要对注册文件进行类似于测试数据的预处理,并且测试和注册发生在相似的信噪比范围内,预处理可以提高SV的性能。在VOiCES数据集的生成和所有噪声条件上获得了相当大的改善。
摘要:Speaker verification (SV) suffers from unsatisfactory performance in far-field scenarios due to environmental noise andthe adverse impact of room reverberation. This work presents a benchmark of multichannel speech enhancement for far-fieldspeaker verification. One approach is a deep neural network-based, and the other is a combination of deep neural network andsignal processing. We integrated a DNN architecture with signal processing techniques to carry out various experiments. Ourapproach is compared to the existing state-of-the-art approaches. We examine the importance of enrollment in pre-processing,which has been largely overlooked in previous studies. Experimental evaluation shows that pre-processing can improve the SVperformance as long as the enrollment files are processed similarly to the test data and that test and enrollment occur within similarSNR ranges. Considerable improvement is obtained on the generated and all the noise conditions of the VOiCES dataset.


【10】 SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of  Self-Supervised Speech Representation Learning

标题:Superb@SLT 2022:自我监督语音表征学习的泛化和效率挑战

链接:https://arxiv.org/abs/2210.08634

* 与cs.SD语音【5】为同一篇

作者:Tzu-hsun Feng,Annie Dong,Ching-Feng Yeh,Shu-wen Yang,Tzu-Quan Lin,Jiatong Shi,Kai-Wei Chang,Zili Huang,Haibin Wu,Xuankai Chang,Shinji Watanabe,Abdelrahman Mohamed,Shang-Wen Li,Hung-yi Lee
机构:National Taiwan University, Taiwan,Meta, USA, Carnegie Mellon University, USA,Johns Hopkins University, USA
备注:Accepted by 2022 SLT Workshop
摘要:我们在SLT 2022上提出了SUPERB挑战,其目的是学习自监督语音表示以获得更好的性能、泛化和效率。该挑战建立在SUPERB基准之上,并实施了度量标准来衡量自监督学习(SSL)表示的计算需求,并评估其在不同SUPERB任务中的推广性和性能。SUPERB性能指标评测全面涵盖了流行的语音处理任务,从语音和说话人识别到音频生成和语义理解。随着SSL技术在语音领域的广泛应用,并取得了令人瞩目的成果,我们设想通过激发任务性能之外的更实用的技术设计来提升SSL技术的影响力。本文总结了14个已提交模型的结果。最后,我们讨论了这些论文的主要发现和SSL研究的未来方向。
摘要:We present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency. The challenge builds upon the SUPERB benchmark and implements metrics to measure the computation requirements of self-supervised learning (SSL) representation and to evaluate its generalizability and performance across the diverse SUPERB tasks. The SUPERB benchmark provides comprehensive coverage of popular speech processing tasks, from speech and speaker recognition to audio generation and semantic understanding. As SSL has gained interest in the speech community and showed promising outcomes, we envision the challenge to uplevel the impact of SSL techniques by motivating more practical designs of techniques beyond task performance. We summarize the results of 14 submitted models in this paper. We also discuss the main findings from those submissions and the future directions of SSL research.


【11】 Robust, General, and Low Complexity Acoustic Scene Classification  Systems and An Effective Visualization for Presenting a Sound Scene Context

标题:稳健、通用、低复杂度的声音场景分类系统和用于呈现声音场景环境的有效可视化

链接:https://arxiv.org/abs/2210.08610

* 与cs.SD语音【6】为同一篇

作者:Lam Pham,Dusan Salovic,Anahid Jalali,Alexander Schindler,Khoa Tran,Canh Vu,Phu X. Nguyen机构:Nguyen is with University ofFPT,  Tran is with Da Nang University
摘要:在本文中,我们提出了一个全面的分析声学场景分类(ASC),任务是识别一个音频记录的场景从其声学签名。特别地,我们首先提出了一个基于开始的低足迹ASC模型,称为ASC基线。然后将提出的ASC基线与MobileNetV 1、MobileNetV 2、VGG 16、VGG 19、ResNet 50 V2、ResNet 152 V2、DenseNet 121、DenseNet 201和Xception的基准和高复杂度网络架构进行了比较。其次,我们提出了一种新的深度神经网络结构,利用残差起始结构和多个核来改进ASC基线。给出了一种新的剩余起始(NRI)模型,进一步评估了模型复杂度和模型精度性能之间的权衡.最后,我们评估了在声音场景记录中发生的声音事件是否有助于提高ASC的准确性,并指出如何通过结合声音场景和声音事件信息来很好地呈现声音场景上下文。我们对各种ASC数据集进行了广泛的实验,包括拥挤场景、IEEE AASP声学场景和事件检测和分类挑战(DCASE)2018年任务1A和1B、2019年任务1A和1B、2020年任务1A、2021年任务1A、2022年任务1。几种不同ASC挑战的实验结果突出了两个主要成就:第一个是提出鲁棒的、通用的和低复杂度的ASC系统,其适合于在广泛的边缘设备和移动设备上的实际应用;二是提出了一种有效的可视化方法,用于全面地呈现声音场景上下文。
摘要:In this paper, we present a comprehensive analysis of Acoustic Scene Classification (ASC), the task of identifying the scene of an audio recording from its acoustic signature. In particular, we firstly propose an inception-based and low footprint ASC model, referred to as the ASC baseline. The proposed ASC baseline is then compared with benchmark and high-complexity network architectures of MobileNetV1, MobileNetV2, VGG16, VGG19, ResNet50V2, ResNet152V2, DenseNet121, DenseNet201, and Xception. Next, we improve the ASC baseline by proposing a novel deep neural network architecture which leverages residual-inception architectures and multiple kernels. Given the novel residual-inception (NRI) model, we further evaluate the trade off between the model complexity and the model accuracy performance. Finally, we evaluate whether sound events occurring in a sound scene recording can help to improve ASC accuracy, then indicate how a sound scene context is well presented by combining both sound scene and sound event information. We conduct extensive experiments on various ASC datasets, including Crowded Scenes, IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 Task 1A and 1B, 2019 Task 1A and 1B, 2020 Task 1A, 2021 Task 1A, 2022 Task 1. The experimental results on several different ASC challenges highlight two main achievements; the first is to propose robust, general, and low complexity ASC systems which are suitable for real-life applications on a wide range of edge devices and mobiles; the second is to propose an effective visualization method for comprehensively presenting a sound scene context.


【12】 A Policy-based Approach to the SpecAugment Method for Low Resource E2E  ASR

标题:一种基于策略的低资源E2E ASR规范调整方法

链接:https://arxiv.org/abs/2210.08520

* 与cs.SD语音【7】为同一篇

作者:Rui Li,Guodong Ma,Dexin Zhao,Ranran Zeng,Xiaoyu Li,Hao Huang
机构:∗ School of Information Science and Engineering, Xinjiang University, Urumqi, China, † China Telecom Beijing Research Institute, Beijing, China, ‡ Xinjiang Provincial Key Laboratory of Multi-lingual Information Technology, Urumqi, China, § Equal contributions
备注:Accepted to APSIPA ASC 2022
摘要:SpecAugment对于HMM和基于E2 E的自动语音识别(ASR)系统都是一种非常有效的数据增强方法。特别是,它还适用于低资源方案。然而,SpecAugment在固定的增强策略中屏蔽了时域或频域的频谱,这可能会给低资源ASR带来相对较小的数据分集。本文提出了一种基于策略的SpecAugment(Policy-SpecAugment)方法来缓解上述问题。其思想是采用增广选择策略和增广参数改变策略来求解固定路径。这些策略是基于验证集的丢失来学习的,验证集的丢失被应用于对应的增强策略。它的目的是鼓励模型学习更多的多样性数据,这是模型相对需要的。在实验中,我们评估了我们的方法在低资源场景下的有效性,即,100小时自由演讲任务。根据结果和分析,我们可以看到,使用我们的建议可以明显缓解上述问题。实验结果表明,与现有的规范增强算法相比,Policy-SpecAugment算法在测试/开发-清洁集上的相对WER降低超过10%,在测试/开发-其他集上的相对WER降低超过5%,在所有测试集上的绝对WER降低超过1%。
摘要:SpecAugment is a very effective data augmentation method for both HMM and E2E-based automatic speech recognition (ASR) systems. Especially, it also works in low-resource scenarios. However, SpecAugment masks the spectrum of time or the frequency domain in a fixed augmentation policy, which may bring relatively less data diversity to the low-resource ASR. In this paper, we propose a policy-based SpecAugment (Policy-SpecAugment) method to alleviate the above problem. The idea is to use the augmentation-select policy and the augmentation-parameter changing policy to solve the fixed way. These policies are learned based on the loss of validation set, which is applied to the corresponding augmentation policies. It aims to encourage the model to learn more diverse data, which the model relatively requires. In experiments, we evaluate the effectiveness of our approach in low-resource scenarios, i.e., the 100 hours librispeech task. According to the results and analysis, we can see that the above issue can be obviously alleviated using our proposal. In addition, the experimental results show that, compared with the state-of-the-art SpecAugment, the proposed Policy-SpecAugment has a relative WER reduction of more than 10% on the Test/Dev-clean set, more than 5% on the Test/Dev-other set, and an absolute WER reduction of more than 1% on all test sets.


【13】 Learning Invariant Representation and Risk Minimized for Unsupervised  Accent Domain Adaptation

标题:无监督重音域自适应的学习不变表征和风险最小化

链接:https://arxiv.org/abs/2210.08182

* 与cs.SD语音【8】为同一篇

作者:Chendong Zhao,Jianzong Wang,Xiaoyang Qu,Haoqian Wang,Jing Xiao
机构:† Ping An Technology (Shenzhen) Co., Ltd., China, ⋆ The Shenzhen International Graduate School, Tsinghua University, China
备注:Accepted to SLT 2022
摘要:针对语音音频的无监督表示学习在语音识别任务中获得了令人印象深刻的性能,特别是当注释语音受限时。然而,无监督范式需要仔细设计,并且对于这些表示获得什么性质知之甚少。不能保证模型学习有价值信息的有意义表示以用于识别。此外,学习的表示对其他域的适应能力仍然需要估计。在这项工作中,我们探索学习领域不变的表示,通过直接映射的语音表示,以其相应的高级语言信息。实验结果表明,学习后的潜台词不仅能捕捉到每个音素的发音特征,而且能提高适应能力,在重音基准测试中的表现大大优于基线。
摘要:Unsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the unsupervised paradigm needs to be carefully designed and little is known about what properties these representations acquire. There is no guarantee that the model learns meaningful representations for valuable information for recognition. Moreover, the adaptation ability of the learned representations to other domains still needs to be estimated. In this work, we explore learning domain-invariant representations via a direct mapping of speech representations to their corresponding high-level linguistic informations. Results prove that the learned latents not only capture the articulatory feature of each phoneme but also enhance the adaptation ability, outperforming the baseline largely on accented benchmarks.


机器翻译,仅供参考