【1】 Fine-grained Early Frequency Attention for Deep Speaker Recognition
标题:用于深度说话人识别的细粒度早期频率注意
链接:https://arxiv.org/abs/2207.10006
作者:Amirhossein Hajavi,Ali Etemad机构:Dept. ECE and Ingenuity Labs Research Institute, Queen’s University, Kingston, Canada备注:Accepted In IJCNN 2022摘要:注意力机制是提高深度模型性能的重要工具,但目前用于说话人识别的注意力机制没有考虑到深度网络输入频谱中的细粒度信息项,如频率单元等,我们提出了一种新的细粒度早期频率注意(FEFA)用于野外说话人识别。一旦集成到深度神经网络中,我们提出的机制通过从网络的早期层获得查询并生成可学习的权重来处理小到为了评估FEFA的性能,我们使用了几个著名的深度模型作为主干网络,并将我们的注意力模块集成到它们的流水线中。在VoxCeleb1数据集上评估了这些网络(有和没有FEFA)的整体性能,在那里我们观察到使用FEFA时有相当大的改善。摘要:Attention mechanisms have emerged as important tools that boost the performance of deep models by allowing them to focus on key parts of learned embeddings. However, current attention mechanisms used in speaker recognition tasks fail to consider fine-grained information items such as frequency bins in input spectral representations used by the deep networks. To address this issue, we propose the novel Fine-grained Early Frequency Attention (FEFA) for speaker recognition in-the-wild. Once integrated into a deep neural network, our proposed mechanism works by obtaining queries from early layers of the network and generating learnable weights to attend to information items as small as the frequency bins in the input spectral representations. To evaluate the performance of FEFA, we use several well-known deep models as backbone networks and integrate our attention module in their pipelines. The overall performance of these networks (with and without FEFA) are evaluated on the VoxCeleb1 dataset, where we observe considerable improvements when FEFA is used.
【2】 Diffsound: Discrete Diffusion Model for Text-to-sound Generation
标题:Diffsound:用于文本到声音生成的离散扩散模型
链接:https://arxiv.org/abs/2207.09983
作者:Dongchao Yang,Jianwei Yu,Helin Wang,Wen Wang,Chao Weng,Yuexian Zou,Dong Yu机构:School of Electronic andComputer Engineering, Peking University备注:Submitted to TASLP2022摘要:产生人类想要的音效是一个重要的课题,然而在这方面的研究很少.在本研究中,我们研究了在文本提示下产生声音的问题,并提出了一个新颖的文本到声音的产生框架,该框架由一个文本编码器,一个矢量量化变分自动编码器组成(VQ-VAE)、解码器和声码器,该框架首先利用解码器将从文本编码器中提取的文本特征在VQ-VAE的帮助下转换为mel谱图,然后用声码器将mel谱图变换成波形.我们发现解码器对mel谱图的生成性能有很大的影响.因此,我们把研究的重点放在设计一个好的解码器上.我们从传统的自回归解码器开始,该方法在以往的声音生成工作中被证明是一种先进的方法.然而,AR解码器总是按顺序逐个预测mel谱图特征点,这就引入了单向偏差和误差累积问题.此外,使用AR解码器,声音生成时间与声音持续时间成线性增长.为了克服AR解码器引入的缺点,我们提出了一种基于离散扩散模型的非自回归解码器Diffsound,该算法首先对mel谱图的所有特征点进行预测,然后再对预测出的特征点进行细化,实验结果表明,本文提出的Diffsound算法不仅能获得较好的文本—文本转换效果,而且能获得较好的预测效果.与AR解码器相比,声音生成结果更好,而且还具有更快的生成速度,例如MOS:3.56 \textitt {v. s} 2.786,并且生成速度比AR解码器快五倍。摘要:Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound generation framework that consists of a text encoder, a Vector Quantized Variational Autoencoder (VQ-VAE), a decoder, and a vocoder. The framework first uses the decoder to transfer the text features extracted from the text encoder to a mel-spectrogram with the help of VQ-VAE, and then the vocoder is used to transform the generated mel-spectrogram into a waveform. We found that the decoder significantly influences the generation performance. Thus, we focus on designing a good decoder in this study. We begin with the traditional autoregressive decoder, which has been proved as a state-of-the-art method in previous sound generation works. However, the AR decoder always predicts the mel-spectrogram tokens one by one in order, which introduces the unidirectional bias and accumulation of errors problems. Moreover, with the AR decoder, the sound generation time increases linearly with the sound duration. To overcome the shortcomings introduced by AR decoders, we propose a non-autoregressive decoder based on the discrete diffusion model, named Diffsound. Specifically, the Diffsound predicts all of the mel-spectrogram tokens in one step and then refines the predicted tokens in the next step, so the best-predicted results can be obtained after several steps. Our experiments show that our proposed Diffsound not only produces better text-to-sound generation results when compared with the AR decoder but also has a faster generation speed, e.g., MOS: 3.56 \textit{v.s} 2.786, and the generation speed is five times faster than the AR decoder.
【3】 When Is TTS Augmentation Through a Pivot Language Useful?
标题:通过Pivot语言增强TTS何时有用?
链接:https://arxiv.org/abs/2207.09889
作者:Nathaniel Robinson,Perez Ogayo,Swetha Gangu,David R. Mortensen,Shinji Watanabe机构:Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA, USA摘要:由于转录的音频数据量很小,为低资源语言开发自动语音识别(ASR)是一项挑战。对于许多这类语言,音频和文本是分开提供的,但音频不能与转录一起提供。使用文本,语音可以通过文本到语音(TTS)系统合成产生。然而,许多低资源语言也没有高质量的TTS系统。我们提出了一种替代方案:通过一个经过训练的TTS系统运行目标语言的文本来生成合成音频。我们研究了在低资源环境中何时以及如何使用这种技术最有效。在我们的实验中,使用几千个合成TTS文本—语音对并复制真实数据来平衡产生最佳结果。我们的研究结果表明,在一组候选枢纽语言上进行搜索可以带来边际改进,令人惊讶的是,语音合成质量的提高会影响ASR的性能,应用这些研究结果,ASR的字符错误减少率(CERR)分别提高了64.5%和45.0%.分别用于两种低资源语言:瓜兰和苏巴。摘要:Developing Automatic Speech Recognition (ASR) for low-resource languages is a challenge due to the small amount of transcribed audio data. For many such languages, audio and text are available separately, but not audio with transcriptions. Using text, speech can be synthetically produced via text-to-speech (TTS) systems. However, many low-resource languages do not have quality TTS systems either. We propose an alternative: produce synthetic audio by running text from the target language through a trained TTS system for a higher-resource pivot language. We investigate when and how this technique is most effective in low-resource settings. In our experiments, using several thousand synthetic TTS text-speech pairs and duplicating authentic data to balance yields optimal results. Our findings suggest that searching over a set of candidate pivot languages can lead to marginal improvements and that, surprisingly, ASR performance can by harmed by increases in measured TTS quality. Application of these findings improves ASR by 64.5\% and 45.0\% character error reduction rate (CERR) respectively for two low-resource languages: Guaran\'i and Suba.
【4】 Improving Data Driven Inverse Text Normalization using Data Augmentation
标题:利用数据扩充改进数据驱动的反向文本规范化
链接:https://arxiv.org/abs/2207.09674
作者:Laxmi Pandey,Debjyoti Paul,Pooja Chitkara,Yutong Pang,Xuedong Zhang,Kjell Schubert,Mark Chou,Shu Liu,Yatharth Saraf机构:University of California Merced†, Meta Inc.摘要:反向文本规范化(ITN)用于转换自动语音识别的口语形式输出传统的手工ITN规则在转录和维护方面可能很复杂。同时,神经建模方法需要与ASR系统在相同或相似领域的高质量大规模口语—书面语对示例这两种方法都需要昂贵而复杂的标注,我们提出了一种数据扩充技术,该技术利用最少的人工注释从域外文本数据有效地生成丰富的口头—书面数字对。我们经验性地证明,使用我们的数据扩充技术训练的ITN模型在所有数值表面(如基数、货币和分数)上始终优于仅使用域内数据训练的ITN模型,总体准确率为14.44%。摘要:Inverse text normalization (ITN) is used to convert the spoken form output of an automatic speech recognition (ASR) system to a written form. Traditional handcrafted ITN rules can be complex to transcribe and maintain. Meanwhile neural modeling approaches require quality large-scale spoken-written pair examples in the same or similar domain as the ASR system (in-domain data), to train. Both these approaches require costly and complex annotations. In this paper, we present a data augmentation technique that effectively generates rich spoken-written numeric pairs from out-of-domain textual data with minimal human annotation. We empirically demonstrate that ITN model trained using our data augmentation technique consistently outperform ITN model trained using only in-domain data across all numeric surfaces like cardinal, currency, and fraction, by an overall accuracy of 14.44%.
【5】 COVID-19 Detection from Respiratory Sounds with Hierarchical Spectrogram Transformers
标题:基于分级谱图Transformers的COVID-19呼吸音检测
链接:https://arxiv.org/abs/2207.09529
作者:Idil Aytekin,Onat Dalmaz,Kaan Gonc,Haydar Ankishan,Emine U Saritas,Ulas Bagci,Haydar Celik,Tolga Cukur机构:and Tolga C¸ ukur∗, Senior Member摘要:对流行的空气传播疾病(如COVID-19)的监测特点是涉及呼吸评估。虽然听诊是症状监测的主流方法,但其诊断效用因需要专门的医院就诊而受到阻碍。基于便携式设备上呼吸音记录的持续远程监测是一种有前途的替代方法,可以帮助筛查COVID-19。在本研究中,我们引入了一种新的深度学习方法,通过咳嗽或呼吸声的音频记录来区分COVID-19患者和健康对照。所提出的方法利用了一种新的分层谱图Transformer HST在声谱图的局部窗口上体现了自我注意机制,在多个国家数据集上的综合验证表明,HST方法优于其他同类方法,在检测COVID-19病例时达到97%以上的受试者工作特征曲线下面积(AUC)。摘要:Monitoring of prevalent airborne diseases such as COVID-19 characteristically involve respiratory assessments. While auscultation is a mainstream method for symptomatic monitoring, its diagnostic utility is hampered by the need for dedicated hospital visits. Continual remote monitoring based on recordings of respiratory sounds on portable devices is a promising alternative, which can assist in screening of COVID-19. In this study, we introduce a novel deep learning approach to distinguish patients with COVID-19 from healthy controls given audio recordings of cough or breathing sounds. The proposed approach leverages a novel hierarchical spectrogram transformer (HST) on spectrogram representations of respiratory sounds. HST embodies self-attention mechanisms over local windows in spectrograms, and window size is progressively grown over model stages to capture local to global context. HST is compared against state-of-the-art conventional and deep-learning baselines. Comprehensive demonstrations on a multi-national dataset indicate that HST outperforms competing methods, achieving over 97% area under the receiver operating characteristic curve (AUC) in detecting COVID-19 cases.
【6】 Towards Transfer Learning of wav2vec 2.0 for Automatic Lyric Transcription
标题:基于wav2vec 2.0的歌词自动转录迁移学习
链接:https://arxiv.org/abs/2207.09747
作者:Longshen Ou,Xiangming Gu,Ye Wang机构:School of Computing, National University of Singapore, ∗Both authors contributed equally to this research.备注:Draft accepted by ISMIR 2022摘要:自动语音识别近年来,由于大规模数据集和自监督学习范式的发展,ASR(ASR)取得了长足的进步然而,作为歌唱领域中的对应问题,自动歌词转录(ALT)的数据有限,歌词的清晰度下降,导致其发展速度较慢。为了填补ALT和ASR之间的性能差距,我们试图利用语音和歌唱之间的相似性,提出了一种基于迁移学习的ALT解决方案,通过探讨不同迁移起点对迁移学习效果的影响,最大限度地提高迁移学习的效果。在各种ALT基准数据集上,我们的方法都比以前的方法有很大的优势.进一步的实验表明,即使在很小的训练数据比例下,我们的方法仍然实现了有竞争力的性能。摘要:Automatic speech recognition (ASR) has progressed significantly in recent years due to large-scale datasets and the paradigm of self-supervised learning (SSL) methods. However, as its counterpart problem in the singing domain, automatic lyric transcription (ALT) suffers from limited data and degraded intelligibility of sung lyrics, which has caused it to develop at a slower pace. To fill in the performance gap between ALT and ASR, we attempt to exploit the similarities between speech and singing. In this work, we propose a transfer-learning-based ALT solution that takes advantage of these similarities by adapting wav2vec 2.0, an SSL ASR model, to the singing domain. We maximize the effectiveness of transfer learning by exploring the influence of different transfer starting points. We further enhance the performance by extending the original CTC model to a hybrid CTC/attention model. Our method surpasses previous approaches by a large margin on various ALT benchmark datasets. Further experiment shows that, with even a tiny proportion of training data, our method still achieves competitive performance.
【7】 Direct and Residual Subspace Decomposition of Spatial Room Impulse Responses
标题:空间房间脉冲响应的直接和残差子空间分解
链接:https://arxiv.org/abs/2207.09733
作者:Thomas Deppisch,Sebastià V. Amengual Garí,Paul Calamia,Jens Ahrens机构:ThomasDeppischandJensAhrensarewiththeChalmersUniversityofTechnology备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing摘要:心理声学实验已经表明,特别是直达声、显著反射并且声学房间响应的后期混响可以对给定房间的听觉感知具有明显的影响。(SRIR)捕获这些属性,并因此用于方向相关的房间声学分析和虚拟声学渲染。本文提出了一种子空间方法,将SRIR分解为直达声和显著反射声组成的直达部分和残差部分,以通过提供对这些组件的单独访问来促进增强的分析和呈现方法。该方法基于广义奇异值分解,将残差分解为噪声,并将其与混响的其他分量分离。它利用噪声估计来识别大的广义奇异值,然后将其归于直接部分。通过在迭代地更新噪声估计的同时从SRIR的末尾向开头推进,该方法能够处理各向异性的慢时变混响声场,不需要确定混响声场的方向。本文提出了一种新的基于小波变换的反射波初至估计方法,与现有方法相比,该方法能更好地分离直达波和残差,对实测信噪比的分析表明,该方法在不同的声学条件下都具有很强的鲁棒性,并给出了一个参考实现。摘要:Psychoacoustic experiments have shown that directional properties of, in particular, the direct sound, salient reflections, and the late reverberation of an acoustic room response can have a distinct influence on the auditory perception of a given room. Spatial room impulse responses (SRIRs) capture those properties and thus are used for direction-dependent room acoustic analysis and virtual acoustic rendering. This work proposes a subspace method that decomposes SRIRs into a direct part, which comprises the direct sound and the salient reflections, and a residual, to facilitate enhanced analysis and rendering methods by providing individual access to these components. The proposed method is based on the generalized singular value decomposition and interprets the residual as noise that is to be separated from the other components of the reverberation. It utilizes a noise estimate to identify large generalized singular values, which are then attributed to the direct part. By advancing from the end of the SRIR toward the beginning while iteratively updating the noise estimate, the method is able to work with anisotropic and slowly time-varying reverberant sound fields. The proposed method does not require direction-of-arrival estimation of reflections and shows an improved separation of the direct part from the residual compared to an existing approach. A case study with measured SRIRs suggests a high robustness of the method under different acoustic conditions. A reference implementation is provided.
【8】 Introducing Auxiliary Text Query-modifier to Content-based Audio Retrieval
标题:在基于内容的音频检索中引入辅助文本查询修饰符
链接:https://arxiv.org/abs/2207.09732
作者:Daiki Takeuchi,Yasunori Ohishi,Daisuke Niizumi,Noboru Harada,Kunio Kashino机构:NTT Corporation, Japan备注:Accepted to Interspeech 2022摘要:在公共网站上可获得的音频数据量正在迅速增长,本文提出了一种基于内容的音频检索方法,该方法通过引入描述查询音频与目标音频之间差异的辅助文本信息,能够检索出与查询音频相似但略有不同的目标音频,而传统的内容检索方法——文本检索方法——不能有效地检索出与查询音频相似但略有不同的目标音频.基于音频的检索局限于与查询音频相似的音频,本文提出的方法通过在共享的潜在空间中嵌入查询音频的基础上嵌入辅助文本查询修饰语来调整检索范围,我们构建了包括两个不同的音频剪辑和描述差别的文本的数据集。实验结果表明,该方法能够更准确地检索出配对音频。我们还基于可视化证实了所提出的方法获得了共享的潜在空间,其中音频差异和对应的文本被表示为相似的嵌入向量。摘要:The amount of audio data available on public websites is growing rapidly, and an efficient mechanism for accessing the desired data is necessary. We propose a content-based audio retrieval method that can retrieve a target audio that is similar to but slightly different from the query audio by introducing auxiliary textual information which describes the difference between the query and target audio. While the range of conventional content-based audio retrieval is limited to audio that is similar to the query audio, the proposed method can adjust the retrieval range by adding an embedding of the auxiliary text query-modifier to the embedding of the query sample audio in a shared latent space. To evaluate our method, we built a dataset comprising two different audio clips and the text that describes the difference. The experimental results show that the proposed method retrieves the paired audio more accurately than the baseline. We also confirmed based on visualization that the proposed method obtains the shared latent space in which the audio difference and the corresponding text are represented as similar embedding vectors.
【9】 ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding
标题:ESPnet-SE++:用于鲁棒语音识别、翻译和理解的语音增强
链接:https://arxiv.org/abs/2207.09514
作者:Yen-Ju Lu,Xuankai Chang,Chenda Li,Wangyou Zhang,Samuele Cornell,Zhaoheng Ni,Yoshiki Masuyama,Brian Yan,Robin Scheibler,Zhong-Qiu Wang,Yu Tsao,Yanmin Qian,Shinji Watanabe机构:Academia Sinica, Taipei, Carnegie Mellon University, USA, Shanghai Jiao Tong University, Shanghai, Italy, Meta AI, USA, LINE Corporation, Japan, Tokyo Metropolitan University, Japan备注:To appear in Interspeech 2022摘要:本文介绍了语音分离与增强技术的最新进展(SSE)集成到ESPnet工具包中。与以前的ESPnet-SE工作相比,增加了许多功能,包括最新的最先进的语音增强模型及其各自的训练和评估方法。重要的是,设计了一个新的接口,可以灵活地将语音增强前端与其他任务结合起来,包括自动语音识别(ASR),语音翻译(ST),口语理解为了展示这种集成,我们在精心设计的合成数据集上进行了噪声混响多通道ST和SLU任务的实验,这些任务可以作为未来研究的基准语料库。我们还使用CHiME-4和WSJ0-2Mix对多通道和单通道SE方法进行了基准测试.结果表明SE前端与后端任务的集成是一个很有前途的研究方向,即使对于ASR以外的任务也是如此,尤其是在多通道场景中.代码可在https://github.com/ESPnet/ESPnet。多通道ST和SLU数据集是这项工作的另一个贡献,在HuggingFace上发布。摘要:This paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit. Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-the-art speech enhancement models with their respective training and evaluation recipes. Importantly, a new interface has been designed to flexibly combine speech enhancement front-ends with other tasks, including automatic speech recognition (ASR), speech translation (ST), and spoken language understanding (SLU). To showcase such integration, we performed experiments on carefully designed synthetic datasets for noisy-reverberant multi-channel ST and SLU tasks, which can be used as benchmark corpora for future research. In addition to these new tasks, we also use CHiME-4 and WSJ0-2Mix to benchmark multi- and single-channel SE approaches. Results show that the integration of SE front-ends with back-end tasks is a promising research direction even for tasks besides ASR, especially in the multi-channel scenario. The code is available online at https://github.com/ESPnet/ESPnet. The multi-channel ST and SLU datasets, which are another contribution of this work, are released on HuggingFace.
【1】 Towards Transfer Learning of wav2vec 2.0 for Automatic Lyric Transcription
标题:基于wav2vec 2.0的歌词自动转录迁移学习
链接:https://arxiv.org/abs/2207.09747
* 与cs.SD语音【6】为同一篇
作者:Longshen Ou,Xiangming Gu,Ye Wang机构:School of Computing, National University of Singapore, ∗Both authors contributed equally to this research.备注:Draft accepted by ISMIR 2022摘要:自动语音识别近年来,由于大规模数据集和自监督学习范式的发展,ASR(ASR)取得了长足的进步然而,作为歌唱领域中的对应问题,自动歌词转录(ALT)的数据有限,歌词的清晰度下降,导致其发展速度较慢。为了填补ALT和ASR之间的性能差距,我们试图利用语音和歌唱之间的相似性,提出了一种基于迁移学习的ALT解决方案,通过探讨不同迁移起点对迁移学习效果的影响,最大限度地提高迁移学习的效果。在各种ALT基准数据集上,我们的方法都比以前的方法有很大的优势.进一步的实验表明,即使在很小的训练数据比例下,我们的方法仍然实现了有竞争力的性能。摘要:Automatic speech recognition (ASR) has progressed significantly in recent years due to large-scale datasets and the paradigm of self-supervised learning (SSL) methods. However, as its counterpart problem in the singing domain, automatic lyric transcription (ALT) suffers from limited data and degraded intelligibility of sung lyrics, which has caused it to develop at a slower pace. To fill in the performance gap between ALT and ASR, we attempt to exploit the similarities between speech and singing. In this work, we propose a transfer-learning-based ALT solution that takes advantage of these similarities by adapting wav2vec 2.0, an SSL ASR model, to the singing domain. We maximize the effectiveness of transfer learning by exploring the influence of different transfer starting points. We further enhance the performance by extending the original CTC model to a hybrid CTC/attention model. Our method surpasses previous approaches by a large margin on various ALT benchmark datasets. Further experiment shows that, with even a tiny proportion of training data, our method still achieves competitive performance.
【2】 Direct and Residual Subspace Decomposition of Spatial Room Impulse Responses
标题:空间房间脉冲响应的直接和残差子空间分解
链接:https://arxiv.org/abs/2207.09733
* 与cs.SD语音【7】为同一篇
作者:Thomas Deppisch,Sebastià V. Amengual Garí,Paul Calamia,Jens Ahrens机构:ThomasDeppischandJensAhrensarewiththeChalmersUniversityofTechnology备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing摘要:心理声学实验已经表明,特别是直达声、显著反射并且声学房间响应的后期混响可以对给定房间的听觉感知具有明显的影响。(SRIR)捕获这些属性,并因此用于方向相关的房间声学分析和虚拟声学渲染。本文提出了一种子空间方法,将SRIR分解为直达声和显著反射声组成的直达部分和残差部分,以通过提供对这些组件的单独访问来促进增强的分析和呈现方法。该方法基于广义奇异值分解,将残差分解为噪声,并将其与混响的其他分量分离。它利用噪声估计来识别大的广义奇异值,然后将其归于直接部分。通过在迭代地更新噪声估计的同时从SRIR的末尾向开头推进,该方法能够处理各向异性的慢时变混响声场,不需要确定混响声场的方向。本文提出了一种新的基于小波变换的反射波初至估计方法,与现有方法相比,该方法能更好地分离直达波和残差,对实测信噪比的分析表明,该方法在不同的声学条件下都具有很强的鲁棒性,并给出了一个参考实现。摘要:Psychoacoustic experiments have shown that directional properties of, in particular, the direct sound, salient reflections, and the late reverberation of an acoustic room response can have a distinct influence on the auditory perception of a given room. Spatial room impulse responses (SRIRs) capture those properties and thus are used for direction-dependent room acoustic analysis and virtual acoustic rendering. This work proposes a subspace method that decomposes SRIRs into a direct part, which comprises the direct sound and the salient reflections, and a residual, to facilitate enhanced analysis and rendering methods by providing individual access to these components. The proposed method is based on the generalized singular value decomposition and interprets the residual as noise that is to be separated from the other components of the reverberation. It utilizes a noise estimate to identify large generalized singular values, which are then attributed to the direct part. By advancing from the end of the SRIR toward the beginning while iteratively updating the noise estimate, the method is able to work with anisotropic and slowly time-varying reverberant sound fields. The proposed method does not require direction-of-arrival estimation of reflections and shows an improved separation of the direct part from the residual compared to an existing approach. A case study with measured SRIRs suggests a high robustness of the method under different acoustic conditions. A reference implementation is provided.
【3】 Introducing Auxiliary Text Query-modifier to Content-based Audio Retrieval
标题:在基于内容的音频检索中引入辅助文本查询修饰符
链接:https://arxiv.org/abs/2207.09732
* 与cs.SD语音【8】为同一篇
作者:Daiki Takeuchi,Yasunori Ohishi,Daisuke Niizumi,Noboru Harada,Kunio Kashino机构:NTT Corporation, Japan备注:Accepted to Interspeech 2022摘要:在公共网站上可获得的音频数据量正在迅速增长,本文提出了一种基于内容的音频检索方法,该方法通过引入描述查询音频与目标音频之间差异的辅助文本信息,能够检索出与查询音频相似但略有不同的目标音频,而传统的内容检索方法——文本检索方法——不能有效地检索出与查询音频相似但略有不同的目标音频.基于音频的检索局限于与查询音频相似的音频,本文提出的方法通过在共享的潜在空间中嵌入查询音频的基础上嵌入辅助文本查询修饰语来调整检索范围,我们构建了包括两个不同的音频剪辑和描述差别的文本的数据集。实验结果表明,该方法能够更准确地检索出配对音频。我们还基于可视化证实了所提出的方法获得了共享的潜在空间,其中音频差异和对应的文本被表示为相似的嵌入向量。摘要:The amount of audio data available on public websites is growing rapidly, and an efficient mechanism for accessing the desired data is necessary. We propose a content-based audio retrieval method that can retrieve a target audio that is similar to but slightly different from the query audio by introducing auxiliary textual information which describes the difference between the query and target audio. While the range of conventional content-based audio retrieval is limited to audio that is similar to the query audio, the proposed method can adjust the retrieval range by adding an embedding of the auxiliary text query-modifier to the embedding of the query sample audio in a shared latent space. To evaluate our method, we built a dataset comprising two different audio clips and the text that describes the difference. The experimental results show that the proposed method retrieves the paired audio more accurately than the baseline. We also confirmed based on visualization that the proposed method obtains the shared latent space in which the audio difference and the corresponding text are represented as similar embedding vectors.
【4】 ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding
标题:ESPnet-SE++:用于鲁棒语音识别、翻译和理解的语音增强
链接:https://arxiv.org/abs/2207.09514
* 与cs.SD语音【9】为同一篇
作者:Yen-Ju Lu,Xuankai Chang,Chenda Li,Wangyou Zhang,Samuele Cornell,Zhaoheng Ni,Yoshiki Masuyama,Brian Yan,Robin Scheibler,Zhong-Qiu Wang,Yu Tsao,Yanmin Qian,Shinji Watanabe机构:Academia Sinica, Taipei, Carnegie Mellon University, USA, Shanghai Jiao Tong University, Shanghai, Italy, Meta AI, USA, LINE Corporation, Japan, Tokyo Metropolitan University, Japan备注:To appear in Interspeech 2022摘要:本文介绍了语音分离与增强技术的最新进展(SSE)集成到ESPnet工具包中。与以前的ESPnet-SE工作相比,增加了许多功能,包括最新的最先进的语音增强模型及其各自的训练和评估方法。重要的是,设计了一个新的接口,可以灵活地将语音增强前端与其他任务结合起来,包括自动语音识别(ASR),语音翻译(ST),口语理解为了展示这种集成,我们在精心设计的合成数据集上进行了噪声混响多通道ST和SLU任务的实验,这些任务可以作为未来研究的基准语料库。我们还使用CHiME-4和WSJ0-2Mix对多通道和单通道SE方法进行了基准测试.结果表明SE前端与后端任务的集成是一个很有前途的研究方向,即使对于ASR以外的任务也是如此,尤其是在多通道场景中.代码可在https://github.com/ESPnet/ESPnet。多通道ST和SLU数据集是这项工作的另一个贡献,在HuggingFace上发布。摘要:This paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit. Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-the-art speech enhancement models with their respective training and evaluation recipes. Importantly, a new interface has been designed to flexibly combine speech enhancement front-ends with other tasks, including automatic speech recognition (ASR), speech translation (ST), and spoken language understanding (SLU). To showcase such integration, we performed experiments on carefully designed synthetic datasets for noisy-reverberant multi-channel ST and SLU tasks, which can be used as benchmark corpora for future research. In addition to these new tasks, we also use CHiME-4 and WSJ0-2Mix to benchmark multi- and single-channel SE approaches. Results show that the integration of SE front-ends with back-end tasks is a promising research direction even for tasks besides ASR, especially in the multi-channel scenario. The code is available online at https://github.com/ESPnet/ESPnet. The multi-channel ST and SLU datasets, which are another contribution of this work, are released on HuggingFace.
【5】 Fine-grained Early Frequency Attention for Deep Speaker Recognition
标题:用于深度说话人识别的细粒度早期频率注意
链接:https://arxiv.org/abs/2207.10006
* 与cs.SD语音【1】为同一篇
作者:Amirhossein Hajavi,Ali Etemad机构:Dept. ECE and Ingenuity Labs Research Institute, Queen’s University, Kingston, Canada备注:Accepted In IJCNN 2022摘要:注意力机制是提高深度模型性能的重要工具,但目前用于说话人识别的注意力机制没有考虑到深度网络输入频谱中的细粒度信息项,如频率单元等,我们提出了一种新的细粒度早期频率注意(FEFA)用于野外说话人识别。一旦集成到深度神经网络中,我们提出的机制通过从网络的早期层获得查询并生成可学习的权重来处理小到为了评估FEFA的性能,我们使用了几个著名的深度模型作为主干网络,并将我们的注意力模块集成到它们的流水线中。在VoxCeleb1数据集上评估了这些网络(有和没有FEFA)的整体性能,在那里我们观察到使用FEFA时有相当大的改善。摘要:Attention mechanisms have emerged as important tools that boost the performance of deep models by allowing them to focus on key parts of learned embeddings. However, current attention mechanisms used in speaker recognition tasks fail to consider fine-grained information items such as frequency bins in input spectral representations used by the deep networks. To address this issue, we propose the novel Fine-grained Early Frequency Attention (FEFA) for speaker recognition in-the-wild. Once integrated into a deep neural network, our proposed mechanism works by obtaining queries from early layers of the network and generating learnable weights to attend to information items as small as the frequency bins in the input spectral representations. To evaluate the performance of FEFA, we use several well-known deep models as backbone networks and integrate our attention module in their pipelines. The overall performance of these networks (with and without FEFA) are evaluated on the VoxCeleb1 dataset, where we observe considerable improvements when FEFA is used.
【6】 Diffsound: Discrete Diffusion Model for Text-to-sound Generation
标题:Diffsound:用于文本到声音生成的离散扩散模型
链接:https://arxiv.org/abs/2207.09983
* 与cs.SD语音【2】为同一篇
作者:Dongchao Yang,Jianwei Yu,Helin Wang,Wen Wang,Chao Weng,Yuexian Zou,Dong Yu机构:School of Electronic andComputer Engineering, Peking University备注:Submitted to TASLP2022摘要:产生人类想要的音效是一个重要的课题,然而在这方面的研究很少.在本研究中,我们研究了在文本提示下产生声音的问题,并提出了一个新颖的文本到声音的产生框架,该框架由一个文本编码器,一个矢量量化变分自动编码器组成(VQ-VAE)、解码器和声码器,该框架首先利用解码器将从文本编码器中提取的文本特征在VQ-VAE的帮助下转换为mel谱图,然后用声码器将mel谱图变换成波形.我们发现解码器对mel谱图的生成性能有很大的影响.因此,我们把研究的重点放在设计一个好的解码器上.我们从传统的自回归解码器开始,该方法在以往的声音生成工作中被证明是一种先进的方法.然而,AR解码器总是按顺序逐个预测mel谱图特征点,这就引入了单向偏差和误差累积问题.此外,使用AR解码器,声音生成时间与声音持续时间成线性增长.为了克服AR解码器引入的缺点,我们提出了一种基于离散扩散模型的非自回归解码器Diffsound,该算法首先对mel谱图的所有特征点进行预测,然后再对预测出的特征点进行细化,实验结果表明,本文提出的Diffsound算法不仅能获得较好的文本—文本转换效果,而且能获得较好的预测效果.与AR解码器相比,声音生成结果更好,而且还具有更快的生成速度,例如MOS:3.56 \textitt {v. s} 2.786,并且生成速度比AR解码器快五倍。摘要:Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound generation framework that consists of a text encoder, a Vector Quantized Variational Autoencoder (VQ-VAE), a decoder, and a vocoder. The framework first uses the decoder to transfer the text features extracted from the text encoder to a mel-spectrogram with the help of VQ-VAE, and then the vocoder is used to transform the generated mel-spectrogram into a waveform. We found that the decoder significantly influences the generation performance. Thus, we focus on designing a good decoder in this study. We begin with the traditional autoregressive decoder, which has been proved as a state-of-the-art method in previous sound generation works. However, the AR decoder always predicts the mel-spectrogram tokens one by one in order, which introduces the unidirectional bias and accumulation of errors problems. Moreover, with the AR decoder, the sound generation time increases linearly with the sound duration. To overcome the shortcomings introduced by AR decoders, we propose a non-autoregressive decoder based on the discrete diffusion model, named Diffsound. Specifically, the Diffsound predicts all of the mel-spectrogram tokens in one step and then refines the predicted tokens in the next step, so the best-predicted results can be obtained after several steps. Our experiments show that our proposed Diffsound not only produces better text-to-sound generation results when compared with the AR decoder but also has a faster generation speed, e.g., MOS: 3.56 \textit{v.s} 2.786, and the generation speed is five times faster than the AR decoder.
【7】 When Is TTS Augmentation Through a Pivot Language Useful?
标题:通过Pivot语言增强TTS何时有用?
链接:https://arxiv.org/abs/2207.09889
* 与cs.SD语音【3】为同一篇
作者:Nathaniel Robinson,Perez Ogayo,Swetha Gangu,David R. Mortensen,Shinji Watanabe机构:Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA, USA摘要:由于转录的音频数据量很小,为低资源语言开发自动语音识别(ASR)是一项挑战。对于许多这类语言,音频和文本是分开提供的,但音频不能与转录一起提供。使用文本,语音可以通过文本到语音(TTS)系统合成产生。然而,许多低资源语言也没有高质量的TTS系统。我们提出了一种替代方案:通过一个经过训练的TTS系统运行目标语言的文本来生成合成音频。我们研究了在低资源环境中何时以及如何使用这种技术最有效。在我们的实验中,使用几千个合成TTS文本—语音对并复制真实数据来平衡产生最佳结果。我们的研究结果表明,在一组候选枢纽语言上进行搜索可以带来边际改进,令人惊讶的是,语音合成质量的提高会影响ASR的性能,应用这些研究结果,ASR的字符错误减少率(CERR)分别提高了64.5%和45.0%.分别用于两种低资源语言:瓜兰和苏巴。摘要:Developing Automatic Speech Recognition (ASR) for low-resource languages is a challenge due to the small amount of transcribed audio data. For many such languages, audio and text are available separately, but not audio with transcriptions. Using text, speech can be synthetically produced via text-to-speech (TTS) systems. However, many low-resource languages do not have quality TTS systems either. We propose an alternative: produce synthetic audio by running text from the target language through a trained TTS system for a higher-resource pivot language. We investigate when and how this technique is most effective in low-resource settings. In our experiments, using several thousand synthetic TTS text-speech pairs and duplicating authentic data to balance yields optimal results. Our findings suggest that searching over a set of candidate pivot languages can lead to marginal improvements and that, surprisingly, ASR performance can by harmed by increases in measured TTS quality. Application of these findings improves ASR by 64.5\% and 45.0\% character error reduction rate (CERR) respectively for two low-resource languages: Guaran\'i and Suba.
【8】 Improving Data Driven Inverse Text Normalization using Data Augmentation
标题:利用数据扩充改进数据驱动的反向文本规范化
链接:https://arxiv.org/abs/2207.09674
* 与cs.SD语音【4】为同一篇
作者:Laxmi Pandey,Debjyoti Paul,Pooja Chitkara,Yutong Pang,Xuedong Zhang,Kjell Schubert,Mark Chou,Shu Liu,Yatharth Saraf机构:University of California Merced†, Meta Inc.摘要:反向文本规范化(ITN)用于转换自动语音识别的口语形式输出传统的手工ITN规则在转录和维护方面可能很复杂。同时,神经建模方法需要与ASR系统在相同或相似领域的高质量大规模口语—书面语对示例这两种方法都需要昂贵而复杂的标注,我们提出了一种数据扩充技术,该技术利用最少的人工注释从域外文本数据有效地生成丰富的口头—书面数字对。我们经验性地证明,使用我们的数据扩充技术训练的ITN模型在所有数值表面(如基数、货币和分数)上始终优于仅使用域内数据训练的ITN模型,总体准确率为14.44%。摘要:Inverse text normalization (ITN) is used to convert the spoken form output of an automatic speech recognition (ASR) system to a written form. Traditional handcrafted ITN rules can be complex to transcribe and maintain. Meanwhile neural modeling approaches require quality large-scale spoken-written pair examples in the same or similar domain as the ASR system (in-domain data), to train. Both these approaches require costly and complex annotations. In this paper, we present a data augmentation technique that effectively generates rich spoken-written numeric pairs from out-of-domain textual data with minimal human annotation. We empirically demonstrate that ITN model trained using our data augmentation technique consistently outperform ITN model trained using only in-domain data across all numeric surfaces like cardinal, currency, and fraction, by an overall accuracy of 14.44%.
【9】 COVID-19 Detection from Respiratory Sounds with Hierarchical Spectrogram Transformers
标题:基于分级谱图Transformers的COVID-19呼吸音检测
链接:https://arxiv.org/abs/2207.09529
* 与cs.SD语音【5】为同一篇
作者:Idil Aytekin,Onat Dalmaz,Kaan Gonc,Haydar Ankishan,Emine U Saritas,Ulas Bagci,Haydar Celik,Tolga Cukur机构:and Tolga C¸ ukur∗, Senior Member摘要:对流行的空气传播疾病(如COVID-19)的监测特点是涉及呼吸评估。虽然听诊是症状监测的主流方法,但其诊断效用因需要专门的医院就诊而受到阻碍。基于便携式设备上呼吸音记录的持续远程监测是一种有前途的替代方法,可以帮助筛查COVID-19。在本研究中,我们引入了一种新的深度学习方法,通过咳嗽或呼吸声的音频记录来区分COVID-19患者和健康对照。所提出的方法利用了一种新的分层谱图Transformer HST在声谱图的局部窗口上体现了自我注意机制,在多个国家数据集上的综合验证表明,HST方法优于其他同类方法,在检测COVID-19病例时达到97%以上的受试者工作特征曲线下面积(AUC)。摘要:Monitoring of prevalent airborne diseases such as COVID-19 characteristically involve respiratory assessments. While auscultation is a mainstream method for symptomatic monitoring, its diagnostic utility is hampered by the need for dedicated hospital visits. Continual remote monitoring based on recordings of respiratory sounds on portable devices is a promising alternative, which can assist in screening of COVID-19. In this study, we introduce a novel deep learning approach to distinguish patients with COVID-19 from healthy controls given audio recordings of cough or breathing sounds. The proposed approach leverages a novel hierarchical spectrogram transformer (HST) on spectrogram representations of respiratory sounds. HST embodies self-attention mechanisms over local windows in spectrograms, and window size is progressively grown over model stages to capture local to global context. HST is compared against state-of-the-art conventional and deep-learning baselines. Comprehensive demonstrations on a multi-national dataset indicate that HST outperforms competing methods, achieving over 97% area under the receiver operating characteristic curve (AUC) in detecting COVID-19 cases.
机器翻译,仅供参考