今天跟大家分享一篇语音相关的论文合集:cs.SD语音39篇,eess.AS音频处理16篇。

cs.SD语音

【1】 Impact of Acoustic Event Tagging on Scene Classification in a Multi-Task  Learning Framework

标题:声事件标注对多任务学习框架中场景分类的影响

链接:https://arxiv.org/abs/2206.13476

作者:Rahil Parikh,Harshavardhan Sundar,Ming Sun,Chao Wang,Spyros Matsoukas
备注:Accepted at ISCA Interspeech 2022
摘要:声学事件是具有明确光谱-时间特征的声音,可以与产生它们的物理对象相关联。声学场景是这些声学事件的集合,没有特定的时间顺序。鉴于事件和场景之间的这种自然联系,人们普遍认为,对事件进行分类的能力必须有助于对场景进行分类。这导致了一些努力,试图利用多任务网络在声学事件标记(AET)和声学场景分类(ASC)方面做得很好。然而,在这些努力中,一项任务的改进并不保证另一项任务的改进,这表明ASC和AET之间存在紧张关系。目前尚不清楚AET的改善是否转化为ASC的改善。我们通过广泛的实证研究探讨了这一难题,并表明在一定条件下,在多任务网络中使用AET作为辅助任务可以持续提高ASC的性能。此外,ASC性能随着AET数据集的大小而进一步提高,并且对AET数据集中事件的选择或事件的数量不敏感。我们的结论是,ASC性能的改善来自于使用AET的正则化效果,而不是来自网络对声学事件的识别能力的提高。
摘要:Acoustic events are sounds with well-defined spectro-temporal characteristics which can be associated with the physical objects generating them. Acoustic scenes are collections of such acoustic events in no specific temporal order. Given this natural linkage between events and scenes, a common belief is that the ability to classify events must help in the classification of scenes. This has led to several efforts attempting to do well on Acoustic Event Tagging (AET) and Acoustic Scene Classification (ASC) using a multi-task network. However, in these efforts, improvement in one task does not guarantee an improvement in the other, suggesting a tension between ASC and AET. It is unclear if improvements in AET translates to improvements in ASC. We explore this conundrum through an extensive empirical study and show that under certain conditions, using AET as an auxiliary task in the multi-task network consistently improves ASC performance. Additionally, ASC performance further improves with the AET data-set size and is not sensitive to the choice of events or the number of events in the AET data-set. We conclude that this improvement in ASC performance comes from the regularization effect of using AET and not from the network's improved ability to discern between acoustic events.


【2】 Is the Language Familiarity Effect gradual? A computational modelling  approach

标题:语言熟悉度的影响是渐进的吗?一种计算建模方法

链接:https://arxiv.org/abs/2206.13415

作者:Maureen de Seyssel,Guillaume Wisniewski,Emmanuel Dupoux
备注:8 pages, 2 figures, accepted at CogSci 2022
摘要:根据语言熟悉度效应(LFE),人们更善于区分母语使用者。尽管文献中对这种认知效应进行了大量研究,但实验仅在有限数量的语言对上进行,其结果仅显示了这种效应的存在,而没有产生可能因语言对而异的渐进测量。在这项工作中,我们表明Thorburn、Feldmand和Schatz(2019)提出的LFE计算模型可以解决这两个限制。在第一个实验中,我们证明了该模型通过复制母语和口音语音的行为发现,能够逐步测量LFE。在第二个实验中,我们评估了大量语言对的LFE,包括许多从未在人类身上测试过的语言对。我们表明,这种效应在广泛的语言中得到了复制,进一步证明了其普遍性。基于LFE的逐步测量,我们还表明,属于同一家族的语言得分较小,支持语言距离对LFE的影响。
摘要:According to the Language Familiarity Effect (LFE), people are better at discriminating between speakers of their native language. Although this cognitive effect was largely studied in the literature, experiments have only been conducted on a limited number of language pairs and their results only show the presence of the effect without yielding a gradual measure that may vary across language pairs. In this work, we show that the computational model of LFE introduced by Thorburn, Feldmand and Schatz (2019) can address these two limitations. In a first experiment, we attest to this model's capacity to obtain a gradual measure of the LFE by replicating behavioural findings on native and accented speech. In a second experiment, we evaluate LFE on a large number of language pairs, including many which have never been tested on humans. We show that the effect is replicated across a wide array of languages, providing further evidence of its universality. Building on the gradual measure of LFE, we also show that languages belonging to the same family yield smaller scores, supporting the idea of an effect of language distance on LFE.


【3】 A Comprehensive Survey on Video Saliency Detection with Auditory  Information: the Audio-visual Consistency Perceptual is the Key!

标题:听觉信息视频显著检测综述:视听一致性感知是关键!

链接:https://arxiv.org/abs/2206.13390

作者:Chenglizhao Chen,Mengke Song,Wenfeng Song,Li Guo,Muwei Jian
摘要:视频显著性检测(VSD)旨在快速定位给定视频片段中最具吸引力的对象/事物/模式。现有的VSD相关作品主要依赖视觉系统,而对音频方面的关注较少,而实际上,我们的音频系统是视觉系统最重要的补充部分。此外,视听显著性检测(AVSD)是模仿人类感知机制的最具代表性的研究课题之一,目前还处于起步阶段,现有的调查论文都没有涉及到它,尤其是从显著性检测的角度。因此,本文的最终目的是对视听融合和显著性检测之间的差距进行广泛的回顾。此外,作为本综述的另一个亮点,我们深入了解了可能直接决定AVSD深度模型性能的关键因素,并声称视听一致性程度(AVC)——一个长期被忽视的问题,可以直接影响在执行显著性检测时使用音频使其视觉对应方受益的有效性。此外,为了使AVC问题对未来的关注者更加实用和有价值,我们新为几乎所有现有的公开可用AVSD数据集配备了额外的帧级AVC标签。基于这些升级的数据集,我们进行了广泛的定量评估,以证明AVC在AVSD任务中的重要性。总之,我们的想法和新设置都是一个方便的平台,提供了初步的准备和指导,所有这些都有助于推动未来的工作,进一步促进最先进的(SOTA)性能。
摘要:Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect, while, actually, our audio system is the most vital complementary part to our visual system. Also, audio-visual saliency detection (AVSD), one of the most representative research topics for mimicking human perceptual mechanisms, is currently in its infancy, and none of the existing survey papers have touched on it, especially from the perspective of saliency detection. Thus, the ultimate goal of this paper is to provide an extensive review to bridge the gap between audio-visual fusion and saliency detection. In addition, as another highlight of this review, we have provided a deep insight into key factors which could directly determine the performances of AVSD deep models, and we claim that the audio-visual consistency degree (AVC) -- a long-overlooked issue, can directly influence the effectiveness of using audio to benefit its visual counterpart when performing saliency detection. Moreover, in order to make the AVC issue more practical and valuable for future followers, we have newly equipped almost all existing publicly available AVSD datasets with additional frame-wise AVC labels. Based on these upgraded datasets, we have conducted extensive quantitative evaluations to ground our claim on the importance of AVC in the AVSD task. In a word, both our ideas and new sets serve as a convenient platform with preliminaries and guidelines, all of which are very potential to facilitate future works in promoting state-of-the-art (SOTA) performance further.


【4】 A two-stage full-band speech enhancement model with effective spectral  compression mapping

标题:一种有效谱压缩映射的两级全频带语音增强模型

链接:https://arxiv.org/abs/2206.13136

作者:Zhongshu Hou,Qinwen Hu,Kai Chen,Jing Lu
摘要:基于深度神经网络(DNN)的宽带语音增强(SE)直接扩展到全频段处理面临着低频分辨率的挑战,这很可能导致模型性能的恶化。在本文中,我们提出了一种可学习的频谱压缩映射(SCM)来有效地压缩高频分量,以便以更有效的方式对其进行处理。通过这样做,该模型可以更加关注低频和中频范围,其中大部分语音功率都集中在低频和中频范围。我们不是在单个网络结构中抑制噪声,而是首先估计频谱幅度掩码,将语音转换为高信噪比(SNR)状态,然后利用后续模型进一步优化预增强信号的实掩码和虚掩码。我们进行了综合实验来验证所提方法的有效性。
摘要:The direct expansion of deep neural network (DNN) based wide-band speech enhancement (SE) to full-band processing faces the challenge of low frequency resolution in low frequency range, which would highly likely lead to deteriorated performance of the model. In this paper, we propose a learnable spectral compression mapping (SCM) to effectively compress the high frequency components so that they can be processed in a more efficient manner. By doing so, the model can pay more attention to low and middle frequency range, where most of the speech power is concentrated. Instead of suppressing noise in a single network structure, we first estimate a spectral magnitude mask, converting the speech to a high signal-to-ratio (SNR) state, and then utilize a subsequent model to further optimize the real and imaginary mask of the pre-enhanced signal. We conduct comprehensive experiments to validate the efficacy of the proposed method.


【5】 TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a  Speech Recognition Baseline

标题:TALCS:一个开源的汉英代码转换语料库和语音识别基线

链接:https://arxiv.org/abs/2206.13135

作者:Chengfei Li,Shuhao Deng,Yaoping Wang,Guangjing Wang,Yaguang Gong,Changbin Chen,Jinfeng Bai备注:accepted by INTERSPEECH 2022
摘要:本文介绍了一种新的汉英码转换语音识别语料库——TALCS语料库,适用于码转换语音识别系统的训练和评价。TALCS语料库来源于好未来教育集团真实的在线一对一英语教学场景,包含大约587小时的16千赫语音样本。据我们所知,TALCS语料库是世界上最大的标记良好的汉英代码转换开源自动语音识别(ASR)数据集。本文将详细介绍录音过程,包括音频采集设备和语料库环境。滑石语料库在许可证下可免费下载1。使用TALCS语料库,我们在两个流行的语音识别工具包中进行ASR实验,以构建一个基线系统,包括ESPnet和Wenet。在TALCS语料库中比较了两种语音识别工具包的混合错误率(MER)性能。实验结果表明,录音和转录的质量是有希望的,基线系统是可行的。
摘要:This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one English teaching scenes in TAL education group, which contains roughly 587 hours of speech sampled at 16 kHz. To our best knowledge, TALCS corpus is the largest well labeled Mandarin-English code-switching open source automatic speech recognition (ASR) dataset in the world. In this paper, we will introduce the recording procedure in detail, including audio capturing devices and corpus environments. And the TALCS corpus is freely available for download under the permissive license1. Using TALCS corpus, we conduct ASR experiments in two popular speech recognition toolkits to make a baseline system, including ESPnet and Wenet. The Mixture Error Rate (MER) performance in the two speech recognition toolkits is compared in TALCS corpus. The experimental results implies that the quality of audio recordings and transcriptions are promising and the baseline system is workable.


【6】 Sequence-level Speaker Change Detection with Difference-based Continuous  Integrate-and-fire

标题:基于差分连续积分点火的序列级说话人变化检测

链接:https://arxiv.org/abs/2206.13110

作者:Zhiyun Fan,Linhao Dong,Meng Cai,Zejun Ma,Bo Xu
备注:Signal Processing Letters 2022
摘要:说话人变化检测是会议和会话等多方交互中的一项重要任务。本文从序列转导的角度研究说话人变化检测任务。具体来说,我们提出了一种新的编码器-解码器框架,直接将输入特征序列转换为说话人身份序列。设计了基于差异的持续集成和触发机制来支持该框架。它通过逐帧积分编码器输出之间的扬声器差异来检测扬声器变化,并根据检测到的扬声器变化将编码器输出传输到分段级扬声器嵌入。整个框架由说话人身份序列监控,这是一个比精确的说话人变化点弱的标签。在AMI和DIHARD-I语料库上的实验表明,我们的序列级方法始终优于使用精确说话人变化标签的强框架级基线。
摘要:Speaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we propose a novel encoder-decoder framework that directly converts the input feature sequence to the speaker identity sequence. The difference-based continuous integrate-and-fire mechanism is designed to support this framework. It detects speaker changes by integrating the speaker difference between the encoder outputs frame-by-frame and transfers encoder outputs to segment-level speaker embeddings according to the detected speaker changes. The whole framework is supervised by the speaker identity sequence, a weaker label than the precise speaker change points. The experiments on the AMI and DIHARD-I corpora show that our sequence-level method consistently outperforms a strong frame-level baseline that uses the precise speaker change labels.


【7】 SpeechEQ: Speech Emotion Recognition based on Multi-scale Unified  Datasets and Multitask Learning

标题:SpeechEQ:基于多尺度统一数据集和多任务学习的语音情感识别

链接:https://arxiv.org/abs/2206.13101

作者:Zuheng Kang,Junqing Peng,Jianzong Wang,Jing Xiao
摘要:语音情感识别(SER)面临许多挑战,但其中一个主要挑战是每个框架都没有统一的标准。在本文中,我们提出了SpeechEQ,这是一个基于多尺度统一度量的统一SER任务的框架。该指标可以通过多任务学习(MTL)进行训练,MTL包括情绪状态类别(EIS)和情绪强度量表(EIS)两个情绪识别任务,以及音位识别和性别识别两个辅助任务。对于这个框架,我们构建了一个汉语SER数据集——SpeechEQ数据集(SEQD)。我们在普通话的公共CASIA和ESD数据集上进行了实验,结果表明,我们的方法比基线方法有较大幅度的提高,准确率分别提高了8.0%和6.5%。在IEMOCAP上对四种情绪类别(即愤怒、快乐、悲伤和中性)的额外实验也表明,所提出的方法实现了78.16%的加权准确率(WA)和77.47%的未加权准确率(UA)的最新水平。
摘要:Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a framework for unifying SER tasks based on a multi-scale unified metric. This metric can be trained by Multitask Learning (MTL), which includes two emotion recognition tasks of Emotion States Category (EIS) and Emotion Intensity Scale (EIS), and two auxiliary tasks of phoneme recognition and gender recognition. For this framework, we build a Mandarin SER dataset - SpeechEQ Dataset (SEQD). We conducted experiments on the public CASIA and ESD datasets in Mandarin, which exhibit that our method outperforms baseline methods by a relatively large margin, yielding 8.0\% and 6.5\% improvement in accuracy respectively. Additional experiments on IEMOCAP with four emotion categories (i.e., angry, happy, sad, and neutral) also show the proposed method achieves a state-of-the-art of both weighted accuracy (WA) of 78.16% and unweighted accuracy (UA) of 77.47%.


【8】 Sound Model Factory: An Integrated System Architecture for Generative  Audio Modelling

标题:声音模型工厂:一个用于生成性音频建模的集成系统架构

链接:https://arxiv.org/abs/2206.13085

作者:Lonce Wyse,Purnima Kamath,Chitralekha Gupta
备注:None
摘要:我们介绍了一种新的数据驱动音频模型设计系统,该系统围绕两种不同的神经网络架构构建,一种是生成对抗网络(GAN)和一种递归神经网络(RNN),它利用每种网络的独特特性来实现两种网络都无法单独解决的系统目标。该系统的目标是生成交互式可控的声音模型,前提是(a)模型应能够合成一系列声音,以及(b)用于导航该声音空间的参数控制规范。声音范围由设计者提供的数据集定义,而导航方式则由数据标签的组合和从GAN学习的潜在空间中选择子流形来定义。我们提出的系统利用了GAN丰富的潜在空间,该空间由声音组成,填满了“之间”的空间“真实数据就像声音。然后,使用来自GAN的这些增强数据来训练RNN,使其能够立即、连续地响应参数变化,并在无限的时间内生成音频。此外,我们开发了一种自组织映射技术,用于“平滑”GAN的潜在空间,从而在音频音色之间进行感知平滑的插值。我们通过用户研究来验证这个过程。该系统有助于提高生成性声音模型设计的技术水平,包括系统配置和组件,以改进插值,并将音频建模能力从音高和打击乐器声音扩展到更复杂的音频纹理空间。
摘要:We introduce a new system for data-driven audio sound model design built around two different neural network architectures, a Generative Adversarial Network(GAN) and a Recurrent Neural Network (RNN), that takes advantage of the unique characteristics of each to achieve the system objectives that neither is capable of addressing alone. The objective of the system is to generate interactively controllable sound models given (a) a range of sounds the model should be able to synthesize, and (b) a specification of the parametric controls for navigating that space of sounds. The range of sounds is defined by a dataset provided by the designer, while the means of navigation is defined by a combination of data labels and the selection of a sub-manifold from the latent space learned by the GAN. Our proposed system takes advantage of the rich latent space of a GAN that consists of sounds that fill out the spaces ''between" real data-like sounds. This augmented data from the GAN is then used to train an RNN for its ability to respond immediately and continuously to parameter changes and to generate audio over unlimited periods of time. Furthermore, we develop a self-organizing map technique for ``smoothing" the latent space of GAN that results in perceptually smooth interpolation between audio timbres. We validate this process through user studies. The system contributes advances to the state of the art for generative sound model design that include system configuration and components for improving interpolation and the expansion of audio modeling capabilities beyond musical pitch and percussive instrument sounds into the more complex space of audio textures.


【9】 Uncertainty Calibration for Deep Audio Classifiers

标题:深度音频分类器的不确定度校正

链接:https://arxiv.org/abs/2206.13071

作者:Tong Ye,Shijing Si,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by InterSpeech 2022, the first two authors contributed equally
摘要:虽然深度神经网络(DNN)在音频分类任务中取得了巨大的成功,但其不确定性校准仍处于探索阶段。一个经过良好校准的模型在其预测确定时应该是准确的,而在可能不准确时则表明其具有很高的不确定性。在这项工作中,我们研究了深度音频分类器的不确定性校准。特别是,我们实证研究了流行的校准方法在音频分类数据集上的性能:(i)蒙特卡罗衰减,(ii)系综,(iii)焦损,(iv)谱归一化高斯过程(SNGP)。为此,我们评估(i-iv)环境声音和音乐流派分类的任务。结果表明,未经校准的深度音频分类器可能过于自信,而SNGP在本文的两个数据集上表现最好,效率非常高。
摘要:Although deep Neural Networks (DNNs) have achieved tremendous success in audio classification tasks, their uncertainty calibration are still under-explored. A well-calibrated model should be accurate when it is certain about its prediction and indicate high uncertainty when it is likely to be inaccurate. In this work, we investigate the uncertainty calibration for deep audio classifiers. In particular, we empirically study the performance of popular calibration methods: (i) Monte Carlo Dropout, (ii) ensemble, (iii) focal loss, and (iv) spectral-normalized Gaussian process (SNGP), on audio classification datasets. To this end, we evaluate (i-iv) for the tasks of environment sound and music genre classification. Results indicate that uncalibrated deep audio classifiers may be over-confident, and SNGP performs the best and is very efficient on the two datasets of this paper.


【10】 Speak Like a Professional: Increasing Speech Intelligibility by  Mimicking Professional Announcer Voice with Voice Conversion

标题:像专业人士一样说话:通过语音转换模仿专业播音员的声音来提高语音清晰度

链接:https://arxiv.org/abs/2206.13021

作者:Tuan Vu Ho,Maori Kobayashi,Masato Akagi
备注:Accepted at INTERSPEECH 2022
摘要:在大多数实际场景中,广播系统必须在噪声环境中传递语音信息,在噪声环境中背景噪声无法消除。局部噪声降低了语音清晰度,增加了听者的听力,从而影响了广播系统的有效性。据报道,在嘈杂的环境中,专业播音员的声音比非专业播音员的声音更清晰、更全面。这一发现表明,语音清晰度可能与专业播音员的说话风格有关,可以使用语音转换方法进行调整。基于这一思想,本文将语音转换方法应用于非专业语音,提出了一种在噪声环境下提高语音清晰度的方法。我们发现,在说话人嵌入平面上,专业播音员和非专业说话人被分为不同的簇。这意味着语音清晰度可以作为说话人个性的一个独立特征来控制。为了检验转换语音在噪声环境中的优势,我们在不同信噪比下使用粉红色噪声中隐藏的测试词进行了实验。客观和主观评价结果证实,在低信噪比条件下,转换语音的语音清晰度高于原始语音。
摘要:In most of practical scenarios, the announcement system must deliver speech messages in a noisy environment, in which the background noise cannot be cancelled out. The local noise reduces speech intelligibility and increases listening effort of the listener, hence hamper the effectiveness of announcement system. There has been reported that voices of professional announcers are clearer and more comprehensive than that of non-expert speakers in noisy environment. This finding suggests that the speech intelligibility might be related to the speaking style of professional announcer, which can be adapted using voice conversion method. Motivated by this idea, this paper proposes a speech intelligibility enhancement in noisy environment by applying voice conversion method on non-professional voice. We discovered that the professional announcers and non-professional speakers are clusterized into different clusters on the speaker embedding plane. This implies that the speech intelligibility can be controlled as an independent feature of speaker individuality. To examine the advantage of converted voice in noisy environment, we experimented using test words masked in pink noise at different SNR levels. The results of objective and subjective evaluations confirm that the speech intelligibility of converted voice is higher than that of original voice in low SNR conditions.


【11】 Annotated Speech Corpus for Low Resource Indian Languages: Awadhi,  Bhojpuri, Braj and Magahi

标题:用于低资源印度语言的标注语音语料库:Awadhi、Bhojpui、Braj和Magahi

链接:https://arxiv.org/abs/2206.12931

作者:Ritesh Kumar,Siddharth Singh,Shyam Ratan,Mohit Raj,Sonal Sinha,bornini lahiri,Vivek Seshadri,Kalika Bali,Atul Kr. Ojha
备注:Speech for Social Good Workshop, 2022, Interspeech 2022
摘要:在本文中,我们讨论了一项正在进行的工作,即使用语言数据收集的现场方法,为四种低资源的印度雅利安语言——阿瓦迪语、博伊普里语、布拉吉语和马加希语开发语音语料库。目前,语料库的总规模约为18小时(每种语言约4-5小时),并通过词性标记、形态特征和普遍依赖关系等语法信息进行转录和注释。我们讨论了用这些语言收集数据的方法,其中大部分是在2019冠状病毒疾病大流行期间完成的,目的之一是为使用这些语言的低收入群体创造一些额外收入。本文还讨论了这些语言的自动语音识别系统的基线实验结果。
摘要:In this paper we discuss an in-progress work on the development of a speech corpus for four low-resource Indo-Aryan languages -- Awadhi, Bhojpuri, Braj and Magahi using the field methods of linguistic data collection. The total size of the corpus currently stands at approximately 18 hours (approx. 4-5 hours each language) and it is transcribed and annotated with grammatical information such as part-of-speech tags, morphological features and Universal dependency relationships. We discuss our methodology for data collection in these languages, most of which was done in the middle of the COVID-19 pandemic, with one of the aims being to generate some additional income for low-income groups speaking these languages. In the paper, we also discuss the results of the baseline experiments for automatic speech recognition system in these languages.


【12】 Data Augmentation for Dementia Detection in Spoken Language

标题:用于口语中痴呆检测的数据增强

链接:https://arxiv.org/abs/2206.12879

作者:Anna Hlédiková,Dominika Woszczyk,Alican Acman,Soteris Demetriou,Björn Schuller
备注:Accepted to INTERSPEECH 2022
摘要:随着我们社会的老龄化,痴呆症是一个日益严重的问题,检测方法往往是侵入性的,而且价格昂贵。最近的深度学习技术可以提供更快的诊断,并已显示出有希望的结果。然而,它们需要大量的标记数据,而这些数据不容易用于痴呆症检测任务。稀疏数据问题的一个有效解决方案是数据扩充,但需要仔细选择精确的方法。迄今为止,还没有针对阿尔茨海默病(AD)数据集的NLP和语音处理进行数据扩充的实证研究。在这项工作中,我们研究了用于AD检测任务的数据增强技术,并在文本和音频域的两种模型上对不同的方法进行了实证评估。我们对这两个域使用基于Transformer的模型,对文本和音频域分别使用SVM和随机森林模型。我们使用传统的以及基于深度学习的方法生成额外的样本,并表明数据增强可以提高基于文本和基于音频的模型的性能,并且这些结果与流行地址集上的最新结果相当,具有精心制作的架构和功能。
摘要:Dementia is a growing problem as our society ages, and detection methods are often invasive and expensive. Recent deep-learning techniques can offer a faster diagnosis and have shown promising results. However, they require large amounts of labelled data which is not easily available for the task of dementia detection. One effective solution to sparse data problems is data augmentation, though the exact methods need to be selected carefully. To date, there has been no empirical study of data augmentation on Alzheimer's disease (AD) datasets for NLP and speech processing. In this work, we investigate data augmentation techniques for the task of AD detection and perform an empirical evaluation of the different approaches on two kinds of models for both the text and audio domains. We use a transformer-based model for both domains, and SVM and Random Forest models for the text and audio domains, respectively. We generate additional samples using traditional as well as deep learning based methods and show that data augmentation improves performance for both the text- and audio-based models and that such results are comparable to state-of-the-art results on the popular ADReSS set, with carefully crafted architectures and features.


【13】 On Comparison of Encoders for Attention based End to End Speech  Recognition in Standalone and Rescoring Mode

标题:基于注意力的端到端语音识别中独立模式和重新评分模式下编码器的比较

链接:https://arxiv.org/abs/2206.12829

作者:Raviraj Joshi,Subodh Kumar
备注:Accepted at SPCOM 2022
摘要:流式自动语音识别(ASR)模型更受欢迎,适用于基于语音的应用。然而,非流模型在查看整个音频上下文时提供了更好的性能。为了在流式应用程序(如语音搜索)中利用非流式模型的优点,它通常用于第二遍重新评分模式。使用流模型生成的候选假设使用非流模型重新评分。在这项工作中,我们评估了Flipkart语音搜索任务中基于非流式注意力的端到端ASR模型,包括独立模式和重新评分模式。这些模型基于Listen-Attend-Spell(LAS)编码器-解码器体系结构。我们对基于LSTM、Transformer和Conformer的不同编码器变体进行了实验。我们比较了这些模型的延迟要求及其性能。总的来说,我们表明Transformer模型提供了可接受的WER和最低的延迟要求。我们报告,在延迟开销低于5ms的情况下,通过第二次通过LAS重新评分,相对WER提高了约16%。我们还强调了CNN前端与转换器架构的重要性,以实现可比的字错误率(WER)。此外,我们观察到,在第二遍重新评分模式中,所有编码器都提供了类似的好处,而在独立文本生成模式中,性能上的差异是显著的。
摘要:The streaming automatic speech recognition (ASR) models are more popular and suitable for voice-based applications. However, non-streaming models provide better performance as they look at the entire audio context. To leverage the benefits of the non-streaming model in streaming applications like voice search, it is commonly used in second pass re-scoring mode. The candidate hypothesis generated using steaming models is re-scored using a non-streaming model. In this work, we evaluate the non-streaming attention-based end-to-end ASR models on the Flipkart voice search task in both standalone and re-scoring modes. These models are based on Listen-Attend-Spell (LAS) encoder-decoder architecture. We experiment with different encoder variations based on LSTM, Transformer, and Conformer. We compare the latency requirements of these models along with their performance. Overall we show that the Transformer model offers acceptable WER with the lowest latency requirements. We report a relative WER improvement of around 16% with the second pass LAS re-scoring with latency overhead under 5ms. We also highlight the importance of CNN front-end with Transformer architecture to achieve comparable word error rates (WER). Moreover, we observe that in the second pass re-scoring mode all the encoders provide similar benefits whereas the difference in performance is prominent in standalone text generation mode.


【14】 Exploiting Transformation Invariance and Equivariance for  Self-supervised Sound Localisation

标题:利用变换不变性和均方差进行自监督声源定位

链接:https://arxiv.org/abs/2206.12772

作者:Jinxiang Liu,Chen Ju,Weidi Xie,Ya Zhang
备注:10 pages,
摘要:我们提出了一个简单而有效的视听表征学习的自监督框架,以定位视频中的声源。为了理解什么使我们能够学习有用的表示,我们系统地研究了数据扩充的效果,并揭示了(1)数据扩充的组成起着关键作用,{\em即}~明确鼓励视听表示对各种变换保持不变{\em变换不变性});(2) 强制几何一致性大大提高了学习表示的质量,{\em即}~检测到的声源应遵循应用于输入视频帧的相同变换({\em变换等变})。大量实验表明,我们的模型在两个声音定位基准(即Flickr SoundNet和VGG sound)上的性能明显优于以前的方法。此外,我们还评估了音频检索和跨模态检索任务。在这两种情况下,我们的自监督模型都表现出了优异的检索性能,甚至可以与有监督的音频检索方法相媲美。这表明所提出的框架学习了强多模态表示,这有利于声音的本地化和推广,以进一步应用。\textit{所有代码都可用}。
摘要:We present a simple yet effective self-supervised framework for audio-visual representation learning, to localize the sound source in videos. To understand what enables to learn useful representations, we systematically investigate the effects of data augmentations, and reveal that (1) composition of data augmentations plays a critical role, {\em i.e.}~explicitly encouraging the audio-visual representations to be invariant to various transformations~({\em transformation invariance}); (2) enforcing geometric consistency substantially improves the quality of learned representations, {\em i.e.}~the detected sound source should follow the same transformation applied on input video frames~({\em transformation equivariance}). Extensive experiments demonstrate that our model significantly outperforms previous methods on two sound localization benchmarks, namely, Flickr-SoundNet and VGG-Sound. Additionally, we also evaluate audio retrieval and cross-modal retrieval tasks. In both cases, our self-supervised models demonstrate superior retrieval performances, even competitive with the supervised approach in audio retrieval. This reveals the proposed framework learns strong multi-modal representations that are beneficial to sound localisation and generalization to further applications. \textit{All codes will be available}.


【15】 Low-resource Accent Classification in Geographically-proximate Settings:  A Forensic and Sociophonetics Perspective

标题:地理邻近环境下的低资源口音分类:法医学和社会语音学视角

链接:https://arxiv.org/abs/2206.12759

作者:Qingcheng Zeng,Dading Chong,Peilin Zhou,Jie Yang
备注:INTERSPEECH 2022
摘要:重音语音识别和重音分类是语音技术中研究相对较少的领域。最近,基于深度学习的方法和基于Transformer的预训练模型在这两个领域都取得了优异的性能。然而,大多数口音分类任务侧重于对不同种类的英语口音进行分类,而对地理位置相近的口音分类关注甚少,尤其是在资源较少的情况下,法医言语科学任务通常会遇到这种情况。在这篇论文中,我们探索了三种主要的口音建模方法,并结合两种不同的分类器,基于从英格兰北部五个城市品种中检索到的105个说话人录音。虽然预训练模型生成的语音表示通常在下游分类中具有更好的性能,但传统的方法如Mel倒谱系数(MFCC)和共振峰测量具有特定的优势。这些结果表明,在数据相对稀缺的法医语音学场景中,一种简单的建模方法和分类器可以与最先进的预训练语音模型作为特征提取器相竞争,这可以在实践中提高对重音信息的更快估计。此外,我们的发现也交叉验证了一种新的量化社会语音变化的方法。
摘要:Accented speech recognition and accent classification are relatively under-explored research areas in speech technology. Recently, deep learning-based methods and Transformer-based pretrained models have achieved superb performances in both areas. However, most accent classification tasks focused on classifying different kinds of English accents and little attention was paid to geographically-proximate accent classification, especially under a low-resource setting where forensic speech science tasks usually encounter. In this paper, we explored three main accent modelling methods combined with two different classifiers based on 105 speaker recordings retrieved from five urban varieties in Northern England. Although speech representations generated from pretrained models generally have better performances in downstream classification, traditional methods like Mel Frequency Cepstral Coefficients (MFCCs) and formant measurements are equipped with specific strengths. These results suggest that in forensic phonetics scenario where data are relatively scarce, a simple modelling method and classifier could be competitive with state-of-the-art pretrained speech models as feature extractors, which could enhance a sooner estimation for the accent information in practices. Besides, our findings also cross-validated a new methodology in quantifying sociophonetic changes.


【16】 TEVR: Improving Speech Recognition by Token Entropy Variance Reduction

标题:TEVR:通过减少令牌熵方差来改善语音识别

链接:https://arxiv.org/abs/2206.12693

作者:Hajo Nils Krabbenhöft,Erhardt Barth
备注:10 pages including 2 pages appendix, 1 figure, 6 tables
摘要:本文提出了一种TEVR语音识别模型,该模型旨在最小化语言模型中标记熵w.r.t.的变化。这充分利用了这样一个事实,即如果语言模型无论如何都能可靠而准确地预测一个标记,那么声学模型就不需要准确地识别它。我们使用9亿个参数对德语ASR模型进行了训练,结果表明,在CommonVoice德语上,TEVR的字错误率非常有竞争力,为3.64%,其字错误率相对降低了16.89%,优于最佳报告结果。我们希望,向社区发布我们经过全面训练的语音识别管道,将在未来带来保护隐私的离线虚拟助理。
摘要:This paper presents TEVR, a speech recognition model designed to minimize the variation in token entropy w.r.t. to the language model. This takes advantage of the fact that if the language model will reliably and accurately predict a token anyway, then the acoustic model doesn't need to be accurate in recognizing it. We train German ASR models with 900 million parameters and show that on CommonVoice German, TEVR scores a very competitive 3.64% word error rate, which outperforms the best reported results by a relative 16.89% reduction in word error rate. We hope that releasing our fully trained speech recognition pipeline to the community will lead to privacy-preserving offline virtual assistants in the future.


【17】 Synthesizing Personalized Non-speech Vocalization from Discrete Speech  Representations

标题:从离散语音表示合成个性化非语音发声

链接:https://arxiv.org/abs/2206.12662

作者:Chin-Cheng Hsu
摘要:我们将非言语发声(NSV)建模作为一项文本到语音任务,并验证了其可行性。具体而言,我们评估了休BERT语音单元在NSV上的语音表现力,并验证了我们的模型能够控制说话人的音色,即使训练数据是说话人很少的镜头。此外,我们证实了记录条件的异质性是NSV建模的主要障碍。最后,我们讨论了对我们的方法的五个改进,以供将来的研究。我们的演示页面上提供了合成NSV的音频示例:https://resemble-ai.github.io/reLaugh.
摘要:We formulated non-speech vocalization (NSV) modeling as a text-to-speech task and verified its viability. Specifically, we evaluated the phonetic expressivity of HUBERT speech units on NSVs and verified our model's ability to control over speaker timbre even though the training data is speaker few-shot. In addition, we substantiated that the heterogeneity in recording conditions is the major obstacle for NSV modeling. Finally, we discussed five improvements over our method for future research. Audio samples of synthesized NSVs are available on our demo page: https://resemble-ai.github.io/reLaugh.


【18】 Distilling a Pretrained Language Model to a Multilingual ASR Model

标题:将预训练语言模型提取为多语言ASR模型

链接:https://arxiv.org/abs/2206.12638

作者:Kwanghee Choi,Hyung-Min Park
备注:Accepted to Interspeech 2022. Official implementation provided in this https URL
摘要:多语言语音数据往往存在语言分布的长尾性,导致性能下降。然而,多语言文本数据更容易获得,从而产生更有用的通用语言模型。因此,我们有动机将嵌入在训练有素的教师文本模型中的丰富知识提取到学生语音模型中。我们提出了一种新的方法,称为提取语言模型到语音模型(提取L2S),该方法将两种不同模态的潜在表示对齐。细微的差异通过收缩机制、最近邻插值和可学习的线性投影层进行处理。我们将蒸馏方法应用于多语言自动语音识别(ASR)任务,证明了该方法的有效性。我们提取了基于转换器的跨语言模型(InfoXLM),同时对每种语言的大规模多语言ASR模型(XLSR-wav2vec 2.0)进行了微调。我们在20种低资源语言的CommonVoice数据集上展示了我们的方法的优越性,这些语言的语音数据少于100小时。
摘要:Multilingual speech data often suffer from long-tailed language distribution, resulting in performance degradation. However, multilingual text data is much easier to obtain, yielding a more useful general language model. Hence, we are motivated to distill the rich knowledge embedded inside a well-trained teacher text model to the student speech model. We propose a novel method called the Distilling a Language model to a Speech model (Distill-L2S), which aligns the latent representations of two different modalities. The subtle differences are handled by the shrinking mechanism, nearest-neighbor interpolation, and a learnable linear projection layer. We demonstrate the effectiveness of our distillation method by applying it to the multilingual automatic speech recognition (ASR) task. We distill the transformer-based cross-lingual language model (InfoXLM) while fine-tuning the large-scale multilingual ASR model (XLSR-wav2vec 2.0) for each language. We show the superiority of our method on 20 low-resource languages of the CommonVoice dataset with less than 100 hours of speech data.


【19】 Self-supervision and Learnable STRFs for Age, Emotion, and Country  Prediction

标题:用于年龄、情绪和国家预测的自我监督和可学习的STRF

链接:https://arxiv.org/abs/2206.12568

作者:Roshan Sharma,Tyler Vuong,Mark Lindsey,Hira Dhamyal,Rita Singh,Bhiksha Raj
备注:None
摘要:这项工作提出了一种多任务方法,用于同时估计2022年ICML表达性发声挑战赛ExVo多任务曲目的年龄、原产国和情感给定的人声突发音频。选择的方法结合了光谱-时间调制和自我监督特征,然后是以多任务范式组织的编码器-解码器网络。我们通过检查独立的任务特定模型和联合模型来评估任务之间的互补性,并探索不同特征集的相对优势。我们还引入了一种简单的分数融合机制,以利用此任务中不同特征集的互补性。我们发现,在光谱-时间感受野模型和HuBERT模型上,稳健的数据预处理结合分数融合,获得了我们的最佳ExVo多任务测试分数0.412。
摘要:This work presents a multitask approach to the simultaneous estimation of age, country of origin, and emotion given vocal burst audio for the 2022 ICML Expressive Vocalizations Challenge ExVo-MultiTask track. The method of choice utilized a combination of spectro-temporal modulation and self-supervised features, followed by an encoder-decoder network organized in a multitask paradigm. We evaluate the complementarity between the tasks posed by examining independent task-specific and joint models, and explore the relative strengths of different feature sets. We also introduce a simple score fusion mechanism to leverage the complementarity of different feature sets for this task.  We find that robust data preprocessing in conjunction with score fusion over spectro-temporal receptive field and HuBERT models achieved our best ExVo-MultiTask test score of 0.412.


【20】 Generating Diverse Vocal Bursts with StyleGAN2 and MEL-Spectrograms

标题:用StyleGAN2和MEL谱图生成不同的发声

链接:https://arxiv.org/abs/2206.12563

作者:Marco Jiralerspong,Gauthier Gidel
备注:To be published at the ICML Expressive Vocalizations Workshop and Competition (ExVo Generate) held in conjunction with the 39th International Conference on Machine Learning
摘要:我们描述了我们在ICML表达性发声比赛中的生成性情感发声突发任务(ExVo Generate)的方法。我们在音频样本预处理版本的mel谱图上训练条件StyleGAN2架构。然后,将模型生成的mel光谱图反转回音频域。因此,从定性和定量的角度来看,我们生成的样本大大改善了比赛提供的所有情绪的基线。更准确地说,即使对于表现最差的情绪(awe),我们也获得了1.76的FAD,而基线为4.81(作为参考,awe的训练/验证集之间的FAD为0.776)。
摘要:We describe our approach for the generative emotional vocal burst task (ExVo Generate) of the ICML Expressive Vocalizations Competition. We train a conditional StyleGAN2 architecture on mel-spectrograms of preprocessed versions of the audio samples. The mel-spectrograms generated by the model are then inverted back to the audio domain. As a result, our generated samples substantially improve upon the baseline provided by the competition from a qualitative and quantitative perspective for all emotions. More precisely, even for our worst-performing emotion (awe), we obtain an FAD of 1.76 compared to the baseline of 4.81 (as a reference, the FAD between the train/validation sets for awe is 0.776).


【21】 Self-supervised Context-aware Style Representation for Expressive Speech  Synthesis

标题:用于表现性语音合成的自监督上下文感知风格表示

链接:https://arxiv.org/abs/2206.12559

作者:Yihan Wu,Xi Wang,Shaofei Zhang,Lei He,Ruihua Song,Jian-Yun Nie
备注:Accepted by Interspeech 2022
摘要:与有声读物合成一样,表达性语音合成对于风格表征学习和预测仍然具有挑战性。从参考音频中派生或从文本中预测样式标记需要大量的标记数据,这些数据的获取成本很高,很难准确定义和注释。在本文中,我们提出了一个新的框架,用于以自我监督的方式从大量的纯文本中进行学习风格表示。它利用情感词汇,使用对比学习和深度聚类。我们进一步将样式表示作为条件嵌入集成到多样式转换器TTS中。与通过预测在同一数据集上训练但带有人类注释的风格标记的多风格TTS相比,我们的方法通过对有声读物语音的域内和域外测试集的主观评估,取得了更好的结果。此外,在隐式语境感知风格表征下,合成音频在长段落中的情感转换更为自然。音频样本可在演示网站上获得。
摘要:Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is costly to acquire and difficult to define and annotate accurately. In this paper, we propose a novel framework for learning style representation from abundant plain text in a self-supervised manner. It leverages an emotion lexicon and uses contrastive learning and deep clustering. We further integrate the style representation as a conditioned embedding in a multi-style Transformer TTS. Comparing with multi-style TTS by predicting style tags trained on the same dataset but with human annotations, our method achieves improved results according to subjective evaluations on both in-domain and out-of-domain test sets in audiobook speech. Moreover, with implicit context-aware style representation, the emotion transition of synthesized audio in a long paragraph appears more natural. The audio samples are available on the demo web.


【22】 Domain Generalization with Relaxed Instance Frequency-wise Normalization  for Multi-device Acoustic Scene Classification

标题:基于松弛实例频域归一化的多设备声场分类

链接:https://arxiv.org/abs/2206.12513

作者:Byeonggeun Kim,Seunghan Yang,Jangho Kim,Hyunsin Park,Juntae Lee,Simyung Chang
备注:Proceedings of INTERSPEECH 2022
摘要:在图像处理中使用二维卷积神经网络(2D CNN)时,可以使用通道统计信息来处理域信息,实例归一化是获得域不变特征的一种很有前景的方法。然而,与图像处理不同,我们分析了音频特征中的域相关信息在频率统计中占主导地位,而不是在信道统计中。基于我们的分析,我们引入了放松的实例频率方向归一化(RFN):一种沿频率轴的即插即用显式归一化模块,可以消除音频特征中的实例特定域差异,同时缓解有用鉴别信息的不必要损失。从经验上看,与以前的声学场景分类领域泛化方法相比,简单地将RFN添加到网络中就可以显示出明显的优势,并提高了对多个音频设备的鲁棒性。特别是,所提出的RFN以明显的优势赢得了DCASE2021挑战任务1a《多设备低复杂度声场景分类》,RFN是我们技术报告的扩展工作。
摘要:While using two-dimensional convolutional neural networks (2D-CNNs) in image processing, it is possible to manipulate domain information using channel statistics, and instance normalization has been a promising way to get domain-invariant features. However, unlike image processing, we analyze that domain-relevant information in an audio feature is dominant in frequency statistics rather than channel statistics. Motivated by our analysis, we introduce Relaxed Instance Frequency-wise Normalization (RFN): a plug-and-play, explicit normalization module along the frequency axis which can eliminate instance-specific domain discrepancy in an audio feature while relaxing undesirable loss of useful discriminative information. Empirically, simply adding RFN to networks shows clear margins compared to previous domain generalization approaches on acoustic scene classification and yields improved robustness for multiple audio devices. Especially, the proposed RFN won the DCASE2021 challenge TASK1A, low-complexity acoustic scene classification with multiple devices, with a clear margin, and RFN is an extended work of our technical report.


【23】 Multitask vocal burst modeling with ResNets and pre-trained  paralinguistic Conformers

标题:基于ResNet和预先训练的副语言一致性的多任务突发发声建模

链接:https://arxiv.org/abs/2206.12494

作者:Josh Belanich,Krishna Somandepalli,Brian Eoff,Brendan Jou
备注:To be published in the ICML Expressive Vocalizations Workshop & Competition 2022 (this https URL)
摘要:本技术报告介绍了我们提交给ICML表达发声研讨会和竞赛多任务跟踪(ExVo multitask)时使用的建模方法。我们首先将不同大小的图像分类模型应用于人声爆发的mel谱图表示,这是声音事件检测文献中的标准。这些模型的结果显示,与基线系统相比,任务指标的调和平均值增加了21.24%,构成了我们团队对多任务跟踪的主要提交。然后,我们试图通过应用一个大型的预先训练的构象模型来描述多任务轨迹中的净空,该模型以前在语音情感识别和面具检测等副语言任务中取得了最新的结果。此外,我们还调查了情感表达子任务、原籍国和年龄预测之间的关系,发现表现最好的模型被训练为单任务模型,质疑问题是否真的从多任务环境中受益。
摘要:This technical report presents the modeling approaches used in our submission to the ICML Expressive Vocalizations Workshop & Competition multitask track (ExVo-MultiTask). We first applied image classification models of various sizes on mel-spectrogram representations of the vocal bursts, as is standard in sound event detection literature. Results from these models show an increase of 21.24% over the baseline system with respect to the harmonic mean of the task metrics, and comprise our team's main submission to the MultiTask track. We then sought to characterize the headroom in the MultiTask track by applying a large pre-trained Conformer model that previously achieved state-of-the-art results on paralinguistic tasks like speech emotion recognition and mask detection. We additionally investigated the relationship between the sub-tasks of emotional expression, country of origin, and age prediction, and discovered that the best performing models are trained as single-task models, questioning whether the problem truly benefits from a multitask setting.


【24】 A Novel Approach For Analysis of Distributed Acoustic Sensing System  Based on Deep Transfer Learning

标题:一种基于深度迁移学习的分布式声学传感系统分析新方法

链接:https://arxiv.org/abs/2206.12484

作者:Ceyhun Efe Kayan,Kivilcim Yuksel Aldogan,Abdurrahman Gumus
摘要:分布式声传感器(DAS)是一种有效的仪器,广泛应用于许多应用领域,用于沿光纤以极高的空间分辨率记录各种事件的信号。为了正确地检测和识别记录的事件,具有高计算要求的高级信号处理算法至关重要。卷积神经网络是提取空间信息的高效工具,非常适合于DAS中的事件识别应用。长短时记忆(LSTM)是处理连续数据的有效工具。在本研究中,我们提出了一种多输入多输出两阶段特征提取方法,该方法将这些神经网络结构的功能与传递学习相结合,以对压电传感器施加到光纤上的振动进行分类。首先,我们从相位OTDR记录中提取差分振幅和相位信息,并将其存储在时空数据矩阵中。然后,在第一阶段,我们使用了最先进的预训练CNN(无密集层)作为特征抽取器。在第二阶段,我们使用LSTMs进一步分析CNN提取的特征。最后,我们使用密集层对提取的特征进行分类。为了观察所使用的CNN架构的效果,我们使用五种最先进的预训练模型(VGG-16、ResNet-50、DenseNet-121、MobileNet和Inception-v3)测试了我们的模型。结果表明,在我们的框架中使用VGG-16体系结构能够在50次训练中获得100%的分类准确率,并且在我们的Phase OTDR数据集上获得了最好的结果。这项研究的结果表明,预先训练的CNN与LSTM相结合,非常适合于分析差异幅度和相位信息,以时空数据矩阵表示,这对于DAS应用中的事件识别操作很有希望。
摘要:Distributed acoustic sensors (DAS) are effective apparatus which are widely used in many application areas for recording signals of various events with very high spatial resolution along the optical fiber. To detect and recognize the recorded events properly, advanced signal processing algorithms with high computational demands are crucial. Convolutional neural networks are highly capable tools for extracting spatial information and very suitable for event recognition applications in DAS. Long-short term memory (LSTM) is an effective instrument for processing sequential data. In this study, we proposed a multi-input multi-output, two stage feature extraction methodology that combines the capabilities of these neural network architectures with transfer learning to classify vibrations applied to an optical fiber by a piezo transducer. First, we extracted the differential amplitude and phase information from the Phase-OTDR recordings and stored them in a temporal-spatial data matrix. Then, we used a state-of-the-art pre-trained CNN without dense layers as a feature extractor in the first stage. In the second stage, we used LSTMs to further analyze the features extracted by the CNN. Finally, we used a dense layer to classify the extracted features. To observe the effect of the utilized CNN architecture, we tested our model with five state-of-the art pre-trained models (VGG-16, ResNet-50, DenseNet-121, MobileNet and Inception-v3). The results show that using the VGG-16 architecture in our framework manages to obtain 100% classification accuracy in 50 trainings and got the best results on our Phase-OTDR dataset. Outcomes of this study indicate that the pre-trained CNNs combined with LSTM are very suitable for the analysis of differential amplitude and phase information, represented in a temporal spatial data matrix which is promising for event recognition operations in DAS applications.


【25】 Burst2Vec: An Adversarial Multi-Task Approach for Predicting Emotion,  Age, and Origin from Vocal Bursts

标题:Burst2Vec:一种预测情绪、年龄和发声来源的对抗性多任务方法

链接:https://arxiv.org/abs/2206.12469

作者:Atijit Anuchitanukul,Lucia Specia
摘要:我们介绍了Burst2Vec,这是一种多任务学习方法,可以从发声爆发中预测情绪、年龄和来源(即母语/语言)。Burst2Vec利用预先训练的语音表示从原始波形中捕获声学信息,并通过对抗性训练引入模型消隐的概念。我们的模型使用预提取的特征比基线的性能提高了30%,在ICML ExVo 2022多任务挑战赛的所有参与者中得分最高。
摘要:We present Burst2Vec, our multi-task learning approach to predict emotion, age, and origin (i.e., native country/language) from vocal bursts. Burst2Vec utilises pre-trained speech representations to capture acoustic information from raw waveforms and incorporates the concept of model debiasing via adversarial training. Our models achieve a relative 30 % performance gain over baselines using pre-extracted features and score the highest amongst all participants in the ICML ExVo 2022 Multi-Task Challenge.


【26】 CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many  Fine-Grained Prosody Transfer

标题:CopyCat2:一种支持多说话人TTS和多对多细粒度韵律传输的单一模型

链接:https://arxiv.org/abs/2206.13443

作者:Sri Karlapati,Penny Karanasou,Mateusz Lajszczak,Ammar Abbas,Alexis Moinet,Peter Makarov,Ray Li,Arent van Korlaar,Simon Slangen,Thomas Drugman
备注:Accepted to be published in the Proceedings of InterSpeech 2022
摘要:在本文中,我们提出了CopyCat2(CC2),这是一种新的模型,能够:a)合成具有不同说话人身份的语音,b)生成具有表达和上下文适当韵律的语音,以及c)在任何一对可见说话人之间进行细粒度的韵律传递。我们通过为不同的任务激活网络的不同部分来实现这一点。我们使用一种新的两阶段训练方法来训练我们的模型。在第一阶段,该模型从语音中学习与说话人无关的词级韵律表示,用于多对多的细粒度韵律转换。在第二阶段,我们学习使用文本中可用的上下文信息预测这些韵律表示,从而使多说话人TTS具有上下文适当的韵律。我们将CC2与两个强基线进行了比较,一个是在具有上下文适当韵律的TTS中,另一个是在细粒度韵律转换中。CC2将基线语音和拷贝合成语音之间的自然度差距缩小了22.79美元\%$。在细粒度韵律迁移评估中,目标-说话人相似度相对提高了33.15 \%$。
摘要:In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually appropriate prosody, and c) transferring prosody at fine-grained level between any pair of seen speakers. We do this by activating distinct parts of the network for different tasks. We train our model using a novel approach to two-stage training. In Stage I, the model learns speaker-independent word-level prosody representations from speech which it uses for many-to-many fine-grained prosody transfer. In Stage II, we learn to predict these prosody representations using the contextual information available in text, thereby, enabling multi-speaker TTS with contextually appropriate prosody. We compare CC2 to two strong baselines, one in TTS with contextually appropriate prosody, and one in fine-grained prosody transfer. CC2 reduces the gap in naturalness between our baseline and copy-synthesised speech by $22.79\%$. In fine-grained prosody transfer evaluations, it obtains a relative improvement of $33.15\%$ in target speaker similarity.


【27】 Audio Similarity is Unreliable as a Proxy for Audio Quality

标题:音频相似性作为音频质量的替代指标并不可靠

链接:https://arxiv.org/abs/2206.13411

作者:Pranay Manocha,Zeyu Jin,Adam Finkelstein
备注:To Appear, Interspeech 2022
摘要:许多音频处理任务需要感知评估。然而,获取“金本位”人类判断的时间和费用限制了此类数据的可用性。大多数应用程序都包含完全引用或其他依赖于干净引用的基于相似性的度量(例如PESQ)。研究人员依靠这些指标来评估和比较各种提议的方法,往往得出结论,微小的、可测量的差异意味着一种方法比另一种方法更有效。本文展示了几个实际的场景,其中相似性度量与人类感知不一致,因为它们:(1)随着干净的引用而变化;(2) 依赖于人类在考虑质量时考虑的属性,(3)对难以察觉的信号水平差异敏感。在这些场景中,我们表明,无参考指标不存在此类缺点,并且与人类感知更好地关联。因此,我们得出结论,相似性作为音频质量的不可靠代理。
摘要:Many audio processing tasks require perceptual assessment. However, the time and expense of obtaining ``gold standard'' human judgments limit the availability of such data. Most applications incorporate full reference or other similarity-based metrics (e.g. PESQ) that depend on a clean reference. Researchers have relied on such metrics to evaluate and compare various proposed methods, often concluding that small, measured differences imply one is more effective than another. This paper demonstrates several practical scenarios where similarity metrics fail to agree with human perception, because they: (1) vary with clean references; (2) rely on attributes that humans factor out when considering quality, and (3) are sensitive to imperceptible signal level differences. In those scenarios, we show that no-reference metrics do not suffer from such shortcomings and correlate better with human perception. We conclude therefore that similarity serves as an unreliable proxy for audio quality.


【28】 Avocodo: Generative Adversarial Network for Artifact-free Vocoder

标题:Avocodo:无伪声编码器的生成性对抗性网络

链接:https://arxiv.org/abs/2206.13404

作者:Taejun Bak,Junmo Lee,Hanbin Bae,Jinhyeok Yang,Jae-Sung Bae,Young-Sun Joo
摘要:基于生成对抗性神经网络(GAN)的神经声码器以其快速的推理速度和轻量级的网络在产生链接:https://arxiv.org/abs/2206.13365高质量语音波形的同时得到了广泛的应用。由于感知上重要的语音成分主要集中在低频段,大多数基于GAN的神经声码器执行多尺度分析,以评估下采样语音波形。这种多尺度分析有助于生成器提高语音清晰度。然而,在初步实验中,我们观察到,聚焦于低频段的多尺度分析会导致意外的伪影,例如混叠和成像伪影,这些伪影会降低合成语音波形的质量。因此,在本文中,我们研究了这些伪影与基于GAN的神经声码器之间的关系,并提出了一种基于GAN的神经声码器,称为Avocodo,它允许合成具有减少伪影的高保真语音。我们介绍了两种用于从不同角度评估波形的鉴别器:协作多频带鉴别器和子频带鉴别器。我们还利用伪正交镜像滤波器组来获得降采样多波段波形,同时避免了混叠。实验结果表明,Avocodo在语音和歌唱语音合成任务上都优于传统的基于GAN的神经声码器,可以合成无伪影的语音。特别是,Avocodo甚至能够再现看不见的扬声器的高质量波形。
摘要:Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency band, most of the GAN-based neural vocoders perform multi-scale analysis that evaluates downsampled speech waveforms. This multi-scale analysis helps the generator improve speech intelligibility. However, in preliminary experiments, we observed that the multi-scale analysis which focuses on the low-frequency band causes unintended artifacts, e.g., aliasing and imaging artifacts, and these artifacts degrade the synthesized speech waveform quality. Therefore, in this paper, we investigate the relationship between these artifacts and GAN-based neural vocoders and propose a GAN-based neural vocoder, called Avocodo, that allows the synthesis of high-fidelity speech with reduced artifacts. We introduce two kinds of discriminators to evaluate waveforms in various perspectives: a collaborative multi-band discriminator and a sub-band discriminator. We also utilize a pseudo quadrature mirror filter bank to obtain downsampled multi-band waveforms while avoiding aliasing. The experimental results show that Avocodo outperforms conventional GAN-based neural vocoders in both speech and singing voice synthesis tasks and can synthesize artifact-free speech. Especially, Avocodo is even capable to reproduce high-quality waveforms of unseen speakers.


【29】 Interpretable Acoustic Representation Learning on Breathing and Speech  Signals for COVID-19 Detection

标题:用于新冠肺炎检测的呼吸和语音信号的可解释声学表示学习

链接:https://arxiv.org/abs/2206.13365

作者:Debottam Dutta,Debarpan Bhattacharya,Sriram Ganapathy,Amir H. Poorjam,Deepak Mittal,Maneesh Singh
摘要:在本文中,我们描述了一种用于2019冠状病毒疾病检测任务的音频信号表示学习方法。原始音频样本使用一组1-D卷积滤波器进行处理,这些滤波器被参数化为余弦调制高斯函数。选择这些核函数可以将滤波器组解释为平滑带通滤波器。过滤后的输出被合并、日志压缩并用于基于自我注意的相关性加权机制。相关性权重强调时频分解中对下游任务重要的关键区域。该模型的后续层由一个循环架构组成,并且该模型针对2019冠状病毒疾病检测任务进行训练。在Coswara数据集上的实验中,我们表明,与基线系统以及其他表示学习方法相比,该模型实现了显著的性能改进。此外,所提出的方法被证明统一适用于语音和呼吸信号以及从更大数据集进行的转移学习。
摘要:In this paper, we describe an approach for representation learning of audio signals for the task of COVID-19 detection. The raw audio samples are processed with a bank of 1-D convolutional filters that are parameterized as cosine modulated Gaussian functions. The choice of these kernels allows the interpretation of the filterbanks as smooth band-pass filters. The filtered outputs are pooled, log-compressed and used in a self-attention based relevance weighting mechanism. The relevance weighting emphasizes the key regions of the time-frequency decomposition that are important for the downstream task. The subsequent layers of the model consist of a recurrent architecture and the models are trained for a COVID-19 detection task. In our experiments on the Coswara data set, we show that the proposed model achieves significant performance improvements over the baseline system as well as other representation learning approaches. Further, the approach proposed is shown to be uniformly applicable for speech and breathing signals and for transfer learning from a larger data set.


【30】 Insights into Deep Non-linear Filters for Improved Multi-channel Speech  Enhancement

标题:对改进多通道语音增强的深度非线性滤波器的见解

链接:https://arxiv.org/abs/2206.13310

作者:Kristina Tesch,Timo Gerkmann
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:使用多个麦克风进行语音增强的关键优势在于,可以使用空间滤波来补充节奏谱处理。在传统设置中,通常分别执行线性空间滤波(波束形成)和单信道后滤波。相比之下,有一种趋势是使用深度神经网络(DNN)来学习空间和节奏谱联合非线性滤波器,这意味着可以潜在地克服线性处理模型以及空间和节奏谱信息单独处理的限制。然而,导致这种用于多通道语音增强的数据驱动滤波器具有良好性能的内部机制尚不清楚。因此,在这项工作中,我们通过仔细控制网络可用的信息源(空间、光谱和时间),分析了DNN实现的非线性空间滤波器的特性及其与时间和光谱处理的相互依赖性。我们证实了非线性空间处理模型的优越性,该模型在具有挑战性的说话人提取场景中,对于低数量的麦克风,其性能优于oracle线性空间滤波器,POLQA分数为0.24。我们的分析表明,特别是光谱信息应与空间信息联合处理,因为这会增加滤波器的空间选择性。然后,我们的系统评估得出了一个简单的网络架构,该架构在说话人提取任务上的表现优于最先进的网络架构,POLQA得分为0.22,CHiME3数据上的POLQA得分为0.32。
摘要:The key advantage of using multiple microphones for speech enhancement is that spatial filtering can be used to complement the tempo-spectral processing. In a traditional setting, linear spatial filtering (beamforming) and single-channel post-filtering are commonly performed separately. In contrast, there is a trend towards employing deep neural networks (DNNs) to learn a joint spatial and tempo-spectral non-linear filter, which means that the restriction of a linear processing model and that of a separate processing of spatial and tempo-spectral information can potentially be overcome. However, the internal mechanisms that lead to good performance of such data-driven filters for multi-channel speech enhancement are not well understood. Therefore, in this work, we analyse the properties of a non-linear spatial filter realized by a DNN as well as its interdependency with temporal and spectral processing by carefully controlling the information sources (spatial, spectral, and temporal) available to the network. We confirm the superiority of a non-linear spatial processing model, which outperforms an oracle linear spatial filter in a challenging speaker extraction scenario for a low number of microphones by 0.24 POLQA score. Our analyses reveal that in particular spectral information should be processed jointly with spatial information as this increases the spatial selectivity of the filter. Our systematic evaluation then leads to a simple network architecture, that outperforms state-of-the-art network architectures on a speaker extraction task by 0.22 POLQA score and by 0.32 POLQA score on the CHiME3 data.


【31】 Wideband Audio Waveform Evaluation Networks: Efficient, Accurate  Estimation of Speech Qualities

标题:宽带音频波形评估网络:高效、准确的语音质量评估

链接:https://arxiv.org/abs/2206.13272

作者:Andrew Catellier,Stephen Voran
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:宽带音频波形评估网络(WAWENET)是一种卷积神经网络,它直接对宽带音频波形进行操作,以产生对这些波形的评估。在目前的工作中,这些评估给出了电信语音的质量(例如,噪音、清晰度、总体语音质量)。Wawenet不是参考网络,因为它们不需要所评估波形的“参考”(原始或未失真)版本。我们最初的WAWEnet出版物介绍了四个WAWEnet,每个WAWEnet都模拟了已建立的全参考语音质量或可懂度估计算法的输出。我们更新了WAWEnet架构,使其更加高效。在这里,我们提供了一个单独的WAWEnet,它密切跟踪七个不同的质量和可理解性值。我们创建了第二个网络,另外跟踪四个主观语音质量维度。我们提供了第三个网络,它只关注主观质量分数,并实现了非常高的一致性。这项工作利用了13种语言334小时的演讲时间、200多万个完整参考目标值和93000多个主观平均意见分数。我们还解释了WAWEnets的操作,并使用信号处理语言确定了其操作的关键:ReLUs战略性地将光谱信息从非直流分量移动到直流分量。96个输出信号的DC值定义了96-D潜在空间中的向量,然后将该向量映射到输入波形的质量或可懂度值。
摘要:Wideband Audio Waveform Evaluation Networks (WAWEnets) are convolutional neural networks that operate directly on wideband audio waveforms in order to produce evaluations of those waveforms. In the present work these evaluations give qualities of telecommunications speech (e.g., noisiness, intelligibility, overall speech quality). WAWEnets are no-reference networks because they do not require ``reference'' (original or undistorted) versions of the waveforms they evaluate. Our initial WAWEnet publication introduced four WAWEnets and each emulated the output of an established full-reference speech quality or intelligibility estimation algorithm.  We have updated the WAWEnet architecture to be more efficient and effective. Here we present a single WAWEnet that closely tracks seven different quality and intelligibility values. We create a second network that additionally tracks four subjective speech quality dimensions. We offer a third network that focuses on just subjective quality scores and achieves very high levels of agreement. This work has leveraged 334 hours of speech in 13 languages, over two million full-reference target values and over 93,000 subjective mean opinion scores.  We also interpret the operation of WAWEnets and identify the key to their operation using the language of signal processing: ReLUs strategically move spectral information from non-DC components into the DC component. The DC values of 96 output signals define a vector in a 96-D latent space and this vector is then mapped to a quality or intelligibility value for the input waveform.


【32】 A Simple Baseline for Domain Adaptation in End to End ASR Systems Using  Synthetic Data

标题:端到端ASR系统中使用合成数据进行域自适应的简单基线

链接:https://arxiv.org/abs/2206.13240

作者:Raviraj Joshi,Anupam Singh
备注:Accepted at ECNLP @ACL 2022
摘要:基于深度学习的端到端语音识别模型一直是自动语音识别(ASR)的主流。这些方法需要大量以音频文本对的形式标记的数据。此外,与传统模型相比,这些模型更容易受到域转移的影响。通常的做法是训练通用ASR模型,然后使用相对较小的数据集将其适应于目标域。我们考虑了一个更极端的领域适应情况,其中只有文本语料库可用。在这项工作中,我们提出了一种用于端到端语音识别模型中域自适应的简单基线技术。我们使用单说话人文本到语音(TTS)引擎将纯文本语料库转换为音频数据。然后使用目标域中的并行数据微调通用ASR模型的最终密集层。我们表明,单说话人合成的TTS数据加上最终的密集层微调,可以合理地提高字错误率。我们使用来自地址和电子商务搜索域的文本数据来显示我们在CTC和基于注意的模型上的低成本基线方法的有效性。
摘要:Automatic Speech Recognition(ASR) has been dominated by deep learning-based end-to-end speech recognition models. These approaches require large amounts of labeled data in the form of audio-text pairs. Moreover, these models are more susceptible to domain shift as compared to traditional models. It is common practice to train generic ASR models and then adapt them to target domains using comparatively smaller data sets. We consider a more extreme case of domain adaptation where text-only corpus is available. In this work, we propose a simple baseline technique for domain adaptation in end-to-end speech recognition models. We convert the text-only corpus to audio data using single speaker Text to Speech (TTS) engine. The parallel data in the target domain is then used to fine-tune the final dense layer of generic ASR models. We show that single speaker synthetic TTS data coupled with final dense layer only fine-tuning provides reasonable improvements in word error rates. We use text data from address and e-commerce search domains to show the effectiveness of our low-cost baseline approach on CTC and attention-based models.


【33】 Conformer Based Elderly Speech Recognition System for Alzheimer's  Disease Detection

标题:基于仿真器的阿尔茨海默病老年语音识别系统

链接:https://arxiv.org/abs/2206.13232

作者:Tianzi Wang,Jiajun Deng,Mengzhe Geng,Zi Ye,Shoukang Hu,Yi Wang,Mingyu Cui,Zengrui Jin,Xunying Liu,Helen Meng
备注:5 pages, 1 figure, accepted by INTERSPEECH 2022
摘要:阿尔茨海默病(AD)的早期诊断对于促进预防护理以延缓进一步进展至关重要。本文介绍了一种基于DementiaBank-Pitt语料库的用于自动广告检测的基于构象的语音识别系统的开发。通过加入一组专门设计的建模特征,包括基于神经架构搜索的领域特定构象超参数自动配置以及参数微调,通过速度扰动和基于SpecAugment的数据扩充训练的基线构象系统得到了显著改进;基于学习隐藏单元贡献的细粒度老年说话人自适应(LHUC);以及基于双通道交叉系统重排序的混合TDNN系统组合。在48名老年人的评估数据中,单词错误率(WER)绝对下降了13.6%(相对下降了34.8%)。利用最终系统的识别输出提取文本特征,获得了基于语音识别的广告检测准确率为91.7%的最佳结果。
摘要:Early diagnosis of Alzheimer's disease (AD) is crucial in facilitating preventive care to delay further progression. This paper presents the development of a state-of-the-art Conformer based speech recognition system built on the DementiaBank Pitt corpus for automatic AD detection. The baseline Conformer system trained with speed perturbation and SpecAugment based data augmentation is significantly improved by incorporating a set of purposefully designed modeling features, including neural architecture search based auto-configuration of domain-specific Conformer hyper-parameters in addition to parameter fine-tuning; fine-grained elderly speaker adaptation using learning hidden unit contributions (LHUC); and two-pass cross-system rescoring based combination with hybrid TDNN systems. An overall word error rate (WER) reduction of 13.6% absolute (34.8% relative) was obtained on the evaluation data of 48 elderly speakers. Using the final systems' recognition outputs to extract textual features, the best-published speech recognition based AD detection accuracy of 91.7% was obtained.


【34】 Detection of Doctored Speech: Towards an End-to-End Parametric  Learn-able Filter Approach

标题:语音篡改检测:一种端到端的参数可学习滤波方法

链接:https://arxiv.org/abs/2206.13066

作者:Rohit Arora
备注:arXiv admin note: text overlap with arXiv:1904.05441 by other authors
摘要:自动说话人验证系统在生物特征识别应用中具有潜在的逻辑控制访问和身份验证功能。如果ASV系统遭到破坏,很多事情都会受到威胁。初步工作分别对这些论文中开发的基于小波和MFCC的最新欺骗检测技术进行了比较分析(Novoselov et al.,2016)(Alam et al.,2016a)。ASVspoof 2015的结果证明我们倾向于基于小波的特征,而不是MFCC特征。在ASVspoof 2019数据库上的实验表明,传统手工制作的功能缺乏可信度,这给了我们更多的理由,让我们有更多的理由朝着使用端到端深层神经网络和最新技术的方向发展。我们使用Sincnet架构作为基线。通过将Sinc层替换为小波散射层和连续小波变换层,我们得到了E2E深度学习模型,分别称为WSTnet和CWTnet。在2019年ASVspoof中对现代欺骗攻击进行评估时,融合模型与传统手工模型和我们的Sincnet基线相比,分别实现了62%和17%的相对改进。CWTnet中使用的最终规模分布和规模数量远远不是当前任务的最佳规模。因此,为了解决这个问题,我们在CWTnet体系结构中用小波反卷积(WD)(Khan和Yener,2018)层替换了CWT层。该层计算类似于CWTnet的离散连续小波变换,但也使用反向传播优化尺度参数。当在2019年ASVspoof数据集上进行评估时,WDnet模型分别比CWTnet和Sincnet模型实现了26%和7%的相对改善。这表明,与CWTnet提取的特征相比,由于只关注最重要和相关的频率区域,因此提取了更多的广义特征。
摘要:The Automatic Speaker Verification systems have potential in biometrics applications for logical control access and authentication. A lot of things happen to be at stake if the ASV system is compromised. The preliminary work presents a comparative analysis of the wavelet and MFCC-based state-of-the-art spoof detection techniques developed in these papers, respectively (Novoselov et al., 2016) (Alam et al., 2016a). The results on ASVspoof 2015 justify our inclination towards wavelet-based features instead of MFCC features. The experiments on the ASVspoof 2019 database show the lack of credibility of the traditional handcrafted features and give us more reason to progress towards using end-to-end deep neural networks and more recent techniques. We use Sincnet architecture as our baseline. We get E2E deep learning models, which we call WSTnet and CWTnet, respectively, by replacing the Sinc layer with the Wavelet Scattering and Continuous wavelet transform layers. The fusion model achieved 62% and 17% relative improvement over traditional handcrafted models and our Sincnet baseline when evaluated on the modern spoofing attacks in ASVspoof 2019.  The final scale distribution and the number of scales used in CWTnet are far from optimal for the task at hand. So to solve this problem, we replaced the CWT layer with a Wavelet Deconvolution(WD) (Khan and Yener, 2018) layer in our CWTnet architecture. This layer calculates the Discrete-Continuous Wavelet Transform similar to the CWTnet but also optimizes the scale parameter using back-propagation. The WDnet model achieved 26% and 7% relative improvement over CWTnet and Sincnet models respectively when evaluated over ASVspoof 2019 dataset. This shows that more generalized features are extracted as compared to the features extracted by CWTnet as only the most important and relevant frequency regions are focused upon.


【35】 Extended U-Net for Speaker Verification in Noisy Environments

标题:噪声环境下说话人确认的扩展U网

链接:https://arxiv.org/abs/2206.13044

作者:Ju-ho Kim,Jungwoo Heo,Hye-jin Shim,Ha-Jin Yu
备注:5 pages, 2 figures, 4 tables, accepted to 2022 Interspeech as a conference paper
摘要:背景噪声是一个众所周知的因素,它通过模糊语音清晰度来降低说话人确认(SV)系统的准确性和可靠性。各种研究使用单独的预训练增强模型作为噪声环境中SV系统的前端模块,这些方法有效地去除了噪声。然而,不适合SV任务的独立增强模型的去噪过程也会扭曲话语中包含的说话人信息。我们认为,增强网络和说话人嵌入提取器应该在噪声条件下针对SV任务进行充分的联合训练,以缓解这一问题。因此,我们提出了一个基于U-Net的集成框架,该框架可以同时优化说话人识别和特征增强损失。此外,我们还分析了直接将U网络用于噪声SV任务的结构局限性,并进一步提出了扩展U网络以减少这些缺陷。我们在各种噪声场景下记录的噪声合成VoxCeleb1测试集和语音开发集上评估了模型。实验结果表明,基于U-Net的完全联合训练框架比基线训练框架更有效,与最近提出的补偿系统相比,扩展的U-Net具有最先进的性能。
摘要:Background noise is a well-known factor that deteriorates the accuracy and reliability of speaker verification (SV) systems by blurring speech intelligibility. Various studies have used separate pretrained enhancement models as the front-end module of the SV system in noisy environments, and these methods effectively remove noises. However, the denoising process of independent enhancement models not tailored to the SV task can also distort the speaker information included in utterances. We argue that the enhancement network and speaker embedding extractor should be fully jointly trained for SV tasks under noisy conditions to alleviate this issue. Therefore, we proposed a U-Net-based integrated framework that simultaneously optimizes speaker identification and feature enhancement losses. Moreover, we analyzed the structural limitations of using U-Net directly for noise SV tasks and further proposed Extended U-Net to reduce these drawbacks. We evaluated the models on the noise-synthesized VoxCeleb1 test set and VOiCES development set recorded in various noisy scenarios. The experimental results demonstrate that the U-Net-based fully joint training framework is more effective than the baseline, and the extended U-Net exhibited state-of-the-art performance versus the recently proposed compensation systems.


【36】 Joint Optimization of Sampling Rate Offsets Based on Entire Signal  Relationship Among Distributed Microphones

标题:基于分布式麦克风整体信号关系的采样率偏差联合优化

链接:https://arxiv.org/abs/2206.13014

作者:Yoshiki Masuyama,Kouei Yamaoka,Nobutaka Ono
备注:5 pages, 2 figures,accepted by Interspeech2022
摘要:在本文中,我们建议同时估计多个设备的所有采样率偏移(SRO)。在分布式麦克风阵列中,SRO是不可避免的,它会降低阵列信号处理的性能。现有的SRO估计方法大多集中于两个话筒的同步。当同步两个以上话筒时,我们选择一个参考话筒,并独立估计每个非参考话筒的SRO。因此,不考虑非参考话筒观察到的信号之间的关系。为了解决这个问题,该方法基于多信道信号的概率模型联合优化所有SRO。SRO和模型参数交替更新,以增加基于辅助函数的对数可能性。该方法的有效性在不同说话人数的混合情况下得到了验证。
摘要:In this paper, we propose to simultaneously estimate all the sampling rate offsets (SROs) of multiple devices. In a distributed microphone array, the SRO is inevitable, which deteriorates the performance of array signal processing. Most of the existing SRO estimation methods focused on synchronizing two microphones. When synchronizing more than two microphones, we select one reference microphone and estimate the SRO of each non-reference microphone independently. Hence, the relationship among signals observed by non-reference microphones is not considered. To address this problem, the proposed method jointly optimizes all SROs based on a probabilistic model of a multichannel signal. The SROs and model parameters are alternately updated to increase the log-likelihood based on an auxiliary function. The effectiveness of the proposed method is validated on mixtures of various numbers of speakers.


【37】 Transport-Oriented Feature Aggregation for Speaker Embedding Learning

标题:用于说话人嵌入学习的面向传输的特征聚合

链接:https://arxiv.org/abs/2206.12857

作者:Yusheng Tian,Jingyu Li,Tan Lee
备注:Accepted for presentation at INTERSPEECH 2022
摘要:为了进行说话人建模,需要将帧级特征聚合到话语级表示中。鉴于基于统计的池方法的成功,我们假设说话人特征在预聚合层输出的统计分布中得到很好的表示,并建议使用面向传输的特征聚合来派生说话人嵌入。聚合表示对基本特征分布的几何结构进行编码,该特征分布预计包含有价值的说话人特定信息,这些信息可能无法由常用的统计度量(如均值和方差)表示。原始的面向传输的特征聚合也扩展到加权帧版本,以纳入注意机制。在Voxceleb数据集上进行的说话人验证实验表明,与统计池及其注意变体相比,该方法有了改进。
摘要:Pooling is needed to aggregate frame-level features into utterance-level representations for speaker modeling. Given the success of statistics-based pooling methods, we hypothesize that speaker characteristics are well represented in the statistical distribution over the pre-aggregation layer's output, and propose to use transport-oriented feature aggregation for deriving speaker embeddings. The aggregated representation encodes the geometric structure of the underlying feature distribution, which is expected to contain valuable speaker-specific information that may not be represented by the commonly used statistical measures like mean and variance. The original transport-oriented feature aggregation is also extended to a weighted-frame version to incorporate the attention mechanism. Experiments on speaker verification with the Voxceleb dataset show improvement over statistics pooling and its attentive variant.


【38】 Meta Auxiliary Learning for Low-resource Spoken Language Understanding

标题:低资源口语理解的元辅助性学习

链接:https://arxiv.org/abs/2206.12774

作者:Yingying Gao,Junlan Feng,Chao Deng,Shilei Zhang
摘要:口语理解(SLU)将自动语音识别(ASR)和自然语言理解(NLU)视为一项统一的任务,通常存在数据匮乏的问题。我们开发了一种基于元辅助学习的ASR和NLU联合训练方法,仅利用丰富的语音数据人工转录来提高低资源SLU任务的性能。这种方法的一个明显优点是,它提供了一个灵活的框架来实现低资源的SLU训练任务,而无需访问任何进一步的语义注释。特别地,将NLU模型作为标签生成网络,从文本中预测意图和时隙标签;多任务网络从语音同步训练ASR任务和SLU任务;并将标签生成网络的预测作为语义目标传递给多任务网络。在公共CATSLU数据集上的实验证明了该算法的有效性,该数据集为后续NLU任务提供了更合适的ASR假设。
摘要:Spoken language understanding (SLU) treats automatic speech recognition (ASR) and natural language understanding (NLU) as a unified task and usually suffers from data scarcity. We exploit an ASR and NLU joint training method based on meta auxiliary learning to improve the performance of low-resource SLU task by only taking advantage of abundant manual transcriptions of speech data. One obvious advantage of such method is that it provides a flexible framework to implement a low-resource SLU training task without requiring access to any further semantic annotations. In particular, a NLU model is taken as label generation network to predict intent and slot tags from texts; a multi-task network trains ASR task and SLU task synchronously from speech; and the predictions of label generation network are delivered to the multi-task network as semantic targets. The efficiency of the proposed algorithm is demonstrated with experiments on the public CATSLU dataset, which produces more suitable ASR hypotheses for the downstream NLU task.


【39】 Predicting within and across language phoneme recognition performance of  self-supervised learning speech pre-trained models

标题:自监督学习语音预训练模型的语内和跨语言音素识别性能预测

链接:https://arxiv.org/abs/2206.12489

作者:Hang Ji,Tanvina Patel,Odette Scharenborg
备注:Submitted to INTERSPEECH 2022
摘要:在这项工作中,我们分析和比较了从不同的冷冻自监督学习(SSL)语音预训练模型中提取的语音表示,分析和比较了这些模型捕获发音特征(AF)信息的能力,以及它们对语言内和跨语言场景的电话识别性能的后续预测。具体而言,我们比较了CPC、wav2vec 2.0和HuBert。首先,执行帧级AF探测任务。随后,实现了用于音素识别任务的电话级端到端ASR系统,并将帧级AF探测任务的性能与电话准确性进行了关联。与传统的语音表示MFCC相比,所有SSL预训练语音表示都捕获了更多的AF信息,并在语言内部和跨语言实现了更好的音素识别性能,HuBert表现最好。帧级AF探测任务可以很好地预测音素识别性能,表明在语音表示中捕获AF信息的重要性。与MFCC相比,在语言内场景中,这些SSL语音预训练模型在AF探测任务上的性能实现了34.4%的最大相对增长,PER最低,为10.2%。在跨语言情况下,26.7%的最大相对增长率也导致了23.0%的最低PER。
摘要:In this work, we analyzed and compared speech representations extracted from different frozen self-supervised learning (SSL) speech pre-trained models on their ability to capture articulatory features (AF) information and their subsequent prediction of phone recognition performance for within and across language scenarios. Specifically, we compared CPC, wav2vec 2.0, and HuBert. First, frame-level AF probing tasks were implemented. Subsequently, phone-level end-to-end ASR systems for phoneme recognition tasks were implemented, and the performance on the frame-level AF probing task and the phone accuracy were correlated. Compared to the conventional speech representation MFCC, all SSL pre-trained speech representations captured more AF information, and achieved better phoneme recognition performance within and across languages, with HuBert performing best. The frame-level AF probing task is a good predictor of phoneme recognition performance, showing the importance of capturing AF information in the speech representations. Compared with MFCC, in the within-language scenario, the performance of these SSL speech pre-trained models on AF probing tasks achieved a maximum relative increase of 34.4%, and it resulted in the lowest PER of 10.2%. In the cross-language scenario, the maximum relative increase of 26.7% also resulted in the lowest PER of 23.0%.
eess.AS音频处理

【1】 Unsupervised Voice Activity Detection by Modeling Source and System  Information using Zero Frequency Filtering

标题:基于零频滤波的信源和系统信息建模的非监督语音活动检测

链接:https://arxiv.org/abs/2206.13420

作者:Eklavya Sarkar,RaviShankar Prasad,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2022
摘要:语音活动检测(VAD)是语音技术应用中一个重要的预处理步骤。该任务包括推导包含发声信息的音频信号的段边界。近年来,研究表明,使用零频率滤波(ZFF)可以提取语音源和声道系统信息,而无需对语音信号进行任何明确的模型假设。本文研究了零频率滤波在联合建模声源和声道系统信息方面的潜力,并提出了两种VAD方法。第一种方法使用由不同零频滤波信号组成的复合信号来划分浊音区域。第二种方法将合成信号作为输入馈送到RVD算法。这些方法与文献中的其他有监督和无监督VAD方法进行了比较,并在Aurora-2数据库上进行了评估,涵盖了一系列SNR(20-5 dB)。我们的研究表明,所提出的基于ZFF的方法的性能与最先进的VAD方法相当,并且对附加的降级和不同的信道特性更具不变性。
摘要:Voice activity detection (VAD) is an important pre-processing step for speech technology applications. The task consists of deriving segment boundaries of audio signals which contain voicing information. In recent years, it has been shown that voice source and vocal tract system information can be extracted using zero-frequency filtering (ZFF) without making any explicit model assumptions about the speech signal. This paper investigates the potential of zero-frequency filtering for jointly modeling voice source and vocal tract system information, and proposes two approaches for VAD. The first approach demarcates voiced regions using a composite signal composed of different zero-frequency filtered signals. The second approach feeds the composite signal as input to the rVAD algorithm. These approaches are compared with other supervised and unsupervised VAD methods in the literature, and are evaluated on the Aurora-2 database, across a range of SNRs (20 to -5 dB). Our studies show that the proposed ZFF-based methods perform comparable to state-of-art VAD methods and are more invariant to added degradation and different channel characteristics.


【2】 Pruned RNN-T for fast, memory-efficient ASR training

标题:经过修剪的RNN-T可实现快速、内存高效的ASR训练

链接:https://arxiv.org/abs/2206.13236

作者:Fangjun Kuang,Liyong Guo,Wei Kang,Long Lin,Mingshuang Luo,Zengwei Yao,Daniel Povey
摘要:用于语音识别的RNN传感器(RNN-T)框架越来越受欢迎,特别是对于部署的实时ASR系统,因为它将高精度与自然流识别相结合。RNN-T的缺点之一是其损失函数的计算速度相对较慢,并且可以使用大量内存。过多的GPU内存使用可能会使在词汇量较大的情况下使用RNN-T损失变得不切实际:例如,对于基于汉字的ASR。我们介绍了一种更快、更高效的RNN-T损失计算方法。我们首先使用编码器和解码器嵌入中线性的简单连接网络获得RNN-T递归的修剪边界;我们可以在不使用太多内存的情况下对此进行评估。然后,我们使用这些修剪边界来评估完整的非线性joiner网络。
摘要:The RNN-Transducer (RNN-T) framework for speech recognition has been growing in popularity, particularly for deployed real-time ASR systems, because it combines high accuracy with naturally streaming recognition. One of the drawbacks of RNN-T is that its loss function is relatively slow to compute, and can use a lot of memory. Excessive GPU memory usage can make it impractical to use RNN-T loss in cases where the vocabulary size is large: for example, for Chinese character-based ASR. We introduce a method for faster and more memory-efficient RNN-T loss computation. We first obtain pruning bounds for the RNN-T recursion using a simple joiner network that is linear in the encoder and decoder embeddings; we can evaluate this without using much memory. We then use those pruning bounds to evaluate the full, non-linear joiner network.


【3】 Conformer Based Elderly Speech Recognition System for Alzheimer's  Disease Detection

标题:基于仿真器的阿尔茨海默病老年语音识别系统

链接:https://arxiv.org/abs/2206.13232

作者:Tianzi Wang,Jiajun Deng,Mengzhe Geng,Zi Ye,Shoukang Hu,Yi Wang,Mingyu Cui,Zengrui Jin,Xunying Liu,Helen Meng
备注:5 pages, 1 figure, accepted by INTERSPEECH 2022
摘要:阿尔茨海默病(AD)的早期诊断对于促进预防护理以延缓进一步进展至关重要。本文介绍了一种基于DementiaBank-Pitt语料库的用于自动广告检测的基于构象的语音识别系统的开发。通过加入一组专门设计的建模特征,包括基于神经架构搜索的领域特定构象超参数自动配置以及参数微调,通过速度扰动和基于SpecAugment的数据扩充训练的基线构象系统得到了显著改进;基于学习隐藏单元贡献的细粒度老年说话人自适应(LHUC);以及基于双通道交叉系统重排序的混合TDNN系统组合。在48名老年人的评估数据中,单词错误率(WER)绝对下降了13.6%(相对下降了34.8%)。利用最终系统的识别输出提取文本特征,获得了基于语音识别的广告检测准确率为91.7%的最佳结果。
摘要:Early diagnosis of Alzheimer's disease (AD) is crucial in facilitating preventive care to delay further progression. This paper presents the development of a state-of-the-art Conformer based speech recognition system built on the DementiaBank Pitt corpus for automatic AD detection. The baseline Conformer system trained with speed perturbation and SpecAugment based data augmentation is significantly improved by incorporating a set of purposefully designed modeling features, including neural architecture search based auto-configuration of domain-specific Conformer hyper-parameters in addition to parameter fine-tuning; fine-grained elderly speaker adaptation using learning hidden unit contributions (LHUC); and two-pass cross-system rescoring based combination with hybrid TDNN systems. An overall word error rate (WER) reduction of 13.6% absolute (34.8% relative) was obtained on the evaluation data of 48 elderly speakers. Using the final systems' recognition outputs to extract textual features, the best-published speech recognition based AD detection accuracy of 91.7% was obtained.


【4】 QbyE-MLPMixer: Query-by-Example Open-Vocabulary Keyword Spotting using  MLPMixer

标题:QbyE-MLPMixer:基于MLPMixer的逐例查询开放词汇关键词识别

链接:https://arxiv.org/abs/2206.13231

作者:Jinmiao Huang,Waseem Gharbieh,Qianhui Wan,Han Suk Shim,Chul Lee
备注:Accepted to INTERSPEECH 2022
摘要:当前的关键字识别系统通常使用大量预定义的关键字进行训练。在开放词汇设置中识别关键字对于个性化智能设备交互至关重要。为了实现这一目标,我们提出了一种基于MLP的纯神经网络,该网络基于MLPMixer,这是一种MLP模型体系结构,可以有效地取代视觉转换器中的注意机制。我们研究了使MLPMixer体系结构适应QbyE开放词汇表关键字发现任务的不同方法。与最先进的RNN和CNN模型的比较表明,我们的方法在具有挑战性的情况下(10dB和6dB环境)在公开可用的Hey-Snips数据集和具有400个扬声器的大型内部数据集上都取得了更好的性能。与基线模型相比,我们提出的模型具有更少的参数和MAC。
摘要:Current keyword spotting systems are typically trained with a large amount of pre-defined keywords. Recognizing keywords in an open-vocabulary setting is essential for personalizing smart device interaction. Towards this goal, we propose a pure MLP-based neural network that is based on MLPMixer - an MLP model architecture that effectively replaces the attention mechanism in Vision Transformers. We investigate different ways of adapting the MLPMixer architecture to the QbyE open-vocabulary keyword spotting task. Comparisons with the state-of-the-art RNN and CNN models show that our method achieves better performance in challenging situations (10dB and 6dB environments) on both the publicly available Hey-Snips dataset and a larger scale internal dataset with 400 speakers. Our proposed model also has a smaller number of parameters and MACs compared to the baseline models.


【5】 Unsupervised Instance Discriminative Learning for Depression Detection  from Speech Signals

标题:无监督实例判别学习用于语音信号中的抑郁检测

链接:https://arxiv.org/abs/2206.13016

作者:Jinhan Wang,Vijay Ravi,Jonathan Flint,Abeer Alwan
摘要:重度抑郁症(MDD)是一种影响数百万人的严重疾病,尽早诊断这种疾病至关重要。从语音信号中检测抑郁症对医生有很大帮助,而且无需任何侵入性操作。由于相关标记数据较少,我们提出了一种改进的实例判别学习(IDL)方法,一种无监督的预训练技术,用于提取增广不变量和实例展开嵌入。在学习增强不变嵌入方面,研究了各种语音数据增强方法,时间掩蔽的性能最好。为了学习实例展开嵌入,我们探索了为训练批次采样实例的方法(不同的基于说话人和随机采样)。研究发现,基于不同说话人的采样比基于随机说话人的采样具有更好的性能,我们假设这一结果是因为相关的说话人信息被保留在嵌入中。此外,我们基于聚类算法提出了一种新的采样策略,即基于伪实例的采样(PIS),以增强嵌入的扩展特性。使用DepAudioNet在DAIC-WOZ(英语)和CONVERGE(普通话)数据集上进行了实验,在未进行预训练的情况下,使用PIS检测MDD相对于基线的情况下,观察到统计学上的显著改善,p值分别为0.0015和0.05。
摘要:Major Depressive Disorder (MDD) is a severe illness that affects millions of people, and it is critical to diagnose this disorder as early as possible. Detecting depression from voice signals can be of great help to physicians and can be done without any invasive procedure. Since relevant labelled data are scarce, we propose a modified Instance Discriminative Learning (IDL) method, an unsupervised pre-training technique, to extract augment-invariant and instance-spread-out embeddings. In terms of learning augment-invariant embeddings, various data augmentation methods for speech are investigated, and time-masking yields the best performance. To learn instance-spread-out embeddings, we explore methods for sampling instances for a training batch (distinct speaker-based and random sampling). It is found that the distinct speaker-based sampling provides better performance than the random one, and we hypothesize that this result is because relevant speaker information is preserved in the embedding. Additionally, we propose a novel sampling strategy, Pseudo Instance-based Sampling (PIS), based on clustering algorithms, to enhance spread-out characteristics of the embeddings. Experiments are conducted with DepAudioNet on DAIC-WOZ (English) and CONVERGE (Mandarin) datasets, and statistically significant improvements, with p-value 0.0015 and 0.05, respectively, are observed using PIS in the detection of MDD relative to the baseline without pre-training.


【6】 Impact of Acoustic Event Tagging on Scene Classification in a Multi-Task  Learning Framework

标题:声事件标注对多任务学习框架中场景分类的影响

链接:https://arxiv.org/abs/2206.13476

* 与cs.SD语音【1】为同一篇

作者:Rahil Parikh,Harshavardhan Sundar,Ming Sun,Chao Wang,Spyros Matsoukas
备注:Accepted at ISCA Interspeech 2022
摘要:声学事件是具有明确光谱-时间特征的声音,可以与产生它们的物理对象相关联。声学场景是这些声学事件的集合,没有特定的时间顺序。鉴于事件和场景之间的这种自然联系,人们普遍认为,对事件进行分类的能力必须有助于对场景进行分类。这导致了一些努力,试图利用多任务网络在声学事件标记(AET)和声学场景分类(ASC)方面做得很好。然而,在这些努力中,一项任务的改进并不保证另一项任务的改进,这表明ASC和AET之间存在紧张关系。目前尚不清楚AET的改善是否转化为ASC的改善。我们通过广泛的实证研究探讨了这一难题,并表明在一定条件下,在多任务网络中使用AET作为辅助任务可以持续提高ASC的性能。此外,ASC性能随着AET数据集的大小而进一步提高,并且对AET数据集中事件的选择或事件的数量不敏感。我们的结论是,ASC性能的改善来自于使用AET的正则化效果,而不是来自网络对声学事件的识别能力的提高。
摘要:Acoustic events are sounds with well-defined spectro-temporal characteristics which can be associated with the physical objects generating them. Acoustic scenes are collections of such acoustic events in no specific temporal order. Given this natural linkage between events and scenes, a common belief is that the ability to classify events must help in the classification of scenes. This has led to several efforts attempting to do well on Acoustic Event Tagging (AET) and Acoustic Scene Classification (ASC) using a multi-task network. However, in these efforts, improvement in one task does not guarantee an improvement in the other, suggesting a tension between ASC and AET. It is unclear if improvements in AET translates to improvements in ASC. We explore this conundrum through an extensive empirical study and show that under certain conditions, using AET as an auxiliary task in the multi-task network consistently improves ASC performance. Additionally, ASC performance further improves with the AET data-set size and is not sensitive to the choice of events or the number of events in the AET data-set. We conclude that this improvement in ASC performance comes from the regularization effect of using AET and not from the network's improved ability to discern between acoustic events.


【7】 Is the Language Familiarity Effect gradual? A computational modelling  approach

标题:语言熟悉度的影响是渐进的吗?一种计算建模方法

链接:https://arxiv.org/abs/2206.13415

* 与cs.SD语音【2】为同一篇

作者:Maureen de Seyssel,Guillaume Wisniewski,Emmanuel Dupoux
备注:8 pages, 2 figures, accepted at CogSci 2022
摘要:根据语言熟悉度效应(LFE),人们更善于区分母语使用者。尽管文献中对这种认知效应进行了大量研究,但实验仅在有限数量的语言对上进行,其结果仅显示了这种效应的存在,而没有产生可能因语言对而异的渐进测量。在这项工作中,我们表明Thorburn、Feldmand和Schatz(2019)提出的LFE计算模型可以解决这两个限制。在第一个实验中,我们证明了该模型通过复制母语和口音语音的行为发现,能够逐步测量LFE。在第二个实验中,我们评估了大量语言对的LFE,包括许多从未在人类身上测试过的语言对。我们表明,这种效应在广泛的语言中得到了复制,进一步证明了其普遍性。基于LFE的逐步测量,我们还表明,属于同一家族的语言得分较小,支持语言距离对LFE的影响。
摘要:According to the Language Familiarity Effect (LFE), people are better at discriminating between speakers of their native language. Although this cognitive effect was largely studied in the literature, experiments have only been conducted on a limited number of language pairs and their results only show the presence of the effect without yielding a gradual measure that may vary across language pairs. In this work, we show that the computational model of LFE introduced by Thorburn, Feldmand and Schatz (2019) can address these two limitations. In a first experiment, we attest to this model's capacity to obtain a gradual measure of the LFE by replicating behavioural findings on native and accented speech. In a second experiment, we evaluate LFE on a large number of language pairs, including many which have never been tested on humans. We show that the effect is replicated across a wide array of languages, providing further evidence of its universality. Building on the gradual measure of LFE, we also show that languages belonging to the same family yield smaller scores, supporting the idea of an effect of language distance on LFE.


【8】 A Comprehensive Survey on Video Saliency Detection with Auditory  Information: the Audio-visual Consistency Perceptual is the Key!

标题:听觉信息视频显著检测综述:视听一致性感知是关键!

链接:https://arxiv.org/abs/2206.13390

* 与cs.SD语音【3】为同一篇

作者:Chenglizhao Chen,Mengke Song,Wenfeng Song,Li Guo,Muwei Jian
摘要:视频显著性检测(VSD)旨在快速定位给定视频片段中最具吸引力的对象/事物/模式。现有的VSD相关作品主要依赖视觉系统,而对音频方面的关注较少,而实际上,我们的音频系统是视觉系统最重要的补充部分。此外,视听显著性检测(AVSD)是模仿人类感知机制的最具代表性的研究课题之一,目前还处于起步阶段,现有的调查论文都没有涉及到它,尤其是从显著性检测的角度。因此,本文的最终目的是对视听融合和显著性检测之间的差距进行广泛的回顾。此外,作为本综述的另一个亮点,我们深入了解了可能直接决定AVSD深度模型性能的关键因素,并声称视听一致性程度(AVC)——一个长期被忽视的问题,可以直接影响在执行显著性检测时使用音频使其视觉对应方受益的有效性。此外,为了使AVC问题对未来的关注者更加实用和有价值,我们新为几乎所有现有的公开可用AVSD数据集配备了额外的帧级AVC标签。基于这些升级的数据集,我们进行了广泛的定量评估,以证明AVC在AVSD任务中的重要性。总之,我们的想法和新设置都是一个方便的平台,提供了初步的准备和指导,所有这些都有助于推动未来的工作,进一步促进最先进的(SOTA)性能。
摘要:Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect, while, actually, our audio system is the most vital complementary part to our visual system. Also, audio-visual saliency detection (AVSD), one of the most representative research topics for mimicking human perceptual mechanisms, is currently in its infancy, and none of the existing survey papers have touched on it, especially from the perspective of saliency detection. Thus, the ultimate goal of this paper is to provide an extensive review to bridge the gap between audio-visual fusion and saliency detection. In addition, as another highlight of this review, we have provided a deep insight into key factors which could directly determine the performances of AVSD deep models, and we claim that the audio-visual consistency degree (AVC) -- a long-overlooked issue, can directly influence the effectiveness of using audio to benefit its visual counterpart when performing saliency detection. Moreover, in order to make the AVC issue more practical and valuable for future followers, we have newly equipped almost all existing publicly available AVSD datasets with additional frame-wise AVC labels. Based on these upgraded datasets, we have conducted extensive quantitative evaluations to ground our claim on the importance of AVC in the AVSD task. In a word, both our ideas and new sets serve as a convenient platform with preliminaries and guidelines, all of which are very potential to facilitate future works in promoting state-of-the-art (SOTA) performance further.


【9】 A two-stage full-band speech enhancement model with effective spectral  compression mapping

标题:一种有效谱压缩映射的两级全频带语音增强模型

链接:https://arxiv.org/abs/2206.13136

* 与cs.SD语音【4】为同一篇

作者:Zhongshu Hou,Qinwen Hu,Kai Chen,Jing Lu
摘要:基于深度神经网络(DNN)的宽带语音增强(SE)直接扩展到全频段处理面临着低频分辨率的挑战,这很可能导致模型性能的恶化。在本文中,我们提出了一种可学习的频谱压缩映射(SCM)来有效地压缩高频分量,以便以更有效的方式对其进行处理。通过这样做,该模型可以更加关注低频和中频范围,其中大部分语音功率都集中在低频和中频范围。我们不是在单个网络结构中抑制噪声,而是首先估计频谱幅度掩码,将语音转换为高信噪比(SNR)状态,然后利用后续模型进一步优化预增强信号的实掩码和虚掩码。我们进行了综合实验来验证所提方法的有效性。
摘要:The direct expansion of deep neural network (DNN) based wide-band speech enhancement (SE) to full-band processing faces the challenge of low frequency resolution in low frequency range, which would highly likely lead to deteriorated performance of the model. In this paper, we propose a learnable spectral compression mapping (SCM) to effectively compress the high frequency components so that they can be processed in a more efficient manner. By doing so, the model can pay more attention to low and middle frequency range, where most of the speech power is concentrated. Instead of suppressing noise in a single network structure, we first estimate a spectral magnitude mask, converting the speech to a high signal-to-ratio (SNR) state, and then utilize a subsequent model to further optimize the real and imaginary mask of the pre-enhanced signal. We conduct comprehensive experiments to validate the efficacy of the proposed method.


【10】 TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a  Speech Recognition Baseline

标题:TALCS:一个开源的汉英代码转换语料库和语音识别基线

链接:https://arxiv.org/abs/2206.13135

* 与cs.SD语音【5】为同一篇

作者:Chengfei Li,Shuhao Deng,Yaoping Wang,Guangjing Wang,Yaguang Gong,Changbin Chen,Jinfeng Bai
备注:accepted by INTERSPEECH 2022
摘要:本文介绍了一种新的汉英码转换语音识别语料库——TALCS语料库,适用于码转换语音识别系统的训练和评价。TALCS语料库来源于好未来教育集团真实的在线一对一英语教学场景,包含大约587小时的16千赫语音样本。据我们所知,TALCS语料库是世界上最大的标记良好的汉英代码转换开源自动语音识别(ASR)数据集。本文将详细介绍录音过程,包括音频采集设备和语料库环境。滑石语料库在许可证下可免费下载1。使用TALCS语料库,我们在两个流行的语音识别工具包中进行ASR实验,以构建一个基线系统,包括ESPnet和Wenet。在TALCS语料库中比较了两种语音识别工具包的混合错误率(MER)性能。实验结果表明,录音和转录的质量是有希望的,基线系统是可行的。
摘要:This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one English teaching scenes in TAL education group, which contains roughly 587 hours of speech sampled at 16 kHz. To our best knowledge, TALCS corpus is the largest well labeled Mandarin-English code-switching open source automatic speech recognition (ASR) dataset in the world. In this paper, we will introduce the recording procedure in detail, including audio capturing devices and corpus environments. And the TALCS corpus is freely available for download under the permissive license1. Using TALCS corpus, we conduct ASR experiments in two popular speech recognition toolkits to make a baseline system, including ESPnet and Wenet. The Mixture Error Rate (MER) performance in the two speech recognition toolkits is compared in TALCS corpus. The experimental results implies that the quality of audio recordings and transcriptions are promising and the baseline system is workable.


【11】 Sequence-level Speaker Change Detection with Difference-based Continuous  Integrate-and-fire

标题:基于差分连续积分点火的序列级说话人变化检测

链接:https://arxiv.org/abs/2206.13110

* 与cs.SD语音【6】为同一篇

作者:Zhiyun Fan,Linhao Dong,Meng Cai,Zejun Ma,Bo Xu
备注:Signal Processing Letters 2022
摘要:说话人变化检测是会议和会话等多方交互中的一项重要任务。本文从序列转导的角度研究说话人变化检测任务。具体来说,我们提出了一种新的编码器-解码器框架,直接将输入特征序列转换为说话人身份序列。设计了基于差异的持续集成和触发机制来支持该框架。它通过逐帧积分编码器输出之间的扬声器差异来检测扬声器变化,并根据检测到的扬声器变化将编码器输出传输到分段级扬声器嵌入。整个框架由说话人身份序列监控,这是一个比精确的说话人变化点弱的标签。在AMI和DIHARD-I语料库上的实验表明,我们的序列级方法始终优于使用精确说话人变化标签的强框架级基线。
摘要:Speaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we propose a novel encoder-decoder framework that directly converts the input feature sequence to the speaker identity sequence. The difference-based continuous integrate-and-fire mechanism is designed to support this framework. It detects speaker changes by integrating the speaker difference between the encoder outputs frame-by-frame and transfers encoder outputs to segment-level speaker embeddings according to the detected speaker changes. The whole framework is supervised by the speaker identity sequence, a weaker label than the precise speaker change points. The experiments on the AMI and DIHARD-I corpora show that our sequence-level method consistently outperforms a strong frame-level baseline that uses the precise speaker change labels.


【12】 SpeechEQ: Speech Emotion Recognition based on Multi-scale Unified  Datasets and Multitask Learning

标题:SpeechEQ:基于多尺度统一数据集和多任务学习的语音情感识别

链接:https://arxiv.org/abs/2206.13101

* 与cs.SD语音【7】为同一篇

作者:Zuheng Kang,Junqing Peng,Jianzong Wang,Jing Xiao
摘要:语音情感识别(SER)面临许多挑战,但其中一个主要挑战是每个框架都没有统一的标准。在本文中,我们提出了SpeechEQ,这是一个基于多尺度统一度量的统一SER任务的框架。该指标可以通过多任务学习(MTL)进行训练,MTL包括情绪状态类别(EIS)和情绪强度量表(EIS)两个情绪识别任务,以及音位识别和性别识别两个辅助任务。对于这个框架,我们构建了一个汉语SER数据集——SpeechEQ数据集(SEQD)。我们在普通话的公共CASIA和ESD数据集上进行了实验,结果表明,我们的方法比基线方法有较大幅度的提高,准确率分别提高了8.0%和6.5%。在IEMOCAP上对四种情绪类别(即愤怒、快乐、悲伤和中性)的额外实验也表明,所提出的方法实现了78.16%的加权准确率(WA)和77.47%的未加权准确率(UA)的最新水平。
摘要:Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a framework for unifying SER tasks based on a multi-scale unified metric. This metric can be trained by Multitask Learning (MTL), which includes two emotion recognition tasks of Emotion States Category (EIS) and Emotion Intensity Scale (EIS), and two auxiliary tasks of phoneme recognition and gender recognition. For this framework, we build a Mandarin SER dataset - SpeechEQ Dataset (SEQD). We conducted experiments on the public CASIA and ESD datasets in Mandarin, which exhibit that our method outperforms baseline methods by a relatively large margin, yielding 8.0\% and 6.5\% improvement in accuracy respectively. Additional experiments on IEMOCAP with four emotion categories (i.e., angry, happy, sad, and neutral) also show the proposed method achieves a state-of-the-art of both weighted accuracy (WA) of 78.16% and unweighted accuracy (UA) of 77.47%.


【13】 Sound Model Factory: An Integrated System Architecture for Generative  Audio Modelling

标题:声音模型工厂:一个用于生成性音频建模的集成系统架构

链接:https://arxiv.org/abs/2206.13085

* 与cs.SD语音【8】为同一篇

作者:Lonce Wyse,Purnima Kamath,Chitralekha Gupta
备注:None
摘要:我们介绍了一种新的数据驱动音频模型设计系统,该系统围绕两种不同的神经网络架构构建,一种是生成对抗网络(GAN)和一种递归神经网络(RNN),它利用每种网络的独特特性来实现两种网络都无法单独解决的系统目标。该系统的目标是生成交互式可控的声音模型,前提是(a)模型应能够合成一系列声音,以及(b)用于导航该声音空间的参数控制规范。声音范围由设计者提供的数据集定义,而导航方式则由数据标签的组合和从GAN学习的潜在空间中选择子流形来定义。我们提出的系统利用了GAN丰富的潜在空间,该空间由声音组成,填满了“之间”的空间“真实数据就像声音。然后,使用来自GAN的这些增强数据来训练RNN,使其能够立即、连续地响应参数变化,并在无限的时间内生成音频。此外,我们开发了一种自组织映射技术,用于“平滑”GAN的潜在空间,从而在音频音色之间进行感知平滑的插值。我们通过用户研究来验证这个过程。该系统有助于提高生成性声音模型设计的技术水平,包括系统配置和组件,以改进插值,并将音频建模能力从音高和打击乐器声音扩展到更复杂的音频纹理空间。
摘要:We introduce a new system for data-driven audio sound model design built around two different neural network architectures, a Generative Adversarial Network(GAN) and a Recurrent Neural Network (RNN), that takes advantage of the unique characteristics of each to achieve the system objectives that neither is capable of addressing alone. The objective of the system is to generate interactively controllable sound models given (a) a range of sounds the model should be able to synthesize, and (b) a specification of the parametric controls for navigating that space of sounds. The range of sounds is defined by a dataset provided by the designer, while the means of navigation is defined by a combination of data labels and the selection of a sub-manifold from the latent space learned by the GAN. Our proposed system takes advantage of the rich latent space of a GAN that consists of sounds that fill out the spaces ''between" real data-like sounds. This augmented data from the GAN is then used to train an RNN for its ability to respond immediately and continuously to parameter changes and to generate audio over unlimited periods of time. Furthermore, we develop a self-organizing map technique for ``smoothing" the latent space of GAN that results in perceptually smooth interpolation between audio timbres. We validate this process through user studies. The system contributes advances to the state of the art for generative sound model design that include system configuration and components for improving interpolation and the expansion of audio modeling capabilities beyond musical pitch and percussive instrument sounds into the more complex space of audio textures.


【14】 Uncertainty Calibration for Deep Audio Classifiers

标题:深度音频分类器的不确定度校正

链接:https://arxiv.org/abs/2206.13071

* 与cs.SD语音【9】为同一篇

作者:Tong Ye,Shijing Si,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by InterSpeech 2022, the first two authors contributed equally
摘要:虽然深度神经网络(DNN)在音频分类任务中取得了巨大的成功,但其不确定性校准仍处于探索阶段。一个经过良好校准的模型在其预测确定时应该是准确的,而在可能不准确时则表明其具有很高的不确定性。在这项工作中,我们研究了深度音频分类器的不确定性校准。特别是,我们实证研究了流行的校准方法在音频分类数据集上的性能:(i)蒙特卡罗衰减,(ii)系综,(iii)焦损,(iv)谱归一化高斯过程(SNGP)。为此,我们评估(i-iv)环境声音和音乐流派分类的任务。结果表明,未经校准的深度音频分类器可能过于自信,而SNGP在本文的两个数据集上表现最好,效率非常高。
摘要:Although deep Neural Networks (DNNs) have achieved tremendous success in audio classification tasks, their uncertainty calibration are still under-explored. A well-calibrated model should be accurate when it is certain about its prediction and indicate high uncertainty when it is likely to be inaccurate. In this work, we investigate the uncertainty calibration for deep audio classifiers. In particular, we empirically study the performance of popular calibration methods: (i) Monte Carlo Dropout, (ii) ensemble, (iii) focal loss, and (iv) spectral-normalized Gaussian process (SNGP), on audio classification datasets. To this end, we evaluate (i-iv) for the tasks of environment sound and music genre classification. Results indicate that uncalibrated deep audio classifiers may be over-confident, and SNGP performs the best and is very efficient on the two datasets of this paper.


【15】 Speak Like a Professional: Increasing Speech Intelligibility by  Mimicking Professional Announcer Voice with Voice Conversion

标题:像专业人士一样说话:通过语音转换模仿专业播音员的声音来提高语音清晰度

链接:https://arxiv.org/abs/2206.13021

* 与cs.SD语音【10】为同一篇

作者:Tuan Vu Ho,Maori Kobayashi,Masato Akagi
备注:Accepted at INTERSPEECH 2022
摘要:在大多数实际场景中,广播系统必须在噪声环境中传递语音信息,在噪声环境中背景噪声无法消除。局部噪声降低了语音清晰度,增加了听者的听力,从而影响了广播系统的有效性。据报道,在嘈杂的环境中,专业播音员的声音比非专业播音员的声音更清晰、更全面。这一发现表明,语音清晰度可能与专业播音员的说话风格有关,可以使用语音转换方法进行调整。基于这一思想,本文将语音转换方法应用于非专业语音,提出了一种在噪声环境下提高语音清晰度的方法。我们发现,在说话人嵌入平面上,专业播音员和非专业说话人被分为不同的簇。这意味着语音清晰度可以作为说话人个性的一个独立特征来控制。为了检验转换语音在噪声环境中的优势,我们在不同信噪比下使用粉红色噪声中隐藏的测试词进行了实验。客观和主观评价结果证实,在低信噪比条件下,转换语音的语音清晰度高于原始语音。
摘要:In most of practical scenarios, the announcement system must deliver speech messages in a noisy environment, in which the background noise cannot be cancelled out. The local noise reduces speech intelligibility and increases listening effort of the listener, hence hamper the effectiveness of announcement system. There has been reported that voices of professional announcers are clearer and more comprehensive than that of non-expert speakers in noisy environment. This finding suggests that the speech intelligibility might be related to the speaking style of professional announcer, which can be adapted using voice conversion method. Motivated by this idea, this paper proposes a speech intelligibility enhancement in noisy environment by applying voice conversion method on non-professional voice. We discovered that the professional announcers and non-professional speakers are clusterized into different clusters on the speaker embedding plane. This implies that the speech intelligibility can be controlled as an independent feature of speaker individuality. To examine the advantage of converted voice in noisy environment, we experimented using test words masked in pink noise at different SNR levels. The results of objective and subjective evaluations confirm that the speech intelligibility of converted voice is higher than that of original voice in low SNR conditions.


【16】 Improving the Training Recipe for a Robust Conformer-based Hybrid Model

标题:一种基于构象的稳健混合模型训练方法的改进

链接:https://arxiv.org/abs/2206.12955

作者:Mohammad Zeineldeen,Jingjing Xu,Christoph Lüscher,Ralf Schlüter,Hermann Ney
备注:Accepted at INTERSPEECH 2022
摘要:说话人自适应对于构建鲁棒的自动语音识别(ASR)系统至关重要。在这项工作中,我们研究了基于特征空间方法的说话人自适应训练(SAT)的各种方法,这些方法适用于交换机300h数据集上基于一致性的声学模型(AM)。我们提出了一种称为加权简单加法(weightedsimple Add)的方法,该方法将加权说话人信息向量添加到一致性AM的多头自我注意模块的输入中。将该方法用于SAT,我们在Hub5'00和Hub5'01的呼叫总部部分分别实现了3.5%和4.5%的相对改善。此外,我们在之前工作的基础上,提出了一种新的、有竞争力的基于一致性的混合AM训练配方。我们扩展并改进了这个配方,在交换机300h Hub5'00数据集上,字错误率(WER)相对提高了11%。我们还通过将参数总数相对减少34%,使该配方有效。
摘要:Speaker adaptation is important to build robust automatic speech recognition (ASR) systems. In this work, we investigate various methods for speaker adaptive training (SAT) based on feature-space approaches for a conformer-based acoustic model (AM) on the Switchboard 300h dataset. We propose a method, called Weighted-Simple-Add, which adds weighted speaker information vectors to the input of the multi-head self-attention module of the conformer AM. Using this method for SAT, we achieve 3.5% and 4.5% relative improvement in terms of WER on the CallHome part of Hub5'00 and Hub5'01 respectively. Moreover, we build on top of our previous work where we proposed a novel and competitive training recipe for a conformer-based hybrid AM. We extend and improve this recipe where we achieve 11% relative improvement in terms of word-error-rate (WER) on Switchboard 300h Hub5'00 dataset. We also make this recipe efficient by reducing the total number of parameters by 34% relative.