今天跟大家分享一篇语音相关的论文合集:cs.SD语音20篇,eess.AS音频处理23篇。

cs.SD语音

【1】 Rethinking Audio-visual Synchronization for Active Speaker Detection

标题:基于有源说话人检测的视听同步再思考

链接:https://arxiv.org/abs/2206.10421

作者:Abudukelimu Wuerkaixi,You Zhang,Zhiyao Duan,Changshui Zhang

机构:⋆ Institute for Artificial Intelligence, Tsinghua University (THUAI), State Key Lab of Intelligent Technologies and Systems, Beijing National Research Center for Information Science and Technology (BNRist)

备注:Accepted by IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2022)

摘要:主动说话人检测(ASD)系统是分析多人会话的重要模块。他们的目标是在任何给定的时间检测视觉场景中哪些说话者在说话或没有人在说话。关于自闭症的现有研究对主动说话人的定义不一致。我们在这项工作中澄清了定义,并要求音频和视频演讲活动之间保持同步。这种定义的澄清是由我们的大量实验推动的,通过实验我们发现,现有的ASD方法无法对视听同步进行建模,并且通常将不同步的视频归类为主动说话。为了解决这个问题,我们提出了一种跨模态对比学习策略,并在有监督的ASD模型的注意模块中应用位置编码来利用同步线索。实验结果表明,我们的模型可以成功地将不同步说话检测为不说话,解决了现有模型的局限性。

摘要:Active speaker detection (ASD) systems are important modules for analyzing multi-talker conversations. They aim to detect which speakers or none are talking in a visual scene at any given time. Existing research on ASD does not agree on the definition of active speakers. We clarify the definition in this work and require synchronization between the audio and visual speaking activities. This clarification of definition is motivated by our extensive experiments, through which we discover that existing ASD methods fail in modeling the audio-visual synchronization and often classify unsynchronized videos as active speaking. To address this problem, we propose a cross-modal contrastive learning strategy and apply positional encoding in attention modules for supervised ASD models to leverage the synchronization cue. Experimental results suggest that our model can successfully detect unsynchronized speaking as not speaking, addressing the limitation of current models.


【2】 Audio-video fusion strategies for active speaker detection in meetings

标题:会议中主动说话人检测的音视频融合策略

链接:https://arxiv.org/abs/2206.10411

作者:Lionel Pibre,Francisco Madrigal,Cyrille Equoy,Frédéric Lerasle,Thomas Pellegrini,Julien Pinquier,Isabelle Ferrané

机构:Fr´ed´eric Lerasle, Isabelle Ferran´e,  IRIT, Universit´e de Toulouse, CNRS, INP Toulouse, UT, Toulouse, France,  LAAS-CNRS, UT, Toulouse, France

摘要:会议是专业环境中的一项常见活动,赋予声乐助理先进的功能以促进会议管理仍然是一项挑战。在这种情况下,像主动说话人检测这样的任务可以为会议参与者之间的交互建模提供有用的见解。受与高级会议助理相关的应用程序上下文的影响,我们希望将音频和视频信息结合起来,以实现最佳性能。在本文中,我们提出了两种不同类型的融合方法来检测主动说话人,通过神经网络将两种视觉模式和一种音频模式相结合。为了进行比较,还使用了经典的无监督音频特征提取方法。我们希望基于嘴唇和面部手势的检测,以每个参与者的面部为中心的视觉数据非常适合检测语音活动。因此,我们的基线系统使用视觉数据,我们选择了一种3D卷积神经网络结构,它可以有效地同时编码外观和运动。为了改进这个系统,我们通过使用CNN或无监督的说话人日记系统处理音频流来补充视觉信息。我们进一步改进了该系统,通过光流运动添加视觉模态信息。我们使用一个公共和最先进的基准:AMI语料库来评估我们的提案。我们分析了每个系统对所进行合并的贡献,以确定某一特定参与者目前是否在发言。我们还讨论了我们得到的结果。此外,我们已经证明,对于我们的应用程序上下文,添加运动信息可以极大地提高性能。最后,我们证明了基于注意的融合在降低标准差的同时提高了性能。

摘要:Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide useful insights to model interaction between meeting participants. Motivated by our application context related to advanced meeting assistant, we want to combine audio and visual information to achieve the best possible performance. In this paper, we propose two different types of fusion for the detection of the active speaker, combining two visual modalities and an audio modality through neural networks. For comparison purpose, classical unsupervised approaches for audio feature extraction are also used. We expect visual data centered on the face of each participant to be very appropriate for detecting voice activity, based on the detection of lip and facial gestures. Thus, our baseline system uses visual data and we chose a 3D Convolutional Neural Network architecture, which is effective for simultaneously encoding appearance and movement. To improve this system, we supplemented the visual information by processing the audio stream with a CNN or an unsupervised speaker diarization system. We have further improved this system by adding visual modality information using motion through optical flow. We evaluated our proposal with a public and state-of-the-art benchmark: the AMI corpus. We analysed the contribution of each system to the merger carried out in order to determine if a given participant is currently speaking. We also discussed the results we obtained. Besides, we have shown that, for our application context, adding motion information greatly improves performance. Finally, we have shown that attention-based fusion improves performance while reducing the standard deviation.


【3】 Joint Analysis of Acoustic Scenes and Sound Events Based on Multitask  Learning with Dynamic Weight Adaptation

标题:基于动态权值自适应多任务学习的声场景与声事件联合分析

链接:https://arxiv.org/abs/2206.10349

作者:Kayo Nada,Keisuke Imoto,Takao Tsuchiya
机构:Doshisha University, Japan.
备注:Submitted to Acoustical Science and Technology
摘要:声场景分类(ASC)和声事件检测(SED)是环境声分析中的主要课题。考虑到声场景和声音事件之间有着密切的联系,以前的一些工作提出了使用基于多任务学习(MTL)的神经网络对声场景和声音事件进行联合分析。传统方法使用具有恒定权重的ASC和SED损失函数的线性组合来训练基于MTL的模型。然而,基于MTL的传统方法的性能在很大程度上取决于ASC和SED损失的权重,很难确定ASC和SED的MTL损失的恒定权重之间的适当平衡。在本文中,我们提出了基于动态权重平均和多焦点损失的ASC和SED的MTL动态权重调整方法,以自动调整学习权重。使用部分2016/2017年TUT声学场景和2016/2017年TUT声音事件进行了评估实验,结果表明,与传统的基于MTL的方法相比,所提出的方法提高了场景分类和事件检测性能。然后,我们研究了ASC和SED任务的学习权重如何随着模型训练的进行而动态调整。
摘要:Acoustic scene classification (ASC) and sound event detection (SED) are major topics in environmental sound analysis. Considering that acoustic scenes and sound events are closely related to each other, the joint analysis of acoustic scenes and sound events using multitask learning (MTL)-based neural networks was proposed in some previous works. Conventional methods train MTL-based models using a linear combination of ASC and SED loss functions with constant weights. However, the performance of conventional MTL-based methods depends strongly on the weights of the ASC and SED losses, and it is difficult to determine the appropriate balance between the constant weights of the losses of MTL of ASC and SED. In this paper, we thus propose dynamic weight adaptation methods for MTL of ASC and SED based on dynamic weight average and multi--focal loss to adjust the learning weights automatically. Evaluation experiments using parts of the TUT Acoustic Scenes 2016/2017 and TUT Sound Events 2016/2017 are conducted, and we show that the proposed methods improve the scene classification and event detection performance characteristics compared with the conventional MTL-based method. We then investigate how the learning weights of ASC and SED tasks dynamically adapt as the model training progresses.


【4】 Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTS

标题:基于DNN的多说话人TTS中的人在环说话人自适应

链接:https://arxiv.org/abs/2206.10256

作者:Kenta Udagawa,Yuki Saito,Hiroshi Saruwatari

机构:Graduate School of Information Science and Technology, The University of Tokyo, Japan.

备注:5 pages, 3 figures, Accepted for INTERSPEECH2022

摘要:提出了一种多说话人文语转换的人在回路说话人自适应方法。采用传统的说话人自适应方法,利用说话人识别任务训练的说话人编码器,从参考语音中提取目标说话人的嵌入向量。然而,当参考语音不可用时,该方法无法获得目标说话人的嵌入向量。我们的方法基于人在回路优化框架,该框架包含用户探索说话人嵌入空间以找到目标说话人的嵌入。该方法使用了一种序列线搜索算法,该算法反复要求用户在嵌入空间的线段上选择一个点。为了有效地从多个刺激中选择最佳语音样本,我们还开发了一个系统,在该系统中,用户可以针对每个音素在多个说话人的语音之间切换,同时循环一个话语。实验结果表明,即使不直接将参考语音作为说话人编码器的输入,该方法在客观和主观评价方面都能达到与传统方法相当的性能。

摘要:This paper proposes a human-in-the-loop speaker-adaptation method for multi-speaker text-to-speech. With a conventional speaker-adaptation method, a target speaker's embedding vector is extracted from his/her reference speech using a speaker encoder trained on a speaker-discriminative task. However, this method cannot obtain an embedding vector for the target speaker when the reference speech is unavailable. Our method is based on a human-in-the-loop optimization framework, which incorporates a user to explore the speaker-embedding space to find the target speaker's embedding. The proposed method uses a sequential line search algorithm that repeatedly asks a user to select a point on a line segment in the embedding space. To efficiently choose the best speech sample from multiple stimuli, we also developed a system in which a user can switch between multiple speakers' voices for each phoneme while looping an utterance. Experimental results indicate that the proposed method can achieve comparable performance to the conventional one in objective and subjective evaluations even if reference speech is not used as the input of a speaker encoder directly.


【5】 Incorporating Voice Instructions in Model-Based Reinforcement Learning  for Self-Driving Cars

标题:自动驾驶汽车模型强化学习中引入语音指令的研究

链接:https://arxiv.org/abs/2206.10249

作者:Mingze Wang,Ziyang Zhang,Grace Hui Yang

机构:InfoSense, Department of Computer Science, Georgetown University, United States

备注:NeurIPS 2021 Workshop on Machine Learning for Autonomous Driving

摘要:该文提出了一种支持自然语言语音指令的新方法,用于指导自动驾驶汽车的深度强化学习(DRL)算法。DRL方法是自主车辆(AV)代理的常用方法。然而,大多数现有的方法都缺乏样本和时间效率,并且缺乏与人类专家的自然沟通渠道。在本文中,新的人类驾驶员如何从人类教练那里学习激励我们研究人在回路学习的新方法,以及为代理人提供更自然、更易于接近的训练界面。我们建议将自然语言语音指令(NLI)纳入基于模型的深度强化学习中,以训练自动驾驶汽车。我们在CARLA模拟器中评估了所提出的方法以及一些最先进的DRL方法。结果表明,NLI可以帮助简化训练过程,显著提高Agent的学习速度。

摘要:This paper presents a novel approach that supports natural language voice instructions to guide deep reinforcement learning (DRL) algorithms when training self-driving cars. DRL methods are popular approaches for autonomous vehicle (AV) agents. However, most existing methods are sample- and time-inefficient and lack a natural communication channel with the human expert. In this paper, how new human drivers learn from human coaches motivates us to study new ways of human-in-the-loop learning and a more natural and approachable training interface for the agents. We propose incorporating natural language voice instructions (NLI) in model-based deep reinforcement learning to train self-driving cars. We evaluate the proposed method together with a few state-of-the-art DRL methods in the CARLA simulator. The results show that NLI can help ease the training process and significantly boost the agents' learning speed.


【6】 Analysis of Self-Supervised Learning and Dimensionality Reduction  Methods in Clustering-Based Active Learning for Speech Emotion Recognition

标题:语音情感识别中基于聚类的主动学习中的自监督学习和降维方法分析

链接:https://arxiv.org/abs/2206.10188

作者:Einari Vaaras,Manu Airaksinen,Okko Räsänen

机构:Unit of Computing Sciences, Tampere University, Finland, Helsinki University Hospital, Helsinki, Finland

备注:To be published in Proc. Interspeech 2022, Incheon, South Korea

摘要:当需要领域专家为复杂的机器学习任务执行数据注释时,为了减少时间和费用,减少注释工作量至关重要。对于没有可用注释的情况,一种方法是利用特征空间的结构进行基于聚类的主动学习(AL)方法。然而,这些方法在很大程度上取决于样本在特征空间中的组织方式以及使用的距离度量。对比预测编码(CPC)等无监督方法有可能用于学习有组织的特征空间,但这些方法通常会创建高维特征,这可能对估计数据密度具有挑战性。在本文中,我们将CPC和多维降维方法相结合,搜索基于聚类的AL的功能实践。我们模拟语音情感识别系统部署的实验表明,特征空间的局部和全局拓扑都可以成功地用于AL,与传统信号特征相比,CPC可用于改善基于聚类的AL性能。此外,我们观察到压缩数据维度不会实质上损害AL性能,并且当注释数量不是很低时,二维特征表示与高维表示实现了类似的AL性能。

摘要:When domain experts are needed to perform data annotation for complex machine-learning tasks, reducing annotation effort is crucial in order to cut down time and expenses. For cases when there are no annotations available, one approach is to utilize the structure of the feature space for clustering-based active learning (AL) methods. However, these methods are heavily dependent on how the samples are organized in the feature space and what distance metric is used. Unsupervised methods such as contrastive predictive coding (CPC) can potentially be used to learn organized feature spaces, but these methods typically create high-dimensional features which might be challenging for estimating data density. In this paper, we combine CPC and multiple dimensionality reduction methods in search of functioning practices for clustering-based AL. Our experiments for simulating speech emotion recognition system deployment show that both the local and global topology of the feature space can be successfully used for AL, and that CPC can be used to improve clustering-based AL performance over traditional signal features. Additionally, we observe that compressing data dimensionality does not harm AL performance substantially, and that 2-D feature representations achieved similar AL performance as higher-dimensional representations when the number of annotations is not very low.


【7】 A Multi-grained based Attention Network for Semi-supervised Sound Event  Detection

标题:基于多粒度注意力网络的半监督声音事件检测

链接:https://arxiv.org/abs/2206.10175

作者:Ying Hu,Xiujuan Zhu,Yunlong Li,Hao Huang,Liang He

机构:Key Laboratory of signal detection and processing in Xinjiang, China, School of Information Science and Engineering, Xinjiang University, Urumqi, China, Department of Electronic Engineering, Tsinghua University, China

备注:None

摘要:由于现实生活中数据的稀缺性和声音事件的多样性,声音事件检测是一项有趣但富有挑战性的任务。提出了一种基于多粒度注意网络的半监督声音事件检测方法。为了获得与声音事件相关的特征表示,设计了一个残差混合卷积(RH-Conv)块来提高普通卷积提取时频特征的能力。此外,还设计了一个多粒度注意(MGA)模块,用于从粗到细学习时间分辨率特征。借助MGA模块,网络可以捕获持续时间短或长的目标事件的特征,从而更准确地确定声音事件的开始和偏移。此外,为了有效提高平均教师(MT)方法的性能,引入了空间移位(SS)模块作为数据扰动机制,以增加数据的多样性。实验结果表明,MGA网络在验证集和公共集上的表现优于已发布的最新竞争对手,分别达到53.27%和56.96%的基于事件的宏F1(EB-F1)分数,0.709和0.739的和声检测分数(PSDS)。

摘要:Sound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. This paper presents a multi-grained based attention network (MGA-Net) for semi-supervised sound event detection. To obtain the feature representations related to sound events, a residual hybrid convolution (RH-Conv) block is designed to boost the vanilla convolution's ability to extract the time-frequency features. Moreover, a multi-grained attention (MGA) module is designed to learn temporal resolution features from coarse-level to fine-level. With the MGA module,the network could capture the characteristics of target events with short- or long-duration, resulting in more accurately determining the onset and offset of sound events. Furthermore, to effectively boost the performance of the Mean Teacher (MT) method, a spatial shift (SS) module as a data perturbation mechanism is introduced to increase the diversity of data. Experimental results show that the MGA-Net outperforms the published state-of-the-art competitors, achieving 53.27% and 56.96% event-based macro F1 (EB-F1) score, 0.709 and 0.739 polyphonic sound detection score (PSDS) on the validation and public set respectively.


【8】 Supervision-Guided Codebooks for Masked Prediction in Speech  Pre-training

标题:语音预训练中监督引导的掩蔽预测码本

链接:https://arxiv.org/abs/2206.10125

作者:Chengyi Wang,Yiming Wang,Yu Wu,Sanyuan Chen,Jinyu Li,Shujie Liu,Furu Wei
机构:Nankai University, Microsoft
备注:To appear in Proc. Interspeech 2022
摘要:近年来,掩蔽预测预训练在语音识别的自监督学习(SSL)中取得了显著的进展。它通常需要以无监督的方式获取码本,从而使其不太准确且难以解释。为了提高自动语音识别(ASR)的性能和预训练效率,我们提出了两种监督引导码书生成方法,一种是使用混合ASR系统进行解码以生成音素级对齐(称为PBERT),另一种是对从端到端CTC模型中提取的监督语音特征进行聚类(称为CTC聚类)。混合模型和CTC模型都是在微调中使用的少量标记语音上进行训练的。实验表明,与各种SSL和自训练基线相比,我们的方法具有显著的优越性,相对功耗降低了17.0%。我们的预训练模型在非ASR语音任务中也表现出良好的可迁移性。
摘要:Recently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition. It usually requires a codebook obtained in an unsupervised way, making it less accurate and difficult to interpret. We propose two supervision-guided codebook generation approaches to improve automatic speech recognition (ASR) performance and also the pre-training efficiency, either through decoding with a hybrid ASR system to generate phoneme-level alignments (named PBERT), or performing clustering on the supervised speech features extracted from an end-to-end CTC model (named CTC clustering). Both the hybrid and CTC models are trained on the same small amount of labeled speech as used in fine-tuning. Experiments demonstrate significant superiority of our methods to various SSL and self-training baselines, with up to 17.0% relative WER reduction. Our pre-trained models also show good transferability in a non-ASR speech task.

【9】 WOLONet: Wave Outlooker for Efficient and High Fidelity Speech Synthesis

标题:WOLONet:高效高保真语音合成的WAVE预测器

链接:https://arxiv.org/abs/2206.09920

作者:Yi Wang,Yi Si
机构:Scybi
摘要:最近,基于GAN的神经声码器,如并行WaveGAN、MelGAN、HiFiGAN和UnivNet,由于其轻量级和并行结构而变得流行,即使在CPU上也能获得高保真度的实时合成波形。HiFiGAN和UnivNet是两个SOTA声码器。尽管质量很高,但仍有改进的余地。本文基于计算机视觉中视觉观察器的结构,采用了类似的思想,提出了一种高效、轻量级的神经声码器WOLONet。在该网络中,我们开发了一种新的轻量级块,该块使用位置变量、通道独立、深度动态卷积核以及正弦激活的动态核权重。为了证明我们方法的有效性和通用性,我们进行了烧蚀研究,以验证我们的新设计,并与典型的GAN声码器进行了主客观比较。结果表明,与HiFiGAN和UnivNet这两种神经SOTA声码器相比,我们的WOLONet在需要较少参数的情况下实现了最佳的生成质量。
摘要:Recently, GAN-based neural vocoders such as Parallel WaveGAN, MelGAN, HiFiGAN, and UnivNet have become popular due to their lightweight and parallel structure, resulting in a real-time synthesized waveform with high fidelity, even on a CPU. HiFiGAN and UnivNet are two SOTA vocoders. Despite their high quality, there is still room for improvement. In this paper, motivated by the structure of Vision Outlooker from computer vision, we adopt a similar idea and propose an effective and lightweight neural vocoder called WOLONet. In this network, we develop a novel lightweight block that uses a location-variable, channel-independent, and depthwise dynamic convolutional kernel with sinusoidally activated dynamic kernel weights. To demonstrate the effectiveness and generalizability of our method, we perform an ablation study to verify our novel design and make a subjective and objective comparison with typical GAN-based vocoders. The results show that our WOLONet achieves the best generation quality while requiring fewer parameters than the two neural SOTA vocoders, HiFiGAN and UnivNet.

【10】 The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic  Speech Recognition

标题:Makerere无线电语料库:用于自动语音识别的卢甘达无线电语料库

链接:https://arxiv.org/abs/2206.09790

作者:Jonathan Mukiibi,Andrew Katumba,Joyce Nakatumba-Nabende,Ali Hussein,Josh Meyer

机构:Makerere University, Ronin Institute, Coqui, Uganda, Egypt, USA

备注:Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022), pages 1945 to 1954 Marseille, 20 to 25 June 2022

摘要:对于资源不足的语言来说,建立一个可用的无线电监测自动语音识别(ASR)系统是一项具有挑战性的任务,然而在无线电是公共通信和讨论的主要媒介的社会中,这一点至关重要。联合国在乌干达的初步努力证明,了解被排除在社交媒体之外的农村人的看法在国家规划中是多么重要。然而,这些努力受到缺乏转录语音数据集的挑战。在本文中,Makerere人工智能研究实验室发布了一个155小时的Luganda无线电语音语料库。据我们所知,这是撒哈拉以南非洲第一个公开的无线电数据集。本文描述了语音语料库的开发,并使用开源语音识别工具包Coqui STT toolkit给出了基线Luganda ASR性能结果。

摘要:Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the United Nations in Uganda have proved how understanding the perceptions of rural people who are excluded from social media is important in national planning. However, these efforts are being challenged by the absence of transcribed speech datasets. In this paper, The Makerere Artificial Intelligence research lab releases a Luganda radio speech corpus of 155 hours. To our knowledge, this is the first publicly available radio dataset in sub-Saharan Africa. The paper describes the development of the voice corpus and presents baseline Luganda ASR performance results using Coqui STT toolkit, an open source speech recognition toolkit.


【11】 GMM based multi-stage Wiener filtering for low SNR speech enhancement

标题:基于GMM的多级维纳滤波用于低信噪比语音增强

链接:https://arxiv.org/abs/2206.09298

作者:Wageesha Manamperi,Prasanga N. Samarasinghe,Thushara D. Abhayapala,Jihui Zhang

机构:The Australian National University, Canberra, Australia

备注:5 pages, 3 figures, submitted to a conference

摘要:本文提出了一种单通道语音增强方法,在低信噪比和非平稳噪声条件下降低噪声并增强语音。具体而言,我们将重点放在使用基于带参数维纳滤波器的多阶段过程的高斯混合模型(GMM)建模噪声。与传统的维纳滤波方法相比,所提出的噪声模型能够估计出更精确的噪声功率谱密度(PSD),并在各种噪声条件下具有更好的泛化能力。仿真结果表明,在低信噪比下,该方法在语音质量(PESQ)和可懂度(STOI)方面都能取得较好的性能。

摘要:This paper proposes a single-channel speech enhancement method to reduce the noise and enhance speech at low signal-to-noise ratio (SNR) levels and non-stationary noise conditions. Specifically, we focus on modeling the noise using a Gaussian mixture model (GMM) based on a multi-stage process with a parametric Wiener filter. The proposed noise model estimates a more accurate noise power spectral density (PSD), and allows for better generalization under various noise conditions compared to traditional Wiener filtering methods. Simulations show that the proposed approach can achieve better performance in terms of speech quality (PESQ) and intelligibility (STOI) at low SNR levels.

【12】 Redundancy Reduction Twins Network: A Training framework for  Multi-output Emotion Regression

标题:冗余双生网络:一种多输出情感回归的训练框架

链接:https://arxiv.org/abs/2206.09142

作者:Xin Jing,Meishu Song,Andreas Triantafyllopoulos,Zijiang Yang,Björn W. Schuller

机构:University of Augsburg

备注:5 pages, accepted by ICML Exvo workshop

摘要:在本文中,我们提出了冗余缩减双生网络(RRTN),这是一种冗余缩减训练框架,通过测量同一网络输出之间的互相关矩阵,并将其与失真版本的样本进行馈送,使其尽可能接近单位矩阵,从而将冗余降至最低。RRTN还应用了一种新的损失函数,即巴洛双胞胎损失函数,以帮助最大化从样本的不同扭曲版本获得的表示的相似性。然而,由于损失的分布可能会导致网络中的性能波动,我们还建议使用受限不确定性权重损失(RUWL)或联合训练来确定损失函数的最佳权重。我们对CNN14提出的最佳方法是在ExVo多任务开发集上获得0.678的CCC对情绪的回归,比普通的CNN14 CCC 0.647增加4.8%,这在95%置信区间(双尾)下取得了显著差异。

摘要:In this paper, we propose the Redundancy Reduction Twins Network (RRTN), a redundancy reduction training framework that minimizes redundancy by measuring the cross-correlation matrix between the outputs of the same network fed with distorted versions of a sample and bringing it as close to the identity matrix as possible. RRTN also applies a new loss function, the Barlow Twins loss function, to help maximize the similarity of representations obtained from different distorted versions of a sample. However, as the distribution of losses can cause performance fluctuations in the network, we also propose the use of a Restrained Uncertainty Weight Loss (RUWL) or joint training to identify the best weights for the loss function. Our best approach on CNN14 with the proposed methodology obtains a CCC over emotion regression of 0.678 on the ExVo Multi-task dev set, a 4.8% increase over a vanilla CNN 14 CCC of 0.647, which achieves a significant difference at the 95% confidence interval (2-tailed).


【13】 Tackling Spoofing-Aware Speaker Verification with Multi-Model Fusion

标题:利用多模型融合解决感知欺骗的说话人确认问题

链接:https://arxiv.org/abs/2206.09131

作者:Haibin Wu,Jiawen Kang,Lingwei Meng,Yang Zhang,Xixin Wu,Zhiyong Wu,Hung-yi Lee,Helen Meng
机构:Graduate Institute of Communication Engineering, National Taiwan University,  Centre for Perceptual and Interactive Intelligence, The Chinese University of Hong Kong,  Human-Computer Communications Laboratory, The Chinese University of Hong Kong
备注:Accepted by Odyssey 2022
摘要:近年来,自动说话人验证(ASV)技术得到了长足的发展。然而,以往的研究表明,最先进的ASV模型极易受到语音欺骗攻击,最近提出的高性能欺骗对抗(CM)模型只关注独立的反欺骗任务,而忽略了后续的说话人验证过程。如何将CM和ASV集成在一起仍然是一个悬而未决的问题。最近出现了一个防欺骗说话人验证(SASV)挑战,理由是当CM和ASV子系统联合优化时,可以提供更好的性能。在挑战的场景下,参与者提出的集成系统需要同时拒绝冒名顶替者说话人和目标说话人的欺骗攻击,这直观有效地符合可靠、欺骗鲁棒ASV系统的期望。这项工作侧重于基于融合的SASV解决方案,并提出了一个多模型融合框架,以利用多个最先进的ASV和CM模型的强大功能。提出的框架将SASV-EER从8.75%大幅提高到1.17%,与SASV挑战中的最佳基线系统相比,相对提高了86%。
摘要:Recent years have witnessed the extraordinary development of automatic speaker verification (ASV). However, previous works show that state-of-the-art ASV models are seriously vulnerable to voice spoofing attacks, and the recently proposed high-performance spoofing countermeasure (CM) models only focus solely on the standalone anti-spoofing tasks, and ignore the subsequent speaker verification process. How to integrate the CM and ASV together remains an open question. A spoofing aware speaker verification (SASV) challenge has recently taken place with the argument that better performance can be delivered when both CM and ASV subsystems are optimized jointly. Under the challenge's scenario, the integrated systems proposed by the participants are required to reject both impostor speakers and spoofing attacks from target speakers, which intuitively and effectively matches the expectation of a reliable, spoofing-robust ASV system. This work focuses on fusion-based SASV solutions and proposes a multi-model fusion framework to leverage the power of multiple state-of-the-art ASV and CM models. The proposed framework vastly improves the SASV-EER from 8.75% to 1.17\%, which is 86% relative improvement compared to the best baseline system in the SASV challenge.

【14】 Boosting Cross-Domain Speech Recognition with Self-Supervision

标题:利用自我监督提高跨域语音识别能力

链接:https://arxiv.org/abs/2206.09783

作者:Han Zhu,Gaofeng Cheng,Jindong Wang,Wenxin Hou,Pengyuan Zhang,Yonghong Yan

机构:andalso with the University of Chinese Academy of Sciences

摘要:由于训练分布和测试分布之间的不匹配,自动语音识别(ASR)的跨域性能可能会受到严重影响。由于目标域通常缺乏标记数据,并且在声学和语言层面上存在域转移,因此对ASR进行无监督域适配(UDA)是一个挑战。之前的工作表明,通过利用未标记数据的自我监督,自我监督学习(SSL)或伪标记(PL)在UDA中是有效的。然而,这些自我监督在不匹配的域分布中也面临性能下降的问题,而以前的工作未能解决这一问题。这项工作提出了一个系统的UDA框架,以充分利用未标记数据,并在预训练和微调范式中进行自我监督。一方面,我们应用持续的预训练和数据回放技术来缓解SSL预训练模型的域不匹配。另一方面,我们提出了一种基于PL技术的域自适应微调方法,并进行了三个独特的改进:首先,我们设计了一种双分支PL方法,以降低对错误伪标签的敏感性;其次,我们设计了一种不确定性感知的置信度过滤策略来提高伪标签的正确性;第三,我们引入了一种两步PL方法来整合目标域语言知识,从而生成更准确的目标域伪标签。在各种跨域场景下的实验结果表明,该方法可以有效地提高跨域性能,并显著优于以前的方法。

摘要:The cross-domain performance of automatic speech recognition (ASR) could be severely hampered due to the mismatch between training and testing distributions. Since the target domain usually lacks labeled data, and domain shifts exist at acoustic and linguistic levels, it is challenging to perform unsupervised domain adaptation (UDA) for ASR. Previous work has shown that self-supervised learning (SSL) or pseudo-labeling (PL) is effective in UDA by exploiting the self-supervisions of unlabeled data. However, these self-supervisions also face performance degradation in mismatched domain distributions, which previous work fails to address. This work presents a systematic UDA framework to fully utilize the unlabeled data with self-supervision in the pre-training and fine-tuning paradigm. On the one hand, we apply continued pre-training and data replay techniques to mitigate the domain mismatch of the SSL pre-trained model. On the other hand, we propose a domain-adaptive fine-tuning approach based on the PL technique with three unique modifications: Firstly, we design a dual-branch PL method to decrease the sensitivity to the erroneous pseudo-labels; Secondly, we devise an uncertainty-aware confidence filtering strategy to improve pseudo-label correctness; Thirdly, we introduce a two-step PL approach to incorporate target domain linguistic knowledge, thus generating more accurate target domain pseudo-labels. Experimental results on various cross-domain scenarios demonstrate that the proposed approach could effectively boost the cross-domain performance and significantly outperform previous approaches.

【15】 An Empirical Analysis on the Vulnerabilities of End-to-End Speech  Segregation Models

标题:端到端语音分离模型脆弱性的实证分析

链接:https://arxiv.org/abs/2206.09556

作者:Rahil Parikh,Gaspar Rochette,Carol Espy-Wilson,Shihab Shamma

机构:Institute for Systems Research, University of Maryland College Park, USA, ENS Paris, PSL Universit´e, France

备注:Accepted at Interspeech 2022

摘要:端到端学习模型在执行语音分离方面表现出了非凡的能力。尽管它们在现实世界中的应用范围很广,但人们对它们用于分组和隔离单个说话者的机制知之甚少。知道和谐度是这些网络对源进行分组的关键线索,在这项工作中,我们对ConvTasnet和DPT网络进行了彻底的调查,以分析它们如何对输入混合进行和谐分析。我们进行烧蚀研究,应用不同通带的低通、高通和带阻滤波器,对隔离最关键的谐波进行经验分析。我们还研究了这些网络如何通过在合成混合物中引入不连续性来决定将哪个输出信道分配给估计源。我们发现,端到端网络非常不稳定,并且在面对人类无法察觉的变形时表现不佳。用频谱图替换这些网络中的编码器会导致整体性能降低,但稳定性更高。这项工作有助于我们理解这些网络依赖于什么信息来进行语音分离,并揭示了泛化错误的两个来源。它还将编码器定位为负责这些错误的网络的一部分,允许使用专家知识或转移学习进行重新设计。

摘要:End-to-end learning models have demonstrated a remarkable capability in performing speech segregation. Despite their wide-scope of real-world applications, little is known about the mechanisms they employ to group and consequently segregate individual speakers. Knowing that harmonicity is a critical cue for these networks to group sources, in this work, we perform a thorough investigation on ConvTasnet and DPT-Net to analyze how they perform a harmonic analysis of the input mixture. We perform ablation studies where we apply low-pass, high-pass, and band-stop filters of varying pass-bands to empirically analyze the harmonics most critical for segregation. We also investigate how these networks decide which output channel to assign to an estimated source by introducing discontinuities in synthetic mixtures. We find that end-to-end networks are highly unstable, and perform poorly when confronted with deformations which are imperceptible to humans. Replacing the encoder in these networks with a spectrogram leads to lower overall performance, but much higher stability. This work helps us to understand what information these network rely on for speech segregation, and exposes two sources of generalization-errors. It also pinpoints the encoder as the part of the network responsible for these errors, allowing for a redesign with expert knowledge or transfer learning.


【16】 Towards Trustworthy Edge Intelligence: Insights from Voice-Activated  Services

标题:走向值得信赖的边缘情报:来自语音激活服务的见解

链接:https://arxiv.org/abs/2206.09523

作者:W. T. Hutiri,A. Y. Ding

机构:Insights from Voice-Activated ServicesWiebke Toussaint HutiriEngineering Systems & ServicesDelft University of TechnologyDelft,  The Netherlands0000-000 2-96 57-9 509Aaron Yi DingEngineering Systems & ServicesDelft University of TechnologyDelft

摘要:在一个监控资本主义的时代,将新兴智能服务的设计锚定在可信的基础上是紧迫而重要的。边缘智能将人工智能和边缘计算领域结合在一起,是智能服务的关键使能技术。因此,值得信赖的边缘智能应该成为优先研究关注的问题。然而,决定边缘情报的可信度并不是一件简单的事。本文研究了语音激活服务的具体应用场景中对可信边缘智能的需求。我们有助于从三个方面加深对新兴边缘智能领域可信度的理解:首先,我们提出了一个统一的可信边缘智能框架,该框架共同考虑了AI和物联网的可信度属性。其次,我们展示了语音激活服务中一个具体案例研究的研究成果,该案例展示了三个重要可信属性之间的相互依赖性:隐私、安全和公平。第三,基于实证和分析结果,我们强调了值得信赖的边缘智能未来重要研究领域面临的挑战和开放性问题。

摘要:In an age of surveillance capitalism, anchoring the design of emerging smart services in trustworthiness is urgent and important. Edge Intelligence, which brings together the fields of AI and Edge computing, is a key enabling technology for smart services. Trustworthy Edge Intelligence should thus be a priority research concern. However, determining what makes Edge Intelligence trustworthy is not straight forward. This paper examines requirements for trustworthy Edge Intelligence in a concrete application scenario of voice-activated services. We contribute to deepening the understanding of trustworthiness in the emerging Edge Intelligence domain in three ways: firstly, we propose a unified framing for trustworthy Edge Intelligence that jointly considers trustworthiness attributes of AI and the IoT. Secondly, we present research outputs of a tangible case study in voice-activated services that demonstrates interdependencies between three important trustworthiness attributes: privacy, security and fairness. Thirdly, based on the empirical and analytical findings, we highlight challenges and open questions that present important future research areas for trustworthy Edge Intelligence.


【17】 Resource-Efficient Separation Transformer

标题:节能型分离Transformer

链接:https://arxiv.org/abs/2206.09507

作者:Cem Subakan,Mirco Ravanelli,Samuele Cornell,Frédéric Lepoutre,François Grondin
机构:Universit´e de Sherbrooke,  Mila-Quebec AI Institute,  Concordia University,  Universit´e de Montr´eal, Universita Politecnica delle Marche
备注:Submitted to IEEE Signal Processing Letters
摘要:Transformer最近在语音分离方面取得了最先进的性能。然而,这些模型的计算要求很高,需要很多可学习的参数。本文探讨了一种基于变换器的语音分离方法,可以降低计算量。我们的主要贡献是开发了资源高效的分离转换器(RE-SepFormer),这是一种基于自我关注的架构,可以从两个方面减少计算负担。首先,在潜在空间中使用非重叠块。其次,它对从每个区块计算出的紧凑潜在摘要进行操作。RE-Sepreformer在流行的WSJ0-2Mix和WHAM上达到了有竞争力的性能!因果和非因果环境中的数据集。值得注意的是,在内存和推理时间方面,它的扩展性明显优于以前基于Transformer和RNN的体系结构,因此更适合处理长混合。
摘要:Transformers have recently achieved state-of-the-art performance in speech separation. These models, however, are computationally-demanding and require a lot of learnable parameters. This paper explores Transformer-based speech separation with a reduced computational cost. Our main contribution is the development of the Resource-Efficient Separation Transformer (RE-SepFormer), a self-attention-based architecture that reduces the computational burden in two ways. First, it uses non-overlapping blocks in the latent space. Second, it operates on compact latent summaries calculated from each chunk. The RE-SepFormer reaches a competitive performance on the popular WSJ0-2Mix and WHAM! datasets in both causal and non-causal settings. Remarkably, it scales significantly better than the previous Transformer and RNN-based architectures in terms of memory and inference-time, making it more suitable for processing long mixtures.


【18】 Transfer Learning for Robust Low-Resource Children's Speech ASR with  Transformers and Source-Filter Warping

标题:基于变换和信源滤波的稳健低资源儿童语音ASR转移学习

链接:https://arxiv.org/abs/2206.09396

作者:Jenthe Thienpondt,Kris Demuynck

机构:IDLab, Department of Electronics and Information Systems, Ghent University - imec, Belgium

备注:proceedings of INTERSPEECH 2022

摘要:众所周知,自动语音识别(ASR)系统在转录儿童语音时会遇到困难。这主要是由于缺乏用于训练鲁棒ASR模型的大型儿童语音语料库,以及使用成人数据训练的系统解码儿童语音时产生的域不匹配。在本文中,我们提出了多项增强来缓解这些问题。首先,我们提出了一种基于语音源滤波模型的数据增强技术,以缩小成人和儿童语音的领域差距。这使我们能够利用成人语音语料库的数据可用性,使这些样本在感知上与儿童语音相似。其次,使用这种增强策略,我们将转移学习应用于成人数据预训练的Transformer模型。该模型遵循最近引入的XLS-R体系结构,这是一种wav2vec 2.0模型,在多个跨语言成人语音语料库上进行预训练,以学习通用和鲁棒的声学框架级表示。在使用成人数据的ASR任务中,采用该模型,再加上所提出的源过滤器扭曲策略和有限数量的域内儿童语音,显著优于之前在PF-STAR英国儿童英语语音语料库上的最新结果,在官方测试集上的WER为4.86%。

摘要:Automatic Speech Recognition (ASR) systems are known to exhibit difficulties when transcribing children's speech. This can mainly be attributed to the absence of large children's speech corpora to train robust ASR models and the resulting domain mismatch when decoding children's speech with systems trained on adult data. In this paper, we propose multiple enhancements to alleviate these issues. First, we propose a data augmentation technique based on the source-filter model of speech to close the domain gap between adult and children's speech. This enables us to leverage the data availability of adult speech corpora by making these samples perceptually similar to children's speech. Second, using this augmentation strategy, we apply transfer learning on a Transformer model pre-trained on adult data. This model follows the recently introduced XLS-R architecture, a wav2vec 2.0 model pre-trained on several cross-lingual adult speech corpora to learn general and robust acoustic frame-level representations. Adopting this model for the ASR task using adult data augmented with the proposed source-filter warping strategy and a limited amount of in-domain children's speech significantly outperforms previous state-of-the-art results on the PF-STAR British English Children's Speech corpus with a 4.86% WER on the official test set.


【19】 Decoupled Federated Learning for ASR with Non-IID Data

标题:基于非IID数据的ASR解耦联合学习

链接:https://arxiv.org/abs/2206.09102

作者:Han Zhu,Jindong Wang,Gaofeng Cheng,Pengyuan Zhang,Yonghong Yan

机构:Key Laboratory of Speech Acoustics and Content Understanding, Institute of Acoustics CAS, China,  University of Chinese Academy of Sciences, China,  Microsoft Research Asia, China

备注:Accepted by Interspeech 2022

摘要:使用联邦学习(FL)的自动语音识别(ASR)可以在不损害隐私的情况下利用来自多个客户端的数据。基于FL的ASR的质量可以通过识别性能、通信和计算成本来衡量。当不同客户端之间的数据不是独立且相同分布的(非IID)时,性能可能会显著降低。在这项工作中,我们使用个性化的FL来解决基于FL的ASR中的非IID问题,它为每个客户学习个性化的模型。具体而言,我们提出了两种针对ASR的个性化FL方法。首先,我们将基于FL的个性化层应用于ASR,它将一些层保留在本地以学习个性化模型。其次,为了降低通信和计算成本,我们提出了解耦联邦学习(DecoupleFL)。一方面,DecoupleFL将计算负担转移到服务器上,从而减少客户端的计算量。另一方面,DecoupleFL通信安全的高级特性,而不是模型参数,从而在模型较大时降低通信成本。实验表明,与FedAvg相比,两种基于FL的个性化ASR方法可以将WER降低2.3%~3.4%。其中,与FedAvg相比,DecoupleFL只需11.4%的通信量和75%的计算量,也明显低于基于个性化层的FL。

摘要:Automatic speech recognition (ASR) with federated learning (FL) makes it possible to leverage data from multiple clients without compromising privacy. The quality of FL-based ASR could be measured by recognition performance, communication and computation costs. When data among different clients are not independently and identically distributed (non-IID), the performance could degrade significantly. In this work, we tackle the non-IID issue in FL-based ASR with personalized FL, which learns personalized models for each client. Concretely, we propose two types of personalized FL approaches for ASR. Firstly, we adapt the personalization layer based FL for ASR, which keeps some layers locally to learn personalization models. Secondly, to reduce the communication and computation costs, we propose decoupled federated learning (DecoupleFL). On one hand, DecoupleFL moves the computation burden to the server, thus decreasing the computation on clients. On the other hand, DecoupleFL communicates secure high-level features instead of model parameters, thus reducing communication cost when models are large. Experiments demonstrate two proposed personalized FL-based ASR approaches could reduce WER by 2.3% - 3.4% compared with FedAvg. Among them, DecoupleFL has only 11.4% communication and 75% computation cost compared with FedAvg, which is also significantly less than the personalization layer based FL.


【20】 Semi-supervised Time Domain Target Speaker Extraction with Attention

标题:基于注意力的半监督时域目标说话人提取

链接:https://arxiv.org/abs/2206.09072

作者:Zhepei Wang,Ritwik Giri,Shrikant Venkataramani,Umut Isik,Jean-Marc Valin,Paris Smaragdis,Mike Goodwin,Arvindh Krishnaswamy

机构:♯University of Illinois at Urbana-Champaign, †Amazon Web Services

摘要:在这项工作中,我们提出了Exformer,一种用于目标说话人提取的时域体系结构。它由一个预先训练好的说话人嵌入网络和一个基于Transformer编码器块的分离网络组成。我们研究了多种将说话人信息与输入混合信息相结合的方法,得到的Exformer结构与以前的时域网络相比具有更好的提取性能。此外,我们还研究了一个两阶段的过程,在预训练的监督模型上使用无参考信号的混合信号来训练模型。实验结果表明,所提出的半监督学习方法提高了有监督基线的性能。

摘要:In this work, we propose Exformer, a time-domain architecture for target speaker extraction. It consists of a pre-trained speaker embedder network and a separator network based on transformer encoder blocks. We study multiple methods to combine speaker information with the input mixture, and the resulting Exformer architecture obtains superior extraction performance compared to prior time-domain networks. Furthermore, we investigate a two-stage procedure to train the model using mixtures without reference signals upon a pre-trained supervised model. Experimental results show that the proposed semi-supervised learning procedure improves the performance of the supervised baselines.


eess.AS音频处理

【1】 Boosting Cross-Domain Speech Recognition with Self-Supervision

标题:利用自我监督提高跨域语音识别能力

链接:https://arxiv.org/abs/2206.09783

* 与cs.SD语音【14】为同一篇

作者:Han Zhu,Gaofeng Cheng,Jindong Wang,Wenxin Hou,Pengyuan Zhang,Yonghong Yan
机构:andalso with the University of Chinese Academy of Sciences
摘要:由于训练分布和测试分布之间的不匹配,自动语音识别(ASR)的跨域性能可能会受到严重影响。由于目标域通常缺乏标记数据,并且在声学和语言层面上存在域转移,因此对ASR进行无监督域适配(UDA)是一个挑战。之前的工作表明,通过利用未标记数据的自我监督,自我监督学习(SSL)或伪标记(PL)在UDA中是有效的。然而,这些自我监督在不匹配的域分布中也面临性能下降的问题,而以前的工作未能解决这一问题。这项工作提出了一个系统的UDA框架,以充分利用未标记数据,并在预训练和微调范式中进行自我监督。一方面,我们应用持续的预训练和数据回放技术来缓解SSL预训练模型的域不匹配。另一方面,我们提出了一种基于PL技术的域自适应微调方法,并进行了三个独特的改进:首先,我们设计了一种双分支PL方法,以降低对错误伪标签的敏感性;其次,我们设计了一种不确定性感知的置信度过滤策略来提高伪标签的正确性;第三,我们引入了一种两步PL方法来整合目标域语言知识,从而生成更准确的目标域伪标签。在各种跨域场景下的实验结果表明,该方法可以有效地提高跨域性能,并显著优于以前的方法。
摘要:The cross-domain performance of automatic speech recognition (ASR) could be severely hampered due to the mismatch between training and testing distributions. Since the target domain usually lacks labeled data, and domain shifts exist at acoustic and linguistic levels, it is challenging to perform unsupervised domain adaptation (UDA) for ASR. Previous work has shown that self-supervised learning (SSL) or pseudo-labeling (PL) is effective in UDA by exploiting the self-supervisions of unlabeled data. However, these self-supervisions also face performance degradation in mismatched domain distributions, which previous work fails to address. This work presents a systematic UDA framework to fully utilize the unlabeled data with self-supervision in the pre-training and fine-tuning paradigm. On the one hand, we apply continued pre-training and data replay techniques to mitigate the domain mismatch of the SSL pre-trained model. On the other hand, we propose a domain-adaptive fine-tuning approach based on the PL technique with three unique modifications: Firstly, we design a dual-branch PL method to decrease the sensitivity to the erroneous pseudo-labels; Secondly, we devise an uncertainty-aware confidence filtering strategy to improve pseudo-label correctness; Thirdly, we introduce a two-step PL approach to incorporate target domain linguistic knowledge, thus generating more accurate target domain pseudo-labels. Experimental results on various cross-domain scenarios demonstrate that the proposed approach could effectively boost the cross-domain performance and significantly outperform previous approaches.


【2】 Multi-channel end-to-end neural network for speech enhancement, source localization, and voice activity detection

标题:用于语音增强、源定位和语音活动检测的多通道端到端神经网络

链接:https://arxiv.org/abs/2206.09728

作者:Yuan Chen,Yicheng Hsu,Mingsian R. Bai

机构:Department of Power Mechanical Engineering, National Tsing Hua University, Taiwan,  Electrical Engineering, National Tsing Hua University, Taiwan

备注:Accepted by ICA2022

摘要:几十年来,语音增强和源定位一直是人们研究的热点,在现实世界中有着广泛的应用。最近,深复卷积递归网络(DCCRN)在单信道系统中取得了令人印象深刻的增强性能。本文提出了一种由波束形成器和新型多通道DCCRN组成的神经波束形成器,用于语音增强和源定位。多信道DCCRN估计的复数滤波器作为波束形成器的权值。此外,还采用了基于一阶段学习的方法进行语音增强和源定位。该网络由多通道DCCRN和辅助网络组成,对声场进行建模,同时最小化无失真响应损失函数。仿真结果表明,该神经波束形成器能有效地增强语音信号,并能很好地保持语音质量。该神经波束形成器还提供了源定位和语音活动检测(VAD)功能。

摘要:Speech enhancement and source localization has been active research for several decades with a wide range of real-world applications. Recently, the Deep Complex Convolution Recurrent network (DCCRN) has yielded impressive enhancement performance for single-channel systems. In this study, a neural beamformer consisting of a beamformer and a novel multi-channel DCCRN is proposed for speech enhancement and source localization. Complex-valued filters estimated by the multi-channel DCCRN serve as the weights of beamformer. In addition, a one-stage learning-based procedure is employed for speech enhancement and source localization. The proposed network composed of the multi-channel DCCRN and the auxiliary network models the sound field, while minimizing the distortionless response loss function. Simulation results show that the proposed neural beamformer is effective in enhancing speech signals, with speech quality well preserved. The proposed neural beamformer also provides source localization and voice activity detection (VAD) functions.


【3】 An Empirical Analysis on the Vulnerabilities of End-to-End Speech  Segregation Models

标题:端到端语音分离模型脆弱性的实证分析

链接:https://arxiv.org/abs/2206.09556

* 与cs.SD语音【15】为同一篇


作者:Rahil Parikh,Gaspar Rochette,Carol Espy-Wilson,Shihab Shamma

机构:Institute for Systems Research, University of Maryland College Park, USA, ENS Paris, PSL Universit´e, France

备注:Accepted at Interspeech 2022

摘要:端到端学习模型在执行语音分离方面表现出了非凡的能力。尽管它们在现实世界中的应用范围很广,但人们对它们用于分组和隔离单个说话者的机制知之甚少。知道和谐度是这些网络对源进行分组的关键线索,在这项工作中,我们对ConvTasnet和DPT网络进行了彻底的调查,以分析它们如何对输入混合进行和谐分析。我们进行烧蚀研究,应用不同通带的低通、高通和带阻滤波器,对隔离最关键的谐波进行经验分析。我们还研究了这些网络如何通过在合成混合物中引入不连续性来决定将哪个输出信道分配给估计源。我们发现,端到端网络非常不稳定,并且在面对人类无法察觉的变形时表现不佳。用频谱图替换这些网络中的编码器会导致整体性能降低,但稳定性更高。这项工作有助于我们理解这些网络依赖于什么信息来进行语音分离,并揭示了泛化错误的两个来源。它还将编码器定位为负责这些错误的网络的一部分,允许使用专家知识或转移学习进行重新设计。

摘要:End-to-end learning models have demonstrated a remarkable capability in performing speech segregation. Despite their wide-scope of real-world applications, little is known about the mechanisms they employ to group and consequently segregate individual speakers. Knowing that harmonicity is a critical cue for these networks to group sources, in this work, we perform a thorough investigation on ConvTasnet and DPT-Net to analyze how they perform a harmonic analysis of the input mixture. We perform ablation studies where we apply low-pass, high-pass, and band-stop filters of varying pass-bands to empirically analyze the harmonics most critical for segregation. We also investigate how these networks decide which output channel to assign to an estimated source by introducing discontinuities in synthetic mixtures. We find that end-to-end networks are highly unstable, and perform poorly when confronted with deformations which are imperceptible to humans. Replacing the encoder in these networks with a spectrogram leads to lower overall performance, but much higher stability. This work helps us to understand what information these network rely on for speech segregation, and exposes two sources of generalization-errors. It also pinpoints the encoder as the part of the network responsible for these errors, allowing for a redesign with expert knowledge or transfer learning.


【4】 A Step Towards Preserving Speakers' Identity While Detecting Depression  Via Speaker Disentanglement

标题:通过说话人解缠检测抑郁的同时保持说话人身份的一步

链接:https://arxiv.org/abs/2206.09530

作者:Vijay Ravi,Jinhan Wang,Jonathan Flint,Abeer Alwan

机构:†Dept . of Electrical and Computer Engineering, University of California, Los Angeles, USA, §Dept. of Psychiatry and Biobehavioral Sciences, University of California, Los Angeles, USA

备注:Accepted to Interspeech 2022

摘要:对于基于语音的心理健康障碍自动诊断来说,保持患者身份是一项挑战。在这篇文章中,我们通过提出抑郁症特征和说话人身份的对抗性分离来解决这个问题。通过最小化抑郁预测损失和最大化说话人预测损失,以说话人身份不变的方式对用于抑郁分类的模型进行训练。该方法的有效性在两个数据集DAIC-WOZ(英语)和CONVERGE(普通话)上进行了验证,其中包含三个特征集(Mel谱图、原始音频信号和wav2vec2.0的最后隐藏状态),使用了一个改进的DepAudioNet模型。通过对抗性训练,与基线相比,抑郁分类在每个特征上都有所改善。具有对抗性学习的Wav2vec2.0功能的性能最好(DAIC-WOZ的F1成绩为69.2%,CONVERGE的成绩为91.5%)。对DepAudioNet模型隐藏状态的类可分性测度(J-ratio)的分析表明,当采用对抗性学习时,后端模型会失去一些说话人的可辨别性,同时会提高抑郁症的可辨别性。这些结果表明,说话人身份的某些组成部分可能对抑郁症的检测没有帮助,最小化它们的影响可以更准确地诊断潜在的障碍,并可以保护说话人的身份。

摘要:Preserving a patient's identity is a challenge for automatic, speech-based diagnosis of mental health disorders. In this paper, we address this issue by proposing adversarial disentanglement of depression characteristics and speaker identity. The model used for depression classification is trained in a speaker-identity-invariant manner by minimizing depression prediction loss and maximizing speaker prediction loss during training. The effectiveness of the proposed method is demonstrated on two datasets - DAIC-WOZ (English) and CONVERGE (Mandarin), with three feature sets (Mel-spectrograms, raw-audio signals, and the last-hidden-state of wav2vec2.0), using a modified DepAudioNet model. With adversarial training, depression classification improves for every feature when compared to the baseline. Wav2vec2.0 features with adversarial learning resulted in the best performance (F1-score of 69.2% for DAIC-WOZ and 91.5% for CONVERGE). Analysis of the class-separability measure (J-ratio) of the hidden states of the DepAudioNet model shows that when adversarial learning is applied, the backend model loses some speaker-discriminability while it improves depression-discriminability. These results indicate that there are some components of speaker identity that may not be useful for depression detection and minimizing their effects provides a more accurate diagnosis of the underlying disorder and can safeguard a speaker's identity.


【5】 Towards Trustworthy Edge Intelligence: Insights from Voice-Activated  Services

标题:走向值得信赖的边缘情报:来自语音激活服务的见解

链接:https://arxiv.org/abs/2206.09523

* 与cs.SD语音【16】为同一篇

作者:W. T. Hutiri,A. Y. Ding
机构:Insights from Voice-Activated ServicesWiebke Toussaint HutiriEngineering Systems & ServicesDelft University of TechnologyDelft,  The Netherlands0000-000 2-96 57-9 509Aaron Yi DingEngineering Systems & ServicesDelft University of TechnologyDelft
摘要:在一个监控资本主义的时代,将新兴智能服务的设计锚定在可信的基础上是紧迫而重要的。边缘智能将人工智能和边缘计算领域结合在一起,是智能服务的关键使能技术。因此,值得信赖的边缘智能应该成为优先研究关注的问题。然而,决定边缘情报的可信度并不是一件简单的事。本文研究了语音激活服务的具体应用场景中对可信边缘智能的需求。我们有助于从三个方面加深对新兴边缘智能领域可信度的理解:首先,我们提出了一个统一的可信边缘智能框架,该框架共同考虑了AI和物联网的可信度属性。其次,我们展示了语音激活服务中一个具体案例研究的研究成果,该案例展示了三个重要可信属性之间的相互依赖性:隐私、安全和公平。第三,基于实证和分析结果,我们强调了值得信赖的边缘智能未来重要研究领域面临的挑战和开放性问题。
摘要:In an age of surveillance capitalism, anchoring the design of emerging smart services in trustworthiness is urgent and important. Edge Intelligence, which brings together the fields of AI and Edge computing, is a key enabling technology for smart services. Trustworthy Edge Intelligence should thus be a priority research concern. However, determining what makes Edge Intelligence trustworthy is not straight forward. This paper examines requirements for trustworthy Edge Intelligence in a concrete application scenario of voice-activated services. We contribute to deepening the understanding of trustworthiness in the emerging Edge Intelligence domain in three ways: firstly, we propose a unified framing for trustworthy Edge Intelligence that jointly considers trustworthiness attributes of AI and the IoT. Secondly, we present research outputs of a tangible case study in voice-activated services that demonstrates interdependencies between three important trustworthiness attributes: privacy, security and fairness. Thirdly, based on the empirical and analytical findings, we highlight challenges and open questions that present important future research areas for trustworthy Edge Intelligence.

【6】 Resource-Efficient Separation Transformer

标题:节能型分离Transformer

链接:https://arxiv.org/abs/2206.09507

* 与cs.SD语音【17】为同一篇

作者:Cem Subakan,Mirco Ravanelli,Samuele Cornell,Frédéric Lepoutre,François Grondin

机构:Universit´e de Sherbrooke,  Mila-Quebec AI Institute,  Concordia University,  Universit´e de Montr´eal, Universita Politecnica delle Marche

备注:Submitted to IEEE Signal Processing Letters

摘要:Transformer最近在语音分离方面取得了最先进的性能。然而,这些模型的计算要求很高,需要很多可学习的参数。本文探讨了一种基于变换器的语音分离方法,可以降低计算量。我们的主要贡献是开发了资源高效的分离转换器(RE-SepFormer),这是一种基于自我关注的架构,可以从两个方面减少计算负担。首先,在潜在空间中使用非重叠块。其次,它对从每个区块计算出的紧凑潜在摘要进行操作。RE-Sepreformer在流行的WSJ0-2Mix和WHAM上达到了有竞争力的性能!因果和非因果环境中的数据集。值得注意的是,在内存和推理时间方面,它的扩展性明显优于以前基于Transformer和RNN的体系结构,因此更适合处理长混合。

摘要:Transformers have recently achieved state-of-the-art performance in speech separation. These models, however, are computationally-demanding and require a lot of learnable parameters. This paper explores Transformer-based speech separation with a reduced computational cost. Our main contribution is the development of the Resource-Efficient Separation Transformer (RE-SepFormer), a self-attention-based architecture that reduces the computational burden in two ways. First, it uses non-overlapping blocks in the latent space. Second, it operates on compact latent summaries calculated from each chunk. The RE-SepFormer reaches a competitive performance on the popular WSJ0-2Mix and WHAM! datasets in both causal and non-causal settings. Remarkably, it scales significantly better than the previous Transformer and RNN-based architectures in terms of memory and inference-time, making it more suitable for processing long mixtures.



【7】 Transfer Learning for Robust Low-Resource Children's Speech ASR with  Transformers and Source-Filter Warping

标题:基于变换和信源滤波的稳健低资源儿童语音ASR转移学习

链接:https://arxiv.org/abs/2206.09396

* 与cs.SD语音【18】为同一篇

作者:Jenthe Thienpondt,Kris Demuynck
机构:IDLab, Department of Electronics and Information Systems, Ghent University - imec, Belgium
备注:proceedings of INTERSPEECH 2022
摘要:众所周知,自动语音识别(ASR)系统在转录儿童语音时会遇到困难。这主要是由于缺乏用于训练鲁棒ASR模型的大型儿童语音语料库,以及使用成人数据训练的系统解码儿童语音时产生的域不匹配。在本文中,我们提出了多项增强来缓解这些问题。首先,我们提出了一种基于语音源滤波模型的数据增强技术,以缩小成人和儿童语音的领域差距。这使我们能够利用成人语音语料库的数据可用性,使这些样本在感知上与儿童语音相似。其次,使用这种增强策略,我们将转移学习应用于成人数据预训练的Transformer模型。该模型遵循最近引入的XLS-R体系结构,这是一种wav2vec 2.0模型,在多个跨语言成人语音语料库上进行预训练,以学习通用和鲁棒的声学框架级表示。在使用成人数据的ASR任务中,采用该模型,再加上所提出的源过滤器扭曲策略和有限数量的域内儿童语音,显著优于之前在PF-STAR英国儿童英语语音语料库上的最新结果,在官方测试集上的WER为4.86%。
摘要:Automatic Speech Recognition (ASR) systems are known to exhibit difficulties when transcribing children's speech. This can mainly be attributed to the absence of large children's speech corpora to train robust ASR models and the resulting domain mismatch when decoding children's speech with systems trained on adult data. In this paper, we propose multiple enhancements to alleviate these issues. First, we propose a data augmentation technique based on the source-filter model of speech to close the domain gap between adult and children's speech. This enables us to leverage the data availability of adult speech corpora by making these samples perceptually similar to children's speech. Second, using this augmentation strategy, we apply transfer learning on a Transformer model pre-trained on adult data. This model follows the recently introduced XLS-R architecture, a wav2vec 2.0 model pre-trained on several cross-lingual adult speech corpora to learn general and robust acoustic frame-level representations. Adopting this model for the ASR task using adult data augmented with the proposed source-filter warping strategy and a limited amount of in-domain children's speech significantly outperforms previous state-of-the-art results on the PF-STAR British English Children's Speech corpus with a 4.86% WER on the official test set.


【8】 Identifying Source Speakers for Voice Conversion based Spoofing Attacks  on Speaker Verification Systems

标题:说话人确认系统中基于语音转换的源说话人识别欺骗攻击

链接:https://arxiv.org/abs/2206.09103

作者:Danwei Cai,Zexin Cai,Ming Li

机构:Department of Electrical and Computer Engineering, Duke University, Durham, USA, Data Science Research Center, Duke Kunshan University, Kunshan, China

摘要:自动说话人验证系统旨在验证语音信号的说话人身份。然而,语音转换系统操纵原始人的语音信号,使其听起来像目标说话人的声音,并欺骗说话人验证系统。大多数基于语音转换的欺骗攻击对策都是为了在说话人验证系统中区分真实语音和欺骗语音而设计的。在本文中,我们研究了源说话人识别问题——根据语音转换语音推断出源说话人的身份。为了进行说话人识别,我们只需在说话人嵌入网络训练期间,将带有说话人身份标签的语音转换语音数据添加到真实语音数据集中。实验结果表明,在使用同一语音转换模型的转换语音进行训练和测试时,源说话人识别是可行的。当测试从一个看不见的语音转换算法转换的语音时,在训练过程中使用更多的语音转换模型可以提高源说话人识别的性能。

摘要:An automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system manipulates the original person's speech signal to make it sound like the target speaker's voice and deceive the speaker verification system. Most countermeasures for voice conversion-based spoofing attacks are designed to discriminate bona fide speech from spoofed speech for speaker verification systems. In this paper, we investigate the problem of source speaker identification -- inferring the identity of the source speaker given the voice converted speech. To perform source speaker identification, we simply add voice-converted speech data with the label of source speaker identity to the genuine speech dataset during speaker embedding network training. Experimental results show the feasibility of source speaker identification when training and testing with converted speeches from the same voice conversion model(s). When testing on converted speeches from an unseen voice conversion algorithm, the performance of source speaker identification improves when more voice conversion models are used during training.


【9】 Decoupled Federated Learning for ASR with Non-IID Data

标题:基于非IID数据的ASR解耦联合学习

链接:https://arxiv.org/abs/2206.09102

* 与cs.SD语音【19】为同一篇

作者:Han Zhu,Jindong Wang,Gaofeng Cheng,Pengyuan Zhang,Yonghong Yan

机构:Key Laboratory of Speech Acoustics and Content Understanding, Institute of Acoustics CAS, China,  University of Chinese Academy of Sciences, China,  Microsoft Research Asia, China

备注:Accepted by Interspeech 2022

摘要:使用联邦学习(FL)的自动语音识别(ASR)可以在不损害隐私的情况下利用来自多个客户端的数据。基于FL的ASR的质量可以通过识别性能、通信和计算成本来衡量。当不同客户端之间的数据不是独立且相同分布的(非IID)时,性能可能会显著降低。在这项工作中,我们使用个性化的FL来解决基于FL的ASR中的非IID问题,它为每个客户学习个性化的模型。具体而言,我们提出了两种针对ASR的个性化FL方法。首先,我们将基于FL的个性化层应用于ASR,它将一些层保留在本地以学习个性化模型。其次,为了降低通信和计算成本,我们提出了解耦联邦学习(DecoupleFL)。一方面,DecoupleFL将计算负担转移到服务器上,从而减少客户端的计算量。另一方面,DecoupleFL通信安全的高级特性,而不是模型参数,从而在模型较大时降低通信成本。实验表明,与FedAvg相比,两种基于FL的个性化ASR方法可以将WER降低2.3%~3.4%。其中,与FedAvg相比,DecoupleFL只需11.4%的通信量和75%的计算量,也明显低于基于个性化层的FL。

摘要:Automatic speech recognition (ASR) with federated learning (FL) makes it possible to leverage data from multiple clients without compromising privacy. The quality of FL-based ASR could be measured by recognition performance, communication and computation costs. When data among different clients are not independently and identically distributed (non-IID), the performance could degrade significantly. In this work, we tackle the non-IID issue in FL-based ASR with personalized FL, which learns personalized models for each client. Concretely, we propose two types of personalized FL approaches for ASR. Firstly, we adapt the personalization layer based FL for ASR, which keeps some layers locally to learn personalization models. Secondly, to reduce the communication and computation costs, we propose decoupled federated learning (DecoupleFL). On one hand, DecoupleFL moves the computation burden to the server, thus decreasing the computation on clients. On the other hand, DecoupleFL communicates secure high-level features instead of model parameters, thus reducing communication cost when models are large. Experiments demonstrate two proposed personalized FL-based ASR approaches could reduce WER by 2.3% - 3.4% compared with FedAvg. Among them, DecoupleFL has only 11.4% communication and 75% computation cost compared with FedAvg, which is also significantly less than the personalization layer based FL.


【10】 Semi-supervised Time Domain Target Speaker Extraction with Attention

标题:基于注意力的半监督时域目标说话人提取

链接:https://arxiv.org/abs/2206.09072

* 与cs.SD语音【20】为同一篇

作者:Zhepei Wang,Ritwik Giri,Shrikant Venkataramani,Umut Isik,Jean-Marc Valin,Paris Smaragdis,Mike Goodwin,Arvindh Krishnaswamy
机构:♯University of Illinois at Urbana-Champaign, †Amazon Web Services
摘要:在这项工作中,我们提出了Exformer,一种用于目标说话人提取的时域体系结构。它由一个预先训练好的说话人嵌入网络和一个基于Transformer编码器块的分离网络组成。我们研究了多种将说话人信息与输入混合信息相结合的方法,得到的Exformer结构与以前的时域网络相比具有更好的提取性能。此外,我们还研究了一个两阶段的过程,在预训练的监督模型上使用无参考信号的混合信号来训练模型。实验结果表明,所提出的半监督学习方法提高了有监督基线的性能。
摘要:In this work, we propose Exformer, a time-domain architecture for target speaker extraction. It consists of a pre-trained speaker embedder network and a separator network based on transformer encoder blocks. We study multiple methods to combine speaker information with the input mixture, and the resulting Exformer architecture obtains superior extraction performance compared to prior time-domain networks. Furthermore, we investigate a two-stage procedure to train the model using mixtures without reference signals upon a pre-trained supervised model. Experimental results show that the proposed semi-supervised learning procedure improves the performance of the supervised baselines.


【11】 NASTAR: Noise Adaptive Speech Enhancement with Target-Conditional  Resampling

标题:NASTAR:基于目标条件重采样的噪声自适应语音增强

链接:https://arxiv.org/abs/2206.09058

作者:Chi-Chang Lee,Cheng-Hung Hu,Yu-Chen Lin,Chu-Song Chen,Hsin-Min Wang,Yu Tsao
机构:Department of Computer Science and Information Engineering, National Taiwan University, Taipei, Taiwan, Research Center for Information Technology Innovation, Academia Sinica, Taipei, Taiwan, Institute of Information Science, Academia Sinica, Taipei, Taiwan
备注:Accepted to Interspeech 2022
摘要:对于基于深度学习的语音增强(SE)系统,训练测试的声学失配会导致显著的性能下降。为了解决不匹配问题,已经推导出了许多噪声自适应策略。在本文中,我们提出了一种新的方法,称为噪声自适应目标条件重采样语音增强(NASTAR),该方法在目标环境中只需一个噪声语音样本(一次)就可以减少失配。NASTAR使用反馈机制,通过噪声提取器和检索模型模拟自适应训练数据。噪声提取器从带噪语音中估计目标噪声,称为伪噪声。噪声检索模型根据噪声语音从噪声信号池中检索相关噪声样本,称为相关队列。将伪噪声和相关队列集联合采样并与源语音语料库混合,以准备噪声自适应的模拟训练数据。实验结果表明,NASTAR可以有效地使用一个带噪语音样本,使SE模型适应目标条件。此外,噪声提取器和噪声检索模型都有助于模型自适应。据我们所知,NASTAR是第一个通过噪声提取和检索执行单次噪声自适应的工作。
摘要:For deep learning-based speech enhancement (SE) systems, the training-test acoustic mismatch can cause notable performance degradation. To address the mismatch issue, numerous noise adaptation strategies have been derived. In this paper, we propose a novel method, called noise adaptive speech enhancement with target-conditional resampling (NASTAR), which reduces mismatches with only one sample (one-shot) of noisy speech in the target environment. NASTAR uses a feedback mechanism to simulate adaptive training data via a noise extractor and a retrieval model. The noise extractor estimates the target noise from the noisy speech, called pseudo-noise. The noise retrieval model retrieves relevant noise samples from a pool of noise signals according to the noisy speech, called relevant-cohort. The pseudo-noise and the relevant-cohort set are jointly sampled and mixed with the source speech corpus to prepare simulated training data for noise adaptation. Experimental results show that NASTAR can effectively use one noisy speech sample to adapt an SE model to a target condition. Moreover, both the noise extractor and the noise retrieval model contribute to model adaptation. To our best knowledge, NASTAR is the first work to perform one-shot noise adaptation through noise extraction and retrieval.


【12】 Rethinking Audio-visual Synchronization for Active Speaker Detection

标题:基于有源说话人检测的视听同步再思考

链接:https://arxiv.org/abs/2206.10421

* 与cs.SD语音【1】为同一篇

作者:Abudukelimu Wuerkaixi,You Zhang,Zhiyao Duan,Changshui Zhang

机构:⋆ Institute for Artificial Intelligence, Tsinghua University (THUAI), State Key Lab of Intelligent Technologies and Systems, Beijing National Research Center for Information Science and Technology (BNRist)

备注:Accepted by IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2022)

摘要:主动说话人检测(ASD)系统是分析多人会话的重要模块。他们的目标是在任何给定的时间检测视觉场景中哪些说话者在说话或没有人在说话。关于自闭症的现有研究对主动说话人的定义不一致。我们在这项工作中澄清了定义,并要求音频和视频演讲活动之间保持同步。这种定义的澄清是由我们的大量实验推动的,通过实验我们发现,现有的ASD方法无法对视听同步进行建模,并且通常将不同步的视频归类为主动说话。为了解决这个问题,我们提出了一种跨模态对比学习策略,并在有监督的ASD模型的注意模块中应用位置编码来利用同步线索。实验结果表明,我们的模型可以成功地将不同步说话检测为不说话,解决了现有模型的局限性。

摘要:Active speaker detection (ASD) systems are important modules for analyzing multi-talker conversations. They aim to detect which speakers or none are talking in a visual scene at any given time. Existing research on ASD does not agree on the definition of active speakers. We clarify the definition in this work and require synchronization between the audio and visual speaking activities. This clarification of definition is motivated by our extensive experiments, through which we discover that existing ASD methods fail in modeling the audio-visual synchronization and often classify unsynchronized videos as active speaking. To address this problem, we propose a cross-modal contrastive learning strategy and apply positional encoding in attention modules for supervised ASD models to leverage the synchronization cue. Experimental results suggest that our model can successfully detect unsynchronized speaking as not speaking, addressing the limitation of current models.


【13】 Audio-video fusion strategies for active speaker detection in meetings

标题:会议中主动说话人检测的音视频融合策略

链接:https://arxiv.org/abs/2206.10411

* 与cs.SD语音【2】为同一篇

作者:Lionel Pibre,Francisco Madrigal,Cyrille Equoy,Frédéric Lerasle,Thomas Pellegrini,Julien Pinquier,Isabelle Ferrané
机构:Fr´ed´eric Lerasle, Isabelle Ferran´e,  IRIT, Universit´e de Toulouse, CNRS, INP Toulouse, UT, Toulouse, France,  LAAS-CNRS, UT, Toulouse, France
摘要:会议是专业环境中的一项常见活动,赋予声乐助理先进的功能以促进会议管理仍然是一项挑战。在这种情况下,像主动说话人检测这样的任务可以为会议参与者之间的交互建模提供有用的见解。受与高级会议助理相关的应用程序上下文的影响,我们希望将音频和视频信息结合起来,以实现最佳性能。在本文中,我们提出了两种不同类型的融合方法来检测主动说话人,通过神经网络将两种视觉模式和一种音频模式相结合。为了进行比较,还使用了经典的无监督音频特征提取方法。我们希望基于嘴唇和面部手势的检测,以每个参与者的面部为中心的视觉数据非常适合检测语音活动。因此,我们的基线系统使用视觉数据,我们选择了一种3D卷积神经网络结构,它可以有效地同时编码外观和运动。为了改进这个系统,我们通过使用CNN或无监督的说话人日记系统处理音频流来补充视觉信息。我们进一步改进了该系统,通过光流运动添加视觉模态信息。我们使用一个公共和最先进的基准:AMI语料库来评估我们的提案。我们分析了每个系统对所进行合并的贡献,以确定某一特定参与者目前是否在发言。我们还讨论了我们得到的结果。此外,我们已经证明,对于我们的应用程序上下文,添加运动信息可以极大地提高性能。最后,我们证明了基于注意的融合在降低标准差的同时提高了性能。
摘要:Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide useful insights to model interaction between meeting participants. Motivated by our application context related to advanced meeting assistant, we want to combine audio and visual information to achieve the best possible performance. In this paper, we propose two different types of fusion for the detection of the active speaker, combining two visual modalities and an audio modality through neural networks. For comparison purpose, classical unsupervised approaches for audio feature extraction are also used. We expect visual data centered on the face of each participant to be very appropriate for detecting voice activity, based on the detection of lip and facial gestures. Thus, our baseline system uses visual data and we chose a 3D Convolutional Neural Network architecture, which is effective for simultaneously encoding appearance and movement. To improve this system, we supplemented the visual information by processing the audio stream with a CNN or an unsupervised speaker diarization system. We have further improved this system by adding visual modality information using motion through optical flow. We evaluated our proposal with a public and state-of-the-art benchmark: the AMI corpus. We analysed the contribution of each system to the merger carried out in order to determine if a given participant is currently speaking. We also discussed the results we obtained. Besides, we have shown that, for our application context, adding motion information greatly improves performance. Finally, we have shown that attention-based fusion improves performance while reducing the standard deviation.


【14】 Joint Analysis of Acoustic Scenes and Sound Events Based on Multitask  Learning with Dynamic Weight Adaptation

标题:基于动态权值自适应多任务学习的声场景与声事件联合分析

链接:https://arxiv.org/abs/2206.10349

* 与cs.SD语音【3】为同一篇

作者:Kayo Nada,Keisuke Imoto,Takao Tsuchiya

机构:Doshisha University, Japan.

备注:Submitted to Acoustical Science and Technology

摘要:声场景分类(ASC)和声事件检测(SED)是环境声分析中的主要课题。考虑到声场景和声音事件之间有着密切的联系,以前的一些工作提出了使用基于多任务学习(MTL)的神经网络对声场景和声音事件进行联合分析。传统方法使用具有恒定权重的ASC和SED损失函数的线性组合来训练基于MTL的模型。然而,基于MTL的传统方法的性能在很大程度上取决于ASC和SED损失的权重,很难确定ASC和SED的MTL损失的恒定权重之间的适当平衡。在本文中,我们提出了基于动态权重平均和多焦点损失的ASC和SED的MTL动态权重调整方法,以自动调整学习权重。使用部分2016/2017年TUT声学场景和2016/2017年TUT声音事件进行了评估实验,结果表明,与传统的基于MTL的方法相比,所提出的方法提高了场景分类和事件检测性能。然后,我们研究了ASC和SED任务的学习权重如何随着模型训练的进行而动态调整。

摘要:Acoustic scene classification (ASC) and sound event detection (SED) are major topics in environmental sound analysis. Considering that acoustic scenes and sound events are closely related to each other, the joint analysis of acoustic scenes and sound events using multitask learning (MTL)-based neural networks was proposed in some previous works. Conventional methods train MTL-based models using a linear combination of ASC and SED loss functions with constant weights. However, the performance of conventional MTL-based methods depends strongly on the weights of the ASC and SED losses, and it is difficult to determine the appropriate balance between the constant weights of the losses of MTL of ASC and SED. In this paper, we thus propose dynamic weight adaptation methods for MTL of ASC and SED based on dynamic weight average and multi--focal loss to adjust the learning weights automatically. Evaluation experiments using parts of the TUT Acoustic Scenes 2016/2017 and TUT Sound Events 2016/2017 are conducted, and we show that the proposed methods improve the scene classification and event detection performance characteristics compared with the conventional MTL-based method. We then investigate how the learning weights of ASC and SED tasks dynamically adapt as the model training progresses.


【15】 Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTS标题:基于DNN的多说话人TTS中的人在环说话人自适应链接:https://arxiv.org/abs/2206.10256

* 与cs.SD语音【4】为同一篇

作者:Kenta Udagawa,Yuki Saito,Hiroshi Saruwatari
机构:Graduate School of Information Science and Technology, The University of Tokyo, Japan.
备注:5 pages, 3 figures, Accepted for INTERSPEECH2022
摘要:提出了一种多说话人文语转换的人在回路说话人自适应方法。采用传统的说话人自适应方法,利用说话人识别任务训练的说话人编码器,从参考语音中提取目标说话人的嵌入向量。然而,当参考语音不可用时,该方法无法获得目标说话人的嵌入向量。我们的方法基于人在回路优化框架,该框架包含用户探索说话人嵌入空间以找到目标说话人的嵌入。该方法使用了一种序列线搜索算法,该算法反复要求用户在嵌入空间的线段上选择一个点。为了有效地从多个刺激中选择最佳语音样本,我们还开发了一个系统,在该系统中,用户可以针对每个音素在多个说话人的语音之间切换,同时循环一个话语。实验结果表明,即使不直接将参考语音作为说话人编码器的输入,该方法在客观和主观评价方面都能达到与传统方法相当的性能。
摘要:This paper proposes a human-in-the-loop speaker-adaptation method for multi-speaker text-to-speech. With a conventional speaker-adaptation method, a target speaker's embedding vector is extracted from his/her reference speech using a speaker encoder trained on a speaker-discriminative task. However, this method cannot obtain an embedding vector for the target speaker when the reference speech is unavailable. Our method is based on a human-in-the-loop optimization framework, which incorporates a user to explore the speaker-embedding space to find the target speaker's embedding. The proposed method uses a sequential line search algorithm that repeatedly asks a user to select a point on a line segment in the embedding space. To efficiently choose the best speech sample from multiple stimuli, we also developed a system in which a user can switch between multiple speakers' voices for each phoneme while looping an utterance. Experimental results indicate that the proposed method can achieve comparable performance to the conventional one in objective and subjective evaluations even if reference speech is not used as the input of a speaker encoder directly.

【16】 Incorporating Voice Instructions in Model-Based Reinforcement Learning  for Self-Driving Cars

标题:自动驾驶汽车模型强化学习中引入语音指令的研究

链接:https://arxiv.org/abs/2206.10249

* 与cs.SD语音【5】为同一篇

作者:Mingze Wang,Ziyang Zhang,Grace Hui Yang

机构:InfoSense, Department of Computer Science, Georgetown University, United States

备注:NeurIPS 2021 Workshop on Machine Learning for Autonomous Driving

摘要:该文提出了一种支持自然语言语音指令的新方法,用于指导自动驾驶汽车的深度强化学习(DRL)算法。DRL方法是自主车辆(AV)代理的常用方法。然而,大多数现有的方法都缺乏样本和时间效率,并且缺乏与人类专家的自然沟通渠道。在本文中,新的人类驾驶员如何从人类教练那里学习激励我们研究人在回路学习的新方法,以及为代理人提供更自然、更易于接近的训练界面。我们建议将自然语言语音指令(NLI)纳入基于模型的深度强化学习中,以训练自动驾驶汽车。我们在CARLA模拟器中评估了所提出的方法以及一些最先进的DRL方法。结果表明,NLI可以帮助简化训练过程,显著提高Agent的学习速度。

摘要:This paper presents a novel approach that supports natural language voice instructions to guide deep reinforcement learning (DRL) algorithms when training self-driving cars. DRL methods are popular approaches for autonomous vehicle (AV) agents. However, most existing methods are sample- and time-inefficient and lack a natural communication channel with the human expert. In this paper, how new human drivers learn from human coaches motivates us to study new ways of human-in-the-loop learning and a more natural and approachable training interface for the agents. We propose incorporating natural language voice instructions (NLI) in model-based deep reinforcement learning to train self-driving cars. We evaluate the proposed method together with a few state-of-the-art DRL methods in the CARLA simulator. The results show that NLI can help ease the training process and significantly boost the agents' learning speed.


【17】 Analysis of Self-Supervised Learning and Dimensionality Reduction  Methods in Clustering-Based Active Learning for Speech Emotion Recognition

标题:语音情感识别中基于聚类的主动学习中的自监督学习和降维方法分析

链接:https://arxiv.org/abs/2206.10188

* 与cs.SD语音【6】为同一篇

作者:Einari Vaaras,Manu Airaksinen,Okko Räsänen

机构:Unit of Computing Sciences, Tampere University, Finland, Helsinki University Hospital, Helsinki, Finland

备注:To be published in Proc. Interspeech 2022, Incheon, South Korea

摘要:当需要领域专家为复杂的机器学习任务执行数据注释时,为了减少时间和费用,减少注释工作量至关重要。对于没有可用注释的情况,一种方法是利用特征空间的结构进行基于聚类的主动学习(AL)方法。然而,这些方法在很大程度上取决于样本在特征空间中的组织方式以及使用的距离度量。对比预测编码(CPC)等无监督方法有可能用于学习有组织的特征空间,但这些方法通常会创建高维特征,这可能对估计数据密度具有挑战性。在本文中,我们将CPC和多维降维方法相结合,搜索基于聚类的AL的功能实践。我们模拟语音情感识别系统部署的实验表明,特征空间的局部和全局拓扑都可以成功地用于AL,与传统信号特征相比,CPC可用于改善基于聚类的AL性能。此外,我们观察到压缩数据维度不会实质上损害AL性能,并且当注释数量不是很低时,二维特征表示与高维表示实现了类似的AL性能。

摘要:When domain experts are needed to perform data annotation for complex machine-learning tasks, reducing annotation effort is crucial in order to cut down time and expenses. For cases when there are no annotations available, one approach is to utilize the structure of the feature space for clustering-based active learning (AL) methods. However, these methods are heavily dependent on how the samples are organized in the feature space and what distance metric is used. Unsupervised methods such as contrastive predictive coding (CPC) can potentially be used to learn organized feature spaces, but these methods typically create high-dimensional features which might be challenging for estimating data density. In this paper, we combine CPC and multiple dimensionality reduction methods in search of functioning practices for clustering-based AL. Our experiments for simulating speech emotion recognition system deployment show that both the local and global topology of the feature space can be successfully used for AL, and that CPC can be used to improve clustering-based AL performance over traditional signal features. Additionally, we observe that compressing data dimensionality does not harm AL performance substantially, and that 2-D feature representations achieved similar AL performance as higher-dimensional representations when the number of annotations is not very low.


【18】 A Multi-grained based Attention Network for Semi-supervised Sound Event  Detection

标题:基于多粒度注意力网络的半监督声音事件检测

链接:https://arxiv.org/abs/2206.10175

* 与cs.SD语音【7】为同一篇

作者:Ying Hu,Xiujuan Zhu,Yunlong Li,Hao Huang,Liang He
机构:Key Laboratory of signal detection and processing in Xinjiang, China, School of Information Science and Engineering, Xinjiang University, Urumqi, China, Department of Electronic Engineering, Tsinghua University, China
备注:None
摘要:由于现实生活中数据的稀缺性和声音事件的多样性,声音事件检测是一项有趣但富有挑战性的任务。提出了一种基于多粒度注意网络的半监督声音事件检测方法。为了获得与声音事件相关的特征表示,设计了一个残差混合卷积(RH-Conv)块来提高普通卷积提取时频特征的能力。此外,还设计了一个多粒度注意(MGA)模块,用于从粗到细学习时间分辨率特征。借助MGA模块,网络可以捕获持续时间短或长的目标事件的特征,从而更准确地确定声音事件的开始和偏移。此外,为了有效提高平均教师(MT)方法的性能,引入了空间移位(SS)模块作为数据扰动机制,以增加数据的多样性。实验结果表明,MGA网络在验证集和公共集上的表现优于已发布的最新竞争对手,分别达到53.27%和56.96%的基于事件的宏F1(EB-F1)分数,0.709和0.739的和声检测分数(PSDS)。
摘要:Sound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. This paper presents a multi-grained based attention network (MGA-Net) for semi-supervised sound event detection. To obtain the feature representations related to sound events, a residual hybrid convolution (RH-Conv) block is designed to boost the vanilla convolution's ability to extract the time-frequency features. Moreover, a multi-grained attention (MGA) module is designed to learn temporal resolution features from coarse-level to fine-level. With the MGA module,the network could capture the characteristics of target events with short- or long-duration, resulting in more accurately determining the onset and offset of sound events. Furthermore, to effectively boost the performance of the Mean Teacher (MT) method, a spatial shift (SS) module as a data perturbation mechanism is introduced to increase the diversity of data. Experimental results show that the MGA-Net outperforms the published state-of-the-art competitors, achieving 53.27% and 56.96% event-based macro F1 (EB-F1) score, 0.709 and 0.739 polyphonic sound detection score (PSDS) on the validation and public set respectively.


【19】 Supervision-Guided Codebooks for Masked Prediction in Speech  Pre-training

标题:语音预训练中监督引导的掩蔽预测码本

链接:https://arxiv.org/abs/2206.10125

* 与cs.SD语音【8】为同一篇

作者:Chengyi Wang,Yiming Wang,Yu Wu,Sanyuan Chen,Jinyu Li,Shujie Liu,Furu Wei

机构:Nankai University, Microsoft

备注:To appear in Proc. Interspeech 2022

摘要:近年来,掩蔽预测预训练在语音识别的自监督学习(SSL)中取得了显著的进展。它通常需要以无监督的方式获取码本,从而使其不太准确且难以解释。为了提高自动语音识别(ASR)的性能和预训练效率,我们提出了两种监督引导码书生成方法,一种是使用混合ASR系统进行解码以生成音素级对齐(称为PBERT),另一种是对从端到端CTC模型中提取的监督语音特征进行聚类(称为CTC聚类)。混合模型和CTC模型都是在微调中使用的少量标记语音上进行训练的。实验表明,与各种SSL和自训练基线相比,我们的方法具有显著的优越性,相对功耗降低了17.0%。我们的预训练模型在非ASR语音任务中也表现出良好的可迁移性。

摘要:Recently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition. It usually requires a codebook obtained in an unsupervised way, making it less accurate and difficult to interpret. We propose two supervision-guided codebook generation approaches to improve automatic speech recognition (ASR) performance and also the pre-training efficiency, either through decoding with a hybrid ASR system to generate phoneme-level alignments (named PBERT), or performing clustering on the supervised speech features extracted from an end-to-end CTC model (named CTC clustering). Both the hybrid and CTC models are trained on the same small amount of labeled speech as used in fine-tuning. Experiments demonstrate significant superiority of our methods to various SSL and self-training baselines, with up to 17.0% relative WER reduction. Our pre-trained models also show good transferability in a non-ASR speech task.

【20】 The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic  Speech Recognition

标题:Makerere无线电语料库:用于自动语音识别的卢甘达无线电语料库

链接:https://arxiv.org/abs/2206.09790

* 与cs.SD语音【10】为同一篇

作者:Jonathan Mukiibi,Andrew Katumba,Joyce Nakatumba-Nabende,Ali Hussein,Josh Meyer

机构:Makerere University, Ronin Institute, Coqui, Uganda, Egypt, USA

备注:Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022), pages 1945 to 1954 Marseille, 20 to 25 June 2022

摘要:对于资源不足的语言来说,建立一个可用的无线电监测自动语音识别(ASR)系统是一项具有挑战性的任务,然而在无线电是公共通信和讨论的主要媒介的社会中,这一点至关重要。联合国在乌干达的初步努力证明,了解被排除在社交媒体之外的农村人的看法在国家规划中是多么重要。然而,这些努力受到缺乏转录语音数据集的挑战。在本文中,Makerere人工智能研究实验室发布了一个155小时的Luganda无线电语音语料库。据我们所知,这是撒哈拉以南非洲第一个公开的无线电数据集。本文描述了语音语料库的开发,并使用开源语音识别工具包Coqui STT toolkit给出了基线Luganda ASR性能结果。

摘要:Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the United Nations in Uganda have proved how understanding the perceptions of rural people who are excluded from social media is important in national planning. However, these efforts are being challenged by the absence of transcribed speech datasets. In this paper, The Makerere Artificial Intelligence research lab releases a Luganda radio speech corpus of 155 hours. To our knowledge, this is the first publicly available radio dataset in sub-Saharan Africa. The paper describes the development of the voice corpus and presents baseline Luganda ASR performance results using Coqui STT toolkit, an open source speech recognition toolkit.

【21】 GMM based multi-stage Wiener filtering for low SNR speech enhancement

标题:基于GMM的多级维纳滤波用于低信噪比语音增强

链接:https://arxiv.org/abs/2206.09298

* 与cs.SD语音【11】为同一篇

作者:Wageesha Manamperi,Prasanga N. Samarasinghe,Thushara D. Abhayapala,Jihui Zhang

机构:The Australian National University, Canberra, Australia

备注:5 pages, 3 figures, submitted to a conference

摘要:本文提出了一种单通道语音增强方法,在低信噪比和非平稳噪声条件下降低噪声并增强语音。具体而言,我们将重点放在使用基于带参数维纳滤波器的多阶段过程的高斯混合模型(GMM)建模噪声。与传统的维纳滤波方法相比,所提出的噪声模型能够估计出更精确的噪声功率谱密度(PSD),并在各种噪声条件下具有更好的泛化能力。仿真结果表明,在低信噪比下,该方法在语音质量(PESQ)和可懂度(STOI)方面都能取得较好的性能。

摘要:This paper proposes a single-channel speech enhancement method to reduce the noise and enhance speech at low signal-to-noise ratio (SNR) levels and non-stationary noise conditions. Specifically, we focus on modeling the noise using a Gaussian mixture model (GMM) based on a multi-stage process with a parametric Wiener filter. The proposed noise model estimates a more accurate noise power spectral density (PSD), and allows for better generalization under various noise conditions compared to traditional Wiener filtering methods. Simulations show that the proposed approach can achieve better performance in terms of speech quality (PESQ) and intelligibility (STOI) at low SNR levels.

【22】 Redundancy Reduction Twins Network: A Training framework for  Multi-output Emotion Regression

标题:冗余双生网络:一种多输出情感回归的训练框架

链接:https://arxiv.org/abs/2206.09142

* 与cs.SD语音【12】为同一篇

作者:Xin Jing,Meishu Song,Andreas Triantafyllopoulos,Zijiang Yang,Björn W. Schuller

机构:University of Augsburg

备注:5 pages, accepted by ICML Exvo workshop

摘要:在本文中,我们提出了冗余缩减双生网络(RRTN),这是一种冗余缩减训练框架,通过测量同一网络输出之间的互相关矩阵,并将其与失真版本的样本进行馈送,使其尽可能接近单位矩阵,从而将冗余降至最低。RRTN还应用了一种新的损失函数,即巴洛双胞胎损失函数,以帮助最大化从样本的不同扭曲版本获得的表示的相似性。然而,由于损失的分布可能会导致网络中的性能波动,我们还建议使用受限不确定性权重损失(RUWL)或联合训练来确定损失函数的最佳权重。我们对CNN14提出的最佳方法是在ExVo多任务开发集上获得0.678的CCC对情绪的回归,比普通的CNN14 CCC 0.647增加4.8%,这在95%置信区间(双尾)下取得了显著差异。

摘要:In this paper, we propose the Redundancy Reduction Twins Network (RRTN), a redundancy reduction training framework that minimizes redundancy by measuring the cross-correlation matrix between the outputs of the same network fed with distorted versions of a sample and bringing it as close to the identity matrix as possible. RRTN also applies a new loss function, the Barlow Twins loss function, to help maximize the similarity of representations obtained from different distorted versions of a sample. However, as the distribution of losses can cause performance fluctuations in the network, we also propose the use of a Restrained Uncertainty Weight Loss (RUWL) or joint training to identify the best weights for the loss function. Our best approach on CNN14 with the proposed methodology obtains a CCC over emotion regression of 0.678 on the ExVo Multi-task dev set, a 4.8% increase over a vanilla CNN 14 CCC of 0.647, which achieves a significant difference at the 95% confidence interval (2-tailed).

【23】 Tackling Spoofing-Aware Speaker Verification with Multi-Model Fusion

标题:利用多模型融合解决感知欺骗的说话人确认问题

链接:https://arxiv.org/abs/2206.09131

* 与cs.SD语音【13】为同一篇

作者:Haibin Wu,Jiawen Kang,Lingwei Meng,Yang Zhang,Xixin Wu,Zhiyong Wu,Hung-yi Lee,Helen Meng

机构:Graduate Institute of Communication Engineering, National Taiwan University,  Centre for Perceptual and Interactive Intelligence, The Chinese University of Hong Kong,  Human-Computer Communications Laboratory, The Chinese University of Hong Kong

备注:Accepted by Odyssey 2022

摘要:近年来,自动说话人验证(ASV)技术得到了长足的发展。然而,以往的研究表明,最先进的ASV模型极易受到语音欺骗攻击,最近提出的高性能欺骗对抗(CM)模型只关注独立的反欺骗任务,而忽略了后续的说话人验证过程。如何将CM和ASV集成在一起仍然是一个悬而未决的问题。最近出现了一个防欺骗说话人验证(SASV)挑战,理由是当CM和ASV子系统联合优化时,可以提供更好的性能。在挑战的场景下,参与者提出的集成系统需要同时拒绝冒名顶替者说话人和目标说话人的欺骗攻击,这直观有效地符合可靠、欺骗鲁棒ASV系统的期望。这项工作侧重于基于融合的SASV解决方案,并提出了一个多模型融合框架,以利用多个最先进的ASV和CM模型的强大功能。提出的框架将SASV-EER从8.75%大幅提高到1.17%,与SASV挑战中的最佳基线系统相比,相对提高了86%。

摘要:Recent years have witnessed the extraordinary development of automatic speaker verification (ASV). However, previous works show that state-of-the-art ASV models are seriously vulnerable to voice spoofing attacks, and the recently proposed high-performance spoofing countermeasure (CM) models only focus solely on the standalone anti-spoofing tasks, and ignore the subsequent speaker verification process. How to integrate the CM and ASV together remains an open question. A spoofing aware speaker verification (SASV) challenge has recently taken place with the argument that better performance can be delivered when both CM and ASV subsystems are optimized jointly. Under the challenge's scenario, the integrated systems proposed by the participants are required to reject both impostor speakers and spoofing attacks from target speakers, which intuitively and effectively matches the expectation of a reliable, spoofing-robust ASV system. This work focuses on fusion-based SASV solutions and proposes a multi-model fusion framework to leverage the power of multiple state-of-the-art ASV and CM models. The proposed framework vastly improves the SASV-EER from 8.75% to 1.17\%, which is 86% relative improvement compared to the best baseline system in the SASV challenge.