今天跟大家分享一篇语音相关的论文合集:cs.SD语音4篇,eess.AS音频处理8篇。

cs.SD语音

【1】 Distance-Based Sound Separation

标题:基于距离的声音分离

链接:https://arxiv.org/abs/2207.00562

作者:Katharine Patterson,Kevin Wilson,Scott Wisdom,John R. Hershey
备注:Accepted for publication at Interspeech 2022
摘要:我们提出了一种新的基于距离的声音分离任务,即仅根据声音与单个麦克风的距离进行分离。在辅助听力设备的环境中,接近度为嘈杂环境中的声音选择提供了一个简单的标准,允许用户专注于与本地对话相关的声音。我们通过训练神经网络在单通道合成混响混合中分离远近声,相对于定义远近边界的阈值距离,证明了该方法的可行性。在一个相邻扬声器和四个远程扬声器的情况下,该模型将近声和远声的比例不变信噪比分别提高了4.4 dB和6.8 dB。
摘要:We propose the novel task of distance-based sound separation, where sounds are separated based only on their distance from a single microphone. In the context of assisted listening devices, proximity provides a simple criterion for sound selection in noisy environments that would allow the user to focus on sounds relevant to a local conversation. We demonstrate the feasibility of this approach by training a neural network to separate near sounds from far sounds in single channel synthetic reverberant mixtures, relative to a threshold distance defining the boundary between near and far. With a single nearby speaker and four distant speakers, the model improves scale-invariant signal to noise ratio by 4.4 dB for near sounds and 6.8 dB for far sounds.


【2】 Toward Low-Cost End-to-End Spoken Language Understanding

标题:走向低成本的端到端口语理解

链接:https://arxiv.org/abs/2207.00352

作者:Marco Dinarelli,Marco Naguib,François Portet
备注:Accepted for publication at Interspeech 2022; Slightly improved (longer) version
摘要:口语理解的最新进展得益于在大型语音语料库上训练的自监督模型。对于法语而言,LeBenchmark项目提供了此类模型,并在包括口语理解在内的多项任务上取得了令人印象深刻的进展。这些进步在计算时间和能源消耗方面具有不可忽视的成本。在本文中,我们比较了几种试图在保持竞争力的同时降低这种成本的学习策略。同时,我们提出了一个广泛的分析,从训练时间和电能消耗的角度衡量我们模型的成本,希望促进一个全面的评估程序。实验在FSC和媒体语料库上进行,结果表明,在保持最先进的性能和使用SSL模型的同时,可以降低学习成本。
摘要:Recent advances in spoken language understanding benefited from Self-Supervised models trained on large speech corpora. For French, the LeBenchmark project has made such models available and has led to impressive progress on several tasks including spoken language understanding. These advances have a non-negligible cost in terms of computation time and energy consumption. In this paper, we compare several learning strategies trying to reduce such cost while keeping competitive performance. At the same time we propose an extensive analysis where we measure the cost of our models in terms of training time and electric energy consumption, hopefully promoting a comprehensive evaluation procedure. The experiments are performed on the FSC and MEDIA corpora, and show that it is possible to reduce the learning cost while maintaining state-of-the-art performance and using SSL models.


【3】 Automatic Evaluation of Speaker Similarity

标题:说话人相似度的自动评价

链接:https://arxiv.org/abs/2207.00344

作者:Deja Kamil,Sanchez Ariadna,Roth Julian,Cotescu Marius
摘要:我们提出了一种新的说话人相似度自动评估方法,该方法与人类感知分数相一致。现代神经文本到语音模型需要大量干净的训练数据,这就是为什么许多解决方案从单说话人模型切换到根据多个不同说话人的示例训练的解决方案。多说话人模型带来了新的可能性,例如更快地创建新语音,但也带来了新的问题-说话人泄漏,即合成示例的说话人身份可能与目标说话人的身份不匹配。目前,发现这个问题的唯一方法是通过昂贵的感知评估。在这项工作中,我们提出了一种自动评估说话人相似性的方法。为此,我们扩展了最近关于说话人验证系统的工作,并评估了不同的指标和说话人嵌入模型如何反映具有隐藏参考和锚定(MUSHRA)分数的多个刺激。我们的实验表明,我们可以训练一个模型,从说话人嵌入中预测说话人相似性MUSHRA分数,准确率为0.96,在话语水平上显著相关,最高可达0.78 Pearson分数。
摘要:We introduce a new automatic evaluation method for speaker similarity assessment, that is consistent with human perceptual scores. Modern neural text-to-speech models require a vast amount of clean training data, which is why many solutions switch from single speaker models to solutions trained on examples from many different speakers. Multi-speaker models bring new possibilities, such as a faster creation of new voices, but also a new problem - speaker leakage, where the speaker identity of a synthesized example might not match those of the target speaker. Currently, the only way to discover this issue is through costly perceptual evaluations. In this work, we propose an automatic method for assessment of speaker similarity. For that purpose, we extend the recent work on speaker verification systems and evaluate how different metrics and speaker embeddings models reflect Multiple Stimuli with Hidden Reference and Anchor (MUSHRA) scores. Our experiments show that we can train a model to predict speaker similarity MUSHRA scores from speaker embeddings with 0.96 accuracy and significant correlation up to 0.78 Pearson score at the utterance level.


【4】 Improving Speech Enhancement through Fine-Grained Speech Characteristics

标题:利用细粒度语音特征提高语音增强效果

链接:https://arxiv.org/abs/2207.00237

作者:Muqiao Yang,Joseph Konan,David Bick,Anurag Kumar,Shinji Watanabe,Bhiksha Raj
摘要:虽然基于深度学习的语音增强系统在改善语音信号质量方面取得了快速进展,但它们仍然可以产生包含伪影的输出,并且听起来可能不自然。我们提出了一种新的语音增强方法,旨在通过优化语音的关键特征来提高增强信号的感知质量和自然度。我们首先确定与语音质量密切相关的关键声学参数(例如抖动、微光和频谱通量),然后提出目标函数,旨在减少纯净语音和增强语音在这些特征方面的差异。全套声学特征是扩展的日内瓦声学参数集(eGeMAPS),其中包括25个与语音感知相关的不同属性。鉴于这些特征计算的不可微性,我们首先构建EGEMap的可微估计器,然后使用它们微调现有的语音增强系统。我们的方法是通用的,可以应用于任何现有的基于深度学习的增强系统,以进一步改善增强的语音信号。在深度噪声抑制挑战数据集上的实验结果表明,我们的方法可以改进最先进的基于深度学习的增强系统。
摘要:While deep learning based speech enhancement systems have made rapid progress in improving the quality of speech signals, they can still produce outputs that contain artifacts and can sound unnatural. We propose a novel approach to speech enhancement aimed at improving perceptual quality and naturalness of enhanced signals by optimizing for key characteristics of speech. We first identify key acoustic parameters that have been found to correlate well with voice quality (e.g. jitter, shimmer, and spectral flux) and then propose objective functions which are aimed at reducing the difference between clean speech and enhanced speech with respect to these features. The full set of acoustic features is the extended Geneva Acoustic Parameter Set (eGeMAPS), which includes 25 different attributes associated with perception of speech. Given the non-differentiable nature of these feature computation, we first build differentiable estimators of the eGeMAPS and then use them to fine-tune existing speech enhancement systems. Our approach is generic and can be applied to any existing deep learning based enhancement systems to further improve the enhanced speech signals. Experimental results conducted on the Deep Noise Suppression (DNS) Challenge dataset shows that our approach can improve the state-of-the-art deep learning based enhancement systems.


eess.AS音频处理

【1】 FitHuBERT: Going Thinner and Deeper for Knowledge Distillation of Speech  Self-Supervised Learning

标题:FitHuBERT:语音自主学习的知识提炼变细变深

链接:https://arxiv.org/abs/2207.00555

作者:Yeonghyeon Lee,Kangwook Jang,Jahyun Goo,Youngmoon Jung,Hoirin Kim
备注:Accepted to Interspeech 2022
摘要:大规模语音自监督学习(SSL)已成为语音处理的主要领域,然而,由于其庞大的规模而产生的计算成本问题给学术界带来了很大的进入障碍。此外,语音SSL模型的现有提取技术通过减少层来压缩模型,这导致语音模式识别任务(如音素识别)的性能下降。在本文中,我们提出了FitHuBERT,与以前的语音SSL蒸馏工作相比,它使几乎所有模型组件的维数都变薄,层更深。此外,我们采用了时间缩减层来加快推理时间,并提出了一种基于提示的提取方法,以减少性能下降。与HuBERT相比,我们的方法将模型的大小减少到23.8%,推理时间减少到35.9%。此外,我们在极好的基准上实现了12.1%的文字错误率和13.3%的音素错误率,这优于以前的工作。
摘要:Large-scale speech self-supervised learning (SSL) has emerged to the main field of speech processing, however, the problem of computational cost arising from its vast size makes a high entry barrier to academia. In addition, existing distillation techniques of speech SSL models compress the model by reducing layers, which induces performance degradation in linguistic pattern recognition tasks such as phoneme recognition (PR). In this paper, we propose FitHuBERT, which makes thinner in dimension throughout almost all model components and deeper in layer compared to prior speech SSL distillation works. Moreover, we employ a time-reduction layer to speed up inference time and propose a method of hint-based distillation for less performance degradation. Our method reduces the model to 23.8% in size and 35.9% in inference time compared to HuBERT. Also, we achieve 12.1% word error rate and 13.3% phoneme error rate on the SUPERB benchmark which is superior than prior work.


【2】 Learning Subject-Invariant Representations from Speech-Evoked EEG Using  Variational Autoencoders

标题:利用变分自动编码器从语音诱发脑电中学习主题不变表征

链接:https://arxiv.org/abs/2207.00323

作者:Lies Bollens,Tom Francart,Hugo Van hamme
备注:None
摘要:脑电图(EEG)是了解大脑如何处理语音的有力方法。为此,线性模型最近已被深度神经网络取代,并产生了有希望的结果。在相关的脑电分类领域中,显式建模主题不变特征提高了跨主题模型的泛化,并提高了分类精度。在这项工作中,我们采用因子分解分层变分自动编码器来利用相同刺激的并行脑电图记录。我们将脑电图建模为两个分离的潜在空间。主题和内容潜空间的主题准确率分别达到98.96%和1.60%,而二元内容分类实验的主题和内容潜空间准确率分别达到51.51%和62.91%。
摘要:The electroencephalogram (EEG) is a powerful method to understand how the brain processes speech. Linear models have recently been replaced for this purpose with deep neural networks and yield promising results. In related EEG classification fields, it is shown that explicitly modeling subject-invariant features improves generalization of models across subjects and benefits classification accuracy. In this work, we adapt factorized hierarchical variational autoencoders to exploit parallel EEG recordings of the same stimuli. We model EEG into two disentangled latent spaces. Subject accuracy reaches 98.96% and 1.60% on respectively the subject and content latent space, whereas binary content classification experiments reach an accuracy of 51.51% and 62.91% on respectively the subject and content latent space.


【3】 Updating Only Encoders Prevents Catastrophic Forgetting of End-to-End  ASR Models

标题:仅更新编码器可防止端到端ASR模型的灾难性遗忘

链接:https://arxiv.org/abs/2207.00216

作者:Yuki Takashima,Shota Horiguchi,Shinji Watanabe,Paola García,Yohei Kawaguchi
备注:Accepted for Interspeech 2022
摘要:在本文中,我们提出了一种增量域自适应技术来防止端到端自动语音识别(ASR)模型的灾难性遗忘。传统方法需要与模型大小相同的额外参数进行优化,并且这些方法很难应用于端到端ASR模型,因为它们有大量的参数。为了解决这个问题,我们首先研究了端到端ASR模型的哪些部分有助于在目标域中实现高精度,同时防止灾难性遗忘。我们使用两种流行的端到端ASR模型对从LibriSpeech数据集到AMI会议语料库的增量域自适应进行了实验,发现仅自适应其编码器的线性层可以防止灾难性遗忘。然后,在这一发现的基础上,我们开发了专注于特定层的元素级参数选择,以进一步减少微调参数的数量。实验结果表明,与从整个模型中选择参数相比,我们的方法始终可以防止灾难性遗忘。
摘要:In this paper, we present an incremental domain adaptation technique to prevent catastrophic forgetting for an end-to-end automatic speech recognition (ASR) model. Conventional approaches require extra parameters of the same size as the model for optimization, and it is difficult to apply these approaches to end-to-end ASR models because they have a huge amount of parameters. To solve this problem, we first investigate which parts of end-to-end ASR models contribute to high accuracy in the target domain while preventing catastrophic forgetting. We conduct experiments on incremental domain adaptation from the LibriSpeech dataset to the AMI meeting corpus with two popular end-to-end ASR models and found that adapting only the linear layers of their encoders can prevent catastrophic forgetting. Then, on the basis of this finding, we develop an element-wise parameter selection focused on specific layers to further reduce the number of fine-tuning parameters. Experimental results show that our approach consistently prevents catastrophic forgetting compared to parameter selection from the whole model.


【4】 SASV Based on Pre-trained ASV System and Integrated Scoring Module

标题:基于预训练ASV系统和集成评分模块的SASV

链接:https://arxiv.org/abs/2207.00150

作者:Yuxiang Zhang,Zhuo Li,Wenchao Wang,Pengyuan Zhang
摘要:基于反欺骗和说话人验证之间存在相关性的假设,提出了一种基于预训练自动说话人验证(ASV)系统和集成评分模块的全分-全集成感知欺骗说话人验证(SASV)系统,并提交给SASV 2022挑战。当前SASV系统中ASV和反欺骗对策(CM)的训练和评分相对独立,忽略了相关性。在本文中,通过利用两个任务之间的相关性,只需在基线预训练的ASV子系统的基础上再训练几层,即可获得集成的SASV系统。预训练ASV系统中的特征用于逻辑访问欺骗语音检测。此外,使用预训练ASV系统提取的说话人嵌入来提高CM的性能。综合评分模块将ASV和反欺骗分支的嵌入作为输入,并通过矩阵运算保持两个任务之间的相关性,以生成综合SASV评分。提交的主要系统在SASV 2022挑战的开发数据集上实现了3.07\%的等错误率(EER),在评估部分实现了4.30\%,比基线系统提高了25%。
摘要:Based on the assumption that there is a correlation between anti-spoofing and speaker verification, a Total-Divide-Total integrated Spoofing-Aware Speaker Verification (SASV) system based on pre-trained automatic speaker verification (ASV) system and integrated scoring module is proposed and submitted to the SASV 2022 Challenge. The training and scoring of ASV and anti-spoofing countermeasure (CM) in current SASV systems are relatively independent, ignoring the correlation. In this paper, by leveraging the correlation between the two tasks, an integrated SASV system can be obtained by simply training a few more layers on the basis of the baseline pre-trained ASV subsystem. The features in pre-trained ASV system are utilized for logical access spoofing speech detection. Further, speaker embeddings extracted by the pre-trained ASV system are used to improve the performance of the CM. The integrated scoring module takes the embeddings of the ASV and anti-spoofing branches as input and preserves the correlation between the two tasks through matrix operations to produce integrated SASV scores. Submitted primary system achieved equal error rate (EER) of 3.07\% on the development dataset of the SASV 2022 Challenge and 4.30\% on the evaluation part, which is a 25\% improvement over the baseline systems.


【5】 Distance-Based Sound Separation

标题:基于距离的声音分离

链接:https://arxiv.org/abs/2207.00562

* 与cs.SD语音【1】为同一篇

作者:Katharine Patterson,Kevin Wilson,Scott Wisdom,John R. Hershey
备注:Accepted for publication at Interspeech 2022
摘要:我们提出了一种新的基于距离的声音分离任务,即仅根据声音与单个麦克风的距离进行分离。在辅助听力设备的环境中,接近度为嘈杂环境中的声音选择提供了一个简单的标准,允许用户专注于与本地对话相关的声音。我们通过训练神经网络在单通道合成混响混合中分离远近声,相对于定义远近边界的阈值距离,证明了该方法的可行性。在一个相邻扬声器和四个远程扬声器的情况下,该模型将近声和远声的比例不变信噪比分别提高了4.4 dB和6.8 dB。
摘要:We propose the novel task of distance-based sound separation, where sounds are separated based only on their distance from a single microphone. In the context of assisted listening devices, proximity provides a simple criterion for sound selection in noisy environments that would allow the user to focus on sounds relevant to a local conversation. We demonstrate the feasibility of this approach by training a neural network to separate near sounds from far sounds in single channel synthetic reverberant mixtures, relative to a threshold distance defining the boundary between near and far. With a single nearby speaker and four distant speakers, the model improves scale-invariant signal to noise ratio by 4.4 dB for near sounds and 6.8 dB for far sounds.


【6】 Toward Low-Cost End-to-End Spoken Language Understanding

标题:走向低成本的端到端口语理解

链接:https://arxiv.org/abs/2207.00352

* 与cs.SD语音【2】为同一篇

作者:Marco Dinarelli,Marco Naguib,François Portet
备注:Accepted for publication at Interspeech 2022; Slightly improved (longer) version
摘要:口语理解的最新进展得益于在大型语音语料库上训练的自监督模型。对于法语而言,LeBenchmark项目提供了此类模型,并在包括口语理解在内的多项任务上取得了令人印象深刻的进展。这些进步在计算时间和能源消耗方面具有不可忽视的成本。在本文中,我们比较了几种试图在保持竞争力的同时降低这种成本的学习策略。同时,我们提出了一个广泛的分析,从训练时间和电能消耗的角度衡量我们模型的成本,希望促进一个全面的评估程序。实验在FSC和媒体语料库上进行,结果表明,在保持最先进的性能和使用SSL模型的同时,可以降低学习成本。
摘要:Recent advances in spoken language understanding benefited from Self-Supervised models trained on large speech corpora. For French, the LeBenchmark project has made such models available and has led to impressive progress on several tasks including spoken language understanding. These advances have a non-negligible cost in terms of computation time and energy consumption. In this paper, we compare several learning strategies trying to reduce such cost while keeping competitive performance. At the same time we propose an extensive analysis where we measure the cost of our models in terms of training time and electric energy consumption, hopefully promoting a comprehensive evaluation procedure. The experiments are performed on the FSC and MEDIA corpora, and show that it is possible to reduce the learning cost while maintaining state-of-the-art performance and using SSL models.


【7】 Automatic Evaluation of Speaker Similarity

标题:说话人相似度的自动评价

链接:https://arxiv.org/abs/2207.00344

* 与cs.SD语音【3】为同一篇

作者:Deja Kamil,Sanchez Ariadna,Roth Julian,Cotescu Marius
摘要:我们提出了一种新的说话人相似度自动评估方法,该方法与人类感知分数相一致。现代神经文本到语音模型需要大量干净的训练数据,这就是为什么许多解决方案从单说话人模型切换到根据多个不同说话人的示例训练的解决方案。多说话人模型带来了新的可能性,例如更快地创建新语音,但也带来了新的问题-说话人泄漏,即合成示例的说话人身份可能与目标说话人的身份不匹配。目前,发现这个问题的唯一方法是通过昂贵的感知评估。在这项工作中,我们提出了一种自动评估说话人相似性的方法。为此,我们扩展了最近关于说话人验证系统的工作,并评估了不同的指标和说话人嵌入模型如何反映具有隐藏参考和锚定(MUSHRA)分数的多个刺激。我们的实验表明,我们可以训练一个模型,从说话人嵌入中预测说话人相似性MUSHRA分数,准确率为0.96,在话语水平上显著相关,最高可达0.78 Pearson分数。
摘要:We introduce a new automatic evaluation method for speaker similarity assessment, that is consistent with human perceptual scores. Modern neural text-to-speech models require a vast amount of clean training data, which is why many solutions switch from single speaker models to solutions trained on examples from many different speakers. Multi-speaker models bring new possibilities, such as a faster creation of new voices, but also a new problem - speaker leakage, where the speaker identity of a synthesized example might not match those of the target speaker. Currently, the only way to discover this issue is through costly perceptual evaluations. In this work, we propose an automatic method for assessment of speaker similarity. For that purpose, we extend the recent work on speaker verification systems and evaluate how different metrics and speaker embeddings models reflect Multiple Stimuli with Hidden Reference and Anchor (MUSHRA) scores. Our experiments show that we can train a model to predict speaker similarity MUSHRA scores from speaker embeddings with 0.96 accuracy and significant correlation up to 0.78 Pearson score at the utterance level.


【8】 Improving Speech Enhancement through Fine-Grained Speech Characteristics

标题:利用细粒度语音特征提高语音增强效果

链接:https://arxiv.org/abs/2207.00237

* 与cs.SD语音【4】为同一篇

作者:Muqiao Yang,Joseph Konan,David Bick,Anurag Kumar,Shinji Watanabe,Bhiksha Raj
摘要:虽然基于深度学习的语音增强系统在改善语音信号质量方面取得了快速进展,但它们仍然可以产生包含伪影的输出,并且听起来可能不自然。我们提出了一种新的语音增强方法,旨在通过优化语音的关键特征来提高增强信号的感知质量和自然度。我们首先确定与语音质量密切相关的关键声学参数(例如抖动、微光和频谱通量),然后提出目标函数,旨在减少纯净语音和增强语音在这些特征方面的差异。全套声学特征是扩展的日内瓦声学参数集(eGeMAPS),其中包括25个与语音感知相关的不同属性。鉴于这些特征计算的不可微性,我们首先构建EGEMap的可微估计器,然后使用它们微调现有的语音增强系统。我们的方法是通用的,可以应用于任何现有的基于深度学习的增强系统,以进一步改善增强的语音信号。在深度噪声抑制挑战数据集上的实验结果表明,我们的方法可以改进最先进的基于深度学习的增强系统。
摘要:While deep learning based speech enhancement systems have made rapid progress in improving the quality of speech signals, they can still produce outputs that contain artifacts and can sound unnatural. We propose a novel approach to speech enhancement aimed at improving perceptual quality and naturalness of enhanced signals by optimizing for key characteristics of speech. We first identify key acoustic parameters that have been found to correlate well with voice quality (e.g. jitter, shimmer, and spectral flux) and then propose objective functions which are aimed at reducing the difference between clean speech and enhanced speech with respect to these features. The full set of acoustic features is the extended Geneva Acoustic Parameter Set (eGeMAPS), which includes 25 different attributes associated with perception of speech. Given the non-differentiable nature of these feature computation, we first build differentiable estimators of the eGeMAPS and then use them to fine-tune existing speech enhancement systems. Our approach is generic and can be applied to any existing deep learning based enhancement systems to further improve the enhanced speech signals. Experimental results conducted on the Deep Noise Suppression (DNS) Challenge dataset shows that our approach can improve the state-of-the-art deep learning based enhancement systems.


机器翻译,仅供参考