今天跟大家分享一篇语音相关的论文合集:cs.SD语音11篇,eess.AS音频处理11篇。

cs.SD语音

【1】 Style Transfer of Audio Effects with Differentiable Signal Processing

标题:基于微分信号处理的音效风格转换

链接:https://arxiv.org/abs/2207.08759

作者:Christian J. Steinmetz,Nicholas J. Bryan,Joshua D. Reiss
备注:Preprint. To appear in the Journal of the Audio Engineering Society
摘要:提出了一个音频效果和风格的生成框架,通过实例将音频效果和风格从一个录音添加到另一个录音中,以简化音频生成过程.我们训练一个深度神经网络来分析输入录音和风格参考录音,并预测用于渲染输出的音频效果的控制参数.与以往的工作相比,我们在我们的框架中集成了音频效果作为可微分算子,通过音频效果执行反向传播,并使用音频域损耗优化端到端。我们使用自使得能够自动控制音频效果而不使用任何标记的或配对的训练数据。我们综述了一系列现有的和新的可微信号处理方法,展示了如何将每一种方法集成到我们的框架中,同时讨论了它们的权衡,我们在语音和音乐任务上评估了我们的方法,证明了我们的方法推广到未看过的记录和甚至推广到不同于训练期间所看到的采样率的采样率。我们的方法产生令人信服的产生式转换结果,从而产生实现可解释性和用户交互的音频效果控制参数。
摘要:We present a framework that can impose the audio effects and production style from one recording to another by example with the goal of simplifying the audio production process. We train a deep neural network to analyze an input recording and a style reference recording, and predict the control parameters of audio effects used to render the output. In contrast to past work, we integrate audio effects as differentiable operators in our framework, perform backpropagation through audio effects, and optimize end-to-end using an audio-domain loss. We use a self-supervised training strategy enabling automatic control of audio effects without the use of any labeled or paired training data. We survey a range of existing and new approaches for differentiable signal processing, showing how each can be integrated into our framework while discussing their trade-offs. We evaluate our approach on both speech and music tasks, demonstrating that our approach generalizes both to unseen recordings and even to sample rates different than those seen during training. Our approach produces convincing production style transfer results with the ability to transform input recordings to produced recordings, yielding audio effect control parameters that enable interpretability and user interaction.


【2】 The Vocal Signature of Social Anxiety: Exploration using  Hypothesis-Testing and Machine-Learning Approaches

标题:社交焦虑的声音特征:使用假设检验和机器学习方法的探索

链接:https://arxiv.org/abs/2207.08534

作者:Or Alon-Ronen,Yosi Shrem,Yossi Keshet,Eva Gilboa-Schechtman
摘要:背景—社交焦虑(SA)是一种常见的衰弱性疾病,即使在亚诊断阈值下也会对生活质量产生负面影响。(ML)方法。方法—参与者形成自发的话语,对拒绝或同意所谓同伴的命令的指令作出反应。结果—我们预测,与低SA组相比,低SA组的语音信号强度和持续时间显著降低。(n=31),高SA(n=32)个个体在拒绝言语中表现出不太自信的语音特征,这一结果仅部分得到经典假设检验方法的支持,ML分析的结果,特别是决策树分类器的结果与SA中的这种语音模式一致。(GP)分类器,我们能够区分高SA和低SA个体(75.6%)准确性和良好(.83 AUC)可分性。我们还预期并发现拒绝和同意话语之间的声音特性差异。结论-我们的研究结果进一步支持了ML方法在精神病理学研究中的有用性,强调了开发自动技术来创建SAD的行为标记的实用性。
摘要:Background - Social anxiety (SA) is a common and debilitating condition, negatively affecting life quality even at sub-diagnostic thresholds. We sought to characterize SA's acoustic signature using hypothesis-testing and machine learning (ML) approaches. Methods - Participants formed spontaneous utterances responding to instructions to refuse or consent to commands of alleged peers. Vocal properties (e.g., intensity and duration) of these utterances were analyzed. Results - Our prediction that, as compared to low-SA (n=31), high-SA (n=32) individuals exhibit a less confident vocal speech signature, especially with respect to refusal utterances, was only partially supported by the classical hypothesis-testing approach. However, the results of the ML analyses and specifically the decision tree classifier were consistent with such speech patterns in SA. Using a Gaussian Process (GP) classifier, we were able to distinguish between high- and low-SA individuals with high (75.6%) accuracy and good (.83 AUC) separability. We also expected and found that vocal properties differentiated between refusal and consent utterances. Conclusions - Our findings provide further support for the usefulness of ML approach for the study of psychopathology, highlighting the utility of developing automatic techniques to create behavioral markers of SAD. Clinically, the simplicity and accessibility of these procedures may encourage people to seek professional help.


【3】 Predictive Neural Speech Coding

标题:预测神经语音编码

链接:https://arxiv.org/abs/2207.08363

作者:Xue Jiang,Xiulian Peng,Huaying Xue,Yuan Zhang,Yan Lu
备注:Submitted to IEEE/ACM TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING (TASLP)
摘要:近年来,神经音频/语音编码已经显示出以比传统方法低得多的比特率提供高质量音频/语音的能力,然而现有的神经音频/语音编码器要么使用声学特征,要么使用学习的盲特征,本文将隐域预测编码引入到矢量量化(VQ)算法中,提出了一种基于隐域预测编码的VQ算法。VAE框架的基础上,提出了一种端到端的低延迟神经语音编码的TF-Codec(TF-Codec)算法,该算法通过对过去量化的潜在帧进行预测,对提取的特征进行编码,进一步消除了时间相关性,提出了一种基于距离—软映射和Gumbel变换的可微分矢量量化算法,该算法在时频输入上引入了一种可学习的压缩机制,以自适应地调整对不同码率下的主频和细节的关注度。在多语言语音数据集上的主观实验结果表明,在40ms的时延下,所提出的1kbps的TF-Codec可以获得比Opus9kbps好得多的质量,3kbps的TF-Codec优于EVS 9.6kbps和Opus12kbps。进行了大量研究以显示这些技术的有效性。
摘要:Neural audio/speech coding has shown its capability to deliver a high quality at much lower bitrates than traditional methods recently. However, existing neural audio/speech codecs employ either acoustic features or learned blind features with a convolutional neural network for encoding, by which there are still temporal redundancies inside encoded features. This paper introduces latent-domain predictive coding into the VQ-VAE framework to fully remove such redundancies and proposes the TF-Codec for low-latency neural speech coding in an end-to-end way. Specifically, the extracted features are encoded conditioned on a prediction from past quantized latent frames so that temporal correlations are further removed. What's more, we introduce a learnable compression on the time-frequency input to adaptively adjust the attention paid on main frequencies and details at different bitrates. A differentiable vector quantization scheme based on distance-to-soft mapping and Gumbel-Softmax is proposed to better model the latent distributions with rate constraint. Subjective results on multilingual speech datasets show that with a latency of 40ms, the proposed TF-Codec at 1kbps can achieve a much better quality than Opus 9kbps and TF-Codec at 3kbps outperforms both EVS 9.6kbps and Opus 12kbps. Numerous studies are conducted to show the effectiveness of these techniques.


【4】 End-to-End Spoken Language Understanding: Performance analyses of a  voice command task in a low resource setting

标题:端到端口语理解:低资源环境下语音命令任务的性能分析

链接:https://arxiv.org/abs/2207.08179

作者:Thierry Desot,François Portet,Michel Vacher
备注:None
摘要:口语理解(SLU)是大多数人机交互系统中的核心任务,随着智能家居、智能手机和智能音箱的出现,SLU已经成为业界的关键技术,在经典的SLU方法中,自动语音识别(ASR)模块将语音信号转录成文本表示,自然语言理解从该文本表示(NLU)模块提取语义信息。最近端到端SLU基于深度神经网络的E2E SLU(E2E SLU)由于受益于ASR和NLU部分的联合优化,从而限制了流水线架构的级联错误效应,因此获得了发展势头。然而,E2E模型用来从语音输入中预测概念和意图的实际语言学性质还知之甚少,我们提出了一项研究,识别E2E模型使用的信号特征和其他语言属性,以执行SLU任务。英语Name(这里是法语)语音命令的测试结果。结果表明,良好的E2E SLU性能并不总是要求完美的ASR能力。此外,结果表明,与管道模型相比,E2E模型在处理背景噪声和语法变化方面具有更好的能力。最后,更细粒度的分析建议E2E模型使用输入信号的音调信息来识别话音命令概念。本文的结果和方法为进一步分析语音处理中的E2E模型提供了一个跳板。
摘要:Spoken Language Understanding (SLU) is a core task in most human-machine interaction systems. With the emergence of smart homes, smart phones and smart speakers, SLU has become a key technology for the industry. In a classical SLU approach, an Automatic Speech Recognition (ASR) module transcribes the speech signal into a textual representation from which a Natural Language Understanding (NLU) module extracts semantic information. Recently End-to-End SLU (E2E SLU) based on Deep Neural Networks has gained momentum since it benefits from the joint optimization of the ASR and the NLU parts, hence limiting the cascade of error effect of the pipeline architecture. However, little is known about the actual linguistic properties used by E2E models to predict concepts and intents from speech input. In this paper, we present a study identifying the signal features and other linguistic properties used by an E2E model to perform the SLU task. The study is carried out in the application domain of a smart home that has to handle non-English (here French) voice commands. The results show that a good E2E SLU performance does not always require a perfect ASR capability. Furthermore, the results show the superior capabilities of the E2E model in handling background noise and syntactic variation compared to the pipeline model. Finally, a finer-grained analysis suggests that the E2E model uses the pitch information of the input signal to identify voice command concepts. The results and methodology outlined in this paper provide a springboard for further analyses of E2E models in speech processing.


【5】 Visually-aware Acoustic Event Detection using Heterogeneous Graphs

标题:基于异构图的视觉感知声学事件检测

链接:https://arxiv.org/abs/2207.07935

作者:Amir Shirian,Krishna Somandepalli,Victor Sanchez,Tanaya Guha
摘要:听觉事件的感知本质上是依赖于听觉和视觉线索的多模态的.现有的许多多模态方法使用模态特定的模型处理每一个模态,我们使用异构图来明确地捕获模态之间的空间和时间关系,并且表示关于潜在信号的详细信息。使用异构图方法来解决视觉感知的声学事件分类问题,异构图是一种简洁、高效和可扩展的数据表示方法,通过异构图,我们示出了在空间和时间尺度上对模态内和模态间关系的有效建模。在AudioSet上的实验表明,该模型具有较高的性能.
摘要:Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-specific models and then fuse the embeddings to encode the joint information. In contrast, we employ heterogeneous graphs to explicitly capture the spatial and temporal relationships between the modalities and represent detailed information about the underlying signal. Using heterogeneous graph approaches to address the task of visually-aware acoustic event classification, which serves as a compact, efficient and scalable way to represent data in the form of graphs. Through heterogeneous graphs, we show efficiently modelling of intra- and inter-modality relationships both at spatial and temporal scales. Our model can easily be adapted to different scales of events through relevant hyperparameters. Experiments on AudioSet, a large benchmark, shows that our model achieves state-of-the-art performance.


【6】 Few-shot bioacoustic event detection at the DCASE 2022 challenge

标题:在DCASE 2022挑战赛中进行的Few-Shot生物声学事件检测

链接:https://arxiv.org/abs/2207.07911

作者:I. Nolasco,S. Singh,E. Vidana-Villa,E. Grout,J. Morford,M. Emmerson,F. Jensens,H. Whitehead,I. Kiskin,A. Strandburg-Peshkin,L. Gill,H. Pamula,V. Lostanlen,V. Morfi,D. Stowell
备注:submitted to DCASE2022 workshop
摘要:Few-Shot声音事件检测是检测声音事件的任务,尽管只具有感兴趣的类别的少数标记的例子,在这类情况下,通常需要注释很长的录音,但专家注释者的时间有限。本文概述了第二版的几个-DCASE 2022挑战中包含的生物声学事件检测任务。详细描述了任务目标、数据集和基线,以及所获得的主要结果和所提交系统的特性。该任务收到了来自15个不同团队的提交,其中13个团队的得分高于基线。评估集的最高F得分为60%,这导致了去年版本的巨大改进。高性能方法使用了原型网络,此外,通过分析每个子集上的结果,我们可以识别系统面临的主要困难,并得出结论,少显示的生物声学声音事件检测仍然是一个开放的挑战。
摘要:Few-shot sound event detection is the task of detecting sound events, despite having only a few labelled examples of the class of interest. This framework is particularly useful in bioacoustics, where often there is a need to annotate very long recordings but the expert annotator time is limited. This paper presents an overview of the second edition of the few-shot bioacoustic sound event detection task included in the DCASE 2022 challenge. A detailed description of the task objectives, dataset, and baselines is presented, together with the main results obtained and characteristics of the submitted systems. This task received submissions from 15 different teams from which 13 scored higher than the baselines. The highest F-score was of 60% on the evaluation set, which leads to a huge improvement over last year's edition. Highly-performing methods made use of prototypical networks, transductive learning, and addressed the variable length of events from all target classes. Furthermore, by analysing results on each of the subsets we can identify the main difficulties that the systems face, and conclude that few-show bioacoustic sound event detection remains an open challenge.


【7】 Improving spatial cues for hearables using a parameterized binaural CDR  estimator

标题:使用参数化双耳CDR估计器改进可听信号的空间线索

链接:https://arxiv.org/abs/2207.08314

作者:Reza Ghanavi,Craig Jin
备注:Accepted by ICA2022. An Australian provisional patent application based on this manuscript has been filed by the University of Sydney
摘要:本文研究了一种基于双耳相干扩散功率比的语音增强方法传统的CDR估计器通常在其公式化中依赖于期望信号和/或扩散噪声场的数学相干模型,提出了一种新的稳健的参数化方向性双耳CDR估计器。该估计器在时频域中计算,并基于双耳麦克风信号之间空间相干函数的几何解释,将新的CDR估计器的双耳性能与三种最先进的CDR估计器在鸡尾酒会上的双耳性能进行了比较。并且已经在诸如PESQ和SRMR之类的若干客观语音质量度量方面显示出改进。我们还讨论了参数化CDR估计器对变化的声音环境的好处,并简要反映了几个非正式的主观使用低延迟实时框架进行评估。
摘要:We investigate a speech enhancement method based on the binaural coherence-to-diffuse power ratio (CDR), which preserves auditory spatial cues for maskers and a broadside target. Conventional CDR estimators typically rely on a mathematical coherence model of the desired signal and/or diffuse noise field in their formulation, which may influence their accuracy in natural environments. This work proposes a new robust and parameterized directional binaural CDR estimator. The estimator is calculated in the time-frequency domain and is based on a geometrical interpretation of the spatial coherence function between the binaural microphone signals. The binaural performance of the new CDR estimator is compared with three state-of-the-art CDR estimators in cocktail-party-like environments and has shown improvements in terms of several objective speech quality metrics such as PESQ and SRMR. We also discuss the benefits of the parameterizable CDR estimator for varying sound environments and briefly reflect on several informal subjective evaluations using a low-latency real-time framework.


【8】 Multi-channel target speech enhancement based on ERB-scaled spatial  coherence features

标题:基于ERB尺度空间相干特征的多通道目标语音增强

链接:https://arxiv.org/abs/2207.08126

作者:Yicheng Hsu,Yonghan Lee,Mingsian R. Bai
备注:Accepted by International Congress on Acoustics (ICA) 2022. arXiv admin note: substantial text overlap with arXiv:2112.05686
摘要:近年来,基于深度学习的语音增强技术受到了广泛的关注,如果充分利用麦克风信号中的空间信息,麦克风阵列在某些恶劣的声学条件下比单麦克风系统具有更大的优势,然而多通道语音增强通常是在短时傅里叶变换中进行的(STFT)域中,这使得增强方法在计算上是昂贵的。提出了一种新的等效矩形带宽(ERB)尺度的空间相干性特征,该特征依赖于两个ERB频带之间的目标说话人活动。在混响环境中的麦克风阵列,包括语音干扰,证明了所提出的系统的有效性,该研究也证明了用ERB尺度的空间特征训练的网络对阵列中麦克风的几何形状和数量的变化是鲁棒的。
摘要:Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageous under some adverse acoustic conditions compared with single-microphone systems. However, multichannel speech enhancement is often performed in the short-time Fourier transform (STFT) domain, which renders the enhancement approach computationally expensive. To remedy this problem, we propose a novel equivalent rectangular bandwidth (ERB)-scaled spatial coherence feature that is dependent on the target speaker activity between two ERB bands. Experiments conducted using a four-microphone array in a reverberant environment, which involved speech interference, demonstrated the efficacy of the proposed system. This study also demonstrated that a network that was trained with the ERB-scaled spatial feature was robust against variations in the geometry and number of the microphones in the array.


【9】 Reducing Geographic Disparities in Automatic Speech Recognition via  Elastic Weight Consolidation

标题:通过弹性权重合并减少自动语音识别中的地理差异

链接:https://arxiv.org/abs/2207.07850

作者:Viet Anh Trinh,Pegah Ghahremani,Brian King,Jasha Droppo,Andreas Stolcke,Roland Maas
备注:Accepted for publication at Interspeech 2022
摘要:我们提出了一种方法来减少地理区域之间的性能差异,而不会降低ASR的总体用户群的性能。一种流行的方法是使用来自ASR模型具有较高单词错误率的区域的数据来微调模型然而,当ASR模型被调整以在这些高-WER区域上获得更好的性能时,其参数从先前的最优值漂移,在本文中,我们利用了弹性重量固结(EWC)正则化损失以识别参数空间中ASR权重可以沿其变化的方向,同时仍能保持对说话人总体的性能。我们的结果表明,EWC可以降低单词错误率(WER)在WER最高的地区相对减少3. 2%,而总体WER相对减少1. 3%。我们还评估了语言和声学模型在ASR公平性中的作用,并提出了一种基于地理区域的聚类算法来识别WER差异。
摘要:We present an approach to reduce the performance disparity between geographic regions without degrading performance on the overall user population for ASR. A popular approach is to fine-tune the model with data from regions where the ASR model has a higher word error rate (WER). However, when the ASR model is adapted to get better performance on these high-WER regions, its parameters wander from the previous optimal values, which can lead to worse performance in other regions. In our proposed method, we utilize the elastic weight consolidation (EWC) regularization loss to identify directions in parameters space along which the ASR weights can vary to improve for high-error regions, while still maintaining performance on the speaker population overall. Our results demonstrate that EWC can reduce the word error rate (WER) in the region with highest WER by 3.2% relative while reducing the overall WER by 1.3% relative. We also evaluate the role of language and acoustic models in ASR fairness and propose a clustering algorithm to identify WER disparities based on geographic region.


【10】 Adversarial Reweighting for Speaker Verification Fairness

标题:说话人确认公平性的对抗式再加权

链接:https://arxiv.org/abs/2207.07776

作者:Minho Jin,Chelsea J. -T. Ju,Zeya Chen,Yi-Chieh Liu,Jasha Droppo,Andreas Stolcke
摘要:我们使用对抗性重新加权来解决说话人确认的性能公平性问题(ARW)方法,对ARW方法进行了改进,使其适用于具有度量学习的说话人确认,并在不同性别和国籍的子组中,而不需要对训练数据中的子组进行注释。对抗网络学习批中每个训练样本的权重,使得主学习器被迫集中于表现不佳的实例。该方法采用最小最大优化算法,提高了说话人确认的公平性,提出了三种不同的ARW公式:累积成对相似性、伪标记和成对加权,并以等错误率衡量它们的性能实验结果表明,两两加权法可以使总的EER降低1.08%,其中男性为1.25%,女性为0.67%,相对EER降低分别为7.7%、10.1%和3.0%,对不同国籍的人群,美国人的EER为1.04%,英国人为0.76%,其他人群为1.22%,性别之间的绝对EER差距从0.70%缩小到0.58%,而各国籍组的标准差则从0.21降至0.19。
摘要:We address performance fairness for speaker verification using the adversarial reweighting (ARW) method. ARW is reformulated for speaker verification with metric learning, and shown to improve results across different subgroups of gender and nationality, without requiring annotation of subgroups in the training data. An adversarial network learns a weight for each training sample in the batch so that the main learner is forced to focus on poorly performing instances. Using a min-max optimization algorithm, this method improves overall speaker verification fairness. We present three different ARWformulations: accumulated pairwise similarity, pseudo-labeling, and pairwise weighting, and measure their performance in terms of equal error rate (EER) on the VoxCeleb corpus. Results show that the pairwise weighting method can achieve 1.08% overall EER, 1.25% for male and 0.67% for female speakers, with relative EER reductions of 7.7%, 10.1% and 3.0%, respectively. For nationality subgroups, the proposed algorithm showed 1.04% EER for US speakers, 0.76% for UK speakers, and 1.22% for all others. The absolute EER gap between gender groups was reduced from 0.70% to 0.58%, while the standard deviation over nationality groups decreased from 0.21 to 0.19.


【11】 Segment-level Metric Learning for Few-shot Bioacoustic Event Detection

标题:基于分段尺度学习的Few-Shot生物声学事件检测

链接:https://arxiv.org/abs/2207.07773

作者:Haohe Liu,Xubo Liu,Xinhao Mei,Qiuqiang Kong,Wenwu Wang,Mark D. Plumbley
备注:2nd place in the DCASE 2022 Challenge Task 5. Submitted to the DCASE 2022 workshop
摘要:生物声学事件检测是一种通过少量实例来检测一个新声音出现时间的任务.以往的方法采用度量学习来建立一个包含不同声音类别的标记部分(也称为正事件)的潜在空间.在本研究中,我们提出了一个分段级的少拍学习框架,在模型优化过程中同时利用正事件和负事件.用负事件训练,其中,正事件的数量大于正事件的数量,这可以提高模型的泛化能力,此外,为了更好地适应新的类别,我们在训练过程中对验证集进行了直推推理.我们在不同的输入特征、训练数据我们的最终系统在DCASE 2022挑战任务5上实现了62.73的F-测度(DCASE2022-T5)验证集的性能测试结果表明,我们提交的系统在DCASE2022-T5中排名第二。本文的代码是完全开源的,位于https://github.com/haoheliu/DCASE_2022_Task_5.
摘要:Few-shot bioacoustic event detection is a task that detects the occurrence time of a novel sound given a few examples. Previous methods employ metric learning to build a latent space with the labeled part of different sound classes, also known as positive events. In this study, we propose a segment-level few-shot learning framework that utilizes both the positive and negative events during model optimization. Training with negative events, which are larger in volume than positive events, can increase the generalization ability of the model. In addition, we use transductive inference on the validation set during training for better adaptation to novel classes. We conduct ablation studies on our proposed method with different setups on input features, training data, and hyper-parameters. Our final system achieves an F-measure of 62.73 on the DCASE 2022 challenge task 5 (DCASE2022-T5) validation set, outperforming the performance of the baseline prototypical network 34.02 by a large margin. Using the proposed method, our submitted system ranks 2nd in DCASE2022-T5. The code of this paper is fully open-sourced at https://github.com/haoheliu/DCASE_2022_Task_5.


eess.AS音频处理

【1】 Improving spatial cues for hearables using a parameterized binaural CDR  estimator

标题:使用参数化双耳CDR估计器改进可听信号的空间线索

链接:https://arxiv.org/abs/2207.08314

* 与cs.SD语音【7】为同一篇

作者:Reza Ghanavi,Craig Jin
备注:Accepted by ICA2022. An Australian provisional patent application based on this manuscript has been filed by the University of Sydney
摘要:本文研究了一种基于双耳相干扩散功率比的语音增强方法传统的CDR估计器通常在其公式化中依赖于期望信号和/或扩散噪声场的数学相干模型,提出了一种新的稳健的参数化方向性双耳CDR估计器。该估计器在时频域中计算,并基于双耳麦克风信号之间空间相干函数的几何解释,将新的CDR估计器的双耳性能与三种最先进的CDR估计器在鸡尾酒会上的双耳性能进行了比较。并且已经在诸如PESQ和SRMR之类的若干客观语音质量度量方面显示出改进。我们还讨论了参数化CDR估计器对变化的声音环境的好处,并简要反映了几个非正式的主观使用低延迟实时框架进行评估。
摘要:We investigate a speech enhancement method based on the binaural coherence-to-diffuse power ratio (CDR), which preserves auditory spatial cues for maskers and a broadside target. Conventional CDR estimators typically rely on a mathematical coherence model of the desired signal and/or diffuse noise field in their formulation, which may influence their accuracy in natural environments. This work proposes a new robust and parameterized directional binaural CDR estimator. The estimator is calculated in the time-frequency domain and is based on a geometrical interpretation of the spatial coherence function between the binaural microphone signals. The binaural performance of the new CDR estimator is compared with three state-of-the-art CDR estimators in cocktail-party-like environments and has shown improvements in terms of several objective speech quality metrics such as PESQ and SRMR. We also discuss the benefits of the parameterizable CDR estimator for varying sound environments and briefly reflect on several informal subjective evaluations using a low-latency real-time framework.


【2】 Multi-channel target speech enhancement based on ERB-scaled spatial  coherence features

标题:基于ERB尺度空间相干特征的多通道目标语音增强

链接:https://arxiv.org/abs/2207.08126

* 与cs.SD语音【8】为同一篇

作者:Yicheng Hsu,Yonghan Lee,Mingsian R. Bai
备注:Accepted by International Congress on Acoustics (ICA) 2022. arXiv admin note: substantial text overlap with arXiv:2112.05686
摘要:近年来,基于深度学习的语音增强技术受到了广泛的关注,如果充分利用麦克风信号中的空间信息,麦克风阵列在某些恶劣的声学条件下比单麦克风系统具有更大的优势,然而多通道语音增强通常是在短时傅里叶变换中进行的(STFT)域中,这使得增强方法在计算上是昂贵的。提出了一种新的等效矩形带宽(ERB)尺度的空间相干性特征,该特征依赖于两个ERB频带之间的目标说话人活动。在混响环境中的麦克风阵列,包括语音干扰,证明了所提出的系统的有效性,该研究也证明了用ERB尺度的空间特征训练的网络对阵列中麦克风的几何形状和数量的变化是鲁棒的。
摘要:Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageous under some adverse acoustic conditions compared with single-microphone systems. However, multichannel speech enhancement is often performed in the short-time Fourier transform (STFT) domain, which renders the enhancement approach computationally expensive. To remedy this problem, we propose a novel equivalent rectangular bandwidth (ERB)-scaled spatial coherence feature that is dependent on the target speaker activity between two ERB bands. Experiments conducted using a four-microphone array in a reverberant environment, which involved speech interference, demonstrated the efficacy of the proposed system. This study also demonstrated that a network that was trained with the ERB-scaled spatial feature was robust against variations in the geometry and number of the microphones in the array.


【3】 Reducing Geographic Disparities in Automatic Speech Recognition via  Elastic Weight Consolidation

标题:通过弹性权重合并减少自动语音识别中的地理差异

链接:https://arxiv.org/abs/2207.07850

* 与cs.SD语音【9】为同一篇

作者:Viet Anh Trinh,Pegah Ghahremani,Brian King,Jasha Droppo,Andreas Stolcke,Roland Maas
备注:Accepted for publication at Interspeech 2022
摘要:我们提出了一种方法来减少地理区域之间的性能差异,而不会降低ASR的总体用户群的性能。一种流行的方法是使用来自ASR模型具有较高单词错误率的区域的数据来微调模型然而,当ASR模型被调整以在这些高-WER区域上获得更好的性能时,其参数从先前的最优值漂移,在本文中,我们利用了弹性重量固结(EWC)正则化损失以识别参数空间中ASR权重可以沿其变化的方向,同时仍能保持对说话人总体的性能。我们的结果表明,EWC可以降低单词错误率(WER)在WER最高的地区相对减少3. 2%,而总体WER相对减少1. 3%。我们还评估了语言和声学模型在ASR公平性中的作用,并提出了一种基于地理区域的聚类算法来识别WER差异。
摘要:We present an approach to reduce the performance disparity between geographic regions without degrading performance on the overall user population for ASR. A popular approach is to fine-tune the model with data from regions where the ASR model has a higher word error rate (WER). However, when the ASR model is adapted to get better performance on these high-WER regions, its parameters wander from the previous optimal values, which can lead to worse performance in other regions. In our proposed method, we utilize the elastic weight consolidation (EWC) regularization loss to identify directions in parameters space along which the ASR weights can vary to improve for high-error regions, while still maintaining performance on the speaker population overall. Our results demonstrate that EWC can reduce the word error rate (WER) in the region with highest WER by 3.2% relative while reducing the overall WER by 1.3% relative. We also evaluate the role of language and acoustic models in ASR fairness and propose a clustering algorithm to identify WER disparities based on geographic region.


【4】 Adversarial Reweighting for Speaker Verification Fairness

标题:说话人确认公平性的对抗式再加权

链接:https://arxiv.org/abs/2207.07776

* 与cs.SD语音【10】为同一篇

作者:Minho Jin,Chelsea J. -T. Ju,Zeya Chen,Yi-Chieh Liu,Jasha Droppo,Andreas Stolcke
摘要:我们使用对抗性重新加权来解决说话人确认的性能公平性问题(ARW)方法,对ARW方法进行了改进,使其适用于具有度量学习的说话人确认,并在不同性别和国籍的子组中,而不需要对训练数据中的子组进行注释。对抗网络学习批中每个训练样本的权重,使得主学习器被迫集中于表现不佳的实例。该方法采用最小最大优化算法,提高了说话人确认的公平性,提出了三种不同的ARW公式:累积成对相似性、伪标记和成对加权,并以等错误率衡量它们的性能实验结果表明,两两加权法可以使总的EER降低1.08%,其中男性为1.25%,女性为0.67%,相对EER降低分别为7.7%、10.1%和3.0%,对不同国籍的人群,美国人的EER为1.04%,英国人为0.76%,其他人群为1.22%,性别之间的绝对EER差距从0.70%缩小到0.58%,而各国籍组的标准差则从0.21降至0.19。
摘要:We address performance fairness for speaker verification using the adversarial reweighting (ARW) method. ARW is reformulated for speaker verification with metric learning, and shown to improve results across different subgroups of gender and nationality, without requiring annotation of subgroups in the training data. An adversarial network learns a weight for each training sample in the batch so that the main learner is forced to focus on poorly performing instances. Using a min-max optimization algorithm, this method improves overall speaker verification fairness. We present three different ARWformulations: accumulated pairwise similarity, pseudo-labeling, and pairwise weighting, and measure their performance in terms of equal error rate (EER) on the VoxCeleb corpus. Results show that the pairwise weighting method can achieve 1.08% overall EER, 1.25% for male and 0.67% for female speakers, with relative EER reductions of 7.7%, 10.1% and 3.0%, respectively. For nationality subgroups, the proposed algorithm showed 1.04% EER for US speakers, 0.76% for UK speakers, and 1.22% for all others. The absolute EER gap between gender groups was reduced from 0.70% to 0.58%, while the standard deviation over nationality groups decreased from 0.21 to 0.19.


【5】 Segment-level Metric Learning for Few-shot Bioacoustic Event Detection

标题:基于分段尺度学习的Few-Shot生物声学事件检测

链接:https://arxiv.org/abs/2207.07773

* 与cs.SD语音【11】为同一篇

作者:Haohe Liu,Xubo Liu,Xinhao Mei,Qiuqiang Kong,Wenwu Wang,Mark D. Plumbley
备注:2nd place in the DCASE 2022 Challenge Task 5. Submitted to the DCASE 2022 workshop
摘要:生物声学事件检测是一种通过少量实例来检测一个新声音出现时间的任务.以往的方法采用度量学习来建立一个包含不同声音类别的标记部分(也称为正事件)的潜在空间.在本研究中,我们提出了一个分段级的少拍学习框架,在模型优化过程中同时利用正事件和负事件.用负事件训练,其中,正事件的数量大于正事件的数量,这可以提高模型的泛化能力,此外,为了更好地适应新的类别,我们在训练过程中对验证集进行了直推推理.我们在不同的输入特征、训练数据我们的最终系统在DCASE 2022挑战任务5上实现了62.73的F-测度(DCASE2022-T5)验证集的性能测试结果表明,我们提交的系统在DCASE2022-T5中排名第二。本文的代码是完全开源的,位于https://github.com/haoheliu/DCASE_2022_Task_5.
摘要:Few-shot bioacoustic event detection is a task that detects the occurrence time of a novel sound given a few examples. Previous methods employ metric learning to build a latent space with the labeled part of different sound classes, also known as positive events. In this study, we propose a segment-level few-shot learning framework that utilizes both the positive and negative events during model optimization. Training with negative events, which are larger in volume than positive events, can increase the generalization ability of the model. In addition, we use transductive inference on the validation set during training for better adaptation to novel classes. We conduct ablation studies on our proposed method with different setups on input features, training data, and hyper-parameters. Our final system achieves an F-measure of 62.73 on the DCASE 2022 challenge task 5 (DCASE2022-T5) validation set, outperforming the performance of the baseline prototypical network 34.02 by a large margin. Using the proposed method, our submitted system ranks 2nd in DCASE2022-T5. The code of this paper is fully open-sourced at https://github.com/haoheliu/DCASE_2022_Task_5.


【6】 Style Transfer of Audio Effects with Differentiable Signal Processing

标题:基于微分信号处理的音效风格转换

链接:https://arxiv.org/abs/2207.08759

* 与cs.SD语音【1】为同一篇

作者:Christian J. Steinmetz,Nicholas J. Bryan,Joshua D. Reiss
备注:Preprint. To appear in the Journal of the Audio Engineering Society
摘要:提出了一个音频效果和风格的生成框架,通过实例将音频效果和风格从一个录音添加到另一个录音中,以简化音频生成过程.我们训练一个深度神经网络来分析输入录音和风格参考录音,并预测用于渲染输出的音频效果的控制参数.与以往的工作相比,我们在我们的框架中集成了音频效果作为可微分算子,通过音频效果执行反向传播,并使用音频域损耗优化端到端。我们使用自使得能够自动控制音频效果而不使用任何标记的或配对的训练数据。我们综述了一系列现有的和新的可微信号处理方法,展示了如何将每一种方法集成到我们的框架中,同时讨论了它们的权衡,我们在语音和音乐任务上评估了我们的方法,证明了我们的方法推广到未看过的记录和甚至推广到不同于训练期间所看到的采样率的采样率。我们的方法产生令人信服的产生式转换结果,从而产生实现可解释性和用户交互的音频效果控制参数。
摘要:We present a framework that can impose the audio effects and production style from one recording to another by example with the goal of simplifying the audio production process. We train a deep neural network to analyze an input recording and a style reference recording, and predict the control parameters of audio effects used to render the output. In contrast to past work, we integrate audio effects as differentiable operators in our framework, perform backpropagation through audio effects, and optimize end-to-end using an audio-domain loss. We use a self-supervised training strategy enabling automatic control of audio effects without the use of any labeled or paired training data. We survey a range of existing and new approaches for differentiable signal processing, showing how each can be integrated into our framework while discussing their trade-offs. We evaluate our approach on both speech and music tasks, demonstrating that our approach generalizes both to unseen recordings and even to sample rates different than those seen during training. Our approach produces convincing production style transfer results with the ability to transform input recordings to produced recordings, yielding audio effect control parameters that enable interpretability and user interaction.


【7】 The Vocal Signature of Social Anxiety: Exploration using  Hypothesis-Testing and Machine-Learning Approaches

标题:社交焦虑的声音特征:使用假设检验和机器学习方法的探索

链接:https://arxiv.org/abs/2207.08534

* 与cs.SD语音【2】为同一篇

作者:Or Alon-Ronen,Yosi Shrem,Yossi Keshet,Eva Gilboa-Schechtman
摘要:背景—社交焦虑(SA)是一种常见的衰弱性疾病,即使在亚诊断阈值下也会对生活质量产生负面影响。(ML)方法。方法—参与者形成自发的话语,对拒绝或同意所谓同伴的命令的指令作出反应。结果—我们预测,与低SA组相比,低SA组的语音信号强度和持续时间显著降低。(n=31),高SA(n=32)个个体在拒绝言语中表现出不太自信的语音特征,这一结果仅部分得到经典假设检验方法的支持,ML分析的结果,特别是决策树分类器的结果与SA中的这种语音模式一致。(GP)分类器,我们能够区分高SA和低SA个体(75.6%)准确性和良好(.83 AUC)可分性。我们还预期并发现拒绝和同意话语之间的声音特性差异。结论-我们的研究结果进一步支持了ML方法在精神病理学研究中的有用性,强调了开发自动技术来创建SAD的行为标记的实用性。
摘要:Background - Social anxiety (SA) is a common and debilitating condition, negatively affecting life quality even at sub-diagnostic thresholds. We sought to characterize SA's acoustic signature using hypothesis-testing and machine learning (ML) approaches. Methods - Participants formed spontaneous utterances responding to instructions to refuse or consent to commands of alleged peers. Vocal properties (e.g., intensity and duration) of these utterances were analyzed. Results - Our prediction that, as compared to low-SA (n=31), high-SA (n=32) individuals exhibit a less confident vocal speech signature, especially with respect to refusal utterances, was only partially supported by the classical hypothesis-testing approach. However, the results of the ML analyses and specifically the decision tree classifier were consistent with such speech patterns in SA. Using a Gaussian Process (GP) classifier, we were able to distinguish between high- and low-SA individuals with high (75.6%) accuracy and good (.83 AUC) separability. We also expected and found that vocal properties differentiated between refusal and consent utterances. Conclusions - Our findings provide further support for the usefulness of ML approach for the study of psychopathology, highlighting the utility of developing automatic techniques to create behavioral markers of SAD. Clinically, the simplicity and accessibility of these procedures may encourage people to seek professional help.


【8】 Predictive Neural Speech Coding

标题:预测神经语音编码

链接:https://arxiv.org/abs/2207.08363

* 与cs.SD语音【3】为同一篇

作者:Xue Jiang,Xiulian Peng,Huaying Xue,Yuan Zhang,Yan Lu
备注:Submitted to IEEE/ACM TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING (TASLP)
摘要:近年来,神经音频/语音编码已经显示出以比传统方法低得多的比特率提供高质量音频/语音的能力,然而现有的神经音频/语音编码器要么使用声学特征,要么使用学习的盲特征,本文将隐域预测编码引入到矢量量化(VQ)算法中,提出了一种基于隐域预测编码的VQ算法。VAE框架的基础上,提出了一种端到端的低延迟神经语音编码的TF-Codec(TF-Codec)算法,该算法通过对过去量化的潜在帧进行预测,对提取的特征进行编码,进一步消除了时间相关性,提出了一种基于距离—软映射和Gumbel变换的可微分矢量量化算法,该算法在时频输入上引入了一种可学习的压缩机制,以自适应地调整对不同码率下的主频和细节的关注度。在多语言语音数据集上的主观实验结果表明,在40ms的时延下,所提出的1kbps的TF-Codec可以获得比Opus9kbps好得多的质量,3kbps的TF-Codec优于EVS 9.6kbps和Opus12kbps。进行了大量研究以显示这些技术的有效性。
摘要:Neural audio/speech coding has shown its capability to deliver a high quality at much lower bitrates than traditional methods recently. However, existing neural audio/speech codecs employ either acoustic features or learned blind features with a convolutional neural network for encoding, by which there are still temporal redundancies inside encoded features. This paper introduces latent-domain predictive coding into the VQ-VAE framework to fully remove such redundancies and proposes the TF-Codec for low-latency neural speech coding in an end-to-end way. Specifically, the extracted features are encoded conditioned on a prediction from past quantized latent frames so that temporal correlations are further removed. What's more, we introduce a learnable compression on the time-frequency input to adaptively adjust the attention paid on main frequencies and details at different bitrates. A differentiable vector quantization scheme based on distance-to-soft mapping and Gumbel-Softmax is proposed to better model the latent distributions with rate constraint. Subjective results on multilingual speech datasets show that with a latency of 40ms, the proposed TF-Codec at 1kbps can achieve a much better quality than Opus 9kbps and TF-Codec at 3kbps outperforms both EVS 9.6kbps and Opus 12kbps. Numerous studies are conducted to show the effectiveness of these techniques.


【9】 End-to-End Spoken Language Understanding: Performance analyses of a  voice command task in a low resource setting

标题:端到端口语理解:低资源环境下语音命令任务的性能分析

链接:https://arxiv.org/abs/2207.08179

* 与cs.SD语音【4】为同一篇

作者:Thierry Desot,François Portet,Michel Vacher
备注:None
摘要:口语理解(SLU)是大多数人机交互系统中的核心任务,随着智能家居、智能手机和智能音箱的出现,SLU已经成为业界的关键技术,在经典的SLU方法中,自动语音识别(ASR)模块将语音信号转录成文本表示,自然语言理解从该文本表示(NLU)模块提取语义信息。最近端到端SLU基于深度神经网络的E2E SLU(E2E SLU)由于受益于ASR和NLU部分的联合优化,从而限制了流水线架构的级联错误效应,因此获得了发展势头。然而,E2E模型用来从语音输入中预测概念和意图的实际语言学性质还知之甚少,我们提出了一项研究,识别E2E模型使用的信号特征和其他语言属性,以执行SLU任务。英语Name(这里是法语)语音命令的测试结果。结果表明,良好的E2E SLU性能并不总是要求完美的ASR能力。此外,结果表明,与管道模型相比,E2E模型在处理背景噪声和语法变化方面具有更好的能力。最后,更细粒度的分析建议E2E模型使用输入信号的音调信息来识别话音命令概念。本文的结果和方法为进一步分析语音处理中的E2E模型提供了一个跳板。
摘要:Spoken Language Understanding (SLU) is a core task in most human-machine interaction systems. With the emergence of smart homes, smart phones and smart speakers, SLU has become a key technology for the industry. In a classical SLU approach, an Automatic Speech Recognition (ASR) module transcribes the speech signal into a textual representation from which a Natural Language Understanding (NLU) module extracts semantic information. Recently End-to-End SLU (E2E SLU) based on Deep Neural Networks has gained momentum since it benefits from the joint optimization of the ASR and the NLU parts, hence limiting the cascade of error effect of the pipeline architecture. However, little is known about the actual linguistic properties used by E2E models to predict concepts and intents from speech input. In this paper, we present a study identifying the signal features and other linguistic properties used by an E2E model to perform the SLU task. The study is carried out in the application domain of a smart home that has to handle non-English (here French) voice commands. The results show that a good E2E SLU performance does not always require a perfect ASR capability. Furthermore, the results show the superior capabilities of the E2E model in handling background noise and syntactic variation compared to the pipeline model. Finally, a finer-grained analysis suggests that the E2E model uses the pitch information of the input signal to identify voice command concepts. The results and methodology outlined in this paper provide a springboard for further analyses of E2E models in speech processing.


【10】 Visually-aware Acoustic Event Detection using Heterogeneous Graphs

标题:基于异构图的视觉感知声学事件检测

链接:https://arxiv.org/abs/2207.07935

* 与cs.SD语音【5】为同一篇

作者:Amir Shirian,Krishna Somandepalli,Victor Sanchez,Tanaya Guha
摘要:听觉事件的感知本质上是依赖于听觉和视觉线索的多模态的.现有的许多多模态方法使用模态特定的模型处理每一个模态,我们使用异构图来明确地捕获模态之间的空间和时间关系,并且表示关于潜在信号的详细信息。使用异构图方法来解决视觉感知的声学事件分类问题,异构图是一种简洁、高效和可扩展的数据表示方法,通过异构图,我们示出了在空间和时间尺度上对模态内和模态间关系的有效建模。在AudioSet上的实验表明,该模型具有较高的性能.
摘要:Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-specific models and then fuse the embeddings to encode the joint information. In contrast, we employ heterogeneous graphs to explicitly capture the spatial and temporal relationships between the modalities and represent detailed information about the underlying signal. Using heterogeneous graph approaches to address the task of visually-aware acoustic event classification, which serves as a compact, efficient and scalable way to represent data in the form of graphs. Through heterogeneous graphs, we show efficiently modelling of intra- and inter-modality relationships both at spatial and temporal scales. Our model can easily be adapted to different scales of events through relevant hyperparameters. Experiments on AudioSet, a large benchmark, shows that our model achieves state-of-the-art performance.


【11】 Few-shot bioacoustic event detection at the DCASE 2022 challenge

标题:在DCASE 2022挑战赛中进行的Few-Shot生物声学事件检测

链接:https://arxiv.org/abs/2207.07911

* 与cs.SD语音【6】为同一篇

作者:I. Nolasco,S. Singh,E. Vidana-Villa,E. Grout,J. Morford,M. Emmerson,F. Jensens,H. Whitehead,I. Kiskin,A. Strandburg-Peshkin,L. Gill,H. Pamula,V. Lostanlen,V. Morfi,D. Stowell
备注:submitted to DCASE2022 workshop
摘要:Few-Shot声音事件检测是检测声音事件的任务,尽管只具有感兴趣的类别的少数标记的例子,在这类情况下,通常需要注释很长的录音,但专家注释者的时间有限。本文概述了第二版的几个-DCASE 2022挑战中包含的生物声学事件检测任务。详细描述了任务目标、数据集和基线,以及所获得的主要结果和所提交系统的特性。该任务收到了来自15个不同团队的提交,其中13个团队的得分高于基线。评估集的最高F得分为60%,这导致了去年版本的巨大改进。高性能方法使用了原型网络,此外,通过分析每个子集上的结果,我们可以识别系统面临的主要困难,并得出结论,少显示的生物声学声音事件检测仍然是一个开放的挑战。
摘要:Few-shot sound event detection is the task of detecting sound events, despite having only a few labelled examples of the class of interest. This framework is particularly useful in bioacoustics, where often there is a need to annotate very long recordings but the expert annotator time is limited. This paper presents an overview of the second edition of the few-shot bioacoustic sound event detection task included in the DCASE 2022 challenge. A detailed description of the task objectives, dataset, and baselines is presented, together with the main results obtained and characteristics of the submitted systems. This task received submissions from 15 different teams from which 13 scored higher than the baselines. The highest F-score was of 60% on the evaluation set, which leads to a huge improvement over last year's edition. Highly-performing methods made use of prototypical networks, transductive learning, and addressed the variable length of events from all target classes. Furthermore, by analysing results on each of the subsets we can identify the main difficulties that the systems face, and conclude that few-show bioacoustic sound event detection remains an open challenge.



机器翻译,仅供参考