【3】 Efficient Transformer-based Speech Enhancement Using Long Frames and STFT Magnitudes
标题:基于变换的长帧和短时傅立叶变换语音增强
链接:https://arxiv.org/abs/2206.11703
* 与cs.SD语音【6】为同一篇
作者:Danilo de Oliveira,Tal Peer,Timo Gerkmann备注:Accepted at Interspeech 2022摘要:SepFormer结构在语音分离方面显示出很好的效果。与其他学习过的编码器模型一样,它使用短帧,因为它们在这些情况下可以获得更好的性能。这导致在输入端出现大量的帧,这是有问题的;由于SepFormer是基于转换器的,因此它的计算复杂性随着序列的延长而急剧增加。在本文中,我们将SepFormer应用于语音增强任务中,并表明通过用幅值短时傅立叶变换(STFT)表示替换学习的编码器特征,我们可以使用长帧而不会影响感知增强性能。我们获得了同等的质量和可理解性评估分数,同时将10秒的话语操作次数减少了大约8倍。摘要:The SepFormer architecture shows very good results in speech separation. Like other learned-encoder models, it uses short frames, as they have been shown to obtain better performance in these cases. This results in a large number of frames at the input, which is problematic; since the SepFormer is transformer-based, its computational complexity drastically increases with longer sequences. In this paper, we employ the SepFormer in a speech enhancement task and show that by replacing the learned-encoder features with a magnitude short-time Fourier transform (STFT) representation, we can use long frames without compromising perceptual enhancement performance. We obtained equivalent quality and intelligibility evaluation scores while reducing the number of operations by a factor of approximately 8 for a 10-second utterance.
【4】 Frequency Dependent Sound Event Detection for DCASE 2022 Challenge Task 4
标题:DCASE 2022挑战任务4的频率相关声音事件检测
链接:https://arxiv.org/abs/2206.11645
作者:Hyeonuk Nam,Seong-Hu Kim,Deokki Min,Byeong-Yun Ko,Seung-Deok Choi,Yong-Hwa Park备注:Technical Reprot submitted for DCASE2022 Challenge Task4摘要:虽然其他领域的许多深度学习方法已应用于声音事件检测(SED),但到目前为止,这些方法的原始领域和SED之间的差异尚未得到适当考虑。由于SED使用两个维度(时间和频率)的音频数据进行输入,因此对这两个维度的全面理解对于在SED上应用其他领域的方法至关重要。以前的工作证明,在SED中,那些在频率维度上寻址的方法尤其强大。通过应用滤波器增强(FilterAugment)和频率动态卷积(frequency dynamic convolution)这两种与频率相关的方法来提高SED性能,我们提交的模型获得了最佳PSDS1(0.4704)和最佳PSDS2(0.8224)。摘要:While many deep learning methods on other domains have been applied to sound event detection (SED), differences between original domains of the methods and SED have not been appropriately considered so far. As SED uses audio data with two dimensions (time and frequency) for input, thorough comprehension on these two dimensions is essential for application of methods from other domains on SED. Previous works proved that methods those address on frequency dimension are especially powerful in SED. By applying FilterAugment and frequency dynamic convolution those are frequency dependent methods proposed to enhance SED performance, our submitted models achieved best PSDS1 of 0.4704 and best PSDS2 of 0.8224.
【5】 Speaker-Independent Microphone Identification in Noisy Conditions
标题:噪声环境下的非特定人麦克风识别
链接:https://arxiv.org/abs/2206.11640
* 与cs.SD语音【7】为同一篇
作者:Antonio Giganti,Luca Cuccovillo,Paolo Bestagini,Patrick Aichroth,Stefano Tubaro
备注:To appear in: Proceedings of the 30th European Signal Processing Conference (EUSIPCO), August 20 -- September 2, 2022, Belgrade, Serbia
摘要:这项工作提出了一种从语音记录中识别源设备的方法,该方法应用基于神经网络的去噪,以减轻使用噪声注入的反取证攻击的影响。通过比较去噪对麦克风分类的三种最新特征的影响,确定它们在应用去噪和不应用去噪的情况下的识别能力,对该方法进行评估。所提出的框架实现了噪声材料的显著性能提高,并且更普遍地,验证了在对噪声记录进行设备识别之前应用去噪的有用性。
摘要:This work proposes a method for source device identification from speech recordings that applies neural-network-based denoising, to mitigate the impact of counter-forensics attacks using noise injection. The method is evaluated by comparing the impact of denoising on three state-of-the-art features for microphone classification, determining their discriminating power with and without denoising being applied. The proposed framework achieves a significant performance increase for noisy material, and more generally, validates the usefulness of applying denoising prior to device identification for noisy recordings.
【6】 Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems
标题:基于双程译码和互适应的端到端整形与混合TDNN ASR系统组合
链接:https://arxiv.org/abs/2206.11596
作者:Mingyu Cui,Jiajun Deng,Shoukang Hu,Xurong Xie,Tianzi Wang,Shujie Hu,Mengzhe Geng,Boyang Xue,Xunying Liu,Helen Meng
备注:It' s accepted to ISCA 2022
摘要:混合和端到端(E2E)自动语音识别(ASR)系统之间的基本建模差异在它们之间创造了巨大的多样性和互补性。本文针对混合TDNN和Conformer E2E ASR系统,研究了基于多通重扫描和交叉自适应的系统组合方法。在多通重排序中,采用最先进的混合LF-MMI训练的CNN-TDNN系统,该系统具有速度扰动、SpecAugment和贝叶斯学习隐藏单元贡献(LHUC)说话人自适应功能,以产生初始N最佳输出,然后由说话人自适应一致性系统使用双向交叉系统分数插值进行重排序。在交叉适应中,CNN-TDNN混合系统适用于一致性系统的1-最佳输出,反之亦然。在300小时的总机语料库上进行的实验表明,使用两种系统组合方法中的任何一种得到的组合系统都优于单独的系统。在NIST Hub5'00、Rt03和Rt02评估数据上,通过多次重新筛选获得的最佳组合系统与独立一致性系统相比,产生了2.5%-3.9%的统计显著字错误率(WER)绝对值(22.5%-28.9%相对值)。
摘要:Fundamental modelling differences between hybrid and end-to-end (E2E) automatic speech recognition (ASR) systems create large diversity and complementarity among them. This paper investigates multi-pass rescoring and cross adaptation based system combination approaches for hybrid TDNN and Conformer E2E ASR systems. In multi-pass rescoring, state-of-the-art hybrid LF-MMI trained CNN-TDNN system featuring speed perturbation, SpecAugment and Bayesian learning hidden unit contributions (LHUC) speaker adaptation was used to produce initial N-best outputs before being rescored by the speaker adapted Conformer system using a 2-way cross system score interpolation. In cross adaptation, the hybrid CNN-TDNN system was adapted to the 1-best output of the Conformer system or vice versa. Experiments on the 300-hour Switchboard corpus suggest that the combined systems derived using either of the two system combination approaches outperformed the individual systems. The best combined system obtained using multi-pass rescoring produced statistically significant word error rate (WER) reductions of 2.5% to 3.9% absolute (22.5% to 28.9% relative) over the stand alone Conformer system on the NIST Hub5'00, Rt03 and Rt02 evaluation data.
【7】 Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
标题:对抗性多任务学习在歌声合成中分离音色和基音的研究
链接:https://arxiv.org/abs/2206.11558
* 与cs.SD语音【8】为同一篇
作者:Tae-Woo Kim,Min-Su Kang,Gyeong-Hoon Lee备注:Accepted to INTERSPEECH 2022摘要:最近,引入了基于深度学习的生成模型来生成唱歌的声音。一种方法是预测由显式语音参数组成的参数声码器特征。这种方法的优点是可以明确区分每个特征的含义。另一种方法是预测神经声码器的mel谱图。然而,参数声码器存在着语音质量的局限性,而且由于音色和基音信息的纠缠,mel谱图特征很难建模。在这项研究中,我们提出了一个多任务学习的歌唱语音合成模型,该模型同时使用两种方法——参数声码器的声学特征和神经声码器的mel谱图。该模型利用参数声码器特征作为辅助特征,可以有效地分离和控制mel谱图的音色和基音分量。此外,在多歌手模型中,采用生成对抗网络框架来提高歌唱声音的质量。实验结果表明,与单任务模型相比,我们提出的模型可以生成更多的自然人声,同时性能优于传统的基于参数声码器的模型。摘要:Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the advantage that the meaning of each feature is explicitly distinguished. Another approach is to predict mel-spectrograms for a neural vocoder. However, parametric vocoders have limitations of voice quality and the mel-spectrogram features are difficult to model because the timbre and pitch information are entangled. In this study, we propose a singing voice synthesis model with multi-task learning to use both approaches -- acoustic features for a parametric vocoder and mel-spectrograms for a neural vocoder. By using the parametric vocoder features as auxiliary features, the proposed model can efficiently disentangle and control the timbre and pitch components of the mel-spectrogram. Moreover, a generative adversarial network framework is applied to improve the quality of singing voices in a multi-singer model. Experimental results demonstrate that our proposed model can generate more natural singing voices than the single-task models, while performing better than the conventional parametric vocoder-based model.
【8】 The SJTU X-LANCE Lab System for CNSRC 2022
标题:用于CNSRC 2022的SJTU X-Lance实验室系统
链接:https://arxiv.org/abs/2206.11699
* 与cs.SD语音【1】为同一篇
作者:Zhengyang Chen,Bei Liu,Bing Han,Leying Zhang,Yanmin Qian摘要:本技术报告描述了CNSRC 2022三条轨道的SJTU X-LANCE实验室系统。在本次挑战中,我们探讨了deep ResNet(deep r-vector)的说话人嵌入建模能力。所有系统仅在Cnceleb训练集上进行训练,我们在CNSRC 2022的三条轨道上使用相同的系统。在这一挑战中,我们的系统在说话人确认任务的固定轨道中排名第一。我们最好的单系统和聚变系统分别达到0.3164和0.2975 minDCF。此外,我们还将ResNet221的结果提交给说话人检索跟踪,获得了0.4626的mAP。摘要:This technical report describes the SJTU X-LANCE Lab system for the three tracks in CNSRC 2022. In this challenge, we explored the speaker embedding modeling ability of deep ResNet (Deeper r-vector). All the systems are only trained on the Cnceleb training set and we use the same systems for the three tracks in CNSRC 2022. In this challenge, our system ranks the first place in the fixed track of speaker verification task. Our best single system and fusion system achieve 0.3164 and 0.2975 minDCF respectively. Besides, we submit the result of ResNet221 to the speaker retrieval track and achieve 0.4626 mAP.
【9】 Towards Green ASR: Lossless 4-bit Quantization of a Hybrid TDNN System on the 300-hr Switchboard Corpus
标题:走向绿色ASR:在300小时交换板语料库上混合TDNN系统的4位无损量化
链接:https://arxiv.org/abs/2206.11643
* 与cs.SD语音【2】为同一篇
作者:Junhao Xu,Shoukang Hu,Xunying Liu,Helen Meng备注:Interspeech 2022 Accepted. arXiv admin note: text overlap with arXiv:2111.14479摘要:最先进的时间自动语音识别(ASR)系统在实际应用中变得越来越复杂和昂贵。本文介绍了一种基于300小时交换机语料库的高性能、低功耗4位量化LF-MMI训练因子时延神经网络(TDNNs)ASR系统的开发。整个系统设计的一个关键特征是考虑不同模型组件对量化误差的细粒度、不同性能敏感性。为此,使用了一组神经结构压缩和混合精度量化方法,以促进最佳因子TDNN权重矩阵子空间维数和量化比特宽度的隐层级自动配置。所提出的技术还用于生成2位混合精度量化Transformer语言模型。对交换机数据进行的实验表明,所提出的神经架构压缩和混合精度量化技术在字错误率(WER)方面始终优于具有可比比特宽度的统一精度量化基线系统。在基线全精度系统(包括TDNN和Transformer组件)上,总体“无损”压缩比为13.6,而WER没有出现统计上的显著增加。摘要:State of the art time automatic speech recognition (ASR) systems are becoming increasingly complex and expensive for practical applications. This paper presents the development of a high performance and low-footprint 4-bit quantized LF-MMI trained factored time delay neural networks (TDNNs) based ASR system on the 300-hr Switchboard corpus. A key feature of the overall system design is to account for the fine-grained, varying performance sensitivity at different model components to quantization errors. To this end, a set of neural architectural compression and mixed precision quantization approaches were used to facilitate hidden layer level auto-configuration of optimal factored TDNN weight matrix subspace dimensionality and quantization bit-widths. The proposed techniques were also used to produce 2-bit mixed precision quantized Transformer language models. Experiments conducted on the Switchboard data suggest that the proposed neural architectural compression and mixed precision quantization techniques consistently outperform the uniform precision quantised baseline systems of comparable bit-widths in terms of word error rate (WER). An overall "lossless" compression ratio of 13.6 was obtained over the baseline full precision system including both the TDNN and Transformer components while incurring no statistically significant WER increase.
【10】 Formant Estimation and Tracking using Probabilistic Heat-Maps
标题:基于概率热图的共振峰估计与跟踪
链接:https://arxiv.org/abs/2206.11632
* 与cs.SD语音【3】为同一篇
作者:Yosi Shrem,Felix Kreuk,Joseph Keshet摘要:共振峰是人类声道声学共振产生的频谱最大值,其准确估计是最基本的语音处理问题之一。最近的研究表明,使用深度学习技术可以准确估计这些频率。然而,当呈现来自不同领域的语音时,这些方法的性能会下降,限制了它们作为通用工具的使用。本文的贡献在于提出了一种新的网络体系结构,它可以在各种不同的说话人和语音域上运行良好。我们提出的模型由一个共享编码器组成,该编码器获取一个谱图作为输入,并输出一个域不变表示。然后,多个解码器进一步处理该表示,每个解码器负责预测不同的共振峰,同时考虑较低的共振峰预测。我们的模型的一个优点是,它基于热图,热图生成的概率分布超过共振峰预测。结果表明,我们提出的模型能够更好地代表各个领域的信号,并能够更好地跟踪和估计共振峰频率。摘要:Formants are the spectral maxima that result from acoustic resonances of the human vocal tract, and their accurate estimation is among the most fundamental speech processing problems. Recent work has been shown that those frequencies can accurately be estimated using deep learning techniques. However, when presented with a speech from a different domain than that in which they have been trained on, these methods exhibit a decline in performance, limiting their usage as generic tools. The contribution of this paper is to propose a new network architecture that performs well on a variety of different speaker and speech domains. Our proposed model is composed of a shared encoder that gets as input a spectrogram and outputs a domain-invariant representation. Then, multiple decoders further process this representation, each responsible for predicting a different formant while considering the lower formant predictions. An advantage of our model is that it is based on heatmaps that generate a probability distribution over formant predictions. Results suggest that our proposed model better represents the signal over various domains and leads to better formant frequency tracking and estimation.
【11】 Restoring speech intelligibility for hearing aid users with deep learning
标题:利用深度学习恢复助听器用户的语音清晰度
链接:https://arxiv.org/abs/2206.11567
* 与cs.SD语音【4】为同一篇
作者:Peter Udo Diehl,Yosef Singer,Hannes Zilly,Uwe Schönfeld,Paul Meyer-Rachner,Mark Berry,Henning Sprekeler,Elias Sprengel,Annett Pudszuhn,Veit M. Hofmann
摘要:全世界近5亿人患有致残性听力损失。虽然助听器可以部分弥补这一点,但在有背景噪声的情况下,很大一部分用户很难理解语音。在这里,我们提出了一种基于深度学习的算法,该算法在保持语音信号的同时选择性地抑制噪声。该算法将助听器用户的语音清晰度恢复到听力正常的对照受试者的水平。它由一个深度网络组成,该网络在一个含噪语音信号的大型自定义数据库上进行训练,并通过神经结构搜索进一步优化,使用一种新的基于深度学习的语音清晰度度量。该网络在一系列人类分级评估中实现了最先进的去噪,概括了不同的噪声类别,与经典的波束形成方法相比,它在单个麦克风上工作。该系统在笔记本电脑上实时运行,表明助听器芯片的大规模部署可能在几年内实现。因此,基于深度学习的去噪技术有可能很快改善数百万听力受损者的生活质量。
摘要:Almost half a billion people world-wide suffer from disabling hearing loss. While hearing aids can partially compensate for this, a large proportion of users struggle to understand speech in situations with background noise. Here, we present a deep learning-based algorithm that selectively suppresses noise while maintaining speech signals. The algorithm restores speech intelligibility for hearing aid users to the level of control subjects with normal hearing. It consists of a deep network that is trained on a large custom database of noisy speech signals and is further optimized by a neural architecture search, using a novel deep learning-based metric for speech intelligibility. The network achieves state-of-the-art denoising on a range of human-graded assessments, generalizes across different noise categories and - in contrast to classic beamforming approaches - operates on a single microphone. The system runs in real time on a laptop, suggesting that large-scale deployment on hearing aid chips could be achieved within a few years. Deep learning-based denoising therefore holds the potential to improve the quality of life of millions of hearing impaired people soon.
【12】 Few-shot Long-Tailed Bird Audio Recognition
标题:Few-Shot长尾鸟音频识别
链接:https://arxiv.org/abs/2206.11260
* 与cs.SD语音【5】为同一篇
作者:Marcos V. Conde,Ui-Jin Choi备注:BirdCLEF 202. Code and models at this https URL摘要:听鸟比看鸟容易。然而,它们在自然界中仍起着至关重要的作用,是环境质量和污染恶化的极好指标。机器学习和卷积神经网络的最新进展使我们能够处理连续的音频数据来检测和分类鸟类的声音。这项技术可以帮助研究人员监测鸟类种群的现状和趋势以及生态系统的生物多样性。我们提出了一种声音检测和分类管道,用于分析复杂的声景记录并识别背景中的鸟巢。我们的方法从弱标签和少量数据中学习,并从声学上识别鸟类物种。我们的解决方案在卡格尔举办的2022年BirdCLEF挑战赛中获得了807支球队中的第18名。摘要:It is easier to hear birds than see them. However, they still play an essential role in nature and are excellent indicators of deteriorating environmental quality and pollution. Recent advances in Machine Learning and Convolutional Neural Networks allow us to process continuous audio data to detect and classify bird sounds. This technology can assist researchers in monitoring bird populations' status and trends and ecosystems' biodiversity. We propose a sound detection and classification pipeline to analyze complex soundscape recordings and identify birdcalls in the background. Our method learns from weak labels and few data and acoustically recognizes the bird species. Our solution achieved 18th place of 807 teams at the BirdCLEF 2022 Challenge hosted on Kaggle.