cs.SD语音,共计11篇,eess.AS音频处理,共计11篇


1.cs.SD语音:

【1】 Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel's Weekly Video Podcasts

标题:默克尔播客语料库:一个由安格拉·默克尔16年每周视频播客汇编而成的多模式数据集

链接:https://arxiv.org/abs/2205.12194

作者:Debjoy Saha,Shravan Nayak,Timo Baumann
备注:Accepted at LREC 2022
摘要:我们介绍默克尔播客语料库,这是一个德语视听文本语料库,收集自德国前总理安吉拉·默克尔16年(几乎每周)的互联网播客。据我们所知,这是第一个德语单语语料库,由大小和时间范围相当的音频、视频和文本模式组成。我们描述了收集和编辑数据所使用的方法,包括下载视频、成绩单和其他元数据、强制对齐、执行主动说话人识别和人脸检测,以最终整理由安吉拉·默克尔所说话语组成的单说话人数据集。拟议的管道是通用的,可用于策划其他类似性质的数据集,如脱口秀内容。通过对该数据集的各种统计分析以及在语音合成和TTS中的应用,我们展示了该数据集的实用性。我们认为,这是对研究界的一项宝贵贡献,尤其是因为它在准备好的演讲和自发的演讲之间有着现实而富有挑战性的材料。
摘要:We introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel. To the best of our knowledge, this is the first single speaker corpus in the German language consisting of audio, visual and text modalities of comparable size and temporal extent. We describe the methods used with which we have collected and edited the data which involves downloading the videos, transcripts and other metadata, forced alignment, performing active speaker recognition and face detection to finally curate the single speaker dataset consisting of utterances spoken by Angela Merkel. The proposed pipeline is general and can be used to curate other datasets of similar nature, such as talk show contents. Through various statistical analyses and applications of the dataset in talking face generation and TTS, we show the utility of the dataset. We argue that it is a valuable contribution to the research community, in particular, due to its realistic and challenging material at the boundary between prepared and spontaneous speech.


【2】 Multi-Level Modeling Units for End-to-End Mandarin Speech Recognition

标题:用于端到端普通话语音识别的多层建模单元

链接:https://arxiv.org/abs/2205.11998

作者:Yuting Yang,Binbin Du,Yuke Li
备注:Submitted to INTERSPEECH2022
摘要:建模单元的选择影响声学建模的性能,在自动语音识别(ASR)中起着重要作用。在普通话场景中,汉字代表意义,但与发音没有直接关系。因此,仅将汉字的书写作为建模单元是不足以捕捉语音特征的。本文提出了一种融合多层次信息的多层次建模单元的汉语语音识别方法。具体而言,编码器块将音节视为建模单元,解码器块处理字符建模单元。在推理过程中,输入的特征序列由编码器块转换为音节序列,然后由解码器块转换为汉字。此过程由统一的端到端模型执行,无需引入其他转换模型。通过引入InterCE辅助任务,我们的方法分别使用Conformer和Transformer主干,在广泛使用的没有语言模型的AISHELL-1基准上取得了具有竞争力的结果,CER分别为4.1%/4.6%和4.6%/5.2%。
摘要:The choice of modeling units affects the performance of the acoustic modeling and plays an important role in automatic speech recognition (ASR). In mandarin scenarios, the Chinese characters represent meaning but are not directly related to the pronunciation. Thus only considering the writing of Chinese characters as modeling units is insufficient to capture speech features. In this paper, we present a novel method involves with multi-level modeling units, which integrates multi-level information for mandarin speech recognition. Specifically, the encoder block considers syllables as modeling units, and the decoder block deals with character modeling units. During inference, the input feature sequences are converted into syllable sequences by the encoder block and then converted into Chinese characters by the decoder block. This process is conducted by a unified end-to-end model without introducing additional conversion models. By introducing InterCE auxiliary task, our method achieves competitive results with CER of 4.1%/4.6% and 4.6%/5.2% on the widely used AISHELL-1 benchmark without a language model, using the Conformer and the Transformer backbones respectively.


【3】 SUSing: SU-net for Singing Voice Synthesis

标题:SUSING:歌唱语音合成的苏网

链接:https://arxiv.org/abs/2205.11841

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)
摘要:歌唱声音合成是一项生成性任务,涉及歌唱模型的多维控制,包括歌词、音高和持续时间,还包括歌手的音色和颤音等歌唱技能。本文提出了一种用于歌唱语音合成的SU网SUSing。将合成歌声视为歌词与乐谱和谱之间的翻译任务。歌词和乐谱信息通过卷积层编码为二维特征表示。二维特征及其频谱通过SU网络以自回归的方式映射到目标频谱。在SU网络中,使用条带池方法代替交替全局池方法来学习频谱中的垂直频率关系和时域中的频率变化。在公共数据集Kiritan上的实验结果表明,该方法可以合成更自然的人声。
摘要:Singing voice synthesis is a generative task that involves multi-dimensional control of the singing model, including lyrics, pitch, and duration, and includes the timbre of the singer and singing skills such as vibrato. In this paper, we proposed SU-net for singing voice synthesis named SUSing. Synthesizing singing voice is treated as a translation task between lyrics and music score and spectrum. The lyrics and music score information is encoded into a two-dimensional feature representation through the convolution layer. The two-dimensional feature and its frequency spectrum are mapped to the target spectrum in an autoregressive manner through a SU-net network. Within the SU-net the stripe pooling method is used to replace the alternate global pooling method to learn the vertical frequency relationship in the spectrum and the changes of frequency in the time domain. The experimental results on the public dataset Kiritan show that the proposed method can synthesize more natural singing voices.


【4】 TDASS: Target Domain Adaptation Speech Synthesis Framework for Multi-speaker Low-Resource TTS

标题:TDASS:面向多说话人低资源TTS的目标域自适应语音合成框架

链接:https://arxiv.org/abs/2205.11824

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)
摘要:近年来,通过文本到语音(TTS)应用合成个性化语音的需求越来越高。但是以前的TTS模型需要大量的目标演讲者演讲来进行训练。这是一项成本很高的任务,很难记录目标说话者的大量话语。语音数据增强是一种解决方案,但会导致合成语音质量低下的问题。为了解决这一问题,提出了一些多扬声器TTS模型。但由于每个说话人的话语量不平衡,导致了语音相似度问题。我们提出了目标域自适应语音合成网络(TDASS)来解决这些问题。TDASS基于Tacotron2模型(高质量TTS模型)的主干,引入了一个自感兴趣的分类器,以减少非目标影响。此外,在分类器中加入了一个特殊的梯度反转层,对目标和非目标进行不同的操作。我们在一个汉语语音语料库上对该模型进行了评估,实验表明,该方法在语音质量和语音相似度方面优于基线方法。
摘要:Recently, synthesizing personalized speech by text-to-speech (TTS) application is highly demanded. But the previous TTS models require a mass of target speaker speeches for training. It is a high-cost task, and hard to record lots of utterances from the target speaker. Data augmentation of the speeches is a solution but leads to the low-quality synthesis speech problem. Some multi-speaker TTS models are proposed to address the issue. But the quantity of utterances of each speaker imbalance leads to the voice similarity problem. We propose the Target Domain Adaptation Speech Synthesis Network (TDASS) to address these issues. Based on the backbone of the Tacotron2 model, which is the high-quality TTS model, TDASS introduces a self-interested classifier for reducing the non-target influence. Besides, a special gradient reversal layer with different operations for target and non-target is added to the classifier. We evaluate the model on a Chinese speech corpus, the experiments show the proposed method outperforms the baseline method in terms of voice quality and voice similarity.


【5】 MetaSID: Singer Identification with Domain Adaptation for Metaverse

标题:MetaSID:Metverse领域适配的歌手识别

链接:https://arxiv.org/abs/2205.11821

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)
摘要:Metaverse将现实世界扩展到了无限的空间。Metaverse将有更多现场音乐会。歌手识别的任务是识别歌曲属于哪个歌手。然而,歌手识别一直存在一个难题,即现场效果的不同。studio版本与live版本不同,训练集和测试集的数据分布不同,分类器性能下降。本文提出了利用领域自适应方法解决歌手识别中的现场效果问题。设计了三种与卷积递归神经网络(CRNN)相结合的域自适应方法,即最大均值差(MMD)、梯度反转(Revgrad)和对比自适应网络(CAN)。MMD是一种基于距离的方法,它增加了域丢失。Revgrad基于这样一种思想,即学习的特征可以表示不同的领域样本。CAN是基于类自适应的,它考虑了源域和目标域类别之间的对应关系。在Artist20公共数据集上的实验结果表明,CRNN-MMD比基线CRNN提高了0.14。CRNN RevGrad的表现优于基线0.21。在唱片集分割上,CRNN-F1测量值为0.83,可以达到最先进的水平。
摘要:Metaverse has stretched the real world into unlimited space. There will be more live concerts in Metaverse. The task of singer identification is to identify the song belongs to which singer. However, there has been a tough problem in singer identification, which is the different live effects. The studio version is different from the live version, the data distribution of the training set and the test set are different, and the performance of the classifier decreases. This paper proposes the use of the domain adaptation method to solve the live effect in singer identification. Three methods of domain adaptation combined with Convolutional Recurrent Neural Network (CRNN) are designed, which are Maximum Mean Discrepancy (MMD), gradient reversal (Revgrad), and Contrastive Adaptation Network (CAN). MMD is a distance-based method, which adds domain loss. Revgrad is based on the idea that learned features can represent different domain samples. CAN is based on class adaptation, it takes into account the correspondence between the categories of the source domain and target domain. Experimental results on the public dataset of Artist20 show that CRNN-MMD leads to an improvement over the baseline CRNN by 0.14. The CRNN-RevGrad outperforms the baseline by 0.21. The CRNN-CAN achieved state of the art with the F1 measure value of 0.83 on album split.


【6】 Singer Identification for Metaverse with Timbral and Middle-Level Perceptual Features

标题:具有音质和中层知觉特征的超时空演唱者识别

链接:https://arxiv.org/abs/2205.11817

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks). arXiv admin note: text overlap with arXiv:2002.06817 by other authors
摘要:Metaverse是一个结合现实和虚拟的互动世界,参与者可以是虚拟化身。任何人都可以在虚拟音乐厅举行音乐会,用户可以通过歌手识别快速识别虚拟偶像背后的真实歌手。大多数歌手识别方法都是使用帧级特征进行处理的。然而,除了歌手的音色之外,音乐框架还包括音乐信息,如旋律、节奏和音调。这意味着音乐信息是噪声,用于使用帧级特征识别歌手。在本文中,我们建议使用另外两个解决此问题的功能,而不仅仅是帧级功能。中级特征,代表音乐的旋律性、节奏稳定性和音调稳定性,能够捕捉音乐的感知特征。用于说话人识别的音色特征代表了歌手的语音特征。此外,我们提出了一种卷积递归神经网络(CRNN)来结合三种特征进行歌手识别。该模型首先融合帧级特征和音色特征,然后将中间级特征与混合特征相结合。在实验中,该方法在Artist20的基准数据集上的平均F1得分为0.81,取得了可比的性能,显著改进了相关工作。
摘要:Metaverse is an interactive world that combines reality and virtuality, where participants can be virtual avatars. Anyone can hold a concert in a virtual concert hall, and users can quickly identify the real singer behind the virtual idol through the singer identification. Most singer identification methods are processed using the frame-level features. However, expect the singer's timbre, the music frame includes music information, such as melodiousness, rhythm, and tonal. It means the music information is noise for using frame-level features to identify the singers. In this paper, instead of only the frame-level features, we propose to use another two features that address this problem. Middle-level feature, which represents the music's melodiousness, rhythmic stability, and tonal stability, and is able to capture the perceptual features of music. The timbre feature, which is used in speaker identification, represents the singers' voice features. Furthermore, we propose a convolutional recurrent neural network (CRNN) to combine three features for singer identification. The model firstly fuses the frame-level feature and timbre feature and then combines middle-level features to the mix features. In experiments, the proposed method achieves comparable performance on an average F1 score of 0.81 on the benchmark dataset of Artist20, which significantly improves related works.


【7】 Deep Learning-based automated classification of Chinese Speech Sound Disorders

标题:基于深度学习的汉语语音障碍自动分类

链接:https://arxiv.org/abs/2205.11748

作者:Yao-Ming Kuo,Shanq-Jang Ruan,Yu-Chin Chen,Ya-Wen Tu
备注:12 pages, 9 figures, journal
摘要:本文描述了一个分析声学数据的系统,以帮助使用计算机诊断和分类儿童言语障碍。分析集中于识别和分类四种不同类型的汉语误解。该研究收集并生成了一个语音语料库,其中包含2540个停音、软腭音、辅音元音和塞擦音样本,这些样本来自90名3-6岁具有正常或病理性发音特征的儿童。每次录音都附有言语治疗领域的详细注释。语音样本的分类是使用三个建立良好的神经网络模型进行图像分类的。特征映射是使用从语音中提取的三组MFCC参数创建的,并聚合成三维数据结构作为模型输入。我们采用六种数据扩充技术来扩充可用数据集,同时避免过度模拟。实验考察了四种不同类别的汉语短语和汉字的可用性。对不同数据子集的实验表明,该系统能够准确地检测出所分析的发音障碍。
摘要:This article describes a system for analyzing acoustic data in order to assist in the diagnosis and classification of children's speech disorders using a computer. The analysis concentrated on identifying and categorizing four distinct types of Chinese misconstructions. The study collected and generated a speech corpus containing 2540 Stopping, Velar, Consonant-vowel, and Affricate samples from 90 children aged 3-6 years with normal or pathological articulatory features. Each recording was accompanied by a detailed annotation from the field of speech therapy. Classification of the speech samples was accomplished using three well-established neural network models for image classification. The feature maps are created using three sets of MFCC parameters extracted from speech sounds and aggregated into a three-dimensional data structure as model input. We employ six techniques for data augmentation in order to augment the available dataset while avoiding over-simulation. The experiments examine the usability of four different categories of Chinese phrases and characters. Experiments with different data subsets demonstrate the system's ability to accurately detect the analyzed pronunciation disorders.


【8】 Adaptive Few-Shot Learning Algorithm for Rare Sound Event Detection

标题:稀有声音事件检测的自适应Few-Shot学习算法

链接:https://arxiv.org/abs/2205.11738

作者:Chendong Zhao,Jianzong Wang,Leilai Li,Xiaoyang Qu,Jing Xiao
备注:Accepted to IJCNN 2022. arXiv admin note: text overlap with arXiv:2110.04474 by other authors摘要:声音事件检测是通过了解周围环境的声音来推断事件。由于罕见声音事件的稀缺性,对于已经掌握了太多先验知识的训练有素的探测器来说,这是一个挑战。同时,很少有镜头学习方法在面对新的有限数据任务时具有良好的泛化能力。最近的方法在这一领域取得了可喜的成果。然而,这些方法独立地处理每个支持示例,忽略了整个任务中其他示例的信息。因此,以前的大多数方法都被限制为为所有测试时任务生成相同的特征嵌入,这对每个输入数据都是不自适应的。在这项工作中,我们提出了一种新的任务自适应模块,该模块易于植入任何基于度量的Few-Shot学习框架中。该模块可以识别与任务相关的特征维度。与基线方法相比,结合我们的模块可以显著提高两个数据集的性能,特别是对于直传传播网络。例如,ESC-50上的5向单发精度为+6.8%,而noiseESC-50上的精度为+5.9%。我们在域不匹配设置中研究了我们的方法,并取得了比以前的方法更好的结果。
摘要:Sound event detection is to infer the event by understanding the surrounding environmental sounds. Due to the scarcity of rare sound events, it becomes challenging for the well-trained detectors which have learned too much prior knowledge. Meanwhile, few-shot learning methods promise a good generalization ability when facing a new limited-data task. Recent approaches have achieved promising results in this field. However, these approaches treat each support example independently, ignoring the information of other examples from the whole task. Because of this, most of previous methods are constrained to generate a same feature embedding for all test-time tasks, which is not adaptive to each inputted data. In this work, we propose a novel task-adaptive module which is easy to plant into any metric-based few-shot learning frameworks. The module could identify the task-relevant feature dimension. Incorporating our module improves the performance considerably on two datasets over baseline methods, especially for the transductive propagation network. Such as +6.8% for 5-way 1-shot accuracy on ESC-50, and +5.9% on noiseESC-50. We investigate our approach in the domain-mismatch setting and also achieve better results than previous methods.


【9】 Defending a Music Recommender Against Hubness-Based Adversarial Attacks

标题:保护音乐推荐器免受基于Hubness的恶意攻击

链接:https://arxiv.org/abs/2205.12032

作者:Katharina Hoedt,Arthur Flexer,Gerhard Widmer
备注:6 pages, to be published in Proceedings of the 19th Sound and Music Computing Conference 2022 (SMC-22)
摘要:对抗性攻击可以极大地降低推荐程序和其他机器学习系统的性能,导致对防御机制的需求增加。我们为攻击提供了一条新的防线,这些攻击利用了在高维数据空间中操作的推荐者的漏洞(所谓的Hubbness问题)。我们使用一种全局数据缩放方法,即相互接近(MP),来保护现实世界中的音乐推荐者,该推荐者以前容易受到攻击,从而夸大了特定歌曲的推荐次数。我们发现,使用MP作为防御可以极大地提高推荐者对一系列攻击的鲁棒性,攻击成功率约为44%(防御前)降至6%(防御后)。此外,对抗性示例仍然能够欺骗防御系统,但代价是音频质量明显降低,平均信噪比降低。
摘要:Adversarial attacks can drastically degrade performance of recommenders and other machine learning systems, resulting in an increased demand for defence mechanisms. We present a new line of defence against attacks which exploit a vulnerability of recommenders that operate in high dimensional data spaces (the so-called hubness problem). We use a global data scaling method, namely Mutual Proximity (MP), to defend a real-world music recommender which previously was susceptible to attacks that inflated the number of times a particular song was recommended. We find that using MP as a defence greatly increases robustness of the recommender against a range of attacks, with success rates of attacks around 44% (before defence) dropping to less than 6% (after defence). Additionally, adversarial examples still able to fool the defended system do so at the price of noticeably lower audio quality as shown by a decreased average SNR.


【10】 PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

标题:PaddleSpeech:一个简单易用的多功能语音工具包

链接:https://arxiv.org/abs/2205.12007

作者:Hui Zhang,Tian Yuan,Junkun Chen,Xintong Li,Renjie Zheng,Yuxin Huang,Xiaojie Chen,Enlei Gong,Zeyu Chen,Xiaoguang Hu,Dianhai Yu,Yanjun Ma,Liang Huang
摘要:PadderSpeech是一个开源的多功能语音工具包。它旨在通过提供易于使用的命令行界面和简单的代码结构,促进语音处理技术的开发和研究。本文描述了PadleSpeech的设计理念和核心架构,以支持几个基本的语音到文本和文本到语音任务。PadleSpeech在各种语音数据集上实现了极具竞争力或最先进的性能,并实现了最流行的方法。它还提供了配方和预训练模型,以快速再现本文中的实验结果。PaddleSpeech可在以下网址公开获取:https://github.com/PaddlePaddle/PaddleSpeech.
摘要:PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleSpeech to support several essential speech-to-text and text-to-speech tasks. PaddleSpeech achieves competitive or state-of-the-art performance on various speech datasets and implements the most popular methods. It also provides recipes and pretrained models to quickly reproduce the experimental results in this paper. PaddleSpeech is publicly avaiable at https://github.com/PaddlePaddle/PaddleSpeech.


【11】 SepIt Approaching a Single Channel Speech Separation Bound

标题:逐个逼近单通道语音分离界

链接:https://arxiv.org/abs/2205.11801

作者:Shahar Lutati,Eliya Nachmani,Lior Wolf
摘要:我们提出了单通道语音分离任务的上界,该上界基于一个关于短语音段性质的假设。使用边界,我们能够表明,虽然最近的方法对少数发言者取得了重大进展,但对五名和十名发言者来说,仍有改进的空间。然后,我们引入了一个深度神经网络SepIt,它可以迭代地改进不同说话人的估计。在测试时,基于我们分析得出的互信息标准,SpeIt对每个测试样本有不同的迭代次数。在一系列广泛的实验中,SepIt对于2、3、5和10个扬声器的性能优于最先进的神经网络。
摘要:We present an upper bound for the Single Channel Speech Separation task, which is based on an assumption regarding the nature of short segments of speech. Using the bound, we are able to show that while the recent methods have made significant progress for a few speakers, there is room for improvement for five and ten speakers. We then introduce a Deep neural network, SepIt, that iteratively improves the different speakers' estimation. At test time, SpeIt has a varying number of iterations per test sample, based on a mutual information criterion that arises from our analysis. In an extensive set of experiments, SepIt outperforms the state-of-the-art neural networks for 2, 3, 5, and 10 speakers.


2.eess.AS音频处理:

【1】 Defending a Music Recommender Against Hubness-Based Adversarial Attacks

标题:保护音乐推荐器免受基于Hubness的恶意攻击

链接:https://arxiv.org/abs/2205.12032

作者:Katharina Hoedt,Arthur Flexer,Gerhard Widmer
备注:6 pages, to be published in Proceedings of the 19th Sound and Music Computing Conference 2022 (SMC-22)
摘要:对抗性攻击可以极大地降低推荐程序和其他机器学习系统的性能,导致对防御机制的需求增加。我们为攻击提供了一条新的防线,这些攻击利用了在高维数据空间中操作的推荐者的漏洞(所谓的Hubbness问题)。我们使用一种全局数据缩放方法,即相互接近(MP),来保护现实世界中的音乐推荐者,该推荐者以前容易受到攻击,从而夸大了特定歌曲的推荐次数。我们发现,使用MP作为防御可以极大地提高推荐者对一系列攻击的鲁棒性,攻击成功率约为44%(防御前)降至6%(防御后)。此外,对抗性示例仍然能够欺骗防御系统,但代价是音频质量明显降低,平均信噪比降低。
摘要:Adversarial attacks can drastically degrade performance of recommenders and other machine learning systems, resulting in an increased demand for defence mechanisms. We present a new line of defence against attacks which exploit a vulnerability of recommenders that operate in high dimensional data spaces (the so-called hubness problem). We use a global data scaling method, namely Mutual Proximity (MP), to defend a real-world music recommender which previously was susceptible to attacks that inflated the number of times a particular song was recommended. We find that using MP as a defence greatly increases robustness of the recommender against a range of attacks, with success rates of attacks around 44% (before defence) dropping to less than 6% (after defence). Additionally, adversarial examples still able to fool the defended system do so at the price of noticeably lower audio quality as shown by a decreased average SNR.


【2】 PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

标题:PaddleSpeech:一个简单易用的多功能语音工具包

链接:https://arxiv.org/abs/2205.12007

作者:Hui Zhang,Tian Yuan,Junkun Chen,Xintong Li,Renjie Zheng,Yuxin Huang,Xiaojie Chen,Enlei Gong,Zeyu Chen,Xiaoguang Hu,Dianhai Yu,Yanjun Ma,Liang Huang
摘要:PadderSpeech是一个开源的多功能语音工具包。它旨在通过提供易于使用的命令行界面和简单的代码结构,促进语音处理技术的开发和研究。本文描述了PadleSpeech的设计理念和核心架构,以支持几个基本的语音到文本和文本到语音任务。PadleSpeech在各种语音数据集上实现了极具竞争力或最先进的性能,并实现了最流行的方法。它还提供了配方和预训练模型,以快速再现本文中的实验结果。PaddleSpeech可在以下网址公开获取:https://github.com/PaddlePaddle/PaddleSpeech.
摘要:PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleSpeech to support several essential speech-to-text and text-to-speech tasks. PaddleSpeech achieves competitive or state-of-the-art performance on various speech datasets and implements the most popular methods. It also provides recipes and pretrained models to quickly reproduce the experimental results in this paper. PaddleSpeech is publicly avaiable at https://github.com/PaddlePaddle/PaddleSpeech.


【3】 SepIt Approaching a Single Channel Speech Separation Bound

标题:逐个逼近单通道语音分离界

链接:https://arxiv.org/abs/2205.11801

作者:Shahar Lutati,Eliya Nachmani,Lior Wolf
摘要:我们提出了单通道语音分离任务的上界,该上界基于一个关于短语音段性质的假设。使用边界,我们能够表明,虽然最近的方法对少数发言者取得了重大进展,但对五名和十名发言者来说,仍有改进的空间。然后,我们引入了一个深度神经网络SepIt,它可以迭代地改进不同说话人的估计。在测试时,基于我们分析得出的互信息标准,SpeIt对每个测试样本有不同的迭代次数。在一系列广泛的实验中,SepIt对于2、3、5和10个扬声器的性能优于最先进的神经网络。
摘要:We present an upper bound for the Single Channel Speech Separation task, which is based on an assumption regarding the nature of short segments of speech. Using the bound, we are able to show that while the recent methods have made significant progress for a few speakers, there is room for improvement for five and ten speakers. We then introduce a Deep neural network, SepIt, that iteratively improves the different speakers' estimation. At test time, SpeIt has a varying number of iterations per test sample, based on a mutual information criterion that arises from our analysis. In an extensive set of experiments, SepIt outperforms the state-of-the-art neural networks for 2, 3, 5, and 10 speakers.


【4】 Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel's Weekly Video Podcasts

标题:默克尔播客语料库:一个由安格拉·默克尔16年每周视频播客汇编而成的多模式数据集

链接:https://arxiv.org/abs/2205.12194

作者:Debjoy Saha,Shravan Nayak,Timo Baumann
备注:Accepted at LREC 2022
摘要:我们介绍默克尔播客语料库,这是一个德语视听文本语料库,收集自德国前总理安吉拉·默克尔16年(几乎每周)的互联网播客。据我们所知,这是第一个德语单语语料库,由大小和时间范围相当的音频、视频和文本模式组成。我们描述了收集和编辑数据所使用的方法,包括下载视频、成绩单和其他元数据、强制对齐、执行主动说话人识别和人脸检测,以最终整理由安吉拉·默克尔所说话语组成的单说话人数据集。拟议的管道是通用的,可用于策划其他类似性质的数据集,如脱口秀内容。通过对该数据集的各种统计分析以及在语音合成和TTS中的应用,我们展示了该数据集的实用性。我们认为,这是对研究界的一项宝贵贡献,尤其是因为它在准备好的演讲和自发的演讲之间有着现实而富有挑战性的材料。
摘要:We introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel. To the best of our knowledge, this is the first single speaker corpus in the German language consisting of audio, visual and text modalities of comparable size and temporal extent. We describe the methods used with which we have collected and edited the data which involves downloading the videos, transcripts and other metadata, forced alignment, performing active speaker recognition and face detection to finally curate the single speaker dataset consisting of utterances spoken by Angela Merkel. The proposed pipeline is general and can be used to curate other datasets of similar nature, such as talk show contents. Through various statistical analyses and applications of the dataset in talking face generation and TTS, we show the utility of the dataset. We argue that it is a valuable contribution to the research community, in particular, due to its realistic and challenging material at the boundary between prepared and spontaneous speech.


【5】 Multi-Level Modeling Units for End-to-End Mandarin Speech Recognition

标题:用于端到端普通话语音识别的多层建模单元

链接:https://arxiv.org/abs/2205.11998

作者:Yuting Yang,Binbin Du,Yuke Li
备注:Submitted to INTERSPEECH2022
摘要:建模单元的选择影响声学建模的性能,在自动语音识别(ASR)中起着重要作用。在普通话场景中,汉字代表意义,但与发音没有直接关系。因此,仅将汉字的书写作为建模单元是不足以捕捉语音特征的。本文提出了一种融合多层次信息的多层次建模单元的汉语语音识别方法。具体而言,编码器块将音节视为建模单元,解码器块处理字符建模单元。在推理过程中,输入的特征序列由编码器块转换为音节序列,然后由解码器块转换为汉字。此过程由统一的端到端模型执行,无需引入其他转换模型。通过引入InterCE辅助任务,我们的方法分别使用Conformer和Transformer主干,在广泛使用的没有语言模型的AISHELL-1基准上取得了具有竞争力的结果,CER分别为4.1%/4.6%和4.6%/5.2%。
摘要:The choice of modeling units affects the performance of the acoustic modeling and plays an important role in automatic speech recognition (ASR). In mandarin scenarios, the Chinese characters represent meaning but are not directly related to the pronunciation. Thus only considering the writing of Chinese characters as modeling units is insufficient to capture speech features. In this paper, we present a novel method involves with multi-level modeling units, which integrates multi-level information for mandarin speech recognition. Specifically, the encoder block considers syllables as modeling units, and the decoder block deals with character modeling units. During inference, the input feature sequences are converted into syllable sequences by the encoder block and then converted into Chinese characters by the decoder block. This process is conducted by a unified end-to-end model without introducing additional conversion models. By introducing InterCE auxiliary task, our method achieves competitive results with CER of 4.1%/4.6% and 4.6%/5.2% on the widely used AISHELL-1 benchmark without a language model, using the Conformer and the Transformer backbones respectively.


【6】 SUSing: SU-net for Singing Voice Synthesis

标题:SUSING:歌唱语音合成的苏网

链接:https://arxiv.org/abs/2205.11841

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)
摘要:歌唱声音合成是一项生成性任务,涉及歌唱模型的多维控制,包括歌词、音高和持续时间,还包括歌手的音色和颤音等歌唱技能。本文提出了一种用于歌唱语音合成的SU网SUSing。将合成歌声视为歌词与乐谱和谱之间的翻译任务。歌词和乐谱信息通过卷积层编码为二维特征表示。二维特征及其频谱通过SU网络以自回归的方式映射到目标频谱。在SU网络中,使用条带池方法代替交替全局池方法来学习频谱中的垂直频率关系和时域中的频率变化。在公共数据集Kiritan上的实验结果表明,该方法可以合成更自然的人声。
摘要:Singing voice synthesis is a generative task that involves multi-dimensional control of the singing model, including lyrics, pitch, and duration, and includes the timbre of the singer and singing skills such as vibrato. In this paper, we proposed SU-net for singing voice synthesis named SUSing. Synthesizing singing voice is treated as a translation task between lyrics and music score and spectrum. The lyrics and music score information is encoded into a two-dimensional feature representation through the convolution layer. The two-dimensional feature and its frequency spectrum are mapped to the target spectrum in an autoregressive manner through a SU-net network. Within the SU-net the stripe pooling method is used to replace the alternate global pooling method to learn the vertical frequency relationship in the spectrum and the changes of frequency in the time domain. The experimental results on the public dataset Kiritan show that the proposed method can synthesize more natural singing voices.


【7】 TDASS: Target Domain Adaptation Speech Synthesis Framework for Multi-speaker Low-Resource TTS

标题:TDASS:面向多说话人低资源TTS的目标域自适应语音合成框架

链接:https://arxiv.org/abs/2205.11824

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)
摘要:近年来,通过文本到语音(TTS)应用合成个性化语音的需求越来越高。但是以前的TTS模型需要大量的目标演讲者演讲来进行训练。这是一项成本很高的任务,很难记录目标说话者的大量话语。语音数据增强是一种解决方案,但会导致合成语音质量低下的问题。为了解决这一问题,提出了一些多扬声器TTS模型。但由于每个说话人的话语量不平衡,导致了语音相似度问题。我们提出了目标域自适应语音合成网络(TDASS)来解决这些问题。TDASS基于Tacotron2模型(高质量TTS模型)的主干,引入了一个自感兴趣的分类器,以减少非目标影响。此外,在分类器中加入了一个特殊的梯度反转层,对目标和非目标进行不同的操作。我们在一个汉语语音语料库上对该模型进行了评估,实验表明,该方法在语音质量和语音相似度方面优于基线方法。
摘要:Recently, synthesizing personalized speech by text-to-speech (TTS) application is highly demanded. But the previous TTS models require a mass of target speaker speeches for training. It is a high-cost task, and hard to record lots of utterances from the target speaker. Data augmentation of the speeches is a solution but leads to the low-quality synthesis speech problem. Some multi-speaker TTS models are proposed to address the issue. But the quantity of utterances of each speaker imbalance leads to the voice similarity problem. We propose the Target Domain Adaptation Speech Synthesis Network (TDASS) to address these issues. Based on the backbone of the Tacotron2 model, which is the high-quality TTS model, TDASS introduces a self-interested classifier for reducing the non-target influence. Besides, a special gradient reversal layer with different operations for target and non-target is added to the classifier. We evaluate the model on a Chinese speech corpus, the experiments show the proposed method outperforms the baseline method in terms of voice quality and voice similarity.


【8】 MetaSID: Singer Identification with Domain Adaptation for Metaverse

标题:MetaSID:Metverse领域适配的歌手识别

链接:https://arxiv.org/abs/2205.11821

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)
摘要:Metaverse将现实世界扩展到了无限的空间。Metaverse将有更多现场音乐会。歌手识别的任务是识别歌曲属于哪个歌手。然而,歌手识别一直存在一个难题,即现场效果的不同。studio版本与live版本不同,训练集和测试集的数据分布不同,分类器性能下降。本文提出了利用领域自适应方法解决歌手识别中的现场效果问题。设计了三种与卷积递归神经网络(CRNN)相结合的域自适应方法,即最大均值差(MMD)、梯度反转(Revgrad)和对比自适应网络(CAN)。MMD是一种基于距离的方法,它增加了域丢失。Revgrad基于这样一种思想,即学习的特征可以表示不同的领域样本。CAN是基于类自适应的,它考虑了源域和目标域类别之间的对应关系。在Artist20公共数据集上的实验结果表明,CRNN-MMD比基线CRNN提高了0.14。CRNN RevGrad的表现优于基线0.21。在唱片集分割上,CRNN-F1测量值为0.83,可以达到最先进的水平。
摘要:Metaverse has stretched the real world into unlimited space. There will be more live concerts in Metaverse. The task of singer identification is to identify the song belongs to which singer. However, there has been a tough problem in singer identification, which is the different live effects. The studio version is different from the live version, the data distribution of the training set and the test set are different, and the performance of the classifier decreases. This paper proposes the use of the domain adaptation method to solve the live effect in singer identification. Three methods of domain adaptation combined with Convolutional Recurrent Neural Network (CRNN) are designed, which are Maximum Mean Discrepancy (MMD), gradient reversal (Revgrad), and Contrastive Adaptation Network (CAN). MMD is a distance-based method, which adds domain loss. Revgrad is based on the idea that learned features can represent different domain samples. CAN is based on class adaptation, it takes into account the correspondence between the categories of the source domain and target domain. Experimental results on the public dataset of Artist20 show that CRNN-MMD leads to an improvement over the baseline CRNN by 0.14. The CRNN-RevGrad outperforms the baseline by 0.21. The CRNN-CAN achieved state of the art with the F1 measure value of 0.83 on album split.


【9】 Singer Identification for Metaverse with Timbral and Middle-Level Perceptual Features

标题:具有音质和中层知觉特征的超时空演唱者识别

链接:https://arxiv.org/abs/2205.11817

作者:Xulong Zhang,Jianzong Wang,Ning Cheng,Jing Xiao
备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks). arXiv admin note: text overlap with arXiv:2002.06817 by other authors
摘要:Metaverse是一个结合现实和虚拟的互动世界,参与者可以是虚拟化身。任何人都可以在虚拟音乐厅举行音乐会,用户可以通过歌手识别快速识别虚拟偶像背后的真实歌手。大多数歌手识别方法都是使用帧级特征进行处理的。然而,除了歌手的音色之外,音乐框架还包括音乐信息,如旋律、节奏和音调。这意味着音乐信息是噪声,用于使用帧级特征识别歌手。在本文中,我们建议使用另外两个解决此问题的功能,而不仅仅是帧级功能。中级特征,代表音乐的旋律性、节奏稳定性和音调稳定性,能够捕捉音乐的感知特征。用于说话人识别的音色特征代表了歌手的语音特征。此外,我们提出了一种卷积递归神经网络(CRNN)来结合三种特征进行歌手识别。该模型首先融合帧级特征和音色特征,然后将中间级特征与混合特征相结合。在实验中,该方法在Artist20的基准数据集上的平均F1得分为0.81,取得了可比的性能,显著改进了相关工作。
摘要:Metaverse is an interactive world that combines reality and virtuality, where participants can be virtual avatars. Anyone can hold a concert in a virtual concert hall, and users can quickly identify the real singer behind the virtual idol through the singer identification. Most singer identification methods are processed using the frame-level features. However, expect the singer's timbre, the music frame includes music information, such as melodiousness, rhythm, and tonal. It means the music information is noise for using frame-level features to identify the singers. In this paper, instead of only the frame-level features, we propose to use another two features that address this problem. Middle-level feature, which represents the music's melodiousness, rhythmic stability, and tonal stability, and is able to capture the perceptual features of music. The timbre feature, which is used in speaker identification, represents the singers' voice features. Furthermore, we propose a convolutional recurrent neural network (CRNN) to combine three features for singer identification. The model firstly fuses the frame-level feature and timbre feature and then combines middle-level features to the mix features. In experiments, the proposed method achieves comparable performance on an average F1 score of 0.81 on the benchmark dataset of Artist20, which significantly improves related works.


【10】 Deep Learning-based automated classification of Chinese Speech Sound Disorders

标题:基于深度学习的汉语语音障碍自动分类

链接:https://arxiv.org/abs/2205.11748

作者:Yao-Ming Kuo,Shanq-Jang Ruan,Yu-Chin Chen,Ya-Wen Tu
备注:12 pages, 9 figures, journal
摘要:本文描述了一个分析声学数据的系统,以帮助使用计算机诊断和分类儿童言语障碍。分析集中于识别和分类四种不同类型的汉语误解。该研究收集并生成了一个语音语料库,其中包含2540个停音、软腭音、辅音元音和塞擦音样本,这些样本来自90名3-6岁具有正常或病理性发音特征的儿童。每次录音都附有言语治疗领域的详细注释。语音样本的分类是使用三个建立良好的神经网络模型进行图像分类的。特征映射是使用从语音中提取的三组MFCC参数创建的,并聚合成三维数据结构作为模型输入。我们采用六种数据扩充技术来扩充可用数据集,同时避免过度模拟。实验考察了四种不同类别的汉语短语和汉字的可用性。对不同数据子集的实验表明,该系统能够准确地检测出所分析的发音障碍。
摘要:This article describes a system for analyzing acoustic data in order to assist in the diagnosis and classification of children's speech disorders using a computer. The analysis concentrated on identifying and categorizing four distinct types of Chinese misconstructions. The study collected and generated a speech corpus containing 2540 Stopping, Velar, Consonant-vowel, and Affricate samples from 90 children aged 3-6 years with normal or pathological articulatory features. Each recording was accompanied by a detailed annotation from the field of speech therapy. Classification of the speech samples was accomplished using three well-established neural network models for image classification. The feature maps are created using three sets of MFCC parameters extracted from speech sounds and aggregated into a three-dimensional data structure as model input. We employ six techniques for data augmentation in order to augment the available dataset while avoiding over-simulation. The experiments examine the usability of four different categories of Chinese phrases and characters. Experiments with different data subsets demonstrate the system's ability to accurately detect the analyzed pronunciation disorders.


【11】 Adaptive Few-Shot Learning Algorithm for Rare Sound Event Detection

标题:稀有声音事件检测的自适应Few-Shot学习算法

链接:https://arxiv.org/abs/2205.11738

作者:Chendong Zhao,Jianzong Wang,Leilai Li,Xiaoyang Qu,Jing Xiao
备注:Accepted to IJCNN 2022. arXiv admin note: text overlap with arXiv:2110.04474 by other authors摘要:声音事件检测是通过了解周围环境的声音来推断事件。由于罕见声音事件的稀缺性,对于已经掌握了太多先验知识的训练有素的探测器来说,这是一个挑战。同时,很少有镜头学习方法在面对新的有限数据任务时具有良好的泛化能力。最近的方法在这一领域取得了可喜的成果。然而,这些方法独立地处理每个支持示例,忽略了整个任务中其他示例的信息。因此,以前的大多数方法都被限制为为所有测试时任务生成相同的特征嵌入,这对每个输入数据都是不自适应的。在这项工作中,我们提出了一种新的任务自适应模块,该模块易于植入任何基于度量的Few-Shot学习框架中。该模块可以识别与任务相关的特征维度。与基线方法相比,结合我们的模块可以显著提高两个数据集的性能,特别是对于直传传播网络。例如,ESC-50上的5向单发精度为+6.8%,而noiseESC-50上的精度为+5.9%。我们在域不匹配设置中研究了我们的方法,并取得了比以前的方法更好的结果。
摘要:Sound event detection is to infer the event by understanding the surrounding environmental sounds. Due to the scarcity of rare sound events, it becomes challenging for the well-trained detectors which have learned too much prior knowledge. Meanwhile, few-shot learning methods promise a good generalization ability when facing a new limited-data task. Recent approaches have achieved promising results in this field. However, these approaches treat each support example independently, ignoring the information of other examples from the whole task. Because of this, most of previous methods are constrained to generate a same feature embedding for all test-time tasks, which is not adaptive to each inputted data. In this work, we propose a novel task-adaptive module which is easy to plant into any metric-based few-shot learning frameworks. The module could identify the task-relevant feature dimension. Incorporating our module improves the performance considerably on two datasets over baseline methods, especially for the transductive propagation network. Such as +6.8% for 5-way 1-shot accuracy on ESC-50, and +5.9% on noiseESC-50. We investigate our approach in the domain-mismatch setting and also achieve better results than previous methods.


机器翻译,仅供参考