今天跟大家分享一篇语音相关的论文合集:cs.SD语音7篇,eess.AS音频处理8篇。本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
【1】 Mel Spectrogram Inversion with Stable Pitch
标题:具有稳定基音的Mel谱图反演
链接:https://arxiv.org/abs/2208.12782
作者:Bruno Di Giorgi,Mark Levy,Richard Sharp备注:7 pages, 5 figures, Proceedings of the 23st International Society for Music Information Retrieval Conference, ISMIR 2022摘要:声码器是能够将音频信号的低维频谱表示(通常是mel频谱图)变换成波形的模型。现代语音生成流水线使用声码器作为其最终组件。近来为语音开发的声码器模型实现了高度的真实性,使得很自然地想知道它们将如何对音乐信号执行。与语音相比,音乐声音纹理的异质性和结构性提出了新的挑战。在这项工作中,我们关注一个特定的伪像,一些为语音设计的声码器模型在应用于音乐时往往表现出:在合成延音音符时感觉到的音高不稳定性。我们认为,这种伪像的特征声音是由于缺乏水平相位相干性,这通常是使用具有对时移不变的模型(如卷积神经网络)的时域目标空间的结果。我们提出了一个新的声码器模型,专门为音乐设计。提高俯仰稳定性的关键是选择由幅度谱和相位梯度组成的平移不变目标空间。我们讨论了启发我们重新制定声码器任务的原因,概述了一个工作示例,并对音乐信号进行了评估。我们的方法使用一种新的谐波误差度量,相对于现有模型,持续音符和和弦的重建分别提高了60%和10%。摘要:Vocoders are models capable of transforming a low-dimensional spectral representation of an audio signal, typically the mel spectrogram, to a waveform. Modern speech generation pipelines use a vocoder as their final component. Recent vocoder models developed for speech achieve a high degree of realism, such that it is natural to wonder how they would perform on music signals. Compared to speech, the heterogeneity and structure of the musical sound texture offers new challenges. In this work we focus on one specific artifact that some vocoder models designed for speech tend to exhibit when applied to music: the perceived instability of pitch when synthesizing sustained notes. We argue that the characteristic sound of this artifact is due to the lack of horizontal phase coherence, which is often the result of using a time-domain target space with a model that is invariant to time-shifts, such as a convolutional neural network. We propose a new vocoder model that is specifically designed for music. Key to improving the pitch stability is the choice of a shift-invariant target space that consists of the magnitude spectrum and the phase gradient. We discuss the reasons that inspired us to re-formulate the vocoder task, outline a working example, and evaluate it on musical signals. Our method results in 60% and 10% improved reconstruction of sustained notes and chords with respect to existing models, using a novel harmonic error metric.
【2】 Spatio-Temporal Representation Learning Enhanced Source Cell-phone Recognition from Speech Recordings
标题:时空表征学习增强的语音源手机识别
链接:https://arxiv.org/abs/2208.12753
作者:Chunyan Zeng,Shixiong Feng,Zhifeng Wang,Xiangkui Wan,Yunfan Chen,Nan Zhao摘要:现有的源手机识别方法缺乏对源设备的长期特征表征,导致源手机相关特征表达不准确,导致识别准确率不足。本文提出了一种基于时空表示学习的源手机识别方法,主要包括两个部分:序列高斯均值矩阵特征的提取和基于时空表示学习的识别模型的构建。在特征提取部分,在分析记录源信号的时间序列表示的基础上,利用高斯混合模型对数据分布的敏感性,提取出具有长、短期表示能力的序列高斯均值矩阵。在模型构建部分,设计了结构化的时空表征学习网络C3D-BiLSTM,充分表征时空信息,结合三维卷积网络和双向长短时记忆网络进行短时谱信息和长时波动信息的表征学习,融合录音源信号的时空特征信息,实现手机的精确识别。在CCNU-Mobile数据集上,对45部手机的闭集识别平均正确率为99.03%,在小样本实验中平均正确率为98.18%,识别性能优于现有的方法.实验结果表明,该方法在多类别手机识别中具有良好的识别性能。摘要:The existing source cell-phone recognition method lacks the long-term feature characterization of the source device, resulting in inaccurate representation of the source cell-phone related features which leads to insufficient recognition accuracy. In this paper, we propose a source cell-phone recognition method based on spatio-temporal representation learning, which includes two main parts: extraction of sequential Gaussian mean matrix features and construction of a recognition model based on spatio-temporal representation learning. In the feature extraction part, based on the analysis of time-series representation of recording source signals, we extract sequential Gaussian mean matrix with long-term and short-term representation ability by using the sensitivity of Gaussian mixture model to data distribution. In the model construction part, we design a structured spatio-temporal representation learning network C3D-BiLSTM to fully characterize the spatio-temporal information, combine 3D convolutional network and bidirectional long short-term memory network for short-term spectral information and long-time fluctuation information representation learning, and achieve accurate recognition of cell-phones by fusing spatio-temporal feature information of recording source signals. The method achieves an average accuracy of 99.03% for the closed-set recognition of 45 cell-phones under the CCNU\_Mobile dataset, and 98.18% in small sample size experiments, with recognition performance better than the existing state-of-the-art methods. The experimental results show that the method exhibits excellent recognition performance in multi-class cell-phones recognition.
【3】 Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages
标题:从公共数据中挖掘音频和文本对改进低资源语言ASR系统的有效性
链接:https://arxiv.org/abs/2208.12666
作者:Kaushal Santosh Bhogale,Abhigyan Raman,Tahir Javed,Sumanth Doddapaneni,Anoop Kunchukuttan,Pratyush Kumar,Mitesh M. Khapra摘要:端到端(E2E)模型已经成为当前最先进的语音识别系统的默认选择。这些模型是在大量的标记数据上训练的,而这些数据对于资源较少的语言来说往往是不可用的。自监督学习和迁移学习等技术有希望,但在训练精确模型方面尚未有效。另一方面,收集不同领域和说话人的标记数据集是非常昂贵的。在这项工作中,我们通过从公共资源,特别是从全印度广播电台的公共档案中'挖掘'印度语言的文本和音频对,展示了一种廉价而有效的替代方法。作为一个关键组件,我们采用Needleman-Wunsch算法将句子与相应的音频片段对齐,给定一个长音频及其转录的PDF,同时对OCR、无关文本和非转录语音造成的错误具有鲁棒性。因此,我们创建了Shrutilipi,这是一个数据集,包含超过6,400小时的标记音频,跨越12种印度语言,总计为4.95M句子。平均而言,Shrutilipi的结果比公开可用的标记数据增加2.3倍。我们通过12种语言的21名人类评估者确定了Shrutilipi的质量。我们还建立了Shrutilipi的多样性方面的代表地区,发言人,并提到命名的实体。值得注意的是,我们表明,将Shrutilipi添加到Wav2Vec模型的训练集中,在IndicSUPERB基准测试中,7种语言的WER平均下降了5.8%。对于基准最多的印地语(7个),平均WER从18.8%下降到13.5%。这种改进扩展到高效模型:我们显示了Conformer模型的WER下降2.3%(比Wav2Vec小10倍)。最后,我们证明了Shrutilipi的多样性,表明用它训练的模型对噪声输入具有更强的鲁棒性。摘要:End-to-end (E2E) models have become the default choice for state-of-the-art speech recognition systems. Such models are trained on large amounts of labelled data, which are often not available for low-resource languages. Techniques such as self-supervised learning and transfer learning hold promise, but have not yet been effective in training accurate models. On the other hand, collecting labelled datasets on a diverse set of domains and speakers is very expensive. In this work, we demonstrate an inexpensive and effective alternative to these approaches by ``mining'' text and audio pairs for Indian languages from public sources, specifically from the public archives of All India Radio. As a key component, we adapt the Needleman-Wunsch algorithm to align sentences with corresponding audio segments given a long audio and a PDF of its transcript, while being robust to errors due to OCR, extraneous text, and non-transcribed speech. We thus create Shrutilipi, a dataset which contains over 6,400 hours of labelled audio across 12 Indian languages totalling to 4.95M sentences. On average, Shrutilipi results in a 2.3x increase over publicly available labelled data. We establish the quality of Shrutilipi with 21 human evaluators across the 12 languages. We also establish the diversity of Shrutilipi in terms of represented regions, speakers, and mentioned named entities. Significantly, we show that adding Shrutilipi to the training set of Wav2Vec models leads to an average decrease in WER of 5.8\% for 7 languages on the IndicSUPERB benchmark. For Hindi, which has the most benchmarks (7), the average WER falls from 18.8% to 13.5%. This improvement extends to efficient models: We show a 2.3% drop in WER for a Conformer model (10x smaller than Wav2Vec). Finally, we demonstrate the diversity of Shrutilipi by showing that the model trained with it is more robust to noisy input.
【4】 Concept-Based Techniques for "Musicologist-friendly" Explanations in a Deep Music Classifier
链接:https://arxiv.org/abs/2208.12485
作者:Francesco Foscarin,Katharina Hoedt,Verena Praher,Arthur Flexer,Gerhard Widmer备注:(to be) published in Proceedings of the 23rd ISMIR 2022摘要:用于解释应用于音乐数据的深度学习系统的当前方法提供了低级特征空间中的结果,通过突出显示频谱图中潜在相关的时间—频率仓或钢琴卷中的时间—音高仓,可以将该时间—音高仓与钢琴卷中的时间—音高仓相关联。这可能很难理解,特别是对于没有技术知识的音乐学家。为了解决这一问题,我们关注基于高层次音乐概念的更人性化的解释。我们的研究以经过训练的系统(事后解释)为目标,探索了两种方法:受监督的一个,其中用户可以定义音乐概念并测试它是否与系统相关;和无人监督的一种,在无人监督的一种中,自动选择包含相关概念的音乐摘录并将其提供给用户进行解释。我们在现有的符号作曲家分类系统上演示了这两种技术,展示了它们的潜力,并强调了它们的内在局限性。摘要:Current approaches for explaining deep learning systems applied to musical data provide results in a low-level feature space, e.g., by highlighting potentially relevant time-frequency bins in a spectrogram or time-pitch bins in a piano roll. This can be difficult to understand, particularly for musicologists without technical knowledge. To address this issue, we focus on more human-friendly explanations based on high-level musical concepts. Our research targets trained systems (post-hoc explanations) and explores two approaches: a supervised one, where the user can define a musical concept and test if it is relevant to the system; and an unsupervised one, where musical excerpts containing relevant concepts are automatically selected and given to the user for interpretation. We demonstrate both techniques on an existing symbolic composer classification system, showcase their potential, and highlight their intrinsic limitations.
【5】 Leveraging Symmetrical Convolutional Transformer Networks for Speech to Singing Voice Style Transfer
标题:利用对称卷积变换网络实现语音到歌唱的语调转换
链接:https://arxiv.org/abs/2208.12410
作者:Shrutina Agarwal,Sriram Ganapathy,Naoya Takahashi备注:accepted to INTERSPEECH 2022摘要:本文提出了一种实现语音风格向演唱风格转换的模型。与以往的基于信号处理的方法不同,这些方法需要高质量的歌唱模板或音素同步,我们探索了一种数据驱动的方法来解决将自然语音转换为歌唱声音的问题。我们开发了一种新的神经网络结构,称为SymNet,它在保持说话人身份和自然度的同时,对输入语音与目标旋律的对齐进行建模。提出的SymNet模型由三种类型的层—卷积层、Transformer层和自注意层的对称堆栈组成。本文还探索了新的数据扩充和生成式损失退火方法来促进模型的训练。实验是在 NUS和NHSS数据集,由语音和歌唱声的平行数据组成。在这些实验中,我们表明所提出的SymNet模型显著地改善了目标重建质量,优于先前发表的方法和基线架构。此外,主观收听测试证实了使用所提出的方法获得的音频的改善的质量(在平均意见得分测量上相对于基线系统的绝对改善为0.37)。摘要:In this paper, we propose a model to perform style transfer of speech to singing voice. Contrary to the previous signal processing-based methods, which require high-quality singing templates or phoneme synchronization, we explore a data-driven approach for the problem of converting natural speech to singing voice. We develop a novel neural network architecture, called SymNet, which models the alignment of the input speech with the target melody while preserving the speaker identity and naturalness. The proposed SymNet model is comprised of symmetrical stack of three types of layers - convolutional, transformer, and self-attention layers. The paper also explores novel data augmentation and generative loss annealing methods to facilitate the model training. Experiments are performed on the NUS and NHSS datasets which consist of parallel data of speech and singing voice. In these experiments, we show that the proposed SymNet model improves the objective reconstruction quality significantly over the previously published methods and baseline architectures. Further, a subjective listening test confirms the improved quality of the audio obtained using the proposed approach (absolute improvement of 0.37 in mean opinion score measure over the baseline system).
【6】 Music Separation Enhancement with Generative Modeling
标题:基于产生式建模的音乐分离增强
链接:https://arxiv.org/abs/2208.12387
作者:Noah Schaffer,Boaz Cogan,Ethan Manilow,Max Morrison,Prem Seetharaman,Bryan Pardo备注:Accepted to ISMIR 2022摘要:尽管近年来取得了显著的进步,但现有技术的音乐分离系统产生具有显著感知缺陷的源估计,诸如添加外来噪声或去除谐波。我们提出了一种后处理模型(使其听起来很好(MSG)后处理器)来增强音乐源分离系统的输出。我们将我们的后处理模型应用于最先进的基于波形和基于谱图的音乐源分离器,包括在训练期间MSG看不到的分离器。我们对源分离器产生的误差的分析表明,波形模型往往会引入更多的高频噪声,而谱图模型往往会丢失瞬变和高频内容。我们引入了客观的度量来量化这两种误差,并表明MSG改善了这两种误差的源重构。众包主观评价表明,人类听众更喜欢已经由MSG后处理的低音和鼓的源估计。摘要:Despite phenomenal progress in recent years, state-of-the-art music separation systems produce source estimates with significant perceptual shortcomings, such as adding extraneous noise or removing harmonics. We propose a post-processing model (the Make it Sound Good (MSG) post-processor) to enhance the output of music source separation systems. We apply our post-processing model to state-of-the-art waveform-based and spectrogram-based music source separators, including a separator unseen by MSG during training. Our analysis of the errors produced by source separators shows that waveform models tend to introduce more high-frequency noise, while spectrogram models tend to lose transients and high frequency content. We introduce objective measures to quantify both kinds of errors and show MSG improves the source reconstruction of both kinds of errors. Crowdsourced subjective evaluations demonstrate that human listeners prefer source estimates of bass and drums that have been post-processed by MSG.
【7】 MuLan: A Joint Embedding of Music Audio and Natural Language
标题:花木兰:音乐、音频和自然语言的联合嵌入
链接:https://arxiv.org/abs/2208.12415
作者:Qingqing Huang,Aren Jansen,Joonseok Lee,Ravi Ganti,Judith Yue Li,Daniel P. W. Ellis备注:To appear in ISMIR 2022摘要:传统上,音乐标签和基于内容的检索系统是使用预定义的本体来构建的,该本体覆盖了一组严格的音乐属性或文本查询。慕兰:第一次尝试将音乐音频直接链接到无约束的自然语言音乐描述的新一代声学模型。慕兰采用了一个双塔的形式,联合音频文本嵌入模型,使用4400万个音乐录音(37万小时)和弱关联的自由形式的文本注释进行训练。通过其与广泛的音乐流派和文本风格(包括传统的音乐标签)的兼容性,所得到的音频—文本表示包含现有的本体,同时逐渐过渡到真正的zero-shot功能。通过迁移学习、zero-shot音乐标注、音乐领域的语言理解和跨模态检索应用等一系列实验,证明了慕兰嵌入的通用性。摘要:Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This paper presents MuLan: a first attempt at a new generation of acoustic models that link music audio directly to unconstrained natural language music descriptions. MuLan takes the form of a two-tower, joint audio-text embedding model trained using 44 million music recordings (370K hours) and weakly-associated, free-form text annotations. Through its compatibility with a wide range of music genres and text styles (including conventional music tags), the resulting audio-text representation subsumes existing ontologies while graduating to true zero-shot functionalities. We demonstrate the versatility of the MuLan embeddings with a range of experiments including transfer learning, zero-shot music tagging, language understanding in the music domain, and cross-modal retrieval applications.
【1】 MuLan: A Joint Embedding of Music Audio and Natural Language
标题:花木兰:音乐、音频和自然语言的联合嵌入
链接:https://arxiv.org/abs/2208.12415
* 与cs.SD语音【7】为同一篇
作者:Qingqing Huang,Aren Jansen,Joonseok Lee,Ravi Ganti,Judith Yue Li,Daniel P. W. Ellis备注:To appear in ISMIR 2022摘要:传统上,音乐标签和基于内容的检索系统是使用预定义的本体来构建的,该本体覆盖了一组严格的音乐属性或文本查询。慕兰:第一次尝试将音乐音频直接链接到无约束的自然语言音乐描述的新一代声学模型。慕兰采用了一个双塔的形式,联合音频文本嵌入模型,使用4400万个音乐录音(37万小时)和弱关联的自由形式的文本注释进行训练。通过其与广泛的音乐流派和文本风格(包括传统的音乐标签)的兼容性,所得到的音频—文本表示包含现有的本体,同时逐渐过渡到真正的zero-shot功能。通过迁移学习、zero-shot音乐标注、音乐领域的语言理解和跨模态检索应用等一系列实验,证明了慕兰嵌入的通用性。摘要:Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This paper presents MuLan: a first attempt at a new generation of acoustic models that link music audio directly to unconstrained natural language music descriptions. MuLan takes the form of a two-tower, joint audio-text embedding model trained using 44 million music recordings (370K hours) and weakly-associated, free-form text annotations. Through its compatibility with a wide range of music genres and text styles (including conventional music tags), the resulting audio-text representation subsumes existing ontologies while graduating to true zero-shot functionalities. We demonstrate the versatility of the MuLan embeddings with a range of experiments including transfer learning, zero-shot music tagging, language understanding in the music domain, and cross-modal retrieval applications.
【2】 Decoding speech from non-invasive brain recordings
标题:从非侵入性脑录音中解码语音
链接:https://arxiv.org/abs/2208.12266
作者:Alexandre Défossez,Charlotte Caucheteux,Jérémy Rapin,Ori Kabeli,Jean-Rémi King备注:15 pages, preprint摘要:从大脑活动中解码语言是医疗保健和神经科学期待已久的目标。由于颅内器械,最近实现了一些重大里程碑:在对基本语言任务的侵入性大脑响应上训练的对象特定的流水线现在开始有效地解码可解释的特征(例如字母、单词、声谱图)。然而,将这种方法扩展到自然语音和非侵入性脑记录仍然是一个重大挑战。在这里,我们提出了一个单一的端到端的架构,该架构在一个大的个体队列中用对比学习来训练,以预测自然语音的自监督表示。我们在四个公共数据集上评估了我们的模型,包括169名志愿者在听自然语音时用脑磁图或脑电图(M/EEG)记录的数据。结果表明,我们的模型可以从3s的MEG信号中识别相应的语音片段,在1,594个不同片段中具有高达72.5%的top-10准确度(和44%的top-1准确度),并且在EEG记录的2,604个片段中具有高达19.1%的top-10准确度,因此允许解码训练集中不存在的短语。模型比较和消融分析表明,这些性能直接受益于我们最初的设计选择,即使用(i)对比目标,(ii)语音的预训练表示和(iii)在几个参与者之间同时训练的公共卷积结构。总之,这些结果描绘了一条有希望的道路,以解码自然语言处理的实时从非侵入性记录的大脑活动。摘要:Decoding language from brain activity is a long-awaited goal in both healthcare and neuroscience. Major milestones have recently been reached thanks to intracranial devices: subject-specific pipelines trained on invasive brain responses to basic language tasks now start to efficiently decode interpretable features (e.g. letters, words, spectrograms). However, scaling this approach to natural speech and non-invasive brain recordings remains a major challenge. Here, we propose a single end-to-end architecture trained with contrastive learning across a large cohort of individuals to predict self-supervised representations of natural speech. We evaluate our model on four public datasets, encompassing 169 volunteers recorded with magneto- or electro-encephalography (M/EEG), while they listened to natural speech. The results show that our model can identify, from 3s of MEG signals, the corresponding speech segment with up to 72.5% top-10 accuracy out of 1,594 distinct segments (and 44% top-1 accuracy), and up to 19.1% out of 2,604 segments for EEG recordings -- hence allowing the decoding of phrases absent from the training set. Model comparison and ablation analyses show that these performances directly benefit from our original design choices, namely the use of (i) a contrastive objective, (ii) pretrained representations of speech and (iii) a common convolutional architecture simultaneously trained across several participants. Together, these results delineate a promising path to decode natural language processing in real time from non-invasive recordings of brain activity.
【3】 Mel Spectrogram Inversion with Stable Pitch
标题:具有稳定基音的Mel谱图反演
链接:https://arxiv.org/abs/2208.12782
* 与cs.SD语音【1】为同一篇
作者:Bruno Di Giorgi,Mark Levy,Richard Sharp备注:7 pages, 5 figures, Proceedings of the 23st International Society for Music Information Retrieval Conference, ISMIR 2022摘要:声码器是能够将音频信号的低维频谱表示(通常是mel频谱图)变换成波形的模型。现代语音生成流水线使用声码器作为其最终组件。近来为语音开发的声码器模型实现了高度的真实性,使得很自然地想知道它们将如何对音乐信号执行。与语音相比,音乐声音纹理的异质性和结构性提出了新的挑战。在这项工作中,我们关注一个特定的伪像,一些为语音设计的声码器模型在应用于音乐时往往表现出:在合成延音音符时感觉到的音高不稳定性。我们认为,这种伪像的特征声音是由于缺乏水平相位相干性,这通常是使用具有对时移不变的模型(如卷积神经网络)的时域目标空间的结果。我们提出了一个新的声码器模型,专门为音乐设计。提高俯仰稳定性的关键是选择由幅度谱和相位梯度组成的平移不变目标空间。我们讨论了启发我们重新制定声码器任务的原因,概述了一个工作示例,并对音乐信号进行了评估。我们的方法使用一种新的谐波误差度量,相对于现有模型,持续音符和和弦的重建分别提高了60%和10%。摘要:Vocoders are models capable of transforming a low-dimensional spectral representation of an audio signal, typically the mel spectrogram, to a waveform. Modern speech generation pipelines use a vocoder as their final component. Recent vocoder models developed for speech achieve a high degree of realism, such that it is natural to wonder how they would perform on music signals. Compared to speech, the heterogeneity and structure of the musical sound texture offers new challenges. In this work we focus on one specific artifact that some vocoder models designed for speech tend to exhibit when applied to music: the perceived instability of pitch when synthesizing sustained notes. We argue that the characteristic sound of this artifact is due to the lack of horizontal phase coherence, which is often the result of using a time-domain target space with a model that is invariant to time-shifts, such as a convolutional neural network. We propose a new vocoder model that is specifically designed for music. Key to improving the pitch stability is the choice of a shift-invariant target space that consists of the magnitude spectrum and the phase gradient. We discuss the reasons that inspired us to re-formulate the vocoder task, outline a working example, and evaluate it on musical signals. Our method results in 60% and 10% improved reconstruction of sustained notes and chords with respect to existing models, using a novel harmonic error metric.
【4】 Spatio-Temporal Representation Learning Enhanced Source Cell-phone Recognition from Speech Recordings
标题:时空表征学习增强的语音源手机识别
链接:https://arxiv.org/abs/2208.12753
* 与cs.SD语音【2】为同一篇
作者:Chunyan Zeng,Shixiong Feng,Zhifeng Wang,Xiangkui Wan,Yunfan Chen,Nan Zhao摘要:现有的源手机识别方法缺乏对源设备的长期特征表征,导致源手机相关特征表达不准确,导致识别准确率不足。本文提出了一种基于时空表示学习的源手机识别方法,主要包括两个部分:序列高斯均值矩阵特征的提取和基于时空表示学习的识别模型的构建。在特征提取部分,在分析记录源信号的时间序列表示的基础上,利用高斯混合模型对数据分布的敏感性,提取出具有长、短期表示能力的序列高斯均值矩阵。在模型构建部分,设计了结构化的时空表征学习网络C3D-BiLSTM,充分表征时空信息,结合三维卷积网络和双向长短时记忆网络进行短时谱信息和长时波动信息的表征学习,融合录音源信号的时空特征信息,实现手机的精确识别。在CCNU-Mobile数据集上,对45部手机的闭集识别平均正确率为99.03%,在小样本实验中平均正确率为98.18%,识别性能优于现有的方法.实验结果表明,该方法在多类别手机识别中具有良好的识别性能。摘要:The existing source cell-phone recognition method lacks the long-term feature characterization of the source device, resulting in inaccurate representation of the source cell-phone related features which leads to insufficient recognition accuracy. In this paper, we propose a source cell-phone recognition method based on spatio-temporal representation learning, which includes two main parts: extraction of sequential Gaussian mean matrix features and construction of a recognition model based on spatio-temporal representation learning. In the feature extraction part, based on the analysis of time-series representation of recording source signals, we extract sequential Gaussian mean matrix with long-term and short-term representation ability by using the sensitivity of Gaussian mixture model to data distribution. In the model construction part, we design a structured spatio-temporal representation learning network C3D-BiLSTM to fully characterize the spatio-temporal information, combine 3D convolutional network and bidirectional long short-term memory network for short-term spectral information and long-time fluctuation information representation learning, and achieve accurate recognition of cell-phones by fusing spatio-temporal feature information of recording source signals. The method achieves an average accuracy of 99.03% for the closed-set recognition of 45 cell-phones under the CCNU\_Mobile dataset, and 98.18% in small sample size experiments, with recognition performance better than the existing state-of-the-art methods. The experimental results show that the method exhibits excellent recognition performance in multi-class cell-phones recognition.
【5】 Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages
标题:从公共数据中挖掘音频和文本对改进低资源语言ASR系统的有效性
链接:https://arxiv.org/abs/2208.12666
* 与cs.SD语音【3】为同一篇
作者:Kaushal Santosh Bhogale,Abhigyan Raman,Tahir Javed,Sumanth Doddapaneni,Anoop Kunchukuttan,Pratyush Kumar,Mitesh M. Khapra摘要:端到端(E2E)模型已经成为当前最先进的语音识别系统的默认选择。这些模型是在大量的标记数据上训练的,而这些数据对于资源较少的语言来说往往是不可用的。自监督学习和迁移学习等技术有希望,但在训练精确模型方面尚未有效。另一方面,收集不同领域和说话人的标记数据集是非常昂贵的。在这项工作中,我们通过从公共资源,特别是从全印度广播电台的公共档案中'挖掘'印度语言的文本和音频对,展示了一种廉价而有效的替代方法。作为一个关键组件,我们采用Needleman-Wunsch算法将句子与相应的音频片段对齐,给定一个长音频及其转录的PDF,同时对OCR、无关文本和非转录语音造成的错误具有鲁棒性。因此,我们创建了Shrutilipi,这是一个数据集,包含超过6,400小时的标记音频,跨越12种印度语言,总计为4.95M句子。平均而言,Shrutilipi的结果比公开可用的标记数据增加2.3倍。我们通过12种语言的21名人类评估者确定了Shrutilipi的质量。我们还建立了Shrutilipi的多样性方面的代表地区,发言人,并提到命名的实体。值得注意的是,我们表明,将Shrutilipi添加到Wav2Vec模型的训练集中,在IndicSUPERB基准测试中,7种语言的WER平均下降了5.8%。对于基准最多的印地语(7个),平均WER从18.8%下降到13.5%。这种改进扩展到高效模型:我们显示了Conformer模型的WER下降2.3%(比Wav2Vec小10倍)。最后,我们证明了Shrutilipi的多样性,表明用它训练的模型对噪声输入具有更强的鲁棒性。摘要:End-to-end (E2E) models have become the default choice for state-of-the-art speech recognition systems. Such models are trained on large amounts of labelled data, which are often not available for low-resource languages. Techniques such as self-supervised learning and transfer learning hold promise, but have not yet been effective in training accurate models. On the other hand, collecting labelled datasets on a diverse set of domains and speakers is very expensive. In this work, we demonstrate an inexpensive and effective alternative to these approaches by ``mining'' text and audio pairs for Indian languages from public sources, specifically from the public archives of All India Radio. As a key component, we adapt the Needleman-Wunsch algorithm to align sentences with corresponding audio segments given a long audio and a PDF of its transcript, while being robust to errors due to OCR, extraneous text, and non-transcribed speech. We thus create Shrutilipi, a dataset which contains over 6,400 hours of labelled audio across 12 Indian languages totalling to 4.95M sentences. On average, Shrutilipi results in a 2.3x increase over publicly available labelled data. We establish the quality of Shrutilipi with 21 human evaluators across the 12 languages. We also establish the diversity of Shrutilipi in terms of represented regions, speakers, and mentioned named entities. Significantly, we show that adding Shrutilipi to the training set of Wav2Vec models leads to an average decrease in WER of 5.8\% for 7 languages on the IndicSUPERB benchmark. For Hindi, which has the most benchmarks (7), the average WER falls from 18.8% to 13.5%. This improvement extends to efficient models: We show a 2.3% drop in WER for a Conformer model (10x smaller than Wav2Vec). Finally, we demonstrate the diversity of Shrutilipi by showing that the model trained with it is more robust to noisy input.
【6】 Concept-Based Techniques for "Musicologist-friendly" Explanations in a Deep Music Classifier
链接:https://arxiv.org/abs/2208.12485
* 与cs.SD语音【4】为同一篇
作者:Francesco Foscarin,Katharina Hoedt,Verena Praher,Arthur Flexer,Gerhard Widmer备注:(to be) published in Proceedings of the 23rd ISMIR 2022摘要:用于解释应用于音乐数据的深度学习系统的当前方法提供了低级特征空间中的结果,通过突出显示频谱图中潜在相关的时间—频率仓或钢琴卷中的时间—音高仓,可以将该时间—音高仓与钢琴卷中的时间—音高仓相关联。这可能很难理解,特别是对于没有技术知识的音乐学家。为了解决这一问题,我们关注基于高层次音乐概念的更人性化的解释。我们的研究以经过训练的系统(事后解释)为目标,探索了两种方法:受监督的一个,其中用户可以定义音乐概念并测试它是否与系统相关;和无人监督的一种,在无人监督的一种中,自动选择包含相关概念的音乐摘录并将其提供给用户进行解释。我们在现有的符号作曲家分类系统上演示了这两种技术,展示了它们的潜力,并强调了它们的内在局限性。摘要:Current approaches for explaining deep learning systems applied to musical data provide results in a low-level feature space, e.g., by highlighting potentially relevant time-frequency bins in a spectrogram or time-pitch bins in a piano roll. This can be difficult to understand, particularly for musicologists without technical knowledge. To address this issue, we focus on more human-friendly explanations based on high-level musical concepts. Our research targets trained systems (post-hoc explanations) and explores two approaches: a supervised one, where the user can define a musical concept and test if it is relevant to the system; and an unsupervised one, where musical excerpts containing relevant concepts are automatically selected and given to the user for interpretation. We demonstrate both techniques on an existing symbolic composer classification system, showcase their potential, and highlight their intrinsic limitations.
【7】 Leveraging Symmetrical Convolutional Transformer Networks for Speech to Singing Voice Style Transfer
标题:利用对称卷积变换网络实现语音到歌唱的语调转换
链接:https://arxiv.org/abs/2208.12410
* 与cs.SD语音【5】为同一篇
作者:Shrutina Agarwal,Sriram Ganapathy,Naoya Takahashi备注:accepted to INTERSPEECH 2022摘要:本文提出了一种实现语音风格向演唱风格转换的模型。与以往的基于信号处理的方法不同,这些方法需要高质量的歌唱模板或音素同步,我们探索了一种数据驱动的方法来解决将自然语音转换为歌唱声音的问题。我们开发了一种新的神经网络结构,称为SymNet,它在保持说话人身份和自然度的同时,对输入语音与目标旋律的对齐进行建模。提出的SymNet模型由三种类型的层—卷积层、Transformer层和自注意层的对称堆栈组成。本文还探索了新的数据扩充和生成式损失退火方法来促进模型的训练。实验是在 NUS和NHSS数据集,由语音和歌唱声的平行数据组成。在这些实验中,我们表明所提出的SymNet模型显著地改善了目标重建质量,优于先前发表的方法和基线架构。此外,主观收听测试证实了使用所提出的方法获得的音频的改善的质量(在平均意见得分测量上相对于基线系统的绝对改善为0.37)。摘要:In this paper, we propose a model to perform style transfer of speech to singing voice. Contrary to the previous signal processing-based methods, which require high-quality singing templates or phoneme synchronization, we explore a data-driven approach for the problem of converting natural speech to singing voice. We develop a novel neural network architecture, called SymNet, which models the alignment of the input speech with the target melody while preserving the speaker identity and naturalness. The proposed SymNet model is comprised of symmetrical stack of three types of layers - convolutional, transformer, and self-attention layers. The paper also explores novel data augmentation and generative loss annealing methods to facilitate the model training. Experiments are performed on the NUS and NHSS datasets which consist of parallel data of speech and singing voice. In these experiments, we show that the proposed SymNet model improves the objective reconstruction quality significantly over the previously published methods and baseline architectures. Further, a subjective listening test confirms the improved quality of the audio obtained using the proposed approach (absolute improvement of 0.37 in mean opinion score measure over the baseline system).
【8】 Music Separation Enhancement with Generative Modeling
标题:基于产生式建模的音乐分离增强
链接:https://arxiv.org/abs/2208.12387
* 与cs.SD语音【6】为同一篇
作者:Noah Schaffer,Boaz Cogan,Ethan Manilow,Max Morrison,Prem Seetharaman,Bryan Pardo备注:Accepted to ISMIR 2022摘要:尽管近年来取得了显著的进步,但现有技术的音乐分离系统产生具有显著感知缺陷的源估计,诸如添加外来噪声或去除谐波。我们提出了一种后处理模型(使其听起来很好(MSG)后处理器)来增强音乐源分离系统的输出。我们将我们的后处理模型应用于最先进的基于波形和基于谱图的音乐源分离器,包括在训练期间MSG看不到的分离器。我们对源分离器产生的误差的分析表明,波形模型往往会引入更多的高频噪声,而谱图模型往往会丢失瞬变和高频内容。我们引入了客观的度量来量化这两种误差,并表明MSG改善了这两种误差的源重构。众包主观评价表明,人类听众更喜欢已经由MSG后处理的低音和鼓的源估计。摘要:Despite phenomenal progress in recent years, state-of-the-art music separation systems produce source estimates with significant perceptual shortcomings, such as adding extraneous noise or removing harmonics. We propose a post-processing model (the Make it Sound Good (MSG) post-processor) to enhance the output of music source separation systems. We apply our post-processing model to state-of-the-art waveform-based and spectrogram-based music source separators, including a separator unseen by MSG during training. Our analysis of the errors produced by source separators shows that waveform models tend to introduce more high-frequency noise, while spectrogram models tend to lose transients and high frequency content. We introduce objective measures to quantify both kinds of errors and show MSG improves the source reconstruction of both kinds of errors. Crowdsourced subjective evaluations demonstrate that human listeners prefer source estimates of bass and drums that have been post-processed by MSG.