今天跟大家分享一篇语音相关的论文合集:cs.SD语音2篇,eess.AS音频处理3篇。本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
【1】 CoBERT: Self-Supervised Speech Representation Learning Through Code Representation Learning
标题:Cobert:基于代码表示学习的自监督语音表示学习
链接:https://arxiv.org/abs/2210.04062
作者:Chutong Meng,Junyi Ao,Tom Ko,Mingxuan Wang,Haizhou Li机构:ByteDance AI Lab, School of Data Science, The Chinese University of Hong Kong, Shenzhen, China备注:Submitted to ICASSP 2023摘要:语音是有限的语音单位集合的表面形式,可以用离散码来表示。我们提出了一种用于自监督语音表示学习的编码BERT(CoBERT)方法。其思想是将话语转换为离散代码序列,并执行代码表示学习,其中我们基于原始语音输入的掩蔽视图来预测代码表示。与教师和学生具有相同模态的现有自我提炼方法不同,我们的目标模型预测来自不同模态的表征。CoBERT在ASR任务上的性能优于最新的最先进的性能,并在SUPERB语音翻译(ST)任务上带来了显著的改进。摘要:Speech is the surface form of a finite set of phonetic units, which can be represented by discrete codes. We propose the Code BERT (CoBERT) approach for self-supervised speech representation learning. The idea is to convert an utterance to a sequence of discrete codes, and perform code representation learning, where we predict the code representations based on a masked view of the original speech input. Unlike the prior self-distillation approaches of which the teacher and the student are of the same modality, our target model predicts representations from a different modality. CoBERT outperforms the most recent state-of-the-art performance on the ASR task and brings significant improvements on the SUPERB speech translation (ST) task.
【2】 Supervised and Unsupervised Learning of Audio Representations for Music Understanding
标题:音乐理解中音频表征的有监督和无监督学习
链接:https://arxiv.org/abs/2210.03799
作者:Matthew C. McCallum,Filip Korzeniowski,Sergio Oramas,Fabien Gouyon,Andreas F. Ehmann摘要:在这项工作中,我们提供了一个广泛的比较分析的策略,为预训练的音频理解模型的几个任务在音乐领域,包括标签的流派,时代,起源,情绪,乐器,关键,音高,声乐特征,节奏和响度。具体而言,我们探讨了预训练数据集(音乐或通用音频)的域和预训练方法(监督或非监督)如何影响所得到的音频嵌入对下游任务的充分性。我们的研究表明,在大规模专家标注的音乐数据集上通过监督学习训练的模型在各种音乐标注任务中都达到了最先进的性能,每个任务都有新颖的内容和词汇。这可以通过包含少于1亿个参数的模型以高效的方式完成,这些模型不需要为下游任务进行微调或重新参数化,使得这种方法对于工业规模的音频目录是实用的。在无监督学习策略类中,我们表明训练数据集的域可以显著影响模型学习的表示的性能。我们发现,将预训练数据集的域限制为音乐允许以较小的批量大小进行训练,同时实现用于音乐理解的无监督学习-在某些情况下,监督学习-的最新技术水平。我们还证实,虽然在许多任务上实现了最先进的性能,但监督学习可以使模型专门化于所提供的监督信息,在某种程度上损害了模型的通用性。摘要:In this work, we provide a broad comparative analysis of strategies for pre-training audio understanding models for several tasks in the music domain, including labelling of genre, era, origin, mood, instrumentation, key, pitch, vocal characteristics, tempo and sonority. Specifically, we explore how the domain of pre-training datasets (music or generic audio) and the pre-training methodology (supervised or unsupervised) affects the adequacy of the resulting audio embeddings for downstream tasks. We show that models trained via supervised learning on large-scale expert-annotated music datasets achieve state-of-the-art performance in a wide range of music labelling tasks, each with novel content and vocabularies. This can be done in an efficient manner with models containing less than 100 million parameters that require no fine-tuning or reparameterization for downstream tasks, making this approach practical for industry-scale audio catalogs. Within the class of unsupervised learning strategies, we show that the domain of the training dataset can significantly impact the performance of representations learned by the model. We find that restricting the domain of the pre-training dataset to music allows for training with smaller batch sizes while achieving state-of-the-art in unsupervised learning -- and in some cases, supervised learning -- for music understanding. We also corroborate that, while achieving state-of-the-art performance on many tasks, supervised learning can cause models to specialize to the supervised information provided, somewhat compromising a model's generality.
【1】 YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding
标题:YFACC:一种基于视觉基础的跨语言关键词定位的Yorúbá语音图像数据集
链接:https://arxiv.org/abs/2210.04600
作者:Kayode Olaleye,Dan Oneata,Herman Kamper机构:MediaLab, E&E Engineering, Stellenbosch University, South Africa, University Politehnica of Bucharest, Romania摘要:视觉接地语音(VGS)模型在与未标记的口述字幕配对的图像上训练。这样的模型可以用于在不可能获得标记数据的环境中构建语音系统,例如用于记录非书面语言。然而,大多数VGS研究都是以英语或其他高资源语言进行的。本文试图解决这一不足。我们收集并发布了一个新的6 k Flickr图片音频字幕的单发言者数据集,该数据集使用Yor\'ub\' a语言--尼日利亚使用的一种真正的低资源语言。我们训练了一个基于注意力的VGS模型,在这个模型中,图像被自动地标记上英语视觉标签,并与Yor\'ub\' a发音配对。这将启用跨语言关键字本地化:在语音中检测并定位书面英语查询。为了量化较小数据集的影响,我们将其与在相似和更多数据上训练的英语系统进行比较。我们希望这个新的数据集将刺激对真正的低资源语言使用VGS模型的研究。摘要:Visually grounded speech (VGS) models are trained on images paired with unlabelled spoken captions. Such models could be used to build speech systems in settings where it is impossible to get labelled data, e.g. for documenting unwritten languages. However, most VGS studies are in English or other high-resource languages. This paper attempts to address this shortcoming. We collect and release a new single-speaker dataset of audio captions for 6k Flickr images in Yor\`ub\'a -- a real low-resource language spoken in Nigeria. We train an attention-based VGS model where images are automatically tagged with English visual labels and paired with Yor\`ub\'a utterances. This enables cross-lingual keyword localisation: a written English query is detected and located in Yor\`ub\'a speech. To quantify the effect of the smaller dataset, we compare to English systems trained on similar and more data. We hope that this new dataset will stimulate research in the use of VGS models for real low-resource languages.
【2】 CoBERT: Self-Supervised Speech Representation Learning Through Code Representation Learning
标题:Cobert:基于代码表示学习的自监督语音表示学习
链接:https://arxiv.org/abs/2210.04062
* 与cs.SD语音【1】为同一篇
作者:Chutong Meng,Junyi Ao,Tom Ko,Mingxuan Wang,Haizhou Li机构:ByteDance AI Lab, School of Data Science, The Chinese University of Hong Kong, Shenzhen, China备注:Submitted to ICASSP 2023摘要:语音是有限的语音单位集合的表面形式,可以用离散码来表示。我们提出了一种用于自监督语音表示学习的编码BERT(CoBERT)方法。其思想是将话语转换为离散代码序列,并执行代码表示学习,其中我们基于原始语音输入的掩蔽视图来预测代码表示。与教师和学生具有相同模态的现有自我提炼方法不同,我们的目标模型预测来自不同模态的表征。CoBERT在ASR任务上的性能优于最新的最先进的性能,并在SUPERB语音翻译(ST)任务上带来了显著的改进。摘要:Speech is the surface form of a finite set of phonetic units, which can be represented by discrete codes. We propose the Code BERT (CoBERT) approach for self-supervised speech representation learning. The idea is to convert an utterance to a sequence of discrete codes, and perform code representation learning, where we predict the code representations based on a masked view of the original speech input. Unlike the prior self-distillation approaches of which the teacher and the student are of the same modality, our target model predicts representations from a different modality. CoBERT outperforms the most recent state-of-the-art performance on the ASR task and brings significant improvements on the SUPERB speech translation (ST) task.
【3】 Supervised and Unsupervised Learning of Audio Representations for Music Understanding
标题:音乐理解中音频表征的有监督和无监督学习
链接:https://arxiv.org/abs/2210.03799
* 与cs.SD语音【2】为同一篇
作者:Matthew C. McCallum,Filip Korzeniowski,Sergio Oramas,Fabien Gouyon,Andreas F. Ehmann摘要:在这项工作中,我们提供了一个广泛的比较分析的策略,为预训练的音频理解模型的几个任务在音乐领域,包括标签的流派,时代,起源,情绪,乐器,关键,音高,声乐特征,节奏和响度。具体而言,我们探讨了预训练数据集(音乐或通用音频)的域和预训练方法(监督或非监督)如何影响所得到的音频嵌入对下游任务的充分性。我们的研究表明,在大规模专家标注的音乐数据集上通过监督学习训练的模型在各种音乐标注任务中都达到了最先进的性能,每个任务都有新颖的内容和词汇。这可以通过包含少于1亿个参数的模型以高效的方式完成,这些模型不需要为下游任务进行微调或重新参数化,使得这种方法对于工业规模的音频目录是实用的。在无监督学习策略类中,我们表明训练数据集的域可以显著影响模型学习的表示的性能。我们发现,将预训练数据集的域限制为音乐允许以较小的批量大小进行训练,同时实现用于音乐理解的无监督学习-在某些情况下,监督学习-的最新技术水平。我们还证实,虽然在许多任务上实现了最先进的性能,但监督学习可以使模型专门化于所提供的监督信息,在某种程度上损害了模型的通用性。摘要:In this work, we provide a broad comparative analysis of strategies for pre-training audio understanding models for several tasks in the music domain, including labelling of genre, era, origin, mood, instrumentation, key, pitch, vocal characteristics, tempo and sonority. Specifically, we explore how the domain of pre-training datasets (music or generic audio) and the pre-training methodology (supervised or unsupervised) affects the adequacy of the resulting audio embeddings for downstream tasks. We show that models trained via supervised learning on large-scale expert-annotated music datasets achieve state-of-the-art performance in a wide range of music labelling tasks, each with novel content and vocabularies. This can be done in an efficient manner with models containing less than 100 million parameters that require no fine-tuning or reparameterization for downstream tasks, making this approach practical for industry-scale audio catalogs. Within the class of unsupervised learning strategies, we show that the domain of the training dataset can significantly impact the performance of representations learned by the model. We find that restricting the domain of the pre-training dataset to music allows for training with smaller batch sizes while achieving state-of-the-art in unsupervised learning -- and in some cases, supervised learning -- for music understanding. We also corroborate that, while achieving state-of-the-art performance on many tasks, supervised learning can cause models to specialize to the supervised information provided, somewhat compromising a model's generality.
机器翻译,仅供参考