cs.SD语音,共计6篇,eess.AS音频处理,共计6篇


1.cs.SD语音:

【1】 Dynamic Recognition of Speakers for Consent Management by Contrastive Embedding Replay

标题:基于对比嵌入重放的说话人动态识别同意管理

链接:https://arxiv.org/abs/2205.08459

作者:Arash Shahmansoori,Utz Roedig
机构: Roedig are with the School of Computer Science and Information Technology, University College Cork
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. The current version includes 32 pages, 10 figures, and 2 tables
摘要:语音助理记录声音,并可以偷听对话。因此,需要一种同意管理机制,以便用户能够表达其希望被记录或不被记录的愿望。可以使用说话人识别实现同意管理;未给予同意的用户注册其语音,随后不会处理这些用户的所有进一步录音。建立基于说话人识别的同意管理是一项具有挑战性的工作,因为该问题具有动态性,需要对大量说话人进行可扩展性,并且需要快速、高精度的说话人识别。本文描述了一个基于说话人识别的同意管理系统,以应对上述挑战。在对记录不同意见的说话人进行训练时,采用完全监督的批量对比学习来学习说话人的潜在等变诱导偏差。不提供同意的演讲者被分组在桶中,桶被持续训练。在训练过程中,这些嵌入被对比学习给桶中的说话者,然后作为分类的重放缓冲区。在训练过程中逐步注册存储桶,并提出了一种新的多步随机抽样对比嵌入重放缓冲区。在每次迭代中,只对桶进行几个步骤的对比训练,并对桶进行重放以逐步进行分类,从而实现快速收敛。描述了一种快速动态地注册和移除桶中扬声器的算法。评估结果表明,该方法为同意管理提供了所需的快速动态解决方案,并且在收敛速度、自适应能力以及推理过程中的验证性能方面优于现有方法。
摘要:Voice assistants record sound and can overhear conversations. Thus, a consent management mechanism is desirable such that users can express their wish to be recorded or not. Consent management can be implemented using speaker recognition; users that do not give consent enrol their voice and all further recordings of these users is subsequently not processed. Building speaker recognition based consent management is challenging due to the dynamic nature of the problem, required scalability for large number of speakers, and need for fast speaker recognition with high accuracy. This paper describes a speaker recognition based consent management system addressing the aforementioned challenges. A fully supervised batch contrastive learning is applied to learn the underlying speaker equivariance inductive bias during the training on the set of speakers noting recording dissent. Speakers that do not provide consent are grouped in buckets which are trained continuously. The embeddings are contrastively learned for speakers in their buckets during training and act later as a replay buffer for classification. The buckets are progressively registered during training and a novel multi-strided random sampling of the contrastive embedding replay buffer is proposed. Buckets are contrastively trained for a few steps only in each iteration and replayed for classification progressively leading to fast convergence. An algorithm for fast and dynamic registration and removal of speakers in buckets is described. The evaluation results show that the proposed approach provides the desired fast and dynamic solution for consent management and outperforms existing approaches in terms of convergence speed and adaptive capabilities as well as verification performance during inference.


【2】 Utterance Weighted Multi-Dilation Temporal Convolutional Networks for Monaural Speech Dereverberation

标题:用于单声道语音去混响的发音加权多伸缩时间卷积网络

链接:https://arxiv.org/abs/2205.08455

作者:William Ravenscroft,Stefan Goetze,Thomas Hain
机构:Department of Computer Science, The University of Sheffield , Sheffield, United Kingdom
备注:Submitted to IWAENC 2022
摘要:语音去冗余是许多语音技术应用中的一个重要阶段。最近在这一领域的研究主要是深度神经网络模型。时间卷积网络(TCN)是一种深度学习模型,用于语音去冗余的序列建模。本文提出了一种加权多重膨胀深度可分离卷积来代替TCN模型中的标准深度可分离卷积。这种提出的卷积使得TCN能够在网络中的每个卷积块上动态地关注其感受野中或多或少的局部信息。结果表明,这种加权多膨胀时间卷积网络(WD-TCN)在不同的模型配置中始终优于TCN,并且使用WD-TCN模型是一种比增加卷积块数更能提高模型性能的参数有效方法。在WHAMR数据集上,与基线TCN相比,最佳性能改进为0.55 dB尺度不变信噪比(SISDR),最佳性能WD-TCN模型达到12.26 dB SISDR。
摘要:Speech dereverberation is an important stage in many speech technology applications. Recent work in this area has been dominated by deep neural network models. Temporal convolutional networks (TCNs) are deep learning models that have been proposed for sequence modelling in the task of dereverberating speech. In this work a weighted multi-dilation depthwise-separable convolution is proposed to replace standard depthwise-separable convolutions in TCN models. This proposed convolution enables the TCN to dynamically focus on more or less local information in its receptive field at each convolutional block in the network. It is shown that this weighted multi-dilation temporal convolutional network (WD-TCN) consistently outperforms the TCN across various model configurations and using the WD-TCN model is a more parameter efficient method to improve the performance of the model than increasing the number of convolutional blocks. The best performance improvement over the baseline TCN is 0.55 dB scale-invariant signal-to-distortion ratio (SISDR) and the best performing WD-TCN model attains 12.26 dB SISDR on the WHAMR dataset.


【3】 SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation

标题:SAMU-XLSR:语义对齐的多通道语音级跨语言语音表示

链接:https://arxiv.org/abs/2205.08180

作者:Sameer Khurana,Antoine Laurent,James Glass
机构:MIT Computer Science and Artificial Intelligence Laboratory, Cambridge, MA, USA, LIUM - Le Mans University, France
摘要:我们提出了SAMU-XLSR:语义一致的多模态话语水平跨语言语音表征学习框架。与之前关于语音表征学习的工作不同,这项工作在声学帧分辨率(10-20ms)下学习多语言上下文语音嵌入,而在句子分辨率(5-10s)下学习多模态(语音文本)多语言语音嵌入,使得嵌入向量空间在不同语言之间语义对齐。我们将最先进的多语言声学帧级语音表示学习模型XLS-R与语言无关的BERT语句嵌入(LaBSE)模型相结合,创建了一个话语级多模态多语言语音编码器SAMU-XLSR。虽然我们只使用多语言转录语音数据训练SAMU-XLSR,但跨语言语音文本和语音关联出现在其学习的表示空间中。为了证实我们的说法,我们将SAMU-XLSR语音编码器与经过预训练的LaBSE文本句子编码器结合使用,用于跨语言语音到文本翻译检索,并将SAMU-XLSR单独用于跨语言语音到语音翻译检索。我们通过跨多个数据集执行多个跨语言文本和语音翻译检索任务来突出这些应用程序。
摘要:We propose the SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation learning framework. Unlike previous works on speech representation learning, which learns multilingual contextual speech embedding at the resolution of an acoustic frame (10-20ms), this work focuses on learning multimodal (speech-text) multilingual speech embedding at the resolution of a sentence (5-10s) such that the embedding vector space is semantically aligned across different languages. We combine state-of-the-art multilingual acoustic frame-level speech representation learning model XLS-R with the Language Agnostic BERT Sentence Embedding (LaBSE) model to create an utterance-level multimodal multilingual speech encoder SAMU-XLSR. Although we train SAMU-XLSR with only multilingual transcribed speech data, cross-lingual speech-text and speech-speech associations emerge in its learned representation space. To substantiate our claims, we use SAMU-XLSR speech encoder in combination with a pre-trained LaBSE text sentence encoder for cross-lingual speech-to-text translation retrieval, and SAMU-XLSR alone for cross-lingual speech-to-speech translation retrieval. We highlight these applications by performing several cross-lingual text and speech translation retrieval tasks across several datasets.


【4】 Perceptual Evaluation on Audio-visual Dataset of 360 Content

标题:360内容视听数据集的感知评价

链接:https://arxiv.org/abs/2205.08007

作者:Randy F Fela,Andréas Pastor,Patrick Le Callet,Nick Zacharov,Toinon Vigier,Søren Forchhammer机构:SenseLab, FORCE Technology, Hørsholm, Denmark, Nantes Universit´e, Ecole Centrale Nantes, CNRS, LS,N, UMR , F-, Nantes, France, Meta Reality Labs – Meta, Paris, France, DTU Electro, Technical University of Denmark, Kgs. Lyngby, Denmark
备注:6 pages, 5 figures, International Conference on Multimedia and Expo 2022
摘要:为了开辟新的可能性来评估全方位媒体格式的多模态感知质量,我们提出了一个新的开源360视听(AV)质量数据集。该数据集包括高质量的360个等矩形(ERP)格式的视频剪辑和高阶的ambisonic(4阶)以及主观分数。对音频、视频和AV进行了三次主观质量实验,并在本文中详细介绍了实验过程。利用主观测试的数据,我们证明了该数据集可用于量化感知的音频、视频和视听质量。并对主观得分的多样性和可辨别性进行了分析。最后,我们研究了我们的数据集如何与音频和视频的各种客观质量指标相关联。本研究结果的证据表明,所提出的数据集有助于未来360内容多模式质量评估的研究。
摘要:To open up new possibilities to assess the multimodal perceptual quality of omnidirectional media formats, we proposed a novel open source 360 audiovisual (AV) quality dataset. The dataset consists of high-quality 360 video clips in equirectangular (ERP) format and higher-order ambisonic (4th order) along with the subjective scores. Three subjective quality experiments were conducted for audio, video, and AV with the procedures detailed in this paper. Using the data from subjective tests, we demonstrated that this dataset can be used to quantify perceived audio, video, and audiovisual quality. The diversity and discriminability of subjective scores were also analyzed. Finally, we investigated how our dataset correlates with various objective quality metrics of audio and video. Evidence from the results of this study implies that the proposed dataset can benefit future studies on multimodal quality evaluation of 360 content.


【5】 Composing General Audio Representation by Fusing Multilayer Features of a Pre-trained Model

标题:融合预先训练模型的多层特征合成通用音频表示

链接:https://arxiv.org/abs/2205.08138

作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Kunio Kashino
机构:NTT Communication Science Laboratories, NTT Corporation, Atsugi, Japan
备注:5 pages, 4 figures and 4 tables. Accepted by EUSIPCO 2022
摘要:许多应用研究依赖于在大规模数据集上预先训练的音频DNN模型作为基本特征提取器,并从最后一层提取特征。在本研究中,我们重点发现,对于某些任务,现有的有监督预训练模型的中间层特征比后期特征更有效。我们提出了一种简单的方法来合成对通用应用有效的特征,包括两个步骤:(1)从中/晚层输出沿时间帧计算特征向量,以及(2)融合它们。这种方法提高了频率和信道信息在下游过程中的效用,并结合了中后期特征对不同任务的有效性。因此,特征向量在一般情况下变得有效。在使用VGISH、PANNs的CNN14和AST对九个下游任务进行的实验中,我们首先表明,这些模型的每一层输出服务于不同的任务。然后,我们证明了所提出的方法显著提高了它们的性能,并使其达到了与最新技术相当的水平。特别是,非语义语音(NOSS)任务的性能大大提高,尤其是在VGGish为+77.1(14.3%至91.4%)的语音命令V2上。
摘要:Many application studies rely on audio DNN models pre-trained on a large-scale dataset as essential feature extractors, and they extract features from the last layers. In this study, we focus on our finding that the middle layer features of existing supervised pre-trained models are more effective than the late layer features for some tasks. We propose a simple approach to compose features effective for general-purpose applications, consisting of two steps: (1) calculating feature vectors along the time frame from middle/late layer outputs, and (2) fusing them. This approach improves the utility of frequency and channel information in downstream processes, and combines the effectiveness of middle and late layer features for different tasks. As a result, the feature vectors become effective for general purposes. In the experiments using VGGish, PANNs' CNN14, and AST on nine downstream tasks, we first show that each layer output of these models serves different tasks. Then, we demonstrate that the proposed approach significantly improves their performance and brings it to a level comparable to that of the state-of-the-art. In particular, the performance of the non-semantic speech (NOSS) tasks greatly improves, especially on Speech commands V2 with VGGish of +77.1 (14.3% to 91.4%).


【6】 Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data

标题:重音语音识别:基准测试、预训练和多样化数据

链接:https://arxiv.org/abs/2205.08014

作者:Alëna Aksënova,Zhehuai Chen,Chung-Cheng Chiu,Daan van Esch,Pavel Golik,Wei Han,Levi King,Bhuvana Ramabhadran,Andrew Rosenberg,Suzan Schwartz,Gary Wang
机构:Google Speech, Google Brain
备注:5 pages, 3 tables
摘要:构建包容性语音识别系统是开发各种语言使用者都能使用的技术的关键一步。因此,ASR系统必须独立于每个人的说话方式来为他们工作。为了实现这一目标,应该有可用的数据集来表示语言的多样性,还应该有对模型配置的理解,这对于实现对所有类型语音的健壮理解是最有帮助的。然而,目前还没有足够的重音语音数据集,对于已有的数据集,需要探索更多的训练方法来提高重音语音识别的质量。在本文中,我们讨论了开发更具包容性的ASR系统的最新进展,即构建代表语言多样性的新数据集的重要性,以及探索新的训练方法以提高所有用户的性能。我们讨论了ASR系统中口音语音基准测试的最新方向,测量了wav2vec 2.0预训练对口音语音识别的影响,并强调了与不同ASR评估相关的语料库。
摘要:Building inclusive speech recognition systems is a crucial step towards developing technologies that speakers of all language varieties can use. Therefore, ASR systems must work for everybody independently of the way they speak. To accomplish this goal, there should be available data sets representing language varieties, and also an understanding of model configuration that is the most helpful in achieving robust understanding of all types of speech. However, there are not enough data sets for accented speech, and for the ones that are already available, more training approaches need to be explored to improve the quality of accented speech recognition. In this paper, we discuss recent progress towards developing more inclusive ASR systems, namely, the importance of building new data sets representing linguistic diversity, and exploring novel training approaches to improve performance for all users. We address recent directions within benchmarking ASR systems for accented speech, measure the effects of wav2vec 2.0 pre-training on accented speech recognition, and highlight corpora relevant for diverse ASR evaluations.


2.eess.AS音频处理:

【1】 Composing General Audio Representation by Fusing Multilayer Features of a Pre-trained Model

标题:融合预先训练模型的多层特征合成通用音频表示

链接:https://arxiv.org/abs/2205.08138

作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Kunio Kashino
机构:NTT Communication Science Laboratories, NTT Corporation, Atsugi, Japan
备注:5 pages, 4 figures and 4 tables. Accepted by EUSIPCO 2022
摘要:许多应用研究依赖于在大规模数据集上预先训练的音频DNN模型作为基本特征提取器,并从最后一层提取特征。在本研究中,我们重点发现,对于某些任务,现有的有监督预训练模型的中间层特征比后期特征更有效。我们提出了一种简单的方法来合成对通用应用有效的特征,包括两个步骤:(1)从中/晚层输出沿时间帧计算特征向量,以及(2)融合它们。这种方法提高了频率和信道信息在下游过程中的效用,并结合了中后期特征对不同任务的有效性。因此,特征向量在一般情况下变得有效。在使用VGISH、PANNs的CNN14和AST对九个下游任务进行的实验中,我们首先表明,这些模型的每一层输出服务于不同的任务。然后,我们证明了所提出的方法显著提高了它们的性能,并使其达到了与最新技术相当的水平。特别是,非语义语音(NOSS)任务的性能大大提高,尤其是在VGGish为+77.1(14.3%至91.4%)的语音命令V2上。
摘要:Many application studies rely on audio DNN models pre-trained on a large-scale dataset as essential feature extractors, and they extract features from the last layers. In this study, we focus on our finding that the middle layer features of existing supervised pre-trained models are more effective than the late layer features for some tasks. We propose a simple approach to compose features effective for general-purpose applications, consisting of two steps: (1) calculating feature vectors along the time frame from middle/late layer outputs, and (2) fusing them. This approach improves the utility of frequency and channel information in downstream processes, and combines the effectiveness of middle and late layer features for different tasks. As a result, the feature vectors become effective for general purposes. In the experiments using VGGish, PANNs' CNN14, and AST on nine downstream tasks, we first show that each layer output of these models serves different tasks. Then, we demonstrate that the proposed approach significantly improves their performance and brings it to a level comparable to that of the state-of-the-art. In particular, the performance of the non-semantic speech (NOSS) tasks greatly improves, especially on Speech commands V2 with VGGish of +77.1 (14.3% to 91.4%).


【2】 Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data

标题:重音语音识别:基准测试、预训练和多样化数据

链接:https://arxiv.org/abs/2205.08014

作者:Alëna Aksënova,Zhehuai Chen,Chung-Cheng Chiu,Daan van Esch,Pavel Golik,Wei Han,Levi King,Bhuvana Ramabhadran,Andrew Rosenberg,Suzan Schwartz,Gary Wang
机构:Google Speech, Google Brain
备注:5 pages, 3 tables
摘要:构建包容性语音识别系统是开发各种语言使用者都能使用的技术的关键一步。因此,ASR系统必须独立于每个人的说话方式来为他们工作。为了实现这一目标,应该有可用的数据集来表示语言的多样性,还应该有对模型配置的理解,这对于实现对所有类型语音的健壮理解是最有帮助的。然而,目前还没有足够的重音语音数据集,对于已有的数据集,需要探索更多的训练方法来提高重音语音识别的质量。在本文中,我们讨论了开发更具包容性的ASR系统的最新进展,即构建代表语言多样性的新数据集的重要性,以及探索新的训练方法以提高所有用户的性能。我们讨论了ASR系统中口音语音基准测试的最新方向,测量了wav2vec 2.0预训练对口音语音识别的影响,并强调了与不同ASR评估相关的语料库。
摘要:Building inclusive speech recognition systems is a crucial step towards developing technologies that speakers of all language varieties can use. Therefore, ASR systems must work for everybody independently of the way they speak. To accomplish this goal, there should be available data sets representing language varieties, and also an understanding of model configuration that is the most helpful in achieving robust understanding of all types of speech. However, there are not enough data sets for accented speech, and for the ones that are already available, more training approaches need to be explored to improve the quality of accented speech recognition. In this paper, we discuss recent progress towards developing more inclusive ASR systems, namely, the importance of building new data sets representing linguistic diversity, and exploring novel training approaches to improve performance for all users. We address recent directions within benchmarking ASR systems for accented speech, measure the effects of wav2vec 2.0 pre-training on accented speech recognition, and highlight corpora relevant for diverse ASR evaluations.


【3】 Dynamic Recognition of Speakers for Consent Management by Contrastive Embedding Replay

标题:基于对比嵌入重放的说话人动态识别同意管理

链接:https://arxiv.org/abs/2205.08459

作者:Arash Shahmansoori,Utz Roedig
机构: Roedig are with the School of Computer Science and Information Technology, University College Cork备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. The current version includes 32 pages, 10 figures, and 2 tables
摘要:语音助理记录声音,并可以偷听对话。因此,需要一种同意管理机制,以便用户能够表达其希望被记录或不被记录的愿望。可以使用说话人识别实现同意管理;未给予同意的用户注册其语音,随后不会处理这些用户的所有进一步录音。建立基于说话人识别的同意管理是一项具有挑战性的工作,因为该问题具有动态性,需要对大量说话人进行可扩展性,并且需要快速、高精度的说话人识别。本文描述了一个基于说话人识别的同意管理系统,以应对上述挑战。在对记录不同意见的说话人进行训练时,采用完全监督的批量对比学习来学习说话人的潜在等变诱导偏差。不提供同意的演讲者被分组在桶中,桶被持续训练。在训练过程中,这些嵌入被对比学习给桶中的说话者,然后作为分类的重放缓冲区。在训练过程中逐步注册存储桶,并提出了一种新的多步随机抽样对比嵌入重放缓冲区。在每次迭代中,只对桶进行几个步骤的对比训练,并对桶进行重放以逐步进行分类,从而实现快速收敛。描述了一种快速动态地注册和移除桶中扬声器的算法。评估结果表明,该方法为同意管理提供了所需的快速动态解决方案,并且在收敛速度、自适应能力以及推理过程中的验证性能方面优于现有方法。
摘要:Voice assistants record sound and can overhear conversations. Thus, a consent management mechanism is desirable such that users can express their wish to be recorded or not. Consent management can be implemented using speaker recognition; users that do not give consent enrol their voice and all further recordings of these users is subsequently not processed. Building speaker recognition based consent management is challenging due to the dynamic nature of the problem, required scalability for large number of speakers, and need for fast speaker recognition with high accuracy. This paper describes a speaker recognition based consent management system addressing the aforementioned challenges. A fully supervised batch contrastive learning is applied to learn the underlying speaker equivariance inductive bias during the training on the set of speakers noting recording dissent. Speakers that do not provide consent are grouped in buckets which are trained continuously. The embeddings are contrastively learned for speakers in their buckets during training and act later as a replay buffer for classification. The buckets are progressively registered during training and a novel multi-strided random sampling of the contrastive embedding replay buffer is proposed. Buckets are contrastively trained for a few steps only in each iteration and replayed for classification progressively leading to fast convergence. An algorithm for fast and dynamic registration and removal of speakers in buckets is described. The evaluation results show that the proposed approach provides the desired fast and dynamic solution for consent management and outperforms existing approaches in terms of convergence speed and adaptive capabilities as well as verification performance during inference.


【4】 Utterance Weighted Multi-Dilation Temporal Convolutional Networks for Monaural Speech Dereverberation

标题:用于单声道语音去混响的发音加权多伸缩时间卷积网络

链接:https://arxiv.org/abs/2205.08455

作者:William Ravenscroft,Stefan Goetze,Thomas Hain
机构:Department of Computer Science, The University of Sheffield , Sheffield, United Kingdom
备注:Submitted to IWAENC 2022
摘要:语音去冗余是许多语音技术应用中的一个重要阶段。最近在这一领域的研究主要是深度神经网络模型。时间卷积网络(TCN)是一种深度学习模型,用于语音去冗余的序列建模。本文提出了一种加权多重膨胀深度可分离卷积来代替TCN模型中的标准深度可分离卷积。这种提出的卷积使得TCN能够在网络中的每个卷积块上动态地关注其感受野中或多或少的局部信息。结果表明,这种加权多膨胀时间卷积网络(WD-TCN)在不同的模型配置中始终优于TCN,并且使用WD-TCN模型是一种比增加卷积块数更能提高模型性能的参数有效方法。在WHAMR数据集上,与基线TCN相比,最佳性能改进为0.55 dB尺度不变信噪比(SISDR),最佳性能WD-TCN模型达到12.26 dB SISDR。
摘要:Speech dereverberation is an important stage in many speech technology applications. Recent work in this area has been dominated by deep neural network models. Temporal convolutional networks (TCNs) are deep learning models that have been proposed for sequence modelling in the task of dereverberating speech. In this work a weighted multi-dilation depthwise-separable convolution is proposed to replace standard depthwise-separable convolutions in TCN models. This proposed convolution enables the TCN to dynamically focus on more or less local information in its receptive field at each convolutional block in the network. It is shown that this weighted multi-dilation temporal convolutional network (WD-TCN) consistently outperforms the TCN across various model configurations and using the WD-TCN model is a more parameter efficient method to improve the performance of the model than increasing the number of convolutional blocks. The best performance improvement over the baseline TCN is 0.55 dB scale-invariant signal-to-distortion ratio (SISDR) and the best performing WD-TCN model attains 12.26 dB SISDR on the WHAMR dataset.


【5】 SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation

标题:SAMU-XLSR:语义对齐的多通道语音级跨语言语音表示

链接:https://arxiv.org/abs/2205.08180

作者:Sameer Khurana,Antoine Laurent,James Glass
机构:MIT Computer Science and Artificial Intelligence Laboratory, Cambridge, MA, USA, LIUM - Le Mans University, France
摘要:我们提出了SAMU-XLSR:语义一致的多模态话语水平跨语言语音表征学习框架。与之前关于语音表征学习的工作不同,这项工作在声学帧分辨率(10-20ms)下学习多语言上下文语音嵌入,而在句子分辨率(5-10s)下学习多模态(语音文本)多语言语音嵌入,使得嵌入向量空间在不同语言之间语义对齐。我们将最先进的多语言声学帧级语音表示学习模型XLS-R与语言无关的BERT语句嵌入(LaBSE)模型相结合,创建了一个话语级多模态多语言语音编码器SAMU-XLSR。虽然我们只使用多语言转录语音数据训练SAMU-XLSR,但跨语言语音文本和语音关联出现在其学习的表示空间中。为了证实我们的说法,我们将SAMU-XLSR语音编码器与经过预训练的LaBSE文本句子编码器结合使用,用于跨语言语音到文本翻译检索,并将SAMU-XLSR单独用于跨语言语音到语音翻译检索。我们通过跨多个数据集执行多个跨语言文本和语音翻译检索任务来突出这些应用程序。
摘要:We propose the SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation learning framework. Unlike previous works on speech representation learning, which learns multilingual contextual speech embedding at the resolution of an acoustic frame (10-20ms), this work focuses on learning multimodal (speech-text) multilingual speech embedding at the resolution of a sentence (5-10s) such that the embedding vector space is semantically aligned across different languages. We combine state-of-the-art multilingual acoustic frame-level speech representation learning model XLS-R with the Language Agnostic BERT Sentence Embedding (LaBSE) model to create an utterance-level multimodal multilingual speech encoder SAMU-XLSR. Although we train SAMU-XLSR with only multilingual transcribed speech data, cross-lingual speech-text and speech-speech associations emerge in its learned representation space. To substantiate our claims, we use SAMU-XLSR speech encoder in combination with a pre-trained LaBSE text sentence encoder for cross-lingual speech-to-text translation retrieval, and SAMU-XLSR alone for cross-lingual speech-to-speech translation retrieval. We highlight these applications by performing several cross-lingual text and speech translation retrieval tasks across several datasets.


【6】 Perceptual Evaluation on Audio-visual Dataset of 360 Content

标题:360内容视听数据集的感知评价

链接:https://arxiv.org/abs/2205.08007

作者:Randy F Fela,Andréas Pastor,Patrick Le Callet,Nick Zacharov,Toinon Vigier,Søren Forchhammer机构:SenseLab, FORCE Technology, Hørsholm, Denmark, Nantes Universit´e, Ecole Centrale Nantes, CNRS, LS,N, UMR , F-, Nantes, France, Meta Reality Labs – Meta, Paris, France, DTU Electro, Technical University of Denmark, Kgs. Lyngby, Denmark
备注:6 pages, 5 figures, International Conference on Multimedia and Expo 2022
摘要:为了开辟新的可能性来评估全方位媒体格式的多模态感知质量,我们提出了一个新的开源360视听(AV)质量数据集。该数据集包括高质量的360个等矩形(ERP)格式的视频剪辑和高阶的ambisonic(4阶)以及主观分数。对音频、视频和AV进行了三次主观质量实验,并在本文中详细介绍了实验过程。利用主观测试的数据,我们证明了该数据集可用于量化感知的音频、视频和视听质量。并对主观得分的多样性和可辨别性进行了分析。最后,我们研究了我们的数据集如何与音频和视频的各种客观质量指标相关联。本研究结果的证据表明,所提出的数据集有助于未来360内容多模式质量评估的研究。
摘要:To open up new possibilities to assess the multimodal perceptual quality of omnidirectional media formats, we proposed a novel open source 360 audiovisual (AV) quality dataset. The dataset consists of high-quality 360 video clips in equirectangular (ERP) format and higher-order ambisonic (4th order) along with the subjective scores. Three subjective quality experiments were conducted for audio, video, and AV with the procedures detailed in this paper. Using the data from subjective tests, we demonstrated that this dataset can be used to quantify perceived audio, video, and audiovisual quality. The diversity and discriminability of subjective scores were also analyzed. Finally, we investigated how our dataset correlates with various objective quality metrics of audio and video. Evidence from the results of this study implies that the proposed dataset can benefit future studies on multimodal quality evaluation of 360 content.


机器翻译,仅供参考