今日论文合集:cs.SD语音12篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily


cs.SD语音

【1】 Maestro-U: Leveraging joint speech-text representation learning for zero  supervised speech ASR

标题:Maestro-U:利用语音-文本联合表征学习实现零监督语音ASR

链接:https://arxiv.org/abs/2210.10027

作者:Zhehuai Chen,Ankur Bapna,Andrew Rosenberg,Yu Zhang,Bhuvana Ramabhadran,Pedro Moreno,Nanxin Chen
机构:Google, Inc.
备注:Accepted by SLT 2022
摘要:训练现有技术的自动语音识别(ASR)模型通常需要大量的转录语音。在这项工作中,我们证明了一个模态匹配的联合语音和文本模型可以用来训练一个大规模的多语言ASR模型,而不需要任何监督(手动转录)的语音。本文探讨了在大规模多语言、零监督语音、真实世界设置中使用联合学习的语音和文本表示,以扩展ASR所覆盖的语言集,仅使用目标语言中的未标记语音和文本。使用FLEURS数据集,我们将任务定义为覆盖$102$种语言,其中转录语音在$52$种语言中可用,并可用于提高剩余$50$中的端到端ASR质量。首先,我们证明了通过将语音表示与字节级文本表示相结合,并使用语言嵌入,我们可以将无监督语音的语言的字符错误率(CER)从64.8%显著降低到30.8%,相对降低了53%。第二,使用南亚语言的子集,我们表明Maestro-U可以促进来自有监督语音的语言的知识转移,即使有限制到没有字形重叠。总体而言,Maestro-U将与Oracle的性能差距缩小了68.5%,并将19种语言的CER降低到15%以下。
摘要:Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model can be leveraged to train a massively multilingual ASR model without any supervised (manually transcribed) speech for some languages. This paper explores the use of jointly learnt speech and text representations in a massively multilingual, zero supervised speech, real-world setting to expand the set of languages covered by ASR with only unlabeled speech and text in the target languages. Using the FLEURS dataset, we define the task to cover $102$ languages, where transcribed speech is available in $52$ of these languages and can be used to improve end-to-end ASR quality on the remaining $50$. First, we show that by combining speech representations with byte-level text representations and use of language embeddings, we can dramatically reduce the Character Error Rate (CER) on languages with no supervised speech from 64.8\% to 30.8\%, a relative reduction of 53\%. Second, using a subset of South Asian languages we show that Maestro-U can promote knowledge transfer from languages with supervised speech even when there is limited to no graphemic overlap. Overall, Maestro-U closes the gap to oracle performance by 68.5\% relative and reduces the CER of 19 languages below 15\%.


【2】 HMM vs. CTC for Automatic Speech Recognition: Comparison Based on  Full-Sum Training from Scratch

标题:基于全和训练的HMM与CTC在自动语音识别中的比较

链接:https://arxiv.org/abs/2210.09951

作者:Tina Raissi,Wei Zhou,Simon Berger,Ralf Schlüter,Hermann Ney
机构:Human Language Technology and Pattern Recognition Group, RWTH Aachen University, Germany, AppTek GmbH, Aachen, Germany
备注:Accepted for Presentation at IEEE SLT 2022
摘要:在这项工作中,我们从零开始比较隐马尔可夫模型(HMM)的序列级交叉熵(全和)训练和用于自动语音识别(ASR)的连接主义时态分类(CTC)拓扑。除了准确性之外,我们还分析了它们在语音信号和转录之间生成高质量时间对准的能力,这对许多后续应用至关重要。此外,我们提出了几种方法来改善从头开始的全和训练的收敛性,通过解决对准建模问题。在Switchboard和LibriSpeech语料库上对CTC、有转移概率和无转移概率的后验HMM以及标准混合HMM进行了系统比较。我们还提供了维特比强制对齐和Baum-Welch全和占用概率的详细分析。
摘要:In this work, we compare from-scratch sequence-level cross-entropy (full-sum) training of Hidden Markov Model (HMM) and Connectionist Temporal Classification (CTC) topologies for automatic speech recognition (ASR). Besides accuracy, we further analyze their capability for generating high-quality time alignment between the speech signal and the transcription, which can be crucial for many subsequent applications. Moreover, we propose several methods to improve convergence of from-scratch full-sum training by addressing the alignment modeling issue. Systematic comparison is conducted on both Switchboard and LibriSpeech corpora across CTC, posterior HMM with and w/o transition probabilities, and standard hybrid HMM. We also provide a detailed analysis of both Viterbi forced-alignment and Baum-Welch full-sum occupation probabilities.


【3】 Mid-attribute speaker generation using optimal-transport-based  interpolation of Gaussian mixture models

标题:基于最优传输的混合高斯模型内插中属性说话人生成

链接:https://arxiv.org/abs/2210.09916

作者:Aya Watanabe,Shinnosuke Takamichi,Yuki Saito,Detai Xin,Hiroshi Saruwatari
机构:The University of Tokyo, Japan.
备注:Submitted to ICASSP 2023. Demo: this https URL
摘要:本文提出了一种在“说话人生成”中综合多个说话人的属性并使其语音特征多样化的方法,“说话人生成”是一个新兴的任务,其目的是合成不存在的说话人的自然发声的语音。传统的基于TacoSpawn的说话人生成方法是利用高斯混合模型(Gaussian mixture models,GMM)来描述说话人嵌入的分布,并考虑说话人的属性。尽管该方法使得能够从知晓说话者属性的GMM中对各种说话者进行采样,但是还不清楚所学习的分布是否能够表示具有中间属性(即,中间属性)。为此,我们提出了一种基于最优传输的方法,该方法对学习的GMM进行插值,以生成具有中间属性(例如,性别中立)的声音。我们通过实验验证了我们的方法,并评估了合成语音的自然度和两个说话者属性的可控性:性别和语言流利性。实验结果表明,该方法能够在不降低语音自然度的前提下,通过连续的标量值来控制说话人的属性。
摘要:In this paper, we propose a method for intermediating multiple speakers' attributes and diversifying their voice characteristics in ``speaker generation,'' an emerging task that aims to synthesize a nonexistent speaker's naturally sounding voice. The conventional TacoSpawn-based speaker generation method represents the distributions of speaker embeddings by Gaussian mixture models (GMMs) conditioned with speaker attributes. Although this method enables the sampling of various speakers from the speaker-attribute-aware GMMs, it is not yet clear whether the learned distributions can represent speakers with an intermediate attribute (i.e., mid-attribute). To this end, we propose an optimal-transport-based method that interpolates the learned GMMs to generate nonexistent speakers with mid-attribute (e.g., gender-neutral) voices. We empirically validate our method and evaluate the naturalness of synthetic speech and the controllability of two speaker attributes: gender and language fluency. The evaluation results show that our method can control the generated speakers' attributes by a continuous scalar value without statistically significant degradation of speech naturalness.


【4】 Spontaneous speech synthesis with linguistic-speech consistency training  using pseudo-filled pauses

标题:利用伪填充停顿进行语言-语音一致性训练的自发语音合成

链接:https://arxiv.org/abs/2210.09815

作者:Yuta Matsunaga,Takaaki Saeki,Shinnosuke Takamichi,Hiroshi Saruwatari
机构:Graduate School of Information Science and Technology, The University of Tokyo, Japan.
备注:Submitted to ICASSP 2023
摘要:本文提出了一种自发语音合成模型的训练方法,保证了合成语音各语言成分的一致性。个性化的自发语音合成旨在再现不流畅的个性,例如填充的停顿。我们的先验模型包括填充停顿预测模型,并且从没有填充停顿的文本合成包括填充停顿的语音。然而,插入填充的停顿降低了合成语音的语言部分的质量。这可能是因为在训练和推理之间填充停顿插入的倾向不同,并且合成模型不能表示在推理中填充停顿和周围音素之间的连接。因此,我们开发了一种语言-语音一致性训练,它保证了有和没有填充停顿的合成语音的语言部分的一致性。所提出的一致性训练不仅利用地面真实填充的停顿,而且利用伪停顿。实验结果表明,该方法提高了合成语音的自然度,提高了整个语音的自然度。
摘要:We propose a training method for spontaneous speech synthesis models that guarantees the consistency of linguistic parts of synthesized speech. Personalized spontaneous speech synthesis aims to reproduce the individuality of disfluency, such as filled pauses. Our prior model includes a filled-pause prediction model and synthesizes filled-pause-included speech from text without filled pauses. However, inserting the filled pauses degrades the quality of the linguistic parts of the synthesized speech. This might be because filled-pause insertion tendencies differ between training and inference, and the synthesis model cannot represent connections between filled pauses and surrounding phonemes in inference. We, therefore, developed a linguistic-speech consistency training that guarantees the consistency of linguistic parts of synthetic speech with and without filled pauses. The proposed consistency training utilizes not only ground-truth-filled pauses but also pseudo ones. Our experiments demonstrate that this method improves the naturalness of the synthetic linguistic speech and the entire predicted-filled-pause-included synthetic speech.


【5】 Discrete Cross-Modal Alignment Enables Zero-Shot Speech Translation

标题:离散跨模式对齐实现零发声语音翻译

链接:https://arxiv.org/abs/2210.09556

作者:Chen Wang,Yuchen Liu,Boxing Chen,Jiajun Zhang,Wei Luo,Zhongqiang Huang,Chengqing Zong机构:National Laboratory of Pattern Recognition, Institute of Automation, CAS, Beijing, China,  Machine Intelligence Technology Lab, Alibaba DAMO Academy
备注:Accepted by the main conference of EMNLP 2022
摘要:端到端语音翻译的目的是在不产生中间转录的情况下将源语言语音翻译成目标语言文本。然而,端到端方法的训练依赖于并行ST数据,而并行ST数据难以获得且昂贵。幸运的是,用于自动语音识别(ASR)和机器翻译(MT)的监督数据通常更容易获得,这使得zero-shot语音翻译成为一个潜在的方向。现有的zero-shot方法无法将语音和文本两种模态对齐到共享的语义空间中,导致性能比有监督的ST方法差得多。为了实现zero-shot ST,提出了一种新的离散跨模态对齐(DCMA)方法,该方法利用一个共享的离散词汇空间来容纳和匹配语音和文本模态。具体地说,我们引入了矢量量化模块,将语音和文本的连续表示离散化为有限的虚拟令牌集,并使用ASR数据将相应的语音和文本映射到共享码本中的同一虚拟令牌。这样,源语言语音可以嵌入到与源语言文本相同的语义空间中,然后可以利用MT模块将源语言文本转换为目标语言文本。在多个语言对上的实验表明,本文的zero-shot ST方法显著提高了SOTA,甚至达到了与强监督ST基线相当的性能.
摘要:End-to-end Speech Translation (ST) aims at translating the source language speech into target language text without generating the intermediate transcriptions. However, the training of end-to-end methods relies on parallel ST data, which are difficult and expensive to obtain. Fortunately, the supervised data for automatic speech recognition (ASR) and machine translation (MT) are usually more accessible, making zero-shot speech translation a potential direction. Existing zero-shot methods fail to align the two modalities of speech and text into a shared semantic space, resulting in much worse performance compared to the supervised ST methods. In order to enable zero-shot ST, we propose a novel Discrete Cross-Modal Alignment (DCMA) method that employs a shared discrete vocabulary space to accommodate and match both modalities of speech and text. Specifically, we introduce a vector quantization module to discretize the continuous representations of speech and text into a finite set of virtual tokens, and use ASR data to map corresponding speech and text to the same virtual token in a shared codebook. This way, source language speech can be embedded in the same semantic space as the source language text, which can be then transformed into target language text with an MT module. Experiments on multiple language pairs demonstrate that our zero-shot ST method significantly improves the SOTA, and even performers on par with the strong supervised ST baselines.


【6】 A Hybrid System of Sound Event Detection Transformer and Frame-wise  Model for DCASE 2022 Task 4

标题:用于DCASE 2022任务4的声事件检测Transformer和框架模型的混合系统

链接:https://arxiv.org/abs/2210.09529

作者:Yiming Li,Zhifang Guo,Zhirong Ye,Xiangdong Wang,Hong Liu,Yueliang Qian,Rui Tao,Long Yan,Kazushige Ouchi
机构:Beijing Key Laboratory of Mobile Computing and Pervasive Device, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China,  University of Chinese Academy of Sciences, Beijing, China,  Toshiba China R&D Center, Beijing, China备注:5 pages, 2 figures, accepted for publication in DCASE2022 Workshop
摘要:本文详细介绍了DCASE 2022 Task 4, 0系统。该系统结合了两种截然不同的模式:端到端声音事件检测Transformer(SEDT)和逐帧模型、度量学习和聚焦损失CNN(MLFL-CNN)。前者是一种基于事件的模型,学习事件级表示,直接预测声音事件类别和边界;后者基于广泛采用的帧分类方案,将每个帧分类为事件类别,并通过阈值化和平滑等后处理获得事件边界。对于SEDT,使用未标记数据进行自监督预训练,并使用在线教师进行半监督学习,在线教师使用指数移动平均(EMA)策略从学生模型更新,并为弱标记和未标记数据生成可靠的伪标记。对于逐帧模型,使用DCASE 2021 Task 4的ICT-TOSHIBA系统。实验结果表明,该混合系统在没有外部数据的验证集上取得了0.420的psds 1和0.783的psds 2,明显优于单个模型。该代码可在www.example.com上获得https://github.com/965694547/Hybrid-system-of-frame-wise-model-and-SEDT。
摘要:In this paper, we describe in detail our system for DCASE 2022 Task4. The system combines two considerably different models: an end-to-end Sound Event Detection Transformer (SEDT) and a frame-wise model, Metric Learning and Focal Loss CNN (MLFL-CNN). The former is an event-wise model which learns event-level representations and predicts sound event categories and boundaries directly, while the latter is based on the widely adopted frame-classification scheme, under which each frame is classified into event categories and event boundaries are obtained by post-processing such as thresholding and smoothing. For SEDT, self-supervised pre-training using unlabeled data is applied, and semi-supervised learning is adopted by using an online teacher, which is updated from the student model using the Exponential Moving Average (EMA) strategy and generates reliable pseudo labels for weakly-labeled and unlabeled data. For the frame-wise model, the ICT-TOSHIBA system of DCASE 2021 Task 4 is used. Experimental results show that the hybrid system considerably outperforms either individual model and achieves psds1 of 0.420 and psds2 of 0.783 on the validation set without external data. The code is available at https://github.com/965694547/Hybrid-system-of-frame-wise-model-and-SEDT.


【7】 SVLDL: Improved Speaker Age Estimation Using Selective Variance Label  Distribution Learning

标题:基于选择性方差标签分布学习的改进说话人年龄估计

链接:https://arxiv.org/abs/2210.09524

作者:Zuheng Kang,Jianzong Wang,Junqing Peng,Jing Xiao
机构:Ping An Technology (Shenzhen) Co., Ltd.
备注:Accepted by SLT 2022. The 2022 IEEE Spoken Language Technology Workshop (SLT 2022)
摘要:从一段讲话中估计年龄是一个经典而富有挑战性的话题。尽管标签分布学习(LDL)可以很好地表示相邻的不可区分的年龄,但是对于每个话语的年龄估计的不确定性因人而异,即,年龄分布的方差是不同的。针对这一问题,提出了一种选择性方差标签分布学习(SVLDL)方法来适应不同年龄分布的方差.此外,该模型采用WavLM作为语音特征提取器,并增加了性别识别的辅助任务,进一步提高了性能。在损失函数上应用两个技巧以增强年龄估计的鲁棒性并改善拟合的年龄分布的质量。大量实验表明,该模型在NIST SRE 08 -10和真实数据集上的各方面性能均达到了最佳水平。
摘要:Estimating age from a single speech is a classic and challenging topic. Although Label Distribution Learning (LDL) can represent adjacent indistinguishable ages well, the uncertainty of the age estimate for each utterance varies from person to person, i.e., the variance of the age distribution is different. To address this issue, we propose selective variance label distribution learning (SVLDL) method to adapt the variance of different age distributions. Furthermore, the model uses WavLM as the speech feature extractor and adds the auxiliary task of gender recognition to further improve the performance. Two tricks are applied on the loss function to enhance the robustness of the age estimation and improve the quality of the fitted age distribution. Extensive experiments show that the model achieves state-of-the-art performance on all aspects of the NIST SRE08-10 and a real-world datasets.


【8】 Personalization of CTC Speech Recognition Models

标题:CTC语音识别模型的个性化

链接:https://arxiv.org/abs/2210.09510

作者:Saket Dingliwal,Monica Sunkara,Srikanth Ronanki,Jeff Farris,Katrin Kirchhoff,Sravan Bodapati
机构:Amazon AWS AI
备注:To appear in SLT 2022
摘要:近年来,基于CTC-Attention Loss联合训练的端到端语音识别模型得到了广泛的应用。在这些模型中,非自回归CTC解码器由于其速度和简单性而经常在推断时使用。然而,这种模型很难个性化,因为它们的条件独立性假设防止来自先前时间步的输出令牌影响未来预测。为了解决这个问题,我们提出了一种新颖的双向方法,首先在预定义的罕见长尾和词汇表外(OOV)单词列表上偏置编码器的注意力,然后在解码过程中使用动态提升和音素对齐网络进一步偏置子词预测。我们在开源VoxPopuli和内部医疗数据集上评估了我们的方法,以显示与强CTC基线相比,特定领域罕见词的F1评分提高了60%。
摘要:End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregressive CTC decoder is often used at inference time due to its speed and simplicity. However, such models are hard to personalize because of their conditional independence assumption that prevents output tokens from previous time steps to influence future predictions. To tackle this, we propose a novel two-way approach that first biases the encoder with attention over a predefined list of rare long-tail and out-of-vocabulary (OOV) words and then uses dynamic boosting and phone alignment network during decoding to further bias the subword predictions. We evaluate our approach on open-source VoxPopuli and in-house medical datasets to showcase a 60% improvement in F1 score on domain-specific rare words over a strong CTC baseline.


【9】 Affective Idiosyncratic Responses to Music

标题:对音乐的情感特质反应

链接:https://arxiv.org/abs/2210.09396

作者:Sky CH-Wang,Evan Li,Oliver Li,Smaranda Muresan,Zhou Yu
机构:◦Department of Computer Science, Columbia University, •Data Science Institute, Columbia University备注:EMNLP 2022 Main Conference; see Github this https URL
摘要:对音乐的情感反应是高度个人化的。尽管人们一致认为特质因素在调节听众对音乐的情感反应方面发挥着关键作用,但事实证明,精确测量这些变量的边际效应是一项挑战。为了弥补这一差距,我们开发了一种计算方法,测量了中国社交音乐平台上超过4.03亿条听众评论对音乐的情感反应。基于音乐心理学的系统和准因果分析研究,我们测试了驱动听众情感反应的音乐、抒情、语境、人口统计学和心理健康效应。最后,受"wng-y\'i-y\' un“社会现象的启发,我们确定了平台用户自我披露的影响因素,他们获得的社会支持,以及披露者用户活动的显著差异。
摘要:Affective responses to music are highly personal. Despite consensus that idiosyncratic factors play a key role in regulating how listeners emotionally respond to music, precisely measuring the marginal effects of these variables has proved challenging. To address this gap, we develop computational methods to measure affective responses to music from over 403M listener comments on a Chinese social music platform. Building on studies from music psychology in systematic and quasi-causal analyses, we test for musical, lyrical, contextual, demographic, and mental health effects that drive listener affective responses. Finally, motivated by the social phenomenon known as w\v{a}ng-y\`i-y\'un, we identify influencing factors of platform user self-disclosures, the social support they receive, and notable differences in discloser user activity.


【10】 Risk of re-identification for shared clinical speech recordings

标题:共享临床语音记录的重新识别风险

链接:https://arxiv.org/abs/2210.09975

作者:Daniela A. Wiepert,Bradley A. Malin,Joseph R. Duffy,Rene L. Utianski,John L. Stricker,David T. Jones,Hugo Botha
机构:Department of Neurology, Rochester, MN, USA, Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, USA, Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN, USA
备注:24 pages, 6 figures
摘要:在医疗保健中利用基于语音的工具需要大型的、精心设计的数据集。这些产品的生产成本很高,导致人们对数据共享的兴趣增加。由于语音可以潜在地标识说话者(即,声纹),共享录音会引起隐私问题。我们使用最先进的说话人识别系统,在不参考人口统计学或元数据的情况下,研究了语音录音的重新识别风险。我们证明了风险与对手必须考虑的比较次数成反比,即,搜索空间。对于较小的搜索空间,风险很高,但随着搜索空间的增长,风险会降低(对于$〈1*10^{6}$的比较,$precision〉0.85$;对于$〉3*10^{8}$的比较,$precision〈0.5$)。接下来,我们将说明语音记录的性质会影响重新识别风险,对于非连接语音(例如,元音延长)更难识别。我们的研究结果表明,说话人识别系统可以在特定的环境中用于重新识别参与者,但在实践中,重新识别的风险似乎很低。
摘要:Large, curated datasets are required to leverage speech-based tools in healthcare. These are costly to produce, resulting in increased interest in data sharing. As speech can potentially identify speakers (i.e., voiceprints), sharing recordings raises privacy concerns. We examine the re-identification risk for speech recordings, without reference to demographic or metadata, using a state-of-the-art speaker recognition system. We demonstrate that the risk is inversely related to the number of comparisons an adversary must consider, i.e., the search space. Risk is high for a small search space but drops as the search space grows ($precision >0.85$ for $<1*10^{6}$ comparisons, $precision <0.5$ for $>3*10^{6}$ comparisons). Next, we show that the nature of a speech recording influences re-identification risk, with non-connected speech (e.g., vowel prolongation) being harder to identify. Our findings suggest that speaker recognition systems can be used to re-identify participants in specific circumstances, but in practice, the re-identification risk appears low.


【11】 Extracting speaker and emotion information from self-supervised speech  models via channel-wise correlations

标题:基于通道相关的自监督语音模型中的说话人和情感信息提取

链接:https://arxiv.org/abs/2210.09513

作者:Themos Stafylakis,Ladislav Mosner,Sofoklis Kakouros,Oldrich Plchot,Lukas Burget,Jan Cernocky
机构:Omilia - Conversational Intelligence, Athens, Greece, University of Helsinki, Finland
备注:Accepted at IEEE-SLT 2022
摘要:从大量未标记数据中对语音表示的自监督学习已经使得能够在若干语音处理任务中获得最新的结果。跨时间聚集这些语音表示通常是通过使用描述性统计量来实现的,并且具体地,使用表示系数的一阶和二阶统计量。在这篇论文中,我们研究了一种从自监督训练模型中提取说话人和情感信息的替代方法,该方法基于表征系数之间的相关性-相关池。我们显示了平均合并的改进,并且当合并方法通过融合进行组合时,进一步获得了收益。该代码可在www.example.com上获得github.com/Lamomal/s3prl_correlation。
摘要:Self-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typically approached by using descriptive statistics, and in particular, using the first- and second-order statistics of representation coefficients. In this paper, we examine an alternative way of extracting speaker and emotion information from self-supervised trained models, based on the correlations between the coefficients of the representations - correlation pooling. We show improvements over mean pooling and further gains when the pooling methods are combined via fusion. The code is available at github.com/Lamomal/s3prl_correlation.


【12】 TorchDIVA: An Extensible Computational Model of Speech Production built  on an Open-Source Machine Learning Library

标题:TorchDIVA:一种基于开源机器学习库的可扩展语音生成计算模型

链接:https://arxiv.org/abs/2210.09334

作者:Sean Kinahan,Julie Liss,Visar Berisha
机构:Arizona State University
摘要:DIVA模型是言语运动控制的计算模型,其将负责言语产生的大脑区域的模拟与人类声道的模型相结合。该模型目前在Matlab Simulink中实现;然而,这并不理想,因为语音技术研究中的大多数开发都是用Python完成的。这意味着Python生态系统中有大量的机器学习工具可以免费获得,但这些工具无法轻松地与DIVA集成。我们展示了TorchDIVA,它是在Python中使用PyTorch张量对DIVA进行的完全重建。DIVA源代码直接从Matlab翻译为Python,内置Simulink信号模块从头开始实现。实施后,通过系统的逐块验证评估每个模块的准确性。显示TorchDIVA模型产生的输出与原始DIVA模型的输出非常匹配,两者之间的差异可忽略不计。我们还提供了一个TorchDIVA作为研究平台的可扩展性示例。TorchDIVA中的语音质量增强是通过与称为DiffWave的现有PyTorch生成式声码器集成来实现的。在人类语音波形上训练改进的DiffWave Mel频谱上采样器,并在TorchDIVA语音产生上调节。结果表明,与基线相比,DiffWave增强的输出中的语音质量度量得到改善。这种增强在原始Matlab实现中很难或不可能实现。这一概念验证证明了TorchDIVA将为研究界带来的价值。研究人员可从以下网址下载新的实施方案:https://github.com/skinahan/DIVA_PyTorch
摘要:The DIVA model is a computational model of speech motor control that combines a simulation of the brain regions responsible for speech production with a model of the human vocal tract. The model is currently implemented in Matlab Simulink; however, this is less than ideal as most of the development in speech technology research is done in Python. This means there is a wealth of machine learning tools which are freely available in the Python ecosystem that cannot be easily integrated with DIVA. We present TorchDIVA, a full rebuild of DIVA in Python using PyTorch tensors. DIVA source code was directly translated from Matlab to Python, and built-in Simulink signal blocks were implemented from scratch. After implementation, the accuracy of each module was evaluated via systematic block-by-block validation. The TorchDIVA model is shown to produce outputs that closely match those of the original DIVA model, with a negligible difference between the two. We additionally present an example of the extensibility of TorchDIVA as a research platform. Speech quality enhancement in TorchDIVA is achieved through an integration with an existing PyTorch generative vocoder called DiffWave. A modified DiffWave mel-spectrum upsampler was trained on human speech waveforms and conditioned on the TorchDIVA speech production. The results indicate improved speech quality metrics in the DiffWave-enhanced output as compared to the baseline. This enhancement would have been difficult or impossible to accomplish in the original Matlab implementation. This proof-of-concept demonstrates the value TorchDIVA will bring to the research community. Researchers can download the new implementation at: https://github.com/skinahan/DIVA_PyTorch


eess.AS音频处理

【1】 Risk of re-identification for shared clinical speech recordings

标题:共享临床语音记录的重新识别风险

链接:https://arxiv.org/abs/2210.09975

* 与cs.SD语音【10】为同一篇

作者:Daniela A. Wiepert,Bradley A. Malin,Joseph R. Duffy,Rene L. Utianski,John L. Stricker,David T. Jones,Hugo Botha
机构:Department of Neurology, Rochester, MN, USA, Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, USA, Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN, USA
备注:24 pages, 6 figures
摘要:在医疗保健中利用基于语音的工具需要大型的、精心设计的数据集。这些产品的生产成本很高,导致人们对数据共享的兴趣增加。由于语音可以潜在地标识说话者(即,声纹),共享录音会引起隐私问题。我们使用最先进的说话人识别系统,在不参考人口统计学或元数据的情况下,研究了语音录音的重新识别风险。我们证明了风险与对手必须考虑的比较次数成反比,即,搜索空间。对于较小的搜索空间,风险很高,但随着搜索空间的增长,风险会降低(对于$〈1*10^{6}$的比较,$precision〉0.85$;对于$〉3*10^{8}$的比较,$precision〈0.5$)。接下来,我们将说明语音记录的性质会影响重新识别风险,对于非连接语音(例如,元音延长)更难识别。我们的研究结果表明,说话人识别系统可以在特定的环境中用于重新识别参与者,但在实践中,重新识别的风险似乎很低。
摘要:Large, curated datasets are required to leverage speech-based tools in healthcare. These are costly to produce, resulting in increased interest in data sharing. As speech can potentially identify speakers (i.e., voiceprints), sharing recordings raises privacy concerns. We examine the re-identification risk for speech recordings, without reference to demographic or metadata, using a state-of-the-art speaker recognition system. We demonstrate that the risk is inversely related to the number of comparisons an adversary must consider, i.e., the search space. Risk is high for a small search space but drops as the search space grows ($precision >0.85$ for $<1*10^{6}$ comparisons, $precision <0.5$ for $>3*10^{6}$ comparisons). Next, we show that the nature of a speech recording influences re-identification risk, with non-connected speech (e.g., vowel prolongation) being harder to identify. Our findings suggest that speaker recognition systems can be used to re-identify participants in specific circumstances, but in practice, the re-identification risk appears low.


【2】 Extracting speaker and emotion information from self-supervised speech  models via channel-wise correlations

标题:基于通道相关的自监督语音模型中的说话人和情感信息提取

链接:https://arxiv.org/abs/2210.09513

* 与cs.SD语音【11】为同一篇

作者:Themos Stafylakis,Ladislav Mosner,Sofoklis Kakouros,Oldrich Plchot,Lukas Burget,Jan Cernocky
机构:Omilia - Conversational Intelligence, Athens, Greece, University of Helsinki, Finland
备注:Accepted at IEEE-SLT 2022
摘要:从大量未标记数据中对语音表示的自监督学习已经使得能够在若干语音处理任务中获得最新的结果。跨时间聚集这些语音表示通常是通过使用描述性统计量来实现的,并且具体地,使用表示系数的一阶和二阶统计量。在这篇论文中,我们研究了一种从自监督训练模型中提取说话人和情感信息的替代方法,该方法基于表征系数之间的相关性-相关池。我们显示了平均合并的改进,并且当合并方法通过融合进行组合时,进一步获得了收益。该代码可在www.example.com上获得github.com/Lamomal/s3prl_correlation。
摘要:Self-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typically approached by using descriptive statistics, and in particular, using the first- and second-order statistics of representation coefficients. In this paper, we examine an alternative way of extracting speaker and emotion information from self-supervised trained models, based on the correlations between the coefficients of the representations - correlation pooling. We show improvements over mean pooling and further gains when the pooling methods are combined via fusion. The code is available at github.com/Lamomal/s3prl_correlation.


【3】 TorchDIVA: An Extensible Computational Model of Speech Production built  on an Open-Source Machine Learning Library

标题:TorchDIVA:一种基于开源机器学习库的可扩展语音生成计算模型

链接:https://arxiv.org/abs/2210.09334

* 与cs.SD语音【12】为同一篇

作者:Sean Kinahan,Julie Liss,Visar Berisha
机构:Arizona State University
摘要:DIVA模型是言语运动控制的计算模型,其将负责言语产生的大脑区域的模拟与人类声道的模型相结合。该模型目前在Matlab Simulink中实现;然而,这并不理想,因为语音技术研究中的大多数开发都是用Python完成的。这意味着Python生态系统中有大量的机器学习工具可以免费获得,但这些工具无法轻松地与DIVA集成。我们展示了TorchDIVA,它是在Python中使用PyTorch张量对DIVA进行的完全重建。DIVA源代码直接从Matlab翻译为Python,内置Simulink信号模块从头开始实现。实施后,通过系统的逐块验证评估每个模块的准确性。显示TorchDIVA模型产生的输出与原始DIVA模型的输出非常匹配,两者之间的差异可忽略不计。我们还提供了一个TorchDIVA作为研究平台的可扩展性示例。TorchDIVA中的语音质量增强是通过与称为DiffWave的现有PyTorch生成式声码器集成来实现的。在人类语音波形上训练改进的DiffWave Mel频谱上采样器,并在TorchDIVA语音产生上调节。结果表明,与基线相比,DiffWave增强的输出中的语音质量度量得到改善。这种增强在原始Matlab实现中很难或不可能实现。这一概念验证证明了TorchDIVA将为研究界带来的价值。研究人员可从以下网址下载新的实施方案:https://github.com/skinahan/DIVA_PyTorch
摘要:The DIVA model is a computational model of speech motor control that combines a simulation of the brain regions responsible for speech production with a model of the human vocal tract. The model is currently implemented in Matlab Simulink; however, this is less than ideal as most of the development in speech technology research is done in Python. This means there is a wealth of machine learning tools which are freely available in the Python ecosystem that cannot be easily integrated with DIVA. We present TorchDIVA, a full rebuild of DIVA in Python using PyTorch tensors. DIVA source code was directly translated from Matlab to Python, and built-in Simulink signal blocks were implemented from scratch. After implementation, the accuracy of each module was evaluated via systematic block-by-block validation. The TorchDIVA model is shown to produce outputs that closely match those of the original DIVA model, with a negligible difference between the two. We additionally present an example of the extensibility of TorchDIVA as a research platform. Speech quality enhancement in TorchDIVA is achieved through an integration with an existing PyTorch generative vocoder called DiffWave. A modified DiffWave mel-spectrum upsampler was trained on human speech waveforms and conditioned on the TorchDIVA speech production. The results indicate improved speech quality metrics in the DiffWave-enhanced output as compared to the baseline. This enhancement would have been difficult or impossible to accomplish in the original Matlab implementation. This proof-of-concept demonstrates the value TorchDIVA will bring to the research community. Researchers can download the new implementation at: https://github.com/skinahan/DIVA_PyTorch


【4】 Maestro-U: Leveraging joint speech-text representation learning for zero  supervised speech ASR

标题:Maestro-U:利用语音-文本联合表征学习实现零监督语音ASR

链接:https://arxiv.org/abs/2210.10027

* 与cs.SD语音【1】为同一篇

作者:Zhehuai Chen,Ankur Bapna,Andrew Rosenberg,Yu Zhang,Bhuvana Ramabhadran,Pedro Moreno,Nanxin Chen
机构:Google, Inc.
备注:Accepted by SLT 2022
摘要:训练现有技术的自动语音识别(ASR)模型通常需要大量的转录语音。在这项工作中,我们证明了一个模态匹配的联合语音和文本模型可以用来训练一个大规模的多语言ASR模型,而不需要任何监督(手动转录)的语音。本文探讨了在大规模多语言、零监督语音、真实世界设置中使用联合学习的语音和文本表示,以扩展ASR所覆盖的语言集,仅使用目标语言中的未标记语音和文本。使用FLEURS数据集,我们将任务定义为覆盖$102$种语言,其中转录语音在$52$种语言中可用,并可用于提高剩余$50$中的端到端ASR质量。首先,我们证明了通过将语音表示与字节级文本表示相结合,并使用语言嵌入,我们可以将无监督语音的语言的字符错误率(CER)从64.8%显著降低到30.8%,相对降低了53%。第二,使用南亚语言的子集,我们表明Maestro-U可以促进来自有监督语音的语言的知识转移,即使有限制到没有字形重叠。总体而言,Maestro-U将与Oracle的性能差距缩小了68.5%,并将19种语言的CER降低到15%以下。
摘要:Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model can be leveraged to train a massively multilingual ASR model without any supervised (manually transcribed) speech for some languages. This paper explores the use of jointly learnt speech and text representations in a massively multilingual, zero supervised speech, real-world setting to expand the set of languages covered by ASR with only unlabeled speech and text in the target languages. Using the FLEURS dataset, we define the task to cover $102$ languages, where transcribed speech is available in $52$ of these languages and can be used to improve end-to-end ASR quality on the remaining $50$. First, we show that by combining speech representations with byte-level text representations and use of language embeddings, we can dramatically reduce the Character Error Rate (CER) on languages with no supervised speech from 64.8\% to 30.8\%, a relative reduction of 53\%. Second, using a subset of South Asian languages we show that Maestro-U can promote knowledge transfer from languages with supervised speech even when there is limited to no graphemic overlap. Overall, Maestro-U closes the gap to oracle performance by 68.5\% relative and reduces the CER of 19 languages below 15\%.


【5】 HMM vs. CTC for Automatic Speech Recognition: Comparison Based on  Full-Sum Training from Scratch

标题:基于全和训练的HMM与CTC在自动语音识别中的比较

链接:https://arxiv.org/abs/2210.09951

* 与cs.SD语音【2】为同一篇

作者:Tina Raissi,Wei Zhou,Simon Berger,Ralf Schlüter,Hermann Ney
机构:Human Language Technology and Pattern Recognition Group, RWTH Aachen University, Germany, AppTek GmbH, Aachen, Germany
备注:Accepted for Presentation at IEEE SLT 2022
摘要:在这项工作中,我们从零开始比较隐马尔可夫模型(HMM)的序列级交叉熵(全和)训练和用于自动语音识别(ASR)的连接主义时态分类(CTC)拓扑。除了准确性之外,我们还分析了它们在语音信号和转录之间生成高质量时间对准的能力,这对许多后续应用至关重要。此外,我们提出了几种方法来改善从头开始的全和训练的收敛性,通过解决对准建模问题。在Switchboard和LibriSpeech语料库上对CTC、有转移概率和无转移概率的后验HMM以及标准混合HMM进行了系统比较。我们还提供了维特比强制对齐和Baum-Welch全和占用概率的详细分析。
摘要:In this work, we compare from-scratch sequence-level cross-entropy (full-sum) training of Hidden Markov Model (HMM) and Connectionist Temporal Classification (CTC) topologies for automatic speech recognition (ASR). Besides accuracy, we further analyze their capability for generating high-quality time alignment between the speech signal and the transcription, which can be crucial for many subsequent applications. Moreover, we propose several methods to improve convergence of from-scratch full-sum training by addressing the alignment modeling issue. Systematic comparison is conducted on both Switchboard and LibriSpeech corpora across CTC, posterior HMM with and w/o transition probabilities, and standard hybrid HMM. We also provide a detailed analysis of both Viterbi forced-alignment and Baum-Welch full-sum occupation probabilities.


【6】 Mid-attribute speaker generation using optimal-transport-based  interpolation of Gaussian mixture models

标题:基于最优传输的混合高斯模型内插中属性说话人生成

链接:https://arxiv.org/abs/2210.09916

* 与cs.SD语音【3】为同一篇

作者:Aya Watanabe,Shinnosuke Takamichi,Yuki Saito,Detai Xin,Hiroshi Saruwatar
i机构:The University of Tokyo, Japan.
备注:Submitted to ICASSP 2023. Demo: this https URL
摘要:本文提出了一种在“说话人生成”中综合多个说话人的属性并使其语音特征多样化的方法,“说话人生成”是一个新兴的任务,其目的是合成不存在的说话人的自然发声的语音。传统的基于TacoSpawn的说话人生成方法是利用高斯混合模型(Gaussian mixture models,GMM)来描述说话人嵌入的分布,并考虑说话人的属性。尽管该方法使得能够从知晓说话者属性的GMM中对各种说话者进行采样,但是还不清楚所学习的分布是否能够表示具有中间属性(即,中间属性)。为此,我们提出了一种基于最优传输的方法,该方法对学习的GMM进行插值,以生成具有中间属性(例如,性别中立)的声音。我们通过实验验证了我们的方法,并评估了合成语音的自然度和两个说话者属性的可控性:性别和语言流利性。实验结果表明,该方法能够在不降低语音自然度的前提下,通过连续的标量值来控制说话人的属性。
摘要:In this paper, we propose a method for intermediating multiple speakers' attributes and diversifying their voice characteristics in ``speaker generation,'' an emerging task that aims to synthesize a nonexistent speaker's naturally sounding voice. The conventional TacoSpawn-based speaker generation method represents the distributions of speaker embeddings by Gaussian mixture models (GMMs) conditioned with speaker attributes. Although this method enables the sampling of various speakers from the speaker-attribute-aware GMMs, it is not yet clear whether the learned distributions can represent speakers with an intermediate attribute (i.e., mid-attribute). To this end, we propose an optimal-transport-based method that interpolates the learned GMMs to generate nonexistent speakers with mid-attribute (e.g., gender-neutral) voices. We empirically validate our method and evaluate the naturalness of synthetic speech and the controllability of two speaker attributes: gender and language fluency. The evaluation results show that our method can control the generated speakers' attributes by a continuous scalar value without statistically significant degradation of speech naturalness.


【7】 Spontaneous speech synthesis with linguistic-speech consistency training  using pseudo-filled pauses

标题:利用伪填充停顿进行语言-语音一致性训练的自发语音合成

链接:https://arxiv.org/abs/2210.09815

* 与cs.SD语音【4】为同一篇

作者:Yuta Matsunaga,Takaaki Saeki,Shinnosuke Takamichi,Hiroshi Saruwatari
机构:Graduate School of Information Science and Technology, The University of Tokyo, Japan.
备注:Submitted to ICASSP 2023
摘要:本文提出了一种自发语音合成模型的训练方法,保证了合成语音各语言成分的一致性。个性化的自发语音合成旨在再现不流畅的个性,例如填充的停顿。我们的先验模型包括填充停顿预测模型,并且从没有填充停顿的文本合成包括填充停顿的语音。然而,插入填充的停顿降低了合成语音的语言部分的质量。这可能是因为在训练和推理之间填充停顿插入的倾向不同,并且合成模型不能表示在推理中填充停顿和周围音素之间的连接。因此,我们开发了一种语言-语音一致性训练,它保证了有和没有填充停顿的合成语音的语言部分的一致性。所提出的一致性训练不仅利用地面真实填充的停顿,而且利用伪停顿。实验结果表明,该方法提高了合成语音的自然度,提高了整个语音的自然度。
摘要:We propose a training method for spontaneous speech synthesis models that guarantees the consistency of linguistic parts of synthesized speech. Personalized spontaneous speech synthesis aims to reproduce the individuality of disfluency, such as filled pauses. Our prior model includes a filled-pause prediction model and synthesizes filled-pause-included speech from text without filled pauses. However, inserting the filled pauses degrades the quality of the linguistic parts of the synthesized speech. This might be because filled-pause insertion tendencies differ between training and inference, and the synthesis model cannot represent connections between filled pauses and surrounding phonemes in inference. We, therefore, developed a linguistic-speech consistency training that guarantees the consistency of linguistic parts of synthetic speech with and without filled pauses. The proposed consistency training utilizes not only ground-truth-filled pauses but also pseudo ones. Our experiments demonstrate that this method improves the naturalness of the synthetic linguistic speech and the entire predicted-filled-pause-included synthetic speech.


【8】 Discrete Cross-Modal Alignment Enables Zero-Shot Speech Translation

标题:离散跨模式对齐实现零发声语音翻译

链接:https://arxiv.org/abs/2210.09556

* 与cs.SD语音【5】为同一篇

作者:Chen Wang,Yuchen Liu,Boxing Chen,Jiajun Zhang,Wei Luo,Zhongqiang Huang,Chengqing Zong机构:National Laboratory of Pattern Recognition, Institute of Automation, CAS, Beijing, China,  Machine Intelligence Technology Lab, Alibaba DAMO Academy
备注:Accepted by the main conference of EMNLP 2022
摘要:端到端语音翻译的目的是在不产生中间转录的情况下将源语言语音翻译成目标语言文本。然而,端到端方法的训练依赖于并行ST数据,而并行ST数据难以获得且昂贵。幸运的是,用于自动语音识别(ASR)和机器翻译(MT)的监督数据通常更容易获得,这使得zero-shot语音翻译成为一个潜在的方向。现有的zero-shot方法无法将语音和文本两种模态对齐到共享的语义空间中,导致性能比有监督的ST方法差得多。为了实现zero-shot ST,提出了一种新的离散跨模态对齐(DCMA)方法,该方法利用一个共享的离散词汇空间来容纳和匹配语音和文本模态。具体地说,我们引入了矢量量化模块,将语音和文本的连续表示离散化为有限的虚拟令牌集,并使用ASR数据将相应的语音和文本映射到共享码本中的同一虚拟令牌。这样,源语言语音可以嵌入到与源语言文本相同的语义空间中,然后可以利用MT模块将源语言文本转换为目标语言文本。在多个语言对上的实验表明,本文的zero-shot ST方法显著提高了SOTA,甚至达到了与强监督ST基线相当的性能.
摘要:End-to-end Speech Translation (ST) aims at translating the source language speech into target language text without generating the intermediate transcriptions. However, the training of end-to-end methods relies on parallel ST data, which are difficult and expensive to obtain. Fortunately, the supervised data for automatic speech recognition (ASR) and machine translation (MT) are usually more accessible, making zero-shot speech translation a potential direction. Existing zero-shot methods fail to align the two modalities of speech and text into a shared semantic space, resulting in much worse performance compared to the supervised ST methods. In order to enable zero-shot ST, we propose a novel Discrete Cross-Modal Alignment (DCMA) method that employs a shared discrete vocabulary space to accommodate and match both modalities of speech and text. Specifically, we introduce a vector quantization module to discretize the continuous representations of speech and text into a finite set of virtual tokens, and use ASR data to map corresponding speech and text to the same virtual token in a shared codebook. This way, source language speech can be embedded in the same semantic space as the source language text, which can be then transformed into target language text with an MT module. Experiments on multiple language pairs demonstrate that our zero-shot ST method significantly improves the SOTA, and even performers on par with the strong supervised ST baselines.


【9】 A Hybrid System of Sound Event Detection Transformer and Frame-wise  Model for DCASE 2022 Task 4

标题:用于DCASE 2022任务4的声事件检测Transformer和框架模型的混合系统

链接:https://arxiv.org/abs/2210.09529

* 与cs.SD语音【6】为同一篇

作者:Yiming Li,Zhifang Guo,Zhirong Ye,Xiangdong Wang,Hong Liu,Yueliang Qian,Rui Tao,Long Yan,Kazushige Ouchi
机构:Beijing Key Laboratory of Mobile Computing and Pervasive Device, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China,  University of Chinese Academy of Sciences, Beijing, China,  Toshiba China R&D Center, Beijing, China
备注:5 pages, 2 figures, accepted for publication in DCASE2022 Workshop
摘要:本文详细介绍了DCASE 2022 Task 4, 0系统。该系统结合了两种截然不同的模式:端到端声音事件检测Transformer(SEDT)和逐帧模型、度量学习和聚焦损失CNN(MLFL-CNN)。前者是一种基于事件的模型,学习事件级表示,直接预测声音事件类别和边界;后者基于广泛采用的帧分类方案,将每个帧分类为事件类别,并通过阈值化和平滑等后处理获得事件边界。对于SEDT,使用未标记数据进行自监督预训练,并使用在线教师进行半监督学习,在线教师使用指数移动平均(EMA)策略从学生模型更新,并为弱标记和未标记数据生成可靠的伪标记。对于逐帧模型,使用DCASE 2021 Task 4的ICT-TOSHIBA系统。实验结果表明,该混合系统在没有外部数据的验证集上取得了0.420的psds 1和0.783的psds 2,明显优于单个模型。该代码可在www.example.com上获得https://github.com/965694547/Hybrid-system-of-frame-wise-model-and-SEDT。
摘要:In this paper, we describe in detail our system for DCASE 2022 Task4. The system combines two considerably different models: an end-to-end Sound Event Detection Transformer (SEDT) and a frame-wise model, Metric Learning and Focal Loss CNN (MLFL-CNN). The former is an event-wise model which learns event-level representations and predicts sound event categories and boundaries directly, while the latter is based on the widely adopted frame-classification scheme, under which each frame is classified into event categories and event boundaries are obtained by post-processing such as thresholding and smoothing. For SEDT, self-supervised pre-training using unlabeled data is applied, and semi-supervised learning is adopted by using an online teacher, which is updated from the student model using the Exponential Moving Average (EMA) strategy and generates reliable pseudo labels for weakly-labeled and unlabeled data. For the frame-wise model, the ICT-TOSHIBA system of DCASE 2021 Task 4 is used. Experimental results show that the hybrid system considerably outperforms either individual model and achieves psds1 of 0.420 and psds2 of 0.783 on the validation set without external data. The code is available at https://github.com/965694547/Hybrid-system-of-frame-wise-model-and-SEDT.


【10】 SVLDL: Improved Speaker Age Estimation Using Selective Variance Label  Distribution Learning

标题:基于选择性方差标签分布学习的改进说话人年龄估计

链接:https://arxiv.org/abs/2210.09524

* 与cs.SD语音【7】为同一篇

作者:Zuheng Kang,Jianzong Wang,Junqing Peng,Jing Xiao
机构:Ping An Technology (Shenzhen) Co., Ltd.
备注:Accepted by SLT 2022. The 2022 IEEE Spoken Language Technology Workshop (SLT 2022)
摘要:从一段讲话中估计年龄是一个经典而富有挑战性的话题。尽管标签分布学习(LDL)可以很好地表示相邻的不可区分的年龄,但是对于每个话语的年龄估计的不确定性因人而异,即,年龄分布的方差是不同的。针对这一问题,提出了一种选择性方差标签分布学习(SVLDL)方法来适应不同年龄分布的方差.此外,该模型采用WavLM作为语音特征提取器,并增加了性别识别的辅助任务,进一步提高了性能。在损失函数上应用两个技巧以增强年龄估计的鲁棒性并改善拟合的年龄分布的质量。大量实验表明,该模型在NIST SRE 08 -10和真实数据集上的各方面性能均达到了最佳水平。
摘要:Estimating age from a single speech is a classic and challenging topic. Although Label Distribution Learning (LDL) can represent adjacent indistinguishable ages well, the uncertainty of the age estimate for each utterance varies from person to person, i.e., the variance of the age distribution is different. To address this issue, we propose selective variance label distribution learning (SVLDL) method to adapt the variance of different age distributions. Furthermore, the model uses WavLM as the speech feature extractor and adds the auxiliary task of gender recognition to further improve the performance. Two tricks are applied on the loss function to enhance the robustness of the age estimation and improve the quality of the fitted age distribution. Extensive experiments show that the model achieves state-of-the-art performance on all aspects of the NIST SRE08-10 and a real-world datasets.


【11】 Personalization of CTC Speech Recognition Models

标题:CTC语音识别模型的个性化

链接:https://arxiv.org/abs/2210.09510

* 与cs.SD语音【8】为同一篇

作者:Saket Dingliwal,Monica Sunkara,Srikanth Ronanki,Jeff Farris,Katrin Kirchhoff,Sravan Bodapati
机构:Amazon AWS AI
备注:To appear in SLT 2022
摘要:近年来,基于CTC-Attention Loss联合训练的端到端语音识别模型得到了广泛的应用。在这些模型中,非自回归CTC解码器由于其速度和简单性而经常在推断时使用。然而,这种模型很难个性化,因为它们的条件独立性假设防止来自先前时间步的输出令牌影响未来预测。为了解决这个问题,我们提出了一种新颖的双向方法,首先在预定义的罕见长尾和词汇表外(OOV)单词列表上偏置编码器的注意力,然后在解码过程中使用动态提升和音素对齐网络进一步偏置子词预测。我们在开源VoxPopuli和内部医疗数据集上评估了我们的方法,以显示与强CTC基线相比,特定领域罕见词的F1评分提高了60%。
摘要:End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregressive CTC decoder is often used at inference time due to its speed and simplicity. However, such models are hard to personalize because of their conditional independence assumption that prevents output tokens from previous time steps to influence future predictions. To tackle this, we propose a novel two-way approach that first biases the encoder with attention over a predefined list of rare long-tail and out-of-vocabulary (OOV) words and then uses dynamic boosting and phone alignment network during decoding to further bias the subword predictions. We evaluate our approach on open-source VoxPopuli and in-house medical datasets to showcase a 60% improvement in F1 score on domain-specific rare words over a strong CTC baseline.


【12】 Affective Idiosyncratic Responses to Music

标题:对音乐的情感特质反应

链接:https://arxiv.org/abs/2210.09396

* 与cs.SD语音【9】为同一篇

作者:Sky CH-Wang,Evan Li,Oliver Li,Smaranda Muresan,Zhou Yu
机构:◦Department of Computer Science, Columbia University, •Data Science Institute, Columbia University
备注:EMNLP 2022 Main Conference; see Github this https URL
摘要:对音乐的情感反应是高度个人化的。尽管人们一致认为特质因素在调节听众对音乐的情感反应方面发挥着关键作用,但事实证明,精确测量这些变量的边际效应是一项挑战。为了弥补这一差距,我们开发了一种计算方法,测量了中国社交音乐平台上超过4.03亿条听众评论对音乐的情感反应。基于音乐心理学的系统和准因果分析研究,我们测试了驱动听众情感反应的音乐、抒情、语境、人口统计学和心理健康效应。最后,受"wng-y\'i-y\' un“社会现象的启发,我们确定了平台用户自我披露的影响因素,他们获得的社会支持,以及披露者用户活动的显著差异。
摘要:Affective responses to music are highly personal. Despite consensus that idiosyncratic factors play a key role in regulating how listeners emotionally respond to music, precisely measuring the marginal effects of these variables has proved challenging. To address this gap, we develop computational methods to measure affective responses to music from over 403M listener comments on a Chinese social music platform. Building on studies from music psychology in systematic and quasi-causal analyses, we test for musical, lyrical, contextual, demographic, and mental health effects that drive listener affective responses. Finally, motivated by the social phenomenon known as w\v{a}ng-y\`i-y\'un, we identify influencing factors of platform user self-disclosures, the social support they receive, and notable differences in discloser user activity.


机器翻译,仅供参考