今天跟大家分享一篇语音相关的论文合集:cs.SD语音9篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily


cs.SD语音

【1】 Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild

标题:野外任意说话人的唇语合成

链接:https://arxiv.org/abs/2209.00642

作者:Sindhu B Hegde,K R Prajwal,Rudrabha Mukhopadhyay,Vinay P Namboodiri,C. V. Jawahar
机构:International Institute of Information, Technology, Hyderabad, India, University of Oxford, United Kingdom, University of Bath
备注:Accepted in ACM-MM 2022, 9 pages, 2 pages supplementary, 7 Figures
摘要:在这项工作中,我们解决了从无声的嘴唇视频生成任何说话者在野外的语音的问题。与之前的工作形成鲜明对比的是,我们的方法(i)不限于固定数量的说话者,(ii)不明确地对领域或词汇表施加约束,以及(iii)处理在野外录制的视频,而不是在实验室设置内。这项任务提出了一系列挑战,其中关键的一个挑战是,所需目标语音的许多特征,如声音、音高和语言内容,不能完全从无声面部视频中推断出来。为了处理这些随机变化,我们提出了一种新的VAE-GAN结构,该结构在变化中学习将嘴唇和语音序列关联起来。在指导训练过程的多个功能强大的识别器的帮助下,我们的生成器学会了合成任何人嘴唇运动的任何声音的语音序列。在多个数据集上的大量实验表明,我们的性能大大优于所有基线。此外,我们的网络可以在特定身份的视频上进行微调,以实现与单扬声器模型相当的性能,后者是在4\x $以上的数据上训练的。我们进行了大量消融研究,以分析我们架构的不同模块的影响。我们还在我们的网站上提供了演示视频,展示了几个定性结果以及代码和经过训练的模型:\url{http://cvit.iiiit.ac.in/research/projects/cvit-projects/lip-to-speech-synthesis}}(唇转语音合成技术)的研究成果
摘要:In this work, we address the problem of generating speech from silent lip videos for any speaker in the wild. In stark contrast to previous works, our method (i) is not restricted to a fixed number of speakers, (ii) does not explicitly impose constraints on the domain or the vocabulary and (iii) deals with videos that are recorded in the wild as opposed to within laboratory settings. The task presents a host of challenges, with the key one being that many features of the desired target speech, like voice, pitch and linguistic content, cannot be entirely inferred from the silent face video. In order to handle these stochastic variations, we propose a new VAE-GAN architecture that learns to associate the lip and speech sequences amidst the variations. With the help of multiple powerful discriminators that guide the training process, our generator learns to synthesize speech sequences in any voice for the lip movements of any person. Extensive experiments on multiple datasets show that we outperform all baselines by a large margin. Further, our network can be fine-tuned on videos of specific identities to achieve a performance comparable to single-speaker models that are trained on $4\times$ more data. We conduct numerous ablation studies to analyze the effect of different modules of our architecture. We also provide a demo video that demonstrates several qualitative results along with the code and trained models on our website: \url{http://cvit.iiit.ac.in/research/projects/cvit-projects/lip-to-speech-synthesis}}


【2】 AccoMontage2: A Complete Harmonization and Accompaniment Arrangement  System

标题:AccoMontage2:一个完整的和声和伴奏编排系统

链接:https://arxiv.org/abs/2209.00353

作者:Li Yi,Haochen Hu,Jingwei Zhao,Gus Xia
机构:Music X Lab, NYU Shanghai,  MBZUAI,  Institute of Data Science, NUS
备注:Accepted by ISMIR 2022
摘要:我们提出了AccoMontage2,一个能够基于主旋律进行完整长度的歌曲和声和伴奏安排的系统。本研究以AccoMontage为基础,以流行歌曲/民歌为研究对象,进行了基于模板的检索方法研究。这项研究的新颖之处有两个方面。首先,我们发明了一个协调模块(AccoMontage没有)。该模块通过优化和平衡三个损耗项来生成结构化和连贯的全长弦级数:对于音符不协和的微观级损失、对于短语模板匹配的中观级损失、以及对于整段连贯性的宏观级损失。其次,我们开发了一个图形用户界面,允许用户选择不同风格的和弦进行和钢琴织体。目前,和弦进行风格包括Pop、R&B和Dark,而钢琴纹理风格包括几个层次的发声密度和节奏复杂度。实验结果表明,本文的协调和排序结果均显著优于基线。最后,我们发布了AccoMontage2在线应用程序,以及组织和弦进行模板作为公共数据集。
摘要:We propose AccoMontage2, a system capable of doing full-length song harmonization and accompaniment arrangement based on a lead melody. Following AccoMontage, this study focuses on generating piano arrangements for popular/folk songs and it carries on the generalized template-based retrieval method. The novelties of this study are twofold. First, we invent a harmonization module (which AccoMontage does not have). This module generates structured and coherent full-length chord progression by optimizing and balancing three loss terms: a micro-level loss for note-wise dissonance, a meso-level loss for phrase-template matching, and a macro-level loss for full piece coherency. Second, we develop a graphical user interface which allows users to select different styles of chord progression and piano texture. Currently, chord progression styles include Pop, R&B, and Dark, while piano texture styles include several levels of voicing density and rhythmic complexity. Experimental results show that both our harmonization and arrangement results significantly outperform the baselines. Lastly, we release AccoMontage2 as an online application as well as the organized chord progression templates as a public dataset.


【3】 Generating Coherent Drum Accompaniment With Fills And Improvisations

标题:用填充物和即兴演奏产生连贯的鼓伴奏

链接:https://arxiv.org/abs/2209.00291

作者:Rishabh Dahale,Vaibhav Talwadker,Preeti Rao,Prateek Verma
机构:Department of Electrical Engineering, Indian Institute of Technology Bombay, India,  Stanford University
备注:8 pages, 7 figures, 23rd International Society for Music Information Retrieval Conference (ISMIR 2022), Bengaluru, India
摘要:创作像音乐这样复杂的艺术作品需要深刻的创造力。随着深度学习的最新进展和Transformers等强大模型的出现,自动音乐生成领域取得了巨大进展。在伴奏生成环境中,即使对于经验丰富的鼓手,在歌曲中的适当位置创建具有适当填充和即兴演奏的连贯鼓模式也是一项具有挑战性的任务。鼓点往往遵循重复的模式,通过节填充或即兴在部分边界。在这部作品中,我们处理了以四种旋律乐器演奏的伴奏音乐为条件的鼓模式生成任务:钢琴、吉他、贝斯和弦乐。我们使用Transformer序列到序列模型来生成以旋律伴奏为条件的基本鼓模式,以发现很大程度上不存在即兴演奏,这可能归因于其在训练数据中预期的相对较低的表示。我们提出了一个新颖的函数来捕捉酒吧相对于其邻居的即兴程度。我们训练一个模型,从旋律伴奏音轨中预测即兴演奏的位置。最后,我们使用一个新颖的BERT启发的填充架构,学习鼓和旋律的结构来填充即兴音乐的元素。
摘要:Creating a complex work of art like music necessitates profound creativity. With recent advancements in deep learning and powerful models such as transformers, there has been huge progress in automatic music generation. In an accompaniment generation context, creating a coherent drum pattern with apposite fills and improvisations at proper locations in a song is a challenging task even for an experienced drummer. Drum beats tend to follow a repetitive pattern through stanzas with fills or improvisation at section boundaries. In this work, we tackle the task of drum pattern generation conditioned on the accompanying music played by four melodic instruments: Piano, Guitar, Bass, and Strings. We use the transformer sequence to sequence model to generate a basic drum pattern conditioned on the melodic accompaniment to find that improvisation is largely absent, attributed possibly to its expectedly relatively low representation in the training data. We propose a novelty function to capture the extent of improvisation in a bar relative to its neighbors. We train a model to predict improvisation locations from the melodic accompaniment tracks. Finally, we use a novel BERT-inspired in-filling architecture, to learn the structure of both the drums and melody to in-fill elements of improvised music.


【4】 Video-Guided Curriculum Learning for Spoken Video Grounding

标题:视听基础课程的视频导学

链接:https://arxiv.org/abs/2209.00277

作者:Yan Xia,Zhou Zhao,Shangwei Ye,Yang Zhao,Haoyuan Li,Yi Ren
机构:Zhejiang University, Hangzhou, China
备注:Accepted by ACM MM 2022
摘要:本文提出了一个新的任务——口语视频基础(SVG),它的目标是从口语描述中定位出所需的视频片段。与使用文本相比,使用音频要求模型直接从原始语音中利用与视频相关的有用音素和音节。此外,我们随机加入环境噪声到这个语音音频中,进一步增加了这个任务的难度,更好地模拟了实际应用。为了从嘈杂的音频中纠正有区别的音素并提取视频相关信息,提出了一种新的视频引导课程学习(VGCL)方法,该方法可以在音频预训练过程中利用重要的视觉感知帮助理解口语并抑制外部噪声。考虑到该模型在推理过程中无法获得真实视频片段,设计了一种课程策略,在预训练阶段将输入视频从真实视频逐步转移到整个视频内容。最后,模型可以学习如何从整个视频剪辑中提取关键视觉信息,以帮助理解口语。此外,我们还收集了第一个基于ActivityNet的大规模口语视频基础数据集,命名为ActivityNet Speech数据集。大量实验表明,视频引导的课程学习能够促进预训练过程,获得一个相互的音频编码器,显著提高口语视频基础任务的成绩。此外,我们还证明了在有噪声的情况下,我们的模型优于用ASR脚本接地视频的方法,进一步证明了我们的课程策略的有效性。
摘要:In this paper, we introduce a new task, spoken video grounding (SVG), which aims to localize the desired video fragments from spoken language descriptions. Compared with using text, employing audio requires the model to directly exploit the useful phonemes and syllables related to the video from raw speech. Moreover, we randomly add environmental noises to this speech audio, further increasing the difficulty of this task and better simulating real applications. To rectify the discriminative phonemes and extract video-related information from noisy audio, we develop a novel video-guided curriculum learning (VGCL) during the audio pre-training process, which can make use of the vital visual perceptions to help understand the spoken language and suppress the external noise. Considering during inference the model can not obtain ground truth video segments, we design a curriculum strategy that gradually shifts the input video from the ground truth to the entire video content during pre-training. Finally, the model can learn how to extract critical visual information from the entire video clip to help understand the spoken language. In addition, we collect the first large-scale spoken video grounding dataset based on ActivityNet, which is named as ActivityNet Speech dataset. Extensive experiments demonstrate our proposed video-guided curriculum learning can facilitate the pre-training process to obtain a mutual audio encoder, significantly promoting the performance of spoken video grounding tasks. Moreover, we prove that in the case of noisy sound, our model outperforms the method that grounding video with ASR transcripts, further demonstrating the effectiveness of our curriculum strategy.


【5】 Attention Enhanced Citrinet for Speech Recognition

标题:用于语音识别的注意力增强型Citrnet

链接:https://arxiv.org/abs/2209.00261

作者:Xianchao Wu
机构:NVIDIA
备注:5 pages, 3 figures
摘要:Citrinet是一种基于卷积连接主义时态分类(CTC)的端到端自动语音识别(ASR)模型。Citrinet采用一维时间通道可分离卷积结合子词编码和压缩激励(SE)技术来获取局部和全局上下文信息,使整个体系结构达到23个块、235个卷积层和46个线性层的深度。这种纯卷积和深度架构使得Critrinet在收敛时相对较慢。本文提出在Citrinet块的卷积模块中引入多头注意和前馈网络,同时保持SE模块和残差模块不变。为了加快速度,我们在每个注意力增强的Citrinet块中删除了8个卷积层,并将23个块减少到13个。在日本CSJ-500h和Magic-1600h数据集上的实验结果表明,与训练时间为80%的Citrinet和训练时间为40%、模型大小为29.8%的Conformer相比,注意增强的Citrinet具有更少的层数和块数,收敛速度更快,字符错误率更低.
摘要:Citrinet is an end-to-end convolutional Connectionist Temporal Classification (CTC) based automatic speech recognition (ASR) model. To capture local and global contextual information, 1D time-channel separable convolutions combined with sub-word encoding and squeeze-and-excitation (SE) are used in Citrinet, making the whole architecture to be as deep as including 23 blocks with 235 convolution layers and 46 linear layers. This pure convolutional and deep architecture makes Critrinet relatively slow at convergence. In this paper, we propose to introduce multi-head attentions together with feed-forward networks in the convolution module in Citrinet blocks while keeping the SE module and residual module unchanged. For speeding up, we remove 8 convolution layers in each attention-enhanced Citrinet block and reduce 23 blocks to 13. Experiments on the Japanese CSJ-500h and Magic-1600h dataset show that the attention-enhanced Citrinet with less layers and blocks and converges faster with lower character error rates than (1) Citrinet with 80\% training time and (2) Conformer with 40\% training time and 29.8\% model size.


【6】 Deep Sparse Conformer for Speech Recognition

标题:一种用于语音识别的深度稀疏整形器

链接:https://arxiv.org/abs/2209.00260

作者:Xianchao Wu
机构:NVIDIA
备注:5 pages, 1 figure
摘要:Conformer利用Transformer对基于内容的全局交互作用的捕捉和卷积神经网络对局部特征的挖掘,在自动语音识别(ASR)中取得了令人印象深刻的效果。在Conformer中,两个具有半步剩余连接的马卡龙状前馈层将多头自注意和卷积模块夹在中间,然后进行后层归一化。我们在两个方向上改进了Conformer的长序列表示能力,\emph{sparser}和\emph{deeper}。我们采用了一种时间复杂度和内存占用率都为$\mathcal{O}(L\text{log}L)$的稀疏自注意机制。当执行残差连接时,使用深度归一化策略,以确保我们训练一百级Conformer块。在日本CSJ-500h数据集上,该深度稀疏Conformer在3个评估集上的CER分别为5.52%、4.03%和4.50%,在12~16、17、50和100层的5个深度稀疏Conformer变体上的CER分别为4.16%、2.84%和3.20%.
摘要:Conformer has achieved impressive results in Automatic Speech Recognition (ASR) by leveraging transformer's capturing of content-based global interactions and convolutional neural network's exploiting of local features. In Conformer, two macaron-like feed-forward layers with half-step residual connections sandwich the multi-head self-attention and convolution modules followed by a post layer normalization. We improve Conformer's long-sequence representation ability in two directions, \emph{sparser} and \emph{deeper}. We adapt a sparse self-attention mechanism with $\mathcal{O}(L\text{log}L)$ in time complexity and memory usage. A deep normalization strategy is utilized when performing residual connections to ensure our training of hundred-level Conformer blocks. On the Japanese CSJ-500h dataset, this deep sparse Conformer achieves respectively CERs of 5.52\%, 4.03\% and 4.50\% on the three evaluation sets and 4.16\%, 2.84\% and 3.20\% when ensembling five deep sparse Conformer variants from 12 to 16, 17, 50, and finally 100 encoder layers.


【7】 What is missing in deep music generation? A study of repetition and  structure in popular music

标题:深度音乐世代缺失了什么?流行音乐中的重复与结构研究

链接:https://arxiv.org/abs/2209.00182

作者:Shuqi Dai,Huiran Yu,Roger B. Dannenberg
机构:Carnegie Mellon University
备注:In Proceedings of the 23rd Int. Society for Music Information Retrieval (ISMIR) 2022
摘要:结构是音乐最本质的方面之一,音乐结构通常通过重复来表示。然而,音乐中的重复和结构的本质仍然没有得到很好的理解,特别是在音乐生成的背景下,音乐信息检索(MIR)技术还有很多有待探索的地方。对中国和美国两个流行音乐数据集的分析说明了重要的音乐结构原则:(1)结构存在于多个分层水平上,(2)歌曲使用重复和有限的词汇,使得单个歌曲不遵循歌曲集合的一般统计,(3)结构与节奏、旋律、和声和可预测性相互作用,以及(4)在歌曲的过程中,重复不是随机的,而是遵循如交叉熵所揭示的一般趋势。这些和其他发现为深度学习音乐生成提供了挑战和机遇,并提出了新的正式音乐标准和评估方法。我们分析了最近的音乐生成系统中的音乐,并将其与我们数据集中的人类创作的音乐进行了比较,通常从结构角度揭示了惊人的差异。
摘要:Structure is one of the most essential aspects of music, and music structure is commonly indicated through repetition. However, the nature of repetition and structure in music is still not well understood, especially in the context of music generation, and much remains to be explored with Music Information Retrieval (MIR) techniques. Analyses of two popular music datasets (Chinese and American) illustrate important music construction principles: (1) structure exists at multiple hierarchical levels, (2) songs use repetition and limited vocabulary so that individual songs do not follow general statistics of song collections, (3) structure interacts with rhythm, melody, harmony, and predictability, and (4) over the course of a song, repetition is not random, but follows a general trend as revealed by cross-entropy. These and other findings offer challenges as well as opportunities for deep-learning music generation and suggest new formal music criteria and evaluation methods. Music from recent music generation systems is analyzed and compared to human-composed music in our datasets, often revealing striking differences from a structural perspective.


【8】 Evaluating generative audio systems and their metrics

标题:评估生成性音频系统及其度量

链接:https://arxiv.org/abs/2209.00130

作者:Ashvala Vinay,Alexander Lerch
机构:Center for Music Technology, Georgia Institute of Technology
备注:Accepted at ISMIR 2022
摘要:近年来,在具有深度生成模型的音频合成方面取得了相当大的进展。然而,最新技术水平很难量化;当报告结果时,不同的研究通常使用不同的评估方法和不同的度量,使得与其他系统的直接比较即使不是不可能也是困难的。此外,在大多数情况下,所报告指标的感知相关性和意义未知,因此无法获得关于实际可用性和音频质量的任何结论性见解。本文介绍了一项研究,调查了最先进的方法并排(i)一组先前提出的客观指标的音频重建,并与(ii)听力研究。结果表明,目前使用的客观评价指标不足以描述当前系统的感知质量。
摘要:Recent years have seen considerable advances in audio synthesis with deep generative models. However, the state-of-the-art is very difficult to quantify; different studies often use different evaluation methodologies and different metrics when reporting results, making a direct comparison to other systems difficult if not impossible. Furthermore, the perceptual relevance and meaning of the reported metrics in most cases unknown, prohibiting any conclusive insights with respect to practical usability and audio quality. This paper presents a study that investigates state-of-the-art approaches side-by-side with (i) a set of previously proposed objective metrics for audio reconstruction, and with (ii) a listening study. The results indicate that currently used objective metrics are insufficient to describe the perceptual quality of current systems.


【9】 Joint Speaker Encoder and Neural Back-end Model for Fully End-to-End  Automatic Speaker Verification with Multiple Enrollment Utterances

标题:多注册语音全端到端自动说话人确认的联合说话人编码和神经网络后端模型

链接:https://arxiv.org/abs/2209.00485

作者:Chang Zeng,Xiaoxiao Miao,Xin Wang,Erica Cooper,Junichi Yamagishi
机构:Senior Member, IEEE
备注:Submitted to TASLP
摘要:传统的自动说话人确认系统通常可以被分解为用于提取说话人嵌入的前端模型(如时间延迟神经网络(TDNN))和用于相似性评分的后端模型(如基于统计的概率线性判别分析(PLDA)或基于神经网络的神经PLDA(NPLDA))。然而,前后端模型的顺序优化可能会导致局部最小值,这在理论上阻碍了整个系统达到最佳优化。尽管已经提出了一些方法来联合优化这两个模型,例如广义端到端(GE2E)模型和NPLDA E2E模型,但是所有这些方法都被设计用于单个登记话语。本文提出了一种新的E2E联合说话人确认方法,该方法专门针对多个注册语音的实际情况而设计。为了充分利用多个注册话语之间的内在联系,我们的模型配备了框架级和话语级注意机制。我们还利用了几种数据增强技术,包括使用MUSAN和RIR数据集的传统噪声增强和用于更好优化的独特的扬声器嵌入级混淆策略。
摘要:Conventional automatic speaker verification systems can usually be decomposed into a front-end model such as time delay neural network (TDNN) for extracting speaker embeddings and a back-end model such as statistics-based probabilistic linear discriminant analysis (PLDA) or neural network-based neural PLDA (NPLDA) for similarity scoring. However, the sequential optimization of the front-end and back-end models may lead to a local minimum, which theoretically prevents the whole system from achieving the best optimization. Although some methods have been proposed for jointly optimizing the two models, such as the generalized end-to-end (GE2E) model and NPLDA E2E model, all of these methods are designed for use with a single enrollment utterance. In this paper, we propose a new E2E joint method for speaker verification especially designed for the practical case of multiple enrollment utterances. In order to leverage the intra-relationship among multiple enrollment utterances, our model comes equipped with frame-level and utterance-level attention mechanisms. We also utilize several data augmentation techniques, including conventional noise augmentation using MUSAN and RIRs datasets and a unique speaker embedding-level mixup strategy for better optimization.


eess.AS音频处理

【1】 diaLogic: Non-Invasive Speaker-Focused Data Acquisition for Team  Behavior Modeling

标题:Dialogic:面向团队行为建模的非侵入性说话人数据采集

链接:https://arxiv.org/abs/2209.00619

作者:Ryan Duke,Alex Doboli
机构:Duke Department of Electrical & Computer EngineeringStony Brook UniversityNY,  Doboli Department of Electrical & Computer EngineeringStony Brook UniversityNY
摘要:本文提出了一个用于模拟团队在解决开放式问题时的行为的人在回路系统—-diaLogic系统。通过从由所获取的语音数据计算的特征中提取的假设来对团队行为建模。这些特征包括说话者交互、说话者情绪、基频以及相应的文本和子句。基于团队行为随时间变化的相似性和差异性,发现了关于不变和差异化情况的假设。为了提供完全自动化的数据采集,diaLogic系统在直观、用户友好的GUI界面中执行。实验表明,该系统在解决问题过程中的团队行为具有很好的性能。
摘要:This paper presents diaLogic system, a Human-In-A-Loop system for modeling the behavior of teams during solving open-ended problems. Team behavior is modeled through the hypotheses extracted from features computed from acquired voice data. These features include speaker interactions, speaker emotions, fundamental frequencies, and the corresponding text and clauses. Hypotheses about the invariant and differentiated situations are found based on the similarities and dissimilarities of the behavior of teams over time. To provide full automation of data acquisition, the diaLogic system is executed within an intuitive, user-friendly GUI interface. Experiments present the performance of the system for a broad set of cases featuring team behavior during problem solving.


【2】 On the potential of jointly-optimised solutions to spoofing attack  detection and automatic speaker verification

标题:联合优化的欺骗攻击检测和自动说话人验证解决方案的潜力

链接:https://arxiv.org/abs/2209.00506

作者:Wanying Ge,Hemlata Tak,Massimiliano Todisco,Nicholas Evans
机构:EURECOM, Sophia Antipolis, France备注:Submitted to IberSPEECH 2022 Conference
摘要:欺骗感知说话人确认(SASV)挑战赛旨在促进联合优化解决方案的研究,以完成传统上分别优化的欺骗检测和说话人确认任务。联合优化的系统具有协同操作的潜力,作为可靠的说话人确认的单一任务的更好执行的解决方案。然而,向SASV 2022提交的23份申请中,没有一份是联合优化的。因此,我们试图确定为什么单独优化的子系统性能最好,或者为什么联合优化不成功。实验结果表明,联合优化算法可以有效提高说话人身份验证系统对欺骗的鲁棒性,但会降低说话人身份验证系统的性能。研究结果表明,电子欺骗检测和说话人验证子系统应该以反映每个子系统提供的信息如何与另一个子系统提供的信息互补的方式来共同优化。进展还可能取决于从更多发言者那里收集数据。
摘要:The spoofing-aware speaker verification (SASV) challenge was designed to promote the study of jointly-optimised solutions to accomplish the traditionally separately-optimised tasks of spoofing detection and speaker verification. Jointly-optimised systems have the potential to operate in synergy as a better performing solution to the single task of reliable speaker verification. However, none of the 23 submissions to SASV 2022 are jointly optimised. We have hence sought to determine why separately-optimised sub-systems perform best or why joint optimisation was not successful. Experiments reported in this paper show that joint optimisation is successful in improving robustness to spoofing but that it degrades speaker verification performance. The findings suggest that spoofing detection and speaker verification sub-systems should be optimised jointly in a manner which reflects the differences in how information provided by each sub-system is complementary to that provided by the other. Progress will also likely depend upon the collection of data from a larger number of speakers.


【3】 Joint Speaker Encoder and Neural Back-end Model for Fully End-to-End  Automatic Speaker Verification with Multiple Enrollment Utterances

标题:多注册语音全端到端自动说话人确认的联合说话人编码和神经网络后端模型

链接:https://arxiv.org/abs/2209.00485

* 与cs.SD语音【9】为同一篇

作者:Chang Zeng,Xiaoxiao Miao,Xin Wang,Erica Cooper,Junichi Yamagishi
机构:Senior Member, IEEE
备注:Submitted to TASLP
摘要:传统的自动说话人确认系统通常可以被分解为用于提取说话人嵌入的前端模型(如时间延迟神经网络(TDNN))和用于相似性评分的后端模型(如基于统计的概率线性判别分析(PLDA)或基于神经网络的神经PLDA(NPLDA))。然而,前后端模型的顺序优化可能会导致局部最小值,这在理论上阻碍了整个系统达到最佳优化。尽管已经提出了一些方法来联合优化这两个模型,例如广义端到端(GE2E)模型和NPLDA E2E模型,但是所有这些方法都被设计用于单个登记话语。本文提出了一种新的E2E联合说话人确认方法,该方法专门针对多个注册语音的实际情况而设计。为了充分利用多个注册话语之间的内在联系,我们的模型配备了框架级和话语级注意机制。我们还利用了几种数据增强技术,包括使用MUSAN和RIR数据集的传统噪声增强和用于更好优化的独特的扬声器嵌入级混淆策略。
摘要:Conventional automatic speaker verification systems can usually be decomposed into a front-end model such as time delay neural network (TDNN) for extracting speaker embeddings and a back-end model such as statistics-based probabilistic linear discriminant analysis (PLDA) or neural network-based neural PLDA (NPLDA) for similarity scoring. However, the sequential optimization of the front-end and back-end models may lead to a local minimum, which theoretically prevents the whole system from achieving the best optimization. Although some methods have been proposed for jointly optimizing the two models, such as the generalized end-to-end (GE2E) model and NPLDA E2E model, all of these methods are designed for use with a single enrollment utterance. In this paper, we propose a new E2E joint method for speaker verification especially designed for the practical case of multiple enrollment utterances. In order to leverage the intra-relationship among multiple enrollment utterances, our model comes equipped with frame-level and utterance-level attention mechanisms. We also utilize several data augmentation techniques, including conventional noise augmentation using MUSAN and RIRs datasets and a unique speaker embedding-level mixup strategy for better optimization.


【4】 Spoofing-Aware Attention based ASV Back-end with Multiple Enrollment  Utterances and a Sampling Strategy for the SASV Challenge 2022

标题:基于欺骗感知注意力的ASV后端多注册话语和SASV 2022挑战赛采样策略

链接:https://arxiv.org/abs/2209.00423

作者:Chang Zeng,Lin Zhang,Meng Liu,Junichi Yamagishi
机构:National Institute of Informatics, Japan ,SOKENDAI, Japan ,Tianjin University, China
备注:Accepted by InterSpeech2022
摘要:当前的自动说话人确认(ASV)系统容易受到表示攻击,为了保护ASV系统,人们提出了几种对抗措施(CMs),以区分真实的和欺骗的说话人。然而,ASV系统和CM通常是独立开发和优化的,而没有考虑它们之间的相互关系。本文提出了一种新的具有欺骗感知能力的ASV后端模块,该模块基于说话人相似度和CM评分计算ASV综合评分。后端模块除了具有两个分数的可学习融合功能外,还具有标度点和前馈自注意两种类型的注意成分,因此也可以同时学习多个注册话语的内部关系信息。此外,设计了一种新的有效的试验抽样策略,用于模拟2022年欺骗感知说话人验证挑战赛(SASV)中引入的新的欺骗感知验证场景。
摘要:Current state-of-the-art automatic speaker verification (ASV) systems are vulnerable to presentation attacks, and several countermeasures (CMs), which distinguish bona fide trials from spoofing ones, have been explored to protect ASV. However, ASV systems and CMs are generally developed and optimized independently without considering their inter-relationship. In this paper, we propose a new spoofing-aware ASV back-end module that efficiently computes a combined ASV score based on speaker similarity and CM score. In addition to the learnable fusion function of the two scores, the proposed back-end module has two types of attention components, scaled-dot and feed-forward self-attention, so that intra-relationship information of multiple enrollment utterances can also be learned at the same time. Moreover, a new effective trials-sampling strategy is designed for simulating new spoofing-aware verification scenarios introduced in the Spoof-Aware Speaker Verification (SASV) challenge 2022.


【5】 Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild

标题:野外任意说话人的唇语合成

链接:https://arxiv.org/abs/2209.00642

* 与cs.SD语音【1】为同一篇

作者:Sindhu B Hegde,K R Prajwal,Rudrabha Mukhopadhyay,Vinay P Namboodiri,C. V. Jawahar
机构:International Institute of Information, Technology, Hyderabad, India, University of Oxford, United Kingdom, University of Bath
备注:Accepted in ACM-MM 2022, 9 pages, 2 pages supplementary, 7 Figures
摘要:在这项工作中,我们解决了从无声的嘴唇视频生成任何说话者在野外的语音的问题。与之前的工作形成鲜明对比的是,我们的方法(i)不限于固定数量的说话者,(ii)不明确地对领域或词汇表施加约束,以及(iii)处理在野外录制的视频,而不是在实验室设置内。这项任务提出了一系列挑战,其中关键的一个挑战是,所需目标语音的许多特征,如声音、音高和语言内容,不能完全从无声面部视频中推断出来。为了处理这些随机变化,我们提出了一种新的VAE-GAN结构,该结构在变化中学习将嘴唇和语音序列关联起来。在指导训练过程的多个功能强大的识别器的帮助下,我们的生成器学会了合成任何人嘴唇运动的任何声音的语音序列。在多个数据集上的大量实验表明,我们的性能大大优于所有基线。此外,我们的网络可以在特定身份的视频上进行微调,以实现与单扬声器模型相当的性能,后者是在4\x $以上的数据上训练的。我们进行了大量消融研究,以分析我们架构的不同模块的影响。我们还在我们的网站上提供了演示视频,展示了几个定性结果以及代码和经过训练的模型:\url{http://cvit.iiiit.ac.in/research/projects/cvit-projects/lip-to-speech-synthesis}}(唇转语音合成技术)的研究成果
摘要:In this work, we address the problem of generating speech from silent lip videos for any speaker in the wild. In stark contrast to previous works, our method (i) is not restricted to a fixed number of speakers, (ii) does not explicitly impose constraints on the domain or the vocabulary and (iii) deals with videos that are recorded in the wild as opposed to within laboratory settings. The task presents a host of challenges, with the key one being that many features of the desired target speech, like voice, pitch and linguistic content, cannot be entirely inferred from the silent face video. In order to handle these stochastic variations, we propose a new VAE-GAN architecture that learns to associate the lip and speech sequences amidst the variations. With the help of multiple powerful discriminators that guide the training process, our generator learns to synthesize speech sequences in any voice for the lip movements of any person. Extensive experiments on multiple datasets show that we outperform all baselines by a large margin. Further, our network can be fine-tuned on videos of specific identities to achieve a performance comparable to single-speaker models that are trained on $4\times$ more data. We conduct numerous ablation studies to analyze the effect of different modules of our architecture. We also provide a demo video that demonstrates several qualitative results along with the code and trained models on our website: \url{http://cvit.iiit.ac.in/research/projects/cvit-projects/lip-to-speech-synthesis}}


【6】 AccoMontage2: A Complete Harmonization and Accompaniment Arrangement  System

标题:AccoMontage2:一个完整的和声和伴奏编排系统

链接:https://arxiv.org/abs/2209.00353

* 与cs.SD语音【2】为同一篇

作者:Li Yi,Haochen Hu,Jingwei Zhao,Gus Xia
机构:Music X Lab, NYU Shanghai,  MBZUAI,  Institute of Data Science, NUS
备注:Accepted by ISMIR 2022
摘要:我们提出了AccoMontage2,一个能够基于主旋律进行完整长度的歌曲和声和伴奏安排的系统。本研究以AccoMontage为基础,以流行歌曲/民歌为研究对象,进行了基于模板的检索方法研究。这项研究的新颖之处有两个方面。首先,我们发明了一个协调模块(AccoMontage没有)。该模块通过优化和平衡三个损耗项来生成结构化和连贯的全长弦级数:对于音符不协和的微观级损失、对于短语模板匹配的中观级损失、以及对于整段连贯性的宏观级损失。其次,我们开发了一个图形用户界面,允许用户选择不同风格的和弦进行和钢琴织体。目前,和弦进行风格包括Pop、R&B和Dark,而钢琴纹理风格包括几个层次的发声密度和节奏复杂度。实验结果表明,本文的协调和排序结果均显著优于基线。最后,我们发布了AccoMontage2在线应用程序,以及组织和弦进行模板作为公共数据集。
摘要:We propose AccoMontage2, a system capable of doing full-length song harmonization and accompaniment arrangement based on a lead melody. Following AccoMontage, this study focuses on generating piano arrangements for popular/folk songs and it carries on the generalized template-based retrieval method. The novelties of this study are twofold. First, we invent a harmonization module (which AccoMontage does not have). This module generates structured and coherent full-length chord progression by optimizing and balancing three loss terms: a micro-level loss for note-wise dissonance, a meso-level loss for phrase-template matching, and a macro-level loss for full piece coherency. Second, we develop a graphical user interface which allows users to select different styles of chord progression and piano texture. Currently, chord progression styles include Pop, R&B, and Dark, while piano texture styles include several levels of voicing density and rhythmic complexity. Experimental results show that both our harmonization and arrangement results significantly outperform the baselines. Lastly, we release AccoMontage2 as an online application as well as the organized chord progression templates as a public dataset.


【7】 Generating Coherent Drum Accompaniment With Fills And Improvisations

标题:用填充物和即兴演奏产生连贯的鼓伴奏

链接:https://arxiv.org/abs/2209.00291

* 与cs.SD语音【3】为同一篇

作者:Rishabh Dahale,Vaibhav Talwadker,Preeti Rao,Prateek Verma
机构:Department of Electrical Engineering, Indian Institute of Technology Bombay, India,  Stanford University
备注:8 pages, 7 figures, 23rd International Society for Music Information Retrieval Conference (ISMIR 2022), Bengaluru, India
摘要:创作像音乐这样复杂的艺术作品需要深刻的创造力。随着深度学习的最新进展和Transformers等强大模型的出现,自动音乐生成领域取得了巨大进展。在伴奏生成环境中,即使对于经验丰富的鼓手,在歌曲中的适当位置创建具有适当填充和即兴演奏的连贯鼓模式也是一项具有挑战性的任务。鼓点往往遵循重复的模式,通过节填充或即兴在部分边界。在这部作品中,我们处理了以四种旋律乐器演奏的伴奏音乐为条件的鼓模式生成任务:钢琴、吉他、贝斯和弦乐。我们使用Transformer序列到序列模型来生成以旋律伴奏为条件的基本鼓模式,以发现很大程度上不存在即兴演奏,这可能归因于其在训练数据中预期的相对较低的表示。我们提出了一个新颖的函数来捕捉酒吧相对于其邻居的即兴程度。我们训练一个模型,从旋律伴奏音轨中预测即兴演奏的位置。最后,我们使用一个新颖的BERT启发的填充架构,学习鼓和旋律的结构来填充即兴音乐的元素。
摘要:Creating a complex work of art like music necessitates profound creativity. With recent advancements in deep learning and powerful models such as transformers, there has been huge progress in automatic music generation. In an accompaniment generation context, creating a coherent drum pattern with apposite fills and improvisations at proper locations in a song is a challenging task even for an experienced drummer. Drum beats tend to follow a repetitive pattern through stanzas with fills or improvisation at section boundaries. In this work, we tackle the task of drum pattern generation conditioned on the accompanying music played by four melodic instruments: Piano, Guitar, Bass, and Strings. We use the transformer sequence to sequence model to generate a basic drum pattern conditioned on the melodic accompaniment to find that improvisation is largely absent, attributed possibly to its expectedly relatively low representation in the training data. We propose a novelty function to capture the extent of improvisation in a bar relative to its neighbors. We train a model to predict improvisation locations from the melodic accompaniment tracks. Finally, we use a novel BERT-inspired in-filling architecture, to learn the structure of both the drums and melody to in-fill elements of improvised music.


【8】 Video-Guided Curriculum Learning for Spoken Video Grounding

标题:视听基础课程的视频导学

链接:https://arxiv.org/abs/2209.00277

* 与cs.SD语音【4】为同一篇

作者:Yan Xia,Zhou Zhao,Shangwei Ye,Yang Zhao,Haoyuan Li,Yi Ren
机构:Zhejiang University, Hangzhou, China
备注:Accepted by ACM MM 2022
摘要:本文提出了一个新的任务——口语视频基础(SVG),它的目标是从口语描述中定位出所需的视频片段。与使用文本相比,使用音频要求模型直接从原始语音中利用与视频相关的有用音素和音节。此外,我们随机加入环境噪声到这个语音音频中,进一步增加了这个任务的难度,更好地模拟了实际应用。为了从嘈杂的音频中纠正有区别的音素并提取视频相关信息,提出了一种新的视频引导课程学习(VGCL)方法,该方法可以在音频预训练过程中利用重要的视觉感知帮助理解口语并抑制外部噪声。考虑到该模型在推理过程中无法获得真实视频片段,设计了一种课程策略,在预训练阶段将输入视频从真实视频逐步转移到整个视频内容。最后,模型可以学习如何从整个视频剪辑中提取关键视觉信息,以帮助理解口语。此外,我们还收集了第一个基于ActivityNet的大规模口语视频基础数据集,命名为ActivityNet Speech数据集。大量实验表明,视频引导的课程学习能够促进预训练过程,获得一个相互的音频编码器,显著提高口语视频基础任务的成绩。此外,我们还证明了在有噪声的情况下,我们的模型优于用ASR脚本接地视频的方法,进一步证明了我们的课程策略的有效性。
摘要:In this paper, we introduce a new task, spoken video grounding (SVG), which aims to localize the desired video fragments from spoken language descriptions. Compared with using text, employing audio requires the model to directly exploit the useful phonemes and syllables related to the video from raw speech. Moreover, we randomly add environmental noises to this speech audio, further increasing the difficulty of this task and better simulating real applications. To rectify the discriminative phonemes and extract video-related information from noisy audio, we develop a novel video-guided curriculum learning (VGCL) during the audio pre-training process, which can make use of the vital visual perceptions to help understand the spoken language and suppress the external noise. Considering during inference the model can not obtain ground truth video segments, we design a curriculum strategy that gradually shifts the input video from the ground truth to the entire video content during pre-training. Finally, the model can learn how to extract critical visual information from the entire video clip to help understand the spoken language. In addition, we collect the first large-scale spoken video grounding dataset based on ActivityNet, which is named as ActivityNet Speech dataset. Extensive experiments demonstrate our proposed video-guided curriculum learning can facilitate the pre-training process to obtain a mutual audio encoder, significantly promoting the performance of spoken video grounding tasks. Moreover, we prove that in the case of noisy sound, our model outperforms the method that grounding video with ASR transcripts, further demonstrating the effectiveness of our curriculum strategy.


【9】 Attention Enhanced Citrinet for Speech Recognition

标题:用于语音识别的注意力增强型Citrnet

链接:https://arxiv.org/abs/2209.00261

* 与cs.SD语音【5】为同一篇

作者:Xianchao Wu
机构:NVIDIA
备注:5 pages, 3 figures
摘要:Citrinet是一种基于卷积连接主义时态分类(CTC)的端到端自动语音识别(ASR)模型。Citrinet采用一维时间通道可分离卷积结合子词编码和压缩激励(SE)技术来获取局部和全局上下文信息,使整个体系结构达到23个块、235个卷积层和46个线性层的深度。这种纯卷积和深度架构使得Critrinet在收敛时相对较慢。本文提出在Citrinet块的卷积模块中引入多头注意和前馈网络,同时保持SE模块和残差模块不变。为了加快速度,我们在每个注意力增强的Citrinet块中删除了8个卷积层,并将23个块减少到13个。在日本CSJ-500h和Magic-1600h数据集上的实验结果表明,与训练时间为80%的Citrinet和训练时间为40%、模型大小为29.8%的Conformer相比,注意增强的Citrinet具有更少的层数和块数,收敛速度更快,字符错误率更低.
摘要:Citrinet is an end-to-end convolutional Connectionist Temporal Classification (CTC) based automatic speech recognition (ASR) model. To capture local and global contextual information, 1D time-channel separable convolutions combined with sub-word encoding and squeeze-and-excitation (SE) are used in Citrinet, making the whole architecture to be as deep as including 23 blocks with 235 convolution layers and 46 linear layers. This pure convolutional and deep architecture makes Critrinet relatively slow at convergence. In this paper, we propose to introduce multi-head attentions together with feed-forward networks in the convolution module in Citrinet blocks while keeping the SE module and residual module unchanged. For speeding up, we remove 8 convolution layers in each attention-enhanced Citrinet block and reduce 23 blocks to 13. Experiments on the Japanese CSJ-500h and Magic-1600h dataset show that the attention-enhanced Citrinet with less layers and blocks and converges faster with lower character error rates than (1) Citrinet with 80\% training time and (2) Conformer with 40\% training time and 29.8\% model size.


【10】 Deep Sparse Conformer for Speech Recognition

标题:一种用于语音识别的深度稀疏整形器

链接:https://arxiv.org/abs/2209.00260

* 与cs.SD语音【6】为同一篇

作者:Xianchao Wu
机构:NVIDIA
备注:5 pages, 1 figure
摘要:Conformer利用Transformer对基于内容的全局交互作用的捕捉和卷积神经网络对局部特征的挖掘,在自动语音识别(ASR)中取得了令人印象深刻的效果。在Conformer中,两个具有半步剩余连接的马卡龙状前馈层将多头自注意和卷积模块夹在中间,然后进行后层归一化。我们在两个方向上改进了Conformer的长序列表示能力,\emph{sparser}和\emph{deeper}。我们采用了一种时间复杂度和内存占用率都为$\mathcal{O}(L\text{log}L)$的稀疏自注意机制。当执行残差连接时,使用深度归一化策略,以确保我们训练一百级Conformer块。在日本CSJ-500h数据集上,该深度稀疏Conformer在3个评估集上的CER分别为5.52%、4.03%和4.50%,在12~16、17、50和100层的5个深度稀疏Conformer变体上的CER分别为4.16%、2.84%和3.20%.
摘要:Conformer has achieved impressive results in Automatic Speech Recognition (ASR) by leveraging transformer's capturing of content-based global interactions and convolutional neural network's exploiting of local features. In Conformer, two macaron-like feed-forward layers with half-step residual connections sandwich the multi-head self-attention and convolution modules followed by a post layer normalization. We improve Conformer's long-sequence representation ability in two directions, \emph{sparser} and \emph{deeper}. We adapt a sparse self-attention mechanism with $\mathcal{O}(L\text{log}L)$ in time complexity and memory usage. A deep normalization strategy is utilized when performing residual connections to ensure our training of hundred-level Conformer blocks. On the Japanese CSJ-500h dataset, this deep sparse Conformer achieves respectively CERs of 5.52\%, 4.03\% and 4.50\% on the three evaluation sets and 4.16\%, 2.84\% and 3.20\% when ensembling five deep sparse Conformer variants from 12 to 16, 17, 50, and finally 100 encoder layers.


【11】 What is missing in deep music generation? A study of repetition and  structure in popular music

标题:深度音乐世代缺失了什么?流行音乐中的重复与结构研究

链接:https://arxiv.org/abs/2209.00182

* 与cs.SD语音【7】为同一篇

作者:Shuqi Dai,Huiran Yu,Roger B. Dannenberg
机构:Carnegie Mellon University
备注:In Proceedings of the 23rd Int. Society for Music Information Retrieval (ISMIR) 2022
摘要:结构是音乐最本质的方面之一,音乐结构通常通过重复来表示。然而,音乐中的重复和结构的本质仍然没有得到很好的理解,特别是在音乐生成的背景下,音乐信息检索(MIR)技术还有很多有待探索的地方。对中国和美国两个流行音乐数据集的分析说明了重要的音乐结构原则:(1)结构存在于多个分层水平上,(2)歌曲使用重复和有限的词汇,使得单个歌曲不遵循歌曲集合的一般统计,(3)结构与节奏、旋律、和声和可预测性相互作用,以及(4)在歌曲的过程中,重复不是随机的,而是遵循如交叉熵所揭示的一般趋势。这些和其他发现为深度学习音乐生成提供了挑战和机遇,并提出了新的正式音乐标准和评估方法。我们分析了最近的音乐生成系统中的音乐,并将其与我们数据集中的人类创作的音乐进行了比较,通常从结构角度揭示了惊人的差异。
摘要:Structure is one of the most essential aspects of music, and music structure is commonly indicated through repetition. However, the nature of repetition and structure in music is still not well understood, especially in the context of music generation, and much remains to be explored with Music Information Retrieval (MIR) techniques. Analyses of two popular music datasets (Chinese and American) illustrate important music construction principles: (1) structure exists at multiple hierarchical levels, (2) songs use repetition and limited vocabulary so that individual songs do not follow general statistics of song collections, (3) structure interacts with rhythm, melody, harmony, and predictability, and (4) over the course of a song, repetition is not random, but follows a general trend as revealed by cross-entropy. These and other findings offer challenges as well as opportunities for deep-learning music generation and suggest new formal music criteria and evaluation methods. Music from recent music generation systems is analyzed and compared to human-composed music in our datasets, often revealing striking differences from a structural perspective.


【12】 Evaluating generative audio systems and their metrics

标题:评估生成性音频系统及其度量

链接:https://arxiv.org/abs/2209.00130

* 与cs.SD语音【8】为同一篇

作者:Ashvala Vinay,Alexander Lerch
机构:Center for Music Technology, Georgia Institute of Technology
备注:Accepted at ISMIR 2022
摘要:近年来,在具有深度生成模型的音频合成方面取得了相当大的进展。然而,最新技术水平很难量化;当报告结果时,不同的研究通常使用不同的评估方法和不同的度量,使得与其他系统的直接比较即使不是不可能也是困难的。此外,在大多数情况下,所报告指标的感知相关性和意义未知,因此无法获得关于实际可用性和音频质量的任何结论性见解。本文介绍了一项研究,调查了最先进的方法并排(i)一组先前提出的客观指标的音频重建,并与(ii)听力研究。结果表明,目前使用的客观评价指标不足以描述当前系统的感知质量。
摘要:Recent years have seen considerable advances in audio synthesis with deep generative models. However, the state-of-the-art is very difficult to quantify; different studies often use different evaluation methodologies and different metrics when reporting results, making a direct comparison to other systems difficult if not impossible. Furthermore, the perceptual relevance and meaning of the reported metrics in most cases unknown, prohibiting any conclusive insights with respect to practical usability and audio quality. This paper presents a study that investigates state-of-the-art approaches side-by-side with (i) a set of previously proposed objective metrics for audio reconstruction, and with (ii) a listening study. The results indicate that currently used objective metrics are insufficient to describe the perceptual quality of current systems.


机器翻译,仅供参考