今天跟大家分享一篇语音相关的论文合集:cs.SD语音17篇,eess.AS音频处理20篇。
cs.SD语音
【1】 Let the paintings play
标题:让这些画发挥作用
链接:https://arxiv.org/abs/2206.14142
作者:Paola Gervasio,Alfio Quarteroni,Daniele Cassani机构:DICATAM, Universita degli Studi di Brescia, Brescia (Italy), MOX, Department of Mathematics, Politecnico di Milano, Milano (Italy) and Institute of, Mathematics, Ecole Polytechnique Federale de Lausanne, Lausanne (Switzerland), (professor emeritus)摘要:在本文中,我们介绍了一种数学方法来提取绘画和音乐曲目之间的相似性。我们的方法是基于绘画和音乐曲目的数字化,通过正交基函数的有限展开(使用傅立叶和小波基)。特定绘画与给定作曲家的音乐曲目样本之间的最佳匹配是通过有限维子空间上的$L^2$投影实现的。本文提供了几个例子,以分析意大利艺术家马塞洛·莫兰迪尼的艺术作品集。最后,我们开发了一个实现上述过程的原始applet,可以从网站免费下载https://github.com/pgerva/playing-paintings.git摘要:In this paper, we introduce a mathematical method to extract similarities between paintings and musical tracks. Our approach is based on the digitalization of both paintings and musical tracks by means of finite expansions in terms of orthogonal basis functions (with both Fourier and wavelet bases). The best fit between a specific painting and a sample of musical tracks from a given composer is achieved via an $L^2$ projection upon a finite-dimensional subspace. Several examples are provided for the analysis of a collection of works of art by the Italian artist Marcello Morandini. Finally, we have developed an original applet that implements the process above and which can be freely downloaded from the site https://github.com/pgerva/playing-paintings.git
【2】 Bengali Common Voice Speech Dataset for Automatic Speech Recognition
标题:用于自动语音识别的孟加拉普通语音数据集
链接:https://arxiv.org/abs/2206.14053
作者:Samiul Alam,Asif Sushmit,Zaowad Abdullah,Shahrin Nakkhatra,MD. Nazmuddoha Ansary,Syed Mobassir Hossen,Sazia Morshed Mehnaz,Tahsin Reasat,Ahmed Imtiaz Humayun机构:Bengali.AI, Michigan State University, RPI, Vanderbilt University, Rice University, equal contribution摘要:孟加拉语是世界上说得最多的语言之一,全球有3亿多人说孟加拉语。尽管孟加拉语语音识别系统很受欢迎,但由于缺乏多样的开源数据集,孟加拉语语音识别系统的开发研究受到阻碍。作为前进的方向,我们已经众包了孟加拉共同语音语音数据集,这是一个句子级自动语音识别语料库。该数据集是在Mozilla Common Voice平台上收集的,是正在进行的一项活动的一部分,这项活动在两个月内收集了超过400小时的数据,并且正在迅速增长。我们的分析表明,与现有最大的开源孟加拉语音数据集OpenSLR孟加拉ASR数据集相比,我们的数据集具有更多的说话人、音素和环境多样性。我们展示了从数据集中获得的见解,并讨论了未来版本中需要解决的关键语言挑战。此外,我们报告了几种自动语音识别(ASR)算法的当前性能,并为未来的研究设定了基准。摘要:Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of diverse open-source datasets. As a way forward, we have crowdsourced the Bengali Common Voice Speech Dataset, which is a sentence-level automatic speech recognition corpus. Collected on the Mozilla Common Voice platform, the dataset is part of an ongoing campaign that has led to the collection of over 400 hours of data in 2 months and is growing rapidly. Our analysis shows that our dataset has more speaker, phoneme, and environmental diversity compared to the OpenSLR Bengali ASR dataset, the largest existing open-source Bengali speech dataset. We present insights obtained from the dataset and discuss key linguistic challenges that need to be addressed in future versions. Additionally, we report the current performance of a few Automatic Speech Recognition (ASR) algorithms and set a benchmark for future research.
【3】 Show Me Your Face, And I'll Tell You How You Speak
标题:把你的脸给我看,我就告诉你怎么说话
链接:https://arxiv.org/abs/2206.14009
作者:Christen Millerdurai,Lotfy Abdel Khaliq,Timon Ulrich摘要:当我们说话时,可以从嘴唇的运动推断出讲话的韵律和内容。在这项工作中,我们探索了唇语合成的任务,即学习仅在说话人嘴唇运动的情况下生成语音,其中我们重点学习在无约束的大词汇量环境中多个说话人的精确唇语映射。我们通过说话人的面部特征(即年龄、性别、种族)捕捉其声音身份,并将其与嘴唇运动一起调节,以生成说话人身份感知语音。为此,我们提出了一种新的方法“Lip2Speech”,通过关键的设计选择,可以在无约束的场景中实现精确的唇语合成。我们还使用定量、定性指标和人类评估进行各种实验和广泛评估。摘要:When we speak, the prosody and content of the speech can be inferred from the movement of our lips. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate speech given only the lip movements of a speaker where we focus on learning accurate lip to speech mappings for multiple speakers in unconstrained, large vocabulary settings. We capture the speaker's voice identity through their facial characteristics, i.e., age, gender, ethnicity and condition them along with the lip movements to generate speaker identity aware speech. To this end, we present a novel method "Lip2Speech", with key design choices to achieve accurate lip to speech synthesis in unconstrained scenarios. We also perform various experiments and extensive evaluation using quantitative, qualitative metrics and human evaluation.
【4】 Attack Agnostic Dataset: Towards Generalization and Stabilization of Audio DeepFake Detection
标题:攻击无关数据集:面向音频DeepFake检测的泛化和稳定化
链接:https://arxiv.org/abs/2206.13979
作者:Piotr Kawa,Marcin Plata,Piotr Syga机构:Wrocław University of Science and Technology, Wrocław, Poland备注:Proceedings of INTERSPEECH 2022摘要:音频深度伪造允许创建高质量、令人信服的话语,因此,由于其潜在的应用,如模仿或虚假新闻,因此构成了威胁。检测这些操作的方法应具有良好的泛化性和稳定性,从而对使用未明确包含在训练中的技术进行的攻击具有鲁棒性。在这项工作中,我们引入了与攻击无关的数据集—两个音频DeepFakes和一个反欺骗数据集的组合,由于攻击的不交叉使用,可以更好地推广检测方法。我们对当前的DeepFake检测方法进行了深入分析,并考虑了不同的音频特征(前端)。此外,我们提出了一种基于LCNN的LFCC和mel谱图前端模型,该模型不仅具有良好的泛化性和稳定性,而且与基于LFCC的模型相比也有改进-我们将所有折叠的标准偏差和EER减少了两倍,最多减少了5%。摘要:Audio DeepFakes allow the creation of high-quality, convincing utterances and therefore pose a threat due to its potential applications such as impersonation or fake news. Methods for detecting these manipulations should be characterized by good generalization and stability leading to robustness against attacks conducted with techniques that are not explicitly included in the training. In this work, we introduce Attack Agnostic Dataset - a combination of two audio DeepFakes and one anti-spoofing datasets that, thanks to the disjoint use of attacks, can lead to better generalization of detection methods. We present a thorough analysis of current DeepFake detection methods and consider different audio features (front-ends). In addition, we propose a model based on LCNN with LFCC and mel-spectrogram front-end, which not only is characterized by a good generalization and stability results but also shows improvement over LFCC-based mode - we decrease standard deviation on all folds and EER in two folds by up to 5%.
【5】 QTI Submission to DCASE 2021: residual normalization for device-imbalanced acoustic scene classification with efficient design
标题:QTI提交给DCASE 2021:高效设计的设备不平衡声场分类的残差归一化
链接:https://arxiv.org/abs/2206.13909
作者:Byeonggeun Kim,Seunghan Yang,Jangho Kim,Simyung Chang机构:Qualcomm AI Research†, Qualcomm Korea YH, Seoul, Republic of Korea, Seoul National University, Seoul, Republic of Korea备注:tech report; won 1st place in DCASE2021 challenge. arXiv admin note: substantial text overlap with arXiv:2111.06531摘要:本技术报告描述了我们提交DCASE2021挑战的TASK1A的详细信息。本课题的目标是在模型复杂性的约束下,为设备不平衡数据集设计一个音频场景分类系统。本报告介绍了实现该目标的四种方法。首先,我们提出了残差归一化,这是一种新的特征归一化方法,它使用实例归一化和一条快捷路径来丢弃不必要的特定于设备的信息,而不会丢失有用的分类信息。其次,我们设计了一个高效的体系结构BC ResNet Mod,它是基线体系结构的一个修改版本,接收域有限。第三,我们利用从一个设备到多个设备的谱图到谱图的转换来增加训练数据。最后,我们使用三种模型压缩方案:剪枝、量化和知识提取来降低模型复杂性。该系统在TAU Urban Acoustic Scenes 2020 Mobile开发数据集中的平均测试精度为76.3%,具有315k参数,压缩到61.0KB非零参数后的平均测试精度为75.3%。摘要:This technical report describes the details of our TASK1A submission of the DCASE2021 challenge. The goal of the task is to design an audio scene classification system for device-imbalanced datasets under the constraints of model complexity. This report introduces four methods to achieve the goal. First, we propose Residual Normalization, a novel feature normalization method that uses instance normalization with a shortcut path to discard unnecessary device-specific information without losing useful information for classification. Second, we design an efficient architecture, BC-ResNet-Mod, a modified version of the baseline architecture with a limited receptive field. Third, we exploit spectrogram-to-spectrogram translation from one to multiple devices to augment training data. Finally, we utilize three model compression schemes: pruning, quantization, and knowledge distillation to reduce model complexity. The proposed system achieves an average test accuracy of 76.3% in TAU Urban Acoustic Scenes 2020 Mobile, development dataset with 315k parameters, and average test accuracy of 75.3% after compression to 61.0KB of non-zero parameters.
【6】 Comparison of Speech Representations for the MOS Prediction System
标题:MOS预报系统中语音表示方法的比较
链接:https://arxiv.org/abs/2206.13817
作者:Aki Kunikoshi,Jaebok Kim,Wonsuk Jun,Kåre Sjölander机构:ReadSpeaker备注:5 pages, 4 figures摘要:为了保证文语转换系统的质量,研究了自动预测听者平均意见分数(MOS)的方法。以前的许多研究侧重于架构方面的进步(如MBNet、LDNet等),以更有效的方式捕获光谱特征与MOS之间的关系,并实现高精度。然而,在泛化能力方面的最优表示在很大程度上仍然是未知的。为此,我们将通过wav2vec框架获得的自监督学习(SSL)特征的性能与光谱特征的性能进行了比较,如光谱图和melspectrogram的幅度。此外,我们建议将SSL功能与我们认为可以为自动MOS保留基本信息的功能相结合,以弥补各自的缺点。我们对从过去的暴风雪和语音转换挑战中收集的大规模听力测试语料库进行了全面的实验。我们发现,wav2vec特征集显示出最好的泛化能力,即使给定的基本事实并不总是可靠的。此外,我们发现这些组合表现最好,并分析了它们如何弥合光谱和wav2vec特征集之间的差距。摘要:Automatic methods to predict Mean Opinion Score (MOS) of listeners have been researched to assure the quality of Text-to-Speech systems. Many previous studies focus on architectural advances (e.g. MBNet, LDNet, etc.) to capture relations between spectral features and MOS in a more effective way and achieved high accuracy. However, the optimal representation in terms of generalization capability still largely remains unknown. To this end, we compare the performance of Self-Supervised Learning (SSL) features obtained by the wav2vec framework to that of spectral features such as magnitude of spectrogram and melspectrogram. Moreover, we propose to combine the SSL features and features which we believe to retain essential information to the automatic MOS to compensate each other for their drawbacks. We conduct comprehensive experiments on a large-scale listening test corpus collected from past Blizzard and Voice Conversion Challenges. We found that the wav2vec feature set showed the best generalization even though the given ground-truth was not always reliable. Furthermore, we found that the combinations performed the best and analyzed how they bridged the gap between spectral and the wav2vec feature sets.
【7】 Personalized Keyword Spotting through Multi-task Learning
标题:基于多任务学习的个性化关键词识别
链接:https://arxiv.org/abs/2206.13708
作者:Seunghan Yang,Byeonggeun Kim,Inseop Chung,Simyung Chang机构:Qualcomm AI Research†, Qualcomm Korea YH, Seoul, Republic of Korea, Seoul National University, Seoul, Republic of Korea备注:Proceedings of INTERSPEECH 2022摘要:关键词发现(KWS)在智能设备上实现基于语音的用户交互方面起着至关重要的作用,而传统的KWS(C-KWS)方法主要用于检测用户不可知的预定义关键词。然而,在实践中,大多数用户交互都来自注册到设备中的目标用户,这促使他们构建个性化的关键字定位。我们设计了两个个性化的KWS任务;(1) 目标用户偏向KWS(TB-KWS)和(2)目标用户专用KWS(TO-KWS)。为了解决这些任务,我们提出了通过多任务学习(PK-MTL)进行个性化关键字发现,该方法包括多任务学习和任务适应。首先,我们介绍了将多任务学习应用于关键词识别和说话人验证,以利用用户信息来实现关键词识别系统。接下来,我们设计了特定于任务的评分函数,以完全适应个性化的KWS任务。我们在传统和个性化场景下评估了我们的框架,结果表明PK-MTL可以显著降低误报率,尤其是在各种实际场景中。摘要:Keyword spotting (KWS) plays an essential role in enabling speech-based user interaction on smart devices, and conventional KWS (C-KWS) approaches have concentrated on detecting user-agnostic pre-defined keywords. However, in practice, most user interactions come from target users enrolled in the device which motivates to construct personalized keyword spotting. We design two personalized KWS tasks; (1) Target user Biased KWS (TB-KWS) and (2) Target user Only KWS (TO-KWS). To solve the tasks, we propose personalized keyword spotting through multi-task learning (PK-MTL) that consists of multi-task learning and task-adaptation. First, we introduce applying multi-task learning on keyword spotting and speaker verification to leverage user information to the keyword spotting system. Next, we design task-specific scoring functions to adapt to the personalized KWS tasks thoroughly. We evaluate our framework on conventional and personalized scenarios, and the results show that PK-MTL can dramatically reduce the false alarm rate, especially in various practical scenarios.
【8】 Domain Agnostic Few-shot Learning for Speaker Verification
标题:用于说话人确认的领域不可知的Few-Shot学习
链接:https://arxiv.org/abs/2206.13700
作者:Seunghan Yang,Debasmit Das,Janghoon Cho,Hyoungwoo Park,Sungrack Yun备注:Proceedings of INTERSPEECH 2022摘要:验证系统的深度学习模型通常无法推广到新用户和新环境,即使它们学习到高度区分的特征。为了解决这个问题,我们提出了一个简单的领域泛化框架,该框架可以学习处理新用户和新领域的分布转移。我们的框架由特定领域和领域聚合网络组成,它们分别是特定领域和组合领域的专家。通过使用这些网络,我们可以在训练阶段模拟新用户和新域的存在,从而最终产生更好的泛化。为了节省内存,我们通过将相似的域聚集在一起来减少特定于域的网络的数量。在对人工生成的噪声域进行广泛评估后,我们可以明确地显示我们的框架的泛化能力。此外,我们将我们提出的方法应用于标准基准上的现有竞争体系结构,这显示了进一步的性能改进。摘要:Deep learning models for verification systems often fail to generalize to new users and new environments, even though they learn highly discriminative features. To address this problem, we propose a few-shot domain generalization framework that learns to tackle distribution shift for new users and new domains. Our framework consists of domain-specific and domain-aggregation networks, which are the experts on specific and combined domains, respectively. By using these networks, we generate episodes that mimic the presence of both novel users and novel domains in the training phase to eventually produce better generalization. To save memory, we reduce the number of domain-specific networks by clustering similar domains together. Upon extensive evaluation on artificially generated noise domains, we can explicitly show generalization ability of our framework. In addition, we apply our proposed methods to the existing competitive architecture on the standard benchmark, which shows further performance improvements.
【9】 Dummy Prototypical Networks for Few-Shot Open-Set Keyword Spotting
标题:Few-Shot开放集关键词识别的虚拟原型网络
链接:https://arxiv.org/abs/2206.13691
作者:Byeonggeun Kim,Seunghan Yang,Inseop Chung,Simyung Chang机构:Qualcomm AI Research†, Qualcomm Korea YH, Seoul, Republic of Korea, Seoul National University, Seoul, Republic of Korea备注:Proceedings of INTERSPEECH 2022摘要:关键字定位是在流式音频中检测关键字的任务。传统的关键字定位以预定义的关键字分类为目标,但在很少的快照(通过示例查询)关键字定位中,例如,给定M-shot支持样本的N向分类,受到了越来越多的关注。此外,在现实场景中,可能存在来自意外类别(开放集)的话语,需要拒绝这些话语,而不是将其归类为N个类别之一。结合这两个需求,我们使用一个新的基准设置splitGSC来解决少数镜头开放集关键字发现问题。我们提出了基于度量学习的已知虚拟原型,以更好地检测开放集,并介绍了一种简单而强大的方法,虚拟原型网络(D-ProtoNets)。与最近提出的splitGSC中的少数镜头开放集识别(FSOSR)方法相比,我们的D-ProtoNets显示出明显的优势。我们还在标准基准上验证了我们的方法,miniImageNet和D-ProtoNets显示了FSOSR中最先进的开放集检测率。摘要:Keyword spotting is the task of detecting a keyword in streaming audio. Conventional keyword spotting targets predefined keywords classification, but there is growing attention in few-shot (query-by-example) keyword spotting, e.g., N-way classification given M-shot support samples. Moreover, in real-world scenarios, there can be utterances from unexpected categories (open-set) which need to be rejected rather than classified as one of the N classes. Combining the two needs, we tackle few-shot open-set keyword spotting with a new benchmark setting, named splitGSC. We propose episode-known dummy prototypes based on metric learning to detect an open-set better and introduce a simple and powerful approach, Dummy Prototypical Networks (D-ProtoNets). Our D-ProtoNets shows clear margins compared to recent few-shot open-set recognition (FSOSR) approaches in the suggested splitGSC. We also verify our method on a standard benchmark, miniImageNet, and D-ProtoNets shows the state-of-the-art open-set detection rate in FSOSR.
【10】 Tiny-Sepformer: A Tiny Time-Domain Transformer Network for Speech Separation
标题:微型隔离器:一种用于语音分离的微型时域变换网络
链接:https://arxiv.org/abs/2206.13689
作者:Jian Luo,Jianzong Wang,Ning Cheng,Edward Xiao,Xulong Zhang,Jing Xiao机构:. Ping An Technology (Shenzhen) Co., Ltd. ,. Aquinas International Academy, CA, USA备注:Accepted by Interspeech 2022摘要:时域变换神经网络在语音分离任务中已经证明了其优越性。然而,这些模型通常具有大量的网络参数,因此经常遇到GPU内存爆炸的问题。在本文中,我们提出了用于语音分离的微型Transformer网络——微型Sepformer。我们提出了两种降低模型参数和内存消耗的技术:(1)卷积注意(CA)块,将vanilla Transformer拆分为两条路径,多头注意和一维深度可分离卷积,(2)参数共享,共享CA块内的层参数。在我们的实验中,Tiny Sepformer可以大大减小模型尺寸,并在WSJ0-2/3Mix数据集上实现与vanilla Sepformer相当的分离性能。摘要:Time-domain Transformer neural networks have proven their superiority in speech separation tasks. However, these models usually have a large number of network parameters, thus often encountering the problem of GPU memory explosion. In this paper, we proposed Tiny-Sepformer, a tiny version of Transformer network for speech separation. We present two techniques to reduce the model parameters and memory consumption: (1) Convolution-Attention (CA) block, spliting the vanilla Transformer to two paths, multi-head attention and 1D depthwise separable convolution, (2) parameter sharing, sharing the layer parameters within the CA block. In our experiments, Tiny-Sepformer could greatly reduce the model size, and achieves comparable separation performance with vanilla Sepformer on WSJ0-2/3Mix datasets.
【11】 ClearBuds: Wireless Binaural Earbuds for Learning-Based Speech Enhancement
标题:ClearBuds:用于基于学习的语音增强的无线双耳耳机
链接:https://arxiv.org/abs/2206.13611
作者:Ishan Chatterjee,Maruchi Kim,Vivek Jayaram,Shyamnath Gollakota,Ira Kemelmacher-Shlizerman,Shwetak Patel,Steven M. Seitz机构:∗Co-primary student authors, Paul G. Allen School of Computer Science & Engineering, University of Washington, Seattle, WA, USA备注:12 pages, Published in Mobisys 2022摘要:我们介绍了ClearBubs,这是第一个利用神经网络来增强来自两个无线耳塞的语音流的硬件和软件系统。无线耳塞的实时语音增强需要高质量的声音分离和背景消除,并在移动电话上实时运行。Clear Buds为盲声源分离和入耳式移动系统的最新深度学习搭建了桥梁,做出了两项关键技术贡献:1)一种新的无线耳塞设计,能够作为同步双耳麦克风阵列运行;2)一种在移动设备上运行的轻型双通道语音增强神经网络。我们的神经网络具有一种新颖的级联结构,它将时域传统神经网络与基于频谱图的频率掩蔽神经网络相结合,以减少音频输出中的伪影。结果表明,我们的无线耳塞实现的同步误差小于64微秒,我们的网络在附带的手机上的运行时间为21.4毫秒。在对8名用户在之前未见过的室内和室外多路径场景中进行的野外评估中,我们的神经网络可以学习空间和声学线索,以执行噪声抑制和背景语音消除。在一项有37名参与者参与的用户研究中,他们花了超过15.4小时对野外采集的1041个音频样本进行评级,我们的系统实现了改进的平均意见评分和背景噪声抑制。包含演示的项目页面:https://clearbuds.cs.washington.edu摘要:We present ClearBuds, the first hardware and software system that utilizes a neural network to enhance speech streamed from two wireless earbuds. Real-time speech enhancement for wireless earbuds requires high-quality sound separation and background cancellation, operating in real-time and on a mobile phone. Clear-Buds bridges state-of-the-art deep learning for blind audio source separation and in-ear mobile systems by making two key technical contributions: 1) a new wireless earbud design capable of operating as a synchronized, binaural microphone array, and 2) a lightweight dual-channel speech enhancement neural network that runs on a mobile device. Our neural network has a novel cascaded architecture that combines a time-domain conventional neural network with a spectrogram-based frequency masking neural network to reduce the artifacts in the audio output. Results show that our wireless earbuds achieve a synchronization error less than 64 microseconds and our network has a runtime of 21.4 milliseconds on an accompanying mobile phone. In-the-wild evaluation with eight users in previously unseen indoor and outdoor multipath scenarios demonstrates that our neural network generalizes to learn both spatial and acoustic cues to perform noise suppression and background speech removal. In a user-study with 37 participants who spent over 15.4 hours rating 1041 audio samples collected in-the-wild, our system achieves improved mean opinion score and background noise suppression. Project page with demos: https://clearbuds.cs.washington.edu
【12】 Expressive, Variable, and Controllable Duration Modelling in TTS
标题:TTS中的表现性、可变性和可控性持续时间建模
链接:https://arxiv.org/abs/2206.14165
作者:Ammar Abbas,Thomas Merritt,Alexis Moinet,Sri Karlapati,Ewa Muszynska,Simon Slangen,Elia Gatti,Thomas Drugman备注:Accepted to be published in the Proceedings of InterSpeech 2022摘要:随着非注意神经文语转换系统的兴起,持续时间建模再次成为一个重要的研究问题。当前的方法在很大程度上依赖于以前的统计参数语音合成技术进行持续时间预测,这对语音的表达性和可变性建模较差。在本文中,我们提出了两种改进持续时间建模的替代方法。首先,我们提出了一个以短语为条件的持续时间模型,该模型改进了预测的持续时间,并提供了更好的暂停建模。我们表明,基于短语的持续时间模型比我们的基线持续时间模型提高了语音的自然度。其次,我们还提出了一个称为Cailiflow的多说话人时长模型,该模型使用归一化流来预测更符合复杂目标时长分布的时长。Cauliflow在自然度方面与我们提出的其他持续时间模型不相上下,同时为相同的提示和可变的表达水平提供可变的持续时间。最后,我们建议以一种新颖的方式对合成语音中的起搏和暂停进行直观控制的参数对菜花进行调节。摘要:Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to relying on previous statistical parametric speech synthesis technology for duration prediction, which poorly models the expressiveness and variability in speech. In this paper, we propose two alternate approaches to improve duration modelling. First, we propose a duration model conditioned on phrasing that improves the predicted durations and provides better modelling of pauses. We show that the duration model conditioned on phrasing improves the naturalness of speech over our baseline duration model. Second, we also propose a multi-speaker duration model called Cauliflow, that uses normalising flows to predict durations that better match the complex target duration distribution. Cauliflow performs on par with our other proposed duration model in terms of naturalness, whilst providing variable durations for the same prompt and variable levels of expressiveness. Lastly, we propose to condition Cauliflow on parameters that provide an intuitive control of the pacing and pausing in the synthesised speech in a novel way.
【13】 RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech Insertion
标题:RetrieverTTS:为基于文本的语音插入建模分解因子
链接:https://arxiv.org/abs/2206.13865
作者:Dacheng Yin,Chuanxin Tang,Yanqing Liu,Xiaoqiang Wang,Zhiyuan Zhao,Yucheng Zhao,Zhiwei Xiong,Sheng Zhao,Chong Luo机构:University of Science and Technology of China, Hefei, China, Microsoft Research Asia, Beijing China, Microsoft Azure Speech, Beijing, China备注:5 pages, 1 figure, 3 tables. Accepted by Interspeech 2022摘要:本文为基于文本的语音插入任务提出了一种新的“分解和编辑”范式,该范式有助于任意长度的语音插入甚至完整句子的生成。在所提出的范式中,语音中的全局和局部因素被显式分解并分别处理,以实现高说话人相似度和连续韵律。具体来说,我们提出用多个标记来表示全局因素,这些标记通过交叉注意操作提取,然后通过链接注意操作注入回。由于全局因素的丰富表示,我们设法以Zero-Shot的方式实现高说话人相似度。此外,我们还引入了一个韵律平滑任务,以使局部韵律因子具有上下文意识,从而实现令人满意的韵律连续性。通过对抗性训练阶段,我们进一步实现了高音质。在主观测试中,我们的方法在自然性和相似性方面都达到了最先进的性能。音频样本可在以下位置找到:https://ydcustc.github.io/retrieverTTS-demo/.摘要:This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrary-length speech insertion and even full sentence generation. In the proposed paradigm, global and local factors in speech are explicitly decomposed and separately manipulated to achieve high speaker similarity and continuous prosody. Specifically, we proposed to represent the global factors by multiple tokens, which are extracted by cross-attention operation and then injected back by link-attention operation. Due to the rich representation of global factors, we manage to achieve high speaker similarity in a zero-shot manner. In addition, we introduce a prosody smoothing task to make the local prosody factor context-aware and therefore achieve satisfactory prosody continuity. We further achieve high voice quality with an adversarial training stage. In the subjective test, our method achieves state-of-the-art performance in both naturalness and similarity. Audio samples can be found at https://ydcustc.github.io/retrieverTTS-demo/.
【14】 Speaker Verification in Multi-Speaker Environments Using Temporal Feature Fusion
标题:基于时间特征融合的多说话人环境下的说话人确认
链接:https://arxiv.org/abs/2206.13808
作者:Ahmad Aloradi,Wolfgang Mack,Mohamed Elminshawi,Emanuël A. P. Habets机构:International Audio Laboratories Erlangen, Am Wolfsmantel , Erlangen, Germany备注:To be presented at EUSIPCO 2022摘要:验证说话人的身份在现代人机界面中至关重要,例如,为了确保隐私保护或启用生物特征认证。经典说话人验证(SV)方法从对说话人语音特征进行编码的话语中估计出一个固定维度的嵌入。验证说话人的声音嵌入是否与所声称的说话人的嵌入足够相似。然而,这种方法假设输入中只存在一个说话人。同时发言的人可能会对表演产生不利影响。为了解决多说话人环境中的SV问题,我们提出了一种基于端到端深度学习的SV系统,用于检测输入中是否存在目标说话人。首先,从参考话语中估计嵌入来表示目标的特征。其次,根据输入混合估计帧级特征。然后将参考嵌入与混合特征进行逐帧融合,以便在帧的基础上区分目标与其他说话人。最后,利用融合后的特征预测目标说话人在语音片段中是否活跃。实验结果表明,该方法在多说话人情况下的性能优于x矢量方法。摘要:Verifying the identity of a speaker is crucial in modern human-machine interfaces, e.g., to ensure privacy protection or to enable biometric authentication. Classical speaker verification (SV) approaches estimate a fixed-dimensional embedding from a speech utterance that encodes the speaker's voice characteristics. A speaker is verified if his/her voice embedding is sufficiently similar to the embedding of the claimed speaker. However, such approaches assume that only a single speaker exists in the input. The presence of concurrent speakers is likely to have detrimental effects on the performance. To address SV in a multi-speaker environment, we propose an end-to-end deep learning-based SV system that detects whether the target speaker exists within an input or not. First, an embedding is estimated from a reference utterance to represent the target's characteristics. Second, frame-level features are estimated from the input mixture. The reference embedding is then fused frame-wise with the mixture's features to allow distinguishing the target from other speakers on a frame basis. Finally, the fused features are used to predict whether the target speaker is active in the speech segment or not. Experimental evaluation shows that the proposed method outperforms the x-vector in multi-speaker conditions.
【15】 Two Methods for Spoofing-Aware Speaker Verification: Multi-Layer Perceptron Score Fusion Model and Integrated Embedding Projector
标题:两种识别欺骗的说话人确认方法:多层感知器分数融合模型和集成嵌入投影仪
链接:https://arxiv.org/abs/2206.13807
作者:Jungwoo Heo,Ju-ho Kim,Hyun-seo Shin机构:School of Computer Science, University of Seoul, Republic of Korea备注:5 pages, 4 figures, 5 tables, accepted to 2022 Interspeech as a conference paper摘要:在过去的十年中,深度神经网络(DNN)的使用极大地提高了自动说话人验证(ASV)的性能。然而,ASV系统很容易被欺骗攻击所中和。因此,防欺骗说话人验证(SASV)挑战旨在通过整合ASV和欺骗对抗(CM)系统,促进能够在考虑欺骗攻击的情况下执行ASV的系统的开发。在本文中,我们提出了两个后端系统:多层感知器分数融合模型(MSFM)和集成嵌入式投影仪(IEP)。MSFM,得分融合后端系统,利用ASV和CM得分以及嵌入,得出SASV得分。另一方面,IEP将ASV和CM嵌入结合到SASV嵌入中,并基于余弦相似度计算最终的SASV分数。我们通过拟议的MSFM和IEP有效地集成了ASV和CM系统,并在SASV 2022挑战的官方评估试验中实现了SASV相等错误率0.56%、1.32%。摘要:The use of deep neural networks (DNN) has dramatically elevated the performance of automatic speaker verification (ASV) over the last decade. However, ASV systems can be easily neutralized by spoofing attacks. Therefore, the Spoofing-Aware Speaker Verification (SASV) challenge is designed and held to promote development of systems that can perform ASV considering spoofing attacks by integrating ASV and spoofing countermeasure (CM) systems. In this paper, we propose two back-end systems: multi-layer perceptron score fusion model (MSFM) and integrated embedding projector (IEP). The MSFM, score fusion back-end system, derived SASV score utilizing ASV and CM scores and embeddings. On the other hand,IEP combines ASV and CM embeddings into SASV embedding and calculates final SASV score based on the cosine similarity. We effectively integrated ASV and CM systems through proposed MSFM and IEP and achieved the SASV equal error rates 0.56%, 1.32% on the official evaluation trials of the SASV 2022 challenge.
【16】 Algorithms for audio inpainting based on probabilistic nonnegative matrix factorization
标题:基于概率非负矩阵分解的音频修复算法
链接:https://arxiv.org/abs/2206.13768
作者:Ondřej Mokrý,Paul Magron,Thomas Oberlin,Cédric Févotte机构:Brno University of Technology, Department of, Telecommunications, Technick´a , Brno, Czech Republic, Universit´e de Lorraine, CNRS, Inria, LORIA, F-, Nancy, France, ISAE-SUPAERO, Universit´e de Toulouse, France摘要:音频修复,即恢复丢失或被遮挡的音频信号样本的任务,通常依赖于稀疏表示或自回归建模。在本文中,我们建议在概率框架下用非负矩阵分解(NMF)构造谱图。首先,我们将缺失样本视为潜在变量,并根据我们是在时间域还是时频域中描述问题,推导出两种用于估计模型参数的期望最大化算法。然后,我们将缺失样本视为参数,并通过推导交替最小化方案来解决这个新问题。我们评估了这些算法在恢复音乐信号中短到中等长度间隔的任务中的潜力。实验表明,所提出的方法具有很好的收敛性,与最先进的音频修复技术相比,具有很强的竞争力。摘要:Audio inpainting, i.e., the task of restoring missing or occluded audio signal samples, usually relies on sparse representations or autoregressive modeling. In this paper, we propose to structure the spectrogram with nonnegative matrix factorization (NMF) in a probabilistic framework. First, we treat the missing samples as latent variables, and derive two expectation-maximization algorithms for estimating the parameters of the model, depending on whether we formulate the problem in the time- or time-frequency domain. Then, we treat the missing samples as parameters, and we address this novel problem by deriving an alternating minimization scheme. We assess the potential of these algorithms for the task of restoring short- to middle-length gaps in music signals. Experiments reveal great convergence properties of the proposed methods, as well as competitive performance when compared to state-of-the-art audio inpainting techniques.
【17】 A Hierarchical Speaker Representation Framework for One-shot Singing Voice Conversion
标题:一种一次歌唱语音转换的层次化说话人表示框架
链接:https://arxiv.org/abs/2206.13762
作者:Xu Li,Shansong Liu,Ying Shan备注:Accepted to INTERSPEECH 2022摘要:通常,歌唱语音转换(SVC)依赖于从说话人查找表(LUT)或说话人识别网络(SRN)提取的嵌入向量来建模说话人身份。然而,歌唱比会话语言包含更多的表达性说话人特征。有人怀疑,单个嵌入向量只能捕获平均和粗粒度的说话人特征,这对于SVC任务来说是不够的。为此,本文提出了一种新的SVC层次说话人表示框架,该框架可以捕获不同粒度的细粒度说话人特征。它由一个上采样流和三个下采样流组成。上采样流将语言特征转换为音频样本,而三个样本中的一个下采样流的操作方向相反。预计每个下采样块的时间统计信息可以代表不同粒度的说话人特征,这将用于上采样块以增强说话人建模。实验结果表明,该方法优于基于LUT和SRN的SVC系统。此外,该系统支持只需几秒钟参考音频的单次SVC。摘要:Typically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speaker identity. However, singing contains more expressive speaker characteristics than conversational speech. It is suspected that a single embedding vector may only capture averaged and coarse-grained speaker characteristics, which is insufficient for the SVC task. To this end, this work proposes a novel hierarchical speaker representation framework for SVC, which can capture fine-grained speaker characteristics at different granularity. It consists of an up-sampling stream and three down-sampling streams. The up-sampling stream transforms the linguistic features into audio samples, while one down-sampling stream of the three operates in the reverse direction. It is expected that the temporal statistics of each down-sampling block can represent speaker characteristics at different granularity, which will be engaged in the up-sampling blocks to enhance the speaker modeling. Experiment results verify that the proposed method outperforms both the LUT and SRN based SVC systems. Moreover, the proposed system supports the one-shot SVC with only a few seconds of reference audio.
【1】 Expressive, Variable, and Controllable Duration Modelling in TTS
标题:TTS中的表现性、可变性和可控性持续时间建模
链接:https://arxiv.org/abs/2206.14165
* 与cs.SD语音【12】为同一篇
作者:Ammar Abbas,Thomas Merritt,Alexis Moinet,Sri Karlapati,Ewa Muszynska,Simon Slangen,Elia Gatti,Thomas Drugman备注:Accepted to be published in the Proceedings of InterSpeech 2022摘要:随着非注意神经文语转换系统的兴起,持续时间建模再次成为一个重要的研究问题。当前的方法在很大程度上依赖于以前的统计参数语音合成技术进行持续时间预测,这对语音的表达性和可变性建模较差。在本文中,我们提出了两种改进持续时间建模的替代方法。首先,我们提出了一个以短语为条件的持续时间模型,该模型改进了预测的持续时间,并提供了更好的暂停建模。我们表明,基于短语的持续时间模型比我们的基线持续时间模型提高了语音的自然度。其次,我们还提出了一个称为Cailiflow的多说话人时长模型,该模型使用归一化流来预测更符合复杂目标时长分布的时长。Cauliflow在自然度方面与我们提出的其他持续时间模型不相上下,同时为相同的提示和可变的表达水平提供可变的持续时间。最后,我们建议以一种新颖的方式对合成语音中的起搏和暂停进行直观控制的参数对菜花进行调节。摘要:Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to relying on previous statistical parametric speech synthesis technology for duration prediction, which poorly models the expressiveness and variability in speech. In this paper, we propose two alternate approaches to improve duration modelling. First, we propose a duration model conditioned on phrasing that improves the predicted durations and provides better modelling of pauses. We show that the duration model conditioned on phrasing improves the naturalness of speech over our baseline duration model. Second, we also propose a multi-speaker duration model called Cauliflow, that uses normalising flows to predict durations that better match the complex target duration distribution. Cauliflow performs on par with our other proposed duration model in terms of naturalness, whilst providing variable durations for the same prompt and variable levels of expressiveness. Lastly, we propose to condition Cauliflow on parameters that provide an intuitive control of the pacing and pausing in the synthesised speech in a novel way.
【2】 RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech Insertion
标题:RetrieverTTS:为基于文本的语音插入建模分解因子
链接:https://arxiv.org/abs/2206.13865
* 与cs.SD语音【13】为同一篇
作者:Dacheng Yin,Chuanxin Tang,Yanqing Liu,Xiaoqiang Wang,Zhiyuan Zhao,Yucheng Zhao,Zhiwei Xiong,Sheng Zhao,Chong Luo机构:University of Science and Technology of China, Hefei, China, Microsoft Research Asia, Beijing China, Microsoft Azure Speech, Beijing, China备注:5 pages, 1 figure, 3 tables. Accepted by Interspeech 2022摘要:本文为基于文本的语音插入任务提出了一种新的“分解和编辑”范式,该范式有助于任意长度的语音插入甚至完整句子的生成。在所提出的范式中,语音中的全局和局部因素被显式分解并分别处理,以实现高说话人相似度和连续韵律。具体来说,我们提出用多个标记来表示全局因素,这些标记通过交叉注意操作提取,然后通过链接注意操作注入回。由于全局因素的丰富表示,我们设法以Zero-Shot的方式实现高说话人相似度。此外,我们还引入了一个韵律平滑任务,以使局部韵律因子具有上下文意识,从而实现令人满意的韵律连续性。通过对抗性训练阶段,我们进一步实现了高音质。在主观测试中,我们的方法在自然性和相似性方面都达到了最先进的性能。音频样本可在以下位置找到:https://ydcustc.github.io/retrieverTTS-demo/.摘要:This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrary-length speech insertion and even full sentence generation. In the proposed paradigm, global and local factors in speech are explicitly decomposed and separately manipulated to achieve high speaker similarity and continuous prosody. Specifically, we proposed to represent the global factors by multiple tokens, which are extracted by cross-attention operation and then injected back by link-attention operation. Due to the rich representation of global factors, we manage to achieve high speaker similarity in a zero-shot manner. In addition, we introduce a prosody smoothing task to make the local prosody factor context-aware and therefore achieve satisfactory prosody continuity. We further achieve high voice quality with an adversarial training stage. In the subjective test, our method achieves state-of-the-art performance in both naturalness and similarity. Audio samples can be found at https://ydcustc.github.io/retrieverTTS-demo/.
【3】 Speaker Verification in Multi-Speaker Environments Using Temporal Feature Fusion
标题:基于时间特征融合的多说话人环境下的说话人确认
链接:https://arxiv.org/abs/2206.13808
* 与cs.SD语音【14】为同一篇
作者:Ahmad Aloradi,Wolfgang Mack,Mohamed Elminshawi,Emanuël A. P. Habets机构:International Audio Laboratories Erlangen, Am Wolfsmantel , Erlangen, Germany备注:To be presented at EUSIPCO 2022摘要:验证说话人的身份在现代人机界面中至关重要,例如,为了确保隐私保护或启用生物特征认证。经典说话人验证(SV)方法从对说话人语音特征进行编码的话语中估计出一个固定维度的嵌入。验证说话人的声音嵌入是否与所声称的说话人的嵌入足够相似。然而,这种方法假设输入中只存在一个说话人。同时发言的人可能会对表演产生不利影响。为了解决多说话人环境中的SV问题,我们提出了一种基于端到端深度学习的SV系统,用于检测输入中是否存在目标说话人。首先,从参考话语中估计嵌入来表示目标的特征。其次,根据输入混合估计帧级特征。然后将参考嵌入与混合特征进行逐帧融合,以便在帧的基础上区分目标与其他说话人。最后,利用融合后的特征预测目标说话人在语音片段中是否活跃。实验结果表明,该方法在多说话人情况下的性能优于x矢量方法。摘要:Verifying the identity of a speaker is crucial in modern human-machine interfaces, e.g., to ensure privacy protection or to enable biometric authentication. Classical speaker verification (SV) approaches estimate a fixed-dimensional embedding from a speech utterance that encodes the speaker's voice characteristics. A speaker is verified if his/her voice embedding is sufficiently similar to the embedding of the claimed speaker. However, such approaches assume that only a single speaker exists in the input. The presence of concurrent speakers is likely to have detrimental effects on the performance. To address SV in a multi-speaker environment, we propose an end-to-end deep learning-based SV system that detects whether the target speaker exists within an input or not. First, an embedding is estimated from a reference utterance to represent the target's characteristics. Second, frame-level features are estimated from the input mixture. The reference embedding is then fused frame-wise with the mixture's features to allow distinguishing the target from other speakers on a frame basis. Finally, the fused features are used to predict whether the target speaker is active in the speech segment or not. Experimental evaluation shows that the proposed method outperforms the x-vector in multi-speaker conditions.
【4】 Two Methods for Spoofing-Aware Speaker Verification: Multi-Layer Perceptron Score Fusion Model and Integrated Embedding Projector
标题:两种识别欺骗的说话人确认方法:多层感知器分数融合模型和集成嵌入投影仪
链接:https://arxiv.org/abs/2206.13807
* 与cs.SD语音【15】为同一篇
作者:Jungwoo Heo,Ju-ho Kim,Hyun-seo Shin机构:School of Computer Science, University of Seoul, Republic of Korea备注:5 pages, 4 figures, 5 tables, accepted to 2022 Interspeech as a conference paper摘要:在过去的十年中,深度神经网络(DNN)的使用极大地提高了自动说话人验证(ASV)的性能。然而,ASV系统很容易被欺骗攻击所中和。因此,防欺骗说话人验证(SASV)挑战旨在通过整合ASV和欺骗对抗(CM)系统,促进能够在考虑欺骗攻击的情况下执行ASV的系统的开发。在本文中,我们提出了两个后端系统:多层感知器分数融合模型(MSFM)和集成嵌入式投影仪(IEP)。MSFM,得分融合后端系统,利用ASV和CM得分以及嵌入,得出SASV得分。另一方面,IEP将ASV和CM嵌入结合到SASV嵌入中,并基于余弦相似度计算最终的SASV分数。我们通过拟议的MSFM和IEP有效地集成了ASV和CM系统,并在SASV 2022挑战的官方评估试验中实现了SASV相等错误率0.56%、1.32%。摘要:The use of deep neural networks (DNN) has dramatically elevated the performance of automatic speaker verification (ASV) over the last decade. However, ASV systems can be easily neutralized by spoofing attacks. Therefore, the Spoofing-Aware Speaker Verification (SASV) challenge is designed and held to promote development of systems that can perform ASV considering spoofing attacks by integrating ASV and spoofing countermeasure (CM) systems. In this paper, we propose two back-end systems: multi-layer perceptron score fusion model (MSFM) and integrated embedding projector (IEP). The MSFM, score fusion back-end system, derived SASV score utilizing ASV and CM scores and embeddings. On the other hand,IEP combines ASV and CM embeddings into SASV embedding and calculates final SASV score based on the cosine similarity. We effectively integrated ASV and CM systems through proposed MSFM and IEP and achieved the SASV equal error rates 0.56%, 1.32% on the official evaluation trials of the SASV 2022 challenge.
【5】 Algorithms for audio inpainting based on probabilistic nonnegative matrix factorization
标题:基于概率非负矩阵分解的音频修复算法
链接:https://arxiv.org/abs/2206.13768
* 与cs.SD语音【16】为同一篇
作者:Ondřej Mokrý,Paul Magron,Thomas Oberlin,Cédric Févotte机构:Brno University of Technology, Department of, Telecommunications, Technick´a , Brno, Czech Republic, Universit´e de Lorraine, CNRS, Inria, LORIA, F-, Nancy, France, ISAE-SUPAERO, Universit´e de Toulouse, France摘要:音频修复,即恢复丢失或被遮挡的音频信号样本的任务,通常依赖于稀疏表示或自回归建模。在本文中,我们建议在概率框架下用非负矩阵分解(NMF)构造谱图。首先,我们将缺失样本视为潜在变量,并根据我们是在时间域还是时频域中描述问题,推导出两种用于估计模型参数的期望最大化算法。然后,我们将缺失样本视为参数,并通过推导交替最小化方案来解决这个新问题。我们评估了这些算法在恢复音乐信号中短到中等长度间隔的任务中的潜力。实验表明,所提出的方法具有很好的收敛性,与最先进的音频修复技术相比,具有很强的竞争力。摘要:Audio inpainting, i.e., the task of restoring missing or occluded audio signal samples, usually relies on sparse representations or autoregressive modeling. In this paper, we propose to structure the spectrogram with nonnegative matrix factorization (NMF) in a probabilistic framework. First, we treat the missing samples as latent variables, and derive two expectation-maximization algorithms for estimating the parameters of the model, depending on whether we formulate the problem in the time- or time-frequency domain. Then, we treat the missing samples as parameters, and we address this novel problem by deriving an alternating minimization scheme. We assess the potential of these algorithms for the task of restoring short- to middle-length gaps in music signals. Experiments reveal great convergence properties of the proposed methods, as well as competitive performance when compared to state-of-the-art audio inpainting techniques.
【6】 A Hierarchical Speaker Representation Framework for One-shot Singing Voice Conversion
标题:一种一次歌唱语音转换的层次化说话人表示框架
链接:https://arxiv.org/abs/2206.13762
* 与cs.SD语音【17】为同一篇
作者:Xu Li,Shansong Liu,Ying Shan备注:Accepted to INTERSPEECH 2022摘要:通常,歌唱语音转换(SVC)依赖于从说话人查找表(LUT)或说话人识别网络(SRN)提取的嵌入向量来建模说话人身份。然而,歌唱比会话语言包含更多的表达性说话人特征。有人怀疑,单个嵌入向量只能捕获平均和粗粒度的说话人特征,这对于SVC任务来说是不够的。为此,本文提出了一种新的SVC层次说话人表示框架,该框架可以捕获不同粒度的细粒度说话人特征。它由一个上采样流和三个下采样流组成。上采样流将语言特征转换为音频样本,而三个样本中的一个下采样流的操作方向相反。预计每个下采样块的时间统计信息可以代表不同粒度的说话人特征,这将用于上采样块以增强说话人建模。实验结果表明,该方法优于基于LUT和SRN的SVC系统。此外,该系统支持只需几秒钟参考音频的单次SVC。摘要:Typically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speaker identity. However, singing contains more expressive speaker characteristics than conversational speech. It is suspected that a single embedding vector may only capture averaged and coarse-grained speaker characteristics, which is insufficient for the SVC task. To this end, this work proposes a novel hierarchical speaker representation framework for SVC, which can capture fine-grained speaker characteristics at different granularity. It consists of an up-sampling stream and three down-sampling streams. The up-sampling stream transforms the linguistic features into audio samples, while one down-sampling stream of the three operates in the reverse direction. It is expected that the temporal statistics of each down-sampling block can represent speaker characteristics at different granularity, which will be engaged in the up-sampling blocks to enhance the speaker modeling. Experiment results verify that the proposed method outperforms both the LUT and SRN based SVC systems. Moreover, the proposed system supports the one-shot SVC with only a few seconds of reference audio.
【7】 Interrelate Training and Searching: A Unified Online Clustering Framework for Speaker Diarization
标题:相关训练和搜索:说话人二元化的统一在线聚类框架
链接:https://arxiv.org/abs/2206.13760
作者:Yifan Chen,Yifan Guo,Qingxuan Li,Gaofeng Cheng,Pengyuan Zhang,Yonghong Yan机构:Key Laboratory of Speech Acoustics and Content Understanding, Institute of Acoustics, Chinese, Academy of Sciences, China,University of Chinese Academy of Sciences, China, Tsinghua University, Beijing, China备注:Accepted by Interspeech 2022摘要:对于在线说话人日记,样本以增量方式到达,样本的总体分布是不可见的。此外,在大多数现有的基于聚类的方法中,嵌入提取器的训练目标并不是专门为聚类而设计的。为了提高在线说话人分类性能,我们提出了一个统一的在线聚类框架,该框架提供了嵌入提取器和聚类算法之间的交互方式。具体来说,该框架由两个高度耦合的部分组成:聚类引导的递归训练(CGRT)和截断波束搜索聚类(TBSC)。CGRT将聚类算法引入到嵌入提取器的训练过程中,不仅可以为嵌入提取器提供聚类感知信息,还可以为后续的聚类过程提供关键参数。通过这些包含度量空间初步信息的参数,TBSC对每个聚类的概率分数进行惩罚,以便以低延迟的在线方式输出更准确的聚类结果。通过上述创新,我们提出的在线聚类系统在AISHELL-4上以2.5s的延迟实现了14.48 \%的DER,衣领为0.25,而离线凝聚层次聚类的DER为14.57 \%。摘要:For online speaker diarization, samples arrive incrementally, and the overall distribution of the samples is invisible. Moreover, in most existing clustering-based methods, the training objective of the embedding extractor is not designed specially for clustering. To improve online speaker diarization performance, we propose a unified online clustering framework, which provides an interactive manner between embedding extractors and clustering algorithms. Specifically, the framework consists of two highly coupled parts: clustering-guided recurrent training (CGRT) and truncated beam searching clustering (TBSC). The CGRT introduces the clustering algorithm into the training process of embedding extractors, which could provide not only cluster-aware information for the embedding extractor, but also crucial parameters for the clustering process afterward. And with these parameters, which contain preliminary information of the metric space, the TBSC penalizes the probability score of each cluster, in order to output more accurate clustering results in online fashion with low latency. With the above innovations, our proposed online clustering system achieves 14.48\% DER with collar 0.25 at 2.5s latency on the AISHELL-4, while the DER of the offline agglomerative hierarchical clustering is 14.57\%.
【8】 Learning from human perception to improve automatic speaker verification in style-mismatched conditions
标题:从人类感知中学习以改进风格不匹配条件下的自动说话人确认
链接:https://arxiv.org/abs/2206.13684
作者:Amber Afshan,Abeer Alwan机构:Department of Electrical and Computer Engineering, University of California Los Angeles, USA备注:To appear in Interspeech, 2022摘要:我们之前的实验表明,人类和机器似乎采用不同的方法来辨别说话人,尤其是在存在说话风格变异的情况下。这些实验检验了阅读与会话的对比。听众在“把说话人说在一起”时关注说话人特有的特质,在“把说话人分开”时关注共享声学空间中的相对距离。然而,自动说话人验证(ASV)系统使用相同的损失函数,而不考虑目标或非目标试验。为了在风格变化的情况下提高ASV的性能,从人类感知中获得的见解被用于设计一个新的训练损失函数,我们称之为“CllrCE损失”。CllrCE loss使用特定于说话人的特性和说话人之间的相对声学距离来训练ASV系统。当使用加州大学洛杉矶分校说话人可变性数据库时,在x向量和调节设置中,与x向量基线相比,CllrCE损失导致EER的相对显著改善1-66%,minDCF的相对改善1-31%和1-56%。使用涉及不同会话言语任务的SITW评估任务,建议的损失与自我注意调节相结合,可使EER和minDCF的相对改善率分别比基线显著提高2-5%和6-12%。在SITW案例中,性能改善仅与调节一致。摘要:Our prior experiments show that humans and machines seem to employ different approaches to speaker discrimination, especially in the presence of speaking style variability. The experiments examined read versus conversational speech. Listeners focused on speaker-specific idiosyncrasies while "telling speakers together", and on relative distances in a shared acoustic space when "telling speakers apart". However, automatic speaker verification (ASV) systems use the same loss function irrespective of target or non-target trials. To improve ASV performance in the presence of style variability, insights learnt from human perception are used to design a new training loss function that we refer to as "CllrCE loss". CllrCE loss uses both speaker-specific idiosyncrasies and relative acoustic distances between speakers to train the ASV system. When using the UCLA speaker variability database, in the x-vector and conditioning setups, CllrCE loss results in significant relative improvements in EER by 1-66%, and minDCF by 1-31% and 1-56%, respectively, when compared to the x-vector baseline. Using the SITW evaluation tasks, which involve different conversational speech tasks, the proposed loss combined with self-attention conditioning results in significant relative improvements in EER by 2-5% and minDCF by 6-12% over baseline. In the SITW case, performance improvements were consistent only with conditioning.
【9】 Attention-based conditioning methods using variable frame rate for style-robust speaker verification
标题:风格稳健说话人验证的基于注意力的变帧速率条件化方法
链接:https://arxiv.org/abs/2206.13680
作者:Amber Afshan,Abeer Alwan机构:Department of Electrical and Computer Engineering, University of California Los Angeles, USA备注:To appear in Interspeech, 2022摘要:我们提出了一种在文本无关说话人验证中提取说话人嵌入的方法,该方法对说话人风格的变化具有鲁棒性。通常,说话人嵌入提取包括训练用于说话人分类的DNN,以及使用瓶颈特征作为说话人表示。这种网络有一个池层,通过计算所有话语帧的统计信息,以相等的权重将帧级特征转换为话语级特征。然而,自关注嵌入执行加权池,使得权重对应于说话人分类任务中帧的重要性。熵可以捕捉由于说话风格变化引起的声音变化。因此,提出了一种基于熵的可变帧速率向量作为自我注意层的外部条件向量,为网络提供可以解决风格效应的信息。这项工作探索了五种不同的调节方法。在12/23个任务中,与x向量基线相比,最佳调节方法(级联选通)在统计学上有显著改善,在使用加州大学洛杉矶分校说话人变异性数据库时,与11/23个任务中的基线相同。在9/23的任务中,它也显著优于无条件的自我注意,在1/23的任务中表现更差。该方法在情景会话的多说话人场景中也有显著的改进。摘要:We propose an approach to extract speaker embeddings that are robust to speaking style variations in text-independent speaker verification. Typically, speaker embedding extraction includes training a DNN for speaker classification and using the bottleneck features as speaker representations. Such a network has a pooling layer to transform frame-level to utterance-level features by calculating statistics over all utterance frames, with equal weighting. However, self-attentive embeddings perform weighted pooling such that the weights correspond to the importance of the frames in a speaker classification task. Entropy can capture acoustic variability due to speaking style variations. Hence, an entropy-based variable frame rate vector is proposed as an external conditioning vector for the self-attention layer to provide the network with information that can address style effects. This work explores five different approaches to conditioning. The best conditioning approach, concatenation with gating, provided statistically significant improvements over the x-vector baseline in 12/23 tasks and was the same as the baseline in 11/23 tasks when using the UCLA speaker variability database. It also significantly outperformed self-attention without conditioning in 9/23 tasks and was worse in 1/23. The method also showed significant improvements in multi-speaker scenarios of SITW.
【10】 Let the paintings play
标题:让这些画发挥作用
链接:https://arxiv.org/abs/2206.14142
* 与cs.SD语音【1】为同一篇
作者:Paola Gervasio,Alfio Quarteroni,Daniele Cassani机构:DICATAM, Universita degli Studi di Brescia, Brescia (Italy), MOX, Department of Mathematics, Politecnico di Milano, Milano (Italy) and Institute of, Mathematics, Ecole Polytechnique Federale de Lausanne, Lausanne (Switzerland), (professor emeritus)摘要:在本文中,我们介绍了一种数学方法来提取绘画和音乐曲目之间的相似性。我们的方法是基于绘画和音乐曲目的数字化,通过正交基函数的有限展开(使用傅立叶和小波基)。特定绘画与给定作曲家的音乐曲目样本之间的最佳匹配是通过有限维子空间上的$L^2$投影实现的。本文提供了几个例子,以分析意大利艺术家马塞洛·莫兰迪尼的艺术作品集。最后,我们开发了一个实现上述过程的原始applet,可以从网站免费下载https://github.com/pgerva/playing-paintings.git摘要:In this paper, we introduce a mathematical method to extract similarities between paintings and musical tracks. Our approach is based on the digitalization of both paintings and musical tracks by means of finite expansions in terms of orthogonal basis functions (with both Fourier and wavelet bases). The best fit between a specific painting and a sample of musical tracks from a given composer is achieved via an $L^2$ projection upon a finite-dimensional subspace. Several examples are provided for the analysis of a collection of works of art by the Italian artist Marcello Morandini. Finally, we have developed an original applet that implements the process above and which can be freely downloaded from the site https://github.com/pgerva/playing-paintings.git
【11】 Bengali Common Voice Speech Dataset for Automatic Speech Recognition
标题:用于自动语音识别的孟加拉普通语音数据集
链接:https://arxiv.org/abs/2206.14053
* 与cs.SD语音【2】为同一篇
作者:Samiul Alam,Asif Sushmit,Zaowad Abdullah,Shahrin Nakkhatra,MD. Nazmuddoha Ansary,Syed Mobassir Hossen,Sazia Morshed Mehnaz,Tahsin Reasat,Ahmed Imtiaz Humayun机构:Bengali.AI, Michigan State University, RPI, Vanderbilt University, Rice University, equal contribution摘要:孟加拉语是世界上说得最多的语言之一,全球有3亿多人说孟加拉语。尽管孟加拉语语音识别系统很受欢迎,但由于缺乏多样的开源数据集,孟加拉语语音识别系统的开发研究受到阻碍。作为前进的方向,我们已经众包了孟加拉共同语音语音数据集,这是一个句子级自动语音识别语料库。该数据集是在Mozilla Common Voice平台上收集的,是正在进行的一项活动的一部分,这项活动在两个月内收集了超过400小时的数据,并且正在迅速增长。我们的分析表明,与现有最大的开源孟加拉语音数据集OpenSLR孟加拉ASR数据集相比,我们的数据集具有更多的说话人、音素和环境多样性。我们展示了从数据集中获得的见解,并讨论了未来版本中需要解决的关键语言挑战。此外,我们报告了几种自动语音识别(ASR)算法的当前性能,并为未来的研究设定了基准。摘要:Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of diverse open-source datasets. As a way forward, we have crowdsourced the Bengali Common Voice Speech Dataset, which is a sentence-level automatic speech recognition corpus. Collected on the Mozilla Common Voice platform, the dataset is part of an ongoing campaign that has led to the collection of over 400 hours of data in 2 months and is growing rapidly. Our analysis shows that our dataset has more speaker, phoneme, and environmental diversity compared to the OpenSLR Bengali ASR dataset, the largest existing open-source Bengali speech dataset. We present insights obtained from the dataset and discuss key linguistic challenges that need to be addressed in future versions. Additionally, we report the current performance of a few Automatic Speech Recognition (ASR) algorithms and set a benchmark for future research.
【12】 Show Me Your Face, And I'll Tell You How You Speak
标题:把你的脸给我看,我就告诉你怎么说话
链接:https://arxiv.org/abs/2206.14009
* 与cs.SD语音【3】为同一篇
作者:Christen Millerdurai,Lotfy Abdel Khaliq,Timon Ulrich摘要:当我们说话时,可以从嘴唇的运动推断出讲话的韵律和内容。在这项工作中,我们探索了唇语合成的任务,即学习仅在说话人嘴唇运动的情况下生成语音,其中我们重点学习在无约束的大词汇量环境中多个说话人的精确唇语映射。我们通过说话人的面部特征(即年龄、性别、种族)捕捉其声音身份,并将其与嘴唇运动一起调节,以生成说话人身份感知语音。为此,我们提出了一种新的方法“Lip2Speech”,通过关键的设计选择,可以在无约束的场景中实现精确的唇语合成。我们还使用定量、定性指标和人类评估进行各种实验和广泛评估。摘要:When we speak, the prosody and content of the speech can be inferred from the movement of our lips. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate speech given only the lip movements of a speaker where we focus on learning accurate lip to speech mappings for multiple speakers in unconstrained, large vocabulary settings. We capture the speaker's voice identity through their facial characteristics, i.e., age, gender, ethnicity and condition them along with the lip movements to generate speaker identity aware speech. To this end, we present a novel method "Lip2Speech", with key design choices to achieve accurate lip to speech synthesis in unconstrained scenarios. We also perform various experiments and extensive evaluation using quantitative, qualitative metrics and human evaluation.
【13】 QTI Submission to DCASE 2021: residual normalization for device-imbalanced acoustic scene classification with efficient design
标题:QTI提交给DCASE 2021:高效设计的设备不平衡声场分类的残差归一化
链接:https://arxiv.org/abs/2206.13909
* 与cs.SD语音【5】为同一篇
作者:Byeonggeun Kim,Seunghan Yang,Jangho Kim,Simyung Chang机构:Qualcomm AI Research†, Qualcomm Korea YH, Seoul, Republic of Korea, Seoul National University, Seoul, Republic of Korea备注:tech report; won 1st place in DCASE2021 challenge. arXiv admin note: substantial text overlap with arXiv:2111.06531摘要:本技术报告描述了我们提交DCASE2021挑战的TASK1A的详细信息。本课题的目标是在模型复杂性的约束下,为设备不平衡数据集设计一个音频场景分类系统。本报告介绍了实现该目标的四种方法。首先,我们提出了残差归一化,这是一种新的特征归一化方法,它使用实例归一化和一条快捷路径来丢弃不必要的特定于设备的信息,而不会丢失有用的分类信息。其次,我们设计了一个高效的体系结构BC ResNet Mod,它是基线体系结构的一个修改版本,接收域有限。第三,我们利用从一个设备到多个设备的谱图到谱图的转换来增加训练数据。最后,我们使用三种模型压缩方案:剪枝、量化和知识提取来降低模型复杂性。该系统在TAU Urban Acoustic Scenes 2020 Mobile开发数据集中的平均测试精度为76.3%,具有315k参数,压缩到61.0KB非零参数后的平均测试精度为75.3%。摘要:This technical report describes the details of our TASK1A submission of the DCASE2021 challenge. The goal of the task is to design an audio scene classification system for device-imbalanced datasets under the constraints of model complexity. This report introduces four methods to achieve the goal. First, we propose Residual Normalization, a novel feature normalization method that uses instance normalization with a shortcut path to discard unnecessary device-specific information without losing useful information for classification. Second, we design an efficient architecture, BC-ResNet-Mod, a modified version of the baseline architecture with a limited receptive field. Third, we exploit spectrogram-to-spectrogram translation from one to multiple devices to augment training data. Finally, we utilize three model compression schemes: pruning, quantization, and knowledge distillation to reduce model complexity. The proposed system achieves an average test accuracy of 76.3% in TAU Urban Acoustic Scenes 2020 Mobile, development dataset with 315k parameters, and average test accuracy of 75.3% after compression to 61.0KB of non-zero parameters.
【14】 Comparison of Speech Representations for the MOS Prediction System
标题:MOS预报系统中语音表示方法的比较
链接:https://arxiv.org/abs/2206.13817
* 与cs.SD语音【6】为同一篇
作者:Aki Kunikoshi,Jaebok Kim,Wonsuk Jun,Kåre Sjölander摘要:为了保证文语转换系统的质量,研究了自动预测听者平均意见分数(MOS)的方法。以前的许多研究侧重于架构方面的进步(如MBNet、LDNet等),以更有效的方式捕获光谱特征与MOS之间的关系,并实现高精度。然而,在泛化能力方面的最优表示在很大程度上仍然是未知的。为此,我们将通过wav2vec框架获得的自监督学习(SSL)特征的性能与光谱特征的性能进行了比较,如光谱图和melspectrogram的幅度。此外,我们建议将SSL功能与我们认为可以为自动MOS保留基本信息的功能相结合,以弥补各自的缺点。我们对从过去的暴风雪和语音转换挑战中收集的大规模听力测试语料库进行了全面的实验。我们发现,wav2vec特征集显示出最好的泛化能力,即使给定的基本事实并不总是可靠的。此外,我们发现这些组合表现最好,并分析了它们如何弥合光谱和wav2vec特征集之间的差距。摘要:Automatic methods to predict Mean Opinion Score (MOS) of listeners have been researched to assure the quality of Text-to-Speech systems. Many previous studies focus on architectural advances (e.g. MBNet, LDNet, etc.) to capture relations between spectral features and MOS in a more effective way and achieved high accuracy. However, the optimal representation in terms of generalization capability still largely remains unknown. To this end, we compare the performance of Self-Supervised Learning (SSL) features obtained by the wav2vec framework to that of spectral features such as magnitude of spectrogram and melspectrogram. Moreover, we propose to combine the SSL features and features which we believe to retain essential information to the automatic MOS to compensate each other for their drawbacks. We conduct comprehensive experiments on a large-scale listening test corpus collected from past Blizzard and Voice Conversion Challenges. We found that the wav2vec feature set showed the best generalization even though the given ground-truth was not always reliable. Furthermore, we found that the combinations performed the best and analyzed how they bridged the gap between spectral and the wav2vec feature sets.
【15】 Exploring linguistic feature and model combination for speech recognition based automatic AD detection
标题:基于语音识别的AD自动检测的语言特征和模型组合研究
链接:https://arxiv.org/abs/2206.13758
作者:Yi Wang,Tianzi Wang,Zi Ye,Lingwei Meng,Shoukang Hu,Xixin Wu,Xunying Liu,Helen Meng机构:The Chinese University of Hong Kong, Hong Kong SAR, China备注:Accepted by INTERSPEECH 2022摘要:阿尔茨海默病(AD)的早期诊断对于促进预防护理和延缓病情发展至关重要。基于语音的自动广告筛查系统为其他临床筛查技术提供了一种非侵入性和更具可扩展性的替代方法。在开发此类系统时,缺乏此类专业数据会导致模型选择和特征学习的不确定性。为此,本文研究了如何使用特征和模型组合方法来提高BERT和Roberta预训练文本编码器在有限数据上进行域微调的鲁棒性,然后将生成的嵌入特征反馈到后端分类器集合中,通过多数投票产生最终的广告检测决策。在ADReSS20挑战数据集上进行的实验表明,在系统开发中使用模型和特征组合可以获得一致的性能改进。在由48名老年人组成的ADReSS20测试集上,分别使用手动和ASR语音记录获得了91.67%和93.75%的最先进AD检测准确率。摘要:Early diagnosis of Alzheimer's disease (AD) is crucial in facilitating preventive care and delay progression. Speech based automatic AD screening systems provide a non-intrusive and more scalable alternative to other clinical screening techniques. Scarcity of such specialist data leads to uncertainty in both model selection and feature learning when developing such systems. To this end, this paper investigates the use of feature and model combination approaches to improve the robustness of domain fine-tuning of BERT and Roberta pre-trained text encoders on limited data, before the resulting embedding features being fed into an ensemble of backend classifiers to produce the final AD detection decision via majority voting. Experiments conducted on the ADReSS20 Challenge dataset suggest consistent performance improvements were obtained using model and feature combination in system development. State-of-the-art AD detection accuracies of 91.67 percent and 93.75 percent were obtained using manual and ASR speech transcripts respectively on the ADReSS20 test set consisting of 48 elderly speakers.
【16】 Personalized Keyword Spotting through Multi-task Learning
标题:基于多任务学习的个性化关键词识别
链接:https://arxiv.org/abs/2206.13708
* 与cs.SD语音【7】为同一篇
作者:Seunghan Yang,Byeonggeun Kim,Inseop Chung,Simyung Chang机构:Qualcomm AI Research†, Qualcomm Korea YH, Seoul, Republic of Korea, Seoul National University, Seoul, Republic of Korea备注:Proceedings of INTERSPEECH 2022摘要:关键词发现(KWS)在智能设备上实现基于语音的用户交互方面起着至关重要的作用,而传统的KWS(C-KWS)方法主要用于检测用户不可知的预定义关键词。然而,在实践中,大多数用户交互都来自注册到设备中的目标用户,这促使他们构建个性化的关键字定位。我们设计了两个个性化的KWS任务;(1) 目标用户偏向KWS(TB-KWS)和(2)目标用户专用KWS(TO-KWS)。为了解决这些任务,我们提出了通过多任务学习(PK-MTL)进行个性化关键字发现,该方法包括多任务学习和任务适应。首先,我们介绍了将多任务学习应用于关键词识别和说话人验证,以利用用户信息来实现关键词识别系统。接下来,我们设计了特定于任务的评分函数,以完全适应个性化的KWS任务。我们在传统和个性化场景下评估了我们的框架,结果表明PK-MTL可以显著降低误报率,尤其是在各种实际场景中。摘要:Keyword spotting (KWS) plays an essential role in enabling speech-based user interaction on smart devices, and conventional KWS (C-KWS) approaches have concentrated on detecting user-agnostic pre-defined keywords. However, in practice, most user interactions come from target users enrolled in the device which motivates to construct personalized keyword spotting. We design two personalized KWS tasks; (1) Target user Biased KWS (TB-KWS) and (2) Target user Only KWS (TO-KWS). To solve the tasks, we propose personalized keyword spotting through multi-task learning (PK-MTL) that consists of multi-task learning and task-adaptation. First, we introduce applying multi-task learning on keyword spotting and speaker verification to leverage user information to the keyword spotting system. Next, we design task-specific scoring functions to adapt to the personalized KWS tasks thoroughly. We evaluate our framework on conventional and personalized scenarios, and the results show that PK-MTL can dramatically reduce the false alarm rate, especially in various practical scenarios.
【17】 Domain Agnostic Few-shot Learning for Speaker Verification
标题:用于说话人确认的领域不可知的Few-Shot学习
链接:https://arxiv.org/abs/2206.13700
* 与cs.SD语音【8】为同一篇
作者:Seunghan Yang,Debasmit Das,Janghoon Cho,Hyoungwoo Park,Sungrack Yun备注:Proceedings of INTERSPEECH 2022摘要:验证系统的深度学习模型通常无法推广到新用户和新环境,即使它们学习到高度区分的特征。为了解决这个问题,我们提出了一个简单的领域泛化框架,该框架可以学习处理新用户和新领域的分布转移。我们的框架由特定领域和领域聚合网络组成,它们分别是特定领域和组合领域的专家。通过使用这些网络,我们可以在训练阶段模拟新用户和新域的存在,从而最终产生更好的泛化。为了节省内存,我们通过将相似的域聚集在一起来减少特定于域的网络的数量。在对人工生成的噪声域进行广泛评估后,我们可以明确地显示我们的框架的泛化能力。此外,我们将我们提出的方法应用于标准基准上的现有竞争体系结构,这显示了进一步的性能改进。摘要:Deep learning models for verification systems often fail to generalize to new users and new environments, even though they learn highly discriminative features. To address this problem, we propose a few-shot domain generalization framework that learns to tackle distribution shift for new users and new domains. Our framework consists of domain-specific and domain-aggregation networks, which are the experts on specific and combined domains, respectively. By using these networks, we generate episodes that mimic the presence of both novel users and novel domains in the training phase to eventually produce better generalization. To save memory, we reduce the number of domain-specific networks by clustering similar domains together. Upon extensive evaluation on artificially generated noise domains, we can explicitly show generalization ability of our framework. In addition, we apply our proposed methods to the existing competitive architecture on the standard benchmark, which shows further performance improvements.
【18】 Dummy Prototypical Networks for Few-Shot Open-Set Keyword Spotting
标题:Few-Shot开放集关键词识别的虚拟原型网络
链接:https://arxiv.org/abs/2206.13691
* 与cs.SD语音【9】为同一篇
作者:Byeonggeun Kim,Seunghan Yang,Inseop Chung,Simyung Chang机构:Qualcomm AI Research†, Qualcomm Korea YH, Seoul, Republic of Korea, Seoul National University, Seoul, Republic of Korea备注:Proceedings of INTERSPEECH 2022摘要:关键字定位是在流式音频中检测关键字的任务。传统的关键字定位以预定义的关键字分类为目标,但在很少的快照(通过示例查询)关键字定位中,例如,给定M-shot支持样本的N向分类,受到了越来越多的关注。此外,在现实场景中,可能存在来自意外类别(开放集)的话语,需要拒绝这些话语,而不是将其归类为N个类别之一。结合这两个需求,我们使用一个新的基准设置splitGSC来解决少数镜头开放集关键字发现问题。我们提出了基于度量学习的已知虚拟原型,以更好地检测开放集,并介绍了一种简单而强大的方法,虚拟原型网络(D-ProtoNets)。与最近提出的splitGSC中的少数镜头开放集识别(FSOSR)方法相比,我们的D-ProtoNets显示出明显的优势。我们还在标准基准上验证了我们的方法,miniImageNet和D-ProtoNets显示了FSOSR中最先进的开放集检测率。摘要:Keyword spotting is the task of detecting a keyword in streaming audio. Conventional keyword spotting targets predefined keywords classification, but there is growing attention in few-shot (query-by-example) keyword spotting, e.g., N-way classification given M-shot support samples. Moreover, in real-world scenarios, there can be utterances from unexpected categories (open-set) which need to be rejected rather than classified as one of the N classes. Combining the two needs, we tackle few-shot open-set keyword spotting with a new benchmark setting, named splitGSC. We propose episode-known dummy prototypes based on metric learning to detect an open-set better and introduce a simple and powerful approach, Dummy Prototypical Networks (D-ProtoNets). Our D-ProtoNets shows clear margins compared to recent few-shot open-set recognition (FSOSR) approaches in the suggested splitGSC. We also verify our method on a standard benchmark, miniImageNet, and D-ProtoNets shows the state-of-the-art open-set detection rate in FSOSR.
【19】 Tiny-Sepformer: A Tiny Time-Domain Transformer Network for Speech Separation
标题:微型隔离器:一种用于语音分离的微型时域变换网络
链接:https://arxiv.org/abs/2206.13689
* 与cs.SD语音【10】为同一篇
作者:Jian Luo,Jianzong Wang,Ning Cheng,Edward Xiao,Xulong Zhang,Jing Xiao机构:. Ping An Technology (Shenzhen) Co., Ltd. ,. Aquinas International Academy, CA, USA备注:Accepted by Interspeech 2022摘要:时域变换神经网络在语音分离任务中已经证明了其优越性。然而,这些模型通常具有大量的网络参数,因此经常遇到GPU内存爆炸的问题。在本文中,我们提出了用于语音分离的微型Transformer网络——微型Sepformer。我们提出了两种降低模型参数和内存消耗的技术:(1)卷积注意(CA)块,将vanilla Transformer拆分为两条路径,多头注意和一维深度可分离卷积,(2)参数共享,共享CA块内的层参数。在我们的实验中,Tiny Sepformer可以大大减小模型尺寸,并在WSJ0-2/3Mix数据集上实现与vanilla Sepformer相当的分离性能。摘要:Time-domain Transformer neural networks have proven their superiority in speech separation tasks. However, these models usually have a large number of network parameters, thus often encountering the problem of GPU memory explosion. In this paper, we proposed Tiny-Sepformer, a tiny version of Transformer network for speech separation. We present two techniques to reduce the model parameters and memory consumption: (1) Convolution-Attention (CA) block, spliting the vanilla Transformer to two paths, multi-head attention and 1D depthwise separable convolution, (2) parameter sharing, sharing the layer parameters within the CA block. In our experiments, Tiny-Sepformer could greatly reduce the model size, and achieves comparable separation performance with vanilla Sepformer on WSJ0-2/3Mix datasets.
【20】 ClearBuds: Wireless Binaural Earbuds for Learning-Based Speech Enhancement
标题:ClearBuds:用于基于学习的语音增强的无线双耳耳机
链接:https://arxiv.org/abs/2206.13611
* 与cs.SD语音【11】为同一篇
作者:Ishan Chatterjee,Maruchi Kim,Vivek Jayaram,Shyamnath Gollakota,Ira Kemelmacher-Shlizerman,Shwetak Patel,Steven M. Seitz机构:∗Co-primary student authors, Paul G. Allen School of Computer Science & Engineering, University of Washington, Seattle, WA, USA备注:12 pages, Published in Mobisys 2022摘要:我们介绍了ClearBubs,这是第一个利用神经网络来增强来自两个无线耳塞的语音流的硬件和软件系统。无线耳塞的实时语音增强需要高质量的声音分离和背景消除,并在移动电话上实时运行。Clear Buds为盲声源分离和入耳式移动系统的最新深度学习搭建了桥梁,做出了两项关键技术贡献:1)一种新的无线耳塞设计,能够作为同步双耳麦克风阵列运行;2)一种在移动设备上运行的轻型双通道语音增强神经网络。我们的神经网络具有一种新颖的级联结构,它将时域传统神经网络与基于频谱图的频率掩蔽神经网络相结合,以减少音频输出中的伪影。结果表明,我们的无线耳塞实现的同步误差小于64微秒,我们的网络在附带的手机上的运行时间为21.4毫秒。在对8名用户在之前未见过的室内和室外多路径场景中进行的野外评估中,我们的神经网络可以学习空间和声学线索,以执行噪声抑制和背景语音消除。在一项有37名参与者参与的用户研究中,他们花了超过15.4小时对野外采集的1041个音频样本进行评级,我们的系统实现了改进的平均意见评分和背景噪声抑制。包含演示的项目页面:https://clearbuds.cs.washington.edu摘要:We present ClearBuds, the first hardware and software system that utilizes a neural network to enhance speech streamed from two wireless earbuds. Real-time speech enhancement for wireless earbuds requires high-quality sound separation and background cancellation, operating in real-time and on a mobile phone. Clear-Buds bridges state-of-the-art deep learning for blind audio source separation and in-ear mobile systems by making two key technical contributions: 1) a new wireless earbud design capable of operating as a synchronized, binaural microphone array, and 2) a lightweight dual-channel speech enhancement neural network that runs on a mobile device. Our neural network has a novel cascaded architecture that combines a time-domain conventional neural network with a spectrogram-based frequency masking neural network to reduce the artifacts in the audio output. Results show that our wireless earbuds achieve a synchronization error less than 64 microseconds and our network has a runtime of 21.4 milliseconds on an accompanying mobile phone. In-the-wild evaluation with eight users in previously unseen indoor and outdoor multipath scenarios demonstrates that our neural network generalizes to learn both spatial and acoustic cues to perform noise suppression and background speech removal. In a user-study with 37 participants who spent over 15.4 hours rating 1041 audio samples collected in-the-wild, our system achieves improved mean opinion score and background noise suppression. Project page with demos: https://clearbuds.cs.washington.edu
机器翻译,仅供参考