cs.SD语音,共计4篇,eess.AS音频处理,共计5篇


1.cs.SD语音:

【1】 Multiple Offsets Multilateration: a new paradigm for sensor network calibration with unsynchronized reference nodes

标题:多偏移多边化法:一种参考节点不同步的传感器网络校准新范式

链接:https://arxiv.org/abs/2205.11299

作者:Luca Ferranti,Kalle Åström,Magnus Oskarsson,Jani Boutellier,Juho Kannala
机构:University of Vaasa, Vaasa, Finland, bLund University, Lund, Sweden, Aalto University, Espoo, Finland备注:accepted to ICASSP2022
摘要:使用波信号测量进行定位在多个应用中,例如GPS系统、声音结构和基于Wifi的定位。从数学上讲,这些问题需要计算接收器和/或发射器的位置以及设备不同步时的时间偏移。在本文中,我们通过引入多偏移量多侧化(MOM)扩展了先前最新的定位公式,MOM是一种新的数学框架,用于计算已知位置的非同步参考发射机的伪距接收机位置。这可以应用于多种情况,例如声音结构和低轨卫星定位。我们从数学上描述了矩量法,确定了网络可解所需的接收器和发射器数量,研究了可能的不同解的数量,导出了基于同伦延拓的稳定解。该解算器对合成音频数据和真实音频数据都具有高效和鲁棒性。
摘要:Positioning using wave signal measurements is used in several applications, such as GPS systems, structure from sound and Wifi based positioning. Mathematically, such problems require the computation of the positions of receivers and/or transmitters as well as time offsets if the devices are unsynchronized. In this paper, we expand the previous state-of-the-art on positioning formulations by introducing Multiple Offsets Multilateration (MOM), a new mathematical framework to compute the receivers positions with pseudoranges from unsynchronized reference transmitters at known positions. This could be applied in several scenarios, for example structure from sound and positioning with LEO satellites. We mathematically describe MOM, determining how many receivers and transmitters are needed for the network to be solvable, a study on the number of possible distinct solutions is presented and stable solvers based on homotopy continuation are derived. The solvers are shown to be efficient and robust to noise both for synthetic and real audio data.


【2】 Calibrate and Refine! A Novel and Agile Framework for ASR-error Robust Intent Detection

标题:校准和改进!一种新颖灵活的ASR错误健壮意图检测框架

链接:https://arxiv.org/abs/2205.11008

作者:Peilin Zhou,Dading Chong,Helin Wang,Qingcheng Zeng
机构:Zhejiang University, Hangzhou, China, Peking University, Shenzhen, China
备注:Submit to INTERSPEECH 2022
摘要:在过去的十年中,基于文本的意图检测得到了快速发展,深度学习技术已经将其基准性能提升到了一个显著的水平。然而,由于环境噪声、独特的语音模式等原因,自动语音识别(ASR)错误在实际应用中不可避免,导致最先进的基于文本的意图检测模型的性能急剧下降。本质上,这种现象是由ASR错误带来的语义漂移造成的,现有的大多数工作倾向于设计新的模型结构以减少其影响,这是以牺牲通用性和灵活性为代价的。与以往的整体模型不同,本文提出了一种新的、灵活的ASR错误鲁棒意图检测框架CR-ID,该框架包含两个即插即用模块,即语义漂移校准模块(SDCM)和音位细化模块(PRM),这两个模型都是不可知的,因此可以很容易地集成到任何现有的意图检测模型中,而无需修改其结构。在SNIPS数据集上的实验结果表明,我们提出的CR-ID框架在ASR输出上取得了有竞争力的性能,优于所有基线方法,这验证了CR-ID可以有效地缓解ASR错误引起的语义漂移。
摘要:The past ten years have witnessed the rapid development of text-based intent detection, whose benchmark performances have already been taken to a remarkable level by deep learning techniques. However, automatic speech recognition (ASR) errors are inevitable in real-world applications due to the environment noise, unique speech patterns and etc, leading to sharp performance drop in state-of-the-art text-based intent detection models. Essentially, this phenomenon is caused by the semantic drift brought by ASR errors and most existing works tend to focus on designing new model structures to reduce its impact, which is at the expense of versatility and flexibility. Different from previous one-piece model, in this paper, we propose a novel and agile framework called CR-ID for ASR error robust intent detection with two plug-and-play modules, namely semantic drift calibration module (SDCM) and phonemic refinement module (PRM), which are both model-agnostic and thus could be easily integrated to any existing intent detection models without modifying their structures. Experimental results on SNIPS dataset show that, our proposed CR-ID framework achieves competitive performance and outperform all the baseline methods on ASR outputs, which verifies that CR-ID can effectively alleviate the semantic drift caused by ASR errors.


【3】 Self-Supervised Speech Representation Learning: A Review

标题:自监督语音表征学习:综述

链接:https://arxiv.org/abs/2205.10643

作者:Abdelrahman Mohamed,Hung-yi Lee,Lasse Borgholt,Jakob D. Havtorn,Joakim Edin,Christian Igel,Katrin Kirchhoff,Shang-Wen Li,Karen Livescu,Lars Maaløe,Tara N. Sainath,Shinji Watanabe
摘要:尽管有监督的深度学习已经彻底改变了语音和音频处理,但它需要为单个任务和应用场景构建专家模型。同样,很难将这一点应用于只有有限标记数据可用的方言和语言。自监督表征学习方法承诺了一个单一的通用模型,这将有利于广泛的任务和领域。这些方法在自然语言处理和计算机视觉领域取得了成功,实现了新的性能水平,同时减少了许多下游场景所需的标签数量。言语表征学习在生成法、对比法和预测法三大类中都有类似的进展。其他方法依赖于多模态数据进行预训练,将文本或视觉数据流与语音混合。尽管自监督语音表征仍然是一个新兴的研究领域,但它与声学单词嵌入和零词汇资源学习密切相关,这两个领域的研究多年来都很活跃。本文综述了自监督语音表征学习的方法及其与其他研究领域的联系。由于当前的许多方法仅将自动语音识别作为一项下游任务,因此我们回顾了最近对学习表示进行基准测试的工作,以将应用扩展到语音识别之外。
摘要:Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios. It is likewise difficult to apply this to dialects and languages for which only limited labeled data is available. Self-supervised representation learning methods promise a single universal model that would benefit a wide variety of tasks and domains. Such methods have shown success in natural language processing and computer vision domains, achieving new levels of performance while reducing the number of labels required for many downstream scenarios. Speech representation learning is experiencing similar progress in three main categories: generative, contrastive, and predictive methods. Other approaches rely on multi-modal data for pre-training, mixing text or visual data streams with speech. Although self-supervised speech representation is still a nascent research area, it is closely related to acoustic word embedding and learning with zero lexical resources, both of which have seen active research for many years. This review presents approaches for self-supervised speech representation learning and their connection to other research areas. Since many current methods focus solely on automatic speech recognition as a downstream task, we review recent efforts on benchmarking learned representations to extend the application beyond speech recognition.


【4】 NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement

标题:NeuralEcho:一种用于声学回波抑制和语音增强的自关注递归神经网络

链接:https://arxiv.org/abs/2205.10401

作者:Meng Yu,Yong Xu,Chunlei Zhang,Shi-Xiong Zhang,Dong Yu
机构:Tencent AI Lab, Bellevue, WA, USA
备注:Submitted to INTERSPEECH 2022
摘要:声回声消除(AEC)在全双工语音通信中起着重要作用,在扬声器回放的情况下,它可以增强前端语音以进行识别。本文提出了一个全深度学习框架,隐式估计回波/噪声和目标语音的二阶统计量,并通过基于注意的递归神经网络联合解决回波和噪声抑制问题。该模型在客观语音质量指标、语音识别准确率和模型复杂度方面优于最新的联合回声消除和语音增强方法F-T-LSTM。我们证明,该模型可以与说话人嵌入一起更好地实现目标语音增强,并进一步为自动增益控制(AGC)任务开发了一个分支,以形成一个一体化的前端语音增强系统。
摘要:Acoustic echo cancellation (AEC) plays an important role in the full-duplex speech communication as well as the front-end speech enhancement for recognition in the conditions when the loudspeaker plays back. In this paper, we present an all-deep-learning framework that implicitly estimates the second order statistics of echo/noise and target speech, and jointly solves echo and noise suppression through an attention based recurrent neural network. The proposed model outperforms the state-of-the-art joint echo cancellation and speech enhancement method F-T-LSTM in terms of objective speech quality metrics, speech recognition accuracy and model complexity. We show that this model can work with speaker embedding for better target speech enhancement and furthermore develop a branch for automatic gain control (AGC) task to form an all-in-one front-end speech enhancement system.


2.eess.AS音频处理:

【1】 NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement

标题:NeuralEcho:一种用于声学回波抑制和语音增强的自关注递归神经网络

链接:https://arxiv.org/abs/2205.10401

作者:Meng Yu,Yong Xu,Chunlei Zhang,Shi-Xiong Zhang,Dong Yu
机构:Tencent AI Lab, Bellevue, WA, USA
备注:Submitted to INTERSPEECH 2022
摘要:声回声消除(AEC)在全双工语音通信中起着重要作用,在扬声器回放的情况下,它可以增强前端语音以进行识别。本文提出了一个全深度学习框架,隐式估计回波/噪声和目标语音的二阶统计量,并通过基于注意的递归神经网络联合解决回波和噪声抑制问题。该模型在客观语音质量指标、语音识别准确率和模型复杂度方面优于最新的联合回声消除和语音增强方法F-T-LSTM。我们证明,该模型可以与说话人嵌入一起更好地实现目标语音增强,并进一步为自动增益控制(AGC)任务开发了一个分支,以形成一个一体化的前端语音增强系统。
摘要:Acoustic echo cancellation (AEC) plays an important role in the full-duplex speech communication as well as the front-end speech enhancement for recognition in the conditions when the loudspeaker plays back. In this paper, we present an all-deep-learning framework that implicitly estimates the second order statistics of echo/noise and target speech, and jointly solves echo and noise suppression through an attention based recurrent neural network. The proposed model outperforms the state-of-the-art joint echo cancellation and speech enhancement method F-T-LSTM in terms of objective speech quality metrics, speech recognition accuracy and model complexity. We show that this model can work with speaker embedding for better target speech enhancement and furthermore develop a branch for automatic gain control (AGC) task to form an all-in-one front-end speech enhancement system.


【2】 Multiple Offsets Multilateration: a new paradigm for sensor network calibration with unsynchronized reference nodes

标题:多偏移多边化法:一种参考节点不同步的传感器网络校准新范式

链接:https://arxiv.org/abs/2205.11299

作者:Luca Ferranti,Kalle Åström,Magnus Oskarsson,Jani Boutellier,Juho Kannala
机构:University of Vaasa, Vaasa, Finland, bLund University, Lund, Sweden, Aalto University, Espoo, Finland
备注:accepted to ICASSP2022
摘要:使用波信号测量进行定位在多个应用中,例如GPS系统、声音结构和基于Wifi的定位。从数学上讲,这些问题需要计算接收器和/或发射器的位置以及设备不同步时的时间偏移。在本文中,我们通过引入多偏移量多侧化(MOM)扩展了先前最新的定位公式,MOM是一种新的数学框架,用于计算已知位置的非同步参考发射机的伪距接收机位置。这可以应用于多种情况,例如声音结构和低轨卫星定位。我们从数学上描述了矩量法,确定了网络可解所需的接收器和发射器数量,研究了可能的不同解的数量,导出了基于同伦延拓的稳定解。该解算器对合成音频数据和真实音频数据都具有高效和鲁棒性。
摘要:Positioning using wave signal measurements is used in several applications, such as GPS systems, structure from sound and Wifi based positioning. Mathematically, such problems require the computation of the positions of receivers and/or transmitters as well as time offsets if the devices are unsynchronized. In this paper, we expand the previous state-of-the-art on positioning formulations by introducing Multiple Offsets Multilateration (MOM), a new mathematical framework to compute the receivers positions with pseudoranges from unsynchronized reference transmitters at known positions. This could be applied in several scenarios, for example structure from sound and positioning with LEO satellites. We mathematically describe MOM, determining how many receivers and transmitters are needed for the network to be solvable, a study on the number of possible distinct solutions is presented and stable solvers based on homotopy continuation are derived. The solvers are shown to be efficient and robust to noise both for synthetic and real audio data.


【3】 Calibrate and Refine! A Novel and Agile Framework for ASR-error Robust Intent Detection

标题:校准和改进!一种新颖灵活的ASR错误健壮意图检测框架

链接:https://arxiv.org/abs/2205.11008

作者:Peilin Zhou,Dading Chong,Helin Wang,Qingcheng Zeng
机构:Zhejiang University, Hangzhou, China, Peking University, Shenzhen, China
备注:Submit to INTERSPEECH 2022
摘要:在过去的十年中,基于文本的意图检测得到了快速发展,深度学习技术已经将其基准性能提升到了一个显著的水平。然而,由于环境噪声、独特的语音模式等原因,自动语音识别(ASR)错误在实际应用中不可避免,导致最先进的基于文本的意图检测模型的性能急剧下降。本质上,这种现象是由ASR错误带来的语义漂移造成的,现有的大多数工作倾向于设计新的模型结构以减少其影响,这是以牺牲通用性和灵活性为代价的。与以往的整体模型不同,本文提出了一种新的、灵活的ASR错误鲁棒意图检测框架CR-ID,该框架包含两个即插即用模块,即语义漂移校准模块(SDCM)和音位细化模块(PRM),这两个模型都是不可知的,因此可以很容易地集成到任何现有的意图检测模型中,而无需修改其结构。在SNIPS数据集上的实验结果表明,我们提出的CR-ID框架在ASR输出上取得了有竞争力的性能,优于所有基线方法,这验证了CR-ID可以有效地缓解ASR错误引起的语义漂移。
摘要:The past ten years have witnessed the rapid development of text-based intent detection, whose benchmark performances have already been taken to a remarkable level by deep learning techniques. However, automatic speech recognition (ASR) errors are inevitable in real-world applications due to the environment noise, unique speech patterns and etc, leading to sharp performance drop in state-of-the-art text-based intent detection models. Essentially, this phenomenon is caused by the semantic drift brought by ASR errors and most existing works tend to focus on designing new model structures to reduce its impact, which is at the expense of versatility and flexibility. Different from previous one-piece model, in this paper, we propose a novel and agile framework called CR-ID for ASR error robust intent detection with two plug-and-play modules, namely semantic drift calibration module (SDCM) and phonemic refinement module (PRM), which are both model-agnostic and thus could be easily integrated to any existing intent detection models without modifying their structures. Experimental results on SNIPS dataset show that, our proposed CR-ID framework achieves competitive performance and outperform all the baseline methods on ASR outputs, which verifies that CR-ID can effectively alleviate the semantic drift caused by ASR errors.


【4】 Self-Supervised Speech Representation Learning: A Review

标题:自监督语音表征学习:综述

链接:https://arxiv.org/abs/2205.10643

作者:Abdelrahman Mohamed,Hung-yi Lee,Lasse Borgholt,Jakob D. Havtorn,Joakim Edin,Christian Igel,Katrin Kirchhoff,Shang-Wen Li,Karen Livescu,Lars Maaløe,Tara N. Sainath,Shinji Watanabe
摘要:尽管有监督的深度学习已经彻底改变了语音和音频处理,但它需要为单个任务和应用场景构建专家模型。同样,很难将这一点应用于只有有限标记数据可用的方言和语言。自监督表征学习方法承诺了一个单一的通用模型,这将有利于广泛的任务和领域。这些方法在自然语言处理和计算机视觉领域取得了成功,实现了新的性能水平,同时减少了许多下游场景所需的标签数量。言语表征学习在生成法、对比法和预测法三大类中都有类似的进展。其他方法依赖于多模态数据进行预训练,将文本或视觉数据流与语音混合。尽管自监督语音表征仍然是一个新兴的研究领域,但它与声学单词嵌入和零词汇资源学习密切相关,这两个领域的研究多年来都很活跃。本文综述了自监督语音表征学习的方法及其与其他研究领域的联系。由于当前的许多方法仅将自动语音识别作为一项下游任务,因此我们回顾了最近对学习表示进行基准测试的工作,以将应用扩展到语音识别之外。
摘要:Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios. It is likewise difficult to apply this to dialects and languages for which only limited labeled data is available. Self-supervised representation learning methods promise a single universal model that would benefit a wide variety of tasks and domains. Such methods have shown success in natural language processing and computer vision domains, achieving new levels of performance while reducing the number of labels required for many downstream scenarios. Speech representation learning is experiencing similar progress in three main categories: generative, contrastive, and predictive methods. Other approaches rely on multi-modal data for pre-training, mixing text or visual data streams with speech. Although self-supervised speech representation is still a nascent research area, it is closely related to acoustic word embedding and learning with zero lexical resources, both of which have seen active research for many years. This review presents approaches for self-supervised speech representation learning and their connection to other research areas. Since many current methods focus solely on automatic speech recognition as a downstream task, we review recent efforts on benchmarking learned representations to extend the application beyond speech recognition.


【5】 Modernizing Open-Set Speech Language Identification

标题:开放式语音语言识别的现代化

链接:https://arxiv.org/abs/2205.10397

作者:Mustafa Eyceoz,Justin Lee,Homayoon Beigi
机构:Dept. of Computer Science, Columbia University, New York, Recognition Technologies, Inc. and Columbia University, New York
备注:7 pages, 6 figures, 3 tables, Technical Report: Recognition Technologies, Inc
摘要:虽然大多数现代语音识别方法都是闭集的,但我们想看看它们是否可以被修改并适应开集问题。当切换到开放集问题时,解决方案可以在音频输入与我们已知的任何语言选项不匹配时拒绝音频输入。我们通过将两种现代最先进的方法应用于封闭集语言识别来处理开放集任务:第一种方法使用带注意的CRNN,第二种方法使用TDNN。除了使用MFCC、对数光谱特征和基音增强输入特征嵌入之外,我们还将尝试两种方法来检测设置外的语言:一种使用阈值,另一种基本上执行验证任务。我们将比较TDNN和CRNN的性能,以及我们的检测方法。
摘要:While most modern speech Language Identification methods are closed-set, we want to see if they can be modified and adapted for the open-set problem. When switching to the open-set problem, the solution gains the ability to reject an audio input when it fails to match any of our known language options. We tackle the open-set task by adapting two modern-day state-of-the-art approaches to closed-set language identification: the first using a CRNN with attention and the second using a TDNN. In addition to enhancing our input feature embeddings using MFCCs, log spectral features, and pitch, we will be attempting two approaches to out-of-set language detection: one using thresholds, and the other essentially performing a verification task. We will compare both the performance of the TDNN and the CRNN, as well as our detection approaches.


机器翻译,仅供参考