本文经arXiv每日学术速递授权转载
标题: ED-sKWS:早期决策尖峰神经网络,用于快速、节能的关键词发现
作者:Zeyang Song,Qianhui Liu,Qu Yang,Yizhou Peng,Haizhou Li
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:关键字定位(KWS)在需要快速和节能响应的边缘计算中至关重要。尖峰神经网络(SNN)非常适合KWS的效率和语音的时间容量。为了进一步降低延迟和能量消耗,本研究引入了ED-sKWS,这是一种基于SNN的KWS模型,具有早期决策机制,可以在语音话语结束之前停止语音处理并输出结果。此外,我们引入了一个累积时间(CT)损失,可以提高预测精度在中间和最后的时间步。为了评估早期决策性能,我们提出了SC-100数据集,包括100个语音命令的开始和结束时间戳注释。在Google Speech Commands v2和我们的SC-100数据集上的实验表明,与没有早期决策机制的SNN模型相比,ED-sKWS保持了61%的时间步长和52%的能耗,确保了快速响应和能源效率。摘要:Keyword Spotting (KWS) is essential in edge computing requiring rapid and energy-efficient responses. Spiking Neural Networks (SNNs) are well-suited for KWS for their efficiency and temporal capacity for speech. To further reduce the latency and energy consumption, this study introduces ED-sKWS, an SNN-based KWS model with an early-decision mechanism that can stop speech processing and output the result before the end of speech utterance. Furthermore, we introduce a Cumulative Temporal (CT) loss that can enhance prediction accuracy at both the intermediate and final timesteps. To evaluate early-decision performance, we present the SC-100 dataset including 100 speech commands with beginning and end timestamp annotation. Experiments on the Google Speech Commands v2 and our SC-100 datasets show that ED-sKWS maintains competitive accuracy with 61% timesteps and 52% energy consumption compared to SNN models without early-decision mechanism, ensuring rapid response and energy efficiency.
【2】 Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and Reaction
标题: 与类人代理人交谈:通过可感知的声学接收和反应进行同理心对话
作者:Haoqiu Yan,Yongxin Zhu,Kai Zheng,Bing Liu,Haoyu Cao,Deqiang Jiang,Linli Xu
备注:9 pages, 3 figures, ACL24 accepted
链接:点击下载PDF文件
摘要:大型语言模型(LLM)增强型代理在人类与AI的通信中变得越来越普遍,从娱乐到专业领域都有巨大的潜力。然而,当前的多模态对话系统忽略了语音中存在的声学信息,这对于理解人类通信的细微差别至关重要。这种疏忽可能导致对发言者意图的误解,导致对话中的反应不一致甚至相互矛盾。为了弥合这一差距,在本文中,我们提出了PerceptiveAgent,一个移情多模态对话系统,旨在通过整合语音模态感知来识别更深层次或更微妙的意义,超越文字的字面解释。PerceptiveAgent以LLM为认知核心,从输入语音中感知声学信息,并根据自然语言描述的说话风格生成移情响应。实验结果表明,PerceptiveAgent在语境理解方面表现出色,在语言意义与说话人的真实感受相反或不一致的情况下,它能准确地识别说话人的真实意图,产生更细致入微和更有表现力的口语对话。代码可在以下位置公开获取: url{https: github.com Haoqiu-Yan PerceptiveAgent}。摘要:Large Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains. However, current multi-modal dialogue systems overlook the acoustic information present in speech, which is crucial for understanding human communication nuances. This oversight can lead to misinterpretations of speakers' intentions, resulting in inconsistent or even contradictory responses within dialogues. To bridge this gap, in this paper, we propose PerceptiveAgent, an empathetic multi-modal dialogue system designed to discern deeper or more subtle meanings beyond the literal interpretations of words through the integration of speech modality perception. Employing LLMs as a cognitive core, PerceptiveAgent perceives acoustic information from input speech and generates empathetic responses based on speaking styles described in natural language. Experimental results indicate that PerceptiveAgent excels in contextual understanding by accurately discerning the speakers' true intentions in scenarios where the linguistic meaning is either contrary to or inconsistent with the speaker's true feelings, producing more nuanced and expressive spoken dialogues. Code is publicly available at: url{https: github.com Haoqiu-Yan PerceptiveAgent}.
【3】 Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
标题: 弥合差距:集成预训练的语音增强和识别模型以实现稳健的语音识别
作者:Kuan-Chen Wang,You-Jin Li,Wei-Lun Chen,Yu-Wen Chen,Yi-Ching Wang,Ping-Cheng Yeh,Chao Zhang,Yu Tsao
链接:点击下载PDF文件
摘要:噪声鲁棒性是在现实世界中应用自动语音识别(ASR)时的关键。一种解决方案涉及使用语音增强(SE)模型作为ASR的前端。然而,基于神经网络(NN)的SE通常会在增强信号中引入伪影并损害ASR性能,特别是当SE和ASR独立训练时。因此,本研究引入了一种简单而有效的SE后处理技术,以解决各种预训练SE和ASR模型之间的差距。提出了一种轻量级的神经网络桥接模块,用于评估语音信号的信号电平信息。随后,利用信号电平信息,应用观测相加技术,有效地减少了SE的缺点。实验结果表明,我们的方法成功地集成了各种预训练的SE和ASR模型,大大提高了ASR的鲁棒性。至关重要的是,在训练或推理阶段不需要ASR或语音内容的先验知识。此外,这种方法的有效性可以扩展到不同的数据集,而不需要对桥接模块进行微调,从而确保效率和提高泛化能力。摘要:Noise robustness is critical when applying automatic speech recognition (ASR) in real-world scenarios. One solution involves the used of speech enhancement (SE) models as the front end of ASR. However, neural network-based (NN-based) SE often introduces artifacts into the enhanced signals and harms ASR performance, particularly when SE and ASR are independently trained. Therefore, this study introduces a simple yet effective SE post-processing technique to address the gap between various pre-trained SE and ASR models. A bridge module, which is a lightweight NN, is proposed to evaluate the signal-level information of the speech signal. Subsequently, using the signal-level information, the observation addition technique is applied to effectively reduce the shortcomings of SE. The experimental results demonstrate the success of our method in integrating diverse pre-trained SE and ASR models, considerably boosting the ASR robustness. Crucially, no prior knowledge of the ASR or speech contents is required during the training or inference stages. Moreover, the effectiveness of this approach extends to different datasets without necessitating the fine-tuning of the bridge module, ensuring efficiency and improved generalization.
【4】 Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
标题: 使用编码器预处理的多语言E2E语音识别快速语言适应
作者:Yosuke Kashiwagi,Hayato Futami,Emiru Tsunoo,Siddhant Arora,Shinji Watanabe
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:端到端多语言语音识别模型通过单个模型处理多种语言,通常包含语言识别以自动检测传入语音的语言。由于常见的情况是语言已经知道,这些模型可以通过使用语言信息作为提示来执行特定于语言的操作,这对于基于注意力的编码器-解码器架构特别有益。然而,通过联合解码和多任务训练来增强识别的连接主义时间分类(CTC)方法通常不包含语言提示,这是由于其条件独立的输出标记。为了克服这一点,我们引入了一个编码器提示技术内的自我调节的CTC框架,使特定语言的适应的CTC模型中的zero-shot的方式。我们的方法已被证明可以将错误平均减少28%,在低资源语言上减少41%。摘要:End-to-end multilingual speech recognition models handle multiple languages through a single model, often incorporating language identification to automatically detect the language of incoming speech. Since the common scenario is where the language is already known, these models can perform as language-specific by using language information as prompts, which is particularly beneficial for attention-based encoder-decoder architectures. However, the Connectionist Temporal Classification (CTC) approach, which enhances recognition via joint decoding and multi-task training, does not normally incorporate language prompts due to its conditionally independent output tokens. To overcome this, we introduce an encoder prompting technique within the self-conditioned CTC framework, enabling language-specific adaptation of the CTC model in a zero-shot manner. Our method has shown to significantly reduce errors by 28% on average and by 41% on low-resource languages.
【5】 Integrating Representational Gestures into Automatically Generated Embodied Explanations and its Effects on Understanding and Interaction Quality
标题: 将代表性手势集成到自动生成的预定描述中及其对理解和交互质量的影响
作者:Amelie Sophie Robrecht,Hendric Voss,Lisa Gottschalk,Stefan Kopp
链接:点击下载PDF文件
摘要:在人类交互中,手势具有多种功能,如标记语音节奏,突出关键元素和补充信息。这些手势也可以在解释性的语境中观察到。然而,手势对虚拟代理提供的解释的影响仍然没有得到充分的研究。用户研究进行了调查如何不同类型的手势影响感知的交互质量和听众的理解。本研究旨在探讨手势在解释中的作用,开发一个具身化的虚拟解释器,整合了节拍手势和图标手势,以增强其自动生成的口头解释。我们的模型结合了由一个学习的语音驱动的合成模块与手动捕获的标志性手势生成的节拍手势,支持代理的口头表达棋盘游戏Quarto!作为一种解释。研究结果表明,无论是单独使用的标志性手势,也不是他们的组合与节拍手势优于基线或节拍的条件下的理解。尽管如此,与以前的研究相比,体现代理显着增强理解。摘要:In human interaction, gestures serve various functions such as marking speech rhythm, highlighting key elements, and supplementing information. These gestures are also observed in explanatory contexts. However, the impact of gestures on explanations provided by virtual agents remains underexplored. A user study was carried out to investigate how different types of gestures influence perceived interaction quality and listener understanding. This study addresses the effect of gestures in explanation by developing an embodied virtual explainer integrating both beat gestures and iconic gestures to enhance its automatically generated verbal explanations. Our model combines beat gestures generated by a learned speech-driven synthesis module with manually captured iconic gestures, supporting the agent's verbal expressions about the board game Quarto! as an explanation scenario. Findings indicate that neither the use of iconic gestures alone nor their combination with beat gestures outperforms the baseline or beat-only conditions in terms of understanding. Nonetheless, compared to prior research, the embodied agent significantly enhances understanding.
【6】 Towards Audio Codec-based Speech Separation
标题: 迈向基于音频编解码器的语音分离
作者:Jia Qi Yip,Shengkui Zhao,Dianwen Ng,Eng Siong Chng,Bin Ma
备注:This paper was accepted by Interspeech 2024, Blue Sky Track
链接:点击下载PDF文件
摘要:神经音频编解码器(NAC)模型的最新改进已经引起了人们对采用预训练的编解码器用于各种语音处理应用以利用从高压缩获得的效率的兴趣,但是这些尚未应用于语音分离(SS)任务。SS可以从高压缩中受益,因为传统SS模型所需的计算使得它们对于许多边缘计算用例来说不切实际。然而,SS是波形掩蔽任务,其中压缩倾向于引入严重影响性能的失真。在这里,我们提出了一个新的任务,音频编解码器为基础的SS,其中SS是在NAC的嵌入空间内执行,并提出了一个新的模型,编解码器,来解决这个任务。在推理中,Codecformer实现了MAC减少52倍,同时产生与Sepformer云部署相当的分离性能。该方法为在实际场景中执行高效SS指明了新的方向。摘要:Recent improvements in neural audio codec (NAC) models have generated interest in adopting pre-trained codecs for a variety of speech processing applications to take advantage of the efficiencies gained from high compression, but these have yet been applied to the speech separation (SS) task. SS can benefit from high compression because the compute required for traditional SS models makes them impractical for many edge computing use cases. However, SS is a waveform-masking task where compression tends to introduce distortions that severely impact performance. Here we propose a novel task of Audio Codec-based SS, where SS is performed within the embedding space of a NAC, and propose a new model, Codecformer, to address this task. At inference, Codecformer achieves a 52x reduction in MAC while producing separation performance comparable to a cloud deployment of Sepformer. This method charts a new direction for performing efficient SS in practical scenarios.
【7】 PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
标题: PSLM:针对低延迟口语对话系统的LLM并行生成文本和语音
作者:Kentaro Mitsui,Koh Mitsuda,Toshiaki Wakatsuki,Yukiya Hono,Kei Sawada
备注:8 pages, 4 figures, 4 tables, demo samples: this https URL
链接:点击下载PDF文件
摘要:同时处理文本和语音的多模态语言模型在口语对话系统中具有应用潜力。然而,当前模型在响应生成延迟方面面临两个主要挑战:(1)生成口头响应需要预先生成书面响应,以及(2)语音序列明显长于文本序列。本研究通过扩展语言模型的输入和输出序列来支持文本和语音的并行生成来解决这些问题。我们对口语问答任务的实验表明,我们的方法在保持响应内容质量的同时提高了延迟。此外,我们表明,延迟可以进一步减少在多个序列中生成语音。演示示例可在https: rinnakk.github.io research publications PSLM上获得。摘要:Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating a spoken response requires the prior generation of a written response, and (2) speech sequences are significantly longer than text sequences. This study addresses these issues by extending the input and output sequences of the language model to support the parallel generation of text and speech. Our experiments on spoken question answering tasks demonstrate that our approach improves latency while maintaining the quality of response content. Additionally, we show that latency can be further reduced by generating speech in multiple sequences. Demo samples are available at https: rinnakk.github.io research publications PSLM.
【8】 JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning
标题: JEN-1 DreamStyler:通过基本参数调整定制音乐概念学习
作者:Boyu Chen,Peike Li,Yao Yao,Alex Wang
链接:点击下载PDF文件
摘要:None摘要:Large models for text-to-music generation have achieved significant progress, facilitating the creation of high-quality and varied musical compositions from provided text prompts. However, input text prompts may not precisely capture user requirements, particularly when the objective is to generate music that embodies a specific concept derived from a designated reference collection. In this paper, we propose a novel method for customized text-to-music generation, which can capture the concept from a two-minute reference music and generate a new piece of music conforming to the concept. We achieve this by fine-tuning a pretrained text-to-music model using the reference music. However, directly fine-tuning all parameters leads to overfitting issues. To address this problem, we propose a Pivotal Parameters Tuning method that enables the model to assimilate the new concept while preserving its original generative capabilities. Additionally, we identify a potential concept conflict when introducing multiple concepts into the pretrained model. We present a concept enhancement strategy to distinguish multiple concepts, enabling the fine-tuned model to generate music incorporating either individual or multiple concepts simultaneously. Since we are the first to work on the customized music generation task, we also introduce a new dataset and evaluation protocol for the new task. Our proposed Jen1-DreamStyler outperforms several baselines in both qualitative and quantitative evaluations. Demos will be available at https: www.jenmusic.ai research#DreamStyler.
【9】 Interface Design for Self-Supervised Speech Models
标题: 自我监督语音模型的接口设计
作者:Yi-Jen Shih,David Harwath
备注:Accepted to Interspeech2024
链接:点击下载PDF文件
摘要:自监督语音(SSL)模型最近已被广泛用于许多下游语音处理任务。一般的使用模式是采用SSL模型作为特征提取器,然后训练下游预测头来解决特定的任务。然而,SSL模型的不同层已经被证明可以捕获不同类型的信息,并且组合它们的方法还没有得到很好的研究。为此,我们扩展了SSL模型利用率的一般框架,提出了连接上游和下游的接口。在这种观点下,通过逐层加权和组合特征的主导技术可以被视为特定的接口。我们提出了几种替代的接口设计,并证明了加权和接口是次优的许多任务。特别是,我们表明,卷积接口的深度尺度与上游模型的深度一致优于许多其他接口设计。摘要:Self-supervised speech (SSL) models have recently become widely adopted for many downstream speech processing tasks. The general usage pattern is to employ SSL models as feature extractors, and then train a downstream prediction head to solve a specific task. However, different layers of SSL models have been shown to capture different types of information, and the methods of combining them are not well studied. To this end, we extend the general framework for SSL model utilization by proposing the interface that connects the upstream and downstream. Under this view, the dominant technique of combining features via a layerwise weighted sum can be regarded as a specific interface. We propose several alternative interface designs and demonstrate that the weighted sum interface is suboptimal for many tasks. In particular, we show that a convolutional interface whose depth scales logarithmically with the depth of the upstream model consistently outperforms many other interface designs.
【10】 A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
标题: 语音合成中基于CWT的Mel谱图增强范式
作者:Guoqiang Hu,Huaning Tan,Ruilai Li
链接:点击下载PDF文件
摘要:声学特征对提高合成语音的质量起着重要的作用。目前,Mel谱图是大多数声学模型中广泛使用的声学特征。然而,由于其傅里叶变换过程中造成的细粒度损失,梅尔频谱图合成的语音清晰度在突变信号中受到影响。为了获得更详细的Mel谱图,我们提出了一种基于连续小波变换(CWT)的Mel谱图增强方法。这种模式引入了一个额外的任务:一个更详细的小波频谱图,它像后处理网络一样,将解码器输出的Mel频谱图作为输入。我们选择Tacotron2和Fastspeech2进行实验验证,以测试自回归(AR)和非自回归(NAR)语音系统,分别。实验结果表明,使用该模型与梅尔频谱图增强范例合成的语音表现出更高的MOS,与基线模型相比,分别提高了0.14和0.09。这些发现为增强范例的普遍性提供了一些验证,因为它们证明了该范例在不同架构中的成功。摘要:Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained loss caused by its Fourier transform process, the clarity of speech synthesised by Mel spectrogram is compromised in mutant signals. In order to obtain a more detailed Mel spectrogram, we propose a Mel spectrogram enhancement paradigm based on the continuous wavelet transform (CWT). This paradigm introduces an additional task: a more detailed wavelet spectrogram, which like the post-processing network takes as input the Mel spectrogram output by the decoder. We choose Tacotron2 and Fastspeech2 for experimental validation in order to test autoregressive (AR) and non-autoregressive (NAR) speech systems, respectively. The experimental results demonstrate that the speech synthesised using the model with the Mel spectrogram enhancement paradigm exhibits higher MOS, with an improvement of 0.14 and 0.09 compared to the baseline model, respectively. These findings provide some validation for the universality of the enhancement paradigm, as they demonstrate the success of the paradigm in different architectures.
【11】 A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
标题: 双任务学习方法微调多语言语义语音编码器以实现口语理解
作者:Gaëlle Laperrière,Sahar Ghannay,Bassam Jabaian,Yannick Estève
备注:In Proceedings of Interspeech 2024
链接:点击下载PDF文件
摘要:自我监督学习被广泛用于有效地表示口语理解的语音,逐渐取代传统的方法。同时,提出了文本SSL模型来编码语言无关的语义。SAMU-XLSR框架利用这些语义信息来丰富多语言语音表示。最近的一项研究调查了SAMU-XLSR领域内的语义丰富,通过专门研究下游transmittance,导致最先进的结果,具有挑战性的SLU任务。本研究的兴趣在于损失的多语言性能和缺乏具体的语义训练,这种专业化在没有任何SLU的含义,在接近的语言。我们还考虑SAMU-XLSR的损失,由于一个单独的SLU微调初始跨语言的能力。因此,本文提出了一种双任务学习方法,以提高SAMU-XLSR语义丰富,同时考虑到遥远的语言多语言和语言移植性实验。摘要:Self-Supervised Learning is vastly used to efficiently represent speech for Spoken Language Understanding, gradually replacing conventional approaches. Meanwhile, textual SSL models are proposed to encode language-agnostic semantics. SAMU-XLSR framework employed this semantic information to enrich multilingual speech representations. A recent study investigated SAMU-XLSR in-domain semantic enrichment by specializing it on downstream transcriptions, leading to state-of-the-art results on a challenging SLU task. This study's interest lies in the loss of multilingual performances and lack of specific-semantics training induced by such specialization in close languages without any SLU implication. We also consider SAMU-XLSR's loss of initial cross-lingual abilities due to a separate SLU fine-tuning. Therefore, this paper proposes a dual task learning approach to improve SAMU-XLSR semantic enrichment while considering distant languages for multilingual and language portability experiments.
【12】 Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4
标题: 基于辅助解码器和最大概率聚合的声音事件检测,以满足DUSE Challenge 2024任务4
作者:Sangwon Son,Jongyeon Park,Hongkook Kim,Sulaiman Vesal,Jeongeun Lim
备注:DCASE 2024 challenge Task4, 4 pages
链接:点击下载PDF文件
摘要:在这份报告中,我们提出了三种新的方法来开发DCASE 2024挑战任务4的声音事件检测(SED)模型。首先,我们提出了一个附加到最终卷积块的辅助解码器,以提高特征提取能力,同时减少对预训练大型模型嵌入的依赖。所提出的辅助解码器独立于主解码器操作,通过在主解码器和辅助解码器损失之间分配不同的权重策略来增强初始训练阶段期间卷积块的性能。接下来,为了解决DESED和MAESTRO数据集之间的时间间隔问题,我们在训练步骤中提出了最大概率聚合(MPA)。建议的MPA方法使模型的输出与MAESTRO数据集中的1 s软标签对齐。最后,我们提出了一个多通道输入功能,采用不同版本的logmel和MFCC功能,以产生时间-频率模式。实验结果表明,这些方法的有效性,在提高SED的性能,通过实现跨不同的数据集和标签类型的平衡增强。最终,这种方法在开发更强大和更灵活的SED模型方面迈出了重要的一步摘要:In this report, we propose three novel methods for developing a sound event detection (SED) model for the DCASE 2024 Challenge Task 4. First, we propose an auxiliary decoder attached to the final convolutional block to improve feature extraction capabilities while reducing dependency on embeddings from pre-trained large models. The proposed auxiliary decoder operates independently from the main decoder, enhancing performance of the convolutional block during the initial training stages by assigning a different weight strategy between main and auxiliary decoder losses. Next, to address the time interval issue between the DESED and MAESTRO datasets, we propose maximum probability aggregation (MPA) during the training step. The proposed MPA method enables the model's output to be aligned with soft labels of 1 s in the MAESTRO dataset. Finally, we propose a multi-channel input feature that employs various versions of logmel and MFCC features to generate time-frequency pattern. The experimental results demonstrate the efficacy of these proposed methods in a view of improving SED performance by achieving a balanced enhancement across different datasets and label types. Ultimately, this approach presents a significant step forward in developing more robust and flexible SED models
【13】 Performant ASR Models for Medical Entities in Accented Speech
标题: 医疗实体的高性能ASB模型在强调言语中
作者:Tejumade Afonja,Tobi Olatunji,Sewade Ogun,Naome A. Etori,Abraham Owodunni,Moshood Yekini
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)的最新进展加速了其在医疗领域的应用,其中它们对重音医学命名实体(NE)(如药物名称,诊断和实验室结果)的性能在很大程度上是未知的。我们在93种非洲口音的临床英语数据集上严格评估了多个ASR模型。我们的分析表明,尽管一些模型实现了较低的整体单词错误率(WER),但临床实体中的错误更高,可能对患者安全构成重大风险。为了实证证明这一点,我们从成绩单中提取临床实体,开发一种新的算法来将ASR预测与这些实体对齐,并计算医疗NE召回,医疗WER和字符错误率。我们的研究结果表明,微调口音的临床语音提高了医疗WER的幅度很大(25- 34%相对),提高其在医疗环境中的实用性。摘要:Recent strides in automatic speech recognition (ASR) have accelerated their application in the medical domain where their performance on accented medical named entities (NE) such as drug names, diagnoses, and lab results, is largely unknown. We rigorously evaluate multiple ASR models on a clinical English dataset of 93 African accents. Our analysis reveals that despite some models achieving low overall word error rates (WER), errors in clinical entities are higher, potentially posing substantial risks to patient safety. To empirically demonstrate this, we extract clinical entities from transcripts, develop a novel algorithm to align ASR predictions with these entities, and compute medical NE Recall, medical WER, and character error rate. Our results show that fine-tuning on accented clinical speech improves medical WER by a wide margin (25-34 % relative), improving their practical applicability in healthcare environments.
【14】 Binaural Selective Attention Model for Target Speaker Extraction
标题: 目标说话人提取的双耳选择性注意模型
作者:Hanyu Meng,Qiquan Zhang,Xiangyu Zhang,Vidhyasaharan Sethu,Eliathamby Ambikairajah
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:人类在鸡尾酒会场景中选择性地聚焦于目标扬声器的显著能力由双耳音频处理促进。本文提出了一种基于滤波求和网络(FaSNet)的双耳时域目标说话人提取模型。受人类选择性听觉的启发,我们提出的模型使用基于多头注意的选择性注意块将目标说话人嵌入分离器。我们还比较了两种双耳交互方法-时域信号的余弦相似性和学习频谱表示中的声道间相关性。我们的实验结果表明,我们提出的模型优于单声道配置和国家的最先进的多通道目标扬声器提取模型,实现了同类最佳的性能与18.52 dB SI-SDR,19.12 dB SDR,和3.05 PESQ的成绩下消声两个扬声器测试配置。摘要:The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extraction model based on the Filter-and-Sum Network (FaSNet). Inspired by human selective hearing, our proposed model introduces target speaker embedding into separators using a multi-head attention-based selective attention block. We also compared two binaural interaction approaches -- the cosine similarity of time-domain signals and inter-channel correlation in learned spectral representations. Our experimental results show that our proposed model outperforms monaural configurations and state-of-the-art multi-channel target speaker extraction models, achieving best-in-class performance with 18.52 dB SI-SDR, 19.12 dB SDR, and 3.05 PESQ scores under anechoic two-speaker test configurations.
【15】 Universal Score-based Speech Enhancement with High Content Preservation
标题: 具有高内容保留的通用基于分数的语音增强
作者:Robin Scheibler,Yusuke Fujita,Yuma Shirahata,Tatsuya Komatsu
备注:5 pages, 5 figures, accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了UNIVERSE++,一种基于分数扩散和对抗训练的通用语音增强方法。具体来说,我们改进了现有的UNIVERSE模型,该模型包含了干净的语音特征提取和扩散。我们的贡献是三方面的。首先,我们对网络架构进行了一些修改,提高了训练稳定性和最终性能。其次,我们引入了一个对抗损失,以促进学习高质量的语音特征。第三,我们提出了一个低秩自适应方案与音素保真度损失,以提高增强语音的内容保存。在实验中,我们训练一个通用的增强模型上的大规模数据集的语音退化的噪声,混响,和各种失真。多个公共基准数据集的结果表明,在广泛的定性和可理解性指标方面,UNIVERSE++与区分基线和生成基线相比都具有优势。摘要:We propose UNIVERSE++, a universal speech enhancement method based on score-based diffusion and adversarial training. Specifically, we improve the existing UNIVERSE model that decouples clean speech feature extraction and diffusion. Our contributions are three-fold. First, we make several modifications to the network architecture, improving training stability and final performance. Second, we introduce an adversarial loss to promote learning high quality speech features. Third, we propose a low-rank adaptation scheme with a phoneme fidelity loss to improve content preservation in the enhanced speech. In the experiments, we train a universal enhancement model on a large scale dataset of speech degraded by noise, reverberation, and various distortions. The results on multiple public benchmark datasets demonstrate that UNIVERSE++ compares favorably to both discriminative and generative baselines for a wide range of qualitative and intelligibility metrics.
eeeess.AS音频处理
【1】 Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4标题: 基于辅助解码器和最大概率聚合的声音事件检测,以满足DUSE Challenge 2024任务4
作者:Sangwon Son,Jongyeon Park,Hongkook Kim,Sulaiman Vesal,Jeongeun Lim
备注:DCASE 2024 challenge Task4, 4 pages
链接:点击下载PDF文件
摘要:在这份报告中,我们提出了三种新的方法来开发DCASE 2024挑战任务4的声音事件检测(SED)模型。首先,我们提出了一个附加到最终卷积块的辅助解码器,以提高特征提取能力,同时减少对预训练大型模型嵌入的依赖。所提出的辅助解码器独立于主解码器操作,通过在主解码器和辅助解码器损失之间分配不同的权重策略来增强初始训练阶段期间卷积块的性能。接下来,为了解决DESED和MAESTRO数据集之间的时间间隔问题,我们在训练步骤中提出了最大概率聚合(MPA)。建议的MPA方法使模型的输出与MAESTRO数据集中的1 s软标签对齐。最后,我们提出了一个多通道输入功能,采用不同版本的logmel和MFCC功能,以产生时间-频率模式。实验结果表明,这些方法的有效性,在提高SED的性能,通过实现跨不同的数据集和标签类型的平衡增强。最终,这种方法在开发更强大和更灵活的SED模型方面迈出了重要的一步摘要:In this report, we propose three novel methods for developing a sound event detection (SED) model for the DCASE 2024 Challenge Task 4. First, we propose an auxiliary decoder attached to the final convolutional block to improve feature extraction capabilities while reducing dependency on embeddings from pre-trained large models. The proposed auxiliary decoder operates independently from the main decoder, enhancing performance of the convolutional block during the initial training stages by assigning a different weight strategy between main and auxiliary decoder losses. Next, to address the time interval issue between the DESED and MAESTRO datasets, we propose maximum probability aggregation (MPA) during the training step. The proposed MPA method enables the model's output to be aligned with soft labels of 1 s in the MAESTRO dataset. Finally, we propose a multi-channel input feature that employs various versions of logmel and MFCC features to generate time-frequency pattern. The experimental results demonstrate the efficacy of these proposed methods in a view of improving SED performance by achieving a balanced enhancement across different datasets and label types. Ultimately, this approach presents a significant step forward in developing more robust and flexible SED models
【2】 Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
标题: 现场说话:基于扩散的声学场景转移到沉浸式语音生成
作者:Miseul Kim,Soo-Whan Chung,Youna Ji,Hong-Goo Kang,Min-Seok Choi
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了生成语音处理中的一个新任务,声学场景转移(AST),其目的是将语音信号的声学场景转移到不同的环境中。AST通过使语音信号背后的声学场景适应所需的环境,承诺在语音感知方面的沉浸式体验。我们提出AST-LDM的AST任务,它产生的语音信号伴随着目标声场景的参考提示。具体来说,AST-LDM是一个潜在的扩散模型条件CLAP嵌入描述目标声学场景中的音频或文本模态。本文的贡献包括介绍AST任务和实现其基线模型。对于AST-LDM,我们强调它的核心框架,这是保持输入语音和生成音频与给定的语音和目标声学环境一致。通过客观和主观测试,验证了该方法的可行性和有效性。摘要:This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments. AST promises an immersive experience in speech perception by adapting the acoustic scene behind speech signals to desired environments. We propose AST-LDM for the AST task, which generates speech signals accompanied by the target acoustic scene of the reference prompt. Specifically, AST-LDM is a latent diffusion model conditioned by CLAP embeddings that describe target acoustic scenes in either audio or text modalities. The contributions of this paper include introducing the AST task and implementing its baseline model. For AST-LDM, we emphasize its core framework, which is to preserve the input speech and generate audio consistently with both the given speech and the target acoustic environment. Experiments, including objective and subjective tests, validate the feasibility and efficacy of our approach.
【3】 Transcribe, Align and Segment: Creating speech datasets for low-resource languages
标题: 转录、对齐和分段:为低资源语言创建语音数据集
作者:Taras Sereda
链接:点击下载PDF文件
摘要:在这项工作中,我们展示了一种具有成本效益的方法,用于生成语音处理任务的训练数据。首先,我们使用最先进的自动语音识别(ASR)模型转录未标记的语音。接下来,我们将生成的成绩单与音频对齐,并对简短的话语进行分割。我们的重点是低资源语言的ASR,如乌克兰语,使用播客作为未标记语音的来源。 我们发布了一个新的数据集UK-PODS,它具有现代乌克兰语会话。它包含超过50个小时的文本音频对以及uk-pods-conformer,这是一个121 M参数的ASR模型,在MCV-10和UK-PODS上训练,并在播客上实现了3倍的字错误率(WER)降低,同时在MCV-10测试分割上保持了可比的WER。数据集UK-PODS https: huggingface.co datasets taras-sereda uk-pods和ASR uk-pods-conformer https: huggingface.co taras-sereda uk-pods-conformer都可以在hugging-face hub上找到。摘要:In this work, we showcase a cost-effective method for generating training data for speech processing tasks. First, we transcribe unlabeled speech using a state-of-the-art Automatic Speech Recognition (ASR) model. Next, we align generated transcripts with the audio and apply segmentation on short utterances. Our focus is on ASR for low-resource languages, such as Ukrainian, using podcasts as a source of unlabeled speech. We release a new dataset UK-PODS that features modern conversational Ukrainian language. It contains over 50 hours of text audio-pairs as well as uk-pods-conformer, a 121 M parameters ASR model that is trained on MCV-10 and UK-PODS and achieves 3x reduction of Word Error Rate (WER) on podcasts comparing to publically available uk-nvidia-citrinet while maintaining comparable WER on MCV-10 test split. Both dataset UK-PODS https: huggingface.co datasets taras-sereda uk-pods and ASR uk-pods-conformer https: huggingface.co taras-sereda uk-pods-conformer are available on the hugging-face hub.
【4】 Challenging margin-based speaker embedding extractors by using the variational information bottleneck
标题: 利用变分信息瓶颈搜索基于边缘的说话人嵌入提取器
作者:Themos Stafylakis,Anna Silnova,Johan Rohdin,Oldrich Plchot,Lukas Burget
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:说话人嵌入提取器通常使用训练说话人的分类损失来训练。在过去的几年中,标准的softmax 交叉熵损失已经被基于边缘的损失所取代,从而显着提高了说话人识别的准确性。出于这样一个事实,即利润率只是减少了目标扬声器在训练过程中的logit,我们考虑一个概率框架,具有类似的效果。变分信息瓶颈提供了一个原则性的机制,使确定性节点的随机性,从而导致隐式减少后验的目标扬声器。我们实验了各种各样的说话人识别基准和评分方法,并报告了与最先进的加性角余量损失所获得的结果相比具有竞争力的结果。摘要:Speaker embedding extractors are typically trained using a classification loss over the training speakers. During the last few years, the standard softmax cross-entropy loss has been replaced by the margin-based losses, yielding significant improvements in speaker recognition accuracy. Motivated by the fact that the margin merely reduces the logit of the target speaker during training, we consider a probabilistic framework that has a similar effect. The variational information bottleneck provides a principled mechanism for making deterministic nodes stochastic, resulting in an implicit reduction of the posterior of the target speaker. We experiment with a wide range of speaker recognition benchmarks and scoring methods and report competitive results to those obtained with the state-of-the-art Additive Angular Margin loss.
【5】 Unsupervised Online Continual Learning for Automatic Speech Recognition
标题: 自动语音识别的无监督在线持续学习
作者:Steven Vander Eeckt,Hugo Van hamme
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)模型适应新的领域会导致灾难性遗忘(CF)以前学习的信息。本文讨论了在线持续学习(OCL)的挑战性背景下的CF,任务作为一个连续的数据流与未知的边界。我们将OCL for ASR扩展到无监督领域,通过利用自训练(ST)来促进无监督适应,使模型能够在不依赖标签和不忘记以前知识的情况下不断适应。通过对两个领域自适应实验中各种OCL和ST方法的比较分析,我们发现UOCL与监督OCL相比遗忘明显减少,使UOCL方法接近监督OCL的性能水平。我们提出的UOCL扩展进一步提高了UOCL的效率。我们的研究结果代表了迈向可持续适应的ASR系统的重要一步,能够在不同的领域利用未标记的数据。摘要:Adapting Automatic Speech Recognition (ASR) models to new domains leads to Catastrophic Forgetting (CF) of previously learned information. This paper addresses CF in the challenging context of Online Continual Learning (OCL), with tasks presented as a continuous data stream with unknown boundaries. We extend OCL for ASR into the unsupervised realm, by leveraging self-training (ST) to facilitate unsupervised adaptation, enabling models to adapt continually without label dependency and without forgetting previous knowledge. Through comparative analysis of various OCL and ST methods across two domain adaptation experiments, we show that UOCL suffers from significantly less forgetting compared to supervised OCL, allowing UOCL methods to approach the performance levels of supervised OCL. Our proposed UOCL extensions further boosts UOCL's efficacy. Our findings represent a significant step towards continually adaptable ASR systems, capable of leveraging unlabeled data across diverse domains.
【6】 Text-aware Speech Separation for Multi-talker Keyword Spotting
标题: 用于多说话者关键词发现的文本感知语音分离
作者:Haoyu Li,Baochen Yang,Yu Xi,Linfeng Yu,Tian Tan,Hao Li,Kai Yu
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:对于嘈杂的环境,确保关键字定位(KWS)系统的鲁棒性至关重要。虽然很多研究都集中在嘈杂的KWS,很少有人注意到多说话人混合语音的情况。与通常的鸡尾酒会问题不同,在鸡尾酒会问题中,使用说话人线索分离多个说话人的语音,这里的关键挑战是基于文本线索提取KWS的目标语音。为了解决这个问题,本文提出了一种新的文本感知的置换确定训练方法的多说话人KWS与基于线索的语音分离前端(TPDT-SS)。我们的研究强调了SS前端的关键作用,并表明将关键词特定的线索纳入这些模型可以大大提高效率。TPDT-SS在解决混合关键字语音中的排列问题方面取得了显着的成功,从而大大提高了后端的性能。此外,微调我们的系统看不见的混合语音的结果在进一步的性能提高。摘要:For noisy environments, ensuring the robustness of keyword spotting (KWS) systems is essential. While much research has focused on noisy KWS, less attention has been paid to multi-talker mixed speech scenarios. Unlike the usual cocktail party problem where multi-talker speech is separated using speaker clues, the key challenge here is to extract the target speech for KWS based on text clues. To address it, this paper proposes a novel Text-aware Permutation Determinization Training method for multi-talker KWS with a clue-based Speech Separation front-end (TPDT-SS). Our research highlights the critical role of SS front-ends and shows that incorporating keyword-specific clues into these models can greatly enhance the effectiveness. TPDT-SS shows remarkable success in addressing permutation problems in mixed keyword speech, thereby greatly boosting the performance of the backend. Additionally, fine-tuning our system on unseen mixed speech results in further performance improvement.
【7】 Exploring Sensing Devices for Heart and Lung Sound Monitoring
标题: 探索心肺声监测的传感设备
作者:Yasaman Torabi,Shahram Shirani,James P. Reilly
链接:点击下载PDF文件
摘要:本文提出了一个全面的审查心肺听诊传感设备,这是有用的理解传感设备的理论方面,以及实用的注意事项,设计新颖的传感设备。设计听诊器的方法之一是使用驻极体电容麦克风(ECM)。在本文中,我们首先介绍了心脏和肺的声学特性,以及听诊器的发展简史。然后,我们讨论了ECM传感器的基本概念和最近的听诊器基于这项技术。针对基于ECM的系统的局限性,我们探索了微机电系统(MEMS)的潜力,特别是专注于压电换能器(PZT)传感器。本文全面回顾了传感技术,强调创新的MEMS为基础的设计,可穿戴心肺听诊在过去的十年。据我们所知,这是第一篇总结ECM和MEMS应用于心肺音分析的论文。保留字:微机电系统;驻极体电容麦克风;可穿戴式传感装置;呼吸听诊;心音描记术;心音;肺音摘要:This paper presents a comprehensive review of cardiorespiratory auscultation sensing devices which is useful for understanding the theoretical aspects of sensing devices, as well as practical notes to design novel sensing devices. One of the methods to design a stethoscope is using electret condenser microphones (ECM). In this paper, we first introduce the acoustic properties of the heart and lungs, as well as a brief history of stethoscope evolution. Then, we discuss the basic concept of ECM sensors and a recent stethoscope based on this technology. In response to the limitations of ECM-based systems, we explore the potential of microelectromechanical systems (MEMS), particularly focusing on piezoelectric transducer (PZT) sensors. This paper comprehensively reviews sensing technologies, emphasizing innovative MEMS-based designs for wearable cardiopulmonary auscultation in the past decade. To our knowledge, this is the first paper to summarize ECM and MEMS applications for heart and lung sound analysis. Keywords: Micro-electro-mechanical Systems (MEMS); Electret Condenser Microphone (ECM); Wearable Sensing Devices; Cardiorespiratory Auscultation; Phonocardiography (PCG); Heart Sound; Lung Sound
【8】 Performant ASR Models for Medical Entities in Accented Speech
标题: 医疗实体的高性能ASB模型在强调言语中
作者:Tejumade Afonja,Tobi Olatunji,Sewade Ogun,Naome A. Etori,Abraham Owodunni,Moshood Yekini
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)的最新进展加速了其在医疗领域的应用,其中它们对重音医学命名实体(NE)(如药物名称,诊断和实验室结果)的性能在很大程度上是未知的。我们在93种非洲口音的临床英语数据集上严格评估了多个ASR模型。我们的分析表明,尽管一些模型实现了较低的整体单词错误率(WER),但临床实体中的错误更高,可能对患者安全构成重大风险。为了实证证明这一点,我们从成绩单中提取临床实体,开发一种新的算法来将ASR预测与这些实体对齐,并计算医疗NE召回,医疗WER和字符错误率。我们的研究结果表明,微调口音的临床语音提高了医疗WER的幅度很大(25- 34%相对),提高其在医疗环境中的实用性。摘要:Recent strides in automatic speech recognition (ASR) have accelerated their application in the medical domain where their performance on accented medical named entities (NE) such as drug names, diagnoses, and lab results, is largely unknown. We rigorously evaluate multiple ASR models on a clinical English dataset of 93 African accents. Our analysis reveals that despite some models achieving low overall word error rates (WER), errors in clinical entities are higher, potentially posing substantial risks to patient safety. To empirically demonstrate this, we extract clinical entities from transcripts, develop a novel algorithm to align ASR predictions with these entities, and compute medical NE Recall, medical WER, and character error rate. Our results show that fine-tuning on accented clinical speech improves medical WER by a wide margin (25-34 % relative), improving their practical applicability in healthcare environments.
【9】 Binaural Selective Attention Model for Target Speaker Extraction
标题: 目标说话人提取的双耳选择性注意模型
作者:Hanyu Meng,Qiquan Zhang,Xiangyu Zhang,Vidhyasaharan Sethu,Eliathamby Ambikairajah
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:双耳音频处理促进了人类在鸡尾酒会场景中选择性地关注目标扬声器的非凡能力。本文提出了一种基于滤波求和网络(FaSNet)的双耳时域目标说话人提取模型。受人类选择性听力的启发,我们提出的模型使用基于多头注意力的选择性注意力块将目标说话者嵌入到分离器中。我们还比较了两种双耳交互方法--时域信号的余弦相似性和学习频谱表示中的声道间相关性。我们的实验结果表明,我们提出的模型优于单声道配置和国家的最先进的多通道目标扬声器提取模型,实现了同类最佳的性能与18.52 dB SI-SDR,19.12 dB SDR,和3.05 PESQ的成绩下消声两个扬声器测试配置。摘要:The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extraction model based on the Filter-and-Sum Network (FaSNet). Inspired by human selective hearing, our proposed model introduces target speaker embedding into separators using a multi-head attention-based selective attention block. We also compared two binaural interaction approaches -- the cosine similarity of time-domain signals and inter-channel correlation in learned spectral representations. Our experimental results show that our proposed model outperforms monaural configurations and state-of-the-art multi-channel target speaker extraction models, achieving best-in-class performance with 18.52 dB SI-SDR, 19.12 dB SDR, and 3.05 PESQ scores under anechoic two-speaker test configurations.
【10】 Universal Score-based Speech Enhancement with High Content Preservation
标题: 具有高内容保留的通用基于分数的语音增强
作者:Robin Scheibler,Yusuke Fujita,Yuma Shirahata,Tatsuya Komatsu
备注:5 pages, 5 figures, accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了UNIVERSE++,一种基于分数扩散和对抗训练的通用语音增强方法。具体来说,我们改进了现有的UNIVERSE模型,该模型包含了干净的语音特征提取和扩散。我们的贡献是三方面的。首先,我们对网络架构进行了一些修改,提高了训练稳定性和最终性能。其次,我们引入了一个对抗损失,以促进学习高质量的语音特征。第三,我们提出了一个低秩自适应方案与音素保真度损失,以提高增强语音的内容保存。在实验中,我们训练一个通用的增强模型上的大规模数据集的语音退化的噪声,混响,和各种失真。在多个公共基准数据集上的结果表明,UNIVERSE++在广泛的定性和可理解性指标方面优于判别和生成基线。摘要:We propose UNIVERSE++, a universal speech enhancement method based on score-based diffusion and adversarial training. Specifically, we improve the existing UNIVERSE model that decouples clean speech feature extraction and diffusion. Our contributions are three-fold. First, we make several modifications to the network architecture, improving training stability and final performance. Second, we introduce an adversarial loss to promote learning high quality speech features. Third, we propose a low-rank adaptation scheme with a phoneme fidelity loss to improve content preservation in the enhanced speech. In the experiments, we train a universal enhancement model on a large scale dataset of speech degraded by noise, reverberation, and various distortions. The results on multiple public benchmark datasets demonstrate that UNIVERSE++ compares favorably to both discriminative and generative baselines for a wide range of qualitative and intelligibility metrics.
【11】 ED-sKWS: Early-Decision Spiking Neural Networks for Rapid,and Energy-Efficient Keyword Spotting
标题: ED-sKWS:早期决策尖峰神经网络,用于快速、节能的关键词发现
作者:Zeyang Song,Qianhui Liu,Qu Yang,Yizhou Peng,Haizhou Li
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:关键字定位(KWS)在需要快速和节能响应的边缘计算中至关重要。尖峰神经网络(SNN)非常适合KWS的效率和语音的时间容量。为了进一步降低延迟和能量消耗,本研究引入了ED-sKWS,这是一种基于SNN的KWS模型,具有早期决策机制,可以在语音话语结束之前停止语音处理并输出结果。此外,我们引入了一个累积时间(CT)损失,可以提高预测精度在中间和最后的时间步。为了评估早期决策性能,我们提出了SC-100数据集,包括100个语音命令的开始和结束时间戳注释。在Google Speech Commands v2和我们的SC-100数据集上的实验表明,与没有早期决策机制的SNN模型相比,ED-sKWS保持了61%的时间步长和52%的能耗,确保了快速响应和能源效率。摘要:Keyword Spotting (KWS) is essential in edge computing requiring rapid and energy-efficient responses. Spiking Neural Networks (SNNs) are well-suited for KWS for their efficiency and temporal capacity for speech. To further reduce the latency and energy consumption, this study introduces ED-sKWS, an SNN-based KWS model with an early-decision mechanism that can stop speech processing and output the result before the end of speech utterance. Furthermore, we introduce a Cumulative Temporal (CT) loss that can enhance prediction accuracy at both the intermediate and final timesteps. To evaluate early-decision performance, we present the SC-100 dataset including 100 speech commands with beginning and end timestamp annotation. Experiments on the Google Speech Commands v2 and our SC-100 datasets show that ED-sKWS maintains competitive accuracy with 61% timesteps and 52% energy consumption compared to SNN models without early-decision mechanism, ensuring rapid response and energy efficiency.
【12】 Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and Reaction
标题: 与类人代理人交谈:通过可感知的声学接收和反应进行同理心对话
作者:Haoqiu Yan,Yongxin Zhu,Kai Zheng,Bing Liu,Haoyu Cao,Deqiang Jiang,Linli Xu
备注:9 pages, 3 figures, ACL24 accepted
链接:点击下载PDF文件
摘要:大型语言模型(LLM)增强型代理在人类与AI的通信中变得越来越普遍,从娱乐到专业领域都有巨大的潜力。然而,当前的多模态对话系统忽略了语音中存在的声学信息,这对于理解人类通信的细微差别至关重要。这种疏忽可能导致对发言者意图的误解,导致对话中的反应不一致甚至相互矛盾。为了弥合这一差距,在本文中,我们提出了PerceptiveAgent,一个移情多模态对话系统,旨在通过整合语音模态感知来识别更深层次或更微妙的意义,超越文字的字面解释。PerceptiveAgent以LLM为认知核心,从输入语音中感知声学信息,并根据自然语言描述的说话风格生成移情响应。实验结果表明,PerceptiveAgent在语境理解方面表现出色,在语言意义与说话人的真实感受相反或不一致的情况下,它能准确地识别说话人的真实意图,产生更细致入微和更有表现力的口语对话。代码可在以下位置公开获取: url{https: github.com Haoqiu-Yan PerceptiveAgent}。摘要:Large Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains. However, current multi-modal dialogue systems overlook the acoustic information present in speech, which is crucial for understanding human communication nuances. This oversight can lead to misinterpretations of speakers' intentions, resulting in inconsistent or even contradictory responses within dialogues. To bridge this gap, in this paper, we propose PerceptiveAgent, an empathetic multi-modal dialogue system designed to discern deeper or more subtle meanings beyond the literal interpretations of words through the integration of speech modality perception. Employing LLMs as a cognitive core, PerceptiveAgent perceives acoustic information from input speech and generates empathetic responses based on speaking styles described in natural language. Experimental results indicate that PerceptiveAgent excels in contextual understanding by accurately discerning the speakers' true intentions in scenarios where the linguistic meaning is either contrary to or inconsistent with the speaker's true feelings, producing more nuanced and expressive spoken dialogues. Code is publicly available at: url{https: github.com Haoqiu-Yan PerceptiveAgent}.
【13】 Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
标题: 弥合差距:集成预训练的语音增强和识别模型以实现稳健的语音识别
作者:Kuan-Chen Wang,You-Jin Li,Wei-Lun Chen,Yu-Wen Chen,Yi-Ching Wang,Ping-Cheng Yeh,Chao Zhang,Yu Tsao
链接:点击下载PDF文件
摘要:噪声鲁棒性是在现实世界中应用自动语音识别(ASR)时的关键。一种解决方案涉及使用语音增强(SE)模型作为ASR的前端。然而,基于神经网络(NN)的SE通常会在增强信号中引入伪影并损害ASR性能,特别是当SE和ASR独立训练时。因此,本研究引入了一种简单而有效的SE后处理技术,以解决各种预训练SE和ASR模型之间的差距。提出了一种轻量级的神经网络桥接模块,用于评估语音信号的信号电平信息。随后,利用信号电平信息,应用观测相加技术,有效地减少了SE的缺点。实验结果表明,我们的方法成功地集成了各种预训练的SE和ASR模型,大大提高了ASR的鲁棒性。至关重要的是,在训练或推理阶段不需要ASR或语音内容的先验知识。此外,这种方法的有效性可以扩展到不同的数据集,而不需要对桥接模块进行微调,从而确保效率和提高泛化能力。摘要:Noise robustness is critical when applying automatic speech recognition (ASR) in real-world scenarios. One solution involves the used of speech enhancement (SE) models as the front end of ASR. However, neural network-based (NN-based) SE often introduces artifacts into the enhanced signals and harms ASR performance, particularly when SE and ASR are independently trained. Therefore, this study introduces a simple yet effective SE post-processing technique to address the gap between various pre-trained SE and ASR models. A bridge module, which is a lightweight NN, is proposed to evaluate the signal-level information of the speech signal. Subsequently, using the signal-level information, the observation addition technique is applied to effectively reduce the shortcomings of SE. The experimental results demonstrate the success of our method in integrating diverse pre-trained SE and ASR models, considerably boosting the ASR robustness. Crucially, no prior knowledge of the ASR or speech contents is required during the training or inference stages. Moreover, the effectiveness of this approach extends to different datasets without necessitating the fine-tuning of the bridge module, ensuring efficiency and improved generalization.
【14】 Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
标题: 使用编码器预处理的多语言E2E语音识别快速语言适应
作者:Yosuke Kashiwagi,Hayato Futami,Emiru Tsunoo,Siddhant Arora,Shinji Watanabe
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:端到端多语言语音识别模型通过单个模型处理多种语言,通常包含语言识别以自动检测传入语音的语言。由于常见的情况是语言已知,因此这些模型可以通过使用语言信息作为提示来执行特定于语言的操作,这对于基于注意力的编码器-解码器架构特别有益。然而,通过联合解码和多任务训练来增强识别的联结主义时间分类(CTC)方法通常不会包含语言提示,因为它的输出标记是条件独立的。为了克服这一点,我们引入了一个编码器提示技术内的自我调节的CTC框架,使特定语言的适应的CTC模型中的zero-shot的方式。我们的方法已被证明可以将错误平均减少28%,在低资源语言上减少41%。摘要:End-to-end multilingual speech recognition models handle multiple languages through a single model, often incorporating language identification to automatically detect the language of incoming speech. Since the common scenario is where the language is already known, these models can perform as language-specific by using language information as prompts, which is particularly beneficial for attention-based encoder-decoder architectures. However, the Connectionist Temporal Classification (CTC) approach, which enhances recognition via joint decoding and multi-task training, does not normally incorporate language prompts due to its conditionally independent output tokens. To overcome this, we introduce an encoder prompting technique within the self-conditioned CTC framework, enabling language-specific adaptation of the CTC model in a zero-shot manner. Our method has shown to significantly reduce errors by 28% on average and by 41% on low-resource languages.
【15】 Integrating Representational Gestures into Automatically Generated Embodied Explanations and its Effects on Understanding and Interaction Quality
标题: 将代表性手势集成到自动生成的预定描述中及其对理解和交互质量的影响
作者:Amelie Sophie Robrecht,Hendric Voss,Lisa Gottschalk,Stefan Kopp
链接:点击下载PDF文件
摘要:在人类交互中,手势具有多种功能,如标记语音节奏,突出关键元素和补充信息。这些手势也可以在解释性的语境中观察到。然而,手势对虚拟代理提供的解释的影响仍然没有得到充分的研究。用户研究进行了调查如何不同类型的手势影响感知的交互质量和听众的理解。本研究旨在探讨手势在解释中的作用,开发一个具身化的虚拟解释器,整合了节拍手势和图标手势,以增强其自动生成的口头解释。我们的模型结合了由一个学习的语音驱动的合成模块与手动捕获的标志性手势生成的节拍手势,支持代理的口头表达棋盘游戏Quarto!作为一种解释。研究结果表明,无论是单独使用的标志性手势,也不是他们的组合与节拍手势优于基线或节拍的条件下的理解。尽管如此,与以前的研究相比,体现代理显着增强理解。摘要:In human interaction, gestures serve various functions such as marking speech rhythm, highlighting key elements, and supplementing information. These gestures are also observed in explanatory contexts. However, the impact of gestures on explanations provided by virtual agents remains underexplored. A user study was carried out to investigate how different types of gestures influence perceived interaction quality and listener understanding. This study addresses the effect of gestures in explanation by developing an embodied virtual explainer integrating both beat gestures and iconic gestures to enhance its automatically generated verbal explanations. Our model combines beat gestures generated by a learned speech-driven synthesis module with manually captured iconic gestures, supporting the agent's verbal expressions about the board game Quarto! as an explanation scenario. Findings indicate that neither the use of iconic gestures alone nor their combination with beat gestures outperforms the baseline or beat-only conditions in terms of understanding. Nonetheless, compared to prior research, the embodied agent significantly enhances understanding.
【16】 Towards Audio Codec-based Speech Separation
标题: 迈向基于音频编解码器的语音分离
作者:Jia Qi Yip,Shengkui Zhao,Dianwen Ng,Eng Siong Chng,Bin Ma
备注:This paper was accepted by Interspeech 2024, Blue Sky Track
链接:点击下载PDF文件
摘要:神经音频编解码器(NAC)模型的最新改进已经引起了人们对采用预训练的编解码器用于各种语音处理应用以利用从高压缩获得的效率的兴趣,但是这些尚未应用于语音分离(SS)任务。SS可以从高压缩中受益,因为传统SS模型所需的计算使得它们对于许多边缘计算用例来说不切实际。然而,SS是波形掩蔽任务,其中压缩倾向于引入严重影响性能的失真。在这里,我们提出了一个新的任务,音频编解码器为基础的SS,其中SS是在NAC的嵌入空间内执行,并提出了一个新的模型,编解码器,来解决这个任务。在推理中,Codecformer实现了MAC减少52倍,同时产生与Sepformer云部署相当的分离性能。该方法为在实际场景中执行高效SS指明了新的方向。摘要:Recent improvements in neural audio codec (NAC) models have generated interest in adopting pre-trained codecs for a variety of speech processing applications to take advantage of the efficiencies gained from high compression, but these have yet been applied to the speech separation (SS) task. SS can benefit from high compression because the compute required for traditional SS models makes them impractical for many edge computing use cases. However, SS is a waveform-masking task where compression tends to introduce distortions that severely impact performance. Here we propose a novel task of Audio Codec-based SS, where SS is performed within the embedding space of a NAC, and propose a new model, Codecformer, to address this task. At inference, Codecformer achieves a 52x reduction in MAC while producing separation performance comparable to a cloud deployment of Sepformer. This method charts a new direction for performing efficient SS in practical scenarios.
【17】 PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
标题: PSLM:针对低延迟口语对话系统的LLM并行生成文本和语音
作者:Kentaro Mitsui,Koh Mitsuda,Toshiaki Wakatsuki,Yukiya Hono,Kei Sawada
备注:8 pages, 4 figures, 4 tables, demo samples: this https URL
链接:点击下载PDF文件
摘要:同时处理文本和语音的多模态语言模型在口语对话系统中具有应用潜力。然而,当前模型在响应生成延迟方面面临两个主要挑战:(1)生成口头响应需要预先生成书面响应,以及(2)语音序列明显长于文本序列。本研究通过扩展语言模型的输入和输出序列来支持文本和语音的并行生成来解决这些问题。我们对口语问答任务的实验表明,我们的方法在保持响应内容质量的同时提高了延迟。此外,我们表明,延迟可以进一步减少在多个序列中生成语音。演示示例可在https: rinnakk.github.io research publications PSLM上获得。摘要:Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating a spoken response requires the prior generation of a written response, and (2) speech sequences are significantly longer than text sequences. This study addresses these issues by extending the input and output sequences of the language model to support the parallel generation of text and speech. Our experiments on spoken question answering tasks demonstrate that our approach improves latency while maintaining the quality of response content. Additionally, we show that latency can be further reduced by generating speech in multiple sequences. Demo samples are available at https: rinnakk.github.io research publications PSLM.
【18】 Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
标题: 在多任务口语理解模型中寻找特定任务的子网络
作者:Hayato Futami,Siddhant Arora,Yosuke Kashiwagi,Emiru Tsunoo,Shinji Watanabe
备注:Accepted to Interspeech2024
链接:点击下载PDF文件
摘要:最近,多任务口语理解(SLU)模型已经出现,旨在解决各种语音处理任务。然而,这些模型通常依赖于大量的参数。此外,他们在适应特定任务的新数据时经常遇到困难,而不会经历灾难性的忘记以前训练过的任务。在这项研究中,我们提出通过神经网络修剪在多任务SLU模型中找到特定于任务的子网络。除了模型压缩之外,我们还希望通过只更新特定于任务的子网络来减轻对先前训练任务的遗忘。我们在最先进的多任务SLU模型“UniverSLU”上进行实验,该模型针对多个任务进行了训练,例如情感识别(ER),意图分类(IC)和自动语音识别(ASR)。我们表明,修剪模型成功地适应了额外的ASR或IC数据,在以前训练的任务上性能下降最小。摘要:Recently, multi-task spoken language understanding (SLU) models have emerged, designed to address various speech processing tasks. However, these models often rely on a large number of parameters. Also, they often encounter difficulties in adapting to new data for a specific task without experiencing catastrophic forgetting of previously trained tasks. In this study, we propose finding task-specific subnetworks within a multi-task SLU model via neural network pruning. In addition to model compression, we expect that the forgetting of previously trained tasks can be mitigated by updating only a task-specific subnetwork. We conduct experiments on top of the state-of-the-art multi-task SLU model UniverSLU'', trained for several tasks such as emotion recognition (ER), intent classification (IC), and automatic speech recognition (ASR). We show that pruned models were successful in adapting to additional ASR or IC data with minimal performance degradation on previously trained tasks.
【19】 JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning
标题: JEN-1 DreamStyler:通过基本参数调整定制音乐概念学习
作者:Boyu Chen,Peike Li,Yao Yao,Alex Wang
链接:点击下载PDF文件
摘要:用于文本到音乐生成的大型模型已经取得了重大进展,促进了从提供的文本提示创建高质量和多样化的音乐作品。然而,输入文本提示可能无法精确地捕获用户需求,特别是当目标是生成体现从指定参考集合导出的特定概念的音乐时。在本文中,我们提出了一种新的方法,定制的文本到音乐的生成,它可以捕获的概念,从两分钟的参考音乐,并生成一个新的音乐作品符合的概念。我们通过使用参考音乐微调预训练的文本到音乐模型来实现这一点。然而,直接微调所有参数会导致过拟合问题。为了解决这个问题,我们提出了一种新的参数调整方法,使模型能够吸收新的概念,同时保持其原有的生成能力。此外,我们在将多个概念引入预训练模型时发现了潜在的概念冲突。我们提出了一个概念增强策略来区分多个概念,使微调模型生成音乐,同时将个人或多个概念。由于我们是第一个致力于定制音乐生成任务的人,因此我们还为新任务引入了新的数据集和评估协议。我们提出的Jen 1-DreamStyler在定性和定量评估方面都优于几个基线。演示将在https: www.jenmusic.ai research#DreamStyler上提供。摘要:Large models for text-to-music generation have achieved significant progress, facilitating the creation of high-quality and varied musical compositions from provided text prompts. However, input text prompts may not precisely capture user requirements, particularly when the objective is to generate music that embodies a specific concept derived from a designated reference collection. In this paper, we propose a novel method for customized text-to-music generation, which can capture the concept from a two-minute reference music and generate a new piece of music conforming to the concept. We achieve this by fine-tuning a pretrained text-to-music model using the reference music. However, directly fine-tuning all parameters leads to overfitting issues. To address this problem, we propose a Pivotal Parameters Tuning method that enables the model to assimilate the new concept while preserving its original generative capabilities. Additionally, we identify a potential concept conflict when introducing multiple concepts into the pretrained model. We present a concept enhancement strategy to distinguish multiple concepts, enabling the fine-tuned model to generate music incorporating either individual or multiple concepts simultaneously. Since we are the first to work on the customized music generation task, we also introduce a new dataset and evaluation protocol for the new task. Our proposed Jen1-DreamStyler outperforms several baselines in both qualitative and quantitative evaluations. Demos will be available at https: www.jenmusic.ai research#DreamStyler.
【20】 Interface Design for Self-Supervised Speech Models
标题: 自我监督语音模型的接口设计
作者:Yi-Jen Shih,David Harwath
备注:Accepted to Interspeech2024
链接:点击下载PDF文件
摘要:自监督语音(SSL)模型最近已被广泛用于许多下游语音处理任务。一般的使用模式是采用SSL模型作为特征提取器,然后训练下游预测头来解决特定的任务。然而,SSL模型的不同层已经被证明可以捕获不同类型的信息,并且组合它们的方法还没有得到很好的研究。为此,我们扩展了SSL模型利用率的一般框架,提出了连接上游和下游的接口。在这种观点下,通过逐层加权和组合特征的主导技术可以被视为特定的接口。我们提出了几种替代的接口设计,并证明了加权和接口是次优的许多任务。特别是,我们表明,卷积接口的深度尺度与上游模型的深度一致优于许多其他接口设计。摘要:Self-supervised speech (SSL) models have recently become widely adopted for many downstream speech processing tasks. The general usage pattern is to employ SSL models as feature extractors, and then train a downstream prediction head to solve a specific task. However, different layers of SSL models have been shown to capture different types of information, and the methods of combining them are not well studied. To this end, we extend the general framework for SSL model utilization by proposing the interface that connects the upstream and downstream. Under this view, the dominant technique of combining features via a layerwise weighted sum can be regarded as a specific interface. We propose several alternative interface designs and demonstrate that the weighted sum interface is suboptimal for many tasks. In particular, we show that a convolutional interface whose depth scales logarithmically with the depth of the upstream model consistently outperforms many other interface designs.
【21】 A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
标题: 语音合成中基于CWT的Mel谱图增强范式
作者:Guoqiang Hu,Huaning Tan,Ruilai Li
链接:点击下载PDF文件
摘要:声学特征对提高合成语音的质量起着重要的作用。目前,Mel谱图是大多数声学模型中广泛使用的声学特征。然而,由于其傅里叶变换过程中造成的细粒度损失,梅尔频谱图合成的语音清晰度在突变信号中受到影响。为了获得更详细的Mel谱图,我们提出了一种基于连续小波变换(CWT)的Mel谱图增强方法。这种模式引入了一个额外的任务:一个更详细的小波频谱图,它像后处理网络一样,将解码器输出的Mel频谱图作为输入。我们选择Tacotron2和Fastspeech2进行实验验证,以测试自回归(AR)和非自回归(NAR)语音系统,分别。实验结果表明,使用该模型与梅尔频谱图增强范例合成的语音表现出更高的MOS,与基线模型相比,分别提高了0.14和0.09。这些发现为增强范例的普遍性提供了一些验证,因为它们证明了该范例在不同架构中的成功。摘要:Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained loss caused by its Fourier transform process, the clarity of speech synthesised by Mel spectrogram is compromised in mutant signals. In order to obtain a more detailed Mel spectrogram, we propose a Mel spectrogram enhancement paradigm based on the continuous wavelet transform (CWT). This paradigm introduces an additional task: a more detailed wavelet spectrogram, which like the post-processing network takes as input the Mel spectrogram output by the decoder. We choose Tacotron2 and Fastspeech2 for experimental validation in order to test autoregressive (AR) and non-autoregressive (NAR) speech systems, respectively. The experimental results demonstrate that the speech synthesised using the model with the Mel spectrogram enhancement paradigm exhibits higher MOS, with an improvement of 0.14 and 0.09 compared to the baseline model, respectively. These findings provide some validation for the universality of the enhancement paradigm, as they demonstrate the success of the paradigm in different architectures.
【22】 A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
标题: 双任务学习方法微调多语言语义语音编码器以实现口语理解
作者:Gaëlle Laperrière,Sahar Ghannay,Bassam Jabaian,Yannick Estève
备注:In Proceedings of Interspeech 2024
链接:点击下载PDF文件
摘要:自我监督学习被广泛用于有效地表示口语理解的语音,逐渐取代传统的方法。同时,提出了文本SSL模型来编码语言无关的语义。SAMU-XLSR框架利用这些语义信息来丰富多语言语音表示。最近的一项研究调查了SAMU-XLSR领域内的语义丰富,通过专门研究下游transmittance,导致最先进的结果,具有挑战性的SLU任务。本研究的兴趣在于损失的多语言性能和缺乏具体的语义训练,这种专业化在没有任何SLU的含义,在接近的语言。我们还考虑SAMU-XLSR的损失,由于一个单独的SLU微调初始跨语言的能力。因此,本文提出了一种双任务学习方法,以提高SAMU-XLSR语义丰富,同时考虑到遥远的语言多语言和语言移植性实验。摘要:Self-Supervised Learning is vastly used to efficiently represent speech for Spoken Language Understanding, gradually replacing conventional approaches. Meanwhile, textual SSL models are proposed to encode language-agnostic semantics. SAMU-XLSR framework employed this semantic information to enrich multilingual speech representations. A recent study investigated SAMU-XLSR in-domain semantic enrichment by specializing it on downstream transcriptions, leading to state-of-the-art results on a challenging SLU task. This study's interest lies in the loss of multilingual performances and lack of specific-semantics training induced by such specialization in close languages without any SLU implication. We also consider SAMU-XLSR's loss of initial cross-lingual abilities due to a separate SLU fine-tuning. Therefore, this paper proposes a dual task learning approach to improve SAMU-XLSR semantic enrichment while considering distant languages for multilingual and language portability experiments.
机器翻译,仅供参考
