今日论文合集:cs.SD语音9篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired  Users using Intermediate ASR Features and Human Memory Models

标题:基于中间ASR特征和人类记忆模型的听障用户非介入性语音清晰度预测
链接:https://arxiv.org/abs/2401.13611
作者:Rhiannon Mogridge,George Close,Robert Sutherland,Thomas Hain,Jon Barker,Stefan Goetze,Anton Ragni
备注:Accepted paper. IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), Seoul, Korea, April 2024
摘要:神经网络已经成功地用于非侵入式语音可懂度预测。最近,人们发现使用来自预训练的自监督和弱监督模型的中间层的特征表示对这项任务特别有用。这项工作结合了使用Whisper ASR解码器层表示作为神经网络输入功能与基于范例的,心理动机的人类记忆模型,以预测助听器用户的人类可懂度评级。大量的性能改善,建立侵入HASPI基线系统,包括增强系统和听众看不见的训练数据,与25.3的均方根误差相比,基线为28.7。
摘要:Neural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work combines the use of Whisper ASR decoder layer representations as neural network input features with an exemplar-based, psychologically motivated model of human memory to predict human intelligibility ratings for hearing-aid users. Substantial performance improvement over an established intrusive HASPI baseline system is found, including on enhancement systems and listeners unseen in the training data, with a root mean squared error of 25.3 compared with the baseline of 28.7.

【2】 A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
标题:多通道语音增强算法的音素尺度评价
链接:https://arxiv.org/abs/2401.13548
作者:Nasser-Eddine Monir,Paul Magron,Romain Serizel
备注:This is the preprint of the paper that we submitted to the Trends in Hearing Journal
摘要:在复杂的声学环境中,语音清晰度受到噪声和混响的挑战,多通道语音增强成为听力损失患者的一种有前途的解决方案。这样的算法通常在话语级进行评估。然而,这种方法忽略了特定音素分析所揭示的颗粒声学细微差别,可能会模糊对其性能的关键见解。本文提出了一个深入的音素尺度评估的3个国家的最先进的多通道语音增强算法。这些算法-- FasNet、MVDR和Tango --在不同的噪声条件和空间设置下进行了广泛的评估,采用了具有实测房间脉冲响应的逼真声学模拟,并利用了双耳听力设置中多个麦克风提供的多样性。该研究强调了细粒度的音素级分析,揭示了虽然像爆破音这样的一些音素受到环境声学的严重影响,并且具有算法处理的挑战性,但其他像鼻音和咝咝声这样的音素在增强后得到了实质性的改善。这些研究表明,在嘈杂的条件下,音素清晰度有了重要的改善,这些见解可以推动更加个性化和音素感知的助听器技术的发展。
摘要:In the intricate acoustic landscapes where speech intelligibility is challenged by noise and reverberation, multichannel speech enhancement emerges as a promising solution for individuals with hearing loss. Such algorithms are commonly evaluated at the utterance level. However, this approach overlooks the granular acoustic nuances revealed by phoneme-specific analysis, potentially obscuring key insights into their performance. This paper presents an in-depth phoneme-scale evaluation of 3 state-of-the-art multichannel speech enhancement algorithms. These algorithms -- FasNet, MVDR, and Tango -- are extensively evaluated across different noise conditions and spatial setups, employing realistic acoustic simulations with measured room impulse responses, and leveraging diversity offered by multiple microphones in a binaural hearing setup. The study emphasizes the fine-grained phoneme-level analysis, revealing that while some phonemes like plosives are heavily impacted by environmental acoustics and challenging to deal with by the algorithms, others like nasals and sibilants see substantial improvements after enhancement. These investigations demonstrate important improvements in phoneme clarity in noisy conditions, with insights that could drive the development of more personalized and phoneme-aware hearing aid technologies.

【3】 SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
标题:SpeechGPT-Gen:可伸缩信息链语音生成
链接:https://arxiv.org/abs/2401.13527
作者:Dong Zhang,Xin Zhang,Jun Zhan,Shimin Li,Yaqian Zhou,Xipeng Qiu
备注:work in progress
摘要:得益于有效的语音建模,当前的语音大语言模型(SLLM)在上下文语音生成和对未知说话人的有效泛化方面表现出了卓越的能力。然而,流行的信息建模过程是由某些冗余,导致语音生成效率低下。我们提出了信息链生成(CoIG),在大规模语音生成的语义和感知信息解耦的方法。在此基础上,我们开发了SpeechGPT-Gen,这是一个在语义和感知信息建模方面具有80亿参数的SLLM。它包括一个基于LLM的自回归模型的语义信息建模和一个非自回归模型采用流匹配的感知信息建模。此外,我们引入了新的方法,注入语义信息的先验分布,以提高流匹配的效率。大量的实验结果表明,SpeechGPT-Gen在zero-shot文本到语音,zero-shot语音转换和语音到语音对话方面表现出色,强调了CoIG在捕捉和建模语音的语义和感知维度方面的卓越能力。代码和模型可在https://github.com/0nutation/SpeechGPT上获得。
摘要:Benefiting from effective speech modeling, current Speech Large Language Models (SLLMs) have demonstrated exceptional capabilities in in-context speech generation and efficient generalization to unseen speakers. However, the prevailing information modeling process is encumbered by certain redundancies, leading to inefficiencies in speech generation. We propose Chain-of-Information Generation (CoIG), a method for decoupling semantic and perceptual information in large-scale speech generation. Building on this, we develop SpeechGPT-Gen, an 8-billion-parameter SLLM efficient in semantic and perceptual information modeling. It comprises an autoregressive model based on LLM for semantic information modeling and a non-autoregressive model employing flow matching for perceptual information modeling. Additionally, we introduce the novel approach of infusing semantic information into the prior distribution to enhance the efficiency of flow matching. Extensive experimental results demonstrate that SpeechGPT-Gen markedly excels in zero-shot text-to-speech, zero-shot voice conversion, and speech-to-speech dialogue, underscoring CoIG's remarkable proficiency in capturing and modeling speech's semantic and perceptual dimensions. Code and models are available at https://github.com/0nutation/SpeechGPT.

【4】 Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific  Input Representation and Diffusion Outpainting
标题:具有乐器特定输入表示和扩散外绘的表现性声学吉他声音合成
链接:https://arxiv.org/abs/2401.13498
作者:Hounsu Kim,Soonbeom Choi,Juhan Nam
备注:Accepted to ICASSP 2024
摘要:合成表演吉他声音是一个极具挑战性的任务,由于复调和表达的高度可变性。最近,深度生成模型在从乐谱合成富有表现力的复调乐器声音方面显示出了很好的结果,通常使用通用的输入。在这项工作中,我们提出了一个富有表现力的原声吉他声音合成模型与定制的输入表示的仪器,我们称之为guitarroll。我们实现了所提出的方法,使用基于扩散的outpainting,可以生成长期一致性的音频。为了克服缺乏吉他/音频配对数据集的问题,我们不仅使用了现有的吉他数据集,还从高质量的基于样本的吉他合成器中收集了数据。通过定量和定性的评估,我们表明,我们提出的模型具有更高的音频质量比基线模型,并产生更逼真的音色比以前的领先工作的声音。
摘要:Synthesizing performing guitar sound is a highly challenging task due to the polyphony and high variability in expression. Recently, deep generative models have shown promising results in synthesizing expressive polyphonic instrument sounds from music scores, often using a generic MIDI input. In this work, we propose an expressive acoustic guitar sound synthesis model with a customized input representation to the instrument, which we call guitarroll. We implement the proposed approach using diffusion-based outpainting which can generate audio with long-term consistency. To overcome the lack of MIDI/audio-paired datasets, we used not only an existing guitar dataset but also collected data from a high quality sample-based guitar synthesizer. Through quantitative and qualitative evaluations, we show that our proposed model has higher audio quality than the baseline model and generates more realistic timbre sounds than the previous leading work.

【5】 SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken  Question Answering
标题:SpeechDPR:面向开放领域口语答疑的端到端口语通道检索
链接:https://arxiv.org/abs/2401.13463
作者:Chyi-Jiunn Lin,Guan-Ting Lin,Yung-Sung Chuang,Wei-Lun Wu,Shang-Wen Li,Abdelrahman Mohamed,Hung-yi Lee,Lin-shan Lee
备注:Accepted at ICASSP 2024
摘要:口语问题分类(SQA)是机器通过在给定的口语段落中找到答案范围来回答用户问题的关键。SQA以前在没有ASR的情况下实现,以避免识别错误和词汇表外(OOV)问题。然而,开放域SQA(openSQA)的现实问题,即机器需要首先从口语存档中检索可能包含答案的段落,从未被考虑过。本文提出了第一个已知的端到端的框架,语音密集通道检索(SpeechDPR),检索组件的openSQA问题。SpeechDPR通过从无监督ASR(UASR)和文本密集检索器(TDR)的级联模型中提取知识来学习一个高级语义表示。不需要人工转录的语音数据。初步实验表明,性能与UASR和TDR的级联模型相当,并且在UASR较差时明显更好,验证了该方法对语音识别错误更鲁棒。
摘要:Spoken Question Answering (SQA) is essential for machines to reply to user's question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and Out-of-Vocabulary (OOV) problems. However, the real-world problem of Open-domain SQA (openSQA), in which the machine needs to first retrieve passages that possibly contain the answer from a spoken archive in addition, was never considered. This paper proposes the first known end-to-end framework, Speech Dense Passage Retriever (SpeechDPR), for the retrieval component of the openSQA problem. SpeechDPR learns a sentence-level semantic representation by distilling knowledge from the cascading model of unsupervised ASR (UASR) and text dense retriever (TDR). No manually transcribed speech data is needed. Initial experiments showed performance comparable to the cascading model of UASR and TDR, and significantly better when UASR was poor, verifying this approach is more robust to speech recognition errors.

【6】 MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion,  ASR Error Detection, and ASR Error Correction
标题:MF-AED-AEC:利用多模式融合、ASR错误检测和ASR错误纠正的语音情感识别
链接:https://arxiv.org/abs/2401.13260
作者:Jiajun He,Xiaohan Shi,Xingfeng Li,Tomoki Toda
备注:Accepted by ICASSP 2024
摘要:语音情感识别(SER)中的流行方法涉及将音频和文本信息两者结合以全面地识别说话者的情感,其中文本通常通过自动语音识别(ASR)获得。这种方法的一个重要问题是,从文本模态的ASR错误可以恶化SER的性能。以前的研究提出了使用辅助ASR错误检测任务,自适应地分配权重的ASR假设中的每个字。然而,这种方法的改进潜力有限,因为它没有解决文本中语义信息的连贯性。此外,不同模态的固有异质性导致其表示之间的分布差距,使其融合具有挑战性。因此,在本文中,我们将两个辅助任务,ASR错误检测(AED)和ASR错误校正(AEC),以提高ASR文本的语义一致性,并进一步引入一种新的多模态融合(MF)方法来学习跨模态的共享表示。我们将我们的方法称为MF-AED-AEC。实验结果表明,MF-AED-AEC模型的性能比基线模型高出4.1%.
摘要:The prevalent approach in speech emotion recognition (SER) involves integrating both audio and textual information to comprehensively identify the speaker's emotion, with the text generally obtained through automatic speech recognition (ASR). An essential issue of this approach is that ASR errors from the text modality can worsen the performance of SER. Previous studies have proposed using an auxiliary ASR error detection task to adaptively assign weights of each word in ASR hypotheses. However, this approach has limited improvement potential because it does not address the coherence of semantic information in the text. Additionally, the inherent heterogeneity of different modalities leads to distribution gaps between their representations, making their fusion challenging. Therefore, in this paper, we incorporate two auxiliary tasks, ASR error detection (AED) and ASR error correction (AEC), to enhance the semantic coherence of ASR text, and further introduce a novel multi-modal fusion (MF) method to learn shared representations across modalities. We refer to our method as MF-AED-AEC. Experimental results indicate that MF-AED-AEC significantly outperforms the baseline model by a margin of 4.1\%.

【7】 TranSentence: Speech-to-speech Translation via Language-agnostic  Sentence-level Speech Encoding without Language-parallel Data
标题:TransSentence:无语言并行数据的语言不可知的句子级语音编码的语音到语音翻译
链接:https://arxiv.org/abs/2401.12992
作者:Seung-Bin Kim,Sang-Hoon Lee,Seong-Whan Lee
备注:Accepted by ICASSP 2024
摘要:虽然语音到语音翻译领域已经取得了重大进展,但传统的模型仍然需要源语言和目标语言之间的语言并行语音数据进行训练。在本文中,我们介绍了TranSentence,一种新的语音到语音翻译没有语言并行语音数据。为了实现这一点,我们首先采用了一种语言不可知论的语音编码,捕捉语音的语义信息,无论语言。然后,我们训练我们的模型,以生成语音的基础上,从一个语言无关的语音级语音编码器,是预先训练与各种语言的编码嵌入。通过这种方法,尽管只在目标语言的单语数据上进行训练,但我们可以在推理阶段使用与语言无关的语音嵌入从源语言语音中生成目标语言语音。此外,我们将TranSentence扩展到多语言语音到语音翻译。实验结果表明,TranSentence模型优于其他模型。
摘要:Although there has been significant advancement in the field of speech-to-speech translation, conventional models still require language-parallel speech data between the source and target languages for training. In this paper, we introduce TranSentence, a novel speech-to-speech translation without language-parallel speech data. To achieve this, we first adopt a language-agnostic sentence-level speech encoding that captures the semantic information of speech, irrespective of language. We then train our model to generate speech based on the encoded embedding obtained from a language-agnostic sentence-level speech encoder that is pre-trained with various languages. With this method, despite training exclusively on the target language's monolingual data, we can generate target language speech in the inference stage using language-agnostic speech embedding from the source language speech. Furthermore, we extend TranSentence to multilingual speech-to-speech translation. The experimental results demonstrate that TranSentence is superior to other models.

【8】 TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition  in Conversation
标题:TelME:教师主导的多模式会话情感识别融合网络
链接:https://arxiv.org/abs/2401.12987
作者:Taeyang Yun,Hyunkuk Lim,Jeonghwan Lee,Min Song
备注:13 pages, 7 figures
摘要:会话中的情感识别(ERC)在使对话系统能够有效地响应用户请求方面起着至关重要的作用。会话中的情感可以通过来自各种模态(诸如音频、视觉和文本)的表示来识别。然而,由于非言语模态对情绪识别的贡献很小,多模态ERC一直被认为是一项具有挑战性的任务。在本文中,我们提出了教师领导的多模态融合网络ERC(TelME)。TelME采用跨模态知识蒸馏,将信息从充当教师的语言模型传递给非语言学生,从而优化弱模态的功效。然后,我们结合多模态功能使用移位融合的方法,其中学生网络支持教师。TelME在MELD中实现了最先进的性能,MELD是ERC的多说话者对话数据集。最后,我们通过额外的实验证明了我们的组件的有效性。
摘要:Emotion Recognition in Conversation (ERC) plays a crucial role in enabling dialogue systems to effectively respond to user requests. The emotions in a conversation can be identified by the representations from various modalities, such as audio, visual, and text. However, due to the weak contribution of non-verbal modalities to recognize emotions, multimodal ERC has always been considered a challenging task. In this paper, we propose Teacher-leading Multimodal fusion network for ERC (TelME). TelME incorporates cross-modal knowledge distillation to transfer information from a language model acting as the teacher to the non-verbal students, thereby optimizing the efficacy of the weak modalities. We then combine multimodal features using a shifting fusion approach in which student networks support the teacher. TelME achieves state-of-the-art performance in MELD, a multi-speaker conversation dataset for ERC. Finally, we demonstrate the effectiveness of our components through additional experiments.


【9】 Locality enhanced dynamic biasing and sampling strategies for contextual  ASR
标题:上下文ASR的局部性增强型动态偏差和采样策略
链接:https://arxiv.org/abs/2401.13146
作者:Md Asif Jalal,Pablo Peso Parada,George Pavlidis,Vasileios Moschopoulos,Karthikeyan Saravanan,Chrysovalantis-Giorgos Kontoulis,Jisi Zhang,Anastasios Drosou,Gil Ho Lee,Jungin Lee,Seokyeong Jung
备注:Accepted for IEEE ASRU 2023
摘要:自动语音识别(ASR)在识别时变稀有短语时仍然面临挑战。上下文偏置(CB)模块偏向ASR模型对这样的上下文相关的短语。在训练过程中,根据采样策略从大量短语中选择一系列偏置短语。在这项工作中,我们首先分析了不同的采样策略,以提供洞察力的训练CB的ASR与相关图之间的偏见嵌入在各个训练阶段。其次,我们引入了邻域注意(NA),将自我注意(SA)定位到最近的相邻帧,以进一步改进CB输出。结果表明,该方法提供了平均25.84%的相对WER提高LibriSpeech集和稀有词的评价相比,基线。
摘要:Automatic Speech Recognition (ASR) still face challenges when recognizing time-variant rare-phrases. Contextual biasing (CB) modules bias ASR model towards such contextually-relevant phrases. During training, a list of biasing phrases are selected from a large pool of phrases following a sampling strategy. In this work we firstly analyse different sampling strategies to provide insights into the training of CB for ASR with correlation plots between the bias embeddings among various training stages. Secondly, we introduce a neighbourhood attention (NA) that localizes self attention (SA) to the nearest neighbouring frames to further refine the CB output. The results show that this proposed approach provides on average a 25.84% relative WER improvement on LibriSpeech sets and rare-word evaluation compared to the baseline.


eess.AS音频处理
【1】 Perceptually-motivated Spatial Audio Codec for Higher-Order Ambisonics  Compression
标题:用于高阶双音压缩的感知激励空间音频编解码器
链接:https://arxiv.org/abs/2401.13401
作者:Christoph Hold,Leo McCormack,Archontis Politis,Ville Pulkki
备注:Accepted for publication in Proceedings of the 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
摘要:基于场景的空间音频格式(诸如高保真度立体声响复制)是回放系统不可知的,并且因此可以有利于将沉浸式音频体验递送到广泛的(潜在未知的)设备。然而,传送高空间分辨率高保真立体声音频所需的通道数量对于低带宽应用来说可能是过高的。因此,本文提出了一种压缩编解码器,这是基于参数高阶方向性音频编码(HO-DirAC)模型。编码器将高阶高保真度立体声(HOA)输入音频下混为减少数量的信号,这些信号伴随有感知激发的场景参数。使用感知音频编码器对下混音频进行编码,而将参数分组到感知频带中、量化并下采样。在解码器侧,低Ambisonic阶数被完全恢复。根据这些参数合成了不完全可恢复的HOA分量。收听测试的结果表明,所提出的参数空间音频编解码器可以改善所采用的感知音频编码器,特别是在低到中高比特率,当应用于五阶HOA信号。
摘要:Scene-based spatial audio formats, such as Ambisonics, are playback system agnostic and may therefore be favoured for delivering immersive audio experiences to a wide range of (potentially unknown) devices. The number of channels required to deliver high spatial resolution Ambisonic audio, however, can be prohibitive for low-bandwidth applications. Therefore, this paper proposes a compression codec, which is based upon the parametric higher-order Directional Audio Coding (HO-DirAC) model. The encoder downmixes the higher-order Ambisonic (HOA) input audio into a reduced number of signals, which are accompanied by perceptually-motivated scene parameters. The downmixed audio is coded using a perceptual audio coder, whereas the parameters are grouped into perceptual bands, quantized, and downsampled. On the decoder side, low Ambisonic orders are fully recovered. Not fully recoverable HOA components are synthesized according to the parameters. The results of a listening test indicate that the proposed parametric spatial audio codec can improve the adopted perceptual audio coder, especially at low to medium-high bitrates, when applied to fifth-order HOA signals.


【2】 SCNet: Sparse Compression Network for Music Source Separation
标题:SCNet:用于音乐源分离的稀疏压缩网络
链接:https://arxiv.org/abs/2401.13276
作者:Weinan Tong,Jiaxu Zhu,Jun Chen,Shiyin Kang,Tao Jiang,Yang Li,Zhiyong Wu,Helen Meng
备注:Accepted by ICASSP 2024
摘要:基于深度学习的方法在音乐源分离方面取得了重大成就。然而,在保持低模型复杂度的同时获得良好的结果仍然是超宽带音乐源分离的挑战。以往的工作要么忽略了子带的差异,或没有充分解决的问题时,产生子带特征的信息丢失。在本文中,我们提出了SCNet,一种新的频域网络,明确地将混合的频谱图分成几个子带,并引入一个基于稀疏性的编码器来模拟不同的频带。我们使用较高的压缩比的子带与较少的信息,以提高信息密度,并专注于建模子带与更多的信息。以这种方式,可以使用较低的计算消耗来显著提高分离性能。实验结果表明,该模型在不使用额外数据的情况下,在MUSDB 18-HQ数据集上实现了9.0 dB的信号失真比(SDR),优于现有的方法。具体来说,SCNet的CPU推理时间仅为HT Demucs的48%,HT Demucs是之前最先进的模型之一。
摘要:Deep learning-based methods have made significant achievements in music source separation. However, obtaining good results while maintaining a low model complexity remains challenging in super wide-band music source separation. Previous works either overlook the differences in subbands or inadequately address the problem of information loss when generating subband features. In this paper, we propose SCNet, a novel frequency-domain network to explicitly split the spectrogram of the mixture into several subbands and introduce a sparsity-based encoder to model different frequency bands. We use a higher compression ratio on subbands with less information to improve the information density and focus on modeling subbands with more information. In this way, the separation performance can be significantly improved using lower computational consumption. Experiment results show that the proposed model achieves a signal to distortion ratio (SDR) of 9.0 dB on the MUSDB18-HQ dataset without using extra data, which outperforms state-of-the-art methods. Specifically, SCNet's CPU inference time is only 48% of HT Demucs, one of the previous state-of-the-art models.

【3】 MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score  Prediction
标题:MOS-FAD:通过自动平均评分预测改进伪音频检测
链接:https://arxiv.org/abs/2401.13249
作者:Wangjin Zhou,Zhengdong Yang,Chenhui Chu,Sheng Li,Raj Dabre,Yi Zhao,Kawahara Tatsuya
备注:Accepted in ICASSP2024
摘要:自动平均意见分数(MOS)预测被用来评估合成语音的质量。这项研究将预测MOS的应用扩展到假音频检测(FAD)的任务中,因为我们期望MOS可以用于评估合成语音与自然人声的接近程度。我们提出了MOS-FAD,其中MOS可以在FAD中的两个关键点上利用:训练数据选择和模型融合。在训练数据选择中,我们证明了MOS能够有效地过滤来自不平衡数据集的样本。在模型融合中,我们的研究结果表明,将MOS作为FAD模型融合的门控机制,提高了整体性能。
摘要:Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD), as we expect that MOS can be used to assess how close synthesized speech is to the natural human voice. We propose MOS-FAD, where MOS can be leveraged at two key points in FAD: training data selection and model fusion. In training data selection, we demonstrate that MOS enables effective filtering of samples from unbalanced datasets. In the model fusion, our results demonstrate that incorporating MOS as a gating mechanism in FAD model fusion enhances overall performance.

【4】 Locality enhanced dynamic biasing and sampling strategies for contextual  ASR
标题:上下文ASR的局部性增强型动态偏差和采样策略
链接:https://arxiv.org/abs/2401.13146
作者:Md Asif Jalal,Pablo Peso Parada,George Pavlidis,Vasileios Moschopoulos,Karthikeyan Saravanan,Chrysovalantis-Giorgos Kontoulis,Jisi Zhang,Anastasios Drosou,Gil Ho Lee,Jungin Lee,Seokyeong Jung
备注:Accepted for IEEE ASRU 2023
摘要:自动语音识别(ASR)在识别时变稀有短语时仍然面临挑战。上下文偏置(CB)模块偏向ASR模型对这样的上下文相关的短语。在训练过程中,根据采样策略从大量短语中选择一系列偏置短语。在这项工作中,我们首先分析了不同的采样策略,以提供洞察力的训练CB的ASR与相关图之间的偏见嵌入在各个训练阶段。其次,我们引入了邻域注意(NA),将自我注意(SA)定位到最近的相邻帧,以进一步改进CB输出。结果表明,该方法提供了平均25.84%的相对WER提高LibriSpeech集和稀有词的评价相比,基线。
摘要:Automatic Speech Recognition (ASR) still face challenges when recognizing time-variant rare-phrases. Contextual biasing (CB) modules bias ASR model towards such contextually-relevant phrases. During training, a list of biasing phrases are selected from a large pool of phrases following a sampling strategy. In this work we firstly analyse different sampling strategies to provide insights into the training of CB for ASR with correlation plots between the bias embeddings among various training stages. Secondly, we introduce a neighbourhood attention (NA) that localizes self attention (SA) to the nearest neighbouring frames to further refine the CB output. The results show that this proposed approach provides on average a 25.84% relative WER improvement on LibriSpeech sets and rare-word evaluation compared to the baseline.

【5】 Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired  Users using Intermediate ASR Features and Human Memory Models
标题:基于中间ASR特征和人类记忆模型的听障用户非介入性语音清晰度预测
链接:https://arxiv.org/abs/2401.13611
作者:Rhiannon Mogridge,George Close,Robert Sutherland,Thomas Hain,Jon Barker,Stefan Goetze,Anton Ragni
备注:Accepted paper. IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), Seoul, Korea, April 2024
摘要:神经网络已经成功地用于非侵入式语音可懂度预测。最近,人们发现使用来自预训练的自监督和弱监督模型的中间层的特征表示对这项任务特别有用。这项工作结合了使用Whisper ASR解码器层表示作为神经网络输入功能与基于范例的,心理动机的人类记忆模型,以预测助听器用户的人类可懂度评级。大量的性能改善,建立侵入HASPI基线系统,包括增强系统和听众看不见的训练数据,与25.3的均方根误差相比,基线为28.7。
摘要:Neural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work combines the use of Whisper ASR decoder layer representations as neural network input features with an exemplar-based, psychologically motivated model of human memory to predict human intelligibility ratings for hearing-aid users. Substantial performance improvement over an established intrusive HASPI baseline system is found, including on enhancement systems and listeners unseen in the training data, with a root mean squared error of 25.3 compared with the baseline of 28.7.


【6】 A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
标题:多通道语音增强算法的音素尺度评估
链接:https://arxiv.org/abs/2401.13548
作者:Nasser-Eddine Monir,Paul Magron,Romain Serizel
备注:This is the preprint of the paper that we submitted to the Trends in Hearing Journal
摘要:在复杂的声学环境中,语音清晰度受到噪声和混响的挑战,多通道语音增强成为听力损失患者的一种有前途的解决方案。这样的算法通常在话语级进行评估。然而,这种方法忽略了特定音素分析所揭示的颗粒声学细微差别,可能会模糊对其性能的关键见解。本文提出了一个深入的音素尺度评估的3个国家的最先进的多通道语音增强算法。这些算法-- FasNet、MVDR和Tango --在不同的噪声条件和空间设置下进行了广泛的评估,采用了具有实测房间脉冲响应的逼真声学模拟,并利用了双耳听力设置中多个麦克风提供的多样性。该研究强调了细粒度的音素级分析,揭示了虽然像爆破音这样的一些音素受到环境声学的严重影响,并且具有算法处理的挑战性,但其他像鼻音和咝咝声这样的音素在增强后得到了实质性的改善。这些研究表明,在嘈杂的条件下,音素清晰度有了重要的改善,这些见解可以推动更加个性化和音素感知的助听器技术的发展。
摘要:In the intricate acoustic landscapes where speech intelligibility is challenged by noise and reverberation, multichannel speech enhancement emerges as a promising solution for individuals with hearing loss. Such algorithms are commonly evaluated at the utterance level. However, this approach overlooks the granular acoustic nuances revealed by phoneme-specific analysis, potentially obscuring key insights into their performance. This paper presents an in-depth phoneme-scale evaluation of 3 state-of-the-art multichannel speech enhancement algorithms. These algorithms -- FasNet, MVDR, and Tango -- are extensively evaluated across different noise conditions and spatial setups, employing realistic acoustic simulations with measured room impulse responses, and leveraging diversity offered by multiple microphones in a binaural hearing setup. The study emphasizes the fine-grained phoneme-level analysis, revealing that while some phonemes like plosives are heavily impacted by environmental acoustics and challenging to deal with by the algorithms, others like nasals and sibilants see substantial improvements after enhancement. These investigations demonstrate important improvements in phoneme clarity in noisy conditions, with insights that could drive the development of more personalized and phoneme-aware hearing aid technologies.

【7】 SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
标题:SpeechGPT-Gen:可伸缩信息链语音生成
链接:https://arxiv.org/abs/2401.13527
作者:Dong Zhang,Xin Zhang,Jun Zhan,Shimin Li,Yaqian Zhou,Xipeng Qiu
备注:work in progress
摘要:得益于有效的语音建模,当前的语音大语言模型(SLLM)在上下文语音生成和对未知说话人的有效泛化方面表现出了卓越的能力。然而,流行的信息建模过程是由某些冗余,导致语音生成效率低下。我们提出了信息链生成(CoIG),在大规模语音生成的语义和感知信息解耦的方法。在此基础上,我们开发了SpeechGPT-Gen,这是一个在语义和感知信息建模方面具有80亿参数的SLLM。它包括一个基于LLM的自回归模型的语义信息建模和一个非自回归模型采用流匹配的感知信息建模。此外,我们引入了新的方法,注入语义信息的先验分布,以提高流匹配的效率。大量的实验结果表明,SpeechGPT-Gen在zero-shot文本到语音,zero-shot语音转换和语音到语音对话方面表现出色,强调了CoIG在捕捉和建模语音的语义和感知维度方面的卓越能力。代码和模型可在https://github.com/0nutation/SpeechGPT上获得。
摘要:Benefiting from effective speech modeling, current Speech Large Language Models (SLLMs) have demonstrated exceptional capabilities in in-context speech generation and efficient generalization to unseen speakers. However, the prevailing information modeling process is encumbered by certain redundancies, leading to inefficiencies in speech generation. We propose Chain-of-Information Generation (CoIG), a method for decoupling semantic and perceptual information in large-scale speech generation. Building on this, we develop SpeechGPT-Gen, an 8-billion-parameter SLLM efficient in semantic and perceptual information modeling. It comprises an autoregressive model based on LLM for semantic information modeling and a non-autoregressive model employing flow matching for perceptual information modeling. Additionally, we introduce the novel approach of infusing semantic information into the prior distribution to enhance the efficiency of flow matching. Extensive experimental results demonstrate that SpeechGPT-Gen markedly excels in zero-shot text-to-speech, zero-shot voice conversion, and speech-to-speech dialogue, underscoring CoIG's remarkable proficiency in capturing and modeling speech's semantic and perceptual dimensions. Code and models are available at https://github.com/0nutation/SpeechGPT.

【8】 Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific  Input Representation and Diffusion Outpainting
标题:具有乐器特定输入表示和扩散外观的表现性声学吉他声音合成
链接:https://arxiv.org/abs/2401.13498
作者:Hounsu Kim,Soonbeom Choi,Juhan Nam
备注:Accepted to ICASSP 2024
摘要:合成表演吉他声音是一个极具挑战性的任务,由于复调和表达的高度可变性。最近,深度生成模型在从乐谱合成富有表现力的复调乐器声音方面显示出了很好的结果,通常使用通用的输入。在这项工作中,我们提出了一个富有表现力的原声吉他声音合成模型与定制的输入表示的仪器,我们称之为guitarroll。我们实现了所提出的方法,使用基于扩散的outpainting,可以生成长期一致性的音频。为了克服缺乏吉他/音频配对数据集的问题,我们不仅使用了现有的吉他数据集,还从高质量的基于样本的吉他合成器中收集了数据。通过定量和定性的评估,我们表明,我们提出的模型具有更高的音频质量比基线模型,并产生更逼真的音色比以前的领先工作的声音。
摘要:Synthesizing performing guitar sound is a highly challenging task due to the polyphony and high variability in expression. Recently, deep generative models have shown promising results in synthesizing expressive polyphonic instrument sounds from music scores, often using a generic MIDI input. In this work, we propose an expressive acoustic guitar sound synthesis model with a customized input representation to the instrument, which we call guitarroll. We implement the proposed approach using diffusion-based outpainting which can generate audio with long-term consistency. To overcome the lack of MIDI/audio-paired datasets, we used not only an existing guitar dataset but also collected data from a high quality sample-based guitar synthesizer. Through quantitative and qualitative evaluations, we show that our proposed model has higher audio quality than the baseline model and generates more realistic timbre sounds than the previous leading work.

【9】 SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken  Question Answering
标题:SpeechDPR:面向开放领域口语答疑的端到端口语通道检索
链接:https://arxiv.org/abs/2401.13463
作者:Chyi-Jiunn Lin,Guan-Ting Lin,Yung-Sung Chuang,Wei-Lun Wu,Shang-Wen Li,Abdelrahman Mohamed,Hung-yi Lee,Lin-shan Lee
备注:Accepted at ICASSP 2024
摘要:口语问题分类(SQA)是机器通过在给定的口语段落中找到答案范围来回答用户问题的关键。SQA以前在没有ASR的情况下实现,以避免识别错误和词汇表外(OOV)问题。然而,开放域SQA(openSQA)的现实问题,即机器需要首先从口语存档中检索可能包含答案的段落,从未被考虑过。本文提出了第一个已知的端到端的框架,语音密集通道检索(SpeechDPR),检索组件的openSQA问题。SpeechDPR通过从无监督ASR(UASR)和文本密集检索器(TDR)的级联模型中提取知识来学习一个高级语义表示。不需要人工转录的语音数据。初步实验表明,性能与UASR和TDR的级联模型相当,并且在UASR较差时明显更好,验证了该方法对语音识别错误更鲁棒。
摘要:Spoken Question Answering (SQA) is essential for machines to reply to user's question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and Out-of-Vocabulary (OOV) problems. However, the real-world problem of Open-domain SQA (openSQA), in which the machine needs to first retrieve passages that possibly contain the answer from a spoken archive in addition, was never considered. This paper proposes the first known end-to-end framework, Speech Dense Passage Retriever (SpeechDPR), for the retrieval component of the openSQA problem. SpeechDPR learns a sentence-level semantic representation by distilling knowledge from the cascading model of unsupervised ASR (UASR) and text dense retriever (TDR). No manually transcribed speech data is needed. Initial experiments showed performance comparable to the cascading model of UASR and TDR, and significantly better when UASR was poor, verifying this approach is more robust to speech recognition errors.


【10】 MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion,  ASR Error Detection, and ASR Error Correction
标题:MF-AED-AEC:利用多模式融合、ASR错误检测和ASR错误纠正的语音情感识别
链接:https://arxiv.org/abs/2401.13260
作者:Jiajun He,Xiaohan Shi,Xingfeng Li,Tomoki Toda
备注:Accepted by ICASSP 2024
摘要:语音情感识别(SER)中的流行方法涉及将音频和文本信息两者结合以全面地识别说话者的情感,其中文本通常通过自动语音识别(ASR)获得。这种方法的一个重要问题是,从文本模态的ASR错误可以恶化SER的性能。以前的研究提出了使用辅助ASR错误检测任务,自适应地分配权重的ASR假设中的每个字。然而,这种方法的改进潜力有限,因为它没有解决文本中语义信息的连贯性。此外,不同模态的固有异质性导致其表示之间的分布差距,使其融合具有挑战性。因此,在本文中,我们将两个辅助任务,ASR错误检测(AED)和ASR错误校正(AEC),以提高ASR文本的语义一致性,并进一步引入一种新的多模态融合(MF)方法来学习跨模态的共享表示。我们将我们的方法称为MF-AED-AEC。实验结果表明,MF-AED-AEC模型的性能比基线模型高出4.1%.
摘要:The prevalent approach in speech emotion recognition (SER) involves integrating both audio and textual information to comprehensively identify the speaker's emotion, with the text generally obtained through automatic speech recognition (ASR). An essential issue of this approach is that ASR errors from the text modality can worsen the performance of SER. Previous studies have proposed using an auxiliary ASR error detection task to adaptively assign weights of each word in ASR hypotheses. However, this approach has limited improvement potential because it does not address the coherence of semantic information in the text. Additionally, the inherent heterogeneity of different modalities leads to distribution gaps between their representations, making their fusion challenging. Therefore, in this paper, we incorporate two auxiliary tasks, ASR error detection (AED) and ASR error correction (AEC), to enhance the semantic coherence of ASR text, and further introduce a novel multi-modal fusion (MF) method to learn shared representations across modalities. We refer to our method as MF-AED-AEC. Experimental results indicate that MF-AED-AEC significantly outperforms the baseline model by a margin of 4.1\%.

【11】 TranSentence: Speech-to-speech Translation via Language-agnostic  Sentence-level Speech Encoding without Language-parallel Data
标题:TransSentence:无语言并行数据的语言不可知的句子级语音编码的语音到语音翻译
链接:https://arxiv.org/abs/2401.12992
作者:Seung-Bin Kim,Sang-Hoon Lee,Seong-Whan Lee
备注:Accepted by ICASSP 2024
摘要:虽然语音到语音翻译领域已经取得了重大进展,但传统的模型仍然需要源语言和目标语言之间的语言并行语音数据进行训练。在本文中,我们介绍了TranSentence,一种新的语音到语音翻译没有语言并行语音数据。为了实现这一点,我们首先采用了一种语言不可知论的语音编码,捕捉语音的语义信息,无论语言。然后,我们训练我们的模型,以生成语音的基础上,从一个语言无关的语音级语音编码器,是预先训练与各种语言的编码嵌入。通过这种方法,尽管只在目标语言的单语数据上进行训练,但我们可以在推理阶段使用与语言无关的语音嵌入从源语言语音中生成目标语言语音。此外,我们将TranSentence扩展到多语言语音到语音翻译。实验结果表明,TranSentence模型优于其他模型。
摘要:Although there has been significant advancement in the field of speech-to-speech translation, conventional models still require language-parallel speech data between the source and target languages for training. In this paper, we introduce TranSentence, a novel speech-to-speech translation without language-parallel speech data. To achieve this, we first adopt a language-agnostic sentence-level speech encoding that captures the semantic information of speech, irrespective of language. We then train our model to generate speech based on the encoded embedding obtained from a language-agnostic sentence-level speech encoder that is pre-trained with various languages. With this method, despite training exclusively on the target language's monolingual data, we can generate target language speech in the inference stage using language-agnostic speech embedding from the source language speech. Furthermore, we extend TranSentence to multilingual speech-to-speech translation. The experimental results demonstrate that TranSentence is superior to other models.


【12】 TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition  in Conversation
标题:TelME:教师主导的多模式会话情感识别融合网络
链接:https://arxiv.org/abs/2401.12987
作者:Taeyang Yun,Hyunkuk Lim,Jeonghwan Lee,Min Song
备注:13 pages, 7 figures
摘要:会话中的情感识别(ERC)在使对话系统能够有效地响应用户请求方面起着至关重要的作用。会话中的情感可以通过来自各种模态(诸如音频、视觉和文本)的表示来识别。然而,由于非言语模态对情绪识别的贡献很小,多模态ERC一直被认为是一项具有挑战性的任务。在本文中,我们提出了教师领导的多模态融合网络ERC(TelME)。TelME采用跨模态知识蒸馏,将信息从充当教师的语言模型传递给非语言学生,从而优化弱模态的功效。然后,我们结合多模态功能使用移位融合的方法,其中学生网络支持教师。TelME在MELD中实现了最先进的性能,MELD是ERC的多说话者对话数据集。最后,我们通过额外的实验证明了我们的组件的有效性。
摘要:Emotion Recognition in Conversation (ERC) plays a crucial role in enabling dialogue systems to effectively respond to user requests. The emotions in a conversation can be identified by the representations from various modalities, such as audio, visual, and text. However, due to the weak contribution of non-verbal modalities to recognize emotions, multimodal ERC has always been considered a challenging task. In this paper, we propose Teacher-leading Multimodal fusion network for ERC (TelME). TelME incorporates cross-modal knowledge distillation to transfer information from a language model acting as the teacher to the non-verbal students, thereby optimizing the efficacy of the weak modalities. We then combine multimodal features using a shifting fusion approach in which student networks support the teacher. TelME achieves state-of-the-art performance in MELD, a multi-speaker conversation dataset for ERC. Finally, we demonstrate the effectiveness of our components through additional experiments.

机器翻译由腾讯交互翻译提供,仅供参考