【1】Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach标题:缺失情态的多通道情感分析:一种知识转移方法链接:https://arxiv.org/abs/2401.10747作者:Weide Liu,Huijing Zhan,Hao Chen,Fengmao Lv备注:5 pages摘要:多模态情感分析旨在识别个人通过视觉,语言和听觉线索表达的情感。然而,大多数现有的研究工作都假设所有模态在训练和测试过程中都是可用的,这使得他们的算法容易受到缺失模态场景的影响。在本文中,我们提出了一种新的知识转移网络,在不同的模态之间进行翻译,以重建丢失的音频模态。此外,我们开发了一个跨模态注意机制,以保留最大的信息重建和观察模态的情感预测。在三个公开的数据集上进行的广泛实验表明,与基线相比,该方法有了显着的改进,并实现了与具有完整多模态监督的先前方法相当的结果。摘要:Multimodal sentiment analysis aims to identify the emotions expressed by individuals through visual, language, and acoustic cues. However, most of the existing research efforts assume that all modalities are available during both training and testing, making their algorithms susceptible to the missing modality scenario. In this paper, we propose a novel knowledge-transfer network to translate between different modalities to reconstruct the missing audio modalities. Moreover, we develop a cross-modality attention mechanism to retain the maximal information of the reconstructed and observed modalities for sentiment prediction. Extensive experiments on three publicly available datasets demonstrate significant improvements over baselines and achieve comparable results to the previous methods with complete multi-modality supervision.
【2】 Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection标题:注意力融合:一种基于Transformer的多模式仇恨语音检测方法链接:https://arxiv.org/abs/2401.10653作者:Atanu Mandal,Gargi Roy,Amit Barman,Indranil Dutta,Sudip Kumar Naskar备注:Accepted in 20th International Conference on Natural Language Processing (ICON)摘要:随着最近社交媒体使用的激增和指数增长,审查社交媒体内容是否存在任何仇恨内容至关重要。自过去十年以来,研究人员一直在努力区分促进仇恨的内容和不促进仇恨的内容。传统上,主要的重点是分析文本内容。然而,最近的研究尝试也开始进入基于音频的内容的识别。然而,研究表明,仅仅依靠音频或基于文本的内容可能是无效的,因为最近的高潮表明,人们经常在他们的演讲和写作中使用讽刺。为了克服这些挑战,我们提出了一种方法来确定是否演讲促进仇恨或不利用音频和文本表示。我们的方法基于Transformer框架,该框架结合了音频和文本采样,并配有我们自己的层“Attentive Fusion”。我们的研究结果超过了以前的最先进的技术,在测试集上获得了令人印象深刻的宏观F1分数0.927。摘要:With the recent surge and exponential growth of social media usage, scrutinizing social media content for the presence of any hateful content is of utmost importance. Researchers have been diligently working since the past decade on distinguishing between content that promotes hatred and content that does not. Traditionally, the main focus has been on analyzing textual content. However, recent research attempts have also commenced into the identification of audio-based content. Nevertheless, studies have shown that relying solely on audio or text-based content may be ineffective, as recent upsurge indicates that individuals often employ sarcasm in their speech and writing. To overcome these challenges, we present an approach to identify whether a speech promotes hate or not utilizing both audio and textual representations. Our methodology is based on the Transformer framework that incorporates both audio and text sampling, accompanied by our very own layer called "Attentive Fusion". The results of our study surpassed previous state-of-the-art techniques, achieving an impressive macro F1 score of 0.927 on the Test Set.
【3】 AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks标题:AAT:适应各种声学识别任务的音频转换器链接:https://arxiv.org/abs/2401.10544作者:Yun Liang,Hai Lin,Shaojian Qiu,Yihang Zhang备注:Preprint version for ICASSP 2024, Korea摘要:最近,Transformers被引入到声学识别领域。它们使用监督学习和半监督学习等方法在大规模数据集上进行预训练,表现出强大的通用性-它可以轻松微调下游任务,并显示出更强大的性能。然而,目前使用的主要微调方法仍然是完全微调,这涉及在训练期间更新所有参数。这不仅会导致大量的内存使用和时间成本,而且还会损害模型的通用性。其他微调方法要么难以解决这个问题,要么无法实现匹配的性能。为此,本文对现有的调优方法进行了综合分析,提出了一种基于适配器调优的高效调优方法,即AAT。其核心思想是冻结音频Transformer模型并插入额外的可学习适配器,有效地获取下游任务知识,而不损害模型的原始通用性。大量的实验表明,我们的方法实现了性能相当,甚至优于完全微调,同时优化只有7.118%的参数。它也证明了优于其他微调方法。摘要:Recently, Transformers have been introduced into the field of acoustics recognition. They are pre-trained on large-scale datasets using methods such as supervised learning and semi-supervised learning, demonstrating robust generality--It fine-tunes easily to downstream tasks and shows more robust performance. However, the predominant fine-tuning method currently used is still full fine-tuning, which involves updating all parameters during training. This not only incurs significant memory usage and time costs but also compromises the model's generality. Other fine-tuning methods either struggle to address this issue or fail to achieve matching performance. Therefore, we conducted a comprehensive analysis of existing fine-tuning methods and proposed an efficient fine-tuning approach based on Adapter tuning, namely AAT. The core idea is to freeze the audio Transformer model and insert extra learnable Adapters, efficiently acquiring downstream task knowledge without compromising the model's original generality. Extensive experiments have shown that our method achieves performance comparable to or even superior to full fine-tuning while optimizing only 7.118% of the parameters. It also demonstrates superiority over other fine-tuning methods. 【4】 Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech标题:用于无词典文本到语音的数据驱动的字素到音素表示链接:https://arxiv.org/abs/2401.10465作者:Abhinav Garg,Jiyeon Kim,Sushil Khyalia,Chanwoo Kim,Dhananjaya Gowda备注:Accepted at ICASSP 2024摘要:在任何现代高质量的文本到语音(TTS)系统中,字素到音素(G2P)都是必不可少的第一步。目前的大多数G2P系统依赖于专家精心制作的词典。这造成了双重问题。首先,词典是使用固定的音素集生成的,通常是ARPABET或IPA,这可能不是表示所有语言音素的最佳方式。其次,制作这样一本专业词典所需的工时非常高。在本文中,我们通过使用自监督学习的最新进展来消除这两个问题,以获得数据驱动的音素表示,而不是固定的表示。我们将我们的无词典方法与利用精心制作的词典的强大基线进行比较。此外,我们表明,我们的数据驱动的无词典方法在平均意见得分(MOS)方面与传统的基于规则或基于词典的神经G2P一样好,甚至略好于传统的基于规则或基于词典的神经G2P,而不使用先前的语言词典或音素集,即没有语言专业知识。摘要:Grapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, usually, ARPABET or IPA, which might not be the most optimal way to represent phonemes for all languages. Secondly, the man-hours required to produce such an expert lexicon are very high. In this paper, we eliminate both of these issues by using recent advances in self-supervised learning to obtain data-driven phoneme representations instead of fixed representations. We compare our lexicon-free approach against strong baselines that utilize a well-crafted lexicon. Furthermore, we show that our data-driven lexicon-free method performs as good or even marginally better than the conventional rule-based or lexicon-based neural G2Ps in terms of Mean Opinion Score (MOS) while using no prior language lexicon or phoneme set, i.e. no linguistic expertise. 【5】 Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis标题:用于高质量语音合成的超轻量级神经差分DSP声码器链接:https://arxiv.org/abs/2401.10460作者:Prabhav Agrawal,Thilo Koehler,Zhiping Xiu,Prashant Serai,Qing He备注:Accepted for ICASSP 2024摘要:神经声码器对原始音频波形进行建模并合成高质量的音频,但即使是像MB-MelGAN和LPCNet这样的高效声码器,也无法在智能眼镜等低端设备上实时运行。基于纯数字信号处理(DSP)的声码器可以经由轻量级快速傅立叶变换(FFT)来实现,并且因此比任何神经声码器快一个数量级。DSP声码器通常由于消耗声道的近似表示的过平滑声学模型预测而获得较低的音频质量。在本文中,我们提出了一个超轻量的差分DSP(DDSP)声码器,使用一个联合优化的声学模型与DSP声码器,学习没有提取的声道频谱特征。该模型实现的音频质量与神经声码器相当,平均MOS为4.36,同时作为DSP声码器是有效的。我们的C++实现,没有任何硬件特定的优化,是在15 MFLOPS,超过MB-MelGAN的340倍的FLOPS,并实现了一个声码器的RTF为0.003和整体RTF为0.044,同时运行在一个2GHz的英特尔至强CPU上的单线程。摘要:Neural vocoders model the raw audio waveform and synthesize high-quality audio, but even the highly efficient ones, like MB-MelGAN and LPCNet, fail to run real-time on a low-end device like a smartglass. A pure digital signal processing (DSP) based vocoder can be implemented via lightweight fast Fourier transforms (FFT), and therefore, is a magnitude faster than any neural vocoder. A DSP vocoder often gets a lower audio quality due to consuming over-smoothed acoustic model predictions of approximate representations for the vocal tract. In this paper, we propose an ultra-lightweight differential DSP (DDSP) vocoder that uses a jointly optimized acoustic model with a DSP vocoder, and learns without an extracted spectral feature for the vocal tract. The model achieves audio quality comparable to neural vocoders with a high average MOS of 4.36 while being efficient as a DSP vocoder. Our C++ implementation, without any hardware-specific optimization, is at 15 MFLOPS, surpasses MB-MelGAN by 340 times in terms of FLOPS, and achieves a vocoder-only RTF of 0.003 and overall RTF of 0.044 while running single-threaded on a 2GHz Intel Xeon CPU.
【6】 Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition标题:语音识别语言建模低阶自适应训练策略及模型稳健性研究链接:https://arxiv.org/abs/2401.10447作者:Yu Yu,Chao-Han Huck Yang,Tuan Dinh,Sungho Ryu,Jari Kolehmainen,Roger Ren,Denis Filimonov,Prashanth G. Shivakumar,Ankur Gandhe,Ariya Rastow,Jia Xu,Ivan Bulyko,Andreas Stolcke摘要:低秩自适应(LoRA)与冻结预训练语言模型(PLM)的使用已经成为内存受限硬件的主流,资源高效的建模方法。在这项研究中,我们首先探索如何通过引入各种LoRA训练策略来提高模型性能,在公共Librispeech数据集上实现相对单词错误率降低3.50%,在消息传递领域的内部数据集上降低3.67%。为了进一步表征基于LoRA的第二遍语音识别模型的稳定性,我们研究了对输入扰动的鲁棒性。这些扰动植根于同音字替换和一种称为N-best Perturbation-based Rescoring Robustness(NPRR)的新度量,两者都旨在测量重新评分模型性能的相对退化。我们的实验结果表明,虽然先进的变种LoRA,如动态排名分配LoRA,导致性能下降,在$1$-最好的扰动,他们减轻了退化,在$N$-最好的扰动。这一发现与完全调优的模型和普通LoRA调优基线进行了比较,表明在使用基于LoRA的自适应以节省计算成本和强大的语言建模时,需要进行全面的选择。摘要:The use of low-rank adaptation (LoRA) with frozen pretrained language models (PLMs) has become increasing popular as a mainstream, resource-efficient modeling approach for memory-constrained hardware. In this study, we first explore how to enhance model performance by introducing various LoRA training strategies, achieving relative word error rate reductions of 3.50\% on the public Librispeech dataset and of 3.67\% on an internal dataset in the messaging domain. To further characterize the stability of LoRA-based second-pass speech recognition models, we examine robustness against input perturbations. These perturbations are rooted in homophone replacements and a novel metric called N-best Perturbation-based Rescoring Robustness (NPRR), both designed to measure the relative degradation in the performance of rescoring models. Our experimental results indicate that while advanced variants of LoRA, such as dynamic rank-allocated LoRA, lead to performance degradation in $1$-best perturbation, they alleviate the degradation in $N$-best perturbation. This finding is in comparison to fully-tuned models and vanilla LoRA tuning baselines, suggesting that a comprehensive selection is needed when using LoRA-based adaptation for compute-cost savings and robust language modeling. 【7】 Large Language Models are Efficient Learners of Noise-Robust Speech Recognition标题:大语言模型是抗噪语音识别的有效学习器链接:https://arxiv.org/abs/2401.10446作者:Yuchen Hu,Chen Chen,Chao-Han Huck Yang,Ruizhe Li,Chao Zhang,Pin-Yu Chen,EnSiong Chng备注:Accepted to ICLR 2024, Spotlight top 5%, 24 pages. This work will be open sourced at: this https URL under MIT license摘要:大型语言模型(LLM)的最新进展促进了自动语音识别(ASR)的生成纠错(GER),其利用LLM丰富的语言知识和强大的推理能力来改善识别结果。最新的工作提出了一个GER基准与HyPorcraft数据集学习映射从ASR N-best假设地面真理转录通过有效的LLM微调,这表明了很大的有效性,但缺乏特异性的噪声鲁棒ASR。在这项工作中,我们将基准扩展到噪声条件,并研究我们是否可以教LLM为GER执行去噪,就像鲁棒ASR所做的那样,其中一个解决方案是将噪声信息作为条件引入LLM。然而,由于跨模态间隙,直接结合来自音频编码器的噪声嵌入可能会损害LLM调谐。为此,本文提出从N-best列表中提取一个语言空间噪声嵌入来表征源语音的噪声状况,从而促进了GER的去噪过程;此外,为了增强其对音频噪声的表征能力,本文设计了一种基于互信息估计的知识提取方法,将音频嵌入中的真实噪声信息提取到我们的语言嵌入中。在各种最新的LLM上的实验表明,我们的方法在训练数据有限的情况下,在字错误率方面获得了高达53.9%的校正改进,实现了新的突破。分析表明,我们的语言空间噪声嵌入可以很好地代表源语音的噪声条件下,现成的LLM显示出强大的语言空间去噪能力。摘要:Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results. The latest work proposes a GER benchmark with HyPoradise dataset to learn the mapping from ASR N-best hypotheses to ground-truth transcription by efficient LLM finetuning, which shows great effectiveness but lacks specificity on noise-robust ASR. In this work, we extend the benchmark to noisy conditions and investigate if we can teach LLMs to perform denoising for GER just like what robust ASR do}, where one solution is introducing noise information as a conditioner into LLM. However, directly incorporating noise embeddings from audio encoder could harm the LLM tuning due to cross-modality gap. To this end, we propose to extract a language-space noise embedding from the N-best list to represent the noise conditions of source speech, which can promote the denoising process in GER. Furthermore, in order to enhance its representation ability of audio noise, we design a knowledge distillation (KD) approach via mutual information estimation to distill the real noise information in audio embeddings to our language embedding. Experiments on various latest LLMs demonstrate our approach achieves a new breakthrough with up to 53.9% correction improvement in terms of word error rate while with limited training data. Analysis shows that our language-space noise embedding can well represent the noise conditions of source speech, under which off-the-shelf LLMs show strong ability of language-space denoising. 【8】 DanceMeld: Unraveling Dance Phrases with Hierarchical Latent Codes for Music-to-Dance Synthesis标题:DanceMeld:用于音乐到舞蹈合成的具有层次化潜在代码的舞蹈短语拆解链接:https://arxiv.org/abs/2401.10242作者:Xin Gao,Li Hu,Peng Zhang,Bang Zhang,Liefeng Bo备注:10 pages, 8 figures摘要:在3D数字人类应用领域,音乐到舞蹈提出了一项具有挑战性的任务。鉴于音乐和舞蹈之间的一对多关系,以前的方法在方法上受到限制,仅依赖于基于音乐节奏的匹配和生成相应的舞蹈动作。在专业领域的舞蹈,一个舞蹈短语包括几个舞蹈姿势和舞蹈动作。舞蹈姿势是由一系列基本的有意义的身体姿势组成的,而舞蹈动作则可以反映舞蹈的节奏、旋律、风格等动态变化。从这些概念中获得灵感,我们引入了一个创新的舞蹈生成管道DanceMeld,它包括两个阶段,即,舞蹈分离阶段和舞蹈生成阶段。在解耦阶段,分层VQ-VAE用于在不同的特征空间级别上解开舞蹈姿势和舞蹈动作,其中底部代码表示舞蹈姿势,并且顶部代码表示舞蹈动作。在生成阶段,我们利用一个扩散模型作为一个先验模型的分布和生成潜在的代码的音乐功能的条件。我们已经通过实验证明了顶层代码和底层代码的代表性能力,使舞蹈姿势和舞蹈动作的明确解耦表达。这种解缠不仅提供了对动作细节、风格和节奏的控制,而且还促进了诸如舞蹈风格转移和舞蹈单元编辑之类的应用。在AIST++数据集上进行了定性和定量实验,证明了该方法的优越性。摘要:In the realm of 3D digital human applications, music-to-dance presents a challenging task. Given the one-to-many relationship between music and dance, previous methods have been limited in their approach, relying solely on matching and generating corresponding dance movements based on music rhythm. In the professional field of choreography, a dance phrase consists of several dance poses and dance movements. Dance poses composed of a series of basic meaningful body postures, while dance movements can reflect dynamic changes such as the rhythm, melody, and style of dance. Taking inspiration from these concepts, we introduce an innovative dance generation pipeline called DanceMeld, which comprising two stages, i.e., the dance decouple stage and the dance generation stage. In the decouple stage, a hierarchical VQ-VAE is used to disentangle dance poses and dance movements in different feature space levels, where the bottom code represents dance poses, and the top code represents dance movements. In the generation stage, we utilize a diffusion model as a prior to model the distribution and generate latent codes conditioned on music features. We have experimentally demonstrated the representational capabilities of top code and bottom code, enabling the explicit decoupling expression of dance poses and dance movements. This disentanglement not only provides control over motion details, styles, and rhythm but also facilitates applications such as dance style transfer and dance unit editing. Our approach has undergone qualitative and quantitative experiments on the AIST++ dataset, demonstrating its superiority over other methods. 【9】 Multilingual acoustic word embeddings for zero-resource languages标题:面向零资源语言的多语种声学单词嵌入链接:https://arxiv.org/abs/2401.10543作者:Christiaan Jacobs,Herman Kamper摘要:这项研究解决了为缺乏标记数据的零资源语言开发语音应用程序的挑战。它特别使用声学词嵌入(AWE)-可变持续时间语音段的固定维度表示-采用多语言传输,其中来自几种资源丰富的语言的标记数据用于相关。该研究引入了一种新的神经网络,它在零资源语言上的性能优于现有的AWE模型。它探讨了选择资源丰富的语言的影响。AWE应用于斯瓦希里语广播中的仇恨言论检测的关键字定位系统,在现实世界中的场景中表现出鲁棒性。此外,新的语义AWE模型改进了语义查询的示例搜索。摘要:This research addresses the challenge of developing speech applications for zero-resource languages that lack labelled data. It specifically uses acoustic word embedding (AWE) -- fixed-dimensional representations of variable-duration speech segments -- employing multilingual transfer, where labelled data from several well-resourced languages are used for pertaining. The study introduces a new neural network that outperforms existing AWE models on zero-resource languages. It explores the impact of the choice of well-resourced languages. AWEs are applied to a keyword-spotting system for hate speech detection in Swahili radio broadcasts, demonstrating robustness in real-world scenarios. Additionally, novel semantic AWE models improve semantic query-by-example search.
【10】 A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement标题:一种实时语音增强的跨谱域两阶段框架链接:https://arxiv.org/abs/2401.10494作者:Yuewei Zhang,Huanbin Zou,Jie Zhu备注:Accepted by ICASSP 2024摘要:两级流水线由于其优于传统的单级方法而在语音增强任务中受到欢迎。目前的两阶段方法通常在第一阶段增强幅度谱,并在第二阶段进一步修改复谱以抑制残余噪声并恢复语音相位。上述整个过程在短时傅立叶变换(STFT)谱域中执行。在本文中,我们重新实现上述第二子过程中的短时离散余弦变换(STDCT)频谱域。原因是我们发现STDCT比STFT具有更好的噪声抑制能力。此外,STDCT的隐式相位确保了更简单和更有效的相位恢复,这在基于STFT的方法中是具有挑战性和计算昂贵的。因此,我们提出了一种新的两阶段框架称为STFT-STDCT频谱融合网络(FDFNet)的语音增强在互谱域。实验结果表明,所提出的FDFNet优于以前的两阶段方法,也表现出优越的性能相比,其他先进的系统。摘要:Two-stage pipeline is popular in speech enhancement tasks due to its superiority over traditional single-stage methods. The current two-stage approaches usually enhance the magnitude spectrum in the first stage, and further modify the complex spectrum to suppress the residual noise and recover the speech phase in the second stage. The above whole process is performed in the short-time Fourier transform (STFT) spectrum domain. In this paper, we re-implement the above second sub-process in the short-time discrete cosine transform (STDCT) spectrum domain. The reason is that we have found STDCT performs greater noise suppression capability than STFT. Additionally, the implicit phase of STDCT ensures simpler and more efficient phase recovery, which is challenging and computationally expensive in the STFT-based methods. Therefore, we propose a novel two-stage framework called the STFT-STDCT spectrum fusion network (FDFNet) for speech enhancement in cross-spectrum domain. Experimental results demonstrate that the proposed FDFNet outperforms the previous two-stage methods and also exhibits superior performance compared to other advanced systems. 【11】 3D Room Geometry Inference from Multichannel Room Impulse Response using Deep Neural Network标题:基于深度神经网络的多通道房间脉冲响应三维房间几何推断链接:https://arxiv.org/abs/2401.10453作者:Inmo Yeon,Jung-Woo Choi备注:None摘要:房间几何推断(RGI)的目的是从测量的房间脉冲响应(RIR)估计房间的形状,并已收到了大量的关注,其在环境感知的音频渲染和虚拟声学表示的真实场地的重要性。在雷达接收机中,利用到达时间差(TDoA)或到达时间(ToA)信息的估计模型已被提出。然而,估计模型应该能够处理更一般的特征和反射之间的复杂关系,以应对各种房间形状和不确定性,例如未知的墙壁数量。在这项研究中,我们提出了一种深度神经网络,可以在不预先假设墙壁形状或数量的情况下估计各种房间形状。该模型由三个子网络组成:特征提取器,参数估计和评估网络,分别从RIR中提取关键特征,估计参数,并评估估计参数的置信度。该网络由使用单个源和球形麦克风阵列在不同形状的房间中模拟的大约40,000个RIR进行训练,并针对不可见形状和尺寸的房间进行测试。该算法在寻找墙壁的真实数量方面达到了几乎完美的精度,并且在房间形状方面显示出可以忽略不计的误差。摘要:Room geometry inference (RGI) aims at estimating room shapes from measured room impulse responses (RIRs) and has received lots of attention for its importance in environment-aware audio rendering and virtual acoustic representation of a real venue. A lot of estimation models utilizing time difference of arrival (TDoA) or time of arrival (ToA) information in RIRs have been proposed. However, an estimation model should be able to handle more general features and complex relations between reflections to cope with various room shapes and uncertainties such as the unknown number of walls. In this study, we propose a deep neural network that can estimate various room shapes without prior assumptions on the shape or number of walls. The proposed model consists of three sub-networks: a feature extractor, parameter estimation, and evaluation networks, which extract key features from RIRs, estimate parameters, and evaluate the confidence of estimated parameters, respectively. The network is trained by about 40,000 RIRs simulated in rooms of different shapes using a single source and spherical microphone array and tested for rooms of unseen shapes and dimensions. The proposed algorithm achieves almost perfect accuracy in finding the true number of walls and shows negligible errors in room shapes.
【12】 Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search标题:基于注意力偏向短语增强波束搜索的语境化自动语音识别链接:https://arxiv.org/abs/2401.10449作者:Yui Sudo,Muhammad Shakeel,Yosuke Fukumoto,Yifan Peng,Shinji Watanabe备注:accepted by ICASSP20224摘要:端到端(E2 E)自动语音识别(ASR)方法表现出显着的性能。然而,由于这样的方法的性能内在地与训练数据中存在的上下文相关联,因此E2 E-ASR方法对于看不见的用户上下文(例如,技术术语、人名和播放列表)。因此,E2 E-ASR方法必须容易地由用户或开发人员上下文化。本文提出了一种基于注意力的上下文偏置方法,该方法可以使用可编辑短语列表(称为偏置列表)进行定制。该方法可以有效地训练相结合的偏见短语索引损失和特殊的令牌来检测输入语音数据中的偏见短语。此外,为了进一步提高推理过程中的上下文化性能,我们提出了一种基于偏置短语索引概率的偏置短语提升(BPB)波束搜索算法。实验结果表明,该方法在Librispeech-960(英语)和我们的内部(日语)数据集上分别提高了偏置列表中目标短语的单词错误率和字符错误率。摘要:End-to-end (E2E) automatic speech recognition (ASR) methods exhibit remarkable performance. However, since the performance of such methods is intrinsically linked to the context present in the training data, E2E-ASR methods do not perform as desired for unseen user contexts (e.g., technical terms, personal names, and playlists). Thus, E2E-ASR methods must be easily contextualized by the user or developer. This paper proposes an attention-based contextual biasing method that can be customized using an editable phrase list (referred to as a bias list). The proposed method can be trained effectively by combining a bias phrase index loss and special tokens to detect the bias phrases in the input speech data. In addition, to improve the contextualization performance during inference further, we propose a bias phrase boosted (BPB) beam search algorithm based on the bias phrase index probability. Experimental results demonstrate that the proposed method consistently improves the word error rate and the character error rate of the target phrases in the bias list on both the Librispeech-960 (English) and our in-house (Japanese) dataset, respectively.
【13】 AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition标题:Agadir:走向阵列几何不可知的定向语音识别链接:https://arxiv.org/abs/2401.10411作者:Ju Lin,Niko Moritz,Yiteng Huang,Ruiming Xie,Ming Sun,Christian Fuegen,Frank Seide备注:Accepted to ICASSP 2024摘要:像智能眼镜这样的可穿戴设备正在接近计算能力,可以无缝地为现场对话生成实时隐藏字幕。我们建立在我们最近推出的具有麦克风阵列的智能眼镜的定向自动语音识别(ASR)的基础上,该智能眼镜将多通道ASR与序列化输出训练相融合,用于佩戴者/对话伙伴消歧以及抑制来自非目标方向和噪声的串扰语音。 当ASR工作是更广泛的系统开发过程的一部分时,随着系统开发的进展,可能会面临麦克风几何形状的变化。 本文旨在使多通道ASR不敏感的麦克风阵列几何形状的有限变化。我们表明,在多个相似几何形状上训练的模型在很大程度上是不可知的,并且可以很好地推广到新的几何形状,只要它们不是太不同。此外,以这种方式训练模型可以将所看到的几何形状的准确性相对提高15%至28%。最后,我们改进了波束形成一个新的非线性约束最小方差准则。摘要:Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise. When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses. This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28\% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion. 【14】 Detecting Post-Stroke Aphasia Via Brain Responses to Speech in a Deep Learning Framework标题:在深度学习框架中通过大脑对语音的反应检测卒中后失语症链接:https://arxiv.org/abs/2401.10291作者:Pieter De Clercq,Corentin Puffay,Jill Kries,Hugo Van Hamme,Maaike Vandermosten,Tom Francart,Jonas Vanthornhout备注:Shared first authors: De Clercq & Puffay摘要:失语症是一种主要由中风引起的语言障碍,传统上使用行为语言测试进行诊断。然而,这些测试是耗时的,需要由训练有素的临床医生手动解释,遭受低生态效度,诊断可能会受到失语症中存在的共病运动和认知问题的影响。在这项研究中,我们介绍了一种用于失语症语音处理障碍的自动筛选工具,该工具依赖于深度学习框架内对语音的时间锁定大脑反应,称为神经跟踪。我们使用在大样本健康参与者上训练的卷积神经网络对故事的声学,分割和语言语音表示进行了脑电图(EEG)响应建模,作为语音的完整神经跟踪模型。随后,我们在一个独立的样本上评估了我们的模型,该样本包括26名失语症(IWA)患者和22名健康对照。我们的研究结果表明,减少跟踪的所有语音表示在IWA。利用神经跟踪措施作为输入的支持向量机分类器,我们证明了在失语症检测在个人水平(85.42%)在一个时间效率的方式(需要9分钟的EEG数据)的高准确性。鉴于其高鲁棒性,时间效率和不可见数据的可推广性,我们的方法在临床应用中具有重要的前景。摘要:Aphasia, a language disorder primarily caused by a stroke, is traditionally diagnosed using behavioral language tests. However, these tests are time-consuming, require manual interpretation by trained clinicians, suffer from low ecological validity, and diagnosis can be biased by comorbid motor and cognitive problems present in aphasia. In this study, we introduce an automated screening tool for speech processing impairments in aphasia that relies on time-locked brain responses to speech, known as neural tracking, within a deep learning framework. We modeled electroencephalography (EEG) responses to acoustic, segmentation, and linguistic speech representations of a story using convolutional neural networks trained on a large sample of healthy participants, serving as a model for intact neural tracking of speech. Subsequently, we evaluated our models on an independent sample comprising 26 individuals with aphasia (IWA) and 22 healthy controls. Our results reveal decreased tracking of all speech representations in IWA. Utilizing a support vector machine classifier with neural tracking measures as input, we demonstrate high accuracy in aphasia detection at the individual level (85.42\%) in a time-efficient manner (requiring 9 minutes of EEG data). Given its high robustness, time efficiency, and generalizability to unseen data, our approach holds significant promise for clinical applications.
eess.AS音频处理【1】 Multilingual acoustic word embeddings for zero-resource languages标题:面向零资源语言的多语种声学单词嵌入链接:https://arxiv.org/abs/2401.10543作者:Christiaan Jacobs,Herman Kamper摘要:None摘要:This research addresses the challenge of developing speech applications for zero-resource languages that lack labelled data. It specifically uses acoustic word embedding (AWE) -- fixed-dimensional representations of variable-duration speech segments -- employing multilingual transfer, where labelled data from several well-resourced languages are used for pertaining. The study introduces a new neural network that outperforms existing AWE models on zero-resource languages. It explores the impact of the choice of well-resourced languages. AWEs are applied to a keyword-spotting system for hate speech detection in Swahili radio broadcasts, demonstrating robustness in real-world scenarios. Additionally, novel semantic AWE models improve semantic query-by-example search. 【2】 A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement标题:一种实时语音增强的跨谱域两阶段框架链接:https://arxiv.org/abs/2401.10494作者:Yuewei Zhang,Huanbin Zou,Jie Zhu备注:Accepted by ICASSP 2024摘要:两级流水线由于其优于传统的单级方法而在语音增强任务中受到欢迎。目前的两阶段方法通常在第一阶段增强幅度谱,并在第二阶段进一步修改复谱以抑制残余噪声并恢复语音相位。上述整个过程在短时傅立叶变换(STFT)谱域中执行。在本文中,我们重新实现上述第二子过程中的短时离散余弦变换(STDCT)频谱域。原因是我们发现STDCT比STFT具有更好的噪声抑制能力。此外,STDCT的隐式相位确保了更简单和更有效的相位恢复,这在基于STFT的方法中是具有挑战性和计算昂贵的。因此,我们提出了一种新的两阶段框架称为STFT-STDCT频谱融合网络(FDFNet)的语音增强在互谱域。实验结果表明,所提出的FDFNet优于以前的两阶段方法,也表现出优越的性能相比,其他先进的系统。摘要:Two-stage pipeline is popular in speech enhancement tasks due to its superiority over traditional single-stage methods. The current two-stage approaches usually enhance the magnitude spectrum in the first stage, and further modify the complex spectrum to suppress the residual noise and recover the speech phase in the second stage. The above whole process is performed in the short-time Fourier transform (STFT) spectrum domain. In this paper, we re-implement the above second sub-process in the short-time discrete cosine transform (STDCT) spectrum domain. The reason is that we have found STDCT performs greater noise suppression capability than STFT. Additionally, the implicit phase of STDCT ensures simpler and more efficient phase recovery, which is challenging and computationally expensive in the STFT-based methods. Therefore, we propose a novel two-stage framework called the STFT-STDCT spectrum fusion network (FDFNet) for speech enhancement in cross-spectrum domain. Experimental results demonstrate that the proposed FDFNet outperforms the previous two-stage methods and also exhibits superior performance compared to other advanced systems.
【3】 3D Room Geometry Inference from Multichannel Room Impulse Response using Deep Neural Network标题:基于深度神经网络的多通道房间脉冲响应三维房间几何推断链接:https://arxiv.org/abs/2401.10453作者:Inmo Yeon,Jung-Woo Choi备注:None摘要:房间几何推断(RGI)的目的是从测量的房间脉冲响应(RIR)估计房间的形状,并已收到了大量的关注,其在环境感知的音频渲染和虚拟声学表示的真实场地的重要性。在雷达接收机中,利用到达时间差(TDoA)或到达时间(ToA)信息的估计模型已被提出。然而,估计模型应该能够处理更一般的特征和反射之间的复杂关系,以应对各种房间形状和不确定性,例如未知的墙壁数量。在这项研究中,我们提出了一种深度神经网络,可以在不预先假设墙壁形状或数量的情况下估计各种房间形状。该模型由三个子网络组成:特征提取器,参数估计和评估网络,分别从RIR中提取关键特征,估计参数,并评估估计参数的置信度。该网络由使用单个源和球形麦克风阵列在不同形状的房间中模拟的大约40,000个RIR进行训练,并针对不可见形状和尺寸的房间进行测试。该算法在寻找墙壁的真实数量方面达到了几乎完美的精度,并且在房间形状方面显示出可以忽略不计的误差。摘要:Room geometry inference (RGI) aims at estimating room shapes from measured room impulse responses (RIRs) and has received lots of attention for its importance in environment-aware audio rendering and virtual acoustic representation of a real venue. A lot of estimation models utilizing time difference of arrival (TDoA) or time of arrival (ToA) information in RIRs have been proposed. However, an estimation model should be able to handle more general features and complex relations between reflections to cope with various room shapes and uncertainties such as the unknown number of walls. In this study, we propose a deep neural network that can estimate various room shapes without prior assumptions on the shape or number of walls. The proposed model consists of three sub-networks: a feature extractor, parameter estimation, and evaluation networks, which extract key features from RIRs, estimate parameters, and evaluate the confidence of estimated parameters, respectively. The network is trained by about 40,000 RIRs simulated in rooms of different shapes using a single source and spherical microphone array and tested for rooms of unseen shapes and dimensions. The proposed algorithm achieves almost perfect accuracy in finding the true number of walls and shows negligible errors in room shapes.
【4】 Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search标题:基于注意力偏向短语增强波束搜索的语境化自动语音识别链接:https://arxiv.org/abs/2401.10449作者:Yui Sudo,Muhammad Shakeel,Yosuke Fukumoto,Yifan Peng,Shinji Watanabe备注:accepted by ICASSP20224摘要:端到端(E2 E)自动语音识别(ASR)方法表现出显着的性能。然而,由于这样的方法的性能内在地与训练数据中存在的上下文相关联,因此E2 E-ASR方法对于看不见的用户上下文(例如,技术术语、人名和播放列表)。因此,E2 E-ASR方法必须容易地由用户或开发人员上下文化。本文提出了一种基于注意力的上下文偏置方法,该方法可以使用可编辑短语列表(称为偏置列表)进行定制。该方法可以有效地训练相结合的偏见短语索引损失和特殊的令牌来检测输入语音数据中的偏见短语。此外,为了进一步提高推理过程中的上下文化性能,我们提出了一种基于偏置短语索引概率的偏置短语提升(BPB)波束搜索算法。实验结果表明,该方法在Librispeech-960(英语)和我们的内部(日语)数据集上分别提高了偏置列表中目标短语的单词错误率和字符错误率。摘要:End-to-end (E2E) automatic speech recognition (ASR) methods exhibit remarkable performance. However, since the performance of such methods is intrinsically linked to the context present in the training data, E2E-ASR methods do not perform as desired for unseen user contexts (e.g., technical terms, personal names, and playlists). Thus, E2E-ASR methods must be easily contextualized by the user or developer. This paper proposes an attention-based contextual biasing method that can be customized using an editable phrase list (referred to as a bias list). The proposed method can be trained effectively by combining a bias phrase index loss and special tokens to detect the bias phrases in the input speech data. In addition, to improve the contextualization performance during inference further, we propose a bias phrase boosted (BPB) beam search algorithm based on the bias phrase index probability. Experimental results demonstrate that the proposed method consistently improves the word error rate and the character error rate of the target phrases in the bias list on both the Librispeech-960 (English) and our in-house (Japanese) dataset, respectively. 【5】 AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition标题:Agadir:走向阵列几何不可知的定向语音识别链接:https://arxiv.org/abs/2401.10411作者:Ju Lin,Niko Moritz,Yiteng Huang,Ruiming Xie,Ming Sun,Christian Fuegen,Frank Seide备注:Accepted to ICASSP 2024摘要:像智能眼镜这样的可穿戴设备正在接近计算能力,可以无缝地为现场对话生成实时隐藏字幕。我们建立在我们最近推出的具有麦克风阵列的智能眼镜的定向自动语音识别(ASR)的基础上,该智能眼镜将多通道ASR与序列化输出训练相融合,用于佩戴者/对话伙伴消歧以及抑制来自非目标方向和噪声的串扰语音。 当ASR工作是更广泛的系统开发过程的一部分时,随着系统开发的进展,可能会面临麦克风几何形状的变化。 本文旨在使多通道ASR不敏感的麦克风阵列几何形状的有限变化。我们表明,在多个相似几何形状上训练的模型在很大程度上是不可知的,并且可以很好地推广到新的几何形状,只要它们不是太不同。此外,以这种方式训练模型可以将所看到的几何形状的准确性相对提高15%至28%。最后,我们改进了波束形成一个新的非线性约束最小方差准则。摘要:Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise. When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses. This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28\% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion.
【6】 Detecting Post-Stroke Aphasia Via Brain Responses to Speech in a Deep Learning Framework标题:深度学习框架中通过大脑对语音的反应检测卒中后失语链接:https://arxiv.org/abs/2401.10291作者:Pieter De Clercq,Corentin Puffay,Jill Kries,Hugo Van Hamme,Maaike Vandermosten,Tom Francart,Jonas Vanthornhout备注:Shared first authors: De Clercq & Puffay摘要:失语症是一种主要由中风引起的语言障碍,传统上使用行为语言测试进行诊断。然而,这些测试是耗时的,需要由训练有素的临床医生手动解释,遭受低生态效度,诊断可能会受到失语症中存在的共病运动和认知问题的影响。在这项研究中,我们介绍了一种用于失语症语音处理障碍的自动筛选工具,该工具依赖于深度学习框架内对语音的时间锁定大脑反应,称为神经跟踪。我们使用在大样本健康参与者上训练的卷积神经网络对故事的声学,分割和语言语音表示进行了脑电图(EEG)响应建模,作为语音的完整神经跟踪模型。随后,我们在一个独立的样本上评估了我们的模型,该样本包括26名失语症(IWA)患者和22名健康对照。我们的研究结果表明,减少跟踪的所有语音表示在IWA。利用神经跟踪措施作为输入的支持向量机分类器,我们证明了在失语症检测在个人水平(85.42%)在一个时间效率的方式(需要9分钟的EEG数据)的高准确性。鉴于其高鲁棒性,时间效率和不可见数据的可推广性,我们的方法在临床应用中具有重要的前景。摘要:Aphasia, a language disorder primarily caused by a stroke, is traditionally diagnosed using behavioral language tests. However, these tests are time-consuming, require manual interpretation by trained clinicians, suffer from low ecological validity, and diagnosis can be biased by comorbid motor and cognitive problems present in aphasia. In this study, we introduce an automated screening tool for speech processing impairments in aphasia that relies on time-locked brain responses to speech, known as neural tracking, within a deep learning framework. We modeled electroencephalography (EEG) responses to acoustic, segmentation, and linguistic speech representations of a story using convolutional neural networks trained on a large sample of healthy participants, serving as a model for intact neural tracking of speech. Subsequently, we evaluated our models on an independent sample comprising 26 individuals with aphasia (IWA) and 22 healthy controls. Our results reveal decreased tracking of all speech representations in IWA. Utilizing a support vector machine classifier with neural tracking measures as input, we demonstrate high accuracy in aphasia detection at the individual level (85.42\%) in a time-efficient manner (requiring 9 minutes of EEG data). Given its high robustness, time efficiency, and generalizability to unseen data, our approach holds significant promise for clinical applications.
【7】 Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach标题:缺失情态的多通道情感分析:一种知识转移方法链接:https://arxiv.org/abs/2401.10747作者:Weide Liu,Huijing Zhan,Hao Chen,Fengmao Lv备注:5 pages摘要:多模态情感分析旨在识别个人通过视觉,语言和听觉线索表达的情感。然而,大多数现有的研究工作都假设所有模态在训练和测试过程中都是可用的,这使得他们的算法容易受到缺失模态场景的影响。在本文中,我们提出了一种新的知识转移网络,在不同的模态之间进行翻译,以重建丢失的音频模态。此外,我们开发了一个跨模态注意机制,以保留最大的信息重建和观察模态的情感预测。在三个公开的数据集上进行的广泛实验表明,与基线相比,该方法有了显着的改进,并实现了与具有完整多模态监督的先前方法相当的结果。摘要:Multimodal sentiment analysis aims to identify the emotions expressed by individuals through visual, language, and acoustic cues. However, most of the existing research efforts assume that all modalities are available during both training and testing, making their algorithms susceptible to the missing modality scenario. In this paper, we propose a novel knowledge-transfer network to translate between different modalities to reconstruct the missing audio modalities. Moreover, we develop a cross-modality attention mechanism to retain the maximal information of the reconstructed and observed modalities for sentiment prediction. Extensive experiments on three publicly available datasets demonstrate significant improvements over baselines and achieve comparable results to the previous methods with complete multi-modality supervision.
【8】 Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection标题:注意力融合:一种基于Transformer的多模式仇恨语音检测方法链接:https://arxiv.org/abs/2401.10653作者:Atanu Mandal,Gargi Roy,Amit Barman,Indranil Dutta,Sudip Kumar Naskar备注:Accepted in 20th International Conference on Natural Language Processing (ICON)摘要:随着最近社交媒体使用的激增和指数增长,审查社交媒体内容是否存在任何仇恨内容至关重要。自过去十年以来,研究人员一直在努力区分促进仇恨的内容和不促进仇恨的内容。传统上,主要的重点是分析文本内容。然而,最近的研究尝试也开始进入基于音频的内容的识别。然而,研究表明,仅仅依靠音频或基于文本的内容可能是无效的,因为最近的高潮表明,人们经常在他们的演讲和写作中使用讽刺。为了克服这些挑战,我们提出了一种方法来确定是否演讲促进仇恨或不利用音频和文本表示。我们的方法基于Transformer框架,该框架结合了音频和文本采样,并配有我们自己的层“Attentive Fusion”。我们的研究结果超过了以前的最先进的技术,在测试集上获得了令人印象深刻的宏观F1分数0.927。摘要:With the recent surge and exponential growth of social media usage, scrutinizing social media content for the presence of any hateful content is of utmost importance. Researchers have been diligently working since the past decade on distinguishing between content that promotes hatred and content that does not. Traditionally, the main focus has been on analyzing textual content. However, recent research attempts have also commenced into the identification of audio-based content. Nevertheless, studies have shown that relying solely on audio or text-based content may be ineffective, as recent upsurge indicates that individuals often employ sarcasm in their speech and writing. To overcome these challenges, we present an approach to identify whether a speech promotes hate or not utilizing both audio and textual representations. Our methodology is based on the Transformer framework that incorporates both audio and text sampling, accompanied by our very own layer called "Attentive Fusion". The results of our study surpassed previous state-of-the-art techniques, achieving an impressive macro F1 score of 0.927 on the Test Set. 【9】 AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks标题:AAT:适应各种声学识别任务的音频Transformer链接:https://arxiv.org/abs/2401.10544作者:Yun Liang,Hai Lin,Shaojian Qiu,Yihang Zhang备注:Preprint version for ICASSP 2024, Korea摘要:最近,Transformers被引入到声学识别领域。它们使用监督学习和半监督学习等方法在大规模数据集上进行预训练,表现出强大的通用性-它可以轻松微调下游任务,并显示出更强大的性能。然而,目前使用的主要微调方法仍然是完全微调,这涉及在训练期间更新所有参数。这不仅会导致大量的内存使用和时间成本,而且还会损害模型的通用性。其他微调方法要么难以解决这个问题,要么无法实现匹配的性能。为此,本文对现有的调优方法进行了综合分析,提出了一种基于适配器调优的高效调优方法,即AAT。其核心思想是冻结音频Transformer模型并插入额外的可学习适配器,有效地获取下游任务知识,而不损害模型的原始通用性。大量的实验表明,我们的方法实现了性能相当,甚至优于完全微调,同时优化只有7.118%的参数。它也证明了优于其他微调方法。摘要:Recently, Transformers have been introduced into the field of acoustics recognition. They are pre-trained on large-scale datasets using methods such as supervised learning and semi-supervised learning, demonstrating robust generality--It fine-tunes easily to downstream tasks and shows more robust performance. However, the predominant fine-tuning method currently used is still full fine-tuning, which involves updating all parameters during training. This not only incurs significant memory usage and time costs but also compromises the model's generality. Other fine-tuning methods either struggle to address this issue or fail to achieve matching performance. Therefore, we conducted a comprehensive analysis of existing fine-tuning methods and proposed an efficient fine-tuning approach based on Adapter tuning, namely AAT. The core idea is to freeze the audio Transformer model and insert extra learnable Adapters, efficiently acquiring downstream task knowledge without compromising the model's original generality. Extensive experiments have shown that our method achieves performance comparable to or even superior to full fine-tuning while optimizing only 7.118% of the parameters. It also demonstrates superiority over other fine-tuning methods.
【10】 Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech标题:用于无词典文本到语音的数据驱动的字素到音素表示链接:https://arxiv.org/abs/2401.10465作者:Abhinav Garg,Jiyeon Kim,Sushil Khyalia,Chanwoo Kim,Dhananjaya Gowda备注:Accepted at ICASSP 2024摘要:在任何现代高质量的文本到语音(TTS)系统中,字素到音素(G2P)都是必不可少的第一步。目前的大多数G2P系统依赖于专家精心制作的词典。这造成了双重问题。首先,词典是使用固定的音素集生成的,通常是ARPABET或IPA,这可能不是表示所有语言音素的最佳方式。其次,制作这样一本专业词典所需的工时非常高。在本文中,我们通过使用自监督学习的最新进展来消除这两个问题,以获得数据驱动的音素表示,而不是固定的表示。我们将我们的无词典方法与利用精心制作的词典的强大基线进行比较。此外,我们表明,我们的数据驱动的无词典方法在平均意见得分(MOS)方面与传统的基于规则或基于词典的神经G2P一样好,甚至略好于传统的基于规则或基于词典的神经G2P,而不使用先前的语言词典或音素集,即没有语言专业知识。摘要:Grapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, usually, ARPABET or IPA, which might not be the most optimal way to represent phonemes for all languages. Secondly, the man-hours required to produce such an expert lexicon are very high. In this paper, we eliminate both of these issues by using recent advances in self-supervised learning to obtain data-driven phoneme representations instead of fixed representations. We compare our lexicon-free approach against strong baselines that utilize a well-crafted lexicon. Furthermore, we show that our data-driven lexicon-free method performs as good or even marginally better than the conventional rule-based or lexicon-based neural G2Ps in terms of Mean Opinion Score (MOS) while using no prior language lexicon or phoneme set, i.e. no linguistic expertise. 【11】 Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis标题:用于高质量语音合成的超轻量级神经差分DSP声码器链接:https://arxiv.org/abs/2401.10460作者:Prabhav Agrawal,Thilo Koehler,Zhiping Xiu,Prashant Serai,Qing He备注:Accepted for ICASSP 2024摘要:神经声码器对原始音频波形进行建模并合成高质量的音频,但即使是像MB-MelGAN和LPCNet这样的高效声码器,也无法在智能眼镜等低端设备上实时运行。基于纯数字信号处理(DSP)的声码器可以经由轻量级快速傅立叶变换(FFT)来实现,并且因此比任何神经声码器快一个数量级。DSP声码器通常由于消耗声道的近似表示的过平滑声学模型预测而获得较低的音频质量。在本文中,我们提出了一个超轻量的差分DSP(DDSP)声码器,使用一个联合优化的声学模型与DSP声码器,学习没有提取的声道频谱特征。该模型实现的音频质量与神经声码器相当,平均MOS为4.36,同时作为DSP声码器是有效的。我们的C++实现,没有任何硬件特定的优化,是在15 MFLOPS,超过MB-MelGAN的340倍的FLOPS,并实现了一个声码器的RTF为0.003和整体RTF为0.044,同时运行在一个2GHz的英特尔至强CPU上的单线程。摘要:Neural vocoders model the raw audio waveform and synthesize high-quality audio, but even the highly efficient ones, like MB-MelGAN and LPCNet, fail to run real-time on a low-end device like a smartglass. A pure digital signal processing (DSP) based vocoder can be implemented via lightweight fast Fourier transforms (FFT), and therefore, is a magnitude faster than any neural vocoder. A DSP vocoder often gets a lower audio quality due to consuming over-smoothed acoustic model predictions of approximate representations for the vocal tract. In this paper, we propose an ultra-lightweight differential DSP (DDSP) vocoder that uses a jointly optimized acoustic model with a DSP vocoder, and learns without an extracted spectral feature for the vocal tract. The model achieves audio quality comparable to neural vocoders with a high average MOS of 4.36 while being efficient as a DSP vocoder. Our C++ implementation, without any hardware-specific optimization, is at 15 MFLOPS, surpasses MB-MelGAN by 340 times in terms of FLOPS, and achieves a vocoder-only RTF of 0.003 and overall RTF of 0.044 while running single-threaded on a 2GHz Intel Xeon CPU.
【12】 Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition标题:语音识别语言建模低阶自适应训练策略及模型稳健性研究链接:https://arxiv.org/abs/2401.10447作者:Yu Yu,Chao-Han Huck Yang,Tuan Dinh,Sungho Ryu,Jari Kolehmainen,Roger Ren,Denis Filimonov,Prashanth G. Shivakumar,Ankur Gandhe,Ariya Rastow,Jia Xu,Ivan Bulyko,Andreas Stolcke摘要:低秩自适应(LoRA)与冻结预训练语言模型(PLM)的使用已经成为内存受限硬件的主流,资源高效的建模方法。在这项研究中,我们首先探索如何通过引入各种LoRA训练策略来提高模型性能,在公共Librispeech数据集上实现相对单词错误率降低3.50%,在消息传递领域的内部数据集上降低3.67%。为了进一步表征基于LoRA的第二遍语音识别模型的稳定性,我们研究了对输入扰动的鲁棒性。这些扰动植根于同音字替换和一种称为N-best Perturbation-based Rescoring Robustness(NPRR)的新度量,两者都旨在测量重新评分模型性能的相对退化。我们的实验结果表明,虽然先进的变种LoRA,如动态排名分配LoRA,导致性能下降,在$1$-最好的扰动,他们减轻了退化,在$N$-最好的扰动。这一发现与完全调优的模型和普通LoRA调优基线进行了比较,表明在使用基于LoRA的自适应以节省计算成本和强大的语言建模时,需要进行全面的选择。摘要:The use of low-rank adaptation (LoRA) with frozen pretrained language models (PLMs) has become increasing popular as a mainstream, resource-efficient modeling approach for memory-constrained hardware. In this study, we first explore how to enhance model performance by introducing various LoRA training strategies, achieving relative word error rate reductions of 3.50\% on the public Librispeech dataset and of 3.67\% on an internal dataset in the messaging domain. To further characterize the stability of LoRA-based second-pass speech recognition models, we examine robustness against input perturbations. These perturbations are rooted in homophone replacements and a novel metric called N-best Perturbation-based Rescoring Robustness (NPRR), both designed to measure the relative degradation in the performance of rescoring models. Our experimental results indicate that while advanced variants of LoRA, such as dynamic rank-allocated LoRA, lead to performance degradation in $1$-best perturbation, they alleviate the degradation in $N$-best perturbation. This finding is in comparison to fully-tuned models and vanilla LoRA tuning baselines, suggesting that a comprehensive selection is needed when using LoRA-based adaptation for compute-cost savings and robust language modeling. 【13】 Large Language Models are Efficient Learners of Noise-Robust Speech Recognition标题:大语言模型是抗噪语音识别的有效学习器链接:https://arxiv.org/abs/2401.10446作者:Yuchen Hu,Chen Chen,Chao-Han Huck Yang,Ruizhe Li,Chao Zhang,Pin-Yu Chen,EnSiong Chng备注:Accepted to ICLR 2024, Spotlight top 5%, 24 pages. This work will be open sourced at: this https URL under MIT license摘要:大型语言模型(LLM)的最新进展促进了自动语音识别(ASR)的生成纠错(GER),其利用LLM丰富的语言知识和强大的推理能力来改善识别结果。最新的工作提出了一个GER基准与HyPorcraft数据集学习映射从ASR N-best假设地面真理转录通过有效的LLM微调,这表明了很大的有效性,但缺乏特异性的噪声鲁棒ASR。在这项工作中,我们将基准扩展到噪声条件,并研究我们是否可以教LLM为GER执行去噪,就像鲁棒ASR所做的那样,其中一个解决方案是将噪声信息作为条件引入LLM。然而,由于跨模态间隙,直接结合来自音频编码器的噪声嵌入可能会损害LLM调谐。为此,本文提出从N-best列表中提取一个语言空间噪声嵌入来表征源语音的噪声状况,从而促进了GER的去噪过程;此外,为了增强其对音频噪声的表征能力,本文设计了一种基于互信息估计的知识提取方法,将音频嵌入中的真实噪声信息提取到我们的语言嵌入中。在各种最新的LLM上的实验表明,我们的方法在训练数据有限的情况下,在字错误率方面获得了高达53.9%的校正改进,实现了新的突破。分析表明,我们的语言空间噪声嵌入可以很好地代表源语音的噪声条件下,现成的LLM显示出强大的语言空间去噪能力。摘要:Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results. The latest work proposes a GER benchmark with HyPoradise dataset to learn the mapping from ASR N-best hypotheses to ground-truth transcription by efficient LLM finetuning, which shows great effectiveness but lacks specificity on noise-robust ASR. In this work, we extend the benchmark to noisy conditions and investigate if we can teach LLMs to perform denoising for GER just like what robust ASR do}, where one solution is introducing noise information as a conditioner into LLM. However, directly incorporating noise embeddings from audio encoder could harm the LLM tuning due to cross-modality gap. To this end, we propose to extract a language-space noise embedding from the N-best list to represent the noise conditions of source speech, which can promote the denoising process in GER. Furthermore, in order to enhance its representation ability of audio noise, we design a knowledge distillation (KD) approach via mutual information estimation to distill the real noise information in audio embeddings to our language embedding. Experiments on various latest LLMs demonstrate our approach achieves a new breakthrough with up to 53.9% correction improvement in terms of word error rate while with limited training data. Analysis shows that our language-space noise embedding can well represent the noise conditions of source speech, under which off-the-shelf LLMs show strong ability of language-space denoising.
【14】 DanceMeld: Unraveling Dance Phrases with Hierarchical Latent Codes for Music-to-Dance Synthesis标题:DanceMeld:用于音乐到舞蹈合成的具有层次化潜在代码的舞蹈短语拆解链接:https://arxiv.org/abs/2401.10242作者:Xin Gao,Li Hu,Peng Zhang,Bang Zhang,Liefeng Bo备注:10 pages, 8 figures摘要:在3D数字人类应用领域,音乐到舞蹈提出了一项具有挑战性的任务。鉴于音乐和舞蹈之间的一对多关系,以前的方法在方法上受到限制,仅依赖于基于音乐节奏的匹配和生成相应的舞蹈动作。在专业领域的舞蹈,一个舞蹈短语包括几个舞蹈姿势和舞蹈动作。舞蹈姿势是由一系列基本的有意义的身体姿势组成的,而舞蹈动作则可以反映舞蹈的节奏、旋律、风格等动态变化。从这些概念中获得灵感,我们引入了一个创新的舞蹈生成管道DanceMeld,它包括两个阶段,即,舞蹈分离阶段和舞蹈生成阶段。在解耦阶段,分层VQ-VAE用于在不同的特征空间级别上解开舞蹈姿势和舞蹈动作,其中底部代码表示舞蹈姿势,并且顶部代码表示舞蹈动作。在生成阶段,我们利用一个扩散模型作为一个先验模型的分布和生成潜在的代码的音乐功能的条件。我们已经通过实验证明了顶层代码和底层代码的代表性能力,使舞蹈姿势和舞蹈动作的明确解耦表达。这种解缠不仅提供了对动作细节、风格和节奏的控制,而且还促进了诸如舞蹈风格转移和舞蹈单元编辑之类的应用。在AIST++数据集上进行了定性和定量实验,证明了该方法的优越性。摘要:In the realm of 3D digital human applications, music-to-dance presents a challenging task. Given the one-to-many relationship between music and dance, previous methods have been limited in their approach, relying solely on matching and generating corresponding dance movements based on music rhythm. In the professional field of choreography, a dance phrase consists of several dance poses and dance movements. Dance poses composed of a series of basic meaningful body postures, while dance movements can reflect dynamic changes such as the rhythm, melody, and style of dance. Taking inspiration from these concepts, we introduce an innovative dance generation pipeline called DanceMeld, which comprising two stages, i.e., the dance decouple stage and the dance generation stage. In the decouple stage, a hierarchical VQ-VAE is used to disentangle dance poses and dance movements in different feature space levels, where the bottom code represents dance poses, and the top code represents dance movements. In the generation stage, we utilize a diffusion model as a prior to model the distribution and generate latent codes conditioned on music features. We have experimentally demonstrated the representational capabilities of top code and bottom code, enabling the explicit decoupling expression of dance poses and dance movements. This disentanglement not only provides control over motion details, styles, and rhythm but also facilitates applications such as dance style transfer and dance unit editing. Our approach has undergone qualitative and quantitative experiments on the AIST++ dataset, demonstrating its superiority over other methods.