本文经arXiv每日学术速递授权转载
标题: 自动语音识别的动态数据修剪
作者:Qiao Xiao,Pingchuan Ma,Adriana Fernandez-Lopez,Boqian Wu,Lu Yin,Stavros Petridis,Mykola Pechenizkiy,Maja Pantic,Decebal Constantin Mocanu,Shiwei Liu
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)最近的成功在很大程度上归功于不断增长的训练数据量。然而,这种趋势使得模型训练成本过高,并提出了计算需求。虽然已经提出了通过识别相关数据的一小部分来缓解这个问题的数据修剪,但它在ASR中的应用几乎没有被探索过,并且现有的工作通常需要大量的开销来实现有意义的结果。为了填补这一空白,本文提出了第一次调查的动态数据修剪的ASR,发现我们可以达到全数据性能的动态选择70%的数据。此外,我们介绍了动态数据修剪ASR(DDP-ASR),它提供了几个细粒度的修剪粒度专门为语音相关的数据集,超越了传统的修剪整个时间序列。我们密集的实验表明,DDP-ASR可以节省高达1.6倍的训练时间,而性能损失可以忽略不计。摘要:The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitively costly and imposed computational demands. While data pruning has been proposed to mitigate this issue by identifying a small subset of relevant data, its application in ASR has been barely explored, and existing works often entail significant overhead to achieve meaningful results. To fill this gap, this paper presents the first investigation of dynamic data pruning for ASR, finding that we can reach the full-data performance by dynamically selecting 70% of data. Furthermore, we introduce Dynamic Data Pruning for ASR (DDP-ASR), which offers several fine-grained pruning granularities specifically tailored for speech-related datasets, going beyond the conventional pruning of entire time sequences. Our intensive experiments show that DDP-ASR can save up to 1.6x training time with negligible performance loss.
【2】 Advancing Airport Tower Command Recognition: Integrating Squeeze-and-Excitation and Broadcasted Residual Learning
标题: 推进机场塔楼命令识别:集成挤压和激励和广播剩余学习
作者:Yuanxi Lin,Tonglin Zhou,Yang Xiao
备注:Accepted by IALP 2024
链接:点击下载PDF文件
摘要:准确识别航空指令对于飞行安全和效率至关重要,因为飞行员必须精确地遵循空中交通管制指令。本文通过提出关键词识别技术,解决了语音命令识别中的挑战,如嘈杂的环境和有限的计算资源。我们创建了一个标准化的机场塔台命令数据集,包括例行和紧急指令。我们用挤压和激励技术和时间帧频率挤压和激励技术增强广播残差学习,从而得到我们的BC-SENet模型。该模型以较少的参数关注关键信息。我们对包括BC-SENet在内的五个关键字定位模型进行的测试显示出卓越的准确性和效率。这些发现强调了我们的模型进步在提高语音命令识别方面的有效性,以确保在嘈杂,高风险环境中的航空安全和效率。此外,BC-SENet在常见的Google Speech Command数据集上显示出相当的性能。摘要:Accurate recognition of aviation commands is vital for flight safety and efficiency, as pilots must follow air traffic control instructions precisely. This paper addresses challenges in speech command recognition, such as noisy environments and limited computational resources, by advancing keyword spotting technology. We create a dataset of standardized airport tower commands, including routine and emergency instructions. We enhance broadcasted residual learning with squeeze-and-excitation and time-frame frequency-wise squeeze-and-excitation techniques, resulting in our BC-SENet model. This model focuses on crucial information with fewer parameters. Our tests on five keyword spotting models, including BC-SENet, demonstrate superior accuracy and efficiency. These findings highlight the effectiveness of our model advancements in improving speech command recognition for aviation safety and efficiency in noisy, high-stakes environments. Additionally, BC-SENet shows comparable performance on the common Google Speech Command dataset.
【3】 Automatic Speech Recognition for Hindi
标题: 印地语自动语音识别
作者:Anish Saha,A. G. Ramakrishnan
链接:点击下载PDF文件
摘要:自动语音识别(ASR)是计算语言学的一个关键领域,专注于开发使计算机能够将口语转换为文本的技术。这个领域结合了语言学和机器学习。ASR模型通过监督学习将语音音频映射到成绩单,需要处理真实和不受限制的文本。文本到语音系统直接使用真实文本,而ASR系统依赖于在大型文本语料库上训练的语言模型。高质量的转录数据对于训练预测模型至关重要。该研究涉及两个主要部分:开发一个Web应用程序和设计一个语音识别的Web界面。该Web应用程序使用JavaScript和Node.js创建,可管理大量音频文件及其转录,促进对ASR转录的协作人工纠正。它使用客户端-服务器架构实时运行。用于语音识别的Web界面记录来自运行Web应用程序的任何设备的16 kHz单声道音频,执行语音活动检测(VAD),并将音频发送到识别引擎。VAD检测人类语音存在,帮助有效的语音处理并减少非语音间隔期间的不必要处理,从而节省VoIP应用中的计算和网络带宽。研究的最后阶段测试了一个神经网络,用于准确地将语音信号与隐马尔可夫模型(HMM)状态对齐。这包括实现一种新的反向传播方法,该方法利用节点协同激活的先验统计数据。摘要:Automatic speech recognition (ASR) is a key area in computational linguistics, focusing on developing technologies that enable computers to convert spoken language into text. This field combines linguistics and machine learning. ASR models, which map speech audio to transcripts through supervised learning, require handling real and unrestricted text. Text-to-speech systems directly work with real text, while ASR systems rely on language models trained on large text corpora. High-quality transcribed data is essential for training predictive models. The research involved two main components: developing a web application and designing a web interface for speech recognition. The web application, created with JavaScript and Node.js, manages large volumes of audio files and their transcriptions, facilitating collaborative human correction of ASR transcripts. It operates in real-time using a client-server architecture. The web interface for speech recognition records 16 kHz mono audio from any device running the web app, performs voice activity detection (VAD), and sends the audio to the recognition engine. VAD detects human speech presence, aiding efficient speech processing and reducing unnecessary processing during non-speech intervals, thus saving computation and network bandwidth in VoIP applications. The final phase of the research tested a neural network for accurately aligning the speech signal to hidden Markov model (HMM) states. This included implementing a novel backpropagation method that utilizes prior statistics of node co-activations.
【4】 Token-Weighted RNN-T for Learning from Flawed Data
标题: 令牌加权RNN-T用于从有缺陷的数据中学习
作者:Gil Keren,Wei Zhou,Ozlem Kalinli
链接:点击下载PDF文件
摘要:ASR模型通常使用交叉熵标准进行训练,以增加目标令牌序列的概率。虽然优化目标序列中所有标记的概率是明智的,但人们可能希望不强调反映转录错误的标记。在这项工作中,我们提出了一种新的令牌加权RNN-T标准,该标准通过特定于令牌的权重来增强RNN-T目标。新目标用于减轻训练数据中的transmittance错误造成的准确性损失,这些错误自然会出现在两种设置中:伪标记和人工注释错误。实验结果表明,使用我们的方法进行半监督学习与伪标签导致一致的准确性提高,高达38%的相对。我们还分析了参考转录中不同水平的WER导致的准确性下降,并表明标记加权RNN-T适合于克服这种下降,恢复64%-99%的准确性损失。摘要:ASR models are commonly trained with the cross-entropy criterion to increase the probability of a target token sequence. While optimizing the probability of all tokens in the target sequence is sensible, one may want to de-emphasize tokens that reflect transcription errors. In this work, we propose a novel token-weighted RNN-T criterion that augments the RNN-T objective with token-specific weights. The new objective is used for mitigating accuracy loss from transcriptions errors in the training data, which naturally appear in two settings: pseudo-labeling and human annotation errors. Experiments results show that using our method for semi-supervised learning with pseudo-labels leads to a consistent accuracy improvement, up to 38% relative. We also analyze the accuracy degradation resulting from different levels of WER in the reference transcription, and show that token-weighted RNN-T is suitable for overcoming this degradation, recovering 64%-99% of the accuracy loss.
【5】 A Study on Synthesizing Expressive Violin Performances: Approaches and Comparisons
标题: 小提琴表现力综合表演的研究:方法与比较
作者:Tzu-Yun Hung,Jui-Te Wu,Yu-Chia Kuo,Yo-Wei Hsiao,Ting-Wei Lin,Li Su
备注:15 pages, 2 figures, 3 tables
链接:点击下载PDF文件
摘要:小提琴演奏中的表现性音乐合成(EMS)是一项具有挑战性的任务,这是由于音乐演奏者之间在表现性音乐术语(EMT)的解释上存在分歧,标记录音的稀缺性以及合成模型的泛化能力有限。这些挑战创造了模型的有效性,多样性产生的结果,和合成系统的可控性之间的权衡,使得有必要进行EMS模型设计的比较研究。本文探讨了两种小提琴EMS方法。端到端方法是对最先进的文本到语音生成器的修改。参数控制的方法是基于一个简单的参数采样过程,可以渲染音符长度和其他参数与MIDI-DDSP兼容。我们研究这两种方法(共三个模型变量),通过客观和主观的实验,并讨论了EMS的几个关键问题的基础上的结果。摘要:Expressive music synthesis (EMS) for violin performance is a challenging task due to the disagreement among music performers in the interpretation of expressive musical terms (EMTs), scarcity of labeled recordings, and limited generalization ability of the synthesis model. These challenges create trade-offs between model effectiveness, diversity of generated results, and controllability of the synthesis system, making it essential to conduct a comparative study on EMS model design. This paper explores two violin EMS approaches. The end-to-end approach is a modification of a state-of-the-art text-to-speech generator. The parameter-controlled approach is based on a simple parameter sampling process that can render note lengths and other parameters compatible with MIDI-DDSP. We study these two approaches (in total, three model variants) through objective and subjective experiments and discuss several key issues of EMS based on the results.
【6】 LLM-Driven Multimodal Opinion Expression Identification
标题: LLM驱动的多模式意见表达识别
作者:Bonian Jia,Huiyao Chen,Yueheng Sun,Meishan Zhang,Min Zhang
备注:6 pages, 3 Figures
链接:点击下载PDF文件
摘要:意见表达识别(OEI)在NLP中对于从语音助理到抑郁症诊断的应用至关重要。这项研究扩展了OEI,包括多模态输入,强调听觉线索的意义,在提供超出文本的能力的情感微妙之处。我们介绍了一种新的多模态OEI(MOEI)任务,整合文本和语音来反映真实世界的场景。利用CMU MOSEI和IEMOCAP数据集,我们构建了CI-MOEI数据集。此外,将文本到语音(TTS)技术应用于MPQA数据集以获得CIM-OEI数据集。我们为OEI任务设计了一个模板,以充分利用大型语言模型(LLM)的生成能力。进一步推进,我们提出了一个LLM驱动的方法STOEI,它结合语音和文本模式来识别意见表达。实验结果表明,MOEI算法显著提高了算法的性能,比现有算法提高了9.20%,并获得了SOTA结果.摘要:Opinion Expression Identification (OEI) is essential in NLP for applications ranging from voice assistants to depression diagnosis. This study extends OEI to encompass multimodal inputs, underlining the significance of auditory cues in delivering emotional subtleties beyond the capabilities of text. We introduce a novel multimodal OEI (MOEI) task, integrating text and speech to mirror real-world scenarios. Utilizing CMU MOSEI and IEMOCAP datasets, we construct the CI-MOEI dataset. Additionally, Text-to-Speech (TTS) technology is applied to the MPQA dataset to obtain the CIM-OEI dataset. We design a template for the OEI task to take full advantage of the generative power of large language models (LLMs). Advancing further, we propose an LLM-driven method STOEI, which combines speech and text modal to identify opinion expressions. Our experiments demonstrate that MOEI significantly improves the performance while our method outperforms existing methods by 9.20 % and obtains SOTA results.
【7】 SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
标题: SC-MoE:统一流媒体和非流媒体代码交换ASB的交换一致者专家混合
作者:Shuaishuai Ye,Shunfei Chen,Xinhui Hu,Xinkang Xu
备注:Accepted by InterSpeech 2024; 5 pages, 2 figures
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了一个基于Switch-Conformer的MoE系统,命名为SC-MoE,用于统一的流和非流代码切换(CS)自动语音识别(ASR),其中我们设计了一个由三个语言专家组成的流MoE层,分别对应于普通话,英语和空白,并配备了一个语言识别(LID)网络与连接时间分类(CTC)的损失作为路由器在编码器的SC-MoE,以实现真正的-时间流CS ASR系统。为了进一步利用嵌入在文本中的语言信息,我们还将MoE层纳入SC-MoE的解码器。此外,我们在编码器和解码器的每个MoE层中引入了路由器,并获得了更好的识别性能。实验结果表明,SC-MoE显着提高CS ASR性能与基线相当的计算效率。摘要:In this work, we propose a Switch-Conformer-based MoE system named SC-MoE for unified streaming and non-streaming code-switching (CS) automatic speech recognition (ASR), where we design a streaming MoE layer consisting of three language experts, which correspond to Mandarin, English, and blank, respectively, and equipped with a language identification (LID) network with a Connectionist Temporal Classification (CTC) loss as a router in the encoder of SC-MoE to achieve a real-time streaming CS ASR system. To further utilize the language information embedded in text, we also incorporate MoE layers into the decoder of SC-MoE. In addition, we introduce routers into every MoE layer of the encoder and the decoder and achieve better recognition performance. Experimental results show that the SC-MoE significantly improves CS ASR performances over baseline with comparable computational efficiency.
【8】 Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
标题: 通过学习单调对齐提高基于LLM的语音合成的鲁棒性
作者:Paarth Neekhara,Shehzeen Hussain,Subhankar Ghosh,Jason Li,Rafael Valle,Rohan Badlani,Boris Ginsburg
备注:Published as a conference paper at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:基于大语言模型(LLM)的文本到语音(TTS)系统在处理大型语音数据集和为新说话人生成自然语音方面表现出了卓越的能力。然而,基于LLM的TTS模型并不鲁棒,因为生成的输出可能包含重复的单词,丢失的单词和未对齐的语音(称为幻觉或注意力错误),特别是当文本包含多次出现的相同标记时。我们在编码器-解码器Transformer模型中研究了这些挑战,发现在这种模型中,某些交叉注意头在训练用于预测给定文本的语音标记时,隐式地学习文本和语音对齐。为了使对齐更加鲁棒,我们提出了利用CTC损失和注意力先验的技术,鼓励单调的交叉注意的文本标记。我们的引导注意力训练技术不引入任何新的可学习参数,并显着提高了基于LLM的TTS模型的鲁棒性。摘要:Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers. However, LLM-based TTS models are not robust as the generated output can contain repeating words, missing words and mis-aligned speech (referred to as hallucinations or attention errors), especially when the text contains multiple occurrences of the same token. We examine these challenges in an encoder-decoder transformer model and find that certain cross-attention heads in such models implicitly learn the text and speech alignment when trained for predicting speech tokens for a given text. To make the alignment more robust, we propose techniques utilizing CTC loss and attention priors that encourage monotonic cross-attention over the text tokens. Our guided attention training technique does not introduce any new learnable parameters and significantly improves robustness of LLM-based TTS models.
【9】 Sequential Editing for Lifelong Training of Speech Recognition Models
标题: 语音识别模型终身训练的序列编辑
作者:Devang Kulshreshtha,Saket Dingliwal,Brady Houston,Nikolaos Pappas,Srikanth Ronanki
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)传统上假设已知域,但从新域添加数据会引起对与现有和新域上的再训练模型相关的计算效率低下的担忧。仅对新领域进行微调会带来灾难性遗忘(CF)的风险。为了解决这个问题,终身学习(LLL)算法已经提出了ASR。先前的研究已经探索了诸如弹性权重合并、知识蒸馏和重放等技术,所有这些技术都需要额外的参数或访问先前的域数据。我们提出顺序模型编辑作为一种新的方法,不断学习新的领域在ASR系统。与以前的方法不同,我们的方法不需要访问以前的数据集或引入额外的参数。我们的研究表明,在微调基线上,单词错误率降低了15%,并且在CommonVoice英语多口音数据集上的效率优于其他LLL技术。摘要:Automatic Speech Recognition (ASR) traditionally assumes known domains, but adding data from a new domain raises concerns about computational inefficiencies linked to retraining models on both existing and new domains. Fine-tuning solely on new domain risks Catastrophic Forgetting (CF). To address this, Lifelong Learning (LLL) algorithms have been proposed for ASR. Prior research has explored techniques such as Elastic Weight Consolidation, Knowledge Distillation, and Replay, all of which necessitate either additional parameters or access to prior domain data. We propose Sequential Model Editing as a novel method to continually learn new domains in ASR systems. Different than previous methods, our approach does not necessitate access to prior datasets or the introduction of extra parameters. Our study demonstrates up to 15% Word Error Rate Reduction (WERR) over fine-tuning baseline, and superior efficiency over other LLL techniques on CommonVoice English multi-accent dataset.
【10】 SonicSense: Object Perception from In-Hand Acoustic Vibration
标题: SonicSense:来自手内声振动的物体感知
作者:Jiaxun Liu,Boyuan Chen
备注:Our project website is at: this http URL
链接:点击下载PDF文件
摘要:我们介绍SonicSense,这是一种硬件和软件的整体设计,通过手内声学振动传感实现丰富的机器人对象感知。虽然以前的研究已经显示出声学传感用于物体感知的有希望的结果,但目前的解决方案仅限于具有简单几何形状和均匀材料的少数物体,单指传感以及对相同物体的混合训练和测试。SonicSense能够区分集装箱库存状态、异质材料预测、3D形状重建以及从83个真实物体中重新识别物体。我们的系统采用了一种简单但有效的启发式探索策略来与物体进行交互,以及基于端到端学习的算法来融合振动信号以推断物体属性。我们的框架强调了手声学振动传感在推进机器人触觉感知的重要性。摘要:We introduce SonicSense, a holistic design of hardware and software to enable rich robot object perception through in-hand acoustic vibration sensing. While previous studies have shown promising results with acoustic sensing for object perception, current solutions are constrained to a handful of objects with simple geometries and homogeneous materials, single-finger sensing, and mixing training and testing on the same objects. SonicSense enables container inventory status differentiation, heterogeneous material prediction, 3D shape reconstruction, and object re-identification from a diverse set of 83 real-world objects. Our system employs a simple but effective heuristic exploration policy to interact with the objects as well as end-to-end learning-based algorithms to fuse vibration signals to infer object properties. Our framework underscores the significance of in-hand acoustic vibration sensing in advancing robot tactile perception.
【11】 FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data
标题: FASA:一种灵活的自动语音对齐器,用于提取高质量对齐儿童语音数据
作者:Dancheng Liu,Jinjun Xiong
备注:4 pages, 1 figure
链接:点击下载PDF文件
摘要:近年来,针对成人语音的自动语音识别(ASR)通过采用深度神经网络(DNN)模型取得了显著进展,但由于儿童语音的独特特征,儿童语音的改进仍然不尽如人意。在成人数据上预先训练的DNN模型通常难以通过微调来概括儿童的讲话,因为缺乏高质量的对齐儿童讲话。当生成数据集时,人工注释是不可扩展的,并且现有的强制对齐工具是不可用的,因为它们对输入数据的质量做出了不切实际的假设。为了解决这些挑战,我们提出了一个新的强制对齐工具,FASA,作为一个灵活的和自动的语音对齐器,从许多现有的嘈杂的儿童的语音数据中提取高质量的对齐儿童的语音数据。我们展示了它的使用CHILDES数据集,并表明,FASA可以提高数据质量的13.6$ times$超过人类注释。摘要:Automatic Speech Recognition (ASR) for adults' speeches has made significant progress by employing deep neural network (DNN) models recently, but improvement in children's speech is still unsatisfactory due to children's speech's distinct characteristics. DNN models pre-trained on adult data often struggle in generalizing children's speeches with fine tuning because of the lack of high-quality aligned children's speeches. When generating datasets, human annotations are not scalable, and existing forced-alignment tools are not usable as they make impractical assumptions about the quality of the input transcriptions. To address these challenges, we propose a new forced-alignment tool, FASA, as a flexible and automatic speech aligner to extract high-quality aligned children's speech data from many of the existing noisy children's speech data. We demonstrate its usage on the CHILDES dataset and show that FASA can improve data quality by 13.6$ times$ over human annotations.
【12】 Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
标题: 使用CNN、双向LSTM和ResNet的尼泊尔语自动语音识别
作者:Manish Dhakal,Arman Chhetri,Aman Kumar Gupta,Prabin Lamichhane,Suraj Pandey,Subarna Shakya
Journal-ref:2022 International Conference on Inventive Computation Technologies (ICICT), pp. 515-521
链接:点击下载PDF文件
摘要:本文提出了一种用于自动语音识别(ASR)的端到端深度学习模型,该模型将尼泊尔语语音转录为文本。该模型在OpenSLR(音频,文本)数据集上进行了训练和测试。大多数音频数据集在两端都有无声的间隙,这些间隙在数据集预处理期间被裁剪,以实现音频帧及其相应文本的更均匀映射。Mel频率倒谱系数(MFCC)用作音频特征以馈送到模型中。具有双向LSTM与ResNet和一维CNN配对的模型在迄今为止已经训练的所有模型(具有LSTM、GRU、CNN和ResNet变体的神经网络)中为该数据集产生最佳结果。该模型使用连接主义时间分类(CTC)函数进行训练过程中的损失计算,并使用CTC波束搜索解码来预测尼泊尔语文本中最有可能的字符序列。在测试数据集上,字符错误率(CER)为17.06%。源代码可在https: github.com manishdhakal ASR-Nepali-using-CNN-BiLSTM-ResNet上获得。摘要:This paper presents an end-to-end deep learning model for Automatic Speech Recognition (ASR) that transcribes Nepali speech to text. The model was trained and tested on the OpenSLR (audio, text) dataset. The majority of the audio dataset have silent gaps at both ends which are clipped during dataset preprocessing for a more uniform mapping of audio frames and their corresponding texts. Mel Frequency Cepstral Coefficients (MFCCs) are used as audio features to feed into the model. The model having Bidirectional LSTM paired with ResNet and one-dimensional CNN produces the best results for this dataset out of all the models (neural networks with variations of LSTM, GRU, CNN, and ResNet) that have been trained so far. This novel model uses Connectionist Temporal Classification (CTC) function for loss calculation during training and CTC beam search decoding for predicting characters as the most likely sequence of Nepali text. On the test dataset, the character error rate (CER) of 17.06 percent has been achieved. The source code is available at: https: github.com manishdhakal ASR-Nepali-using-CNN-BiLSTM-ResNet.
【13】 A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
标题: 基于vits 2的多扬声器多语言语音克隆系统for limmits 2024挑战
作者:Xiaopeng Wang,Yi Lu,Xin Qi,Zhiyong Wang,Yuankun Xie,Shuchen Shi,Ruibo Fu
链接:点击下载PDF文件
摘要:本文介绍了一个语音合成系统的LIMMITS'24挑战赛的发展,主要集中在轨道2。这项挑战的目标是建立一个具有语音克隆能力的多说话者、多语言的印度文语转换系统,涵盖七种印度语言,包括男性和女性。该系统使用挑战数据进行训练,并针对目标扬声器上的Few-Shot语音克隆进行微调。评估包括所有七种语言的单语言和跨语言合成,主观测试评估自然度和说话者相似性。我们的系统使用VITS2架构,增强了多语言ID和BERT模型,以提高上下文语言的理解。在轨道1中,不允许使用额外的数据,我们的模型实现了4.02的说话人相似度得分。在允许使用额外数据的音轨2中,它获得了4.17的说话者相似度得分。摘要:This paper presents the development of a speech synthesis system for the LIMMITS'24 Challenge, focusing primarily on Track 2. The objective of the challenge is to establish a multi-speaker, multi-lingual Indic Text-to-Speech system with voice cloning capabilities, covering seven Indian languages with both male and female speakers. The system was trained using challenge data and fine-tuned for few-shot voice cloning on target speakers. Evaluation included both mono-lingual and cross-lingual synthesis across all seven languages, with subjective tests assessing naturalness and speaker similarity. Our system uses the VITS2 architecture, augmented with a multi-lingual ID and a BERT model to enhance contextual language comprehension. In Track 1, where no additional data usage was permitted, our model achieved a Speaker Similarity score of 4.02. In Track 2, which allowed the use of extra data, it attained a Speaker Similarity score of 4.17.
【14】 MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
标题: MSR-86 K:一个不断发展的多语言数据库,包含86,300小时的转录音频,用于语音识别研究
作者:Song Li,Yongbin You,Xuezhi Wang,Zhengkun Tian,Ke Ding,Guanglu Wan
备注:Accepted by InterSpeech 2024
链接:点击下载PDF文件
摘要:最近,以ChatGPT为例的多语言人工智能助手获得了极大的欢迎。作为人机交互的重要门户,多语言自动语音识别(ASR)也受到了极大的关注,像Whisper这样的系统就是明证。然而,训练数据的专有性质阻碍了研究人员研究多语言ASR的努力。本文介绍了MSR-86 K,一个不断发展的,大规模的语音识别研究的多语种语料库。该语料库来源于YouTube上的公开视频,包括15种语言和总计86,300小时的转录ASR数据。我们还介绍了如何使用MSR-86 K语料库和其他开源语料库来训练一个强大的多语言ASR模型,与Whisper竞争。MSR-86 K将在HuggingFace上公开发布,我们相信这样一个大型语料库将为多语言ASR的研究铺平新的道路。摘要:Recently, multilingual artificial intelligence assistants, exemplified by ChatGPT, have gained immense popularity. As a crucial gateway to human-computer interaction, multilingual automatic speech recognition (ASR) has also garnered significant attention, as evidenced by systems like Whisper. However, the proprietary nature of the training data has impeded researchers' efforts to study multilingual ASR. This paper introduces MSR-86K, an evolving, large-scale multilingual corpus for speech recognition research. The corpus is derived from publicly accessible videos on YouTube, comprising 15 languages and a total of 86,300 hours of transcribed ASR data. We also introduce how to use the MSR-86K corpus and other open-source corpora to train a robust multilingual ASR model that is competitive with Whisper. MSR-86K will be publicly released on HuggingFace, and we believe that such a large corpus will pave new avenues for research in multilingual ASR.
【15】 On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
标题: 语音分类模型的校准:基于能量的模型研究的见解
作者:Yaqian Hao,Chenguang Hu,Yingying Gao,Shilei Zhang,Junlan Feng
链接:点击下载PDF文件
摘要:对于语音分类任务,深度学习模型通常可以实现高准确度,但在校准方面存在不足,表现为分类器表现出过度自信。校准的重要性在于它在保证深度学习系统中决策的可靠性方面的关键作用。本研究探讨了基于能量的模型在校准语音分类任务的置信度的有效性,通过训练一个联合EBM集成的判别和生成模型,从而提高分类器的校准和减轻过度自信。对三个语音分类任务进行了实验评估:年龄,情感和语言识别。我们的研究结果突出了EBM在校准语音分类模型方面的竞争力。这项研究强调了EBM在语音分类任务中的潜力,证明了它们在不牺牲准确性的情况下增强校准的能力。摘要:For speech classification tasks, deep learning models often achieve high accuracy but exhibit shortcomings in calibration, manifesting as classifiers exhibiting overconfidence. The significance of calibration lies in its critical role in guaranteeing the reliability of decision-making within deep learning systems. This study explores the effectiveness of Energy-Based Models in calibrating confidence for speech classification tasks by training a joint EBM integrating a discriminative and a generative model, thereby enhancing the classifiers calibration and mitigating overconfidence. Experimental evaluations conducted on three speech classification tasks specifically: age, emotion, and language recognition. Our findings highlight the competitive performance of EBMs in calibrating the speech classification models. This research emphasizes the potential of EBMs in speech classification tasks, demonstrating their ability to enhance calibration without sacrificing accuracy.
【16】 E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
标题: E2 TTC:令人尴尬的简单完全非自回归Zero-ShotTTC
作者:Sefik Emre Eskimez,Xiaofei Wang,Manthan Thakker,Canrun Li,Chung-Hsien Tsai,Zhen Xiao,Hemin Yang,Zirun Zhu,Min Tang,Xu Tan,Yanqing Liu,Sheng Zhao,Naoyuki Kanda
链接:点击下载PDF文件
摘要:本文介绍了一种完全非自回归zero-shot的文本到语音转换系统,它提供了人类水平的自然度和最先进的说话人相似度和可懂度。在E2 TTS框架中,文本输入被转换为具有填充符标记的字符序列。然后基于音频填充任务训练基于流匹配的梅尔频谱图生成器。与许多以前的作品不同,它不需要额外的组件(例如,持续时间模型,字素到音素)或复杂技术(例如,单调比对搜索)。尽管简单,E2 TTS实现了最先进的zero-shot TTS功能,可与之前的作品(包括Voicebox和NaturalSpeech 3)相媲美或超越。E2 TTS的简单性还允许输入表示的灵活性。我们提出了几个变种的E2 TTS,以提高推理过程中的可用性。有关演示示例,请访问https: aka.ms e2tts 。摘要:This paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-art speaker similarity and intelligibility. In the E2 TTS framework, the text input is converted into a character sequence with filler tokens. The flow-matching-based mel spectrogram generator is then trained based on the audio infilling task. Unlike many previous works, it does not require additional components (e.g., duration model, grapheme-to-phoneme) or complex techniques (e.g., monotonic alignment search). Despite its simplicity, E2 TTS achieves state-of-the-art zero-shot TTS capabilities that are comparable to or surpass previous works, including Voicebox and NaturalSpeech 3. The simplicity of E2 TTS also allows for flexibility in the input representation. We propose several variants of E2 TTS to improve usability during inference. See https: aka.ms e2tts for demo samples.
【17】 Fish Tracking, Counting, and Behaviour Analysis in Digital Aquaculture: A Comprehensive Review
标题: 数字水族馆中的鱼类追踪、计数和行为分析:全面评论
作者:Meng Cui,Xubo Liu,Haohe Liu,Jinzheng Zhao,Daoliang Li,Wenwu Wang
链接:点击下载PDF文件
摘要:数字水产养殖利用先进的技术和数据驱动的方法,提供了比传统水产养殖实践更大的好处。鱼类跟踪、计数和行为分析是数字水产养殖的重要组成部分,对于优化生产效率、提高鱼类福利和改善资源管理至关重要。以往的审查侧重于单一模式,限制了它们全面应对这些任务中遇到的各种挑战的能力。本文对水产养殖数字化技术的现状进行了全面分析,包括基于视觉、基于声学和基于生物传感器的方法。我们研究这些方法的优点,局限性和应用,突出最近的进展,并确定关键的研究差距。缺乏全面的鱼类数据集和缺乏统一的评价标准,使得难以比较不同技术的性能,被认为是阻碍这一领域取得进展的主要障碍。为了克服当前的局限性,提高鱼类监测系统的准确性、鲁棒性和效率,我们探索了多模态数据融合和深度学习等新兴技术的潜力。此外,我们通过提供现有数据集的摘要来为该领域做出贡献,这些数据集可用于鱼类跟踪,计数和行为分析。概述了未来的研究方向,强调需要全面的数据集和评估标准,以促进技术之间的有意义的比较,并促进其在现实世界的水产养殖环境中的实际实施。摘要:Digital aquaculture leverages advanced technologies and data-driven methods, providing substantial benefits over traditional aquaculture practices. Fish tracking, counting, and behaviour analysis are crucial components of digital aquaculture, which are essential for optimizing production efficiency, enhancing fish welfare, and improving resource management. Previous reviews have focused on single modalities, limiting their ability to address the diverse challenges encountered in these tasks comprehensively. This review provides a comprehensive analysis of the current state of aquaculture digital technologies, including vision-based, acoustic-based, and biosensor-based methods. We examine the advantages, limitations, and applications of these methods, highlighting recent advancements and identifying critical research gaps. The scarcity of comprehensive fish datasets and the lack of unified evaluation standards, which make it difficult to compare the performance of different technologies, are identified as major obstacles hindering progress in this field. To overcome current limitations and improve the accuracy, robustness, and efficiency of fish monitoring systems, we explore the potential of emerging technologies such as multimodal data fusion and deep learning. Additionally, we contribute to the field by providing a summary of existing datasets available for fish tracking, counting, and behaviour analysis. Future research directions are outlined, emphasizing the need for comprehensive datasets and evaluation standards to facilitate meaningful comparisons between technologies and promote their practical implementation in real-world aquaculture settings.
eess.AS音频处理
【1】 MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research标题: MSR-86 K:一个不断发展的多语言数据库,包含86,300小时的转录音频,用于语音识别研究
作者:Song Li,Yongbin You,Xuezhi Wang,Zhengkun Tian,Ke Ding,Guanglu Wan
备注:Accepted by InterSpeech 2024
链接:点击下载PDF文件
摘要:最近,以ChatGPT为例的多语言人工智能助手获得了极大的欢迎。作为人机交互的重要门户,多语言自动语音识别(ASR)也受到了极大的关注,像Whisper这样的系统就是明证。然而,训练数据的专有性质阻碍了研究人员研究多语言ASR的努力。本文介绍了MSR-86 K,一个不断发展的,大规模的语音识别研究的多语种语料库。该语料库来源于YouTube上的公开视频,包括15种语言和总计86,300小时的转录ASR数据。我们还介绍了如何使用MSR-86 K语料库和其他开源语料库来训练一个强大的多语言ASR模型,与Whisper竞争。MSR-86 K将在HuggingFace上公开发布,我们相信这样一个大型语料库将为多语言ASR的研究铺平新的道路。摘要:Recently, multilingual artificial intelligence assistants, exemplified by ChatGPT, have gained immense popularity. As a crucial gateway to human-computer interaction, multilingual automatic speech recognition (ASR) has also garnered significant attention, as evidenced by systems like Whisper. However, the proprietary nature of the training data has impeded researchers' efforts to study multilingual ASR. This paper introduces MSR-86K, an evolving, large-scale multilingual corpus for speech recognition research. The corpus is derived from publicly accessible videos on YouTube, comprising 15 languages and a total of 86,300 hours of transcribed ASR data. We also introduce how to use the MSR-86K corpus and other open-source corpora to train a robust multilingual ASR model that is competitive with Whisper. MSR-86K will be publicly released on HuggingFace, and we believe that such a large corpus will pave new avenues for research in multilingual ASR.
【2】 On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
标题: 语音分类模型的校准:基于能量的模型研究的见解
作者:Yaqian Hao,Chenguang Hu,Yingying Gao,Shilei Zhang,Junlan Feng
链接:点击下载PDF文件
摘要:对于语音分类任务,深度学习模型通常可以实现高准确度,但在校准方面存在不足,表现为分类器表现出过度自信。校准的重要性在于它在保证深度学习系统中决策的可靠性方面的关键作用。本研究探讨了基于能量的模型在校准语音分类任务的置信度的有效性,通过训练一个联合EBM集成的判别和生成模型,从而提高分类器的校准和减轻过度自信。对三个语音分类任务进行了实验评估:年龄,情感和语言识别。我们的研究结果突出了EBM在校准语音分类模型方面的竞争力。这项研究强调了EBM在语音分类任务中的潜力,证明了它们在不牺牲准确性的情况下增强校准的能力。摘要:For speech classification tasks, deep learning models often achieve high accuracy but exhibit shortcomings in calibration, manifesting as classifiers exhibiting overconfidence. The significance of calibration lies in its critical role in guaranteeing the reliability of decision-making within deep learning systems. This study explores the effectiveness of Energy-Based Models in calibrating confidence for speech classification tasks by training a joint EBM integrating a discriminative and a generative model, thereby enhancing the classifiers calibration and mitigating overconfidence. Experimental evaluations conducted on three speech classification tasks specifically: age, emotion, and language recognition. Our findings highlight the competitive performance of EBMs in calibrating the speech classification models. This research emphasizes the potential of EBMs in speech classification tasks, demonstrating their ability to enhance calibration without sacrificing accuracy.
【3】 E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
标题: E2 TTC:令人尴尬的简单完全非自回归Zero-ShotTTC
作者:Sefik Emre Eskimez,Xiaofei Wang,Manthan Thakker,Canrun Li,Chung-Hsien Tsai,Zhen Xiao,Hemin Yang,Zirun Zhu,Min Tang,Xu Tan,Yanqing Liu,Sheng Zhao,Naoyuki Kanda
链接:点击下载PDF文件
摘要:本文介绍了一种完全非自回归zero-shot的文本到语音转换系统,它提供了人类水平的自然度和最先进的说话人相似度和可懂度。在E2 TTS框架中,文本输入被转换为具有填充符标记的字符序列。然后基于音频填充任务训练基于流匹配的梅尔频谱图生成器。与许多以前的作品不同,它不需要额外的组件(例如,持续时间模型,字素到音素)或复杂技术(例如,单调比对搜索)。尽管简单,E2 TTS实现了最先进的zero-shot TTS功能,可与之前的作品(包括Voicebox和NaturalSpeech 3)相媲美或超越。E2 TTS的简单性还允许输入表示的灵活性。我们提出了几个变种的E2 TTS,以提高推理过程中的可用性。有关演示示例,请访问https: aka.ms e2tts 。摘要:This paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-art speaker similarity and intelligibility. In the E2 TTS framework, the text input is converted into a character sequence with filler tokens. The flow-matching-based mel spectrogram generator is then trained based on the audio infilling task. Unlike many previous works, it does not require additional components (e.g., duration model, grapheme-to-phoneme) or complex techniques (e.g., monotonic alignment search). Despite its simplicity, E2 TTS achieves state-of-the-art zero-shot TTS capabilities that are comparable to or surpass previous works, including Voicebox and NaturalSpeech 3. The simplicity of E2 TTS also allows for flexibility in the input representation. We propose several variants of E2 TTS to improve usability during inference. See https: aka.ms e2tts for demo samples.
【4】 Dynamic Data Pruning for Automatic Speech Recognition
标题: 自动语音识别的动态数据修剪
作者:Qiao Xiao,Pingchuan Ma,Adriana Fernandez-Lopez,Boqian Wu,Lu Yin,Stavros Petridis,Mykola Pechenizkiy,Maja Pantic,Decebal Constantin Mocanu,Shiwei Liu
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)最近的成功在很大程度上归功于不断增长的训练数据量。然而,这种趋势使得模型训练成本过高,并提出了计算需求。虽然已经提出了通过识别相关数据的一小部分来缓解这个问题的数据修剪,但它在ASR中的应用几乎没有被探索过,并且现有的工作通常需要大量的开销来实现有意义的结果。为了填补这一空白,本文提出了第一次调查的动态数据修剪的ASR,发现我们可以达到全数据性能的动态选择70%的数据。此外,我们介绍了动态数据修剪ASR(DDP-ASR),它提供了几个细粒度的修剪粒度专门为语音相关的数据集,超越了传统的修剪整个时间序列。我们密集的实验表明,DDP-ASR可以节省高达1.6倍的训练时间,而性能损失可以忽略不计。摘要:The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitively costly and imposed computational demands. While data pruning has been proposed to mitigate this issue by identifying a small subset of relevant data, its application in ASR has been barely explored, and existing works often entail significant overhead to achieve meaningful results. To fill this gap, this paper presents the first investigation of dynamic data pruning for ASR, finding that we can reach the full-data performance by dynamically selecting 70% of data. Furthermore, we introduce Dynamic Data Pruning for ASR (DDP-ASR), which offers several fine-grained pruning granularities specifically tailored for speech-related datasets, going beyond the conventional pruning of entire time sequences. Our intensive experiments show that DDP-ASR can save up to 1.6x training time with negligible performance loss.
【5】 Advancing Airport Tower Command Recognition: Integrating Squeeze-and-Excitation and Broadcasted Residual Learning
标题: 推进机场塔楼命令识别:集成挤压和激励和广播剩余学习
作者:Yuanxi Lin,Tonglin Zhou,Yang Xiao
备注:Accepted by IALP 2024
链接:点击下载PDF文件
摘要:准确识别航空指令对于飞行安全和效率至关重要,因为飞行员必须精确地遵循空中交通管制指令。本文通过提出关键词识别技术,解决了语音命令识别中的挑战,如嘈杂的环境和有限的计算资源。我们创建了一个标准化的机场塔台命令数据集,包括例行和紧急指令。我们用挤压和激励技术和时间帧频率挤压和激励技术增强广播残差学习,从而得到我们的BC-SENet模型。该模型以较少的参数关注关键信息。我们对包括BC-SENet在内的五个关键字定位模型进行的测试显示出卓越的准确性和效率。这些发现强调了我们的模型进步在提高语音命令识别方面的有效性,以确保在嘈杂,高风险环境中的航空安全和效率。此外,BC-SENet在常见的Google Speech Command数据集上显示出相当的性能。摘要:Accurate recognition of aviation commands is vital for flight safety and efficiency, as pilots must follow air traffic control instructions precisely. This paper addresses challenges in speech command recognition, such as noisy environments and limited computational resources, by advancing keyword spotting technology. We create a dataset of standardized airport tower commands, including routine and emergency instructions. We enhance broadcasted residual learning with squeeze-and-excitation and time-frame frequency-wise squeeze-and-excitation techniques, resulting in our BC-SENet model. This model focuses on crucial information with fewer parameters. Our tests on five keyword spotting models, including BC-SENet, demonstrate superior accuracy and efficiency. These findings highlight the effectiveness of our model advancements in improving speech command recognition for aviation safety and efficiency in noisy, high-stakes environments. Additionally, BC-SENet shows comparable performance on the common Google Speech Command dataset.
【6】 Automatic Speech Recognition for Hindi
标题: 印地语自动语音识别
作者:Anish Saha,A. G. Ramakrishnan
链接:点击下载PDF文件
摘要:自动语音识别(ASR)是计算语言学的一个关键领域,专注于开发使计算机能够将口语转换为文本的技术。这个领域结合了语言学和机器学习。ASR模型通过监督学习将语音音频映射到成绩单,需要处理真实和不受限制的文本。文本到语音系统直接使用真实文本,而ASR系统依赖于在大型文本语料库上训练的语言模型。高质量的转录数据对于训练预测模型至关重要。该研究涉及两个主要部分:开发一个Web应用程序和设计一个语音识别的Web界面。该Web应用程序使用JavaScript和Node.js创建,可管理大量音频文件及其转录,促进对ASR转录的协作人工纠正。它使用客户端-服务器架构实时运行。用于语音识别的Web界面记录来自运行Web应用程序的任何设备的16 kHz单声道音频,执行语音活动检测(VAD),并将音频发送到识别引擎。VAD检测人类语音存在,帮助有效的语音处理并减少非语音间隔期间的不必要处理,从而节省VoIP应用中的计算和网络带宽。研究的最后阶段测试了一个神经网络,用于准确地将语音信号与隐马尔可夫模型(HMM)状态对齐。这包括实现一种新的反向传播方法,该方法利用节点协同激活的先验统计数据。摘要:Automatic speech recognition (ASR) is a key area in computational linguistics, focusing on developing technologies that enable computers to convert spoken language into text. This field combines linguistics and machine learning. ASR models, which map speech audio to transcripts through supervised learning, require handling real and unrestricted text. Text-to-speech systems directly work with real text, while ASR systems rely on language models trained on large text corpora. High-quality transcribed data is essential for training predictive models. The research involved two main components: developing a web application and designing a web interface for speech recognition. The web application, created with JavaScript and Node.js, manages large volumes of audio files and their transcriptions, facilitating collaborative human correction of ASR transcripts. It operates in real-time using a client-server architecture. The web interface for speech recognition records 16 kHz mono audio from any device running the web app, performs voice activity detection (VAD), and sends the audio to the recognition engine. VAD detects human speech presence, aiding efficient speech processing and reducing unnecessary processing during non-speech intervals, thus saving computation and network bandwidth in VoIP applications. The final phase of the research tested a neural network for accurately aligning the speech signal to hidden Markov model (HMM) states. This included implementing a novel backpropagation method that utilizes prior statistics of node co-activations.
【7】 Token-Weighted RNN-T for Learning from Flawed Data
标题: 令牌加权RNN-T用于从有缺陷的数据中学习
作者:Gil Keren,Wei Zhou,Ozlem Kalinli
链接:点击下载PDF文件
摘要:ASR模型通常使用交叉熵标准进行训练,以增加目标令牌序列的概率。虽然优化目标序列中所有标记的概率是明智的,但人们可能希望不强调反映转录错误的标记。在这项工作中,我们提出了一种新的令牌加权RNN-T标准,该标准通过特定于令牌的权重来增强RNN-T目标。新目标用于减轻训练数据中的transmittance错误造成的准确性损失,这些错误自然会出现在两种设置中:伪标记和人工注释错误。实验结果表明,使用我们的方法进行半监督学习与伪标签导致一致的准确性提高,高达38%的相对。我们还分析了参考转录中不同水平的WER导致的准确性下降,并表明标记加权RNN-T适合于克服这种下降,恢复64%-99%的准确性损失。摘要:ASR models are commonly trained with the cross-entropy criterion to increase the probability of a target token sequence. While optimizing the probability of all tokens in the target sequence is sensible, one may want to de-emphasize tokens that reflect transcription errors. In this work, we propose a novel token-weighted RNN-T criterion that augments the RNN-T objective with token-specific weights. The new objective is used for mitigating accuracy loss from transcriptions errors in the training data, which naturally appear in two settings: pseudo-labeling and human annotation errors. Experiments results show that using our method for semi-supervised learning with pseudo-labels leads to a consistent accuracy improvement, up to 38% relative. We also analyze the accuracy degradation resulting from different levels of WER in the reference transcription, and show that token-weighted RNN-T is suitable for overcoming this degradation, recovering 64%-99% of the accuracy loss.
【8】 A Study on Synthesizing Expressive Violin Performances: Approaches and Comparisons
标题: 小提琴表现力综合表演的研究:方法与比较
作者:Tzu-Yun Hung,Jui-Te Wu,Yu-Chia Kuo,Yo-Wei Hsiao,Ting-Wei Lin,Li Su
备注:15 pages, 2 figures, 3 tables
链接:点击下载PDF文件
摘要:小提琴演奏中的表现性音乐合成(EMS)是一项具有挑战性的任务,这是由于音乐演奏者之间在表现性音乐术语(EMT)的解释上存在分歧,标记录音的稀缺性以及合成模型的泛化能力有限。这些挑战创造了模型的有效性,多样性产生的结果,和合成系统的可控性之间的权衡,使得有必要进行EMS模型设计的比较研究。本文探讨了两种小提琴EMS方法。端到端方法是对最先进的文本到语音生成器的修改。参数控制的方法是基于一个简单的参数采样过程,可以渲染音符长度和其他参数与MIDI-DDSP兼容。我们研究这两种方法(共三个模型变量),通过客观和主观的实验,并讨论了EMS的几个关键问题的基础上的结果。摘要:Expressive music synthesis (EMS) for violin performance is a challenging task due to the disagreement among music performers in the interpretation of expressive musical terms (EMTs), scarcity of labeled recordings, and limited generalization ability of the synthesis model. These challenges create trade-offs between model effectiveness, diversity of generated results, and controllability of the synthesis system, making it essential to conduct a comparative study on EMS model design. This paper explores two violin EMS approaches. The end-to-end approach is a modification of a state-of-the-art text-to-speech generator. The parameter-controlled approach is based on a simple parameter sampling process that can render note lengths and other parameters compatible with MIDI-DDSP. We study these two approaches (in total, three model variants) through objective and subjective experiments and discuss several key issues of EMS based on the results.
【9】 LLM-Driven Multimodal Opinion Expression Identification
标题: LLM驱动的多模式意见表达识别
作者:Bonian Jia,Huiyao Chen,Yueheng Sun,Meishan Zhang,Min Zhang
备注:6 pages, 3 Figures
链接:点击下载PDF文件
摘要:意见表达识别(OEI)在NLP中对于从语音助理到抑郁症诊断的应用至关重要。这项研究扩展了OEI,包括多模态输入,强调听觉线索的意义,在提供超出文本的能力的情感微妙之处。我们介绍了一种新的多模态OEI(MOEI)任务,整合文本和语音来反映真实世界的场景。利用CMU MOSEI和IEMOCAP数据集,我们构建了CI-MOEI数据集。此外,将文本到语音(TTS)技术应用于MPQA数据集以获得CIM-OEI数据集。我们为OEI任务设计了一个模板,以充分利用大型语言模型(LLM)的生成能力。进一步推进,我们提出了一个LLM驱动的方法STOEI,它结合语音和文本模式来识别意见表达。实验结果表明,MOEI算法显著提高了算法的性能,比现有算法提高了9.20%,并获得了SOTA结果.摘要:Opinion Expression Identification (OEI) is essential in NLP for applications ranging from voice assistants to depression diagnosis. This study extends OEI to encompass multimodal inputs, underlining the significance of auditory cues in delivering emotional subtleties beyond the capabilities of text. We introduce a novel multimodal OEI (MOEI) task, integrating text and speech to mirror real-world scenarios. Utilizing CMU MOSEI and IEMOCAP datasets, we construct the CI-MOEI dataset. Additionally, Text-to-Speech (TTS) technology is applied to the MPQA dataset to obtain the CIM-OEI dataset. We design a template for the OEI task to take full advantage of the generative power of large language models (LLMs). Advancing further, we propose an LLM-driven method STOEI, which combines speech and text modal to identify opinion expressions. Our experiments demonstrate that MOEI significantly improves the performance while our method outperforms existing methods by 9.20 % and obtains SOTA results.
【10】 Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification
标题: 探索方言识别中的分布外检测基于能量的模型
作者:Yaqian Hao,Chenguang Hu,Yingying Gao,Shilei Zhang,Junlan Feng
链接:点击下载PDF文件
摘要:方言的多样性给在特定语言模式上训练的模型带来了挑战,使得它们在面对看不见的或分布外(OOD)数据时容易出错。本文提出了一种新的边缘增强联合能量模型(MEJEM),专门用于方言中的OOD检测。通过结合生成模型和能量裕度损失,我们的方法旨在提高方言识别系统的鲁棒性。此外,我们探索了两个OOD分数OOD方言检测,我们的研究结果最终证明,能源分数优于softmax分数。利用锐度感知最小化来优化联合模型的训练过程,我们通过最小化损失和锐度来增强模型的泛化能力。方言识别任务的实验验证了基于能量的模型的有效性,并提供了宝贵的见解,他们的表现。摘要:The diverse nature of dialects presents challenges for models trained on specific linguistic patterns, rendering them susceptible to errors when confronted with unseen or out-of-distribution (OOD) data. This study introduces a novel margin-enhanced joint energy model (MEJEM) tailored specifically for OOD detection in dialects. By integrating a generative model and the energy margin loss, our approach aims to enhance the robustness of dialect identification systems. Furthermore, we explore two OOD scores for OOD dialect detection, and our findings conclusively demonstrate that the energy score outperforms the softmax score. Leveraging Sharpness-Aware Minimization to optimize the training process of the joint model, we enhance model generalization by minimizing both loss and sharpness. Experiments conducted on dialect identification tasks validate the efficacy of Energy-Based Models and provide valuable insights into their performance.
【11】 SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
标题: SC-MoE:统一流媒体和非流媒体代码交换ASB的交换一致者专家混合
作者:Shuaishuai Ye,Shunfei Chen,Xinhui Hu,Xinkang Xu
备注:Accepted by InterSpeech 2024; 5 pages, 2 figures
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了一个基于Switch-Conformer的MoE系统,命名为SC-MoE,用于统一的流和非流代码切换(CS)自动语音识别(ASR),其中我们设计了一个由三个语言专家组成的流MoE层,分别对应于普通话,英语和空白,并配备了一个语言识别(LID)网络与连接时间分类(CTC)的损失作为路由器在编码器的SC-MoE,以实现真正的-时间流CS ASR系统。为了进一步利用嵌入在文本中的语言信息,我们还将MoE层纳入SC-MoE的解码器。此外,我们在编码器和解码器的每个MoE层中引入了路由器,并获得了更好的识别性能。实验结果表明,SC-MoE显着提高CS ASR性能与基线相当的计算效率。摘要:In this work, we propose a Switch-Conformer-based MoE system named SC-MoE for unified streaming and non-streaming code-switching (CS) automatic speech recognition (ASR), where we design a streaming MoE layer consisting of three language experts, which correspond to Mandarin, English, and blank, respectively, and equipped with a language identification (LID) network with a Connectionist Temporal Classification (CTC) loss as a router in the encoder of SC-MoE to achieve a real-time streaming CS ASR system. To further utilize the language information embedded in text, we also incorporate MoE layers into the decoder of SC-MoE. In addition, we introduce routers into every MoE layer of the encoder and the decoder and achieve better recognition performance. Experimental results show that the SC-MoE significantly improves CS ASR performances over baseline with comparable computational efficiency.
【12】 Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
标题: 通过学习单调对齐提高基于LLM的语音合成的鲁棒性
作者:Paarth Neekhara,Shehzeen Hussain,Subhankar Ghosh,Jason Li,Rafael Valle,Rohan Badlani,Boris Ginsburg
备注:Published as a conference paper at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:基于大语言模型(LLM)的文本到语音(TTS)系统在处理大型语音数据集和为新说话人生成自然语音方面表现出了卓越的能力。然而,基于LLM的TTS模型并不鲁棒,因为生成的输出可能包含重复的单词,丢失的单词和未对齐的语音(称为幻觉或注意力错误),特别是当文本包含多次出现的相同标记时。我们在编码器-解码器Transformer模型中研究了这些挑战,发现在这种模型中,某些交叉注意头在训练用于预测给定文本的语音标记时,隐式地学习文本和语音对齐。为了使对齐更加鲁棒,我们提出了利用CTC损失和注意力先验的技术,鼓励单调的交叉注意的文本标记。我们的引导注意力训练技术不引入任何新的可学习参数,并显着提高了基于LLM的TTS模型的鲁棒性。摘要:Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers. However, LLM-based TTS models are not robust as the generated output can contain repeating words, missing words and mis-aligned speech (referred to as hallucinations or attention errors), especially when the text contains multiple occurrences of the same token. We examine these challenges in an encoder-decoder transformer model and find that certain cross-attention heads in such models implicitly learn the text and speech alignment when trained for predicting speech tokens for a given text. To make the alignment more robust, we propose techniques utilizing CTC loss and attention priors that encourage monotonic cross-attention over the text tokens. Our guided attention training technique does not introduce any new learnable parameters and significantly improves robustness of LLM-based TTS models.
【13】 Sequential Editing for Lifelong Training of Speech Recognition Models
标题: 语音识别模型终身训练的序列编辑
作者:Devang Kulshreshtha,Saket Dingliwal,Brady Houston,Nikolaos Pappas,Srikanth Ronanki
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:自动语音识别(ASR)传统上假设已知域,但从新域添加数据会引起对与现有和新域上的再训练模型相关的计算效率低下的担忧。仅对新领域进行微调会带来灾难性遗忘(CF)的风险。为了解决这个问题,终身学习(LLL)算法已经提出了ASR。先前的研究已经探索了诸如弹性权重合并、知识蒸馏和重放等技术,所有这些技术都需要额外的参数或访问先前的域数据。我们提出顺序模型编辑作为一种新的方法,不断学习新的领域在ASR系统。与以前的方法不同,我们的方法不需要访问以前的数据集或引入额外的参数。我们的研究表明,在微调基线上,单词错误率降低了15%,并且在CommonVoice英语多口音数据集上的效率优于其他LLL技术。摘要:Automatic Speech Recognition (ASR) traditionally assumes known domains, but adding data from a new domain raises concerns about computational inefficiencies linked to retraining models on both existing and new domains. Fine-tuning solely on new domain risks Catastrophic Forgetting (CF). To address this, Lifelong Learning (LLL) algorithms have been proposed for ASR. Prior research has explored techniques such as Elastic Weight Consolidation, Knowledge Distillation, and Replay, all of which necessitate either additional parameters or access to prior domain data. We propose Sequential Model Editing as a novel method to continually learn new domains in ASR systems. Different than previous methods, our approach does not necessitate access to prior datasets or the introduction of extra parameters. Our study demonstrates up to 15% Word Error Rate Reduction (WERR) over fine-tuning baseline, and superior efficiency over other LLL techniques on CommonVoice English multi-accent dataset.
【14】 SonicSense: Object Perception from In-Hand Acoustic Vibration
标题: SonicSense:来自手内声振动的物体感知
作者:Jiaxun Liu,Boyuan Chen
备注:Our project website is at: this http URL
链接:点击下载PDF文件
摘要:我们介绍SonicSense,这是一种硬件和软件的整体设计,通过手内声学振动传感实现丰富的机器人对象感知。虽然以前的研究已经显示出声学传感用于物体感知的有希望的结果,但目前的解决方案仅限于具有简单几何形状和均匀材料的少数物体,单指传感以及对相同物体的混合训练和测试。SonicSense能够区分集装箱库存状态、异质材料预测、3D形状重建以及从83个真实物体中重新识别物体。我们的系统采用了一种简单但有效的启发式探索策略来与物体进行交互,以及基于端到端学习的算法来融合振动信号以推断物体属性。我们的框架强调了手声学振动传感在推进机器人触觉感知的重要性。摘要:We introduce SonicSense, a holistic design of hardware and software to enable rich robot object perception through in-hand acoustic vibration sensing. While previous studies have shown promising results with acoustic sensing for object perception, current solutions are constrained to a handful of objects with simple geometries and homogeneous materials, single-finger sensing, and mixing training and testing on the same objects. SonicSense enables container inventory status differentiation, heterogeneous material prediction, 3D shape reconstruction, and object re-identification from a diverse set of 83 real-world objects. Our system employs a simple but effective heuristic exploration policy to interact with the objects as well as end-to-end learning-based algorithms to fuse vibration signals to infer object properties. Our framework underscores the significance of in-hand acoustic vibration sensing in advancing robot tactile perception.
【15】 FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data
标题: FASA:一种灵活的自动语音对齐器,用于提取高质量对齐儿童语音数据
作者:Dancheng Liu,Jinjun Xiong
备注:4 pages, 1 figure
链接:点击下载PDF文件
摘要:近年来,通过采用深度神经网络(DNN)模型,成人语音的自动语音识别(ASR)取得了重大进展,但由于儿童语音的独特特征,儿童语音的改善仍然不尽如人意。由于缺乏高质量的对齐儿童语音,根据成人数据预训练的DNN模型在通过微调概括儿童语音时经常遇到困难。在生成数据集时,人类注释是不可扩展的,并且现有的强制对齐工具也不可用,因为它们对输入转录的质量做出了不切实际的假设。为了解决这些挑战,我们提出了一种新的强制对齐工具FASA,作为一种灵活且自动的语音对齐器,可以从许多现有的嘈杂儿童语音数据中提取高质量对齐的儿童语音数据。我们展示了它的使用CHILDES数据集,并表明,FASA可以提高数据质量的13.6$ times$超过人类注释。摘要:Automatic Speech Recognition (ASR) for adults' speeches has made significant progress by employing deep neural network (DNN) models recently, but improvement in children's speech is still unsatisfactory due to children's speech's distinct characteristics. DNN models pre-trained on adult data often struggle in generalizing children's speeches with fine tuning because of the lack of high-quality aligned children's speeches. When generating datasets, human annotations are not scalable, and existing forced-alignment tools are not usable as they make impractical assumptions about the quality of the input transcriptions. To address these challenges, we propose a new forced-alignment tool, FASA, as a flexible and automatic speech aligner to extract high-quality aligned children's speech data from many of the existing noisy children's speech data. We demonstrate its usage on the CHILDES dataset and show that FASA can improve data quality by 13.6$ times$ over human annotations.
【16】 Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
标题: 使用CNN、双向LSTM和ResNet的尼泊尔语自动语音识别
作者:Manish Dhakal,Arman Chhetri,Aman Kumar Gupta,Prabin Lamichhane,Suraj Pandey,Subarna Shakya
Journal-ref:2022 International Conference on Inventive Computation Technologies (ICICT), pp. 515-521
链接:点击下载PDF文件
摘要:本文提出了一种用于自动语音识别(ASR)的端到端深度学习模型,该模型将尼泊尔语语音转录为文本。该模型在OpenSLR(音频,文本)数据集上进行了训练和测试。大多数音频数据集在两端都有无声的间隙,这些间隙在数据集预处理期间被裁剪,以实现音频帧及其相应文本的更均匀映射。Mel频率倒谱系数(MFCC)用作音频特征以馈送到模型中。具有双向LSTM与ResNet和一维CNN配对的模型在迄今为止已经训练的所有模型(具有LSTM、GRU、CNN和ResNet变体的神经网络)中为该数据集产生最佳结果。该模型使用连接主义时间分类(CTC)函数进行训练过程中的损失计算,并使用CTC波束搜索解码来预测尼泊尔语文本中最有可能的字符序列。在测试数据集上,字符错误率(CER)为17.06%。源代码可在https: github.com manishdhakal ASR-Nepali-using-CNN-BiLSTM-ResNet上获得。摘要:This paper presents an end-to-end deep learning model for Automatic Speech Recognition (ASR) that transcribes Nepali speech to text. The model was trained and tested on the OpenSLR (audio, text) dataset. The majority of the audio dataset have silent gaps at both ends which are clipped during dataset preprocessing for a more uniform mapping of audio frames and their corresponding texts. Mel Frequency Cepstral Coefficients (MFCCs) are used as audio features to feed into the model. The model having Bidirectional LSTM paired with ResNet and one-dimensional CNN produces the best results for this dataset out of all the models (neural networks with variations of LSTM, GRU, CNN, and ResNet) that have been trained so far. This novel model uses Connectionist Temporal Classification (CTC) function for loss calculation during training and CTC beam search decoding for predicting characters as the most likely sequence of Nepali text. On the test dataset, the character error rate (CER) of 17.06 percent has been achieved. The source code is available at: https: github.com manishdhakal ASR-Nepali-using-CNN-BiLSTM-ResNet.
【17】 A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
标题: 基于vits 2的多扬声器多语言语音克隆系统for limmits 2024挑战
作者:Xiaopeng Wang,Yi Lu,Xin Qi,Zhiyong Wang,Yuankun Xie,Shuchen Shi,Ruibo Fu
链接:点击下载PDF文件
摘要:本文介绍了一个语音合成系统的LIMMITS'24挑战赛的发展,主要集中在轨道2。这项挑战的目标是建立一个具有语音克隆能力的多说话者、多语言的印度文语转换系统,涵盖七种印度语言,包括男性和女性。该系统使用挑战数据进行训练,并针对目标扬声器上的Few-Shot语音克隆进行微调。评估包括所有七种语言的单语言和跨语言合成,主观测试评估自然度和说话者相似性。我们的系统使用VITS2架构,增强了多语言ID和BERT模型,以提高上下文语言的理解。在轨道1中,不允许使用额外的数据,我们的模型实现了4.02的说话人相似度得分。在允许使用额外数据的音轨2中,它获得了4.17的说话者相似度得分。摘要:This paper presents the development of a speech synthesis system for the LIMMITS'24 Challenge, focusing primarily on Track 2. The objective of the challenge is to establish a multi-speaker, multi-lingual Indic Text-to-Speech system with voice cloning capabilities, covering seven Indian languages with both male and female speakers. The system was trained using challenge data and fine-tuned for few-shot voice cloning on target speakers. Evaluation included both mono-lingual and cross-lingual synthesis across all seven languages, with subjective tests assessing naturalness and speaker similarity. Our system uses the VITS2 architecture, augmented with a multi-lingual ID and a BERT model to enhance contextual language comprehension. In Track 1, where no additional data usage was permitted, our model achieved a Speaker Similarity score of 4.02. In Track 2, which allowed the use of extra data, it attained a Speaker Similarity score of 4.17.
【18】 Fish Tracking, Counting, and Behaviour Analysis in Digital Aquaculture: A Comprehensive Review
标题: 数字水族馆中的鱼类追踪、计数和行为分析:全面评论
作者:Meng Cui,Xubo Liu,Haohe Liu,Jinzheng Zhao,Daoliang Li,Wenwu Wang
链接:点击下载PDF文件
摘要:数字水产养殖利用先进的技术和数据驱动的方法,提供了比传统水产养殖实践更大的好处。鱼类跟踪、计数和行为分析是数字水产养殖的重要组成部分,对于优化生产效率、提高鱼类福利和改善资源管理至关重要。以往的审查侧重于单一模式,限制了它们全面应对这些任务中遇到的各种挑战的能力。本文对水产养殖数字化技术的现状进行了全面分析,包括基于视觉、基于声学和基于生物传感器的方法。我们研究这些方法的优点,局限性和应用,突出最近的进展,并确定关键的研究差距。缺乏全面的鱼类数据集和缺乏统一的评价标准,使得难以比较不同技术的性能,被认为是阻碍这一领域取得进展的主要障碍。为了克服当前的局限性,提高鱼类监测系统的准确性、鲁棒性和效率,我们探索了多模态数据融合和深度学习等新兴技术的潜力。此外,我们通过提供现有数据集的摘要来为该领域做出贡献,这些数据集可用于鱼类跟踪,计数和行为分析。概述了未来的研究方向,强调需要全面的数据集和评估标准,以促进技术之间的有意义的比较,并促进其在现实世界的水产养殖环境中的实际实施。摘要:Digital aquaculture leverages advanced technologies and data-driven methods, providing substantial benefits over traditional aquaculture practices. Fish tracking, counting, and behaviour analysis are crucial components of digital aquaculture, which are essential for optimizing production efficiency, enhancing fish welfare, and improving resource management. Previous reviews have focused on single modalities, limiting their ability to address the diverse challenges encountered in these tasks comprehensively. This review provides a comprehensive analysis of the current state of aquaculture digital technologies, including vision-based, acoustic-based, and biosensor-based methods. We examine the advantages, limitations, and applications of these methods, highlighting recent advancements and identifying critical research gaps. The scarcity of comprehensive fish datasets and the lack of unified evaluation standards, which make it difficult to compare the performance of different technologies, are identified as major obstacles hindering progress in this field. To overcome current limitations and improve the accuracy, robustness, and efficiency of fish monitoring systems, we explore the potential of emerging technologies such as multimodal data fusion and deep learning. Additionally, we contribute to the field by providing a summary of existing datasets available for fish tracking, counting, and behaviour analysis. Future research directions are outlined, emphasizing the need for comprehensive datasets and evaluation standards to facilitate meaningful comparisons between technologies and promote their practical implementation in real-world aquaculture settings.
机器翻译,仅供参考
