今日论文合集:cs.SD语音10篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】LastResort at SemEval-2024 Task 3: Exploring Multimodal Emotion Cause  Pair Extraction as Sequence Labelling Task

标题:LastResort在SemEval—2024任务3:探索多模态情绪原因对提取作为序列标签任务的提取

链接:https://arxiv.org/abs/2404.02088

作者:Suyash Vardhan Mathur,Akshett Rai Jindal,Hardik Mittal,Manish Shrivastava

摘要:对话是人类交流的最自然形式,其中每个话语都可以涵盖各种可能的情感。虽然对文本中的情感检测已经做了大量的工作,但对找到所述情感的原因,特别是在多模态设置中,做的工作相对较少。SemEval 2024引入了对话中的多模态情感原因分析任务,其目的是提取涉及多模态(文本,音频和视觉模态)的对话中个体话语中反映的情感以及作为情感原因的相应话语。在本文中,我们提出了将此任务作为话语标记和序列标记问题来处理的模型,并对这些模型进行了比较研究,包括使用不同编码器的基线,使用BiLSTM添加会话的上下文信息,最后添加CRF层以尝试更有效地对相邻话语之间的相互依赖关系进行建模。在该任务的官方排行榜中,我们的架构排名第8,在排行榜上获得了0.1759的F1分数。

摘要:Conversation is the most natural form of human communication, where each utterance can range over a variety of possible emotions. While significant work has been done towards the detection of emotions in text, relatively little work has been done towards finding the cause of the said emotions, especially in multimodal settings. SemEval 2024 introduces the task of Multimodal Emotion Cause Analysis in Conversations, which aims to extract emotions reflected in individual utterances in a conversation involving multiple modalities (textual, audio, and visual modalities) along with the corresponding utterances that were the cause for the emotion. In this paper, we propose models that tackle this task as an utterance labeling and a sequence labeling problem and perform a comparative study of these models, involving baselines using different encoders, using BiLSTM for adding contextual information of the conversation, and finally adding a CRF layer to try to model the inter-dependencies between adjacent utterances more effectively. In the official leaderboard for the task, our architecture was ranked 8th, achieving an F1-score of 0.1759 on the leaderboard.


【2】 SPMamba: State-space model is all you need in speech separation
标题:SPMamba:状态空间模型是语音分离中所需的全部
链接:https://arxiv.org/abs/2404.02063
作者:Kai Li,Guo Chen
备注:Technical Report. Work in progress. Code is available at this https URL
摘要:在语音分离中,基于CNN和Transformer的模型都表现出了强大的分离能力,在研究界引起了极大的关注。然而,基于CNN的方法对于长序列音频具有有限的建模能力,导致次优分离性能。相反,基于变换器的方法由于其高计算复杂度而在实际应用中受到限制。值得注意的是,在计算机视觉中,基于Mamba的方法因其强大的性能和降低的计算要求而闻名。在本文中,我们提出了一个网络架构的语音分离使用的状态空间模型,即SPMamba。我们采用TF-GridNet模型作为基础框架,并将其Transformer组件替换为双向Mamba模块,旨在捕获更广泛的上下文信息。我们的实验结果揭示了一个重要的作用,在性能方面的曼巴为基础的模型。SPMamba在基于Librispeech构建的数据集中展示了卓越的性能,与现有的分离模型相比具有显着的优势。值得注意的是,SPMamba实现了分离质量的大幅提高,与TF-GridNet相比,SI-SNRi增强了2.42 dB。SPMamba的源代码可在https://github.com/JusperLee/SPMamba上公开访问。
摘要:In speech separation, both CNN- and Transformer-based models have demonstrated robust separation capabilities, garnering significant attention within the research community. However, CNN-based methods have limited modelling capability for long-sequence audio, leading to suboptimal separation performance. Conversely, Transformer-based methods are limited in practical applications due to their high computational complexity. Notably, within computer vision, Mamba-based methods have been celebrated for their formidable performance and reduced computational requirements. In this paper, we propose a network architecture for speech separation using a state-space model, namely SPMamba. We adopt the TF-GridNet model as the foundational framework and substitute its Transformer component with a bidirectional Mamba module, aiming to capture a broader range of contextual information. Our experimental results reveal an important role in the performance aspects of Mamba-based models. SPMamba demonstrates superior performance with a significant advantage over existing separation models in a dataset built on Librispeech. Notably, SPMamba achieves a substantial improvement in separation quality, with a 2.42 dB enhancement in SI-SNRi compared to the TF-GridNet. The source code for SPMamba is publicly accessible at https://github.com/JusperLee/SPMamba .


【3】 Africa-Centric Self-Supervised Pre-Training for Multilingual Speech  Representation in a Sub-Saharan Context
标题:撒哈拉以南地区多语言语音表征的非洲中心自监督预训练
链接:https://arxiv.org/abs/2404.02000
作者:Antoine Caubrière,Elodie Gauthier
备注:To appear in AfricaNLP 2024
摘要:我们提出了第一个专门针对非洲语音训练的自监督多语言语音模型。该模型从撒哈拉以南非洲21种语言和方言的近60,000小时未标记的语音片段中学习。在FLEURS-102数据集的SSA子集上,我们基于HuBERT$_{base}$(0.09B)架构的方法对于ASR下游任务显示出具有竞争力的结果,与FLEURS基准中提出的w2 v-bert-51(0.6B)预训练模型相比,同时通过使用7倍的数据和6倍的参数更有效。此外,在LID下游任务的背景下,我们的方法比FLEURS基线准确度高出22%以上。
摘要:We present the first self-supervised multilingual speech model trained exclusively on African speech. The model learned from nearly 60 000 hours of unlabeled speech segments in 21 languages and dialects spoken in sub-Saharan Africa. On the SSA subset of the FLEURS-102 dataset, our approach based on a HuBERT$_{base}$ (0.09B) architecture shows competitive results, for ASR downstream task, compared to the w2v-bert-51 (0.6B) pre-trained model proposed in the FLEURS benchmark, while being more efficient by using 7x less data and 6x less parameters. Furthermore, in the context of a LID downstream task, our approach outperforms FLEURS baselines accuracy by over 22\%.


【4】 Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
标题:临床试验中的Zero-Shot多语言说话人验证
链接:https://arxiv.org/abs/2404.01981
作者:Ali Akram,Marija Stanojevic,Malikeh Ehghaghi,Jekaterina Novikova
摘要:由于临床试验涉及大量的临床医生、患者和数据收集环境,因此收集高质量的数据是一项重大挑战。在临床试验中,根据患者的语音数据对患者进行评估,以检测和监测认知和心理健康障碍。我们建议使用这些语音记录来验证入组患者的身份,并识别和排除试图在同一试验中多次入组的个人。由于临床研究通常在不同的国家进行,因此创建一个可以在不需要额外开发工作的情况下以不同语言执行说话人验证的系统势在必行。我们通过招募和测试讲英语,德语,丹麦语,西班牙语和阿拉伯语的语言障碍患者来评估预训练的TitaNet,ECAPA-TDNN和SpeakerNet模型。我们的研究结果表明,测试模型可以有效地推广到临床扬声器,欧洲语言的EER不到2.7%,阿拉伯语的EER为8.26%。这是为认知和心理健康临床试验开发更通用、更高效的说话人验证系统的重要一步,该系统可用于各种语言和方言,大大减少了开发多种语言说话人验证系统所需的工作量。我们还评估了语音任务和参与试验的扬声器数量如何影响性能,并表明语音任务的类型影响模型的性能。
摘要:Due to the substantial number of clinicians, patients, and data collection environments involved in clinical trials, gathering data of superior quality poses a significant challenge. In clinical trials, patients are assessed based on their speech data to detect and monitor cognitive and mental health disorders. We propose using these speech recordings to verify the identities of enrolled patients and identify and exclude the individuals who try to enroll multiple times in the same trial. Since clinical studies are often conducted across different countries, creating a system that can perform speaker verification in diverse languages without additional development effort is imperative. We evaluate pre-trained TitaNet, ECAPA-TDNN, and SpeakerNet models by enrolling and testing with speech-impaired patients speaking English, German, Danish, Spanish, and Arabic languages. Our results demonstrate that tested models can effectively generalize to clinical speakers, with less than 2.7% EER for European Languages and 8.26% EER for Arabic. This represents a significant step in developing more versatile and efficient speaker verification systems for cognitive and mental health clinical trials that can be used across a wide range of languages and dialects, substantially reducing the effort required to develop speaker verification systems for multiple languages. We also evaluate how speech tasks and number of speakers involved in the trial influence the performance and show that the type of speech tasks impacts the model performance.

【5】 T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
标题:T—VSL:文本引导的混合视觉声源定位
链接:https://arxiv.org/abs/2404.01751
作者:Tanvir Mahmud,Yapeng Tian,Diana Marculescu
备注:Tech report. Accepted in CVPR-2024
摘要:视觉声源定位在识别视频中每个声源的语义区域方面提出了重大挑战。现有的自监督和弱监督源定位方法难以准确区分每个发声对象的语义区域,特别是在多源混合中。这些方法通常依赖于视听对应作为指导,这可能导致在复杂的多源定位场景中性能大幅下降。在训练过程中无法获得多源混合中的单个源声音,这加剧了学习有效的视听对应以进行定位的难度。为了解决这一限制,在本文中,我们建议使用三模态联合嵌入模型(例如,AudioCLIP)来解开多源混合中的语义视听源对应。我们的框架,被称为T-VSL,开始预测类的声音实体的混合物。随后,每个发声源的文本表示被用作指导,以从多源混合物中解开细粒度的视听源对应关系,利用三模态AudioCLIP嵌入。这种方法使我们的框架,以处理灵活的来源数量,并表现出有前途的zero-shot transferability看不见的类在测试时间。在MUSIC,VGGSound和VGGSound-Instruments数据集上进行的大量实验表明,与最先进的方法相比,性能有了显着提高。
摘要:Visual sound source localization poses a significant challenge in identifying the semantic region of each sounding source within a video. Existing self-supervised and weakly supervised source localization methods struggle to accurately distinguish the semantic regions of each sounding object, particularly in multi-source mixtures. These methods often rely on audio-visual correspondence as guidance, which can lead to substantial performance drops in complex multi-source localization scenarios. The lack of access to individual source sounds in multi-source mixtures during training exacerbates the difficulty of learning effective audio-visual correspondence for localization. To address this limitation, in this paper, we propose incorporating the text modality as an intermediate feature guide using tri-modal joint embedding models (e.g., AudioCLIP) to disentangle the semantic audio-visual source correspondence in multi-source mixtures. Our framework, dubbed T-VSL, begins by predicting the class of sounding entities in mixtures. Subsequently, the textual representation of each sounding source is employed as guidance to disentangle fine-grained audio-visual source correspondence from multi-source mixtures, leveraging the tri-modal AudioCLIP embedding. This approach enables our framework to handle a flexible number of sources and exhibits promising zero-shot transferability to unseen classes during test time. Extensive experiments conducted on the MUSIC, VGGSound, and VGGSound-Instruments datasets demonstrate significant performance improvements over state-of-the-art methods.


【6】 Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
标题:基于双模语义相似度的弱监督音频分离
链接:https://arxiv.org/abs/2404.01740
作者:Tanvir Mahmud,Saeed Amizadeh,Kazuhito Koishida,Diana Marculescu
备注:Tech report. Accepted in ICLR-2024
摘要:在训练期间不访问单源声音数据的情况下,多源音频混合中的有条件声音分离是一个长期存在的挑战。现有的基于混合和分离的方法遭受显着的性能下降与多源训练混合物由于缺乏监督信号的单源分离情况下,在训练过程中。然而,在语言条件音频分离的情况下,我们确实可以访问训练数据中每个音频混合的相应文本描述,这些文本描述可以被视为语言模态中音频样本的(粗略)表示。为此,在本文中,我们提出了一个通用的双模态分离框架,它可以增强现有的无监督框架,以分离目标模态中的单源信号(即,音频)在调节模态中使用容易分离的对应信号(即,语言),而无需在训练期间访问目标模态中的单源样本。我们根据经验表明,如果我们可以访问两种模态之间的预训练联合嵌入模型(即,CLAP)。此外,我们建议将我们的框架纳入两个基本方案,以提高分离性能。首先,我们证明了我们提出的方法通过减少训练样本和测试样本之间的分布偏移,显着提高了纯无监督基线的性能。特别是,我们证明了我们的框架在信号失真比(SDR)方面可以实现71%的提升,达到97.5%的监督学习性能。其次,我们表明,如果我们通过我们提出的弱监督框架来增强监督学习,我们可以将监督学习本身的性能进一步提高17%,从而为音频分离提供强大的半监督框架。
摘要:Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop with multi-source training mixtures due to the lack of supervision signal for single source separation cases during training. However, in the case of language-conditional audio separation, we do have access to corresponding text descriptions for each audio mixture in our training data, which can be seen as (rough) representations of the audio samples in the language modality. To this end, in this paper, we propose a generic bi-modal separation framework which can enhance the existing unsupervised frameworks to separate single-source signals in a target modality (i.e., audio) using the easily separable corresponding signals in the conditioning modality (i.e., language), without having access to single-source samples in the target modality during training. We empirically show that this is well within reach if we have access to a pretrained joint embedding model between the two modalities (i.e., CLAP). Furthermore, we propose to incorporate our framework into two fundamental scenarios to enhance separation performance. First, we show that our proposed methodology significantly improves the performance of purely unsupervised baselines by reducing the distribution shift between training and test samples. In particular, we show that our framework can achieve 71% boost in terms of Signal-to-Distortion Ratio (SDR) over the baseline, reaching 97.5% of the supervised learning performance. Second, we show that we can further improve the performance of the supervised learning itself by 17% if we augment it by our proposed weakly-supervised framework, that enables a powerful semi-supervised framework for audio separation.


【7】 Voice EHR: Introducing Multimodal Audio Data for Health
标题:语音EHR:引入多模态音频数据
链接:https://arxiv.org/abs/2404.01620
作者:James Anibal,Hannah Huth,Ming Li,Lindsey Hazen,Yen Minh Lam,Nguyen Thi Thu Hang,Michael Kleinman,Shelley Ost,Christopher Jackson,Laura Sprabery,Cheran Elangovan,Balaji Krishnaiah,Lee Akst,Ioan Lina,Iqbal Elyazar,Lenny Ekwati,Stefan Jansen,Richard Nduwayezu,Charisse Garcia,Jeffrey Plum,Jacqueline Brenner,Miranda Song,Emily Ricotta,David Clifton,C. Louise Thwaites,Yael Bensoussan,Bradford Wood
备注:18 pages, 2 figures, 7 tables
摘要:在音频数据上训练的大型人工智能模型可能有潜力快速对患者进行分类,增强医疗决策,并通过早期检测来改善结果。现有技术依赖于有限的数据集,使用高收入英语国家昂贵的记录设备。这对资源受限、高容量环境中的部署提出了挑战,在这些环境中,音频数据可能会产生深远的影响。本报告介绍了一种新的数据类型和相应的收集系统,该系统仅使用移动/Web应用程序通过引导性问题捕获健康数据。该应用最终产生音频电子健康记录(语音EHR),其可以包含来自常规语音/呼吸特征、语音模式和具有语义含义的语言的复杂的健康生物标志物-补偿单峰临床数据集的典型限制。本报告介绍了一个全球合作伙伴联盟,介绍了用于数据收集的应用程序,并展示了信息语音EHR在推进音频AI的可扩展性和多样性方面的潜力。
摘要:Large AI models trained on audio data may have the potential to rapidly classify patients, enhancing medical decision-making and potentially improving outcomes through early detection. Existing technologies depend on limited datasets using expensive recording equipment in high-income, English-speaking countries. This challenges deployment in resource-constrained, high-volume settings where audio data may have a profound impact. This report introduces a novel data type and a corresponding collection system that captures health data through guided questions using only a mobile/web application. This application ultimately results in an audio electronic health record (voice EHR) which may contain complex biomarkers of health from conventional voice/respiratory features, speech patterns, and language with semantic meaning - compensating for the typical limitations of unimodal clinical datasets. This report introduces a consortium of partners for global work, presents the application used for data collection, and showcases the potential of informative voice EHR to advance the scalability and diversity of audio AI.

【8】 Transforming LLMs into Cross-modal and Cross-lingual RetrievalSystems
标题:将法学硕士转化为跨模式和跨语言检索系统
链接:https://arxiv.org/abs/2404.01616
作者:Frank Palma Gomez,Ramon Sanabria,Yun-hsuan Sung,Daniel Cer,Siddharth Dalmia,Gustavo Hernandez Abrego
摘要:大型语言模型(LLM)是在纯文本数据上训练的,远远超出了具有配对语音和文本数据的语言。同时,基于双编码器(DE)的检索系统将查询和文档投影到同一个嵌入空间中,并在检索和双文本挖掘方面取得了成功。为了在多种语言中匹配语音和文本,我们建议使用LLM来初始化多模态DE检索系统。与传统方法不同,我们的系统在LLM预训练期间不需要语音数据,并且可以利用LLM的多语言文本理解功能来匹配检索训练期间看不到的语言中的语音和文本。我们的多模态基于LLM的检索系统能够匹配102种语言的语音和文本,尽管只训练了21种语言。我们的系统比以前的系统在所有102种语言上都有更好的表现。我们在这些语言中平均实现了10%的Recall@1绝对改进。此外,我们的模型演示了跨语言语音和文本匹配,这是进一步增强了现成的机器翻译数据。
摘要:Large language models (LLMs) are trained on text-only data that go far beyond the languages with paired speech and text data. At the same time, Dual Encoder (DE) based retrieval systems project queries and documents into the same embedding space and have demonstrated their success in retrieval and bi-text mining. To match speech and text in many languages, we propose using LLMs to initialize multi-modal DE retrieval systems. Unlike traditional methods, our system doesn't require speech data during LLM pre-training and can exploit LLM's multilingual text understanding capabilities to match speech and text in languages unseen during retrieval training. Our multi-modal LLM-based retrieval system is capable of matching speech and text in 102 languages despite only training on 21 languages. Our system outperforms previous systems trained explicitly on all 102 languages. We achieve a 10% absolute improvement in Recall@1 averaged across these languages. Additionally, our model demonstrates cross-lingual speech and text matching, which is further enhanced by readily available machine translation data.

【9】 Audio Simulation for Sound Source Localization in Virtual Evironment
标题:虚拟环境中声源定位的音频仿真
链接:https://arxiv.org/abs/2404.01611
作者:Yi Di Yuan,Swee Liang Wong,Jonathan Pan
备注:2024 IEEE World Forum on Public Safety Technology
摘要:信号缺失环境下的非视距定位是一个具有挑战性的问题。由于混响性质,这种主要室内场景中的声学方法遇到困难。在这项研究中,我们的目标是通过利用物理接地的声音传播模拟和机器学习方法将声源定位到虚拟环境中的特定位置。该过程试图克服数据不足的问题,以将声源定位到其发生的位置,特别是在事后定位中。我们使用音频Transformer频谱图方法实现0.786+/- 0.0136 F1分数。
摘要:Non-line-of-sight localization in signal-deprived environments is a challenging yet pertinent problem. Acoustic methods in such predominantly indoor scenarios encounter difficulty due to the reverberant nature. In this study, we aim to locate sound sources to specific locations within a virtual environment by leveraging physically grounded sound propagation simulations and machine learning methods. This process attempts to overcome the issue of data insufficiency to localize sound sources to their location of occurrence especially in post-event localization. We achieve 0.786+/- 0.0136 F1-score using an audio transformer spectrogram approach.

【10】 Transfer Learning from Whisper for Microscopic Intelligibility  Prediction
标题:基于Whisper的迁移学习在微观可懂度预测中的应用
链接:https://arxiv.org/abs/2404.01737
作者:Paul Best,Santiago Cuervo,Ricard Marxer
摘要:宏观可懂度模型预测预期的人类单词错误率为一个给定的语音噪声刺激。相比之下,微观可懂度模型旨在对听者的感知进行细粒度的预测,例如预测语音或词汇反应。最先进的宏观模型使用来自大规模深度学习模型的迁移学习进行语音处理,而这种方法很少用于微观建模。在本文中,我们研究了Whisper的迁移学习的使用,Whisper是一种最先进的自动语音识别深度学习模型,用于词汇响应水平的微观可懂度预测。即使在zero-shot设置中,我们的方法也优于所考虑的基线,并且当微调以预测听众的响应时,产生高达66%的相对改进。我们的研究结果展示了基于大规模深度学习的微观可懂度预测方法的前景。摘要:Macroscopic intelligibility models predict the expected human word-error-rate for a given speech-in-noise stimulus. In contrast, microscopic intelligibility models aim to make fine-grained predictions about listeners' perception, e.g. predicting phonetic or lexical responses. State-of-the-art macroscopic models use transfer learning from large scale deep learning models for speech processing, whereas such methods have rarely been used for microscopic modeling. In this paper, we study the use of transfer learning from Whisper, a state-of-the-art deep learning model for automatic speech recognition, for microscopic intelligibility prediction at the level of lexical responses. Our method outperforms the considered baselines, even in a zero-shot setup, and yields a relative improvement of up to 66\% when fine-tuned to predict listeners' responses. Our results showcase the promise of large scale deep learning based methods for microscopic intelligibility prediction.

eess.AS音频处理
【1】 Transfer Learning from Whisper for Microscopic Intelligibility  Prediction
标题:基于Whisper的迁移学习在微观可懂度预测中的应用
链接:https://arxiv.org/abs/2404.01737
作者:Paul Best,Santiago Cuervo,Ricard Marxer
摘要:宏观可懂度模型预测预期的人类单词错误率为一个给定的语音噪声刺激。相比之下,微观可懂度模型旨在对听者的感知进行细粒度的预测,例如预测语音或词汇反应。最先进的宏观模型使用来自大规模深度学习模型的迁移学习进行语音处理,而这种方法很少用于微观建模。在本文中,我们研究了Whisper的迁移学习的使用,Whisper是一种最先进的自动语音识别深度学习模型,用于词汇响应水平的微观可懂度预测。即使在zero-shot设置中,我们的方法也优于所考虑的基线,并且当微调以预测听众的响应时,产生高达66%的相对改进。我们的研究结果展示了基于大规模深度学习的微观可懂度预测方法的前景。
摘要:Macroscopic intelligibility models predict the expected human word-error-rate for a given speech-in-noise stimulus. In contrast, microscopic intelligibility models aim to make fine-grained predictions about listeners' perception, e.g. predicting phonetic or lexical responses. State-of-the-art macroscopic models use transfer learning from large scale deep learning models for speech processing, whereas such methods have rarely been used for microscopic modeling. In this paper, we study the use of transfer learning from Whisper, a state-of-the-art deep learning model for automatic speech recognition, for microscopic intelligibility prediction at the level of lexical responses. Our method outperforms the considered baselines, even in a zero-shot setup, and yields a relative improvement of up to 66\% when fine-tuned to predict listeners' responses. Our results showcase the promise of large scale deep learning based methods for microscopic intelligibility prediction.


【2】 Effective internal language model training and fusion for factorized  transducer mode
l标题:基于因子分解传感器模型的有效内部语言模型训练与融合
链接:https://arxiv.org/abs/2404.01716
作者:Jinxi Guo,Niko Moritz,Yingyi Ma,Frank Seide,Chunyang Wu,Jay Mahadeokar,Ozlem Kalinli,Christian Fuegen,Mike Seltzer
备注:Accepted to ICASSP 2024
摘要:神经传感器的内部语言模型(internal language model,ILM)已经得到了广泛的研究。在大多数先前的工作中,它主要用于估计ILM分数,随后在推理过程中被减去,以促进与外部语言模型的集成。最近,已经提出了各种因子分解的换能器模型,其明确地包含用于非空白标记预测的独立内部语言模型。然而,即使采用因子分解的换能器模型,与浅融合相比,也观察到有限的改善。在本文中,我们提出了一种新的ILM训练和解码策略的因式分解传感器模型,有效地结合了空白,声学和ILM分数。我们的实验表明,在LibriSpeech数据集上使用经过良好训练的ILM和所提出的解码策略时,标准解码方法的相对改进为17%。此外,当与外部LM融合增强的强RNN-T基线相比时,所提出的模型在一般集合上产生了5.5%的相对改进,并且对于稀有词减少了8.9%的WER。所提出的模型可以在不依赖外部语言模型的情况下实现卓越的性能,从而使其在生产用例中非常高效。为了进一步提高性能,我们提出了一种新的和内存效率的ILM融合意识的最小字错误率(MWER)的训练方法,提高ILM集成显着。
摘要:The internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracted during inference to facilitate improved integration with external language models. Recently, various of factorized transducer models have been proposed, which explicitly embrace a standalone internal language model for non-blank token prediction. However, even with the adoption of factorized transducer models, limited improvement has been observed compared to shallow fusion. In this paper, we propose a novel ILM training and decoding strategy for factorized transducer models, which effectively combines the blank, acoustic and ILM scores. Our experiments show a 17% relative improvement over the standard decoding method when utilizing a well-trained ILM and the proposed decoding strategy on LibriSpeech datasets. Furthermore, when compared to a strong RNN-T baseline enhanced with external LM fusion, the proposed model yields a 5.5% relative improvement on general-sets and an 8.9% WER reduction for rare words. The proposed model can achieve superior performance without relying on external language models, rendering it highly efficient for production use-cases. To further improve the performance, we propose a novel and memory-efficient ILM-fusion-aware minimum word error rate (MWER) training method which improves ILM integration significantly.

【3】 LastResort at SemEval-2024 Task 3: Exploring Multimodal Emotion Cause  Pair Extraction as Sequence Labelling Task
标题:LastResort在SemEval—2024任务3:探索多模态情绪原因对提取作为序列标签任务的提取
链接:https://arxiv.org/abs/2404.02088
作者:Suyash Vardhan Mathur,Akshett Rai Jindal,Hardik Mittal,Manish Shrivastava
摘要:对话是人类交流的最自然形式,其中每个话语都可以涵盖各种可能的情感。虽然对文本中的情感检测已经做了大量的工作,但对找到所述情感的原因,特别是在多模态设置中,做的工作相对较少。SemEval 2024引入了对话中的多模态情感原因分析任务,其目的是提取涉及多模态(文本,音频和视觉模态)的对话中个体话语中反映的情感以及作为情感原因的相应话语。在本文中,我们提出了将此任务作为话语标记和序列标记问题来处理的模型,并对这些模型进行了比较研究,包括使用不同编码器的基线,使用BiLSTM添加会话的上下文信息,最后添加CRF层以尝试更有效地对相邻话语之间的相互依赖关系进行建模。在该任务的官方排行榜中,我们的架构排名第8,在排行榜上获得了0.1759的F1分数。
摘要:Conversation is the most natural form of human communication, where each utterance can range over a variety of possible emotions. While significant work has been done towards the detection of emotions in text, relatively little work has been done towards finding the cause of the said emotions, especially in multimodal settings. SemEval 2024 introduces the task of Multimodal Emotion Cause Analysis in Conversations, which aims to extract emotions reflected in individual utterances in a conversation involving multiple modalities (textual, audio, and visual modalities) along with the corresponding utterances that were the cause for the emotion. In this paper, we propose models that tackle this task as an utterance labeling and a sequence labeling problem and perform a comparative study of these models, involving baselines using different encoders, using BiLSTM for adding contextual information of the conversation, and finally adding a CRF layer to try to model the inter-dependencies between adjacent utterances more effectively. In the official leaderboard for the task, our architecture was ranked 8th, achieving an F1-score of 0.1759 on the leaderboard.


【4】 SPMamba: State-space model is all you need in speech separation
标题:SPMamba:状态空间模型是语音分离中所需的全部
链接:https://arxiv.org/abs/2404.02063
作者:Kai Li,Guo Chen
备注:Technical Report. Work in progress. Code is available at this https URL
摘要:在语音分离中,基于CNN和Transformer的模型都表现出了强大的分离能力,在研究界引起了极大的关注。然而,基于CNN的方法对于长序列音频具有有限的建模能力,导致次优分离性能。相反,基于变换器的方法由于其高计算复杂度而在实际应用中受到限制。值得注意的是,在计算机视觉中,基于Mamba的方法因其强大的性能和降低的计算要求而闻名。在本文中,我们提出了一个网络架构的语音分离使用的状态空间模型,即SPMamba。我们采用TF-GridNet模型作为基础框架,并将其Transformer组件替换为双向Mamba模块,旨在捕获更广泛的上下文信息。我们的实验结果揭示了一个重要的作用,在性能方面的曼巴为基础的模型。SPMamba在基于Librispeech构建的数据集中展示了卓越的性能,与现有的分离模型相比具有显着的优势。值得注意的是,SPMamba实现了分离质量的大幅提高,与TF-GridNet相比,SI-SNRi增强了2.42 dB。SPMamba的源代码可在https://github.com/JusperLee/SPMamba上公开访问。
摘要:In speech separation, both CNN- and Transformer-based models have demonstrated robust separation capabilities, garnering significant attention within the research community. However, CNN-based methods have limited modelling capability for long-sequence audio, leading to suboptimal separation performance. Conversely, Transformer-based methods are limited in practical applications due to their high computational complexity. Notably, within computer vision, Mamba-based methods have been celebrated for their formidable performance and reduced computational requirements. In this paper, we propose a network architecture for speech separation using a state-space model, namely SPMamba. We adopt the TF-GridNet model as the foundational framework and substitute its Transformer component with a bidirectional Mamba module, aiming to capture a broader range of contextual information. Our experimental results reveal an important role in the performance aspects of Mamba-based models. SPMamba demonstrates superior performance with a significant advantage over existing separation models in a dataset built on Librispeech. Notably, SPMamba achieves a substantial improvement in separation quality, with a 2.42 dB enhancement in SI-SNRi compared to the TF-GridNet. The source code for SPMamba is publicly accessible at https://github.com/JusperLee/SPMamba .

【5】 Africa-Centric Self-Supervised Pre-Training for Multilingual Speech  Representation in a Sub-Saharan Context
标题:撒哈拉以南地区多语言语音表征的非洲中心自监督预训练
链接:https://arxiv.org/abs/2404.02000
作者:Antoine Caubrière,Elodie Gauthier
备注:To appear in AfricaNLP 2024
摘要:我们提出了第一个专门针对非洲语音训练的自监督多语言语音模型。该模型从撒哈拉以南非洲21种语言和方言的近60,000小时未标记的语音片段中学习。在FLEURS-102数据集的SSA子集上,我们基于HuBERT$_{base}$(0.09B)架构的方法对于ASR下游任务显示出具有竞争力的结果,与FLEURS基准中提出的w2 v-bert-51(0.6B)预训练模型相比,同时通过使用7倍的数据和6倍的参数更有效。此外,在LID下游任务的背景下,我们的方法比FLEURS基线准确度高出22%以上。
摘要:We present the first self-supervised multilingual speech model trained exclusively on African speech. The model learned from nearly 60 000 hours of unlabeled speech segments in 21 languages and dialects spoken in sub-Saharan Africa. On the SSA subset of the FLEURS-102 dataset, our approach based on a HuBERT$_{base}$ (0.09B) architecture shows competitive results, for ASR downstream task, compared to the w2v-bert-51 (0.6B) pre-trained model proposed in the FLEURS benchmark, while being more efficient by using 7x less data and 6x less parameters. Furthermore, in the context of a LID downstream task, our approach outperforms FLEURS baselines accuracy by over 22\%.


【6】 Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
标题:临床试验中的Zero-Shot多语言说话人验证
链接:https://arxiv.org/abs/2404.01981
作者:Ali Akram,Marija Stanojevic,Malikeh Ehghaghi,Jekaterina Novikova
摘要:由于临床试验涉及大量的临床医生、患者和数据收集环境,因此收集高质量的数据是一项重大挑战。在临床试验中,根据患者的语音数据对患者进行评估,以检测和监测认知和心理健康障碍。我们建议使用这些语音记录来验证入组患者的身份,并识别和排除试图在同一试验中多次入组的个人。由于临床研究通常在不同的国家进行,因此创建一个可以在不需要额外开发工作的情况下以不同语言执行说话人验证的系统势在必行。我们通过招募和测试讲英语,德语,丹麦语,西班牙语和阿拉伯语的语言障碍患者来评估预训练的TitaNet,ECAPA-TDNN和SpeakerNet模型。我们的研究结果表明,测试模型可以有效地推广到临床扬声器,欧洲语言的EER不到2.7%,阿拉伯语的EER为8.26%。这是为认知和心理健康临床试验开发更通用、更高效的说话人验证系统的重要一步,该系统可用于各种语言和方言,大大减少了开发多种语言说话人验证系统所需的工作量。我们还评估了语音任务和参与试验的扬声器数量如何影响性能,并表明语音任务的类型影响模型的性能。
摘要:Due to the substantial number of clinicians, patients, and data collection environments involved in clinical trials, gathering data of superior quality poses a significant challenge. In clinical trials, patients are assessed based on their speech data to detect and monitor cognitive and mental health disorders. We propose using these speech recordings to verify the identities of enrolled patients and identify and exclude the individuals who try to enroll multiple times in the same trial. Since clinical studies are often conducted across different countries, creating a system that can perform speaker verification in diverse languages without additional development effort is imperative. We evaluate pre-trained TitaNet, ECAPA-TDNN, and SpeakerNet models by enrolling and testing with speech-impaired patients speaking English, German, Danish, Spanish, and Arabic languages. Our results demonstrate that tested models can effectively generalize to clinical speakers, with less than 2.7% EER for European Languages and 8.26% EER for Arabic. This represents a significant step in developing more versatile and efficient speaker verification systems for cognitive and mental health clinical trials that can be used across a wide range of languages and dialects, substantially reducing the effort required to develop speaker verification systems for multiple languages. We also evaluate how speech tasks and number of speakers involved in the trial influence the performance and show that the type of speech tasks impacts the model performance.


【7】 T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
标题:T—VSL:文本引导的混合视觉声源定位
链接:https://arxiv.org/abs/2404.01751
作者:Tanvir Mahmud,Yapeng Tian,Diana Marculescu
备注:Tech report. Accepted in CVPR-2024
摘要:视觉声源定位在识别视频中每个声源的语义区域方面提出了重大挑战。现有的自监督和弱监督源定位方法难以准确区分每个发声对象的语义区域,特别是在多源混合中。这些方法通常依赖于视听对应作为指导,这可能导致在复杂的多源定位场景中性能大幅下降。在训练过程中无法获得多源混合中的单个源声音,这加剧了学习有效的视听对应以进行定位的难度。为了解决这一限制,在本文中,我们建议使用三模态联合嵌入模型(例如,AudioCLIP)来解开多源混合中的语义视听源对应。我们的框架,被称为T-VSL,开始预测类的声音实体的混合物。随后,每个发声源的文本表示被用作指导,以从多源混合物中解开细粒度的视听源对应关系,利用三模态AudioCLIP嵌入。这种方法使我们的框架,以处理灵活的来源数量,并表现出有前途的zero-shot transferability看不见的类在测试时间。在MUSIC,VGGSound和VGGSound-Instruments数据集上进行的大量实验表明,与最先进的方法相比,性能有了显着提高。
摘要:Visual sound source localization poses a significant challenge in identifying the semantic region of each sounding source within a video. Existing self-supervised and weakly supervised source localization methods struggle to accurately distinguish the semantic regions of each sounding object, particularly in multi-source mixtures. These methods often rely on audio-visual correspondence as guidance, which can lead to substantial performance drops in complex multi-source localization scenarios. The lack of access to individual source sounds in multi-source mixtures during training exacerbates the difficulty of learning effective audio-visual correspondence for localization. To address this limitation, in this paper, we propose incorporating the text modality as an intermediate feature guide using tri-modal joint embedding models (e.g., AudioCLIP) to disentangle the semantic audio-visual source correspondence in multi-source mixtures. Our framework, dubbed T-VSL, begins by predicting the class of sounding entities in mixtures. Subsequently, the textual representation of each sounding source is employed as guidance to disentangle fine-grained audio-visual source correspondence from multi-source mixtures, leveraging the tri-modal AudioCLIP embedding. This approach enables our framework to handle a flexible number of sources and exhibits promising zero-shot transferability to unseen classes during test time. Extensive experiments conducted on the MUSIC, VGGSound, and VGGSound-Instruments datasets demonstrate significant performance improvements over state-of-the-art methods.


【8】 Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
标题:基于双模语义相似度的弱监督音频分离
链接:https://arxiv.org/abs/2404.01740
作者:Tanvir Mahmud,Saeed Amizadeh,Kazuhito Koishida,Diana Marculescu
备注:Tech report. Accepted in ICLR-2024
摘要:在训练期间不访问单源声音数据的情况下,多源音频混合中的有条件声音分离是一个长期存在的挑战。现有的基于混合和分离的方法遭受显着的性能下降与多源训练混合物由于缺乏监督信号的单源分离情况下,在训练过程中。然而,在语言条件音频分离的情况下,我们确实可以访问训练数据中每个音频混合的相应文本描述,这些文本描述可以被视为语言模态中音频样本的(粗略)表示。为此,在本文中,我们提出了一个通用的双模态分离框架,它可以增强现有的无监督框架,以分离目标模态中的单源信号(即,音频)在调节模态中使用容易分离的对应信号(即,语言),而无需在训练期间访问目标模态中的单源样本。我们根据经验表明,如果我们可以访问两种模态之间的预训练联合嵌入模型(即,CLAP)。此外,我们建议将我们的框架纳入两个基本方案,以提高分离性能。首先,我们证明了我们提出的方法通过减少训练样本和测试样本之间的分布偏移,显着提高了纯无监督基线的性能。特别是,我们证明了我们的框架在信号失真比(SDR)方面可以实现71%的提升,达到97.5%的监督学习性能。其次,我们表明,如果我们通过我们提出的弱监督框架来增强监督学习,我们可以将监督学习本身的性能进一步提高17%,从而为音频分离提供强大的半监督框架。
摘要:Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop with multi-source training mixtures due to the lack of supervision signal for single source separation cases during training. However, in the case of language-conditional audio separation, we do have access to corresponding text descriptions for each audio mixture in our training data, which can be seen as (rough) representations of the audio samples in the language modality. To this end, in this paper, we propose a generic bi-modal separation framework which can enhance the existing unsupervised frameworks to separate single-source signals in a target modality (i.e., audio) using the easily separable corresponding signals in the conditioning modality (i.e., language), without having access to single-source samples in the target modality during training. We empirically show that this is well within reach if we have access to a pretrained joint embedding model between the two modalities (i.e., CLAP). Furthermore, we propose to incorporate our framework into two fundamental scenarios to enhance separation performance. First, we show that our proposed methodology significantly improves the performance of purely unsupervised baselines by reducing the distribution shift between training and test samples. In particular, we show that our framework can achieve 71% boost in terms of Signal-to-Distortion Ratio (SDR) over the baseline, reaching 97.5% of the supervised learning performance. Second, we show that we can further improve the performance of the supervised learning itself by 17% if we augment it by our proposed weakly-supervised framework, that enables a powerful semi-supervised framework for audio separation.


【9】 Release of Pre-Trained Models for the Japanese Language
标题:日语预训练模型的发布
链接:https://arxiv.org/abs/2404.01657
作者:Kei Sawada,Tianyu Zhao,Makoto Shing,Kentaro Mitsui,Akio Kaga,Yukiya Hono,Toshiaki Wakatsuki,Koh Mitsuda
备注:9 pages, 1 figure, 5 tables, accepted for LREC-COLING 2024. Models are publicly available at this https URL
摘要:人工智能民主化旨在创造一个普通人可以利用人工智能技术的世界。为了实现这一目标,许多研究机构都试图向公众提供他们的成果。特别是,在大规模数据上训练的大型预训练模型显示出前所未有的潜力,它们的发布产生了重大影响。然而,大多数已发布的模型专门用于英语,因此,非英语社区的人工智能民主化明显滞后。为了缩小AI访问的差距,我们发布了生成预训练Transformer(GPT),对比语言和图像预训练(CLIP),稳定扩散和隐藏单元双向编码器表示来自Transformers(HuBERT)日语预训练。通过提供这些模型,用户可以自由地与符合日本文化价值观的人工智能进行交互,并确保日本文化的身份,从而增强人工智能的民主化。此外,实验表明,专门针对日语的预训练模型可以有效地在日语任务中实现高性能。
摘要:AI democratization aims to create a world in which the average person can utilize AI techniques. To achieve this goal, numerous research institutes have attempted to make their results accessible to the public. In particular, large pre-trained models trained on large-scale data have shown unprecedented potential, and their release has had a significant impact. However, most of the released models specialize in the English language, and thus, AI democratization in non-English-speaking communities is lagging significantly. To reduce this gap in AI access, we released Generative Pre-trained Transformer (GPT), Contrastive Language and Image Pre-training (CLIP), Stable Diffusion, and Hidden-unit Bidirectional Encoder Representations from Transformers (HuBERT) pre-trained in Japanese. By providing these models, users can freely interface with AI that aligns with Japanese cultural values and ensures the identity of Japanese culture, thus enhancing the democratization of AI. Additionally, experiments showed that pre-trained models specialized for Japanese can efficiently achieve high performance in Japanese tasks.


【10】 Voice EHR: Introducing Multimodal Audio Data for Health
标题:语音EHR:引入多模态音频数据
链接:https://arxiv.org/abs/2404.01620
作者:James Anibal,Hannah Huth,Ming Li,Lindsey Hazen,Yen Minh Lam,Nguyen Thi Thu Hang,Michael Kleinman,Shelley Ost,Christopher Jackson,Laura Sprabery,Cheran Elangovan,Balaji Krishnaiah,Lee Akst,Ioan Lina,Iqbal Elyazar,Lenny Ekwati,Stefan Jansen,Richard Nduwayezu,Charisse Garcia,Jeffrey Plum,Jacqueline Brenner,Miranda Song,Emily Ricotta,David Clifton,C. Louise Thwaites,Yael Bensoussan,Bradford Wood
备注:18 pages, 2 figures, 7 tables
摘要:在音频数据上训练的大型人工智能模型可能有潜力快速对患者进行分类,增强医疗决策,并通过早期检测来改善结果。现有技术依赖于有限的数据集,使用高收入英语国家昂贵的记录设备。这对资源受限、高容量环境中的部署提出了挑战,在这些环境中,音频数据可能会产生深远的影响。本报告介绍了一种新的数据类型和相应的收集系统,该系统仅使用移动/Web应用程序通过引导性问题捕获健康数据。该应用最终产生音频电子健康记录(语音EHR),其可以包含来自常规语音/呼吸特征、语音模式和具有语义含义的语言的复杂的健康生物标志物-补偿单峰临床数据集的典型限制。本报告介绍了一个全球合作伙伴联盟,介绍了用于数据收集的应用程序,并展示了信息语音EHR在推进音频AI的可扩展性和多样性方面的潜力。
摘要:Large AI models trained on audio data may have the potential to rapidly classify patients, enhancing medical decision-making and potentially improving outcomes through early detection. Existing technologies depend on limited datasets using expensive recording equipment in high-income, English-speaking countries. This challenges deployment in resource-constrained, high-volume settings where audio data may have a profound impact. This report introduces a novel data type and a corresponding collection system that captures health data through guided questions using only a mobile/web application. This application ultimately results in an audio electronic health record (voice EHR) which may contain complex biomarkers of health from conventional voice/respiratory features, speech patterns, and language with semantic meaning - compensating for the typical limitations of unimodal clinical datasets. This report introduces a consortium of partners for global work, presents the application used for data collection, and showcases the potential of informative voice EHR to advance the scalability and diversity of audio AI.

【11】 Transforming LLMs into Cross-modal and Cross-lingual RetrievalSystems
标题:将法学硕士转化为跨模式和跨语言检索系统
链接:https://arxiv.org/abs/2404.01616
作者:Frank Palma Gomez,Ramon Sanabria,Yun-hsuan Sung,Daniel Cer,Siddharth Dalmia,Gustavo Hernandez Abrego
摘要:大型语言模型(LLM)是在纯文本数据上训练的,远远超出了具有配对语音和文本数据的语言。同时,基于双编码器(DE)的检索系统将查询和文档投影到同一个嵌入空间中,并在检索和双文本挖掘方面取得了成功。为了在多种语言中匹配语音和文本,我们建议使用LLM来初始化多模态DE检索系统。与传统方法不同,我们的系统在LLM预训练期间不需要语音数据,并且可以利用LLM的多语言文本理解功能来匹配检索训练期间看不到的语言中的语音和文本。我们的多模态基于LLM的检索系统能够匹配102种语言的语音和文本,尽管只训练了21种语言。我们的系统比以前的系统在所有102种语言上都有更好的表现。我们在这些语言中平均实现了10%的Recall@1绝对改进。此外,我们的模型演示了跨语言语音和文本匹配,这是进一步增强了现成的机器翻译数据。
摘要:Large language models (LLMs) are trained on text-only data that go far beyond the languages with paired speech and text data. At the same time, Dual Encoder (DE) based retrieval systems project queries and documents into the same embedding space and have demonstrated their success in retrieval and bi-text mining. To match speech and text in many languages, we propose using LLMs to initialize multi-modal DE retrieval systems. Unlike traditional methods, our system doesn't require speech data during LLM pre-training and can exploit LLM's multilingual text understanding capabilities to match speech and text in languages unseen during retrieval training. Our multi-modal LLM-based retrieval system is capable of matching speech and text in 102 languages despite only training on 21 languages. Our system outperforms previous systems trained explicitly on all 102 languages. We achieve a 10% absolute improvement in Recall@1 averaged across these languages. Additionally, our model demonstrates cross-lingual speech and text matching, which is further enhanced by readily available machine translation data.

【12】 Audio Simulation for Sound Source Localization in Virtual Evironment
标题:虚拟环境中声源定位的音频仿真
链接:https://arxiv.org/abs/2404.01611
作者:Yi Di Yuan,Swee Liang Wong,Jonathan Pan
备注:2024 IEEE World Forum on Public Safety Technology
摘要:信号缺失环境下的非视距定位是一个具有挑战性的问题。由于混响性质,这种主要室内场景中的声学方法遇到困难。在这项研究中,我们的目标是通过利用物理接地的声音传播模拟和机器学习方法将声源定位到虚拟环境中的特定位置。该过程试图克服数据不足的问题,以将声源定位到其发生的位置,特别是在事后定位中。我们使用音频Transformer频谱图方法实现0.786+/- 0.0136 F1分数。
摘要:Non-line-of-sight localization in signal-deprived environments is a challenging yet pertinent problem. Acoustic methods in such predominantly indoor scenarios encounter difficulty due to the reverberant nature. In this study, we aim to locate sound sources to specific locations within a virtual environment by leveraging physically grounded sound propagation simulations and machine learning methods. This process attempts to overcome the issue of data insufficiency to localize sound sources to their location of occurrence especially in post-event localization. We achieve 0.786+/- 0.0136 F1-score using an audio transformer spectrogram approach.

机器翻译由腾讯交互翻译提供,仅供参考