今日论文合集:cs.SD语音29篇,eess.AS音频处理32篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Disentangled Representation Learning for Environment-agnostic Speaker Recognition
标题: 用于环境不可知的说话人识别的分离表示学习
作者:KiHyun Nam,Hee-Soo Heo,Jee-weon Jung,Joon Son Chung
备注:Interspeech 2024. The official webpage can be found at this https URL
链接:点击下载PDF文件
摘要:这项工作提出了一个框架的基础上,功能解开学习扬声器嵌入是强大的环境变化。我们的框架利用自动编码器作为解缠器,将输入扬声器嵌入划分为与扬声器和其他残留信息相关的组件。我们采用了一组目标函数,以确保自动编码器的代码表示-用于细化嵌入-只压缩扬声器的特性。我们展示了我们的框架的多功能性,通过其与任何现有的扬声器嵌入提取器的兼容性,不需要结构修改或适应集成。我们验证我们的框架的有效性,将其纳入两个常用的嵌入提取器,并在各种基准进行实验。结果显示,性能提高高达16%。我们发布了我们的代码为这项工作提供https: github.com kaistmm voxceleb-disentangler摘要:This work presents a framework based on feature disentanglement to learn speaker embeddings that are robust to environmental variations. Our framework utilises an auto-encoder as a disentangler, dividing the input speaker embedding into components related to the speaker and other residual information. We employ a group of objective functions to ensure that the auto-encoder's code representation - used as the refined embedding - condenses only the speaker characteristics. We show the versatility of our framework through its compatibility with any existing speaker embedding extractor, requiring no structural modifications or adaptations for integration. We validate the effectiveness of our framework by incorporating it into two popularly used embedding extractors and conducting experiments across various benchmarks. The results show a performance improvement of up to 16%. We release our code for this work to be available https: github.com kaistmm voxceleb-disentangler

【2】 Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
标题: 第二届eXplainable AI for the Arts(XAIxArts)国际研讨会会议录
作者:Nick Bryan-Kinns,Corey Ford,Shuoyang Zheng,Helen Kennedy,Alan Chamberlain,Makayla Lewis,Drew Hemment,Zijin Li,Qiong Wu,Lanxi Xiao,Gus Xia,Jeba Rezwana,Michael Clemens,Gabriel Vigliensoni
链接:点击下载PDF文件
摘要:第二届可解释人工智能艺术(XAIxArts)国际研讨会汇集了HCI,交互设计,人工智能,可解释人工智能(XAI)和数字艺术的研究人员社区,以探索XAI对艺术的作用。第16届ACM创意与认知会议(C&C 2024),芝加哥,美国摘要:This second international workshop on explainable AI for the Arts (XAIxArts) brought together a community of researchers in HCI, Interaction Design, AI, explainable AI (XAI), and digital arts to explore the role of XAI for the Arts. Workshop held at the 16th ACM Conference on Creativity and Cognition (C&C 2024), Chicago, USA.

【3】 A Review of Common Online Speaker Diarization Methods
标题: 常用在线发言人拨号方法回顾
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages
链接:点击下载PDF文件
摘要:发言人日记提供了一个问题的答案“谁发言时?“获取音频文件。这些信息可用于完成音频转录,以供进一步处理。大多数发言人日记系统假设音频文件作为一个整体是可用的。然而,存在在音频片段到达之后立即需要扬声器标签的场景。具有相应低延迟的说话者日记被称为在线说话者日记。本文件提供了一个概述。首先简要介绍了在线演讲者日志化的历史。接下来给出了用于训练和评估的分类法和数据集。在接下来的章节中,我们将详细讨论在线日志化方法和系统。最后,本文提出了在线演讲者日志化领域未来研究需要解决的问题。摘要:Speaker diarization provides the answer to the question "who spoke when?" for an audio file. This information can be used to complete audio transcripts for further processing steps. Most speaker diarization systems assume that the audio file is available as a whole. However, there are scenarios in which the speaker labels are needed immediately after the arrival of an audio segment. Speaker diarization with a correspondingly low latency is referred to as online speaker diarization. This paper provides an overview. First the history of online speaker diarization is briefly presented. Next a taxonomy and datasets for training and evaluation are given. In the sections that follow, online diarization methods and systems are discussed in detail. This paper concludes with the presentation of challenges that still need to be solved by future research in the field of online speaker diarization.

【4】 LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
标题: LARP:冷启动播放列表延续的语言音频关系预训练
作者:Rebecca Salganik,Xiaohao Liu,Yunshan Ma,Jian Kang,Tat-Seng Chua
链接:点击下载PDF文件
摘要:随着在线音乐消费越来越多地转向基于播放列表的收听,播放列表延续的任务,其中算法建议歌曲以个性化和音乐上连贯的方式扩展播放列表,对于音乐流媒体的成功至关重要。目前,许多现有的播放列表连续方法依赖于协同过滤方法来执行推荐。然而,这些方法将难以推荐缺乏交互数据的歌曲,这就是所谓的冷启动问题。目前的方法,这一挑战设计复杂的机制,从稀疏的协作数据提取关系信号,并将它们集成到内容表示。然而,这些方法将内容表示学习排除在范围之外,并利用可能与特定音乐设置的分布或格式不一致的冻结的、预先训练的内容模型。此外,即使是最先进的音乐内容模块也是(1)与冷启动设置不兼容的,或者(2)不能有效地集成跨模态和关系信号。在本文中,我们引入LARP,多模态冷启动播放列表延续模型,有效地克服这些限制。LARP是一个三阶段的对比学习框架,它将多模态和关系信号整合到其学习的表征中。我们的框架使用越来越多的阶段的任务特定的抽象:内轨道(语言音频)对比损失,轨道轨道对比损失,和轨道播放列表对比损失。两个公开的数据集上的实验结果表明,在单模态和多模态模型的播放列表连续在冷启动设置的LARP的功效。代码和数据集发布于:https: github.com Rsalganik1123 LARP。摘要:As online music consumption increasingly shifts towards playlist-based listening, the task of playlist continuation, in which an algorithm suggests songs to extend a playlist in a personalized and musically cohesive manner, has become vital to the success of music streaming. Currently, many existing playlist continuation approaches rely on collaborative filtering methods to perform recommendation. However, such methods will struggle to recommend songs that lack interaction data, an issue known as the cold-start problem. Current approaches to this challenge design complex mechanisms for extracting relational signals from sparse collaborative data and integrating them into content representations. However, these approaches leave content representation learning out of scope and utilize frozen, pre-trained content models that may not be aligned with the distribution or format of a specific musical setting. Furthermore, even the musical state-of-the-art content modules are either (1) incompatible with the cold-start setting or (2) unable to effectively integrate cross-modal and relational signals. In this paper, we introduce LARP, a multi-modal cold-start playlist continuation model, to effectively overcome these limitations. LARP is a three-stage contrastive learning framework that integrates both multi-modal and relational signals into its learned representations. Our framework uses increasing stages of task-specific abstraction: within-track (language-audio) contrastive loss, track-track contrastive loss, and track-playlist contrastive loss. Experimental results on two publicly available datasets demonstrate the efficacy of LARP over uni-modal and multi-modal models for playlist continuation in a cold-start setting. Code and dataset are released at: https: github.com Rsalganik1123 LARP.

【5】 DASB -- Discrete Audio and Speech Benchmark
标题: DASB --离散音频和语音基准
作者:Pooneh Mousavi,Luca Della Libera,Jarod Duret,Artem Ploujnikov,Cem Subakan,Mirco Ravanelli
备注:9 pages, 5 tables
链接:点击下载PDF文件
摘要:离散音频令牌最近获得了相当大的关注,因为它们有可能连接音频和语言处理,从而能够创建现代多模态大型语言模型。理想的音频标记必须有效地保留语音和语义内容以及语言信息、说话者身份和其他细节。虽然最近已经提出了几种类型的音频令牌,但由于现有研究中的评估设置不一致,因此确定各种任务的最佳令牌化器具有挑战性。为了解决这一差距,我们发布了离散音频和语音基准测试(DASB),这是一个全面的排行榜,用于在各种区分性任务中对离散音频令牌进行基准测试,包括语音识别,说话人识别和验证,情感识别,关键字定位和意图分类,以及语音增强,分离和文本到语音等生成任务。我们的研究结果表明,平均而言,语义标记在大多数判别和生成任务中的表现优于压缩标记。然而,语义标记和标准连续表示之间的性能差距仍然很大,突出了在这一领域进一步研究的必要性。摘要:Discrete audio tokens have recently gained considerable attention for their potential to connect audio and language processing, enabling the creation of modern multimodal large language models. Ideal audio tokens must effectively preserve phonetic and semantic content along with paralinguistic information, speaker identity, and other details. While several types of audio tokens have been recently proposed, identifying the optimal tokenizer for various tasks is challenging due to the inconsistent evaluation settings in existing studies. To address this gap, we release the Discrete Audio and Speech Benchmark (DASB), a comprehensive leaderboard for benchmarking discrete audio tokens across a wide range of discriminative tasks, including speech recognition, speaker identification and verification, emotion recognition, keyword spotting, and intent classification, as well as generative tasks such as speech enhancement, separation, and text-to-speech. Our results show that, on average, semantic tokens outperform compression tokens across most discriminative and generative tasks. However, the performance gap between semantic tokens and standard continuous representations remains substantial, highlighting the need for further research in this field.

【6】 SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
标题: SimSeamless:FBK参加IWSYS 2024同步语音翻译
作者:Sara Papi,Marco Gaido,Matteo Negri,Luisa Bentivogli
链接:点击下载PDF文件
摘要:本文介绍了FBK参与IWITH 2024同声传译评估活动的情况。对于今年提交的语音到文本翻译(ST)子轨道,我们提出了SimulSeamless,这是通过在其中等配置中结合AlignAtt和M4 T来实现的。无故障的M4 T模型是“现成的”,其同步推理是通过采用AlignAtt来实现的,AlignAtt是一种基于交叉注意的SimulST策略,可以在不对同步任务的底层模型进行任何重新训练或调整的情况下应用。我们参与了所有共享任务语言(英语- {德语,日语,中文}和捷克语- 英语),与去年的提交相比,取得了可接受的甚至更好的结果。SimulSeamless涵盖143种源语言和200种目标语言,发布于:https: github.com hlt-mt FBK-fairseq 。摘要:This paper describes the FBK's participation in the Simultaneous Translation Evaluation Campaign at IWSLT 2024. For this year's submission in the speech-to-text translation (ST) sub-track, we propose SimulSeamless, which is realized by combining AlignAtt and SeamlessM4T in its medium configuration. The SeamlessM4T model is used "off-the-shelf" and its simultaneous inference is enabled through the adoption of AlignAtt, a SimulST policy based on cross-attention that can be applied without any retraining or adaptation of the underlying model for the simultaneous task. We participated in all the Shared Task languages (English- {German, Japanese, Chinese}, and Czech- English), achieving acceptable or even better results compared to last year's submissions. SimulSeamless, covering more than 143 source languages and 200 target languages, is released at: https: github.com hlt-mt FBK-fairseq .

【7】 A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
标题: 用于视听深度伪造检测的单类学习多流融合方法
作者:Kyungbok Lee,You Zhang,Zhiyao Duan
链接:点击下载PDF文件
摘要:本文解决了开发一个强大的音频-视觉深度伪造检测模型的挑战。在实际用例中,新一代算法不断涌现,并且在检测方法的开发过程中不会遇到这些算法。这就要求该方法具有泛化能力。此外,为了确保检测方法的可信度,模型解释视频中哪些线索表明它是假的是有益的。出于这些考虑,我们提出了一种多流融合方法,将单类学习作为表示级正则化技术。我们通过扩展和重新分割现有的FakeAVCeleb数据集来创建一个新的基准来研究视听deepfake检测的泛化问题。该基准测试包含四个类别的假视频(真实音频-假视觉、假音频-假视觉、假音频-真实视觉和不同步视频)。实验结果表明,与基线模型相比,我们的方法在四个测试集上平均提高了7.31%的模型对不可见攻击的检测。此外,我们提出的框架提供了可解释性,表明模型将哪种模态识别为假模态。摘要:This paper addresses the challenge of developing a robust audio-visual deepfake detection model. In practical use cases, new generation algorithms are continually emerging, and these algorithms are not encountered during the development of detection methods. This calls for the generalization ability of the method. Additionally, to ensure the credibility of detection methods, it is beneficial for the model to interpret which cues from the video indicate it is fake. Motivated by these considerations, we then propose a multi-stream fusion approach with one-class learning as a representation-level regularization technique. We study the generalization problem of audio-visual deepfake detection by creating a new benchmark by extending and re-splitting the existing FakeAVCeleb dataset. The benchmark contains four categories of fake video(Real Audio-Fake Visual, Fake Audio-Fake Visual, Fake Audio-Real Visual, and unsynchronized video). The experimental results show that our approach improves the model's detection of unseen attacks by an average of 7.31% across four test sets, compared to the baseline model. Additionally, our proposed framework offers interpretability, indicating which modality the model identifies as fake.

【8】 Improved Remixing Process for Domain Adaptation-Based Speech Enhancement by Mitigating Data Imbalance in Signal-to-Noise Ratio
标题: 通过缓解信号噪比中的数据不平衡来改进基于域自适应的语音增强的重新混合过程
作者:Li Li,Shogo Seki
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:RemixIT和Remixed 2 Remixed是基于域自适应的语音增强(DASE)方法,其使用在完全监督下训练的教师模型,通过重新混合教师模型的输出来生成伪配对数据。用于增强真实世界记录信号的学生模型使用没有地面实况的伪配对数据进行训练。由于噪声信号是在自然环境中记录的,因此数据集不可避免地会在某些声学特性中出现数据不平衡,从而导致代表性不足的数据的性能低于标准。在监督学习中固有平衡的信噪比(SNR)就是一个很好的例子。在本文中,我们使用CHiME-7 UDASE任务的数据集提供了经验证据,证明伪数据的SNR对模型性能有显着影响,突出了DASE中平衡SNR的重要性。此外,我们建议采用课程学习来涵盖广泛的SNR,以提高代表性不足的数据的性能。摘要:RemixIT and Remixed2Remixed are domain adaptation-based speech enhancement (DASE) methods that use a teacher model trained in full supervision to generate pseudo-paired data by remixing the outputs of the teacher model. The student model for enhancing real-world recorded signals is trained using the pseudo-paired data without ground truth. Since the noisy signals are recorded in natural environments, the dataset inevitably suffers data imbalance in some acoustic properties, leading to subpar performance for the underrepresented data. The signal-to-noise ratio (SNR), inherently balanced in supervised learning, is a prime example. In this paper, we provide empirical evidence that the SNR of pseudo data has a significant impact on model performance using the dataset of the CHiME-7 UDASE task, highlighting the importance of balanced SNR in DASE. Furthermore, we propose adopting curriculum learning to encompass a broad range of SNRs to boost performance for underrepresented data.

【9】 Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control
标题: 空中交通管制中的关节与顺序说话者角色检测和自动语音识别
作者:Alexander Blatt,Aravind Krishnan,Dietrich Klakow
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:将空中交通管制(ATC)数据用于下游自然语言处理任务需要预处理步骤。关键步骤是通过自动语音识别(ASR)和说话人日志化(Speaker Diarization)转录数据,分别使用说话人角色检测(SRD)将转录本分为飞行员和空中交通管制员(ATCO)转录本。虽然传统方法分别承担这些任务,但我们提出了一种基于变压器的联合ASR-SRD系统,该系统在依赖标准ASR架构的同时联合解决了这两项任务。我们比较这个联合系统对两个级联的方法ASR和SRD多ATC数据集。我们的研究表明,在这种情况下,我们的联合系统可以优于两个传统的方法,在这种情况下,其他架构是优选的。我们还评估了声学和词汇差异如何影响所有架构,并展示了如何克服它们的联合架构。摘要:Utilizing air-traffic control (ATC) data for downstream natural-language processing tasks requires preprocessing steps. Key steps are the transcription of the data via automatic speech recognition (ASR) and speaker diarization, respectively speaker role detection (SRD) to divide the transcripts into pilot and air-traffic controller (ATCO) transcripts. While traditional approaches take on these tasks separately, we propose a transformer-based joint ASR-SRD system that solves both tasks jointly while relying on a standard ASR architecture. We compare this joint system against two cascaded approaches for ASR and SRD on multiple ATC datasets. Our study shows in which cases our joint system can outperform the two traditional approaches and in which cases the other architectures are preferable. We additionally evaluate how acoustic and lexical differences influence all architectures and show how to overcome them for our joint architecture.

【10】 Automated Bioacoustic Monitoring for South African Bird Species on Unlabeled Data
标题: 基于未标记数据的南非鸟类自动生物声学监测
作者:Michael Doell,Dominik Kuehn,Vanessa Suessle,Matthew J. Burnett,Colleen T. Downs,Andreas Weinmann,Elke Hergenroether
Journal-ref:International Conferences in Central Europe on Computer Graphics, Visualization and Computer Vision 2024
链接:点击下载PDF文件
摘要:基于被动声监测(PAM)记录的生物多样性监测分析是耗时的,并受到记录中存在背景噪声的挑战。现有的声音事件检测(SED)模型只适用于某些鸟类,进一步模型的开发需要标记数据。开发的框架自动提取标记的数据,从现有的平台上选定的鸟类物种。标记的数据被嵌入到录音中,包括环境声音和噪声,并用于训练卷积递归神经网络(CRNN)模型。模型进行了评估,在城市夸祖鲁-纳塔尔栖息地记录的未加工的真实世界的数据。自适应SED-CRNN模型的F1得分为0.73,证明了其在嘈杂的真实世界条件下的效率。所提出的方法,自动提取标记的数据,为选定的鸟类物种,使PAM容易适应其他物种和栖息地,为未来的保护项目。摘要:Analyses for biodiversity monitoring based on passive acoustic monitoring (PAM) recordings is time-consuming and challenged by the presence of background noise in recordings. Existing models for sound event detection (SED) worked only on certain avian species and the development of further models required labeled data. The developed framework automatically extracted labeled data from available platforms for selected avian species. The labeled data were embedded into recordings, including environmental sounds and noise, and were used to train convolutional recurrent neural network (CRNN) models. The models were evaluated on unprocessed real world data recorded in urban KwaZulu-Natal habitats. The Adapted SED-CRNN model reached a F1 score of 0.73, demonstrating its efficiency under noisy, real-world conditions. The proposed approach to automatically extract labeled data for chosen avian species enables an easy adaption of PAM to other species and habitats for future conservation projects.

【11】 ManWav: The First Manchu ASR Model
标题: ManWav:第一个满语ASC模型
作者:Jean Seo,Minha Kang,Sungjoo Byun,Sangah Lee
备注:ACL2024Field Matters
链接:点击下载PDF文件
摘要:本研究旨在解决高资源和极低资源语言之间自动语音识别(ASR)研究中日益扩大的差距,特别关注满语,一种极度濒危的语言。满语体现了边缘化语言社区在获得最先进技术方面所面临的挑战。作为开创性的努力,我们引入了有史以来第一个满语ASR模型ManWav,利用Wav 2 Vec 2-XLSR-53。第一个满语ASR的结果是有希望的,特别是在使用我们的增强数据进行训练时。与使用原始数据微调的相同基础模型相比,使用增强数据微调的Wav 2 Vec 2-XLSR-53显示CER下降0.02,WER下降0.13。摘要:This study addresses the widening gap in Automatic Speech Recognition (ASR) research between high resource and extremely low resource languages, with a particular focus on Manchu, a critically endangered language. Manchu exemplifies the challenges faced by marginalized linguistic communities in accessing state-of-the-art technologies. In a pioneering effort, we introduce the first-ever Manchu ASR model ManWav, leveraging Wav2Vec2-XLSR-53. The results of the first Manchu ASR is promising, especially when trained with our augmented data. Wav2Vec2-XLSR-53 fine-tuned with augmented data demonstrates a 0.02 drop in CER and 0.13 drop in WER compared to the same base model fine-tuned with original data.

【12】 Children's Speech Recognition through Discrete Token Enhancement
标题: 通过离散令牌增强实现儿童语音识别
作者:Vrunda N. Sukhadia,Shammur Absar Chowdhury
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:儿童的语音识别被认为是一个低资源的任务,主要是由于缺乏公开的数据。造成这种数据稀缺的原因有几个,包括昂贵的数据收集和注释过程,以及数据隐私等。将语音信号转换为不携带敏感信息但捕获语言和声学信息的离散令牌可能是隐私问题的解决方案。在这项研究中,我们调查的离散语音令牌集成到儿童的语音识别系统作为输入,而不会显着降低ASR性能。此外,我们探索了创建这些离散标签的单视图和多视图策略。此外,我们测试了模型的泛化能力与看不见的域和出生数据集。结果表明,儿童的离散令牌ASR实现了几乎相同的性能,参数减少了约83%。摘要:Children's speech recognition is considered a low-resource task mainly due to the lack of publicly available data. There are several reasons for such data scarcity, including expensive data collection and annotation processes, and data privacy, among others. Transforming speech signals into discrete tokens that do not carry sensitive information but capture both linguistic and acoustic information could be a solution for privacy concerns. In this study, we investigate the integration of discrete speech tokens into children's speech recognition systems as input without significantly degrading the ASR performance. Additionally, we explored single-view and multi-view strategies for creating these discrete labels. Furthermore, we tested the models for generalization capabilities with unseen domain and nativity dataset. Results reveal that the discrete token ASR for children achieves nearly equivalent performance with an approximate 83% reduction in parameters.

【13】 Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
标题: 直接通过Gumbel Softmax估计器的双峰神经架构搜索用于视听深度伪造检测
作者:Aravinda Reddy PN,Raghavendra Ramachandra,Krothapalli Sreenivasa Rao,Pabitra Mitra,Vinod Rathod
链接:点击下载PDF文件
摘要:Deepfakes是生物识别认证的主要安全风险。这项技术创造了逼真的假视频,可以冒充真人,欺骗依赖面部特征和语音模式进行识别的系统。现有的多模式deepfake检测器依赖于传统的融合方法,例如多数规则和集成投票,这些方法通常难以适应不断变化的数据特征和复杂模式。在本文中,我们介绍了直通Gumbel-Softmax(STGS)框架,提供了一个全面的方法来搜索多模态融合模型架构。使用两级搜索方法,该框架优化了网络结构,参数和性能。最初,从骨干网络中有效地识别关键特征,而在细胞结构中,加权融合操作集成了来自各种来源的信息。通过改变温度和采样时间等参数,得到了一个最大限度地提高分类性能的架构。在FakeAVCeleb和SWAN-DF数据集上的实验结果表明,以最小的模型参数实现了令人印象深刻的AUC值94.4%。摘要:Deepfakes are a major security risk for biometric authentication. This technology creates realistic fake videos that can impersonate real people, fooling systems that rely on facial features and voice patterns for identification. Existing multimodal deepfake detectors rely on conventional fusion methods, such as majority rule and ensemble voting, which often struggle to adapt to changing data characteristics and complex patterns. In this paper, we introduce the Straight-through Gumbel-Softmax (STGS) framework, offering a comprehensive approach to search multimodal fusion model architectures. Using a two-level search approach, the framework optimizes the network architecture, parameters, and performance. Initially, crucial features were efficiently identified from backbone networks, whereas within the cell structure, a weighted fusion operation integrated information from various sources. An architecture that maximizes the classification performance is derived by varying parameters such as temperature and sampling time. The experimental results on the FakeAVCeleb and SWAN-DF datasets demonstrated an impressive AUC value 94.4 % achieved with minimal model parameters.

【14】 Transferable speech-to-text large language model alignment module
标题: 可移植语音到文本大型语言模型对齐模块
作者:Boyong Wu,Chao Yan,Haoran Pu
备注:Accepted by InterSpeech 2024; 5 pages, 2 figures
链接:点击下载PDF文件
摘要:通过利用大型语言模型(LLM)和语音基础模型的强大功能,最先进的语音-文本双峰作品可以实现具有挑战性的任务,如口语翻译(ST)和问答(SQA)。在本文中,我们利用Whisper编码器和预训练的Yi-6 B的能力。实验结果表明,模态对齐可以实现一层模块和语音文本多任务语料库的百小时。我们在推理过程中进一步将Yi-6 B与Yi-6 B-Chat的人类偏好对齐版本交换,并发现对齐能力也适用。此外,通过奇异值分解(SVD)揭示的对齐子空间也意味着线性对齐子空间是稀疏的,这使得连接其他特征(如声纹或视频)以扩展模态的可能性。摘要:By leveraging the power of Large Language Models(LLMs) and speech foundation models, state of the art speech-text bimodal works can achieve challenging tasks like spoken translation(ST) and question answering(SQA) altogether with much simpler architectures. In this paper, we utilize the capability of Whisper encoder and pre-trained Yi-6B. Empirical results reveal that modal alignment can be achieved with one layer module and hundred hours of speech-text multitask corpus. We further swap the Yi-6B with human preferences aligned version of Yi-6B-Chat during inference, and discover that the alignment capability is applicable as well. In addition, the alignment subspace revealed by singular value decomposition(SVD) also implies linear alignment subspace is sparse, which leaves the possibility to concatenate other features like voice-print or video to expand modality.

【15】 SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
标题: SD-Eval:口语对话理解超越言语的基准数据集
作者:Junyi Ao,Yuancheng Wang,Xiaohai Tian,Dekun Chen,Jun Zhang,Lu Lu,Yuxuan Wang,Haizhou Li,Zhizheng Wu
链接:点击下载PDF文件
摘要:语音包含丰富的信息,包括但不限于内容、语言和环境信息。语音的这种综合性质对通信产生了重大影响,对人机交互至关重要。面向聊天的大型语言模型(LLM)以其通用的辅助功能而闻名,已经发展到可以处理包括语音在内的多模态输入。虽然这些模型可以熟练地识别和分析语音,但它们通常无法产生适当的响应。我们认为,这是由于缺乏任务定义和模型开发的原则,这需要开源数据集和适合模型评估的指标。为了弥合这一差距,我们提出了SD-Eval,这是一个基准数据集,旨在对口语对话的理解和生成进行多维评估。SD-Eval侧重于语言和环境信息,包括7,303个话语,相当于8.76小时的语音数据。这些数据来自八个公共数据集,代表四个方面:情感,口音,年龄和背景声音。为了评估SD-Eval基准数据集,我们实现了三种不同的模型,并按照与SD-Eval类似的过程构建了一个训练集。训练集包含1,052.72小时的语音数据和724.4k话语。我们还使用客观评估方法(例如BLEU和ROUGE),主观评估和基于LLM的指标对生成的响应进行综合评估。以语言和环境信息为条件的模型在客观和主观测量中都优于它们的同行。此外,实验表明,基于LLM的指标显示出更高的相关性与人类的评价相比,传统的指标。我们在https: github.com amphionspace SD-Eval上开源了SD-Eval。摘要:Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information. This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction. Chat-Oriented Large Language Models (LLMs), known for their general-purpose assistance capabilities, have evolved to handle multi-modal inputs, including speech. Although these models can be adept at recognizing and analyzing speech, they often fall short of generating appropriate responses. We argue that this is due to the lack of principles on task definition and model development, which requires open-source datasets and metrics suitable for model evaluation. To bridge the gap, we present SD-Eval, a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation. SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound. To assess the SD-Eval benchmark dataset, we implement three different models and construct a training set following a similar process as SD-Eval. The training set contains 1,052.72 hours of speech data and 724.4k utterances. We also conduct a comprehensive evaluation using objective evaluation methods (e.g. BLEU and ROUGE), subjective evaluations and LLM-based metrics for the generated responses. Models conditioned with paralinguistic and environmental information outperform their counterparts in both objective and subjective measures. Moreover, experiments demonstrate LLM-based metrics show a higher correlation with human evaluation compared to traditional metrics. We open-source SD-Eval at https: github.com amphionspace SD-Eval.

【16】 Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
标题: 通过优化的音频编码通过大型语言模型增强自动音频字幕
作者:Jizhong Liu,Gang Li,Junbo Zhang,Heinrich Dinkel,Yongqing Wang,Zhiyong Yan,Yujun Wang,Bin Wang
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自动音频字幕(AAC)是一种以自然语言描述音频内容的音频到文本的任务。最近,大型语言模型(LLM)的进步,以及音频编码器训练方法的改进,为改进AAC开辟了可能性。因此,本文从三个方面对AAC的增强进行了探索:1)通过一致性集成蒸馏(CED)的预训练音频编码器来提高声学标记的有效性,并使用查询Transformer(Q-Former)弥补了LLM和压缩声学标记的模态差距; 2)我们研究了使用具有7 B参数的Llama 2作为解码器的优点; 3)另一个预训练的LLM校正由训练数据不足和注释歧义引起的文本错误。音频编码器和文本解码器都通过-Base(LoRA)进行了优化。实验表明,这些改进都是有效的。我们的方法获得了33.0 SPIDER-FL分数,优于DCASE 2023任务6A的获胜者。摘要:Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improvements in training approaches for audio encoders, have opened up possibilities for improving AAC. Thus, we explore enhancing AAC from three aspects: 1) a pre-trained audio encoder via consistent ensemble distillation (CED) is used to improve the effectivity of acoustic tokens, with a querying transformer (Q-Former) bridging the modality gap to LLM and compress acoustic tokens; 2) we investigate the advantages of using a Llama 2 with 7B parameters as the decoder; 3) another pre-trained LLM corrects text errors caused by insufficient training data and annotation ambiguities. Both the audio encoder and text decoder are optimized by -Base (LoRA). Experiments show that each of these enhancements is effective. Our method obtains a 33.0 SPIDEr-FL score, outperforming the winner of DCASE 2023 Task 6A.

【17】 Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
标题: 利用尖峰神经网络进行全局-局部卷积以实现节能关键词发现
作者:Shuai Wang,Dehao Zhang,Kexin Shi,Yuchen Wang,Wenjie Wei,Jibin Wu,Malu Zhang
链接:点击下载PDF文件
摘要:得益于深度神经网络(DNN),关键词识别(KWS)的准确性已经取得了实质性的进展。然而,由于KWS系统通常在边缘设备上实现,因此除了性能之外,能源效率也成为一项关键要求。在这里,我们利用尖峰神经网络的能量效率,并提出了一个端到端的轻量级KWS模型。该模型由两个创新模块组成:1)全局-局部尖峰卷积(GLSC)模块和2)瓶颈PLIF模块。与手工特征提取方法相比,GLSC模块实现了更稀疏、更节能、性能更好的语音特征提取。Bottleneck-PLIF模块进一步处理来自GLSC的信号,目的是用更少的参数实现更高的精度。在Google语音命令数据集(V1和V2)上进行了大量的实验。结果表明,我们的方法取得了竞争力的性能基于SNN的KWS模型与较少的参数。摘要:Thanks to Deep Neural Networks (DNNs), the accuracy of Keyword Spotting (KWS) has made substantial progress. However, as KWS systems are usually implemented on edge devices, energy efficiency becomes a critical requirement besides performance. Here, we take advantage of spiking neural networks' energy efficiency and propose an end-to-end lightweight KWS model. The model consists of two innovative modules: 1) Global-Local Spiking Convolution (GLSC) module and 2) Bottleneck-PLIF module. Compared to the hand-crafted feature extraction methods, the GLSC module achieves speech feature extraction that is sparser, more energy-efficient, and yields better performance. The Bottleneck-PLIF module further processes the signals from GLSC with the aim to achieve higher accuracy with fewer parameters. Extensive experiments are conducted on the Google Speech Commands Dataset (V1 and V2). The results show our method achieves competitive performance among SNN-based KWS models with fewer parameters.

【18】 CONMOD: Controllable Neural Frame-based Modulation Effects
标题: CONMOD:可控的基于神经框架的调制效应
作者:Gyubin Lee,Hounsu Kim,Junwon Lee,Juhan Nam
链接:点击下载PDF文件
摘要:深度学习模型已经广泛用于对LFO驱动的音频效果进行建模,例如相位器和镶边。虽然现有的神经架构表现出高质量的仿真的个人影响,他们不具备通过控制参数来操纵输出的能力。为了解决这个问题,我们引入了基于可控神经帧的调制效应(CONMOD),这是一个单一的黑盒模型,它以逐帧的方式模拟各种LFO驱动的效应,提供对LFO频率和反馈参数的控制。此外,该模型能够学习两种不同相位器效果的连续嵌入空间,使我们能够在效果之间进行引导并实现创造性的输出。我们的模型优于以前的工作,同时具有可控性和通用性,提供了机会,以提高现代LFO驱动的音频效果的创造力。摘要:Deep learning models have seen widespread use in modelling LFO-driven audio effects, such as phaser and flanger. Although existing neural architectures exhibit high-quality emulation of individual effects, they do not possess the capability to manipulate the output via control parameters. To address this issue, we introduce Controllable Neural Frame-based Modulation Effects (CONMOD), a single black-box model which emulates various LFO-driven effects in a frame-wise manner, offering control over LFO frequency and feedback parameters. Additionally, the model is capable of learning the continuous embedding space of two distinct phaser effects, enabling us to steer between effects and achieve creative outputs. Our model outperforms previous work while possessing both controllability and universality, presenting opportunities to enhance creativity in modern LFO-driven audio effects.

【19】 Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
标题: 基于扩散的生成建模和区分性指导,用于流媒体语音增强
作者:Chenda Li,Samuele Cornell,Shinji Watanabe,Yanmin Qian
链接:点击下载PDF文件
摘要:基于扩散的生成模型(DGMs)最近在语音增强研究(SE)中引起了人们的注意,因为以前的工作显示出显着的泛化能力。然而,DGM也是计算密集型的,因为它们通常需要在反向扩散过程(RDP)中进行多次迭代,使得它们对于流式SE系统是不切实际的。在本文中,我们建议在RDP的第一步中使用来自判别模型的判别分数。这些判别分数仅需要一个前向传递与判别模型的多个RDP步骤,从而大大减少了计算。这种方法还允许性能改进。我们表明,我们可以权衡之间的生成和判别能力的步骤数与判别分数的增加。此外,我们提出了一种新的流式时域生成模型,算法延迟为50 ms,与离线模型相比,它没有显着的性能下降。摘要:Diffusion-based generative models (DGMs) have recently attracted attention in speech enhancement research (SE) as previous works showed a remarkable generalization capability. However, DGMs are also computationally intensive, as they usually require many iterations in the reverse diffusion process (RDP), making them impractical for streaming SE systems. In this paper, we propose to use discriminative scores from discriminative models in the first steps of the RDP. These discriminative scores require only one forward pass with the discriminative model for multiple RDP steps, thus greatly reducing computations. This approach also allows for performance improvements. We show that we can trade off between generative and discriminative capabilities as the number of steps with the discriminative score increases. Furthermore, we propose a novel streamable time-domain generative model with an algorithmic latency of 50 ms, which has no significant performance degradation compared to offline models.

【20】 Automatic Voice Classification Of Autistic Subjects
标题: 自闭症受试者的自动语音分类
作者:Jessica Vacca,Natascia Brondino,Fabio Dell'Acqua,Anna Vizziello,Pietro Savazzi
备注:Accepted for publication at EAI BODYNETS 2023 2024 - 18th EAI International Conference on Body Area Networks: Intelligent Edge Cloud for Dependable Globally Connected BAN, February 5-6, 2024 Milan, Italy
链接:点击下载PDF文件
摘要:自闭症谱系障碍(ASD)描述了一组被归类为神经发育障碍的异质性疾病。虽然ASD的潜在机制尚未完全了解,但最近的文献集中在多种遗传和 或环境风险因素上。症状的异质性,特别是在这种情况下的轻度形式,可能是临床医生的挑战。在这项工作中,提出了一种自动语音分类算法来表征最能区分自闭症的韵律元素,以支持传统的诊断。所提出的算法的性能进行评估,通过测试的分类算法在一个数据集组成的录音讲话,自闭症和非自闭症患者的主题之间收集。摘要:Autism Spectrum Disorders (ASD) describe a heterogeneous set of conditions classified as neurodevelopmental disorders. Although the mechanisms underlying ASD are not yet fully understood, more recent literature focused on multiple genetics and or environmental risk factors. Heterogeneity of symptoms, especially in milder forms of this condition, could be a challenge for the clinician. In this work, an automatic speech classification algorithm is proposed to characterize the prosodic elements that best distinguish autism, to support the traditional diagnosis. The performance of the proposed algorithm is evaluted by testing the classification algorithms on a dataset composed of recorded speeches, collected among both autustic and non autistic subjects.

【21】 Online Domain-Incremental Learning Approach to Classify Acoustic Scenes in All Locations
标题: 在线领域增量学习方法对所有地点的声学场景进行分类
作者:Manjunath Mulimani,Annamaria Mesaros
备注:Accepted to EUSIPCO 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种方法在线域增量学习的声学场景分类从一系列不同的位置。简单地在一系列不同的位置上训练深度学习模型会导致忘记以前学到的知识。在这项工作中,我们只使用一些样本来纠正模型的Batch Normalization层的统计数据,以从新位置学习声学场景,而无需任何过度训练。实验进行声学场景从11个不同的位置,与包含声学场景从6个位置的初始任务和其余5个增量任务,每个代表的声学场景从不同的位置。所提出的方法优于基于微调的方法,并实现了48.8%的平均准确率后,学习的最后一个任务的顺序,而不会忘记声学场景从以前学习的位置。摘要:In this paper, we propose a method for online domain-incremental learning of acoustic scene classification from a sequence of different locations. Simply training a deep learning model on a sequence of different locations leads to forgetting of previously learned knowledge. In this work, we only correct the statistics of the Batch Normalization layers of a model using a few samples to learn the acoustic scenes from a new location without any excessive training. Experiments are performed on acoustic scenes from 11 different locations, with an initial task containing acoustic scenes from 6 locations and the remaining 5 incremental tasks each representing the acoustic scenes from a different location. The proposed approach outperforms fine-tuning based methods and achieves an average accuracy of 48.8% after learning the last task in sequence without forgetting acoustic scenes from the previously learned locations.

【22】 Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
标题: 通过非负矩阵分解和探测进行可解释的按设计音频分割
作者:Martin Lebourdais,Théo Mariotte,Antonio Almudévar,Marie Tahon,Alfonso Ortega
备注:Accepted at Interspeech 2024, 5 pages, 2 figures, 3 tables
链接:点击下载PDF文件
摘要:音频分割是许多语音技术的关键任务,其中大多数基于神经网络,通常被认为是黑箱,具有高水平的性能。然而,在许多领域,其中包括健康或法医学,不仅需要良好的性能,而且还需要对输出决策的解释。直接从潜在表征导出的解释需要满足“好”的属性,例如信息量、紧凑性或模块性,才是可解释的。在这篇文章中,我们提出了一个基于非负矩阵分解(NMF)的可解释的设计音频分割模型,这是一个很好的候选人的设计可解释的表示。本文表明,我们的模型达到了良好的分割性能,并提出了深入的分析,从非负矩阵提取的潜在表示。所提出的方法打开了新的视角,根据“好”的属性对可解释的表示的评价。摘要:Audio segmentation is a key task for many speech technologies, most of which are based on neural networks, usually considered as black boxes, with high-level performances. However, in many domains, among which health or forensics, there is not only a need for good performance but also for explanations about the output decision. Explanations derived directly from latent representations need to satisfy "good" properties, such as informativeness, compactness, or modularity, to be interpretable. In this article, we propose an explainable-by-design audio segmentation model based on non-negative matrix factorization (NMF) which is a good candidate for the design of interpretable representations. This paper shows that our model reaches good segmentation performances, and presents deep analyses of the latent representation extracted from the non-negative matrix. The proposed approach opens new perspectives toward the evaluation of interpretable representations according to "good" properties.

【23】 Medical Spoken Named Entity Recognition
标题: 医疗口语命名实体识别
作者:Khai Le-Duc
备注:Preprint, 40 pages
链接:点击下载PDF文件
摘要:口语命名实体识别(NER)的目的是从语音中提取命名实体,并将其分类为人,位置,组织等类型。在这项工作中,我们提出了VietMed-NER -医疗领域的第一个口语NER数据集。据我们所知,就实体类型的数量而言,我们的真实世界数据集是世界上最大的口语NER数据集,具有18种不同的类型。其次,我们使用各种最先进的预训练模型呈现基线结果:仅编码器和序列到序列。我们发现,预训练的多语言模型XLM-R在参考文本和ASR输出上都优于所有单语言模型。同样在一般情况下,编码器执行NER任务比序列到序列模型。通过简单的翻译,文字记录不仅适用于越南语,也适用于其他语言。所有的代码、数据和模型都在这里公开发布:https: github.com leduckhai MultiMed摘要:Spoken Named Entity Recognition (NER) aims to extracting named entities from speech and categorizing them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our best knowledge, our real-world dataset is the largest spoken NER dataset in the world in terms of the number of entity types, featuring 18 distinct types. Secondly, we present baseline results using various state-of-the-art pre-trained models: encoder-only and sequence-to-sequence. We found that pre-trained multilingual models XLM-R outperformed all monolingual models on both reference text and ASR output. Also in general, encoders perform better than sequence-to-sequence models for the NER task. By simply translating, the transcript is applicable not just to Vietnamese but to other languages as well. All code, data and models are made publicly available here: https: github.com leduckhai MultiMed

【24】 Pushing the Limit of Sound Event Detection with Multi-Dilated Frequency Dynamic Convolution
标题: 利用多频动态卷积突破声音事件检测的极限
作者:Hyeonuk Nam,Yong-Hwa Park
链接:点击下载PDF文件
摘要:频率动态卷积(FDY conv)是声音事件检测(SED)领域的一个里程碑,但由于多个基核,它涉及模型大小的大幅增加。在这项工作中,我们提出了部分频率动态卷积(PFD conv),它连接静态传统的2D卷积分支输出和动态FDY conv分支输出,以尽量减少模型大小的增加,同时保持性能。此外,我们提出了多扩张频率动态卷积(MDFD conv),它集成了多个扩张频率动态卷积(DFD conv)分支与不同的扩张大小集和一个静态分支在一个单一的卷积模块,实现了3.2%的改善复调声音检测分数(PSDS)FDY conv。广泛消融研究的拟议方法进一步增强了对FDY conv变体的理解和可用性。摘要:Frequency dynamic convolution (FDY conv) has been a milestone in the sound event detection (SED) field, but it involves a substantial increase in model size due to multiple basis kernels. In this work, we propose partial frequency dynamic convolution (PFD conv), which concatenates static conventional 2D convolution branch output and dynamic FDY conv branch output in order to minimize model size increase while maintaining the performance. Additionally, we propose multi-dilated frequency dynamic convolution (MDFD conv), which integrates multiple dilated frequency dynamic convolution (DFD conv) branches with different dilation size sets and a static branch within a single convolution module, achieving a 3.2% improvement in polyphonic sound detection score (PSDS) over FDY conv. Proposed methods with extensive ablation studies further enhance understanding and usability of FDY conv variants.

【25】 CEC: A Noisy Label Detection Method for Speaker Recognition
标题: NEC:一种用于说话人识别的有噪标签检测方法
作者:Yao Shen,Yingying Gao,Yaqian Hao,Chenguang Hu,Fulin Zhang,Junlan Feng,Shilei Zhang
备注:interspeech 2024
链接:点击下载PDF文件
摘要:噪声标签是不可避免的,即使是在注释良好的数据集中。噪声标签的检测对于提高说话人识别模型的鲁棒性具有重要意义。在本文中,我们提出了一种新的噪声标签检测方法的基础上两个新的统计指标:连续不一致计数(CIC)和总不一致计数(TIC)。这些指标是通过跨时期计数(CEC)计算的,分别对应于训练的早期和后期阶段。此外,我们根据预测结果将样本分为三类:不一致的样本,硬样本和容易的样本。在训练过程中,我们逐渐增加硬样本更新模型参数的难度,防止噪声标签被过度拟合。与对比方案相比,我们的方法不仅在说话人确认方面取得了最好的性能,而且在噪声标签检测方面也表现出色。摘要:Noisy labels are inevitable, even in well-annotated datasets. The detection of noisy labels is of significant importance to enhance the robustness of speaker recognition models. In this paper, we propose a novel noisy label detection approach based on two new statistical metrics: Continuous Inconsistent Counting (CIC) and Total Inconsistent Counting (TIC). These metrics are calculated through Cross-Epoch Counting (CEC) and correspond to the early and late stages of training, respectively. Additionally, we categorize samples based on their prediction results into three categories: inconsistent samples, hard samples, and easy samples. During training, we gradually increase the difficulty of hard samples to update model parameters, preventing noisy labels from being overfitted. Compared to contrastive schemes, our approach not only achieves the best performance in speaker verification but also excels in noisy label detection.

【26】 Audio Fingerprinting with Holographic Reduced Representations
标题: 具有全息简化表示的音频指纹识别
作者:Yusuke Fujita,Tatsuya Komatsu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:提出了一种基于全息简化表示的音频指纹模型。所提出的方法减少了存储的指纹的数量,而传统的神经音频指纹识别需要许多指纹的每个音轨,以实现高精度和时间分辨率。我们利用HRR通过循环卷积和求和将多个指纹聚合成一个复合指纹,从而减少与原始指纹具有相同维度空间的指纹。我们的搜索方法有效地找到一个组合指纹,其中存在一个查询指纹。利用HRR的逆运算,它可以恢复组合指纹内的相对位置,保持原始的时间分辨率。实验表明,该方法可以减少指纹的数量,同时保持适当的精度下降的时间分辨率,优于简单的抽取和求和为基础的聚合方法。摘要:This paper proposes an audio fingerprinting model with holographic reduced representation (HRR). The proposed method reduces the number of stored fingerprints, whereas conventional neural audio fingerprinting requires many fingerprints for each audio track to achieve high accuracy and time resolution. We utilize HRR to aggregate multiple fingerprints into a composite fingerprint via circular convolution and summation, resulting in fewer fingerprints with the same dimensional space as the original. Our search method efficiently finds a combined fingerprint in which a query fingerprint exists. Using HRR's inverse operation, it can recover the relative position within a combined fingerprint, retaining the original time resolution. Experiments show that our method can reduce the number of fingerprints with modest accuracy degradation while maintaining the time resolution, outperforming simple decimation and summation-based aggregation methods.

【27】 Articulatory Encodec: Vocal Tract Kinematics as a Codec for Speech
标题: 关节编码编码器:作为言语编码器的声道运动学
作者:Cheol Jun Cho,Peter Wu,Tejas S. Prabhune,Dhruv Agarwal,Gopala K. Anumanchipalli
链接:点击下载PDF文件
摘要:声道发音是一个自然的、有根据的言语产生控制空间。发音器官的时空协调与声源相结合,形成可理解的语音,以实现有效的口语交流。基于语音的这一生理基础,我们提出了一种新的神经语音编解码框架--发音编码器。发音编码器包括从语音音频推断发音特征的发音分析模型和从发音特征合成语音音频的发音合成模型。发音特征是发音器官的运动轨迹和声源特征,是语音产生的实际物理界面,具有直观的可解释性和可控性。一个额外的扬声器身份编码器与发音合成器联合训练,以告知各个扬声器的语音纹理。通过对大规模语音数据的训练,我们实现了一个完全可理解的,高质量的发音合成器,它可以推广到看不见的说话者。此外,扬声器嵌入有效地从发音中解脱出来,这使得能够实现口音保持的zero-shot语音转换。据我们所知,这是第一次演示的通用,高性能的发音推理和合成,建议作为一个强大的语音编码系统的框架。摘要:Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spoken communication. Based on this physiological grounding of speech, we propose a new framework of neural encoding-decoding of speech -- articulatory encodec. The articulatory encodec comprises an articulatory analysis model that infers articulatory features from speech audio, and an articulatory synthesis model that synthesizes speech audio from articulatory features. The articulatory features are kinematic traces of vocal tract articulators and source features, which are intuitively interpretable and controllable, being the actual physical interface of speech production. An additional speaker identity encoder is jointly trained with the articulatory synthesizer to inform the voice texture of individual speakers. By training on large-scale speech data, we achieve a fully intelligible, high-quality articulatory synthesizer that generalizes to unseen speakers. Furthermore, the speaker embedding is effectively disentangled from articulations, which enables accent-perserving zero-shot voice conversion. To the best of our knowledge, this is the first demonstration of universal, high-performance articulatory inference and synthesis, suggesting the proposed framework as a powerful coding system of speech.

【28】 Self-Train Before You Transcribe
标题: 抄写前自我训练
作者:Robert Flynn,Anton Ragni
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:当训练域和测试域之间存在不匹配时,当前的语音识别系统表现出显着的性能下降。自我训练方法,例如嘈杂的学生教师培训,可以帮助解决这个问题,并使模型能够在这种领域转变下进行调整。然而,自我训练通常需要收集未标记的目标域数据。对于不切实际的环境,我们研究了对测试集中的录音进行嘈杂的学生教师培训作为测试时适应方法的好处。类似于语言建模中的动态评估方法,这使得跨话语边界的信息传递成为可能,并且作为域适应的方法发挥作用。一系列的域内和域外数据集用于实验,证明高达32.2%的相对增益。有趣的是,我们的方法显示出比利用单独适应数据的典型自我训练设置更大的增益。摘要:When there is a mismatch between the training and test domains, current speech recognition systems show significant performance degradation. Self-training methods, such as noisy student teacher training, can help address this and enable the adaptation of models under such domain shifts. However, self-training typically requires a collection of unlabelled target domain data. For settings where this is not practical, we investigate the benefit of performing noisy student teacher training on recordings in the test set as a test-time adaptation approach. Similarly to the dynamic evaluation approach in language modelling, this enables the transfer of information across utterance boundaries and functions as a method of domain adaptation. A range of in-domain and out-of-domain datasets are used for experiments demonstrating large relative gains of up to 32.2%. Interestingly, our method showed larger gains than the typical self-training setup that utilises separate adaptation data.

【29】 Automatic Speech Recognition for Biomedical Data in Bengali Language
标题: 孟加拉语生物医学数据的自动语音识别
作者:Shariar Kabir,Nazmun Nahar,Shyamasree Saha,Mamunur Rashid
链接:点击下载PDF文件
摘要:本文介绍了一个原型的自动语音识别(ASR)系统,专门为孟加拉生物医学数据的发展。孟加拉语ASR的最新进展令人鼓舞,但缺乏特定领域的数据限制了实际医疗保健ASR模型的创建。该项目通过开发针对孟加拉语医学术语(如症状,严重程度和疾病)的ASR系统来弥合这一差距,包括两种主要方言:孟加拉语和Sylheti。我们训练和评估两个流行的ASR框架上全面的46小时孟加拉语医学语料库。我们的核心目标是为数字健康应用创建可部署的健康领域ASR系统,最终提高医疗保健领域非技术用户的可访问性。摘要:This paper presents the development of a prototype Automatic Speech Recognition (ASR) system specifically designed for Bengali biomedical data. Recent advancements in Bengali ASR are encouraging, but a lack of domain-specific data limits the creation of practical healthcare ASR models. This project bridges this gap by developing an ASR system tailored for Bengali medical terms like symptoms, severity levels, and diseases, encompassing two major dialects: Bengali and Sylheti. We train and evaluate two popular ASR frameworks on a comprehensive 46-hour Bengali medical corpus. Our core objective is to create deployable health-domain ASR systems for digital health applications, ultimately increasing accessibility for non-technical users in the healthcare sector.



eess.AS音频处理
【1】 Decoding Vocal Articulations from Acoustic Latent Representations
标题: 从声学潜在表示解码声乐清晰度
作者:Mateo Cámara,Fernando Marcos,José Luis Blanco
备注:Presented at AES Europe 2024 in Madrid
链接:点击下载PDF文件
摘要:我们提出了一种新的神经编码器系统的声学发音反转。我们利用粉红长号语音合成器,揭示发音参数(例如舌头位置和声带配置)。我们的系统被设计用于识别负责产生包含在神经潜在表示中的特定声学特征的发音特征。为了生成必要的潜在嵌入,我们采用了两种主要方法。第一个是一个自我监督的变分自动编码器,从头开始训练,在解码器阶段重建输入信号。我们用一个称为“投影仪”的子网络来调节它的瓶颈层,它解码语音合成器的参数。 第二种方法使用了两个预训练模型:EnCodec和Wav2Vec。它们消除了从头开始训练编码过程的需要,使我们能够专注于训练投影仪网络。这种方法的目的是探索这些现有的模型在声学发音反转的背景下的潜力。通过重用预训练的模型,我们大大简化了数据处理管道,提高了效率并减少了计算开销。 我们项目的主要目标是证明这些神经架构可以有效地封装声学和发音特征。这种基于预测的方法比传统的基于声学特征的参数优化方法快得多。我们通过预测六个不同的参数来验证我们的模型,并使用合成器和人类生成的声音用客观和ViSQOL主观等效度量对其进行评估。实验结果表明,当输入到合成器中时,预测的参数可以产生类似于人类的元音声音。我们提供了数据集,代码和详细的研究结果,以支持在这一领域的未来研究。摘要:We present a novel neural encoder system for acoustic-to-articulatory inversion. We leverage the Pink Trombone voice synthesizer that reveals articulatory parameters (e.g tongue position and vocal cord configuration). Our system is designed to identify the articulatory features responsible for producing specific acoustic characteristics contained in a neural latent representation. To generate the necessary latent embeddings, we employed two main methodologies. The first was a self-supervised variational autoencoder trained from scratch to reconstruct the input signal at the decoder stage. We conditioned its bottleneck layer with a subnetwork called the "projector," which decodes the voice synthesizer's parameters. The second methodology utilized two pretrained models: EnCodec and Wav2Vec. They eliminate the need to train the encoding process from scratch, allowing us to focus on training the projector network. This approach aimed to explore the potential of these existing models in the context of acoustic-to-articulatory inversion. By reusing the pretrained models, we significantly simplified the data processing pipeline, increasing efficiency and reducing computational overhead. The primary goal of our project was to demonstrate that these neural architectures can effectively encapsulate both acoustic and articulatory features. This prediction-based approach is much faster than traditional methods focused on acoustic feature-based parameter optimization. We validated our models by predicting six different parameters and evaluating them with objective and ViSQOL subjective-equivalent metric using both synthesizer- and human-generated sounds. The results show that the predicted parameters can generate human-like vowel sounds when input into the synthesizer. We provide the dataset, code, and detailed findings to support future research in this field.

【2】 CONMOD: Controllable Neural Frame-based Modulation Effects
标题: CONMOD:可控的基于神经框架的调制效应
作者:Gyubin Lee,Hounsu Kim,Junwon Lee,Juhan Nam
链接:点击下载PDF文件
摘要:深度学习模型已经广泛用于对LFO驱动的音频效果进行建模,例如相位器和镶边。虽然现有的神经架构表现出高质量的仿真的个人影响,他们不具备通过控制参数来操纵输出的能力。为了解决这个问题,我们引入了基于可控神经帧的调制效应(CONMOD),这是一个单一的黑盒模型,它以逐帧的方式模拟各种LFO驱动的效应,提供对LFO频率和反馈参数的控制。此外,该模型能够学习两种不同相位器效果的连续嵌入空间,使我们能够在效果之间进行引导并实现创造性的输出。我们的模型优于以前的工作,同时具有可控性和通用性,提供了机会,以提高现代LFO驱动的音频效果的创造力。摘要:Deep learning models have seen widespread use in modelling LFO-driven audio effects, such as phaser and flanger. Although existing neural architectures exhibit high-quality emulation of individual effects, they do not possess the capability to manipulate the output via control parameters. To address this issue, we introduce Controllable Neural Frame-based Modulation Effects (CONMOD), a single black-box model which emulates various LFO-driven effects in a frame-wise manner, offering control over LFO frequency and feedback parameters. Additionally, the model is capable of learning the continuous embedding space of two distinct phaser effects, enabling us to steer between effects and achieve creative outputs. Our model outperforms previous work while possessing both controllability and universality, presenting opportunities to enhance creativity in modern LFO-driven audio effects.

【3】 Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
标题: 基于扩散的生成建模和区分性指导,用于流媒体语音增强
作者:Chenda Li,Samuele Cornell,Shinji Watanabe,Yanmin Qian
链接:点击下载PDF文件
摘要:基于扩散的生成模型(DGMs)最近在语音增强研究(SE)中引起了人们的注意,因为以前的工作显示出显着的泛化能力。然而,DGM也是计算密集型的,因为它们通常需要在反向扩散过程(RDP)中进行多次迭代,使得它们对于流式SE系统是不切实际的。在本文中,我们建议在RDP的第一步中使用来自判别模型的判别分数。这些判别分数仅需要一个前向传递与判别模型的多个RDP步骤,从而大大减少了计算。这种方法还允许性能改进。我们表明,我们可以权衡之间的生成和判别能力的步骤数与判别分数的增加。此外,我们提出了一种新的流式时域生成模型,算法延迟为50 ms,与离线模型相比,它没有显着的性能下降。摘要:Diffusion-based generative models (DGMs) have recently attracted attention in speech enhancement research (SE) as previous works showed a remarkable generalization capability. However, DGMs are also computationally intensive, as they usually require many iterations in the reverse diffusion process (RDP), making them impractical for streaming SE systems. In this paper, we propose to use discriminative scores from discriminative models in the first steps of the RDP. These discriminative scores require only one forward pass with the discriminative model for multiple RDP steps, thus greatly reducing computations. This approach also allows for performance improvements. We show that we can trade off between generative and discriminative capabilities as the number of steps with the discriminative score increases. Furthermore, we propose a novel streamable time-domain generative model with an algorithmic latency of 50 ms, which has no significant performance degradation compared to offline models.

【4】 Automatic Voice Classification Of Autistic Subjects
标题: 自闭症受试者的自动语音分类
作者:Jessica Vacca,Natascia Brondino,Fabio Dell'Acqua,Anna Vizziello,Pietro Savazzi
备注:Accepted for publication at EAI BODYNETS 2023 2024 - 18th EAI International Conference on Body Area Networks: Intelligent Edge Cloud for Dependable Globally Connected BAN, February 5-6, 2024 Milan, Italy
链接:点击下载PDF文件
摘要:自闭症谱系障碍(ASD)描述了一组被归类为神经发育障碍的异质性疾病。虽然ASD的潜在机制尚未完全了解,但最近的文献集中在多种遗传和 或环境风险因素上。症状的异质性,特别是在这种情况下的轻度形式,可能是临床医生的挑战。在这项工作中,提出了一种自动语音分类算法来表征最能区分自闭症的韵律元素,以支持传统的诊断。所提出的算法的性能进行评估,通过测试的分类算法在一个数据集组成的录音讲话,自闭症和非自闭症患者的主题之间收集。摘要:Autism Spectrum Disorders (ASD) describe a heterogeneous set of conditions classified as neurodevelopmental disorders. Although the mechanisms underlying ASD are not yet fully understood, more recent literature focused on multiple genetics and or environmental risk factors. Heterogeneity of symptoms, especially in milder forms of this condition, could be a challenge for the clinician. In this work, an automatic speech classification algorithm is proposed to characterize the prosodic elements that best distinguish autism, to support the traditional diagnosis. The performance of the proposed algorithm is evaluted by testing the classification algorithms on a dataset composed of recorded speeches, collected among both autustic and non autistic subjects.

【5】 Online Domain-Incremental Learning Approach to Classify Acoustic Scenes in All Locations
标题: 在线领域增量学习方法对所有地点的声学场景进行分类
作者:Manjunath Mulimani,Annamaria Mesaros
备注:Accepted to EUSIPCO 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种方法在线域增量学习的声学场景分类从一系列不同的位置。简单地在一系列不同的位置上训练深度学习模型会导致忘记以前学到的知识。在这项工作中,我们只使用一些样本来纠正模型的Batch Normalization层的统计数据,以从新位置学习声学场景,而无需任何过度训练。实验进行声学场景从11个不同的位置,与包含声学场景从6个位置的初始任务和其余5个增量任务,每个代表的声学场景从不同的位置。所提出的方法优于基于微调的方法,并实现了48.8%的平均准确率后,学习的最后一个任务的顺序,而不会忘记声学场景从以前学习的位置。摘要:In this paper, we propose a method for online domain-incremental learning of acoustic scene classification from a sequence of different locations. Simply training a deep learning model on a sequence of different locations leads to forgetting of previously learned knowledge. In this work, we only correct the statistics of the Batch Normalization layers of a model using a few samples to learn the acoustic scenes from a new location without any excessive training. Experiments are performed on acoustic scenes from 11 different locations, with an initial task containing acoustic scenes from 6 locations and the remaining 5 incremental tasks each representing the acoustic scenes from a different location. The proposed approach outperforms fine-tuning based methods and achieves an average accuracy of 48.8% after learning the last task in sequence without forgetting acoustic scenes from the previously learned locations.

【6】 Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
标题: 通过非负矩阵分解和探测进行可解释的按设计音频分割
作者:Martin Lebourdais,Théo Mariotte,Antonio Almudévar,Marie Tahon,Alfonso Ortega
备注:Accepted at Interspeech 2024, 5 pages, 2 figures, 3 tables
链接:点击下载PDF文件
摘要:音频分割是许多语音技术的关键任务,其中大多数基于神经网络,通常被认为是黑箱,具有高水平的性能。然而,在许多领域,其中包括健康或法医学,不仅需要良好的性能,而且还需要对输出决策的解释。直接从潜在表征导出的解释需要满足“好”的属性,例如信息量、紧凑性或模块性,才是可解释的。在这篇文章中,我们提出了一个基于非负矩阵分解(NMF)的可解释的设计音频分割模型,这是一个很好的候选人的设计可解释的表示。本文表明,我们的模型达到了良好的分割性能,并提出了深入的分析,从非负矩阵提取的潜在表示。所提出的方法打开了新的视角,根据“好”的属性对可解释的表示的评价。摘要:Audio segmentation is a key task for many speech technologies, most of which are based on neural networks, usually considered as black boxes, with high-level performances. However, in many domains, among which health or forensics, there is not only a need for good performance but also for explanations about the output decision. Explanations derived directly from latent representations need to satisfy "good" properties, such as informativeness, compactness, or modularity, to be interpretable. In this article, we propose an explainable-by-design audio segmentation model based on non-negative matrix factorization (NMF) which is a good candidate for the design of interpretable representations. This paper shows that our model reaches good segmentation performances, and presents deep analyses of the latent representation extracted from the non-negative matrix. The proposed approach opens new perspectives toward the evaluation of interpretable representations according to "good" properties.

【7】 Medical Spoken Named Entity Recognition
标题: 医疗口语命名实体识别
作者:Khai Le-Duc
备注:Preprint, 40 pages
链接:点击下载PDF文件
摘要:口语命名实体识别(NER)的目的是从语音中提取命名实体,并将其分类为人,位置,组织等类型。在这项工作中,我们提出了VietMed-NER -医疗领域的第一个口语NER数据集。据我们所知,就实体类型的数量而言,我们的真实世界数据集是世界上最大的口语NER数据集,具有18种不同的类型。其次,我们使用各种最先进的预训练模型呈现基线结果:仅编码器和序列到序列。我们发现,预训练的多语言模型XLM-R在参考文本和ASR输出上都优于所有单语言模型。同样在一般情况下,编码器执行NER任务比序列到序列模型。通过简单的翻译,文字记录不仅适用于越南语,也适用于其他语言。所有的代码、数据和模型都在这里公开发布:https: github.com leduckhai MultiMed摘要:Spoken Named Entity Recognition (NER) aims to extracting named entities from speech and categorizing them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our best knowledge, our real-world dataset is the largest spoken NER dataset in the world in terms of the number of entity types, featuring 18 distinct types. Secondly, we present baseline results using various state-of-the-art pre-trained models: encoder-only and sequence-to-sequence. We found that pre-trained multilingual models XLM-R outperformed all monolingual models on both reference text and ASR output. Also in general, encoders perform better than sequence-to-sequence models for the NER task. By simply translating, the transcript is applicable not just to Vietnamese but to other languages as well. All code, data and models are made publicly available here: https: github.com leduckhai MultiMed

【8】 Pushing the Limit of Sound Event Detection with Multi-Dilated Frequency Dynamic Convolution
标题: 利用多频动态卷积突破声音事件检测的极限
作者:Hyeonuk Nam,Yong-Hwa Park
链接:点击下载PDF文件
摘要:频率动态卷积(FDY conv)是声音事件检测(SED)领域的一个里程碑,但由于多个基核,它涉及模型大小的大幅增加。在这项工作中,我们提出了部分频率动态卷积(PFD conv),它连接静态传统的2D卷积分支输出和动态FDY conv分支输出,以尽量减少模型大小的增加,同时保持性能。此外,我们提出了多扩张频率动态卷积(MDFD conv),它集成了多个扩张频率动态卷积(DFD conv)分支与不同的扩张大小集和一个静态分支在一个单一的卷积模块,实现了3.2%的改善复调声音检测分数(PSDS)FDY conv。广泛消融研究的拟议方法进一步增强了对FDY conv变体的理解和可用性。摘要:Frequency dynamic convolution (FDY conv) has been a milestone in the sound event detection (SED) field, but it involves a substantial increase in model size due to multiple basis kernels. In this work, we propose partial frequency dynamic convolution (PFD conv), which concatenates static conventional 2D convolution branch output and dynamic FDY conv branch output in order to minimize model size increase while maintaining the performance. Additionally, we propose multi-dilated frequency dynamic convolution (MDFD conv), which integrates multiple dilated frequency dynamic convolution (DFD conv) branches with different dilation size sets and a static branch within a single convolution module, achieving a 3.2% improvement in polyphonic sound detection score (PSDS) over FDY conv. Proposed methods with extensive ablation studies further enhance understanding and usability of FDY conv variants.

【9】 CEC: A Noisy Label Detection Method for Speaker Recognition
标题: NEC:一种用于说话人识别的有噪标签检测方法
作者:Yao Shen,Yingying Gao,Yaqian Hao,Chenguang Hu,Fulin Zhang,Junlan Feng,Shilei Zhang
备注:interspeech 2024
链接:点击下载PDF文件
摘要:噪声标签是不可避免的,即使是在注释良好的数据集中。噪声标签的检测对于提高说话人识别模型的鲁棒性具有重要意义。在本文中,我们提出了一种新的噪声标签检测方法的基础上两个新的统计指标:连续不一致计数(CIC)和总不一致计数(TIC)。这些指标是通过跨时期计数(CEC)计算的,分别对应于训练的早期和后期阶段。此外,我们根据预测结果将样本分为三类:不一致的样本,硬样本和容易的样本。在训练过程中,我们逐渐增加硬样本更新模型参数的难度,防止噪声标签被过度拟合。与对比方案相比,我们的方法不仅在说话人确认方面取得了最好的性能,而且在噪声标签检测方面也表现出色。摘要:Noisy labels are inevitable, even in well-annotated datasets. The detection of noisy labels is of significant importance to enhance the robustness of speaker recognition models. In this paper, we propose a novel noisy label detection approach based on two new statistical metrics: Continuous Inconsistent Counting (CIC) and Total Inconsistent Counting (TIC). These metrics are calculated through Cross-Epoch Counting (CEC) and correspond to the early and late stages of training, respectively. Additionally, we categorize samples based on their prediction results into three categories: inconsistent samples, hard samples, and easy samples. During training, we gradually increase the difficulty of hard samples to update model parameters, preventing noisy labels from being overfitted. Compared to contrastive schemes, our approach not only achieves the best performance in speaker verification but also excels in noisy label detection.

【10】 Audio Fingerprinting with Holographic Reduced Representations
标题: 具有全息简化表示的音频指纹识别
作者:Yusuke Fujita,Tatsuya Komatsu
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:提出了一种基于全息简化表示的音频指纹模型。所提出的方法减少了存储的指纹的数量,而传统的神经音频指纹识别需要许多指纹的每个音轨,以实现高精度和时间分辨率。我们利用HRR通过循环卷积和求和将多个指纹聚合成一个复合指纹,从而减少与原始指纹具有相同维度空间的指纹。我们的搜索方法有效地找到一个组合指纹,其中存在一个查询指纹。利用HRR的逆运算,它可以恢复组合指纹内的相对位置,保持原始的时间分辨率。实验表明,该方法可以减少指纹的数量,同时保持适当的精度下降的时间分辨率,优于简单的抽取和求和为基础的聚合方法。摘要:This paper proposes an audio fingerprinting model with holographic reduced representation (HRR). The proposed method reduces the number of stored fingerprints, whereas conventional neural audio fingerprinting requires many fingerprints for each audio track to achieve high accuracy and time resolution. We utilize HRR to aggregate multiple fingerprints into a composite fingerprint via circular convolution and summation, resulting in fewer fingerprints with the same dimensional space as the original. Our search method efficiently finds a combined fingerprint in which a query fingerprint exists. Using HRR's inverse operation, it can recover the relative position within a combined fingerprint, retaining the original time resolution. Experiments show that our method can reduce the number of fingerprints with modest accuracy degradation while maintaining the time resolution, outperforming simple decimation and summation-based aggregation methods.

【11】 Articulatory Encodec: Vocal Tract Kinematics as a Codec for Speech
标题: 关节编码编码器:作为言语编码器的声道运动学
作者:Cheol Jun Cho,Peter Wu,Tejas S. Prabhune,Dhruv Agarwal,Gopala K. Anumanchipalli
链接:点击下载PDF文件
摘要:声道发音是一个自然的、有根据的言语产生控制空间。发音器官的时空协调与声源相结合,形成可理解的语音,以实现有效的口语交流。基于语音的这一生理基础,我们提出了一种新的神经语音编解码框架--发音编码器。发音编码器包括从语音音频推断发音特征的发音分析模型和从发音特征合成语音音频的发音合成模型。发音特征是发音器官的运动轨迹和声源特征,是语音产生的实际物理界面,具有直观的可解释性和可控性。一个额外的扬声器身份编码器与发音合成器联合训练,以告知各个扬声器的语音纹理。通过对大规模语音数据的训练,我们实现了一个完全可理解的,高质量的发音合成器,它可以推广到看不见的说话者。此外,扬声器嵌入有效地从发音中解脱出来,这使得能够实现口音保持的zero-shot语音转换。据我们所知,这是第一次演示的通用,高性能的发音推理和合成,建议作为一个强大的语音编码系统的框架。摘要:Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spoken communication. Based on this physiological grounding of speech, we propose a new framework of neural encoding-decoding of speech -- articulatory encodec. The articulatory encodec comprises an articulatory analysis model that infers articulatory features from speech audio, and an articulatory synthesis model that synthesizes speech audio from articulatory features. The articulatory features are kinematic traces of vocal tract articulators and source features, which are intuitively interpretable and controllable, being the actual physical interface of speech production. An additional speaker identity encoder is jointly trained with the articulatory synthesizer to inform the voice texture of individual speakers. By training on large-scale speech data, we achieve a fully intelligible, high-quality articulatory synthesizer that generalizes to unseen speakers. Furthermore, the speaker embedding is effectively disentangled from articulations, which enables accent-perserving zero-shot voice conversion. To the best of our knowledge, this is the first demonstration of universal, high-performance articulatory inference and synthesis, suggesting the proposed framework as a powerful coding system of speech.

【12】 Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
标题: 语音语言模型的指令数据生成和无监督适应
作者:Vahid Noroozi,Zhehuai Chen,Somshubra Majumdar,Steve Huang,Jagadeesh Balam,Boris Ginsburg
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了三种方法来生成合成样本,训练和评估多模态大型语言模型能够处理文本和语音输入。解决稀缺的样本包含两种模态,合成数据生成出现作为一个关键的策略,以提高此类系统的性能,并促进语音和文本域之间的跨模态关系的建模。我们的过程采用大型语言模型来生成文本组件和文本到语音系统来生成语音组件。所提出的方法为扩展这些模型的训练数据集提供了一种实用有效的手段。实验结果表明,在实现文本和语音的综合理解的进展。我们还强调了使用未标记的语音数据生成合成样本的潜力,这些样本的质量与可用的transmittance相当,使这些模型能够扩展到更多的语言。摘要:In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. Addressing the scarcity of samples containing both modalities, synthetic data generation emerges as a crucial strategy to enhance the performance of such systems and facilitate the modeling of cross-modal relationships between the speech and text domains. Our process employs large language models to generate textual components and text-to-speech systems to generate speech components. The proposed methods offer a practical and effective means to expand the training dataset for these models. Experimental results show progress in achieving an integrated understanding of text and speech. We also highlight the potential of using unlabeled speech data to generate synthetic samples comparable in quality to those with available transcriptions, enabling the expansion of these models to more languages.

【13】 Self-Train Before You Transcribe
标题: 抄写前自我训练
作者:Robert Flynn,Anton Ragni
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:当训练域和测试域之间存在不匹配时,当前的语音识别系统表现出显著的性能下降。自我训练方法,如嘈杂的学生教师培训,可以帮助解决这个问题,并使模型适应这种领域的变化。然而,自我训练通常需要收集未标记的目标域数据。对于设置,这是不切实际的,我们调查的好处进行嘈杂的学生教师培训的录音在测试集作为一个测试时间的适应方法。类似于语言建模中的动态评估方法,这使得跨话语边界的信息传递成为可能,并且作为域适应的方法发挥作用。一系列的域内和域外数据集用于实验,证明高达32.2%的相对增益。有趣的是,我们的方法显示出比利用单独适应数据的典型自我训练设置更大的增益。摘要:When there is a mismatch between the training and test domains, current speech recognition systems show significant performance degradation. Self-training methods, such as noisy student teacher training, can help address this and enable the adaptation of models under such domain shifts. However, self-training typically requires a collection of unlabelled target domain data. For settings where this is not practical, we investigate the benefit of performing noisy student teacher training on recordings in the test set as a test-time adaptation approach. Similarly to the dynamic evaluation approach in language modelling, this enables the transfer of information across utterance boundaries and functions as a method of domain adaptation. A range of in-domain and out-of-domain datasets are used for experiments demonstrating large relative gains of up to 32.2%. Interestingly, our method showed larger gains than the typical self-training setup that utilises separate adaptation data.

【14】 Automatic Speech Recognition for Biomedical Data in Bengali Language
标题: 孟加拉语生物医学数据的自动语音识别
作者:Shariar Kabir,Nazmun Nahar,Shyamasree Saha,Mamunur Rashid
链接:点击下载PDF文件
摘要:本文介绍了一个原型的自动语音识别(ASR)系统,专门为孟加拉生物医学数据的发展。孟加拉语ASR的最新进展令人鼓舞,但缺乏特定领域的数据限制了实际医疗保健ASR模型的创建。该项目通过开发针对孟加拉语医学术语(如症状,严重程度和疾病)的ASR系统来弥合这一差距,包括两种主要方言:孟加拉语和Sylheti。我们训练和评估两个流行的ASR框架上全面的46小时孟加拉语医学语料库。我们的核心目标是为数字健康应用创建可部署的健康领域ASR系统,最终提高医疗保健领域非技术用户的可访问性。摘要:This paper presents the development of a prototype Automatic Speech Recognition (ASR) system specifically designed for Bengali biomedical data. Recent advancements in Bengali ASR are encouraging, but a lack of domain-specific data limits the creation of practical healthcare ASR models. This project bridges this gap by developing an ASR system tailored for Bengali medical terms like symptoms, severity levels, and diseases, encompassing two major dialects: Bengali and Sylheti. We train and evaluate two popular ASR frameworks on a comprehensive 46-hour Bengali medical corpus. Our core objective is to create deployable health-domain ASR systems for digital health applications, ultimately increasing accessibility for non-technical users in the healthcare sector.

【15】 Disentangled Representation Learning for Environment-agnostic Speaker Recognition
标题: 用于环境不可知的说话人识别的分离表示学习
作者:KiHyun Nam,Hee-Soo Heo,Jee-weon Jung,Joon Son Chung
备注:Interspeech 2024. The official webpage can be found at this https URL
链接:点击下载PDF文件
摘要:这项工作提出了一个框架的基础上,功能解开学习扬声器嵌入是强大的环境变化。我们的框架利用自动编码器作为解缠器,将输入扬声器嵌入划分为与扬声器和其他残留信息相关的组件。我们采用了一组目标函数,以确保自动编码器的代码表示-用于细化嵌入-只压缩扬声器的特性。我们展示了我们的框架的多功能性,通过其与任何现有的扬声器嵌入提取器的兼容性,不需要结构修改或适应集成。我们验证我们的框架的有效性,将其纳入两个常用的嵌入提取器,并在各种基准进行实验。结果显示,性能提高高达16%。我们发布了我们的代码为这项工作提供https: github.com kaistmm voxceleb-disentangler摘要:This work presents a framework based on feature disentanglement to learn speaker embeddings that are robust to environmental variations. Our framework utilises an auto-encoder as a disentangler, dividing the input speaker embedding into components related to the speaker and other residual information. We employ a group of objective functions to ensure that the auto-encoder's code representation - used as the refined embedding - condenses only the speaker characteristics. We show the versatility of our framework through its compatibility with any existing speaker embedding extractor, requiring no structural modifications or adaptations for integration. We validate the effectiveness of our framework by incorporating it into two popularly used embedding extractors and conducting experiments across various benchmarks. The results show a performance improvement of up to 16%. We release our code for this work to be available https: github.com kaistmm voxceleb-disentangler

【16】 Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
标题: 第二届eXplainable AI for the Arts(XAIxArts)国际研讨会会议录
作者:Nick Bryan-Kinns,Corey Ford,Shuoyang Zheng,Helen Kennedy,Alan Chamberlain,Makayla Lewis,Drew Hemment,Zijin Li,Qiong Wu,Lanxi Xiao,Gus Xia,Jeba Rezwana,Michael Clemens,Gabriel Vigliensoni
链接:点击下载PDF文件
摘要:第二届可解释人工智能艺术(XAIxArts)国际研讨会汇集了HCI,交互设计,人工智能,可解释人工智能(XAI)和数字艺术的研究人员社区,以探索XAI对艺术的作用。第16届ACM创意与认知会议(C&C 2024),芝加哥,美国摘要:This second international workshop on explainable AI for the Arts (XAIxArts) brought together a community of researchers in HCI, Interaction Design, AI, explainable AI (XAI), and digital arts to explore the role of XAI for the Arts. Workshop held at the 16th ACM Conference on Creativity and Cognition (C&C 2024), Chicago, USA.

【17】 A Review of Common Online Speaker Diarization Methods
标题: 常用在线发言人拨号方法回顾
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages
链接:点击下载PDF文件
摘要:发言人日记提供了一个问题的答案“谁发言时?“获取音频文件。这些信息可用于完成音频转录,以供进一步处理。大多数发言人日记系统假设音频文件作为一个整体是可用的。然而,存在在音频片段到达之后立即需要扬声器标签的场景。具有相应低延迟的说话者日记被称为在线说话者日记。本文件提供了一个概述。首先简要介绍了在线演讲者日志化的历史。接下来给出了用于训练和评估的分类法和数据集。在接下来的部分中,将详细讨论在线日志化方法和系统。最后,本文提出了在线演讲者日志化领域未来研究需要解决的问题。摘要:Speaker diarization provides the answer to the question "who spoke when?" for an audio file. This information can be used to complete audio transcripts for further processing steps. Most speaker diarization systems assume that the audio file is available as a whole. However, there are scenarios in which the speaker labels are needed immediately after the arrival of an audio segment. Speaker diarization with a correspondingly low latency is referred to as online speaker diarization. This paper provides an overview. First the history of online speaker diarization is briefly presented. Next a taxonomy and datasets for training and evaluation are given. In the sections that follow, online diarization methods and systems are discussed in detail. This paper concludes with the presentation of challenges that still need to be solved by future research in the field of online speaker diarization.

【18】 LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
标题: LARP:冷启动播放列表延续的语言音频关系预训练
作者:Rebecca Salganik,Xiaohao Liu,Yunshan Ma,Jian Kang,Tat-Seng Chua
链接:点击下载PDF文件
摘要:随着在线音乐消费越来越多地转向基于播放列表的收听,播放列表延续的任务,其中算法建议歌曲以个性化和音乐上连贯的方式扩展播放列表,对于音乐流媒体的成功至关重要。目前,许多现有的播放列表连续方法依赖于协同过滤方法来执行推荐。然而,这些方法将难以推荐缺乏交互数据的歌曲,这就是所谓的冷启动问题。目前的方法,这一挑战设计复杂的机制,从稀疏的协作数据提取关系信号,并将它们集成到内容表示。然而,这些方法将内容表示学习排除在范围之外,并利用可能与特定音乐设置的分布或格式不一致的冻结的、预先训练的内容模型。此外,即使是最先进的音乐内容模块也是(1)与冷启动设置不兼容的,或者(2)不能有效地集成跨模态和关系信号。在本文中,我们引入LARP,多模态冷启动播放列表延续模型,有效地克服这些限制。LARP是一个三阶段的对比学习框架,它将多模态和关系信号整合到其学习的表征中。我们的框架使用越来越多的阶段的任务特定的抽象:内轨道(语言音频)对比损失,轨道轨道对比损失,和轨道播放列表对比损失。两个公开的数据集上的实验结果表明,在单模态和多模态模型的播放列表连续在冷启动设置的LARP的功效。代码和数据集发布于:https: github.com Rsalganik1123 LARP。摘要:As online music consumption increasingly shifts towards playlist-based listening, the task of playlist continuation, in which an algorithm suggests songs to extend a playlist in a personalized and musically cohesive manner, has become vital to the success of music streaming. Currently, many existing playlist continuation approaches rely on collaborative filtering methods to perform recommendation. However, such methods will struggle to recommend songs that lack interaction data, an issue known as the cold-start problem. Current approaches to this challenge design complex mechanisms for extracting relational signals from sparse collaborative data and integrating them into content representations. However, these approaches leave content representation learning out of scope and utilize frozen, pre-trained content models that may not be aligned with the distribution or format of a specific musical setting. Furthermore, even the musical state-of-the-art content modules are either (1) incompatible with the cold-start setting or (2) unable to effectively integrate cross-modal and relational signals. In this paper, we introduce LARP, a multi-modal cold-start playlist continuation model, to effectively overcome these limitations. LARP is a three-stage contrastive learning framework that integrates both multi-modal and relational signals into its learned representations. Our framework uses increasing stages of task-specific abstraction: within-track (language-audio) contrastive loss, track-track contrastive loss, and track-playlist contrastive loss. Experimental results on two publicly available datasets demonstrate the efficacy of LARP over uni-modal and multi-modal models for playlist continuation in a cold-start setting. Code and dataset are released at: https: github.com Rsalganik1123 LARP.

【19】 DASB -- Discrete Audio and Speech Benchmark
标题: DASB --离散音频和语音基准
作者:Pooneh Mousavi,Luca Della Libera,Jarod Duret,Artem Ploujnikov,Cem Subakan,Mirco Ravanelli
备注:9 pages, 5 tables
链接:点击下载PDF文件
摘要:离散音频令牌最近获得了相当大的关注,因为它们有可能连接音频和语言处理,从而能够创建现代多模态大型语言模型。理想的音频标记必须有效地保留语音和语义内容以及语言信息、说话者身份和其他细节。虽然最近已经提出了几种类型的音频标记,但由于现有研究中的评估设置不一致,因此确定各种任务的最佳标记器具有挑战性。为了解决这一差距,我们发布了离散音频和语音基准测试(DASB),这是一个全面的排行榜,用于在各种区分性任务中对离散音频令牌进行基准测试,包括语音识别,说话人识别和验证,情感识别,关键字定位和意图分类,以及语音增强,分离和文本到语音等生成任务。我们的研究结果表明,平均而言,语义标记在大多数判别和生成任务中的表现优于压缩标记。然而,语义标记和标准连续表示之间的性能差距仍然很大,突出了在这一领域进一步研究的必要性。摘要:Discrete audio tokens have recently gained considerable attention for their potential to connect audio and language processing, enabling the creation of modern multimodal large language models. Ideal audio tokens must effectively preserve phonetic and semantic content along with paralinguistic information, speaker identity, and other details. While several types of audio tokens have been recently proposed, identifying the optimal tokenizer for various tasks is challenging due to the inconsistent evaluation settings in existing studies. To address this gap, we release the Discrete Audio and Speech Benchmark (DASB), a comprehensive leaderboard for benchmarking discrete audio tokens across a wide range of discriminative tasks, including speech recognition, speaker identification and verification, emotion recognition, keyword spotting, and intent classification, as well as generative tasks such as speech enhancement, separation, and text-to-speech. Our results show that, on average, semantic tokens outperform compression tokens across most discriminative and generative tasks. However, the performance gap between semantic tokens and standard continuous representations remains substantial, highlighting the need for further research in this field.

【20】 SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
标题: SimSeamless:FBK参加IWSYS 2024同步语音翻译
作者:Sara Papi,Marco Gaido,Matteo Negri,Luisa Bentivogli
链接:点击下载PDF文件
摘要:本文介绍了FBK参与IWITH 2024同声传译评估活动的情况。对于今年提交的语音到文本翻译(ST)子轨道,我们提出了SimulSeamless,这是通过在其中等配置中结合AlignAtt和M4 T来实现的。无故障的M4 T模型是“现成的”,其同步推理是通过采用AlignAtt来实现的,AlignAtt是一种基于交叉注意的SimulST策略,可以在不对同步任务的底层模型进行任何重新训练或调整的情况下应用。我们参与了所有共享任务语言(英语- {德语,日语,中文}和捷克语- 英语),与去年的提交相比,取得了可接受的甚至更好的结果。SimulSeamless涵盖143种源语言和200种目标语言,发布于:https: github.com hlt-mt FBK-fairseq 。摘要:This paper describes the FBK's participation in the Simultaneous Translation Evaluation Campaign at IWSLT 2024. For this year's submission in the speech-to-text translation (ST) sub-track, we propose SimulSeamless, which is realized by combining AlignAtt and SeamlessM4T in its medium configuration. The SeamlessM4T model is used "off-the-shelf" and its simultaneous inference is enabled through the adoption of AlignAtt, a SimulST policy based on cross-attention that can be applied without any retraining or adaptation of the underlying model for the simultaneous task. We participated in all the Shared Task languages (English- {German, Japanese, Chinese}, and Czech- English), achieving acceptable or even better results compared to last year's submissions. SimulSeamless, covering more than 143 source languages and 200 target languages, is released at: https: github.com hlt-mt FBK-fairseq .

【21】 A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
标题: 用于视听深度伪造检测的单类学习多流融合方法
作者:Kyungbok Lee,You Zhang,Zhiyao Duan
链接:点击下载PDF文件
摘要:本文解决了开发一个强大的音频-视觉深度伪造检测模型的挑战。在实际用例中,新一代算法不断涌现,并且在检测方法的开发过程中不会遇到这些算法。这就要求该方法具有泛化能力。此外,为了确保检测方法的可信度,模型解释视频中哪些线索表明它是假的是有益的。出于这些考虑,我们提出了一种多流融合方法,将单类学习作为表示级正则化技术。我们通过扩展和重新分割现有的FakeAVCeleb数据集来创建一个新的基准来研究视听deepfake检测的泛化问题。该基准测试包含四个类别的假视频(真实音频-假视觉、假音频-假视觉、假音频-真实视觉和不同步视频)。实验结果表明,与基线模型相比,我们的方法在四个测试集上平均提高了7.31%的模型对不可见攻击的检测。此外,我们提出的框架提供了可解释性,表明模型将哪种模态识别为假模态。摘要:This paper addresses the challenge of developing a robust audio-visual deepfake detection model. In practical use cases, new generation algorithms are continually emerging, and these algorithms are not encountered during the development of detection methods. This calls for the generalization ability of the method. Additionally, to ensure the credibility of detection methods, it is beneficial for the model to interpret which cues from the video indicate it is fake. Motivated by these considerations, we then propose a multi-stream fusion approach with one-class learning as a representation-level regularization technique. We study the generalization problem of audio-visual deepfake detection by creating a new benchmark by extending and re-splitting the existing FakeAVCeleb dataset. The benchmark contains four categories of fake video(Real Audio-Fake Visual, Fake Audio-Fake Visual, Fake Audio-Real Visual, and unsynchronized video). The experimental results show that our approach improves the model's detection of unseen attacks by an average of 7.31% across four test sets, compared to the baseline model. Additionally, our proposed framework offers interpretability, indicating which modality the model identifies as fake.

【22】 Seamless Language Expansion: Enhancing Multilingual Mastery in Self-Supervised Models
标题: 无缝语言扩展:增强自我监督模型中的多语言掌握
作者:Jing Xu,Minglin Wu,Xixin Wu,Helen Meng
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自监督(SSL)模型在各种下游任务中表现出出色的性能。然而,它们通常是为有限的语言开发的,并且在现实世界中可能会遇到新的语言。为每种新语言开发SSL模型的成本很高。因此,找出如何有效地使现有的SSL模型适应新的语言而不损害其原有的能力至关重要。我们提出了适应方法,将LoRA集成到现有的SSL模型中,以扩展新的语言。我们还开发了保存策略,包括数据组合和重新聚类,以保留现有语言的能力。应用于mHuBERT,我们研究了它们在语音再合成任务上的有效性。实验结果表明,我们的自适应方法使mHuBERT适用于一个新的语言(普通话)的MOS值增加约1.6和WER的相对值降低高达61.72%。此外,我们的保存策略确保现有语言和新语言的性能保持不变。摘要:Self-supervised (SSL) models have shown great performance in various downstream tasks. However, they are typically developed for limited languages, and may encounter new languages in real-world. Developing a SSL model for each new language is costly. Thus, it is vital to figure out how to efficiently adapt existed SSL models to a new language without impairing its original abilities. We propose adaptation methods which integrate LoRA to existed SSL models to extend new language. We also develop preservation strategies which include data combination and re-clustering to retain abilities on existed languages. Applied to mHuBERT, we investigate their effectiveness on speech re-synthesis task. Experiments show that our adaptation methods enable mHuBERT to be applied to a new language (Mandarin) with MOS value increased about 1.6 and the relative value of WER reduced up to 61.72%. Also, our preservation strategies ensure that the performance on both existed and new languages remains intact.

【23】 Improved Remixing Process for Domain Adaptation-Based Speech Enhancement by Mitigating Data Imbalance in Signal-to-Noise Ratio
标题: 通过缓解信号噪比中的数据不平衡来改进基于域自适应的语音增强的重新混合过程
作者:Li Li,Shogo Seki
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:RemixIT和Remixed 2 Remixed是基于域自适应的语音增强(DASE)方法,其使用在完全监督下训练的教师模型,通过重新混合教师模型的输出来生成伪配对数据。用于增强真实世界记录信号的学生模型使用没有地面实况的伪配对数据进行训练。由于噪声信号是在自然环境中记录的,因此数据集不可避免地会在某些声学特性中出现数据不平衡,从而导致代表性不足的数据的性能低于标准。在监督学习中固有平衡的信噪比(SNR)就是一个很好的例子。在本文中,我们使用CHiME-7 UDASE任务的数据集提供了经验证据,证明伪数据的SNR对模型性能有显着影响,突出了DASE中平衡SNR的重要性。此外,我们建议采用课程学习来涵盖广泛的SNR,以提高代表性不足的数据的性能。摘要:RemixIT and Remixed2Remixed are domain adaptation-based speech enhancement (DASE) methods that use a teacher model trained in full supervision to generate pseudo-paired data by remixing the outputs of the teacher model. The student model for enhancing real-world recorded signals is trained using the pseudo-paired data without ground truth. Since the noisy signals are recorded in natural environments, the dataset inevitably suffers data imbalance in some acoustic properties, leading to subpar performance for the underrepresented data. The signal-to-noise ratio (SNR), inherently balanced in supervised learning, is a prime example. In this paper, we provide empirical evidence that the SNR of pseudo data has a significant impact on model performance using the dataset of the CHiME-7 UDASE task, highlighting the importance of balanced SNR in DASE. Furthermore, we propose adopting curriculum learning to encompass a broad range of SNRs to boost performance for underrepresented data.

【24】 Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control
标题: 空中交通管制中的关节与顺序说话者角色检测和自动语音识别
作者:Alexander Blatt,Aravind Krishnan,Dietrich Klakow
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:将空中交通管制(ATC)数据用于下游自然语言处理任务需要预处理步骤。关键步骤是通过自动语音识别(ASR)和说话人日志化(Speaker Diarization)转录数据,分别使用说话人角色检测(SRD)将转录本分为飞行员和空中交通管制员(ATCO)转录本。虽然传统方法分别承担这些任务,但我们提出了一种基于变压器的联合ASR-SRD系统,该系统在依赖标准ASR架构的同时联合解决了这两项任务。我们比较这个联合系统对两个级联的方法ASR和SRD多ATC数据集。我们的研究表明,在这种情况下,我们的联合系统可以优于两个传统的方法,在这种情况下,其他架构是优选的。我们还评估了声学和词汇差异如何影响所有架构,并展示了如何克服它们的联合架构。摘要:Utilizing air-traffic control (ATC) data for downstream natural-language processing tasks requires preprocessing steps. Key steps are the transcription of the data via automatic speech recognition (ASR) and speaker diarization, respectively speaker role detection (SRD) to divide the transcripts into pilot and air-traffic controller (ATCO) transcripts. While traditional approaches take on these tasks separately, we propose a transformer-based joint ASR-SRD system that solves both tasks jointly while relying on a standard ASR architecture. We compare this joint system against two cascaded approaches for ASR and SRD on multiple ATC datasets. Our study shows in which cases our joint system can outperform the two traditional approaches and in which cases the other architectures are preferable. We additionally evaluate how acoustic and lexical differences influence all architectures and show how to overcome them for our joint architecture.

【25】 Automated Bioacoustic Monitoring for South African Bird Species on Unlabeled Data
标题: 基于未标记数据的南非鸟类自动生物声学监测
作者:Michael Doell,Dominik Kuehn,Vanessa Suessle,Matthew J. Burnett,Colleen T. Downs,Andreas Weinmann,Elke Hergenroether
Journal-ref:International Conferences in Central Europe on Computer Graphics, Visualization and Computer Vision 2024
链接:点击下载PDF文件
摘要:基于被动声监测(PAM)记录的生物多样性监测分析是耗时的,并受到记录中存在背景噪声的挑战。现有的声音事件检测(SED)模型只适用于某些鸟类,进一步模型的开发需要标记数据。开发的框架自动提取标记的数据,从现有的平台上选定的鸟类物种。标记的数据被嵌入到录音中,包括环境声音和噪声,并用于训练卷积递归神经网络(CRNN)模型。模型进行了评估,在城市夸祖鲁-纳塔尔栖息地记录的未加工的真实世界的数据。自适应SED-CRNN模型的F1得分为0.73,证明了其在嘈杂的真实世界条件下的效率。所提出的方法,自动提取标记的数据,为选定的鸟类物种,使PAM容易适应其他物种和栖息地,为未来的保护项目。摘要:Analyses for biodiversity monitoring based on passive acoustic monitoring (PAM) recordings is time-consuming and challenged by the presence of background noise in recordings. Existing models for sound event detection (SED) worked only on certain avian species and the development of further models required labeled data. The developed framework automatically extracted labeled data from available platforms for selected avian species. The labeled data were embedded into recordings, including environmental sounds and noise, and were used to train convolutional recurrent neural network (CRNN) models. The models were evaluated on unprocessed real world data recorded in urban KwaZulu-Natal habitats. The Adapted SED-CRNN model reached a F1 score of 0.73, demonstrating its efficiency under noisy, real-world conditions. The proposed approach to automatically extract labeled data for chosen avian species enables an easy adaption of PAM to other species and habitats for future conservation projects.

【26】 ManWav: The First Manchu ASR Model
标题: ManWav:第一个满语ASC模型
作者:Jean Seo,Minha Kang,Sungjoo Byun,Sangah Lee
备注:ACL2024Field Matters
链接:点击下载PDF文件
摘要:本研究旨在解决高资源和极低资源语言之间自动语音识别(ASR)研究中日益扩大的差距,特别关注满语,一种极度濒危的语言。满语体现了边缘化语言社区在获得最先进技术方面所面临的挑战。作为开创性的努力,我们引入了有史以来第一个满语ASR模型ManWav,利用Wav 2 Vec 2-XLSR-53。第一个满语ASR的结果是有希望的,特别是在使用我们的增强数据进行训练时。与使用原始数据微调的相同基础模型相比,使用增强数据微调的Wav 2 Vec 2-XLSR-53显示CER下降0.02,WER下降0.13。摘要:This study addresses the widening gap in Automatic Speech Recognition (ASR) research between high resource and extremely low resource languages, with a particular focus on Manchu, a critically endangered language. Manchu exemplifies the challenges faced by marginalized linguistic communities in accessing state-of-the-art technologies. In a pioneering effort, we introduce the first-ever Manchu ASR model ManWav, leveraging Wav2Vec2-XLSR-53. The results of the first Manchu ASR is promising, especially when trained with our augmented data. Wav2Vec2-XLSR-53 fine-tuned with augmented data demonstrates a 0.02 drop in CER and 0.13 drop in WER compared to the same base model fine-tuned with original data.

【27】 Children's Speech Recognition through Discrete Token Enhancement
标题: 通过离散令牌增强实现儿童语音识别
作者:Vrunda N. Sukhadia,Shammur Absar Chowdhury
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:儿童的语音识别被认为是一个低资源的任务,主要是由于缺乏公开的数据。造成这种数据稀缺的原因有几个,包括昂贵的数据收集和注释过程,以及数据隐私等。将语音信号转换为不携带敏感信息但捕获语言和声学信息的离散令牌可能是隐私问题的解决方案。在这项研究中,我们调查的离散语音令牌集成到儿童的语音识别系统作为输入,而不会显着降低ASR性能。此外,我们探索了创建这些离散标签的单视图和多视图策略。此外,我们测试了模型的泛化能力与看不见的域和出生数据集。结果表明,儿童的离散令牌ASR实现了几乎相同的性能,参数减少了约83%。摘要:Children's speech recognition is considered a low-resource task mainly due to the lack of publicly available data. There are several reasons for such data scarcity, including expensive data collection and annotation processes, and data privacy, among others. Transforming speech signals into discrete tokens that do not carry sensitive information but capture both linguistic and acoustic information could be a solution for privacy concerns. In this study, we investigate the integration of discrete speech tokens into children's speech recognition systems as input without significantly degrading the ASR performance. Additionally, we explored single-view and multi-view strategies for creating these discrete labels. Furthermore, we tested the models for generalization capabilities with unseen domain and nativity dataset. Results reveal that the discrete token ASR for children achieves nearly equivalent performance with an approximate 83% reduction in parameters.

【28】 Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
标题: 直接通过Gumbel Softmax估计器的双峰神经架构搜索用于视听深度伪造检测
作者:Aravinda Reddy PN,Raghavendra Ramachandra,Krothapalli Sreenivasa Rao,Pabitra Mitra,Vinod Rathod
链接:点击下载PDF文件
摘要:Deepfakes是生物识别认证的主要安全风险。这项技术创造了逼真的假视频,可以冒充真人,欺骗依赖面部特征和语音模式进行识别的系统。现有的多模式deepfake检测器依赖于传统的融合方法,例如多数规则和集成投票,这些方法通常难以适应不断变化的数据特征和复杂模式。在本文中,我们介绍了直通Gumbel-Softmax(STGS)框架,提供了一个全面的方法来搜索多模态融合模型架构。使用两级搜索方法,该框架优化了网络结构,参数和性能。最初,从骨干网络中有效地识别关键特征,而在细胞结构中,加权融合操作集成了来自各种来源的信息。通过改变温度和采样时间等参数,得到了一个最大限度地提高分类性能的架构。在FakeAVCeleb和SWAN-DF数据集上的实验结果表明,以最小的模型参数实现了令人印象深刻的AUC值94.4%。摘要:Deepfakes are a major security risk for biometric authentication. This technology creates realistic fake videos that can impersonate real people, fooling systems that rely on facial features and voice patterns for identification. Existing multimodal deepfake detectors rely on conventional fusion methods, such as majority rule and ensemble voting, which often struggle to adapt to changing data characteristics and complex patterns. In this paper, we introduce the Straight-through Gumbel-Softmax (STGS) framework, offering a comprehensive approach to search multimodal fusion model architectures. Using a two-level search approach, the framework optimizes the network architecture, parameters, and performance. Initially, crucial features were efficiently identified from backbone networks, whereas within the cell structure, a weighted fusion operation integrated information from various sources. An architecture that maximizes the classification performance is derived by varying parameters such as temperature and sampling time. The experimental results on the FakeAVCeleb and SWAN-DF datasets demonstrated an impressive AUC value 94.4 % achieved with minimal model parameters.

【29】 Transferable speech-to-text large language model alignment module
标题: 可移植语音到文本大型语言模型对齐模块
作者:Boyong Wu,Chao Yan,Haoran Pu
备注:Accepted by InterSpeech 2024; 5 pages, 2 figures
链接:点击下载PDF文件
摘要:通过利用大型语言模型(LLM)和语音基础模型的强大功能,最先进的语音-文本双峰作品可以实现具有挑战性的任务,如口语翻译(ST)和问答(SQA)。在本文中,我们利用Whisper编码器和预训练的Yi-6 B的能力。实验结果表明,模态对齐可以实现一层模块和语音文本多任务语料库的百小时。我们在推理过程中进一步将Yi-6 B与Yi-6 B-Chat的人类偏好对齐版本交换,并发现对齐能力也适用。此外,通过奇异值分解(SVD)揭示的对齐子空间也意味着线性对齐子空间是稀疏的,这使得连接其他特征(如声纹或视频)以扩展模态的可能性。摘要:By leveraging the power of Large Language Models(LLMs) and speech foundation models, state of the art speech-text bimodal works can achieve challenging tasks like spoken translation(ST) and question answering(SQA) altogether with much simpler architectures. In this paper, we utilize the capability of Whisper encoder and pre-trained Yi-6B. Empirical results reveal that modal alignment can be achieved with one layer module and hundred hours of speech-text multitask corpus. We further swap the Yi-6B with human preferences aligned version of Yi-6B-Chat during inference, and discover that the alignment capability is applicable as well. In addition, the alignment subspace revealed by singular value decomposition(SVD) also implies linear alignment subspace is sparse, which leaves the possibility to concatenate other features like voice-print or video to expand modality.

【30】 SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
标题: SD-Eval:口语对话理解超越言语的基准数据集
作者:Junyi Ao,Yuancheng Wang,Xiaohai Tian,Dekun Chen,Jun Zhang,Lu Lu,Yuxuan Wang,Haizhou Li,Zhizheng Wu
链接:点击下载PDF文件
摘要:语音包含丰富的信息,包括但不限于内容、语言和环境信息。语音的这种综合性质对通信产生了重大影响,对人机交互至关重要。面向聊天的大型语言模型(LLM)以其通用的辅助功能而闻名,已经发展到可以处理包括语音在内的多模态输入。虽然这些模型可以熟练地识别和分析语音,但它们通常无法产生适当的响应。我们认为,这是由于缺乏任务定义和模型开发的原则,这需要开源数据集和适合模型评估的指标。为了弥合这一差距,我们提出了SD-Eval,这是一个基准数据集,旨在对口语对话的理解和生成进行多维评估。SD-Eval侧重于语言和环境信息,包括7,303个话语,相当于8.76小时的语音数据。这些数据来自八个公共数据集,代表四个方面:情感,口音,年龄和背景声音。为了评估SD-Eval基准数据集,我们实现了三种不同的模型,并按照与SD-Eval类似的过程构建了一个训练集。训练集包含1,052.72小时的语音数据和724.4k话语。我们还使用客观评估方法(例如BLEU和ROUGE),主观评估和基于LLM的指标对生成的响应进行综合评估。以语言和环境信息为条件的模型在客观和主观测量中都优于它们的同行。此外,实验表明,基于LLM的指标显示出更高的相关性与人类的评价相比,传统的指标。我们在https: github.com amphionspace SD-Eval上开源了SD-Eval。摘要:Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information. This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction. Chat-Oriented Large Language Models (LLMs), known for their general-purpose assistance capabilities, have evolved to handle multi-modal inputs, including speech. Although these models can be adept at recognizing and analyzing speech, they often fall short of generating appropriate responses. We argue that this is due to the lack of principles on task definition and model development, which requires open-source datasets and metrics suitable for model evaluation. To bridge the gap, we present SD-Eval, a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation. SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound. To assess the SD-Eval benchmark dataset, we implement three different models and construct a training set following a similar process as SD-Eval. The training set contains 1,052.72 hours of speech data and 724.4k utterances. We also conduct a comprehensive evaluation using objective evaluation methods (e.g. BLEU and ROUGE), subjective evaluations and LLM-based metrics for the generated responses. Models conditioned with paralinguistic and environmental information outperform their counterparts in both objective and subjective measures. Moreover, experiments demonstrate LLM-based metrics show a higher correlation with human evaluation compared to traditional metrics. We open-source SD-Eval at https: github.com amphionspace SD-Eval.

【31】 Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
标题: 通过优化的音频编码通过大型语言模型增强自动音频字幕
作者:Jizhong Liu,Gang Li,Junbo Zhang,Heinrich Dinkel,Yongqing Wang,Zhiyong Yan,Yujun Wang,Bin Wang
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自动音频字幕(AAC)是一种以自然语言描述音频内容的音频到文本的任务。最近,大型语言模型(LLM)的进步,以及音频编码器训练方法的改进,为改进AAC开辟了可能性。因此,本文从三个方面对AAC的增强进行了探索:1)通过一致性集成蒸馏(CED)的预训练音频编码器来提高声学标记的有效性,并使用查询Transformer(Q-Former)弥补了LLM和压缩声学标记的模态差距; 2)我们研究了使用具有7 B参数的Llama 2作为解码器的优点; 3)另一个预训练的LLM校正由训练数据不足和注释歧义引起的文本错误。音频编码器和文本解码器都通过-Base(LoRA)进行了优化。实验表明,这些改进都是有效的。我们的方法获得了33.0 SPIDER-FL分数,优于DCASE 2023任务6A的获胜者。摘要:Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improvements in training approaches for audio encoders, have opened up possibilities for improving AAC. Thus, we explore enhancing AAC from three aspects: 1) a pre-trained audio encoder via consistent ensemble distillation (CED) is used to improve the effectivity of acoustic tokens, with a querying transformer (Q-Former) bridging the modality gap to LLM and compress acoustic tokens; 2) we investigate the advantages of using a Llama 2 with 7B parameters as the decoder; 3) another pre-trained LLM corrects text errors caused by insufficient training data and annotation ambiguities. Both the audio encoder and text decoder are optimized by -Base (LoRA). Experiments show that each of these enhancements is effective. Our method obtains a 33.0 SPIDEr-FL score, outperforming the winner of DCASE 2023 Task 6A.

【32】 Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
标题: 利用尖峰神经网络进行全局-局部卷积以实现节能关键词发现
作者:Shuai Wang,Dehao Zhang,Kexin Shi,Yuchen Wang,Wenjie Wei,Jibin Wu,Malu Zhang
链接:点击下载PDF文件
摘要:得益于深度神经网络(DNN),关键词识别(KWS)的准确性已经取得了实质性的进展。然而,由于KWS系统通常在边缘设备上实现,因此除了性能之外,能源效率也成为一项关键要求。在这里,我们利用尖峰神经网络的能量效率,并提出了一个端到端的轻量级KWS模型。该模型由两个创新模块组成:1)全局-局部尖峰卷积(GLSC)模块和2)瓶颈PLIF模块。与手工特征提取方法相比,GLSC模块实现了更稀疏、更节能、性能更好的语音特征提取。Bottleneck-PLIF模块进一步处理来自GLSC的信号,目的是用更少的参数实现更高的精度。在Google语音命令数据集(V1和V2)上进行了大量的实验。结果表明,我们的方法取得了竞争力的性能基于SNN的KWS模型与较少的参数。摘要:Thanks to Deep Neural Networks (DNNs), the accuracy of Keyword Spotting (KWS) has made substantial progress. However, as KWS systems are usually implemented on edge devices, energy efficiency becomes a critical requirement besides performance. Here, we take advantage of spiking neural networks' energy efficiency and propose an end-to-end lightweight KWS model. The model consists of two innovative modules: 1) Global-Local Spiking Convolution (GLSC) module and 2) Bottleneck-PLIF module. Compared to the hand-crafted feature extraction methods, the GLSC module achieves speech feature extraction that is sparser, more energy-efficient, and yields better performance. The Bottleneck-PLIF module further processes the signals from GLSC with the aim to achieve higher accuracy with fewer parameters. Extensive experiments are conducted on the Google Speech Commands Dataset (V1 and V2). The results show our method achieves competitive performance among SNN-based KWS models with fewer parameters.


机器翻译,仅供参考