今日论文合集:cs.SD语音11篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载

cs.SD语音

【1】Communication-Efficient Personalized Federated Learning for  Speech-to-Text Tasks

标题:用于语音到文本任务的通信高效的个性化联合学习

链接:https://arxiv.org/abs/2401.10070

作者:Yichao Du,Zhirui Zhang,Linan Yue,Xu Huang,Yuqing Zhang,Tong Xu,Linli Xu,Enhong Chen

备注:ICASSP 2024

摘要:为了保护隐私和满足法律法规,联邦学习(FL)在训练语音到文本(S2 T)系统方面获得了极大的关注,包括自动语音识别(ASR)和语音翻译(ST)。然而,通常使用的FL方法(即,\textsc{FedAvg})存在大量的通信开销和客户端数据异构性导致的性能下降。为了解决这些问题,提出了一种个性化的联邦S2 T框架,该框架引入了\textsc{FedLoRA},一个轻量级的LoRA模块,用于客户端调优和与服务器的交互,以最大限度地减少通信开销,和\textsc{FedMem},这是一个全局模型,配备了$k$-最近邻($k$NN)分类器,可捕获特定于客户端的分布变化,以实现个性化并克服数据异构性。基于Conformer和Whisper骨干模型在CoVoST和GigaSpeech基准上进行的大量实验表明,我们的方法显着降低了所有S2 T任务的通信开销,并有效地个性化全局模型,以克服数据异构性。

摘要:To protect privacy and meet legal regulations, federated learning (FL) has gained significant attention for training speech-to-text (S2T) systems, including automatic speech recognition (ASR) and speech translation (ST). However, the commonly used FL approach (i.e., \textsc{FedAvg}) in S2T tasks typically suffers from extensive communication overhead due to multi-round interactions based on the whole model and performance degradation caused by data heterogeneity among clients.To address these issues, we propose a personalized federated S2T framework that introduces \textsc{FedLoRA}, a lightweight LoRA module for client-side tuning and interaction with the server to minimize communication overhead, and \textsc{FedMem}, a global model equipped with a $k$-nearest-neighbor ($k$NN) classifier that captures client-specific distributional shifts to achieve personalization and overcome data heterogeneity. Extensive experiments based on Conformer and Whisper backbone models on CoVoST and GigaSpeech benchmarks show that our approach significantly reduces the communication overhead on all S2T tasks and effectively personalizes the global model to overcome data heterogeneity.


【2】 Developing an AI-based Integrated System for Bee Health Evaluation
标题:基于人工智能的蜜蜂健康综合评价系统的开发
链接:https://arxiv.org/abs/2401.09988
作者:Andrew Liang
摘要:蜜蜂授粉约占世界粮食供应的三分之一,但由于杀虫剂和害虫等多种因素,蜂群数量在过去十年中惊人地下降了近40%。监测蜂箱的传统方法,如人工检查,是主观的,破坏性的,耗时的。为了克服这些局限性,人工智能已被用于评估蜂巢健康。然而,以前的研究缺乏端到端的解决方案,主要依赖于来自单一来源的数据,无论是蜜蜂图像还是声音。本研究介绍一个综合系统,包括蜜蜂目标检测和健康评估。此外,它利用视觉和音频信号的组合来分析蜜蜂的行为。基于注意力的多模态神经网络(AMNN)被开发用于自适应地关注来自每种类型的信号的关键特征,以进行准确的蜜蜂健康评估。AMNN的总体准确率达到92.61%,超过了现有的8个单信号卷积神经网络和递归神经网络。它比最好的基于图像的模型高出32.51%,比最好的基于声音的模型高出13.98%,同时保持了有效的处理时间。此外,它还提高了预测的鲁棒性,在所有四种评估的健康状况中,F1得分均高于90%。该研究还表明,音频信号比图像更可靠地评估蜜蜂健康。通过将AMNN与图像和声音数据无缝集成在一个全面的蜜蜂健康监测系统中,这种方法为蜜蜂疾病的早期检测和蜂群的保护提供了一种更有效和非侵入性的解决方案。
摘要:Honey bees pollinate about one-third of the world's food supply, but bee colonies have alarmingly declined by nearly 40% over the past decade due to several factors, including pesticides and pests. Traditional methods for monitoring beehives, such as human inspection, are subjective, disruptive, and time-consuming. To overcome these limitations, artificial intelligence has been used to assess beehive health. However, previous studies have lacked an end-to-end solution and primarily relied on data from a single source, either bee images or sounds. This study introduces a comprehensive system consisting of bee object detection and health evaluation. Additionally, it utilized a combination of visual and audio signals to analyze bee behaviors. An Attention-based Multimodal Neural Network (AMNN) was developed to adaptively focus on key features from each type of signal for accurate bee health assessment. The AMNN achieved an overall accuracy of 92.61%, surpassing eight existing single-signal Convolutional Neural Networks and Recurrent Neural Networks. It outperformed the best image-based model by 32.51% and the top sound-based model by 13.98% while maintaining efficient processing times. Furthermore, it improved prediction robustness, attaining an F1-score higher than 90% across all four evaluated health conditions. The study also shows that audio signals are more reliable than images for assessing bee health. By seamlessly integrating AMNN with image and sound data in a comprehensive bee health monitoring system, this approach provides a more efficient and non-invasive solution for the early detection of bee diseases and the preservation of bee colonies.

【3】 Attention-Based Recurrent Neural Network For Automatic Behavior Laying  Hen Recognition
标题:基于注意力的递归神经网络用于蛋鸡行为自动识别
链接:https://arxiv.org/abs/2401.09880
作者:Fréjus A. A. Laleye,Mikaël A. Mousse
摘要:现代家禽养殖的兴趣之一是蛋鸡的发声,其中包含非常有用的健康行为信息。这些信息被用作健康和福利指标,帮助育种者更好地监测蛋鸡,这涉及到早期发现问题,以便进行快速和更有效的干预。在这项工作中,我们专注于声音分析识别的蛋鸡的呼叫类型,以提出一个强大的系统表征他们的行为,更好地监测。为此,我们首先收集并标注蛋鸡叫声信号,然后基于时域和频域特征的组合设计了最佳声学表征。然后,我们使用这些特征来建立基于递归神经网络的多标签分类模型,为表征蛋鸡行为的发声分配语义类。结果显示,基于时域和频域特征组合的模型的整体性能获得了最高的F1分数(F1=92.75),使用频域特征的模型增益为17%,文献中比较的方法增益为8%。
摘要:One of the interests of modern poultry farming is the vocalization of laying hens which contain very useful information on health behavior. This information is used as health and well-being indicators that help breeders better monitor laying hens, which involves early detection of problems for rapid and more effective intervention. In this work, we focus on the sound analysis for the recognition of the types of calls of the laying hens in order to propose a robust system of characterization of their behavior for a better monitoring. To do this, we first collected and annotated laying hen call signals, then designed an optimal acoustic characterization based on the combination of time and frequency domain features. We then used these features to build the multi-label classification models based on recurrent neural network to assign a semantic class to the vocalization that characterize the laying hen behavior. The results show an overall performance with our model based on the combination of time and frequency domain features that obtained the highest F1-score (F1=92.75) with a gain of 17% on the models using the frequency domain features and of 8% on the compared approaches from the litterature.


【4】 On the Audio Hallucinations in Large Audio-Video Language Models
标题:论大型视听语言模型中的幻听现象
链接:https://arxiv.org/abs/2401.09774
作者:Taichi Nishimura,Shota Nakada,Masayoshi Kondo
备注:6 pages
摘要:大型音视频语言模型可以生成视频和音频的描述。然而,它们有时会忽略音频内容,仅依靠视觉信息制作音频描述。本文将其称为音频幻觉,并在大型音频视频语言模型中对其进行分析。我们通过询问音频信息收集了1,000个句子,并对它们是否包含幻觉进行注释。如果一个句子是幻觉,我们也会对幻觉的类型进行分类。结果表明,332个句子是幻觉与不同的趋势观察到的名词和动词的幻觉类型。基于此,我们解决了一个任务的音频幻觉分类使用预训练的音频文本模型在zero-shot和微调设置。我们的实验结果表明,zero-shot模型实现了更高的性能(52.2%,在F1)比随机(40.3%)和微调模型达到87.9%,优于zero-shot模型。
摘要:Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on visual information. This paper refers to this as audio hallucinations and analyzes them in large audio-video language models. We gather 1,000 sentences by inquiring about audio information and annotate them whether they contain hallucinations. If a sentence is hallucinated, we also categorize the type of hallucination. The results reveal that 332 sentences are hallucinated with distinct trends observed in nouns and verbs for each hallucination type. Based on this, we tackle a task of audio hallucination classification using pre-trained audio-text models in the zero-shot and fine-tuning settings. Our experimental results reveal that the zero-shot models achieve higher performance (52.2% in F1) than the random (40.3%) and the fine-tuning models achieve 87.9%, outperforming the zero-shot models.

【5】 SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech  Recognition
标题:SlideAVSR:用于视听语音识别的纸质讲解视频数据集
链接:https://arxiv.org/abs/2401.09759
作者:Hao Wang,Shuhei Kurita,Shuichiro Shimizu,Daisuke Kawahara
摘要:视听语音识别(AVSR)是自动语音识别(ASR)的多模态扩展,使用视频作为音频的补充。在AVSR中,相当大的努力已经针对面部特征的数据集,如唇读,而他们往往在更广泛的背景下评估图像理解能力不足。在本文中,我们构建了SlideAVSR,AVSR数据集使用科学论文解释视频。SlideAVSR提供了一个新的基准,模型在演示记录的幻灯片上转录语音话语和文本。众所周知,在没有参考文本的情况下,经常出现在论文解释中的技术术语很难转录,因此我们的SlideAVSR数据集突出了AVSR问题的一个新方面。作为一个简单而有效的基线,我们提出了DocWhisper,一个AVSR模型,可以引用幻灯片中的文本信息,并确认其在SlideAVSR上的有效性。
摘要:Audio-visual speech recognition (AVSR) is a multimodal extension of automatic speech recognition (ASR), using video as a complement to audio. In AVSR, considerable efforts have been directed at datasets for facial features such as lip-readings, while they often fall short in evaluating the image comprehension capabilities in broader contexts. In this paper, we construct SlideAVSR, an AVSR dataset using scientific paper explanation videos. SlideAVSR provides a new benchmark where models transcribe speech utterances with texts on the slides on the presentation recordings. As technical terminologies that are frequent in paper explanations are notoriously challenging to transcribe without reference texts, our SlideAVSR dataset spotlights a new aspect of AVSR problems. As a simple yet effective baseline, we propose DocWhisper, an AVSR model that can refer to textual information from slides, and confirm its effectiveness on SlideAVSR.

【6】 Improving Speaker-independent Speech Emotion Recognition Using Dynamic  Joint Distribution Adaptation
标题:动态联合分布自适应改进非特定人语音情感识别
链接:https://arxiv.org/abs/2401.09752
作者:Cheng Lu,Yuan Zong,Hailun Lian,Yan Zhao,Björn Schuller,Wenming Zheng
备注:Accepted by ICASSP 2024
摘要:在与说话人无关的语音情感识别中,训练和测试样本是从不同的说话人中收集的,这导致了来自不同说话人的数据的特征分布之间的多域移位挑战。因此,当训练好的模型面对来自新说话者的数据时,其性能往往会下降。为了解决这个问题,我们提出了一个动态联合分布自适应(DJDA)的多源域自适应框架下的方法。DJDA首先利用联合分布自适应(Joint Distribution Adaptation,JDA),包括边缘分布自适应(Margin Distribution Adaptation,MDA)和条件分布自适应(Conditional Distribution Adaptation,CDA),更精确地测量不同说话人引起的多域分布偏移。这有助于消除情感特征中的说话人偏见,允许从粗级到细级学习区分性和说话人不变的语音情感特征。此外,我们量化的MDA和CDA的适应JDA中的贡献,通过使用一个动态的平衡因素的基础上$\mathcal{A}$-距离,促进有效地处理未知的分布遇到的数据,从新的发言人。实验结果表明,我们的DJDA相比,其他国家的最先进的(SOTA)方法的优越性能。
摘要:In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new speakers, its performance tends to degrade. To address the issue, we propose a Dynamic Joint Distribution Adaptation (DJDA) method under the framework of multi-source domain adaptation. DJDA firstly utilizes joint distribution adaptation (JDA), involving marginal distribution adaptation (MDA) and conditional distribution adaptation (CDA), to more precisely measure the multi-domain distribution shifts caused by different speakers. This helps eliminate speaker bias in emotion features, allowing for learning discriminative and speaker-invariant speech emotion features from coarse-level to fine-level. Furthermore, we quantify the adaptation contributions of MDA and CDA within JDA by using a dynamic balance factor based on $\mathcal{A}$-Distance, promoting to effectively handle the unknown distributions encountered in data from new speakers. Experimental results demonstrate the superior performance of our DJDA as compared to other state-of-the-art (SOTA) methods.


【7】 MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
标题:MLAAD:多语言音频反欺骗数据集
链接:https://arxiv.org/abs/2401.09512
作者:Nicolas M. Müller,Piotr Kawa,Wei Herng Choong,Edresson Casanova,Eren Gölge,Thorsten Müller,Piotr Syga,Philip Sperl,Konstantin Böttinger
备注:Submitted to IJCNN 2024
摘要:文本到语音(TTS)技术带来了显着的优势,例如为那些有语言障碍的人提供语音,但也可以实现音频深度伪造和欺骗。前者误导个人,可能传播错误信息,而后者破坏语音生物识别安全系统。基于人工智能的检测可以通过自动区分真实和伪造的录音来帮助解决这些挑战。然而,这些模型的效果仅与其训练数据一样好,由于反欺骗数据库中的英语和中文音频过于集中,目前训练数据受到严重限制,从而限制了其全球有效性。作为回应,本文提出了多语言音频反欺骗数据集(MLAAD),使用52个TTS模型创建,包括19种不同的架构,以23种不同的语言生成160.1小时的合成语音。我们使用MLAAD训练和评估了三种最先进的深度伪造检测模型,并观察到MLAAD在用作训练资源时表现出优于InTheWild或FakeOrReal等可比数据集的性能。此外,与著名的ASVspoof 2019数据集相比,MLAAD被证明是一种补充资源。在八个数据集的测试中,MLAAD和ASVspoof 2019交替优于对方,都在四个数据集上表现出色。通过发布MLAAD并通过交互式网络服务器访问经过训练的模型,我们的目标是使反欺骗技术民主化,使其超越专家领域,从而为全球打击音频欺骗和深度伪造的努力做出贡献。
摘要:Text-to-Speech (TTS) technology brings significant advantages, such as giving a voice to those with speech impairments, but also enables audio deepfakes and spoofs. The former mislead individuals and may propagate misinformation, while the latter undermine voice biometric security systems. AI-based detection can help to address these challenges by automatically differentiating between genuine and fabricated voice recordings. However, these models are only as good as their training data, which currently is severely limited due to an overwhelming concentration on English and Chinese audio in anti-spoofing databases, thus restricting its worldwide effectiveness. In response, this paper presents the Multi-Language Audio Anti-Spoof Dataset (MLAAD), created using 52 TTS models, comprising 19 different architectures, to generate 160.1 hours of synthetic voice in 23 different languages. We train and evaluate three state-of-the-art deepfake detection models with MLAAD, and observe that MLAAD demonstrates superior performance over comparable datasets like InTheWild or FakeOrReal when used as a training resource. Furthermore, in comparison with the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, both excelling on four datasets. By publishing MLAAD and making trained models accessible via an interactive webserver , we aim to democratize antispoofing technology, making it accessible beyond the realm of specialists, thus contributing to global efforts against audio spoofing and deepfakes.

【8】 Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from  their voices
标题:Voxeleb-ESP:从西班牙名人的声音中检测他们的初步实验
链接:https://arxiv.org/abs/2401.09441
作者:Beltrán Labrador,Manuel Otero-Gonzalez,Alicia Lozano-Diez,Daniel Ramos,Doroteo T. Toledano,Joaquin Gonzalez-Rodriguez
摘要:本文介绍了VoxCeleb-ESP,一个指向YouTube视频的指针和时间戳集合,用于创建一个新的说话人识别数据集。VoxCeleb-ESP捕捉真实世界的场景,结合不同的说话风格,噪音和通道失真。它包括160名西班牙名人跨越不同类别,确保在西班牙的年龄组和地理区域的代表性分布。我们为说话人识别任务提供了两个说话人试验列表,每个列表分别具有相同视频或不同视频的目标试验,并伴随着ResNet预训练模型的跨语言评估。初步的说话人识别结果表明,在VoxCeleb-ESP的检测任务的复杂性相当于原来的和更大的VoxCeleb在英语。VoxCeleb-ESP为西班牙语提供了全面而多样化的数据集,有助于扩展说话人识别基准。
摘要:This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles, noises, and channel distortions. It includes 160 Spanish celebrities spanning various categories, ensuring a representative distribution across age groups and geographic regions in Spain. We provide two speaker trial lists for speaker identification tasks, each of them with same-video or different-video target trials respectively, accompanied by a cross-lingual evaluation of ResNet pretrained models. Preliminary speaker identification results suggest that the complexity of the detection task in VoxCeleb-ESP is equivalent to that of the original and much larger VoxCeleb in English. VoxCeleb-ESP contributes to the expansion of speaker recognition benchmarks with a comprehensive and diverse dataset for the Spanish language.


【9】 Multilingual Visual Speech Recognition with a Single Model by Learning  with Discrete Visual Speech Units
标题:基于离散视觉单元学习的单模型多语言视觉语音识别
链接:https://arxiv.org/abs/2401.09802
作者:Minsu Kim,Jeong Hun Yeo,Jeongsoo Choi,Se Jin Park,Yong Man Ro
摘要:本文首次研究了单模型的多语言视觉语音识别问题。由于大量的多语言建模的视觉数据需要巨大的计算成本,我们提出了一种新的策略,处理与视觉语音单元。最近的成功的音频语音单元的动机,建议的视觉语音单元是通过离散化的视觉语音特征提取的自监督视觉语音模型。为了正确捕捉多语言视觉语音,我们首先在5,512小时的多语言视听数据上训练自监督视觉语音模型。通过分析,我们证实了视觉言语单位主要包含视位信息,而抑制了非语言信息。通过使用视觉语音单元作为我们的系统的输入,我们预训练模型,以预测相应的文本输出的海量多语种数据,通过合并多个VSR数据库构建。由于输入和输出都是离散的,与标准VSR训练相比,我们可以大大提高训练效率。具体而言,输入数据大小减少到原始视频输入的0.016%。为了补充语音识别中视觉信息的不足,我们应用课程学习,其中系统的输入从视听语音单元开始,并逐渐改变为视觉语音单元。在预训练之后,模型在连续特征上进行微调。我们通过使用单个训练模型实现与以前的特定语言VSR模型相当的性能,从而设置了新的最先进的多语言VSR性能。
摘要:This paper explores sentence-level Multilingual Visual Speech Recognition with a single model for the first time. As the massive multilingual modeling of visual data requires huge computational costs, we propose a novel strategy, processing with visual speech units. Motivated by the recent success of the audio speech unit, the proposed visual speech unit is obtained by discretizing the visual speech features extracted from the self-supervised visual speech model. To correctly capture multilingual visual speech, we first train the self-supervised visual speech model on 5,512 hours of multilingual audio-visual data. Through analysis, we verify that the visual speech units mainly contain viseme information while suppressing non-linguistic information. By using the visual speech units as the inputs of our system, we pre-train the model to predict corresponding text outputs on massive multilingual data constructed by merging several VSR databases. As both the inputs and outputs are discrete, we can greatly improve the training efficiency compared to the standard VSR training. Specifically, the input data size is reduced to 0.016% of the original video inputs. In order to complement the insufficient visual information in speech recognition, we apply curriculum learning where the inputs of the system begin with audio-visual speech units and gradually change to visual speech units. After pre-training, the model is finetuned on continuous features. We set new state-of-the-art multilingual VSR performances by achieving comparable performances to the previous language-specific VSR models, with a single trained model.


【10】 Parameter Selection for Analyzing Conversations with Autism Spectrum  Disorder
标题:自闭症谱系障碍对话分析的参数选择
链接:https://arxiv.org/abs/2401.09717
作者:Tahiya Chowdhury,Veronica Romero,Amanda Stent
备注:5 pages, 4 tables, Proceedings of INTERSPEECH 2023
摘要:自闭症谱系障碍(ASD)的诊断是一项复杂而具有挑战性的任务,因为它依赖于心理学家对干扰行为的分析,而不是使用生化诊断。在本文中,我们提出了一种建模方法,ASD诊断分析声学/韵律和语言特征提取的诊断对话之间的心理学家和儿童谁是典型的发展(TD)或有ASD。我们比较了不同功能在一系列会话任务中的贡献。我们专注于寻找一个最小的参数集的自闭症儿童的会话行为的特点。由于ASD是通过对话互动来诊断的,除了分析儿童的行为外,我们还调查了心理学家的对话行为是否在诊断组之间存在差异。我们的研究结果可以促进自闭症儿童对话数据的细粒度分析,以支持诊断和干预。
摘要:The diagnosis of autism spectrum disorder (ASD) is a complex, challenging task as it depends on the analysis of interactional behaviors by psychologists rather than the use of biochemical diagnostics. In this paper, we present a modeling approach to ASD diagnosis by analyzing acoustic/prosodic and linguistic features extracted from diagnostic conversations between a psychologist and children who either are typically developing (TD) or have ASD. We compare the contributions of different features across a range of conversation tasks. We focus on finding a minimal set of parameters that characterize conversational behaviors of children with ASD. Because ASD is diagnosed through conversational interaction, in addition to analyzing the behavior of the children, we also investigate whether the psychologist's conversational behaviors vary across diagnostic groups. Our results can facilitate fine-grained analysis of conversation data for children with ASD to support diagnosis and intervention.

【11】 An Empirical Study on the Impact of Positional Encoding in  Transformer-based Monaural Speech Enhancemen
t标题:位置编码对基于Transformer的单声道语音增强影响的实验研究
链接:https://arxiv.org/abs/2401.09686
作者:Qiquan Zhang,Meng Ge,Hongxu Zhu,Eliathamby Ambikairajah,Qi Song,Zhaoheng Ni,Haizhou Li备注:ICASSP 2024
摘要:Transformer体系结构使得语音增强的最新进展成为可能。由于Transformers是位置agostic的,因此位置编码是事实上的标准组件,用于使Transformers能够区分序列中元素的顺序。然而,目前还不清楚位置编码如何确切地影响基于Transformer架构的语音增强。在本文中,我们进行了全面的实证研究,评估五个位置编码方法,即,正弦和学习的绝对位置嵌入(APE)、T5-RPE、KERPLE以及无位置编码的Transformer(No-Pos),跨因果和非因果配置。我们进行了广泛的语音增强实验,涉及频谱映射和掩蔽方法。我们的研究结果表明,位置编码是不是很有帮助的模型在一个因果配置,这表明因果注意可能隐含的位置信息。在非因果配置中,模型显著受益于位置编码的使用。此外,我们发现,在四个位置嵌入,相对位置嵌入优于APE。
摘要:Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encoding exactly impacts speech enhancement based on Transformer architectures. In this paper, we perform a comprehensive empirical study evaluating five positional encoding methods, i.e., Sinusoidal and learned absolute position embedding (APE), T5-RPE, KERPLE, as well as the Transformer without positional encoding (No-Pos), across both causal and noncausal configurations. We conduct extensive speech enhancement experiments, involving spectral mapping and masking methods. Our findings establish that positional encoding is not quite helpful for the models in a causal configuration, which indicates that causal attention may implicitly incorporate position information. In a noncausal configuration, the models significantly benefit from the use of positional encoding. In addition, we find that among the four position embeddings, relative position embeddings outperform APEs.

eess.AS音频处理
【1】 FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
标题:FreGrad:轻量级快速频率感知扩散声码器
链接:https://arxiv.org/abs/2401.10032
作者:Tan Dat Nguyen,Ji-Hoon Kim,Youngjoon Jang,Jaehun Kim,Joon Son Chung
备注:Accepted to ICASSP 2024
摘要:本文的目标是生成逼真的音频与轻量级和快速扩散为基础的声码器,名为FreGrad。我们的框架包括以下三个主要组成部分:(1)我们采用离散小波变换将复杂的波形分解成子带小波,这有助于FreGrad在简单而简洁的特征空间上操作,(2)我们设计了频率感知的扩张卷积,提高了频率感知,从而产生具有准确频率信息的语音,以及(3)我们引入了一系列技巧来提高所提出的模型的生成质量。在我们的实验中,FreGrad实现了比我们的基线快3.7倍的训练时间和快2.2倍的推理速度,同时将模型大小减少了0.6倍(只有1.78 M个参数),而不牺牲输出质量。音频样本可在https://mm.kaist.ac.kr/projects/FreGrad上获得。
摘要:The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps FreGrad to operate on a simple and concise feature space, (2) We design a frequency-aware dilated convolution that elevates frequency awareness, resulting in generating speech with accurate frequency information, and (3) We introduce a bag of tricks that boosts the generation quality of the proposed model. In our experiments, FreGrad achieves 3.7 times faster training time and 2.2 times faster inference speed compared to our baseline while reducing the model size by 0.6 times (only 1.78M parameters) without sacrificing the output quality. Audio samples are available at: https://mm.kaist.ac.kr/projects/FreGrad.

【2】 Multilingual Visual Speech Recognition with a Single Model by Learning  with Discrete Visual Speech Units
标题:基于离散视觉单元学习的单模型多语言视觉语音识别
链接:https://arxiv.org/abs/2401.09802
作者:Minsu Kim,Jeong Hun Yeo,Jeongsoo Choi,Se Jin Park,Yong Man Ro
摘要:本文首次研究了单模型的多语言视觉语音识别问题。由于大量的多语言建模的视觉数据需要巨大的计算成本,我们提出了一种新的策略,处理与视觉语音单元。最近的成功的音频语音单元的动机,建议的视觉语音单元是通过离散化的视觉语音特征提取的自监督视觉语音模型。为了正确捕捉多语言视觉语音,我们首先在5,512小时的多语言视听数据上训练自监督视觉语音模型。通过分析,我们证实了视觉言语单位主要包含视位信息,而抑制了非语言信息。通过使用视觉语音单元作为我们的系统的输入,我们预训练模型,以预测相应的文本输出的海量多语种数据,通过合并多个VSR数据库构建。由于输入和输出都是离散的,与标准VSR训练相比,我们可以大大提高训练效率。具体而言,输入数据大小减少到原始视频输入的0.016%。为了补充语音识别中视觉信息的不足,我们应用课程学习,其中系统的输入从视听语音单元开始,并逐渐改变为视觉语音单元。在预训练之后,模型在连续特征上进行微调。我们通过使用单个训练模型实现与以前的特定语言VSR模型相当的性能,从而设置了新的最先进的多语言VSR性能。
摘要:This paper explores sentence-level Multilingual Visual Speech Recognition with a single model for the first time. As the massive multilingual modeling of visual data requires huge computational costs, we propose a novel strategy, processing with visual speech units. Motivated by the recent success of the audio speech unit, the proposed visual speech unit is obtained by discretizing the visual speech features extracted from the self-supervised visual speech model. To correctly capture multilingual visual speech, we first train the self-supervised visual speech model on 5,512 hours of multilingual audio-visual data. Through analysis, we verify that the visual speech units mainly contain viseme information while suppressing non-linguistic information. By using the visual speech units as the inputs of our system, we pre-train the model to predict corresponding text outputs on massive multilingual data constructed by merging several VSR databases. As both the inputs and outputs are discrete, we can greatly improve the training efficiency compared to the standard VSR training. Specifically, the input data size is reduced to 0.016% of the original video inputs. In order to complement the insufficient visual information in speech recognition, we apply curriculum learning where the inputs of the system begin with audio-visual speech units and gradually change to visual speech units. After pre-training, the model is finetuned on continuous features. We set new state-of-the-art multilingual VSR performances by achieving comparable performances to the previous language-specific VSR models, with a single trained model.

【3】 Parameter Selection for Analyzing Conversations with Autism Spectrum  Disorder
标题:自闭症谱系障碍对话分析的参数选择
链接:https://arxiv.org/abs/2401.09717
作者:Tahiya Chowdhury,Veronica Romero,Amanda Stent
备注:5 pages, 4 tables, Proceedings of INTERSPEECH 2023
摘要:自闭症谱系障碍(ASD)的诊断是一项复杂而具有挑战性的任务,因为它依赖于心理学家对干扰行为的分析,而不是使用生化诊断。在本文中,我们提出了一种建模方法,ASD诊断分析声学/韵律和语言特征提取的诊断对话之间的心理学家和儿童谁是典型的发展(TD)或有ASD。我们比较了不同功能在一系列会话任务中的贡献。我们专注于寻找一个最小的参数集的自闭症儿童的会话行为的特点。由于ASD是通过对话互动来诊断的,除了分析儿童的行为外,我们还调查了心理学家的对话行为是否在诊断组之间存在差异。我们的研究结果可以促进自闭症儿童对话数据的细粒度分析,以支持诊断和干预。
摘要:The diagnosis of autism spectrum disorder (ASD) is a complex, challenging task as it depends on the analysis of interactional behaviors by psychologists rather than the use of biochemical diagnostics. In this paper, we present a modeling approach to ASD diagnosis by analyzing acoustic/prosodic and linguistic features extracted from diagnostic conversations between a psychologist and children who either are typically developing (TD) or have ASD. We compare the contributions of different features across a range of conversation tasks. We focus on finding a minimal set of parameters that characterize conversational behaviors of children with ASD. Because ASD is diagnosed through conversational interaction, in addition to analyzing the behavior of the children, we also investigate whether the psychologist's conversational behaviors vary across diagnostic groups. Our results can facilitate fine-grained analysis of conversation data for children with ASD to support diagnosis and intervention.


【4】 An Empirical Study on the Impact of Positional Encoding in  Transformer-based Monaural Speech Enhancement
标题:位置编码对基于Transformer的单声道语音增强影响的实验研究
链接:https://arxiv.org/abs/2401.09686
作者:Qiquan Zhang,Meng Ge,Hongxu Zhu,Eliathamby Ambikairajah,Qi Song,Zhaoheng Ni,Haizhou Li备注:ICASSP 2024
摘要:Transformer体系结构使得语音增强的最新进展成为可能。由于Transformers是位置agostic的,因此位置编码是事实上的标准组件,用于使Transformers能够区分序列中元素的顺序。然而,目前还不清楚位置编码如何确切地影响基于Transformer架构的语音增强。在本文中,我们进行了全面的实证研究,评估五个位置编码方法,即,正弦和学习的绝对位置嵌入(APE)、T5-RPE、KERPLE以及无位置编码的Transformer(No-Pos),跨因果和非因果配置。我们进行了广泛的语音增强实验,涉及频谱映射和掩蔽方法。我们的研究结果表明,位置编码是不是很有帮助的模型在一个因果配置,这表明因果注意可能隐含的位置信息。在非因果配置中,模型显著受益于位置编码的使用。此外,我们发现,在四个位置嵌入,相对位置嵌入优于APE。
摘要:Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encoding exactly impacts speech enhancement based on Transformer architectures. In this paper, we perform a comprehensive empirical study evaluating five positional encoding methods, i.e., Sinusoidal and learned absolute position embedding (APE), T5-RPE, KERPLE, as well as the Transformer without positional encoding (No-Pos), across both causal and noncausal configurations. We conduct extensive speech enhancement experiments, involving spectral mapping and masking methods. Our findings establish that positional encoding is not quite helpful for the models in a causal configuration, which indicates that causal attention may implicitly incorporate position information. In a noncausal configuration, the models significantly benefit from the use of positional encoding. In addition, we find that among the four position embeddings, relative position embeddings outperform APEs.

【5】 Communication-Efficient Personalized Federated Learning for  Speech-to-Text Tasks
标题:用于语音到文本任务的通信高效的个性化联合学习
链接:https://arxiv.org/abs/2401.10070
作者:Yichao Du,Zhirui Zhang,Linan Yue,Xu Huang,Yuqing Zhang,Tong Xu,Linli Xu,Enhong Chen
备注:ICASSP 2024
摘要:为了保护隐私和满足法律法规,联邦学习(FL)在训练语音到文本(S2 T)系统方面获得了极大的关注,包括自动语音识别(ASR)和语音翻译(ST)。然而,通常使用的FL方法(即,\textsc{FedAvg})存在大量的通信开销和客户端数据异构性导致的性能下降。为了解决这些问题,提出了一种个性化的联邦S2 T框架,该框架引入了\textsc{FedLoRA},一个轻量级的LoRA模块,用于客户端调优和与服务器的交互,以最大限度地减少通信开销,和\textsc{FedMem},这是一个全局模型,配备了$k$-最近邻($k$NN)分类器,可捕获特定于客户端的分布变化,以实现个性化并克服数据异构性。基于Conformer和Whisper骨干模型在CoVoST和GigaSpeech基准上进行的大量实验表明,我们的方法显着降低了所有S2 T任务的通信开销,并有效地个性化全局模型,以克服数据异构性。
摘要:To protect privacy and meet legal regulations, federated learning (FL) has gained significant attention for training speech-to-text (S2T) systems, including automatic speech recognition (ASR) and speech translation (ST). However, the commonly used FL approach (i.e., \textsc{FedAvg}) in S2T tasks typically suffers from extensive communication overhead due to multi-round interactions based on the whole model and performance degradation caused by data heterogeneity among clients.To address these issues, we propose a personalized federated S2T framework that introduces \textsc{FedLoRA}, a lightweight LoRA module for client-side tuning and interaction with the server to minimize communication overhead, and \textsc{FedMem}, a global model equipped with a $k$-nearest-neighbor ($k$NN) classifier that captures client-specific distributional shifts to achieve personalization and overcome data heterogeneity. Extensive experiments based on Conformer and Whisper backbone models on CoVoST and GigaSpeech benchmarks show that our approach significantly reduces the communication overhead on all S2T tasks and effectively personalizes the global model to overcome data heterogeneity.

【6】 Towards Hierarchical Spoken Language Dysfluency Modeling
标题:面向层次化口语流利性建模
链接:https://arxiv.org/abs/2401.10015
作者:Jiachen Lian,Gopala Anumanchipalli
备注:2024 EACL Long (main conference). arXiv admin note: substantial text overlap with arXiv:2312.12810
摘要:言语不流利建模是言语治疗和语言学习的瓶颈。然而,没有人工智能解决方案来系统地解决这个问题。我们首先提出定义不流利语音和不流利语音建模的概念。然后,我们提出了分层无约束的不流畅建模(H-UDM)的方法,解决了不流畅的转录和检测,以消除大量的手动注释的需要。此外,我们还引入了一个名为VCTK++的模拟不流利数据集,以增强H-UDM在语音转录方面的能力。我们的实验结果表明,我们提出的方法在转录和检测任务的有效性和鲁棒性。
摘要:Speech dysfluency modeling is the bottleneck for both speech therapy and language learning. However, there is no AI solution to systematically tackle this problem. We first propose to define the concept of dysfluent speech and dysfluent speech modeling. We then present Hierarchical Unconstrained Dysfluency Modeling (H-UDM) approach that addresses both dysfluency transcription and detection to eliminate the need for extensive manual annotation. Furthermore, we introduce a simulated dysfluent dataset called VCTK++ to enhance the capabilities of H-UDM in phonetic transcription. Our experimental results demonstrate the effectiveness and robustness of our proposed methods in both transcription and detection tasks.


【7】 Developing an AI-based Integrated System for Bee Health Evaluation
标题:基于人工智能的蜜蜂健康综合评价系统的开发
链接:https://arxiv.org/abs/2401.09988
作者:Andrew Liang
摘要:蜜蜂授粉约占世界粮食供应的三分之一,但由于杀虫剂和害虫等多种因素,蜂群数量在过去十年中惊人地下降了近40%。监测蜂箱的传统方法,如人工检查,是主观的,破坏性的,耗时的。为了克服这些局限性,人工智能已被用于评估蜂巢健康。然而,以前的研究缺乏端到端的解决方案,主要依赖于来自单一来源的数据,无论是蜜蜂图像还是声音。本研究介绍一个综合系统,包括蜜蜂目标检测和健康评估。此外,它利用视觉和音频信号的组合来分析蜜蜂的行为。基于注意力的多模态神经网络(AMNN)被开发用于自适应地关注来自每种类型的信号的关键特征,以进行准确的蜜蜂健康评估。AMNN的总体准确率达到92.61%,超过了现有的8个单信号卷积神经网络和递归神经网络。它比最好的基于图像的模型高出32.51%,比最好的基于声音的模型高出13.98%,同时保持了有效的处理时间。此外,它还提高了预测的鲁棒性,在所有四种评估的健康状况中,F1得分均高于90%。该研究还表明,音频信号比图像更可靠地评估蜜蜂健康。通过将AMNN与图像和声音数据无缝集成在一个全面的蜜蜂健康监测系统中,这种方法为蜜蜂疾病的早期检测和蜂群的保护提供了一种更有效和非侵入性的解决方案。
摘要:Honey bees pollinate about one-third of the world's food supply, but bee colonies have alarmingly declined by nearly 40% over the past decade due to several factors, including pesticides and pests. Traditional methods for monitoring beehives, such as human inspection, are subjective, disruptive, and time-consuming. To overcome these limitations, artificial intelligence has been used to assess beehive health. However, previous studies have lacked an end-to-end solution and primarily relied on data from a single source, either bee images or sounds. This study introduces a comprehensive system consisting of bee object detection and health evaluation. Additionally, it utilized a combination of visual and audio signals to analyze bee behaviors. An Attention-based Multimodal Neural Network (AMNN) was developed to adaptively focus on key features from each type of signal for accurate bee health assessment. The AMNN achieved an overall accuracy of 92.61%, surpassing eight existing single-signal Convolutional Neural Networks and Recurrent Neural Networks. It outperformed the best image-based model by 32.51% and the top sound-based model by 13.98% while maintaining efficient processing times. Furthermore, it improved prediction robustness, attaining an F1-score higher than 90% across all four evaluated health conditions. The study also shows that audio signals are more reliable than images for assessing bee health. By seamlessly integrating AMNN with image and sound data in a comprehensive bee health monitoring system, this approach provides a more efficient and non-invasive solution for the early detection of bee diseases and the preservation of bee colonies.

【8】 Attention-Based Recurrent Neural Network For Automatic Behavior Laying  Hen Recognition
标题:基于注意力的递归神经网络用于蛋鸡行为自动识别
链接:https://arxiv.org/abs/2401.09880
作者:Fréjus A. A. Laleye,Mikaël A. Mousse
摘要:现代家禽养殖的兴趣之一是蛋鸡的发声,其中包含非常有用的健康行为信息。这些信息被用作健康和福利指标,帮助育种者更好地监测蛋鸡,这涉及到早期发现问题,以便进行快速和更有效的干预。在这项工作中,我们专注于声音分析识别的蛋鸡的呼叫类型,以提出一个强大的系统表征他们的行为,更好地监测。为此,我们首先收集并标注蛋鸡叫声信号,然后基于时域和频域特征的组合设计了最佳声学表征。然后,我们使用这些特征来建立基于递归神经网络的多标签分类模型,为表征蛋鸡行为的发声分配语义类。结果显示,基于时域和频域特征组合的模型的整体性能获得了最高的F1分数(F1=92.75),使用频域特征的模型增益为17%,文献中比较的方法增益为8%。
摘要:One of the interests of modern poultry farming is the vocalization of laying hens which contain very useful information on health behavior. This information is used as health and well-being indicators that help breeders better monitor laying hens, which involves early detection of problems for rapid and more effective intervention. In this work, we focus on the sound analysis for the recognition of the types of calls of the laying hens in order to propose a robust system of characterization of their behavior for a better monitoring. To do this, we first collected and annotated laying hen call signals, then designed an optimal acoustic characterization based on the combination of time and frequency domain features. We then used these features to build the multi-label classification models based on recurrent neural network to assign a semantic class to the vocalization that characterize the laying hen behavior. The results show an overall performance with our model based on the combination of time and frequency domain features that obtained the highest F1-score (F1=92.75) with a gain of 17% on the models using the frequency domain features and of 8% on the compared approaches from the litterature.

【9】 On the Audio Hallucinations in Large Audio-Video Language Models
标题:论大型视听语言模型中的幻听现象
链接:https://arxiv.org/abs/2401.09774
作者:Taichi Nishimura,Shota Nakada,Masayoshi Kondo
备注:6 pages
摘要:大型音视频语言模型可以生成视频和音频的描述。然而,它们有时会忽略音频内容,仅依靠视觉信息制作音频描述。本文将其称为音频幻觉,并在大型音频视频语言模型中对其进行分析。我们通过询问音频信息收集了1,000个句子,并对它们是否包含幻觉进行注释。如果一个句子是幻觉,我们也会对幻觉的类型进行分类。结果表明,332个句子是幻觉与不同的趋势观察到的名词和动词的幻觉类型。基于此,我们解决了一个任务的音频幻觉分类使用预训练的音频文本模型在zero-shot和微调设置。我们的实验结果表明,zero-shot模型实现了更高的性能(52.2%,在F1)比随机(40.3%)和微调模型达到87.9%,优于zero-shot模型。
摘要:Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on visual information. This paper refers to this as audio hallucinations and analyzes them in large audio-video language models. We gather 1,000 sentences by inquiring about audio information and annotate them whether they contain hallucinations. If a sentence is hallucinated, we also categorize the type of hallucination. The results reveal that 332 sentences are hallucinated with distinct trends observed in nouns and verbs for each hallucination type. Based on this, we tackle a task of audio hallucination classification using pre-trained audio-text models in the zero-shot and fine-tuning settings. Our experimental results reveal that the zero-shot models achieve higher performance (52.2% in F1) than the random (40.3%) and the fine-tuning models achieve 87.9%, outperforming the zero-shot models.


【10】 SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech  Recognition
标题:SlideAVSR:用于视听语音识别的纸质讲解视频数据集
链接:https://arxiv.org/abs/2401.09759
作者:Hao Wang,Shuhei Kurita,Shuichiro Shimizu,Daisuke Kawahara
摘要:视听语音识别(AVSR)是自动语音识别(ASR)的多模态扩展,使用视频作为音频的补充。在AVSR中,相当大的努力已经针对面部特征的数据集,如唇读,而他们往往在更广泛的背景下评估图像理解能力不足。在本文中,我们构建了SlideAVSR,AVSR数据集使用科学论文解释视频。SlideAVSR提供了一个新的基准,模型在演示记录的幻灯片上转录语音话语和文本。众所周知,在没有参考文本的情况下,经常出现在论文解释中的技术术语很难转录,因此我们的SlideAVSR数据集突出了AVSR问题的一个新方面。作为一个简单而有效的基线,我们提出了DocWhisper,一个AVSR模型,可以引用幻灯片中的文本信息,并确认其在SlideAVSR上的有效性。
摘要:Audio-visual speech recognition (AVSR) is a multimodal extension of automatic speech recognition (ASR), using video as a complement to audio. In AVSR, considerable efforts have been directed at datasets for facial features such as lip-readings, while they often fall short in evaluating the image comprehension capabilities in broader contexts. In this paper, we construct SlideAVSR, an AVSR dataset using scientific paper explanation videos. SlideAVSR provides a new benchmark where models transcribe speech utterances with texts on the slides on the presentation recordings. As technical terminologies that are frequent in paper explanations are notoriously challenging to transcribe without reference texts, our SlideAVSR dataset spotlights a new aspect of AVSR problems. As a simple yet effective baseline, we propose DocWhisper, an AVSR model that can refer to textual information from slides, and confirm its effectiveness on SlideAVSR.


【11】 Improving Speaker-independent Speech Emotion Recognition Using Dynamic  Joint Distribution Adaptation
标题:动态联合分布自适应改进非特定人语音情感识别
链接:https://arxiv.org/abs/2401.09752
作者:Cheng Lu,Yuan Zong,Hailun Lian,Yan Zhao,Björn Schuller,Wenming Zheng
备注:Accepted by ICASSP 2024
摘要:在与说话人无关的语音情感识别中,训练和测试样本是从不同的说话人中收集的,这导致了来自不同说话人的数据的特征分布之间的多域移位挑战。因此,当训练好的模型面对来自新说话者的数据时,其性能往往会下降。为了解决这个问题,我们提出了一个动态联合分布自适应(DJDA)的多源域自适应框架下的方法。DJDA首先利用联合分布自适应(Joint Distribution Adaptation,JDA),包括边缘分布自适应(Margin Distribution Adaptation,MDA)和条件分布自适应(Conditional Distribution Adaptation,CDA),更精确地测量不同说话人引起的多域分布偏移。这有助于消除情感特征中的说话人偏见,允许从粗级到细级学习区分性和说话人不变的语音情感特征。此外,我们量化的MDA和CDA的适应JDA中的贡献,通过使用一个动态的平衡因素的基础上$\mathcal{A}$-距离,促进有效地处理未知的分布遇到的数据,从新的发言人。实验结果表明,我们的DJDA相比,其他国家的最先进的(SOTA)方法的优越性能。
摘要:In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new speakers, its performance tends to degrade. To address the issue, we propose a Dynamic Joint Distribution Adaptation (DJDA) method under the framework of multi-source domain adaptation. DJDA firstly utilizes joint distribution adaptation (JDA), involving marginal distribution adaptation (MDA) and conditional distribution adaptation (CDA), to more precisely measure the multi-domain distribution shifts caused by different speakers. This helps eliminate speaker bias in emotion features, allowing for learning discriminative and speaker-invariant speech emotion features from coarse-level to fine-level. Furthermore, we quantify the adaptation contributions of MDA and CDA within JDA by using a dynamic balance factor based on $\mathcal{A}$-Distance, promoting to effectively handle the unknown distributions encountered in data from new speakers. Experimental results demonstrate the superior performance of our DJDA as compared to other state-of-the-art (SOTA) methods.


【12】 MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
标题:MLAAD:多语言音频反欺骗数据集
链接:https://arxiv.org/abs/2401.09512
作者:Nicolas M. Müller,Piotr Kawa,Wei Herng Choong,Edresson Casanova,Eren Gölge,Thorsten Müller,Piotr Syga,Philip Sperl,Konstantin Böttinger
备注:Submitted to IJCNN 2024
摘要:文本到语音(TTS)技术带来了显着的优势,例如为那些有语言障碍的人提供语音,但也可以实现音频深度伪造和欺骗。前者误导个人,可能传播错误信息,而后者破坏语音生物识别安全系统。基于人工智能的检测可以通过自动区分真实和伪造的录音来帮助解决这些挑战。然而,这些模型的效果仅与其训练数据一样好,由于反欺骗数据库中的英语和中文音频过于集中,目前训练数据受到严重限制,从而限制了其全球有效性。作为回应,本文提出了多语言音频反欺骗数据集(MLAAD),使用52个TTS模型创建,包括19种不同的架构,以23种不同的语言生成160.1小时的合成语音。我们使用MLAAD训练和评估了三种最先进的深度伪造检测模型,并观察到MLAAD在用作训练资源时表现出优于InTheWild或FakeOrReal等可比数据集的性能。此外,与著名的ASVspoof 2019数据集相比,MLAAD被证明是一种补充资源。在八个数据集的测试中,MLAAD和ASVspoof 2019交替优于对方,都在四个数据集上表现出色。通过发布MLAAD并通过交互式网络服务器访问经过训练的模型,我们的目标是使反欺骗技术民主化,使其超越专家领域,从而为全球打击音频欺骗和深度伪造的努力做出贡献。
摘要:Text-to-Speech (TTS) technology brings significant advantages, such as giving a voice to those with speech impairments, but also enables audio deepfakes and spoofs. The former mislead individuals and may propagate misinformation, while the latter undermine voice biometric security systems. AI-based detection can help to address these challenges by automatically differentiating between genuine and fabricated voice recordings. However, these models are only as good as their training data, which currently is severely limited due to an overwhelming concentration on English and Chinese audio in anti-spoofing databases, thus restricting its worldwide effectiveness. In response, this paper presents the Multi-Language Audio Anti-Spoof Dataset (MLAAD), created using 52 TTS models, comprising 19 different architectures, to generate 160.1 hours of synthetic voice in 23 different languages. We train and evaluate three state-of-the-art deepfake detection models with MLAAD, and observe that MLAAD demonstrates superior performance over comparable datasets like InTheWild or FakeOrReal when used as a training resource. Furthermore, in comparison with the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, both excelling on four datasets. By publishing MLAAD and making trained models accessible via an interactive webserver , we aim to democratize antispoofing technology, making it accessible beyond the realm of specialists, thus contributing to global efforts against audio spoofing and deepfakes.

【13】 Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from  their voices
标题:Voxeleb-ESP:从西班牙名人的声音中检测他们的初步实验
链接:https://arxiv.org/abs/2401.09441
作者:Beltrán Labrador,Manuel Otero-Gonzalez,Alicia Lozano-Diez,Daniel Ramos,Doroteo T. Toledano,Joaquin Gonzalez-Rodriguez
摘要:本文介绍了VoxCeleb-ESP,一个指向YouTube视频的指针和时间戳集合,用于创建一个新的说话人识别数据集。VoxCeleb-ESP捕捉真实世界的场景,结合不同的说话风格,噪音和通道失真。它包括160名西班牙名人跨越不同类别,确保在西班牙的年龄组和地理区域的代表性分布。我们为说话人识别任务提供了两个说话人试验列表,每个列表分别具有相同视频或不同视频的目标试验,并伴随着ResNet预训练模型的跨语言评估。初步的说话人识别结果表明,在VoxCeleb-ESP的检测任务的复杂性相当于原来的和更大的VoxCeleb在英语。VoxCeleb-ESP为西班牙语提供了全面而多样化的数据集,有助于扩展说话人识别基准。
摘要:This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles, noises, and channel distortions. It includes 160 Spanish celebrities spanning various categories, ensuring a representative distribution across age groups and geographic regions in Spain. We provide two speaker trial lists for speaker identification tasks, each of them with same-video or different-video target trials respectively, accompanied by a cross-lingual evaluation of ResNet pretrained models. Preliminary speaker identification results suggest that the complexity of the detection task in VoxCeleb-ESP is equivalent to that of the original and much larger VoxCeleb in English. VoxCeleb-ESP contributes to the expansion of speaker recognition benchmarks with a comprehensive and diverse dataset for the Spanish language.


机器翻译由腾讯交互翻译提供,仅供参考