本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2412.09467
摘要:随着人工智能技术的快速发展,Deepfake技术在音频领域的应用逐渐增多,由此带来的安全隐患也越来越广泛。特别是在金融和社会安全领域,deepfake音频的滥用引起了严重关注。为了应对这一挑战,本研究提出了一种基于多频道注意机制(MFCA)和2D离散余弦变换(DCT)的音频深度伪造检测方法。该方法通过将音频信号处理成melspectrogram,利用MobileNet V2提取深度特征,结合MFCA模块对音频信号中不同频率通道进行加权,能够有效捕捉音频信号中的细粒度频域特征,增强对虚假音频的分类能力。实验结果表明,与传统方法相比,本文提出的模型在准确率、精确率、召回率、F1评分等指标上表现出明显的优势。特别是在复杂音频场景下,该方法表现出更强的鲁棒性和泛化能力,为音频深度伪造检测提供了新的思路,具有重要的实际应用价值。未来将探索更先进的音频检测技术和优化策略,进一步提高音频deepfake检测的准确性和泛化能力。
摘要:With the rapid development of artificial intelligence technology, theapplication of deepfake technology in the audio field has gradually increased,resulting in a wide range of security risks. Especially in the financial andsocial security fields, the misuse of deepfake audios has raised seriousconcerns. To address this challenge, this study proposes an audio deepfakedetection method based on multi-frequency channel attention mechanism (MFCA)and 2D discrete cosine transform (DCT). By processing the audio signal into amelspectrogram, using MobileNet V2 to extract deep features, and combining itwith the MFCA module to weight different frequency channels in the audiosignal, this method can effectively capture the fine-grained frequency domainfeatures in the audio signal and enhance the Classification capability of fakeaudios. Experimental results show that compared with traditional methods, themodel proposed in this study shows significant advantages in accuracy,precision,recall, F1 score and other indicators. Especially in complex audioscenarios, this method shows stronger robustness and generalizationcapabilities and provides a new idea for audio deepfake detection and hasimportant practical application value. In the future, more advanced audiodetection technologies and optimization strategies will be explored to furtherimprove the accuracy and generalization capabilities of audio deepfakedetection.
标题:具有显式桥梁和检索增强的多模式音乐生成
链接:https://arxiv.org/abs/2412.09428
摘要:多模态音乐生成旨在从不同的输入模态(包括文本、视频和图像)生成音乐。现有的方法使用一个共同的嵌入空间进行多模态融合。尽管它们在其他模态中是有效的,但它们在多模态音乐生成中的应用面临着数据稀缺、弱跨模态对齐和有限可控性的挑战。本文通过使用文本和音乐的显式桥梁进行多模式对齐来解决这些问题。我们介绍了一种新的方法命名为视觉音乐桥(VMB)。具体而言,多模态音乐描述模型将视觉输入转换为详细的文本描述,以提供文本桥;双轨音乐检索模块,其结合了广泛的和有针对性的检索策略,以提供音乐桥并使用户能够控制。最后,我们设计了一个基于这两个桥梁的音乐生成框架。我们进行视频到音乐,图像到音乐,文本到音乐,可控的音乐生成任务的实验,以及可控性的实验。结果表明,VMB显着提高了音乐质量,模态和定制对齐相比,以前的方法。VMB为可解释性和表现力的多模态音乐生成设定了新的标准,并在各种多媒体领域中应用。演示和代码可在https://github.com/wbs2788/VMB上获得。
摘要:Multimodal music generation aims to produce music from diverse inputmodalities, including text, videos, and images. Existing methods use a commonembedding space for multimodal fusion. Despite their effectiveness in othermodalities, their application in multimodal music generation faces challengesof data scarcity, weak cross-modal alignment, and limited controllability. Thispaper addresses these issues by using explicit bridges of text and music formultimodal alignment. We introduce a novel method named Visuals Music Bridge(VMB). Specifically, a Multimodal Music Description Model converts visualinputs into detailed textual descriptions to provide the text bridge; aDual-track Music Retrieval module that combines broad and targeted retrievalstrategies to provide the music bridge and enable user control. Finally, wedesign an Explicitly Conditioned Music Generation framework to generate musicbased on the two bridges. We conduct experiments on video-to-music,image-to-music, text-to-music, and controllable music generation tasks, alongwith experiments on controllability. The results demonstrate that VMBsignificantly enhances music quality, modality, and customization alignmentcompared to previous methods. VMB sets a new standard for interpretable andexpressive multimodal music generation with applications in various multimediafields. Demos and code are available at https://github.com/wbs2788/VMB.
标题:基于视频和音频输入的多模式情绪分析
链接:https://arxiv.org/abs/2412.09317
备注:Presented as a full paper in the 15th International Conference on Emerging Ubiquitous Systems and Pervasive Networks (EUSPN 2024) October 28-30, 2024, Leuven, Belgium
摘要:尽管目前有大量的研究致力于从视频和音频中进行情感分析,但找到最好的模型,提供最高的准确率仍然被认为是该领域研究人员的挑战。本文的主要目的是证明的可用性的情感识别模型,视频和音频输入。用于训练模型的数据集是用于音频的CREMA-D数据集和用于视频的RAVDESS数据集。使用的微调模型是:Facebook/wav 2 vec 2-large用于音频,Google/vivit-b-16 x2-kinetics 400用于视频。在决策框架中利用了由两个先前模型生成的每个情感的概率的平均值。在结果不一致之后,如果其中一个模型获得更高的准确性,则创建另一个测试框架。采用的方法有加权平均法、置信度阈值法、基于置信度的动态加权法和基于规则的逻辑法。这种有限的方法给出了令人鼓舞的结果,使未来的研究这些方法可行。
摘要:Despite the abundance of current researches working on the sentiment analysisfrom videos and audios, finding the best model that gives the highest accuracyrate is still considered a challenge for researchers in this field. The mainobjective of this paper is to prove the usability of emotion recognition modelsthat take video and audio inputs. The datasets used to train the models are theCREMA-D dataset for audio and the RAVDESS dataset for video. The fine-tunedmodels that been used are: Facebook/wav2vec2-large for audio and theGoogle/vivit-b-16x2-kinetics400 for video. The avarage of the probabilities foreach emotion generated by the two previous models is utilized in the decisionmaking framework. After disparity in the results, if one of the models getsmuch higher accuracy, another test framework is created. The methods used arethe Weighted Average method, the Confidence Level Threshold method, the DynamicWeighting Based on Confidence method, and the Rule-Based Logic method. Thislimited approach gives encouraging results that make future research into thesemethods viable.
标题:语音隐私保护中说话人对抗性扰动的产生和消除
链接:https://arxiv.org/abs/2412.09195
备注:6 pages, 3 figures, published to IEEE SLT Workshop 2024
摘要:众所周知,神经网络容易受到通过对输入数据进行细微扰动而进行的对抗性攻击。语音隐私保护的最新发展已经显示了相同技术的积极用例,以利用对抗网络生成的附加扰动信号来隐藏说话者的语音属性。本文研究了可逆性属性,其中生成对抗性扰动的实体被授权移除它们并恢复原始语音(例如,演讲者本人)。调查人员也可以使用类似的技术来对语音保护的语音进行去匿名化,以在安全和法医分析中恢复罪犯的身份。在该设置中,假设扰动生成模块在去除处理中是已知的。为此,提出了扰动生成和去除模块的联合训练。在LibriSpeech数据集上的实验结果表明,该方法可以从匿名语音中预测出添加到原始语音中的细微扰动,同时达到隐私保护的目的。通过从匿名样本中去除这些扰动,可以恢复原始语音。音频示例可以在\url{https://voiceprivacy.github.io/Perturbation-Generation-Removal/}中找到。
摘要:Neural networks are commonly known to be vulnerable to adversarial attacksmounted through subtle perturbation on the input data. Recent development invoice-privacy protection has shown the positive use cases of the same techniqueto conceal speaker's voice attribute with additive perturbation signalgenerated by an adversarial network. This paper examines the reversibilityproperty where an entity generating the adversarial perturbations is authorizedto remove them and restore original speech (e.g., the speaker him/herself). Asimilar technique could also be used by an investigator to deanonymize avoice-protected speech to restore criminals' identities in security andforensic analysis. In this setting, the perturbation generative module isassumed to be known in the removal process. To this end, a joint training ofperturbation generation and removal modules is proposed. Experimental resultson the LibriSpeech dataset demonstrated that the subtle perturbations added tothe original speech can be predicted from the anonymized speech while achievingthe goal of privacy protection. By removing these perturbations from theanonymized sample, the original speech can be restored. Audio samples can befound in \url{https://voiceprivacy.github.io/Perturbation-Generation-Removal/}.
标题:YingSound:具有多模式思维链控制的视频引导音效生成
链接:https://arxiv.org/abs/2412.09168
备注:16 pages, 4 figures
摘要:为产品级视频生成声音效果,其中只有少量的标记数据可用于各种场景,需要在Few-Shot设置中生成高质量的声音。为了解决现实场景中有限的标记数据的挑战,我们引入了YingSound,这是一种为视频引导的声音生成而设计的基础模型,支持在Few-Shot设置中生成高质量的音频。具体来说,YingSound包括两个主要模块。第一个模块使用一个条件流匹配Transformer,以实现有效的语义对齐的声音生成跨音频和视觉模态。本模块旨在构建一个可学习的视听聚合器(AVA),在多个阶段将高分辨率的视觉特征与相应的音频特征集成在一起。第二个模块的开发与建议的多模态视听链的思想(CoT)的方法,以产生更好的声音效果,在Few-Shot设置。最后,介绍了一个包含各种真实场景的行业标准视频到音频(V2 A)数据集。我们表明,YingSound通过自动评估和人类研究,在不同的条件输入中有效地生成高质量的同步声音。项目页面:\url{https://giantailab.github.io/yingsound/}
摘要:Generating sound effects for product-level videos, where only a small amountof labeled data is available for diverse scenes, requires the production ofhigh-quality sounds in few-shot settings. To tackle the challenge of limitedlabeled data in real-world scenes, we introduce YingSound, a foundation modeldesigned for video-guided sound generation that supports high-quality audiogeneration in few-shot settings. Specifically, YingSound consists of two majormodules. The first module uses a conditional flow matching transformer toachieve effective semantic alignment in sound generation across audio andvisual modalities. This module aims to build a learnable audio-visualaggregator (AVA) that integrates high-resolution visual features withcorresponding audio features at multiple stages. The second module is developedwith a proposed multi-modal visual-audio chain-of-thought (CoT) approach togenerate finer sound effects in few-shot settings. Finally, anindustry-standard video-to-audio (V2A) dataset that encompasses variousreal-world scenarios is presented. We show that YingSound effectively generateshigh-quality synchronized sounds across diverse conditional inputs throughautomated evaluations and human studies. Project Page:\url{https://giantailab.github.io/yingsound/}
标题:语音取证:走向全面的合成语音数据集建立和分析
链接:https://arxiv.org/abs/2412.09032
备注:None
摘要:由于错误信息和身份冒充的风险,从真实语音中检测合成语音变得越来越重要。虽然已经开发了各种用于合成语音分析的数据集,但它们通常集中在特定领域,限制了它们在综合研究中的实用性。为了填补这一空白,我们提出了Speech-Forensics数据集,它广泛覆盖了真实的、合成的和部分伪造的语音样本,其中包括由不同高质量算法合成的多个片段。此外,我们提出了一个TEMPORAL语音定位网络,称为测试,旨在同时执行真实性检测,多个假段定位,和合成算法识别,没有任何复杂的后处理。TEST有效地集成了LSTM和Transformer,以提取更强大的时间语音表示,并利用多尺度金字塔特征的密集预测来估计合成跨度。我们的模型在话语水平上实现了83.55%的平均mAP和5.25%的EER。在段水平上,它达到了1.07%的EER和92.19%的F1得分。这些结果突出了该模型的综合分析的合成语音的鲁棒能力,在这一领域的未来研究和实际应用提供了一个有前途的途径。
摘要:Detecting synthetic from real speech is increasingly crucial due to the risksof misinformation and identity impersonation. While various datasets forsynthetic speech analysis have been developed, they often focus on specificareas, limiting their utility for comprehensive research. To fill this gap, wepropose the Speech-Forensics dataset by extensively covering authentic,synthetic, and partially forged speech samples that include multiple segmentssynthesized by different high-quality algorithms. Moreover, we propose aTEmporal Speech LocalizaTion network, called TEST, aiming at simultaneouslyperforming authenticity detection, multiple fake segments localization, andsynthesis algorithms recognition, without any complex post-processing. TESTeffectively integrates LSTM and Transformer to extract more powerful temporalspeech representations and utilizes dense prediction on multi-scale pyramidfeatures to estimate the synthetic spans. Our model achieves an average mAP of83.55% and an EER of 5.25% at the utterance level. At the segment level, itattains an EER of 1.07% and a 92.19% F1 score. These results highlight themodel's robust capability for a comprehensive analysis of synthetic speech,offering a promising avenue for future research and practical applications inthis field.
标题:配音:迈向高质量和情感可控的电影配音
链接:https://arxiv.org/abs/2412.08988
备注:Under review
摘要:给定一段文本、一个视频剪辑和一个参考音频,电影配音任务的目标是生成与视频一致的语音,同时克隆所需的语音。现有的方法有两个主要的缺陷:(1)他们努力同时保持视听同步,实现清晰的发音;(2)他们缺乏表达用户定义的情感的能力。为了解决这些问题,我们提出了一种情感可控的配音架构,允许用户指定情感类型和情感强度,同时满足高质量的唇同步和发音。具体来说,我们首先设计了唇相关韵律对齐(LPA),它侧重于学习唇运动和韵律变化之间的内在一致性,通过持续时间水平的对比学习,以纳入合理的对齐。然后,我们设计了发音增强(PE)策略,通过有效的一致性融合视频级音素序列,以提高语音可懂度。其次,说话人身份自适应模块的目标是解码声学先验和注入说话人风格嵌入。在此基础上,提出了基于流的用户情绪控制(FUEC)的合成波形流匹配预测网络的声学先验条件。在这个过程中,FUEC基于用户的情感指令,通过积极和消极的引导机制来确定梯度方向和引导尺度,该机制侧重于放大期望的情感而抑制其他情感。在三个基准数据集上的大量实验结果表明,与几种最先进的方法相比,该方法具有良好的性能。
摘要:Given a piece of text, a video clip, and a reference audio, the movie dubbingtask aims to generate speech that aligns with the video while cloning thedesired voice. The existing methods have two primary deficiencies: (1) Theystruggle to simultaneously hold audio-visual sync and achieve clearpronunciation; (2) They lack the capacity to express user-defined emotions. Toaddress these problems, we propose EmoDubber, an emotion-controllable dubbingarchitecture that allows users to specify emotion type and emotional intensitywhile satisfying high-quality lip sync and pronunciation. Specifically, wefirst design Lip-related Prosody Aligning (LPA), which focuses on learning theinherent consistency between lip motion and prosody variation by duration levelcontrastive learning to incorporate reasonable alignment. Then, we designPronunciation Enhancing (PE) strategy to fuse the video-level phoneme sequencesby efficient conformer to improve speech intelligibility. Next, the speakeridentity adapting module aims to decode acoustics prior and inject the speakerstyle embedding. After that, the proposed Flow-based User Emotion Controlling(FUEC) is used to synthesize waveform by flow matching prediction networkconditioned on acoustics prior. In this process, the FUEC determines thegradient direction and guidance scale based on the user's emotion instructionsby the positive and negative guidance mechanism, which focuses on amplifyingthe desired emotion while suppressing others. Extensive experimental results onthree benchmark datasets demonstrate favorable performance compared to severalstate-of-the-art methods.
标题:用音乐解释图形符号LDM:科尼利厄斯·卡杜(Cornelius Cardew)广告的人工智能即兴创作
链接:https://arxiv.org/abs/2412.08944
备注:None
摘要:这项工作提出了一种新颖的作曲和即兴创作音乐的方法,灵感来自科尼利厄斯·卡杜(Cornelius Cardew)的音乐创作,使用人工智能来连接图形符号和音乐表达。通过利用OpenAI的ChatGPT来解释Treatise的抽象视觉元素,我们将这些图形图像转换为描述性文本提示。然后,这些提示被输入到MusicLDM中,这是一个为音乐生成而设计的预训练潜在扩散模型。我们引入了一种名为“outpainting”的技术,它将AI生成的音乐部分重叠,以创建无缝和有凝聚力的作品。我们展示了表演和解释图形乐谱的新视角,展示了人工智能如何将视觉刺激转化为声音,并扩展了当代/实验音乐创作的可能性。音乐作品可在https://bit.ly/TreatiseAI
摘要:This work presents a novel method for composing and improvising musicinspired by Cornelius Cardew's Treatise, using AI to bridge graphic notationand musical expression. By leveraging OpenAI's ChatGPT to interpret theabstract visual elements of Treatise, we convert these graphical images intodescriptive textual prompts. These prompts are then input into MusicLDM, apre-trained latent diffusion model designed for music generation. We introducea technique called "outpainting," which overlaps sections of AI-generated musicto create a seamless and cohesive composition. We demostrate a new perspectiveon performing and interpreting graphic scores, showing how AI can transformvisual stimuli into sound and expand the creative possibilities incontemporary/experimental music composition. Musical pieces are available athttps://bit.ly/TreatiseAI
标题:单耳语音增强的复循环一致扩散模型
链接:https://arxiv.org/abs/2412.08856
备注:AAAI 2025
摘要:本文提出了一种基于扩散模型的单声道语音增强方法。我们的方法结合了语音频谱的幅度和相位的两个扩散网络的单独估计。在整个扩散过程中,噪声剪辑从现实世界的噪声干扰逐渐添加到干净的语音频谱和噪声感知的反向过程中提出了学习如何生成干净的语音频谱和噪声频谱。此外,为了充分利用幅度和相位之间的内在关系,我们引入了一个复杂的周期一致性(CCC)机制,使用估计的幅度来映射相位,反之亦然。我们在一个相位感知的语音增强扩散模型(SEDM)内实现了该算法。我们在公共数据集上进行了广泛的实验,以证明我们方法的有效性,强调了利用相位和幅度信息之间的内在关系来增强语音的显着好处。与传统扩散模型的比较表明了SEDM的优越性。
摘要:In this paper, we present a novel diffusion model-based monaural speechenhancement method. Our approach incorporates the separate estimation of speechspectra's magnitude and phase in two diffusion networks. Throughout thediffusion process, noise clips from real-world noise interferences are addedgradually to the clean speech spectra and a noise-aware reverse process isproposed to learn how to generate both clean speech spectra and noise spectra.Furthermore, to fully leverage the intrinsic relationship between magnitude andphase, we introduce a complex-cycle-consistent (CCC) mechanism that uses theestimated magnitude to map the phase, and vice versa. We implement thisalgorithm within a phase-aware speech enhancement diffusion model (SEDM). Weconduct extensive experiments on public datasets to demonstrate theeffectiveness of our method, highlighting the significant benefits ofexploiting the intrinsic relationship between phase and magnitude informationto enhance speech. The comparison to conventional diffusion models demonstratesthe superiority of SEDM.
标题:利用动态注意机制进行基于情绪越南语言语的抑郁症诊断
链接:https://arxiv.org/abs/2412.08683
备注:9 Page, 5 Figures
摘要:重度抑郁症是一种普遍而严重的心理健康状况,对您的情绪,思想,行动和对世界的整体感知产生负面影响。由于抑郁症的症状不明显,确定一个人是否抑郁是很复杂的。然而,他们的声音可能是我们可以识别抑郁症迹象的因素之一。抑郁的人会表现出不适、悲伤,他们可能说话缓慢、颤抖,声音中失去情感。在这项研究中,我们提出了动态卷积块注意力模块(Dynamic-CBAM),用于在注意力-GRU网络中通过分析人类的音频信号来分类情感。根据结果,我们可以诊断哪些患者患有抑郁症或容易患抑郁症,以便尽快开始治疗和预防。该研究深入研究了实现Attention-GRU深度学习架构所涉及的复杂计算步骤。通过实验,该模型在VNEMOS数据集中取得了令人印象深刻的识别率,未加权准确率(UA)为0.87,加权准确率(WA)为0.86,F1率为0.87。培训代码发布在https://github.com/fiyud/Emotional-Vietnamese-Speech-Based-Depression-Diagnosis-Using-Dynamic-Attention-Mechanism
摘要:Major depressive disorder is a prevalent and serious mental health conditionthat negatively impacts your emotions, thoughts, actions, and overallperception of the world. It is complicated to determine whether a person isdepressed due to the symptoms of depression not apparent. However, their voicecan be one of the factor from which we can acknowledge signs of depression.People who are depressed express discomfort, sadness and they may speak slowly,trembly, and lose emotion in their voices. In this study, we proposed theDynamic Convolutional Block Attention Module (Dynamic-CBAM) to utilized with inan Attention-GRU Network to classify the emotions by analyzing the audio signalof humans. Based on the results, we can diagnose which patients are depressedor prone to depression then so that treatment and prevention can be started assoon as possible. The research delves into the intricate computational stepsinvolved in implementing a Attention-GRU deep learning architecture. Throughexperimentation, the model has achieved an impressive recognition withUnweighted Accuracy (UA) rate of 0.87 and 0.86 Weighted Accuracy (WA) rate andF1 rate of 0.87 in the VNEMOS dataset. Training code is released inhttps://github.com/fiyud/Emotional-Vietnamese-Speech-Based-Depression-Diagnosis-Using-Dynamic-Attention-Mechanism
标题:利用非峰值CIC丢失和深度语言后注入来增强代码转换ASB
链接:https://arxiv.org/abs/2412.08651
备注:SLT 2024
摘要:语码转换--多语言使用者在对话过程中交替切换语言--由于声学和语义混淆的现象,仍然对端到端(E2 E)自动语音识别(ASR)系统提出了重大挑战。这个问题的出现是因为ASR系统难以有效地处理语言的快速交替,这通常会导致显着的性能下降。我们的主要贡献至少有三个方面:首先,我们将语言识别(LID)信息纳入编码器的几个中间层,旨在用更详细的语言信息丰富输出嵌入。其次,通过语言边界对齐丢失的新颖应用,使后续的ASR模块能够更有效地利用内部语言后验子的知识。第三,我们探讨了使用语言后验子来促进共享编码器和特定语言编码器之间的深度交互的可行性。通过对SEAME语料库的综合实验,我们已经验证了我们提出的方法优于现有技术的方法,基于解纠缠的混合专家(D-MoE),进一步提高了编码器的敏锐度的语言。
摘要:Code-switching-where multilingual speakers alternately switch betweenlanguages during conversations-still poses significant challenges to end-to-end(E2E) automatic speech recognition (ASR) systems due to phenomena of bothacoustic and semantic confusion. This issue arises because ASR systems struggleto handle the rapid alternation of languages effectively, which often leads tosignificant performance degradation. Our main contributions are at leastthreefold: First, we incorporate language identification (LID) information intoseveral intermediate layers of the encoder, aiming to enrich output embeddingswith more detailed language information. Secondly, through the novelapplication of language boundary alignment loss, the subsequent ASR modules areenabled to more effectively utilize the knowledge of internal languageposteriors. Third, we explore the feasibility of using language posteriors tofacilitate deep interaction between shared encoder and language-specificencoders. Through comprehensive experiments on the SEAME corpus, we haveverified that our proposed method outperforms the prior-art method, disentanglebased mixture-of-experts (D-MoE), further enhancing the acuity of the encoderto languages.
标题:用于压缩学习的学习压缩
链接:https://arxiv.org/abs/2412.09405
备注:Accepted as paper to 2025 IEEE Data Compression Conference
摘要:现代传感器产生越来越丰富的高分辨率数据流。由于资源限制,机器学习系统通过分辨率降低丢弃了绝大多数信息。压缩域学习允许模型对紧凑的潜在表示进行操作,从而在相同的预算下实现更高的有效分辨率。然而,现有的压缩系统对于压缩学习并不理想。线性变换编码和端到端学习压缩系统降低了比特率,但不会均匀地降低维度;因此,它们不会有意义地提高效率。生成式自动编码器降低了维度,但它们的对抗性或感知性目标导致了显著的信息丢失。为了解决这些限制,我们引入了WaLLoC(小波学习有损压缩),一种神经编解码器架构,它将线性变换编码与非线性降维自编码器相结合。WaLLoC在可逆小波包变换之间夹有一个浅的、不对称的自动编码器和熵瓶颈。在几个关键指标上,WaLLoC优于最先进的潜在扩散模型中使用的自动编码器。WaLLoC不需要感知或对抗性损失来表示高频细节,从而提供与RGB图像和立体声音频以外的模式的兼容性。WaLLoC的编码器几乎完全由线性运算组成,因此非常高效,适用于移动计算、遥感和直接从压缩数据中学习。我们展示了WaLLoC在多个任务中进行压缩域学习的能力,包括图像分类,着色,文档理解和音乐源分离。我们的代码、实验和预训练的音频和图像编解码器可在https://ut-sysml.org/walloc上获得
摘要:Modern sensors produce increasingly rich streams of high-resolution data. Dueto resource constraints, machine learning systems discard the vast majority ofthis information via resolution reduction. Compressed-domain learning allowsmodels to operate on compact latent representations, allowing higher effectiveresolution for the same budget. However, existing compression systems are notideal for compressed learning. Linear transform coding and end-to-end learnedcompression systems reduce bitrate, but do not uniformly reduce dimensionality;thus, they do not meaningfully increase efficiency. Generative autoencodersreduce dimensionality, but their adversarial or perceptual objectives lead tosignificant information loss. To address these limitations, we introduce WaLLoC(Wavelet Learned Lossy Compression), a neural codec architecture that combineslinear transform coding with nonlinear dimensionality-reducing autoencoders.WaLLoC sandwiches a shallow, asymmetric autoencoder and entropy bottleneckbetween an invertible wavelet packet transform. Across several key metrics,WaLLoC outperforms the autoencoders used in state-of-the-art latent diffusionmodels. WaLLoC does not require perceptual or adversarial losses to representhigh-frequency detail, providing compatibility with modalities beyond RGBimages and stereo audio. WaLLoC's encoder consists almost entirely of linearoperations, making it exceptionally efficient and suitable for mobilecomputing, remote sensing, and learning directly from compressed data. Wedemonstrate WaLLoC's capability for compressed-domain learning across severaltasks, including image classification, colorization, document understanding,and music source separation. Our code, experiments, and pre-trained audio andimage codecs are available at https://ut-sysml.org/walloc
标题:CSSinger:基于条件变分自动编码器的端到端块流媒体歌唱语音合成系统
链接:https://arxiv.org/abs/2412.08918
备注:Accepted by AAAI2025
摘要:歌唱声音合成(SVS)的目标是生成具有高保真度和表现力的歌唱声音。传统的SVS系统通常利用声学模型将乐谱转换成声学特征,然后利用声码器重建歌唱声音。最近表明,端到端建模在SVS和文本到语音(TTS)领域是有效的。因此,在这项工作中,我们提出了一个完全端到端的SVS方法,以及一个分块流推理,以解决实际使用的延迟问题。请注意,这是第一次尝试使用VAE中的潜在表示来完全实现端到端流式音频合成。我们已经做出了具体的改进,以提高使用潜在表示的流SVS的性能。实验结果表明,该方法在流SVS和TTS任务中实现了具有高表现力和音高准确性的合成音频。
摘要:Singing Voice Synthesis (SVS) {aims} to generate singing voices {of high}fidelity and expressiveness. {Conventional SVS systems usually utilize} anacoustic model to transform a music score into acoustic features, {followed bya vocoder to reconstruct the} singing voice. It was recently shown thatend-to-end modeling is effective in the fields of SVS and Text to Speech (TTS).In this work, we thus present a fully end-to-end SVS method together with achunkwise streaming inference to address the latency issue for practicalusages. Note that this is the first attempt to fully implement end-to-endstreaming audio synthesis using latent representations in VAE. We have madespecific improvements to enhance the performance of streaming SVS using latentrepresentations. Experimental results demonstrate that the proposed methodachieves synthesized audio with high expressiveness and pitch accuracy in bothstreaming SVS and TTS tasks.
标题:利用非峰值CIC丢失和深度语言后注入来增强代码转换ASB
链接:https://arxiv.org/abs/2412.08651
备注:SLT 2024
摘要:语码转换--多语言使用者在对话过程中交替切换语言--由于声学和语义混淆的现象,仍然对端到端(E2 E)自动语音识别(ASR)系统提出了重大挑战。这个问题的出现是因为ASR系统难以有效地处理语言的快速交替,这通常会导致显着的性能下降。我们的主要贡献至少有三个方面:首先,我们将语言识别(LID)信息纳入编码器的几个中间层,旨在用更详细的语言信息丰富输出嵌入。其次,通过语言边界对齐损失的新应用,使得后续的ASR模块能够更有效地利用内部语言后验子的知识。第三,我们探讨了使用语言后验子来促进共享编码器和特定语言编码器之间的深度交互的可行性。通过对SEAME语料库的综合实验,我们已经验证了我们提出的方法优于现有技术的方法,基于解纠缠的混合专家(D-MoE),进一步提高了编码器的敏锐度的语言。
摘要:Code-switching-where multilingual speakers alternately switch betweenlanguages during conversations-still poses significant challenges to end-to-end(E2E) automatic speech recognition (ASR) systems due to phenomena of bothacoustic and semantic confusion. This issue arises because ASR systems struggleto handle the rapid alternation of languages effectively, which often leads tosignificant performance degradation. Our main contributions are at leastthreefold: First, we incorporate language identification (LID) information intoseveral intermediate layers of the encoder, aiming to enrich output embeddingswith more detailed language information. Secondly, through the novelapplication of language boundary alignment loss, the subsequent ASR modules areenabled to more effectively utilize the knowledge of internal languageposteriors. Third, we explore the feasibility of using language posteriors tofacilitate deep interaction between shared encoder and language-specificencoders. Through comprehensive experiments on the SEAME corpus, we haveverified that our proposed method outperforms the prior-art method, disentanglebased mixture-of-experts (D-MoE), further enhancing the acuity of the encoderto languages.
标题:Audios不撒谎:用于音频深度伪造检测的多频通道注意机制
链接:https://arxiv.org/abs/2412.09467
摘要:随着人工智能技术的快速发展,Deepfake技术在音频领域的应用逐渐增多,由此带来的安全隐患也越来越广泛。特别是在金融和社会安全领域,deepfake音频的滥用引起了严重关注。为了应对这一挑战,本研究提出了一种基于多频道注意机制(MFCA)和2D离散余弦变换(DCT)的音频深度伪造检测方法。该方法通过将音频信号处理成melspectrogram,利用MobileNet V2提取深度特征,结合MFCA模块对音频信号中不同频率通道进行加权,能够有效捕捉音频信号中的细粒度频域特征,增强对虚假音频的分类能力。实验结果表明,与传统方法相比,本文提出的模型在准确率、精确率、召回率、F1评分等指标上表现出明显的优势。特别是在复杂音频场景下,该方法表现出更强的鲁棒性和泛化能力,为音频深度伪造检测提供了新的思路,具有重要的实际应用价值。未来将探索更先进的音频检测技术和优化策略,进一步提高音频deepfake检测的准确性和泛化能力。
摘要:With the rapid development of artificial intelligence technology, theapplication of deepfake technology in the audio field has gradually increased,resulting in a wide range of security risks. Especially in the financial andsocial security fields, the misuse of deepfake audios has raised seriousconcerns. To address this challenge, this study proposes an audio deepfakedetection method based on multi-frequency channel attention mechanism (MFCA)and 2D discrete cosine transform (DCT). By processing the audio signal into amelspectrogram, using MobileNet V2 to extract deep features, and combining itwith the MFCA module to weight different frequency channels in the audiosignal, this method can effectively capture the fine-grained frequency domainfeatures in the audio signal and enhance the Classification capability of fakeaudios. Experimental results show that compared with traditional methods, themodel proposed in this study shows significant advantages in accuracy,precision,recall, F1 score and other indicators. Especially in complex audioscenarios, this method shows stronger robustness and generalizationcapabilities and provides a new idea for audio deepfake detection and hasimportant practical application value. In the future, more advanced audiodetection technologies and optimization strategies will be explored to furtherimprove the accuracy and generalization capabilities of audio deepfakedetection.
标题:具有显式桥梁和检索增强的多模式音乐生成
链接:https://arxiv.org/abs/2412.09428
摘要:多模态音乐生成旨在从不同的输入模态(包括文本、视频和图像)生成音乐。现有的方法使用一个共同的嵌入空间进行多模态融合。尽管它们在其他模态中是有效的,但它们在多模态音乐生成中的应用面临着数据稀缺、弱跨模态对齐和有限可控性的挑战。本文通过使用文本和音乐的多模态对齐明确的桥梁来解决这些问题。我们介绍了一种新的方法命名为视觉音乐桥(VMB)。具体而言,多模态音乐描述模型将视觉输入转换为详细的文本描述,以提供文本桥;双轨音乐检索模块,其结合了广泛的和有针对性的检索策略,以提供音乐桥并使用户能够控制。最后,我们设计了一个基于这两个桥梁的音乐生成框架。我们进行了视频到音乐、图像到音乐、文本到音乐和可控音乐生成任务的实验,以及可控性实验。结果表明,VMB显着提高了音乐质量,模态和定制对齐相比,以前的方法。VMB为可解释性和表现力的多模态音乐生成设定了新的标准,并在各种多媒体领域中应用。演示和代码可在https://github.com/wbs2788/VMB上获得。
摘要:Multimodal music generation aims to produce music from diverse inputmodalities, including text, videos, and images. Existing methods use a commonembedding space for multimodal fusion. Despite their effectiveness in othermodalities, their application in multimodal music generation faces challengesof data scarcity, weak cross-modal alignment, and limited controllability. Thispaper addresses these issues by using explicit bridges of text and music formultimodal alignment. We introduce a novel method named Visuals Music Bridge(VMB). Specifically, a Multimodal Music Description Model converts visualinputs into detailed textual descriptions to provide the text bridge; aDual-track Music Retrieval module that combines broad and targeted retrievalstrategies to provide the music bridge and enable user control. Finally, wedesign an Explicitly Conditioned Music Generation framework to generate musicbased on the two bridges. We conduct experiments on video-to-music,image-to-music, text-to-music, and controllable music generation tasks, alongwith experiments on controllability. The results demonstrate that VMBsignificantly enhances music quality, modality, and customization alignmentcompared to previous methods. VMB sets a new standard for interpretable andexpressive multimodal music generation with applications in various multimediafields. Demos and code are available at https://github.com/wbs2788/VMB.
标题:基于视频和音频输入的多模式情绪分析
链接:https://arxiv.org/abs/2412.09317
备注:Presented as a full paper in the 15th International Conference on Emerging Ubiquitous Systems and Pervasive Networks (EUSPN 2024) October 28-30, 2024, Leuven, Belgium
摘要:尽管目前有大量的研究致力于从视频和音频中进行情感分析,但找到最好的模型,提供最高的准确率仍然被认为是该领域研究人员的挑战。本文的主要目的是证明的可用性的情感识别模型,视频和音频输入。用于训练模型的数据集是用于音频的CREMA-D数据集和用于视频的RAVDESS数据集。使用的微调模型是:Facebook/wav 2 vec 2-large用于音频,Google/vivit-b-16 x2-kinetics 400用于视频。在决策框架中利用了由两个先前模型生成的每个情感的概率的平均值。在结果不一致之后,如果其中一个模型获得更高的准确性,则创建另一个测试框架。采用的方法有加权平均法、置信度阈值法、基于置信度的动态加权法和基于规则的逻辑法。这种有限的方法给出了令人鼓舞的结果,使未来的研究这些方法可行。
摘要:Despite the abundance of current researches working on the sentiment analysisfrom videos and audios, finding the best model that gives the highest accuracyrate is still considered a challenge for researchers in this field. The mainobjective of this paper is to prove the usability of emotion recognition modelsthat take video and audio inputs. The datasets used to train the models are theCREMA-D dataset for audio and the RAVDESS dataset for video. The fine-tunedmodels that been used are: Facebook/wav2vec2-large for audio and theGoogle/vivit-b-16x2-kinetics400 for video. The avarage of the probabilities foreach emotion generated by the two previous models is utilized in the decisionmaking framework. After disparity in the results, if one of the models getsmuch higher accuracy, another test framework is created. The methods used arethe Weighted Average method, the Confidence Level Threshold method, the DynamicWeighting Based on Confidence method, and the Rule-Based Logic method. Thislimited approach gives encouraging results that make future research into thesemethods viable.
标题:语音隐私保护中说话人对抗性扰动的产生和消除
链接:https://arxiv.org/abs/2412.09195
备注:6 pages, 3 figures, published to IEEE SLT Workshop 2024
摘要:众所周知,神经网络容易受到通过对输入数据进行细微扰动而进行的对抗性攻击。语音隐私保护的最新发展已经显示了相同技术的积极用例,以利用对抗网络生成的附加扰动信号来隐藏说话者的语音属性。本文研究了可逆性属性,其中生成对抗性扰动的实体被授权移除它们并恢复原始语音(例如,演讲者本人)。调查人员也可以使用类似的技术来对语音保护的语音进行去匿名化,以在安全和法医分析中恢复罪犯的身份。在该设置中,假设扰动生成模块在去除处理中是已知的。为此,提出了扰动生成和去除模块的联合训练。在LibriSpeech数据集上的实验结果表明,该方法可以从匿名语音中预测出添加到原始语音中的细微扰动,同时达到隐私保护的目的。通过从匿名样本中去除这些扰动,可以恢复原始语音。音频示例可以在\url{https://voiceprivacy.github.io/Perturbation-Generation-Removal/}中找到。
摘要:Neural networks are commonly known to be vulnerable to adversarial attacksmounted through subtle perturbation on the input data. Recent development invoice-privacy protection has shown the positive use cases of the same techniqueto conceal speaker's voice attribute with additive perturbation signalgenerated by an adversarial network. This paper examines the reversibilityproperty where an entity generating the adversarial perturbations is authorizedto remove them and restore original speech (e.g., the speaker him/herself). Asimilar technique could also be used by an investigator to deanonymize avoice-protected speech to restore criminals' identities in security andforensic analysis. In this setting, the perturbation generative module isassumed to be known in the removal process. To this end, a joint training ofperturbation generation and removal modules is proposed. Experimental resultson the LibriSpeech dataset demonstrated that the subtle perturbations added tothe original speech can be predicted from the anonymized speech while achievingthe goal of privacy protection. By removing these perturbations from theanonymized sample, the original speech can be restored. Audio samples can befound in \url{https://voiceprivacy.github.io/Perturbation-Generation-Removal/}.
标题:YingSound:具有多模式思维链控制的视频引导音效生成
链接:https://arxiv.org/abs/2412.09168
备注:16 pages, 4 figures
摘要:为产品级视频生成声音效果,其中只有少量的标记数据可用于各种场景,需要在Few-Shot设置中生成高质量的声音。为了解决现实场景中有限的标记数据的挑战,我们引入了YingSound,这是一种为视频引导的声音生成而设计的基础模型,支持在Few-Shot设置中生成高质量的音频。具体来说,YingSound包括两个主要模块。第一个模块使用一个条件流匹配Transformer,以实现有效的语义对齐的声音生成跨音频和视觉模态。本模块旨在构建一个可学习的视听聚合器(AVA),在多个阶段将高分辨率的视觉特征与相应的音频特征集成在一起。第二个模块的开发与建议的多模态视听链的思想(CoT)的方法,以产生更好的声音效果,在Few-Shot设置。最后,介绍了一个包含各种真实场景的行业标准视频到音频(V2 A)数据集。我们表明,YingSound通过自动评估和人类研究,在各种条件输入中有效地生成高质量的同步声音。项目页面:\url{https://giantailab.github.io/yingsound/}
摘要:Generating sound effects for product-level videos, where only a small amountof labeled data is available for diverse scenes, requires the production ofhigh-quality sounds in few-shot settings. To tackle the challenge of limitedlabeled data in real-world scenes, we introduce YingSound, a foundation modeldesigned for video-guided sound generation that supports high-quality audiogeneration in few-shot settings. Specifically, YingSound consists of two majormodules. The first module uses a conditional flow matching transformer toachieve effective semantic alignment in sound generation across audio andvisual modalities. This module aims to build a learnable audio-visualaggregator (AVA) that integrates high-resolution visual features withcorresponding audio features at multiple stages. The second module is developedwith a proposed multi-modal visual-audio chain-of-thought (CoT) approach togenerate finer sound effects in few-shot settings. Finally, anindustry-standard video-to-audio (V2A) dataset that encompasses variousreal-world scenarios is presented. We show that YingSound effectively generateshigh-quality synchronized sounds across diverse conditional inputs throughautomated evaluations and human studies. Project Page:\url{https://giantailab.github.io/yingsound/}
标题:语音取证:走向全面的合成语音数据集建立和分析
链接:https://arxiv.org/abs/2412.09032
备注:None
摘要:由于错误信息和身份冒充的风险,从真实语音中检测合成语音变得越来越重要。虽然已经开发了各种用于合成语音分析的数据集,但它们通常集中在特定领域,限制了它们在综合研究中的实用性。为了填补这一空白,我们提出了语音取证数据集,广泛覆盖真实、合成和部分伪造的语音样本,其中包括由不同高质量算法合成的多个片段。此外,我们提出了一个TEMPORAL语音定位网络,称为测试,旨在同时执行真实性检测,多个假段定位,和合成算法识别,没有任何复杂的后处理。TEST有效地集成了LSTM和Transformer,以提取更强大的时间语音表示,并利用多尺度金字塔特征的密集预测来估计合成跨度。我们的模型在话语级别实现了83.55%的平均mAP和5.25%的EER。在段水平上,它达到了1.07%的EER和92.19%的F1得分。这些结果突出了该模型的综合分析的合成语音的鲁棒能力,在这一领域的未来研究和实际应用提供了一个有前途的途径。
摘要:Detecting synthetic from real speech is increasingly crucial due to the risksof misinformation and identity impersonation. While various datasets forsynthetic speech analysis have been developed, they often focus on specificareas, limiting their utility for comprehensive research. To fill this gap, wepropose the Speech-Forensics dataset by extensively covering authentic,synthetic, and partially forged speech samples that include multiple segmentssynthesized by different high-quality algorithms. Moreover, we propose aTEmporal Speech LocalizaTion network, called TEST, aiming at simultaneouslyperforming authenticity detection, multiple fake segments localization, andsynthesis algorithms recognition, without any complex post-processing. TESTeffectively integrates LSTM and Transformer to extract more powerful temporalspeech representations and utilizes dense prediction on multi-scale pyramidfeatures to estimate the synthetic spans. Our model achieves an average mAP of83.55% and an EER of 5.25% at the utterance level. At the segment level, itattains an EER of 1.07% and a 92.19% F1 score. These results highlight themodel's robust capability for a comprehensive analysis of synthetic speech,offering a promising avenue for future research and practical applications inthis field.
标题:配音:迈向高质量和情感可控的电影配音
链接:https://arxiv.org/abs/2412.08988
备注:Under review
摘要:给定一段文本、一个视频剪辑和一个参考音频,电影配音任务的目标是生成与视频一致的语音,同时克隆所需的语音。现有的方法有两个主要的缺陷:(1)他们努力同时保持视听同步,实现清晰的发音;(2)他们缺乏表达用户定义的情感的能力。为了解决这些问题,我们提出了一种情感可控的配音架构,允许用户指定情感类型和情感强度,同时满足高质量的唇同步和发音。具体来说,我们首先设计了唇相关韵律对齐(LPA),它侧重于学习唇运动和韵律变化之间的内在一致性,通过持续时间水平的对比学习,以纳入合理的对齐。然后,我们设计了发音增强(PE)策略,通过有效的一致性融合视频级音素序列,以提高语音可懂度。其次,说话人身份自适应模块的目标是解码声学先验和注入说话人风格嵌入。在此基础上,提出了基于流的用户情绪控制(FUEC)的合成波形流匹配预测网络的声学先验条件。在这个过程中,FUEC基于用户的情感指令,通过积极和消极的引导机制来确定梯度方向和引导尺度,该机制侧重于放大期望的情感而抑制其他情感。在三个基准数据集上的大量实验结果表明,与几种最先进的方法相比,该方法具有良好的性能。
摘要:Given a piece of text, a video clip, and a reference audio, the movie dubbingtask aims to generate speech that aligns with the video while cloning thedesired voice. The existing methods have two primary deficiencies: (1) Theystruggle to simultaneously hold audio-visual sync and achieve clearpronunciation; (2) They lack the capacity to express user-defined emotions. Toaddress these problems, we propose EmoDubber, an emotion-controllable dubbingarchitecture that allows users to specify emotion type and emotional intensitywhile satisfying high-quality lip sync and pronunciation. Specifically, wefirst design Lip-related Prosody Aligning (LPA), which focuses on learning theinherent consistency between lip motion and prosody variation by duration levelcontrastive learning to incorporate reasonable alignment. Then, we designPronunciation Enhancing (PE) strategy to fuse the video-level phoneme sequencesby efficient conformer to improve speech intelligibility. Next, the speakeridentity adapting module aims to decode acoustics prior and inject the speakerstyle embedding. After that, the proposed Flow-based User Emotion Controlling(FUEC) is used to synthesize waveform by flow matching prediction networkconditioned on acoustics prior. In this process, the FUEC determines thegradient direction and guidance scale based on the user's emotion instructionsby the positive and negative guidance mechanism, which focuses on amplifyingthe desired emotion while suppressing others. Extensive experimental results onthree benchmark datasets demonstrate favorable performance compared to severalstate-of-the-art methods.
标题:用音乐解释图形符号LDM:科尼利厄斯·卡杜(Cornelius Cardew)广告的人工智能即兴创作
链接:https://arxiv.org/abs/2412.08944
备注:None
摘要:这项工作提出了一种新颖的作曲和即兴创作音乐的方法,灵感来自科尼利厄斯·卡杜(Cornelius Cardew)的即兴创作,使用人工智能来连接图形符号和音乐表达。通过利用OpenAI的ChatGPT来解释抽象的视觉元素,我们将这些图形图像转换为描述性的文本提示。然后,这些提示被输入到MusicLDM中,这是一个为音乐生成而设计的预训练潜在扩散模型。我们引入了一种名为“outpainting”的技术,它将AI生成的音乐部分重叠,以创建无缝和有凝聚力的作品。我们展示了表演和解释图形乐谱的新视角,展示了人工智能如何将视觉刺激转化为声音,并扩展了当代/实验音乐创作的可能性。音乐作品可在https://bit.ly/TreatiseAI
摘要:This work presents a novel method for composing and improvising musicinspired by Cornelius Cardew's Treatise, using AI to bridge graphic notationand musical expression. By leveraging OpenAI's ChatGPT to interpret theabstract visual elements of Treatise, we convert these graphical images intodescriptive textual prompts. These prompts are then input into MusicLDM, apre-trained latent diffusion model designed for music generation. We introducea technique called "outpainting," which overlaps sections of AI-generated musicto create a seamless and cohesive composition. We demostrate a new perspectiveon performing and interpreting graphic scores, showing how AI can transformvisual stimuli into sound and expand the creative possibilities incontemporary/experimental music composition. Musical pieces are available athttps://bit.ly/TreatiseAI
标题:单耳语音增强的复循环一致扩散模型
链接:https://arxiv.org/abs/2412.08856
备注:AAAI 2025
摘要:本文提出了一种基于扩散模型的单声道语音增强方法。我们的方法结合了语音频谱的幅度和相位的两个扩散网络的单独估计。在整个扩散过程中,噪声剪辑从现实世界的噪声干扰逐渐添加到干净的语音频谱和噪声感知的反向过程中提出了学习如何生成干净的语音频谱和噪声频谱。此外,为了充分利用幅度和相位之间的内在关系,我们引入了一个复杂的周期一致性(CCC)机制,使用估计的幅度来映射相位,反之亦然。我们在一个相位感知的语音增强扩散模型(SEDM)内实现了该算法。我们在公共数据集上进行了广泛的实验,以证明我们的方法的有效性,突出了利用相位和幅度信息之间的内在关系来增强语音的显着好处。与传统扩散模型的比较表明了SEDM的优越性。
摘要:In this paper, we present a novel diffusion model-based monaural speechenhancement method. Our approach incorporates the separate estimation of speechspectra's magnitude and phase in two diffusion networks. Throughout thediffusion process, noise clips from real-world noise interferences are addedgradually to the clean speech spectra and a noise-aware reverse process isproposed to learn how to generate both clean speech spectra and noise spectra.Furthermore, to fully leverage the intrinsic relationship between magnitude andphase, we introduce a complex-cycle-consistent (CCC) mechanism that uses theestimated magnitude to map the phase, and vice versa. We implement thisalgorithm within a phase-aware speech enhancement diffusion model (SEDM). Weconduct extensive experiments on public datasets to demonstrate theeffectiveness of our method, highlighting the significant benefits ofexploiting the intrinsic relationship between phase and magnitude informationto enhance speech. The comparison to conventional diffusion models demonstratesthe superiority of SEDM.
标题:利用动态注意机制进行基于情绪越南语言语的抑郁症诊断
链接:https://arxiv.org/abs/2412.08683
备注:9 Page, 5 Figures
摘要:重度抑郁症是一种普遍而严重的心理健康状况,对您的情绪,思想,行动和对世界的整体感知产生负面影响。由于抑郁症的症状不明显,确定一个人是否抑郁是很复杂的。然而,他们的声音可能是我们可以识别抑郁症迹象的因素之一。抑郁的人会表现出不适、悲伤,他们可能说话缓慢、颤抖,声音中失去情感。在这项研究中,我们提出了动态卷积块注意力模块(Dynamic-CBAM),用于在注意力-GRU网络中通过分析人类的音频信号来分类情感。根据结果,我们可以诊断哪些患者患有抑郁症或容易患抑郁症,以便尽快开始治疗和预防。该研究深入研究了实现Attention-GRU深度学习架构所涉及的复杂计算步骤。通过实验,该模型在VNEMOS数据集中取得了令人印象深刻的识别率,未加权准确率(UA)为0.87,加权准确率(WA)为0.86,F1率为0.87。培训代码发布在https://github.com/fiyud/Emotional-Vietnamese-Speech-Based-Depression-Diagnosis-Using-Dynamic-Attention-Mechanism
摘要:Major depressive disorder is a prevalent and serious mental health conditionthat negatively impacts your emotions, thoughts, actions, and overallperception of the world. It is complicated to determine whether a person isdepressed due to the symptoms of depression not apparent. However, their voicecan be one of the factor from which we can acknowledge signs of depression.People who are depressed express discomfort, sadness and they may speak slowly,trembly, and lose emotion in their voices. In this study, we proposed theDynamic Convolutional Block Attention Module (Dynamic-CBAM) to utilized with inan Attention-GRU Network to classify the emotions by analyzing the audio signalof humans. Based on the results, we can diagnose which patients are depressedor prone to depression then so that treatment and prevention can be started assoon as possible. The research delves into the intricate computational stepsinvolved in implementing a Attention-GRU deep learning architecture. Throughexperimentation, the model has achieved an impressive recognition withUnweighted Accuracy (UA) rate of 0.87 and 0.86 Weighted Accuracy (WA) rate andF1 rate of 0.87 in the VNEMOS dataset. Training code is released inhttps://github.com/fiyud/Emotional-Vietnamese-Speech-Based-Depression-Diagnosis-Using-Dynamic-Attention-Mechanism
