本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2412.11978
备注:Accepted at COLING 2025 main conference
摘要:虽然众包是促进和扩大语音数据收集的既定解决方案,但非专家的参与需要协议来确保最终数据质量。为了降低这些基本控制的成本,本文研究了使用语音基础模型(SFM)自动化的验证过程中,第一次检查数据采集的成本/质量权衡。在法国、德国和韩国数据上进行的实验表明,基于SFM的验证有可能减少对人工验证的依赖,从而在不降低最终数据质量的情况下节省超过40.0%的成本。这些发现为更高效、更具成本效益且可扩展的语音数据采集开辟了新的机会。
摘要:While crowdsourcing is an established solution for facilitating and scalingthe collection of speech data, the involvement of non-experts necessitatesprotocols to ensure final data quality. To reduce the costs of these essentialcontrols, this paper investigates the use of Speech Foundation Models (SFMs) toautomate the validation process, examining for the first time the cost/qualitytrade-off in data acquisition. Experiments conducted on French, German, andKorean data demonstrate that SFM-based validation has the potential to reducereliance on human validation, resulting in an estimated cost saving of over40.0% without degrading final data quality. These findings open newopportunities for more efficient, cost-effective, and scalable speech dataacquisition.
标题: auttrainer:一个用于计算机试镜任务的模块化和可扩展深度学习工具包
链接:https://arxiv.org/abs/2412.11943
摘要:这项工作介绍了autrainer的关键操作原理,autrainer是我们用于计算机试听任务的新深度学习训练框架。autrainer是一个基于PyTorch的工具包,可以对各种不同的计算机试听任务进行快速、可重复且易于扩展的培训。具体来说,autrainer提供低代码训练,并支持广泛的神经网络以及预处理例程。在这项工作中,我们提出了一个概述其内部工作原理和关键功能。
摘要:This work introduces the key operating principles for autrainer, our new deeplearning training framework for computer audition tasks. autrainer is aPyTorch-based toolkit that allows for rapid, reproducible, and easilyextensible training on a variety of different computer audition tasks.Concretely, autrainer offers low-code training and supports a wide range ofneural networks as well as preprocessing routines. In this work, we present anoverview of its inner workings and key capabilities.
标题: AudioCIL:用于音频类-多场景增量学习的Python工作空间
链接:https://arxiv.org/abs/2412.11907
摘要:深度学习凭借其强大的音频特征提取能力,在音频信号处理方面取得了重大成功。通常,这些方法依赖于静态的、预先收集的大规模数据集进行训练,在固定数量的类上表现良好。然而,现实世界的特点是不断变化,新的音频类从流媒体中出现,或者由于隐私而临时可用。音频环境的这种动态性质需要模型能够在不丢弃现有信息的情况下为新类增量地学习新知识。将增量学习引入音频信号处理领域,即,音频课程增量学习(AuCIL)是一项有意义的努力。我们提出了一个名为AudioCIL的工具箱,以使音频信号处理算法与真实世界的场景保持一致,并加强音频类增量学习的研究。
摘要:Deep learning, with its robust aotomatic feature extraction capabilities, hasdemonstrated significant success in audio signal processing. Typically, thesemethods rely on static, pre-collected large-scale datasets for training,performing well on a fixed number of classes. However, the real world ischaracterized by constant change, with new audio classes emerging fromstreaming or temporary availability due to privacy. This dynamic nature ofaudio environments necessitates models that can incrementally learn newknowledge for new classes without discarding existing information. Introducingincremental learning to the field of audio signal processing, i.e., AudioClass-Incremental Learning (AuCIL), is a meaningful endeavor. We propose such atoolbox named AudioCIL to align audio signal processing algorithms withreal-world scenarios and strengthen research in audio class-incrementallearning.
标题: 多语言音频的自发语音和脚本语音的分类
链接:https://arxiv.org/abs/2412.11896
备注:Accepted to IEEE Spoken Language Technology Workshop 2024
摘要:区分讲稿和自发言语是更好地理解言语风格如何影响言语处理研究的重要工具。它还可以通过更好地分割大型录制语音目录来改善媒体用户的推荐系统和发现体验。本文解决了建立一个分类器,以及在不同的格式和语言的概括的挑战。我们系统地评估了从传统的、手工制作的声学和韵律特征到先进的音频Transformers的模型,利用大型的、多语言的专有播客数据集进行训练和验证。我们将每个模型在11个语言组中的表现进行了分解,以评估跨语言偏见。我们的实验分析扩展到公开可用的数据集,以评估模型对非播客领域的普适性。我们的研究结果表明,基于transformer的模型始终优于传统的基于特征的技术,在区分各种语言的脚本和自发语音方面实现了最先进的性能。
摘要:Distinguishing scripted from spontaneous speech is an essential tool forbetter understanding how speech styles influence speech processing research. Itcan also improve recommendation systems and discovery experiences for mediausers through better segmentation of large recorded speech catalogues. Thispaper addresses the challenge of building a classifier that generalises wellacross different formats and languages. We systematically evaluate modelsranging from traditional, handcrafted acoustic and prosodic features toadvanced audio transformers, utilising a large, multilingual proprietarypodcast dataset for training and validation. We break down the performance ofeach model across 11 language groups to evaluate cross-lingual biases. Ourexperimental analysis extends to publicly available datasets to assess themodels' generalisability to non-podcast domains. Our results indicate thattransformer-based models consistently outperform traditional feature-basedtechniques, achieving state-of-the-art performance in distinguishing betweenscripted and spontaneous speech across various languages.
标题: ProsodyFM:用于可理解语音合成的无监督短语和语调控制
链接:https://arxiv.org/abs/2412.11795
备注:Accepted by AAAI 2025
摘要:韵律包含了丰富的信息,超出了字面意思的话,这是至关重要的语音的可理解性。目前的模型在措辞和语调方面仍然存在不足,它们在合成具有复杂结构的长句时不仅会丢失或放错停顿,而且会产生不自然的语调。我们提出ProsodyFM,韵律感知的文本到语音合成(TTS)模型与流匹配(FM)骨干,旨在提高韵律的措辞和语调方面。ProsodyFM引入了两个关键组件:一个短语中断编码器,用于捕获初始短语中断位置,然后是一个持续时间预测器,用于灵活调整中断持续时间;以及一个终端语调编码器,它集成了一组语调形状标记,并结合了一个新颖的音高处理器,用于对人类感知的语调变化进行更鲁棒的建模。ProsodyFM是在没有明确韵律标签的情况下训练的,但可以发现广泛的中断持续时间和语调模式。实验结果表明,ProsodyFM可以有效地改善韵律的措辞和语调方面,从而提高整体的可懂度相比,四个国家的最先进的(SOTA)模型。分布外实验表明,这种韵律改进可以进一步带来ProsodyFM优越的泛化能力,看不见的复杂句子和扬声器。我们的案例研究直观地说明了强大的和细粒度的可控性的ProsodyFM的措辞和语调。
摘要:Prosody contains rich information beyond the literal meaning of words, whichis crucial for the intelligibility of speech. Current models still fall shortin phrasing and intonation; they not only miss or misplace breaks whensynthesizing long sentences with complex structures but also produce unnaturalintonation. We propose ProsodyFM, a prosody-aware text-to-speech synthesis(TTS) model with a flow-matching (FM) backbone that aims to enhance thephrasing and intonation aspects of prosody. ProsodyFM introduces two keycomponents: a Phrase Break Encoder to capture initial phrase break locations,followed by a Duration Predictor for the flexible adjustment of breakdurations; and a Terminal Intonation Encoder which integrates a set ofintonation shape tokens combined with a novel Pitch Processor for more robustmodeling of human-perceived intonation change. ProsodyFM is trained with noexplicit prosodic labels and yet can uncover a broad spectrum of breakdurations and intonation patterns. Experimental results demonstrate thatProsodyFM can effectively improve the phrasing and intonation aspects ofprosody, thereby enhancing the overall intelligibility compared to fourstate-of-the-art (SOTA) models. Out-of-distribution experiments show that thisprosody improvement can further bring ProsodyFM superior generalizability forunseen complex sentences and speakers. Our case study intuitively illustratesthe powerful and fine-grained controllability of ProsodyFM over phrasing andintonation.
标题: 它是Chug吗?以数据驱动的方式理解吉他音调描述
链接:https://arxiv.org/abs/2412.11769
备注:Accepted for publication at the 3rd Workshop on NLP for Music and Audio (NLP4MusA 2024)
摘要:自然语言通常用于描述乐器音色,例如“温暖”或“沉重”的声音。由于这些描述符基于人类感知,因此对于哪些声学特征对应于给定的形容词可能存在分歧。在这项工作中,我们追求一种数据驱动的方法,以进一步了解这些形容词的背景下,吉他音。我们的主要贡献是一个音色形容词的数据集,通过处理乐器音频的单个片段来构建,通过EQ和失真等效果的调整来产生不同的音色。众包专家为每个片段获得形容词注释,以完成成对比较和标记任务。我们检查了数据集,并揭示了形容词评级和突出实例之间的相关性,其中数据与关于光谱特征和音色形容词的流行理论相矛盾,这表明需要对音色进行更细致入微的数据驱动理解。
摘要:Natural language is commonly used to describe instrument timbre, such as a"warm" or "heavy" sound. As these descriptors are based on human perception,there can be disagreement over which acoustic features correspond to a givenadjective. In this work, we pursue a data-driven approach to further ourunderstanding of such adjectives in the context of guitar tone. Our maincontribution is a dataset of timbre adjectives, constructed by processingsingle clips of instrument audio to produce varied timbres through adjustmentsin EQ and effects such as distortion. Adjective annotations are obtained foreach clip by crowdsourcing experts to complete a pairwise comparison and alabeling task. We examine the dataset and reveal correlations between adjectiveratings and highlight instances where the data contradicts prevailing theorieson spectral features and timbral adjectives, suggesting a need for a morenuanced, data-driven understanding of timbre.
标题: 用于增强视听Zero-Shot学习的差异感知注意力网络
链接:https://arxiv.org/abs/2412.11715
摘要:视听零镜头学习(Audio-visual Zero-Shot Learning,简称ZRL)由于能够识别看不见的类别并在视频分类任务中表现良好而受到了广泛关注。然而,模态不平衡的(G)CNOL导致过度依赖的最佳模态,降低了不可见的类的区分能力。一些研究试图通过修改参数梯度来解决这一问题,但仍然存在两个挑战:(a)质量差异,即模式为同一概念提供不同数量和质量的信息。(b)内容差异,其中模态内的样本贡献差异很大。为了应对这些挑战,我们提出了一个差异感知注意力网络(DAAN)增强视听视听网络。我们的方法引入了一个质量离散缓解注意力(QDMA)单元,以最大限度地减少冗余信息的高质量模态和对比样本级梯度调制(CSGM)块,以调整梯度幅度和平衡内容差异。我们量化模态的贡献,整合优化和收敛速度更精确的梯度调制CSGM。实验表明,DAAN在基准数据集上实现了最先进的性能,消融研究验证了各个模块的有效性。
摘要:Audio-visual Zero-Shot Learning (ZSL) has attracted significant attention forits ability to identify unseen classes and perform well in video classificationtasks. However, modal imbalance in (G)ZSL leads to over-reliance on the optimalmodality, reducing discriminative capabilities for unseen classes. Some studieshave attempted to address this issue by modifying parameter gradients, but twochallenges still remain: (a) Quality discrepancies, where modalities offerdiffering quantities and qualities of information for the same concept. (b)Content discrepancies, where sample contributions within a modality varysignificantly. To address these challenges, we propose a Discrepancy-AwareAttention Network (DAAN) for Enhanced Audio-Visual ZSL. Our approach introducesa Quality-Discrepancy Mitigation Attention (QDMA) unit to minimize redundantinformation in the high-quality modality and a Contrastive Sample-levelGradient Modulation (CSGM) block to adjust gradient magnitudes and balancecontent discrepancies. We quantify modality contributions by integratingoptimization and convergence rate for more precise gradient modulation in CSGM.Experiments demonstrates DAAN achieves state-of-the-art performance onbenchmark datasets, with ablation studies validating the effectiveness ofindividual modules.
标题: 音频深度伪造检测的连续学习中基于区域的优化
链接:https://arxiv.org/abs/2412.11551
备注:Accepted by AAAI 2025
摘要:语音合成和语音转换的快速发展带来了便利,但也带来了新的安全风险,迫切需要有效的音频深度伪造检测。尽管目前的模型表现良好,但当面对现实世界中deepfake的多样性和不断变化的性质时,它们的有效性就会降低。为了解决这个问题,我们提出了一种名为基于区域的优化(RegO)的持续学习方法,用于音频deepfake检测。具体来说,我们使用Fisher信息矩阵来测量真实和虚假音频检测的重要神经元区域,将它们分为四个区域。首先,我们直接对不太重要的区域进行微调,以快速适应新任务。接下来,我们对仅对真实音频检测重要的区域并行地应用梯度优化,并且对仅对虚假音频检测重要的区域在正交方向上应用梯度优化。对于对两者都很重要的区域,我们使用基于样本比例的自适应梯度优化。这种区域自适应优化确保了记忆稳定性和学习可塑性之间的适当权衡。此外,为了解决旧任务中冗余神经元的增加,我们进一步引入Ebbinghaus遗忘机制来释放它们,从而提高模型学习更多广义判别特征的能力。实验结果表明,我们的方法在EER方面比最先进的用于音频deepfake检测的持续学习方法RWM提高了21.3%。此外,RegO的有效性超出了音频deepfake检测领域,在其他任务中显示出潜在的意义,例如图像识别。该代码可在https://github.com/cyjie429/RegO上获得
摘要:Rapid advancements in speech synthesis and voice conversion bring conveniencebut also new security risks, creating an urgent need for effective audiodeepfake detection. Although current models perform well, their effectivenessdiminishes when confronted with the diverse and evolving nature of real-worlddeepfakes. To address this issue, we propose a continual learning method namedRegion-Based Optimization (RegO) for audio deepfake detection. Specifically, weuse the Fisher information matrix to measure important neuron regions for realand fake audio detection, dividing them into four regions. First, we directlyfine-tune the less important regions to quickly adapt to new tasks. Next, weapply gradient optimization in parallel for regions important only to realaudio detection, and in orthogonal directions for regions important only tofake audio detection. For regions that are important to both, we use sampleproportion-based adaptive gradient optimization. This region-adaptiveoptimization ensures an appropriate trade-off between memory stability andlearning plasticity. Additionally, to address the increase of redundant neuronsfrom old tasks, we further introduce the Ebbinghaus forgetting mechanism torelease them, thereby promoting the capability of the model to learn moregeneralized discriminative features. Experimental results show our methodachieves a 21.3% improvement in EER over the state-of-the-art continuallearning approach RWM for audio deepfake detection. Moreover, the effectivenessof RegO extends beyond the audio deepfake detection domain, showing potentialsignificance in other tasks, such as image recognition. The code is availableat https://github.com/cyjie429/RegO
标题: Whisper-GPT:混合表示音频大语言模型
链接:https://arxiv.org/abs/2412.11449
备注:6 pages, 3 figures. 50th International Conference on Acoustics, Speech and Signal Processing, Hyderabad, India
摘要:我们建议WHISPER-GPT:一个用于语音和音乐的生成式大型语言模型(LLM),允许我们同时处理连续音频表示和离散令牌,作为单个架构的一部分。生成音频、语音和音乐模型已经出现了巨大的激增,这些模型利用了从神经压缩算法(例如ENCODEC)中导出的离散音频令牌。然而,这种方法的一个主要缺点是处理上下文长度。如果必须考虑下一个令牌预测的各种频率的所有音频内容,那么高保真生成架构就会崩溃。通过将连续音频表示(如频谱图)和离散声学令牌相结合,我们保留了两全其美:在单个令牌中具有特定时间实例音频所需的所有信息,但允许LLM预测未来令牌,以允许采样和离散空间提供的其他好处。我们展示了我们的架构如何提高困惑和负对数似然分数为下一个令牌预测相比,基于令牌的LLM语音和音乐。
摘要:We propose WHISPER-GPT: A generative large language model (LLM) for speechand music that allows us to work with continuous audio representations anddiscrete tokens simultaneously as part of a single architecture. There has beena huge surge in generative audio, speech, and music models that utilizediscrete audio tokens derived from neural compression algorithms, e.g. ENCODEC.However, one of the major drawbacks of this approach is handling the contextlength. It blows up for high-fidelity generative architecture if one has toaccount for all the audio contents at various frequencies for the next tokenprediction. By combining continuous audio representation like the spectrogramand discrete acoustic tokens, we retain the best of both worlds: Have all theinformation needed from the audio at a specific time instance in a singletoken, yet allow LLM to predict the future token to allow for sampling andother benefits discrete space provides. We show how our architecture improvesthe perplexity and negative log-likelihood scores for the next token predictioncompared to a token-based LLM for speech and music.
标题: Sonicmesh:利用声信号增强视觉受损环境中的3D人体网格重建
链接:https://arxiv.org/abs/2412.11325
摘要:从2D RGB图像进行3D人体网格重建(HMR)在光照不良、隐私问题或遮挡的环境中面临挑战。RGB成像的这些弱点可以通过广泛可用、易于部署并且能够穿透障碍物的声学信号来补充。然而,没有现有的方法有效地结合声信号与RGB数据的鲁棒3D HMR。主要挑战包括声学信号生成的低分辨率图像以及缺乏专用处理骨干。我们介绍SonicMesh,一种新的方法结合声学信号与RGB图像重建三维人体网格。为了解决低分辨率和缺乏专用的处理骨干在由声信号生成的图像的挑战,我们修改了现有的方法,HRNet,有效的特征提取。我们还集成了一种通用的特征嵌入技术,以提高跨维特征对齐的精度,使SonicMesh能够实现高精度。实验结果表明,SonicMesh准确地重建三维人体网格在具有挑战性的环境,如闭塞,非视线的情况下,和光线不足。
摘要:3D Human Mesh Reconstruction (HMR) from 2D RGB images faces challenges inenvironments with poor lighting, privacy concerns, or occlusions. Theseweaknesses of RGB imaging can be complemented by acoustic signals, which arewidely available, easy to deploy, and capable of penetrating obstacles.However, no existing methods effectively combine acoustic signals with RGB datafor robust 3D HMR. The primary challenges include the low-resolution imagesgenerated by acoustic signals and the lack of dedicated processing backbones.We introduce SonicMesh, a novel approach combining acoustic signals with RGBimages to reconstruct 3D human mesh. To address the challenges of lowresolution and the absence of dedicated processing backbones in imagesgenerated by acoustic signals, we modify an existing method, HRNet, foreffective feature extraction. We also integrate a universal feature embeddingtechnique to enhance the precision of cross-dimensional feature alignment,enabling SonicMesh to achieve high accuracy. Experimental results demonstratethat SonicMesh accurately reconstructs 3D human mesh in challengingenvironments such as occlusions, non-line-of-sight scenarios, and poorlighting.
标题: 流媒体语音的高效耳语
链接:https://arxiv.org/abs/2412.11272
摘要:以OpenAI的Whisper为例的语音基础模型,由于其出色的准确性和适应性,已成为语音理解的领导者。然而,它们的使用主要集中在处理预先录制的音频上,对流式语音的有效处理仍处于起步阶段。这一限制背后的几个核心挑战是:(1)这些模型是针对长的固定长度音频输入(通常为30秒)进行训练的。(2)对这样的输入进行编码涉及通过多个Transformer层处理多达1,500个令牌。(3)生成输出需要不规则的和计算量大的波束搜索。因此,在资源受限的边缘设备上进行流式语音处理比许多其他人工智能任务(包括文本生成)要求更高。为了应对这些挑战,我们引入了Whisper-T,这是一个结合了模型和系统级优化的创新框架:(1)Hush words,附加到输入中的短可学习音频片段,防止过度处理并减少模型中的幻觉。(2)波束修剪可随时间调整流式音频缓冲区,利用中间解码结果显著加快处理速度。(3)CPU/GPU流水线在编码和解码阶段之间动态分配资源,通过适应音频输入、模型特性和硬件的变化来优化性能。我们在基于ARM的平台上评估了Whisper-T,该平台具有4-12个CPU内核和10-30个GPU内核,延迟降低了1.6倍-4.7倍,每个字的延迟低至0.5秒,精度损失最小。此外,在MacBook Air上,Whisper-T每个字保持约1秒的延迟,同时仅消耗7瓦的总系统功率。
摘要:Speech foundation models, exemplified by OpenAI's Whisper, have emerged asleaders in speech understanding thanks to their exceptional accuracy andadaptability. However, their usage largely focuses on processing pre-recordedaudio, with the efficient handling of streaming speech still in its infancy.Several core challenges underlie this limitation: (1) These models are trainedfor long, fixed-length audio inputs (typically 30 seconds). (2) Encoding suchinputs involves processing up to 1,500 tokens through numerous transformerlayers. (3) Generating outputs requires an irregular and computationally heavybeam search. Consequently, streaming speech processing on edge devices withconstrained resources is more demanding than many other AI tasks, includingtext generation. To address these challenges, we introduce Whisper-T, aninnovative framework combining both model and system-level optimizations: (1)Hush words, short learnable audio segments appended to inputs, preventover-processing and reduce hallucinations in the model. (2) Beam pruning alignsstreaming audio buffers over time, leveraging intermediate decoding results tosignificantly speed up the process. (3) CPU/GPU pipelining dynamicallydistributes resources between encoding and decoding stages, optimizingperformance by adapting to variations in audio input, model characteristics,and hardware. We evaluate Whisper-T on ARM-based platforms with 4-12 CPU coresand 10-30 GPU cores, demonstrating latency reductions of 1.6x-4.7x, achievingper-word delays as low as 0.5 seconds with minimal accuracy loss. Additionally,on a MacBook Air, Whisper-T maintains approximately 1-second latency per wordwhile consuming just 7 Watts of total system power.
标题: Hanprome:修改后的韩语发音表达
链接:https://arxiv.org/abs/2412.11090
备注:21 pages
摘要:韩文是作为一种语音字母表创建的,众所周知,在现有字母表中,字母和发音之间具有最好的1:1对应关系。在本文中,我们探讨了修改的基本形式,并使用它作为一种语音符号的可能性。这种方法的核心概念是保持字母表的基本形式,只修改笔画的形状,而不是字母本身。据我们所知,以前没有任何语言试图通过改变字母笔画的形状来表达与原始字母不同的字母发音,本文可能是这个方向的第一次尝试。
摘要:Hangeul was created as a phonetic alphabet and is known to have the best 1:1correspondence between letters and pronunciation among existing alphabets. Inthis paper, we examine the possibility of modifying the basic form of Hangeuland using it as a kind of phonetic symbol. The core concept of this approach isto preserve the basic form of the alphabet, modifying only the shape of astroke rather than the letter itself. To the best of our knowledge, no previousattempts in any language have been made to express pronunciations of analphabet different from the original simply by changing the shape of thealphabet strokes, and this paper is probably the first attempt in thisdirection.
标题: 作曲家对人工智能音乐工具的评价:以人为本的设计见解
链接:https://arxiv.org/abs/2412.10968
备注:Accepted to NeurIPS 2024 Workshop on Generative AI and Creativity: A dialogue between machine learning researchers and creative professionals in Vancouver, Canada
摘要:我们提出了一项研究,探讨了以用户为中心的设计在开发生成AI(GenAI)工具进行音乐创作中的作用。通过与专业作曲家的半结构化访谈,我们收集了关于创造变化的新生成模型的见解,强调了对信任,透明度和道德设计的关注。这些发现有助于形成一个反馈回路,指导对模型的改进,强调可追溯性,透明度和可解释性。他们还揭示了新的创新领域,包括可控性的新功能以及GenAI模型的伦理和实际实施的研究问题。
摘要:We present a study that explores the role of user-centred design indeveloping Generative AI (GenAI) tools for music composition. Throughsemi-structured interviews with professional composers, we gathered insights ona novel generative model for creating variations, highlighting concerns aroundtrust, transparency, and ethical design. The findings helped form a feedbackloop, guiding improvements to the model that emphasised traceability,transparency and explainability. They also revealed new areas for innovation,including novel features for controllability and research questions on theethical and practical implementation of GenAI models.
标题: 使用深度神经网络对语音中波斯语孤立数字的鲁棒识别
链接:https://arxiv.org/abs/2412.10857
备注:15 pages, submitted to journal
摘要:近年来,人工智能(AI)在语音识别应用中取得了显着进展。与数字系统的基于语音的交互,特别是人工智能驱动的数字识别,已成为一个重要的应用。然而,现有的基于神经网络的方法往往忽视了噪声的影响,导致在嘈杂环境中的准确性降低。这项研究解决了在嘈杂的环境中识别孤立的波斯语口语数字(0到9)的挑战,特别是区分语音相似的数字。所提出的方法,这是专为非特定人识别,结合了剩余卷积神经网络和双向门控递归单元在一个混合结构的波斯数字识别。该方法使用单词单元代替音素单元作为输入。从FARSDIGIT 1数据库的51个发言人的音频数据使用各种噪声增强后,梅尔频率倒谱系数(MFCC)技术用于特征提取。实验结果表明,所提出的方法的有效性与98.53%,96.10%和95.9%的识别准确率的训练,验证和测试,分别。在噪声环境中,该方法比基于音素单元的波斯数字LSTM方法平均性能提高了26.88%。此外,该方法的准确性是7.61%,优于梅尔尺度二维根倒谱系数(MTDRCC)特征提取技术与MLP模型在测试数据的相同的数据集。
摘要:In recent years, artificial intelligence (AI) has advanced significantly inspeech recognition applications. Speech-based interaction with digital systems,particularly AI-driven digit recognition, has emerged as a prominentapplication. However, existing neural network-based methods often neglect theimpact of noise, leading to reduced accuracy in noisy environments. This studytackles the challenge of recognizing the isolated spoken Persian numbers (zeroto nine), particularly distinguishing phonetically similar numbers, in noisyenvironments. The proposed method, which is designed for speaker-independentrecognition, combines residual convolutional neural network and bidirectionalgated recurrent unit in a hybrid structure for Persian number recognition. Thismethod employs word units as input instead of phoneme units. Audio data from 51speakers of FARSDIGIT1 database are utilized after augmentation using variousnoises, and the Mel-Frequency Cepstral Coefficients (MFCC) technique isemployed for feature extraction. The experimental results show the proposedmethod efficacy with 98.53%, 96.10%, and 95.9% recognition accuracy fortraining, validation, and test, respectively. In the noisy environment, theproposed method exhibits an average performance improvement of 26.88% overphoneme unit-based LSTM method for Persian numbers. In addition, the accuracyof the proposed method is 7.61% better than that of the Mel-scale Two DimensionRoot Cepstrum Coefficients (MTDRCC) feature extraction technique along with MLPmodel in the test data for the same dataset.
标题: 使用深度一类支持量数据描述的工业机器中基于音频的异常检测
链接:https://arxiv.org/abs/2412.10792
备注:To be published in 2025 IEEE Symposium Series on Computational Intelligence
摘要:工业设备的频繁故障和故障已经促使人们越来越关注利用具有成本效益且易于部署的传感器(例如麦克风)来对机械进行有效的状态监测。麦克风为广泛使用的状态监测传感器提供了一种低成本的替代方案,其高带宽和检测其他传感器可能灵敏度较低的细微异常的能力。在这项研究中,我们调查故障的工业机器,以评估和比较不同机器类型和故障条件下的异常检测性能。机器声音的Log-Mel谱图被用作输入,并且使用两种不同方法的曲线下面积(AUC)得分来评估性能:基线密集自动编码器(AE)和具有不同子空间维度的一类深度支持向量数据描述(deep SVDD)。我们在MIMII声音数据集上的结果表明,子空间维度为2的深度SVDD方法提供了卓越的异常检测性能,与基线模型的0.82,0.72和0.64相比,6 dB,0 dB和-6 dB信噪比(SNR)的平均AUC分数分别为0.84,0.80和0.69。此外,深度SVDD所需的可训练参数比基线密集AE少7.4倍,强调了其在有效性和计算效率方面的优势。
摘要:The frequent breakdowns and malfunctions of industrial equipment have drivenincreasing interest in utilizing cost-effective and easy-to-deploy sensors,such as microphones, for effective condition monitoring of machinery.Microphones offer a low-cost alternative to widely used condition monitoringsensors with their high bandwidth and capability to detect subtle anomaliesthat other sensors might have less sensitivity. In this study, we investigatemalfunctioning industrial machines to evaluate and compare anomaly detectionperformance across different machine types and fault conditions. Log-Melspectrograms of machinery sound are used as input, and the performance isevaluated using the area under the curve (AUC) score for two different methods:baseline dense autoencoder (AE) and one-class deep Support Vector DataDescription (deep SVDD) with different subspace dimensions. Our results overthe MIMII sound dataset demonstrate that the deep SVDD method with a subspacedimension of 2 provides superior anomaly detection performance, achievingaverage AUC scores of 0.84, 0.80, and 0.69 for 6 dB, 0 dB, and -6 dBsignal-to-noise ratios (SNRs), respectively, compared to 0.82, 0.72, and 0.64for the baseline model. Moreover, deep SVDD requires 7.4 times fewer trainableparameters than the baseline dense AE, emphasizing its advantage in botheffectiveness and computational efficiency.
标题: VinTAGE:联合视频和文本调节,实现整体音频生成
链接:https://arxiv.org/abs/2412.10768
摘要:音频生成的最新进展集中在文本到音频(T2 A)和视频到音频(V2 A)任务上。然而,T2 A或V2 A方法不能生成整体声音(屏幕外和屏幕外)。这是因为T2 A无法生成与屏幕对象对齐的声音,而V2 A无法生成语义完整的声音(缺少屏幕外的声音)。在这项工作中,我们解决了整体音频生成的任务:给定视频和文本提示,我们的目标是生成与视频在时间上同步并与文本和视频在语义上对齐的屏上和屏外声音。以前的联合文本和视频到音频生成的方法往往受到模态偏见,有利于一个模态。为了克服这一限制,我们引入了VINTAGe,一个基于流的Transformer模型,它联合考虑文本和视频来指导音频生成。我们的框架包括两个关键组成部分:视觉文本编码器和联合VT-SiT模型。为了减少模态偏差并提高生成质量,我们采用预训练的单模态文本到音频和视频到音频生成模型进行额外的指导。由于缺乏适当的基准测试,我们还介绍了VINTAGe-Bench,这是一个包含636个视频-文本-音频对的数据集,其中包含屏幕外和屏幕外的声音。我们在VINTAGe-Bench上的综合实验表明,联合文本和视觉交互对于整体音频生成是必要的。此外,VinTAGe在VGGSound基准测试中取得了最先进的结果。我们的源代码和预训练模型将被发布。演示版可在以下网址获得:https://www.youtube.com/watch? v=QmqWhUjPkJI。
摘要:Recent advances in audio generation have focused on text-to-audio (T2A) andvideo-to-audio (V2A) tasks. However, T2A or V2A methods cannot generateholistic sounds (onscreen and off-screen). This is because T2A cannot generatesounds aligning with onscreen objects, while V2A cannot generate semanticallycomplete (offscreen sounds missing). In this work, we address the task ofholistic audio generation: given a video and a text prompt, we aim to generateboth onscreen and offscreen sounds that are temporally synchronized with thevideo and semantically aligned with text and video. Previous approaches forjoint text and video-to-audio generation often suffer from modality bias,favoring one modality over the other. To overcome this limitation, we introduceVinTAGe, a flow-based transformer model that jointly considers text and videoto guide audio generation. Our framework comprises two key components: aVisual-Text Encoder and a Joint VT-SiT model. To reduce modality bias andimprove generation quality, we employ pretrained uni-modal text-to-audio andvideo-to-audio generation models for additional guidance. Due to the lack ofappropriate benchmarks, we also introduce VinTAGe-Bench, a dataset of 636video-text-audio pairs containing both onscreen and offscreen sounds. Ourcomprehensive experiments on VinTAGe-Bench demonstrate that joint text andvisual interaction is necessary for holistic audio generation. Furthermore,VinTAGe achieves state-of-the-art results on the VGGSound benchmark. Our sourcecode and pre-trained models will be released. Demo is available at:https://www.youtube.com/watch?v=QmqWhUjPkJI.
标题: 日语ASB多语言模型的有效适应
链接:https://arxiv.org/abs/2412.10705
摘要:本研究探讨了微调多语言ASR(自动语音识别)模型,特别是OpenAI的Whisper-Tiny,以提高日语的性能。虽然像Whisper这样的多语言模型提供了多功能性,但它们在特定语言中往往缺乏精确性。相反,像ReazonSpeech这样的单语模型在特定语言的任务中表现出色,但适应性较差。使用日本特定的数据集和低秩自适应(LoRA)以及端到端(E2 E)训练,我们对Whisper-Tiny进行了微调,以弥合这一差距。我们的研究结果表明,微调将Whisper-Tiny的字符错误率(CER)从LoRA的32.7降低到20.8,并通过端到端微调降低到14.7,超过了Whisper-Base的CER 20.2。然而,特定领域术语的挑战仍然存在,突出表明需要专门的数据集。这些发现表明,微调多语言模型可以实现强大的语言特定性能,同时保持其灵活性。这种方法提供了一种可扩展的解决方案,用于在资源受限的环境和具有复杂书写系统的语言(如日语)中改进ASR。
摘要:This study explores fine-tuning multilingual ASR (Automatic SpeechRecognition) models, specifically OpenAI's Whisper-Tiny, to improve performancein Japanese. While multilingual models like Whisper offer versatility, theyoften lack precision in specific languages. Conversely, monolingual models likeReazonSpeech excel in language-specific tasks but are less adaptable. UsingJapanese-specific datasets and Low-Rank Adaptation (LoRA) along with end-to-end(E2E) training, we fine-tuned Whisper-Tiny to bridge this gap. Our results showthat fine-tuning reduced Whisper-Tiny's Character Error Rate (CER) from 32.7 to20.8 with LoRA and to 14.7 with end-to-end fine-tuning, surpassingWhisper-Base's CER of 20.2. However, challenges with domain-specific termsremain, highlighting the need for specialized datasets. These findingsdemonstrate that fine-tuning multilingual models can achieve stronglanguage-specific performance while retaining their flexibility. This approachprovides a scalable solution for improving ASR in resource-constrainedenvironments and languages with complex writing systems like Japanese.
标题: 隐藏的回声在音频到音频生成乐器模型中的训练中幸存下来
链接:https://arxiv.org/abs/2412.10649
备注:8 pages, 11 Figures, Proceedings of 2025 AAAI Workshop on AI for Music
摘要:随着生成技术在音频领域的普及,人们越来越有兴趣追溯这些复杂的模型,以了解它们如何利用训练数据来合成新的示例,以确保它们使用正确授权的数据,并阐明它们的黑箱行为。在本文中,我们表明,如果无法感知的回声隐藏在训练数据中,各种各样的音频到音频架构(可区分数字信号处理(DDSP),实时音频变分自动编码器(RAVE)和“舞蹈扩散”)将在其输出中再现这些回声。隐藏一个单一的回声是特别强大的所有架构,但我们也显示出有希望的结果隐藏更长的时间扩展回声模式,以增加信息容量。我们的结论表明,回声使他们的方式进入微调模型,他们生存的混合/去混合,他们生存在训练过程中的音高变化增强。因此,这个简单的,经典的水印思想显示了标记生成音频模型的重要承诺。
摘要:As generative techniques pervade the audio domain, there has been increasinginterest in tracing back through these complicated models to understand howthey draw on their training data to synthesize new examples, both to ensurethat they use properly licensed data and also to elucidate their black boxbehavior. In this paper, we show that if imperceptible echoes are hidden in thetraining data, a wide variety of audio to audio architectures (differentiabledigital signal processing (DDSP), Realtime Audio Variational autoEncoder(RAVE), and ``Dance Diffusion'') will reproduce these echoes in their outputs.Hiding a single echo is particularly robust across all architectures, but wealso show promising results hiding longer time spread echo patterns for anincreased information capacity. We conclude by showing that echoes make theirway into fine tuned models, that they survive mixing/demixing, and that theysurvive pitch shift augmentation during training. Hence, this simple, classicalidea in watermarking shows significant promise for tagging generative audiomodels.
标题: 临界点、脉搏弹性和调性张力:关于临界点产生的实证研究
链接:https://arxiv.org/abs/2412.10481
备注:International Society for Music Information Retrieval Conference, Oct 2017, Suzhou, China, China
摘要:引爆点是指音乐作品中的关键转折点。这项研究提出了第一步,定量和系统地描述音乐性质的临界点。定时信息和计算得出的音调张力值,对应于不和谐音,距离键,和谐波运动相比,在Ashkenazy的六个肖邦马祖卡的录音,由35名听众确定的临界点。分析表明,所有流行的临界点,但一个可以解释为统计上显着的时间偏差或变化点,至少在三个张力参数之一。
摘要:Tipping points are moments of change that characterise crucial turning pointsin a piece of music. This study presents a first step towards quantitativelyand systematically describing the musical properties of tipping points. Timinginformation and computationally-derived tonal tension values which correspondto dissonance, distance from key, and harmonic motion are compared to tippingpoints in Ashkenazy's recordings of six Chopin Mazurkas, as identified by 35listeners. The analysis shows that all popular tipping points but one could beexplained by statistically significant timing deviations or changepoints in atleast one of the three tension parameters.
标题: Mel频率倒谱系数的比较分析和基于子波的音频信号处理用于口语中的情感检测和心理健康评估
链接:https://arxiv.org/abs/2412.10469
摘要:技术和心理健康的交叉刺激了评估情绪健康的创新方法,特别是通过应用于音频数据分析的计算技术。本研究探讨了卷积神经网络(CNN)和长短期记忆(LSTM)模型在小波提取特征和梅尔频率倒谱系数(MFCC)上的应用,以从口语中进行情感检测。数据增强技术,特征提取,归一化和模型训练进行评估模型的分类情感状态的性能。结果表明,CNN模型的准确率为61%,而LSTM模型的准确率为56%。这两种模型在预测特定情绪(如惊讶和愤怒)方面表现出更好的性能,利用了不同的音频特征,如音高和速度变化。建议包括进一步探索先进的数据增强技术,结合特征提取方法,以及将语言分析与语音特征相结合,以提高心理健康诊断的准确性。标准化数据集收集和共享的合作,建议促进情感计算和心理健康护理干预措施的进步。
摘要:The intersection of technology and mental health has spurred innovativeapproaches to assessing emotional well-being, particularly throughcomputational techniques applied to audio data analysis. This study exploresthe application of Convolutional Neural Network (CNN) and Long Short-TermMemory (LSTM) models on wavelet extracted features and Mel-frequency CepstralCoefficients (MFCCs) for emotion detection from spoken speech. Dataaugmentation techniques, feature extraction, normalization, and model trainingwere conducted to evaluate the models' performance in classifying emotionalstates. Results indicate that the CNN model achieved a higher accuracy of 61%compared to the LSTM model's accuracy of 56%. Both models demonstrated betterperformance in predicting specific emotions such as surprise and anger,leveraging distinct audio features like pitch and speed variations.Recommendations include further exploration of advanced data augmentationtechniques, combined feature extraction methods, and the integration oflinguistic analysis with speech characteristics for improved accuracy in mentalhealth diagnostics. Collaboration for standardized dataset collection andsharing is recommended to foster advancements in affective computing and mentalhealth care interventions.
标题: 通过视听内容的文本情感描述丰富多模式情感分析
链接:https://arxiv.org/abs/2412.10460
备注:None
摘要:多模态情感分析(MSA)是一个重要的研究前沿,旨在通过融合文本,音频和视觉数据来全面揭示人类的情感。然而,在音频和视频表达中辨别微妙的情感细微差别构成了一个巨大的挑战,特别是当各个部分的情感极性看起来相似时。在本文中,我们的目标是突出情感相关的属性的音频和视觉模态,以促进多模态融合的背景下,在视觉-音频场景中的细微差别的情感变化。为此,我们引入DEVA,一个渐进的融合框架,建立在文本情感描述的基础上,旨在强调视听内容的情感特征。DEVA采用情感描述生成器(EDG)将原始音频和视频数据转化为文本化的情感描述,从而放大它们的情感特征。然后将这些描述与源数据集成,以产生更丰富的增强功能。此外,DEVA结合了文本引导渐进融合模块(TPF),利用不同层次的文本作为核心模态指南。该模块逐步融合视听次要模态,以减轻文本和视听模态之间的差异。在广泛使用的情感分析基准数据集(包括MOSI,MOSEI和CH-SIMS)上的实验结果强调了与最先进的模型相比的显着增强。此外,细粒度的情绪实验证实了强大的敏感性DEVA微妙的情绪变化。
摘要:Multimodal Sentiment Analysis (MSA) stands as a critical research frontier,seeking to comprehensively unravel human emotions by amalgamating text, audio,and visual data. Yet, discerning subtle emotional nuances within audio andvideo expressions poses a formidable challenge, particularly when emotionalpolarities across various segments appear similar. In this paper, our objectiveis to spotlight emotion-relevant attributes of audio and visual modalities tofacilitate multimodal fusion in the context of nuanced emotional shifts invisual-audio scenarios. To this end, we introduce DEVA, a progressive fusionframework founded on textual sentiment descriptions aimed at accentuatingemotional features of visual-audio content. DEVA employs an EmotionalDescription Generator (EDG) to transmute raw audio and visual data intotextualized sentiment descriptions, thereby amplifying their emotionalcharacteristics. These descriptions are then integrated with the source data toyield richer, enhanced features. Furthermore, DEVA incorporates the Text-guidedProgressive Fusion Module (TPF), leveraging varying levels of text as a coremodality guide. This module progressively fuses visual-audio minor modalitiesto alleviate disparities between text and visual-audio modalities. Experimentalresults on widely used sentiment analysis benchmark datasets, including MOSI,MOSEI, and CH-SIMS, underscore significant enhancements compared tostate-of-the-art models. Moreover, fine-grained emotion experiments corroboratethe robust sensitivity of DEVA to subtle emotional variations.
标题: 在心理健康中利用音频和文本模式:LLM绩效研究
链接:https://arxiv.org/abs/2412.10417
摘要:心理健康障碍在世界范围内日益普遍,迫切需要创新工具来支持早期诊断和干预。本研究探讨了大型语言模型(LLM)在多模态心理健康诊断中的潜力,特别是通过文本和音频模式检测抑郁症和创伤后应激障碍。使用E-DAIC数据集,我们比较了文本和音频模态,以研究LLM是否可以在音频输入下表现得同样好或更好。我们进一步研究了两种模式的整合,以确定这是否可以提高诊断准确性,这通常会导致性能指标的改善。我们的分析专门利用定制的指标;模态优先度评分和不一致解决评分来评估组合模态如何影响模型性能。Gemini 1.5 Pro模型在使用组合模式时在二元抑郁分类中获得了最高分数,在整个数据集中评估的F1分数为0.67,平衡准确度(BA)为77.4%。这些结果比文本模态的性能提高了3.1%,比音频模态提高了2.7%,突出了整合模态以提高诊断准确性的有效性。值得注意的是,所有结果都是在zero-shot推断中获得的,突出了模型的鲁棒性,而不需要特定于任务的微调。为了探索不同配置对模型性能的影响,我们使用zero-shot和Few-Shot提示进行二进制,严重性和多类任务,检查提示变化对性能的影响。结果显示,Gemini 1.5 Pro在文本和音频模式中以及GPT-4 o mini在文本模式中的模型在多个任务中的平衡准确性和F1得分方面往往超过其他模型。
摘要:Mental health disorders are increasingly prevalent worldwide, creating anurgent need for innovative tools to support early diagnosis and intervention.This study explores the potential of Large Language Models (LLMs) in multimodalmental health diagnostics, specifically for detecting depression and PostTraumatic Stress Disorder through text and audio modalities. Using the E-DAICdataset, we compare text and audio modalities to investigate whether LLMs canperform equally well or better with audio inputs. We further examine theintegration of both modalities to determine if this can enhance diagnosticaccuracy, which generally results in improved performance metrics. Our analysisspecifically utilizes custom-formulated metrics; Modal Superiority Score andDisagreement Resolvement Score to evaluate how combined modalities influencemodel performance. The Gemini 1.5 Pro model achieves the highest scores inbinary depression classification when using the combined modality, with an F1score of 0.67 and a Balanced Accuracy (BA) of 77.4%, assessed across the fulldataset. These results represent an increase of 3.1% over its performance withthe text modality and 2.7% over the audio modality, highlighting theeffectiveness of integrating modalities to enhance diagnostic accuracy.Notably, all results are obtained in zero-shot inferring, highlighting therobustness of the models without requiring task-specific fine-tuning. Toexplore the impact of different configurations on model performance, we conductbinary, severity, and multiclass tasks using both zero-shot and few-shotprompts, examining the effects of prompt variations on performance. The resultsreveal that models such as Gemini 1.5 Pro in text and audio modalities, andGPT-4o mini in the text modality, often surpass other models in balancedaccuracy and F1 scores across multiple tasks.
标题: SpeechPrune:用于语音信息检索的上下文感知令牌修剪
链接:https://arxiv.org/abs/2412.12009
备注:Project page and dataset is available at this https URL
摘要:我们介绍了语音信息检索(SIR),语音大语言模型(语音LLM)的一个新的长上下文任务,并提出了SPIRAL,一个1,012样本的基准测试模型的能力,从大约90秒的口语输入提取关键细节。虽然目前的语音LLM擅长于短形式的任务,但它们难以满足较长音频序列的计算和代表性需求。为了解决这一限制,我们提出了SpeechPrune,这是一种无需训练的令牌修剪策略,它使用语音-文本相似性和近似注意力分数来有效地丢弃不相关的令牌。在SPIRAL中,SpeechPrune在修剪率为20%的情况下,分别比原始模型和随机修剪模型实现了29%和高达47%的准确性提高。SpeechPrune甚至可以在80%的修剪水平下保持网络性能。这种方法突出了令牌级修剪的潜力,有效和可扩展的长形式语音理解。
摘要:We introduce Speech Information Retrieval (SIR), a new long-context task forSpeech Large Language Models (Speech LLMs), and present SPIRAL, a 1,012-samplebenchmark testing models' ability to extract critical details fromapproximately 90-second spoken inputs. While current Speech LLMs excel atshort-form tasks, they struggle with the computational and representationaldemands of longer audio sequences. To address this limitation, we proposeSpeechPrune, a training-free token pruning strategy that uses speech-textsimilarity and approximated attention scores to efficiently discard irrelevanttokens. In SPIRAL, SpeechPrune achieves accuracy improvements of 29% and up to47% over the original model and the random pruning model at a pruning rate of20%, respectively. SpeechPrune can maintain network performance even at apruning level of 80%. This approach highlights the potential of token-levelpruning for efficient and scalable long-form speech understanding.
标题: 一种轻量级且稳健的语音盲宽带到全带扩展方法
链接:https://arxiv.org/abs/2412.11392
备注:prelimnary version, content and author list might still change
摘要:在资源受限的环境中,如低带宽语音传输或低复杂度声码,减少语音带宽是常见的做法。我们提出了一个轻量级的和强大的方法来扩展宽带语音信号的带宽,这是在语音编码背景下开发的经典方法的启发。最终的模型只有$\sim $~K个参数,复杂度为~140 MFLOPS(或~70 MMACS)。该模型的帧大小为10 ms,预测时间仅为0.27 ms,非常适合常见的宽带语音编解码器。我们通过将该模型与Opus SILK语音编解码器(1.5版本)配对来评估该模型的鲁棒性,并在P.808 DCR收听测试中验证,该模型将质量从6 kb/s显著提高到12 kb/s。我们还证明,Opus 1.5连同所提出的带宽扩展在9 kb/s满足3GPP EVS在9.6 kb/s的质量和Opus 1.4在18 kb/s的质量,这表明盲带宽扩展可以满足经典的引导带宽扩展的质量。
摘要:Reducing the bandwidth of speech is common practice in resource constrainedenvironments like low-bandwidth speech transmission or low-complexity vocoding.We propose a lightweight and robust method for extending the bandwidth ofwideband speech signals that is inspired by classical methods developed in thespeech coding context. The resulting model has just $\sim 370$~K parameters anda complexity of ~140 MFLOPS (or ~70 MMACS). With a frame size of 10 ms and alookahead of just 0.27 ms the model is well-suited for common wideband speechcodecs. We evaluate the model's robustness by pairing it with the Opus SILKspeech codec (1.5 release) and verify in a P.808 DCR listening test that itsignificantly improves quality from 6 to 12 kb/s. We also demonstrate that Opus1.5 together with the proposed bandwidth extension at 9 kb/s meets the qualityof 3GPP EVS at 9.6 kb/s and that of Opus 1.4 at 18 kb/s showing that the blindbandwidth extension can meet the quality of classical guided bandwidthextensions.
标题: 自动语音识别的音译Zero-Shot域自适应
链接:https://arxiv.org/abs/2412.11185
摘要:自动语音识别模型的性能往往在训练数据未覆盖的域上退化。领域适配可以解决这个问题,假设目标语言的目标领域数据的可用性。然而,这种假设在许多现实世界的应用中并不成立。为了使域适应更适用,我们解决的问题,zero-shot域自适应(MANDA),目标领域的数据是不可用的目标语言。相反,我们从另一种源语言中转移目标领域知识,其中目标领域数据更容易访问。为此,我们首先执行跨语言预训练(XLPT)以跨语言共享领域知识,然后使用目标语言微调来构建最终模型。这种做法的一个挑战是,预先训练的知识可能会在微调过程中被遗忘,从而导致次优的适应性能。为了解决这个问题,我们提出了音译的CANDDA来实现一致的预训练和微调标签,从而最大限度地保留预训练的知识。实验结果表明,与wav2vec 2.0基线相比,音译后的WADA相对降低了9.2%的单词错误率。此外,音译的ARMDA始终优于自监督ARMDA,并与监督ARMDA表现相当,证明了基于音译的预训练标签的优越性。
摘要:The performance of automatic speech recognition models often degenerates ondomains not covered by the training data. Domain adaptation can address thisissue, assuming the availability of the target domain data in the targetlanguage. However, such assumption does not stand in many real-worldapplications. To make domain adaptation more applicable, we address the problemof zero-shot domain adaptation (ZSDA), where target domain data is unavailablein the target language. Instead, we transfer the target domain knowledge fromanother source language where the target domain data is more accessible. To dothat, we first perform cross-lingual pre-training (XLPT) to share domainknowledge across languages, then use target language fine-tuning to build thefinal model. One challenge in this practice is that the pre-trained knowledgecan be forgotten during fine-tuning, resulting in sub-optimal adaptationperformance. To address this issue, we propose transliterated ZSDA to achieveconsistent pre-training and fine-tuning labels, leading to maximum preservationof the pre-trained knowledge. Experimental results show that transliteratedZSDA relatively decreases the word error rate by 9.2% compared with a wav2vec2.0 baseline. Moreover, transliterated ZSDA consistently outperformsself-supervised ZSDA and performs on par with supervised ZSDA, proving thesuperiority of transliteration-based pre-training labels.
标题: MASV:具有全球和本地背景的演讲者验证Mamba
链接:https://arxiv.org/abs/2412.10989
摘要:像卷积神经网络和Transformers这样的深度学习模型在语音验证方面表现出了令人印象深刻的能力,在研究界引起了相当大的关注。然而,基于CNN的方法难以有效地对长序列音频进行建模,从而导致次优的验证性能。另一方面,基于变换器的方法经常受到高计算需求的阻碍,限制了它们的实用性。本文提出了MASV模型,一种新的架构,将Mamba模块集成到ECAPA-TDNN框架。通过引入本地上下文双向Mamba和Tri-Mamba块,该模型有效地捕获音频序列中的全局和本地上下文。实验结果表明,MASV模型大大提高了验证性能,超过现有的模型在准确性和效率。
摘要:Deep learning models like Convolutional Neural Networks and transformers haveshown impressive capabilities in speech verification, gaining considerableattention in the research community. However, CNN-based approaches strugglewith modeling long-sequence audio effectively, resulting in suboptimalverification performance. On the other hand, transformer-based methods areoften hindered by high computational demands, limiting their practicality. Thispaper presents the MASV model, a novel architecture that integrates the Mambamodule into the ECAPA-TDNN framework. By introducing the Local ContextBidirectional Mamba and Tri-Mamba block, the model effectively captures bothglobal and local context within audio sequences. Experimental resultsdemonstrate that the MASV model substantially enhances verificationperformance, surpassing existing models in both accuracy and efficiency.
标题: SpeechPrune:用于语音信息检索的上下文感知令牌修剪
链接:https://arxiv.org/abs/2412.12009
备注:Project page and dataset is available at this https URL
摘要:我们介绍了语音信息检索(SIR),语音大语言模型(语音LLM)的一个新的长上下文任务,并提出了SPIRAL,一个1,012样本的基准测试模型的能力,从大约90秒的口语输入提取关键细节。虽然目前的语音LLM擅长于短形式的任务,但它们难以满足较长音频序列的计算和代表性需求。为了解决这一限制,我们提出了SpeechPrune,这是一种无需训练的令牌修剪策略,它使用语音-文本相似性和近似注意力分数来有效地丢弃不相关的令牌。在SPIRAL中,SpeechPrune在修剪率为20%的情况下,分别比原始模型和随机修剪模型实现了29%和高达47%的准确性提高。SpeechPrune甚至可以在80%的修剪水平下保持网络性能。这种方法突出了令牌级修剪的潜力,有效和可扩展的长形式语音理解。
摘要:We introduce Speech Information Retrieval (SIR), a new long-context task forSpeech Large Language Models (Speech LLMs), and present SPIRAL, a 1,012-samplebenchmark testing models' ability to extract critical details fromapproximately 90-second spoken inputs. While current Speech LLMs excel atshort-form tasks, they struggle with the computational and representationaldemands of longer audio sequences. To address this limitation, we proposeSpeechPrune, a training-free token pruning strategy that uses speech-textsimilarity and approximated attention scores to efficiently discard irrelevanttokens. In SPIRAL, SpeechPrune achieves accuracy improvements of 29% and up to47% over the original model and the random pruning model at a pruning rate of20%, respectively. SpeechPrune can maintain network performance even at apruning level of 80%. This approach highlights the potential of token-levelpruning for efficient and scalable long-form speech understanding.
标题: 一种轻量级且稳健的语音盲宽带到全带扩展方法
链接:https://arxiv.org/abs/2412.11392
备注:prelimnary version, content and author list might still change
摘要:在资源受限的环境中,如低带宽语音传输或低复杂度声码,减少语音带宽是常见的做法。我们提出了一个轻量级的和强大的方法来扩展宽带语音信号的带宽,这是在语音编码背景下开发的经典方法的启发。最终的模型只有$\sim $~K个参数,复杂度为~140 MFLOPS(或~70 MMACS)。该模型的帧大小为10 ms,预测时间仅为0.27 ms,非常适合常见的宽带语音编解码器。我们通过将该模型与Opus SILK语音编解码器(1.5版本)配对来评估该模型的鲁棒性,并在P.808 DCR收听测试中验证,该模型将质量从6 kb/s显著提高到12 kb/s。我们还证明,Opus 1.5连同所提出的带宽扩展在9 kb/s满足3GPP EVS在9.6 kb/s的质量和Opus 1.4在18 kb/s的质量,这表明盲带宽扩展可以满足经典的引导带宽扩展的质量。
摘要:Reducing the bandwidth of speech is common practice in resource constrainedenvironments like low-bandwidth speech transmission or low-complexity vocoding.We propose a lightweight and robust method for extending the bandwidth ofwideband speech signals that is inspired by classical methods developed in thespeech coding context. The resulting model has just $\sim 370$~K parameters anda complexity of ~140 MFLOPS (or ~70 MMACS). With a frame size of 10 ms and alookahead of just 0.27 ms the model is well-suited for common wideband speechcodecs. We evaluate the model's robustness by pairing it with the Opus SILKspeech codec (1.5 release) and verify in a P.808 DCR listening test that itsignificantly improves quality from 6 to 12 kb/s. We also demonstrate that Opus1.5 together with the proposed bandwidth extension at 9 kb/s meets the qualityof 3GPP EVS at 9.6 kb/s and that of Opus 1.4 at 18 kb/s showing that the blindbandwidth extension can meet the quality of classical guided bandwidthextensions.
标题: 自动语音识别的音译Zero-Shot域自适应
链接:https://arxiv.org/abs/2412.11185
摘要:自动语音识别模型的性能往往在训练数据未覆盖的域上退化。领域适配可以解决这个问题,假设目标语言的目标领域数据的可用性。然而,这种假设在许多现实世界的应用中并不成立。为了使域适应更适用,我们解决的问题,zero-shot域自适应(MANDA),目标领域的数据是不可用的目标语言。相反,我们从另一种源语言中转移目标领域知识,其中目标领域数据更容易访问。为此,我们首先执行跨语言预训练(XLPT)以跨语言共享领域知识,然后使用目标语言微调来构建最终模型。这种做法的一个挑战是,预先训练的知识可能会在微调过程中被遗忘,从而导致次优的适应性能。为了解决这个问题,我们提出了音译的CANDDA来实现一致的预训练和微调标签,从而最大限度地保留预训练的知识。实验结果表明,与wav2vec 2.0基线相比,音译后的WADA相对降低了9.2%的单词错误率。此外,音译的ARMDA始终优于自监督ARMDA,并与监督ARMDA表现相当,证明了基于音译的预训练标签的优越性。
摘要:The performance of automatic speech recognition models often degenerates ondomains not covered by the training data. Domain adaptation can address thisissue, assuming the availability of the target domain data in the targetlanguage. However, such assumption does not stand in many real-worldapplications. To make domain adaptation more applicable, we address the problemof zero-shot domain adaptation (ZSDA), where target domain data is unavailablein the target language. Instead, we transfer the target domain knowledge fromanother source language where the target domain data is more accessible. To dothat, we first perform cross-lingual pre-training (XLPT) to share domainknowledge across languages, then use target language fine-tuning to build thefinal model. One challenge in this practice is that the pre-trained knowledgecan be forgotten during fine-tuning, resulting in sub-optimal adaptationperformance. To address this issue, we propose transliterated ZSDA to achieveconsistent pre-training and fine-tuning labels, leading to maximum preservationof the pre-trained knowledge. Experimental results show that transliteratedZSDA relatively decreases the word error rate by 9.2% compared with a wav2vec2.0 baseline. Moreover, transliterated ZSDA consistently outperformsself-supervised ZSDA and performs on par with supervised ZSDA, proving thesuperiority of transliteration-based pre-training labels.
标题: MASV:具有全球和本地背景的演讲者验证Mamba
链接:https://arxiv.org/abs/2412.10989
摘要:像卷积神经网络和Transformers这样的深度学习模型在语音验证方面表现出了令人印象深刻的能力,在研究界引起了相当大的关注。然而,基于CNN的方法难以有效地对长序列音频进行建模,从而导致次优的验证性能。另一方面,基于变换器的方法经常受到高计算需求的阻碍,限制了它们的实用性。本文提出了MASV模型,一种新的架构,将Mamba模块集成到ECAPA-TDNN框架。通过引入局部上下文双向Mamba和Tri-Mamba块,该模型有效地捕获音频序列中的全局和局部上下文。实验结果表明,MASV模型大大提高了验证性能,超过现有的模型在准确性和效率。
摘要:Deep learning models like Convolutional Neural Networks and transformers haveshown impressive capabilities in speech verification, gaining considerableattention in the research community. However, CNN-based approaches strugglewith modeling long-sequence audio effectively, resulting in suboptimalverification performance. On the other hand, transformer-based methods areoften hindered by high computational demands, limiting their practicality. Thispaper presents the MASV model, a novel architecture that integrates the Mambamodule into the ECAPA-TDNN framework. By introducing the Local ContextBidirectional Mamba and Tri-Mamba block, the model effectively captures bothglobal and local context within audio sequences. Experimental resultsdemonstrate that the MASV model substantially enhances verificationperformance, surpassing existing models in both accuracy and efficiency.
标题: 语音基础模型和众包实现高效、高质量的数据收集
链接:https://arxiv.org/abs/2412.11978
备注:Accepted at COLING 2025 main conference
摘要:虽然众包是促进和扩大语音数据收集的既定解决方案,但非专家的参与需要协议来确保最终数据质量。为了降低这些基本控制的成本,本文研究了使用语音基础模型(SFM)自动化的验证过程中,第一次检查数据采集的成本/质量权衡。在法国、德国和韩国数据上进行的实验表明,基于SFM的验证有可能减少对人工验证的依赖,从而在不降低最终数据质量的情况下节省超过40.0%的成本。这些发现为更高效、更具成本效益和可扩展的语音数据采集提供了新的机会。
摘要:While crowdsourcing is an established solution for facilitating and scalingthe collection of speech data, the involvement of non-experts necessitatesprotocols to ensure final data quality. To reduce the costs of these essentialcontrols, this paper investigates the use of Speech Foundation Models (SFMs) toautomate the validation process, examining for the first time the cost/qualitytrade-off in data acquisition. Experiments conducted on French, German, andKorean data demonstrate that SFM-based validation has the potential to reducereliance on human validation, resulting in an estimated cost saving of over40.0% without degrading final data quality. These findings open newopportunities for more efficient, cost-effective, and scalable speech dataacquisition.
标题: auttrainer:一个用于计算机试镜任务的模块化和可扩展深度学习工具包
链接:https://arxiv.org/abs/2412.11943
摘要:这项工作介绍了autrainer的关键操作原理,autrainer是我们用于计算机试听任务的新深度学习训练框架。autrainer是一个基于PyTorch的工具包,允许对各种不同的计算机试听任务进行快速,可重复和易于扩展的培训。具体来说,autrainer提供低代码训练,并支持广泛的神经网络以及预处理例程。在这项工作中,我们提出了一个概述其内部工作原理和关键功能。
摘要:This work introduces the key operating principles for autrainer, our new deeplearning training framework for computer audition tasks. autrainer is aPyTorch-based toolkit that allows for rapid, reproducible, and easilyextensible training on a variety of different computer audition tasks.Concretely, autrainer offers low-code training and supports a wide range ofneural networks as well as preprocessing routines. In this work, we present anoverview of its inner workings and key capabilities.
标题: AudioCIL:用于音频类-多场景增量学习的Python工作空间
链接:https://arxiv.org/abs/2412.11907
摘要:深度学习凭借其强大的音频特征提取能力,在音频信号处理方面取得了重大成功。通常,这些方法依赖于静态的、预先收集的大规模数据集进行训练,在固定数量的类上表现良好。然而,现实世界的特点是不断变化,新的音频类从流媒体中出现,或者由于隐私而临时可用。音频环境的这种动态性质需要模型能够在不丢弃现有信息的情况下为新类增量地学习新知识。将增量学习引入音频信号处理领域,即,音频类增量学习(AuCIL),是一个有意义的努力。我们提出了一个名为AudioCIL的工具箱,以使音频信号处理算法与真实世界的场景保持一致,并加强音频类增量学习的研究。
摘要:Deep learning, with its robust aotomatic feature extraction capabilities, hasdemonstrated significant success in audio signal processing. Typically, thesemethods rely on static, pre-collected large-scale datasets for training,performing well on a fixed number of classes. However, the real world ischaracterized by constant change, with new audio classes emerging fromstreaming or temporary availability due to privacy. This dynamic nature ofaudio environments necessitates models that can incrementally learn newknowledge for new classes without discarding existing information. Introducingincremental learning to the field of audio signal processing, i.e., AudioClass-Incremental Learning (AuCIL), is a meaningful endeavor. We propose such atoolbox named AudioCIL to align audio signal processing algorithms withreal-world scenarios and strengthen research in audio class-incrementallearning.
标题: 多语言音频的自发语音和脚本语音的分类
链接:https://arxiv.org/abs/2412.11896
备注:Accepted to IEEE Spoken Language Technology Workshop 2024
摘要:区分讲稿和自发言语是更好地理解言语风格如何影响言语处理研究的重要工具。它还可以通过更好地分割大型录制语音目录来改善媒体用户的推荐系统和发现体验。本文解决了建立一个分类器,以及在不同的格式和语言的概括的挑战。我们系统地评估了从传统的、手工制作的声学和韵律特征到先进的音频Transformers的模型,利用大型的、多语言的专有播客数据集进行训练和验证。我们将每个模型在11个语言组中的表现进行了分解,以评估跨语言偏见。我们的实验分析扩展到公开可用的数据集,以评估模型对非播客领域的普适性。我们的研究结果表明,基于transformer的模型始终优于传统的基于特征的技术,在区分各种语言的脚本和自发语音方面实现了最先进的性能。
摘要:Distinguishing scripted from spontaneous speech is an essential tool forbetter understanding how speech styles influence speech processing research. Itcan also improve recommendation systems and discovery experiences for mediausers through better segmentation of large recorded speech catalogues. Thispaper addresses the challenge of building a classifier that generalises wellacross different formats and languages. We systematically evaluate modelsranging from traditional, handcrafted acoustic and prosodic features toadvanced audio transformers, utilising a large, multilingual proprietarypodcast dataset for training and validation. We break down the performance ofeach model across 11 language groups to evaluate cross-lingual biases. Ourexperimental analysis extends to publicly available datasets to assess themodels' generalisability to non-podcast domains. Our results indicate thattransformer-based models consistently outperform traditional feature-basedtechniques, achieving state-of-the-art performance in distinguishing betweenscripted and spontaneous speech across various languages.
标题: ProsodyFM:用于可理解语音合成的无监督短语和语调控制
链接:https://arxiv.org/abs/2412.11795
备注:Accepted by AAAI 2025
摘要:韵律包含了丰富的信息,超出了字面意思的话,这是至关重要的语音的可理解性。目前的模型在措辞和语调方面仍然存在不足,它们在合成具有复杂结构的长句时不仅会丢失或放错停顿,而且会产生不自然的语调。我们提出ProsodyFM,韵律感知的文本到语音合成(TTS)模型与流匹配(FM)骨干,旨在提高韵律的措辞和语调方面。ProsodyFM引入了两个关键组件:一个短语中断编码器,用于捕获初始短语中断位置,然后是一个持续时间预测器,用于灵活调整中断持续时间;以及一个终端语调编码器,它集成了一组语调形状标记,并结合了一个新颖的音高处理器,用于对人类感知的语调变化进行更鲁棒的建模。ProsodyFM是在没有明确韵律标签的情况下训练的,但可以发现广泛的中断持续时间和语调模式。实验结果表明,ProsodyFM可以有效地改善韵律的措辞和语调方面,从而提高整体的可懂度相比,四个国家的最先进的(SOTA)模型。分布外实验表明,这种韵律改进可以进一步带来ProsodyFM优越的泛化能力,看不见的复杂句子和扬声器。我们的案例研究直观地说明了强大的和细粒度的可控性的ProsodyFM的措辞和语调。
摘要:Prosody contains rich information beyond the literal meaning of words, whichis crucial for the intelligibility of speech. Current models still fall shortin phrasing and intonation; they not only miss or misplace breaks whensynthesizing long sentences with complex structures but also produce unnaturalintonation. We propose ProsodyFM, a prosody-aware text-to-speech synthesis(TTS) model with a flow-matching (FM) backbone that aims to enhance thephrasing and intonation aspects of prosody. ProsodyFM introduces two keycomponents: a Phrase Break Encoder to capture initial phrase break locations,followed by a Duration Predictor for the flexible adjustment of breakdurations; and a Terminal Intonation Encoder which integrates a set ofintonation shape tokens combined with a novel Pitch Processor for more robustmodeling of human-perceived intonation change. ProsodyFM is trained with noexplicit prosodic labels and yet can uncover a broad spectrum of breakdurations and intonation patterns. Experimental results demonstrate thatProsodyFM can effectively improve the phrasing and intonation aspects ofprosody, thereby enhancing the overall intelligibility compared to fourstate-of-the-art (SOTA) models. Out-of-distribution experiments show that thisprosody improvement can further bring ProsodyFM superior generalizability forunseen complex sentences and speakers. Our case study intuitively illustratesthe powerful and fine-grained controllability of ProsodyFM over phrasing andintonation.
标题: 它是Chug吗?以数据驱动的方式理解吉他音调描述
链接:https://arxiv.org/abs/2412.11769
备注:Accepted for publication at the 3rd Workshop on NLP for Music and Audio (NLP4MusA 2024)
摘要:自然语言通常用于描述乐器音色,例如“温暖”或“沉重”的声音。由于这些描述符基于人类感知,因此对于哪些声学特征对应于给定的形容词可能存在分歧。在这项工作中,我们追求一种数据驱动的方法,以进一步了解这些形容词的背景下,吉他音。我们的主要贡献是一个音色形容词的数据集,通过处理乐器音频的单个片段来构建,通过EQ和失真等效果的调整来产生不同的音色。众包专家为每个片段获得形容词注释,以完成成对比较和标记任务。我们检查了数据集,揭示了形容词评级之间的相关性,并强调了数据与有关光谱特征和音色形容词的流行理论相矛盾的实例,这表明需要对音色进行更细致入微、数据驱动的理解。
摘要:Natural language is commonly used to describe instrument timbre, such as a"warm" or "heavy" sound. As these descriptors are based on human perception,there can be disagreement over which acoustic features correspond to a givenadjective. In this work, we pursue a data-driven approach to further ourunderstanding of such adjectives in the context of guitar tone. Our maincontribution is a dataset of timbre adjectives, constructed by processingsingle clips of instrument audio to produce varied timbres through adjustmentsin EQ and effects such as distortion. Adjective annotations are obtained foreach clip by crowdsourcing experts to complete a pairwise comparison and alabeling task. We examine the dataset and reveal correlations between adjectiveratings and highlight instances where the data contradicts prevailing theorieson spectral features and timbral adjectives, suggesting a need for a morenuanced, data-driven understanding of timbre.
标题: 用于增强视听Zero-Shot学习的差异感知注意力网络
链接:https://arxiv.org/abs/2412.11715
摘要:视听零镜头学习(Audio-visual Zero-Shot Learning,简称ZRL)由于能够识别看不见的类别并在视频分类任务中表现良好而受到了广泛关注。然而,模态不平衡的(G)CNOL导致过度依赖的最佳模态,降低了不可见的类的区分能力。一些研究试图通过修改参数梯度来解决这一问题,但仍然存在两个挑战:(a)质量差异,即模式为同一概念提供不同数量和质量的信息。(b)内容差异,其中模态内的样本贡献差异很大。为了应对这些挑战,我们提出了一个差异感知注意力网络(DAAN)增强视听视听网络。我们的方法引入了一个质量离散缓解注意力(QDMA)单元,以最大限度地减少冗余信息的高质量模态和对比样本级梯度调制(CSGM)块,以调整梯度幅度和平衡内容差异。我们量化模态的贡献,整合优化和收敛速度更精确的梯度调制CSGM。实验表明,DAAN在基准数据集上实现了最先进的性能,消融研究验证了各个模块的有效性。
摘要:Audio-visual Zero-Shot Learning (ZSL) has attracted significant attention forits ability to identify unseen classes and perform well in video classificationtasks. However, modal imbalance in (G)ZSL leads to over-reliance on the optimalmodality, reducing discriminative capabilities for unseen classes. Some studieshave attempted to address this issue by modifying parameter gradients, but twochallenges still remain: (a) Quality discrepancies, where modalities offerdiffering quantities and qualities of information for the same concept. (b)Content discrepancies, where sample contributions within a modality varysignificantly. To address these challenges, we propose a Discrepancy-AwareAttention Network (DAAN) for Enhanced Audio-Visual ZSL. Our approach introducesa Quality-Discrepancy Mitigation Attention (QDMA) unit to minimize redundantinformation in the high-quality modality and a Contrastive Sample-levelGradient Modulation (CSGM) block to adjust gradient magnitudes and balancecontent discrepancies. We quantify modality contributions by integratingoptimization and convergence rate for more precise gradient modulation in CSGM.Experiments demonstrates DAAN achieves state-of-the-art performance onbenchmark datasets, with ablation studies validating the effectiveness ofindividual modules.
标题: 音频深度伪造检测的连续学习中基于区域的优化
链接:https://arxiv.org/abs/2412.11551
备注:Accepted by AAAI 2025
摘要:语音合成和语音转换的快速发展带来了便利,但也带来了新的安全风险,迫切需要有效的音频深度伪造检测。尽管目前的模型表现良好,但当面对现实世界中deepfake的多样性和不断变化的性质时,它们的有效性就会降低。为了解决这个问题,我们提出了一种名为基于区域的优化(RegO)的持续学习方法,用于音频deepfake检测。具体来说,我们使用Fisher信息矩阵来测量真实和虚假音频检测的重要神经元区域,将它们分为四个区域。首先,我们直接对不太重要的区域进行微调,以快速适应新任务。接下来,我们对仅对真实音频检测重要的区域并行地应用梯度优化,并且对仅对虚假音频检测重要的区域在正交方向上应用梯度优化。对于对两者都很重要的区域,我们使用基于样本比例的自适应梯度优化。这种区域自适应优化确保了记忆稳定性和学习可塑性之间的适当权衡。此外,为了解决旧任务中冗余神经元的增加,我们进一步引入Ebbinghaus遗忘机制来释放它们,从而提高模型学习更多广义判别特征的能力。实验结果表明,我们的方法在EER方面比最先进的用于音频deepfake检测的持续学习方法RWM提高了21.3%。此外,RegO的有效性超出了音频deepfake检测领域,在其他任务中显示出潜在的意义,例如图像识别。该代码可在https://github.com/cyjie429/RegO上获得
摘要:Rapid advancements in speech synthesis and voice conversion bring conveniencebut also new security risks, creating an urgent need for effective audiodeepfake detection. Although current models perform well, their effectivenessdiminishes when confronted with the diverse and evolving nature of real-worlddeepfakes. To address this issue, we propose a continual learning method namedRegion-Based Optimization (RegO) for audio deepfake detection. Specifically, weuse the Fisher information matrix to measure important neuron regions for realand fake audio detection, dividing them into four regions. First, we directlyfine-tune the less important regions to quickly adapt to new tasks. Next, weapply gradient optimization in parallel for regions important only to realaudio detection, and in orthogonal directions for regions important only tofake audio detection. For regions that are important to both, we use sampleproportion-based adaptive gradient optimization. This region-adaptiveoptimization ensures an appropriate trade-off between memory stability andlearning plasticity. Additionally, to address the increase of redundant neuronsfrom old tasks, we further introduce the Ebbinghaus forgetting mechanism torelease them, thereby promoting the capability of the model to learn moregeneralized discriminative features. Experimental results show our methodachieves a 21.3% improvement in EER over the state-of-the-art continuallearning approach RWM for audio deepfake detection. Moreover, the effectivenessof RegO extends beyond the audio deepfake detection domain, showing potentialsignificance in other tasks, such as image recognition. The code is availableat https://github.com/cyjie429/RegO
标题: 迈向新加坡及其他地区的言语基金会模式
链接:https://arxiv.org/abs/2412.11538
摘要:本技术报告介绍了MERaLiON语音编码器,这是一种基础模型,旨在支持广泛的下游语音应用。作为新加坡国家多模态大型语言模型计划的一部分,MERALiON语音编码器专为满足新加坡和周边东南亚地区的语音处理需求而设计。该模型目前主要支持英语,包括新加坡所说的各种语言。我们正在积极扩展我们的数据集,以便在后续版本中逐步覆盖其他语言。MERALiON语音编码器使用基于掩蔽语言建模的自监督学习方法,在20万小时的未标记语音数据上从头开始进行预训练。我们在下面详细描述我们的训练过程和超参数调整实验。我们的评估表明,语音识别的自发和新加坡语音基准的改进,同时保持竞争力的其他国家的最先进的语音编码器在其他十个语音任务。我们致力于发布我们的模型,支持在新加坡和其他地方进行更广泛的研究工作。
摘要:This technical report describes the MERaLiON Speech Encoder, a foundationmodel designed to support a wide range of downstream speech applications.Developed as part of Singapore's National Multimodal Large Language ModelProgramme, the MERaLiON Speech Encoder is tailored to address the speechprocessing needs in Singapore and the surrounding Southeast Asian region. Themodel currently supports mainly English, including the variety spoken inSingapore. We are actively expanding our datasets to gradually cover otherlanguages in subsequent releases. The MERaLiON Speech Encoder was pre-trainedfrom scratch on 200K hours of unlabelled speech data using a self-supervisedlearning approach based on masked language modelling. We describe our trainingprocedure and hyperparameter tuning experiments in detail below. Our evaluationdemonstrates improvements to spontaneous and Singapore speech benchmarks forspeech recognition, while remaining competitive to other state-of-the-artspeech encoders across ten other speech tasks. We commit to releasing ourmodel, supporting broader research endeavours, both in Singapore and beyond.
标题: Whisper-GPT:混合表示音频大语言模型
链接:https://arxiv.org/abs/2412.11449
备注:6 pages, 3 figures. 50th International Conference on Acoustics, Speech and Signal Processing, Hyderabad, India
摘要:我们建议WHISPER-GPT:一个用于语音和音乐的生成式大型语言模型(LLM),允许我们同时处理连续音频表示和离散令牌,作为单个架构的一部分。生成音频、语音和音乐模型已经出现了巨大的激增,这些模型利用了从神经压缩算法(例如ENCODEC)中导出的离散音频令牌。然而,这种方法的一个主要缺点是处理上下文长度。如果必须考虑下一个令牌预测的各种频率的所有音频内容,那么高保真生成架构就会崩溃。通过将连续音频表示(如频谱图)和离散声学令牌相结合,我们保留了两全其美:在单个令牌中具有特定时间实例音频所需的所有信息,但允许LLM预测未来令牌,以允许采样和离散空间提供的其他好处。我们展示了我们的架构如何提高困惑和负对数似然分数为下一个令牌预测相比,基于令牌的LLM语音和音乐。
摘要:We propose WHISPER-GPT: A generative large language model (LLM) for speechand music that allows us to work with continuous audio representations anddiscrete tokens simultaneously as part of a single architecture. There has beena huge surge in generative audio, speech, and music models that utilizediscrete audio tokens derived from neural compression algorithms, e.g. ENCODEC.However, one of the major drawbacks of this approach is handling the contextlength. It blows up for high-fidelity generative architecture if one has toaccount for all the audio contents at various frequencies for the next tokenprediction. By combining continuous audio representation like the spectrogramand discrete acoustic tokens, we retain the best of both worlds: Have all theinformation needed from the audio at a specific time instance in a singletoken, yet allow LLM to predict the future token to allow for sampling andother benefits discrete space provides. We show how our architecture improvesthe perplexity and negative log-likelihood scores for the next token predictioncompared to a token-based LLM for speech and music.
标题: Sonicmesh:利用声信号增强视觉受损环境中的3D人体网格重建
链接:https://arxiv.org/abs/2412.11325
摘要:从2D RGB图像进行3D人体网格重建(HMR)在光照不良、隐私问题或遮挡的环境中面临挑战。RGB成像的这些弱点可以通过广泛可用、易于部署并且能够穿透障碍物的声学信号来补充。然而,没有现有的方法有效地结合声信号与RGB数据的鲁棒3D HMR。主要挑战包括声学信号生成的低分辨率图像以及缺乏专用处理骨干。我们介绍SonicMesh,一种新的方法结合声学信号与RGB图像重建三维人体网格。为了解决低分辨率和缺乏专用的处理骨干在由声信号生成的图像的挑战,我们修改了现有的方法,HRNet,有效的特征提取。我们还集成了一种通用的特征嵌入技术,以提高跨维特征对齐的精度,使SonicMesh能够实现高精度。实验结果表明,SonicMesh准确地重建三维人体网格在具有挑战性的环境,如闭塞,非视线的情况下,和光线不足。
摘要:3D Human Mesh Reconstruction (HMR) from 2D RGB images faces challenges inenvironments with poor lighting, privacy concerns, or occlusions. Theseweaknesses of RGB imaging can be complemented by acoustic signals, which arewidely available, easy to deploy, and capable of penetrating obstacles.However, no existing methods effectively combine acoustic signals with RGB datafor robust 3D HMR. The primary challenges include the low-resolution imagesgenerated by acoustic signals and the lack of dedicated processing backbones.We introduce SonicMesh, a novel approach combining acoustic signals with RGBimages to reconstruct 3D human mesh. To address the challenges of lowresolution and the absence of dedicated processing backbones in imagesgenerated by acoustic signals, we modify an existing method, HRNet, foreffective feature extraction. We also integrate a universal feature embeddingtechnique to enhance the precision of cross-dimensional feature alignment,enabling SonicMesh to achieve high accuracy. Experimental results demonstratethat SonicMesh accurately reconstructs 3D human mesh in challengingenvironments such as occlusions, non-line-of-sight scenarios, and poorlighting.
标题: 流媒体语音的高效耳语
链接:https://arxiv.org/abs/2412.11272
摘要:以OpenAI的Whisper为例的语音基础模型,由于其出色的准确性和适应性,已成为语音理解的领导者。然而,它们的使用主要集中在处理预先录制的音频上,对流式语音的有效处理仍处于起步阶段。这一限制背后的几个核心挑战是:(1)这些模型是针对长的固定长度音频输入(通常为30秒)进行训练的。(2)对这样的输入进行编码涉及通过多个Transformer层处理多达1,500个令牌。(3)生成输出需要不规则的和计算量大的波束搜索。因此,在资源受限的边缘设备上进行流式语音处理比许多其他人工智能任务(包括文本生成)要求更高。为了应对这些挑战,我们引入了Whisper-T,这是一个结合了模型和系统级优化的创新框架:(1)Hush words,附加到输入中的短可学习音频片段,防止过度处理并减少模型中的幻觉。(2)波束修剪可随时间调整流式音频缓冲区,利用中间解码结果显著加快处理速度。(3)CPU/GPU流水线在编码和解码阶段之间动态分配资源,通过适应音频输入、模型特性和硬件的变化来优化性能。我们在基于ARM的平台上评估了Whisper-T,该平台具有4-12个CPU内核和10-30个GPU内核,延迟降低了1.6倍-4.7倍,每个字的延迟低至0.5秒,精度损失最小。此外,在MacBook Air上,Whisper-T每个字保持约1秒的延迟,同时仅消耗7瓦的总系统功率。
摘要:Speech foundation models, exemplified by OpenAI's Whisper, have emerged asleaders in speech understanding thanks to their exceptional accuracy andadaptability. However, their usage largely focuses on processing pre-recordedaudio, with the efficient handling of streaming speech still in its infancy.Several core challenges underlie this limitation: (1) These models are trainedfor long, fixed-length audio inputs (typically 30 seconds). (2) Encoding suchinputs involves processing up to 1,500 tokens through numerous transformerlayers. (3) Generating outputs requires an irregular and computationally heavybeam search. Consequently, streaming speech processing on edge devices withconstrained resources is more demanding than many other AI tasks, includingtext generation. To address these challenges, we introduce Whisper-T, aninnovative framework combining both model and system-level optimizations: (1)Hush words, short learnable audio segments appended to inputs, preventover-processing and reduce hallucinations in the model. (2) Beam pruning alignsstreaming audio buffers over time, leveraging intermediate decoding results tosignificantly speed up the process. (3) CPU/GPU pipelining dynamicallydistributes resources between encoding and decoding stages, optimizingperformance by adapting to variations in audio input, model characteristics,and hardware. We evaluate Whisper-T on ARM-based platforms with 4-12 CPU coresand 10-30 GPU cores, demonstrating latency reductions of 1.6x-4.7x, achievingper-word delays as low as 0.5 seconds with minimal accuracy loss. Additionally,on a MacBook Air, Whisper-T maintains approximately 1-second latency per wordwhile consuming just 7 Watts of total system power.
标题: Hanprome:修改后的韩语发音表达
链接:https://arxiv.org/abs/2412.11090
备注:21 pages
摘要:韩语是作为一种语音字母表创建的,并且已知在现有字母表中字母和发音之间具有最好的1:1对应关系。在本文中,我们探讨了修改的基本形式,并使用它作为一种语音符号的可能性。这种方法的核心概念是保持字母表的基本形式,只修改笔画的形状,而不是字母本身。据我们所知,以前没有任何语言试图通过改变字母笔画的形状来表达与原始字母不同的字母发音,本文可能是这个方向的第一次尝试。
摘要:Hangeul was created as a phonetic alphabet and is known to have the best 1:1correspondence between letters and pronunciation among existing alphabets. Inthis paper, we examine the possibility of modifying the basic form of Hangeuland using it as a kind of phonetic symbol. The core concept of this approach isto preserve the basic form of the alphabet, modifying only the shape of astroke rather than the letter itself. To the best of our knowledge, no previousattempts in any language have been made to express pronunciations of analphabet different from the original simply by changing the shape of thealphabet strokes, and this paper is probably the first attempt in thisdirection.
标题: 作曲家对人工智能音乐工具的评价:以人为本的设计见解
链接:https://arxiv.org/abs/2412.10968
备注:Accepted to NeurIPS 2024 Workshop on Generative AI and Creativity: A dialogue between machine learning researchers and creative professionals in Vancouver, Canada
摘要:我们提出了一项研究,探讨了以用户为中心的设计在开发生成AI(GenAI)工具进行音乐创作中的作用。通过与专业作曲家的半结构化访谈,我们收集了关于创造变化的新生成模型的见解,强调了对信任,透明度和道德设计的关注。这些发现有助于形成一个反馈回路,指导对模型的改进,强调可追溯性,透明度和可解释性。他们还揭示了新的创新领域,包括可控性的新功能以及GenAI模型的伦理和实际实施的研究问题。
摘要:We present a study that explores the role of user-centred design indeveloping Generative AI (GenAI) tools for music composition. Throughsemi-structured interviews with professional composers, we gathered insights ona novel generative model for creating variations, highlighting concerns aroundtrust, transparency, and ethical design. The findings helped form a feedbackloop, guiding improvements to the model that emphasised traceability,transparency and explainability. They also revealed new areas for innovation,including novel features for controllability and research questions on theethical and practical implementation of GenAI models.
标题: 使用深度神经网络对语音中波斯语孤立数字的鲁棒识别
链接:https://arxiv.org/abs/2412.10857
备注:15 pages, submitted to journal
摘要:近年来,人工智能(AI)在语音识别应用中取得了显着进展。与数字系统的基于语音的交互,特别是人工智能驱动的数字识别,已成为一个重要的应用。然而,现有的基于神经网络的方法往往忽略了噪声的影响,导致在噪声环境中的准确性降低。这项研究解决了在嘈杂的环境中识别孤立的波斯语口语数字(0到9)的挑战,特别是区分语音相似的数字。所提出的方法是针对与说话者无关的识别而设计的,将残余卷积神经网络和双向门控递归单元结合在混合结构中用于波斯数字识别。该方法使用单词单元代替音素单元作为输入。从FARSDIGIT 1数据库的51个发言人的音频数据使用各种噪声增强后,梅尔频率倒谱系数(MFCC)技术用于特征提取。实验结果表明,所提出的方法的有效性与98.53%,96.10%和95.9%的识别准确率的训练,验证和测试,分别。在噪声环境中,该方法比基于音素单元的波斯数字LSTM方法平均性能提高了26.88%。此外,该方法的准确性是7.61%,优于梅尔尺度二维根倒谱系数(MTDRCC)特征提取技术与MLP模型在测试数据的相同的数据集。
摘要:In recent years, artificial intelligence (AI) has advanced significantly inspeech recognition applications. Speech-based interaction with digital systems,particularly AI-driven digit recognition, has emerged as a prominentapplication. However, existing neural network-based methods often neglect theimpact of noise, leading to reduced accuracy in noisy environments. This studytackles the challenge of recognizing the isolated spoken Persian numbers (zeroto nine), particularly distinguishing phonetically similar numbers, in noisyenvironments. The proposed method, which is designed for speaker-independentrecognition, combines residual convolutional neural network and bidirectionalgated recurrent unit in a hybrid structure for Persian number recognition. Thismethod employs word units as input instead of phoneme units. Audio data from 51speakers of FARSDIGIT1 database are utilized after augmentation using variousnoises, and the Mel-Frequency Cepstral Coefficients (MFCC) technique isemployed for feature extraction. The experimental results show the proposedmethod efficacy with 98.53%, 96.10%, and 95.9% recognition accuracy fortraining, validation, and test, respectively. In the noisy environment, theproposed method exhibits an average performance improvement of 26.88% overphoneme unit-based LSTM method for Persian numbers. In addition, the accuracyof the proposed method is 7.61% better than that of the Mel-scale Two DimensionRoot Cepstrum Coefficients (MTDRCC) feature extraction technique along with MLPmodel in the test data for the same dataset.
标题: 使用深度一类支持量数据描述的工业机器中基于音频的异常检测
链接:https://arxiv.org/abs/2412.10792
备注:To be published in 2025 IEEE Symposium Series on Computational Intelligence
摘要:工业设备的频繁故障和故障已经促使人们越来越关注利用具有成本效益且易于部署的传感器(例如麦克风)来对机械进行有效的状态监测。麦克风为广泛使用的状态监测传感器提供了一种低成本的替代方案,其高带宽和检测其他传感器可能灵敏度较低的细微异常的能力。在这项研究中,我们调查故障的工业机器,以评估和比较不同机器类型和故障条件下的异常检测性能。机器声音的Log-Mel谱图被用作输入,并且使用两种不同方法的曲线下面积(AUC)得分来评估性能:基线密集自动编码器(AE)和具有不同子空间维度的一类深度支持向量数据描述(deep SVDD)。我们在MIMII声音数据集上的结果表明,子空间维度为2的深度SVDD方法提供了卓越的异常检测性能,与基线模型的0.82,0.72和0.64相比,6 dB,0 dB和-6 dB信噪比(SNR)的平均AUC分数分别为0.84,0.80和0.69。此外,深度SVDD需要比基线密集AE少7.4倍的可训练参数,强调其在有效性和计算效率方面的优势。
摘要:The frequent breakdowns and malfunctions of industrial equipment have drivenincreasing interest in utilizing cost-effective and easy-to-deploy sensors,such as microphones, for effective condition monitoring of machinery.Microphones offer a low-cost alternative to widely used condition monitoringsensors with their high bandwidth and capability to detect subtle anomaliesthat other sensors might have less sensitivity. In this study, we investigatemalfunctioning industrial machines to evaluate and compare anomaly detectionperformance across different machine types and fault conditions. Log-Melspectrograms of machinery sound are used as input, and the performance isevaluated using the area under the curve (AUC) score for two different methods:baseline dense autoencoder (AE) and one-class deep Support Vector DataDescription (deep SVDD) with different subspace dimensions. Our results overthe MIMII sound dataset demonstrate that the deep SVDD method with a subspacedimension of 2 provides superior anomaly detection performance, achievingaverage AUC scores of 0.84, 0.80, and 0.69 for 6 dB, 0 dB, and -6 dBsignal-to-noise ratios (SNRs), respectively, compared to 0.82, 0.72, and 0.64for the baseline model. Moreover, deep SVDD requires 7.4 times fewer trainableparameters than the baseline dense AE, emphasizing its advantage in botheffectiveness and computational efficiency.
标题: VinTAGE:联合视频和文本调节,实现整体音频生成
链接:https://arxiv.org/abs/2412.10768
摘要:音频生成的最新进展集中在文本到音频(T2 A)和视频到音频(V2 A)任务上。然而,T2 A或V2 A方法不能生成整体声音(屏幕外和屏幕外)。这是因为T2 A无法生成与屏幕上对象对齐的声音,而V2 A无法生成语义完整的声音(缺少屏幕外的声音)。在这项工作中,我们解决了整体音频生成的任务:给定视频和文本提示,我们的目标是生成与视频在时间上同步并与文本和视频在语义上对齐的屏上和屏外声音。以前的联合文本和视频到音频生成的方法往往受到模态偏见,有利于一个模态。为了克服这一限制,我们引入了VINTAGe,一个基于流的Transformer模型,它联合考虑文本和视频来指导音频生成。我们的框架包括两个关键组成部分:视觉文本编码器和联合VT-SiT模型。为了减少模态偏差并提高生成质量,我们采用预训练的单模态文本到音频和视频到音频生成模型进行额外的指导。由于缺乏适当的基准测试,我们还介绍了VINTAGe-Bench,这是一个包含636个视频-文本-音频对的数据集,其中包含屏幕外和屏幕外的声音。我们在VINTAGe-Bench上的综合实验表明,联合文本和视觉交互对于整体音频生成是必要的。此外,VinTAGe在VGGSound基准测试中取得了最先进的结果。我们的源代码和预训练模型将被发布。演示版可在以下网址获得:https://www.youtube.com/watch? v=QmqWhUjPkJI。
摘要:Recent advances in audio generation have focused on text-to-audio (T2A) andvideo-to-audio (V2A) tasks. However, T2A or V2A methods cannot generateholistic sounds (onscreen and off-screen). This is because T2A cannot generatesounds aligning with onscreen objects, while V2A cannot generate semanticallycomplete (offscreen sounds missing). In this work, we address the task ofholistic audio generation: given a video and a text prompt, we aim to generateboth onscreen and offscreen sounds that are temporally synchronized with thevideo and semantically aligned with text and video. Previous approaches forjoint text and video-to-audio generation often suffer from modality bias,favoring one modality over the other. To overcome this limitation, we introduceVinTAGe, a flow-based transformer model that jointly considers text and videoto guide audio generation. Our framework comprises two key components: aVisual-Text Encoder and a Joint VT-SiT model. To reduce modality bias andimprove generation quality, we employ pretrained uni-modal text-to-audio andvideo-to-audio generation models for additional guidance. Due to the lack ofappropriate benchmarks, we also introduce VinTAGe-Bench, a dataset of 636video-text-audio pairs containing both onscreen and offscreen sounds. Ourcomprehensive experiments on VinTAGe-Bench demonstrate that joint text andvisual interaction is necessary for holistic audio generation. Furthermore,VinTAGe achieves state-of-the-art results on the VGGSound benchmark. Our sourcecode and pre-trained models will be released. Demo is available at:https://www.youtube.com/watch?v=QmqWhUjPkJI.
标题: 日语ASB多语言模型的有效适应
链接:https://arxiv.org/abs/2412.10705
摘要:本研究探讨了微调多语言ASR(自动语音识别)模型,特别是OpenAI的Whisper-Tiny,以提高日语的性能。虽然像Whisper这样的多语言模型提供了多功能性,但它们在特定语言中往往缺乏精确性。相反,像ReazonSpeech这样的单语模型在特定语言的任务中表现出色,但适应性较差。使用日本特定的数据集和低秩自适应(LoRA)以及端到端(E2 E)训练,我们对Whisper-Tiny进行了微调,以弥合这一差距。我们的研究结果表明,微调将Whisper-Tiny的字符错误率(CER)从LoRA的32.7降低到20.8,并通过端到端微调降低到14.7,超过了Whisper-Base的CER 20.2。然而,特定领域术语的挑战仍然存在,突出表明需要专门的数据集。这些发现表明,微调多语言模型可以实现强大的语言特定性能,同时保持其灵活性。这种方法提供了一种可扩展的解决方案,用于在资源受限的环境和具有复杂书写系统的语言(如日语)中改进ASR。
摘要:This study explores fine-tuning multilingual ASR (Automatic SpeechRecognition) models, specifically OpenAI's Whisper-Tiny, to improve performancein Japanese. While multilingual models like Whisper offer versatility, theyoften lack precision in specific languages. Conversely, monolingual models likeReazonSpeech excel in language-specific tasks but are less adaptable. UsingJapanese-specific datasets and Low-Rank Adaptation (LoRA) along with end-to-end(E2E) training, we fine-tuned Whisper-Tiny to bridge this gap. Our results showthat fine-tuning reduced Whisper-Tiny's Character Error Rate (CER) from 32.7 to20.8 with LoRA and to 14.7 with end-to-end fine-tuning, surpassingWhisper-Base's CER of 20.2. However, challenges with domain-specific termsremain, highlighting the need for specialized datasets. These findingsdemonstrate that fine-tuning multilingual models can achieve stronglanguage-specific performance while retaining their flexibility. This approachprovides a scalable solution for improving ASR in resource-constrainedenvironments and languages with complex writing systems like Japanese.
标题: 隐藏的回声在音频到音频生成乐器模型中的训练中幸存下来
链接:https://arxiv.org/abs/2412.10649
备注:8 pages, 11 Figures, Proceedings of 2025 AAAI Workshop on AI for Music
摘要:随着生成技术在音频领域的普及,人们越来越有兴趣追溯这些复杂的模型,以了解它们如何利用训练数据来合成新的示例,以确保它们使用正确授权的数据,并阐明它们的黑箱行为。在本文中,我们表明,如果无法感知的回声隐藏在训练数据中,各种各样的音频到音频架构(可区分数字信号处理(DDSP),实时音频变分自动编码器(RAVE)和“舞蹈扩散”)将在其输出中再现这些回声。隐藏一个单一的回声是特别强大的所有架构,但我们也显示出有希望的结果隐藏更长的时间扩展回声模式,以增加信息容量。我们的结论表明,回声使他们的方式进入微调模型,他们生存的混合/去混合,他们生存在训练过程中的音高变化增强。因此,这个简单的,经典的水印思想显示了标记生成音频模型的重要承诺。
摘要:As generative techniques pervade the audio domain, there has been increasinginterest in tracing back through these complicated models to understand howthey draw on their training data to synthesize new examples, both to ensurethat they use properly licensed data and also to elucidate their black boxbehavior. In this paper, we show that if imperceptible echoes are hidden in thetraining data, a wide variety of audio to audio architectures (differentiabledigital signal processing (DDSP), Realtime Audio Variational autoEncoder(RAVE), and ``Dance Diffusion'') will reproduce these echoes in their outputs.Hiding a single echo is particularly robust across all architectures, but wealso show promising results hiding longer time spread echo patterns for anincreased information capacity. We conclude by showing that echoes make theirway into fine tuned models, that they survive mixing/demixing, and that theysurvive pitch shift augmentation during training. Hence, this simple, classicalidea in watermarking shows significant promise for tagging generative audiomodels.
标题: 临界点、脉搏弹性和调性张力:关于临界点产生的实证研究
链接:https://arxiv.org/abs/2412.10481
备注:International Society for Music Information Retrieval Conference, Oct 2017, Suzhou, China, China
摘要:引爆点是指音乐作品中的关键转折点。这项研究提出了第一步,定量和系统地描述音乐性质的临界点。定时信息和计算得出的音调张力值,对应于不和谐音,距离键,和谐波运动相比,在Ashkenazy的六个肖邦马祖卡的录音,由35名听众确定的临界点。分析表明,所有流行的临界点,但一个可以解释为统计上显着的时间偏差或变化点,至少在三个张力参数之一。
摘要:Tipping points are moments of change that characterise crucial turning pointsin a piece of music. This study presents a first step towards quantitativelyand systematically describing the musical properties of tipping points. Timinginformation and computationally-derived tonal tension values which correspondto dissonance, distance from key, and harmonic motion are compared to tippingpoints in Ashkenazy's recordings of six Chopin Mazurkas, as identified by 35listeners. The analysis shows that all popular tipping points but one could beexplained by statistically significant timing deviations or changepoints in atleast one of the three tension parameters.
标题: Mel频率倒谱系数的比较分析和基于子波的音频信号处理用于口语中的情感检测和心理健康评估
链接:https://arxiv.org/abs/2412.10469
摘要:技术和心理健康的交叉刺激了评估情绪健康的创新方法,特别是通过应用于音频数据分析的计算技术。本研究探讨了卷积神经网络(CNN)和长短期记忆(LSTM)模型在小波提取特征和Mel频率倒谱系数(MFCC)的口语情感检测中的应用。数据增强技术,特征提取,归一化和模型训练进行评估模型的分类情感状态的性能。结果表明,CNN模型的准确率为61%,而LSTM模型的准确率为56%。这两种模型在预测特定情绪(如惊讶和愤怒)方面表现出更好的性能,利用了不同的音频特征,如音高和速度变化。建议包括进一步探索先进的数据增强技术,结合特征提取方法,以及将语言分析与语音特征相结合,以提高心理健康诊断的准确性。标准化数据集收集和共享的合作,建议促进情感计算和心理健康护理干预措施的进步。
摘要:The intersection of technology and mental health has spurred innovativeapproaches to assessing emotional well-being, particularly throughcomputational techniques applied to audio data analysis. This study exploresthe application of Convolutional Neural Network (CNN) and Long Short-TermMemory (LSTM) models on wavelet extracted features and Mel-frequency CepstralCoefficients (MFCCs) for emotion detection from spoken speech. Dataaugmentation techniques, feature extraction, normalization, and model trainingwere conducted to evaluate the models' performance in classifying emotionalstates. Results indicate that the CNN model achieved a higher accuracy of 61%compared to the LSTM model's accuracy of 56%. Both models demonstrated betterperformance in predicting specific emotions such as surprise and anger,leveraging distinct audio features like pitch and speed variations.Recommendations include further exploration of advanced data augmentationtechniques, combined feature extraction methods, and the integration oflinguistic analysis with speech characteristics for improved accuracy in mentalhealth diagnostics. Collaboration for standardized dataset collection andsharing is recommended to foster advancements in affective computing and mentalhealth care interventions.
标题: 通过视听内容的文本情感描述丰富多模式情感分析
链接:https://arxiv.org/abs/2412.10460
备注:None
摘要:多模态情感分析(MSA)是一个重要的研究前沿,旨在通过融合文本,音频和视觉数据来全面揭示人类的情感。然而,在音频和视频表达中辨别微妙的情感细微差别构成了一个巨大的挑战,特别是当各个部分的情感极性看起来相似时。在本文中,我们的目标是突出情感相关的属性的音频和视觉模态,以促进多模态融合的背景下,在视觉-音频场景中的细微差别的情感变化。为此,我们引入DEVA,一个渐进的融合框架,建立在文本情感描述的基础上,旨在强调视听内容的情感特征。DEVA采用情感描述生成器(EDG)将原始音频和视频数据转化为文本化的情感描述,从而放大它们的情感特征。然后将这些描述与源数据集成,以产生更丰富的增强功能。此外,DEVA还集成了文本引导渐进融合模块(TPF),利用不同级别的文本作为核心模态指南。该模块逐步融合视听次要模态,以减轻文本和视听模态之间的差异。在广泛使用的情感分析基准数据集(包括MOSI,MOSEI和CH-SIMS)上的实验结果强调了与最先进的模型相比的显着增强。此外,细粒度的情绪实验证实了强大的敏感性DEVA微妙的情绪变化。
摘要:Multimodal Sentiment Analysis (MSA) stands as a critical research frontier,seeking to comprehensively unravel human emotions by amalgamating text, audio,and visual data. Yet, discerning subtle emotional nuances within audio andvideo expressions poses a formidable challenge, particularly when emotionalpolarities across various segments appear similar. In this paper, our objectiveis to spotlight emotion-relevant attributes of audio and visual modalities tofacilitate multimodal fusion in the context of nuanced emotional shifts invisual-audio scenarios. To this end, we introduce DEVA, a progressive fusionframework founded on textual sentiment descriptions aimed at accentuatingemotional features of visual-audio content. DEVA employs an EmotionalDescription Generator (EDG) to transmute raw audio and visual data intotextualized sentiment descriptions, thereby amplifying their emotionalcharacteristics. These descriptions are then integrated with the source data toyield richer, enhanced features. Furthermore, DEVA incorporates the Text-guidedProgressive Fusion Module (TPF), leveraging varying levels of text as a coremodality guide. This module progressively fuses visual-audio minor modalitiesto alleviate disparities between text and visual-audio modalities. Experimentalresults on widely used sentiment analysis benchmark datasets, including MOSI,MOSEI, and CH-SIMS, underscore significant enhancements compared tostate-of-the-art models. Moreover, fine-grained emotion experiments corroboratethe robust sensitivity of DEVA to subtle emotional variations.
标题: 在心理健康中利用音频和文本模式:LLM绩效研究
链接:https://arxiv.org/abs/2412.10417
摘要:心理健康障碍在世界范围内日益普遍,迫切需要创新工具来支持早期诊断和干预。本研究探讨了大型语言模型(LLM)在多模态心理健康诊断中的潜力,特别是通过文本和音频模式检测抑郁症和创伤后应激障碍。使用E-DAIC数据集,我们比较了文本和音频模态,以研究LLM是否可以在音频输入下表现得同样好或更好。我们进一步研究了两种模式的整合,以确定这是否可以提高诊断准确性,这通常会导致性能指标的改善。我们的分析专门利用定制的指标;模态优先度评分和不一致解决评分来评估组合模态如何影响模型性能。当使用组合方式时,Gemini 1.5 Pro模型在二元抑郁分类中获得了最高分数,在整个数据集中评估的F1分数为0.67,平衡准确度(BA)为77.4%。这些结果比文本模态的性能提高了3.1%,比音频模态提高了2.7%,突出了整合模态以提高诊断准确性的有效性。值得注意的是,所有结果都是在zero-shot推断中获得的,突出了模型的鲁棒性,而不需要特定于任务的微调。为了探索不同配置对模型性能的影响,我们使用zero-shot和Few-Shot提示进行二进制,严重性和多类任务,检查提示变化对性能的影响。结果显示,Gemini 1.5 Pro在文本和音频模式中以及GPT-4 o mini在文本模式中的模型在多个任务中的平衡准确性和F1得分方面往往超过其他模型。
摘要:Mental health disorders are increasingly prevalent worldwide, creating anurgent need for innovative tools to support early diagnosis and intervention.This study explores the potential of Large Language Models (LLMs) in multimodalmental health diagnostics, specifically for detecting depression and PostTraumatic Stress Disorder through text and audio modalities. Using the E-DAICdataset, we compare text and audio modalities to investigate whether LLMs canperform equally well or better with audio inputs. We further examine theintegration of both modalities to determine if this can enhance diagnosticaccuracy, which generally results in improved performance metrics. Our analysisspecifically utilizes custom-formulated metrics; Modal Superiority Score andDisagreement Resolvement Score to evaluate how combined modalities influencemodel performance. The Gemini 1.5 Pro model achieves the highest scores inbinary depression classification when using the combined modality, with an F1score of 0.67 and a Balanced Accuracy (BA) of 77.4%, assessed across the fulldataset. These results represent an increase of 3.1% over its performance withthe text modality and 2.7% over the audio modality, highlighting theeffectiveness of integrating modalities to enhance diagnostic accuracy.Notably, all results are obtained in zero-shot inferring, highlighting therobustness of the models without requiring task-specific fine-tuning. Toexplore the impact of different configurations on model performance, we conductbinary, severity, and multiclass tasks using both zero-shot and few-shotprompts, examining the effects of prompt variations on performance. The resultsreveal that models such as Gemini 1.5 Pro in text and audio modalities, andGPT-4o mini in the text modality, often surpass other models in balancedaccuracy and F1 scores across multiple tasks.
