本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
链接:https://arxiv.org/abs/2502.04328
摘要:大型语言模型的最新进展,特别是在GPT-4 o之后,引发了人们对开发能够理解更多模态的全模态模型的兴趣。虽然已经出现了一些开源替代方案,但在性能方面仍明显落后于专门的单一模式模型。在本文中,我们介绍了Ola,一种全模态语言模型,与专业同行相比,它在图像,视频和音频理解方面具有竞争力的性能。Ola的核心设计在于其渐进式模态对齐策略,该策略渐进地扩展了语言模型的支持模态。我们的训练管道从最独特的模态开始:图像和文本,然后使用连接语言和音频知识的语音数据以及连接所有模态的视频数据逐步扩展模型的技能集。渐进式学习管道还使我们能够保持相对较小的跨模态对齐数据,从而使从现有视觉语言模型开发全模态变得简单且成本更低。此外,为了解锁像GPT-4 o这样的高级交互体验,我们进一步设计了一种用于流媒体语音生成的逐行解码解决方案。大量的实验表明,Ola在所有模态上都超越了现有的开放式全模态LLM,同时与类似尺寸的最先进的专业模型相比,实现了极具竞争力的性能。我们的目标是使Ola成为一个完全开放的全模态理解解决方案,以推动这一新兴领域的未来研究。模型权重、代码和数据在https://github.com/Ola-Omni/Ola上开源。
摘要:Recent advances in large language models, particularly following GPT-4o, havesparked increasing interest in developing omni-modal models capable ofunderstanding more modalities. While some open-source alternatives haveemerged, there is still a notable lag behind specialized single-modality modelsin performance. In this paper, we present Ola, an Omni-modal language modelthat achieves competitive performance across image, video, and audiounderstanding compared to specialized counterparts. The core design of Ola liesin its progressive modality alignment strategy that extends the supportingmodality of the language model progressively. Our training pipeline begins withthe most distinct modalities: image and text, then gradually expands the skillsets of the model using speech data that connects language and audio knowledge,and video data that connects all modalities. The progressive learning pipelinealso enables us to maintain a relatively small size of the cross-modalalignment data, making developing omni-modal from existing vision-languagemodels easy and less costly. Moreover, to unlock an advanced interactiveexperience like GPT-4o, we further design a sentence-wise decoding solution forstreaming speech generation. Extensive experiments demonstrate that Olasurpasses existing open omni-modal LLMs across all modalities while achievinghighly competitive performance compared to state-of-the-art specialized modelsof similar sizes. We aim to make Ola a fully open omni-modal understandingsolution to advance future research in this emerging field. Model weights,code, and data are open-sourced at https://github.com/Ola-Omni/Ola.
标题:XAttnMark:通过交叉注意力学习稳健的音频水印
链接:https://arxiv.org/abs/2502.04230
备注:24 pages, 10 figures
摘要:生成式音频合成和编辑技术的迅速普及引发了人们对版权侵权、数据来源以及通过deepfake音频传播错误信息的严重担忧。水印通过将不可察觉、可识别和可追踪的标记嵌入音频内容中提供了一种主动解决方案。虽然最近基于神经网络的水印方法(如WavMark和AudioSeal)提高了鲁棒性和质量,但它们很难同时实现鲁棒性检测和准确归因。本文介绍了交叉注意鲁棒音频水印(XAttnMark),它通过利用生成器和检测器之间的部分参数共享,用于有效消息检索的交叉注意机制和用于改进消息分发的时间调节模块来弥合这一差距。此外,我们提出了一种心理声学对齐的时间频率掩蔽损失,可以捕获细粒度的听觉掩蔽效果,从而增强水印的不可感知性。我们的方法在检测和归因方面都实现了最先进的性能,对各种音频转换表现出卓越的鲁棒性,包括具有强大编辑强度的挑战性生成编辑。该项目的网页可在https://liuyixin-louis.github.io/xattnmark/上查阅。
摘要:The rapid proliferation of generative audio synthesis and editingtechnologies has raised significant concerns about copyright infringement, dataprovenance, and the spread of misinformation through deepfake audio.Watermarking offers a proactive solution by embedding imperceptible,identifiable, and traceable marks into audio content. While recent neuralnetwork-based watermarking methods like WavMark and AudioSeal have improvedrobustness and quality, they struggle to achieve both robust detection andaccurate attribution simultaneously. This paper introduces Cross-AttentionRobust Audio Watermark (XAttnMark), which bridges this gap by leveragingpartial parameter sharing between the generator and the detector, across-attention mechanism for efficient message retrieval, and a temporalconditioning module for improved message distribution. Additionally, we proposea psychoacoustic-aligned temporal-frequency masking loss that capturesfine-grained auditory masking effects, enhancing watermark imperceptibility.Our approach achieves state-of-the-art performance in both detection andattribution, demonstrating superior robustness against a wide range of audiotransformations, including challenging generative editing with strong editingstrength. The project webpage is available athttps://liuyixin-louis.github.io/xattnmark/.
标题:数据驱动的双麦克风现场声吸收测量方法
链接:https://arxiv.org/abs/2502.04143
备注:41 pages, 8 figures
摘要:本文提出了一种基于数据驱动的方法,利用神经网络和双传声器测量方法对有限多孔板的吸声系数进行了估计。1D卷积网络从在两个麦克风位置处测量的声压之间的复值传递函数预测吸声系数。该网络的训练和验证的边界元模型使用Delany-Bazley-Miki模型生成的数值数据,证明了各种数值样本的准确预测。该方法进行了实验验证与折流矩形样品的纤维材料,其中样品尺寸和源高度是不同的。结果表明,神经网络提供了可靠的预测,使用传统的双麦克风的方法,如果样品是无限的多孔材料的现场吸声的可能性。由该网络得到的垂直入射吸声系数与理论计算值和阻抗管中的吸声系数进行了比较。所提出的方法具有很好的前景,估计吸声材料安装后,在现实的操作条件下的吸声系数。
摘要:This work presents a data-driven approach to estimating the sound absorptioncoefficient of an infinite porous slab using a neural network and atwo-microphone measurement on a finite porous sample. A 1D-convolutionalnetwork predicts the sound absorption coefficient from the complex-valuedtransfer function between the sound pressure measured at the two microphonepositions. The network is trained and validated with numerical data generatedby a boundary element model using the Delany-Bazley-Miki model, demonstratingaccurate predictions for various numerical samples. The method isexperimentally validated with baffled rectangular samples of a fibrousmaterial, where sample size and source height are varied. The results show thatthe neural network offers the possibility to reliably predict the in-situ soundabsorption of a porous material using the traditional two-microphone method asif the sample were infinite. The normal-incidence sound absorption coefficientobtained by the network compares well with that obtained theoretically and inan impedance tube. The proposed method has promising perspectives forestimating the sound absorption coefficient of acoustic materials afterinstallation and in realistic operational conditions.
标题:迈向跨维度和类别模型的统一音乐情感识别
链接:https://arxiv.org/abs/2502.03979
摘要:音乐情感识别(MER)中最重要的挑战之一来自以下事实:情感标签在关于情感表示的数据集之间可以是异构的,包括分类的(例如,快乐、悲伤)与维度标签(例如,价-唤醒)。在本文中,我们提出了一个统一的多任务学习框架,该框架结合了这两种类型的标签,因此能够在多个数据集上进行训练。该框架使用有效的输入表示,其组合音乐特征(即,键和和弦)和MERT嵌入。此外,知识蒸馏用于将在单个数据集上训练的教师模型的知识转移到学生模型,从而增强其跨多个任务进行概括的能力。为了验证我们提出的框架,我们对各种数据集进行了广泛的实验,包括MTG-Jamendo、DEAM、PMEmo和EmoMusic。根据我们的实验结果,包括音乐功能,多任务学习和知识蒸馏显着提高性能。特别是,我们的模型优于最先进的模型,包括来自MTG-Jamendo数据集的MediaEval 2021竞赛的最佳模型。我们的工作通过允许在一个统一的框架中组合分类和维度情感标签,从而实现跨数据集的训练,为MER做出了重大贡献。
摘要:One of the most significant challenges in Music Emotion Recognition (MER)comes from the fact that emotion labels can be heterogeneous across datasetswith regard to the emotion representation, including categorical (e.g., happy,sad) versus dimensional labels (e.g., valence-arousal). In this paper, wepresent a unified multitask learning framework that combines these two types oflabels and is thus able to be trained on multiple datasets. This framework usesan effective input representation that combines musical features (i.e., key andchords) and MERT embeddings. Moreover, knowledge distillation is employed totransfer the knowledge of teacher models trained on individual datasets to astudent model, enhancing its ability to generalize across multiple tasks. Tovalidate our proposed framework, we conducted extensive experiments on avariety of datasets, including MTG-Jamendo, DEAM, PMEmo, and EmoMusic.According to our experimental results, the inclusion of musical features,multitask learning, and knowledge distillation significantly enhancesperformance. In particular, our model outperforms the state-of-the-art models,including the best-performing model from the MediaEval 2021 competition on theMTG-Jamendo dataset. Our work makes a significant contribution to MER byallowing the combination of categorical and dimensional emotion labels in oneunified framework, thus enabling training across datasets.
标题:UniForm:用于音频视频生成的统一扩散Transformer
链接:https://arxiv.org/abs/2502.03897
摘要:作为一种自然的多模态内容,音频视频提供了身临其境的感官体验。因此,音频-视频生成系统具有巨大的潜力。然而,现有的基于扩散的研究主要采用相对独立的模块来生成每个模态,缺乏对共享权重生成模块的探索。这种方法可能未充分利用音频和视觉模态之间的内在相关性,可能导致次优的生成质量。为了解决这个问题,我们提出了统一的形式,一个统一的扩散Transformer,旨在提高跨模态的一致性。通过连接听觉和视觉信息,UniForm学会在统一的潜在空间内同时生成音频和视频,从而促进创建高质量和对齐良好的视听对。大量的实验表明,我们的方法在联合音视频生成,音频引导的视频生成,视频引导的音频生成任务的优越性能。我们的演示可在https://uniform-t2av.github.io/上获得。
摘要:As a natural multimodal content, audible video delivers an immersive sensoryexperience. Consequently, audio-video generation systems have substantialpotential. However, existing diffusion-based studies mainly employ relativelyindependent modules for generating each modality, which lack exploration ofshared-weight generative modules. This approach may under-use the intrinsiccorrelations between audio and visual modalities, potentially resulting insub-optimal generation quality. To address this, we propose UniForm, a unifieddiffusion transformer designed to enhance cross-modal consistency. Byconcatenating auditory and visual information, UniForm learns to generate audioand video simultaneously within a unified latent space, facilitating thecreation of high-quality and well-aligned audio-visual pairs. Extensiveexperiments demonstrate the superior performance of our method in jointaudio-video generation, audio-guided video generation, and video-guided audiogeneration tasks. Our demos are available at https://uniform-t2av.github.io/.
标题:Llasa:基于Llama的语音合成的扩展训练时间和推理时间计算
链接:https://arxiv.org/abs/2502.04128
摘要:基于文本的大型语言模型(LLM)的最新进展,特别是GPT系列和o1模型,已经证明了扩展训练时间和推理时间计算的有效性。然而,利用LLM的当前最先进的TTS系统通常是多级的,需要单独的模型(例如,LLM之后的扩散模型),这使得在训练或测试期间是否缩放特定模型的决定变得复杂。本文的主要贡献如下:首先,我们研究了语音合成中训练时间和推理时间计算的尺度问题。其次,我们提出了一个简单的框架Llasa语音合成,采用单层矢量量化(VQ)编解码器和一个单一的Transformer架构,完全符合标准LLM,如Llama。我们的实验表明,Llasa的缩放训练时间计算始终提高了合成语音的自然度,并能够生成更复杂和准确的韵律模式。此外,从缩放推理时间计算的角度来看,我们在搜索过程中采用语音理解模型作为验证者,发现缩放推理时间计算将采样模式向特定验证者的偏好转移,从而提高情感表达力,音色一致性和内容准确性。此外,我们还公开了TTS模型(1B,3B,8B)和编解码器模型的检查点和训练代码。
摘要:Recent advances in text-based large language models (LLMs), particularly inthe GPT series and the o1 model, have demonstrated the effectiveness of scalingboth training-time and inference-time compute. However, currentstate-of-the-art TTS systems leveraging LLMs are often multi-stage, requiringseparate models (e.g., diffusion models after LLM), complicating the decisionof whether to scale a particular model during training or testing. This workmakes the following contributions: First, we explore the scaling of train-timeand inference-time compute for speech synthesis. Second, we propose a simpleframework Llasa for speech synthesis that employs a single-layer vectorquantizer (VQ) codec and a single Transformer architecture to fully align withstandard LLMs such as Llama. Our experiments reveal that scaling train-timecompute for Llasa consistently improves the naturalness of synthesized speechand enables the generation of more complex and accurate prosody patterns.Furthermore, from the perspective of scaling inference-time compute, we employspeech understanding models as verifiers during the search, finding thatscaling inference-time compute shifts the sampling modes toward the preferencesof specific verifiers, thereby improving emotional expressiveness, timbreconsistency, and content accuracy. In addition, we released the checkpoint andtraining code for our TTS model (1B, 3B, 8B) and codec model publiclyavailable.
标题:DiTAR:语音生成的扩散Transformer自回归建模
链接:https://arxiv.org/abs/2502.03930
备注:16 pages, 8 figures
摘要:最近的几项研究试图通过结合扩散和自回归模型来自动生成连续语音表示,而不需要离散语音标记,但它们经常面临计算负荷过大或次优结果的挑战。在这项工作中,我们提出了扩散Transformer自回归建模(DiTAR),一个基于补丁的自回归框架相结合的语言模型与扩散Transformer。这种方法显著增强了自回归模型对连续令牌的有效性,并降低了计算需求。DiTAR采用分治策略生成补丁,其中语言模型处理聚合的补丁嵌入,扩散Transformer随后基于语言模型的输出生成下一个补丁。对于推理,我们建议定义温度作为在反向扩散ODE过程中引入噪声的时间点,以平衡多样性和确定性。我们还在广泛的扩展分析中表明,DiTAR具有出色的可扩展性。在zero-shot语音生成中,DiTAR在鲁棒性、说话者相似性和自然度方面实现了最先进的性能。
摘要:Several recent studies have attempted to autoregressively generate continuousspeech representations without discrete speech tokens by combining diffusionand autoregressive models, yet they often face challenges with excessivecomputational loads or suboptimal outcomes. In this work, we propose DiffusionTransformer Autoregressive Modeling (DiTAR), a patch-based autoregressiveframework combining a language model with a diffusion transformer. Thisapproach significantly enhances the efficacy of autoregressive models forcontinuous tokens and reduces computational demands. DiTAR utilizes adivide-and-conquer strategy for patch generation, where the language modelprocesses aggregated patch embeddings and the diffusion transformersubsequently generates the next patch based on the output of the languagemodel. For inference, we propose defining temperature as the time point ofintroducing noise during the reverse diffusion ODE to balance diversity anddeterminism. We also show in the extensive scaling analysis that DiTAR hassuperb scalability. In zero-shot speech generation, DiTAR achievesstate-of-the-art performance in robustness, speaker similarity, andnaturalness.
标题:用于音频深度伪造检测的SSL模型的全面分层分析
链接:https://arxiv.org/abs/2502.03559
备注:13 pages, 3 figures, 3 tables. Accepted to NAACL Findings 2025
摘要:本文对不同背景下的音频deepfake检测的自监督学习(SSL)模型进行了全面的逐层分析,包括多语言数据集(英语,中文,西班牙语),部分,歌曲和基于场景的deepfake场景。通过系统地评估不同Transformer层的贡献,我们发现了对模型行为和性能的重要见解。我们的研究结果表明,较低的层始终提供最具鉴别力的特征,而较高的层捕获较少的相关信息。值得注意的是,所有模型即使在采用减少的层数时也能实现具有竞争力的等错误率(EER)分数。这表明我们可以通过仅利用几个较低的层来降低计算成本并提高检测深度伪造的推理速度。这项工作增强了我们对deepfake检测中SSL模型的理解,提供了适用于各种语言和上下文环境的宝贵见解。我们的训练模型和代码是公开的:https://github.com/Yaselley/SSL_Layerwise_Deepfake。
摘要:This paper conducts a comprehensive layer-wise analysis of self-supervisedlearning (SSL) models for audio deepfake detection across diverse contexts,including multilingual datasets (English, Chinese, Spanish), partial, song, andscene-based deepfake scenarios. By systematically evaluating the contributionsof different transformer layers, we uncover critical insights into modelbehavior and performance. Our findings reveal that lower layers consistentlyprovide the most discriminative features, while higher layers capture lessrelevant information. Notably, all models achieve competitive equal error rate(EER) scores even when employing a reduced number of layers. This indicatesthat we can reduce computational costs and increase the inference speed ofdetecting deepfakes by utilizing only a few lower layers. This work enhancesour understanding of SSL models in deepfake detection, offering valuableinsights applicable across varied linguistic and contextual settings. Ourtrained models and code are publicly available:https://github.com/Yaselley/SSL_Layerwise_Deepfake.
标题:使用声学语音和特征选择的痴呆症分类
链接:https://arxiv.org/abs/2502.03484
摘要:痴呆症是一组综合征的总称,这些综合征影响认知功能,如记忆,思维,推理和执行日常任务的能力。随着人口老龄化,痴呆症患者的数量正在增加,据估计,每年有超过1000万人患上痴呆症。痴呆症是逐渐发展的,病人越早得到帮助和支持,他们保持功能能力的机会就越大。因此,早期诊断痴呆症非常重要。近年来,基于自然口语的机器学习模型已被开发用于痴呆症的早期诊断。这些方法已被证明是用户友好的,具有成本效益的,可扩展的,并能够提供非常快速的诊断。这项研究利用众所周知的ADReSS挑战数据集对健康对照和阿尔茨海默病患者进行分类。该数据集包含从健康对照组和痴呆症患者收集的以厨房场景为特征的图片描述任务的语音记录。与大多数研究不同,这项研究没有将音频记录分割成活跃的语音片段;相反,从整个录音中提取声学特征。该研究采用岭线性回归,极端最小学习机和线性支持向量机机器学习模型来计算基于模型输出的特征重要性得分。岭模型在Leave-One-Subject-Out交叉验证中表现最好,分类准确率为87.8%。EMLM模型,被证明是有效的交叉验证和一个单独的测试数据集的分类,准确率分别为85.3%和79.2%。与使用相同数据集和声学特征提取进行痴呆症诊断的其他研究相比,该研究的结果名列前茅。
摘要:Dementia is a general term for a group of syndromes that affect cognitivefunctions such as memory, thinking, reasoning, and the ability to perform dailytasks. The number of dementia patients is increasing as the population ages,and it is estimated that over 10 million people develop dementia each year.Dementia progresses gradually, and the sooner a patient receives help andsupport, the better their chances of maintaining their functional abilities.For this reason, early diagnosis of dementia is important. In recent years,machine learning models based on naturally spoken language have been developedfor the early diagnosis of dementia. These methods have proven to beuser-friendly, cost-effective, scalable, and capable of providing extremelyfast diagnoses. This study utilizes the well-known ADReSS challenge dataset forclassifying healthy controls and Alzheimer's patients. The dataset containsspeech recordings from a picture description task featuring a kitchen scene,collected from both healthy controls and dementia patients. Unlike moststudies, this research does not segment the audio recordings into active speechsegments; instead, acoustic features are extracted from entire recordings. Thestudy employs Ridge linear regression, Extreme Minimal Learning Machine, andLinear Support Vector Machine machine learning models to compute featureimportance scores based on model outputs. The Ridge model performed best inLeave-One-Subject-Out cross-validation, achieving a classification accuracy of87.8%. The EMLM model, proved to be effective in both cross-validation and theclassification of a separate test dataset, with accuracies of 85.3% and 79.2%,respectively. The study's results rank among the top compared to other studiesusing the same dataset and acoustic feature extraction for dementia diagnosis.
标题:Llasa:基于Llama的语音合成的扩展训练时间和推理时间计算
链接:https://arxiv.org/abs/2502.04128
摘要:基于文本的大型语言模型(LLM)的最新进展,特别是GPT系列和o1模型,已经证明了扩展训练时间和推理时间计算的有效性。然而,利用LLM的当前最先进的TTS系统通常是多级的,需要单独的模型(例如,LLM之后的扩散模型),这使得在训练或测试期间是否缩放特定模型的决定变得复杂。本文的主要贡献如下:首先,我们研究了语音合成中训练时间和推理时间计算的尺度问题。其次,我们提出了一个简单的框架Llasa语音合成,采用单层矢量量化(VQ)编解码器和一个单一的Transformer架构,完全符合标准LLM,如Llama。我们的实验表明,Llasa的缩放训练时间计算始终提高了合成语音的自然度,并能够生成更复杂和准确的韵律模式。此外,从扩展推理时计算的角度来看,我们在搜索过程中使用语音理解模型作为验证者,发现扩展推理时计算将采样模式转向特定验证者的偏好,从而提高情感表达力、音色一致性和内容准确性。此外,我们还公开了TTS模型(1B,3B,8B)和编解码器模型的检查点和训练代码。
摘要:Recent advances in text-based large language models (LLMs), particularly inthe GPT series and the o1 model, have demonstrated the effectiveness of scalingboth training-time and inference-time compute. However, currentstate-of-the-art TTS systems leveraging LLMs are often multi-stage, requiringseparate models (e.g., diffusion models after LLM), complicating the decisionof whether to scale a particular model during training or testing. This workmakes the following contributions: First, we explore the scaling of train-timeand inference-time compute for speech synthesis. Second, we propose a simpleframework Llasa for speech synthesis that employs a single-layer vectorquantizer (VQ) codec and a single Transformer architecture to fully align withstandard LLMs such as Llama. Our experiments reveal that scaling train-timecompute for Llasa consistently improves the naturalness of synthesized speechand enables the generation of more complex and accurate prosody patterns.Furthermore, from the perspective of scaling inference-time compute, we employspeech understanding models as verifiers during the search, finding thatscaling inference-time compute shifts the sampling modes toward the preferencesof specific verifiers, thereby improving emotional expressiveness, timbreconsistency, and content accuracy. In addition, we released the checkpoint andtraining code for our TTS model (1B, 3B, 8B) and codec model publiclyavailable.
标题:迈向可解释的欺骗语音归因和检测:一种用于特征语音合成器组件的概率方法
链接:https://arxiv.org/abs/2502.04049
备注:Submitted to Computer Speech and Language
摘要:我们提出了一个可解释的概率框架,通过将其分解为概率属性嵌入来表征欺骗语音。与原始的高维对策嵌入,缺乏可解释性,所提出的概率属性嵌入的目的是检测特定的语音合成器组件,通过高层次的属性和它们相应的值表示。我们使用这些概率嵌入与四个分类器后端来解决两个下游任务:欺骗检测和欺骗攻击归因。前者是众所周知的善意欺骗检测任务,而后者则试图识别欺骗话语的源方法(生成器)。我们还使用Shapley值(机器学习中广泛使用的技术)来量化每个任务中每个属性值对决策过程的相对贡献。ASVspoof 2019数据集的结果证明了持续时间和转换建模在欺骗检测中的重要作用;以及波形生成和扬声器建模在欺骗攻击归因中的重要作用。在检测任务中,概率属性嵌入的均衡准确率为99.7美元,等误率为0.22美元,与原始属性嵌入的均衡准确率和等误率(EER)(分别为99.9美元和0.22美元)相当。同样,在归因任务中,我们的嵌入实现了$90.23\%$平衡准确度和$2.07\%$ EER,而原始嵌入的准确度和EER分别为$90.16\%$和$2.11\%$。这些结果表明,所提出的框架是内在的设计和能够实现的原始CM嵌入的性能相媲美的解释。
摘要:We propose an explainable probabilistic framework for characterizing spoofedspeech by decomposing it into probabilistic attribute embeddings. Unlike rawhigh-dimensional countermeasure embeddings, which lack interpretability, theproposed probabilistic attribute embeddings aim to detect specific speechsynthesizer components, represented through high-level attributes and theircorresponding values. We use these probabilistic embeddings with fourclassifier back-ends to address two downstream tasks: spoofing detection andspoofing attack attribution. The former is the well-known bonafide-spoofdetection task, whereas the latter seeks to identify the source method(generator) of a spoofed utterance. We additionally use Shapley values, awidely used technique in machine learning, to quantify the relativecontribution of each attribute value to the decision-making process in eachtask. Results on the ASVspoof2019 dataset demonstrate the substantial role ofduration and conversion modeling in spoofing detection; and waveform generationand speaker modeling in spoofing attack attribution. In the detection task, theprobabilistic attribute embeddings achieve $99.7\%$ balanced accuracy and$0.22\%$ equal error rate (EER), closely matching the performance of rawembeddings ($99.9\%$ balanced accuracy and $0.22\%$ EER). Similarly, in theattribution task, our embeddings achieve $90.23\%$ balanced accuracy and$2.07\%$ EER, compared to $90.16\%$ and $2.11\%$ with raw embeddings. Theseresults demonstrate that the proposed framework is both inherently explainableby design and capable of achieving performance comparable to raw CM embeddings.
标题:DiTAR:语音生成的扩散Transformer自回归建模
链接:https://arxiv.org/abs/2502.03930
备注:16 pages, 8 figures
摘要:最近的几项研究试图通过结合扩散和自回归模型来自动生成连续语音表示,而不需要离散语音标记,但它们经常面临计算负荷过大或次优结果的挑战。在这项工作中,我们提出了扩散Transformer自回归建模(DiTAR),一个基于补丁的自回归框架相结合的语言模型与扩散Transformer。这种方法显著增强了自回归模型对连续令牌的有效性,并降低了计算需求。DiTAR采用分治策略生成补丁,其中语言模型处理聚合的补丁嵌入,扩散Transformer随后基于语言模型的输出生成下一个补丁。对于推理,我们建议定义温度作为在反向扩散ODE过程中引入噪声的时间点,以平衡多样性和确定性。我们还在广泛的扩展分析中表明,DiTAR具有出色的可扩展性。在zero-shot语音生成中,DiTAR在鲁棒性、说话人相似性和自然度方面实现了最先进的性能。
摘要:Several recent studies have attempted to autoregressively generate continuousspeech representations without discrete speech tokens by combining diffusionand autoregressive models, yet they often face challenges with excessivecomputational loads or suboptimal outcomes. In this work, we propose DiffusionTransformer Autoregressive Modeling (DiTAR), a patch-based autoregressiveframework combining a language model with a diffusion transformer. Thisapproach significantly enhances the efficacy of autoregressive models forcontinuous tokens and reduces computational demands. DiTAR utilizes adivide-and-conquer strategy for patch generation, where the language modelprocesses aggregated patch embeddings and the diffusion transformersubsequently generates the next patch based on the output of the languagemodel. For inference, we propose defining temperature as the time point ofintroducing noise during the reverse diffusion ODE to balance diversity anddeterminism. We also show in the extensive scaling analysis that DiTAR hassuperb scalability. In zero-shot speech generation, DiTAR achievesstate-of-the-art performance in robustness, speaker similarity, andnaturalness.
标题:基于独立分量提取的盲Capon Beamformer:单参数算法,
链接:https://arxiv.org/abs/2502.03871
备注:None
摘要:我们考虑了线性传感器阵列的相移混合模型在盲源提取的背景下。我们推导出盲Capon波束形成器,该波束形成器寻求输出独立于混合信号中的其他信号的方向。该算法是基于独立分量提取和施加正交约束,由于它优化只有一个实值参数相关的到达角。推导了平均干扰信号比的Cram\'er-Rao下界。该算法和界相比,传统的盲和波达方向估计+波束形成方法,显示在提取精度方面的改进。最后给出了在低混响室内的频域说话人提取中的应用。
摘要:We consider a phase-shift mixing model for linear sensor arrays in thecontext of blind source extraction. We derive a blind Capon beamformer thatseeks the direction where the output is independent of the other signals in themixture. The algorithm is based on Independent Component Extraction and imposesan orthogonal constraint, thanks to which it optimizes only one real-valuedparameter related to the angle of arrival. The Cram\'er-Rao lower bound for themean interference-to-signal ratio is derived. The algorithm and the bound arecompared with conventional blind and direction-of-arrivalestimation+beamforming methods, showing improvements in terms of extractionaccuracy. An application is demonstrated in frequency-domain speaker extractionin a low-reverberation room.
标题:用于音频深度伪造检测的SSL模型的全面分层分析
链接:https://arxiv.org/abs/2502.03559
备注:13 pages, 3 figures, 3 tables. Accepted to NAACL Findings 2025
摘要:本文对不同背景下的音频deepfake检测的自监督学习(SSL)模型进行了全面的逐层分析,包括多语言数据集(英语,中文,西班牙语),部分,歌曲和基于场景的deepfake场景。通过系统地评估不同Transformer层的贡献,我们发现了对模型行为和性能的重要见解。我们的研究结果表明,较低的层始终提供最具鉴别力的特征,而较高的层捕获较少的相关信息。值得注意的是,所有模型即使在采用减少的层数时也能实现具有竞争力的等错误率(EER)分数。这表明我们可以通过仅利用几个较低的层来降低计算成本并提高检测深度伪造的推理速度。这项工作增强了我们对deepfake检测中SSL模型的理解,提供了适用于各种语言和上下文环境的有价值的见解。我们的训练模型和代码是公开的:https://github.com/Yaselley/SSL_Layerwise_Deepfake。
摘要:This paper conducts a comprehensive layer-wise analysis of self-supervisedlearning (SSL) models for audio deepfake detection across diverse contexts,including multilingual datasets (English, Chinese, Spanish), partial, song, andscene-based deepfake scenarios. By systematically evaluating the contributionsof different transformer layers, we uncover critical insights into modelbehavior and performance. Our findings reveal that lower layers consistentlyprovide the most discriminative features, while higher layers capture lessrelevant information. Notably, all models achieve competitive equal error rate(EER) scores even when employing a reduced number of layers. This indicatesthat we can reduce computational costs and increase the inference speed ofdetecting deepfakes by utilizing only a few lower layers. This work enhancesour understanding of SSL models in deepfake detection, offering valuableinsights applicable across varied linguistic and contextual settings. Ourtrained models and code are publicly available:https://github.com/Yaselley/SSL_Layerwise_Deepfake.
标题:使用声学语音和特征选择的痴呆症分类
链接:https://arxiv.org/abs/2502.03484
摘要:痴呆症是一组综合征的总称,这些综合征影响认知功能,如记忆,思维,推理和执行日常任务的能力。随着人口老龄化,痴呆症患者的数量正在增加,据估计,每年有超过1000万人患上痴呆症。痴呆症是逐渐发展的,病人越早得到帮助和支持,他们保持功能能力的机会就越大。因此,早期诊断痴呆症非常重要。近年来,基于自然口语的机器学习模型已被开发用于痴呆症的早期诊断。这些方法已被证明是用户友好的,具有成本效益的,可扩展的,并能够提供非常快速的诊断。这项研究利用众所周知的ADReSS挑战数据集对健康对照和阿尔茨海默病患者进行分类。该数据集包含从健康对照组和痴呆症患者收集的以厨房场景为特征的图片描述任务的语音记录。与大多数研究不同,这项研究没有将音频记录分割成活跃的语音片段;相反,从整个录音中提取声学特征。该研究采用岭线性回归,极端最小学习机和线性支持向量机机器学习模型来计算基于模型输出的特征重要性得分。岭模型在Leave-One-Subject-Out交叉验证中表现最好,分类准确率为87.8%。EMLM模型,被证明是有效的交叉验证和一个单独的测试数据集的分类,准确率分别为85.3%和79.2%。与使用相同数据集和声学特征提取进行痴呆症诊断的其他研究相比,该研究的结果名列前茅。
摘要:Dementia is a general term for a group of syndromes that affect cognitivefunctions such as memory, thinking, reasoning, and the ability to perform dailytasks. The number of dementia patients is increasing as the population ages,and it is estimated that over 10 million people develop dementia each year.Dementia progresses gradually, and the sooner a patient receives help andsupport, the better their chances of maintaining their functional abilities.For this reason, early diagnosis of dementia is important. In recent years,machine learning models based on naturally spoken language have been developedfor the early diagnosis of dementia. These methods have proven to beuser-friendly, cost-effective, scalable, and capable of providing extremelyfast diagnoses. This study utilizes the well-known ADReSS challenge dataset forclassifying healthy controls and Alzheimer's patients. The dataset containsspeech recordings from a picture description task featuring a kitchen scene,collected from both healthy controls and dementia patients. Unlike moststudies, this research does not segment the audio recordings into active speechsegments; instead, acoustic features are extracted from entire recordings. Thestudy employs Ridge linear regression, Extreme Minimal Learning Machine, andLinear Support Vector Machine machine learning models to compute featureimportance scores based on model outputs. The Ridge model performed best inLeave-One-Subject-Out cross-validation, achieving a classification accuracy of87.8%. The EMLM model, proved to be effective in both cross-validation and theclassification of a separate test dataset, with accuracies of 85.3% and 79.2%,respectively. The study's results rank among the top compared to other studiesusing the same dataset and acoustic feature extraction for dementia diagnosis.
标题:Ola:通过渐进式情态对齐突破全情态语言模型的前沿
链接:https://arxiv.org/abs/2502.04328
摘要:大型语言模型的最新进展,特别是在GPT-4 o之后,引发了人们对开发能够理解更多模态的全模态模型的兴趣。虽然已经出现了一些开源替代方案,但在性能方面仍明显落后于专门的单一模式模型。在本文中,我们介绍了Ola,一种全模态语言模型,与专业同行相比,它在图像,视频和音频理解方面具有竞争力的性能。Ola的核心设计在于其渐进式模态对齐策略,该策略渐进地扩展了语言模型的支持模态。我们的训练管道从最独特的模态开始:图像和文本,然后使用连接语言和音频知识的语音数据以及连接所有模态的视频数据逐步扩展模型的技能集。渐进式学习管道还使我们能够保持相对较小的跨模态对齐数据,从而使从现有视觉语言模型开发全模态变得简单且成本更低。此外,为了解锁像GPT-4 o这样的高级交互体验,我们进一步设计了一种用于流媒体语音生成的逐行解码解决方案。广泛的实验表明,Ola在所有模态上都超越了现有的开放式全模态LLM,同时与类似尺寸的最先进的专业模型相比,实现了极具竞争力的性能。我们的目标是使Ola成为一个完全开放的全模态理解解决方案,以推动这一新兴领域的未来研究。模型权重、代码和数据在https://github.com/Ola-Omni/Ola上开源。
摘要:Recent advances in large language models, particularly following GPT-4o, havesparked increasing interest in developing omni-modal models capable ofunderstanding more modalities. While some open-source alternatives haveemerged, there is still a notable lag behind specialized single-modality modelsin performance. In this paper, we present Ola, an Omni-modal language modelthat achieves competitive performance across image, video, and audiounderstanding compared to specialized counterparts. The core design of Ola liesin its progressive modality alignment strategy that extends the supportingmodality of the language model progressively. Our training pipeline begins withthe most distinct modalities: image and text, then gradually expands the skillsets of the model using speech data that connects language and audio knowledge,and video data that connects all modalities. The progressive learning pipelinealso enables us to maintain a relatively small size of the cross-modalalignment data, making developing omni-modal from existing vision-languagemodels easy and less costly. Moreover, to unlock an advanced interactiveexperience like GPT-4o, we further design a sentence-wise decoding solution forstreaming speech generation. Extensive experiments demonstrate that Olasurpasses existing open omni-modal LLMs across all modalities while achievinghighly competitive performance compared to state-of-the-art specialized modelsof similar sizes. We aim to make Ola a fully open omni-modal understandingsolution to advance future research in this emerging field. Model weights,code, and data are open-sourced at https://github.com/Ola-Omni/Ola.
标题:XAttnMark:通过交叉注意力学习稳健的音频水印
链接:https://arxiv.org/abs/2502.04230
备注:24 pages, 10 figures
摘要:生成式音频合成和编辑技术的迅速普及引发了人们对版权侵权、数据来源以及通过deepfake音频传播错误信息的严重担忧。水印通过将难以察觉、可识别且可追踪的标记嵌入音频内容中,提供了一种主动的解决方案。虽然最近的基于神经网络的水印方法,如WavMark和AudioSeal,提高了鲁棒性和质量,但它们很难同时实现鲁棒检测和准确属性。本文介绍了交叉注意鲁棒音频水印(XAttnMark),它通过利用生成器和检测器之间的部分参数共享,用于有效消息检索的交叉注意机制和用于改进消息分发的时间调节模块来弥合这一差距。此外,我们提出了一种心理声学对齐的时间频率掩蔽损失,可以捕获细粒度的听觉掩蔽效果,从而增强水印的不可感知性。我们的方法在检测和归因方面都实现了最先进的性能,对各种音频转换表现出卓越的鲁棒性,包括具有强大编辑强度的挑战性生成编辑。该项目的网页可在https://liuyixin-louis.github.io/xattnmark/上查阅。
摘要:The rapid proliferation of generative audio synthesis and editingtechnologies has raised significant concerns about copyright infringement, dataprovenance, and the spread of misinformation through deepfake audio.Watermarking offers a proactive solution by embedding imperceptible,identifiable, and traceable marks into audio content. While recent neuralnetwork-based watermarking methods like WavMark and AudioSeal have improvedrobustness and quality, they struggle to achieve both robust detection andaccurate attribution simultaneously. This paper introduces Cross-AttentionRobust Audio Watermark (XAttnMark), which bridges this gap by leveragingpartial parameter sharing between the generator and the detector, across-attention mechanism for efficient message retrieval, and a temporalconditioning module for improved message distribution. Additionally, we proposea psychoacoustic-aligned temporal-frequency masking loss that capturesfine-grained auditory masking effects, enhancing watermark imperceptibility.Our approach achieves state-of-the-art performance in both detection andattribution, demonstrating superior robustness against a wide range of audiotransformations, including challenging generative editing with strong editingstrength. The project webpage is available athttps://liuyixin-louis.github.io/xattnmark/.
标题:数据驱动的双麦克风现场声吸收测量方法
链接:https://arxiv.org/abs/2502.04143
备注:41 pages, 8 figures
摘要:本文提出了一种基于数据驱动的方法,利用神经网络和双传声器测量方法对有限多孔板的吸声系数进行了估计。1D卷积网络从在两个麦克风位置处测量的声压之间的复值传递函数预测吸声系数。该网络的训练和验证的边界元模型使用Delany-Bazley-Miki模型生成的数值数据,证明了各种数值样本的准确预测。该方法进行了实验验证与折流矩形样品的纤维材料,其中样品尺寸和源高度是不同的。结果表明,神经网络提供了可靠的预测,使用传统的双麦克风的方法,如果样品是无限的多孔材料的现场吸声的可能性。由该网络得到的垂直入射吸声系数与理论计算值和阻抗管中的吸声系数进行了比较。所提出的方法具有很好的前景,估计吸声材料安装后,在现实的操作条件下的吸声系数。
摘要:This work presents a data-driven approach to estimating the sound absorptioncoefficient of an infinite porous slab using a neural network and atwo-microphone measurement on a finite porous sample. A 1D-convolutionalnetwork predicts the sound absorption coefficient from the complex-valuedtransfer function between the sound pressure measured at the two microphonepositions. The network is trained and validated with numerical data generatedby a boundary element model using the Delany-Bazley-Miki model, demonstratingaccurate predictions for various numerical samples. The method isexperimentally validated with baffled rectangular samples of a fibrousmaterial, where sample size and source height are varied. The results show thatthe neural network offers the possibility to reliably predict the in-situ soundabsorption of a porous material using the traditional two-microphone method asif the sample were infinite. The normal-incidence sound absorption coefficientobtained by the network compares well with that obtained theoretically and inan impedance tube. The proposed method has promising perspectives forestimating the sound absorption coefficient of acoustic materials afterinstallation and in realistic operational conditions.
标题:迈向跨维度和类别模型的统一音乐情感识别
链接:https://arxiv.org/abs/2502.03979
摘要:音乐情感识别(MER)中最重要的挑战之一来自以下事实:情感标签在关于情感表示的数据集之间可以是异构的,包括分类的(例如,快乐、悲伤)与维度标签(例如,效价激发)。在本文中,我们提出了一个统一的多任务学习框架,该框架结合了这两种类型的标签,因此能够在多个数据集上进行训练。该框架使用有效的输入表示,其组合音乐特征(即,键和和弦)和MERT嵌入。此外,知识蒸馏用于将在单个数据集上训练的教师模型的知识转移到学生模型,从而增强其跨多个任务进行概括的能力。为了验证我们提出的框架,我们在各种数据集上进行了广泛的实验,包括MTG-Jamendo,DEAM,PMEmo和PMEMO Music。根据我们的实验结果,包括音乐功能,多任务学习和知识蒸馏显着提高性能。特别是,我们的模型优于最先进的模型,包括来自MTG-Jamendo数据集的MediaEval 2021竞赛的最佳模型。我们的工作通过允许在一个统一的框架中组合分类和维度情感标签,从而实现跨数据集的训练,为MER做出了重大贡献。
摘要:One of the most significant challenges in Music Emotion Recognition (MER)comes from the fact that emotion labels can be heterogeneous across datasetswith regard to the emotion representation, including categorical (e.g., happy,sad) versus dimensional labels (e.g., valence-arousal). In this paper, wepresent a unified multitask learning framework that combines these two types oflabels and is thus able to be trained on multiple datasets. This framework usesan effective input representation that combines musical features (i.e., key andchords) and MERT embeddings. Moreover, knowledge distillation is employed totransfer the knowledge of teacher models trained on individual datasets to astudent model, enhancing its ability to generalize across multiple tasks. Tovalidate our proposed framework, we conducted extensive experiments on avariety of datasets, including MTG-Jamendo, DEAM, PMEmo, and EmoMusic.According to our experimental results, the inclusion of musical features,multitask learning, and knowledge distillation significantly enhancesperformance. In particular, our model outperforms the state-of-the-art models,including the best-performing model from the MediaEval 2021 competition on theMTG-Jamendo dataset. Our work makes a significant contribution to MER byallowing the combination of categorical and dimensional emotion labels in oneunified framework, thus enabling training across datasets.
标题:UniForm:用于音频视频生成的统一扩散Transformer
链接:https://arxiv.org/abs/2502.03897
摘要:作为一种自然的多模态内容,音频视频提供了身临其境的感官体验。因此,音频-视频生成系统具有巨大的潜力。然而,现有的基于扩散的研究主要采用相对独立的模块来生成每个模态,缺乏对共享权重生成模块的探索。这种方法可能未充分利用音频和视觉模态之间的内在相关性,可能导致次优的生成质量。为了解决这个问题,我们提出了统一的形式,一个统一的扩散Transformer,旨在提高跨模态的一致性。通过连接听觉和视觉信息,UniForm学会在统一的潜在空间内同时生成音频和视频,从而促进创建高质量和对齐良好的视听对。大量的实验表明,我们的方法在联合音视频生成,音频引导的视频生成,视频引导的音频生成任务的优越性能。我们的演示可在https://uniform-t2av.github.io/上获得。
摘要:As a natural multimodal content, audible video delivers an immersive sensoryexperience. Consequently, audio-video generation systems have substantialpotential. However, existing diffusion-based studies mainly employ relativelyindependent modules for generating each modality, which lack exploration ofshared-weight generative modules. This approach may under-use the intrinsiccorrelations between audio and visual modalities, potentially resulting insub-optimal generation quality. To address this, we propose UniForm, a unifieddiffusion transformer designed to enhance cross-modal consistency. Byconcatenating auditory and visual information, UniForm learns to generate audioand video simultaneously within a unified latent space, facilitating thecreation of high-quality and well-aligned audio-visual pairs. Extensiveexperiments demonstrate the superior performance of our method in jointaudio-video generation, audio-guided video generation, and video-guided audiogeneration tasks. Our demos are available at https://uniform-t2av.github.io/.
