本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
链接:https://arxiv.org/abs/2503.22605
摘要:说话人合成已成为计算机图形学和多媒体领域的一个重要研究领域,然而大多数现有的方法往往难以平衡生成质量和计算效率。在本文中,我们提出了一种新的方法,利用音频分解平面(音频平面)为基础的高斯飞溅高品质和实时说话头生成。为了对动态讲话头部进行建模,需要4D体积表示。然而,直接存储密集的4D网格是不切实际的,这是由于高成本和缺乏较长持续时间的可扩展性。我们克服了这一挑战与建议的音频平面,其中的4D体积表示被分解成音频独立的空间平面和音频相关的平面。这为说话头部提供了紧凑且可解释的特征表示,促进了更精确的音频感知空间编码和增强的音频驱动的唇部动态建模。为了进一步改善语音动态,我们开发了一种动态飞溅方法,帮助网络更有效地专注于对嘴部区域的动态建模。大量的实验表明,通过将这些创新与强大的高斯飞溅相结合,我们的方法能够实时合成高度逼真的谈话视频,同时确保精确的音频-嘴唇同步。合成结果可在https://sstzal.github.io/Audio-Plane/上获得。
摘要:Talking head synthesis has become a key research area in computer graphics and multimedia, yet most existing methods often struggle to balance generation quality with computational efficiency. In this paper, we present a novel approach that leverages an Audio Factorization Plane (Audio-Plane) based Gaussian Splatting for high-quality and real-time talking head generation. For modeling a dynamic talking head, 4D volume representation is needed. However, directly storing a dense 4D grid is impractical due to the high cost and lack of scalability for longer durations. We overcome this challenge with the proposed Audio-Plane, where the 4D volume representation is decomposed into audio-independent space planes and audio-dependent planes. This provides a compact and interpretable feature representation for talking head, facilitating more precise audio-aware spatial encoding and enhanced audio-driven lip dynamic modeling. To further improve speech dynamics, we develop a dynamic splatting method that helps the network more effectively focus on modeling the dynamics of the mouth region. Extensive experiments demonstrate that by integrating these innovations with the powerful Gaussian Splatting, our method is capable of synthesizing highly realistic talking videos in real time while ensuring precise audio-lip synchronization. Synthesized results are available in https://sstzal.github.io/Audio-Plane/.
【2】 Cross-Technology Generalization in Synthesized Speech Detection: Evaluating AST Models with Modern Voice Generators
标题: 合成语音检测中的跨技术概括:使用现代语音发生器评估AST模型
链接:https://arxiv.org/abs/2503.22503
备注:10 pages, 5 figures
摘要:本文评估了音频频谱图Transformer(AST)的合成语音检测架构,重点是在整个现代语音生成技术的推广。使用差异化的增强策略,该模型在对ElevenLabs、NotebookLM和Minimax AI语音生成器进行测试时,总体EER达到0.91%。值得注意的是,在仅使用来自单一技术的102个样本进行训练后,该模型表现出强大的跨技术泛化能力,在完全不可见的语音生成器上实现了3.3%的EER。这项工作建立了快速适应新兴合成技术的基准,并提供了证据表明,基于transformer的架构可以识别不同神经语音合成方法中的常见伪影,从而有助于更强大的语音验证系统。
摘要:This paper evaluates the Audio Spectrogram Transformer (AST) architecture for synthesized speech detection, with focus on generalization across modern voice generation technologies. Using differentiated augmentation strategies, the model achieves 0.91% EER overall when tested against ElevenLabs, NotebookLM, and Minimax AI voice generators. Notably, after training with only 102 samples from a single technology, the model demonstrates strong cross-technology generalization, achieving 3.3% EER on completely unseen voice generators. This work establishes benchmarks for rapid adaptation to emerging synthesis technologies and provides evidence that transformer-based architectures can identify common artifacts across different neural voice synthesis methods, contributing to more robust speech verification systems.
【3】 DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
标题: DeepAudio-V1:面向多模式多阶段端到端视频到语音和音频生成
链接:https://arxiv.org/abs/2503.22265
备注:11 pages, 5 figures
摘要:目前,高质量的同步音频是使用各种多模态联合学习框架,利用视频和可选的文本输入合成的。在视频到音频基准测试中,视频到音频质量、语义对齐和视听同步被有效地实现。然而,在现实场景中,语音和音频往往同时共存于视频中,并且在给定视频和文本条件下同步语音和音频的端到端生成还没有得到很好的研究。因此,我们提出了一个端到端的多模态生成框架,同时产生语音和音频的基础上的视频和文本的条件。此外,用于从视频生成语音的视频到音频(V2 A)模型的优点仍然不清楚。所提出的框架DeepAudio由视频到音频(V2 A)模块、文本到语音(TTS)模块和动态混合模态融合(MoF)模块组成。在评估中,提出的端到端框架实现了最先进的性能上的视频音频基准,视频语音基准,和文本语音基准。详细地说,我们的框架在视频-音频和文本-语音基准测试中与最先进的模型进行了比较,并在视频-语音基准测试中超过了最先进的模型,WER为16.57%至3.15%(+80.99%),SPK-SIM 78.30%至89.38%(+14.15%),EMO-SIM 66.24%至75.56%(+14.07%),MCD 8.59至7.98(+7.10%),MCD SL 11.05至9.40(+14.93%)。
摘要:Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio benchmarks, video-to-audio quality, semantic alignment, and audio-visual synchronization are effectively achieved. However, in real-world scenarios, speech and audio often coexist in videos simultaneously, and the end-to-end generation of synchronous speech and audio given video and text conditions are not well studied. Therefore, we propose an end-to-end multi-modal generation framework that simultaneously produces speech and audio based on video and text conditions. Furthermore, the advantages of video-to-audio (V2A) models for generating speech from videos remain unclear. The proposed framework, DeepAudio, consists of a video-to-audio (V2A) module, a text-to-speech (TTS) module, and a dynamic mixture of modality fusion (MoF) module. In the evaluation, the proposed end-to-end framework achieves state-of-the-art performance on the video-audio benchmark, video-speech benchmark, and text-speech benchmark. In detail, our framework achieves comparable results in the comparison with state-of-the-art models for the video-audio and text-speech benchmarks, and surpassing state-of-the-art models in the video-speech benchmark, with WER 16.57% to 3.15% (+80.99%), SPK-SIM 78.30% to 89.38% (+14.15%), EMO-SIM 66.24% to 75.56% (+14.07%), MCD 8.59 to 7.98 (+7.10%), MCD SL 11.05 to 9.40 (+14.93%) across a variety of dubbing settings.
【4】 DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos
标题: DeepSound-V1:开始逐步思考从视频生成音频
链接:https://arxiv.org/abs/2503.22208
备注:11 pages, 6 figures
摘要:目前,高质量的同步音频是使用各种多模态联合学习框架从视频和可选的文本输入合成的。然而,视觉域和生成的音频域之间的精确对准仍然远远不能令人满意。一个关键因素是在开源视频音频和文本音频基准测试中缺乏足够的时间和语义对齐注释。因此,我们提出了一个从视频生成音频的框架,利用多模态大型语言模型(MLLM)的内部思想链(CoT)来实现逐步推理,而无需额外的注释。此外,一个相应的多模态推理数据集的构建,以促进学习的初始推理的音频生成。在实验中,我们证明了所提出的框架的有效性,减少错位(画外音)生成的音频和实现有竞争力的性能相比,各种国家的最先进的模型。评估结果表明,该方法优于国家的最先进的方法在多个指标。具体而言,F DP aSST指标降低高达10.07%,F DP AN N s指标降低高达11.62%,F DV GG指标降低高达38.61%。此外,IS指标提高了4.95%,IB评分指标提高了6.39%,DeSync指标降低了0.89%。
摘要:Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far from satisfactory. One key factor is the lack of sufficient temporal and semantic alignment annotations in open-source video-audio and text-audio benchmarks. Therefore, we propose a framework for audio generation from videos, leveraging the internal chain-of-thought (CoT) of a multi-modal large language model (MLLM) to enable step-by-step reasoning without requiring additional annotations. Additionally, a corresponding multi-modal reasoning dataset is constructed to facilitate the learning of initial reasoning in audio generation. In the experiments, we demonstrate the effectiveness of the proposed framework in reducing misalignment (voice-over) in generated audio and achieving competitive performance compared to various state-of-the-art models. The evaluation results show that the proposed method outperforms state-of-the-art approaches across multiple metrics. Specifically, the F DP aSST indicator is reduced by up to 10.07%, the F DP AN N s indicator by up to 11.62%, and the F DV GG indicator by up to 38.61%. Furthermore, the IS indicator improves by up to 4.95%, the IB-score indicator increases by up to 6.39%, and the DeSync indicator is reduced by up to 0.89%.
【5】 Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization
标题: 通过多步CoT类引导和组合偏好优化提高流匹配V2 A模型的生成质量
链接:https://arxiv.org/abs/2503.22200
备注:10 pages, 4 figures
摘要:从视频和文本提示创建高质量的声音效果需要在语义和时间上精确调整视觉和音频域,以及专业音频生成的分步指导。然而,目前最先进的视频引导音频生成模型通常无法为一般和专业用例生成高质量音频。为了应对这一挑战,我们引入了一个多阶段,多模式,端到端的生成框架,具有类似于链的指导学习,称为执行链(CoP)。首先,我们采用基于变压器的网络架构,旨在实现CoP指导,从而生成通用和专业音频。其次,我们实施了一个多阶段的培训框架,遵循分步指导,以确保产生高质量的声音效果。第三,我们开发了一个由视频引导的CoP多模态数据集,以支持逐步生成声音效果。评估结果突出了所提出的多阶段CoP生成框架与各种数据集上的最先进模型相比的优势,FAD为0.79至0.74(+6.33%),CLIP 16.12至17.70(+9.80%)在VGGSound上,SI-SDR 1.98dB至3.35dB(+69.19%),PianoYT-2 h的MOS 2.94至3.49(+18.71%),Piano-10 h的SI-SDR 2.22dB至3.21dB(+44.59%),MOS 3.07至3.42(+11.40%)。
摘要:Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-step guidance for professional audio generation. However, current state-of-the-art video-guided audio generation models often fall short of producing high-quality audio for both general and specialized use cases. To address this challenge, we introduce a multi-stage, multi-modal, end-to-end generative framework with Chain-of-Thought-like (CoT-like) guidance learning, termed Chain-of-Perform (CoP). First, we employ a transformer-based network architecture designed to achieve CoP guidance, enabling the generation of both general and professional audio. Second, we implement a multi-stage training framework that follows step-by-step guidance to ensure the generation of high-quality sound effects. Third, we develop a CoP multi-modal dataset, guided by video, to support step-by-step sound effects generation. Evaluation results highlight the advantages of the proposed multi-stage CoP generative framework compared to the state-of-the-art models on a variety of datasets, with FAD 0.79 to 0.74 (+6.33%), CLIP 16.12 to 17.70 (+9.80%) on VGGSound, SI-SDR 1.98dB to 3.35dB (+69.19%), MOS 2.94 to 3.49(+18.71%) on PianoYT-2h, and SI-SDR 2.22dB to 3.21dB (+44.59%), MOS 3.07 to 3.42 (+11.40%) on Piano-10h.
【6】 Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
标题: 通过负条件反射潜在扩散模型增强舞蹈到音乐的生成
链接:https://arxiv.org/abs/2503.22138
摘要:条件扩散模型由于其令人印象深刻的跨模态合成结果而受到越来越多的关注,其中条件输入和生成的输出之间的强对齐可以通过训练具有交叉注意机制的时间条件U-Net来实现。在本文中,我们专注于生成与给定的舞蹈视频的节奏的视觉线索同步的音乐的问题。考虑到双向指导更有利于训练扩散模型,我们建议通过采用正节奏信息和负节奏信息(PN-扩散)作为条件来提高生成的音乐的质量及其与舞蹈视频的同步,其中设计了双扩散和反向过程。具体来说,为了训练一个顺序的多模态U-网络结构,PN-扩散包括一个噪声预测目标的积极条件和一个额外的噪声预测目标的消极条件。为了准确地定义和选择积极和消极的条件反射,我们巧妙地利用舞蹈视频中的时间相关性,通过分别向前和向后播放来捕捉积极和消极的节奏线索。通过在舞蹈音乐节拍对齐和生成音乐的质量方面对输入输出对应性进行主观和客观评估,在AIST++和TikTok舞蹈视频数据集上的实验结果表明,我们的模型优于SOTA舞蹈音乐生成模型。
摘要:Conditional diffusion models have gained increasing attention since their impressive results for cross-modal synthesis, where the strong alignment between conditioning input and generated output can be achieved by training a time-conditioned U-Net augmented with cross-attention mechanism. In this paper, we focus on the problem of generating music synchronized with rhythmic visual cues of the given dance video. Considering that bi-directional guidance is more beneficial for training a diffusion model, we propose to enhance the quality of generated music and its synchronization with dance videos by adopting both positive rhythmic information and negative ones (PN-Diffusion) as conditions, where a dual diffusion and reverse processes is devised. Specifically, to train a sequential multi-modal U-Net structure, PN-Diffusion consists of a noise prediction objective for positive conditioning and an additional noise prediction objective for negative conditioning. To accurately define and select both positive and negative conditioning, we ingeniously utilize temporal correlations in dance videos, capturing positive and negative rhythmic cues by playing them forward and backward, respectively. Through subjective and objective evaluations of input-output correspondence in terms of dance-music beat alignment and the quality of generated music, experimental results on the AIST++ and TikTok dance video datasets demonstrate that our model outperforms SOTA dance-to-music generation models.
【7】 Tune It Up: Music Genre Transfer and Prediction
标题: 调音:音乐类型转移和预测
链接:https://arxiv.org/abs/2503.22008
摘要:深度生成模型已用于图像的风格转换任务。在这项研究中,我们调整和改进CycleGAN模型来执行爵士乐和经典流派的音乐风格转移。通过这样做,我们的目标是轻松生成新的歌曲,涵盖不同音乐流派的音乐,并减少这些过程中所需的安排。我们训练并使用音乐流派分类器来评估迁移模型的性能。为此,我们使用多层感知器算法获得了87.7%的准确率。为了改善我们的风格迁移基线,我们在模型中添加了辅助鉴别器和三重损失。根据我们的实验,我们获得了最好的准确率为69.4%,在爵士乐经典的任务和39.3%,在经典的爵士乐任务与我们开发的体裁分类。我们还进行了主观实验,结果表明,我们的传输模型的整体性能是好的,它设法保存输入的旋律上传输的输出。我们的代码可以在https://github.com/ fidansamet/tune-it-up上找到
摘要:Deep generative models have been used in style transfer tasks for images. In this study, we adapt and improve CycleGAN model to perform music style transfer on Jazz and Classic genres. By doing so, we aim to easily generate new songs, cover music to different music genres and reduce the arrangements needed in those processes. We train and use music genre classifier to assess the performance of the transfer models. To that end, we obtain 87.7% accuracy with Multi-layer Perceptron algorithm. To improve our style transfer baseline, we add auxiliary discriminators and triplet loss to our model. According to our experiments, we obtain the best accuracies as 69.4% in Jazz to Classic task and 39.3% in Classic to Jazz task with our developed genre classifier. We also run a subjective experiment and results of it show that the overall performance of our transfer model is good and it manages to conserve melody of inputs on the transferred outputs. Our code is available at https://github.com/ fidansamet/tune-it-up
【8】 Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
标题: 分层标签传播:AudioSet标签的依赖于模型大小的性能提升器
链接:https://arxiv.org/abs/2503.21826
备注:None
摘要:AudioSet是音频标记中最常用和最大的数据集之一,包含大约200万个音频样本,这些样本被手动标记为527个事件类别,并组织成一个本体。然而,注释包含不一致性,特别是根据本体应该被标记为积极的类别经常被错误地标记为消极的。为了解决这个问题,我们应用分层标签传播(HLP),它将标签传播到本体层次结构,导致每个音频片段的正标签平均从1.98增加到2.39,并影响527个类中的109个。我们的研究结果表明,HLP在各种模型架构中提供了性能优势,包括卷积神经网络(PANN的CNN6和ConvNeXT)和Transformers(PaSST),较小的模型显示出更多的改进。最后,在另一个广泛使用的数据集FSD 50K上,在AudioSet上使用HLP训练的模型始终优于在没有HLP的情况下训练的模型。我们的源代码将在GitHub上提供。
摘要:AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology. However, the annotations contain inconsistencies, particularly where categories that should be labeled as positive according to the ontology are frequently mislabeled as negative. To address this issue, we apply Hierarchical Label Propagation (HLP), which propagates labels up the ontology hierarchy, resulting in a mean increase in positive labels per audio clip from 1.98 to 2.39 and affecting 109 out of the 527 classes. Our results demonstrate that HLP provides performance benefits across various model architectures, including convolutional neural networks (PANN's CNN6 and ConvNeXT) and transformers (PaSST), with smaller models showing more improvements. Finally, on FSD50K, another widely used dataset, models trained on AudioSet with HLP consistently outperformed those trained without HLP. Our source code will be made available on GitHub.
【9】 Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
标题: 制造一些噪音:使用声音令牌实现LLM音频推理和生成
链接:https://arxiv.org/abs/2503.22275
备注:5 pages, 2 figures, Accepted at ICASSP 2025
摘要:由于音频的连续性和由此产生的高采样率,将音频理解和生成集成到大型语言模型(LLM)中仍然具有挑战性。在这里,我们介绍了一种新的方法,将变分量化与条件流匹配相结合,将音频转换为0.23kpbs的超低比特率离散令牌,从而实现与LLM中的文本令牌的无缝集成。我们使用低秩自适应(LoRA)对预训练的基于文本的LLM进行了微调,以评估其在实现真正的多模态功能方面的有效性,即,音频理解和生成。我们的标记器在具有不同声学事件的各种数据集上优于传统的VQ-VAE。尽管通过音频标记化丢失了大量的细粒度细节,但我们用离散标记训练的多模态LLM在音频理解方面取得了具有竞争力的结果,尽管音频生成很差。我们的研究结果强调了需要更大,更多样化的数据集和改进的评估指标,以提高多模态LLM性能。
摘要:Integrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. Here, we introduce a novel approach that combines Variational Quantization with Conditional Flow Matching to convert audio into ultra-low bitrate discrete tokens of 0.23kpbs, allowing for seamless integration with text tokens in LLMs. We fine-tuned a pretrained text-based LLM using Low-Rank Adaptation (LoRA) to assess its effectiveness in achieving true multimodal capabilities, i.e., audio comprehension and generation. Our tokenizer outperforms a traditional VQ-VAE across various datasets with diverse acoustic events. Despite the substantial loss of fine-grained details through audio tokenization, our multimodal LLM trained with discrete tokens achieves competitive results in audio comprehension with state-of-the-art methods, though audio generation is poor. Our results highlight the need for larger, more diverse datasets and improved evaluation metrics to advance multimodal LLM performance.
【10】 Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
标题: 声音场景空间语义分割的基线系统和评估表
链接:https://arxiv.org/abs/2503.22088
备注:5 pages
摘要:沉浸式通信已经取得了重大进展,特别是随着沉浸式语音和音频服务编解码器的发布。为了进一步实现这一目标,DCASE 2025挑战赛最近引入了一项声音场景空间语义分割任务(S5),重点是检测和分离空间声音场景中的声音事件。在本文中,我们将探索解决S5任务的方法。具体来说,我们提出了基线S5系统,结合音频标记(AT)和标签查询源分离(LSS)模型。我们研究了两种基于ResUNet架构的LSS方法:a)为每个检测到的事件提取单个源,b)同时查询多个源。由于S5中的每个分离的源由其声音事件类标签标识,因此我们提出了新的类感知度量来同时评估声源和标签。一阶立体混响空间音频的实验结果表明,所提出的系统的有效性,并确认的度量的有效性。
摘要:Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introduced a task for spatial semantic segmentation of sound scenes (S5), which focuses on detecting and separating sound events in spatial sound scenes. In this paper, we explore methods for addressing the S5 task. Specifically, we present baseline S5 systems that combine audio tagging (AT) and label-queried source separation (LSS) models. We investigate two LSS approaches based on the ResUNet architecture: a) extracting a single source for each detected event and b) querying multiple sources concurrently. Since each separated source in S5 is identified by its sound event class label, we propose new class-aware metrics to evaluate both the sound sources and labels simultaneously. Experimental results on first-order ambisonics spatial audio demonstrate the effectiveness of the proposed systems and confirm the efficacy of the metrics.
【11】 Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling for Barrier-free Communication
标题: 伸出援手:通过MLLM驱动的手部建模进行半免训练的提示语音识别,实现无障碍沟通
链接:https://arxiv.org/abs/2503.21785
摘要:提示语音(CS)是一种创新的视觉交流系统,将唇读与手势编码相结合,旨在提高听力障碍者的有效沟通。自动CS识别(ACSR)是指人工智能驱动的自动识别CS中的手势和嘴唇运动并将其转换为文本的过程。然而,以前的工作往往依赖于复杂的融合模块和训练技术。此外,由于CS中的数据量有限,手部特征的提取以及识别建模一直处于低水平,这大大限制了ACSR的有效性。为了解决这个问题,我们创新性地探索了多模态大语言模型(MLLM)在CS中识别手形和位置的能力。更准确地说,我们提出了一个新的半培训免费范式ACSR,命名为STF-ACSR。该方法通过中文CS提示模块(CCSPM)实现手部动作的zero-shot识别,CCSPM提供了免训练的关键帧过滤和基于MLLM的定制提示工程。然后,它集成到唇读模型使用最低限度的融合模块(MFM)的识别结果,有效地实现了卓越的识别结果。此外,特别是对于这项研究,我们通过记录来自8名听力障碍患者的额外数据,补充了现有的6名听力正常CS患者的数据集,从而形成了一个新的混合数据集。大量的实验表明,STF-ACSR显着优于以往的方法对正常和听力受损的数据。实现和检查点可在https://github.com/DennisHgj/STF_ACSR上获得。
摘要:Cued Speech (CS) is an innovative visual communication system that integrates lip-reading with hand coding, designed to enhance effective communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) refers to the AI-driven process of automatically recognizing hand gestures and lip movements in CS, converting them into text. However, previous work often relies on complex fusion modules and training techniques. Additionally, due to the limited amount of data in CS, the extraction of hand features, as well as recognition modeling, has consistently been subpar, significantly limiting the effectiveness of ACSR. To address this issue, we have innovatively explored the capabilities of Multimodal large language models (MLLMs) in recognizing hand shapes and positions in CS. More precisely, we propose a new Semi Training-Free paradigm for ACSR, named STF-ACSR. This approach leverages zero-shot recognition of hand movements through the Chinese CS Prompt Module (CCSPM), which equipped a training-free keyframe filtering and customized prompt engineering based on MLLM. It then integrates the recognition results into the lip-reading model using a Minimalist Fusion Module (MFM), effectively achieving superior recognition results. Furthermore, specifically for this study, we have supplemented the existing dataset of 6 normal hearing CS cuers by recording additional data from 8 cuers with hearing impairments, resulting in a new mixed dataset. Extensive experiments have demonstrated that STF-ACSR significantly outperforms previous methods on both normal and hearing-impaired data. Implementation and checkpoints are available at https://github.com/DennisHgj/STF_ACSR.
【1】 Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
标题: 制造一些噪音:使用声音令牌实现LLM音频推理和生成
链接:https://arxiv.org/abs/2503.22275
备注:5 pages, 2 figures, Accepted at ICASSP 2025
摘要:由于音频的连续性和由此产生的高采样率,将音频理解和生成集成到大型语言模型(LLM)中仍然具有挑战性。在这里,我们介绍了一种新的方法,将变分量化与条件流匹配相结合,将音频转换为0.23kpbs的超低比特率离散令牌,从而实现与LLM中的文本令牌的无缝集成。我们使用低秩自适应(LoRA)对预训练的基于文本的LLM进行了微调,以评估其在实现真正的多模态功能方面的有效性,即,音频理解和生成。我们的标记器在具有不同声学事件的各种数据集上优于传统的VQ-VAE。尽管通过音频标记化丢失了大量的细粒度细节,但我们用离散标记训练的多模态LLM在音频理解方面取得了具有竞争力的结果,尽管音频生成很差。我们的研究结果强调了需要更大,更多样化的数据集和改进的评估指标,以提高多模态LLM性能。
摘要:Integrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. Here, we introduce a novel approach that combines Variational Quantization with Conditional Flow Matching to convert audio into ultra-low bitrate discrete tokens of 0.23kpbs, allowing for seamless integration with text tokens in LLMs. We fine-tuned a pretrained text-based LLM using Low-Rank Adaptation (LoRA) to assess its effectiveness in achieving true multimodal capabilities, i.e., audio comprehension and generation. Our tokenizer outperforms a traditional VQ-VAE across various datasets with diverse acoustic events. Despite the substantial loss of fine-grained details through audio tokenization, our multimodal LLM trained with discrete tokens achieves competitive results in audio comprehension with state-of-the-art methods, though audio generation is poor. Our results highlight the need for larger, more diverse datasets and improved evaluation metrics to advance multimodal LLM performance.
【2】 M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
标题: M2 D2:探索CLAP之外的通用音频语言表示
链接:https://arxiv.org/abs/2503.22104
备注:15 pages, 7 figures, 13 tables, under review at an IEEE journal
摘要:对比语言-音频预训练(CLAP)通过在共同特征空间中对齐音频和文本来解决音频-语言任务,例如音频-文本检索。虽然CLAP解决了一般的音频语言任务,但其音频功能在音频任务中并不通用。相比之下,自监督学习(SSL)模型学习在各种音频任务中表现良好的通用音频功能。我们追求可以广泛用于音频应用的表示学习,并假设学习通用音频特征和CLAP特征的方法应该实现我们的目标,我们称之为通用音频语言表示。为了实现我们的假设,我们提出了M2 D2,这是第二代掩码建模二人组(M2 D),它结合了SSL M2 D和CLAP。M2 D2在两阶段训练过程中使用两种模态(音频和文本)学习两种类型的特征。它还在CLAP训练中利用了先进的基于LLM的句子嵌入,以实现强大的语义监督。在第一阶段,M2 D2从M2 D和CLAP学习可概括的音频特征,其中CLAP将特征与基于LLM的精细语义嵌入对齐。在第二阶段,它使用从基于LLM的嵌入中学习的音频特征来学习CLAP特征。通过这些预训练阶段,M2 D2应该增强其音频和CLAP功能的通用性和性能。实验验证了M2 D2实现了有效的通用音频语言表示,突出显示了AudioSet的SOTA微调mAP为49.0,音乐任务中的SOTA性能以及音频语言任务中的顶级性能。
摘要:Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its audio features do not generalize well in audio tasks. In contrast, self-supervised learning (SSL) models learn general-purpose audio features that perform well in diverse audio tasks. We pursue representation learning that can be widely used in audio applications and hypothesize that a method that learns both general audio features and CLAP features should achieve our goal, which we call a general-purpose audio-language representation. To implement our hypothesis, we propose M2D2, a second-generation masked modeling duo (M2D) that combines an SSL M2D and CLAP. M2D2 learns two types of features using two modalities (audio and text) in a two-stage training process. It also utilizes advanced LLM-based sentence embeddings in CLAP training for powerful semantic supervision. In the first stage, M2D2 learns generalizable audio features from M2D and CLAP, where CLAP aligns the features with the fine LLM-based semantic embeddings. In the second stage, it learns CLAP features using the audio features learned from the LLM-based embeddings. Through these pre-training stages, M2D2 should enhance generalizability and performance in its audio and CLAP features. Experiments validated that M2D2 achieves effective general-purpose audio-language representation, highlighted with SOTA fine-tuning mAP of 49.0 for AudioSet, SOTA performance in music tasks, and top-level performance in audio-language tasks.
【3】 Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
标题: 声音场景空间语义分割的基线系统和评估表
链接:https://arxiv.org/abs/2503.22088
备注:5 pages
摘要:沉浸式通信已经取得了重大进展,特别是随着沉浸式语音和音频服务编解码器的发布。为了进一步实现这一目标,DCASE 2025挑战赛最近引入了一项声音场景空间语义分割任务(S5),重点是检测和分离空间声音场景中的声音事件。在本文中,我们将探索解决S5任务的方法。具体来说,我们提出了基线S5系统,结合音频标记(AT)和标签查询源分离(LSS)模型。我们研究了两种基于ResUNet架构的LSS方法:a)为每个检测到的事件提取单个源,b)同时查询多个源。由于S5中的每个分离的源由其声音事件类标签标识,因此我们提出了新的类感知度量来同时评估声源和标签。一阶立体混响空间音频的实验结果表明,所提出的系统的有效性,并确认的度量的有效性。
摘要:Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introduced a task for spatial semantic segmentation of sound scenes (S5), which focuses on detecting and separating sound events in spatial sound scenes. In this paper, we explore methods for addressing the S5 task. Specifically, we present baseline S5 systems that combine audio tagging (AT) and label-queried source separation (LSS) models. We investigate two LSS approaches based on the ResUNet architecture: a) extracting a single source for each detected event and b) querying multiple sources concurrently. Since each separated source in S5 is identified by its sound event class label, we propose new class-aware metrics to evaluate both the sound sources and labels simultaneously. Experimental results on first-order ambisonics spatial audio demonstrate the effectiveness of the proposed systems and confirm the efficacy of the metrics.
【4】 Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling for Barrier-free Communication
标题: 伸出援手:通过MLLM驱动的手部建模进行半免训练的提示语音识别,实现无障碍沟通
链接:https://arxiv.org/abs/2503.21785
摘要:提示语音(CS)是一种创新的视觉沟通系统,将唇读与手势编码相结合,旨在提高听力障碍者的有效沟通。自动CS识别(ACSR)是指人工智能驱动的自动识别CS中的手势和嘴唇运动的过程,并将其转换为文本。然而,以前的工作往往依赖于复杂的融合模块和训练技术。此外,由于CS中的数据量有限,手部特征的提取以及识别建模一直处于低水平,这大大限制了ACSR的有效性。为了解决这个问题,我们创新性地探索了多模态大语言模型(MLLM)在CS中识别手形和位置的能力。更准确地说,我们提出了一个新的半培训免费范式ACSR,命名为STF-ACSR。该方法通过中文CS提示模块(CCSPM)实现手部动作的zero-shot识别,CCSPM提供了免训练的关键帧过滤和基于MLLM的定制提示工程。然后,它集成到唇读模型使用最低限度的融合模块(MFM)的识别结果,有效地实现了卓越的识别结果。此外,特别是对于这项研究,我们通过记录来自8名听力障碍患者的额外数据,补充了现有的6名听力正常CS患者的数据集,从而形成了一个新的混合数据集。大量的实验表明,STF-ACSR显着优于以往的方法对正常和听力受损的数据。实现和检查点可在https://github.com/DennisHgj/STF_ACSR上获得。
摘要:Cued Speech (CS) is an innovative visual communication system that integrates lip-reading with hand coding, designed to enhance effective communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) refers to the AI-driven process of automatically recognizing hand gestures and lip movements in CS, converting them into text. However, previous work often relies on complex fusion modules and training techniques. Additionally, due to the limited amount of data in CS, the extraction of hand features, as well as recognition modeling, has consistently been subpar, significantly limiting the effectiveness of ACSR. To address this issue, we have innovatively explored the capabilities of Multimodal large language models (MLLMs) in recognizing hand shapes and positions in CS. More precisely, we propose a new Semi Training-Free paradigm for ACSR, named STF-ACSR. This approach leverages zero-shot recognition of hand movements through the Chinese CS Prompt Module (CCSPM), which equipped a training-free keyframe filtering and customized prompt engineering based on MLLM. It then integrates the recognition results into the lip-reading model using a Minimalist Fusion Module (MFM), effectively achieving superior recognition results. Furthermore, specifically for this study, we have supplemented the existing dataset of 6 normal hearing CS cuers by recording additional data from 8 cuers with hearing impairments, resulting in a new mixed dataset. Extensive experiments have demonstrated that STF-ACSR significantly outperforms previous methods on both normal and hearing-impaired data. Implementation and checkpoints are available at https://github.com/DennisHgj/STF_ACSR.
【5】 DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
标题: DeepAudio-V1:面向多模式多阶段端到端视频到语音和音频生成
链接:https://arxiv.org/abs/2503.22265
备注:11 pages, 5 figures
摘要:目前,高质量的同步音频是使用各种多模态联合学习框架,利用视频和可选的文本输入合成的。在视频到音频基准测试中,视频到音频质量、语义对齐和视听同步被有效地实现。然而,在现实场景中,语音和音频往往同时共存于视频中,并且在给定视频和文本条件下同步语音和音频的端到端生成还没有得到很好的研究。因此,我们提出了一个端到端的多模态生成框架,同时产生语音和音频的基础上的视频和文本的条件。此外,用于从视频生成语音的视频到音频(V2A)模型的优点仍然不清楚。所提出的框架DeepAudio由视频到音频(V2A)模块、文本到语音(TTS)模块和动态混合模态融合(MoF)模块组成。在评估中,提出的端到端框架实现了最先进的性能上的视频音频基准,视频语音基准,和文本语音基准。详细地说,我们的框架在视频-音频和文本-语音基准测试中与最先进的模型进行了比较,并在视频-语音基准测试中超过了最先进的模型,WER为16.57%至3.15%(+80.99%),SPK-SIM 78.30%至89.38%(+14.15%),EMO-SIM 66.24%至75.56%(+14.07%),MCD 8.59至7.98(+7.10%),MCD SL 11.05至9.40(+14.93%)。
摘要:Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio benchmarks, video-to-audio quality, semantic alignment, and audio-visual synchronization are effectively achieved. However, in real-world scenarios, speech and audio often coexist in videos simultaneously, and the end-to-end generation of synchronous speech and audio given video and text conditions are not well studied. Therefore, we propose an end-to-end multi-modal generation framework that simultaneously produces speech and audio based on video and text conditions. Furthermore, the advantages of video-to-audio (V2A) models for generating speech from videos remain unclear. The proposed framework, DeepAudio, consists of a video-to-audio (V2A) module, a text-to-speech (TTS) module, and a dynamic mixture of modality fusion (MoF) module. In the evaluation, the proposed end-to-end framework achieves state-of-the-art performance on the video-audio benchmark, video-speech benchmark, and text-speech benchmark. In detail, our framework achieves comparable results in the comparison with state-of-the-art models for the video-audio and text-speech benchmarks, and surpassing state-of-the-art models in the video-speech benchmark, with WER 16.57% to 3.15% (+80.99%), SPK-SIM 78.30% to 89.38% (+14.15%), EMO-SIM 66.24% to 75.56% (+14.07%), MCD 8.59 to 7.98 (+7.10%), MCD SL 11.05 to 9.40 (+14.93%) across a variety of dubbing settings.
【6】 Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
标题: 通过负条件反射潜在扩散模型增强舞蹈到音乐的生成
链接:https://arxiv.org/abs/2503.22138
摘要:条件扩散模型由于其令人印象深刻的跨模态合成结果而受到越来越多的关注,其中条件输入和生成的输出之间的强对齐可以通过训练具有交叉注意机制的时间条件U-Net来实现。在本文中,我们专注于生成与给定的舞蹈视频的节奏的视觉线索同步的音乐的问题。考虑到双向指导更有利于训练扩散模型,我们建议通过采用正节奏信息和负节奏信息(PN-扩散)作为条件来提高生成的音乐的质量及其与舞蹈视频的同步,其中设计了双扩散和反向过程。具体来说,为了训练一个顺序的多模态U-网络结构,PN-扩散包括一个噪声预测目标的积极条件和一个额外的噪声预测目标的消极条件。为了准确地定义和选择积极和消极的条件反射,我们巧妙地利用舞蹈视频中的时间相关性,通过分别向前和向后播放来捕捉积极和消极的节奏线索。通过在舞蹈音乐节拍对齐和生成音乐的质量方面对输入输出对应性进行主观和客观评估,在AIST++和TikTok舞蹈视频数据集上的实验结果表明,我们的模型优于SOTA舞蹈音乐生成模型。
摘要:Conditional diffusion models have gained increasing attention since their impressive results for cross-modal synthesis, where the strong alignment between conditioning input and generated output can be achieved by training a time-conditioned U-Net augmented with cross-attention mechanism. In this paper, we focus on the problem of generating music synchronized with rhythmic visual cues of the given dance video. Considering that bi-directional guidance is more beneficial for training a diffusion model, we propose to enhance the quality of generated music and its synchronization with dance videos by adopting both positive rhythmic information and negative ones (PN-Diffusion) as conditions, where a dual diffusion and reverse processes is devised. Specifically, to train a sequential multi-modal U-Net structure, PN-Diffusion consists of a noise prediction objective for positive conditioning and an additional noise prediction objective for negative conditioning. To accurately define and select both positive and negative conditioning, we ingeniously utilize temporal correlations in dance videos, capturing positive and negative rhythmic cues by playing them forward and backward, respectively. Through subjective and objective evaluations of input-output correspondence in terms of dance-music beat alignment and the quality of generated music, experimental results on the AIST++ and TikTok dance video datasets demonstrate that our model outperforms SOTA dance-to-music generation models.
【7】 Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
标题: 分层标签传播:AudioSet标签的依赖于模型大小的性能提升器
链接:https://arxiv.org/abs/2503.21826
备注:None
摘要:AudioSet是音频标记中最常用和最大的数据集之一,包含大约200万个音频样本,这些样本被手动标记为527个事件类别,并组织成一个本体。然而,注释包含不一致性,特别是根据本体应该被标记为积极的类别经常被错误地标记为消极的。为了解决这个问题,我们应用分层标签传播(HLP),它将标签传播到本体层次结构,导致每个音频片段的正标签平均从1.98增加到2.39,并影响527个类中的109个。我们的研究结果表明,HLP在各种模型架构中提供了性能优势,包括卷积神经网络(PANN的CNN6和ConvNeXT)和Transformers(PaSST),较小的模型显示出更多的改进。最后,在另一个广泛使用的数据集FSD 50K上,在AudioSet上使用HLP训练的模型始终优于在没有HLP的情况下训练的模型。我们的源代码将在GitHub上提供。
摘要:AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology. However, the annotations contain inconsistencies, particularly where categories that should be labeled as positive according to the ontology are frequently mislabeled as negative. To address this issue, we apply Hierarchical Label Propagation (HLP), which propagates labels up the ontology hierarchy, resulting in a mean increase in positive labels per audio clip from 1.98 to 2.39 and affecting 109 out of the 527 classes. Our results demonstrate that HLP provides performance benefits across various model architectures, including convolutional neural networks (PANN's CNN6 and ConvNeXT) and transformers (PaSST), with smaller models showing more improvements. Finally, on FSD50K, another widely used dataset, models trained on AudioSet with HLP consistently outperformed those trained without HLP. Our source code will be made available on GitHub.
