今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Binamix -- A Python Library for Generating Binaural Audio Datasets
标题: Binamix --用于生成双耳音频数据集的Python库
链接:https://arxiv.org/abs/2505.01369
作者: Dan Barry,  Davoud Shariat Panah,  Alessandro Ragano,  Jan Skoglund,  Andrew Hines 
备注:Accepted to the 158th Audio Engineering Society Convention, 2025
摘要:在虚拟现实、沉浸式媒体和空间音频研究等应用中对空间音频的需求不断增长,需要强大的解决方案来生成用于测试和验证的双耳音频数据集。Binamix是一个开源Python库,旨在使用广泛的SADIE II数据库促进程序化双耳混合,该数据库提供20个受试者的头部相关脉冲响应(HRIR)和双耳房间脉冲响应(BRIR)数据。Binamix库为创建大规模空间音频数据集提供了灵活且可重复的框架,使其成为编解码器评估,音频质量指标开发和机器学习模型训练的宝贵资源。一系列预构建的示例脚本、实用程序函数和可视化图进一步简化了自定义管道创建的过程。本文概述了该库的功能,包括双耳渲染,脉冲响应插值,以及各种扬声器布局的多声道混合。这些工具利用修改的Delaunay三角测量技术,在数据中不存在所需角度的情况下实现精确的HRIR/BRIR插值。通过支持方位角、仰角、受试者脉冲响应(IR)、扬声器布局、混音控制等各种参数,该库使研究人员能够为任何下游目的创建大型双耳数据集。Binamix通过为双耳渲染和数据集生成提供开源解决方案,使研究人员和开发人员能够通过可再现的方法来推进空间音频应用。我们在Apache 2.0许可证下发布该库,网址为https://github.com/QxLabIreland/Binamix/
摘要:The increasing demand for spatial audio in applications such as virtual reality, immersive media, and spatial audio research necessitates robust solutions to generate binaural audio data sets for use in testing and validation. Binamix is an open-source Python library designed to facilitate programmatic binaural mixing using the extensive SADIE II Database, which provides Head Related Impulse Response (HRIR) and Binaural Room Impulse Response (BRIR) data for 20 subjects. The Binamix library provides a flexible and repeatable framework for creating large-scale spatial audio datasets, making it an invaluable resource for codec evaluation, audio quality metric development, and machine learning model training. A range of pre-built example scripts, utility functions, and visualization plots further streamline the process of custom pipeline creation. This paper presents an overview of the library's capabilities, including binaural rendering, impulse response interpolation, and multi-track mixing for various speaker layouts. The tools utilize a modified Delaunay triangulation technique to achieve accurate HRIR/BRIR interpolation where desired angles are not present in the data. By supporting a wide range of parameters such as azimuth, elevation, subject Impulse Responses (IRs), speaker layouts, mixing controls, and more, the library enables researchers to create large binaural datasets for any downstream purpose. Binamix empowers researchers and developers to advance spatial audio applications with reproducible methodologies by offering an open-source solution for binaural rendering and dataset generation. We release the library under the Apache 2.0 License at https://github.com/QxLabIreland/Binamix/


【2】 FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and  Flow Matching based Voice Enhancing

标题: FlowDubber:基于LLM的语义感知学习和基于流匹配的语音增强的电影配音
链接:https://arxiv.org/abs/2505.01263
作者: Gaoxiang Cong,  Liang Li,  Jiadong Pan,  Zhedong Zhang,  Amin Beheshti,  Anton van den Hengel,  Yuankai Qi,  Qingming Huang 
摘要:电影配音旨在将脚本转换为在时间和情感方面与给定电影剪辑保持一致的语音,同时保留给定简短参考音频的声音音色。现有的方法主要集中在降低字错误率,而忽略了对口型和音质的重要性。为了解决这些问题,我们提出了一个大的语言模型(LLM)为基础的流匹配架构配音,命名为FlowDubber,它实现了高质量的视听同步和发音,通过将一个大的语音语言模型和双对比对齐,同时实现更好的声学质量,通过建议的语音增强流匹配比以前的作品。首先,我们引入Qwen2.5作为LLM的骨干,从电影脚本和参考音频中学习上下文序列。然后,所提出的语义感知学习的重点是捕捉LLM语义知识的音素水平。其次,双对比对齐(DCA)增强了嘴唇运动的相互对齐,减少了相似音素可能混淆的歧义。最后,提出的基于流的语音增强(FVE)从两个方面改善声学质量,即引入基于LLM的声学流匹配指导来增强清晰度,并在通过梯度矢量场预测将噪声恢复到梅尔频谱图时使用仿射风格来增强身份。大量的实验表明,我们的方法优于几个国家的最先进的方法在两个主要的基准。这些演示可在{\href{https://galaxyyop.github.io/LLM-Flow-Dubber/}{\textcolor{red}{https://galaxyyop.github.io/LLM-Flow-Dubber/}}获得。
摘要:Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing the word error rate while ignoring the importance of lip-sync and acoustic quality. To address these issues, we propose a large language model (LLM) based flow matching architecture for dubbing, named FlowDubber, which achieves high-quality audio-visual sync and pronunciation by incorporating a large speech language model and dual contrastive aligning while achieving better acoustic quality via the proposed voice-enhanced flow matching than previous works. First, we introduce Qwen2.5 as the backbone of LLM to learn the in-context sequence from movie scripts and reference audio. Then, the proposed semantic-aware learning focuses on capturing LLM semantic knowledge at the phoneme level. Next, dual contrastive aligning (DCA) boosts mutual alignment with lip movement, reducing ambiguities where similar phonemes might be confused. Finally, the proposed Flow-based Voice Enhancing (FVE) improves acoustic quality in two aspects, which introduces an LLM-based acoustics flow matching guidance to strengthen clarity and uses affine style prior to enhance identity when recovering noise into mel-spectrograms via gradient vector field prediction. Extensive experiments demonstrate that our method outperforms several state-of-the-art methods on two primary benchmarks. The demos are available at {\href{https://galaxycong.github.io/LLM-Flow-Dubber/}{\textcolor{red}{https://galaxycong.github.io/LLM-Flow-Dubber/}}}.


【3】 CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via  Fine-Grained Alignment

标题: CAV-MAE同步:通过细粒度对齐改进对比视听面罩自动编码器
链接:https://arxiv.org/abs/2505.01237
作者: Edson Araujo,  Andrew Rouditchenko,  Yuan Gong,  Saurabhchand Bhati,  Samuel Thomas,  Brian Kingsbury,  Leonid Karlinsky,  Rogerio Feris,  James R. Glass 
备注:To be published at CVPR 2025, code available at this https URL
摘要:视听学习的最新进展显示出跨模态学习表征的可喜成果。然而,大多数方法依赖于全局音频表示,无法捕获细粒度的时间对应与视觉帧。此外,现有方法在尝试联合学习重构和跨模态对齐时,经常与冲突的优化目标作斗争。在这项工作中,我们提出了CAV-MAE Sync作为原始CAV-MAE框架的简单而有效的扩展,用于自我监督的视听学习。我们解决了三个关键挑战:首先,我们通过将音频视为与视频帧对齐的时间序列来解决模态之间的粒度失配,而不是使用全局表示。其次,我们通过专用的全局令牌分离对比和重建目标来解决相互冲突的优化目标。第三,我们通过引入可学习的寄存器令牌,减少补丁令牌上的语义负载来改进空间定位。我们在AudioSet、VGG Sound和ADE 20 K Sound数据集上对所提出的方法进行了评估,这些数据集用于zero-shot检索、分类和定位任务,展示了最先进的性能,并优于更复杂的架构。
摘要:Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences with visual frames. Additionally, existing methods often struggle with conflicting optimization objectives when trying to jointly learn reconstruction and cross-modal alignment. In this work, we propose CAV-MAE Sync as a simple yet effective extension of the original CAV-MAE framework for self-supervised audio-visual learning. We address three key challenges: First, we tackle the granularity mismatch between modalities by treating audio as a temporal sequence aligned with video frames, rather than using global representations. Second, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens. Third, we improve spatial localization by introducing learnable register tokens that reduce semantic load on patch tokens. We evaluate the proposed approach on AudioSet, VGG Sound, and the ADE20K Sound dataset on zero-shot retrieval, classification and localization tasks demonstrating state-of-the-art performance and outperforming more complex architectures.


【4】 SMSAT: A Multimodal Acoustic Dataset and Deep Contrastive Learning  Framework for Affective and Physiological Modeling of Spiritual Meditation

标题: SMSAT:用于精神冥想情感和生理建模的多模式声学数据集和深度对比学习框架
链接:https://arxiv.org/abs/2505.00839
作者: Ahmad Suleman,  Yazeed Alkhrijah,  Misha Urooj Khan,  Hareem Khan,  Muhammad Abdullah Husnain Ali Faiz,  Mohamad A. Alawad,  Zeeshan Kaleem,  Guan Gui 
摘要:了解听觉刺激如何影响情绪和生理状态是推进情感计算和心理健康技术的基础。在本文中,我们提出了一个多模态评估的情感和生理影响的三个听觉条件,即精神冥想(SM),音乐(M),和自然的沉默(NS),使用一套全面的生物特征信号的措施。为了便于这种分析,我们介绍了精神,音乐,沉默声学时间序列(SMSAT)数据集,一种新的基准,包括声学时间序列(ATS)的信号记录下控制暴露协议,仔细注意人口的多样性和实验的一致性。为了对听觉诱导状态进行建模,我们开发了一种基于对比学习的SMSAT音频编码器,该编码器从ATS数据中提取高度区分的嵌入,在类间和类内评估中实现了99.99%的分类准确率。此外,我们提出了冷静分析模型(CAM),这是一个深度学习框架,集成了25个手工制作和学习的特征,用于在听觉条件下进行情感状态分类,实现了99.99%的分类准确率。相比之下,成对t检验显示,通过方差分析进行的SM分析之间的心脏反应特征(CRC)存在显着偏差,从而诱导更显着的生理波动。与现有的最先进的方法报告的准确性高达90%相比,所提出的模型表现出显着的性能增益(高达99%)。这项工作为压力监测、心理健康和基于音频的治疗干预中的情感计算应用提供了一个经过验证的多模态数据集和一个可扩展的深度学习框架。
摘要:Understanding how auditory stimuli influence emotional and physiological states is fundamental to advancing affective computing and mental health technologies. In this paper, we present a multimodal evaluation of the affective and physiological impacts of three auditory conditions, that is, spiritual meditation (SM), music (M), and natural silence (NS), using a comprehensive suite of biometric signal measures. To facilitate this analysis, we introduce the Spiritual, Music, Silence Acoustic Time Series (SMSAT) dataset, a novel benchmark comprising acoustic time series (ATS) signals recorded under controlled exposure protocols, with careful attention to demographic diversity and experimental consistency. To model the auditory induced states, we develop a contrastive learning based SMSAT audio encoder that extracts highly discriminative embeddings from ATS data, achieving 99.99% classification accuracy in interclass and intraclass evaluations. Furthermore, we propose the Calmness Analysis Model (CAM), a deep learning framework integrating 25 handcrafted and learned features for affective state classification across auditory conditions, attaining robust 99.99% classification accuracy. In contrast, pairwise t tests reveal significant deviations in cardiac response characteristics (CRC) between SM analysis via ANOVA inducing more significant physiological fluctuations. Compared to existing state of the art methods reporting accuracies up to 90%, the proposed model demonstrates substantial performance gains (up to 99%). This work contributes a validated multimodal dataset and a scalable deep learning framework for affective computing applications in stress monitoring, mental well-being, and therapeutic audio-based interventions.


【5】 GVPT -- A software for guided visual pitch tracking

标题: GVPT --引导视觉音调跟踪软件
链接:https://arxiv.org/abs/2505.00750
作者: Hyunjin Cho,  Farhad Tabasi,  Jeremy D. Greenlee,  Rahul Singh 
摘要:GVPT(Guided Visual Pitch Tracking)是一个公开的实时音高跟踪软件,旨在使用视觉反馈指导和评估音高控制。该系统是为临床和研究应用而开发的,它可以呈现各种视觉目标音高轮廓,并实时叠加受试者的音高,以促进准确的声乐再现。GVPT支持难度修改,会话记录和精确的音高跟踪。该软件可以在实验和治疗环境中进行语音音高控制练习。
摘要:GVPT (Guided visual pitch tracking) is a publicly available, real-time pitch tracking software designed to guide and evaluate vocal pitch control using visual feedback. Developed for clinical and research applications, the system presents various visual target pitch contour and overlays the subject's pitch in real-time to promote accurate vocal reproduction. GVPT supports difficulty modification, session logging, and precise pitch tracking. The software enables voice pitch control exercise in both experimental and therapeutic settings.


【6】 How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement  in Distant Microphone Scenarios

标题: 消除多少回响?远距离麦克风场景中的低延迟单通道语音增强
链接:https://arxiv.org/abs/2505.01338
作者: Satvik Venkatesh,  Philip Coleman,  Arthur Benilov,  Simon Brown,  Selim Sheta,  Frederic Roskam 
备注:Published in ICASSP 2025
摘要:去混响是语音增强中的一个重要子任务,可以提高语音信号的可懂度和质量。然而,它仍然具有挑战性,因为混响与信号高度相关。此外,单通道SE文献主要集中在混响时间短(通常低于1秒)的房间,较小的房间(体积低于1000立方米)和相对较短的距离(高达2米)。在本文中,我们将探索在远距离麦克风场景下(例如5至10米)的实时低延迟单通道SE,并将重点放在会议室和剧院,具有较大的房间尺寸和混响时间。这样的设置对于诸如演讲演示、戏剧和增强舞台声学的应用是有用的。首先,我们表明,单通道SE在这种具有挑战性的情况下是可行的。其次,我们研究了房间体积和混响时间之间的关系,并证明了它的重要性时,随机模拟房间脉冲响应。最后,我们表明,对于衰减时间短的去混响,在衰减房间的传递函数之前保留早期反射可以提高整体信号质量。
摘要:Dereverberation is an important sub-task of Speech Enhancement (SE) to improve the signal's intelligibility and quality. However, it remains challenging because the reverberation is highly correlated with the signal. Furthermore, the single-channel SE literature has predominantly focused on rooms with short reverb times (typically under 1 second), smaller rooms (under volumes of 1000 cubic meters) and relatively short distances (up to 2 meters). In this paper, we explore real-time low-latency single-channel SE under distant microphone scenarios, such as 5 to 10 meters, and focus on conference rooms and theatres, with larger room dimensions and reverberation times. Such a setup is useful for applications such as lecture demonstrations, drama, and to enhance stage acoustics. First, we show that single-channel SE in such challenging scenarios is feasible. Second, we investigate the relationship between room volume and reverberation time, and demonstrate its importance when randomly simulating room impulse responses. Lastly, we show that for dereverberation with short decay times, preserving early reflections before decaying the transfer function of the room improves overall signal quality.


eess.AS音频处理

【1】 How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement  in Distant Microphone Scenarios

标题: 消除多少回响?远距离麦克风场景中的低延迟单通道语音增强
链接:https://arxiv.org/abs/2505.01338
作者: Satvik Venkatesh,  Philip Coleman,  Arthur Benilov,  Simon Brown,  Selim Sheta,  Frederic Roskam 
备注:Published in ICASSP 2025
摘要:去混响是语音增强中的一个重要子任务,可以提高语音信号的可懂度和质量。然而,它仍然具有挑战性,因为混响与信号高度相关。此外,单通道SE文献主要集中在混响时间短(通常低于1秒)的房间,较小的房间(体积低于1000立方米)和相对较短的距离(高达2米)。在本文中,我们将探索在远距离麦克风场景下(例如5至10米)的实时低延迟单通道SE,并将重点放在会议室和剧院,具有较大的房间尺寸和混响时间。这样的设置对于诸如演讲演示、戏剧和增强舞台声学的应用是有用的。首先,我们表明,单通道SE在这种具有挑战性的情况下是可行的。其次,我们研究了房间体积和混响时间之间的关系,并证明了它的重要性时,随机模拟房间脉冲响应。最后,我们表明,对于衰减时间短的去混响,在衰减房间的传递函数之前保留早期反射可以提高整体信号质量。
摘要:Dereverberation is an important sub-task of Speech Enhancement (SE) to improve the signal's intelligibility and quality. However, it remains challenging because the reverberation is highly correlated with the signal. Furthermore, the single-channel SE literature has predominantly focused on rooms with short reverb times (typically under 1 second), smaller rooms (under volumes of 1000 cubic meters) and relatively short distances (up to 2 meters). In this paper, we explore real-time low-latency single-channel SE under distant microphone scenarios, such as 5 to 10 meters, and focus on conference rooms and theatres, with larger room dimensions and reverberation times. Such a setup is useful for applications such as lecture demonstrations, drama, and to enhance stage acoustics. First, we show that single-channel SE in such challenging scenarios is feasible. Second, we investigate the relationship between room volume and reverberation time, and demonstrate its importance when randomly simulating room impulse responses. Lastly, we show that for dereverberation with short decay times, preserving early reflections before decaying the transfer function of the room improves overall signal quality.


【2】 Physics-Informed Neural Network-Driven Sparse Field Discretization  Method for Near-Field Acoustic Holography

标题: 物理信息神经网络驱动的近场声全息稀疏场离散化方法
链接:https://arxiv.org/abs/2505.00897
作者: Xinmeng Luan,  Mirco Pezzoli,  Fabio Antonacci,  Augusto Sarti 
备注:12 pages, 7 figures
摘要:我们提出了物理信息神经网络驱动的稀疏场离散化方法(PINN-SFD),这是一种新的自监督,物理信息深度学习方法,用于解决近场声全息(NAH)问题。与现有的NAH深度学习方法不同,这些方法主要由大型数据集监督,我们的方法不需要训练阶段,并且它是物理信息。波传播场被离散成稀疏区域,这一过程称为场离散化,其包括一系列源平面集合,以解决逆问题。我们的方法采用离散Kirchhoff-Helmholtz积分作为波的传播模型。通过结合虚拟平面,在实际声源附近实施额外的约束,从而改善重建过程。使用物理信息神经网络(PINN)进行优化,其中基于物理的约束被集成到损失函数中,以考虑直接(从等效源平面到全息平面)和附加(从虚拟平面到全息平面)波传播路径。此外,稀疏性是强制的等效源的速度。我们对各种矩形和小提琴顶板进行了全面验证,涵盖了广泛的振动模式,表明PINN-SFD始终优于传统的压缩等效源方法(C-ESM),特别是在复杂振动模式的重建精度方面。值得注意的是,与C-ESM相比,该方法对正则化参数的敏感性降低。
摘要:We propose the Physics-Informed Neural Network-driven Sparse Field Discretization method (PINN-SFD), a novel self-supervised, physics-informed deep learning approach for addressing the Near-Field Acoustic Holography (NAH) problem. Unlike existing deep learning methods for NAH, which are predominantly supervised by large datasets, our approach does not require a training phase and it is physics-informed. The wave propagation field is discretized into sparse regions, a process referred to as field discretization, which includes a series of set of source planes, to address the inverse problem. Our method employs the discretized Kirchhoff-Helmholtz integral as the wave propagation model. By incorporating virtual planes, additional constraints are enforced near the actual sound source, improving the reconstruction process. Optimization is carried out using Physics-Informed Neural Networks (PINNs), where physics-based constraints are integrated into the loss functions to account for both direct (from equivalent source plane to hologram plane) and additional (from virtual planes to hologram plane) wave propagation paths. Additionally, sparsity is enforced on the velocity of the equivalent sources. Our comprehensive validation across various rectangular and violin top plates, covering a wide range of vibrational modes, demonstrates that PINN-SFD consistently outperforms the conventional Compressive-Equivalent Source Method (C-ESM), particularly in terms of reconstruction accuracy for complex vibrational patterns. Significantly, this method demonstrates reduced sensitivity to regularization parameters compared to C-ESM.


【3】 Binamix -- A Python Library for Generating Binaural Audio Datasets

标题: Binamix --用于生成双耳音频数据集的Python库
链接:https://arxiv.org/abs/2505.01369
作者: Dan Barry,  Davoud Shariat Panah,  Alessandro Ragano,  Jan Skoglund,  Andrew Hines 
备注:Accepted to the 158th Audio Engineering Society Convention, 2025
摘要:在虚拟现实、沉浸式媒体和空间音频研究等应用中对空间音频的需求不断增长,需要强大的解决方案来生成用于测试和验证的双耳音频数据集。Binamix是一个开源Python库,旨在使用广泛的SADIE II数据库促进程序化双耳混合,该数据库提供20个受试者的头部相关脉冲响应(HRIR)和双耳房间脉冲响应(BRIR)数据。Binamix库为创建大规模空间音频数据集提供了灵活且可重复的框架,使其成为编解码器评估,音频质量指标开发和机器学习模型训练的宝贵资源。一系列预构建的示例脚本、实用程序函数和可视化图进一步简化了自定义管道创建的过程。本文概述了该库的功能,包括双耳渲染,脉冲响应插值,以及各种扬声器布局的多声道混合。这些工具利用修改的Delaunay三角测量技术,在数据中不存在所需角度的情况下实现精确的HRIR/BRIR插值。通过支持方位角、仰角、受试者脉冲响应(IR)、扬声器布局、混音控制等各种参数,该库使研究人员能够为任何下游目的创建大型双耳数据集。Binamix通过为双耳渲染和数据集生成提供开源解决方案,使研究人员和开发人员能够通过可再现的方法来推进空间音频应用。我们在Apache 2.0许可证下发布该库,网址为https://github.com/QxLabIreland/Binamix/
摘要:The increasing demand for spatial audio in applications such as virtual reality, immersive media, and spatial audio research necessitates robust solutions to generate binaural audio data sets for use in testing and validation. Binamix is an open-source Python library designed to facilitate programmatic binaural mixing using the extensive SADIE II Database, which provides Head Related Impulse Response (HRIR) and Binaural Room Impulse Response (BRIR) data for 20 subjects. The Binamix library provides a flexible and repeatable framework for creating large-scale spatial audio datasets, making it an invaluable resource for codec evaluation, audio quality metric development, and machine learning model training. A range of pre-built example scripts, utility functions, and visualization plots further streamline the process of custom pipeline creation. This paper presents an overview of the library's capabilities, including binaural rendering, impulse response interpolation, and multi-track mixing for various speaker layouts. The tools utilize a modified Delaunay triangulation technique to achieve accurate HRIR/BRIR interpolation where desired angles are not present in the data. By supporting a wide range of parameters such as azimuth, elevation, subject Impulse Responses (IRs), speaker layouts, mixing controls, and more, the library enables researchers to create large binaural datasets for any downstream purpose. Binamix empowers researchers and developers to advance spatial audio applications with reproducible methodologies by offering an open-source solution for binaural rendering and dataset generation. We release the library under the Apache 2.0 License at https://github.com/QxLabIreland/Binamix/


【4】 FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and  Flow Matching based Voice Enhancing

标题: FlowDubber:基于LLM的语义感知学习和基于流匹配的语音增强的电影配音
链接:https://arxiv.org/abs/2505.01263
作者: Gaoxiang Cong,  Liang Li,  Jiadong Pan,  Zhedong Zhang,  Amin Beheshti,  Anton van den Hengel,  Yuankai Qi,  Qingming Huang 
摘要:电影配音旨在将脚本转换为在时间和情感方面与给定电影剪辑保持一致的语音,同时保留给定简短参考音频的声音音色。现有的方法主要集中在降低字错误率,而忽略了对口型和音质的重要性。为了解决这些问题,我们提出了一个大的语言模型(LLM)为基础的流匹配架构配音,命名为FlowDubber,它实现了高质量的视听同步和发音,通过将一个大的语音语言模型和双对比对齐,同时实现更好的声学质量,通过建议的语音增强流匹配比以前的作品。首先,我们引入Qwen2.5作为LLM的骨干,从电影脚本和参考音频中学习上下文序列。然后,所提出的语义感知学习的重点是捕捉LLM语义知识的音素水平。其次,双对比对齐(DCA)增强了嘴唇运动的相互对齐,减少了相似音素可能混淆的歧义。最后,提出的基于流的语音增强(FVE)从两个方面改善声学质量,即引入基于LLM的声学流匹配指导来增强清晰度,并在通过梯度矢量场预测将噪声恢复到梅尔频谱图时使用仿射风格来增强身份。大量的实验表明,我们的方法优于几个国家的最先进的方法在两个主要的基准。这些演示可在{\href {https://galaxyyop.github.io/LLM-Flow-Dubber/}{\textcolor {red}{https://galaxyyop.github.io/LLM-Flow-Dubber/}}获得。
摘要:Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing the word error rate while ignoring the importance of lip-sync and acoustic quality. To address these issues, we propose a large language model (LLM) based flow matching architecture for dubbing, named FlowDubber, which achieves high-quality audio-visual sync and pronunciation by incorporating a large speech language model and dual contrastive aligning while achieving better acoustic quality via the proposed voice-enhanced flow matching than previous works. First, we introduce Qwen2.5 as the backbone of LLM to learn the in-context sequence from movie scripts and reference audio. Then, the proposed semantic-aware learning focuses on capturing LLM semantic knowledge at the phoneme level. Next, dual contrastive aligning (DCA) boosts mutual alignment with lip movement, reducing ambiguities where similar phonemes might be confused. Finally, the proposed Flow-based Voice Enhancing (FVE) improves acoustic quality in two aspects, which introduces an LLM-based acoustics flow matching guidance to strengthen clarity and uses affine style prior to enhance identity when recovering noise into mel-spectrograms via gradient vector field prediction. Extensive experiments demonstrate that our method outperforms several state-of-the-art methods on two primary benchmarks. The demos are available at {\href{https://galaxycong.github.io/LLM-Flow-Dubber/}{\textcolor{red}{https://galaxycong.github.io/LLM-Flow-Dubber/}}}.


【5】 CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via  Fine-Grained Alignment

标题: CAV-MAE同步:通过细粒度对齐改进对比视听面罩自动编码器
链接:https://arxiv.org/abs/2505.01237
作者: Edson Araujo,  Andrew Rouditchenko,  Yuan Gong,  Saurabhchand Bhati,  Samuel Thomas,  Brian Kingsbury,  Leonid Karlinsky,  Rogerio Feris,  James R. Glass 
备注:To be published at CVPR 2025, code available at this https URL
摘要:视听学习的最新进展显示出跨模态学习表征的可喜成果。然而,大多数方法依赖于全局音频表示,无法捕获细粒度的时间对应与视觉帧。此外,现有方法在尝试联合学习重构和跨模态对齐时,经常与冲突的优化目标作斗争。在这项工作中,我们提出了CAV-MAE Sync作为原始CAV-MAE框架的简单而有效的扩展,用于自我监督的视听学习。我们解决了三个关键挑战:首先,我们通过将音频视为与视频帧对齐的时间序列来解决模态之间的粒度失配,而不是使用全局表示。其次,我们通过专用的全局令牌分离对比和重建目标来解决相互冲突的优化目标。第三,我们通过引入可学习的寄存器令牌,减少补丁令牌上的语义负载来改进空间定位。我们在AudioSet、VGG Sound和ADE 20 K Sound数据集上对所提出的方法进行了评估,这些数据集用于zero-shot检索、分类和定位任务,展示了最先进的性能,并优于更复杂的架构。
摘要:Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences with visual frames. Additionally, existing methods often struggle with conflicting optimization objectives when trying to jointly learn reconstruction and cross-modal alignment. In this work, we propose CAV-MAE Sync as a simple yet effective extension of the original CAV-MAE framework for self-supervised audio-visual learning. We address three key challenges: First, we tackle the granularity mismatch between modalities by treating audio as a temporal sequence aligned with video frames, rather than using global representations. Second, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens. Third, we improve spatial localization by introducing learnable register tokens that reduce semantic load on patch tokens. We evaluate the proposed approach on AudioSet, VGG Sound, and the ADE20K Sound dataset on zero-shot retrieval, classification and localization tasks demonstrating state-of-the-art performance and outperforming more complex architectures.


【6】 SMSAT: A Multimodal Acoustic Dataset and Deep Contrastive Learning  Framework for Affective and Physiological Modeling of Spiritual Meditation

标题: SMSAT:用于精神冥想情感和生理建模的多模式声学数据集和深度对比学习框架
链接:https://arxiv.org/abs/2505.00839
作者: Ahmad Suleman,  Yazeed Alkhrijah,  Misha Urooj Khan,  Hareem Khan,  Muhammad Abdullah Husnain Ali Faiz,  Mohamad A. Alawad,  Zeeshan Kaleem,  Guan Gui 
摘要:了解听觉刺激如何影响情绪和生理状态是推进情感计算和心理健康技术的基础。在本文中,我们提出了一个多模态评估的情感和生理影响的三个听觉条件,即精神冥想(SM),音乐(M),和自然的沉默(NS),使用一套全面的生物特征信号的措施。为了便于这种分析,我们介绍了精神,音乐,沉默声学时间序列(SMSAT)数据集,一种新的基准,包括声学时间序列(ATS)的信号记录下控制暴露协议,仔细注意人口的多样性和实验的一致性。为了对听觉诱导状态进行建模,我们开发了一种基于对比学习的SMSAT音频编码器,该编码器从ATS数据中提取高度区分的嵌入,在类间和类内评估中实现了99.99%的分类准确率。此外,我们提出了冷静分析模型(CAM),这是一个深度学习框架,集成了25个手工制作和学习的特征,用于在听觉条件下进行情感状态分类,实现了99.99%的分类准确率。相比之下,成对t检验显示,通过方差分析进行的SM分析之间的心脏反应特征(CRC)存在显着偏差,从而诱导更显着的生理波动。与现有的最先进的方法报告的准确性高达90%相比,所提出的模型表现出显着的性能增益(高达99%)。这项工作为压力监测、心理健康和基于音频的治疗干预中的情感计算应用提供了一个经过验证的多模态数据集和一个可扩展的深度学习框架。
摘要:Understanding how auditory stimuli influence emotional and physiological states is fundamental to advancing affective computing and mental health technologies. In this paper, we present a multimodal evaluation of the affective and physiological impacts of three auditory conditions, that is, spiritual meditation (SM), music (M), and natural silence (NS), using a comprehensive suite of biometric signal measures. To facilitate this analysis, we introduce the Spiritual, Music, Silence Acoustic Time Series (SMSAT) dataset, a novel benchmark comprising acoustic time series (ATS) signals recorded under controlled exposure protocols, with careful attention to demographic diversity and experimental consistency. To model the auditory induced states, we develop a contrastive learning based SMSAT audio encoder that extracts highly discriminative embeddings from ATS data, achieving 99.99% classification accuracy in interclass and intraclass evaluations. Furthermore, we propose the Calmness Analysis Model (CAM), a deep learning framework integrating 25 handcrafted and learned features for affective state classification across auditory conditions, attaining robust 99.99% classification accuracy. In contrast, pairwise t tests reveal significant deviations in cardiac response characteristics (CRC) between SM analysis via ANOVA inducing more significant physiological fluctuations. Compared to existing state of the art methods reporting accuracies up to 90%, the proposed model demonstrates substantial performance gains (up to 99%). This work contributes a validated multimodal dataset and a scalable deep learning framework for affective computing applications in stress monitoring, mental well-being, and therapeutic audio-based interventions.


【7】 GVPT -- A software for guided visual pitch tracking

标题: GVPT --引导视觉音调跟踪软件
链接:https://arxiv.org/abs/2505.00750
作者: Hyunjin Cho,  Farhad Tabasi,  Jeremy D. Greenlee,  Rahul Singh 
摘要:GVPT(Guided Visual Pitch Tracking)是一个公开的实时音高跟踪软件,旨在使用视觉反馈指导和评估音高控制。该系统是为临床和研究应用而开发的,它可以呈现各种视觉目标音高轮廓,并实时叠加受试者的音高,以促进准确的声乐再现。GVPT支持难度修改,会话记录和精确的音高跟踪。该软件可以在实验和治疗环境中进行语音音高控制练习。
摘要:GVPT (Guided visual pitch tracking) is a publicly available, real-time pitch tracking software designed to guide and evaluate vocal pitch control using visual feedback. Developed for clinical and research applications, the system presents various visual target pitch contour and overlays the subject's pitch in real-time to promote accurate vocal reproduction. GVPT supports difficulty modification, session logging, and precise pitch tracking. The software enables voice pitch control exercise in both experimental and therapeutic settings.


机器翻译由腾讯交互翻译提供,仅供参考