本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
【1】 Past, Present, and Future of Spatial Audio and Room Acoustics
链接:https://arxiv.org/abs/2503.12948
备注:Accepted to International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2025
摘要:空间音频和室内声学的研究旨在通过模拟声音在空间中的行为的物理和心理声学来创建沉浸式音频体验。在这一研究领域的漫长历史中,各种关键技术都是基于理论进步和实践创新而开发的。我们突出的历史成就,倡议活动,最近的进展,并在空间音频记录和再现,室内声学模拟,建模,分析和控制的研究领域的未来展望。
摘要:The study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have been developed based both on theoretical advancements and practical innovations. We highlight historical achievements, initiative activities, recent advancements, and future outlooks in the research area of spatial audio recording and reproduction, and room acoustic simulation, modeling, analysis, and control.
标题:FNSE-SBGAN:使用Schrodinger桥和生成对抗网络的远场语音增强
链接:https://arxiv.org/abs/2503.12936
备注:13 pages, 6 figures
摘要:目前神经语音增强的主要方法依赖于使用模拟的远场噪声混响语音(混合)和干净语音对的纯监督深度学习。然而,这些训练的模型往往表现出有限的泛化到真实记录的混合物。为了解决这个问题,本研究直接在真实混合物上研究训练增强模型。具体来说,我们重新审视单通道远场到近场语音增强(FNSE)任务,重点是现实世界的数据,其特征在于低信噪比(SNR),高混响,中高频衰减。我们提出了FNSE-SBGAN,这是一个新的框架,它将基于薛定谔桥(SB)的扩散模型与生成对抗网络(GANs)集成在一起。我们的方法在各种指标和主观评估方面实现了最先进的性能,与远场信号相比,字符错误率(CER)显著降低了14.58%。实验结果表明,FNSE-SBGAN保留了优越的主观质量,并建立了一个新的基准,为现实世界的远场语音增强。此外,我们引入了一个新的评估框架,利用矩阵秩分析在时间-频率域,提供系统的见解模型的性能和揭示不同的生成方法的优点和缺点。
摘要:The current dominant approach for neural speech enhancement relies on purely-supervised deep learning using simulated pairs of far-field noisy-reverberant speech (mixtures) and clean speech. However, these trained models often exhibit limited generalizability to real-recorded mixtures. To address this issue, this study investigates training enhancement models directly on real mixtures. Specifically, we revisit the single-channel far-field to near-field speech enhancement (FNSE) task, focusing on real-world data characterized by low signal-to-noise ratio (SNR), high reverberation, and mid-to-high frequency attenuation. We propose FNSE-SBGAN, a novel framework that integrates a Schrodinger Bridge (SB)-based diffusion model with generative adversarial networks (GANs). Our approach achieves state-of-the-art performance across various metrics and subjective evaluations, significantly reducing the character error rate (CER) by up to 14.58% compared to far-field signals. Experimental results demonstrate that FNSE-SBGAN preserves superior subjective quality and establishes a new benchmark for real-world far-field speech enhancement. Additionally, we introduce a novel evaluation framework leveraging matrix rank analysis in the time-frequency domain, providing systematic insights into model performance and revealing the strengths and weaknesses of different generative methods.
标题: 自适应混合专家学习以实现稳健的音频欺骗检测
链接:https://arxiv.org/abs/2503.12010
备注:Submitted to Eusipco 2025, 5 pages, 1 figure, 5 table
摘要:在音频欺骗检测中,大多数研究依赖于干净的数据集,使得模型容易受到真实世界的后处理攻击,例如通道压缩和噪声。为了克服这一挑战,我们提出了自适应混合专家学习(AMEL)框架,该框架通过利用特定于攻击的知识和动态适应不同的攻击条件来增强弹性。具体来说,AMEL利用了通过低秩自适应(LoRA)进行微调的攻击特定专家(ASE),使每个专家能够针对特定的后处理模式,同时只需要完全微调所需参数的1.12%。此外,我们引入动态专家聚合(DEA),自适应地选择和整合专家知识,以提高欺骗检测的鲁棒性。实验结果表明,AMEL显着提高了鲁棒性,提高噪声弹性和表现出更大的适应性,以前看不见的后处理方法相比,依赖于完全微调的模型。此外,我们的框架在各种混合攻击下的性能优于单个专家和简单平均集成,证明了其在管理复杂的现实世界条件下的卓越鲁棒性和适应性。
摘要:In audio spoofing detection, most studies rely on clean datasets, making models susceptible to real-world post-processing attacks, such as channel compression and noise. To overcome this challenge, we propose the Adaptive Mixture of Experts Learning (AMEL) framework, which enhances resilience by leveraging attack-specific knowledge and adapting dynamically to varied attack conditions. Specifically, AMEL utilizes Attack-Specific Experts (ASE) fine-tuned with Low-Rank Adaptation (LoRA), enabling each expert to target specific post-processing patterns while requiring only 1.12\% of the parameters needed for full fine-tuning. Furthermore, we introduce Dynamic Expert Aggregation (DEA), which adaptively selects and integrates expert knowledge to enhance the robustness of spoofing detection. Experimental results demonstrate that AMEL significantly enhances robustness by improving noise resilience and exhibiting greater adaptability to previously unseen post-processing methods compared to models relying on full fine-tuning. Additionally, our framework outperforms both single expert and simple average ensemble under various mixed attacks, demonstrating its superior robustness and adaptability in managing complex, real-world conditions.
标题:AV-Surf:表面增强的几何感知新视图声学合成
链接:https://arxiv.org/abs/2503.12806
摘要:在复杂的真实世界环境中准确地建模声音传播对于新视图声学合成(NVAS)至关重要。虽然以前的研究已经利用视觉感知来估计空间声学,但在声学建模中结合使用来自3D表示的表面法线和结构细节一直未被充分探索。考虑到它们对声波反射和传播的直接影响,表面法线应与结构细节联合建模,以实现精确的空间声学。在本文中,我们提出了一种表面增强的几何感知方法NVAS,以提高空间声学建模。为了实现这一点,我们利用几何先验,如图像,深度图,表面法线,和点云使用3D高斯溅射(3DGS)为基础的框架。我们引入了一个基于双交叉注意的Transformer,将几何约束集成到频率查询中,以了解发射器的周围环境。此外,我们设计了一个基于ConvNeXT的频谱特征处理网络,称为频谱细化网络(SRN),以合成逼真的双耳音频。RWAVS和SoundSpace数据集上的实验结果突出了我们方法的必要性,因为它在新视图声学合成中超越了现有方法。
摘要:Accurately modeling sound propagation with complex real-world environments is essential for Novel View Acoustic Synthesis (NVAS). While previous studies have leveraged visual perception to estimate spatial acoustics, the combined use of surface normal and structural details from 3D representations in acoustic modeling has been underexplored. Given their direct impact on sound wave reflections and propagation, surface normals should be jointly modeled with structural details to achieve accurate spatial acoustics. In this paper, we propose a surface-enhanced geometry-aware approach for NVAS to improve spatial acoustic modeling. To achieve this, we exploit geometric priors, such as image, depth map, surface normals, and point clouds obtained using a 3D Gaussian Splatting (3DGS) based framework. We introduce a dual cross-attention-based transformer integrating geometrical constraints into frequency query to understand the surroundings of the emitter. Additionally, we design a ConvNeXt-based spectral features processing network called Spectral Refinement Network (SRN) to synthesize realistic binaural audio. Experimental results on the RWAVS and SoundSpace datasets highlight the necessity of our approach, as it surpasses existing methods in novel view acoustic synthesis.
【5】 Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
标题:通过音频引导视觉收敛对齐的鲁棒视听分割
链接:https://arxiv.org/abs/2503.12847
备注:Accepted by CVPR2025
【6】 Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
标题:动态推导和消除:具有增强音频语义的视听分段
链接:https://arxiv.org/abs/2503.12840
备注:Accepted by CVPR2025
【7】 A General Close-loop Predictive Coding Framework for Auditory Working Memory
标题:听觉工作记忆的通用闭环预测编码框架
链接:https://arxiv.org/abs/2503.12506
标题: 域不变语音分离的上下文感知两步训练方案
链接:https://arxiv.org/abs/2503.12589
摘要:语音分离试图从多讲话语音混合物中分离单独的语音信号。尽管取得了很大进展,但在合成数据上训练良好的系统通常会在域外数据(如真实世界的语音混合)上遇到性能下降。为了解决这个问题,我们引入了一种新的上下文感知,两阶段的语音分离模型的训练计划。在该训练方案中,传统的端到端架构被替换为包含上下文提取器和分离器的框架。这两个模块被逐步训练,以模拟听觉系统的语音分离过程。我们通过对合成和真实世界语音混合物的跨域实验来评估所提出的训练方案,并证明我们的新方案有效地提高了不同域的分离质量,而无需自适应,如通过信号质量指标和字错误率(WER)所测量的。此外,对真实测试集的消融研究强调,上下文信息,包括来自预训练SSL模型的音素和单词表示,可以作为分离模型的有效域不变训练目标。
摘要:Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences performance degradation on out-of-domain data, such as real-world speech mixtures. To address this, we introduce a novel context-aware, two-stage training scheme for speech separation models. In this training scheme, the conventional end-to-end architecture is replaced with a framework that contains a context extractor and a segregator. The two modules are trained step by step to simulate the speech separation process of an auditory system. We evaluate the proposed training scheme through cross-domain experiments on both synthetic and real-world speech mixtures, and demonstrate that our new scheme effectively boosts separation quality across different domains without adaptation, as measured by signal quality metrics and word error rate (WER). Additionally, an ablation study on the real test set highlights that the context information, including phoneme and word representations from pretrained SSL models, serves as effective domain invariant training targets for separation models.
标题: 小夜曲:基于音频填充的演唱风格转换框架
链接:https://arxiv.org/abs/2503.12388
备注:Preprint under review
摘要:我们提出了小夜曲,一个新的框架的歌唱风格转换(SSC)的任务。虽然歌手身份转换在过去的几年里取得了很大的进展,转换歌手的演唱风格一直是一个未探索的研究领域。我们发现三个主要的挑战,在SSC:建模的目标风格,解开源风格,并保留源旋律。为了对目标歌唱风格进行建模,我们使用音频填充任务,通过使用掩蔽的目标梅尔频谱图的补充以及解开的声学特征的流匹配模型来预测目标梅尔频谱图的掩蔽段。另一方面,为了解开源歌唱风格,我们使用循环训练方法,其中我们使用合成转换样本作为源输入,并重建原始源梅尔声谱图作为目标。最后,为了更好地保留源旋律,我们研究了使用基于源滤波器的声码器的后处理模块,并使用原始的F0模式重新合成转换后的波形。我们的研究结果表明,小夜曲框架可以处理广义SSC任务的最佳整体相似性得分,特别是在建模呼吸和混合演唱风格。此外,虽然重新合成与原来的F0模式减轻了走调唱歌,提高了自然度,我们发现,由于没有改变F0模式到目标风格的相似性略有权衡。
摘要:We propose Serenade, a novel framework for the singing style conversion (SSC) task. Although singer identity conversion has made great strides in the previous years, converting the singing style of a singer has been an unexplored research area. We find three main challenges in SSC: modeling the target style, disentangling source style, and retaining the source melody. To model the target singing style, we use an audio infilling task by predicting a masked segment of the target mel-spectrogram with a flow-matching model using the complement of the masked target mel-spectrogram along with disentangled acoustic features. On the other hand, to disentangle the source singing style, we use a cyclic training approach, where we use synthetic converted samples as source inputs and reconstruct the original source mel-spectrogram as a target. Finally, to retain the source melody better, we investigate a post-processing module using a source-filter-based vocoder and resynthesize the converted waveforms using the original F0 patterns. Our results showed that the Serenade framework can handle generalized SSC tasks with the best overall similarity score, especially in modeling breathy and mixed singing styles. Moreover, although resynthesizing with the original F0 patterns alleviated out-of-tune singing and improved naturalness, we found a slight tradeoff in similarity due to not changing the F0 patterns into the target style.
标题: 视听情感识别的弱互补关系处理
链接:https://arxiv.org/abs/2503.12261
备注:Submission to valence arousal track of 8th ABAW competition. arXiv admin note: substantial text overlap with arXiv:2403.13659
摘要:多模态情感识别最近引起了情感计算的极大兴趣,因为它具有超越孤立的单峰方法的巨大潜力。音频和视觉模态是视频中两个主要的非接触通道,它们通常被期望彼此具有互补关系。然而,音频和视频通道可能并不总是彼此互补,导致较差的音频-视频特征表示,从而降低系统的性能。在本文中,我们提出了一个灵活的视听融合模型,可以适应弱互补关系,使用门控注意机制。具体来说,我们扩展了递归联合交叉注意模型,通过在每次迭代中引入门控机制来控制输入特征和关注特征之间的信息流,这取决于它们互补关系的强度。例如,如果模态表现出强互补关系,则选通机制选择交叉参与的特征,否则选择非参与的特征。为了进一步提高系统的性能,我们进一步引入了阶段门控机制,用于控制每次迭代的门控输出之间的信息流。因此,所提出的模型提高了系统的性能,即使当音频和视觉模态不具有很强的互补关系,通过增加更多的灵活性,以递归联合交叉注意机制。该模型已在具有挑战性的Affwild 2数据集上进行了评估,并显着优于最先进的融合方法。
摘要:Multimodal emotion recognition has recently drawn a lot of interest in affective computing as it has immense potential to outperform isolated unimodal approaches. Audio and visual modalities are two predominant contact-free channels in videos, which are often expected to carry a complementary relationship with each other. However, audio and visual channels may not always be complementary with each other, resulting in poor audio-visual feature representations, thereby degrading the performance of the system. In this paper, we propose a flexible audio-visual fusion model that can adapt to weak complementary relationships using a gated attention mechanism. Specifically, we extend the recursive joint cross-attention model by introducing gating mechanism in every iteration to control the flow of information between the input features and the attended features depending on the strength of their complementary relationship. For instance, if the modalities exhibit strong complementary relationships, the gating mechanism chooses cross-attended features, otherwise non-attended features. To further improve the performance of the system, we further introduce stage gating mechanism, which is used to control the flow of information across the gated outputs of each iteration. Therefore, the proposed model improves the performance of the system even when the audio and visual modalities do not have a strong complementary relationship with each other by adding more flexibility to the recursive joint cross attention mechanism. The proposed model has been evaluated on the challenging Affwild2 dataset and significantly outperforms the state-of-the-art fusion approaches.
标题:DiffGAP:对比空间中用于弥合跨模型差距的轻量级扩散模块
链接:https://arxiv.org/abs/2503.12131
摘要:最近在跨模态理解和生成方面的工作,特别是通过CLAP(对比音频预训练)和CAVP(对比视听预训练)等模型,通过单一的对比损失显着增强了文本,视频和音频嵌入的对齐。然而,这些方法往往忽略了双向的相互作用和固有的噪音存在于每一个模态,这可能会严重影响跨模态集成的质量和效率。为了解决这个问题,我们引入了DiffGAP,一种新的方法,将一个轻量级的生成模块内的对比空间。具体来说,我们的DiffGAP采用了双向扩散过程,以更有效地桥接跨模态间隙。这涉及到以音频嵌入为条件的文本和视频嵌入的去噪过程,反之亦然,从而促进了更细致和鲁棒的跨模态交互。我们在VGGSound和AudioCaps数据集上的实验结果表明,DiffGAP显著提高了视频/文本-音频生成和检索任务的性能,证实了其在增强跨模态理解和生成能力方面的有效性。
摘要:Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignment of text, video, and audio embeddings via a single contrastive loss. However, these methods often overlook the bidirectional interactions and inherent noises present in each modality, which can crucially impact the quality and efficacy of cross-modal integration. To address this limitation, we introduce DiffGAP, a novel approach incorporating a lightweight generative module within the contrastive space. Specifically, our DiffGAP employs a bidirectional diffusion process tailored to bridge the cross-modal gap more effectively. This involves a denoising process on text and video embeddings conditioned on audio embeddings and vice versa, thus facilitating a more nuanced and robust cross-modal interaction. Our experimental results on VGGSound and AudioCaps datasets demonstrate that DiffGAP significantly improves performance in video/text-audio generation and retrieval tasks, confirming its effectiveness in enhancing cross-modal understanding and generation capabilities.
标题: 通过低比特率神经编解码器和预训练表示的通用语音令牌学习
链接:https://arxiv.org/abs/2503.12115
备注:Accepted by IEEE Journal of Selected Topics in Signal Processing(JSTSP)
摘要:当前的大型语音语言模型主要基于来自自监督学习表示的离散化的语义令牌和来自神经编解码器的声学令牌,遵循语义建模和声学合成范例。然而,语义标记丢弃了对自然口语通信很重要的说话者的非语言属性,而基于语义标记的声学合成在恢复非语言细节方面具有限制,并且遭受鲁棒性问题,特别是当提示和目标之间存在域间隙时。本文统一了两种类型的令牌,并提出了UniCodec,一个通用的语音令牌学习,封装语音的所有语义,包括语言和非语言信息,到一个紧凑的和语义分离的统一令牌。这样一个统一的标记不仅有利于语音语言模型的理解与语言暗示,而且有助于高质量的语音生成输出。利用低比特率神经编解码器来学习全局和局部尺度上的此类解纠缠离散表示,并从自监督学习的特征中提取知识。多语言数据集上的广泛评估表明,它在生成自然,表达和长期一致的输出质量的有效性,以及保留在几个语音处理任务中的非语言属性。
摘要:Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, following a semantic-modeling and acoustic-synthesis paradigm. However, semantic tokens discard paralinguistic attributes of speakers that is important for natural spoken communication, while prompt-based acoustic synthesis from semantic tokens has limits in recovering paralinguistic details and suffers from robustness issues, especially when there are domain gaps between the prompt and the target. This paper unifies two types of tokens and proposes the UniCodec, a universal speech token learning that encapsulates all semantics of speech, including linguistic and paralinguistic information, into a compact and semantically-disentangled unified token. Such a unified token can not only benefit speech language models in understanding with paralinguistic hints but also help speech generation with high-quality output. A low-bitrate neural codec is leveraged to learn such disentangled discrete representations at global and local scales, with knowledge distilled from self-supervised learned features. Extensive evaluations on multilingual datasets demonstrate its effectiveness in generating natural, expressive and long-term consistent output quality with paralinguistic attributes well preserved in several speech processing tasks.
标题: 表现性音乐数据处理和生成
链接:https://arxiv.org/abs/2503.11896
备注:7 pages, 4 figures
摘要:音乐表现力和连贯性在音乐创作和演奏中不可或缺,但在现代人工智能生成模型中往往被忽视。在这项工作中,我们介绍了一种基于数据处理技术,捕捉音乐表现的表现力。这种源自韦伯定律的技术反映了人类听觉的真实性,并在训练输入中保留了音乐的微妙性和表现力。为了促进音乐的连贯性,我们在神经网络中基于概率链规则对音乐数据中的多个参数(如音高、持续时间、速度等)之间的输出相互依赖性进行建模。在实践中,我们将多输出序列模型分解为单输出子模型,并将先前采样的输出条件化到后续子模型上,以诱导条件分布。最后,提出了一种基于输出熵的序列选择方法。熵序列被设置为一个标准,以选择可预测的和稳定的世代,这是进一步研究的信息审美措施的背景下,量化音乐的乐趣和信息增益沿音乐的倾向。
摘要:Musical expressivity and coherence are indispensable in music composition and performance, while often neglected in modern AI generative models. In this work, we introduce a listening-based data-processing technique that captures the expressivity in musical performance. This technique derived from Weber's law reflects the human perceptual truth of listening and preserves musical subtlety and expressivity in the training input. To facilitate musical coherence, we model the output interdependencies among multiple arguments in the music data such as pitch, duration, velocity, etc. in the neural networks based on the probabilistic chain rule. In practice, we decompose the multi-output sequential model into single-output submodels and condition previously sampled outputs on the subsequent submodels to induce conditional distributions. Finally, to select eligible sequences from all generations, a tentative measure based on the output entropy was proposed. The entropy sequence is set as a criterion to select predictable and stable generations, which is further studied under the context of informational aesthetic measures to quantify musical pleasure and information gain along the music tendency.
标题:空间音频和房间声学的过去、现在和未来
链接:https://arxiv.org/abs/2503.12948
备注:Accepted to International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2025
摘要:空间音频和室内声学的研究旨在通过模拟声音在空间中的行为的物理和心理声学来创建沉浸式音频体验。在这一研究领域的漫长历史中,各种关键技术都是基于理论进步和实践创新而开发的。我们突出的历史成就,倡议活动,最近的进展,并在空间音频记录和再现,室内声学模拟,建模,分析和控制的研究领域的未来展望。
摘要:The study of spatial audio and room acoustics aims to create immersive audioexperiences by modeling the physics and psychoacoustics of how sound behaves inspace. In the long history of this research area, various key technologies havebeen developed based both on theoretical advancements and practicalinnovations. We highlight historical achievements, initiative activities,recent advancements, and future outlooks in the research area of spatial audiorecording and reproduction, and room acoustic simulation, modeling, analysis,and control.
标题:FNSE-SBGAN:使用Schrodinger桥和生成对抗网络的远场语音增强
链接:https://arxiv.org/abs/2503.12936
备注:13 pages, 6 figures
摘要:目前神经语音增强的主要方法依赖于使用模拟的远场噪声混响语音(混合)和干净语音对的纯监督深度学习。然而,这些训练的模型往往表现出有限的泛化到真实记录的混合物。为了解决这个问题,本研究直接在真实混合物上研究训练增强模型。具体来说,我们重新审视单通道远场到近场语音增强(FNSE)任务,重点是现实世界的数据,其特征在于低信噪比(SNR),高混响,中高频衰减。我们提出了FNSE-SBGAN,这是一个新的框架,它将基于薛定谔桥(SB)的扩散模型与生成对抗网络(GANs)集成在一起。我们的方法在各种指标和主观评估方面实现了最先进的性能,与远场信号相比,字符错误率(CER)显著降低了14.58%。实验结果表明,FNSE-SBGAN保留了优越的主观质量,并建立了一个新的基准,为现实世界的远场语音增强。此外,我们引入了一个新的评估框架,利用矩阵秩分析在时间-频率域,提供系统的见解模型的性能和揭示不同的生成方法的优点和缺点。
摘要:The current dominant approach for neural speech enhancement relies onpurely-supervised deep learning using simulated pairs of far-fieldnoisy-reverberant speech (mixtures) and clean speech. However, these trainedmodels often exhibit limited generalizability to real-recorded mixtures. Toaddress this issue, this study investigates training enhancement modelsdirectly on real mixtures. Specifically, we revisit the single-channelfar-field to near-field speech enhancement (FNSE) task, focusing on real-worlddata characterized by low signal-to-noise ratio (SNR), high reverberation, andmid-to-high frequency attenuation. We propose FNSE-SBGAN, a novel frameworkthat integrates a Schrodinger Bridge (SB)-based diffusion model with generativeadversarial networks (GANs). Our approach achieves state-of-the-art performanceacross various metrics and subjective evaluations, significantly reducing thecharacter error rate (CER) by up to 14.58% compared to far-field signals.Experimental results demonstrate that FNSE-SBGAN preserves superior subjectivequality and establishes a new benchmark for real-world far-field speechenhancement. Additionally, we introduce a novel evaluation framework leveragingmatrix rank analysis in the time-frequency domain, providing systematicinsights into model performance and revealing the strengths and weaknesses ofdifferent generative methods.
标题:自适应混合专家学习以实现稳健的音频欺骗检测
链接:https://arxiv.org/abs/2503.12010
备注:Submitted to Eusipco 2025, 5 pages, 1 figure, 5 table
摘要:在音频欺骗检测中,大多数研究依赖于干净的数据集,使得模型容易受到真实世界的后处理攻击,例如通道压缩和噪声。为了克服这一挑战,我们提出了自适应混合专家学习(AMEL)框架,该框架通过利用特定于攻击的知识和动态适应不同的攻击条件来增强弹性。具体来说,AMEL利用了通过低秩自适应(LoRA)进行微调的攻击特定专家(ASE),使每个专家能够针对特定的后处理模式,同时只需要完全微调所需参数的1.12%。此外,我们引入动态专家聚合(DEA),自适应地选择和整合专家知识,以提高欺骗检测的鲁棒性。实验结果表明,AMEL显着提高了鲁棒性,提高噪声弹性和表现出更大的适应性,以前看不见的后处理方法相比,依赖于完全微调的模型。此外,我们的框架在各种混合攻击下的性能优于单个专家和简单平均集成,证明了其在管理复杂的现实世界条件下的卓越鲁棒性和适应性。
摘要:In audio spoofing detection, most studies rely on clean datasets, makingmodels susceptible to real-world post-processing attacks, such as channelcompression and noise. To overcome this challenge, we propose the AdaptiveMixture of Experts Learning (AMEL) framework, which enhances resilience byleveraging attack-specific knowledge and adapting dynamically to varied attackconditions. Specifically, AMEL utilizes Attack-Specific Experts (ASE)fine-tuned with Low-Rank Adaptation (LoRA), enabling each expert to targetspecific post-processing patterns while requiring only 1.12\% of the parametersneeded for full fine-tuning. Furthermore, we introduce Dynamic ExpertAggregation (DEA), which adaptively selects and integrates expert knowledge toenhance the robustness of spoofing detection. Experimental results demonstratethat AMEL significantly enhances robustness by improving noise resilience andexhibiting greater adaptability to previously unseen post-processing methodscompared to models relying on full fine-tuning. Additionally, our frameworkoutperforms both single expert and simple average ensemble under various mixedattacks, demonstrating its superior robustness and adaptability in managingcomplex, real-world conditions.
标题:AV-Surf:表面增强的几何感知新视图声学合成
链接:https://arxiv.org/abs/2503.12806
摘要:在复杂的真实世界环境中准确地建模声音传播对于新视图声学合成(NVAS)至关重要。虽然以前的研究已经利用视觉感知来估计空间声学,但在声学建模中结合使用来自3D表示的表面法线和结构细节一直未被充分探索。考虑到它们对声波反射和传播的直接影响,表面法线应与结构细节联合建模,以实现精确的空间声学。在本文中,我们提出了一种表面增强的几何感知方法NVAS,以提高空间声学建模。为了实现这一点,我们利用几何先验,如图像,深度图,表面法线,和点云使用3D高斯溅射(3DGS)为基础的框架。我们引入了一个基于双交叉注意的Transformer,将几何约束集成到频率查询中,以了解发射器的周围环境。此外,我们设计了一个基于ConvNeXT的频谱特征处理网络,称为频谱细化网络(SRN),以合成逼真的双耳音频。RWAVS和SoundSpace数据集上的实验结果突出了我们方法的必要性,因为它在新视图声学合成中超越了现有方法。
摘要:Accurately modeling sound propagation with complex real-world environments isessential for Novel View Acoustic Synthesis (NVAS). While previous studies haveleveraged visual perception to estimate spatial acoustics, the combined use ofsurface normal and structural details from 3D representations in acousticmodeling has been underexplored. Given their direct impact on sound wavereflections and propagation, surface normals should be jointly modeled withstructural details to achieve accurate spatial acoustics. In this paper, wepropose a surface-enhanced geometry-aware approach for NVAS to improve spatialacoustic modeling. To achieve this, we exploit geometric priors, such as image,depth map, surface normals, and point clouds obtained using a 3D GaussianSplatting (3DGS) based framework. We introduce a dual cross-attention-basedtransformer integrating geometrical constraints into frequency query tounderstand the surroundings of the emitter. Additionally, we design aConvNeXt-based spectral features processing network called Spectral RefinementNetwork (SRN) to synthesize realistic binaural audio. Experimental results onthe RWAVS and SoundSpace datasets highlight the necessity of our approach, asit surpasses existing methods in novel view acoustic synthesis.
标题:域不变语音分离的上下文感知两步训练方案
链接:https://arxiv.org/abs/2503.12589
摘要:语音分离试图从多讲话语音混合物中分离单独的语音信号。尽管取得了很大进展,但在合成数据上训练良好的系统通常会在域外数据(如真实世界的语音混合)上遇到性能下降。为了解决这个问题,我们引入了一种新的上下文感知,两阶段的语音分离模型的训练计划。在该训练方案中,传统的端到端架构被替换为包含上下文提取器和分离器的框架。这两个模块被逐步训练,以模拟听觉系统的语音分离过程。我们通过对合成和真实世界语音混合物的跨域实验来评估所提出的训练方案,并证明我们的新方案有效地提高了不同域的分离质量,而无需自适应,如通过信号质量指标和字错误率(WER)所测量的。此外,对真实测试集的消融研究强调,上下文信息,包括来自预训练SSL模型的音素和单词表示,可以作为分离模型的有效域不变训练目标。
摘要:Speech separation seeks to isolate individual speech signals from amulti-talk speech mixture. Despite much progress, a system well-trained onsynthetic data often experiences performance degradation on out-of-domain data,such as real-world speech mixtures. To address this, we introduce a novelcontext-aware, two-stage training scheme for speech separation models. In thistraining scheme, the conventional end-to-end architecture is replaced with aframework that contains a context extractor and a segregator. The two modulesare trained step by step to simulate the speech separation process of anauditory system. We evaluate the proposed training scheme through cross-domainexperiments on both synthetic and real-world speech mixtures, and demonstratethat our new scheme effectively boosts separation quality across differentdomains without adaptation, as measured by signal quality metrics and worderror rate (WER). Additionally, an ablation study on the real test sethighlights that the context information, including phoneme and wordrepresentations from pretrained SSL models, serves as effective domaininvariant training targets for separation models.
标题:小夜曲:基于音频填充的演唱风格转换框架
链接:https://arxiv.org/abs/2503.12388
备注:Preprint under review
摘要:我们提出了小夜曲,一个新的框架的歌唱风格转换(SSC)的任务。虽然歌手身份转换在过去的几年里取得了很大的进展,转换歌手的演唱风格一直是一个未探索的研究领域。我们发现三个主要的挑战,在SSC:建模的目标风格,解开源风格,并保留源旋律。为了对目标歌唱风格进行建模,我们使用音频填充任务,通过使用掩蔽的目标梅尔频谱图的补充以及解开的声学特征的流匹配模型来预测目标梅尔频谱图的掩蔽段。另一方面,为了解开源歌唱风格,我们使用循环训练方法,其中我们使用合成转换样本作为源输入,并重建原始源梅尔声谱图作为目标。最后,为了更好地保留源旋律,我们研究了使用基于源滤波器的声码器的后处理模块,并使用原始的F0模式重新合成转换后的波形。我们的研究结果表明,小夜曲框架可以处理广义SSC任务的最佳整体相似性得分,特别是在建模呼吸和混合演唱风格。此外,虽然重新合成与原来的F0模式减轻了走调唱歌,提高了自然度,我们发现,由于没有改变F0模式到目标风格的相似性略有权衡。
摘要:We propose Serenade, a novel framework for the singing style conversion (SSC)task. Although singer identity conversion has made great strides in theprevious years, converting the singing style of a singer has been an unexploredresearch area. We find three main challenges in SSC: modeling the target style,disentangling source style, and retaining the source melody. To model thetarget singing style, we use an audio infilling task by predicting a maskedsegment of the target mel-spectrogram with a flow-matching model using thecomplement of the masked target mel-spectrogram along with disentangledacoustic features. On the other hand, to disentangle the source singing style,we use a cyclic training approach, where we use synthetic converted samples assource inputs and reconstruct the original source mel-spectrogram as a target.Finally, to retain the source melody better, we investigate a post-processingmodule using a source-filter-based vocoder and resynthesize the convertedwaveforms using the original F0 patterns. Our results showed that the Serenadeframework can handle generalized SSC tasks with the best overall similarityscore, especially in modeling breathy and mixed singing styles. Moreover,although resynthesizing with the original F0 patterns alleviated out-of-tunesinging and improved naturalness, we found a slight tradeoff in similarity dueto not changing the F0 patterns into the target style.
标题:视听情感识别的弱互补关系处理
链接:https://arxiv.org/abs/2503.12261
备注:Submission to valence arousal track of 8th ABAW competition. arXiv admin note: substantial text overlap with arXiv:2403.13659
摘要:多模态情感识别最近引起了情感计算的极大兴趣,因为它具有超越孤立的单峰方法的巨大潜力。音频和视觉模态是视频中两个主要的非接触通道,它们通常被期望彼此具有互补关系。然而,音频和视频通道可能并不总是彼此互补,导致较差的音频-视频特征表示,从而降低系统的性能。在本文中,我们提出了一个灵活的视听融合模型,可以适应弱互补关系,使用门控注意机制。具体来说,我们扩展了递归联合交叉注意模型,通过在每次迭代中引入门控机制来控制输入特征和关注特征之间的信息流,这取决于它们互补关系的强度。例如,如果模态表现出强互补关系,则选通机制选择交叉参与的特征,否则选择非参与的特征。为了进一步提高系统的性能,我们进一步引入了阶段门控机制,用于控制每次迭代的门控输出之间的信息流。因此,所提出的模型提高了系统的性能,即使当音频和视觉模态不具有很强的互补关系,通过增加更多的灵活性,以递归联合交叉注意机制。该模型已在具有挑战性的Affwild 2数据集上进行了评估,并显着优于最先进的融合方法。
摘要:Multimodal emotion recognition has recently drawn a lot of interest inaffective computing as it has immense potential to outperform isolated unimodalapproaches. Audio and visual modalities are two predominant contact-freechannels in videos, which are often expected to carry a complementaryrelationship with each other. However, audio and visual channels may not alwaysbe complementary with each other, resulting in poor audio-visual featurerepresentations, thereby degrading the performance of the system. In thispaper, we propose a flexible audio-visual fusion model that can adapt to weakcomplementary relationships using a gated attention mechanism. Specifically, weextend the recursive joint cross-attention model by introducing gatingmechanism in every iteration to control the flow of information between theinput features and the attended features depending on the strength of theircomplementary relationship. For instance, if the modalities exhibit strongcomplementary relationships, the gating mechanism chooses cross-attendedfeatures, otherwise non-attended features. To further improve the performanceof the system, we further introduce stage gating mechanism, which is used tocontrol the flow of information across the gated outputs of each iteration.Therefore, the proposed model improves the performance of the system even whenthe audio and visual modalities do not have a strong complementary relationshipwith each other by adding more flexibility to the recursive joint crossattention mechanism. The proposed model has been evaluated on the challengingAffwild2 dataset and significantly outperforms the state-of-the-art fusionapproaches.
标题:DiffGAP:对比空间中用于弥合跨模型差距的轻量级扩散模块
链接:https://arxiv.org/abs/2503.12131
摘要:最近在跨模态理解和生成方面的工作,特别是通过CLAP(对比音频预训练)和CAVP(对比视听预训练)等模型,通过单一的对比损失显着增强了文本,视频和音频嵌入的对齐。然而,这些方法往往忽略了双向的相互作用和固有的噪音存在于每一个模态,这可能会严重影响跨模态集成的质量和效率。为了解决这个问题,我们引入了DiffGAP,一种新的方法,将一个轻量级的生成模块内的对比空间。具体来说,我们的DiffGAP采用了双向扩散过程,以更有效地桥接跨模态间隙。这涉及到以音频嵌入为条件的文本和视频嵌入的去噪过程,反之亦然,从而促进了更细致和鲁棒的跨模态交互。我们在VGGSound和AudioCaps数据集上的实验结果表明,DiffGAP显著提高了视频/文本-音频生成和检索任务的性能,证实了其在增强跨模态理解和生成能力方面的有效性。
摘要:Recent works in cross-modal understanding and generation, notably throughmodels like CLAP (Contrastive Language-Audio Pretraining) and CAVP (ContrastiveAudio-Visual Pretraining), have significantly enhanced the alignment of text,video, and audio embeddings via a single contrastive loss. However, thesemethods often overlook the bidirectional interactions and inherent noisespresent in each modality, which can crucially impact the quality and efficacyof cross-modal integration. To address this limitation, we introduce DiffGAP, anovel approach incorporating a lightweight generative module within thecontrastive space. Specifically, our DiffGAP employs a bidirectional diffusionprocess tailored to bridge the cross-modal gap more effectively. This involvesa denoising process on text and video embeddings conditioned on audioembeddings and vice versa, thus facilitating a more nuanced and robustcross-modal interaction. Our experimental results on VGGSound and AudioCapsdatasets demonstrate that DiffGAP significantly improves performance invideo/text-audio generation and retrieval tasks, confirming its effectivenessin enhancing cross-modal understanding and generation capabilities.
标题:通过低比特率神经编解码器和预训练表示的通用语音令牌学习
链接:https://arxiv.org/abs/2503.12115
备注:Accepted by IEEE Journal of Selected Topics in Signal Processing(JSTSP)
摘要:当前的大型语音语言模型主要基于来自自监督学习表示的离散化的语义令牌和来自神经编解码器的声学令牌,遵循语义建模和声学合成范例。然而,语义标记丢弃了对自然口语通信很重要的说话者的非语言属性,而基于语义标记的声学合成在恢复非语言细节方面具有限制,并且遭受鲁棒性问题,特别是当提示和目标之间存在域间隙时。本文统一了两种类型的令牌,并提出了UniCodec,一个通用的语音令牌学习,封装语音的所有语义,包括语言和非语言信息,到一个紧凑的和语义分离的统一令牌。这样一个统一的标记不仅有利于语音语言模型的理解与语言暗示,而且有助于高质量的语音生成输出。利用低比特率神经编解码器来学习全局和局部尺度上的此类解纠缠离散表示,并从自监督学习的特征中提取知识。多语言数据集上的广泛评估表明,它在生成自然,表达和长期一致的输出质量的有效性,以及保留在几个语音处理任务中的非语言属性。
摘要:Current large speech language models are mainly based on semantic tokens fromdiscretization of self-supervised learned representations and acoustic tokensfrom a neural codec, following a semantic-modeling and acoustic-synthesisparadigm. However, semantic tokens discard paralinguistic attributes ofspeakers that is important for natural spoken communication, while prompt-basedacoustic synthesis from semantic tokens has limits in recovering paralinguisticdetails and suffers from robustness issues, especially when there are domaingaps between the prompt and the target. This paper unifies two types of tokensand proposes the UniCodec, a universal speech token learning that encapsulatesall semantics of speech, including linguistic and paralinguistic information,into a compact and semantically-disentangled unified token. Such a unifiedtoken can not only benefit speech language models in understanding withparalinguistic hints but also help speech generation with high-quality output.A low-bitrate neural codec is leveraged to learn such disentangled discreterepresentations at global and local scales, with knowledge distilled fromself-supervised learned features. Extensive evaluations on multilingualdatasets demonstrate its effectiveness in generating natural, expressive andlong-term consistent output quality with paralinguistic attributes wellpreserved in several speech processing tasks.
标题:表现性音乐数据处理和生成
链接:https://arxiv.org/abs/2503.11896
备注:7 pages, 4 figures
摘要:音乐表现力和连贯性在音乐创作和演奏中不可或缺,但在现代人工智能生成模型中往往被忽视。在这项工作中,我们介绍了一种基于数据处理技术,捕捉音乐表现的表现力。这种源自韦伯定律的技术反映了人类听觉的真实性,并在训练输入中保留了音乐的微妙性和表现力。为了促进音乐的连贯性,我们在神经网络中基于概率链规则对音乐数据中的多个参数(如音高、持续时间、速度等)之间的输出相互依赖性进行建模。在实践中,我们将多输出序列模型分解为单输出子模型,并将先前采样的输出条件化到后续子模型上,以诱导条件分布。最后,提出了一种基于输出熵的序列选择方法。熵序列被设置为一个标准,以选择可预测的和稳定的世代,这是进一步研究的信息审美措施的背景下,量化音乐的乐趣和信息增益沿音乐的倾向。
摘要:Musical expressivity and coherence are indispensable in music composition andperformance, while often neglected in modern AI generative models. In thiswork, we introduce a listening-based data-processing technique that capturesthe expressivity in musical performance. This technique derived from Weber'slaw reflects the human perceptual truth of listening and preserves musicalsubtlety and expressivity in the training input. To facilitate musicalcoherence, we model the output interdependencies among multiple arguments inthe music data such as pitch, duration, velocity, etc. in the neural networksbased on the probabilistic chain rule. In practice, we decompose themulti-output sequential model into single-output submodels and conditionpreviously sampled outputs on the subsequent submodels to induce conditionaldistributions. Finally, to select eligible sequences from all generations, atentative measure based on the output entropy was proposed. The entropysequence is set as a criterion to select predictable and stable generations,which is further studied under the context of informational aesthetic measuresto quantify musical pleasure and information gain along the music tendency.
