微信公众号:arXiv_Daily
cs.SD语音
【1】A Conditioned UNet for Music Source Separation
标题:用于音乐源分离的条件化UNet
链接:https://arxiv.org/abs/2512.15532
摘要:在本文中,我们提出了一个有条件的UNet音乐源分离(MSS)。MSS通常由多输出神经网络(通常为UNets)执行,每个输出表示来自预定义乐器词汇表的特定词干。相比之下,条件MSS网络接受与感兴趣的干相关的音频查询,以及从中提取干的信号。因此,不需要严格的词汇表,这使得MSS中的任务更加现实。由于缺乏合适的数据,这些任务的条件方法的潜力在某种程度上被隐藏了,最近MoisesDb数据集解决了这个问题。最近的一种方法Banquet使用了这个数据集,在更大的词汇表上看到了有希望的结果。Banquet使用Bandsplit RNN而不是UNet,作者指出UNet不应该适合有条件的MSS。我们反对这种说法,并提出QSCNet,一种新的有条件的UNet的MSS,集成了网络调节元素的稀疏压缩网络MSS。我们发现QSCNet在几个MSS任务上的性能优于Banquet超过1dB SNR,而使用的参数数量不到一半。
摘要:In this paper we propose a conditioned UNet for Music Source Separation (MSS). MSS is generally performed by multi-output neural networks, typically UNets, with each output representing a particular stem from a predefined instrument vocabulary. In contrast, conditioned MSS networks accept an audio query related to a stem of interest alongside the signal from which that stem is to be extracted. Thus, a strict vocabulary is not required and this enables more realistic tasks in MSS. The potential of conditioned approaches for such tasks has been somewhat hidden due to a lack of suitable data, an issue recently addressed with the MoisesDb dataset. A recent method, Banquet, employs this dataset with promising results seen on larger vocabularies. Banquet uses Bandsplit RNN rather than a UNet and the authors state that UNets should not be suitable for conditioned MSS. We counter this argument and propose QSCNet, a novel conditioned UNet for MSS that integrates network conditioning elements in the Sparse Compressed Network for MSS. We find QSCNet to outperform Banquet by over 1dB SNR on a couple of MSS tasks, while using less than half the number of parameters.
【2】Time-Varying Audio Effect Modeling by End-to-End Adversarial Training
标题:通过端到端对抗训练进行时变音效建模
链接:https://arxiv.org/abs/2512.15313
备注:Submitted for review to the Journal of the Audio Engineering Society (JAES). Accompanying website: https://ybourdin.github.io/sptvmod
摘要:深度学习已经成为音频效果建模的标准方法,但严格的黑盒建模对于时变系统仍然存在问题。与时不变效应不同,在具有内部调制的设备上训练模型通常需要记录或提取控制信号,以确保标准损失函数所需的时间对准。本文介绍了一种生成对抗网络(GAN)框架,仅使用输入-输出音频记录来对此类效果进行建模,无需提取调制信号。我们提出了一种通过两阶段策略训练的卷积递归架构:初始对抗阶段允许模型在没有严格相位约束的情况下学习调制行为的分布,然后是监督微调阶段,其中状态预测网络(SPN)估计将模型与目标同步所需的初始内部状态。此外,一个新的客观度量的基础上啁啾序列信号的调制精度进行量化。一个老式的硬件相位器建模实验表明,该方法的能力,捕捉时变动态在一个完全黑盒的上下文中。
摘要:Deep learning has become a standard approach for the modeling of audio effects, yet strictly black-box modeling remains problematic for time-varying systems. Unlike time-invariant effects, training models on devices with internal modulation typically requires the recording or extraction of control signals to ensure the time-alignment required by standard loss functions. This paper introduces a Generative Adversarial Network (GAN) framework to model such effects using only input-output audio recordings, removing the need for modulation signal extraction. We propose a convolutional-recurrent architecture trained via a two-stage strategy: an initial adversarial phase allows the model to learn the distribution of the modulation behavior without strict phase constraints, followed by a supervised fine-tuning phase where a State Prediction Network (SPN) estimates the initial internal states required to synchronize the model with the target. Additionally, a new objective metric based on chirp-train signals is developed to quantify modulation accuracy. Experiments modeling a vintage hardware phaser demonstrate the method's ability to capture time-varying dynamics in a fully black-box context.
【3】O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization
标题:O-EENC-SD:用于说话者拨号化的高效在线端到端神经集群
链接:https://arxiv.org/abs/2512.15229
摘要:我们介绍O-EENC-SD:一个基于EEND-EDA的端到端在线说话人日志化系统,采用了一种新的基于RNN的拼接机制进行在线预测。特别是,我们开发了一种新的质心细化解码器,其有用性是通过严格的消融研究进行评估。与现有方法相比,我们的系统具有关键优势:与无监督聚类方法相比,无超参数解决方案,以及与计算成本高的当前在线端到端方法相比,更有效的替代方案。我们证明,O-EENC-SD是有竞争力的两个扬声器的对话电话语音域的最先进的,测试的CallHome数据集。我们的研究结果表明,O-EENC-SD提供了DER和复杂性之间的一个很好的权衡,即使在没有重叠的独立块上工作,也使系统非常高效。
摘要:We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness is assessed through a rigorous ablation study. Our system provides key advantages over existing methods: a hyperparameter-free solution compared to unsupervised clustering approaches, and a more efficient alternative to current online end-to-end methods, which are computationally costly. We demonstrate that O-EENC-SD is competitive with the state of the art in the two-speaker conversational telephone speech domain, as tested on the CallHome dataset. Our results show that O-EENC-SD provides a great trade-off between DER and complexity, even when working on independent chunks with no overlap, making the system extremely efficient.
【4】BEAT2AASIST model with layer fusion for ESDD 2026 Challenge
标题:BEAT2AASIST模型与ESDD 2026挑战赛的层融合
链接:https://arxiv.org/abs/2512.15180
备注:3 pages, 1 figure, challenge paper
摘要:音频生成的最新进展增加了现实环境声音操纵的风险,促使ESDD 2026挑战赛成为环境声音深度伪造检测(ESDD)的第一个大规模基准。我们提出了BEAT 2AASIST,它扩展了BEATs-AASIST,通过沿频率或信道维度拆分BEAT导出的表示,并使用双AASIST分支处理它们。为了丰富特征表示,我们使用级联、CNN门控和SE门控策略将前k个Transformer层融合。此外,基于声码器的数据增强应用,以提高对看不见的欺骗方法的鲁棒性。官方测试集上的实验结果表明,所提出的方法在整个挑战轨道上取得了有竞争力的性能。
摘要:Recent advances in audio generation have increased the risk of realistic environmental sound manipulation, motivating the ESDD 2026 Challenge as the first large-scale benchmark for Environmental Sound Deepfake Detection (ESDD). We propose BEAT2AASIST which extends BEATs-AASIST by splitting BEATs-derived representations along frequency or channel dimension and processing them with dual AASIST branches. To enrich feature representations, we incorporate top-k transformer layer fusion using concatenation, CNN-gated, and SE-gated strategies. In addition, vocoder-based data augmentation is applied to improve robustness against unseen spoofing methods. Experimental results on the official test sets demonstrate that the proposed approach achieves competitive performance across the challenge tracks.
【5】Synaspot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy
标题:Synaspot:一个轻量级的流媒体多模式框架,用于关键词识别与音频文本协同
链接:https://arxiv.org/abs/2512.15124
摘要:连续语音流中的开放词汇表关键词识别(KWS)在广泛的实际应用中具有重要的实用价值。虽然人们越来越注意不同模式在知识和预警系统中的作用,但它们的有效性已得到承认。然而,多模式集成增加的参数成本和端到端部署的限制限制了这种模型的实用性。为了解决这些挑战,我们提出了一个轻量级的,流多模态框架。首先,我们专注于多模态注册功能和减少特定于说话人(声纹)的信息,在语音注册提取说话人无关的特征。其次,我们有效地融合语音和文本特征。最后,我们引入了一个流解码框架,它只需要编码器提取特征,然后用我们的三种模态表示进行数学解码。LibriPhase和WenetPrase上的实验证明了我们的模型的性能。与现有的流媒体方法相比,我们的方法实现了更好的性能与显着更少的参数。
摘要:Open-vocabulary keyword spotting (KWS) in continuous speech streams holds significant practical value across a wide range of real-world applications. While increasing attention has been paid to the role of different modalities in KWS, their effectiveness has been acknowledged. However, the increased parameter cost from multimodal integration and the constraints of end-to-end deployment have limited the practical applicability of such models. To address these challenges, we propose a lightweight, streaming multi-modal framework. First, we focus on multimodal enrollment features and reduce speaker-specific (voiceprint) information in the speech enrollment to extract speaker-irrelevant characteristics. Second, we effectively fuse speech and text features. Finally, we introduce a streaming decoding framework that only requires the encoder to extract features, which are then mathematically decoded with our three modal representations. Experiments on LibriPhase and WenetPrase demonstrate the performance of our model. Compared to existing streaming approaches, our method achieves better performance with significantly fewer parameters.
【6】Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
标题:自适应多模式人员识别:处理缺失模式的稳健框架
链接:https://arxiv.org/abs/2512.14961
备注:10 pages and 8 tables
摘要:人员识别系统通常依赖于音频,视觉或行为线索,但现实世界的条件经常导致丢失或降级的模态。为了解决这一挑战,我们提出了一个三模态的人识别框架,集成语音,面部和手势模态,同时保持强大的模态损失。我们的方法利用多任务学习来独立处理每种模态,然后通过交叉注意和门控融合机制来促进跨模态的交互。此外,置信加权融合策略动态适应缺失和低质量数据,即使在单峰或双峰场景下也能确保最佳分类。我们在CANDOR上评估我们的方法,CANDOR是一个新引入的基于访谈的多模态数据集,我们首次进行基准测试。我们的研究结果表明,所提出的三模态系统在人员识别任务上达到了99.18%的Top-1准确率,优于传统的单峰和后期融合方法。此外,我们在VoxCeleb 1数据集上评估了我们的模型作为基准,并在双峰模式下达到了99.92%的准确率。此外,我们表明,即使在一个或两个模态不可用时,我们的系统也保持了很高的准确性,使其成为现实世界中人员识别应用的鲁棒解决方案。这项工作的代码和数据是公开的。
摘要:Person recognition systems often rely on audio, visual, or behavioral cues, but real-world conditions frequently result in missing or degraded modalities. To address this challenge, we propose a Trimodal person identification framework that integrates voice, face, and gesture modalities, while remaining robust to modality loss. Our approach leverages multi-task learning to process each modality independently, followed by a cross-attention and gated fusion mechanisms to facilitate interaction across modalities. Moreover, a confidence-weighted fusion strategy dynamically adapts to missing and low-quality data, ensuring optimal classification even in Unimodal or Bimodal scenarios. We evaluate our method on CANDOR, a newly introduced interview-based multimodal dataset, which we benchmark for the first time. Our results demonstrate that the proposed Trimodal system achieves 99.18% Top-1 accuracy on person identification tasks, outperforming conventional Unimodal and late-fusion approaches. In addition, we evaluate our model on the VoxCeleb1 dataset as a benchmark and reach 99.92% accuracy in Bimodal mode. Moreover, we show that our system maintains high accuracy even when one or two modalities are unavailable, making it a robust solution for real-world person recognition applications. The code and data for this work are publicly available.
【7】TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
标题:TalkVerse:一分钟音频驱动视频生成民主化
链接:https://arxiv.org/abs/2512.14938
备注:open-sourced single-person full-body talking video generation dataset, training code and checkpoints
摘要:我们介绍了TalkVerse,这是一个大规模的开放语料库,用于单人,音频驱动的谈话视频生成,旨在实现公平,可重复的方法比较。虽然目前最先进的系统依赖于封闭的数据或计算繁重的模型,但TalkVerse提供了230万个高分辨率(720 p/1080 p)音频视频同步剪辑,总计6.3k小时。这些都是通过透明的管道从超过60 k小时的视频中策划的,包括场景切换检测,美学评估,严格的视听同步检查以及全面的注释,包括2D骨架和结构化的视觉/音频风格字幕。利用TalkVerse,我们提出了一个基于Wan2.2- 5 B构建的可重复的5 B DiT基线。通过利用具有高下采样率的视频VAE和具有运动帧上下文的滑动窗口机制,我们的模型实现了具有低漂移的分钟长的生成。它提供了与14 B Wan-S2 V型号相当的对口型同步和视觉质量,但推理成本低10倍。为了增强长视频中的故事讲述,我们整合了一个MLLM导演,根据音频和视觉线索重写提示。此外,我们的模型支持zero-shot视频配音通过控制潜在噪声注入。我们开源了数据集、训练配方和5 B检查点,以降低音频驱动的人类视频生成研究的障碍。项目页面:https://zhenzhiwang.github.io/talkverse/
摘要:We introduce TalkVerse, a large-scale, open corpus for single-person, audio-driven talking video generation designed to enable fair, reproducible comparison across methods. While current state-of-the-art systems rely on closed data or compute-heavy models, TalkVerse offers 2.3 million high-resolution (720p/1080p) audio-video synchronized clips totaling 6.3k hours. These are curated from over 60k hours of video via a transparent pipeline that includes scene-cut detection, aesthetic assessment, strict audio-visual synchronization checks, and comprehensive annotations including 2D skeletons and structured visual/audio-style captions. Leveraging TalkVerse, we present a reproducible 5B DiT baseline built on Wan2.2-5B. By utilizing a video VAE with a high downsampling ratio and a sliding window mechanism with motion-frame context, our model achieves minute-long generation with low drift. It delivers comparable lip-sync and visual quality to the 14B Wan-S2V model but with 10$\times$ lower inference cost. To enhance storytelling in long videos, we integrate an MLLM director to rewrite prompts based on audio and visual cues. Furthermore, our model supports zero-shot video dubbing via controlled latent noise injection. We open-source the dataset, training recipes, and 5B checkpoints to lower barriers for research in audio-driven human video generation. Project Page: https://zhenzhiwang.github.io/talkverse/
【8】Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
标题:音频多重挑战:口语对话系统对自然人类互动的多轮评估
链接:https://arxiv.org/abs/2512.14865
摘要:端到端(E2 E)口语对话系统正在越来越多地取代基于语音的人机交互的级联管道,直接处理原始音频而无需中间转录。现有的基准主要评估这些模型的合成语音和单轮任务,离开现实的多轮会话能力的探索。我们介绍音频MultiChallenge,一个开源的基准评估E2 E口语对话系统下的自然多轮交互模式。建立在基于文本的MultiChallenge框架,它评估推理记忆,指令保留和自我连贯性,我们引入了一个新的轴语音编辑,测试鲁棒性,中间话语语音修复和回溯。我们进一步增强了每个轴的音频模态,例如为推理记忆引入Audio-Cue挑战,需要回忆环境声音和语义内容之外的非语言信号。我们通过一个混合音频原生代理和人在回路的管道,策划了来自47个扬声器的452个对话,其中有1,712个实例特定的主题,该管道暴露了大规模的模型故障,同时保留了在无脚本的人类语音中发现的自然不流利。我们对专有和开源模型的评估显示,即使是前沿模型也难以达到我们的基准,Gemini 3 Pro Preview(Thinking)是我们性能最高的模型,通过率为54.65%。错误分析表明,模型在我们的新轴上最常失败,并且自相干性随着音频上下文的延长而降低。这些失败反映了在自然口语对话中跟踪编辑、音频提示和远程上下文的困难。Audio MultiChallenge提供了一个可重复的测试平台来量化它们,并推动音频原生多回合交互功能的改进。
摘要:End-to-end (E2E) spoken dialogue systems are increasingly replacing cascaded pipelines for voice-based human-AI interaction, processing raw audio directly without intermediate transcription. Existing benchmarks primarily evaluate these models on synthetic speech and single-turn tasks, leaving realistic multi-turn conversational ability underexplored. We introduce Audio MultiChallenge, an open-source benchmark to evaluate E2E spoken dialogue systems under natural multi-turn interaction patterns. Building on the text-based MultiChallenge framework, which evaluates Inference Memory, Instruction Retention, and Self Coherence, we introduce a new axis Voice Editing that tests robustness to mid-utterance speech repairs and backtracking. We further augment each axis to the audio modality, such as introducing Audio-Cue challenges for Inference Memory that require recalling ambient sounds and paralinguistic signals beyond semantic content. We curate 452 conversations from 47 speakers with 1,712 instance-specific rubrics through a hybrid audio-native agentic and human-in-the-loop pipeline that exposes model failures at scale while preserving natural disfluencies found in unscripted human speech. Our evaluation of proprietary and open-source models reveals that even frontier models struggle on our benchmark, with Gemini 3 Pro Preview (Thinking), our highest-performing model achieving a 54.65% pass rate. Error analysis shows that models fail most often on our new axes and that Self Coherence degrades with longer audio context. These failures reflect difficulty of tracking edits, audio cues, and long-range context in natural spoken dialogue. Audio MultiChallenge provides a reproducible testbed to quantify them and drive improvements in audio-native multi-turn interaction capability.
【9】Improving Underwater Acoustic Classification Through Learnable Gabor Filter Convolution and Attention Mechanisms
标题:通过可学习的Gabor过滤卷积和注意力机制改进水下声学分类
链接:https://arxiv.org/abs/2512.14714
摘要:水声目标的远程检测和分类是环境监测和防御的关键。然而,船舶辐射和环境水下噪声的复杂性对精确的信号处理提出了重大挑战。虽然机器学习的最新进展提高了分类准确性,但数据集可用性有限和缺乏标准化实验等问题阻碍了泛化和鲁棒性。本文介绍了GSE ResNeXt,这是一种深度学习架构,它将可学习的Gabor卷积层与通过挤压和激发注意力机制增强的ResNeXt主干集成在一起。Gabor滤波器用作二维自适应带通滤波器,扩展了特征通道表示。它与通道注意力的结合提高了训练稳定性和收敛性,同时增强了模型提取区分特征的能力。该模型进行评估的三个分类任务的复杂性增加。特别是,训练和测试数据之间的时间差异的影响进行了探讨,揭示了船舶和传感器之间的距离显着影响性能。结果表明,GSE ResNeXt在分类性能方面始终优于Xception,ResNet和MobileNetV2等基线模型。关于稳定性和收敛性,在模型的初始层中添加Gabor卷积表示训练时间减少了28%。这些结果强调了信号处理策略在不同环境条件下,特别是在数据有限的水下声学分类情况下,提高模型的可靠性和通用性的重要性。未来的发展应侧重于减轻环境因素对输入信号的影响。
摘要:Remotely detecting and classifying underwater acoustic targets is critical for environmental monitoring and defence. However, the complex nature of ship-radiated and environmental underwater noise poses significant challenges to accurate signal processing. While recent advancements in machine learning have improved classification accuracy, issues such as limited dataset availability and a lack of standardised experimentation hinder generalisation and robustness. This paper introduces GSE ResNeXt, a deep learning architecture integrating learnable Gabor convolutional layers with a ResNeXt backbone enhanced by squeeze-and-excitation attention mechanisms. The Gabor filters serve as two-dimensional adaptive band-pass filters, extending the feature channel representation. Its combination with channel attention improves training stability and convergence while enhancing the model's ability to extract discriminative features. The model is evaluated on three classification tasks of increasing complexity. In particular, the impact of temporal differences between the training and testing data is explored, revealing that the distance between the vessel and sensor significantly affects performance. Results show that, GSE ResNeXt consistently outperforms baseline models like Xception, ResNet, and MobileNetV2, in terms of classification performance. Regarding stability and convergence, the addition of Gabor convolutions in the initial layers of the model represents a 28% reduction in training time. These results emphasise the importance of signal processing strategies in improving the reliability and generalisation of models under different environmental conditions, especially in data-limited underwater acoustic classification scenarios. Future developments should focus on mitigating the impact of environmental factors on input signals.
【1】On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation
标题:关于使用自我监督的表示学习进行说话人二元化和分离
链接:https://arxiv.org/abs/2512.15224
备注:accepted at ASRU25
摘要:在过去的几年里,诸如wav2vec2.0和WavLM之类的自监督语音模型已经被证明可以显着提高许多下游语音任务的性能,特别是在低资源环境中。尽管如此,对诸如说话人日记和语音分离等任务的评估仍然有限。本文研究了最近的自我监督的语音表示这两个扬声器身份相关的任务的质量,突出了现有文献中的差距,源于现有基准的限制,特别是缺乏多样性的评估数据集和多样性的下游系统相关的日志和分离。
摘要:Self-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in low-resource settings, over the past few years. Despite this, evaluations on tasks such as Speaker Diarization and Speech Separation remain limited. This paper investigates the quality of recent self-supervised speech representations on these two speaker identity-related tasks, highlighting gaps in the current literature that stem from limitations in the existing benchmarks, particularly the lack of diversity in evaluation datasets and variety in downstream systems associated to both diarization and separation.
【2】A Conditioned UNet for Music Source Separation
标题:用于音乐源分离的条件化UNet
链接:https://arxiv.org/abs/2512.15532
摘要:在本文中,我们提出了一个有条件的UNet音乐源分离(MSS)。MSS通常由多输出神经网络(通常为UNets)执行,每个输出表示来自预定义乐器词汇表的特定词干。相比之下,条件MSS网络接受与感兴趣的干相关的音频查询,以及从中提取干的信号。因此,不需要严格的词汇表,这使得MSS中的任务更加现实。由于缺乏合适的数据,这些任务的条件方法的潜力在某种程度上被隐藏了,最近MoisesDb数据集解决了这个问题。最近的一种方法Banquet使用了这个数据集,在更大的词汇表上看到了有希望的结果。Banquet使用Bandsplit RNN而不是UNet,作者指出UNet不应该适合有条件的MSS。我们反对这种说法,并提出QSCNet,一种新的有条件的UNet的MSS,集成了网络调节元素的稀疏压缩网络MSS。我们发现QSCNet在几个MSS任务上的性能优于Banquet超过1dB SNR,而使用的参数数量不到一半。
摘要:In this paper we propose a conditioned UNet for Music Source Separation (MSS). MSS is generally performed by multi-output neural networks, typically UNets, with each output representing a particular stem from a predefined instrument vocabulary. In contrast, conditioned MSS networks accept an audio query related to a stem of interest alongside the signal from which that stem is to be extracted. Thus, a strict vocabulary is not required and this enables more realistic tasks in MSS. The potential of conditioned approaches for such tasks has been somewhat hidden due to a lack of suitable data, an issue recently addressed with the MoisesDb dataset. A recent method, Banquet, employs this dataset with promising results seen on larger vocabularies. Banquet uses Bandsplit RNN rather than a UNet and the authors state that UNets should not be suitable for conditioned MSS. We counter this argument and propose QSCNet, a novel conditioned UNet for MSS that integrates network conditioning elements in the Sparse Compressed Network for MSS. We find QSCNet to outperform Banquet by over 1dB SNR on a couple of MSS tasks, while using less than half the number of parameters.
【3】Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
标题:自适应多模式人员识别:处理缺失模式的稳健框架
链接:https://arxiv.org/abs/2512.14961
备注:10 pages and 8 tables
摘要:人员识别系统通常依赖于音频,视觉或行为线索,但现实世界的条件经常导致丢失或降级的模态。为了解决这一挑战,我们提出了一个三模态的人识别框架,集成语音,面部和手势模态,同时保持强大的模态损失。我们的方法利用多任务学习来独立处理每种模态,然后通过交叉注意和门控融合机制来促进跨模态的交互。此外,置信加权融合策略动态适应缺失和低质量数据,即使在单峰或双峰场景下也能确保最佳分类。我们在CANDOR上评估我们的方法,CANDOR是一个新引入的基于访谈的多模态数据集,我们首次进行基准测试。我们的研究结果表明,所提出的三模态系统在人员识别任务上达到了99.18%的Top-1准确率,优于传统的单峰和后期融合方法。此外,我们在VoxCeleb 1数据集上评估了我们的模型作为基准,并在双峰模式下达到了99.92%的准确率。此外,我们表明,即使在一个或两个模态不可用时,我们的系统也保持了很高的准确性,使其成为现实世界中人员识别应用的鲁棒解决方案。这项工作的代码和数据是公开的。
摘要:Person recognition systems often rely on audio, visual, or behavioral cues, but real-world conditions frequently result in missing or degraded modalities. To address this challenge, we propose a Trimodal person identification framework that integrates voice, face, and gesture modalities, while remaining robust to modality loss. Our approach leverages multi-task learning to process each modality independently, followed by a cross-attention and gated fusion mechanisms to facilitate interaction across modalities. Moreover, a confidence-weighted fusion strategy dynamically adapts to missing and low-quality data, ensuring optimal classification even in Unimodal or Bimodal scenarios. We evaluate our method on CANDOR, a newly introduced interview-based multimodal dataset, which we benchmark for the first time. Our results demonstrate that the proposed Trimodal system achieves 99.18% Top-1 accuracy on person identification tasks, outperforming conventional Unimodal and late-fusion approaches. In addition, we evaluate our model on the VoxCeleb1 dataset as a benchmark and reach 99.92% accuracy in Bimodal mode. Moreover, we show that our system maintains high accuracy even when one or two modalities are unavailable, making it a robust solution for real-world person recognition applications. The code and data for this work are publicly available.
机器翻译由腾讯交互翻译提供,仅供参考
