微信公众号:arXiv_Daily
cs.SD语音
【1】A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data
标题:通过语音和言语数据预测抑郁和焦虑的多模式Bayesian网络
链接:https://arxiv.org/abs/2512.07741
摘要:在精神病评估过程中,临床医生不仅要观察患者报告的内容,还要观察重要的非语言体征,如语调、语速、流畅性、反应性和肢体语言。权衡和整合这些不同的信息源是一项具有挑战性的任务,也是智能驱动工具支持的一个很好的候选者-然而,这尚未在临床上实现。在这里,我们认为,可以使用贝叶斯网络建模来解决几个重要的障碍。为了证明这一点,我们评估了一个从大规模数据集(30,135个独立说话者)中的语音和语音特征预测抑郁和焦虑症状的模型。除了对条件和症状的表现(对于抑郁症,焦虑ROC-AUC= 0.842,0.831 ECE= 0.018,0.015;核心个体症状ROC-AUC>0.74),我们评估了人口统计公平性,并调查了不同输入模态类型之间的整合和冗余。临床有用的指标和接受精神卫生服务用户进行了探讨。当提供足够丰富和大规模的多模态数据流,并指定代表常见的精神状况的症状,而不是障碍水平,这样的模型是一个原则性的方法,用于建立强大的评估支持工具:提供临床相关的输出在一个透明的和可解释的格式,是直接服从专家的临床监督。
摘要:During psychiatric assessment, clinicians observe not only what patients report, but important nonverbal signs such as tone, speech rate, fluency, responsiveness, and body language. Weighing and integrating these different information sources is a challenging task and a good candidate for support by intelligence-driven tools - however this is yet to be realized in the clinic. Here, we argue that several important barriers to adoption can be addressed using Bayesian network modelling. To demonstrate this, we evaluate a model for depression and anxiety symptom prediction from voice and speech features in large-scale datasets (30,135 unique speakers). Alongside performance for conditions and symptoms (for depression, anxiety ROC-AUC=0.842,0.831 ECE=0.018,0.015; core individual symptom ROC-AUC>0.74), we assess demographic fairness and investigate integration across and redundancy between different input modality types. Clinical usefulness metrics and acceptability to mental health service users are explored. When provided with sufficiently rich and large-scale multimodal data streams and specified to represent common mental conditions at the symptom rather than disorder level, such models are a principled approach for building robust assessment support tools: providing clinically-relevant outputs in a transparent and explainable format that is directly amenable to expert clinical supervision.
【2】Incorporating Structure and Chord Constraints in Symbolic Transformer-based Melodic Harmonization
标题:基于符号转换器的旋律和谐中的共鸣结构和和弦约束
链接:https://arxiv.org/abs/2512.07627
备注:Proceedings of the 6th Conference on AI Music Creativity (AIMC 2025), Brussels, Belgium, September 10th-12th
摘要:Transformer架构在符号音乐的生成方面提供了显着的优势;它们将用户偏好纳入其生成内容的能力正在许多方面进行研究。本文研究了在旋律协调中包含预定义的和弦约束,即,其中,在特定位置处的期望和弦与旋律一起被提供作为输入,并且自回归Transformer模型需要将该和弦并入其生成的和声中。涉及这种约束的特点进行了讨论,并提出了一种算法来处理这项任务。该算法称为B*,它结合了波束搜索和A* 以及回溯的各个方面,以迫使预训练的Transformers在正确的小节内的正确开始位置满足和弦约束。该算法是蛮力,并在最坏的情况下,具有指数的复杂性,然而,本文是第一次尝试突出的困难的问题,并提出了一种算法,提供了许多改进的可能性,因为它适应了参与的mathematics。
摘要:Transformer architectures offer significant advantages regarding the generation of symbolic music; their capabilities for incorporating user preferences toward what they generate is being studied under many aspects. This paper studies the inclusion of predefined chord constraints in melodic harmonization, i.e., where a desired chord at a specific location is provided along with the melody as inputs and the autoregressive transformer model needs to incorporate the chord in the harmonization that it generates. The peculiarities of involving such constraints is discussed and an algorithm is proposed for tackling this task. This algorithm is called B* and it combines aspects of beam search and A* along with backtracking to force pretrained transformers to satisfy the chord constraints, at the correct onset position within the correct bar. The algorithm is brute-force and has exponential complexity in the worst case; however, this paper is a first attempt to highlight the difficulties of the problem and proposes an algorithm that offers many possibilities for improvements since it accommodates the involvement of heuristics.
【3】MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection
标题:MultiAPI Spoof:用于语音反欺骗检测的多API数据集和本地注意力网络
链接:https://arxiv.org/abs/2512.07352
摘要:现有的语音反欺骗基准依赖于一组狭窄的公共模型,这与商业系统采用不同的,通常是专有的API的真实场景存在很大的差距。为了解决这个问题,我们引入了MultiAPI Spoof,这是一个多API音频反欺骗数据集,包括由30个不同的API生成的约230小时的合成语音,包括商业服务,开源模型和在线平台。基于此数据集,我们定义了API跟踪任务,从而能够将欺骗音频细粒度地归因于其生成源。我们进一步提出了Nes 2Net-LA,这是Nes 2Net的一个局部注意力增强变体,它改进了局部上下文建模和细粒度欺骗特征提取。实验表明,Nes 2Net-LA实现了最先进的性能,并提供了卓越的鲁棒性,特别是在各种不可见的欺骗条件下。代码\footnote{https://github.com/XuepingZhang/MultiAPI-Spoof}和数据集\footnote{https://xuepingzhang.github.io/MultiAPI-Spoof-Dataset/}已经发布。
摘要:Existing speech anti-spoofing benchmarks rely on a narrow set of public models, creating a substantial gap from real-world scenarios in which commercial systems employ diverse, often proprietary APIs. To address this issue, we introduce MultiAPI Spoof, a multi-API audio anti-spoofing dataset comprising about 230 hours of synthetic speech generated by 30 distinct APIs, including commercial services, open-source models, and online platforms. Based on this dataset, we define the API tracing task, enabling fine-grained attribution of spoofed audio to its generation source. We further propose Nes2Net-LA, a local-attention enhanced variant of Nes2Net that improves local context modeling and fine-grained spoofing feature extraction. Experiments show that Nes2Net-LA achieves state-of-the-art performance and offers superior robustness, particularly under diverse and unseen spoofing conditions. Code \footnote{https://github.com/XuepingZhang/MultiAPI-Spoof} and dataset \footnote{https://xuepingzhang.github.io/MultiAPI-Spoof-Dataset/} have released.
【4】DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
标题:DeepAgent:用于鲁棒的多模式Deepfake检测的双流多Agent融合
链接:https://arxiv.org/abs/2512.07351
摘要:合成媒体的使用越来越多,特别是deepfakes,是数字内容验证的一个新挑战。虽然最近的研究使用音频和视觉信息,大多数集成这些线索在一个单一的模型,这仍然容易受到模态失配,噪声和操纵。为了解决这一差距,我们提出了DeepAgent,这是一种先进的多智能体协作框架,同时结合了视觉和音频模式,以有效检测deepfakes。DeepAgent由两个互补的代理组成。Agent-1使用精简的基于AlexNet的CNN检查每个视频,以识别deepfake操纵的符号,而Agent-2通过结合声学特征,Whisper的音频传输和EasyOCR的图像帧读取序列来检测视听不一致。他们的决策通过随机森林元分类器进行融合,该元分类器通过利用每个代理学习的不同决策边界来提高最终性能。本研究使用三个基准数据集来评估所提出的框架,以证明组件级和融合的性能。Agent-1在Celeb-DF和FakeAVCeleb组合数据集上的测试准确率为94.35%。在FakeAVCeleb数据集上,Agent-2和最终元分类器的准确率分别为93.69%和81.56%。此外,DeepFakeTIMIT上的跨数据集验证证实了元分类器的鲁棒性,最终准确率达到97.49%,并表明跨不同数据集的强大能力。这些发现证实了基于层次的融合通过减轻单个模态的弱点来增强鲁棒性,并证明了多智能体方法在解决deepfakes中不同类型的操纵方面的有效性。
摘要:The increasing use of synthetic media, particularly deepfakes, is an emerging challenge for digital content verification. Although recent studies use both audio and visual information, most integrate these cues within a single model, which remains vulnerable to modality mismatches, noise, and manipulation. To address this gap, we propose DeepAgent, an advanced multi-agent collaboration framework that simultaneously incorporates both visual and audio modalities for the effective detection of deepfakes. DeepAgent consists of two complementary agents. Agent-1 examines each video with a streamlined AlexNet-based CNN to identify the symbols of deepfake manipulation, while Agent-2 detects audio-visual inconsistencies by combining acoustic features, audio transcriptions from Whisper, and frame-reading sequences of images through EasyOCR. Their decisions are fused through a Random Forest meta-classifier that improves final performance by taking advantage of the different decision boundaries learned by each agent. This study evaluates the proposed framework using three benchmark datasets to demonstrate both component-level and fused performance. Agent-1 achieves a test accuracy of 94.35% on the combined Celeb-DF and FakeAVCeleb datasets. On the FakeAVCeleb dataset, Agent-2 and the final meta-classifier attain accuracies of 93.69% and 81.56%, respectively. In addition, cross-dataset validation on DeepFakeTIMIT confirms the robustness of the meta-classifier, which achieves a final accuracy of 97.49%, and indicates a strong capability across diverse datasets. These findings confirm that hierarchy-based fusion enhances robustness by mitigating the weaknesses of individual modalities and demonstrate the effectiveness of a multi-agent approach in addressing diverse types of manipulations in deepfakes.
【5】Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits
标题:通过视频编辑后的条件音频生成进行连贯的视听编辑
链接:https://arxiv.org/abs/2512.07209
摘要:我们介绍了一种新颖的管道,用于联合视听编辑,增强了编辑的视频及其伴随的音频之间的一致性。我们的方法首先应用最先进的视频编辑技术来生成目标视频,然后执行音频编辑以与视觉变化保持一致。为了实现这一点,我们提出了一个新的视频到音频生成模型的条件下的源音频,目标视频,和文本提示。我们扩展了模型架构,将条件音频输入,并提出了一个数据增强策略,提高了训练效率。此外,我们的模型根据编辑的复杂性动态调整源音频的影响,并尽可能保留原始音频结构。实验结果表明,我们的方法优于现有的方法在保持视听对齐和内容完整性。
摘要:We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then performs audio editing to align with the visual changes. To achieve this, we present a new video-to-audio generation model that conditions on the source audio, target video, and a text prompt. We extend the model architecture to incorporate conditional audio input and propose a data augmentation strategy that improves training efficiency. Furthermore, our model dynamically adjusts the influence of the source audio based on the complexity of the edits, preserving the original audio structure where possible. Experimental results demonstrate that our method outperforms existing approaches in maintaining audio-visual alignment and content integrity.
【6】JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
标题:JEPA作为神经标记器:使用密度自适应注意力学习鲁棒的语音表示
链接:https://arxiv.org/abs/2512.07168
备注:UniReps: Unifying Representations in Neural Models (NeurIPS 2025 Workshop)
摘要:我们引入了一个两阶段的自监督框架,该框架将联合嵌入预测架构(JEPA)与密度自适应注意力机制(DAAM)相结合,用于学习鲁棒的语音表示。阶段1使用JEPA和DAAM通过潜在空间中的掩蔽预测来学习语义音频特征,与波形重构完全解耦。阶段2利用这些表示来使用有限标量量化(FSQ)和混合基数打包方案进行有效的令牌化,然后使用HiFi-GAN解码器进行高保真波形重构。该模型通过在JEPA编码器中引入基于高斯混合的密度自适应选通,实现了自适应时域特征选择,并在2.5Hz的低帧率下发现了语音的层次结构。由此产生的令牌(47.5令牌/秒)提供了一种可逆的,高度压缩的,语言模型友好的表示,与现有的神经音频编解码器相比具有竞争力,并且通常更有效。
摘要:We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM to learn semantic audio features via masked prediction in latent space, fully decoupled from waveform reconstruction. Stage~2 leverages these representations for efficient tokenization using Finite Scalar Quantization (FSQ) and a mixed-radix packing scheme, followed by high-fidelity waveform reconstruction with a HiFi-GAN decoder. By integrating Gaussian mixture-based density-adaptive gating into the JEPA encoder, the model performs adaptive temporal feature selection and discovers hierarchical speech structure at a low frame rate of 2.5~Hz. The resulting tokens (47.5 tokens/sec) provide a reversible, highly compressed, and language-model-friendly representation that is competitive with, and often more efficient than, existing neural audio codecs.
【7】Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition
标题:多口音普通话干声乐歌唱数据集:歌唱口音识别的基准
链接:https://arxiv.org/abs/2512.07005
备注:Accepted by ACMMM 2025
摘要:与语音口音研究相比,歌唱口音研究的探索不足,主要是由于缺乏合适的数据集。现有的歌唱数据集往往遭受细节损失,经常导致声乐乐器分离过程。此外,他们往往缺乏地区口音注释。为了解决这个问题,我们引入了多口音普通话干声歌唱数据集(MADVSD)。MADVSD包括超过670小时的干声乐录音,来自中国九个不同地区的4,206名母语为普通话的人。除了每个参与者用他们的母语录制三首流行歌曲的音频外,他们还录制了涵盖所有普通话元音和完整八度范围的语音练习。我们通过唱歌口音识别的基准实验验证了MADVSD,证明了它在唱歌环境中评估最先进的语音模型的实用性。此外,我们还利用MADVSD独特的语音练习,探讨了方言对歌唱口音的影响,并分析了元音在重音变化中的作用。
摘要:Singing accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal-instrumental separation process. Additionally, they often lack regional accent annotations. To address this, we introduce the Multi-Accent Mandarin Dry-Vocal Singing Dataset (MADVSD). MADVSD comprises over 670 hours of dry vocal recordings from 4,206 native Mandarin speakers across nine distinct Chinese regions. In addition to each participant recording audio of three popular songs in their native accent, they also recorded phonetic exercises covering all Mandarin vowels and a full octave range. We validated MADVSD through benchmark experiments in singing accent recognition, demonstrating its utility for evaluating state-of-the-art speech models in singing contexts. Furthermore, we explored dialectal influences on singing accent and analyzed the role of vowels in accentual variations, leveraging MADVSD's unique phonetic exercises.
【8】Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
标题:基于多模式大基础模型的歌唱音色受欢迎程度评估
链接:https://arxiv.org/abs/2512.06999
备注:Accepted to ACMMM 2025 oral
摘要:自动歌唱评估对于教育和娱乐至关重要。然而,现有的系统面临着两个基本的限制:依赖于参考轨道,这扼杀了创造性的表达,并简化了复杂的性能到非诊断分数的基础上完全音高和节奏。我们提倡从歧视性评估转向描述性评估,为无参考、多维度的评估创建一个完整的生态系统。首先,我们介绍Sing-MD,这是一个由专家在四个维度上注释的大规模数据集:呼吸控制,音色质量,情感表达和声乐技巧。我们的分析揭示了专家之间的显着注释不一致,挑战传统的基于准确性的指标的有效性。其次,针对多模态大型语言模型(MLLM)在分析全长歌曲时的内存限制,我们提出了VocalVerse。这种高效的混合架构利用轻量级声学编码器来模拟全局性能特征和长期依赖关系。第三,为了解决自动化度量的缺点,我们建立了H-TPR(Human-in-the-loop Tiered Perceptual Ranking)基准,该基准评估模型生成感知有效排名的能力,而不是预测嘈杂的地面实况得分。
摘要:Automated singing assessment is crucial for education and entertainment. However, existing systems face two fundamental limitations: reliance on reference tracks, which stifles creative expression, and the simplification of complex performances into non-diagnostic scores based solely on pitch and rhythm. We advocate for a shift from discriminative to descriptive evaluation, creating a complete ecosystem for reference-free, multi-dimensional assessment. First, we introduce Sing-MD, a large-scale dataset annotated by experts across four dimensions: breath control, timbre quality, emotional expression, and vocal technique. Our analysis reveals significant annotation inconsistencies among experts, challenging the validity of traditional accuracy-based metrics. Second, addressing the memory limitations of Multimodal Large Language Models (MLLMs) in analyzing full-length songs, we propose VocalVerse. This efficient hybrid architecture leverages a lightweight acoustic encoder to model global performance features and long-term dependencies. Third, to address automated metric shortcomings, we establish the H-TPR (Human-in-the-loop Tiered Perceptual Ranking) benchmark, which evaluates a model's ability to generate perceptually valid rankings rather than predicting noisy ground-truth scores.
【9】What Needs to be Known in Order to Perform a Meaningful Scientific Comparison Between Animal Communications and Human Spoken Language
标题:为了对动物交流和人类口语进行有意义的科学比较,需要了解什么
链接:https://arxiv.org/abs/2512.06890
备注:5 pages, 1 figure, Proc. Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR-24), Kos, Greece, 6 Sept. 2024
摘要:人类口语一直是科学研究的主题,特别是关于言语产生的机制。同样,动物交流的研究也有大量的文献,其中许多研究都集中在发声上。最近,人们对比较动物交流和人类语言越来越感兴趣。然而,这里提出,这种比较需要评估一组最低限度的临界现象:i)发声器官的自由度的数量,ii)独立地控制这些自由度的能力,iii)进行通信的声学环境的特性,iv)所生成的声音的感知显著性,v)声音对比的程度,vi)组合性的存在/不存在,以及vii)所得到的通信的信息速率。
摘要:Human spoken language has long been the subject of scientific investigation, particularly with regard to the mechanisms underpinning speech production. Likewise, the study of animal communications has a substantial literature, with many studies focusing on vocalisation. More recently, there has been growing interest in comparing animal communications and human speech. However, it is proposed here that such a comparison necessitates the appraisal of a minimum set of critical phenomena: i) the number of degrees-of-freedom of the vocal apparatus, ii) the ability to control those degrees-of-freedom independently, iii) the properties of the acoustic environment in which communication takes place, iv) the perceptual salience of the generated sounds, v) the degree to which sounds are contrastive, vi) the presence/absence of compositionality, and vii) the information rate(s) of the resulting communications.
【10】XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association
标题:XM-ALIGN:面向脸语音关联的统一跨模式嵌入对齐
链接:https://arxiv.org/abs/2512.06757
备注:FAME 2026 Technical Report
摘要:本文介绍了我们的解决方案,XM-ALIGN(统一跨模态嵌入对齐框架),在ICASSP 2026年提出的FAME挑战。我们的框架结合了显式和隐式对齐机制,显着提高跨模态验证性能在“听到”和“听不到”的语言。通过从人脸和语音编码器中提取特征嵌入,并使用共享分类器对其进行联合优化,我们采用均方误差(MSE)作为嵌入对齐损失,以确保模态之间的紧密对齐。此外,在模型训练期间应用数据增强策略以增强泛化。实验结果表明,我们的方法在MAV-Celeb数据集上表现出优越的性能。该代码将在https://github.com/PunkMale/XM-ALIGN上发布。
摘要:This paper introduces our solution, XM-ALIGN (Unified Cross-Modal Embedding Alignment Framework), proposed for the FAME challenge at ICASSP 2026. Our framework combines explicit and implicit alignment mechanisms, significantly improving cross-modal verification performance in both "heard" and "unheard" languages. By extracting feature embeddings from both face and voice encoders and jointly optimizing them using a shared classifier, we employ mean squared error (MSE) as the embedding alignment loss to ensure tight alignment between modalities. Additionally, data augmentation strategies are applied during model training to enhance generalization. Experimental results show that our approach demonstrates superior performance on the MAV-Celeb dataset. The code will be released at https://github.com/PunkMale/XM-ALIGN.
【11】Hankel-FNO: Fast Underwater Acoustic Charting Via Physics-Encoded Fourier Neural Operator
标题:Hankel-FNO:通过物理编码傅里叶神经运算符快速水下声学图表
链接:https://arxiv.org/abs/2512.06417
摘要:快速准确的水声制图对于环境感知传感器布局优化和自主车辆路径规划等下游任务至关重要。传统的方法依赖于计算昂贵而精确的数值求解器,这对于大规模或实时应用来说是不可扩展的。虽然基于深度学习的代理模型可以加速这些计算,但它们通常受到限制,例如固定分辨率约束或依赖显式偏微分方程公式。这些问题阻碍了它们在不同环境中的适用性和通用性。我们提出了Hankel-FNO,一种基于傅立叶神经运算符(FNO)的模型,用于高效准确的声学制图。通过结合声传播知识和水深测量,我们的方法具有较高的精度,同时保持较高的计算速度。结果表明,Hankel-FNO在速度上优于传统求解器,在精度上优于数据驱动的替代方案,特别是在长期预测方面。实验表明,该模型的适应性,以最小的微调不同的环境和声源设置。
摘要:Fast and accurate underwater acoustic charting is crucial for downstream tasks such as environment-aware sensor placement optimization and autonomous vehicle path planning. Conventional methods rely on computationally expensive while accurate numerical solvers, which are not scalable for large-scale or real-time applications. Although deep learning-based surrogate models can accelerate these computations, they often suffer from limitations such as fixed-resolution constraints or dependence on explicit partial differential equation formulations. These issues hinder their applicability and generalization across diverse environments. We propose Hankel-FNO, a Fourier Neural Operator (FNO)-based model for efficient and accurate acoustic charting. By incorporating sound propagation knowledge and bathymetry, our method has high accuracy while maintaining high computational speed. Results demonstrate that Hankel-FNO outperforms traditional solvers in speed and surpasses data-driven alternatives in accuracy, especially in long-range predictions. Experiments show the model's adaptability to diverse environments and sound source settings with minimal fine-tuning.
【12】Protecting Bystander Privacy via Selective Hearing in LALMs
标题:通过LALM中的选择性听证保护旁观者隐私
链接:https://arxiv.org/abs/2512.06380
备注:Dataset: https://huggingface.co/datasets/BrianatCambridge/SelectiveHearingBench
摘要:大型音频语言模型(LALM)越来越多地部署在现实环境中,它们不可避免地从附近无意的旁观者那里捕获语音,从而提高了现有基准和防御措施在很大程度上忽视的隐私风险。我们介绍了SH-Bench,这是第一个旨在评估选择性听力的基准:模型在拒绝处理或透露有关偶然旁观者言语的信息的同时关注预期主要发言者的能力。SH-Bench包含3,968个跨越真实世界和合成场景的多扬声器音频混合,以及77 k个在一般和选择性操作模式下探测模型的多项选择题。我们提出了选择效能(SE),一个统一的度量捕获多说话人的理解和旁观者隐私保护。我们对最先进的开源和专有LALM的评估揭示了大量的隐私泄露,强大的音频理解未能转化为对旁观者隐私的选择性保护。为了弥补这一差距,我们引入了旁观者隐私微调(BPFT),这是一个训练管道,可以教模型拒绝旁观者相关的查询,而不会降低主说话者的理解能力。BPFT产生实质性的收益,比Gemini 2.5 Pro提高SE高达15.9%,表明选择性听力是可学习的,但在当前的LALM中还远未实现。SH-Bench和BPFT为测量和改善音频基础模型中的旁观者隐私提供了第一个系统框架。
摘要:Large audio language models (LALMs) are increasingly deployed in real-world settings where they inevitably capture speech from unintended nearby bystanders, raising privacy risks that existing benchmarks and defences largely overlook. We introduce SH-Bench, the first benchmark designed to evaluate selective hearing: a model's ability to attend to an intended main speaker while refusing to process or reveal information about incidental bystander speech. SH-Bench contains 3,968 multi-speaker audio mixtures spanning both real-world and synthetic scenarios, paired with 77k multiple-choice questions that probe models under general and selective operating modes. We propose Selective Efficacy (SE), a unified metric capturing both multi-speaker comprehension and bystander-privacy protection. Our evaluation of state-of-the-art open-source and proprietary LALMs reveals substantial privacy leakage, with strong audio understanding failing to translate into selective protection of bystander privacy. To mitigate this gap, we introduce Bystander Privacy Fine-Tuning (BPFT), a training pipeline that teaches models to refuse bystander-related queries without degrading main-speaker comprehension. BPFT yields substantial gains which improve SE by up to 15.9% over Gemini 2.5 Pro, demonstrating that selective hearing is learnable but far from achieved in current LALMs. SH-Bench and BPFT provide the first systematic framework for measuring and improving bystander privacy in audio foundation models.
【13】Who Will Top the Charts? Multimodal Music Popularity Prediction via Adaptive Fusion of Modality Experts and Temporal Engagement Modeling
标题:谁将位居排行榜榜首?通过情态专家和时间参与建模的自适应融合的多情态音乐流行度预测
链接:https://arxiv.org/abs/2512.06259
备注:8 pages
摘要:预测一首歌曲在发行前的商业成功仍然是音乐行业的一个开放和关键的研究挑战。音乐流行的早期预测为战略决策、创意规划和营销提供了信息。现有的方法有四个局限性:(i)音频和歌词中的时间动态被平均掉;(ii)歌词被表示为一袋单词,忽略了作曲结构和情感语义;(iii)艺术家和歌曲级别的历史表现被忽略;以及(iv)多模式融合方法依赖于简单的特征级联,导致对齐不良的共享表示。为了解决这些限制,我们引入了GAMENet,这是一种用于音乐流行度预测的端到端多模式深度学习架构。GAMENet通过自适应门控机制集成了音频、歌词和社交元数据的特定模态专家。我们使用了通过OnionEnsembleAENet处理的Music 4AllOnion的音频特征,OnionEnsembleAENet是一个为鲁棒特征提取而设计的自动编码器网络;通过大型语言模型管道获得的歌词嵌入;以及新引入的职业轨迹动态(CTD)功能,可以捕获多年的艺术家职业发展势头和歌曲级别的轨迹统计数据。使用Music 4All数据集(113 k首曲目),之前在MIR任务中进行了探索,但没有进行流行度预测,GAMENet在R^2方面比直接多模态特征拼接提高了12%。Spotify音频描述符单独产生0.13的R^2。整合聚合CTD特征将其增加到0.69,并且从时间CTD特征获得额外的7%增益。我们使用SpotGenTrack Popularity Dataset(100 k曲目)进一步验证了鲁棒性,比之前的基线提高了16%。广泛的消融证实了模型的有效性和每种方式的独特贡献。
摘要:Predicting a song's commercial success prior to its release remains an open and critical research challenge for the music industry. Early prediction of music popularity informs strategic decisions, creative planning, and marketing. Existing methods suffer from four limitations:(i) temporal dynamics in audio and lyrics are averaged away; (ii) lyrics are represented as a bag of words, disregarding compositional structure and affective semantics; (iii) artist- and song-level historical performance is ignored; and (iv) multimodal fusion approaches rely on simple feature concatenation, resulting in poorly aligned shared representations. To address these limitations, we introduce GAMENet, an end-to-end multimodal deep learning architecture for music popularity prediction. GAMENet integrates modality-specific experts for audio, lyrics, and social metadata through an adaptive gating mechanism. We use audio features from Music4AllOnion processed via OnionEnsembleAENet, a network of autoencoders designed for robust feature extraction; lyric embeddings derived through a large language model pipeline; and newly introduced Career Trajectory Dynamics (CTD) features that capture multi-year artist career momentum and song-level trajectory statistics. Using the Music4All dataset (113k tracks), previously explored in MIR tasks but not popularity prediction, GAMENet achieves a 12% improvement in R^2 over direct multimodal feature concatenation. Spotify audio descriptors alone yield an R^2 of 0.13. Integrating aggregate CTD features increases this to 0.69, with an additional 7% gain from temporal CTD features. We further validate robustness using the SpotGenTrack Popularity Dataset (100k tracks), achieving a 16% improvement over the previous baseline. Extensive ablations confirm the model's effectiveness and the distinct contribution of each modality.
【14】Technical Report of Nomi Team in the Environmental Sound Deepfake Detection Challenge 2026
标题:Nomi团队在2026年环境之声Deepfake检测挑战赛中的技术报告
链接:https://arxiv.org/abs/2512.06041
摘要:本文介绍了我们在ICASSP 2026环境声音深度伪造检测(ESDD)挑战赛中的工作。该挑战基于由各种合成环境声音组成的大规模EnvSDD数据集。我们专注于解决看不见的发电机和低资源的黑盒场景的复杂性,提出了一个音频文本交叉注意力模型。单独和组合的文本-音频模型的实验证明了在挑战基线(BEAT +AASIST模型)上具有竞争力的EER改进。
摘要:This paper presents our work for the ICASSP 2026 Environmental Sound Deepfake Detection (ESDD) Challenge. The challenge is based on the large-scale EnvSDD dataset that consists of various synthetic environmental sounds. We focus on addressing the complexities of unseen generators and low-resource black-box scenarios by proposing an audio-text cross-attention model. Experiments with individual and combined text-audio models demonstrate competitive EER improvements over the challenge baseline (BEATs+AASIST model).
【15】Physics-Guided Deepfake Detection for Voice Authentication Systems
标题:语音认证系统的物理引导Deepfake检测
链接:https://arxiv.org/abs/2512.06040
摘要:部署在网络边缘的语音认证系统面临双重威胁:a)复杂的deepfake合成攻击和b)分布式联邦学习协议中的控制平面中毒。我们提出了一个框架,将物理引导的deepfake检测与边缘学习中的不确定性感知相结合。该框架融合了可解释的物理特征建模声道动态与来自自监督学习模块的表示。然后,通过多模态Ensemble架构处理表示,然后通过贝叶斯系综提供不确定性估计。对音频样本进行基于物理的特征评估和不确定性估计,使我们提出的框架能够对先进的deepfake攻击和复杂的控制平面中毒保持鲁棒性,从而解决网络语音认证的完整威胁模型。
摘要:Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. We present a framework coupling physics-guided deepfake detection with uncertainty-aware in edge learning. The framework fuses interpretable physics features modeling vocal tract dynamics with representations coming from a self-supervised learning module. The representations are then processed via a Multi-Modal Ensemble Architecture, followed by a Bayesian ensemble providing uncertainty estimates. Incorporating physics-based characteristics evaluations and uncertainty estimates of audio samples allows our proposed framework to remain robust to both advanced deepfake attacks and sophisticated control-plane poisoning, addressing the complete threat model for networked voice authentication.
【16】DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
标题:DreamFoley:用于高保真视频到音频生成的可扩展VLM
链接:https://arxiv.org/abs/2512.06022
备注:10 pages; Bytedance
摘要:视频生成的最新进展已经在视觉内容保真度方面取得了显著的改进。然而,同步音频的缺乏严重破坏了沉浸式体验,并限制了这些技术的实际应用。为了应对这一挑战,一些开创性的工作已经探索了扩散Transformer架构,用于生成合理的视频同步音频,包括Kling-foley,HunyuanVideo-foley和Thinksound。与现有的作品不同,我们引入了一个自回归音频生成架构(DreamFoley),利用大型视觉语言模型(VLM)的能力,共同模拟视频,音频和文本模态之间的顺序交互。我们的方法具有双视觉编码器模块,有效地捕捉音频对齐和文本对齐的视觉功能。此外,我们采用了一个残差矢量量化音频标记器与延迟模式生成方案,以平衡训练效率和音频质量之间的权衡。此外,我们将无分类器指导策略引入到VLM中,以引导生成的音频质量。此外,我们还建立了一个有效的数据生产管道,以扩展音频-视频-文本三重收集。最后,进行了大量的实验来验证我们的模型的有效性,在流行的基准测试中取得了良好的性能。我们希望这项研究的结果为未来的视频到音频生成研究提供了坚实的基础。我们还将之前缺失的视听文本描述从公共基准中发布,旨在方便后续研究人员进行更方便有效的评估和比较。
摘要:Recent advances in video generation have achieved remarkable improvements in visual content fidelity. However, the absence of synchronized audio severely undermines immersive experience and restricts practical applications of these technologies. To address this challenge, several pioneering works have explored diffusion transformer architectures for generating plausible video-synchronized audio, including Kling-foley, HunyuanVideo-foley and Thinksound. Distinct from existing works, we introduce an autoregressive audio generation architecture (DreamFoley) that harnesses the capabilities of large vision-language models (VLMs) to jointly model sequential interactions among video, audio, and text modalities. Our approach features a dual-visual encoder module that effectively captures both audio-aligned and text-aligned visual features. Additionally, we employ a Residual Vector Quantization audio tokenizer with a delay-pattern generation scheme to balance the trade-off between training efficiency and audio quality. Moreover, we introduce the classifier-free guidance strategy into VLMs to bootstrap generated audio quality. Furthermore, we establish an efficient data production pipeline to scale audio-video-text triple collection. Finally, extensive experiments are conducted to validate the effectiveness of our model, achieving promising performance across popular benchmarks. We hope that the findings in this study provide a strong foundation for future video-to-audio generation research. We also release the previously missing audio-visual textual descriptions from the public benchmark, aiming to facilitate subsequent researchers in conducting more convenient and effective evaluations and comparisons.
【17】Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
标题:具有扩散源先验的无监督单通道音频分离
链接:https://arxiv.org/abs/2512.07226
备注:15 pages, 31 figures, accepted by The 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026)
摘要:单声道音频分离的目的是从单声道混合中分离出单独的源。大多数现有的方法依赖于监督学习与合成生成的配对数据。然而,在现实世界的场景中获得高质量的配对数据往往是困难的。这种数据稀缺性会在不可见的条件下降低模型性能,并限制泛化能力。为此,在这项工作中,我们从无监督的角度来处理这个问题,将其视为概率逆问题。我们的方法只需要在单个源上训练的扩散先验。然后通过重建引导迭代地将初始状态引导到解来实现分离。重要的是,我们引入了一个专门为分离设计的高级逆问题求解器,它可以缓解逆去噪过程中扩散先验和重建指导之间的干扰所引起的梯度冲突。这种设计确保了各个源的高质量和平衡的分离性能。此外,我们发现,初始化的去噪过程与增强的混合,而不是纯高斯噪声提供了一个信息的起点,显着提高了最终的性能。为了进一步增强音频先验建模,我们设计了一种新的基于时频注意力的网络架构,该架构具有强大的音频建模能力。总的来说,这些改进带来了显着的性能增益,在语音声音事件,声音事件和语音分离任务中得到了验证。
摘要:Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated paired data. However, obtaining high-quality paired data in real-world scenarios is often difficult. This data scarcity can degrade model performance under unseen conditions and limit generalization ability. To this end, in this work, we approach this problem from an unsupervised perspective, framing it as a probabilistic inverse problem. Our method requires only diffusion priors trained on individual sources. Separation is then achieved by iteratively guiding an initial state toward the solution through reconstruction guidance. Importantly, we introduce an advanced inverse problem solver specifically designed for separation, which mitigates gradient conflicts caused by interference between the diffusion prior and reconstruction guidance during inverse denoising. This design ensures high-quality and balanced separation performance across individual sources. Additionally, we find that initializing the denoising process with an augmented mixture instead of pure Gaussian noise provides an informative starting point that significantly improves the final performance. To further enhance audio prior modeling, we design a novel time-frequency attention-based network architecture that demonstrates strong audio modeling capability. Collectively, these improvements lead to significant performance gains, as validated across speech-sound event, sound event, and speech separation tasks.
【18】Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation
标题:有辱人格的声音:通过输入操纵实现稳健语音转换的全面概述
链接:https://arxiv.org/abs/2512.06304
摘要:身份、口音、风格和情感是人类语言的基本组成部分。语音转换(VC)技术处理两个输入说话者的语音信号和诸如提示和情感标签的其他模态的辅助信息。它在保持语言内容的同时,将准语言特征从一个变换为另一个。最近,VC模型在生成质量和个性化功能方面都取得了快速进步。这些发展已经吸引了相当多的关注,不同的应用,包括隐私保护,声纹再现死者,和构音障碍的语音恢复。然而,由于训练数据干净,这些模型只学习非鲁棒性特征。随后,当处理真实世界场景中的降级输入语音时,包括额外的噪声,混响,对抗性攻击,甚至微小的扰动,它会导致不令人满意的性能。因此,它需要强大的部署,特别是在现实环境中。虽然最新的研究试图找到潜在的攻击和对策的VC系统,仍然有一个显着的差距,在输入操作下的VC模型是如何强大的全面理解。这里也提出了许多问题:例如,不同形式的输入退化攻击在多大程度上改变了VC模型的预期输出?是否有优化这些攻击和防御策略的潜力?为了回答这些问题,我们从输入操作的角度对现有的攻击和防御方法进行分类,并在四个维度上评估降级输入语音的影响,包括可懂度,自然度,音色相似性和主观感知。最后,我们概述了开放的问题和未来的方向。
摘要:Identity, accent, style, and emotions are essential components of human speech. Voice conversion (VC) techniques process the speech signals of two input speakers and other modalities of auxiliary information such as prompts and emotion tags. It changes para-linguistic features from one to another, while maintaining linguistic contents. Recently, VC models have made rapid advancements in both generation quality and personalization capabilities. These developments have attracted considerable attention for diverse applications, including privacy preservation, voice-print reproduction for the deceased, and dysarthric speech recovery. However, these models only learn non-robust features due to the clean training data. Subsequently, it results in unsatisfactory performances when dealing with degraded input speech in real-world scenarios, including additional noise, reverberation, adversarial attacks, or even minor perturbation. Hence, it demands robust deployments, especially in real-world settings. Although latest researches attempt to find potential attacks and countermeasures for VC systems, there remains a significant gap in the comprehensive understanding of how robust the VC model is under input manipulation. here also raises many questions: For instance, to what extent do different forms of input degradation attacks alter the expected output of VC models? Is there potential for optimizing these attack and defense strategies? To answer these questions, we classify existing attack and defense methods from the perspective of input manipulation and evaluate the impact of degraded input speech across four dimensions, including intelligibility, naturalness, timbre similarity, and subjective perception. Finally, we outline open issues and future directions.
【19】KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
标题:KidSpeak:用于儿童语音识别和筛查的通用多用途LLM
链接:https://arxiv.org/abs/2512.05994
摘要:随着基于对话和扩散的人工智能的快速发展,人工智能在教育服务中的应用越来越多,从评分和评估工具到为学生提供有针对性支持的个性化学习系统。然而,这种适应性尚未完全扩展到儿童语音领域,现有的模型往往由于依赖于为清晰,清晰的成人语音设计的数据集而失败。儿童,特别是那些处于早期发育阶段或患有言语和语言疾病的儿童,提出了当前人工智能模型和数据集无法处理的独特挑战。为了解决这个问题,我们介绍了KidSpeak,一个多任务语音增强的基础模型,能够生成和区分专门针对儿童的语音模式的任务。我们的框架采用了一个两阶段的训练过程,将语音知识融入语音编码器,在四个独立的任务中实现了87%的平均准确率。此外,认识到可扩展的人类注释和现有的语音对齐工具的局限性,我们提出了灵活和自动语音对齐器(FASA),并利用该方法来构建高质量的数据集进行训练和评估。这种新颖的对齐工具显着提高了从噪声数据中对齐儿童语音的质量,与人类注释相比,数据质量提高了13.6倍,如CHILDES数据集所示。据我们所知,KidSpeak和FASA代表了第一个为儿童语音和语言治疗设计的综合解决方案,提供了多功能语音LLM和强大的对齐工具。
摘要:With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for students. However, this adaptability has yet to fully extend to the domain of children's speech, where existing models often fail due to their reliance on datasets designed for clear, articulate adult speech. Children, particularly those in early developmental stages or with speech and language pathologies, present unique challenges that current AI models and datasets are ill-equipped to handle. To address this, we introduce KidSpeak, a multi-task speech-enhanced Foundation Model capable of both generative and discriminative tasks specifically tailored to children's speech patterns. Our framework employs a two-stage training process that incorporates phonetic knowledge into the speech encoder, achieving an average accuracy of 87% across four separate tasks. Furthermore, recognizing the limitations of scalable human annotation and existing speech alignment tools, we propose the Flexible and Automatic Speech Aligner (FASA) and leverage the method to construct high quality datasets for training and evaluation. This novel alignment tool significantly improves the quality of aligned children's speech from noisy data, enhancing data quality by 13.6x compared to human annotations, as demonstrated on the CHILDES dataset. To the best of our knowledge, KidSpeak and FASA represent the first comprehensive solution designed for speech and language therapy in children, offering both a multi-purpose speech LLM and a robust alignment tool.
【1】Introduction to Ambisonics, Part 1: The Part With No Math
标题:Ambisonics入门,第1部分:没有数学的部分
链接:https://arxiv.org/abs/2512.07570
摘要:本文档是高保真度立体声的两部分介绍的第1部分,针对希望实际使用高保真度立体声的读者。我们在这一部分省略了深入的技术细节,重点是帮助读者对底层概念有直观的理解。我们解释了什么是高保真度立体声信号,如何获得它们,可以对它们进行什么操作,以及如何将它们重现给听众。我们提供了各种音频示例来说明这一问题。第2部分介绍了立体混响是提供在一个单独的文件,并针对读者谁愿意了解的数学细节。
摘要:The present document is Part 1 of a 2-part introduction to ambisonics and aims at readers who would like to work practically with ambisonics. We leave out deep technical details in this part and focus on helping the reader to develop an intuitive understanding of the underlying concept. We explain what ambisonic signals are, how they can be obtained, what manipulations can be applied to them, and how they can be reproduced to a listener. We provide a variety of audio examples that illustrate the matter. Part 2 of this introduction into ambisonics is provided in a separate document and aims at readers who would like to understand the mathematical details.
【2】Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
标题:具有扩散源先验的无监督单通道音频分离
链接:https://arxiv.org/abs/2512.07226
备注:15 pages, 31 figures, accepted by The 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026)
摘要:单声道音频分离的目的是从单声道混合中分离出单独的源。大多数现有的方法依赖于监督学习与合成生成的配对数据。然而,在现实世界的场景中获得高质量的配对数据往往是困难的。这种数据稀缺性会在不可见的条件下降低模型性能,并限制泛化能力。为此,在这项工作中,我们从无监督的角度来处理这个问题,将其视为概率逆问题。我们的方法只需要在单个源上训练的扩散先验。然后通过重建引导迭代地将初始状态引导到解来实现分离。重要的是,我们引入了一个专门为分离设计的高级逆问题求解器,它可以缓解逆去噪过程中扩散先验和重建指导之间的干扰所引起的梯度冲突。这种设计确保了各个源的高质量和平衡的分离性能。此外,我们发现,初始化的去噪过程与增强的混合,而不是纯高斯噪声提供了一个信息的起点,显着提高了最终的性能。为了进一步增强音频先验建模,我们设计了一种新的基于时频注意力的网络架构,该架构具有强大的音频建模能力。总的来说,这些改进带来了显着的性能增益,在语音声音事件,声音事件和语音分离任务中得到了验证。
摘要:Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated paired data. However, obtaining high-quality paired data in real-world scenarios is often difficult. This data scarcity can degrade model performance under unseen conditions and limit generalization ability. To this end, in this work, we approach this problem from an unsupervised perspective, framing it as a probabilistic inverse problem. Our method requires only diffusion priors trained on individual sources. Separation is then achieved by iteratively guiding an initial state toward the solution through reconstruction guidance. Importantly, we introduce an advanced inverse problem solver specifically designed for separation, which mitigates gradient conflicts caused by interference between the diffusion prior and reconstruction guidance during inverse denoising. This design ensures high-quality and balanced separation performance across individual sources. Additionally, we find that initializing the denoising process with an augmented mixture instead of pure Gaussian noise provides an informative starting point that significantly improves the final performance. To further enhance audio prior modeling, we design a novel time-frequency attention-based network architecture that demonstrates strong audio modeling capability. Collectively, these improvements lead to significant performance gains, as validated across speech-sound event, sound event, and speech separation tasks.
【3】Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation
标题:有辱人格的声音:通过输入操纵实现稳健语音转换的全面概述
链接:https://arxiv.org/abs/2512.06304
摘要:身份、口音、风格和情感是人类语言的基本组成部分。语音转换(VC)技术处理两个输入说话者的语音信号和诸如提示和情感标签的其他模态的辅助信息。它在保持语言内容的同时,将准语言特征从一个变换为另一个。最近,VC模型在生成质量和个性化功能方面都取得了快速进步。这些发展已经吸引了相当多的关注,不同的应用,包括隐私保护,声纹再现死者,和构音障碍的语音恢复。然而,由于训练数据干净,这些模型只学习非鲁棒性特征。随后,当处理真实世界场景中的降级输入语音时,包括额外的噪声,混响,对抗性攻击,甚至微小的扰动,它会导致不令人满意的性能。因此,它需要强大的部署,特别是在现实环境中。虽然最新的研究试图找到潜在的攻击和对策的VC系统,仍然有一个显着的差距,在输入操作下的VC模型是如何强大的全面理解。这里也提出了许多问题:例如,不同形式的输入退化攻击在多大程度上改变了VC模型的预期输出?是否有优化这些攻击和防御策略的潜力?为了回答这些问题,我们从输入操作的角度对现有的攻击和防御方法进行分类,并在四个维度上评估降级输入语音的影响,包括可懂度,自然度,音色相似性和主观感知。最后,我们概述了开放的问题和未来的方向。
摘要:Identity, accent, style, and emotions are essential components of human speech. Voice conversion (VC) techniques process the speech signals of two input speakers and other modalities of auxiliary information such as prompts and emotion tags. It changes para-linguistic features from one to another, while maintaining linguistic contents. Recently, VC models have made rapid advancements in both generation quality and personalization capabilities. These developments have attracted considerable attention for diverse applications, including privacy preservation, voice-print reproduction for the deceased, and dysarthric speech recovery. However, these models only learn non-robust features due to the clean training data. Subsequently, it results in unsatisfactory performances when dealing with degraded input speech in real-world scenarios, including additional noise, reverberation, adversarial attacks, or even minor perturbation. Hence, it demands robust deployments, especially in real-world settings. Although latest researches attempt to find potential attacks and countermeasures for VC systems, there remains a significant gap in the comprehensive understanding of how robust the VC model is under input manipulation. here also raises many questions: For instance, to what extent do different forms of input degradation attacks alter the expected output of VC models? Is there potential for optimizing these attack and defense strategies? To answer these questions, we classify existing attack and defense methods from the perspective of input manipulation and evaluate the impact of degraded input speech across four dimensions, including intelligibility, naturalness, timbre similarity, and subjective perception. Finally, we outline open issues and future directions.
【4】KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
标题:KidSpeak:用于儿童语音识别和筛查的通用多用途LLM
链接:https://arxiv.org/abs/2512.05994
摘要:随着基于对话和扩散的人工智能的快速发展,人工智能在教育服务中的应用越来越多,从评分和评估工具到为学生提供有针对性支持的个性化学习系统。然而,这种适应性尚未完全扩展到儿童语音领域,现有的模型往往由于依赖于为清晰,清晰的成人语音设计的数据集而失败。儿童,特别是那些处于早期发育阶段或患有言语和语言疾病的儿童,提出了当前人工智能模型和数据集无法处理的独特挑战。为了解决这个问题,我们介绍了KidSpeak,一个多任务语音增强的基础模型,能够生成和区分专门针对儿童的语音模式的任务。我们的框架采用了一个两阶段的训练过程,将语音知识融入语音编码器,在四个独立的任务中实现了87%的平均准确率。此外,认识到可扩展的人类注释和现有的语音对齐工具的局限性,我们提出了灵活和自动语音对齐器(FASA),并利用该方法来构建高质量的数据集进行训练和评估。这种新颖的对齐工具显着提高了从噪声数据中对齐儿童语音的质量,与人类注释相比,数据质量提高了13.6倍,如CHILDES数据集所示。据我们所知,KidSpeak和FASA代表了第一个为儿童语音和语言治疗设计的综合解决方案,提供了多功能语音LLM和强大的对齐工具。
摘要:With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for students. However, this adaptability has yet to fully extend to the domain of children's speech, where existing models often fail due to their reliance on datasets designed for clear, articulate adult speech. Children, particularly those in early developmental stages or with speech and language pathologies, present unique challenges that current AI models and datasets are ill-equipped to handle. To address this, we introduce KidSpeak, a multi-task speech-enhanced Foundation Model capable of both generative and discriminative tasks specifically tailored to children's speech patterns. Our framework employs a two-stage training process that incorporates phonetic knowledge into the speech encoder, achieving an average accuracy of 87% across four separate tasks. Furthermore, recognizing the limitations of scalable human annotation and existing speech alignment tools, we propose the Flexible and Automatic Speech Aligner (FASA) and leverage the method to construct high quality datasets for training and evaluation. This novel alignment tool significantly improves the quality of aligned children's speech from noisy data, enhancing data quality by 13.6x compared to human annotations, as demonstrated on the CHILDES dataset. To the best of our knowledge, KidSpeak and FASA represent the first comprehensive solution designed for speech and language therapy in children, offering both a multi-purpose speech LLM and a robust alignment tool.
【5】Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
标题:低资源语言的高效ASB:利用跨语言无标签数据
链接:https://arxiv.org/abs/2512.07277
备注:Accepted in AACL IJCNLP 2025
摘要:低资源语言的自动语音识别仍然从根本上受到最先进模型所需的标记数据和计算资源的稀缺性的限制。我们对低资源语言的跨语言连续预训练进行了系统的研究,使用波斯语-阿拉伯语(波斯语,阿拉伯语和乌尔都语)作为我们的主要案例研究。我们的方法表明,战略利用未标记的语音数据可以有效地弥合资源差距,而不牺牲识别精度。我们通过一个可扩展的未标记数据收集管道构建了一个3,000小时的多语言语料库,并采用有针对性的持续预训练结合形态感知标记化来开发一个300 M参数模型,其性能可与5倍大的系统相媲美。我们的模型在波斯语上的表现优于Whisper Large v3(1.5B参数),并且在阿拉伯语和乌尔都语上取得了有竞争力的结果,尽管使用了更少的参数和更少的标记数据。这些发现挑战了普遍的假设,即ASR质量主要与模型大小有关,相反,数据相关性和战略预训练是低资源场景的更关键因素。这项工作为包容性语音技术提供了一条切实可行的途径,为代表性不足的语言提供了有效的ASR,而不依赖于大规模的计算基础设施或专有数据集。
摘要:Automatic speech recognition for low-resource languages remains fundamentally constrained by the scarcity of labeled data and computational resources required by state-of-the-art models. We present a systematic investigation into cross-lingual continuous pretraining for low-resource languages, using Perso-Arabic languages (Persian, Arabic, and Urdu) as our primary case study. Our approach demonstrates that strategic utilization of unlabeled speech data can effectively bridge the resource gap without sacrificing recognition accuracy. We construct a 3,000-hour multilingual corpus through a scalable unlabeled data collection pipeline and employ targeted continual pretraining combined with morphologically-aware tokenization to develop a 300M parameter model that achieves performance comparable to systems 5 times larger. Our model outperforms Whisper Large v3 (1.5B parameters) on Persian and achieves competitive results on Arabic and Urdu despite using significantly fewer parameters and substantially less labeled data. These findings challenge the prevailing assumption that ASR quality scales primarily with model size, revealing instead that data relevance and strategic pretraining are more critical factors for low-resource scenarios. This work provides a practical pathway toward inclusive speech technology, enabling effective ASR for underrepresented languages without dependence on massive computational infrastructure or proprietary datasets.
【6】TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation
标题:TeluguST-46:Teluguu-英语语音翻译的基准数据库和综合评估
链接:https://arxiv.org/abs/2512.07265
备注:Submitted to AACL IJCNLP 2025
摘要:尽管泰卢固语被超过8000万人使用,但这种形态丰富的语言的语音翻译研究仍然严重不足。我们通过开发一个高质量的泰卢固语-英语语音翻译基准,从46小时的人工验证的CSTD语料库数据(30小时/8小时/8小时训练/开发/测试分裂)来解决这个差距。我们对级联与端到端架构的系统比较表明,虽然IndicWhisper + IndicMT由于大量的泰卢固语特定训练数据而实现了最高性能,但微调的无源M4 T模型尽管使用了显著较少的泰卢固语特定训练数据,但仍表现出显着的竞争力。这一发现表明,通过仔细的超参数调整和足够的并行数据(可能少于100小时),端到端系统可以在低资源环境中实现与级联方法相当的性能。我们的度量可靠性研究评估BLEU,METEOR,ChrF++,ROUGE-L,TER和BERTScore对人类的判断表明,传统的指标提供了更好的质量比BERTScore的区分泰卢固语-英语翻译。这项工作提供了三个关键的贡献:一个可重复的泰卢固语-英语基准,在低资源的情况下竞争性的端到端的性能潜力的经验证据,并在形态复杂的语言对自动评估的实际指导。
摘要:Despite Telugu being spoken by over 80 million people, speech translation research for this morphologically rich language remains severely underexplored. We address this gap by developing a high-quality Telugu--English speech translation benchmark from 46 hours of manually verified CSTD corpus data (30h/8h/8h train/dev/test split). Our systematic comparison of cascaded versus end-to-end architectures shows that while IndicWhisper + IndicMT achieves the highest performance due to extensive Telugu-specific training data, finetuned SeamlessM4T models demonstrate remarkable competitiveness despite using significantly less Telugu-specific training data. This finding suggests that with careful hyperparameter tuning and sufficient parallel data (potentially less than 100 hours), end-to-end systems can achieve performance comparable to cascaded approaches in low-resource settings. Our metric reliability study evaluating BLEU, METEOR, ChrF++, ROUGE-L, TER, and BERTScore against human judgments reveals that traditional metrics provide better quality discrimination than BERTScore for Telugu--English translation. The work delivers three key contributions: a reproducible Telugu--English benchmark, empirical evidence of competitive end-to-end performance potential in low-resource scenarios, and practical guidance for automatic evaluation in morphologically complex language pairs.
【7】JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
标题:JEPA作为神经标记器:使用密度自适应注意力学习鲁棒的语音表示
链接:https://arxiv.org/abs/2512.07168
备注:UniReps: Unifying Representations in Neural Models (NeurIPS 2025 Workshop)
摘要:我们引入了一个两阶段的自监督框架,该框架将联合嵌入预测架构(JEPA)与密度自适应注意力机制(DAAM)相结合,用于学习鲁棒的语音表示。阶段1使用JEPA和DAAM通过潜在空间中的掩蔽预测来学习语义音频特征,与波形重构完全解耦。阶段2利用这些表示来使用有限标量量化(FSQ)和混合基数打包方案进行有效的令牌化,然后使用HiFi-GAN解码器进行高保真波形重构。该模型通过在JEPA编码器中引入基于高斯混合的密度自适应选通,实现了自适应时域特征选择,并在2.5Hz的低帧率下发现了语音的层次结构。由此产生的令牌(47.5令牌/秒)提供了一种可逆的,高度压缩的,语言模型友好的表示,与现有的神经音频编解码器相比具有竞争力,并且通常更有效。
摘要:We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM to learn semantic audio features via masked prediction in latent space, fully decoupled from waveform reconstruction. Stage~2 leverages these representations for efficient tokenization using Finite Scalar Quantization (FSQ) and a mixed-radix packing scheme, followed by high-fidelity waveform reconstruction with a HiFi-GAN decoder. By integrating Gaussian mixture-based density-adaptive gating into the JEPA encoder, the model performs adaptive temporal feature selection and discovers hierarchical speech structure at a low frame rate of 2.5~Hz. The resulting tokens (47.5 tokens/sec) provide a reversible, highly compressed, and language-model-friendly representation that is competitive with, and often more efficient than, existing neural audio codecs.
【8】Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
标题:用于统一语音增强和分离的轻量级Wasserstein视听模型
链接:https://arxiv.org/abs/2512.06689
备注:Accepted to ASRU 2025
摘要:语音增强(SE)和语音分离(SS)传统上被视为语音处理中的不同任务。然而,现实世界的音频通常涉及背景噪声和重叠的扬声器,因此需要一个统一的解决方案。虽然最近的方法试图将SE和SS集成到多级架构中,但这些方法通常涉及复杂的参数密集型模型,并依赖于监督训练,限制了可扩展性和泛化能力。在这项工作中,我们提出了UniVoiceLite,一个轻量级的和无监督的视听框架,统一SE和SS在一个单一的模型。UniVoiceLite利用嘴唇运动和面部身份线索来指导语音提取,并采用Wasserstein距离正则化来稳定潜在空间,而不需要配对的噪声干净数据。实验结果表明,UniVoiceLite实现了强大的性能,在嘈杂和多扬声器的情况下,结合效率与鲁棒的泛化。源代码可在https://github.com/jisoo-o/UniVoiceLite上获得。
摘要:Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivating the need for a unified solution. While recent approaches have attempted to integrate SE and SS within multi-stage architectures, these approaches typically involve complex, parameter-heavy models and rely on supervised training, limiting scalability and generalization. In this work, we propose UniVoiceLite, a lightweight and unsupervised audio-visual framework that unifies SE and SS within a single model. UniVoiceLite leverages lip motion and facial identity cues to guide speech extraction and employs Wasserstein distance regularization to stabilize the latent space without requiring paired noisy-clean data. Experimental results demonstrate that UniVoiceLite achieves strong performance in both noisy and multi-speaker scenarios, combining efficiency with robust generalization. The source code is available at https://github.com/jisoo-o/UniVoiceLite.
【9】Technical Report of Nomi Team in the Environmental Sound Deepfake Detection Challenge 2026
标题:Nomi团队在2026年环境之声Deepfake检测挑战赛中的技术报告
链接:https://arxiv.org/abs/2512.06041
摘要:本文介绍了我们在ICASSP 2026环境声音深度伪造检测(ESDD)挑战赛中的工作。该挑战基于由各种合成环境声音组成的大规模EnvSDD数据集。我们专注于解决看不见的发电机和低资源的黑盒场景的复杂性,提出了一个音频文本交叉注意力模型。单独和组合的文本-音频模型的实验证明了在挑战基线(BEAT +AASIST模型)上具有竞争力的EER改进。
摘要:This paper presents our work for the ICASSP 2026 Environmental Sound Deepfake Detection (ESDD) Challenge. The challenge is based on the large-scale EnvSDD dataset that consists of various synthetic environmental sounds. We focus on addressing the complexities of unseen generators and low-resource black-box scenarios by proposing an audio-text cross-attention model. Experiments with individual and combined text-audio models demonstrate competitive EER improvements over the challenge baseline (BEATs+AASIST model).
【10】Physics-Guided Deepfake Detection for Voice Authentication Systems
标题:语音认证系统的物理引导Deepfake检测
链接:https://arxiv.org/abs/2512.06040
摘要:部署在网络边缘的语音认证系统面临双重威胁:a)复杂的deepfake合成攻击和b)分布式联邦学习协议中的控制平面中毒。我们提出了一个框架,将物理引导的deepfake检测与边缘学习中的不确定性感知相结合。该框架融合了可解释的物理特征建模声道动态与来自自监督学习模块的表示。然后,通过多模态Ensemble架构处理表示,然后通过贝叶斯系综提供不确定性估计。对音频样本进行基于物理的特征评估和不确定性估计,使我们提出的框架能够对先进的deepfake攻击和复杂的控制平面中毒保持鲁棒性,从而解决网络语音认证的完整威胁模型。
摘要:Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. We present a framework coupling physics-guided deepfake detection with uncertainty-aware in edge learning. The framework fuses interpretable physics features modeling vocal tract dynamics with representations coming from a self-supervised learning module. The representations are then processed via a Multi-Modal Ensemble Architecture, followed by a Bayesian ensemble providing uncertainty estimates. Incorporating physics-based characteristics evaluations and uncertainty estimates of audio samples allows our proposed framework to remain robust to both advanced deepfake attacks and sophisticated control-plane poisoning, addressing the complete threat model for networked voice authentication.
机器翻译由腾讯交互翻译提供,仅供参考
