今日论文合集:cs.SD语音8篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Emovectors: assessing emotional content in jazz improvisations for creativity evaluation
标题:格式载体:评估爵士乐即兴演奏中的情感内容以进行创造力评估
链接:https://arxiv.org/abs/2512.08812

作者:Anna Jordanous
备注:Presented at IEEE Big Data 2025 3rd Workshop on AI Music Generation (AIMG 2025). https://www.intellisky.org/workshops/AIMG2025/workshop_AIMG2025.html
摘要:音乐即兴是迷人的研究,本质上是一个创造性的过程的现场演示。在爵士乐中,音乐家经常在预定义的和弦进行中即兴创作(leadsheets)。我们如何评价爵士乐即兴创作的创造力?我们能在当前基于LLM的生成系统的创造力自动化指标中捕获这一点吗?情感投入的表现与即兴创作的创造力密切相关。分析音乐音频,我们能检测到情感参与吗?这项研究假设,如果即兴创作包含更多的情感负载内容的证据,它更有可能被认为是创造性的。提出了一种基于嵌入的方法来捕获音乐即兴创作中的情感内容,使用与情感相关的音乐特征的心理学基础分类。分析所产生的“情感载体”来测试上述假设,并在多个即兴表演中进行比较。以这种可量化的方式捕捉情感内容可以为创造力评估提供新的指标,这些指标可以大规模应用。
摘要:Music improvisation is fascinating to study, being essentially a live demonstration of a creative process. In jazz, musicians often improvise across predefined chord progressions (leadsheets). How do we assess the creativity of jazz improvisations? And can we capture this in automated metrics for creativity for current LLM-based generative systems? Demonstration of emotional involvement is closely linked with creativity in improvisation. Analysing musical audio, can we detect emotional involvement? This study hypothesises that if an improvisation contains more evidence of emotion-laden content, it is more likely to be recognised as creative. An embeddings-based method is proposed for capturing the emotional content in musical improvisations, using a psychologically-grounded classification of musical characteristics associated with emotions. Resulting 'emovectors' are analysed to test the above hypothesis, comparing across multiple improvisations. Capturing emotional content in this quantifiable way can contribute towards new metrics for creativity evaluation that can be applied at scale.


【2】DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components
标题:DFALLM:通过优化音频LLM组件实现可推广的多任务Deepfake检测
链接:https://arxiv.org/abs/2512.08403

作者:Yupei Li,Li Wang,Yuxiang Wang,Lei Wang,Rizhao Cai,Jie Shi,Björn W. Schuller,Zhizheng Wu
摘要:音频deepfake检测最近因其对安全性和可靠性的影响而引起了公众的关注。传统的深度学习方法已被广泛应用于这一任务,但在面对新出现的欺骗技术和更多任务(如欺骗属性识别而不是简单的二进制分类)时,往往缺乏通用性。原则上,大型语言模型(LLM)被认为具有所需的泛化能力。然而,先前对音频LLM(ALLM)的研究表明,即使有足够的数据可用,音频deepfake检测性能也存在泛化瓶颈。因此,本研究调查的模型架构,并检查ALLM的主要组成部分,即音频编码器和基于文本的LLM的影响。我们的实验表明,音频编码器和基于文本的LLM的仔细选择和组合对于释放ALLM的深度伪造检测潜力至关重要。我们进一步提出了一种ALLM结构,能够将deepfake检测能力推广到域外欺骗测试和其他deepfake任务,例如欺骗定位和欺骗属性识别。我们提出的模型架构在多个数据集(包括ASVSpoof2019,InTheWild和Demopage)上实现了最先进的(SOTA)性能,平均准确率高达95.76%,并且与SOTA音频理解模型相比,在其他deepfake检测任务(如归因和本地化)中表现出竞争力。数据和代码在补充材料中提供。
摘要:Audio deepfake detection has recently garnered public concern due to its implications for security and reliability. Traditional deep learning methods have been widely applied to this task but often lack generalisability when confronted with newly emerging spoofing techniques and more tasks such as spoof attribution recognition rather than simple binary classification. In principle, Large Language Models (LLMs) are considered to possess the needed generalisation capabilities. However, previous research on Audio LLMs (ALLMs) indicates a generalization bottleneck in audio deepfake detection performance, even when sufficient data is available. Consequently, this study investigates the model architecture and examines the effects of the primary components of ALLMs, namely the audio encoder and the text-based LLM. Our experiments demonstrate that the careful selection and combination of audio encoders and text-based LLMs are crucial for unlocking the deepfake detection potential of ALLMs. We further propose an ALLM structure capable of generalizing deepfake detection abilities to out-of-domain spoofing tests and other deepfake tasks, such as spoof positioning and spoof attribution recognition. Our proposed model architecture achieves state-of-the-art (SOTA) performance across multiple datasets, including ASVSpoof2019, InTheWild, and Demopage, with accuracy reaching up to 95.76% on average, and exhibits competitive capabilities in other deepfake detection tasks such as attribution, and localisation compared to SOTA audio understanding models. Data and codes are provided in supplementary materials.


【3】PAVAS: Physics-Aware Video-to-Audio Synthesis
标题:PAVAS:物理感知的视频到音频合成
链接:https://arxiv.org/abs/2512.08282

作者:Oh Hyun-Bin,Yuhta Takida,Toshimitsu Uesaka,Tae-Hyun Oh,Yuki Mitsufuji
摘要:视频到音频(V2 A)生成的最新进展已经实现了令人印象深刻的感知质量和时间同步,但大多数模型仍然是外观驱动的,在不考虑塑造真实世界声音的物理因素的情况下捕获视觉-声学相关性。我们提出了物理感知视频到音频合成(PAVAS),一种通过物理驱动的音频适配器(Phy-Adapter)将物理推理纳入基于潜在扩散的V2 A生成的方法。适配器接收由物理参数估计器(PPE)估计的对象级物理参数,PPE使用视觉语言模型(VLM)来推断移动对象质量和基于分割的动态3D重建模块来恢复其运动轨迹以用于速度计算。这些物理线索使模型能够合成反映潜在物理因素的声音。为了评估物理真实性,我们策划了VGG-Impact,这是一个专注于对象与对象交互的基准测试,并引入了音频物理相关系数(APCC),这是一个衡量物理和听觉属性之间一致性的评估指标。综合实验表明,PAVAS产生物理上合理和感知上连贯的音频,在定量和定性评估方面优于现有的V2 A模型。访问https://physics-aware-video-to-audio-synthesis.github.io获取演示视频。
摘要:Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing visual-acoustic correlations without considering the physical factors that shape real-world sounds. We present Physics-Aware Video-to-Audio Synthesis (PAVAS), a method that incorporates physical reasoning into a latent diffusion-based V2A generation through the Physics-Driven Audio Adapter (Phy-Adapter). The adapter receives object-level physical parameters estimated by the Physical Parameter Estimator (PPE), which uses a Vision-Language Model (VLM) to infer the moving-object mass and a segmentation-based dynamic 3D reconstruction module to recover its motion trajectory for velocity computation. These physical cues enable the model to synthesize sounds that reflect underlying physical factors. To assess physical realism, we curate VGG-Impact, a benchmark focusing on object-object interactions, and introduce Audio-Physics Correlation Coefficient (APCC), an evaluation metric that measures consistency between physical and auditory attributes. Comprehensive experiments show that PAVAS produces physically plausible and perceptually coherent audio, outperforming existing V2A models in both quantitative and qualitative evaluations. Visit https://physics-aware-video-to-audio-synthesis.github.io for demo videos.


【4】SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
标题:SpeechQualityLLM:基于LLM的多模态语音质量评估
链接:https://arxiv.org/abs/2512.08238

作者:Mahathir Monjur,Shahriar Nirjon
备注:9 pages, 5 figures, 8 tables
摘要:客观的语音质量评估是电话、VoIP和流媒体系统的核心,在这些系统中,必须大规模监控和优化大量降级的音频。PESQ和POLQA等经典指标近似于人类平均意见分数(MOS),但需要仔细控制条件和昂贵的听力测试,而NISQA等基于学习的模型则从波形或频谱图中回归MOS和多个感知维度,实现与主观评级的高度相关性,但仍然保持刚性:它们不支持交互式自然语言查询,也不提供文本依据。在这项工作中,我们介绍了SpeechQualityLLM,一个多模式语音质量问答(QA)系统,它将音频编码器与语言模型耦合在一起,并使用基于模板的问答对在NISQA语料库上进行训练,这些问答对涵盖了单端(仅降级)和双端(降级加干净参考)设置中的整体MOS和四个感知维度(噪音,着色,不连续性和响度)。我们的系统不是直接回归分数,而是监督生成文本答案,从中解析数字预测并使用标准回归和排名指标进行评估;在NISQA剪辑上,双端模型的MOS平均绝对误差(MAE)为0.41,Pearson相关系数为0.86,在维度任务上具有竞争力。除了这些量化收益,它还提供了一个灵活的自然语言界面,其中语言模型充当音频质量专家:从业者可以查询退化的任意方面,促使模型模拟不同的听众配置文件以捕获人类的变化并产生多样化但合理的判断,而不是单一的确定性分数,从而减少对大规模众包测试及其货币成本的依赖。
摘要:Objective speech quality assessment is central to telephony, VoIP, and streaming systems, where large volumes of degraded audio must be monitored and optimized at scale. Classical metrics such as PESQ and POLQA approximate human mean opinion scores (MOS) but require carefully controlled conditions and expensive listening tests, while learning-based models such as NISQA regress MOS and multiple perceptual dimensions from waveforms or spectrograms, achieving high correlation with subjective ratings yet remaining rigid: they do not support interactive, natural-language queries and do not natively provide textual rationales. In this work, we introduce SpeechQualityLLM, a multimodal speech quality question-answering (QA) system that couples an audio encoder with a language model and is trained on the NISQA corpus using template-based question-answer pairs covering overall MOS and four perceptual dimensions (noisiness, coloration, discontinuity, and loudness) in both single-ended (degraded only) and double-ended (degraded plus clean reference) setups. Instead of directly regressing scores, our system is supervised to generate textual answers from which numeric predictions are parsed and evaluated with standard regression and ranking metrics; on held-out NISQA clips, the double-ended model attains a MOS mean absolute error (MAE) of 0.41 with Pearson correlation of 0.86, with competitive performance on dimension-wise tasks. Beyond these quantitative gains, it offers a flexible natural-language interface in which the language model acts as an audio quality expert: practitioners can query arbitrary aspects of degradations, prompt the model to emulate different listener profiles to capture human variability and produce diverse but plausible judgments rather than a single deterministic score, and thereby reduce reliance on large-scale crowdsourced tests and their monetary cost.


【5】Error-Resilient Semantic Communication for Speech Transmission over Packet-Loss Networks
标题:分组丢失网络上语音传输的抗错误语义通信
链接:https://arxiv.org/abs/2512.08203

作者:Zhuohang Han,Jincheng Dai,Shengshi Yao,Junyi Wang,Yanlong Li,Kai Niu,Wenjun Xu,Ping Zhang
备注:submitted to IEEE in Nov. 2025
摘要:无线网络上的实时语音通信仍然具有挑战性,因为传统的信道保护机制在严格的带宽和延迟约束下不能有效地对抗分组丢失。语义通信已经成为一个很有前途的范例,以提高语音传输的鲁棒性,通过联合信源信道编码(JSCC)。然而,由于其与现有数字通信系统的不兼容性,其跨层设计阻碍了实际部署。在这种情况下,语音通信的鲁棒性因此主要通过对无线网络上的分组丢失的错误恢复来评估。为了解决这些挑战,我们提出了一个基于生成潜在先验的弹性语音语义通信框架,在生成潜在空间中执行弹性语音编码。生成潜在先验在接收器侧实现高质量的分组丢失隐藏(PLC),很好地平衡了语义一致性和重建保真度。此外,一个集成的错误恢复机制的设计,以减轻错误传播,提高PLC的有效性。与传统的分组级前向纠错(FEC)策略相比,我们的新方法在动态无线网络中实现了增强的鲁棒性,同时显着减少冗余开销。在LibriSpeech数据集上的实验结果表明,Glaris始终优于现有的容错编解码器,在保持与现有系统无缝兼容的同时实现了JSCC级别的鲁棒性,并且在传输效率和语音重建质量之间取得了良好的平衡。
摘要:Real-time speech communication over wireless networks remains challenging, as conventional channel protection mechanisms cannot effectively counter packet loss under stringent bandwidth and latency constraints. Semantic communication has emerged as a promising paradigm for enhancing the robustness of speech transmission by means of joint source-channel coding (JSCC). However, its cross-layer design hinders practical deployment due to the incompatibility with existing digital communication systems. In this case, the robustness of speech communication is consequently evaluated primarily by the error-resilience to packet loss over wireless networks. To address these challenges, we propose \emph{Glaris}, a generative latent-prior-based resilient speech semantic communication framework that performs resilient speech coding in the generative latent space. Generative latent priors enable high-quality packet loss concealment (PLC) at the receiver side, well-balancing semantic consistency and reconstruction fidelity. Additionally, an integrated error resilience mechanism is designed to mitigate the error propagation and improve the effectiveness of PLC. Compared with traditional packet-level forward error correction (FEC) strategies, our new method achieves enhanced robustness over dynamic wireless networks while reducing redundancy overhead significantly. Experimental results on the LibriSpeech dataset demonstrate that \emph{Glaris} consistently outperforms existing error-resilient codecs, achieving JSCC-level robustness while maintaining seamless compatibility with existing systems, and it also strikes a favorable balance between transmission efficiency and speech reconstruction quality.


【6】Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
标题:超越统一模型:面向服务的低延迟、上下文感知语音合成方法
链接:https://arxiv.org/abs/2512.08006

作者:Mahta Fetrat,Donya Navabi,Zahra Dehghanian,Morteza Abolghasemi,Hamid R. Rabiee
摘要:轻量级、实时的文本到语音系统对于可访问性至关重要。然而,最有效的TTS模型往往依赖于轻量级的音素,与上下文相关的挑战斗争。相比之下,具有更深的语言理解的更高级的音素生成器通常会产生高的计算成本,这阻碍了实时性能。   本文研究了G2P辅助TTS系统中音素化质量和推理速度之间的权衡,介绍了一个实用的框架来弥合这一差距。我们提出了轻量级的策略,上下文感知音素化和面向服务的TTS架构,执行这些模块作为独立的服务。这种设计将大量上下文感知组件从核心TTS引擎中分离出来,有效地打破了延迟障碍,并实现了高质量音素化模型的实时使用。实验结果证实,该系统提高了发音的声音和语言的准确性,同时保持实时响应,使其非常适合离线和终端设备的TTS应用。
摘要:Lightweight, real-time text-to-speech systems are crucial for accessibility. However, the most efficient TTS models often rely on lightweight phonemizers that struggle with context-dependent challenges. In contrast, more advanced phonemizers with a deeper linguistic understanding typically incur high computational costs, which prevents real-time performance.   This paper examines the trade-off between phonemization quality and inference speed in G2P-aided TTS systems, introducing a practical framework to bridge this gap. We propose lightweight strategies for context-aware phonemization and a service-oriented TTS architecture that executes these modules as independent services. This design decouples heavy context-aware components from the core TTS engine, effectively breaking the latency barrier and enabling real-time use of high-quality phonemization models. Experimental results confirm that the proposed system improves pronunciation soundness and linguistic accuracy while maintaining real-time responsiveness, making it well-suited for offline and end-device TTS applications.


【7】LocaGen: Sub-Sample Time-Delay Learning for Beam Localization
标题:LocaGen:用于束定位的子样本延时学习
链接:https://arxiv.org/abs/2512.07872

作者:Ishaan Kunwar,Henry Cantor,Tyler Rizzo,Ayaan Qayyum
备注:7 pages
摘要:LocaGen的目标是提高音频信号在2-D波束定位问题中的定位性能。LocaGen通过在模拟生成的真实合成数据上训练的机器学习模型来减少采样量化误差。该系统提高了来自三个麦克风的阵列的音频波束的到达方向(DOA)和精确位置估计的准确性。我们证明了LocaGen的低功耗嵌入式系统的效率,提高了定位精度,在实时资源使用量的最小增加。LocaGen被证明可以将DOA误差降低约67%,即使在音频处理中麦克风阵列仅为10 kHz。
摘要:The goal of LocaGen is to improve the localization performance of audio signals in the 2-D beam localization problem. LocaGen reduces sampling quantization errors through machine learning models trained on realistic synthetic data generated by a simulation. The system increases the accuracy of both direction-of-arrival (DOA) and precise location estimation of an audio beam from an array of three microphones. We demonstrate LocaGen's efficacy on a low-powered embedded system with an increased localization accuracy with a minimal increase in real-time resource usage. LocaGen was demonstrated to reduce DOA error by approximately 67% even with a microphone array of only 10 kHz in audio processing.


【8】AudioScene: Integrating Object-Event Audio into 3D Scenes
标题:AudioScene:将对象事件音频集成到3D场景中
链接:https://arxiv.org/abs/2512.07845

作者:Shuaihang Yuan,Congcong Wen,Muhammad Shafique,Anthony Tzes,Yi Fang
摘要:音频分析的快速发展强调了其在人机交互、环境监测和公共安全方面的巨大潜力;然而,现有的仅音频数据集往往缺乏空间背景。为了解决这一差距,我们提出了两个新的音频空间场景数据集,AudioScanNet和AudioRoboTHOR,旨在探索3D环境中的音频任务。通过将音频片段与空间对齐的3D场景相集成,我们的数据集可以研究音频信号如何与空间背景相互作用。为了将音频事件与相应的空间信息相关联,我们利用大型语言模型的常识推理能力,并辅以严格的人工验证,与纯手动注释相比,这种方法提供了更大的可扩展性,同时保持了高标准的准确性,完整性和多样性,通过注释者之间的协议和两个基准任务上的性能进行量化,基于音频的3D视觉接地和基于音频的机器人零击导航。结果突出了当前以听觉为中心的方法的局限性,并强调了我们的数据集在推进音频引导空间学习方面的实际挑战和意义。
摘要:The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two novel audiospatial scene datasets, AudioScanNet and AudioRoboTHOR, designed to explore audioconditioned tasks within 3D environments. By integrating audio clips with spatially aligned 3D scenes, our datasets enable research on how audio signals interact with spatial context. To associate audio events with corresponding spatial information, we leverage the common sense reasoning ability of large language models and supplement them with rigorous human verification, This approach offers greater scalability compared to purely manual annotation while maintaining high standards of accuracy, completeness, and diversity, quantified through inter annotator agreement and performance on two benchmark tasks audio based 3D visual grounding and audio based robotic zeroshot navigation. The results highlight the limitations of current audiocentric methods and underscore the practical challenges and significance of our datasets in advancing audio guided spatial learning.


eess.AS音频处理


【1】BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge
标题:但ESDD 2026挑战赛中的环境声音Deepfake检测系统
链接:https://arxiv.org/abs/2512.08319

作者:Junyi Peng,Lin Zhang,Jin Li,Oldrich Plchot,Jan Cernocky
摘要:本文介绍了BUT提交给ESDD 2026挑战赛的申请,特别关注第一场:使用看不见的发生器进行环境声音深度伪造检测。为了解决推广到由看不见的合成算法生成的音频的关键挑战,我们提出了一个强大的集成框架,利用不同的自监督学习(SSL)模型。我们对通用音频SSL模型(包括BEAT,EAT和Dasheng)和语音特定SSL进行了全面分析。这些前端与轻量级的多头分解注意力(MHFA)后端相结合,以捕获有区别的表示。此外,我们引入了一个基于分布不确定性建模的特征域增强策略,以增强模型对不可见光谱失真的鲁棒性。所有模型都是在官方EnvSDD数据上训练的,不使用任何外部资源。实验结果表明,我们的方法的有效性:我们最好的单一系统实现的等错误率(EER)的0.00\%,4.60\%,和4.80\%的发展,进展(轨道1),和最终评估集,分别。融合系统进一步提高了泛化能力,在相同的分区上产生的EER分别为0.00%、3.52%和4.38%。
摘要:This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the critical challenge of generalizing to audio generated by unseen synthesis algorithms, we propose a robust ensemble framework leveraging diverse Self-Supervised Learning (SSL) models. We conduct a comprehensive analysis of general audio SSL models (including BEATs, EAT, and Dasheng) and speech-specific SSLs. These front-ends are coupled with a lightweight Multi-Head Factorized Attention (MHFA) back-end to capture discriminative representations. Furthermore, we introduce a feature domain augmentation strategy based on distribution uncertainty modeling to enhance model robustness against unseen spectral distortions. All models are trained exclusively on the official EnvSDD data, without using any external resources. Experimental results demonstrate the effectiveness of our approach: our best single system achieved Equal Error Rates (EER) of 0.00\%, 4.60\%, and 4.80\% on the Development, Progress (Track 1), and Final Evaluation sets, respectively. The fusion system further improved generalization, yielding EERs of 0.00\%, 3.52\%, and 4.38\% across the same partitions.


【2】An Adaptive Method for Target Curve Selection
标题:目标曲线选择的自适应方法
链接:https://arxiv.org/abs/2512.08313

作者:Gabriele Ravizza,Julián Villegas,Christer P. Volk,Tore Stegenborg-Andersen,Yan Pei
备注:8 pages,6 figures. Accepted for presentation at the Audio Engineering Society (AES) International Conference on Headphone Technology, 2025
摘要:在本文中,我们介绍了适应的“交互式差分进化”(IDE)算法的音频域的任务,确定消费者之间的首选过耳耳机频率响应目标。该方法是基于使用自适应配对评级听力测试范例(配对比较规模)的数据收集。详细说明了IDE算法及其参数。此外,收集的数据从三个听实验,超过20个消费者,算法的性能在这个未经测试的域的基础上,两个收敛措施进行了研究。结果表明,该方法可以收敛,并可以减轻任务的“提取”的频率响应偏好从未经训练的消费者。
摘要:In this paper, we introduce an adaptation of the "Interactive Differential Evolution" (IDE) algorithm to the audio domain for the task of identifying the preferred over-the-ear headphone frequency response target among consumers. The method is based on data collection using an adaptive paired rating listening test paradigm (paired comparison with a scale). The IDE algorithm and its parameters are explained in detail. Additionally, data collected from three listening experiments with more than 20 consumers is presented, and the algorithm's performance in this untested domain is investigated on the basis of two convergence measures. The results indicate that this method can converge and may ease the task of 'extracting' frequency response preference from untrained consumers.


【3】Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
标题:超越统一模型:面向服务的低延迟、上下文感知语音合成方法
链接:https://arxiv.org/abs/2512.08006

作者:Mahta Fetrat,Donya Navabi,Zahra Dehghanian,Morteza Abolghasemi,Hamid R. Rabiee
摘要:轻量级、实时的文本到语音系统对于可访问性至关重要。然而,最有效的TTS模型往往依赖于轻量级的音素,与上下文相关的挑战斗争。相比之下,具有更深的语言理解的更高级的音素生成器通常会产生高的计算成本,这阻碍了实时性能。   本文研究了G2P辅助TTS系统中音素化质量和推理速度之间的权衡,介绍了一个实用的框架来弥合这一差距。我们提出了轻量级的策略,上下文感知音素化和面向服务的TTS架构,执行这些模块作为独立的服务。这种设计将大量上下文感知组件从核心TTS引擎中分离出来,有效地打破了延迟障碍,并实现了高质量音素化模型的实时使用。实验结果证实,该系统提高了发音的声音和语言的准确性,同时保持实时响应,使其非常适合离线和终端设备的TTS应用。
摘要:Lightweight, real-time text-to-speech systems are crucial for accessibility. However, the most efficient TTS models often rely on lightweight phonemizers that struggle with context-dependent challenges. In contrast, more advanced phonemizers with a deeper linguistic understanding typically incur high computational costs, which prevents real-time performance.   This paper examines the trade-off between phonemization quality and inference speed in G2P-aided TTS systems, introducing a practical framework to bridge this gap. We propose lightweight strategies for context-aware phonemization and a service-oriented TTS architecture that executes these modules as independent services. This design decouples heavy context-aware components from the core TTS engine, effectively breaking the latency barrier and enabling real-time use of high-quality phonemization models. Experimental results confirm that the proposed system improves pronunciation soundness and linguistic accuracy while maintaining real-time responsiveness, making it well-suited for offline and end-device TTS applications.


【4】LocaGen: Sub-Sample Time-Delay Learning for Beam Localization
标题:LocaGen:用于束定位的子样本延时学习
链接:https://arxiv.org/abs/2512.07872

作者:Ishaan Kunwar,Henry Cantor,Tyler Rizzo,Ayaan Qayyum
备注:7 pages
摘要:LocaGen的目标是提高音频信号在2-D波束定位问题中的定位性能。LocaGen通过在模拟生成的真实合成数据上训练的机器学习模型来减少采样量化误差。该系统提高了来自三个麦克风的阵列的音频波束的到达方向(DOA)和精确位置估计的准确性。我们证明了LocaGen的低功耗嵌入式系统的效率,提高了定位精度,在实时资源使用量的最小增加。LocaGen被证明可以将DOA误差降低约67%,即使在音频处理中麦克风阵列仅为10 kHz。
摘要:The goal of LocaGen is to improve the localization performance of audio signals in the 2-D beam localization problem. LocaGen reduces sampling quantization errors through machine learning models trained on realistic synthetic data generated by a simulation. The system increases the accuracy of both direction-of-arrival (DOA) and precise location estimation of an audio beam from an array of three microphones. We demonstrate LocaGen's efficacy on a low-powered embedded system with an increased localization accuracy with a minimal increase in real-time resource usage. LocaGen was demonstrated to reduce DOA error by approximately 67% even with a microphone array of only 10 kHz in audio processing.


【5】AudioScene: Integrating Object-Event Audio into 3D Scenes
标题:AudioScene:将对象事件音频集成到3D场景中
链接:https://arxiv.org/abs/2512.07845

作者:Shuaihang Yuan,Congcong Wen,Muhammad Shafique,Anthony Tzes,Yi Fang
摘要:音频分析的快速发展强调了其在人机交互、环境监测和公共安全方面的巨大潜力;然而,现有的仅音频数据集往往缺乏空间背景。为了解决这一差距,我们提出了两个新的音频空间场景数据集,AudioScanNet和AudioRoboTHOR,旨在探索3D环境中的音频任务。通过将音频片段与空间对齐的3D场景相集成,我们的数据集可以研究音频信号如何与空间背景相互作用。为了将音频事件与相应的空间信息相关联,我们利用大型语言模型的常识推理能力,并辅以严格的人工验证,与纯手动注释相比,这种方法提供了更大的可扩展性,同时保持了高标准的准确性,完整性和多样性,通过注释者之间的协议和两个基准任务上的性能进行量化,基于音频的3D视觉接地和基于音频的机器人零击导航。结果突出了当前以听觉为中心的方法的局限性,并强调了我们的数据集在推进音频引导空间学习方面的实际挑战和意义。
摘要:The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two novel audiospatial scene datasets, AudioScanNet and AudioRoboTHOR, designed to explore audioconditioned tasks within 3D environments. By integrating audio clips with spatially aligned 3D scenes, our datasets enable research on how audio signals interact with spatial context. To associate audio events with corresponding spatial information, we leverage the common sense reasoning ability of large language models and supplement them with rigorous human verification, This approach offers greater scalability compared to purely manual annotation while maintaining high standards of accuracy, completeness, and diversity, quantified through inter annotator agreement and performance on two benchmark tasks audio based 3D visual grounding and audio based robotic zeroshot navigation. The results highlight the limitations of current audiocentric methods and underscore the practical challenges and significance of our datasets in advancing audio guided spatial learning.


机器翻译由腾讯交互翻译提供,仅供参考