微信公众号:arXiv_Daily
cs.SD语音
【1】Reciprocal Latent Fields for Precomputed Sound Propagation
标题:预先计算的声音传播的相互潜场
链接:https://arxiv.org/abs/2602.06937
备注:Temporary pre-print, will be updated. In review at a conference
摘要:真实的声音传播对于沉浸在虚拟场景中至关重要,但物理精确的基于波的模拟对于实时应用来说仍然在计算上受到限制。波编码方法通过预先计算给定场景的脉冲响应并将其压缩成一组标量声学参数来解决这一限制,这些标量声学参数在具有许多源-接收器对的大型环境中可能达到难以管理的大小。我们介绍了互易潜场(RLF),一个内存效率的框架编码和预测这些声学参数。RLF框架采用用对称函数解码的可训练潜在嵌入的体积网格,确保声学互易性。我们研究了各种解码器,并表明利用黎曼度量学习可以更好地再现复杂场景中的声学现象。实验验证表明,RLF保持复制质量,同时减少了几个数量级的内存占用。此外,类似MUSHRA的主观听力测试表明,通过RLF呈现的声音在感知上与地面真实模拟无法区分。
摘要:Realistic sound propagation is essential for immersion in a virtual scene, yet physically accurate wave-based simulations remain computationally prohibitive for real-time applications. Wave coding methods address this limitation by precomputing and compressing impulse responses of a given scene into a set of scalar acoustic parameters, which can reach unmanageable sizes in large environments with many source-receiver pairs. We introduce Reciprocal Latent Fields (RLF), a memory-efficient framework for encoding and predicting these acoustic parameters. The RLF framework employs a volumetric grid of trainable latent embeddings decoded with a symmetric function, ensuring acoustic reciprocity. We study a variety of decoders and show that leveraging Riemannian metric learning leads to a better reproduction of acoustic phenomena in complex scenes. Experimental validation demonstrates that RLF maintains replication quality while reducing the memory footprint by several orders of magnitude. Furthermore, a MUSHRA-like subjective listening test indicates that sound rendered via RLF is perceptually indistinguishable from ground-truth simulations.
【2】DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
标题:DynFOA:通过条件扩散生成动态和声复杂的360度视频的一阶立体声
链接:https://arxiv.org/abs/2602.06846
摘要:空间音频对于创建引人入胜的沉浸式360度视频体验至关重要。然而,从复杂声学场景中的360度视频生成逼真的空间音频(诸如一阶立体混响(FOA))仍然具有挑战性。现有的方法往往忽略了360度场景的动态特性和声学复杂性,未能充分考虑动态声源,并且忽略了受场景几何形状和材料影响的复杂环境效应,例如遮挡、反射和混响。我们提出了DynFOA,一个基于动态声学感知和条件扩散的框架,用于从360度视频中生成高保真FOA。DynFOA首先通过视频编码器进行视觉处理,检测和定位多个动态声源,估计它们的深度和语义,并使用3D高斯溅射重建场景几何和材质。这种重建技术基于重建的3D场景的几何形状和材料以及收听者的视角来精确地对遮挡、反射和混响进行建模。音频编码器然后捕获空间运动和时间4D声源轨迹以微调基于扩散的FOA生成器。微调的FOA发生器实时调整空间线索,确保在听众头部旋转和复杂环境变化期间保持一致的方向保真度。广泛的评估表明,DynFOA在空间精度、声学保真度和分布匹配等指标上始终优于现有方法,同时还改善了用户体验。因此,DynFOA为VR和沉浸式媒体应用提供了一种强大且可扩展的方法来渲染逼真的动态空间音频。
摘要:Spatial audio is crucial for creating compelling immersive 360-degree video experiences. However, generating realistic spatial audio, such as first-order ambisonics (FOA), from 360-degree videos in complex acoustic scenes remains challenging. Existing methods often overlook the dynamic nature and acoustic complexity of 360-degree scenes, fail to fully account for dynamic sound sources, and neglect complex environmental effects such as occlusion, reflections, and reverberation, which are influenced by scene geometries and materials. We propose DynFOA, a framework based on dynamic acoustic perception and conditional diffusion, for generating high-fidelity FOA from 360-degree videos. DynFOA first performs visual processing via a video encoder, which detects and localizes multiple dynamic sound sources, estimates their depth and semantics, and reconstructs the scene geometry and materials using a 3D Gaussian Splatting. This reconstruction technique accurately models occlusion, reflections, and reverberation based on the geometries and materials of the reconstructed 3D scene and the listener's viewpoint. The audio encoder then captures the spatial motion and temporal 4D sound source trajectories to fine-tune the diffusion-based FOA generator. The fine-tuned FOA generator adjusts spatial cues in real time, ensuring consistent directional fidelity during listener head rotation and complex environmental changes. Extensive evaluations demonstrate that DynFOA consistently outperforms existing methods across metrics such as spatial accuracy, acoustic fidelity, and distribution matching, while also improving the user experience. Therefore, DynFOA provides a robust and scalable approach to rendering realistic dynamic spatial audio for VR and immersive media applications.
【3】AI-Generated Music Detection in Broadcast Monitoring
标题:广播监控中的人工智能生成音乐检测
链接:https://arxiv.org/abs/2602.06823
摘要:人工智能音乐生成器已经发展到它们的输出通常与人类作品无法区分的地步。虽然检测方法已经出现,但它们通常是在具有干净的全长曲目的音乐流媒体环境中设计和验证的。然而,广播音频提出了一个不同的挑战:音乐以简短的摘录出现,通常被主导语音掩盖,现有检测器在这种情况下失败。在这项工作中,我们介绍了AI-OpenBMAT,这是第一个为广播式AI音乐检测量身定制的数据集。它包含3,294个一分钟音频摘录(54.9小时),遵循真实电视音频的持续时间模式和响度关系,将人造制作音乐与Suno v3.5生成的风格匹配的延续相结合。我们对CNN基线和最先进的SpectTTTra模型进行基准测试,以评估SNR和持续时间的鲁棒性,并在完整的广播场景中进行评估。在所有设置中,在流媒体场景中表现出色的模型都会受到严重影响,当音乐处于背景或持续时间较短时,F1分数会下降到60%以下。这些结果强调了语音掩蔽和短音乐长度是AI音乐检测的关键挑战,并将AI-OpenBMAT定位为开发能够满足工业广播要求的检测器的基准。
摘要:AI music generators have advanced to the point where their outputs are often indistinguishable from human compositions. While detection methods have emerged, they are typically designed and validated in music streaming contexts with clean, full-length tracks. Broadcast audio, however, poses a different challenge: music appears as short excerpts, often masked by dominant speech, conditions under which existing detectors fail. In this work, we introduce AI-OpenBMAT, the first dataset tailored to broadcast-style AI-music detection. It contains 3,294 one-minute audio excerpts (54.9 hours) that follow the duration patterns and loudness relations of real television audio, combining human-made production music with stylistically matched continuations generated with Suno v3.5. We benchmark a CNN baseline and state-of-the-art SpectTTTra models to assess SNR and duration robustness, and evaluate on a full broadcast scenario. Across all settings, models that excel in streaming scenarios suffer substantial degradation, with F1-scores dropping below 60% when music is in the background or has a short duration. These results highlight speech masking and short music length as critical open challenges for AI music detection, and position AI-OpenBMAT as a benchmark for developing detectors capable of meeting industrial broadcast requirements.
【4】Hierarchical Activity Recognition and Captioning from Long-Form Audio
标题:从长格式音频中进行分层活动识别和字幕
链接:https://arxiv.org/abs/2602.06765
备注:Accepted by ICASSP 2026
摘要:真实世界音频中的复杂活动会在较长的持续时间内展开,并表现出层次结构,但大多数先前的工作都集中在短片段和孤立的事件上。为了弥合这一差距,我们引入了MultiAct,这是一个新的数据集和基准,用于从长格式音频中多层次结构化理解人类活动。MultiAct包括在三个语义级别(活动、子活动和事件)上注释的长时间厨房录音,并配有细粒度的标题和高级摘要。我们进一步提出了一个统一的分层模型,共同执行分类,检测,序列预测和多分辨率字幕。MultiAct上的实验建立了强大的基线,并揭示了对长格式音频的层次和组成结构进行建模的关键挑战。未来工作的一个有希望的方向是探索更适合捕捉长格式音频中复杂的长距离关系的方法。
摘要:Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge this gap, we introduce MultiAct, a new dataset and benchmark for multi-level structured understanding of human activities from long-form audio. MultiAct comprises long-duration kitchen recordings annotated at three semantic levels (activities, sub-activities and events) and paired with fine-grained captions and high-level summaries. We further propose a unified hierarchical model that jointly performs classification, detection, sequence prediction and multi-resolution captioning. Experiments on MultiAct establish strong baselines and reveal key challenges in modelling hierarchical and compositional structure of long-form audio. A promising direction for future work is the exploration of methods better suited to capturing the complex, long-range relationships in long-form audio.
【5】Scaling Speech Tokenizers with Diffusion Autoencoders
标题:使用扩散自动编码器缩放语音令牌器
链接:https://arxiv.org/abs/2602.06602
备注:ICLR 2026
摘要:语音标记器是语音语言模型的基础,然而现有方法面临两个主要挑战:(1)平衡用于理解的编码语义与用于重建的声学之间的权衡,以及(2)实现低比特率和低标记率。我们提出了语音扩散标记器(SiTok),这是一种扩散自动编码器,通过监督学习联合学习语义丰富的表示,并通过扩散实现高保真音频重建。我们将SiTok扩展到1.6 B参数,并在200万小时的语音上对其进行训练。实验表明,SiTok在理解,重建和生成任务方面优于强大的基线,其令牌速率极低,为12.5 $ Hz,比特率为200比特每秒。
摘要:Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understanding and acoustics for reconstruction, and (2) achieving low bit rates and low token rates. We propose Speech Diffusion Tokenizer (SiTok), a diffusion autoencoder that jointly learns semantic-rich representations through supervised learning and enables high-fidelity audio reconstruction with diffusion. We scale SiTok to 1.6B parameters and train it on 2 million hours of speech. Experiments show that SiTok outperforms strong baselines on understanding, reconstruction and generation tasks, at an extremely low token rate of $12.5$ Hz and a bit-rate of 200 bits-per-second.
【6】EMG-to-Speech with Fewer Channels
标题:使用更少的渠道进行EMG到言语
链接:https://arxiv.org/abs/2602.06460
摘要:表面肌电图(EMG)是一种很有前途的无声语音接口模式,但其有效性在很大程度上取决于传感器的位置和通道的可用性。在这项工作中,我们调查的贡献,个人和合并的肌电通道的语音重建性能。我们的研究结果表明,虽然某些EMG通道是单独更多的信息,最高的性能产生的子集,利用渠道之间的互补关系。我们还分析了通道消融下的音素分类精度,并观察到可解释的模式,反映了底层肌肉的解剖作用。为了解决通道减少导致的性能下降,我们使用随机通道丢弃对完整的8通道数据进行预训练,并在减少的通道子集上对其进行微调。对于4 - 6个通道设置,微调始终优于从头开始的训练,最佳退出策略取决于通道的数量。这些结果表明,从传感器减少的性能下降,可以减轻通过预训练和通道感知设计,支持轻量级和实用的基于EMG的无声语音系统的发展。
摘要:Surface electromyography (EMG) is a promising modality for silent speech interfaces, but its effectiveness depends heavily on sensor placement and channel availability. In this work, we investigate the contribution of individual and combined EMG channels to speech reconstruction performance. Our findings reveal that while certain EMG channels are individually more informative, the highest performance arises from subsets that leverage complementary relationships among channels. We also analyzed phoneme classification accuracy under channel ablations and observed interpretable patterns reflecting the anatomical roles of the underlying muscles. To address performance degradation from channel reduction, we pretrained models on full 8-channel data using random channel dropout and fine-tuned them on reduced-channel subsets. Fine-tuning consistently outperformed training from scratch for 4 - 6 channel settings, with the best dropout strategy depending on the number of channels. These results suggest that performance degradation from sensor reduction can be mitigated through pretraining and channel-aware design, supporting the development of lightweight and practical EMG-based silent speech systems.
【7】Misophonia Trigger Sound Detection on Synthetic Soundscapes Using a Hybrid Model with a Frozen Pre-Trained CNN and a Time-Series Module
标题:厌音症触发合成音景上的声音检测,使用具有冷冻预训练CNN和时间序列模块的混合模型
链接:https://arxiv.org/abs/2602.06271
备注:13 pages, 3 figures. Submitted to IJCNN 2026
摘要:恐音症是一种障碍,其特征是对特定的日常声音(触发声音)的耐受性降低,这些声音可以引起强烈的负面情绪反应,如愤怒,恐慌或焦虑。这些反应会严重损害日常功能和生活质量。选择性检测触发声音的辅助技术可以帮助减少痛苦并改善福祉。在这项研究中,我们调查的声音事件检测(SED),以本地化的时间间隔的触发声音在连续的环境音频作为基础的一步,这样的辅助支持。出于现实世界的恐音症数据的稀缺性,我们生成合成音景适合使用音频合成技术的恐音症触发声音检测。然后,我们使用基于混合CNN的模型执行触发声音检测任务。这些模型将使用冻结的预训练CNN骨干的特征提取与可训练的时间序列模块相结合,例如门控递归单元(GRU),长短期记忆(LSTM),回声状态网络(ESN)及其双向变体。检测性能使用常见的SED指标进行评估,包括复调声音检测分数1(PSDS 1)。在多类触发SED任务中,双向时间建模始终提高了检测性能,双向GRU(BiGRU)实现了最佳的整体准确性。值得注意的是,双向ESN(BiESN)通过仅优化读出而获得有竞争力的性能,同时需要数量级更少的可训练参数。我们进一步模拟用户个性化通过一个Few-Shot“吃的声音”检测任务,最多五个支持剪辑,其中BiGRU和BiESN进行比较。在这种严格的适应设置,BiESN显示出强大和稳定的性能,这表明轻量级的时间模块是有希望的个性化misophonia触发SED。
摘要:Misophonia is a disorder characterized by a decreased tolerance to specific everyday sounds (trigger sounds) that can evoke intense negative emotional responses such as anger, panic, or anxiety. These reactions can substantially impair daily functioning and quality of life. Assistive technologies that selectively detect trigger sounds could help reduce distress and improve well-being. In this study, we investigate sound event detection (SED) to localize intervals of trigger sounds in continuous environmental audio as a foundational step toward such assistive support. Motivated by the scarcity of real-world misophonia data, we generate synthetic soundscapes tailored to misophonia trigger sound detection using audio synthesis techniques. Then, we perform trigger sound detection tasks using hybrid CNN-based models. The models combine feature extraction using a frozen pre-trained CNN backbone with a trainable time-series module such as gated recurrent units (GRUs), long short-term memories (LSTMs), echo state networks (ESNs), and their bidirectional variants. The detection performance is evaluated using common SED metrics, including Polyphonic Sound Detection Score 1 (PSDS1). On the multi-class trigger SED task, bidirectional temporal modeling consistently improves detection performance, with Bidirectional GRU (BiGRU) achieving the best overall accuracy. Notably, the Bidirectional ESN (BiESN) attains competitive performance while requiring orders of magnitude fewer trainable parameters by optimizing only the readout. We further simulate user personalization via a few-shot "eating sound" detection task with at most five support clips, in which BiGRU and BiESN are compared. In this strict adaptation setting, BiESN shows robust and stable performance, suggesting that lightweight temporal modules are promising for personalized misophonia trigger SED.
【1】The Combination of Several Decorrelation Methods to Improve Acoustic Feedback Cancellation
标题:几种去相关方法的组合改善声反馈抵消
链接:https://arxiv.org/abs/2602.06921
摘要:本文扩展了声反馈抵消系统,结合多种去相关方法。基线系统是基于频域卡尔曼滤波器实现的多延迟结构。所提出的扩展包括可变时间延迟线、预测、失真补偿和简化的混响模型。每个扩展进行了分析,并定义了一个实用的参数范围。 虽然现有的文献往往集中在一个单一的扩展,如预测,来描述一个最佳的系统,这项工作表明,每个单独的扩展有助于性能的提高。此外,所有提议的扩展的组合会产生一个更优越的系统。使用公开的数据集进行评估,通过系统距离度量和客观语音质量度量PSEQ评估性能。
摘要:This paper extends an acoustic feedback cancellation system by incorporating multiple decorrelation methods. The baseline system is based on a frequency-domain Kalman filter implemented in a multi-delay structure. The proposed extensions include a variable time delay line, prediction, distortion compensation, and a simplified reverberation model. Each extension is analyzed, and a practical parameter range is defined. While existing literature often focuses on a single extension, such as prediction, to describe an optimal system, this work demonstrates that each individual extension contributes to performance improvements. Furthermore, the combination of all proposed extensions results in a superior system. The evaluation is conducted using publicly available datasets, with performance assessed through system distance metrics and the objective speech quality measure PSEQ.
【2】Automatic Detection and Analysis of Singing Mistakes for Music Pedagogy
标题:音乐教学学中演唱错误的自动检测与分析
链接:https://arxiv.org/abs/2602.06917
备注:Under Review at Transactions of Audio Speech and Language Processing
摘要:音频分析中机器学习的进步为技术增强的音乐教育开辟了新的可能性。本文介绍了一个框架,自动歌唱错误检测的背景下,音乐教学,支持一个新的策划数据集。该数据集包括同步的教师学习者声乐录音,并带有标记学习者所犯不同类型错误的注释。使用这个数据集,我们开发了不同的深度学习模型来进行错误检测并对其进行基准测试。为了比较错误检测系统的效率,提出了一种新的评估方法。实验结果表明,基于学习的方法优于基于规则的方法。错误的系统研究和跨教师的研究揭示了音乐教学法,可用于各种音乐应用的见解。这部作品为音乐教育学的研究开辟了新的方向。代码和数据集是公开的。
摘要:The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing mistake detection in the context of music pedagogy, supported by a newly curated dataset. The dataset comprises synchronized teacher learner vocal recordings, with annotations marking different types of mistakes made by learners. Using this dataset, we develop different deep learning models for mistake detection and benchmark them. To compare the efficacy of mistake detection systems, a new evaluation methodology is proposed. Experiments indicate that the proposed learning-based methods are superior to rule-based methods. A systematic study of errors and a cross-teacher study reveal insights into music pedagogy that can be utilised for various music applications. This work sets out new directions of research in music pedagogy. The codes and dataset are publicly available.
【3】B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
标题:B-GRPO:基于批量组相对策略优化的无监督语音情感识别
链接:https://arxiv.org/abs/2602.06290
备注:Accepted by ICASSP2026
摘要:无监督语音情感识别(SER)的研究重点是解决情感语音的数据稀疏性和标注偏差问题。强化学习(RL)是一种很有前途的方法,它通过基于规则或基于模型的验证功能而不是人工注释来增强性能。将学习过程中的样本选择看作一个长期过程,将是否选择样本看作决策行为,从而实现了RL在SER中对样本质量的度量。提出了一种改进的组相对策略优化(GRPO),使之适用于分类问题,其将一批中的样本作为一组,并使用这些样本的平均奖励作为基线来计算优势。我们提出了自我奖励函数和教师奖励函数,而不是像GRPO中那样使用可验证的奖励函数,以鼓励模型产生高信心的输出。实验结果表明,该方法提高了19.8%的基线没有RL的性能。
摘要:Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising method which enhances the performance through rule-based or model-based verification functions rather than human annotations. We treat the sample selection during the learning process as a long-term procedure and whether to select a sample as the action to make policy, thus achieving the application of RL to measure sample quality in SER. We propose a modified Group Relative Policy Optimization (GRPO) to adapt it to classification problems, which takes the samples in a batch as a group and uses the average reward of these samples as the baseline to calculate the advantage. And rather than using a verifiable reward function as in GRPO, we put forward self-reward functions and teacher-reward functions to encourage the model to produce high-confidence outputs. Experiments indicate that the proposed method improves the performance of baseline without RL by 19.8%.
【4】From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding
标题:从幻觉到清晰度:超低比特率神经语音编码的语言模型驱动的损失
链接:https://arxiv.org/abs/2602.06213
备注:To appear in ICASSP 2026. Demo wavs, code, and checkpoints (currently) availble at https://github.com/stet-stet/lmloss-icassp2026
摘要:“音素幻觉(Phoneme Hallucinations,PH)”通常发生在低比特率的基于DNN的编解码器中。这是生成解码器试图从丢失一些语义信息的过度压缩的令牌合成合理的输出。在这项工作中,我们提出了语言模型驱动的损失(LM损失),并表明他们可以减轻PH值比语义蒸馏(SD)的目标在非常低的比特率设置。拟议的LM损失建立在预先训练的语言模型上,以将语音与文本相关联。当地面实况成绩单不可用时,我们建议修改一个流行的自动语音识别(ASR)模型,耳语,比较解码的话语对ASR推断的输入语音的转录。否则,我们建议使用定时文本正则化器(TTR)比较WavLM表示的解码话语对BERT表示的地面实况transmittance。我们测试和比较LM损失对SD目标,使用参考编解码器的三阶段训练方案后,几个流行的编解码器设计。主观和客观的评估得出结论,LM损失可以提供更强的指导,从自我监督的语音表示中提取语义信息,提高人类感知的语义坚持,同时保持整体输出质量。演示示例、代码和检查点可在线获取。
摘要:``Phoneme Hallucinations (PH)'' commonly occur in low-bitrate DNN-based codecs. It is the generative decoder's attempt to synthesize plausible outputs from excessively compressed tokens missing some semantic information. In this work, we propose language model-driven losses (LM loss) and show they may alleviate PHs better than a semantic distillation (SD) objective in very-low-bitrate settings. The proposed LM losses build upon language models pretrained to associate speech with text. When ground-truth transcripts are unavailable, we propose to modify a popular automatic speech recognition (ASR) model, Whisper, to compare the decoded utterance against the ASR-inferred transcriptions of the input speech. Else, we propose to use the timed-text regularizer (TTR) to compare WavLM representations of the decoded utterance against BERT representations of the ground-truth transcriptions. We test and compare LM losses against an SD objective, using a reference codec whose three-stage training regimen was designed after several popular codecs. Subjective and objective evaluations conclude that LM losses may provide stronger guidance to extract semantic information from self-supervised speech representations, boosting human-perceived semantic adherence while preserving overall output quality. Demo samples, code, and checkpoints are available online.
【5】STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs
标题:STACodec:用于平衡音频编解码器中声学保真度和语义信息的语义令牌分配
链接:https://arxiv.org/abs/2602.06180
备注:ICASSP 2026
摘要:神经音频编解码器广泛用于音频压缩,并且可以集成到基于令牌的语言模型中。传统的编解码器很好地保留了声学细节,但缺乏语义信息。最近的混合编解码器试图通过蒸馏来合并语义信息,但这通常会降低重建性能,从而难以实现两者。为了解决这一限制,我们引入了STACodec,这是一种统一的编解码器,它通过语义令牌分配(STA)将来自自监督学习(SSL)模型的语义信息集成到残差矢量量化(RVQ-1)的第一层中。为了进一步消除对基于SSL的语义标记器的依赖并提高推理过程中的效率,我们提出了一个语义预蒸馏(SPD)模块,该模块直接预测语义标记,以便在推理过程中分配给第一个RVQ层。实验结果表明,STACodec优于现有的混合编解码器在音频重建和下游的语义任务,表现出更好的声学保真度和语义能力之间的平衡。
摘要:Neural audio codecs are widely used for audio compression and can be integrated into token-based language models. Traditional codecs preserve acoustic details well but lack semantic information. Recent hybrid codecs attempt to incorporate semantic information through distillation, but this often degrades reconstruction performance, making it difficult to achieve both. To address this limitation, we introduce STACodec, a unified codec that integrates semantic information from self-supervised learning (SSL) models into the first layer of residual vector quantization (RVQ-1) via semantic token assignment (STA). To further eliminate reliance on SSL-based semantic tokenizers and improve efficiency during inference, we propose a semantic pre-distillation (SPD) module, which predicts semantic tokens directly for assignment to the first RVQ layer during inference. Experimental results show that STACodec outperforms existing hybrid codecs in both audio reconstruction and downstream semantic tasks, demonstrating a better balance between acoustic fidelity and semantic capability.
【6】Reciprocal Latent Fields for Precomputed Sound Propagation
标题:预先计算的声音传播的相互潜场
链接:https://arxiv.org/abs/2602.06937
备注:Temporary pre-print, will be updated. In review at a conference
摘要:逼真的声音传播对于沉浸在虚拟场景中是必不可少的,但物理上精确的基于波的模拟对于实时应用仍然在计算上是禁止的。波编码方法通过预先计算给定场景的脉冲响应并将其压缩成一组标量声学参数来解决这一限制,这些标量声学参数在具有许多源-接收器对的大型环境中可能达到难以管理的大小。我们介绍了互易潜场(RLF),一个内存效率的框架编码和预测这些声学参数。RLF框架采用用对称函数解码的可训练潜在嵌入的体积网格,确保声学互易性。我们研究了各种解码器,并表明利用黎曼度量学习可以更好地再现复杂场景中的声学现象。实验验证表明,RLF保持复制质量,同时减少了几个数量级的内存占用。此外,类似MUSHRA的主观听力测试表明,通过RLF呈现的声音在感知上与地面真实模拟无法区分。
摘要:Realistic sound propagation is essential for immersion in a virtual scene, yet physically accurate wave-based simulations remain computationally prohibitive for real-time applications. Wave coding methods address this limitation by precomputing and compressing impulse responses of a given scene into a set of scalar acoustic parameters, which can reach unmanageable sizes in large environments with many source-receiver pairs. We introduce Reciprocal Latent Fields (RLF), a memory-efficient framework for encoding and predicting these acoustic parameters. The RLF framework employs a volumetric grid of trainable latent embeddings decoded with a symmetric function, ensuring acoustic reciprocity. We study a variety of decoders and show that leveraging Riemannian metric learning leads to a better reproduction of acoustic phenomena in complex scenes. Experimental validation demonstrates that RLF maintains replication quality while reducing the memory footprint by several orders of magnitude. Furthermore, a MUSHRA-like subjective listening test indicates that sound rendered via RLF is perceptually indistinguishable from ground-truth simulations.
【7】AI-Generated Music Detection in Broadcast Monitoring
标题:广播监控中的人工智能生成音乐检测
链接:https://arxiv.org/abs/2602.06823
摘要:人工智能音乐生成器已经发展到它们的输出通常与人类作品无法区分的地步。虽然检测方法已经出现,但它们通常是在具有干净的全长曲目的音乐流媒体环境中设计和验证的。然而,广播音频提出了一个不同的挑战:音乐以简短的摘录出现,通常被主导语音掩盖,现有检测器在这种情况下失败。在这项工作中,我们介绍了AI-OpenBMAT,这是第一个为广播式AI音乐检测量身定制的数据集。它包含3,294个一分钟的音频摘录(54.9小时),遵循真实电视音频的持续时间模式和响度关系,将人造制作音乐与Suno v3.5生成的风格匹配的延续相结合。我们对CNN基线和最先进的SpectTTTra模型进行基准测试,以评估SNR和持续时间的鲁棒性,并在完整的广播场景中进行评估。在所有设置中,在流媒体场景中表现出色的模型都会受到严重影响,当音乐处于背景或持续时间较短时,F1分数会下降到60%以下。这些结果强调了语音掩蔽和短音乐长度是AI音乐检测的关键挑战,并将AI-OpenBMAT定位为开发能够满足工业广播要求的检测器的基准。
摘要:AI music generators have advanced to the point where their outputs are often indistinguishable from human compositions. While detection methods have emerged, they are typically designed and validated in music streaming contexts with clean, full-length tracks. Broadcast audio, however, poses a different challenge: music appears as short excerpts, often masked by dominant speech, conditions under which existing detectors fail. In this work, we introduce AI-OpenBMAT, the first dataset tailored to broadcast-style AI-music detection. It contains 3,294 one-minute audio excerpts (54.9 hours) that follow the duration patterns and loudness relations of real television audio, combining human-made production music with stylistically matched continuations generated with Suno v3.5. We benchmark a CNN baseline and state-of-the-art SpectTTTra models to assess SNR and duration robustness, and evaluate on a full broadcast scenario. Across all settings, models that excel in streaming scenarios suffer substantial degradation, with F1-scores dropping below 60% when music is in the background or has a short duration. These results highlight speech masking and short music length as critical open challenges for AI music detection, and position AI-OpenBMAT as a benchmark for developing detectors capable of meeting industrial broadcast requirements.
【8】Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
标题:在波浪之间阅读:使用句子间音频特征的稳健主题分割
链接:https://arxiv.org/abs/2602.06647
备注:Accepted to IEEE ICASSP 2026
摘要:在线视频和播客等口语内容通常跨越多个主题,这使得自动主题分割对于用户导航和下游应用程序至关重要。然而,目前的方法没有充分利用声学特征,留下了改进的空间。我们提出了一种多模态的方法,微调文本编码器和暹罗音频编码器,捕捉句子边界周围的声学线索。在大规模YouTube视频数据集上进行的实验显示,与纯文本和多模式基线相比,性能有了显着提高。我们的模型还证明对ASR噪声更具弹性,并且在葡萄牙语,德语和英语的三个额外数据集上的表现优于更大的纯文本基线,强调了学习的声学特征对于鲁棒主题分割的价值。
摘要:Spoken content, such as online videos and podcasts, often spans multiple topics, which makes automatic topic segmentation essential for user navigation and downstream applications. However, current methods do not fully leverage acoustic features, leaving room for improvement. We propose a multi-modal approach that fine-tunes both a text encoder and a Siamese audio encoder, capturing acoustic cues around sentence boundaries. Experiments on a large-scale dataset of YouTube videos show substantial gains over text-only and multi-modal baselines. Our model also proves more resilient to ASR noise and outperforms a larger text-only baseline on three additional datasets in Portuguese, German, and English, underscoring the value of learned acoustic features for robust topic segmentation.
【9】Scaling Speech Tokenizers with Diffusion Autoencoders
标题:使用扩散自动编码器缩放语音令牌器
链接:https://arxiv.org/abs/2602.06602
备注:ICLR 2026
摘要:语音标记器是语音语言模型的基础,然而现有方法面临两个主要挑战:(1)平衡用于理解的编码语义与用于重建的声学之间的权衡,以及(2)实现低比特率和低标记率。我们提出了语音扩散标记器(SiTok),这是一种扩散自动编码器,通过监督学习联合学习语义丰富的表示,并通过扩散实现高保真音频重建。我们将SiTok扩展到1.6 B参数,并在200万小时的语音上对其进行训练。实验表明,SiTok在理解,重建和生成任务方面优于强大的基线,其令牌速率极低,为12.5 $ Hz,比特率为200比特每秒。
摘要:Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understanding and acoustics for reconstruction, and (2) achieving low bit rates and low token rates. We propose Speech Diffusion Tokenizer (SiTok), a diffusion autoencoder that jointly learns semantic-rich representations through supervised learning and enables high-fidelity audio reconstruction with diffusion. We scale SiTok to 1.6B parameters and train it on 2 million hours of speech. Experiments show that SiTok outperforms strong baselines on understanding, reconstruction and generation tasks, at an extremely low token rate of $12.5$ Hz and a bit-rate of 200 bits-per-second.
【10】Misophonia Trigger Sound Detection on Synthetic Soundscapes Using a Hybrid Model with a Frozen Pre-Trained CNN and a Time-Series Module
标题:厌音症触发合成音景上的声音检测,使用具有冷冻预训练CNN和时间序列模块的混合模型
链接:https://arxiv.org/abs/2602.06271
备注:13 pages, 3 figures. Submitted to IJCNN 2026
摘要:恐音症是一种障碍,其特征是对特定的日常声音(触发声音)的耐受性降低,这些声音可以引起强烈的负面情绪反应,如愤怒,恐慌或焦虑。这些反应会严重损害日常功能和生活质量。选择性检测触发声音的辅助技术可以帮助减少痛苦并改善福祉。在这项研究中,我们调查的声音事件检测(SED),以本地化的时间间隔的触发声音在连续的环境音频作为基础的一步,这样的辅助支持。出于现实世界的恐音症数据的稀缺性,我们生成合成音景适合使用音频合成技术的恐音症触发声音检测。然后,我们使用基于混合CNN的模型执行触发声音检测任务。这些模型将使用冻结的预训练CNN骨干的特征提取与可训练的时间序列模块相结合,例如门控递归单元(GRU),长短期记忆(LSTM),回声状态网络(ESN)及其双向变体。检测性能使用常见的SED指标进行评估,包括复调声音检测分数1(PSDS 1)。在多类触发SED任务中,双向时间建模始终提高了检测性能,双向GRU(BiGRU)实现了最佳的整体准确性。值得注意的是,双向ESN(BiESN)通过仅优化读出而获得有竞争力的性能,同时需要数量级更少的可训练参数。我们进一步模拟用户个性化通过一个Few-Shot“吃的声音”检测任务,最多五个支持剪辑,其中BiGRU和BiESN进行比较。在这种严格的适应设置,BiESN显示出强大和稳定的性能,这表明轻量级的时间模块是有希望的个性化misophonia触发SED。
摘要:Misophonia is a disorder characterized by a decreased tolerance to specific everyday sounds (trigger sounds) that can evoke intense negative emotional responses such as anger, panic, or anxiety. These reactions can substantially impair daily functioning and quality of life. Assistive technologies that selectively detect trigger sounds could help reduce distress and improve well-being. In this study, we investigate sound event detection (SED) to localize intervals of trigger sounds in continuous environmental audio as a foundational step toward such assistive support. Motivated by the scarcity of real-world misophonia data, we generate synthetic soundscapes tailored to misophonia trigger sound detection using audio synthesis techniques. Then, we perform trigger sound detection tasks using hybrid CNN-based models. The models combine feature extraction using a frozen pre-trained CNN backbone with a trainable time-series module such as gated recurrent units (GRUs), long short-term memories (LSTMs), echo state networks (ESNs), and their bidirectional variants. The detection performance is evaluated using common SED metrics, including Polyphonic Sound Detection Score 1 (PSDS1). On the multi-class trigger SED task, bidirectional temporal modeling consistently improves detection performance, with Bidirectional GRU (BiGRU) achieving the best overall accuracy. Notably, the Bidirectional ESN (BiESN) attains competitive performance while requiring orders of magnitude fewer trainable parameters by optimizing only the readout. We further simulate user personalization via a few-shot "eating sound" detection task with at most five support clips, in which BiGRU and BiESN are compared. In this strict adaptation setting, BiESN shows robust and stable performance, suggesting that lightweight temporal modules are promising for personalized misophonia trigger SED.
机器翻译由腾讯交互翻译提供,仅供参考
