微信公众号:arXiv_Daily
cs.SD语音
【1】Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
标题:伪倒谱:基于Mel的神经声码器的音调修改
链接:https://arxiv.org/pdf/2512.16519v1
摘要:本文介绍了一种基于倒谱的基音修正方法,该方法可以应用于任何梅尔频谱表示。因此,该方法与任何基于mel的声码器兼容,而不需要对模型进行任何额外的训练或改变。这是通过直接修改倒谱特征空间以将谐波结构移位到期望目标来实现的。通过伪逆梅尔变换计算频谱图幅度,然后通过应用DCT转换为倒谱。在这个域中,倒谱峰值被移位而不必估计其位置,并且通过应用IDCT和梅尔滤波器组来重新计算修改后的梅尔。这些音高移位的梅尔频谱图特征可以用任何兼容的声码器转换成语音。所提出的方法进行了实验验证与客观和主观的各种国家的最先进的神经声码器的指标,以及与传统的音调修改方法的比较。摘要:This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any additional training or changes to the model. This is achieved by directly modifying the cepstrum feature space in order to shift the harmonic structure to the desired target. The spectrogram magnitude is computed via the pseudo-inverse mel transform, then converted to the cepstrum by applying DCT. In this domain, the cepstral peak is shifted without having to estimate its position and the modified mel is recomputed by applying IDCT and mel-filterbank. These pitch-shifted mel-spectrogram features can be converted to speech with any compatible vocoder. The proposed method is validated experimentally with objective and subjective metrics on various state-of-the-art neural vocoders as well as in comparison with traditional pitch modification methods.
【2】DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
标题:DPDFNet:通过双路径RNN提升DeepLayer Net 2
链接:https://arxiv.org/pdf/2512.16420v1
备注:我们提出了DPDFNet,这是一种因果单通道语音增强模型,它在编码器中使用双路径块扩展了DeepFilterNet 2架构,在保留原始增强框架的同时加强了长距离时间和跨频带建模。此外,我们证明,增加一个损失的组成部分,以减轻过度衰减的增强语音,结合微调阶段量身定制的“永远在线”的应用程序,导致整体模型性能的大幅改善。为了将我们提出的架构与各种因果开源模型进行比较,我们创建了一个新的评估集,其中包括12种语言的长时间低SNR记录,这些记录跨越了日常噪声场景,比常用的基准测试更好地反映了现实世界的条件。在这个评估集上,DPDFNet提供了优于其他因果开源模型的性能,包括一些更大,计算要求更高的模型。我们还提出了一个整体的度量名为PRISM,复合,规模归一化聚合的侵入性和非侵入性的指标,这表明了明确的可扩展性与双路径块的数量。我们通过在Ceva-NeuPro-Nano边缘NPU上部署DPDFNet进一步证明了设备上的可行性。结果表明,我们的第二大模型DPDFNet-4在NPN 32上实现了实时性能,在NPN 64上运行速度更快,证实了在严格的嵌入式功耗和延迟限制下可以保持最先进的质量。摘要:We present DPDFNet, a causal single-channel speech enhancement model that extends DeepFilterNet2 architecture with dual-path blocks in the encoder, strengthening long-range temporal and cross-band modeling while preserving the original enhancement framework. In addition, we demonstrate that adding a loss component to mitigate over-attenuation in the enhanced speech, combined with a fine-tuning phase tailored for "always-on" applications, leads to substantial improvements in overall model performance. To compare our proposed architecture with a variety of causal open-source models, we created a new evaluation set comprising long, low-SNR recordings in 12 languages across everyday noise scenarios, better reflecting real-world conditions than commonly used benchmarks. On this evaluation set, DPDFNet delivers superior performance to other causal open-source models, including some that are substantially larger and more computationally demanding. We also propose an holistic metric named PRISM, a composite, scale-normalized aggregate of intrusive and non-intrusive metrics, which demonstrates clear scalability with the number of dual-path blocks. We further demonstrate on-device feasibility by deploying DPDFNet on Ceva-NeuPro-Nano edge NPUs. Results indicate that DPDFNet-4, our second-largest model, achieves real-time performance on NPN32 and runs even faster on NPN64, confirming that state-of-the-art quality can be sustained within strict embedded power and latency constraints.
【3】Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
标题:听着翻译:言语情态融入法学硕士的有效性
链接:https://arxiv.org/pdf/2512.16378v1
备注:Project available at https:github.comsarapapihearing2translate
摘要:随着大型语言模型(LLM)扩展到文本之外,将语音作为一种原生模态进行集成,从而产生了SpeechLLM,其目的是直接翻译口语,从而绕过传统的基于转录的管道。然而,这种集成是否比现有的级联架构提高了语音到文本的翻译质量,仍然是一个悬而未决的问题。我们提出了听力翻译,这是第一个全面的测试套件,对5个最先进的SpeechLLM进行了严格的基准测试,针对16个强大的直接和级联系统,这些系统将领先的语音基础模型(SFM)与多语言LLM相结合。我们的分析涵盖了16个基准测试,13个语言对和9个具有挑战性的条件,包括不流利,嘈杂和长形式的语音。在这次广泛的评估中,我们发现级联系统总体上仍然是最可靠的,而当前的SpeechLLM仅在选定的设置中匹配级联,而SFM落后于两者,突出表明在模型或管道中集成LLM对于高质量的语音翻译至关重要。摘要:As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which aim to translate spoken language directly, thereby bypassing traditional transcription-based pipelines. Whether this integration improves speech-to-text translation quality over established cascaded architectures, however, remains an open question. We present Hearing to Translate, the first comprehensive test suite rigorously benchmarking 5 state-of-the-art SpeechLLMs against 16 strong direct and cascade systems that couple leading speech foundation models (SFM), with multilingual LLMs. Our analysis spans 16 benchmarks, 13 language pairs, and 9 challenging conditions, including disfluent, noisy, and long-form speech. Across this extensive evaluation, we find that cascaded systems remain the most reliable overall, while current SpeechLLMs only match cascades in selected settings and SFMs lag behind both, highlighting that integrating an LLM, either within the model or in a pipeline, is essential for high-quality speech translation.【4】CogSR: Semantic-Aware Speech Super-Resolution via Chain-of-Thought Guided Flow Matching
标题:CogSR:通过思想链引导流匹配实现语义感知语音超分辨率
链接:https://arxiv.org/pdf/2512.16304v1
备注:7 pages
摘要:将语音超分辨率(SR)应用于采样率极低的录音是数字存档和调查音频恢复的关键挑战。在这些场景中,输入缺乏必要的声学线索。因此,现有的生成模型经常失败;没有足够的上下文,它们会产生语音内容的幻觉,根据概率而不是意义来猜测单词。 为了解决这个问题,我们提出了CogSR,这是一个专门为高精度离线恢复而设计的框架。我们的方法将重点从简单的信号映射转移到认知重建。通过整合一个大的音频语言模型,我们采用了思想链推理作为语义锚,而明确的声学先验确保扬声器的身份保持一致。这将指导整流主干合成高频细节,这些细节不仅真实,而且在语言上准确。评估表明,CogSR有效地消除了严重退化制度中的模糊性,使其成为恢复高价值遗产和监控音频的强大解决方案。摘要:Applying speech super-resolution (SR) to recordings with severely low sampling rates is a critical challenge in digital archiving and investigative audio recovery. In these scenarios, the input lacks essential acoustic cues. Consequently, existing generative models often fail; without sufficient context, they hallucinate phonetic content, guessing words based on probability rather than meaning. To address this, we propose CogSR, a framework designed specifically for high-precision, offline restoration. Our approach shifts the focus from simple signal mapping to cognitive reconstruction. By integrating a Large Audio-Language Model, we employ Chain-of-Thought reasoning to act as a semantic anchor, while explicit acoustic priors ensure the speaker's identity remains consistent. This guides a Rectified Flow backbone to synthesize high-frequency details that are not only realistic but linguistically accurate. Evaluations show that CogSR effectively eliminates ambiguity in severe degradation regimes, making it a robust solution for restoring high-value legacy and surveillance audio.
【5】Domain-Agnostic Causal-Aware Audio Transformer for Infant Cry Classification
标题:用于婴儿哭声分类的领域不可知的CASEARCH感知音频Transformer
链接:https://arxiv.org/pdf/2512.16271v1
备注:This paper has been published in the IEEE proceedings of the 8th International Conference of Computer and Informatics Engineering (IC2IE)
摘要:婴儿哭声语言学的准确和可解释的分类对于早期发现新生儿窘迫和临床决策支持至关重要。然而,许多现有的深度学习方法依赖于相关性驱动的声学表示,这使得它们容易受到噪声,虚假线索和记录环境中的域偏移的影响。我们提出了DACH-TIC,一个领域不可知的婴儿感知分层音频Transformer强大的婴儿哭声分类。该模型在一个统一的框架内集成了因果注意、层次表征学习、多任务监督和对抗域泛化。 DACH-TIC采用了一个结构化的Transformer骨干,具有本地令牌级和全局语义编码器,通过因果注意掩蔽和受控扰动训练来增强,以近似反事实的声学变化。领域对抗目标促进了环境不变的表示,而多任务学习则共同优化了哭泣类型识别、痛苦强度估计和因果相关性预测。该模型在Baby Chillanto和Donate-a-Cry数据集上进行评估,并使用ESC-50环境噪声叠加进行域增强。 实验结果表明,DACH-TIC的性能优于最先进的基线,包括HTS-AT和SE-ResNet Transformer,准确率提高了2.6%,宏观F1分数提高了2.2分,同时增强了因果保真度。该模型有效地推广到看不见的声学环境,域性能差距仅为2.4%,表明其适用于现实世界的新生儿声学监测系统。摘要:Accurate and interpretable classification of infant cry paralinguistics is essential for early detection of neonatal distress and clinical decision support. However, many existing deep learning methods rely on correlation-driven acoustic representations, which makes them vulnerable to noise, spurious cues, and domain shifts across recording environments. We propose DACH-TIC, a Domain-Agnostic Causal-Aware Hierarchical Audio Transformer for robust infant cry classification. The model integrates causal attention, hierarchical representation learning, multi-task supervision, and adversarial domain generalization within a unified framework. DACH-TIC employs a structured transformer backbone with local token-level and global semantic encoders, augmented by causal attention masking and controlled perturbation training to approximate counterfactual acoustic variations. A domain-adversarial objective promotes environment-invariant representations, while multi-task learning jointly optimizes cry type recognition, distress intensity estimation, and causal relevance prediction. The model is evaluated on the Baby Chillanto and Donate-a-Cry datasets, with ESC-50 environmental noise overlays for domain augmentation. Experimental results show that DACH-TIC outperforms state-of-the-art baselines, including HTS-AT and SE-ResNet Transformer, achieving improvements of 2.6 percent in accuracy and 2.2 points in macro-F1 score, alongside enhanced causal fidelity. The model generalizes effectively to unseen acoustic environments, with a domain performance gap of only 2.4 percent, demonstrating its suitability for real-world neonatal acoustic monitoring systems.【6】From Minutes to Days: Scaling Intracranial Speech Decoding with Supervised Pretraining
标题:从几分钟到几天:通过有监督的预训练来扩展脑内言语解码
链接:https://arxiv.org/pdf/2512.15830v1
备注:Linnea Evanson and Mingfang (Lucy) Zhang are joint first authors. Pierre Bourdillon and Jean-Rémi King are joint last authors
摘要:从大脑活动中解码语音通常依赖于在短期和高度受控的实验中收集的有限的神经记录。在这里,我们引入了一个框架,利用来自接受临床监测的患者的为期一周的颅内和音频记录,有效地将训练数据集的大小增加了两个数量级以上。通过这种预训练,我们的对比学习模型大大优于仅在经典实验数据上训练的模型,其增益与数据集大小呈对数线性关系。对学习表征的分析表明,虽然大脑活动代表了语音特征,但其全局结构在很大程度上会在不同的日子里漂移,这突出了对明确说明跨天变化的模型的需求。总的来说,我们的方法开辟了一条可扩展的道路,在现实生活和受控任务设置中解码和建模大脑表征。摘要:Decoding speech from brain activity has typically relied on limited neural recordings collected during short and highly controlled experiments. Here, we introduce a framework to leverage week-long intracranial and audio recordings from patients undergoing clinical monitoring, effectively increasing the training dataset size by over two orders of magnitude. With this pretraining, our contrastive learning model substantially outperforms models trained solely on classic experimental data, with gains that scale log-linearly with dataset size. Analysis of the learned representations reveals that, while brain activity represents speech features, its global structure largely drifts across days, highlighting the need for models that explicitly account for cross-day variability. Overall, our approach opens a scalable path toward decoding and modeling brain representations in both real-life and controlled task settings.
【1】BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
标题:BEST-STD 2.0:用于口语检测的平衡高效语音令牌器
链接:https://arxiv.org/pdf/2512.16395v1
备注:Submitted to ICASSP 2026
摘要:快速、准确的语音内容检索对于语音搜索等应用至关重要。按实例查询口语术语检测(STD)涉及在给定口语查询的情况下从音频数据库检索匹配片段。使用离散语音表示的基于令牌的STD系统使得能够进行有效的搜索,但是在对噪声和混响的鲁棒性以及低效的令牌利用方面存在困难。我们通过提出一种噪声和混响增强的训练策略来解决这些挑战,以提高标记器的鲁棒性。此外,我们引入了基于传输的最优正则化,以确保均衡的令牌使用并提高令牌效率。为了进一步加快检索速度,我们采用了基于TF-IDF的搜索机制。实证评估表明,该方法优于STD基线在各种失真水平,同时保持高的搜索效率。摘要:Fast and accurate spoken content retrieval is vital for applications such as voice search. Query-by-Example Spoken Term Detection (STD) involves retrieving matching segments from an audio database given a spoken query. Token-based STD systems, which use discrete speech representations, enable efficient search but struggle with robustness to noise and reverberation, and with inefficient token utilization. We address these challenges by proposing a noise and reverberation-augmented training strategy to improve tokenizer robustness. In addition, we introduce optimal transport-based regularization to ensure balanced token usage and enhance token efficiency. To further speed up retrieval, we adopt a TF-IDF-based search mechanism. Empirical evaluations demonstrate that the proposed method outperforms STD baselines across various distortion levels while maintaining high search efficiency.
【2】Learning Recursive Attenuation Filters Under Noisy Conditions
标题:在噪音条件下学习循环衰减过滤器
链接:https://arxiv.org/pdf/2512.16318v1
备注:Submitted to the Journal of Audio Engineering Society
摘要:递归是滤波器和音频系统设计中的基本概念。特别地,使用延迟网络的人工混响系统依赖于递归路径来控制回声密度和模态分量的衰减速率。可微分数字信号处理框架已经显示出在给定目标房间脉冲响应的情况下自动调谐递归和非递归元件两者的前景。这是通过基于能量衰减或频谱图差异将梯度下降应用于损失函数来完成的。然而,这些表示对背景噪声高度敏感,背景噪声在实际测量中无处不在,产生伪损耗最小值并导致不正确的衰减。本文讨论了目标噪声时反馈延迟网络递归衰减滤波器的调谐问题。我们研究了与不同优化目标相关的损失景观,并提出了一种方法,确保在低信噪比条件下正确的最小值。我们证明了所提出的方法的有效性,通过对80个单独的优化实例的统计分析。结果表明,明确建模的噪声恢复正确的最小值。此外,我们确定的灵敏度衰减滤波器的参数调谐到扰动的频率无关的参数。这些研究结果提供了更强大的和可重复的基于梯度的反馈延迟网络的优化实用指南。摘要:Recursion is a fundamental concept in the design of filters and audio systems. In particular, artificial reverberation systems that use delay networks depend on recursive paths to control both echo density and the decay rate of modal components. The differentiable digital signal processing framework has shown promise in automatically tuning both recursive and non-recursive elements given a target room impulse response. This is done by applying gradient descent to loss functions based on energy-decay or spectrogram differences. However, these representations are highly sensitive to background noise, which is ubiquitous in real measurements, producing spurious loss minima and leading to incorrect attenuation. This paper addresses the problem of tuning recursive attenuation filters of a feedback delay network when targets are noisy. We examine the loss landscape associated with different optimization objectives and propose a method that ensures correct minima under low signal-to-noise conditions. We demonstrate the effectiveness of the proposed approach through statistical analysis on 80 individual optimization examples. The results reveal that explicitly modeling the noise restores correct minima. Furthermore, we identify the sensitivity of attenuation filter parameters tuning to perturbations in frequency-independent parameters. These findings provide practical guidelines for more robust and reproducible gradient-based optimization of feedback delay networks.【3】Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
标题:伪倒谱:基于Mel的神经声码器的音调修改
链接:https://arxiv.org/pdf/2512.16519v1
摘要:本文介绍了一种基于倒谱的基音修正方法,该方法可以应用于任何梅尔频谱表示。因此,该方法与任何基于mel的声码器兼容,而不需要对模型进行任何额外的训练或改变。这是通过直接修改倒谱特征空间以将谐波结构移位到期望目标来实现的。通过伪逆梅尔变换计算频谱图幅度,然后通过应用DCT转换为倒谱。在这个域中,倒谱峰值被移位而不必估计其位置,并且通过应用IDCT和梅尔滤波器组来重新计算修改后的梅尔。这些音高移位的梅尔频谱图特征可以用任何兼容的声码器转换成语音。所提出的方法进行了实验验证与客观和主观的各种国家的最先进的神经声码器的指标,以及与传统的音调修改方法的比较。摘要:This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any additional training or changes to the model. This is achieved by directly modifying the cepstrum feature space in order to shift the harmonic structure to the desired target. The spectrogram magnitude is computed via the pseudo-inverse mel transform, then converted to the cepstrum by applying DCT. In this domain, the cepstral peak is shifted without having to estimate its position and the modified mel is recomputed by applying IDCT and mel-filterbank. These pitch-shifted mel-spectrogram features can be converted to speech with any compatible vocoder. The proposed method is validated experimentally with objective and subjective metrics on various state-of-the-art neural vocoders as well as in comparison with traditional pitch modification methods.
【4】Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
标题:海报:使用多任务学习识别隐藏在耳中的私有密钥以实现可靠的无声语音接口
链接:https://arxiv.org/pdf/2512.16518v1
备注:UbiComp Poster 2025
摘要:无声语音接口(SSI)允许免提输入,而不会发出声音,但大多数SSI系统不会验证说话者身份。我们介绍了HEar-ID,它使用消费者主动降噪耳塞来捕获低频“耳语”音频和高频超声波反射。来自两个流的特征通过一个共享的编码器,产生嵌入,该嵌入提供用于用户身份验证的对比分支和用于无声拼写识别的SSI头。这种设计支持50个单词的解码,同时可靠地拒绝冒名顶替者,所有这些都在单一型号的商品耳塞上。实验结果表明,HEar-ID具有较强的拼写准确性和鲁棒性。摘要:Silent speech interface (SSI) enables hands-free input without audible vocalization, but most SSI systems do not verify speaker identity. We present HEar-ID, which uses consumer active noise-canceling earbuds to capture low-frequency "whisper" audio and high-frequency ultrasonic reflections. Features from both streams pass through a shared encoder, producing embeddings that feed a contrastive branch for user authentication and an SSI head for silent spelling recognition. This design supports decoding of 50 words while reliably rejecting impostors, all on commodity earbuds with a single model. Experiments demonstrate that HEar-ID achieves strong spelling accuracy and robust authentication.
机器翻译由腾讯交互翻译提供,仅供参考
