今日论文合集:cs.SD语音6篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Next Tokens Denoising for Speech Synthesis
标题:下一个语音合成的令牌去噪
链接:https://arxiv.org/abs/2507.22746

作者:Yanqing Liu, Ruiqing Xue, Chong Zhang, Yufei Liu, Gang Wang, Bohan Li, Yao Qian, Lei He, Shujie Liu, Sheng Zhao
摘要:虽然扩散和自回归(AR)模型具有显着先进的生成建模,但它们各自存在明显的局限性。依赖于因果注意的AR模型无法利用未来的上下文,并且生成速度缓慢。相反,扩散模型则要处理键值(KV)缓存。为了克服这些挑战,我们引入了Dragon-FM,这是一种新颖的文本到语音(TTS)设计,它将AR和流匹配相结合。该模型以每秒12.5个令牌的紧凑速率以块的形式处理48 kHz音频编解码器令牌。这种设计使AR建模跨块,确保全局一致性,而块内的并行流匹配有助于快速迭代去噪。因此,所提出的模型可以跨块利用KV缓存,并将未来的上下文合并到每个块中。此外,它还连接了连续和离散特征建模,证明连续AR流匹配可以用有限的标量量化器预测离散令牌。这种高效的编解码器和快速块自回归架构也使得所提出的模型对于生成扩展内容特别有效。在播客数据集上进行的演示实验表明,该方法能够有效地生成高质量的zero-shot播客。
摘要:While diffusion and autoregressive (AR) models have significantly advanced generative modeling, they each present distinct limitations. AR models, which rely on causal attention, cannot exploit future context and suffer from slow generation speeds. Conversely, diffusion models struggle with key-value (KV) caching. To overcome these challenges, we introduce Dragon-FM, a novel text-to-speech (TTS) design that unifies AR and flow-matching. This model processes 48 kHz audio codec tokens in chunks at a compact 12.5 tokens per second rate. This design enables AR modeling across chunks, ensuring global coherence, while parallel flow-matching within chunks facilitates fast iterative denoising. Consequently, the proposed model can utilize KV-cache across chunks and incorporate future context within each chunk. Furthermore, it bridges continuous and discrete feature modeling, demonstrating that continuous AR flow-matching can predict discrete tokens with finite scalar quantizers. This efficient codec and fast chunk-autoregressive architecture also makes the proposed model particularly effective for generating extended content. Experiment for demos of our work} on podcast datasets demonstrate its capability to efficiently generate high-quality zero-shot podcasts.


【2】Adaptive Duration Model for Text Speech Alignment
标题:文本语音对齐的自适应持续时间模型
链接:https://arxiv.org/abs/2507.22612

作者:Junjie Cao
备注:4 pages, 3 figures, 2 tables
摘要:语音到文本对齐是神经文本到语音(TTS)模型的关键组成部分。自回归TTS模型通常使用注意力机制来在线学习这些对齐。然而,这些对齐往往是脆弱的,并且经常不能推广到长话语和域外文本,导致单词丢失或重复。大多数非自回归端到端TTS模型依赖于从外部源提取的持续时间,使用额外的持续时间模型进行对齐。在本文中,我们提出了一个新的持续时间预测框架,可以给妥协音素级的持续时间分布与给定的文本。在我们的实验中,提出的持续时间模型具有更精确的预测和条件适应能力相比,以前的基线模型。从数值上看,它在对齐精度上有大约11.3个百分点的提高,并且使zero-shot TTS模型的性能对提示音频和输入音频之间的失配更鲁棒。
摘要:Speech-to-text alignment is a critical component of neural text to-speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line. However, these alignments tend to be brittle and often fail to generalize to long utterances and out-of-domain text, leading to missing or repeating words. Most non-autoregressive end to-end TTS models rely on durations extracted from external sources, using additional duration models for alignment. In this paper, we propose a novel duration prediction framework that can give compromising phoneme-level duration distribution with given text. In our experiments, the proposed duration model has more precise prediction and condition adaptation ability compared to previous baseline models. Numerically, it has roughly a 11.3 percents immprovement on alignment accuracy, and makes the performance of zero-shot TTS models more robust to the mismatch between prompt audio and input audio.


【3】Prediction of acoustic field in 1-D uniform duct with varying mean flow and temperature using neural networks
标题:用神经网络预测平均流量和温度变化的一维均匀管道中的声学场
链接:https://arxiv.org/abs/2507.22370

作者:Veerababu Dharanalakota, Prasanta K. Ghosh
备注:22 pages
摘要:受物理定律约束的神经网络作为另一种数值工具出现。本文推导了一维管道中声传播的控制方程。将问题转化为无约束优化问题,并使用神经网络进行求解。采用传统的龙格-库塔法对声状态变量声压和质点速度进行了预测和验证。研究了温度梯度对声场的影响。利用机器学习技术,如迁移学习和自动微分声学应用程序的演示。
摘要:Neural networks constrained by the physical laws emerged as an alternate numerical tool. In this paper, the governing equation that represents the propagation of sound inside a one-dimensional duct carrying a heterogeneous medium is derived. The problem is converted into an unconstrained optimization problem and solved using neural networks. Both the acoustic state variables: acoustic pressure and particle velocity are predicted and validated with the traditional Runge-Kutta solver. The effect of the temperature gradient on the acoustic field is studied. Utilization of machine learning techniques such as transfer learning and automatic differentiation for acoustic applications is demonstrated.


【4】A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection
标题:用于增强声音事件定位和检测的两步学习框架
链接:https://arxiv.org/abs/2507.22322

作者:Hogeon Yu
备注:5pages, 2figures
摘要:声音事件定位和检测(SELD)在空间音频处理中至关重要,使系统能够检测声音事件并估计其3D方向。现有的SELD方法使用单分支或双分支架构:单分支模型共享SED和DoA表示,导致优化冲突,而双分支模型分离任务但限制信息交换。为了解决这个问题,我们提出了一个两步学习框架。首先,我们引入了一个tracwise重新排序格式,以保持时间的一致性,防止跨轨道的事件重新排序。接下来,我们训练SED和DoA网络以防止干扰并确保特定于任务的特征学习。最后,我们有效地融合DoA和SED功能,以提高SELD性能,更好的空间和事件表示。在2023年DCASE挑战任务3数据集上的实验验证了我们的框架,显示了它克服单分支和双分支限制并改进事件分类和本地化的能力。
摘要:Sound Event Localization and Detection (SELD) is crucial in spatial audio processing, enabling systems to detect sound events and estimate their 3D directions. Existing SELD methods use single- or dual-branch architectures: single-branch models share SED and DoA representations, causing optimization conflicts, while dual-branch models separate tasks but limit information exchange. To address this, we propose a two-step learning framework. First, we introduce a tracwise reordering format to maintain temporal consistency, preventing event reassignments across tracks. Next, we train SED and DoA networks to prevent interference and ensure task-specific feature learning. Finally, we effectively fuse DoA and SED features to enhance SELD performance with better spatial and event representation. Experiments on the 2023 DCASE challenge Task 3 dataset validate our framework, showing its ability to overcome single- and dual-branch limitations and improve event classification and localization.


【5】Quantum-Inspired Audio Unlearning: Towards Privacy-Preserving Voice Biometrics
标题:量子启发的音频遗忘:迈向保护隐私的语音生物识别技术
链接:https://arxiv.org/abs/2507.22208

作者:Shreyansh Pathak, Sonu Shreshtha, Richa Singh, Mayank Vatsa
备注:9 pages, 2 figures, 5 tables, Accepted at IJCB 2025 (Osaka, Japan)
摘要:语音认证和音频生物识别系统的广泛采用大大增加了与敏感语音数据相关的隐私漏洞。要遵守GDPR的被遗忘权和印度的DPDP法案等隐私法规,就必须从已经训练好的生物识别模型中有针对性地有效删除个人特定的语音签名。现有的为视觉数据设计的去学习方法不能充分处理音频信号的顺序、时间和高维性质,导致无效或不完整的说话者和口音擦除。为了解决这个问题,我们引入了QPAudioEraser,这是一个量子启发的音频遗忘框架。我们的our-phase方法包括:(1)使用相消干涉来取消目标特征的权重初始化,(2)基于叠加的标签转换,掩盖类身份,(3)不确定性最大化量子损失函数,以及(4)纠缠启发的相关权重混合,以保留模型知识。ResNet 18、ViT和CNN架构在AudioMNIST、Speech Commands、LibriSpeech和Speech Accent Archive数据集上的综合评估验证了QPAudioEraser的卓越性能。该框架实现了目标数据的完全擦除(0%遗忘精度),同时对模型实用性的影响最小,保留数据的性能下降低至0.05%。QPAudioEraser在单类、多类、顺序和重音级别擦除场景中始终超越传统基线,将所提出的方法确立为一种强大的隐私保护解决方案。
摘要:The widespread adoption of voice-enabled authentication and audio biometric systems have significantly increased privacy vulnerabilities associated with sensitive speech data. Compliance with privacy regulations such as GDPR's right to be forgotten and India's DPDP Act necessitates targeted and efficient erasure of individual-specific voice signatures from already-trained biometric models. Existing unlearning methods designed for visual data inadequately handle the sequential, temporal, and high-dimensional nature of audio signals, leading to ineffective or incomplete speaker and accent erasure. To address this, we introduce QPAudioEraser, a quantum-inspired audio unlearning framework. Our our-phase approach involves: (1) weight initialization using destructive interference to nullify target features, (2) superposition-based label transformations that obscure class identity, (3) an uncertainty-maximizing quantum loss function, and (4) entanglement-inspired mixing of correlated weights to retain model knowledge. Comprehensive evaluations with ResNet18, ViT, and CNN architectures across AudioMNIST, Speech Commands, LibriSpeech, and Speech Accent Archive datasets validate QPAudioEraser's superior performance. The framework achieves complete erasure of target data (0% Forget Accuracy) while incurring minimal impact on model utility, with a performance degradation on retained data as low as 0.05%. QPAudioEraser consistently surpasses conventional baselines across single-class, multi-class, sequential, and accent-level erasure scenarios, establishing the proposed approach as a robust privacy-preserving solution.


【6】A k-space approach to modeling multi-channel parametric array loudspeaker systems
标题:多通道参数阵列扬声器系统建模的k空间方法
链接:https://arxiv.org/abs/2507.22628

作者:Tao Zhuang, Longbiao He, Feng Niu, Jia-Xin Zhong, Jing Lu
摘要:多通道参量阵列扬声器(MCPAL)系统提供了增强的灵活性,并有望在现实世界的应用中产生高度定向的音频波束。然而,由于涉及复杂的非线性行为和多通道信号处理,有效和准确地预测其产生的声场仍然是一个重大挑战。为了克服这个障碍,我们提出了一个k空间的方法来建模任意MCPAL系统安排在一个有障碍的平面表面。在我们的方法中,首先使用角谱方法求解线性超声场,随后在k空间中有效地计算准线性音频声场。通过利用三维快速傅立叶变换,我们的方法不仅实现了高的计算和存储效率,但也保持精度,而不依赖于傍轴近似。对于典型的配置研究,所提出的方法表现出超过四个数量级的速度相比,直接集成方法。我们提出的方法为模拟和设计先进的MCPAL系统铺平了道路。
摘要:Multi-channel parametric array loudspeaker (MCPAL) systems offer enhanced flexibility and promise for generating highly directional audio beams in real-world applications. However, efficient and accurate prediction of their generated sound fields remains a major challenge due to the complex nonlinear behavior and multi-channel signal processing involved. To overcome this obstacle, we propose a k-space approach for modeling arbitrary MCPAL systems arranged on a baffled planar surface. In our method, the linear ultrasound field is first solved using the angular spectrum approach, and the quasilinear audio sound field is subsequently computed efficiently in k-space. By leveraging three-dimensional fast Fourier transforms, our approach not only achieves high computational and memory efficiency but also maintains accuracy without relying on the paraxial approximation. For typical configurations studied, the proposed method demonstrates a speed-up of more than four orders of magnitude compared to the direct integration method. Our proposed approach paved the way for simulating and designing advanced MCPAL systems.


eess.AS音频处理


【1】A k-space approach to modeling multi-channel parametric array loudspeaker systems
标题:多通道参数阵列扬声器系统建模的k空间方法
链接:https://arxiv.org/abs/2507.22628

作者:Tao Zhuang, Longbiao He, Feng Niu, Jia-Xin Zhong, Jing Lu
摘要:多通道参量阵列扬声器(MCPAL)系统提供了增强的灵活性,并有望在现实世界的应用中产生高度定向的音频波束。然而,由于涉及复杂的非线性行为和多通道信号处理,有效和准确地预测其产生的声场仍然是一个重大挑战。为了克服这个障碍,我们提出了一个k空间的方法来建模任意MCPAL系统安排在一个有障碍的平面表面。在我们的方法中,首先使用角谱方法求解线性超声场,随后在k空间中有效地计算准线性音频声场。通过利用三维快速傅立叶变换,我们的方法不仅实现了高的计算和存储效率,但也保持精度,而不依赖于傍轴近似。对于典型的配置研究,所提出的方法表现出超过四个数量级的速度相比,直接集成方法。我们提出的方法为模拟和设计先进的MCPAL系统铺平了道路。
摘要:Multi-channel parametric array loudspeaker (MCPAL) systems offer enhanced flexibility and promise for generating highly directional audio beams in real-world applications. However, efficient and accurate prediction of their generated sound fields remains a major challenge due to the complex nonlinear behavior and multi-channel signal processing involved. To overcome this obstacle, we propose a k-space approach for modeling arbitrary MCPAL systems arranged on a baffled planar surface. In our method, the linear ultrasound field is first solved using the angular spectrum approach, and the quasilinear audio sound field is subsequently computed efficiently in k-space. By leveraging three-dimensional fast Fourier transforms, our approach not only achieves high computational and memory efficiency but also maintains accuracy without relying on the paraxial approximation. For typical configurations studied, the proposed method demonstrates a speed-up of more than four orders of magnitude compared to the direct integration method. Our proposed approach paved the way for simulating and designing advanced MCPAL systems.


【2】Modeling Multi-Level Hearing Loss for Speech Intelligibility Prediction
标题:用于语音可理解度预测的多级别听力损失建模
链接:https://arxiv.org/abs/2507.22599

作者:XIAJIE ZHOU, Candy Olivia Mawalim, Masashi Unoki
备注:5 pages, 2 figures, to appear in WASPAA 2025
摘要:听力损失的各种感知后果严重阻碍了言语交流,但标准的临床测听,这是集中在基于阈值的频率灵敏度,不足以捕捉频率和时间分辨率的缺陷。为了解决这个问题,我们提出了一种语音清晰度预测方法,明确模拟听觉退化,根据听力损失的严重程度,通过扩大耳蜗滤波器和应用低通调制滤波的时间包络。语音信号随后分析使用频谱时间调制(STM)表示,这反映了听觉分辨率损失如何改变底层的调制结构。此外,归一化互相关(NCC)矩阵量化干净的语音和噪声中的语音的STM表示之间的相似性。利用这些语义信息特征来训练基于视觉变换器的回归模型,该回归模型集成STM图和NCC嵌入以估计语音可懂度分数。在Clarity Prediction Challenge语料库上的评估表明,该方法在轻度和中度至重度听力损失组中的性能优于助听器语音感知指数v2(HASPI v2),轻度组的相对均方根误差降低了16.5%,中度至重度组降低了6.1%。这些结果突出了显式建模特定的频率和时间分辨率退化,以提高语音清晰度预测和提供可解释性的听觉失真的重要性。
摘要:The diverse perceptual consequences of hearing loss severely impede speech communication, but standard clinical audiometry, which is focused on threshold-based frequency sensitivity, does not adequately capture deficits in frequency and temporal resolution. To address this limitation, we propose a speech intelligibility prediction method that explicitly simulates auditory degradations according to hearing loss severity by broadening cochlear filters and applying low-pass modulation filtering to temporal envelopes. Speech signals are subsequently analyzed using the spectro-temporal modulation (STM) representations, which reflect how auditory resolution loss alters the underlying modulation structure. In addition, normalized cross-correlation (NCC) matrices quantify the similarity between the STM representations of clean speech and speech in noise. These auditory-informed features are utilized to train a Vision Transformer-based regression model that integrates the STM maps and NCC embeddings to estimate speech intelligibility scores. Evaluations on the Clarity Prediction Challenge corpus show that the proposed method outperforms the Hearing-Aid Speech Perception Index v2 (HASPI v2) in both mild and moderate-to-severe hearing loss groups, with a relative root mean squared error reduction of 16.5% for the mild group and a 6.1% reduction for the moderate-to-severe group. These results highlight the importance of explicitly modeling listener-specific frequency and temporal resolution degradations to improve speech intelligibility prediction and provide interpretability in auditory distortions.


【3】The Risks and Detection of Overestimated Privacy Protection in Voice Anonymisation
标题:语音匿名化中高估隐私保护的风险及检测
链接:https://arxiv.org/abs/2507.22534

作者:Michele Panariello, Sarina Meyer, Pierre Champion, Xiaoxiao Miao, Massimiliano Todisco, Ngoc Thang Vu, Nicholas Evans
备注:Accepted at SPSC 2025 - 5th Symposium on Security and Privacy in Speech Communication
摘要:语音匿名化的目的是在语音记录中隐藏说话者的语音身份。隐私保护通常是从使用说话人验证系统来重新识别匿名化后的说话人的难度来估计的。因此,业绩评估取决于核查模式和匿名系统。因此,当核查系统训练不足,可能数据不匹配时,就有可能高估隐私保护。在本文中,我们展示了高估匿名化性能的潜在风险,并展示了文献中报道的夸大性能的例子。对于我们确定的最差情况,性能相对高估了74%。然后,我们引入了一种方法来检测性能评估时可能是不值得信赖的,并表明它可以识别所有高估的情况下提出的文件。我们的解决方案作为2024年语音隐私挑战评估工具包的一个分支公开提供。
摘要:Voice anonymisation aims to conceal the voice identity of speakers in speech recordings. Privacy protection is usually estimated from the difficulty of using a speaker verification system to re-identify the speaker post-anonymisation. Performance assessments are therefore dependent on the verification model as well as the anonymisation system. There is hence potential for privacy protection to be overestimated when the verification system is poorly trained, perhaps with mismatched data. In this paper, we demonstrate the insidious risk of overestimating anonymisation performance and show examples of exaggerated performance reported in the literature. For the worst case we identified, performance is overestimated by 74% relative. We then introduce a means to detect when performance assessment might be untrustworthy and show that it can identify all overestimation scenarios presented in the paper. Our solution is openly available as a fork of the 2024 VoicePrivacy Challenge evaluation toolkit.


【4】Tiny Noise-Robust Voice Activity Detector for Voice Assistants
标题:用于语音助理的微型噪音稳健语音活动检测器
链接:https://arxiv.org/abs/2507.22157

作者:Hamed Jafarzadeh Asl, Mahsa Ghazvini Nejad, Amin Edraki, Masoud Asgharian, Vahid Partovi Nia
备注:Hamed Jafarzadeh Asl and Mahsa Ghazvini Nejad contributed equally to   this work
摘要:背景噪声背景下的语音端点检测(VAD)一直是语音处理中的一个难题。准确的VAD对于自动语音识别、语音转文本、对话代理等至关重要,其中噪音会严重降低性能。现代应用包括专门安装在诸如蜂窝电话、智能眼镜、耳塞等的人工智能(AIoT)设备上的语音助理,其中语音信号包括背景噪声。因此,VAD模块由于其实际的设备上的限制而必须保持轻量化。现有的模型通常在不同的声学环境中与低信噪比作斗争。简单的VAD通常在干净的环境中检测人类声音,但在嘈杂的环境中很难检测人类声音。我们提出了一个噪声鲁棒VAD,包括一个轻量级的VAD,与数据预处理和后处理添加模块来处理背景噪声。这种方法显著提高了VAD在噪声环境中的精度,并且既不需要更大的模型,也不需要微调。实验结果表明,与基线相比,我们的方法取得了显着的改进,特别是在背景噪声干扰较强的环境中。这种改进的VAD还改进了干净语音检测。
摘要:Voice Activity Detection (VAD) in the presence of background noise remains a challenging problem in speech processing. Accurate VAD is essential in automatic speech recognition, voice-to-text, conversational agents, etc, where noise can severely degrade the performance. A modern application includes the voice assistant, specially mounted on Artificial Intelligence of Things (AIoT) devices such as cell phones, smart glasses, earbuds, etc, where the voice signal includes background noise. Therefore, VAD modules must remain light-weight due to their practical on-device limitation. The existing models often struggle with low signal-to-noise ratios across diverse acoustic environments. A simple VAD often detects human voice in a clean environment, but struggles to detect the human voice in noisy conditions. We propose a noise-robust VAD that comprises a light-weight VAD, with data pre-processing and post-processing added modules to handle the background noise. This approach significantly enhances the VAD accuracy in noisy environments and requires neither a larger model, nor fine-tuning. Experimental results demonstrate that our approach achieves a notable improvement compared to baselines, particularly in environments with high background noise interference. This modified VAD additionally improving clean speech detection.


【5】Scaling and Distilling Transformer Models for sEMG
标题:sEMG的缩放和提炼Transformer模型
链接:https://arxiv.org/abs/2507.22094

作者:Nicholas Mehlman, Jean-Christophe Gagnon-Audet, Michael Shvartsman, Kelvin Niu, Alexander H. Miller, Shagun Sodhani
备注:Accepted at TMLR 2025 (this https URL), 11 pages
摘要:表面肌电图(sEMG)信号通过提供对肌肉活动的洞察,为开发创新的人机界面提供了一条有前途的途径。然而,有限的训练数据量和部署过程中的计算约束,限制了扩大模型规模解决表面肌电信号任务的调查。在本文中,我们证明了香草Transformer模型可以有效地扩展到sEMG数据,并提高跨用户性能高达110 M参数,超过了其他sEMG研究中研究的模型大小制度(通常<10 M参数)。我们发现,> 100 M参数的模型可以有效地提炼成小50倍的模型,性能损失最小(<1.5%绝对值)。这将产生适用于现实环境中复杂实时sEMG任务的高效和富有表现力的模型。
摘要:Surface electromyography (sEMG) signals offer a promising avenue for developing innovative human-computer interfaces by providing insights into muscular activity. However, the limited volume of training data and computational constraints during deployment have restricted the investigation of scaling up the model size for solving sEMG tasks. In this paper, we demonstrate that vanilla transformer models can be effectively scaled up on sEMG data and yield improved cross-user performance up to 110M parameters, surpassing the model size regime investigated in other sEMG research (usually <10M parameters). We show that >100M-parameter models can be effectively distilled into models 50x smaller with minimal loss of performance (<1.5% absolute). This results in efficient and expressive models suitable for complex real-time sEMG tasks in real-world environments.


【6】Next Tokens Denoising for Speech Synthesis
标题:下一个语音合成的令牌去噪
链接:https://arxiv.org/abs/2507.22746

作者:Yanqing Liu, Ruiqing Xue, Chong Zhang, Yufei Liu, Gang Wang, Bohan Li, Yao Qian, Lei He, Shujie Liu, Sheng Zhao
摘要:虽然扩散和自回归(AR)模型具有显着先进的生成建模,但它们各自存在明显的局限性。依赖于因果注意的AR模型无法利用未来的上下文,并且生成速度缓慢。相反,扩散模型则要处理键值(KV)缓存。为了克服这些挑战,我们引入了Dragon-FM,这是一种新颖的文本到语音(TTS)设计,它将AR和流匹配相结合。该模型以每秒12.5个令牌的紧凑速率以块的形式处理48 kHz音频编解码器令牌。这种设计使AR建模跨块,确保全局一致性,而块内的并行流匹配有助于快速迭代去噪。因此,所提出的模型可以跨块利用KV缓存,并将未来的上下文合并到每个块中。此外,它还连接了连续和离散特征建模,证明连续AR流匹配可以用有限的标量量化器预测离散令牌。这种高效的编解码器和快速块自回归架构也使得所提出的模型对于生成扩展内容特别有效。在播客数据集上进行的演示实验表明,该方法能够有效地生成高质量的zero-shot播客。
摘要:While diffusion and autoregressive (AR) models have significantly advanced generative modeling, they each present distinct limitations. AR models, which rely on causal attention, cannot exploit future context and suffer from slow generation speeds. Conversely, diffusion models struggle with key-value (KV) caching. To overcome these challenges, we introduce Dragon-FM, a novel text-to-speech (TTS) design that unifies AR and flow-matching. This model processes 48 kHz audio codec tokens in chunks at a compact 12.5 tokens per second rate. This design enables AR modeling across chunks, ensuring global coherence, while parallel flow-matching within chunks facilitates fast iterative denoising. Consequently, the proposed model can utilize KV-cache across chunks and incorporate future context within each chunk. Furthermore, it bridges continuous and discrete feature modeling, demonstrating that continuous AR flow-matching can predict discrete tokens with finite scalar quantizers. This efficient codec and fast chunk-autoregressive architecture also makes the proposed model particularly effective for generating extended content. Experiment for demos of our work} on podcast datasets demonstrate its capability to efficiently generate high-quality zero-shot podcasts.


【7】Adaptive Duration Model for Text Speech Alignment
标题:文本语音对齐的自适应持续时间模型
链接:https://arxiv.org/abs/2507.22612

作者:Junjie Cao
备注:4 pages, 3 figures, 2 tables
摘要:语音到文本对齐是神经文本到语音(TTS)模型的关键组成部分。自回归TTS模型通常使用注意力机制来在线学习这些对齐。然而,这些对齐往往是脆弱的,并且经常不能推广到长话语和域外文本,导致单词丢失或重复。大多数非自回归端到端TTS模型依赖于从外部源提取的持续时间,使用额外的持续时间模型进行对齐。在本文中,我们提出了一个新的持续时间预测框架,可以给妥协音素级的持续时间分布与给定的文本。实验结果表明,与已有的基线模型相比,该模型具有更高的预测精度和条件自适应能力。从数值上看,它在对齐精度上有大约11.3个百分点的提高,并且使zero-shot TTS模型的性能对提示音频和输入音频之间的失配更鲁棒。
摘要:Speech-to-text alignment is a critical component of neural text to-speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line. However, these alignments tend to be brittle and often fail to generalize to long utterances and out-of-domain text, leading to missing or repeating words. Most non-autoregressive end to-end TTS models rely on durations extracted from external sources, using additional duration models for alignment. In this paper, we propose a novel duration prediction framework that can give compromising phoneme-level duration distribution with given text. In our experiments, the proposed duration model has more precise prediction and condition adaptation ability compared to previous baseline models. Numerically, it has roughly a 11.3 percents immprovement on alignment accuracy, and makes the performance of zero-shot TTS models more robust to the mismatch between prompt audio and input audio.


【8】Prediction of acoustic field in 1-D uniform duct with varying mean flow and temperature using neural networks
标题:用神经网络预测平均流量和温度变化的一维均匀管道中的声学场
链接:https://arxiv.org/abs/2507.22370

作者:Veerababu Dharanalakota, Prasanta K. Ghosh
备注:22 pages
摘要:受物理定律约束的神经网络作为另一种数值工具出现。本文推导了一维管道中声传播的控制方程。将问题转化为无约束优化问题,并使用神经网络进行求解。采用传统的龙格-库塔法对声状态变量声压和质点速度进行了预测和验证。研究了温度梯度对声场的影响。利用机器学习技术,如迁移学习和自动微分声学应用程序的演示。
摘要:Neural networks constrained by the physical laws emerged as an alternate numerical tool. In this paper, the governing equation that represents the propagation of sound inside a one-dimensional duct carrying a heterogeneous medium is derived. The problem is converted into an unconstrained optimization problem and solved using neural networks. Both the acoustic state variables: acoustic pressure and particle velocity are predicted and validated with the traditional Runge-Kutta solver. The effect of the temperature gradient on the acoustic field is studied. Utilization of machine learning techniques such as transfer learning and automatic differentiation for acoustic applications is demonstrated.


【9】A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection
标题:用于增强声音事件定位和检测的两步学习框架
链接:https://arxiv.org/abs/2507.22322

作者:Hogeon Yu
备注:5pages, 2figures
摘要:声音事件定位和检测(SELD)在空间音频处理中至关重要,使系统能够检测声音事件并估计其3D方向。现有的SELD方法使用单分支或双分支架构:单分支模型共享SED和DoA表示,导致优化冲突,而双分支模型分离任务但限制信息交换。为了解决这个问题,我们提出了一个两步学习框架。首先,我们引入了一个tracwise重新排序格式,以保持时间的一致性,防止跨轨道的事件重新排序。接下来,我们训练SED和DoA网络以防止干扰并确保特定于任务的特征学习。最后,我们有效地融合DoA和SED功能,以提高SELD性能,更好的空间和事件表示。在2023年DCASE挑战任务3数据集上的实验验证了我们的框架,显示了它克服单分支和双分支限制并改进事件分类和本地化的能力。
摘要:Sound Event Localization and Detection (SELD) is crucial in spatial audio processing, enabling systems to detect sound events and estimate their 3D directions. Existing SELD methods use single- or dual-branch architectures: single-branch models share SED and DoA representations, causing optimization conflicts, while dual-branch models separate tasks but limit information exchange. To address this, we propose a two-step learning framework. First, we introduce a tracwise reordering format to maintain temporal consistency, preventing event reassignments across tracks. Next, we train SED and DoA networks to prevent interference and ensure task-specific feature learning. Finally, we effectively fuse DoA and SED features to enhance SELD performance with better spatial and event representation. Experiments on the 2023 DCASE challenge Task 3 dataset validate our framework, showing its ability to overcome single- and dual-branch limitations and improve event classification and localization.


【10】Quantum-Inspired Audio Unlearning: Towards Privacy-Preserving Voice Biometrics
标题:量子启发的音频遗忘:迈向保护隐私的语音生物识别技术
链接:https://arxiv.org/abs/2507.22208

作者: Shreyansh Pathak, Sonu Shreshtha, Richa Singh, Mayank Vatsa
备注:9 pages, 2 figures, 5 tables, Accepted at IJCB 2025 (Osaka, Japan)
摘要:语音认证和音频生物识别系统的广泛采用大大增加了与敏感语音数据相关的隐私漏洞。要遵守GDPR的被遗忘权和印度的DPDP法案等隐私法规,就必须从已经训练好的生物识别模型中有针对性地有效删除个人特定的语音签名。现有的为视觉数据设计的去学习方法不能充分处理音频信号的顺序、时间和高维性质,导致无效或不完整的说话者和口音擦除。为了解决这个问题,我们引入了QPAudioEraser,这是一个量子启发的音频遗忘框架。我们的our-phase方法包括:(1)使用相消干涉来取消目标特征的权重初始化,(2)基于叠加的标签转换,掩盖类身份,(3)不确定性最大化量子损失函数,以及(4)纠缠启发的相关权重混合,以保留模型知识。ResNet 18、ViT和CNN架构在AudioMNIST、Speech Commands、LibriSpeech和Speech Accent Archive数据集上的综合评估验证了QPAudioEraser的卓越性能。该框架实现了目标数据的完全擦除(0%遗忘精度),同时对模型实用性的影响最小,保留数据的性能下降低至0.05%。QPAudioEraser在单类、多类、顺序和重音级别擦除场景中始终超越传统基线,将所提出的方法确立为一种强大的隐私保护解决方案。
摘要:The widespread adoption of voice-enabled authentication and audio biometric systems have significantly increased privacy vulnerabilities associated with sensitive speech data. Compliance with privacy regulations such as GDPR's right to be forgotten and India's DPDP Act necessitates targeted and efficient erasure of individual-specific voice signatures from already-trained biometric models. Existing unlearning methods designed for visual data inadequately handle the sequential, temporal, and high-dimensional nature of audio signals, leading to ineffective or incomplete speaker and accent erasure. To address this, we introduce QPAudioEraser, a quantum-inspired audio unlearning framework. Our our-phase approach involves: (1) weight initialization using destructive interference to nullify target features, (2) superposition-based label transformations that obscure class identity, (3) an uncertainty-maximizing quantum loss function, and (4) entanglement-inspired mixing of correlated weights to retain model knowledge. Comprehensive evaluations with ResNet18, ViT, and CNN architectures across AudioMNIST, Speech Commands, LibriSpeech, and Speech Accent Archive datasets validate QPAudioEraser's superior performance. The framework achieves complete erasure of target data (0% Forget Accuracy) while incurring minimal impact on model utility, with a performance degradation on retained data as low as 0.05%. QPAudioEraser consistently surpasses conventional baselines across single-class, multi-class, sequential, and accent-level erasure scenarios, establishing the proposed approach as a robust privacy-preserving solution.


机器翻译由腾讯交互翻译提供,仅供参考