今日论文合集:cs.SD语音13篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
标题:RST Steer-TTC:细粒度且免训练的可通过激活引导控制的文本转语音
链接:https://arxiv.org/abs/2508.03543

作者:Tianxin Xie, Shan Yang, Chenxing Li, Dong Yu, Li Liu
摘要:文语转换(TTS)近年来取得了很大的进展。然而,大多数现有的TTS系统只提供粗糙和僵硬的情感控制,通常通过离散的情感标签或精心制作的和详细的情感文本提示,使得细粒度的情感操纵要么无法访问或不稳定。这些模型还需要大量高质量的数据集进行训练。为了解决这些局限性,我们提出了一种新的无训练的方法--TTS,通过激活转向来实现细粒度的语音情感控制(转换、插值、擦除)。我们首先经验性地观察到,修改基于流匹配的TTS模型内的内部激活的子集可以有效地改变合成语音的情感基调。在此基础上,我们开发了一种免训练的高效算法,包括激活提取、情感令牌搜索和推理时间转向,这些算法可以无缝集成到各种预训练模型中(例如,F5-TTS、CosyVoice 2和E2-TTS)。此外,为了获得有效的导向向量,我们构建了一个有不同发言者的策划情感语音数据集。大量的实验表明,语音转向TTS能够对语音情感进行细粒度的、可解释的和连续的控制,性能优于最先进的(SOTA)。据我们所知,这是第一个在TTS中实现无训练和连续细粒度情感控制的方法。
摘要:Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emotional text prompt, making fine-grained emotion manipulation either inaccessible or unstable. These models also require extensive, high-quality datasets for training. To address these limitations, we propose EmoSteer-TTS, a novel training-free approach, to achieve fine-grained speech emotion control (conversion, interpolation, erasure) by activation steering. We first empirically observe that modifying a subset of the internal activations within a flow matching-based TTS model can effectively alter the emotional tone of synthesized speech. Building on this insight, we then develop a training-free and efficient algorithm, including activation extraction, emotional token searching, and inference-time steering, which can be seamlessly integrated into a wide range of pretrained models (e.g., F5-TTS, CosyVoice2, and E2-TTS). In addition, to derive effective steering vectors, we construct a curated emotional speech dataset with diverse speakers. Extensive experiments demonstrate that EmoSteer-TTS enables fine-grained, interpretable, and continuous control over speech emotion, outperforming the state-of-the-art (SOTA). To the best of our knowledge, this is the first method that achieves training-free and continuous fine-grained emotion control in TTS.


【2】READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
标题:阅读:实时高效的同步扩散,用于音频驱动的会说话的头部生成
链接:https://arxiv.org/abs/2508.03457

作者:Haotian Wang, Yuzhe Weng, Jun Du, Haoran Xu, Xiaoyan Wu, Shan He, Bing Yin, Cong Liu, Jianqing Gao, Qingfeng Liu
备注:9 pages
摘要:扩散模型的引入为音频驱动的说话头部生成领域带来了重大进展。然而,极慢的推理速度严重限制了基于扩散的说话头生成模型的实际实现。在这项研究中,我们提出了READ,第一个实时扩散变压器为基础的说话头生成框架。我们的方法首先通过时间VAE学习时空高度压缩的视频潜在空间,显着减少令牌计数以加速生成。为了实现更好的视听对齐在这个压缩的潜在空间内,一个预先训练的语音自动编码器(SpeechAE),提出了生成时间压缩的语音潜在代码对应的视频潜在空间。这些潜在的表示,然后通过精心设计的音频到视频扩散Transformer(A2 V-DiT)的骨干有效的说话头合成建模。此外,为了确保时间的一致性和加速推理的扩展代,我们提出了一种新的异步噪声调度器(ANS)的训练和推理过程中,我们的框架。ANS在潜在空间中利用异步添加噪声和异步运动引导生成,确保生成的视频剪辑的一致性。实验结果表明,READ优于国家的最先进的方法,通过生成具有显着降低运行时间的有竞争力的讲话头部视频,实现质量和速度之间的最佳平衡,同时保持鲁棒的度量稳定性,在长时间生成。
摘要:The introduction of diffusion models has brought significant advances to the field of audio-driven talking head generation. However, the extremely slow inference speed severely limits the practical implementation of diffusion-based talking head generation models. In this study, we propose READ, the first real-time diffusion-transformer-based talking head generation framework. Our approach first learns a spatiotemporal highly compressed video latent space via a temporal VAE, significantly reducing the token count to accelerate generation. To achieve better audio-visual alignment within this compressed latent space, a pre-trained Speech Autoencoder (SpeechAE) is proposed to generate temporally compressed speech latent codes corresponding to the video latent space. These latent representations are then modeled by a carefully designed Audio-to-Video Diffusion Transformer (A2V-DiT) backbone for efficient talking head synthesis. Furthermore, to ensure temporal consistency and accelerated inference in extended generation, we propose a novel asynchronous noise scheduler (ANS) for both the training and inference process of our framework. The ANS leverages asynchronous add-noise and asynchronous motion-guided generation in the latent space, ensuring consistency in generated video clips. Experimental results demonstrate that READ outperforms state-of-the-art methods by generating competitive talking head videos with significantly reduced runtime, achieving an optimal balance between quality and speed while maintaining robust metric stability in long-time generation.


【3】SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
标题:SonicMaster:迈向可控的一体化音乐修复和掌握
链接:https://arxiv.org/abs/2508.03448

作者:Jan Melechovsky, Ambuj Mehrish, Dorien Herremans
摘要:音乐录音通常会遇到音频质量问题,例如过度混响、失真、削波、音调不平衡和狭窄的立体声图像,特别是在没有专业设备或专业知识的非专业环境中创建时。这些问题通常使用单独的专用工具和手动调整来纠正。在本文中,我们介绍了SonicMaster,这是第一个用于音乐恢复和母带的统一生成模型,它通过基于文本的控制解决了广泛的音频伪影。SonicMaster以自然语言指令为条件,以应用有针对性的增强功能,或者可以在自动模式下进行一般恢复。为了训练这个模型,我们构建了SonicMaster数据集,这是一个大型的成对降级和高质量音轨数据集,通过模拟常见的降级类型,其中19个降级函数属于5个增强组:均衡,动态,混响,幅度和立体声。我们的方法利用流匹配生成训练范式来学习音频转换,该转换将降级的输入映射到由文本提示引导的清洁的、掌握的版本。客观的音频质量指标表明,SonicMaster显着提高了所有工件类别的声音质量。此外,主观听力测试证实,听众更喜欢SonicMaster的增强输出,而不是原来的降级音频,突出了我们的统一方法的有效性。
摘要:Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment or expertise. These problems are typically corrected using separate specialized tools and manual adjustments. In this paper, we introduce SonicMaster, the first unified generative model for music restoration and mastering that addresses a broad spectrum of audio artifacts with text-based control. SonicMaster is conditioned on natural language instructions to apply targeted enhancements, or can operate in an automatic mode for general restoration. To train this model, we construct the SonicMaster dataset, a large dataset of paired degraded and high-quality tracks by simulating common degradation types with nineteen degradation functions belonging to five enhancements groups: equalization, dynamics, reverb, amplitude, and stereo. Our approach leverages a flow-matching generative training paradigm to learn an audio transformation that maps degraded inputs to their cleaned, mastered versions guided by text prompts. Objective audio quality metrics demonstrate that SonicMaster significantly improves sound quality across all artifact categories. Furthermore, subjective listening tests confirm that listeners prefer SonicMaster's enhanced outputs over the original degraded audio, highlighting the effectiveness of our unified approach.


【4】When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
标题:当好的声音变得敌对时:用良性输入越狱的音频模型
链接:https://arxiv.org/abs/2508.03365

作者:Bodam Kim, Hiskias Dingeto, Taeyoun Kwon, Dasol Choi, DongGeon Lee, Haon Park, JaeHoon Lee, Jongho Shin
摘要:随着大型语言模型越来越多地融入日常生活,音频已成为人机交互的关键接口。然而,这种便利性也引入了新的漏洞,使音频成为对手的潜在攻击面。我们的研究引入了WhisperInject,这是一个两阶段的对抗性音频攻击框架,可以操纵最先进的音频语言模型来生成有害内容。我们的方法在音频输入中使用对人类听众保持良性的不可察觉的扰动。第一阶段使用一种新的基于奖励的优化方法,投影梯度下降强化学习(RL-PGD),引导目标模型规避其自身的安全协议并生成有害的本地响应。然后,这种原生有害响应将作为阶段2有效载荷注入的目标,在该阶段中,我们使用投影梯度下降(PGD)来优化嵌入到良性音频载体中的细微扰动,例如天气查询或问候消息。在严格的StrongRESIDENT、LlamaGuard以及人体评估安全评估框架下进行验证,我们的实验表明,Qwen2.5-Omni-3B、Qwen2.5-Omni-7 B和Phi-4-Multimodal的成功率超过86%。我们的工作展示了一类新的实际的音频原生威胁,超越了理论上的利用,揭示了一种操纵人工智能行为的可行和隐蔽的方法。
摘要:As large language models become increasingly integrated into daily life, audio has emerged as a key interface for human-AI interaction. However, this convenience also introduces new vulnerabilities, making audio a potential attack surface for adversaries. Our research introduces WhisperInject, a two-stage adversarial audio attack framework that can manipulate state-of-the-art audio language models to generate harmful content. Our method uses imperceptible perturbations in audio inputs that remain benign to human listeners. The first stage uses a novel reward-based optimization method, Reinforcement Learning with Projected Gradient Descent (RL-PGD), to guide the target model to circumvent its own safety protocols and generate harmful native responses. This native harmful response then serves as the target for Stage 2, Payload Injection, where we use Projected Gradient Descent (PGD) to optimize subtle perturbations that are embedded into benign audio carriers, such as weather queries or greeting messages. Validated under the rigorous StrongREJECT, LlamaGuard, as well as Human Evaluation safety evaluation framework, our experiments demonstrate a success rate exceeding 86% across Qwen2.5-Omni-3B, Qwen2.5-Omni-7B, and Phi-4-Multimodal. Our work demonstrates a new class of practical, audio-native threats, moving beyond theoretical exploits to reveal a feasible and covert method for manipulating AI behavior.


【5】MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
标题:MiSTR:基于变换器的韵律预测和神经相位重建的多模态iEEG到语音合成
链接:https://arxiv.org/abs/2508.03166

作者:Mohammed Salah Al-Radhi, Géza Németh, Branislav Gerazov
备注:5 pages, 2 figures, 1 table. Accepted for presentation at Interspeech 2025
摘要:从颅内脑电(iEEG)信号的语音合成提供了一个有前途的途径,恢复严重的语言障碍的个人的沟通。然而,由于特征表示、韵律建模和相位重建的限制,实现可理解和自然的语音仍然具有挑战性。我们介绍了MiSTR,这是一个深度学习框架,它集成了:1)基于小波的特征提取,以捕获iEEG信号的细粒度时间,频谱和神经生理表示,2)基于transformer的解码器,用于韵律感知频谱图预测,以及3)神经相位声码器,通过自适应频谱校正实现谐波一致性。在公共iEEG数据集上进行评估,MiSTR实现了最先进的语音清晰度,重建和原始Mel频谱图之间的平均Pearson相关性为0.91,优于现有的神经语音合成基线。
摘要:Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challenging due to limitations in feature representation, prosody modeling, and phase reconstruction. We introduce MiSTR, a deep-learning framework that integrates: 1) Wavelet-based feature extraction to capture fine-grained temporal, spectral, and neurophysiological representations of iEEG signals, 2) A Transformer-based decoder for prosody-aware spectrogram prediction, and 3) A neural phase vocoder enforcing harmonic consistency via adaptive spectral correction. Evaluated on a public iEEG dataset, MiSTR achieves state-of-the-art speech intelligibility, with a mean Pearson correlation of 0.91 between reconstructed and original Mel spectrograms, improving over existing neural speech synthesis baselines.


【6】Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
标题:使用人类反馈的强化学习微调文本到语音扩散模型
链接:https://arxiv.org/abs/2508.03123

作者:Jingyi Chen, Ju Seung Byun, Micha Elsner, Pichao Wang, Andrew Perrault
备注:4 pages, 1 figure, INTERSPEECH 2025. arXiv admin note: text overlap   with arXiv:2405.14632
摘要:扩散模型产生高保真语音,但由于长的去噪步骤和建模语调和节奏的挑战,对于实时使用是低效的。为了改善这一点,我们提出了扩散损失导向的政策优化(DLPO),一个RLHF框架的TTS扩散模型。DLPO将原始训练损失集成到奖励函数中,保留生成能力,同时减少效率低下。使用自然度分数作为反馈,DLPO将奖励优化与扩散模型的结构相结合,从而提高语音质量。我们在WaveGrad 2(一种基于非自回归扩散的TTS模型)上评估DLPO。结果显示,在客观指标(UTMOS 3.65,NISQA 4.02)和主观评价,DLPO音频首选67%的时间显着改善。这些发现表明DLPO在实时、资源有限的环境中具有高效、高质量扩散TTS的潜力。
摘要:Diffusion models produce high-fidelity speech but are inefficient for real-time use due to long denoising steps and challenges in modeling intonation and rhythm. To improve this, we propose Diffusion Loss-Guided Policy Optimization (DLPO), an RLHF framework for TTS diffusion models. DLPO integrates the original training loss into the reward function, preserving generative capabilities while reducing inefficiencies. Using naturalness scores as feedback, DLPO aligns reward optimization with the diffusion model's structure, improving speech quality. We evaluate DLPO on WaveGrad 2, a non-autoregressive diffusion-based TTS model. Results show significant improvements in objective metrics (UTMOS 3.65, NISQA 4.02) and subjective evaluations, with DLPO audio preferred 67\% of the time. These findings demonstrate DLPO's potential for efficient, high-quality diffusion TTS in real-time, resource-limited settings.


【7】TF-MLPNet: Tiny Real-Time Neural Speech Separation
标题:TF-MLPNet:微小的实时神经语音分离
链接:https://arxiv.org/abs/2508.03047

作者:Malek Itani, Tuochao Chen, Shyamnath Gollakota
备注:The 6th Clarity Workshop on Improving Speech-in-Noise for Hearing Devices (Clarity 2025)
摘要:听觉设备上的语音分离可以实现变革性的增强和增强的听觉能力。然而,最先进的语音分离网络无法在为听觉设计的小型低功耗神经加速器上实时运行,因为它们的计算能力有限。我们提出了TF-MLPNet,这是第一个能够在这种低功耗加速器上实时运行的语音分离网络,同时优于现有的盲语音分离和目标语音提取流模型。我们的网络在时频域中运行,使用沿信道和频率维度交替的全连接层堆栈处理频率序列,并使用卷积层独立处理每个频率点的时间序列。结果表明,我们的混合精度量化感知训练(QAT)模型可以在GAP 9处理器上实时处理6 ms的音频块,与之前的语音分离模型相比,运行时间减少了3.5- 4倍。
摘要:Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelerators designed for hearables, due to their limited compute capabilities. We present TF-MLPNet, the first speech separation network capable of running in real-time on such low-power accelerators while outperforming existing streaming models for blind speech separation and target speech extraction. Our network operates in the time-frequency domain, processing frequency sequences with stacks of fully connected layers that alternate along the channel and frequency dimensions, and independently processing the time sequence at each frequency bin using convolutional layers. Results show that our mixed-precision quantization-aware trained (QAT) model can process 6 ms audio chunks in real-time on the GAP9 processor, achieving a 3.5-4x runtime reduction compared to prior speech separation models.


【8】Neural Speech Extraction with Human Feedback
标题:利用人类反馈的神经语音提取
链接:https://arxiv.org/abs/2508.03041

作者:Malek Itani, Ashton Graves, Sefik Emre Eskimez, Shyamnath Gollakota
备注:Interspeech 2025
摘要:我们提出了第一个神经目标语音提取(TSE)系统,使用人的反馈迭代细化。我们的方法允许用户标记TSE输出的特定部分,生成编辑掩码。然后,细化系统改进标记的部分,同时保留未标记的区域。由于人类标记的错误的大规模数据集很难收集,我们使用各种自动掩蔽函数生成合成数据集,并对每个数据集进行训练。评估表明,使用基于噪声功率的掩蔽(以dBFS为单位)和概率阈值训练的模型表现最好,与人类注释保持一致。在一项有22名参与者的研究中,用户显示出对精炼输出的偏好超过基线TSE。我们的研究结果表明,人在环细化是提高神经语音提取性能的一种很有前途的方法。
摘要:We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then improves the marked sections while preserving unmarked regions. Since large-scale datasets of human-marked errors are difficult to collect, we generate synthetic datasets using various automated masking functions and train models on each. Evaluations show that models trained with noise power-based masking (in dBFS) and probabilistic thresholding perform best, aligning with human annotations. In a study with 22 participants, users showed a preference for refined outputs over baseline TSE. Our findings demonstrate that human-in-the-loop refinement is a promising approach for improving the performance of neural speech extraction.


【9】How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
标题:听起来怎么样?室内场景的材料控制多峰声学轮廓生成
链接:https://arxiv.org/abs/2508.02905

作者:Mahnoor Fatima Saad, Ziad Al-Halah
备注:Accepted to ICCV 2025. Project Page: this https URL
摘要:在一个铺着地毯的地板和墙上贴着隔音砖的录音室里,声音会有什么变化?我们介绍了材料控制的声学配置文件生成的任务,其中,给定一个室内场景具有特定的视听特性,目标是在推理时间根据用户定义的材料配置生成目标声学配置文件。我们解决这个任务与一种新的编码器-解码器的方法,编码场景的关键属性从视听观察和生成目标房间脉冲响应(RIR)的条件下,由用户提供的材料规格。我们的模型能够生成不同的RIR的基础上动态定义的各种材料配置在推理时间。为了支持这一任务,我们创建了一个新的基准,声学仙境数据集,旨在开发和评估材料感知RIR预测方法在不同的和具有挑战性的设置。我们的研究结果表明,该模型有效地编码材料信息,并产生高保真RIR,优于几个基线和国家的最先进的方法。
摘要:How would the sound in a studio change with a carpeted floor and acoustic tiles on the walls? We introduce the task of material-controlled acoustic profile generation, where, given an indoor scene with specific audio-visual characteristics, the goal is to generate a target acoustic profile based on a user-defined material configuration at inference time. We address this task with a novel encoder-decoder approach that encodes the scene's key properties from an audio-visual observation and generates the target Room Impulse Response (RIR) conditioned on the material specifications provided by the user. Our model enables the generation of diverse RIRs based on various material configurations defined dynamically at inference time. To support this task, we create a new benchmark, the Acoustic Wonderland Dataset, designed for developing and evaluating material-aware RIR prediction methods under diverse and challenging settings. Our results demonstrate that the proposed model effectively encodes material information and generates high-fidelity RIRs, outperforming several baselines and state-of-the-art methods.


【10】Adaptive Knowledge Distillation for Device-Directed Speech Detection
标题:设备指导语音检测的自适应知识提炼
链接:https://arxiv.org/abs/2508.02801

作者: Hyung Gun Chi, Florian Pesce, Wonil Chang, Oggi Rudovic, Arturo Argueta, Stefan Braun, Vineet Garg, Ahmed Hussen Abdelaziz
备注:5 pages, 2 figures, Interspeech accepted
摘要:设备导向语音检测(DDSD)是一种二进制分类任务,它将用户对语音助理(VA)的查询与背景语音或侧面对话分开。这对于实现自然的用户体验非常重要。为此,我们提出了知识蒸馏(KD),以提高DDSD的准确性,同时确保有效的部署。具体来说,我们介绍了一种新的自适应KD方法,该方法从ASR大型预训练声学编码器(教师)的一般表示中传输知识。我们应用特定于任务的适配器,在(冻结)教师编码器的顶部,与DDSD上的学生模型一起训练。我们证明了所提出的自适应KD在关键字和无关键字(后续)调用方面优于没有蒸馏的学生模型,在相等错误率方面分别提高了+26%和+19%。我们还表明,这种方法概括了整个Transformer和一致性为基础的模型架构。
摘要:Device-directed speech detection (DDSD) is a binary classification task that separates the user's queries to a voice assistant (VA) from background speech or side conversations. This is important for achieving naturalistic user experience. To this end, we propose knowledge distillation (KD) to enhance DDSD accuracy while ensuring efficient deployment. Specifically, we introduce a novel adaptive KD method that transfers knowledge from general representations of an ASR large pre-trained acoustic encoder (teacher). We apply task-specific adapters, on top of the (frozen) teacher encoder, trained jointly with the student model on DDSD. We demonstrate that the proposed adaptive KD outperforms the student model without distillation in the keyword and keyword-free (follow-up) invocations, with an improvement of +26% and +19% in terms of Equal Error Rate, respectively. We also show that this approach generalizes across the transformer and conformer-based model architectures.


【11】DeepGB-TB: A Risk-Balanced Cross-Attention Gradient-Boosted Convolutional Network for Rapid, Interpretable Tuberculosis Screening
标题:DeepGB-TB:一个风险平衡的交叉注意力受试者推动的卷积网络,用于快速、可解释的结核病筛查
链接:https://arxiv.org/abs/2508.02741

作者:Zhixiang Lu, Yulong Li, Feilong Tang, Zhengyong Jiang, Chong Li, Mian Zhou, Tenglong Li, Jionglong Su
摘要:大规模结核病筛查受到传统诊断的高成本和操作复杂性的限制,因此需要人工智能解决方案。我们提出了DeepGB-TB,这是一种非侵入性系统,仅使用咳嗽音频和基本人口统计数据即可立即分配结核病风险评分。该模型将用于音频处理的轻量级一维卷积神经网络与用于表格特征的梯度提升决策树相结合。它的主要创新是跨模态双向交叉注意模块(CM-BCA),该模块在模态之间迭代地交换显著线索,模仿临床医生整合症状和风险因素的方式。为了满足最大限度地减少漏诊病例的临床优先级,我们设计了一个结核病风险平衡损失(TRBL),对假阴性预测施加更强的惩罚,从而减少高风险的错误分类。DeepGB-TB在7个国家收集的1,105名患者的多样化数据集上进行了评估,AUROC为0.903,F1得分为0.851,代表了最新的技术水平。其计算效率可以直接在普通移动设备上进行实时离线推理,非常适合低资源环境。重要的是,该系统产生了临床验证的解释,促进了一线卫生工作者的信任和采用。通过将人工智能创新与公共卫生对速度、可负担性和可靠性的要求相结合,DeepGB-TB为推进全球结核病控制提供了一种工具。
摘要:Large-scale tuberculosis (TB) screening is limited by the high cost and operational complexity of traditional diagnostics, creating a need for artificial-intelligence solutions. We propose DeepGB-TB, a non-invasive system that instantly assigns TB risk scores using only cough audio and basic demographic data. The model couples a lightweight one-dimensional convolutional neural network for audio processing with a gradient-boosted decision tree for tabular features. Its principal innovation is a Cross-Modal Bidirectional Cross-Attention module (CM-BCA) that iteratively exchanges salient cues between modalities, emulating the way clinicians integrate symptoms and risk factors. To meet the clinical priority of minimizing missed cases, we design a Tuberculosis Risk-Balanced Loss (TRBL) that places stronger penalties on false-negative predictions, thereby reducing high-risk misclassifications. DeepGB-TB is evaluated on a diverse dataset of 1,105 patients collected across seven countries, achieving an AUROC of 0.903 and an F1-score of 0.851, representing a new state of the art. Its computational efficiency enables real-time, offline inference directly on common mobile devices, making it ideal for low-resource settings. Importantly, the system produces clinically validated explanations that promote trust and adoption by frontline health workers. By coupling AI innovation with public-health requirements for speed, affordability, and reliability, DeepGB-TB offers a tool for advancing global TB control.


【12】Fast Algorithm for Moving Sound Source
标题:移动光源的快速算法
链接:https://arxiv.org/abs/2508.03065

作者:Dong Yang
摘要:现代基于神经网络的语音处理系统需要抗混响能力,依赖于大量的混响数据进行训练。现有的方法通过对静态系统进行采样或补充测量数据来模拟动态场景,但很难模拟符合物理规律的运动数据。针对运动场景下语音增强模型训练数据不足的问题,提出了Yang的运动时空采样重构理论,实现了对运动引起的连续时变混响的高效仿真。该方法突破了传统静态图像源方法(ISM)在时变系统中的局限性,将运动图像源的脉冲响应分解为线性时不变调制和离散时变分数延迟,建立了符合物理要求的运动声场模型。基于运动位移的带限特性,采用分层采样策略:对低阶图像采用高采样率以保留细节,对高阶图像采用低采样率以降低复杂度,并结合快速合成结构进行实时仿真。实验表明,与开源模型GSound相比,该理论更准确地还原了运动场景下的幅度和相位变化,解决了运动声源数据仿真的行业难题。它为语音增强模型提供了高质量的动态训练数据,提高了多通道端到端语音跟踪算法的鲁棒性。
摘要:Modern neural network-based speech processing systems need reverberation resistance, relying on large amounts of reverberation data for training. Existing methods simulate dynamic scenarios by sampling static systems or supplement with measured data, but struggle to simulate motion data conforming to physical laws. To address insufficient training data for speech enhancement models in moving scenarios, this paper proposes Yang's motion spatio-temporal sampling reconstruction theory, enabling efficient simulation of motion-induced continuous time-varying reverberation. It breaks through the limitations of traditional static Image-Source Method (ISM) in time-varying systems by decomposing the moving image source's impulse response into linear time-invariant modulation and discrete time-varying fractional delay, establishing a physics-compliant moving sound field model. Based on the band-limited nature of motion displacement, a hierarchical sampling strategy is adopted: high sampling rates for low-order images to retain details, and low rates for high-order ones to reduce complexity, combined with a fast synthesis architecture for real-time simulation. Experiments show that compared to open-source model GSound, the theory more accurately restores amplitude and phase changes in moving scenarios, solving the industry challenge of motion sound source data simulation. It provides high-quality dynamic training data for speech enhancement models and improves the robustness of multi-channel end-to-end voice tracking algorithms.


【13】SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
标题:SecoustiCodec:跨模式对齐流媒体单Codecbook语音编解码器
链接:https://arxiv.org/abs/2508.02849

作者:Chunyu Qiang, Haoyu Wang, Cheng Gong, Tianrui Wang, Ruibo Fu, Tao Wang, Ruilong Chen, Jiangyan Yi, Zhengqi Wen, Chen Zhang, Longbiao Wang, Jianwu Dang, Jianhua Tao
摘要:语音编解码器是统一语音和文本语言模型的关键桥梁。现有的编解码器方法在语义编码中面临若干挑战,诸如残留的非语言信息(例如,音色、情感)、语义完整性不足、重建能力有限以及缺乏对流媒体的支持。为了解决这些挑战,我们提出了SecoustiCodec,一个跨模态对齐的低比特率流语音编解码器,在一个单一的码本空间中解开语义和语义信息。为了保证语义的完整性和重建的逼真度,引入了语义编码来弥合语义编码和声学编码之间的信息鸿沟。提出了一种基于变分自编码器(VAE)和有限标量量化(FSQ)的仅限语义的高效量化方法。这种方法在保持高码本利用率的同时,消除了令牌的长尾分布问题。提出了一种基于对比学习的语义解纠缠方法,该方法将文本和语音在一个联合多模态的帧级空间中对齐,有效地去除了语义编码中的非语言信息。为了保证算法的鲁棒稳定收敛,提出了一种基于声学约束的多级优化策略。Figure~\ref{fig:pesq_kbps_below_2kbps}显示SecoustiCodec在0.27/1 kbps时实现了1.77/2.58的SOTA(最先进)重建质量(PESQ)。SecoustiCodec的代码和模型权重将在同行评审过程完成后开源。我们开源了SecoustiCodec的演示、代码和模型权重。
摘要:Speech codecs serve as a crucial bridge in unifying speech and text language models. Existing codec methods face several challenges in semantic encoding, such as residual paralinguistic information (e.g., timbre, emotion), insufficient semantic completeness, limited reconstruction capability, and lack of support for streaming. To address these challenges, we propose SecoustiCodec, a cross-modal aligned low-bitrate streaming speech codec that disentangles semantic and paralinguistic information in a single-codebook space. To ensure semantic completeness and reconstruction fidelity, paralinguistic encoding is introduced to bridge the information gap between semantic and acoustic encoding. A semantic-only efficient quantization method based on VAE (Variational Autoencoder) and FSQ (Finite Scalar Quantization) is proposed. This approach alleviates the long-tail distribution problem of tokens while maintaining high codebook utilization. A semantic disentanglement method based on contrastive learning is proposed, which aligns text and speech in a joint multimodal frame-level space, effectively removing paralinguistic information from semantic encoding. An acoustic-constrained multi-stage optimization strategy is proposed to ensure robust and stable convergence. Figure~\ref{fig:pesq_kbps_below_2kbps} shows SecoustiCodec achieves SOTA (state-of-the-art) reconstruction quality (PESQ) of 1.77/2.58 at 0.27/1 kbps. The code and model weights for SecoustiCodec will be open-sourced upon the completion of the peer-review process. We've open-sourced SecoustiCodec's demo, code, and model weights.


eess.AS音频处理


【1】PatchDSU: Uncertainty Modeling for Out of Distribution Generalization in Keyword Spotting
标题:PatchDSU:关键词发现中分布不确定性建模
链接:https://arxiv.org/abs/2508.03190

作者:Bronya Roni Chernyak, Yael Segal, Yosi Shrem, Joseph Keshet
备注:This work has been submitted to the IEEE for possible publication
摘要:深度学习模型擅长许多任务,但依赖于训练和测试数据遵循相同分布的假设。这种假设在现实世界的语音系统中往往不成立,因为在现实世界中,由于环境、录音条件和说话者多样性的变化,分布偏移是常见的。   不确定性域移位(DSU)方法根据输入特征统计来增加各层神经网络的输入。它通过假设特征统计遵循多元高斯分布并使用来自该分布的采样特征替代输入来解决域外泛化问题。虽然对计算机视觉有效,但由于数据的性质,将DSU应用于语音带来了挑战。与静态视觉数据不同,语音是一种时间信号,通常由频谱图表示-频率随时间的变化。这种表示不能被视为一个简单的图像,并且当应用于整个输入时,所产生的稀疏性可能导致倾斜的特征统计。   为了解决关键字定位中的分布问题,我们提出了PatchDSU,它通过将输入拆分为补丁并独立地增强每个补丁来扩展DSU。我们评估了PatchDSU和DSU以及Google Speech Commands,Librispeech和TED-LIUM上的其他方法。此外,我们评估了白高斯和MUSAN音乐噪声条件下的性能。我们还通过分析模型在未经训练的数据集上的性能来探索域外泛化。总的来说,在大多数情况下,PatchDSU和DSU都优于其他方法。值得注意的是,与其他方法相比,PatchDSU在评估的场景中表现出更一致的改进。
摘要:Deep learning models excel at many tasks but rely on the assumption that training and test data follow the same distribution. This assumption often does not hold in real-world speech systems, where distribution shifts are common due to varying environments, recording conditions, and speaker diversity.   The method of Domain Shifts with Uncertainty (DSU) augments the input of each neural network layer based on the input feature statistics. It addresses the problem of out-of-domain generalization by assuming feature statistics follow a multivariate Gaussian distribution and substitutes the input with sampled features from this distribution. While effective for computer vision, applying DSU to speech presents challenges due to the nature of the data. Unlike static visual data, speech is a temporal signal commonly represented by a spectrogram - the change of frequency over time. This representation cannot be treated as a simple image, and the resulting sparsity can lead to skewed feature statistics when applied to the entire input.   To tackle out-of-distribution issues in keyword spotting, we propose PatchDSU, which extends DSU by splitting the input into patches and independently augmenting each patch. We evaluated PatchDSU and DSU alongside other methods on the Google Speech Commands, Librispeech, and TED-LIUM. Additionally, we evaluated performance under white Gaussian and MUSAN music noise conditions. We also explored out-of-domain generalization by analyzing model performance on datasets they were not trained on. Overall, in most cases, both PatchDSU and DSU outperform other methods. Notably, PatchDSU demonstrates more consistent improvements across the evaluated scenarios compared to other approaches.


【2】Kernel ridge regression based sound field estimation using a rigid spherical microphone array
标题:使用刚性球形麦克风阵列基于核岭回归的声学场估计
链接:https://arxiv.org/abs/2508.03087

作者:Ryo Matsuda, Juliano G. C. Ribeiro, Hitoshi Akiyama, Jorge Trevino
备注:This paper has been accepted to the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:提出了一种基于核岭回归的刚性球形麦克风阵列声场估计方法。具有物理约束的核函数的核岭回归,以及进一步具有适于观察到的声场的核函数,已经被证明是强大的工具。然而,这样的方法通常假设开放球体麦克风阵列配置,即,在观测或估计区域内不存在散射体。或者,一些方法假设散射体的存在,并试图通过最小二乘公式来消除它们的影响。即使这样,这些方法通常不包括散射体的边界条件,这是不假定是已知的。相反,我们利用的事实,散射体在这里是一个刚性的球体。这意味着,虚拟散射源位置和边界条件都是明确定义的。在此基础上,我们制定了核岭回归框架内的散射声场,并提出了一种新的声场表示结合边界约束。通过数值仿真和实际实验证明了该方法的有效性,使用新开发的球形麦克风阵列。
摘要:We propose a sound field estimation method based on kernel ridge regression using a rigid spherical microphone array. Kernel ridge regression with physically constrained kernel functions, and further with kernel functions adapted to observed sound fields, have proven to be powerful tools. However, such methods generally assume an open-sphere microphone array configuration, i.e., no scatterers exist within the observation or estimation region. Alternatively, some approaches assume the presence of scatterers and attempt to eliminate their influence through a least-squares formulation. Even then, these methods typically do not incorporate the boundary conditions of the scatterers, which are not presumed to be known. In contrast, we exploit the fact the scatterer here is a rigid sphere. Meaning, both the virtual scattering source locations and the boundary conditions are well-defined. Based on this, we formulate the scattered sound field within the kernel ridge regression framework and propose a novel sound field representation incorporating a boundary constraint. The effectiveness of the proposed method is demonstrated through numerical simulations and real-world experiments using a newly developed spherical microphone array.


【3】Fast Algorithm for Moving Sound Source
标题:移动光源的快速算法
链接:https://arxiv.org/abs/2508.03065

作者:Dong Yang
摘要:现代基于神经网络的语音处理系统需要抗混响能力,依赖于大量的混响数据进行训练。现有的方法通过对静态系统进行采样或补充测量数据来模拟动态场景,但很难模拟符合物理规律的运动数据。针对运动场景下语音增强模型训练数据不足的问题,提出了Yang的运动时空采样重构理论,实现了对运动引起的连续时变混响的高效仿真。该方法突破了传统静态图像源方法(ISM)在时变系统中的局限性,将运动图像源的脉冲响应分解为线性时不变调制和离散时变分数延迟,建立了符合物理要求的运动声场模型。基于运动位移的带限特性,采用分层采样策略:对低阶图像采用高采样率以保留细节,对高阶图像采用低采样率以降低复杂度,并结合快速合成结构进行实时仿真。实验表明,与开源模型GSound相比,该理论更准确地还原了运动场景下的幅度和相位变化,解决了运动声源数据仿真的行业难题。它为语音增强模型提供了高质量的动态训练数据,提高了多通道端到端语音跟踪算法的鲁棒性。
摘要:Modern neural network-based speech processing systems need reverberation resistance, relying on large amounts of reverberation data for training. Existing methods simulate dynamic scenarios by sampling static systems or supplement with measured data, but struggle to simulate motion data conforming to physical laws. To address insufficient training data for speech enhancement models in moving scenarios, this paper proposes Yang's motion spatio-temporal sampling reconstruction theory, enabling efficient simulation of motion-induced continuous time-varying reverberation. It breaks through the limitations of traditional static Image-Source Method (ISM) in time-varying systems by decomposing the moving image source's impulse response into linear time-invariant modulation and discrete time-varying fractional delay, establishing a physics-compliant moving sound field model. Based on the band-limited nature of motion displacement, a hierarchical sampling strategy is adopted: high sampling rates for low-order images to retain details, and low rates for high-order ones to reduce complexity, combined with a fast synthesis architecture for real-time simulation. Experiments show that compared to open-source model GSound, the theory more accurately restores amplitude and phase changes in moving scenarios, solving the industry challenge of motion sound source data simulation. It provides high-quality dynamic training data for speech enhancement models and improves the robustness of multi-channel end-to-end voice tracking algorithms.


【4】Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model
标题:使用神经音频编解码器作为基础模型的喉咙麦克风在噪音中的实时语音增强
链接:https://arxiv.org/abs/2508.02974

作者:Julien Hauret, Thomas Joubaud, Éric Bavu
备注:2 pages, 2 figures
摘要:我们提出了一个实时的语音增强演示使用语音捕获的喉咙麦克风。该演示旨在展示完整的管道,从录音到基于深度学习的后处理,用于在嘈杂环境中使用身体传导麦克风捕获的语音。喉咙麦克风记录皮肤振动,从而自然衰减外部噪音,但这种鲁棒性是以降低音频带宽为代价的。为了应对这一挑战,我们在Vibravox上微调了Kyutai的Mimi--一种支持实时推理的神经音频编解码器,Vibravox是一个包含成对空气传导和喉咙麦克风录音的数据集。我们将这种增强策略与最先进的模型进行比较,并展示了其优越的性能。推理在交互式界面中运行,允许用户切换增强,可视化频谱图和监控处理延迟。
摘要:We present a real-time speech enhancement demo using speech captured with a throat microphone. This demo aims to showcase the complete pipeline, from recording to deep learning-based post-processing, for speech captured in noisy environments with a body-conducted microphone. The throat microphone records skin vibrations, which naturally attenuate external noise, but this robustness comes at the cost of reduced audio bandwidth. To address this challenge, we fine-tune Kyutai's Mimi--a neural audio codec supporting real-time inference--on Vibravox, a dataset containing paired air-conducted and throat microphone recordings. We compare this enhancement strategy against state-of-the-art models and demonstrate its superior performance. The inference runs in an interactive interface that allows users to toggle enhancement, visualize spectrograms, and monitor processing latency.


【5】SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
标题:SecoustiCodec:跨模式对齐流媒体单Codecbook语音编解码器
链接:https://arxiv.org/abs/2508.02849

作者:Chunyu Qiang, Haoyu Wang, Cheng Gong, Tianrui Wang, Ruibo Fu, Tao Wang, Ruilong Chen, Jiangyan Yi, Zhengqi Wen, Chen Zhang, Longbiao Wang, Jianwu Dang, Jianhua Tao
摘要:语音编解码器是统一语音和文本语言模型的关键桥梁。现有的编解码器方法在语义编码中面临若干挑战,诸如残留的非语言信息(例如,音色、情感)、语义完整性不足、重建能力有限以及缺乏对流式传输的支持。为了解决这些挑战,我们提出了SecoustiCodec,一个跨模态对齐的低比特率流语音编解码器,在一个单一的码本空间中解开语义和语义信息。为了保证语义的完整性和重建的逼真度,引入了语义编码来弥合语义编码和声学编码之间的信息鸿沟。提出了一种基于变分自编码器(VAE)和有限标量量化(FSQ)的仅限语义的高效量化方法。这种方法在保持高码本利用率的同时,消除了令牌的长尾分布问题。提出了一种基于对比学习的语义解纠缠方法,该方法将文本和语音在一个联合多模态的帧级空间中对齐,有效地去除了语义编码中的非语言信息。为了保证算法的鲁棒稳定收敛,提出了一种基于声学约束的多级优化策略。Figure~\ref{fig:pesq_kbps_below_2kbps}显示SecoustiCodec在0.27/1 kbps时实现了1.77/2.58的SOTA(最先进)重建质量(PESQ)。SecoustiCodec的代码和模型权重将在同行评审过程完成后开源。我们开源了SecoustiCodec的演示、代码和模型权重。
摘要:Speech codecs serve as a crucial bridge in unifying speech and text language models. Existing codec methods face several challenges in semantic encoding, such as residual paralinguistic information (e.g., timbre, emotion), insufficient semantic completeness, limited reconstruction capability, and lack of support for streaming. To address these challenges, we propose SecoustiCodec, a cross-modal aligned low-bitrate streaming speech codec that disentangles semantic and paralinguistic information in a single-codebook space. To ensure semantic completeness and reconstruction fidelity, paralinguistic encoding is introduced to bridge the information gap between semantic and acoustic encoding. A semantic-only efficient quantization method based on VAE (Variational Autoencoder) and FSQ (Finite Scalar Quantization) is proposed. This approach alleviates the long-tail distribution problem of tokens while maintaining high codebook utilization. A semantic disentanglement method based on contrastive learning is proposed, which aligns text and speech in a joint multimodal frame-level space, effectively removing paralinguistic information from semantic encoding. An acoustic-constrained multi-stage optimization strategy is proposed to ensure robust and stable convergence. Figure~\ref{fig:pesq_kbps_below_2kbps} shows SecoustiCodec achieves SOTA (state-of-the-art) reconstruction quality (PESQ) of 1.77/2.58 at 0.27/1 kbps. The code and model weights for SecoustiCodec will be open-sourced upon the completion of the peer-review process. We've open-sourced SecoustiCodec's demo, code, and model weights.


【6】SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
标题:SonicMaster:迈向可控的一体化音乐修复和掌握
链接:https://arxiv.org/abs/2508.03448

作者:jan melechovsky, Ambuj Mehrish, Dorien Herremans
摘要:音乐录音通常会遇到音频质量问题,例如过度混响、失真、削波、音调不平衡和狭窄的立体声图像,特别是在没有专业设备或专业知识的非专业环境中创建时。这些问题通常使用单独的专用工具和手动调整来纠正。在本文中,我们介绍了SonicMaster,这是第一个用于音乐恢复和母带的统一生成模型,它通过基于文本的控制解决了广泛的音频伪影。SonicMaster以自然语言指令为条件来应用有针对性的增强功能,或者可以在自动模式下运行以进行一般恢复。为了训练这个模型,我们构建了SonicMaster数据集,这是一个大型的成对降级和高质量音轨数据集,通过模拟常见的降级类型,其中19个降级函数属于5个增强组:均衡,动态,混响,幅度和立体声。我们的方法利用流匹配生成训练范式来学习音频转换,该转换将降级的输入映射到由文本提示引导的清洁的、掌握的版本。客观的音频质量指标表明,SonicMaster显着提高了所有工件类别的声音质量。此外,主观听力测试证实,听众更喜欢SonicMaster的增强输出,而不是原来的降级音频,突出了我们的统一方法的有效性。
摘要:Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment or expertise. These problems are typically corrected using separate specialized tools and manual adjustments. In this paper, we introduce SonicMaster, the first unified generative model for music restoration and mastering that addresses a broad spectrum of audio artifacts with text-based control. SonicMaster is conditioned on natural language instructions to apply targeted enhancements, or can operate in an automatic mode for general restoration. To train this model, we construct the SonicMaster dataset, a large dataset of paired degraded and high-quality tracks by simulating common degradation types with nineteen degradation functions belonging to five enhancements groups: equalization, dynamics, reverb, amplitude, and stereo. Our approach leverages a flow-matching generative training paradigm to learn an audio transformation that maps degraded inputs to their cleaned, mastered versions guided by text prompts. Objective audio quality metrics demonstrate that SonicMaster significantly improves sound quality across all artifact categories. Furthermore, subjective listening tests confirm that listeners prefer SonicMaster's enhanced outputs over the original degraded audio, highlighting the effectiveness of our unified approach.


【7】When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
标题:当好的声音变得敌对时:用良性输入越狱的音频模型
链接:https://arxiv.org/abs/2508.03365

作者:Bodam Kim, Hiskias Dingeto, Taeyoun Kwon, Dasol Choi, DongGeon Lee, Haon Park, JaeHoon Lee, Jongho Shin
摘要:随着大型语言模型越来越多地融入日常生活,音频已成为人机交互的关键接口。然而,这种便利性也引入了新的漏洞,使音频成为对手的潜在攻击面。我们的研究引入了WhisperInject,这是一个两阶段的对抗性音频攻击框架,可以操纵最先进的音频语言模型来生成有害内容。我们的方法在音频输入中使用对人类听众保持良性的不可察觉的扰动。第一阶段使用一种新的基于奖励的优化方法,投影梯度下降强化学习(RL-PGD),引导目标模型规避其自身的安全协议并生成有害的本地响应。然后,这种原生有害响应将作为阶段2有效载荷注入的目标,在该阶段中,我们使用投影梯度下降(PGD)来优化嵌入到良性音频载体中的细微扰动,例如天气查询或问候消息。在严格的StrongRESIDENT、LlamaGuard以及人体评估安全评估框架下进行验证,我们的实验表明,Qwen2.5-Omni-3B、Qwen2.5-Omni-7 B和Phi-4-Multimodal的成功率超过86%。我们的工作展示了一类新的实际的音频原生威胁,超越了理论上的利用,揭示了一种操纵人工智能行为的可行和隐蔽的方法。
摘要:As large language models become increasingly integrated into daily life, audio has emerged as a key interface for human-AI interaction. However, this convenience also introduces new vulnerabilities, making audio a potential attack surface for adversaries. Our research introduces WhisperInject, a two-stage adversarial audio attack framework that can manipulate state-of-the-art audio language models to generate harmful content. Our method uses imperceptible perturbations in audio inputs that remain benign to human listeners. The first stage uses a novel reward-based optimization method, Reinforcement Learning with Projected Gradient Descent (RL-PGD), to guide the target model to circumvent its own safety protocols and generate harmful native responses. This native harmful response then serves as the target for Stage 2, Payload Injection, where we use Projected Gradient Descent (PGD) to optimize subtle perturbations that are embedded into benign audio carriers, such as weather queries or greeting messages. Validated under the rigorous StrongREJECT, LlamaGuard, as well as Human Evaluation safety evaluation framework, our experiments demonstrate a success rate exceeding 86% across Qwen2.5-Omni-3B, Qwen2.5-Omni-7B, and Phi-4-Multimodal. Our work demonstrates a new class of practical, audio-native threats, moving beyond theoretical exploits to reveal a feasible and covert method for manipulating AI behavior.


【8】MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
标题:MiSTR:基于变换器的韵律预测和神经相位重建的多模态iEEG到语音合成
链接:https://arxiv.org/abs/2508.03166

作者:Mohammed Salah Al-Radhi, Géza Németh, Branislav Gerazov
备注:5 pages, 2 figures, 1 table. Accepted for presentation at Interspeech 2025
摘要:从颅内脑电(iEEG)信号的语音合成提供了一个有前途的途径,恢复严重的语言障碍的个人的沟通。然而,由于特征表示、韵律建模和相位重建的限制,实现可理解和自然的语音仍然具有挑战性。我们介绍了MiSTR,这是一个深度学习框架,它集成了:1)基于小波的特征提取,以捕获iEEG信号的细粒度时间,频谱和神经生理表示,2)基于transformer的解码器,用于韵律感知频谱图预测,以及3)神经相位声码器,通过自适应频谱校正实现谐波一致性。在公共iEEG数据集上进行评估,MiSTR实现了最先进的语音清晰度,重建和原始Mel频谱图之间的平均Pearson相关性为0.91,优于现有的神经语音合成基线。
摘要:Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challenging due to limitations in feature representation, prosody modeling, and phase reconstruction. We introduce MiSTR, a deep-learning framework that integrates: 1) Wavelet-based feature extraction to capture fine-grained temporal, spectral, and neurophysiological representations of iEEG signals, 2) A Transformer-based decoder for prosody-aware spectrogram prediction, and 3) A neural phase vocoder enforcing harmonic consistency via adaptive spectral correction. Evaluated on a public iEEG dataset, MiSTR achieves state-of-the-art speech intelligibility, with a mean Pearson correlation of 0.91 between reconstructed and original Mel spectrograms, improving over existing neural speech synthesis baselines.


【9】Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
标题:使用人类反馈的强化学习微调文本到语音扩散模型
链接:https://arxiv.org/abs/2508.03123

作者:Jingyi Chen, Ju Seung Byun, Micha Elsner, Pichao Wang, Andrew Perrault
备注:4 pages, 1 figure, INTERSPEECH 2025. arXiv admin note: text overlap   with arXiv:2405.14632
摘要:扩散模型产生高保真语音,但由于长的去噪步骤和建模语调和节奏的挑战,对于实时使用是低效的。为了改善这一点,我们提出了扩散损失导向的政策优化(DLPO),一个RLHF框架的TTS扩散模型。DLPO将原始训练损失集成到奖励函数中,保留生成能力,同时减少效率低下。使用自然度分数作为反馈,DLPO将奖励优化与扩散模型的结构相结合,从而提高语音质量。我们在WaveGrad 2(一种基于非自回归扩散的TTS模型)上评估DLPO。结果显示,在客观指标(UTMOS 3.65,NISQA 4.02)和主观评价,DLPO音频首选67%的时间显着改善。这些发现表明DLPO在实时、资源有限的环境中具有高效、高质量扩散TTS的潜力。
摘要:Diffusion models produce high-fidelity speech but are inefficient for real-time use due to long denoising steps and challenges in modeling intonation and rhythm. To improve this, we propose Diffusion Loss-Guided Policy Optimization (DLPO), an RLHF framework for TTS diffusion models. DLPO integrates the original training loss into the reward function, preserving generative capabilities while reducing inefficiencies. Using naturalness scores as feedback, DLPO aligns reward optimization with the diffusion model's structure, improving speech quality. We evaluate DLPO on WaveGrad 2, a non-autoregressive diffusion-based TTS model. Results show significant improvements in objective metrics (UTMOS 3.65, NISQA 4.02) and subjective evaluations, with DLPO audio preferred 67\% of the time. These findings demonstrate DLPO's potential for efficient, high-quality diffusion TTS in real-time, resource-limited settings.


【10】TF-MLPNet: Tiny Real-Time Neural Speech Separation
标题:TF-MLPNet:微小的实时神经语音分离
链接:https://arxiv.org/abs/2508.03047

作者:Malek Itani, Tuochao Chen, Shyamnath Gollakota
备注:The 6th Clarity Workshop on Improving Speech-in-Noise for Hearing Devices (Clarity 2025)
摘要:听觉设备上的语音分离可以实现变革性的增强和增强的听觉能力。然而,最先进的语音分离网络无法在为听觉设计的小型低功耗神经加速器上实时运行,因为它们的计算能力有限。我们提出了TF-MLPNet,这是第一个能够在这种低功耗加速器上实时运行的语音分离网络,同时优于现有的盲语音分离和目标语音提取流模型。我们的网络在时频域中运行,使用沿信道和频率维度交替的全连接层堆栈处理频率序列,并使用卷积层独立处理每个频率点的时间序列。结果表明,我们的混合精度量化感知训练(QAT)模型可以在GAP 9处理器上实时处理6 ms的音频块,与之前的语音分离模型相比,运行时间减少了3.5- 4倍。
摘要:Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelerators designed for hearables, due to their limited compute capabilities. We present TF-MLPNet, the first speech separation network capable of running in real-time on such low-power accelerators while outperforming existing streaming models for blind speech separation and target speech extraction. Our network operates in the time-frequency domain, processing frequency sequences with stacks of fully connected layers that alternate along the channel and frequency dimensions, and independently processing the time sequence at each frequency bin using convolutional layers. Results show that our mixed-precision quantization-aware trained (QAT) model can process 6 ms audio chunks in real-time on the GAP9 processor, achieving a 3.5-4x runtime reduction compared to prior speech separation models.


【11】Neural Speech Extraction with Human Feedback
标题:利用人类反馈的神经语音提取
链接:https://arxiv.org/abs/2508.03041

作者:Malek Itani, Ashton Graves, Sefik Emre Eskimez, Shyamnath Gollakota
备注:Interspeech 2025
摘要:我们提出了第一个神经目标语音提取(TSE)系统,使用人的反馈迭代细化。我们的方法允许用户标记TSE输出的特定部分,生成编辑掩码。然后,细化系统改进标记的部分,同时保留未标记的区域。由于人类标记的错误的大规模数据集很难收集,我们使用各种自动掩蔽函数生成合成数据集,并对每个数据集进行训练。评估表明,使用基于噪声功率的掩蔽(以dBFS为单位)和概率阈值训练的模型表现最好,与人类注释保持一致。在一项有22名参与者的研究中,用户显示出对精炼输出的偏好超过基线TSE。我们的研究结果表明,人在回路细化是一个很有前途的方法,提高神经语音提取的性能。
摘要:We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then improves the marked sections while preserving unmarked regions. Since large-scale datasets of human-marked errors are difficult to collect, we generate synthetic datasets using various automated masking functions and train models on each. Evaluations show that models trained with noise power-based masking (in dBFS) and probabilistic thresholding perform best, aligning with human annotations. In a study with 22 participants, users showed a preference for refined outputs over baseline TSE. Our findings demonstrate that human-in-the-loop refinement is a promising approach for improving the performance of neural speech extraction.


【12】How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
标题:听起来怎么样?室内场景的材料控制多峰声学轮廓生成
链接:https://arxiv.org/abs/2508.02905

作者:Mahnoor Fatima Saad, Ziad Al-Halah
备注:Accepted to ICCV 2025. Project Page: this https URL
摘要:在一个铺着地毯的地板和墙上贴着隔音砖的录音室里,声音会有什么变化?我们介绍了材料控制的声学配置文件生成的任务,其中,给定一个室内场景具有特定的视听特性,目标是在推理时间根据用户定义的材料配置生成目标声学配置文件。我们解决这个任务与一种新的编码器-解码器的方法,编码场景的关键属性从视听观察和生成目标房间脉冲响应(RIR)的条件下,由用户提供的材料规格。我们的模型能够生成不同的RIR的基础上动态定义的各种材料配置在推理时间。为了支持这一任务,我们创建了一个新的基准,声学仙境数据集,旨在开发和评估材料感知RIR预测方法在不同的和具有挑战性的设置。我们的研究结果表明,该模型有效地编码材料信息,并产生高保真RIR,优于几个基线和国家的最先进的方法。
摘要:How would the sound in a studio change with a carpeted floor and acoustic tiles on the walls? We introduce the task of material-controlled acoustic profile generation, where, given an indoor scene with specific audio-visual characteristics, the goal is to generate a target acoustic profile based on a user-defined material configuration at inference time. We address this task with a novel encoder-decoder approach that encodes the scene's key properties from an audio-visual observation and generates the target Room Impulse Response (RIR) conditioned on the material specifications provided by the user. Our model enables the generation of diverse RIRs based on various material configurations defined dynamically at inference time. To support this task, we create a new benchmark, the Acoustic Wonderland Dataset, designed for developing and evaluating material-aware RIR prediction methods under diverse and challenging settings. Our results demonstrate that the proposed model effectively encodes material information and generates high-fidelity RIRs, outperforming several baselines and state-of-the-art methods.


【13】Adaptive Knowledge Distillation for Device-Directed Speech Detection
标题:设备指导语音检测的自适应知识提炼
链接:https://arxiv.org/abs/2508.02801

作者: Hyung Gun Chi, Florian Pesce, Wonil Chang, Oggi Rudovic, Arturo Argueta, Stefan Braun, Vineet Garg, Ahmed Hussen Abdelaziz
备注:5 pages, 2 figures, Interspeech accepted
摘要:设备导向语音检测(DDSD)是一种二进制分类任务,它将用户对语音助理(VA)的查询与背景语音或侧面对话分开。这对于实现自然的用户体验非常重要。为此,我们提出了知识蒸馏(KD),以提高DDSD的准确性,同时确保有效的部署。具体来说,我们介绍了一种新的自适应KD方法,该方法从ASR大型预训练声学编码器(教师)的一般表示中传输知识。我们应用特定于任务的适配器,在(冻结)教师编码器的顶部,与DDSD上的学生模型一起训练。我们证明了所提出的自适应KD在关键字和无关键字(后续)调用方面优于没有蒸馏的学生模型,在相等错误率方面分别提高了+26%和+19%。我们还表明,这种方法概括了整个Transformer和一致性为基础的模型架构。
摘要:Device-directed speech detection (DDSD) is a binary classification task that separates the user's queries to a voice assistant (VA) from background speech or side conversations. This is important for achieving naturalistic user experience. To this end, we propose knowledge distillation (KD) to enhance DDSD accuracy while ensuring efficient deployment. Specifically, we introduce a novel adaptive KD method that transfers knowledge from general representations of an ASR large pre-trained acoustic encoder (teacher). We apply task-specific adapters, on top of the (frozen) teacher encoder, trained jointly with the student model on DDSD. We demonstrate that the proposed adaptive KD outperforms the student model without distillation in the keyword and keyword-free (follow-up) invocations, with an improvement of +26% and +19% in terms of Equal Error Rate, respectively. We also show that this approach generalizes across the transformer and conformer-based model architectures.


机器翻译由腾讯交互翻译提供,仅供参考