今日论文合集:cs.SD语音15篇,eess.AS音频处理20篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Sounding that Object: Interactive Object-Aware Image to Audio Generation
标题: 探测该对象:交互式对象感知图像到音频生成
链接:https://arxiv.org/abs/2506.04214
作者: Tingle Li,  Baihe Huang,  Xiaobin Zhuang,  Dongya Jia,  Jiawei Chen,  Yuping Wang,  Zhuo Chen,  Gopala Anumanchipalli,  Yuxuan Wang 
备注:ICML 2025
摘要:为复杂的视听场景生成准确的声音是一项挑战,特别是在存在多个对象和声源的情况下。在本文中,我们提出了一个{\em交互式对象感知音频生成}模型,该模型将声音生成建立在用户选择的图像中的视觉对象上。我们的方法将以对象为中心的学习集成到一个条件潜在扩散模型中,该模型通过多模态注意力学习将图像区域与其相应的声音相关联。在测试时,我们的模型采用图像分割,让用户交互式地生成声音在{\em对象}的水平。我们从理论上验证了我们的注意力机制在功能上近似于测试时的分割掩码,确保生成的音频与选定的对象对齐。定量和定性评估表明,我们的模型优于基线,实现了对象及其相关声音之间的更好对齐。项目页面:https://tinglok.netlify.app/files/avobject/
摘要:Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds sound generation in user-selected visual objects within images. Our method integrates object-centric learning into a conditional latent diffusion model, which learns to associate image regions with their corresponding sounds through multi-modal attention. At test time, our model employs image segmentation to allow users to interactively generate sounds at the {\em object} level. We theoretically validate that our attention mechanism functionally approximates test-time segmentation masks, ensuring the generated audio aligns with selected objects. Quantitative and qualitative evaluations show that our model outperforms baselines, achieving better alignment between objects and their associated sounds. Project page: https://tinglok.netlify.app/files/avobject/


【2】 UniCUE: Unified Recognition and Generation Framework for Chinese Cued  Speech Video-to-Speech Generation

标题: UniCUE:中文提示语音视频到语音生成的统一识别和生成框架
链接:https://arxiv.org/abs/2506.04134
作者: Jinting Wang,  Shan Yang,  Li Liu 
备注:10 pages, 10 figures
摘要:提示语音(CS)通过手编码增强唇读,为听障者提供精确的语音感知支持。CS视频到语音生成(CSV 2S)任务旨在将听障者的CS视觉表达(CS视频)转换为可理解的语音信号。从CS视频直接生成语音(称为单CSV 2S)由于CS数据不足而产生差的性能。目前的研究主要集中在CS识别(CSR),它将视频内容转换为语言文本。基于此,CSV 2S的一个直接方法是将CSR与文本到语音系统相结合。这种组合架构依赖于文本作为中间媒介逐步跨模态对齐,这可能会导致错误传播和语音和视频动态之间的时间错位。为了解决这些挑战,我们提出了一种新的方法,直接从CS视频生成语音,而不依赖于中间文本。在此基础上,我们提出了UniCUE,第一个统一的框架CSV 2S,其核心创新在于整合的CSR任务,提供细粒度的视觉语义信息,以促进语音生成从CS视频。更确切地说,(1)一种新的细粒度语义对齐池,以确保视觉特征和语音内容之间的精确映射;(2)一个VisioPhonetic适配器,以桥接跨任务表示,确保两个不同任务之间的无缝兼容性(即,CSV 2S和CSR);(3)提出了一种基于姿态感知的视觉处理器,以增强CS视频中唇和手运动之间的细粒度时空相关性。在我们新建立的中文CS数据集(14 cuers 1:8 hearing-impaired and 6 norm-hearing)上的实验表明,与单一的CSV 2S相比,我们的UniCUE显著降低了78.3%的单词错误率,并提高了32%的唇语同步。
摘要:Cued Speech (CS) enhances lipreading through hand coding, providing precise speech perception support for the hearing-impaired. CS Video-to-Speech generation (CSV2S) task aims to convert the CS visual expressions (CS videos) of hearing-impaired individuals into comprehensible speech signals. Direct generation of speech from CS video (called single CSV2S) yields poor performance due to insufficient CS data. Current research mostly focuses on CS Recognition (CSR), which convert video content into linguistic text. Based on this, one straightforward way of CSV2S is to combine CSR with a Text-to-Speech system. This combined architecture relies on text as an intermediate medium for stepwise cross-modal alignment, which may lead to error propagation and temporal misalignment between speech and video dynamics. To address these challenges, we propose a novel approach that directly generates speech from CS videos without relying on intermediate text. Building upon this, we propose UniCUE, the first unified framework for CSV2S, whose core innovation lies in the integration of the CSR task that provides fine-grained visual-semantic information to facilitate speech generation from CS videos. More precisely, (1) a novel fine-grained semantic alignment pool to ensure precise mapping between visual features and speech contents; (2) a VisioPhonetic adapter to bridge cross-task representations, ensuring seamless compatibility between two distinct tasks (i.e., CSV2S and CSR); (3) a pose-aware visual processor is introduced to enhance fine-grained spatiotemporal correlations between lip and hand movements in CS video. Experiments on our new established Chinese CS dataset (14 cuers1: 8 hearing-impaired and 6 normal-hearing) show that our UniCUE significantly reduces Word Error Rate by 78.3% and improves lip-speech synchronization by 32% compared to the single CSV2S.


【3】 A Novel Data Augmentation Approach for Automatic Speaking Assessment on  Opinion Expressions

标题: 一种新的数据增强方法,用于对意见表达进行自动口语评估
链接:https://arxiv.org/abs/2506.04077
作者: Chung-Chun Wang,  Jhen-Ke Lin,  Hao-Chien Lu,  Hong-Yun Lin,  Berlin Chen 
备注:submitted to the ISCA SLaTE-2025 Workshop
摘要:意见表达的自动口语评估(ASA)通常受到标记录音稀缺的阻碍,这限制了即时多样性并破坏了评分可靠性。为了解决这一挑战,我们提出了一种新的训练范式,利用大型语言模型(LLM)来生成给定熟练程度的各种响应,通过说话者感知的文本到语音合成将响应转换为合成语音,并采用动态重要性损失来基于合成语音和真实语音之间的特征分布差异自适应地重新加权训练实例。随后,多模态大语言模型集成对齐的文本特征与语音信号直接预测熟练度分数。在LTTC数据集上进行的实验表明,我们的方法优于依赖于真实数据或传统增强的方法,有效地缓解了低资源约束,并使ASA能够在具有跨模态信息的意见表达上。
摘要:Automated speaking assessment (ASA) on opinion expressions is often hampered by the scarcity of labeled recordings, which restricts prompt diversity and undermines scoring reliability. To address this challenge, we propose a novel training paradigm that leverages a large language models (LLM) to generate diverse responses of a given proficiency level, converts responses into synthesized speech via speaker-aware text-to-speech synthesis, and employs a dynamic importance loss to adaptively reweight training instances based on feature distribution differences between synthesized and real speech. Subsequently, a multimodal large language model integrates aligned textual features with speech signals to predict proficiency scores directly. Experiments conducted on the LTTC dataset show that our approach outperforms methods relying on real data or conventional augmentation, effectively mitigating low-resource constraints and enabling ASA on opinion expressions with cross-modal information.


【4】 Acoustically Precise Hesitation Tagging Is Essential for End-to-End  Verbatim Transcription Systems

标题: 声学精确的犹豫标记对于端到端逐字转录系统至关重要
链接:https://arxiv.org/abs/2506.04076
作者: Jhen-Ke Lin,  Hao-Chien Lu,  Chung-Chun Wang,  Hong-Yun Lin,  Berlin Chen 
备注:submitted to the ISCA SLaTE-2025 Workshop
摘要:用于自动口语评估的逐字记录需要准确捕捉不流利的内容,这对于错误分析和反馈等下游任务至关重要。然而,许多ASR系统丢弃或概括犹豫,丢失重要的声学细节。我们使用低秩自适应(LoRA)在Speak & Improve 2025语料库上微调Whisper模型,而无需依赖外部音频训练数据。我们比较了三种注释方案:删除犹豫(纯),通用标签(丰富),声学精确的填料推断双子座2.0闪光灯从现有的音频转录对(额外)。我们的挑战系统实现了6.47% WER(纯)和5.81% WER(额外)。挑战后的实验表明,微调耳语大V3涡轮与“额外”计划产生了5.5%的WER,11.3%的相对改善“纯”计划(6.2% WER)。这表明,明确的,现实的填充停顿标记显着提高ASR的准确性逐字L2语音转录。
摘要:Verbatim transcription for automatic speaking assessment demands accurate capture of disfluencies, crucial for downstream tasks like error analysis and feedback. However, many ASR systems discard or generalize hesitations, losing important acoustic details. We fine-tune Whisper models on the Speak & Improve 2025 corpus using low-rank adaptation (LoRA), without recourse to external audio training data. We compare three annotation schemes: removing hesitations (Pure), generic tags (Rich), and acoustically precise fillers inferred by Gemini 2.0 Flash from existing audio-transcript pairs (Extra). Our challenge system achieved 6.47% WER (Pure) and 5.81% WER (Extra). Post-challenge experiments reveal that fine-tuning Whisper Large V3 Turbo with the "Extra" scheme yielded a 5.5% WER, an 11.3% relative improvement over the "Pure" scheme (6.2% WER). This demonstrates that explicit, realistic filled-pause labeling significantly enhances ASR accuracy for verbatim L2 speech transcription.


【5】 A Statistics-Driven Differentiable Approach for Sound Texture Synthesis  and Analysis

标题: 一种统计驱动的声音纹理合成与分析的可区分方法
链接:https://arxiv.org/abs/2506.04073
作者: Esteban Gutiérrez,  Frederic Font,  Xavier Serra,  Lonce Wyse 
备注:Accepted to the 28th International Conference on Digital Audio Effects (DAFx 2025) to be held in Ancona, Italy. 8 pages, one diagram and 5 tables
摘要:在这项工作中,我们介绍了TexStat,一种新的损失函数,专门设计用于分析和合成纹理的声音,其特征在于随机结构和感知平稳性。从McDermott和Simoncelli的统计和感知框架中汲取灵感,TexStat识别属于同一纹理类别的信号之间的相似性,而不依赖于时间结构。我们还建议使用TexStat作为验证度量一起Frechet音频距离(FAD)来评估纹理声音合成模型。除了TexStat,我们提出了TexEnv,一个有效的,轻量级的和可区分的纹理声音合成器,通过对过滤后的噪声施加幅度包络来生成音频。我们进一步将这些组件集成到TexDSP中,TexDSP是一种为纹理声音量身定制的DDSP启发的生成模型。通过对各种纹理声音类型的广泛实验,我们证明了TexStat在感知上是有意义的,时不变的,并且对噪声具有鲁棒性,这些特征使其既可以作为生成任务的损失函数,也可以作为验证指标。所有工具和代码都是作为开源贡献提供的,我们的PyTorch实现是高效的,可区分的,高度可配置的,使其能够在生成任务中使用,并作为感知基础的评估指标。
摘要:In this work, we introduce TexStat, a novel loss function specifically designed for the analysis and synthesis of texture sounds characterized by stochastic structure and perceptual stationarity. Drawing inspiration from the statistical and perceptual framework of McDermott and Simoncelli, TexStat identifies similarities between signals belonging to the same texture category without relying on temporal structure. We also propose using TexStat as a validation metric alongside Frechet Audio Distances (FAD) to evaluate texture sound synthesis models. In addition to TexStat, we present TexEnv, an efficient, lightweight and differentiable texture sound synthesizer that generates audio by imposing amplitude envelopes on filtered noise. We further integrate these components into TexDSP, a DDSP-inspired generative model tailored for texture sounds. Through extensive experiments across various texture sound types, we demonstrate that TexStat is perceptually meaningful, time-invariant, and robust to noise, features that make it effective both as a loss function for generative tasks and as a validation metric. All tools and code are provided as open-source contributions and our PyTorch implementations are efficient, differentiable, and highly configurable, enabling its use in both generative tasks and as a perceptually grounded evaluation metric.


【6】 Towards Better Disentanglement in Non-Autoregressive Zero-Shot  Expressive Voice Conversion

标题: 在非自回归Zero-Shot表达性语音转换中实现更好的解纠缠
链接:https://arxiv.org/abs/2506.04013
作者: Seymanur Akti,  Tuan Nam Nguyen,  Alexander Waibel 
备注:Accepted to Interspeech 2025
摘要:表达性语音转换的目的是将说话人身份和表达属性从目标语音转换为给定的源语音。在这项工作中,我们改进了一个具有条件变分自编码器的自监督非自回归框架,重点是减少源音色泄漏和改善语言声学解纠缠,以实现更好的风格传递。为了最大限度地减少风格泄漏,我们使用多语言离散语音单元进行内容表示,并通过基于增强的相似性损失和混合风格层规范化来加强嵌入。为了增强表现力转移,我们通过交叉注意将局部F0信息结合起来,并提取富含全局音高和能量特征的风格嵌入。实验表明,我们的模型在情感和说话人相似性方面优于基线,表现出优越的风格适应性和减少源风格泄漏。
摘要:Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework with a conditional variational autoencoder, focusing on reducing source timbre leakage and improving linguistic-acoustic disentanglement for better style transfer. To minimize style leakage, we use multilingual discrete speech units for content representation and reinforce embeddings with augmentation-based similarity loss and mix-style layer normalization. To enhance expressivity transfer, we incorporate local F0 information via cross-attention and extract style embeddings enriched with global pitch and energy features. Experiments show our model outperforms baselines in emotion and speaker similarity, demonstrating superior style adaptation and reduced source style leakage.


【7】 From Spikes to Speech: NeuroVoc -- A Biologically Plausible Vocoder  Framework for Auditory Perception and Cochlear Implant Simulation

标题: 从尖峰到语音:NeuroVoc --一种用于听觉感知和角膜植入物模拟的生物学上合理的声码器框架
链接:https://arxiv.org/abs/2506.03959
作者: Jacob de Nobel,  Jeroen J. Briaire,  Thomas H.W. Baeck,  Anna V. Kononova,  Johan H.M. Frijns 
备注:43 Pages, 11 Figures, 2 Tables
摘要:我们提出了NeuroVoc,一个灵活的模型无关的声码器框架,使用逆傅立叶变换从模拟的神经活动模式重建声波波形。该系统将直接的信号处理应用于神经图表示,即来自听觉神经纤维模型的时频分箱输出。至关重要的是,模型架构是模块化的,允许轻松替换或修改底层听觉模型。这种灵活性消除了在模拟人工耳蜗(CI)用户的听觉感知时对语音编码策略特定的声码器实现的需要。它还允许直接比较正常听力(NH)和电听力(EH)模型,如本研究所示。声码器保留每个模型的独特特征;例如,NH模型比EH模型更忠实地保留谐波结构。我们使用在线数字噪声(DIN)测试评估了噪声中的感知可懂度,参与者完成了三个测试条件:一个是标准语音,两个是使用NH和EH模型的语音编码语音。标准DIN测试和EH声码组在统计学上等同于NH和CI听众的临床报告数据。平均而言,NH和EH声码组与标准测试相比,SRT分别增加了2.4 dB和7.1 dB。这些研究结果表明,虽然发生了一些退化,声码器可以重建清晰的语音在两个听力模型,并准确地反映了降低语音噪声性能CI用户所经历的。
摘要:We present NeuroVoc, a flexible model-agnostic vocoder framework that reconstructs acoustic waveforms from simulated neural activity patterns using an inverse Fourier transform. The system applies straightforward signal processing to neurogram representations, time-frequency binned outputs from auditory nerve fiber models. Crucially, the model architecture is modular, allowing for easy substitution or modification of the underlying auditory models. This flexibility eliminates the need for speech-coding-strategy-specific vocoder implementations when simulating auditory perception in cochlear implant (CI) users. It also allows direct comparisons between normal hearing (NH) and electrical hearing (EH) models, as demonstrated in this study. The vocoder preserves distinctive features of each model; for example, the NH model retains harmonic structure more faithfully than the EH model. We evaluated perceptual intelligibility in noise using an online Digits-in-Noise (DIN) test, where participants completed three test conditions: one with standard speech, and two with vocoded speech using the NH and EH models. Both the standard DIN test and the EH-vocoded groups were statistically equivalent to clinically reported data for NH and CI listeners. On average, the NH and EH vocoded groups increased SRT compared to the standard test by 2.4 dB and 7.1 dB, respectively. These findings show that, although some degradation occurs, the vocoder can reconstruct intelligible speech under both hearing models and accurately reflects the reduced speech-in-noise performance experienced by CI users.


【8】 Brain-tuned Speech Models Better Reflect Speech Processing Stages in the  Brain

标题: 大脑调谐语音模型更好地反映大脑中的语音处理阶段
链接:https://arxiv.org/abs/2506.03832
作者: Omer Moussa,  Mariya Toneva 
备注:Proceedings of Interspeech 2025
摘要:预训练的自监督语音模型在语音任务中表现出色,但并不反映人类语音处理的层次结构,因为它们在中间层编码丰富的语义,而在后期层编码较差的语义。最近的研究表明,大脑调整(使用人脑记录微调模型)可以提高语音模型的语义理解。在这里,我们研究如何以及大脑调谐模型进一步反映大脑的语音处理的中间阶段。我们发现,大脑调整模型的后期层在与语义语言区域的对齐方面大大优于预训练模型。进一步的逐层探测表明,早期的层仍然致力于低级别的声学特征,而后期的层在复杂的高级别任务中变得最好。这些发现表明,大脑调整的模型不仅表现更好,而且表现出从声学到语义表示的良好定义的分层处理,使它们能够更好地为人类语音处理建模。
摘要:Pretrained self-supervised speech models excel in speech tasks but do not reflect the hierarchy of human speech processing, as they encode rich semantics in middle layers and poor semantics in late layers. Recent work showed that brain-tuning (fine-tuning models using human brain recordings) improves speech models' semantic understanding. Here, we examine how well brain-tuned models further reflect the brain's intermediate stages of speech processing. We find that late layers of brain-tuned models substantially improve over pretrained models in their alignment with semantic language regions. Further layer-wise probing reveals that early layers remain dedicated to low-level acoustic features, while late layers become the best at complex high-level tasks. These findings show that brain-tuned models not only perform better but also exhibit a well-defined hierarchical processing going from acoustic to semantic representations, making them better model organisms for human speech processing.


【9】 Conformer-based Ultrasound-to-Speech Conversion

标题: 基于适形器的超声到语音转换
链接:https://arxiv.org/abs/2506.03831
作者: Ibrahim Ibrahimov,  Zainkó Csaba,  Gábor Gosztolya 
备注:accepted to Interspeech 2025
摘要:深度神经网络在实现无声语音接口的超声波到语音转换任务方面表现出了很大的潜力。在这项工作中,我们应用了两种基于Conformer的DNN架构(Base和一种使用bi-LSTM)来完成这项任务。使用来自Ultrasuite-Tal 80数据集的四个扬声器的数据训练特定于扬声器的模型,同时使用HiFi-GAN声码器将生成的mel频谱图合成为音频波形。与标准2D-CNN基线相比,客观测量(MSE和mel倒谱失真)显示两种模型均无统计学显著改善。然而,MUSHRA听力测试显示,采用bi-LSTM的Conformer提供了更好的感知质量,而Conformer Base由于其更简单的架构,与基线的性能相匹配,训练时间快了3倍。这些发现表明,基于Conformer的模型,特别是具有bi-LSTM的Conformer,为CNN提供了一种有前途的替代方案,用于超声到语音转换。
摘要:Deep neural networks have shown promising potential for ultrasound-to-speech conversion task towards Silent Speech Interfaces. In this work, we applied two Conformer-based DNN architectures (Base and one with bi-LSTM) for this task. Speaker-specific models were trained on the data of four speakers from the Ultrasuite-Tal80 dataset, while the generated mel spectrograms were synthesized to audio waveform using a HiFi-GAN vocoder. Compared to a standard 2D-CNN baseline, objective measurements (MSE and mel cepstral distortion) showed no statistically significant improvement for either model. However, a MUSHRA listening test revealed that Conformer with bi-LSTM provided better perceptual quality, while Conformer Base matched the performance of the baseline along with a 3x faster training time due to its simpler architecture. These findings suggest that Conformer-based models, especially the Conformer with bi-LSTM, offer a promising alternative to CNNs for ultrasound-to-speech conversion.


【10】 MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech  Recognition

标题: MFLA:流语音识别的单调有限前瞻注意力
链接:https://arxiv.org/abs/2506.03722
作者: Yinfeng Xia,  Huiyan Li,  Chenyang Le,  Manhong Wang,  Yutao Sun,  Xingyang Ma,  Yanmin Qian 
备注:Accepted by Interspeech 2025
摘要:应用像Whisper这样的大型预训练语音模型,在降低各种语音任务的训练成本方面表现出了希望。然而,将这些模型集成到流媒体系统中仍然是一个挑战。本文提出了一种新的前缀到前缀的训练框架,通过微调的耳语流识别。我们引入连续集成和消防机制,建立一个准单调的连续语音序列和离散文本标记之间的对齐。此外,我们设计了单调有限前瞻注意,允许每个令牌参加无限左上下文和有限右上下文的语音序列。我们还采用了wait-k解码策略来简化解码过程,同时确保训练和测试之间的一致性。我们的理论分析和实验表明,这种方法实现了一个可控的延迟和质量之间的权衡,使其适用于各种流媒体应用。
摘要:Applying large pre-trained speech models like Whisper has shown promise in reducing training costs for various speech tasks. However, integrating these models into streaming systems remains a challenge. This paper presents a novel prefix-to-prefix training framework for streaming recognition by fine-tuning the Whisper. We introduce the Continuous Integrate-and-Fire mechanism to establish a quasi-monotonic alignment between continuous speech sequences and discrete text tokens. Additionally, we design Monotonic Finite Look-ahead Attention, allowing each token to attend to infinite left-context and finite right-context from the speech sequences. We also employ the wait-k decoding strategy to simplify the decoding process while ensuring consistency between training and testing. Our theoretical analysis and experiments demonstrate that this approach achieves a controllable trade-off between latency and quality, making it suitable for various streaming applications.


【11】 Efficient Data Selection for Domain Adaptation of ASR Using  Pseudo-Labels and Multi-Stage Filtering

标题: 使用伪标签和多阶段过滤进行ASB域自适应的高效数据选择
链接:https://arxiv.org/abs/2506.03681
作者: Pradeep Rangappa,  Andres Carofilis,  Jeena Prakash,  Shashi Kumar,  Sergio Burdisso,  Srikanth Madikeri,  Esau Villatoro-Tello,  Bidisha Sharma,  Petr Motlicek,  Kadri Hacioglu,  Shankar Venkatesan,  Saurabh Vyas,  Andreas Stolcke 
备注:Accepted at Interspeech 2025, Netherlands
摘要:针对特定领域微调预训练的ASR模型对于具有有限标记数据和计算资源的小型组织来说具有挑战性。在这里,我们探索了不同的数据选择管道,并提出了一种强大的方法,通过过滤使用Whisper(编码器-解码器)和Zipformer(换能器)模型生成的伪标签来改善ASR自适应。我们的方法集成了多种选择策略-包括单词错误率(WER)预测,命名实体识别(NER)和字符错误率(CER)分析-以提取高质量的训练片段。我们使用7500小时的基线对Whisper和Zipformer的方法进行了评估,并将其与基于CER的方法进行了比较,该方法依赖于来自三个ASR系统的假设。对7500小时的伪标记呼叫中心数据进行微调,实现了12.3%的WER,而我们的过滤将数据集减少到100小时(1.4%),具有类似的性能;在Fisher English上观察到类似的趋势。
摘要:Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English.


【12】 Comparative Analysis of Fast and High-Fidelity Neural Vocoders for  Low-Latency Streaming Synthesis in Resource-Constrained Environments

标题: 资源受限环境中用于低延迟流媒体合成的快速和高保真神经声码器的比较分析
链接:https://arxiv.org/abs/2506.03554
作者: Reo Yoneyama,  Masaya Kawamura,  Ryo Terashima,  Ryuichi Yamamoto,  Tomoki Toda 
备注:Accepted to Interspeech 2025
摘要:在实时语音合成中,神经声码器通常需要通过因果处理和流式传输进行低延迟合成。然而,流式传输引入了批量合成中不存在的低效率,诸如有限的并行性、帧间依赖性管理和参数加载开销。本文提出了多流Wavehax(MS-Wavehax),一个有效的神经声码器的低延迟流,通过扩展的混叠的神经声码器Wavehax与多流分解。我们分析了延迟吞吐量权衡在CPU的环境中,并确定流神经声码器的关键瓶颈。我们的研究结果为优化块大小和设计适合特定应用需求和硬件限制的声码器提供了实用的见解。此外,我们的主观评估表明,MS-Wavehax提供高语音质量的因果关系和非因果关系的条件下,同时非常紧凑,易于部署在资源受限的环境中。
摘要:In real-time speech synthesis, neural vocoders often require low-latency synthesis through causal processing and streaming. However, streaming introduces inefficiencies absent in batch synthesis, such as limited parallelism, inter-frame dependency management, and parameter loading overhead. This paper proposes multi-stream Wavehax (MS-Wavehax), an efficient neural vocoder for low-latency streaming, by extending the aliasing-free neural vocoder Wavehax with multi-stream decomposition. We analyze the latency-throughput trade-off in a CPU-only environment and identify key bottlenecks in streaming neural vocoders. Our findings provide practical insights for optimizing chunk sizes and designing vocoders tailored to specific application demands and hardware constraints. Furthermore, our subjective evaluations show that MS-Wavehax delivers high speech quality under causal and non-causal conditions while being remarkably compact and easily deployable in resource-constrained environments.


【13】 Local Equivariance Error-Based Metrics for Evaluating  Sampling-Frequency-Independent Property of Neural Network

标题: 基于局部等方差误差的神经网络采样频率无关性的插值器
链接:https://arxiv.org/abs/2506.03550
作者: Kanami Imamura,  Tomohiko Nakamura,  Norihiro Takamune,  Kohei Yatabe,  Hiroshi Saruwatari 
备注:5 pages, 4 figures, accepted for European Signal Processing Conference 2025 (EUSIPCO 2025)
摘要:基于深度神经网络(DNN)的音频信号处理方法通常仅在单个采样频率(SF)下进行训练,因此需要信号恢复来处理未经训练的SF。然而,最近的研究表明,信号恢复会降低未经训练的SF的性能。这个问题被忽视了,因为大多数研究只评估训练有素的SF的表现。在本文中,为了评估DNN对SF变化的鲁棒性,我们将其称为SF独立(SFI)属性,我们提出了三个度量来量化基于局部等方差误差(LEE)的SFI属性。LEE测量DNN对输入变换的鲁棒性。通过将信号重构作为输入变换,我们将LEE推广到度量音频源分离方法对信号重构的鲁棒性。所提出的指标被构造来量化特定网络组件中负责预测时频掩码的SFI属性。音乐源分离的实验表明,在未经训练的SF所提出的指标和性能下降之间有很强的相关性。
摘要:Audio signal processing methods based on deep neural networks (DNNs) are typically trained only at a single sampling frequency (SF) and therefore require signal resampling to handle untrained SFs. However, recent studies have shown that signal resampling can degrade performance with untrained SFs. This problem has been overlooked because most studies evaluate only the performance at trained SFs. In this paper, to assess the robustness of DNNs to SF changes, which we refer to as the SF-independent (SFI) property, we propose three metrics to quantify the SFI property on the basis of local equivariance error (LEE). LEE measures the robustness of DNNs to input transformations. By using signal resampling as input transformation, we extend LEE to measure the robustness of audio source separation methods to signal resampling. The proposed metrics are constructed to quantify the SFI property in specific network components responsible for predicting time-frequency masks. Experiments on music source separation demonstrated a strong correlation between the proposed metrics and performance degradation at untrained SFs.


【14】 BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and  Weight Indexing

标题: BitTTC:使用1.58位量化和权重索引的高度紧凑的文本到语音
链接:https://arxiv.org/abs/2506.03515
作者: Masaya Kawamura,  Takuya Hasumi,  Yuma Shirahata,  Ryuichi Yamamoto 
备注:Accepted to INTERSPEECH 2025
摘要:本文提出了一种高度紧凑,轻量级的文本到语音(TTS)模型的设备上的应用。为了减小模型的大小,所提出的模型引入了两种技术。首先,我们引入了量化感知训练(QAT),它在训练期间将模型参数量化到低至1.58位。在这种情况下,大多数32位模型参数被量化为三进制值{-1,0,1}。其次,我们提出了一种方法命名为权重索引。在这种方法中,我们将一组1.58位权重保存为单个int 8索引。这允许模型参数的有效存储,即使在以8位为单位处理值的硬件上。实验结果表明,该方法实现了83%的减少模型的大小,同时优于基线的相似的模型大小没有量化的合成质量。
摘要:This paper proposes a highly compact, lightweight text-to-speech (TTS) model for on-device applications. To reduce the model size, the proposed model introduces two techniques. First, we introduce quantization-aware training (QAT), which quantizes model parameters during training to as low as 1.58-bit. In this case, most of 32-bit model parameters are quantized to ternary values {-1, 0, 1}. Second, we propose a method named weight indexing. In this method, we save a group of 1.58-bit weights as a single int8 index. This allows for efficient storage of model parameters, even on hardware that treats values in units of 8-bit. Experimental results demonstrate that the proposed method achieved 83 % reduction in model size, while outperforming the baseline of similar model size without quantization in synthesis quality.


【15】 Towards Source Attribution of Singing Voice Deepfake with Multimodal  Foundation Models

标题: 利用多模式基础模型研究演唱声音Deepfake的源属性
链接:https://arxiv.org/abs/2506.03364
作者: Orchid Chetia Phukan,  Girish,  Mohd Mujtaba Akhtar,  Swarup Ranjan Behera,  Priyabrata Mallick,  Pailla Balakrishna Reddy,  Arun Balaji Buduru,  Rajesh Sharma 
备注:Accepted to INTERSPEECH 2025
摘要:在这项工作中,我们介绍了歌唱声音Deepfake源属性(SVDSA)的任务。我们假设,多模态基础模型(MMFM),如ImageBind,WavelageBind,将是最有效的SVDSA,因为它们更好地捕捉微妙的源特定的特征,如独特的音色,音高操纵,或合成工件的每个歌声deepfake源由于其跨模态预训练。我们的实验与MMFM,语音基础模型和音乐基础模型验证的假设,MMFM是最有效的SVDSA。此外,受相关研究的启发,我们还探讨了融合基础模型(FM)的改进SVDSA。为此,我们提出了一个新的框架,COFFE采用的FMF距离作为新的损失函数的有效融合。通过COFFE与MMFM的交响乐,我们达到了最高的性能相比,所有的个人FM和基线融合方法。
摘要:In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they are better equipped for capturing subtle source-specific characteristics-such as unique timbre, pitch manipulation, or synthesis artifacts of each singing voice deepfake source due to their cross-modality pre-training. Our experiments with MMFMs, speech foundation models and music foundation models verify the hypothesis that MMFMs are the most effective for SVDSA. Furthermore, inspired from related research, we also explore fusion of foundation models (FMs) for improved SVDSA. To this end, we propose a novel framework, COFFE which employs Chernoff Distance as novel loss function for effective fusion of FMs. Through COFFE with the symphony of MMFMs, we attain the topmost performance in comparison to all the individual FMs and baseline fusion methods.


eess.AS音频处理


【1】 HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset

标题: HiFiRTS-2:大规模高带宽语音数据集
链接:https://arxiv.org/abs/2506.04152
作者: Ryan Langman,  Xuesong Yang,  Paarth Neekhara,  Shehzeen Hussain,  Edresson Casanova,  Evelina Bakhturina,  Jason Li 
备注:Submitted to Interspeech 2025
摘要:本文介绍了HiFiTTS-2,一个大规模的语音数据集设计的高带宽语音合成。该数据集来自LibriVox有声读物,包含大约36.7k小时的英语语音(22.05 kHz训练)和31.7k小时的44.1 kHz训练。我们提出了我们的数据处理流水线,包括带宽估计,分割,文本预处理,和多说话人检测。该数据集伴随着由我们的管道生成的详细话语和有声读物元数据,使研究人员能够应用数据质量过滤器来使数据集适应各种用例。实验结果表明,我们的数据管道和产生的数据集可以促进高质量,zero-shot文本到语音(TTS)模型在高带宽的训练。
摘要:This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximately 36.7k hours of English speech for 22.05 kHz training, and 31.7k hours for 44.1 kHz training. We present our data processing pipeline, including bandwidth estimation, segmentation, text preprocessing, and multi-speaker detection. The dataset is accompanied by detailed utterance and audiobook metadata generated by our pipeline, enabling researchers to apply data quality filters to adapt the dataset to various use cases. Experimental results demonstrate that our data pipeline and resulting dataset can facilitate the training of high-quality, zero-shot text-to-speech (TTS) models at high bandwidths.


【2】 Sound Field Reconstruction Using Physics-Informed Boundary Integral  Networks

标题: 使用物理知识的边界积分网络重建声学场
链接:https://arxiv.org/abs/2506.03917
作者: Stefano Damiano,  Toon van Waterschoot 
备注:Accepted for publication at EUSIPCO 2025
摘要:声场重建是指仅使用有限的测量集合来估计空间的任意区域上的声压场的问题。物理信息的神经网络已被采用来解决这个问题,通过将控制偏微分方程,无论是亥姆霍兹或波动方程的训练损失函数。在这项工作中,我们介绍了一个边界积分网络的声场重建。基于Kirchhoff-Helmholtz边界积分方程对给定空间区域内的声场进行建模,采用浅层神经网络反演所考虑区域边界上的声压分布,从而能够准确地反演其内部的声压。假设测量麦克风的位置已知,我们通过最小化那些位置处的估计压力和测量压力之间的均方误差来训练模型。实验结果表明,该模型优于现有的物理信息的数据驱动技术。
摘要:Sound field reconstruction refers to the problem of estimating the acoustic pressure field over an arbitrary region of space, using only a limited set of measurements. Physics-informed neural networks have been adopted to solve the problem by incorporating in the training loss function the governing partial differential equation, either the Helmholtz or the wave equation. In this work, we introduce a boundary integral network for sound field reconstruction. Relying on the Kirchhoff-Helmholtz boundary integral equation to model the sound field in a given region of space, we employ a shallow neural network to retrieve the pressure distribution on the boundary of the considered domain, enabling to accurately retrieve the acoustic pressure inside of it. Assuming the positions of measurement microphones are known, we train the model by minimizing the mean squared error between the estimated and measured pressure at those locations. Experimental results indicate that the proposed model outperforms existing physics-informed data-driven techniques.


【3】 Tone recognition in low-resource languages of North-East India: peeling  the layers of SSL-based speech models

标题: 印度东北部低资源语言中的语气识别:剥离基于SSL的语音模型的分层
链接:https://arxiv.org/abs/2506.03606
作者: Parismita Gogoi,  Sishir Kalita,  Wendy Lalhminghlui,  Viyazonuo Terhiija,  Moakala Tzudir,  Priyankoo Sarmah,  S. R. M. Prasanna 
备注:Accepted in Interspeech2025
摘要:本研究探讨了使用自监督学习(SSL)模型在印度东北部的三种低资源语言中进行音调识别:Angami,Ao和Mizo。我们评估了四个Wav2vec2.0基础模型,这些模型在音调和非音调语言上进行了预训练。我们分析了所有三种语言的层之间的音调性能,并比较了不同的模型。我们的研究结果表明,声调识别效果最好的Mizo和最差的Angami。SSL模型的中间层对于音调识别是最重要的,无论预训练语言是什么,即音调还是非音调。我们还发现声调、声调类型和方言变异对声调识别有影响。这些发现为基于SSL的音调语言嵌入的优点和缺点提供了有用的见解,并强调了在低资源环境中提高音调识别的潜力。源代码可以在GitHub 1上找到。
摘要:This study explores the use of self-supervised learning (SSL) models for tone recognition in three low-resource languages from North Eastern India: Angami, Ao, and Mizo. We evaluate four Wav2vec2.0 base models that were pre-trained on both tonal and non-tonal languages. We analyze tone-wise performance across the layers for all three languages and compare the different models. Our results show that tone recognition works best for Mizo and worst for Angami. The middle layers of the SSL models are the most important for tone recognition, regardless of the pre-training language, i.e. tonal or non-tonal. We have also found that the tone inventory, tone types, and dialectal variations affect tone recognition. These findings provide useful insights into the strengths and weaknesses of SSL-based embeddings for tonal languages and highlight the potential for improving tone recognition in low-resource settings. The source code is available at GitHub 1 .


【4】 BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and  Weight Indexing

标题: BitTTC:使用1.58位量化和权重索引的高度紧凑的文本到语音
链接:https://arxiv.org/abs/2506.03515
作者: Masaya Kawamura,  Takuya Hasumi,  Yuma Shirahata,  Ryuichi Yamamoto 
备注:Accepted to INTERSPEECH 2025
摘要:本文提出了一种高度紧凑,轻量级的文本到语音(TTS)模型的设备上的应用。为了减小模型的大小,所提出的模型引入了两种技术。首先,我们引入了量化感知训练(QAT),它在训练期间将模型参数量化到低至1.58位。在这种情况下,大多数32位模型参数被量化为三进制值{-1,0,1}。其次,我们提出了一种方法命名为权重索引。在这种方法中,我们将一组1.58位权重保存为单个int 8索引。这允许模型参数的有效存储,即使在以8位为单位处理值的硬件上。实验结果表明,该方法实现了83%的减少模型的大小,同时优于基线的相似的模型大小没有量化的合成质量。
摘要:This paper proposes a highly compact, lightweight text-to-speech (TTS) model for on-device applications. To reduce the model size, the proposed model introduces two techniques. First, we introduce quantization-aware training (QAT), which quantizes model parameters during training to as low as 1.58-bit. In this case, most of 32-bit model parameters are quantized to ternary values {-1, 0, 1}. Second, we propose a method named weight indexing. In this method, we save a group of 1.58-bit weights as a single int8 index. This allows for efficient storage of model parameters, even on hardware that treats values in units of 8-bit. Experimental results demonstrate that the proposed method achieved 83 % reduction in model size, while outperforming the baseline of similar model size without quantization in synthesis quality.


【5】 A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations

标题: 音频Deepfake解释的数据驱动的基于扩散的方法
链接:https://arxiv.org/abs/2506.03425
作者: Petr Grinberg,  Ankur Kumar,  Surya Koppisetti,  Gaurav Bharaj 
备注:5 pages, 3 figures, accepted at Interspeech 2025
摘要:在音频深度伪造检测的背景下评估可解释性技术(例如SHAP和LRP)是具有挑战性的,因为缺乏清晰的地面实况注释。在我们能够获得基本事实的情况下,我们发现这些方法很难提供准确的解释。在这项工作中,我们提出了一种新的数据驱动方法来识别deepfake音频中的伪影区域。我们考虑成对的真实音频和声码音频,并使用时频表示的差异作为地面实况解释。然后,差信号用作监督以训练扩散模型,以暴露给定声码音频中的深度伪影。在VocV4和LibriSeVoc数据集上的实验结果表明,我们的方法在定性和定量方面都优于传统的可解释性技术。
摘要:Evaluating explainability techniques, such as SHAP and LRP, in the context of audio deepfake detection is challenging due to lack of clear ground truth annotations. In the cases when we are able to obtain the ground truth, we find that these methods struggle to provide accurate explanations. In this work, we propose a novel data-driven approach to identify artifact regions in deepfake audio. We consider paired real and vocoded audio, and use the difference in time-frequency representation as the ground-truth explanation. The difference signal then serves as a supervision to train a diffusion model to expose the deepfake artifacts in a given vocoded audio. Experimental results on the VocV4 and LibriSeVoc datasets demonstrate that our method outperforms traditional explainability techniques, both qualitatively and quantitatively.


【6】 HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in  Hyperbolic Space for Speech Emotion Recognition

标题: Hyzens:在双曲空间中对齐异类语音预训练表示以进行语音情感识别
链接:https://arxiv.org/abs/2506.03403
作者: Orchid Chetia Phukan,  Girish,  Mohd Mujtaba Akhtar,  Swarup Ranjan Behera,  Pailla Balakrishna Reddy,  Arun Balaji Buduru,  Rajesh Sharma 
备注:Accepted to INTERSPEECH 2025
摘要:来自神经音频编解码器(如EnCodec)的基于压缩的表示(CBR)捕获复杂的声学特征,如音高和音色,而来自预先训练的模型的基于表示学习的表示(RLR)用于语音表示学习,如WavLM编码高级语义和韵律信息。以往的语音情感识别(SER)的研究都是在这两方面进行的,但CBR和RLR的融合还没有被研究。在这项研究中,我们解决了这一差距,并调查融合的RLRs和CBR,并假设他们将更有效地提供补充信息。为此,我们提出,HYDROGEN,一个新的框架,融合的表示,将它们转换为双曲空间。使用HYPERTIES,通过x矢量(RLR)和声音流(CBR)的融合,我们实现了与单个表示以及RLR和CBR的同质融合相比的最佳性能,并报告SOTA。
摘要:Compression-based representations (CBRs) from neural audio codecs such as EnCodec capture intricate acoustic features like pitch and timbre, while representation-learning-based representations (RLRs) from pre-trained models trained for speech representation learning such as WavLM encode high-level semantic and prosodic information. Previous research on Speech Emotion Recognition (SER) has explored both, however, fusion of CBRs and RLRs haven't been explored yet. In this study, we solve this gap and investigate the fusion of RLRs and CBRs and hypothesize they will be more effective by providing complementary information. To this end, we propose, HYFuse, a novel framework that fuses the representations by transforming them to hyperbolic space. With HYFuse, through fusion of x-vector (RLR) and Soundstream (CBR), we achieve the top performance in comparison to individual representations as well as the homogeneous fusion of RLRs and CBRs and report SOTA.


【7】 SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through  Audio-Visual Alignment with Cascaded Cross-Transformer

标题: SNIFR:通过级联交叉Transformer的视听对齐来增强细粒度儿童有害内容检测
链接:https://arxiv.org/abs/2506.03378
作者: Orchid Chetia Phukan,  Mohd Mujtaba Akhtar,  Girish,  Swarup Ranjan Behera,  Abu Osama Siddiqui,  Sarthak Jain,  Priyabrata Mallick,  Jaya Sai Kiran Patibandla,  Pailla Balakrishna Reddy,  Arun Balaji Buduru,  Rajesh Sharma 
备注:Accepted to INTERSPEECH 2025
摘要:随着视频分享平台在过去十年中的发展,儿童观看人数激增,这增加了对暴力或露骨场景等有害内容的精确检测的需求。恶意用户通过在最小帧中嵌入不安全内容来利用审核系统以逃避检测。虽然先前的研究集中在视觉线索和先进的细粒度检测,音频功能仍然没有得到充分的探索。在这项研究中,我们嵌入音频线索与视觉细粒度的儿童有害内容检测,并引入SNIFR,一种新的框架,有效的对齐。SNIFR采用Transformer编码器进行模态内交互,然后采用级联交叉Transformer进行模态间对齐。我们的方法比单峰和基线融合方法具有更好的性能,开创了一个新的最先进的技术。
摘要:As video-sharing platforms have grown over the past decade, child viewership has surged, increasing the need for precise detection of harmful content like violence or explicit scenes. Malicious users exploit moderation systems by embedding unsafe content in minimal frames to evade detection. While prior research has focused on visual cues and advanced such fine-grained detection, audio features remain underexplored. In this study, we embed audio cues with visual for fine-grained child harmful content detection and introduce SNIFR, a novel framework for effective alignment. SNIFR employs a transformer encoder for intra-modality interaction, followed by a cascaded cross-transformer for inter-modality alignment. Our approach achieves superior performance over unimodal and baseline fusion methods, setting a new state-of-the-art.


【8】 Towards Source Attribution of Singing Voice Deepfake with Multimodal  Foundation Models

标题: 利用多模式基础模型研究演唱声音Deepfake的源属性
链接:https://arxiv.org/abs/2506.03364
作者: Orchid Chetia Phukan,  Girish,  Mohd Mujtaba Akhtar,  Swarup Ranjan Behera,  Priyabrata Mallick,  Pailla Balakrishna Reddy,  Arun Balaji Buduru,  Rajesh Sharma 
备注:Accepted to INTERSPEECH 2025
摘要:在这项工作中,我们介绍了歌唱声音Deepfake源属性(SVDSA)的任务。我们假设,多模态基础模型(MMFM),如ImageBind,WavelageBind,将是最有效的SVDSA,因为它们更好地捕捉微妙的源特定的特征,如独特的音色,音高操纵,或合成工件的每个歌声deepfake源由于其跨模态预训练。我们的实验与MMFM,语音基础模型和音乐基础模型验证的假设,MMFM是最有效的SVDSA。此外,受相关研究的启发,我们还探讨了融合基础模型(FM)的改进SVDSA。为此,我们提出了一个新的框架,COFFE采用的FMF距离作为新的损失函数的有效融合。通过COFFE与MMFM的交响乐,我们达到了最高的性能相比,所有的个人FM和基线融合方法。
摘要:In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they are better equipped for capturing subtle source-specific characteristics-such as unique timbre, pitch manipulation, or synthesis artifacts of each singing voice deepfake source due to their cross-modality pre-training. Our experiments with MMFMs, speech foundation models and music foundation models verify the hypothesis that MMFMs are the most effective for SVDSA. Furthermore, inspired from related research, we also explore fusion of foundation models (FMs) for improved SVDSA. To this end, we propose a novel framework, COFFE which employs Chernoff Distance as novel loss function for effective fusion of FMs. Through COFFE with the symphony of MMFMs, we attain the topmost performance in comparison to all the individual FMs and baseline fusion methods.


【9】 Sounding that Object: Interactive Object-Aware Image to Audio Generation

标题: 探测该对象:交互式对象感知图像到音频生成
链接:https://arxiv.org/abs/2506.04214
作者: Tingle Li,  Baihe Huang,  Xiaobin Zhuang,  Dongya Jia,  Jiawei Chen,  Yuping Wang,  Zhuo Chen,  Gopala Anumanchipalli,  Yuxuan Wang 
备注:ICML 2025
摘要:为复杂的视听场景生成准确的声音是一项挑战,特别是在存在多个对象和声源的情况下。在本文中,我们提出了一个{\em交互式对象感知音频生成}模型,该模型将声音生成建立在用户选择的图像中的视觉对象上。我们的方法将以对象为中心的学习集成到一个条件潜在扩散模型中,该模型通过多模态注意力学习将图像区域与其相应的声音相关联。在测试时,我们的模型采用图像分割,让用户交互式地生成声音在{\em对象}的水平。我们从理论上验证了我们的注意力机制在功能上近似于测试时的分割掩码,确保生成的音频与选定的对象对齐。定量和定性评估表明,我们的模型优于基线,实现了对象及其相关声音之间的更好对齐。项目页面:https://tinglok.netlify.app/files/avobject/
摘要:Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds sound generation in user-selected visual objects within images. Our method integrates object-centric learning into a conditional latent diffusion model, which learns to associate image regions with their corresponding sounds through multi-modal attention. At test time, our model employs image segmentation to allow users to interactively generate sounds at the {\em object} level. We theoretically validate that our attention mechanism functionally approximates test-time segmentation masks, ensuring the generated audio aligns with selected objects. Quantitative and qualitative evaluations show that our model outperforms baselines, achieving better alignment between objects and their associated sounds. Project page: https://tinglok.netlify.app/files/avobject/


【10】 UniCUE: Unified Recognition and Generation Framework for Chinese Cued  Speech Video-to-Speech Generation

标题: UniCUE:中文提示语音视频到语音生成的统一识别和生成框架
链接:https://arxiv.org/abs/2506.04134
作者: Jinting Wang,  Shan Yang,  Li Liu 
备注:10 pages, 10 figures
摘要:提示语音(CS)通过手编码增强唇读,为听障者提供精确的语音感知支持。CS视频到语音生成(CSV 2S)任务旨在将听障者的CS视觉表达(CS视频)转换为可理解的语音信号。从CS视频直接生成语音(称为单CSV 2S)由于CS数据不足而产生差的性能。目前的研究主要集中在CS识别(CSR),它将视频内容转换为语言文本。基于此,CSV 2S的一个直接方法是将CSR与文本到语音系统相结合。这种组合架构依赖于文本作为中间媒介逐步跨模态对齐,这可能会导致错误传播和语音和视频动态之间的时间错位。为了解决这些挑战,我们提出了一种新的方法,直接从CS视频生成语音,而不依赖于中间文本。在此基础上,我们提出了UniCUE,第一个统一的框架CSV 2S,其核心创新在于整合的CSR任务,提供细粒度的视觉语义信息,以促进语音生成从CS视频。更确切地说,(1)一种新的细粒度语义对齐池,以确保视觉特征和语音内容之间的精确映射;(2)一个VisioPhonetic适配器,以桥接跨任务表示,确保两个不同任务之间的无缝兼容性(即,CSV 2S和CSR);(3)提出了一种基于姿态感知的视觉处理器,以增强CS视频中唇和手运动之间的细粒度时空相关性。在我们新建立的中文CS数据集(14 cuers 1:8 hearing-impaired and 6 norm-hearing)上的实验表明,与单一的CSV 2S相比,我们的UniCUE显著降低了78.3%的单词错误率,并提高了32%的唇语同步。
摘要:Cued Speech (CS) enhances lipreading through hand coding, providing precise speech perception support for the hearing-impaired. CS Video-to-Speech generation (CSV2S) task aims to convert the CS visual expressions (CS videos) of hearing-impaired individuals into comprehensible speech signals. Direct generation of speech from CS video (called single CSV2S) yields poor performance due to insufficient CS data. Current research mostly focuses on CS Recognition (CSR), which convert video content into linguistic text. Based on this, one straightforward way of CSV2S is to combine CSR with a Text-to-Speech system. This combined architecture relies on text as an intermediate medium for stepwise cross-modal alignment, which may lead to error propagation and temporal misalignment between speech and video dynamics. To address these challenges, we propose a novel approach that directly generates speech from CS videos without relying on intermediate text. Building upon this, we propose UniCUE, the first unified framework for CSV2S, whose core innovation lies in the integration of the CSR task that provides fine-grained visual-semantic information to facilitate speech generation from CS videos. More precisely, (1) a novel fine-grained semantic alignment pool to ensure precise mapping between visual features and speech contents; (2) a VisioPhonetic adapter to bridge cross-task representations, ensuring seamless compatibility between two distinct tasks (i.e., CSV2S and CSR); (3) a pose-aware visual processor is introduced to enhance fine-grained spatiotemporal correlations between lip and hand movements in CS video. Experiments on our new established Chinese CS dataset (14 cuers1: 8 hearing-impaired and 6 normal-hearing) show that our UniCUE significantly reduces Word Error Rate by 78.3% and improves lip-speech synchronization by 32% compared to the single CSV2S.


【11】 A Novel Data Augmentation Approach for Automatic Speaking Assessment on  Opinion Expressions

标题: 一种新的数据增强方法,用于对意见表达进行自动口语评估
链接:https://arxiv.org/abs/2506.04077
作者: Chung-Chun Wang,  Jhen-Ke Lin,  Hao-Chien Lu,  Hong-Yun Lin,  Berlin Chen 
备注:submitted to the ISCA SLaTE-2025 Workshop
摘要:意见表达的自动口语评估(ASA)通常受到标记录音稀缺的阻碍,这限制了即时多样性并破坏了评分可靠性。为了解决这一挑战,我们提出了一种新的训练范式,利用大型语言模型(LLM)来生成给定熟练程度的各种响应,通过说话者感知的文本到语音合成将响应转换为合成语音,并采用动态重要性损失来基于合成语音和真实语音之间的特征分布差异自适应地重新加权训练实例。随后,多模态大语言模型集成对齐的文本特征与语音信号直接预测熟练度分数。在LTTC数据集上进行的实验表明,我们的方法优于依赖于真实数据或传统增强的方法,有效地缓解了低资源约束,并使ASA能够在具有跨模态信息的意见表达上。
摘要:Automated speaking assessment (ASA) on opinion expressions is often hampered by the scarcity of labeled recordings, which restricts prompt diversity and undermines scoring reliability. To address this challenge, we propose a novel training paradigm that leverages a large language models (LLM) to generate diverse responses of a given proficiency level, converts responses into synthesized speech via speaker-aware text-to-speech synthesis, and employs a dynamic importance loss to adaptively reweight training instances based on feature distribution differences between synthesized and real speech. Subsequently, a multimodal large language model integrates aligned textual features with speech signals to predict proficiency scores directly. Experiments conducted on the LTTC dataset show that our approach outperforms methods relying on real data or conventional augmentation, effectively mitigating low-resource constraints and enabling ASA on opinion expressions with cross-modal information.


【12】 Acoustically Precise Hesitation Tagging Is Essential for End-to-End  Verbatim Transcription Systems

标题: 声学精确的犹豫标记对于端到端逐字转录系统至关重要
链接:https://arxiv.org/abs/2506.04076
作者: Jhen-Ke Lin,  Hao-Chien Lu,  Chung-Chun Wang,  Hong-Yun Lin,  Berlin Chen 
备注:submitted to the ISCA SLaTE-2025 Workshop
摘要:用于自动口语评估的逐字记录需要准确捕捉不流利的内容,这对于错误分析和反馈等下游任务至关重要。然而,许多ASR系统丢弃或概括犹豫,丢失重要的声学细节。我们使用低秩自适应(LoRA)在Speak & Improve 2025语料库上微调Whisper模型,而无需依赖外部音频训练数据。我们比较了三种注释方案:删除犹豫(纯),通用标签(丰富),声学精确的填料推断双子座2.0闪光灯从现有的音频转录对(额外)。我们的挑战系统实现了6.47% WER(纯)和5.81% WER(额外)。挑战后的实验表明,微调耳语大V3涡轮与“额外”计划产生了5.5%的WER,11.3%的相对改善“纯”计划(6.2% WER)。这表明,明确的,现实的填充停顿标记显着提高ASR的准确性逐字L2语音转录。
摘要:Verbatim transcription for automatic speaking assessment demands accurate capture of disfluencies, crucial for downstream tasks like error analysis and feedback. However, many ASR systems discard or generalize hesitations, losing important acoustic details. We fine-tune Whisper models on the Speak & Improve 2025 corpus using low-rank adaptation (LoRA), without recourse to external audio training data. We compare three annotation schemes: removing hesitations (Pure), generic tags (Rich), and acoustically precise fillers inferred by Gemini 2.0 Flash from existing audio-transcript pairs (Extra). Our challenge system achieved 6.47% WER (Pure) and 5.81% WER (Extra). Post-challenge experiments reveal that fine-tuning Whisper Large V3 Turbo with the "Extra" scheme yielded a 5.5% WER, an 11.3% relative improvement over the "Pure" scheme (6.2% WER). This demonstrates that explicit, realistic filled-pause labeling significantly enhances ASR accuracy for verbatim L2 speech transcription.


【13】 A Statistics-Driven Differentiable Approach for Sound Texture Synthesis  and Analysis

标题: 一种统计驱动的声音纹理合成与分析的可区分方法
链接:https://arxiv.org/abs/2506.04073
作者: Esteban Gutiérrez,  Frederic Font,  Xavier Serra,  Lonce Wyse 
备注:Accepted to the 28th International Conference on Digital Audio Effects (DAFx 2025) to be held in Ancona, Italy. 8 pages, one diagram and 5 tables
摘要:在这项工作中,我们介绍了TexStat,一种新的损失函数,专门设计用于分析和合成纹理的声音,其特征在于随机结构和感知平稳性。从McDermott和Simoncelli的统计和感知框架中汲取灵感,TexStat识别属于同一纹理类别的信号之间的相似性,而不依赖于时间结构。我们还建议使用TexStat作为验证度量一起Frechet音频距离(FAD)来评估纹理声音合成模型。除了TexStat,我们提出了TexEnv,一个有效的,轻量级的和可区分的纹理声音合成器,通过对过滤后的噪声施加幅度包络来生成音频。我们进一步将这些组件集成到TexDSP中,TexDSP是一种为纹理声音量身定制的DDSP启发的生成模型。通过对各种纹理声音类型的广泛实验,我们证明了TexStat在感知上是有意义的,时不变的,并且对噪声具有鲁棒性,这些特征使其既可以作为生成任务的损失函数,也可以作为验证指标。所有工具和代码都是作为开源贡献提供的,我们的PyTorch实现是高效的,可区分的,高度可配置的,使其能够在生成任务中使用,并作为感知基础的评估指标。
摘要:In this work, we introduce TexStat, a novel loss function specifically designed for the analysis and synthesis of texture sounds characterized by stochastic structure and perceptual stationarity. Drawing inspiration from the statistical and perceptual framework of McDermott and Simoncelli, TexStat identifies similarities between signals belonging to the same texture category without relying on temporal structure. We also propose using TexStat as a validation metric alongside Frechet Audio Distances (FAD) to evaluate texture sound synthesis models. In addition to TexStat, we present TexEnv, an efficient, lightweight and differentiable texture sound synthesizer that generates audio by imposing amplitude envelopes on filtered noise. We further integrate these components into TexDSP, a DDSP-inspired generative model tailored for texture sounds. Through extensive experiments across various texture sound types, we demonstrate that TexStat is perceptually meaningful, time-invariant, and robust to noise, features that make it effective both as a loss function for generative tasks and as a validation metric. All tools and code are provided as open-source contributions and our PyTorch implementations are efficient, differentiable, and highly configurable, enabling its use in both generative tasks and as a perceptually grounded evaluation metric.


【14】 The mutual exclusivity bias of bilingual visually grounded speech models

标题: 双语视觉基础语音模型的相互排他性偏见
链接:https://arxiv.org/abs/2506.04037
作者: Dan Oneata,  Leanne Nortje,  Yevgen Matusevych,  Herman Kamper 
备注:Interspeech 2025
摘要:互斥性(ME)是一种策略,其中一个新的词与一个新的对象,而不是一个熟悉的,促进儿童的语言学习。最近的研究发现,在用成对图像训练英语语音的视觉接地语音(VGS)模型中存在ME偏见。但ME也在双语儿童中进行了研究,由于跨语言歧义,他们可能较少使用它。我们使用在英语,法语和荷兰语的组合上训练的双语VGS模型来计算探索这种模式。我们发现,双语模型一般表现出较弱的ME偏置比单语模型,虽然存在例外。分析表明,双语模型的组合视觉嵌入对于熟悉的数据具有较小的方差,部分解释了新概念和熟悉概念之间的混淆增加。我们还提供了新的见解,为什么我的偏见存在于VGS模型摆在首位。代码和数据:https://github.com/danoneata/me-vgs
摘要:Mutual exclusivity (ME) is a strategy where a novel word is associated with a novel object rather than a familiar one, facilitating language learning in children. Recent work has found an ME bias in a visually grounded speech (VGS) model trained on English speech with paired images. But ME has also been studied in bilingual children, who may employ it less due to cross-lingual ambiguity. We explore this pattern computationally using bilingual VGS models trained on combinations of English, French, and Dutch. We find that bilingual models generally exhibit a weaker ME bias than monolingual models, though exceptions exist. Analyses show that the combined visual embeddings of bilingual models have a smaller variance for familiar data, partly explaining the increase in confusion between novel and familiar concepts. We also provide new insights into why the ME bias exists in VGS models in the first place. Code and data: https://github.com/danoneata/me-vgs


【15】 Towards Better Disentanglement in Non-Autoregressive Zero-Shot  Expressive Voice Conversion

标题: 在非自回归Zero-Shot表达性语音转换中实现更好的解纠缠
链接:https://arxiv.org/abs/2506.04013
作者: Seymanur Akti,  Tuan Nam Nguyen,  Alexander Waibel 
备注:Accepted to Interspeech 2025
摘要:表达性语音转换的目的是将说话人身份和表达属性从目标语音转换为给定的源语音。在这项工作中,我们改进了一个具有条件变分自编码器的自监督非自回归框架,重点是减少源音色泄漏和改善语言声学解纠缠,以实现更好的风格传递。为了最大限度地减少风格泄漏,我们使用多语言离散语音单元进行内容表示,并通过基于增强的相似性损失和混合风格层规范化来加强嵌入。为了增强表现力转移,我们通过交叉注意将局部F0信息结合起来,并提取富含全局音高和能量特征的风格嵌入。实验表明,我们的模型在情感和说话人相似性方面优于基线,表现出优越的风格适应性和减少源风格泄漏。
摘要:Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework with a conditional variational autoencoder, focusing on reducing source timbre leakage and improving linguistic-acoustic disentanglement for better style transfer. To minimize style leakage, we use multilingual discrete speech units for content representation and reinforce embeddings with augmentation-based similarity loss and mix-style layer normalization. To enhance expressivity transfer, we incorporate local F0 information via cross-attention and extract style embeddings enriched with global pitch and energy features. Experiments show our model outperforms baselines in emotion and speaker similarity, demonstrating superior style adaptation and reduced source style leakage.


【16】 Brain-tuned Speech Models Better Reflect Speech Processing Stages in the  Brain

标题: 大脑调谐语音模型更好地反映大脑中的语音处理阶段
链接:https://arxiv.org/abs/2506.03832
作者: Omer Moussa,  Mariya Toneva 
备注:Proceedings of Interspeech 2025
摘要:预训练的自监督语音模型在语音任务中表现出色,但并不反映人类语音处理的层次结构,因为它们在中间层编码丰富的语义,而在后期层编码较差的语义。最近的研究表明,大脑调整(使用人脑记录微调模型)可以提高语音模型的语义理解。在这里,我们研究如何以及大脑调谐模型进一步反映大脑的语音处理的中间阶段。我们发现,大脑调整模型的后期层在与语义语言区域的对齐方面比预训练模型有了显着改善。进一步的逐层探测表明,早期的层仍然致力于低级别的声学特征,而后期的层在复杂的高级别任务中变得最好。这些发现表明,大脑调整的模型不仅表现更好,而且表现出从声学到语义表示的良好定义的分层处理,使它们能够更好地为人类语音处理建模。
摘要:Pretrained self-supervised speech models excel in speech tasks but do not reflect the hierarchy of human speech processing, as they encode rich semantics in middle layers and poor semantics in late layers. Recent work showed that brain-tuning (fine-tuning models using human brain recordings) improves speech models' semantic understanding. Here, we examine how well brain-tuned models further reflect the brain's intermediate stages of speech processing. We find that late layers of brain-tuned models substantially improve over pretrained models in their alignment with semantic language regions. Further layer-wise probing reveals that early layers remain dedicated to low-level acoustic features, while late layers become the best at complex high-level tasks. These findings show that brain-tuned models not only perform better but also exhibit a well-defined hierarchical processing going from acoustic to semantic representations, making them better model organisms for human speech processing.


【17】 MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech  Recognition

标题: MFLA:流语音识别的单调有限前瞻注意力
链接:https://arxiv.org/abs/2506.03722
作者: Yinfeng Xia,  Huiyan Li,  Chenyang Le,  Manhong Wang,  Yutao Sun,  Xingyang Ma,  Yanmin Qian 
备注:Accepted by Interspeech 2025
摘要:应用像Whisper这样的大型预训练语音模型,在降低各种语音任务的训练成本方面表现出了希望。然而,将这些模型集成到流媒体系统中仍然是一个挑战。本文提出了一种新的前缀到前缀的训练框架,通过微调的耳语流识别。我们引入连续集成和消防机制,建立一个准单调的连续语音序列和离散文本标记之间的对齐。此外,我们设计了单调有限前瞻注意,允许每个令牌参加无限左上下文和有限右上下文的语音序列。我们还采用了wait-k解码策略来简化解码过程,同时确保训练和测试之间的一致性。我们的理论分析和实验表明,这种方法实现了一个可控的延迟和质量之间的权衡,使其适用于各种流媒体应用。
摘要:Applying large pre-trained speech models like Whisper has shown promise in reducing training costs for various speech tasks. However, integrating these models into streaming systems remains a challenge. This paper presents a novel prefix-to-prefix training framework for streaming recognition by fine-tuning the Whisper. We introduce the Continuous Integrate-and-Fire mechanism to establish a quasi-monotonic alignment between continuous speech sequences and discrete text tokens. Additionally, we design Monotonic Finite Look-ahead Attention, allowing each token to attend to infinite left-context and finite right-context from the speech sequences. We also employ the wait-k decoding strategy to simplify the decoding process while ensuring consistency between training and testing. Our theoretical analysis and experiments demonstrate that this approach achieves a controllable trade-off between latency and quality, making it suitable for various streaming applications.


【18】 Efficient Data Selection for Domain Adaptation of ASR Using  Pseudo-Labels and Multi-Stage Filtering

标题: 使用伪标签和多阶段过滤进行ASB域自适应的高效数据选择
链接:https://arxiv.org/abs/2506.03681
作者: Pradeep Rangappa,  Andres Carofilis,  Jeena Prakash,  Shashi Kumar,  Sergio Burdisso,  Srikanth Madikeri,  Esau Villatoro-Tello,  Bidisha Sharma,  Petr Motlicek,  Kadri Hacioglu,  Shankar Venkatesan,  Saurabh Vyas,  Andreas Stolcke 
备注:Accepted at Interspeech 2025, Netherlands
摘要:针对特定领域微调预训练的ASR模型对于具有有限标记数据和计算资源的小型组织来说具有挑战性。在这里,我们探索了不同的数据选择管道,并提出了一种强大的方法,通过过滤使用Whisper(编码器-解码器)和Zipformer(换能器)模型生成的伪标签来改善ASR自适应。我们的方法集成了多种选择策略-包括单词错误率(WER)预测,命名实体识别(NER)和字符错误率(CER)分析-以提取高质量的训练片段。我们使用7500小时的基线对Whisper和Zipformer的方法进行了评估,并将其与基于CER的方法进行了比较,该方法依赖于来自三个ASR系统的假设。对7500小时的伪标记呼叫中心数据进行微调,实现了12.3%的WER,而我们的过滤将数据集减少到100小时(1.4%),具有类似的性能;在Fisher English上观察到类似的趋势。
摘要:Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English.


【19】 Comparative Analysis of Fast and High-Fidelity Neural Vocoders for  Low-Latency Streaming Synthesis in Resource-Constrained Environments

标题: 资源受限环境中用于低延迟流媒体合成的快速和高保真神经声码器的比较分析
链接:https://arxiv.org/abs/2506.03554
作者: Reo Yoneyama,  Masaya Kawamura,  Ryo Terashima,  Ryuichi Yamamoto,  Tomoki Toda 
备注:Accepted to Interspeech 2025
摘要:在实时语音合成中,神经声码器通常需要通过因果处理和流式传输进行低延迟合成。然而,流式传输引入了批量合成中不存在的低效率,诸如有限的并行性、帧间依赖性管理和参数加载开销。本文提出了多流Wavehax(MS-Wavehax),一个有效的神经声码器的低延迟流,通过扩展的混叠的神经声码器Wavehax与多流分解。我们分析了延迟吞吐量权衡在CPU的环境中,并确定流神经声码器的关键瓶颈。我们的研究结果为优化块大小和设计适合特定应用需求和硬件限制的声码器提供了实用的见解。此外,我们的主观评估表明,MS-Wavehax提供高语音质量的因果关系和非因果关系的条件下,同时非常紧凑,易于部署在资源受限的环境中。
摘要:In real-time speech synthesis, neural vocoders often require low-latency synthesis through causal processing and streaming. However, streaming introduces inefficiencies absent in batch synthesis, such as limited parallelism, inter-frame dependency management, and parameter loading overhead. This paper proposes multi-stream Wavehax (MS-Wavehax), an efficient neural vocoder for low-latency streaming, by extending the aliasing-free neural vocoder Wavehax with multi-stream decomposition. We analyze the latency-throughput trade-off in a CPU-only environment and identify key bottlenecks in streaming neural vocoders. Our findings provide practical insights for optimizing chunk sizes and designing vocoders tailored to specific application demands and hardware constraints. Furthermore, our subjective evaluations show that MS-Wavehax delivers high speech quality under causal and non-causal conditions while being remarkably compact and easily deployable in resource-constrained environments.


【20】 Local Equivariance Error-Based Metrics for Evaluating  Sampling-Frequency-Independent Property of Neural Network

标题: 基于局部等方差误差的神经网络采样频率无关性的插值器
链接:https://arxiv.org/abs/2506.03550
作者: Kanami Imamura,  Tomohiko Nakamura,  Norihiro Takamune,  Kohei Yatabe,  Hiroshi Saruwatari 
备注:5 pages, 4 figures, accepted for European Signal Processing Conference 2025 (EUSIPCO 2025)
摘要:基于深度神经网络(DNN)的音频信号处理方法通常仅在单个采样频率(SF)下进行训练,因此需要信号恢复来处理未经训练的SF。然而,最近的研究表明,信号恢复会降低未经训练的SF的性能。这个问题被忽视了,因为大多数研究只评估训练有素的SF的表现。在本文中,为了评估DNN对SF变化的鲁棒性,我们将其称为SF独立(SFI)属性,我们提出了三个度量来量化基于局部等方差误差(LEE)的SFI属性。LEE测量DNN对输入变换的鲁棒性。通过将信号重构作为输入变换,我们将LEE推广到度量音频源分离方法对信号重构的鲁棒性。所提出的指标被构造来量化特定网络组件中负责预测时频掩码的SFI属性。音乐源分离的实验表明,在未经训练的SF所提出的指标和性能下降之间有很强的相关性。
摘要:Audio signal processing methods based on deep neural networks (DNNs) are typically trained only at a single sampling frequency (SF) and therefore require signal resampling to handle untrained SFs. However, recent studies have shown that signal resampling can degrade performance with untrained SFs. This problem has been overlooked because most studies evaluate only the performance at trained SFs. In this paper, to assess the robustness of DNNs to SF changes, which we refer to as the SF-independent (SFI) property, we propose three metrics to quantify the SFI property on the basis of local equivariance error (LEE). LEE measures the robustness of DNNs to input transformations. By using signal resampling as input transformation, we extend LEE to measure the robustness of audio source separation methods to signal resampling. The proposed metrics are constructed to quantify the SFI property in specific network components responsible for predicting time-frequency masks. Experiments on music source separation demonstrated a strong correlation between the proposed metrics and performance degradation at untrained SFs.


机器翻译由腾讯交互翻译提供,仅供参考