今日论文合集:cs.SD语音28篇,eess.AS音频处理35篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Masked Self-distilled Transducer-based Keyword Spotting with  Semi-autoregressive Decoding
标题: 基于半自回归解码的掩蔽自蒸馏传感器的关键词发现
链接:https://arxiv.org/abs/2505.24820
作者: Yu Xi,  Xiaoyu Gu,  Haoyu Li,  Jun Song,  Bo Zheng,  Kai Yu 
摘要:基于RNN-T的自回归解码关键词发现(KWS)由于其流结构和优越的性能而受到关注。然而,RNN-T中预测网络的简单性带来了过拟合问题,特别是在具有挑战性的场景下,导致性能下降。在本文中,我们提出了一种掩蔽自蒸馏(MSD)训练策略,避免RNN-Ts过度依赖预测网络来减轻过拟合。这种训练实现了掩码非自回归(NAR)解码,其在KWS解码期间完全掩蔽RNN-T预测器输出。此外,我们提出了一种半自回归(SAR)解码方法,结合AR和NAR解码的优点。我们在多个KWS数据集上的实验表明,MSD训练有效地消除了过拟合。该方法在保留AR解码优越性能的同时,利用NAR解码的过拟合抑制,取得了良好的效果。
摘要:RNN-T-based keyword spotting (KWS) with autoregressive decoding~(AR) has gained attention due to its streaming architecture and superior performance. However, the simplicity of the prediction network in RNN-T poses an overfitting issue, especially under challenging scenarios, resulting in degraded performance. In this paper, we propose a masked self-distillation (MSD) training strategy that avoids RNN-Ts overly relying on prediction networks to alleviate overfitting. Such training enables masked non-autoregressive (NAR) decoding, which fully masks the RNN-T predictor output during KWS decoding. In addition, we propose a semi-autoregressive (SAR) decoding approach to integrate the advantages of AR and NAR decoding. Our experiments across multiple KWS datasets demonstrate that MSD training effectively alleviates overfitting. The SAR decoding method preserves the superior performance of AR decoding while benefits from the overfitting suppression of NAR decoding, achieving excellent results.


【2】 Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic  Dialect Identification

标题: 语音转换提高阿拉伯语口语识别的跨域稳健性
链接:https://arxiv.org/abs/2505.24713
作者: Badr M. Abdullah,  Matthew Baas,  Bernd Möbius,  Dietrich Klakow 
备注:Accepted in Interspeech 2025
摘要:阿拉伯语方言识别(ADI)系统对于大规模数据收集管道至关重要,这些管道能够为阿拉伯语变体开发包容性语音技术。然而,当前ADI系统的可靠性受限于对域外语音的较差推广。在本文中,我们提出了一种基于语音转换的有效方法来训练ADI模型,该模型具有最先进的性能,并显着提高了跨域场景的鲁棒性。在新收集的跨越四个不同领域的真实世界测试集上进行评估,我们的方法在各个领域的准确性上得到了高达+34.1%的一致改进。此外,我们提出了我们的方法的分析,并证明语音转换有助于减轻ADI数据集中的扬声器偏见。我们发布了强大的ADI模型和跨域评估数据集,以支持阿拉伯语包容性语音技术的开发。
摘要:Arabic dialect identification (ADI) systems are essential for large-scale data collection pipelines that enable the development of inclusive speech technologies for Arabic language varieties. However, the reliability of current ADI systems is limited by poor generalization to out-of-domain speech. In this paper, we present an effective approach based on voice conversion for training ADI models that achieves state-of-the-art performance and significantly improves robustness in cross-domain scenarios. Evaluated on a newly collected real-world test set spanning four different domains, our approach yields consistent improvements of up to +34.1% in accuracy across domains. Furthermore, we present an analysis of our approach and demonstrate that voice conversion helps mitigate the speaker bias in the ADI dataset. We release our robust ADI model and cross-domain evaluation dataset to support the development of inclusive speech technologies for Arabic.


【3】 Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing  Cross-Lingual Transfer in Low-Resource Scenarios

标题: 使用音素增强CoT的语音转文本翻译:在低资源场景中增强跨语言转换
链接:https://arxiv.org/abs/2505.24691
作者: Gerard I. Gállego,  Oriol Pareras,  Martí Cortada Garcia,  Lucas Takanori,  Javier Hernando 
备注:Accepted at Interspeech 2025
摘要:我们提出了一种语音到文本翻译(S2TT)的方法,将音素表示集成到一个思想链(CoT)框架,以提高在低资源和零资源设置的翻译。通过引入音素识别作为中间步骤,我们增强了跨语言迁移,即使是没有标记语音数据的语言也可以进行翻译。我们的系统建立在一个多语言的LLM,我们扩展到处理语音和音素。培训遵循课程学习战略,逐步引入更复杂的任务。在多语言S2TT基准测试上的实验表明,音素增强CoT在低资源条件下提高了翻译质量,并实现了零资源翻译,同时对高资源性能略有影响。尽管存在这种权衡,但我们的研究结果表明,基于音素的CoT是使S2TT在不同语言中更容易使用的有希望的一步。
摘要:We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing phoneme recognition as an intermediate step, we enhance cross-lingual transfer, enabling translation even for languages with no labeled speech data. Our system builds on a multilingual LLM, which we extend to process speech and phonemes. Training follows a curriculum learning strategy that progressively introduces more complex tasks. Experiments on multilingual S2TT benchmarks show that phoneme-augmented CoT improves translation quality in low-resource conditions and enables zero-resource translation, while slightly impacting high-resource performance. Despite this trade-off, our findings demonstrate that phoneme-based CoT is a promising step toward making S2TT more accessible across diverse languages.


【4】 MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised  Domain Adaptation in ASR

标题: MSDA:结合伪标记和自监督实现ASR中的无监督域自适应
链接:https://arxiv.org/abs/2505.24656
作者: Dimitrios Damianos,  Georgios Paraskevopoulos,  Alexandros Potamianos 
摘要:在这项工作中,我们研究了自动语音识别(ASR)的Meta PL无监督域自适应框架。我们介绍了一个多阶段域自适应管道(MSDA),一个样本效率,两阶段的适应方法,集成了自监督学习与半监督技术。MSDA旨在增强ASR模型的鲁棒性和通用性,使其更适应不同的条件。它对于像希腊语这样的低资源语言以及标记数据稀缺或嘈杂的弱监督场景特别有效。通过大量的实验,我们证明了Meta PL可以有效地应用于ASR任务,实现最先进的结果,显着优于最先进的方法,并提供更强大的解决方案,在ASR的无监督域自适应。我们的消融强调了在自我监督与自我训练相结合时利用级联方法的必要性。
摘要:In this work, we investigate the Meta PL unsupervised domain adaptation framework for Automatic Speech Recognition (ASR). We introduce a Multi-Stage Domain Adaptation pipeline (MSDA), a sample-efficient, two-stage adaptation approach that integrates self-supervised learning with semi-supervised techniques. MSDA is designed to enhance the robustness and generalization of ASR models, making them more adaptable to diverse conditions. It is particularly effective for low-resource languages like Greek and in weakly supervised scenarios where labeled data is scarce or noisy. Through extensive experiments, we demonstrate that Meta PL can be applied effectively to ASR tasks, achieving state-of-the-art results, significantly outperforming state-of-the-art methods, and providing more robust solutions for unsupervised domain adaptation in ASR. Our ablations highlight the necessity of utilizing a cascading approach when combining self-supervision with self-training.


【5】 ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis  Optimization for Speech Multi-Metric Estimation

标题: ARECHO:通过基于链的假设优化的语音多指标估计的自回归评估
链接:https://arxiv.org/abs/2505.24518
作者: Jiatong Shi,  Yifan Cheng,  Bo-Hao Su,  Hye-jin Shim,  Jinchuan Tian,  Samuele Cornell,  Yiwen Zhao,  Siddhant Arora,  Shinji Watanabe 
摘要:语音信号分析提出了重大挑战,特别是在语音质量评估和分析等任务中,其目标是预测多个感知和客观指标。例如,PESQ(语音质量的感知评估),STOI(短时客观可懂度)和MOS(平均意见得分)等指标都可以捕获语音质量的不同方面。然而,这些度量通常具有不同的尺度、假设和依赖性,使得联合估计不平凡。为了解决这些问题,我们介绍了ARECHO(自回归评估通过基于链的假设优化),一个基于链的,多功能的语音评估系统,基于自回归依赖模型。ARECHO有三个关键创新:(1)全面的语音信息标记化管道;(2)明确捕获度量间依赖关系的动态分类器链;(3)增强推理可靠性的两步置信度解码算法。实验表明,ARECHO显着优于基线框架在不同的评估方案,包括增强语音分析,语音生成评估,和嘈杂的语音评估。此外,它的动态依赖建模通过捕获度量间关系提高了可解释性。
摘要:Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships.


【6】 MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging  LLM Embedded Knowledge

标题: MELT:利用LLM嵌入式知识实现自动化多模式情感数据注释
链接:https://arxiv.org/abs/2505.24493
作者: Xin Jing,  Jiadong Wang,  Iosif Tsangko,  Andreas Triantafyllopoulos,  Björn W. Schuller 
摘要:虽然语音情感识别(SER)在深度学习方面取得了显著进展,但注释仍然是一个主要障碍。人工标注不仅成本高,而且容易出现不一致性,标注者往往有不同的偏好,可能缺乏必要的上下文知识,这可能导致标签的变化和不准确。与此同时,大型语言模型(LLM)已经成为注释文本数据的可扩展替代方案。然而,LLM在没有人类监督的情况下执行情感语音数据注释的潜力还有待深入研究。为了解决这些问题,我们应用GPT-4 o来注释从情景喜剧《老友记》中收集的多模态数据集,只使用文本线索作为输入。通过制作结构化的文本提示,我们的方法利用了GPT-4 o在训练过程中积累的知识,展示了它可以在不直接访问多模态输入的情况下生成准确且与上下文相关的注释。因此,我们提出了MELT,一个由GPT-4 o完全注释的多模态情感数据集。我们通过微调四个自监督学习(SSL)主干并评估情感数据集上的语音情感识别性能来证明MELT的有效性。此外,我们的主观实验结果表明,一致的性能改善SER。
摘要:Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to inconsistencies annotators often have different preferences and may lack the necessary contextual knowledge, which can lead to varied and inaccurate labels. Meanwhile, Large Language Models (LLMs) have emerged as a scalable alternative for annotating text data. However, the potential of LLMs to perform emotional speech data annotation without human supervision has yet to be thoroughly investigated. To address these problems, we apply GPT-4o to annotate a multimodal dataset collected from the sitcom Friends, using only textual cues as inputs. By crafting structured text prompts, our methodology capitalizes on the knowledge GPT-4o has accumulated during its training, showcasing that it can generate accurate and contextually relevant annotations without direct access to multimodal inputs. Therefore, we propose MELT, a multimodal emotion dataset fully annotated by GPT-4o. We demonstrate the effectiveness of MELT by fine-tuning four self-supervised learning (SSL) backbones and assessing speech emotion recognition performance across emotion datasets. Additionally, our subjective experiments\' results demonstrate a consistence performance improvement on SER.


【7】 Rehearsal with Auxiliary-Informed Sampling for Audio Deepfake Detection

标题: 音频深度伪造检测的辅助抽样排练
链接:https://arxiv.org/abs/2505.24486
作者: Falih Gozi Febrinanto,  Kristen Moore,  Chandra Thapa,  Jiangang Ma,  Vidya Saikrishna,  Feng Xia 
备注:Accepted by Interspeech 2025
摘要:现有的音频deepfake检测框架在面对新的deepfake攻击时性能会下降。基于排练的持续学习(CL)使用有限的一组旧数据样本更新模型,有助于保留先验知识,同时融入新信息。然而,现有的排练技术不能有效地捕捉音频特征的多样性,引入偏见,增加了遗忘的风险。为了应对这一挑战,我们提出了一种基于排练的CL音频深度伪造检测方法-。RAIS采用标签生成网络来生成辅助标签,指导存储缓冲区的不同样本选择。大量的实验表明,RAIS优于最先进的方法,在五个经验中实现了1.953%的平均等错误率(EER)。该代码可从以下网址获得:https://github.com/falihgoz/RAIS。
摘要:The performance of existing audio deepfake detection frameworks degrades when confronted with new deepfake attacks. Rehearsal-based continual learning (CL), which updates models using a limited set of old data samples, helps preserve prior knowledge while incorporating new information. However, existing rehearsal techniques don't effectively capture the diversity of audio characteristics, introducing bias and increasing the risk of forgetting. To address this challenge, we propose Rehearsal with Auxiliary-Informed Sampling (RAIS), a rehearsal-based CL approach for audio deepfake detection. RAIS employs a label generation network to produce auxiliary labels, guiding diverse sample selection for the memory buffer. Extensive experiments show RAIS outperforms state-of-the-art methods, achieving an average Equal Error Rate (EER) of 1.953 % across five experiences. The code is available at: https://github.com/falihgoz/RAIS.


【8】 SuPseudo: A Pseudo-supervised Learning Method for Neural Speech  Enhancement in Far-field Speech Recognition

标题: SuNorton:远场语音识别中神经语音增强的伪监督学习方法
链接:https://arxiv.org/abs/2505.24450
作者: Longjie Luo,  Lin Li,  Qingyang Hong 
备注:Accepted by InterSpeech 2025
摘要:由于在真实记录的远场会话数据集中缺乏目标语音注释,语音增强(SE)模型通常在模拟数据上训练。然而,训练后的模型在现实世界中的表现往往很差,阻碍了它们在远场语音识别中的应用。为了解决这个问题,我们(a)提出了直接声音估计(DSE)来估计SE的真实记录数据的预言直接声音;(b)提出了一种新的伪监督学习方法SuPseudo,它利用DSE估计作为伪标签,使SE模型能够直接从真实记录数据中学习和适应,从而提高其泛化能力。此外,一个名为FARNET的SE模型被设计为充分利用SuPseudo。在MISP 2023语料库上的实验证明了SuPseudo的有效性,我们的系统明显优于以前的最先进的。我们的方法的演示可以在https://EeLLJ.github.io/SuPseudo/上找到。
摘要:Due to the lack of target speech annotations in real-recorded far-field conversational datasets, speech enhancement (SE) models are typically trained on simulated data. However, the trained models often perform poorly in real-world conditions, hindering their application in far-field speech recognition. To address the issue, we (a) propose direct sound estimation (DSE) to estimate the oracle direct sound of real-recorded data for SE; and (b) present a novel pseudo-supervised learning method, SuPseudo, which leverages DSE-estimates as pseudo-labels and enables SE models to directly learn from and adapt to real-recorded data, thereby improving their generalization capability. Furthermore, an SE model called FARNET is designed to fully utilize SuPseudo. Experiments on the MISP2023 corpus demonstrate the effectiveness of SuPseudo, and our system significantly outperforms the previous state-of-the-art. A demo of our method can be found at https://EeLLJ.github.io/SuPseudo/.


【9】 Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the  MISP-Meeting Challenge

标题: MISP会议挑战中AVSR任务的基于伪标签的神经语音增强
链接:https://arxiv.org/abs/2505.24446
作者: Longjie Luo,  Shenghui Lu,  Lin Li,  Qingyang Hong 
备注:Accepted by InterSpeech 2025
摘要:本文介绍了我们的系统MISP会议挑战轨道2。主要的困难在于数据集,其中包含强背景噪声,混响,重叠的语音和不同的会议主题。为了解决这些问题,我们(a)设计了G-SpatialNet,一个语音增强(SE)模型来改善引导源分离(GSS)信号;(b)提出了TLS,一个包括时间对齐、电平对齐和信噪比滤波的框架,用于为真实记录的远场音频数据生成信号级伪标签,从而促进SE模型的训练;以及(c)探索微调策略、数据增强和多模态信息,以提高预先训练的自动语音识别(ASR)模型在会议场景中的性能。最后,我们的系统在Dev和Eval集上分别实现了5.44%和9.52%的字符错误率(CER),比基线分别提高了64.8%和52.6%,获得第二名。
摘要:This paper presents our system for the MISP-Meeting Challenge Track 2. The primary difficulty lies in the dataset, which contains strong background noise, reverberation, overlapping speech, and diverse meeting topics. To address these issues, we (a) designed G-SpatialNet, a speech enhancement (SE) model to improve Guided Source Separation (GSS) signals; (b) proposed TLS, a framework comprising time alignment, level alignment, and signal-to-noise ratio filtering, to generate signal-level pseudo labels for real-recorded far-field audio data, thereby facilitating SE models' training; and (c) explored fine-tuning strategies, data augmentation, and multimodal information to enhance the performance of pre-trained Automatic Speech Recognition (ASR) models in meeting scenarios. Finally, our system achieved character error rates (CERs) of 5.44% and 9.52% on the Dev and Eval sets, respectively, with relative improvements of 64.8% and 52.6% over the baseline, securing second place.


【10】 SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization

标题: Switch Codec:具有稀疏量化的高保真神经音频编解码器
链接:https://arxiv.org/abs/2505.24437
作者: Jin Wang,  Wenbin Jiang,  Xiangbo Wang 
备注:5 pages,4 figures
摘要:我们提出了一个通用的高保真神经音频压缩算法,可以压缩语音,音乐,和一般的音频低于3 kbps的带宽。虽然当前最先进的音频编解码器在音频压缩方面表现出色,但当嵌入空间急剧减少时,其有效性显著下降,这对应于更高的压缩。为了解决这个问题,我们提出了残差专家矢量量化(REVQ),它显着扩展了可用的嵌入空间,提高了性能,同时几乎不牺牲带宽。此外,我们引入了一个策略,以确保巨大的嵌入空间可以得到充分利用。此外,我们提出了一个基于STFT的控制器来指导生成器产生难以区分的频谱图。我们证明,所提出的方法优于基线方法,通过详细的消融。
摘要:We present a universal high-fidelity neural audio compression algorithm that can compress speech, music, and general audio below 3 kbps bandwidth. Although current state-of-the-art audio codecs excel in audio compression, their effectiveness significantly declines when embedding space is sharply reduced, which corresponds to higher compression. To address this problem, we propose Residual Experts Vector Quantization (REVQ), which significantly expands the available embedding space and improves the performance while hardly sacrificing the bandwidth. Furthermore, we introduce a strategy to ensure that the vast embedding space can be fully utilized. Additionally, we propose a STFT-based discriminator to guide the generator in producing indistinguishable spectrograms. We demonstrate that the proposed approach outperforms baseline methods through detailed ablations.


【11】 DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture  Switching for Speech Codec

标题: DS-编解码器:语音编解码器的双阶段训练,采用双镜像架构切换
链接:https://arxiv.org/abs/2505.24314
作者: Peijie Chen,  Wenhao Guan,  Kaidi Wang,  Weijie Wu,  Hukai Huang,  Qingyang Hong,  Lin Li 
备注:Accepted to Interspeech 2025
摘要:神经语音编解码器对于推进文本到语音(TTS)系统至关重要。随着近年来大型语言模型在文本生成方面的成功,开发高质量的语音标记器变得越来越重要。本文介绍了DS-Codec,一种新型的神经语音编解码器,具有镜像和非镜像架构切换的双阶段训练框架,旨在实现卓越的语音重建。我们进行了广泛的实验和消融研究,以评估我们的训练策略的有效性,并比较两种架构的性能。我们的研究结果表明,镜像结构显着提高了学习的码本的鲁棒性,和训练策略之间的优势平衡镜像和非镜像结构,导致改善高保真语音重建。
摘要:Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech tokenizers has become increasingly important. This paper introduces DS-Codec, a novel neural speech codec featuring a dual-stage training framework with mirror and non-mirror architectures switching, designed to achieve superior speech reconstruction. We conduct extensive experiments and ablation studies to evaluate the effectiveness of our training strategy and compare the performance of the two architectures. Our results show that the mirrored structure significantly enhances the robustness of the learned codebooks, and the training strategy balances the advantages between mirrored and non-mirrored structures, leading to improved high-fidelity speech reconstruction.


【12】 Discl-VC: Disentangled Discrete Tokens and In-Context Learning for  Controllable Zero-Shot Voice Conversion

标题: Discl-VC:分离离散令牌和上下文内学习以实现可控Zero-Shot语音转换
链接:https://arxiv.org/abs/2505.24291
作者: Kaidi Wang,  Wenhao Guan,  Ziyue Jiang,  Hukai Huang,  Peijie Chen,  Weijie Wu,  Qingyang Hong,  Lin Li 
摘要:目前,zero-shot语音转换系统能够合成看不见的说话者的语音。然而,大多数现有方法难以准确地复制源说话者的说话风格或模仿目标说话者的独特说话风格,从而限制了语音转换的可控性。在这项工作中,我们提出了Discl-VC,一种新的语音转换框架,从自我监督的语音表示中解开内容和韵律信息,并通过上下文学习与流匹配Transformer合成目标说话人的声音。为了能够精确控制生成的语音的韵律,我们引入了一个掩码生成Transformer,它基于提示以非自回归的方式预测离散的韵律标记。实验结果表明,该方法在zero-shot语音转换中具有良好的性能,在合成语音的韵律控制中具有较高的精度。
摘要:Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive speaking style of the target speaker, thereby limiting the controllability of voice conversion. In this work, we propose Discl-VC, a novel voice conversion framework that disentangles content and prosody information from self-supervised speech representations and synthesizes the target speaker's voice through in-context learning with a flow matching transformer. To enable precise control over the prosody of generated speech, we introduce a mask generative transformer that predicts discrete prosody tokens in a non-autoregressive manner based on prompts. Experimental results demonstrate the superior performance of Discl-VC in zero-shot voice conversion and its remarkable accuracy in prosody control for synthesized speech.


【13】 Dynamic Context-Aware Streaming Pretrained Language Model For Inverse  Text Normalization

标题: 用于反向文本规范化的动态上下文感知流预训练语言模型
链接:https://arxiv.org/abs/2505.24229
作者: Luong Ho,  Khanh Le,  Vinh Pham,  Bao Nguyen,  Tan Tran,  Duc Chau 
备注:Accepted to INTERSPEECH 2025
摘要:反向文本规范化(ITN)对于将口语自动语音识别(ASR)输出转换为格式良好的书面文本,增强可读性和可用性至关重要。尽管流媒体ITN很重要,但由于准确性、效率和适应性方面的挑战,特别是在资源匮乏和上下文有限的场景中,流媒体ITN在流媒体ASR中的集成在很大程度上尚未得到探索。在本文中,我们为ITN引入了一个流预训练语言模型,利用预训练语言表示来提高鲁棒性。为了解决流约束,我们在训练和推理过程中提出了动态上下文感知,从而实现自适应块大小调整和右上下文信息的集成。实验结果表明,我们的方法实现了与非流式ITN相当的准确性,并超过了越南数据集上现有的流式ITN模型,同时保持低延迟,确保无缝集成到ASR系统中。
摘要:Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integration of streaming ITN within streaming ASR remains largely unexplored due to challenges in accuracy, efficiency, and adaptability, particularly in low-resource and limited-context scenarios. In this paper, we introduce a streaming pretrained language model for ITN, leveraging pretrained linguistic representations for improved robustness. To address streaming constraints, we propose Dynamic Context-Aware during training and inference, enabling adaptive chunk size adjustments and the integration of right-context information. Experimental results demonstrate that our method achieves accuracy comparable to non-streaming ITN and surpasses existing streaming ITN models on a Vietnamese dataset, all while maintaining low latency, ensuring seamless integration into ASR systems.


【14】 Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with  Data Augmentation and LID-Aware CTC

标题: 改进ML-SURB 2.0上的多语言语音模型:使用数据增强和LID-Aware CTC进行微调
链接:https://arxiv.org/abs/2505.24200
作者: Qingzheng Wang,  Jiancheng Sun,  Yifan Peng,  Shinji Watanabe 
摘要:使用自监督或监督预训练语音基础模型(SFM)的多语言语音处理在语言识别(LID)和自动语音识别(ASR)等任务上取得了很好的性能。然而,这些模型在微调期间与有限的资源作斗争。本文通过探索多种适应SFM的策略,包括冻结上游训练,部分微调和低秩适应,增强了ML-SUPERB 2.0上的多语言LID和ASR。此外,我们采用数据增强来缓解Few-Shot设置中的性能差距,并引入LID连接时间分类(CTC)损失进行正则化。我们的方法在ML-SUPERB 2.0上实现了LID准确性14%的相对提高和ASR CER 30%的相对降低,在Interspeech 2025 ML-SUPERB 2.0挑战赛中获得第二名。
摘要:Multilingual speech processing with self-supervised or supervised pre-trained Speech Foundation Models (SFM) has achieved strong performance on tasks like Language Identification (LID) and Automatic Speech Recognition (ASR). However, these models struggle with limited resources during fine-tuning. This paper enhances multilingual LID and ASR on ML-SUPERB 2.0 by exploring multiple strategies for adapting SFMs, including frozen upstream training, partial fine-tuning, and low-rank adaptation. Furthermore, we employ data augmentation to mitigate performance gaps in few-shot settings and introduce LID Connectionist Temporal Classification (CTC) loss for regularization. Our approach achieves a 14% relative improvement in LID accuracy and a 30% relative reduction in ASR CER over the baseline on ML-SUPERB 2.0, securing second place in the Interspeech 2025 ML-SUPERB 2.0 Challenge.


【15】 FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing  System

标题: QUALureSense:在始终在线的音频传感系统中保护说话者属性
链接:https://arxiv.org/abs/2505.24115
作者: Bhawana Chhaglani,  Sarmistha Sarna Gomasta,  Yuvraj Agarwal,  Jeremy Gummeson,  Prashant Shenoy 
摘要:音频是一种丰富的感测模态,可用于各种人类活动识别任务。然而,智能手机和带有始终开启麦克风的智能扬声器的无处不在的性质导致了许多隐私问题,并且对部署这些基于音频的传感系统缺乏信任。本文解决了这个关键的挑战,保护用户隐私时,使用音频传感应用程序,同时保持实用程序。虽然以前的工作主要集中在保护可恢复的语音内容,我们表明,敏感的扬声器特定的属性,如年龄和性别仍然可以推断掩蔽语音后,并提出了一个全面的隐私评估框架,以评估这个扬声器属性泄漏。我们设计并实现了一个开源库,提供了一组可推广的隐私感知音频功能,可用于广泛的传感应用。我们提出了一个自适应的特定于任务的功能选择算法,优化的隐私效用成本权衡的基础上的应用程序的要求。通过我们广泛的评估,我们证明了在各种传感任务中的高实用性。我们的系统在保护用户特定隐私方面比现有的隐私技术高出60.6%。这项工作提供了一个基础框架,通过启用有效的隐私感知音频分类系统来确保对音频感测的信任。
摘要:Audio is a rich sensing modality that is useful for a variety of human activity recognition tasks. However, the ubiquitous nature of smartphones and smart speakers with always-on microphones has led to numerous privacy concerns and a lack of trust in deploying these audio-based sensing systems. This paper addresses this critical challenge of preserving user privacy when using audio for sensing applications while maintaining utility. While prior work focuses primarily on protecting recoverable speech content, we show that sensitive speaker-specific attributes such as age and gender can still be inferred after masking speech and propose a comprehensive privacy evaluation framework to assess this speaker attribute leakage. We design and implement FeatureSense, an open-source library that provides a set of generalizable privacy-aware audio features that can be used for wide range of sensing applications. We present an adaptive task-specific feature selection algorithm that optimizes the privacy-utility-cost trade-off based on the application requirements. Through our extensive evaluation, we demonstrate the high utility of FeatureSense across a diverse set of sensing tasks. Our system outperforms existing privacy techniques by 60.6% in preserving user-specific privacy. This work provides a foundational framework for ensuring trust in audio sensing by enabling effective privacy-aware audio classification systems.


【16】 Acoustic Classification of Maritime Vessels using Learnable Filterbanks

标题: 使用可学习滤片组对船舶进行声学分类
链接:https://arxiv.org/abs/2505.23964
作者: Jonas Elsborg,  Tejs Vegge,  Arghya Bhowmik 
备注:9 pages, 5 figures, 2 tables
摘要:基于声学特征可靠地监测和识别海上船只由于不同记录场景的可变性而变得复杂。一个强大的分类框架必须能够概括不同的声学环境和可变的源传感器距离。为此,我们提出了一个深度学习模型,该模型在不同的记录场景中具有强大的性能。使用可训练的频谱前端和时间特征编码器来学习Gabor滤波器组,该模型可以动态地强调不同的频率分量。在格鲁吉亚海峡的VTUAD水听器记录上进行训练后,我们的模型CATFISH在不同的源传感器距离上实现了最先进的96.63%的测试准确率,超过了之前的基准超过12个百分点。我们提出的模型,证明我们的架构选择,分析学习的Gabor滤波器,并进行消融研究传感器数据融合和基于注意力的池。
摘要:Reliably monitoring and recognizing maritime vessels based on acoustic signatures is complicated by the variability of different recording scenarios. A robust classification framework must be able to generalize across diverse acoustic environments and variable source-sensor distances. To this end, we present a deep learning model with robust performance across different recording scenarios. Using a trainable spectral front-end and temporal feature encoder to learn a Gabor filterbank, the model can dynamically emphasize different frequency components. Trained on the VTUAD hydrophone recordings from the Strait of Georgia, our model, CATFISH, achieves a state-of-the-art 96.63 % percent test accuracy across varying source-sensor distances, surpassing the previous benchmark by over 12 percentage points. We present the model, justify our architectural choices, analyze the learned Gabor filters, and perform ablation studies on sensor data fusion and attention-based pooling.


【17】 Patient-Aware Feature Alignment for Robust Lung Sound  Classification:Cohesion-Separation and Global Alignment Losses

标题: 用于鲁棒肺音分类的患者感知特征对齐:内聚分离和全局对齐损失
链接:https://arxiv.org/abs/2505.23834
作者: Seung Gyu Jeong,  Seong Eun Kim 
备注:Accepted INTERSPEECH 2025
摘要:肺音分级对呼吸系统疾病的早期诊断至关重要。然而,生物医学信号往往表现出患者间的差异,甚至在具有相同症状的患者之间,需要考虑个体差异的学习方法。我们提出了一个患者感知特征对齐(PAFA)框架,其中包含两个新的损失,患者凝聚分离损失(PCSL)和全局患者对齐损失(GPAL)。PCSL将同一患者的特征聚类,同时将这些特征与其他患者分离,以捕获患者的变异性,而GPAL将每个患者的质心绘制到全局中心,防止特征空间碎片化。在ICBHI数据集上,该方法取得了较好的效果,四类分类的分类率为64.84%,两类分类的分类率为72.08%.这些发现突出了PAFA捕捉个性化模式的能力,并在不同的患者群中展示了性能增益,为以患者为中心的医疗保健提供了更广泛的应用。
摘要:Lung sound classification is vital for early diagnosis of respiratory diseases. However, biomedical signals often exhibit inter-patient variability even among patients with the same symptoms, requiring a learning approach that considers individual differences. We propose a Patient-Aware Feature Alignment (PAFA) framework with two novel losses, Patient Cohesion-Separation Loss (PCSL) and Global Patient Alignment Loss (GPAL). PCSL clusters features of the same patient while separating those from other patients to capture patient variability, whereas GPAL draws each patient's centroid toward a global center, preventing feature space fragmentation. Our method achieves outstanding results on the ICBHI dataset with a score of 64.84\% for four-class and 72.08\% for two-class classification. These findings highlight PAFA's ability to capture individualized patterns and demonstrate performance gains in distinct patient clusters, offering broader applications for patient-centered healthcare.


【18】 SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks  via Watermarking

标题: SpeechVerification:强大的声学指纹,防止通过水印进行的篡改攻击
链接:https://arxiv.org/abs/2505.23821
作者: Lingfeng Yao (1),  Chenpei Huang (1),  Shengyao Wang (2),  Junpei Xue (2),  Hanqing Guo (3),  Jiang Liu (2),  Xun Chen (4),  Miao Pan (1) ((1) University of Houston (2) Waseda University (3) University of Hawaii at Mānoa (4) Independent Researcher) 
摘要:随着社交媒体的蓬勃发展,恶意篡改的公共言论,特别是来自有影响力的人物的言论,严重影响了社会稳定和公众信任。现有的语音篡改检测方法仍然存在不足:它们要么依赖于外部参考数据,要么对攻击不敏感,对良性操作(如压缩和重定向)不鲁棒。为了应对这些挑战,我们引入了SpeechVerifer,仅使用发布的语音本身来主动验证语音完整性,即,而不需要任何外部参考。受音频指纹和水印的启发,SpeechVerifier可以(i)有效地检测篡改攻击,(ii)对良性操作具有鲁棒性,(iii)仅基于发布的语音验证完整性。简言之,SpeechVerifier利用多尺度特征提取来捕获跨不同时间分辨率的语音特征。然后,它采用对比学习来生成指纹,可以检测不同粒度的修改。这些指纹被设计为对良性操作具有鲁棒性,但在发生恶意篡改时会显示出显著的变化。为了能够以自包含的方式进行语音验证,然后通过分段水印将生成的指纹嵌入到语音信号中。在没有外部参考的情况下,SpeechVerifier可以从发布的音频中检索指纹,并与嵌入的水印进行检查,以验证语音的完整性。大量的实验结果表明,提出的语音验证器是有效的检测篡改攻击和稳健的良性操作。
摘要:With the surge of social media, maliciously tampered public speeches, especially those from influential figures, have seriously affected social stability and public trust. Existing speech tampering detection methods remain insufficient: they either rely on external reference data or fail to be both sensitive to attacks and robust to benign operations, such as compression and resampling. To tackle these challenges, we introduce SpeechVerifer to proactively verify speech integrity using only the published speech itself, i.e., without requiring any external references. Inspired by audio fingerprinting and watermarking, SpeechVerifier can (i) effectively detect tampering attacks, (ii) be robust to benign operations and (iii) verify the integrity only based on published speeches. Briefly, SpeechVerifier utilizes multiscale feature extraction to capture speech features across different temporal resolutions. Then, it employs contrastive learning to generate fingerprints that can detect modifications at varying granularities. These fingerprints are designed to be robust to benign operations, but exhibit significant changes when malicious tampering occurs. To enable speech verification in a self-contained manner, the generated fingerprints are then embedded into the speech signal by segment-wise watermarking. Without external references, SpeechVerifier can retrieve the fingerprint from the published audio and check it with the embedded watermark to verify the integrity of the speech. Extensive experimental results demonstrate that the proposed SpeechVerifier is effective in detecting tampering attacks and robust to benign operations.


【19】 Learning Normal Patterns in Musical Loops

标题: 学习音乐循环中的正常模式
链接:https://arxiv.org/abs/2505.23784
作者: Shayan Dadman,  Bernt Arild Bremdal,  Børre Bang,  Rune Dalmo 
备注:27 pages, 10 figures
摘要:本文介绍了一个无监督的框架,通过异常检测技术检测音乐样本(循环)中的音频模式,解决音乐信息检索(MIR)的挑战。现有的方法往往受到手工制作的功能,特定领域的限制,或依赖于迭代的用户交互的依赖。我们通过将深度特征提取与无监督异常检测相结合的架构来解决这些限制。我们的方法利用预先训练的分层令牌语义音频Transformer(HTS-AT),与特征融合机制(FFM)配对,从可变长度的音频循环生成表示。这些嵌入使用单类深度支持向量数据描述(Deep SVDD)进行处理,该描述通过将标准音频模式映射到紧凑的潜在超球体来学习标准音频模式。对策划的低音和吉他数据集的评估将标准和残差自动编码器变体与隔离森林(IF)和主成分分析(PCA)方法等基线进行比较。结果表明,我们的Deep SVDD模型,特别是残差自动编码器变体,提供了改进的异常分离,特别是对于较大的变化。这项研究为处理不同的音频样本提供了一种灵活的、完全无监督的解决方案,克服了以前的结构和输入限制,同时通过基于距离的潜在空间评分实现了有效的模式识别。
摘要:This paper introduces an unsupervised framework for detecting audio patterns in musical samples (loops) through anomaly detection techniques, addressing challenges in music information retrieval (MIR). Existing methods are often constrained by reliance on handcrafted features, domain-specific limitations, or dependence on iterative user interaction. We address these limitations through an architecture combining deep feature extraction with unsupervised anomaly detection. Our approach leverages a pre-trained Hierarchical Token-semantic Audio Transformer (HTS-AT), paired with a Feature Fusion Mechanism (FFM), to generate representations from variable-length audio loops. These embeddings are processed using one-class Deep Support Vector Data Description (Deep SVDD), which learns normative audio patterns by mapping them to a compact latent hypersphere. Evaluations on curated bass and guitar datasets compare standard and residual autoencoder variants against baselines like Isolation Forest (IF) and and principle component analysis (PCA) methods. Results show our Deep SVDD models, especially the residual autoencoder variant, deliver improved anomaly separation, particularly for larger variations. This research contributes a flexible, fully unsupervised solution for processing diverse audio samples, overcoming previous structural and input limitations while enabling effective pattern identification through distance-based latent space scoring.


【20】 4,500 Seconds: Small Data Training Approaches for Deep UAV Audio  Classification

标题: 4,500秒:深度无人机音频分类的小数据训练方法
链接:https://arxiv.org/abs/2505.23782
作者: Andrew P. Berg,  Qian Zhang,  Mia Y. Wang 
备注:Accepted at the 14th International Conference on Data Science, Technology, and Applications (DATA), 2025
摘要:无人机的使用预计将在未来十年激增,这就需要加强安全措施,以防止侵犯领空和安全威胁。本研究探讨了无人机分类的深度学习方法,重点关注数据稀缺性这一关键问题。为了研究这一点,我们选择使用总共4,500秒的音频样本来训练模型,这些音频样本均匀分布在9类数据集中。我们利用参数有效微调(PEFT)和数据增强来缓解数据稀缺性。本文实现并比较了卷积神经网络(CNN)和基于注意力的Transformers的使用。我们的研究结果表明,CNN的准确率比Transformers高1- 2%,同时仍然具有更高的计算效率。然而,这些早期的发现指出了使用Transformers模型的潜力;这表明,通过更多的数据和进一步的优化,它们可以胜过CNN。未来的工作旨在扩大数据集,以更好地了解这些方法之间的权衡。
摘要:Unmanned aerial vehicle (UAV) usage is expected to surge in the coming decade, raising the need for heightened security measures to prevent airspace violations and security threats. This study investigates deep learning approaches to UAV classification focusing on the key issue of data scarcity. To investigate this we opted to train the models using a total of 4,500 seconds of audio samples, evenly distributed across a 9-class dataset. We leveraged parameter efficient fine-tuning (PEFT) and data augmentations to mitigate the data scarcity. This paper implements and compares the use of convolutional neural networks (CNNs) and attention-based transformers. Our results show that, CNNs outperform transformers by 1-2\% accuracy, while still being more computationally efficient. These early findings, however, point to potential in using transformers models; suggesting that with more data and further optimizations they could outperform CNNs. Future works aims to upscale the dataset to better understand the trade-offs between these approaches.


【21】 Unified AI for Accurate Audio Anomaly Detection

标题: 统一人工智能用于准确的音频异常检测
链接:https://arxiv.org/abs/2505.23781
作者: Hamideh Khaleghpour,  Brett McKinney 
备注:6 pages, 14 figures. Based on original research. Submitted to arXiv for public preprint
摘要:本文提出了一个统一的AI框架,通过集成先进的降噪,特征提取和机器学习建模技术,高精度的音频异常检测。该方法结合了频谱减法和自适应滤波来增强音频质量,然后使用传统方法(如MFCC)和深度嵌入预训练模型(如OpenL3)进行特征提取。建模管道结合了经典模型(SVM,随机森林),深度学习架构(CNN)和集成方法,以提高鲁棒性和准确性。在包括TORGO和LibriSpeech在内的基准数据集上进行评估,该框架在精确度,召回率和模糊语音与正常语音的分类方面表现出优异的性能。这项工作解决了噪声环境和实时应用中的挑战,并为基于音频的异常检测提供了一个可扩展的解决方案。
摘要:This paper presents a unified AI framework for high-accuracy audio anomaly detection by integrating advanced noise reduction, feature extraction, and machine learning modeling techniques. The approach combines spectral subtraction and adaptive filtering to enhance audio quality, followed by feature extraction using traditional methods like MFCCs and deep embeddings from pre-trained models such as OpenL3. The modeling pipeline incorporates classical models (SVM, Random Forest), deep learning architectures (CNNs), and ensemble methods to boost robustness and accuracy. Evaluated on benchmark datasets including TORGO and LibriSpeech, the proposed framework demonstrates superior performance in precision, recall, and classification of slurred vs. normal speech. This work addresses challenges in noisy environments and real-time applications and provides a scalable solution for audio-based anomaly detection.


【22】 More-than-Human Storytelling: Designing Longitudinal Narrative  Engagements with Generative AI

标题: 比人类更多的讲故事:用生成性人工智能设计纵向叙事参与
链接:https://arxiv.org/abs/2505.23780
作者: Émilie Fabre,  Katie Seaborn,  Shuta Koiwai,  Mizuki Watanabe,  Paul Riesch 
备注:CHI EA '25
摘要:与生成式AI(GenAI)讲故事代理的纵向参与是一个及时但不太明确的领域。我们使用“Dreamsmithy”探索了多代人的体验,这是一个日常的梦想制作应用程序,参与者(N = 28)每天与AI叙述者“Makoto”共同创作故事。通过为期两周的日记研究捕捉反思和互动。自反性主题分析揭示了诸如“摇摆不定的矛盾心理”和“社会时间关系”等主题,突出了随着时间的推移,个人和人工智能叙述者之间出现的复杂动态。研究结果表明,虽然人们欣赏个人笔记,反思的机会和人工智能的创造力,但叙事连贯性和控制的局限性偶尔会导致挫折感。这些结果强调了GenAI在纵向讲故事方面的潜力,但也提出了关于用户代理和道德的关键问题。我们为开发自适应的、超越人类的讲故事系统提供了初步的经验见解和设计考虑。
摘要:Longitudinal engagement with generative AI (GenAI) storytelling agents is a timely but less charted domain. We explored multi-generational experiences with "Dreamsmithy," a daily dream-crafting app, where participants (N = 28) co-created stories with AI narrator "Makoto" every day. Reflections and interactions were captured through a two-week diary study. Reflexive thematic analysis revealed themes likes "oscillating ambivalence" and "socio-chronological bonding," highlighting the complex dynamics that emerged between individuals and the AI narrator over time. Findings suggest that while people appreciated the personal notes, opportunities for reflection, and AI creativity, limitations in narrative coherence and control occasionally caused frustration. The results underscore the potential of GenAI for longitudinal storytelling, but also raise critical questions about user agency and ethics. We contribute initial empirical insights and design considerations for developing adaptive, more-than-human storytelling systems.


【23】 Identifying Primary Stress Across Related Languages and Dialects with  Transformer-based Speech Encoder Models

标题: 使用基于转换器的语音编码器模型识别相关语言和方言的主要压力
链接:https://arxiv.org/abs/2505.24571
作者: Nikola Ljubešić,  Ivan Porupski,  Peter Rupnik 
备注:Accepted to InterSpeech2025
摘要:由于重音在编码意义和辅助言语理解中的作用,自动化主重音识别一直是一个活跃的研究领域。以前的研究主要依赖于传统的声学特征和英语数据集。在本文中,我们研究的方法微调预训练的Transformer模型与音频帧分类头。我们的实验使用了一个新的克罗地亚语训练数据集,测试集包括克罗地亚语、塞尔维亚语、Chakavian方言和斯洛文尼亚语。通过比较SVM分类器使用传统的声学功能与微调语音Transformer,我们证明了Transformer的优势,全面,实现了近乎完美的结果克罗地亚和塞尔维亚,与10点的性能下降,更遥远的Chakavian和斯洛文尼亚。最后,我们表明,只有几百个多音节训练词就足以获得出色的表现。我们在许可证下发布数据集和模型。
摘要:Automating primary stress identification has been an active research field due to the role of stress in encoding meaning and aiding speech comprehension. Previous studies relied mainly on traditional acoustic features and English datasets. In this paper, we investigate the approach of fine-tuning a pre-trained transformer model with an audio frame classification head. Our experiments use a new Croatian training dataset, with test sets in Croatian, Serbian, the Chakavian dialect, and Slovenian. By comparing an SVM classifier using traditional acoustic features with the fine-tuned speech transformer, we demonstrate the transformer's superiority across the board, achieving near-perfect results for Croatian and Serbian, with a 10-point performance drop for the more distant Chakavian and Slovenian. Finally, we show that only a few hundred multi-syllabic training words suffice for strong performance. We release our datasets and model under permissive licenses.


【24】 Pretraining Multi-Speaker Identification for Neural Speaker Diarization

标题: 用于神经说话人二进制化的预训练多说话人识别
链接:https://arxiv.org/abs/2505.24545
作者: Shota Horiguchi,  Atsushi Ando,  Marc Delcroix,  Naohiro Tawara 
备注:Accepted to Interspeech 2025
摘要:端到端说话人日记化通过并行地联合估计多个说话人的语音活动来实现精确的语音感知日记化。这种方法需要大量的数据,需要大量的标记会话数据,而这些数据不能完全从真实数据集中获得。为了解决这个问题,大规模模拟数据通常用于预训练,但它需要巨大的存储和I/O容量,并且模拟与真实对话非常相似的数据仍然具有挑战性。在本文中,我们提出了一个预训练模型,以确定多个扬声器从输入完全重叠的混合物作为一种替代预训练的日记模型。这种方法消除了准备大规模模拟数据集的需要,同时利用大规模说话人识别数据集进行训练。通过全面的实验,我们证明了该方法能够实现高度准确但轻量级的本地日志化模型,而无需模拟会话数据。
摘要:End-to-end speaker diarization enables accurate overlap-aware diarization by jointly estimating multiple speakers' speech activities in parallel. This approach is data-hungry, requiring a large amount of labeled conversational data, which cannot be fully obtained from real datasets alone. To address this issue, large-scale simulated data is often used for pretraining, but it requires enormous storage and I/O capacity, and simulating data that closely resembles real conversations remains challenging. In this paper, we propose pretraining a model to identify multiple speakers from an input fully overlapped mixture as an alternative to pretraining a diarization model. This method eliminates the need to prepare a large-scale simulated dataset while leveraging large-scale speaker recognition datasets for training. Through comprehensive experiments, we demonstrate that the proposed method enables a highly accurate yet lightweight local diarization model without simulated conversational data.


【25】 When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from  Human to Animal and Designed Sounds

标题: 当人类咆哮和鸟类说话时:从人类到动物的高保真声音转换和设计声音
链接:https://arxiv.org/abs/2505.24336
作者: Minsu Kang,  Seolhee Lee,  Choonghyeon Lee,  Namhyun Cho 
备注:INTERSPEECH 2025 accepted
摘要:人类到非人类语音转换(H2 NH-VC)将人类语音转换为动物或设计的发声。与之前专注于狗的声音和16或22.05kHz音频转换的研究不同,这项工作涉及更广泛的非语音声音,包括自然声音(狮子咆哮,鸟鸣)和设计的声音(合成咆哮)。为了适应各种非语音声音的生成和44.1kHz的高质量音频转换,我们引入了一个预处理管道和一个改进的基于CVAE的H2 NH-VC模型,两者都针对人类和非人类语音进行了优化。实验结果表明,该方法在质量、自然度和相似度MOS方面优于基线,实现了不同非人音色的有效语音转换。演示示例可在https://nc-ai.github.io/speech/publications/nonhuman-vc/上获得
摘要:Human to non-human voice conversion (H2NH-VC) transforms human speech into animal or designed vocalizations. Unlike prior studies focused on dog-sounds and 16 or 22.05kHz audio transformation, this work addresses a broader range of non-speech sounds, including natural sounds (lion-roars, birdsongs) and designed voice (synthetic growls). To accomodate generation of diverse non-speech sounds and 44.1kHz high-quality audio transformation, we introduce a preprocessing pipeline and an improved CVAE-based H2NH-VC model, both optimized for human and non-human voices. Experimental results showed that the proposed method outperformed baselines in quality, naturalness, and similarity MOS, achieving effective voice conversion across diverse non-human timbres. Demo samples are available at https://nc-ai.github.io/speech/publications/nonhuman-vc/


【26】 A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a  Rater's Shadowing and Sequence-to-sequence Voice Conversion

标题: 基于感知的L2语音可懂度指标:利用评分者的影子和序列到序列语音转换
链接:https://arxiv.org/abs/2505.24304
作者: Haopeng Geng,  Daisuke Saito,  Nobuaki Minematsu 
备注:Accepted by Interspeech 2025
摘要:二语语音可懂度的评估对于有效的计算机辅助语言学习(CALL)至关重要。传统的基于ASR的方法通常专注于本地相似性,这可能无法捕获人类听众感知的实际可理解性。相比之下,我们的工作介绍了一种新的,基于感知的L2语音清晰度指标,利用一个本地评分员的阴影数据内的序列到序列(seq2seq)的语音转换框架。通过整合对齐机制和声学特征重建,我们的方法模拟了本族听众的听觉感知,识别L2语音中可能导致理解困难的片段。客观和主观的评估表明,我们的方法更接近本地的判断比传统的基于ASR的指标,提供了一个有前途的新方向的CALL系统在全球范围内,多语言环境。
摘要:Evaluating L2 speech intelligibility is crucial for effective computer-assisted language learning (CALL). Conventional ASR-based methods often focus on native-likeness, which may fail to capture the actual intelligibility perceived by human listeners. In contrast, our work introduces a novel, perception based L2 speech intelligibility indicator that leverages a native rater's shadowing data within a sequence-to-sequence (seq2seq) voice conversion framework. By integrating an alignment mechanism and acoustic feature reconstruction, our approach simulates the auditory perception of native listeners, identifying segments in L2 speech that are likely to cause comprehension difficulties. Both objective and subjective evaluations indicate that our method aligns more closely with native judgments than traditional ASR-based metrics, offering a promising new direction for CALL systems in a global, multilingual contexts.


【27】 Probing the Robustness Properties of Neural Speech Codecs

标题: 探索神经语音编解码器的鲁棒性
链接:https://arxiv.org/abs/2505.24248
作者: Wei-Cheng Tseng,  David Harwath 
备注:Interspeech 2025
摘要:神经语音编解码器彻底改变了语音编码,在保持音频保真度的同时实现了更高的压缩。除了压缩之外,它们还作为标记化策略出现,实现了对语音的语言建模,并在各种语音处理任务中推动范式转变。尽管有这些进步,但它们在嘈杂环境中的鲁棒性仍然没有得到充分的探索,这引起了人们对它们在现实世界中的推广的担忧。在这项工作中,我们系统地评估了各种噪声条件下的神经语音编解码器,揭示了它们的鲁棒性的重要差异。我们进一步研究了它们的线性特性,揭示了非线性失真,这部分解释了观测到的鲁棒性变化。最后,我们分析了它们的频率响应,以确定影响音频保真度的因素。我们的研究结果为编解码器行为和未来的编解码器设计提供了重要的见解,并强调了噪声鲁棒性对其现实世界集成的重要性。
摘要:Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategies, enabling language modeling on speech and driving paradigm shifts across various speech processing tasks. Despite these advancements, their robustness in noisy environments remains underexplored, raising concerns about their generalization to real-world scenarios. In this work, we systematically evaluate neural speech codecs under various noise conditions, revealing non-trivial differences in their robustness. We further examine their linearity properties, uncovering non-linear distortions which partly explain observed variations in robustness. Lastly, we analyze their frequency response to identify factors affecting audio fidelity. Our findings provide critical insights into codec behavior and future codec design, as well as emphasizing the importance of noise robustness for their real-world integration.


【28】 Can Emotion Fool Anti-spoofing?

标题: 情感可以愚弄反欺骗吗?
链接:https://arxiv.org/abs/2505.23962
作者: Aurosweta Mahapatra,  Ismail Rasim Ulgen,  Abinay Reddy Naini,  Carlos Busso,  Berrak Sisman 
备注:Accepted to Interspeech 2025
摘要:传统的反欺骗侧重于基于合成语音的模型和数据集,这些语音大多处于中性状态,忽略了各种情绪变化。因此,它们对高质量、情感表达的合成语音的鲁棒性是不确定的。我们解决这个问题,通过引入语音欺骗TTS,一个语料库的情感文本到语音样本。我们的分析表明,现有的反欺骗模型与情感合成语音作斗争,暴露了针对情感的攻击的风险。即使在情绪数据上进行训练,由于对情绪方面的关注有限,模型表现不佳,并且显示出情绪之间的性能差异。这突出了在数据集和方法中需要以情感为中心的反欺骗范例。我们提出了GEM,一个门控集成的情感专用模型与语音情感识别门控网络。GEM在所有情绪和中性状态下都能有效执行,提高了对欺骗攻击的防御能力。我们发布了Spoof-TTS数据集:https://emospoof-tts.github.io/Dataset/
摘要:Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness against high-quality, emotionally expressive synthetic speech is uncertain. We address this by introducing EmoSpoof-TTS, a corpus of emotional text-to-speech samples. Our analysis shows existing anti-spoofing models struggle with emotional synthetic speech, exposing risks of emotion-targeted attacks. Even trained on emotional data, the models underperform due to limited focus on emotional aspect and show performance disparities across emotions. This highlights the need for emotion-focused anti-spoofing paradigm in both dataset and methodology. We propose GEM, a gated ensemble of emotion-specialized models with a speech emotion recognition gating network. GEM performs effectively across all emotions and neutral state, improving defenses against spoofing attacks. We release the EmoSpoof-TTS Dataset: https://emospoof-tts.github.io/Dataset/


eess.AS音频处理


【1】 "Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking

链接:https://arxiv.org/abs/2505.24736
作者: Julio Cesar Cavalcanti,  Gabriel Skantze 
备注:Accepted to Interspeech 2025
摘要:对话中的话轮转换受到普遍的限制,但也有很大的差异。本研究使用大量的美国英语会话数据集(Fisher),考察了人口统计学(性别、年龄、教育)和个人因素如何影响话轮转换。我们分析了过渡地板偏移(TFO),并发现显着的话者间的变化。性别和年龄的影响虽小,但却很显著,女性和年长者的偏移量略短,而教育则没有影响。轻松的话题与较短的TFO相关。然而,个体差异有更大的影响,由一个强大的特质和一个更强大的“二元”成分驱动-二元体中的说话者彼此相似,而不是他们在不同的二元体中相似。这表明,二元关系和联合活动是TFO的最强决定因素,超过了人口的影响。
摘要:Turn-taking in dialogue follows universal constraints but also varies significantly. This study examines how demographic (sex, age, education) and individual factors shape turn-taking using a large dataset of US English conversations (Fisher). We analyze Transition Floor Offset (TFO) and find notable interspeaker variation. Sex and age have small but significant effects female speakers and older individuals exhibit slightly shorter offsets - while education shows no effect. Lighter topics correlate with shorter TFOs. However, individual differences have a greater impact, driven by a strong idiosyncratic and an even stronger "dyadosyncratic" component - speakers in a dyad resemble each other more than they resemble themselves in different dyads. This suggests that the dyadic relationship and joint activity are the strongest determinants of TFO, outweighing demographic influences.


【2】 A Composite Predictive-Generative Approach to Monaural Universal Speech  Enhancement

标题: 单耳通用语音增强的复合预测生成方法
链接:https://arxiv.org/abs/2505.24576
作者: Jie Zhang,  Haoyin Yan,  Xiaofei Li 
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing
摘要:有希望设计出一种能够抑制各种失真并提高语音质量的单一模型,即,通用语音增强(USE)。与基于监督学习的预测方法相比,基于扩散的生成模型显示出更大的潜力,这是由于具有严重受损信息的退化语音的生成能力。然而,在非常不利的条件下可能会引入伪影,并且扩散模型通常由于许多推理步骤而遭受沉重的计算负担。为了联合利用预测和生成的优势,克服各自的缺陷,在这项工作中,我们提出了一个通用的语音增强模型称为PGUSE相结合的预测和生成建模。我们的模型由两个分支组成:预测分支直接从退化信号中预测干净样本,而生成分支优化扩散模型的去噪目标。我们利用输出融合和截断扩散方案来有效地集成预测和生成建模,其中前者直接结合来自两个分支的结果,后者用来自预测分支的初始估计修改反向扩散过程。在几个数据集上进行的大量实验验证了所提出的模型优于最先进的基线,证明了预测和生成建模相结合的互补性和优势。
摘要:It is promising to design a single model that can suppress various distortions and improve speech quality, i.e., universal speech enhancement (USE). Compared to supervised learning-based predictive methods, diffusion-based generative models have shown greater potential due to the generative capacities from degraded speech with severely damaged information. However, artifacts may be introduced in highly adverse conditions, and diffusion models often suffer from a heavy computational burden due to many steps for inference. In order to jointly leverage the superiority of prediction and generation and overcome the respective defects, in this work we propose a universal speech enhancement model called PGUSE by combining predictive and generative modeling. Our model consists of two branches: the predictive branch directly predicts clean samples from degraded signals, while the generative branch optimizes the denoising objective of diffusion models. We utilize the output fusion and truncated diffusion scheme to effectively integrate predictive and generative modeling, where the former directly combines results from both branches and the latter modifies the reverse diffusion process with initial estimates from the predictive branch. Extensive experiments on several datasets verify the superiority of the proposed model over state-of-the-art baselines, demonstrating the complementarity and benefits of combining predictive and generative modeling.


【3】 Identifying Primary Stress Across Related Languages and Dialects with  Transformer-based Speech Encoder Models

标题: 使用基于转换器的语音编码器模型识别相关语言和方言的主要压力
链接:https://arxiv.org/abs/2505.24571
作者: Nikola Ljubešić,  Ivan Porupski,  Peter Rupnik 
备注:Accepted to InterSpeech2025
摘要:由于重音在编码意义和辅助言语理解中的作用,自动化主重音识别一直是一个活跃的研究领域。以前的研究主要依赖于传统的声学特征和英语数据集。在本文中,我们研究的方法微调预训练的Transformer模型与音频帧分类头。我们的实验使用了一个新的克罗地亚语训练数据集,测试集包括克罗地亚语、塞尔维亚语、Chakavian方言和斯洛文尼亚语。通过比较SVM分类器使用传统的声学功能与微调语音Transformer,我们证明了Transformer的优势,全面,实现了近乎完美的结果克罗地亚和塞尔维亚,与10点的性能下降,更遥远的Chakavian和斯洛文尼亚。最后,我们表明,只有几百个多音节训练词就足以获得出色的表现。我们在许可证下发布数据集和模型。
摘要:Automating primary stress identification has been an active research field due to the role of stress in encoding meaning and aiding speech comprehension. Previous studies relied mainly on traditional acoustic features and English datasets. In this paper, we investigate the approach of fine-tuning a pre-trained transformer model with an audio frame classification head. Our experiments use a new Croatian training dataset, with test sets in Croatian, Serbian, the Chakavian dialect, and Slovenian. By comparing an SVM classifier using traditional acoustic features with the fine-tuned speech transformer, we demonstrate the transformer's superiority across the board, achieving near-perfect results for Croatian and Serbian, with a 10-point performance drop for the more distant Chakavian and Slovenian. Finally, we show that only a few hundred multi-syllabic training words suffice for strong performance. We release our datasets and model under permissive licenses.


【4】 Pretraining Multi-Speaker Identification for Neural Speaker Diarization

标题: 用于神经说话人二进制化的预训练多说话人识别
链接:https://arxiv.org/abs/2505.24545
作者: Shota Horiguchi,  Atsushi Ando,  Marc Delcroix,  Naohiro Tawara 
备注:Accepted to Interspeech 2025
摘要:端到端说话人日记化通过并行地联合估计多个说话人的语音活动来实现精确的语音感知日记化。这种方法需要大量的数据,需要大量的标记会话数据,而这些数据不能完全从真实数据集中获得。为了解决这个问题,大规模模拟数据通常用于预训练,但它需要巨大的存储和I/O容量,并且模拟与真实对话非常相似的数据仍然具有挑战性。在本文中,我们提出了一个预训练模型,以确定多个扬声器从输入完全重叠的混合物作为一种替代预训练的日记模型。这种方法消除了准备大规模模拟数据集的需要,同时利用大规模说话人识别数据集进行训练。通过全面的实验,我们证明了该方法能够实现高度准确但轻量级的本地日志化模型,而无需模拟会话数据。
摘要:End-to-end speaker diarization enables accurate overlap-aware diarization by jointly estimating multiple speakers' speech activities in parallel. This approach is data-hungry, requiring a large amount of labeled conversational data, which cannot be fully obtained from real datasets alone. To address this issue, large-scale simulated data is often used for pretraining, but it requires enormous storage and I/O capacity, and simulating data that closely resembles real conversations remains challenging. In this paper, we propose pretraining a model to identify multiple speakers from an input fully overlapped mixture as an alternative to pretraining a diarization model. This method eliminates the need to prepare a large-scale simulated dataset while leveraging large-scale speaker recognition datasets for training. Through comprehensive experiments, we demonstrate that the proposed method enables a highly accurate yet lightweight local diarization model without simulated conversational data.


【5】 Speech Token Prediction via Compressed-to-fine Language Modeling for  Speech Generation

标题: 通过压缩到精细语言建模的语音标记预测语音生成
链接:https://arxiv.org/abs/2505.24496
作者: Wenrui Liu,  Qian Chen,  Wen Wang,  Yafeng Chen,  Jin Xu,  Zhifang Guo,  Guanrou Yang,  Weiqin Li,  Xiaoda Yang,  Tao Jin,  Minghui Fang,  Jialong Zuo,  Bai Jionghao,  Zemin Liu 
摘要:神经音频编解码器作为语音标记器,在语音生成领域显示出巨大的潜力。然而,为了确保高保真音频重建,神经音频编解码器通常将音频编码为长序列的语音令牌,这对长上下文建模中的下游语言模型构成了重大挑战。我们观察到,语音令牌序列表现出短程依赖性:由于文本到语音(TTS)任务中的文本和语音之间的单调对齐,当前令牌的预测主要依赖于其本地上下文,而远程令牌对当前令牌预测的贡献较小,并且通常包含冗余信息。受此观察的启发,我们提出了一种\textbf{压缩到精细语言建模}方法来解决神经编解码器语言模型中长序列语音令牌的挑战:(1)\textbf{细粒度初始和短程信息}:我们的方法在预测期间保留提示和本地令牌,以确保文本对齐和非语言信息的完整性;(2)\textbf{压缩远程上下文}:我们的方法压缩长距离令牌跨度到紧凑的表示,以减少冗余信息,同时保留基本的语义。对各种神经音频编解码器和下游语言模型的广泛实验验证了所提出的方法的有效性和通用性,突出了令牌压缩在改善神经编解码器语言模型中的语音生成方面的重要性。音频样本的演示将在https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM上提供。
摘要:Neural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio into long sequences of speech tokens, posing a significant challenge for downstream language models in long-context modeling. We observe that speech token sequences exhibit short-range dependency: due to the monotonic alignment between text and speech in text-to-speech (TTS) tasks, the prediction of the current token primarily relies on its local context, while long-range tokens contribute less to the current token prediction and often contain redundant information. Inspired by this observation, we propose a \textbf{compressed-to-fine language modeling} approach to address the challenge of long sequence speech tokens within neural codec language models: (1) \textbf{Fine-grained Initial and Short-range Information}: Our approach retains the prompt and local tokens during prediction to ensure text alignment and the integrity of paralinguistic information; (2) \textbf{Compressed Long-range Context}: Our approach compresses long-range token spans into compact representations to reduce redundant information while preserving essential semantics. Extensive experiments on various neural audio codecs and downstream language models validate the effectiveness and generalizability of the proposed approach, highlighting the importance of token compression in improving speech generation within neural codec language models. The demo of audio samples will be available at https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM.


【6】 When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from  Human to Animal and Designed Sounds

标题: 当人类咆哮和鸟类说话时:从人类到动物的高保真声音转换和设计声音
链接:https://arxiv.org/abs/2505.24336
作者: Minsu Kang,  Seolhee Lee,  Choonghyeon Lee,  Namhyun Cho 
备注:INTERSPEECH 2025 accepted
摘要:人类到非人类语音转换(H2 NH-VC)将人类语音转换为动物或设计的发声。与之前专注于狗的声音和16或22.05kHz音频转换的研究不同,这项工作涉及更广泛的非语音声音,包括自然声音(狮子咆哮,鸟鸣)和设计的声音(合成咆哮)。为了适应各种非语音声音的生成和44.1kHz的高质量音频转换,我们引入了一个预处理管道和一个改进的基于CVAE的H2 NH-VC模型,两者都针对人类和非人类语音进行了优化。实验结果表明,该方法在质量、自然度和相似度MOS方面优于基线,实现了不同非人音色的有效语音转换。演示示例可在https://nc-ai.github.io/speech/publications/nonhuman-vc/上获得
摘要:Human to non-human voice conversion (H2NH-VC) transforms human speech into animal or designed vocalizations. Unlike prior studies focused on dog-sounds and 16 or 22.05kHz audio transformation, this work addresses a broader range of non-speech sounds, including natural sounds (lion-roars, birdsongs) and designed voice (synthetic growls). To accomodate generation of diverse non-speech sounds and 44.1kHz high-quality audio transformation, we introduce a preprocessing pipeline and an improved CVAE-based H2NH-VC model, both optimized for human and non-human voices. Experimental results showed that the proposed method outperformed baselines in quality, naturalness, and similarity MOS, achieving effective voice conversion across diverse non-human timbres. Demo samples are available at https://nc-ai.github.io/speech/publications/nonhuman-vc/


【7】 A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a  Rater's Shadowing and Sequence-to-sequence Voice Conversion

标题: 基于感知的L2语音可懂度指标:利用评分者的影子和序列到序列语音转换
链接:https://arxiv.org/abs/2505.24304
作者: Haopeng Geng,  Daisuke Saito,  Nobuaki Minematsu 
备注:Accepted by Interspeech 2025
摘要:二语语音可懂度的评估对于有效的计算机辅助语言学习(CALL)至关重要。传统的基于ASR的方法通常专注于本地相似性,这可能无法捕获人类听众感知的实际可理解性。相比之下,我们的工作介绍了一种新的,基于感知的L2语音清晰度指标,利用一个本地评分员的阴影数据内的序列到序列(seq2seq)的语音转换框架。通过整合对齐机制和声学特征重建,我们的方法模拟了本族听众的听觉感知,识别L2语音中可能导致理解困难的片段。客观和主观的评估表明,我们的方法更接近本地的判断比传统的基于ASR的指标,提供了一个有前途的新方向的CALL系统在全球范围内,多语言环境。
摘要:Evaluating L2 speech intelligibility is crucial for effective computer-assisted language learning (CALL). Conventional ASR-based methods often focus on native-likeness, which may fail to capture the actual intelligibility perceived by human listeners. In contrast, our work introduces a novel, perception based L2 speech intelligibility indicator that leverages a native rater's shadowing data within a sequence-to-sequence (seq2seq) voice conversion framework. By integrating an alignment mechanism and acoustic feature reconstruction, our approach simulates the auditory perception of native listeners, identifying segments in L2 speech that are likely to cause comprehension difficulties. Both objective and subjective evaluations indicate that our method aligns more closely with native judgments than traditional ASR-based metrics, offering a promising new direction for CALL systems in a global, multilingual contexts.


【8】 Probing the Robustness Properties of Neural Speech Codecs

标题: 探索神经语音编解码器的鲁棒性
链接:https://arxiv.org/abs/2505.24248
作者: Wei-Cheng Tseng,  David Harwath 
备注:Interspeech 2025
摘要:神经语音编解码器彻底改变了语音编码,在保持音频保真度的同时实现了更高的压缩。除了压缩之外,它们还作为标记化策略出现,实现了对语音的语言建模,并在各种语音处理任务中推动范式转变。尽管有这些进步,但它们在嘈杂环境中的鲁棒性仍然没有得到充分的探索,这引起了人们对它们在现实世界中的推广的担忧。在这项工作中,我们系统地评估了各种噪声条件下的神经语音编解码器,揭示了它们的鲁棒性的重要差异。我们进一步研究了它们的线性特性,揭示了非线性失真,这部分解释了观测到的鲁棒性变化。最后,我们分析了它们的频率响应,以确定影响音频保真度的因素。我们的研究结果为编解码器行为和未来的编解码器设计提供了重要的见解,并强调了噪声鲁棒性对其现实世界集成的重要性。
摘要:Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategies, enabling language modeling on speech and driving paradigm shifts across various speech processing tasks. Despite these advancements, their robustness in noisy environments remains underexplored, raising concerns about their generalization to real-world scenarios. In this work, we systematically evaluate neural speech codecs under various noise conditions, revealing non-trivial differences in their robustness. We further examine their linearity properties, uncovering non-linear distortions which partly explain observed variations in robustness. Lastly, we analyze their frequency response to identify factors affecting audio fidelity. Our findings provide critical insights into codec behavior and future codec design, as well as emphasizing the importance of noise robustness for their real-world integration.


【9】 MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly  Speech Recognition

标题: MOPSA:用于老年语音识别的基于预算专家的说话人适应混合
链接:https://arxiv.org/abs/2505.24224
作者: Chengxi Deng,  Xurong Xie,  Shujie Hu,  Mengzhe Geng,  Yicong Jiang,  Jiankun Zhao,  Jiajun Deng,  Guinan Li,  Youjun Chen,  Huimeng Wang,  Haoning Xu,  Mingyu Cui,  Xunying Liu 
备注:Accepted by Interspeech 2025
摘要:本文提出了一种新的基于混合专家的说话人自适应方法(MOPSA)用于老年人语音识别。它允许zero-shot,实时适应看不见的扬声器,并利用领域知识量身定制的老年扬声器。使用K-means得到的前K个最有特色的说话人提示聚类作为专家。一个路由器网络被训练成动态地组合群集的专家组。老年人扬声器之间的声学和语言水平的变化建模使用单独的编码器和解码器提示耳语。在英语DementiaBank Pitt和粤语JCCOCC MoCA老年人语音数据集上的实验表明,在线MOPSA自适应优于说话者无关(SI)模型,其统计学显着的单词错误率(WER)或字符错误率(CER)绝对减少了0.86%和1.47%(相对减少了4.21%和5.40%)。实时因素(RTF)的速度比高达16.12倍,获得了离线批处理模式的适应。
摘要:This paper proposes a novel Mixture of Prompt-Experts based Speaker Adaptation approach (MOPSA) for elderly speech recognition. It allows zero-shot, real-time adaptation to unseen speakers, and leverages domain knowledge tailored to elderly speakers. Top-K most distinctive speaker prompt clusters derived using K-means serve as experts. A router network is trained to dynamically combine clustered prompt-experts. Acoustic and language level variability among elderly speakers are modelled using separate encoder and decoder prompts for Whisper. Experiments on the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets suggest that online MOPSA adaptation outperforms the speaker-independent (SI) model by statistically significant word error rate (WER) or character error rate (CER) reductions of 0.86% and 1.47% absolute (4.21% and 5.40% relative). Real-time factor (RTF) speed-up ratios of up to 16.12 times are obtained over offline batch-mode adaptation.


【10】 Fine-tune Before Structured Pruning: Towards Compact and Accurate  Self-Supervised Models for Speaker Diarization

标题: 结构化修剪前的微调:用于说话人日志化的紧凑而精确的自监督模型
链接:https://arxiv.org/abs/2505.24111
作者: Jiangyu Han,  Federico Landini,  Johan Rohdin,  Anna Silnova,  Mireia Diez,  Jan Cernocky,  Lukas Burget 
备注:Accepted by INTERSPEECH 2025
摘要:在构建说话者日记系统时,可以有效利用WavLM等自监督学习(SSL)模型,但通常体积大且速度慢,限制了它们在资源受限场景中的使用。以前的研究已经探索了压缩技术,但通常在高修剪率的性能下降的代价。在这项工作中,我们提出了压缩SSL模型,通过结构化修剪引入知识蒸馏。与现有的工作不同的是,我们强调了在修剪之前对SSL模型进行微调的重要性。在远场单通道AMI,AISHELL-4和AliMeeting数据集上的实验表明,我们的方法可以去除高达80%的WavLM Base+和WavLM Large的冗余参数,而不会降低性能。修剪后,Base+和Large模型在单个GPU上的推理速度分别提高了4.0倍和2.6倍。我们的源代码是公开的。
摘要:Self-supervised learning (SSL) models like WavLM can be effectively utilized when building speaker diarization systems but are often large and slow, limiting their use in resource constrained scenarios. Previous studies have explored compression techniques, but usually for the price of degraded performance at high pruning ratios. In this work, we propose to compress SSL models through structured pruning by introducing knowledge distillation. Different from the existing works, we emphasize the importance of fine-tuning SSL models before pruning. Experiments on far-field single-channel AMI, AISHELL-4, and AliMeeting datasets show that our method can remove redundant parameters of WavLM Base+ and WavLM Large by up to 80% without any performance degradation. After pruning, the inference speeds on a single GPU for the Base+ and Large models are 4.0 and 2.6 times faster, respectively. Our source code is publicly available.


【11】 Can Emotion Fool Anti-spoofing?

标题: 情感可以愚弄反欺骗吗?
链接:https://arxiv.org/abs/2505.23962
作者: Aurosweta Mahapatra,  Ismail Rasim Ulgen,  Abinay Reddy Naini,  Carlos Busso,  Berrak Sisman 
备注:Accepted to Interspeech 2025
摘要:传统的反欺骗侧重于基于合成语音的模型和数据集,这些语音大多处于中性状态,忽略了各种情绪变化。因此,它们对高质量、情感表达的合成语音的鲁棒性是不确定的。我们解决这个问题,通过引入语音欺骗TTS,一个语料库的情感文本到语音样本。我们的分析表明,现有的反欺骗模型与情感合成语音作斗争,暴露了针对情感的攻击的风险。即使在情绪数据上进行训练,由于对情绪方面的关注有限,模型表现不佳,并且显示出情绪之间的性能差异。这突出了在数据集和方法中需要以情感为中心的反欺骗范例。我们提出了GEM,一个门控集成的情感专用模型与语音情感识别门控网络。GEM在所有情绪和中性状态下都能有效执行,提高了对欺骗攻击的防御能力。我们发布了Spoof-TTS数据集:https://emospoof-tts.github.io/Dataset/
摘要:Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness against high-quality, emotionally expressive synthetic speech is uncertain. We address this by introducing EmoSpoof-TTS, a corpus of emotional text-to-speech samples. Our analysis shows existing anti-spoofing models struggle with emotional synthetic speech, exposing risks of emotion-targeted attacks. Even trained on emotional data, the models underperform due to limited focus on emotional aspect and show performance disparities across emotions. This highlights the need for emotion-focused anti-spoofing paradigm in both dataset and methodology. We propose GEM, a gated ensemble of emotion-specialized models with a speech emotion recognition gating network. GEM performs effectively across all emotions and neutral state, improving defenses against spoofing attacks. We release the EmoSpoof-TTS Dataset: https://emospoof-tts.github.io/Dataset/


【12】 Masked Self-distilled Transducer-based Keyword Spotting with  Semi-autoregressive Decoding

标题: 基于半自回归解码的掩蔽自蒸馏传感器的关键词发现
链接:https://arxiv.org/abs/2505.24820
作者: Yu Xi,  Xiaoyu Gu,  Haoyu Li,  Jun Song,  Bo Zheng,  Kai Yu 
摘要:基于RNN-T的自回归解码关键词发现(KWS)由于其流结构和优越的性能而受到关注。然而,RNN-T中预测网络的简单性带来了过拟合问题,特别是在具有挑战性的场景下,导致性能下降。在本文中,我们提出了一种掩蔽自蒸馏(MSD)训练策略,避免RNN-Ts过度依赖预测网络来减轻过拟合。这种训练实现了掩码非自回归(NAR)解码,其在KWS解码期间完全掩蔽RNN-T预测器输出。此外,我们提出了一种半自回归(SAR)解码方法,结合AR和NAR解码的优点。我们在多个KWS数据集上的实验表明,MSD训练有效地消除了过拟合。该方法在保留AR解码优越性能的同时,利用NAR解码的过拟合抑制,取得了良好的效果。
摘要:RNN-T-based keyword spotting (KWS) with autoregressive decoding~(AR) has gained attention due to its streaming architecture and superior performance. However, the simplicity of the prediction network in RNN-T poses an overfitting issue, especially under challenging scenarios, resulting in degraded performance. In this paper, we propose a masked self-distillation (MSD) training strategy that avoids RNN-Ts overly relying on prediction networks to alleviate overfitting. Such training enables masked non-autoregressive (NAR) decoding, which fully masks the RNN-T predictor output during KWS decoding. In addition, we propose a semi-autoregressive (SAR) decoding approach to integrate the advantages of AR and NAR decoding. Our experiments across multiple KWS datasets demonstrate that MSD training effectively alleviates overfitting. The SAR decoding method preserves the superior performance of AR decoding while benefits from the overfitting suppression of NAR decoding, achieving excellent results.


【13】 Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic  Dialect Identification

标题: 语音转换提高阿拉伯语口语识别的跨域稳健性
链接:https://arxiv.org/abs/2505.24713
作者: Badr M. Abdullah,  Matthew Baas,  Bernd Möbius,  Dietrich Klakow 
备注:Accepted in Interspeech 2025
摘要:阿拉伯语方言识别(ADI)系统对于大规模数据收集管道至关重要,这些管道能够为阿拉伯语变体开发包容性语音技术。然而,当前ADI系统的可靠性受限于对域外语音的较差推广。在本文中,我们提出了一种基于语音转换的有效方法来训练ADI模型,该模型具有最先进的性能,并显着提高了跨域场景的鲁棒性。在新收集的跨越四个不同领域的真实世界测试集上进行评估,我们的方法在各个领域的准确性上得到了高达+34.1%的一致改进。此外,我们提出了我们的方法的分析,并证明语音转换有助于减轻ADI数据集中的扬声器偏见。我们发布了强大的ADI模型和跨域评估数据集,以支持阿拉伯语包容性语音技术的开发。
摘要:Arabic dialect identification (ADI) systems are essential for large-scale data collection pipelines that enable the development of inclusive speech technologies for Arabic language varieties. However, the reliability of current ADI systems is limited by poor generalization to out-of-domain speech. In this paper, we present an effective approach based on voice conversion for training ADI models that achieves state-of-the-art performance and significantly improves robustness in cross-domain scenarios. Evaluated on a newly collected real-world test set spanning four different domains, our approach yields consistent improvements of up to +34.1% in accuracy across domains. Furthermore, we present an analysis of our approach and demonstrate that voice conversion helps mitigate the speaker bias in the ADI dataset. We release our robust ADI model and cross-domain evaluation dataset to support the development of inclusive speech technologies for Arabic.


【14】 Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing  Cross-Lingual Transfer in Low-Resource Scenarios

标题: 使用音素增强CoT的语音转文本翻译:在低资源场景中增强跨语言转换
链接:https://arxiv.org/abs/2505.24691
作者: Gerard I. Gállego,  Oriol Pareras,  Martí Cortada Garcia,  Lucas Takanori,  Javier Hernando 
备注:Accepted at Interspeech 2025
摘要:我们提出了一种语音到文本翻译(S2TT)的方法,将音素表示集成到一个思想链(CoT)框架,以提高在低资源和零资源设置的翻译。通过引入音素识别作为中间步骤,我们增强了跨语言迁移,即使是没有标记语音数据的语言也可以进行翻译。我们的系统建立在一个多语言的LLM,我们扩展到处理语音和音素。培训遵循课程学习战略,逐步引入更复杂的任务。在多语言S2TT基准测试上的实验表明,音素增强CoT在低资源条件下提高了翻译质量,并实现了零资源翻译,同时对高资源性能略有影响。尽管存在这种权衡,但我们的研究结果表明,基于音素的CoT是使S2TT在不同语言中更容易使用的有希望的一步。
摘要:We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing phoneme recognition as an intermediate step, we enhance cross-lingual transfer, enabling translation even for languages with no labeled speech data. Our system builds on a multilingual LLM, which we extend to process speech and phonemes. Training follows a curriculum learning strategy that progressively introduces more complex tasks. Experiments on multilingual S2TT benchmarks show that phoneme-augmented CoT improves translation quality in low-resource conditions and enables zero-resource translation, while slightly impacting high-resource performance. Despite this trade-off, our findings demonstrate that phoneme-based CoT is a promising step toward making S2TT more accessible across diverse languages.


【15】 MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised  Domain Adaptation in ASR

标题: MSDA:结合伪标记和自监督实现ASR中的无监督域自适应
链接:https://arxiv.org/abs/2505.24656
作者: Dimitrios Damianos,  Georgios Paraskevopoulos,  Alexandros Potamianos 
摘要:在这项工作中,我们研究了自动语音识别(ASR)的Meta PL无监督域自适应框架。我们介绍了一个多阶段域自适应管道(MSDA),一个样本效率,两阶段的适应方法,集成了自监督学习与半监督技术。MSDA旨在增强ASR模型的鲁棒性和通用性,使其更适应不同的条件。它对于像希腊语这样的低资源语言以及标记数据稀缺或嘈杂的弱监督场景特别有效。通过大量的实验,我们证明了Meta PL可以有效地应用于ASR任务,实现最先进的结果,显着优于最先进的方法,并提供更强大的解决方案,在ASR的无监督域自适应。我们的消融强调了在自我监督与自我训练相结合时,利用级联方法的必要性。
摘要:In this work, we investigate the Meta PL unsupervised domain adaptation framework for Automatic Speech Recognition (ASR). We introduce a Multi-Stage Domain Adaptation pipeline (MSDA), a sample-efficient, two-stage adaptation approach that integrates self-supervised learning with semi-supervised techniques. MSDA is designed to enhance the robustness and generalization of ASR models, making them more adaptable to diverse conditions. It is particularly effective for low-resource languages like Greek and in weakly supervised scenarios where labeled data is scarce or noisy. Through extensive experiments, we demonstrate that Meta PL can be applied effectively to ASR tasks, achieving state-of-the-art results, significantly outperforming state-of-the-art methods, and providing more robust solutions for unsupervised domain adaptation in ASR. Our ablations highlight the necessity of utilizing a cascading approach when combining self-supervision with self-training.


【16】 ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis  Optimization for Speech Multi-Metric Estimation

标题: ARECHO:通过基于链的假设优化的语音多指标估计的自回归评估
链接:https://arxiv.org/abs/2505.24518
作者: Jiatong Shi,  Yifan Cheng,  Bo-Hao Su,  Hye-jin Shim,  Jinchuan Tian,  Samuele Cornell,  Yiwen Zhao,  Siddhant Arora,  Shinji Watanabe 
摘要:语音信号分析提出了重大挑战,特别是在语音质量评估和分析等任务中,其目标是预测多个感知和客观指标。例如,PESQ(语音质量的感知评估),STOI(短时客观可懂度)和MOS(平均意见得分)等指标都可以捕获语音质量的不同方面。然而,这些度量通常具有不同的尺度、假设和依赖性,使得联合估计不平凡。为了解决这些问题,我们介绍了ARECHO(自回归评估通过基于链的假设优化),一个基于链的,多功能的语音评估系统,基于自回归依赖模型。ARECHO有三个关键创新:(1)全面的语音信息标记化管道;(2)明确捕获度量间依赖关系的动态分类器链;(3)增强推理可靠性的两步置信度解码算法。实验表明,ARECHO显着优于基线框架在不同的评估方案,包括增强语音分析,语音生成评估,和嘈杂的语音评估。此外,它的动态依赖建模通过捕获度量间关系提高了可解释性。
摘要:Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships.


【17】 MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging  LLM Embedded Knowledge

标题: MELT:利用LLM嵌入式知识实现自动化多模式情感数据注释
链接:https://arxiv.org/abs/2505.24493
作者: Xin Jing,  Jiadong Wang,  Iosif Tsangko,  Andreas Triantafyllopoulos,  Björn W. Schuller 
摘要:虽然语音情感识别(SER)在深度学习方面取得了显著进展,但注释仍然是一个主要障碍。人工标注不仅成本高,而且容易出现不一致性,标注者往往有不同的偏好,可能缺乏必要的上下文知识,这可能导致标签的变化和不准确。与此同时,大型语言模型(LLM)已经成为注释文本数据的可扩展替代方案。然而,LLM在没有人类监督的情况下执行情感语音数据注释的潜力还有待深入研究。为了解决这些问题,我们应用GPT-4 o来注释从情景喜剧《老友记》中收集的多模态数据集,只使用文本线索作为输入。通过制作结构化的文本提示,我们的方法利用了GPT-4 o在训练过程中积累的知识,展示了它可以在不直接访问多模态输入的情况下生成准确且与上下文相关的注释。因此,我们提出了MELT,一个由GPT-4 o完全注释的多模态情感数据集。我们通过微调四个自监督学习(SSL)主干并评估情感数据集上的语音情感识别性能来证明MELT的有效性。此外,我们的主观实验结果表明,SER的性能得到了一致的改善。
摘要:Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to inconsistencies annotators often have different preferences and may lack the necessary contextual knowledge, which can lead to varied and inaccurate labels. Meanwhile, Large Language Models (LLMs) have emerged as a scalable alternative for annotating text data. However, the potential of LLMs to perform emotional speech data annotation without human supervision has yet to be thoroughly investigated. To address these problems, we apply GPT-4o to annotate a multimodal dataset collected from the sitcom Friends, using only textual cues as inputs. By crafting structured text prompts, our methodology capitalizes on the knowledge GPT-4o has accumulated during its training, showcasing that it can generate accurate and contextually relevant annotations without direct access to multimodal inputs. Therefore, we propose MELT, a multimodal emotion dataset fully annotated by GPT-4o. We demonstrate the effectiveness of MELT by fine-tuning four self-supervised learning (SSL) backbones and assessing speech emotion recognition performance across emotion datasets. Additionally, our subjective experiments\' results demonstrate a consistence performance improvement on SER.


【18】 Rehearsal with Auxiliary-Informed Sampling for Audio Deepfake Detection

标题: 音频深度伪造检测的辅助抽样排练
链接:https://arxiv.org/abs/2505.24486
作者: Falih Gozi Febrinanto,  Kristen Moore,  Chandra Thapa,  Jiangang Ma,  Vidya Saikrishna,  Feng Xia 
备注:Accepted by Interspeech 2025
摘要:现有的音频deepfake检测框架在面对新的deepfake攻击时性能会下降。基于排练的持续学习(CL)使用有限的一组旧数据样本更新模型,有助于保留先验知识,同时融入新信息。然而,现有的排练技术不能有效地捕捉音频特征的多样性,引入偏见,增加了遗忘的风险。为了应对这一挑战,我们提出了一种基于排练的CL音频深度伪造检测方法-。RAIS采用标签生成网络来生成辅助标签,指导存储缓冲区的不同样本选择。大量的实验表明,RAIS优于最先进的方法,在五个经验中实现了1.953%的平均等错误率(EER)。该代码可从以下网址获得:https://github.com/falihgoz/RAIS。
摘要:The performance of existing audio deepfake detection frameworks degrades when confronted with new deepfake attacks. Rehearsal-based continual learning (CL), which updates models using a limited set of old data samples, helps preserve prior knowledge while incorporating new information. However, existing rehearsal techniques don't effectively capture the diversity of audio characteristics, introducing bias and increasing the risk of forgetting. To address this challenge, we propose Rehearsal with Auxiliary-Informed Sampling (RAIS), a rehearsal-based CL approach for audio deepfake detection. RAIS employs a label generation network to produce auxiliary labels, guiding diverse sample selection for the memory buffer. Extensive experiments show RAIS outperforms state-of-the-art methods, achieving an average Equal Error Rate (EER) of 1.953 % across five experiences. The code is available at: https://github.com/falihgoz/RAIS.


【19】 SuPseudo: A Pseudo-supervised Learning Method for Neural Speech  Enhancement in Far-field Speech Recognition

标题: SuNorton:远场语音识别中神经语音增强的伪监督学习方法
链接:https://arxiv.org/abs/2505.24450
作者: Longjie Luo,  Lin Li,  Qingyang Hong 
备注:Accepted by InterSpeech 2025
摘要:None
摘要:Due to the lack of target speech annotations in real-recorded far-field conversational datasets, speech enhancement (SE) models are typically trained on simulated data. However, the trained models often perform poorly in real-world conditions, hindering their application in far-field speech recognition. To address the issue, we (a) propose direct sound estimation (DSE) to estimate the oracle direct sound of real-recorded data for SE; and (b) present a novel pseudo-supervised learning method, SuPseudo, which leverages DSE-estimates as pseudo-labels and enables SE models to directly learn from and adapt to real-recorded data, thereby improving their generalization capability. Furthermore, an SE model called FARNET is designed to fully utilize SuPseudo. Experiments on the MISP2023 corpus demonstrate the effectiveness of SuPseudo, and our system significantly outperforms the previous state-of-the-art. A demo of our method can be found at https://EeLLJ.github.io/SuPseudo/.


【20】 Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the  MISP-Meeting Challenge

标题: MISP会议挑战中AVSR任务的基于伪标签的神经语音增强
链接:https://arxiv.org/abs/2505.24446
作者: Longjie Luo,  Shenghui Lu,  Lin Li,  Qingyang Hong 
备注:Accepted by InterSpeech 2025
摘要:本文介绍了我们的系统MISP会议挑战轨道2。主要的困难在于数据集,其中包含强背景噪声,混响,重叠的语音和不同的会议主题。为了解决这些问题,我们(a)设计了G-SpatialNet,一个语音增强(SE)模型来改善引导源分离(GSS)信号;(b)提出了TLS,一个包括时间对齐、电平对齐和信噪比滤波的框架,用于为真实记录的远场音频数据生成信号级伪标签,从而促进SE模型的训练;以及(c)探索微调策略、数据增强和多模态信息,以提高预先训练的自动语音识别(ASR)模型在会议场景中的性能。最后,我们的系统在Dev和Eval集上分别实现了5.44%和9.52%的字符错误率(CER),比基线分别提高了64.8%和52.6%,获得第二名。
摘要:This paper presents our system for the MISP-Meeting Challenge Track 2. The primary difficulty lies in the dataset, which contains strong background noise, reverberation, overlapping speech, and diverse meeting topics. To address these issues, we (a) designed G-SpatialNet, a speech enhancement (SE) model to improve Guided Source Separation (GSS) signals; (b) proposed TLS, a framework comprising time alignment, level alignment, and signal-to-noise ratio filtering, to generate signal-level pseudo labels for real-recorded far-field audio data, thereby facilitating SE models' training; and (c) explored fine-tuning strategies, data augmentation, and multimodal information to enhance the performance of pre-trained Automatic Speech Recognition (ASR) models in meeting scenarios. Finally, our system achieved character error rates (CERs) of 5.44% and 9.52% on the Dev and Eval sets, respectively, with relative improvements of 64.8% and 52.6% over the baseline, securing second place.


【21】 SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization

标题: Switch Codec:具有稀疏量化的高保真神经音频编解码器
链接:https://arxiv.org/abs/2505.24437
作者: Jin Wang,  Wenbin Jiang,  Xiangbo Wang 
备注:5 pages,4 figures
摘要:我们提出了一个通用的高保真神经音频压缩算法,可以压缩语音,音乐,和一般的音频低于3 kbps的带宽。虽然当前最先进的音频编解码器在音频压缩方面表现出色,但当嵌入空间急剧减少时,其有效性显著下降,这对应于更高的压缩。为了解决这个问题,我们提出了残差专家矢量量化(REVQ),它显着扩展了可用的嵌入空间,提高了性能,同时几乎不牺牲带宽。此外,我们引入了一个策略,以确保巨大的嵌入空间可以得到充分利用。此外,我们提出了一个基于STFT的控制器来指导生成器产生难以区分的频谱图。我们证明,所提出的方法优于基线方法,通过详细的消融。
摘要:We present a universal high-fidelity neural audio compression algorithm that can compress speech, music, and general audio below 3 kbps bandwidth. Although current state-of-the-art audio codecs excel in audio compression, their effectiveness significantly declines when embedding space is sharply reduced, which corresponds to higher compression. To address this problem, we propose Residual Experts Vector Quantization (REVQ), which significantly expands the available embedding space and improves the performance while hardly sacrificing the bandwidth. Furthermore, we introduce a strategy to ensure that the vast embedding space can be fully utilized. Additionally, we propose a STFT-based discriminator to guide the generator in producing indistinguishable spectrograms. We demonstrate that the proposed approach outperforms baseline methods through detailed ablations.


【22】 Fewer Hallucinations, More Verification: A Three-Stage LLM-Based  Framework for ASR Error Correction

标题: 更少的幻觉,更多的验证:一个基于LLM的三阶段ASB错误纠正框架
链接:https://arxiv.org/abs/2505.24347
作者: Yangui Fang,  Baixu Cheng,  Jing Peng,  Xu Li,  Yu Xi,  Chengwei Zhang,  Guohui Zhong 
摘要:自动语音识别(ASR)纠错的目的是纠正识别错误,同时保留准确的文本。尽管传统方法表现出中等的有效性,但LLM提供了一种无需训练和标记数据的范例。然而,直接使用LLM会遇到幻觉问题,这可能会导致正确文本的修改。为了解决这个问题,我们提出了可靠的LLM校正框架(RLLM-CF),它包括三个阶段:(1)错误预检测,(2)思想链子任务迭代校正,(3)推理过程验证。我们的方法的优点是,它不需要额外的信息或微调的模型,并确保在多遍编程下的LLM校正的正确性。在AISHELL-1、AISHELL-2和Librispeech上的实验表明,通过我们的框架增强的GPT-4 o模型在CER/WER上实现了21%、11%、9%和11.4%的相对减少。
摘要:Automatic Speech Recognition (ASR) error correction aims to correct recognition errors while preserving accurate text. Although traditional approaches demonstrate moderate effectiveness, LLMs offer a paradigm that eliminates the need for training and labeled data. However, directly using LLMs will encounter hallucinations problem, which may lead to the modification of the correct text. To address this problem, we propose the Reliable LLM Correction Framework (RLLM-CF), which consists of three stages: (1) error pre-detection, (2) chain-of-thought sub-tasks iterative correction, and (3) reasoning process verification. The advantage of our method is that it does not require additional information or fine-tuning of the model, and ensures the correctness of the LLM correction under multi-pass programming. Experiments on AISHELL-1, AISHELL-2, and Librispeech show that the GPT-4o model enhanced by our framework achieves 21%, 11%, 9%, and 11.4% relative reductions in CER/WER.


【23】 DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture  Switching for Speech Codec

标题: DS-编解码器:语音编解码器的双阶段训练,采用双镜像架构切换
链接:https://arxiv.org/abs/2505.24314
作者: Peijie Chen,  Wenhao Guan,  Kaidi Wang,  Weijie Wu,  Hukai Huang,  Qingyang Hong,  Lin Li 
备注:Accepted to Interspeech 2025
摘要:神经语音编解码器对于推进文本到语音(TTS)系统至关重要。随着近年来大型语言模型在文本生成方面的成功,开发高质量的语音标记器变得越来越重要。本文介绍了DS-Codec,一种新型的神经语音编解码器,具有镜像和非镜像架构切换的双阶段训练框架,旨在实现卓越的语音重建。我们进行了广泛的实验和消融研究,以评估我们的训练策略的有效性,并比较两种架构的性能。我们的研究结果表明,镜像结构显着提高了学习的码本的鲁棒性,和训练策略之间的优势平衡镜像和非镜像结构,导致改善高保真语音重建。
摘要:Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech tokenizers has become increasingly important. This paper introduces DS-Codec, a novel neural speech codec featuring a dual-stage training framework with mirror and non-mirror architectures switching, designed to achieve superior speech reconstruction. We conduct extensive experiments and ablation studies to evaluate the effectiveness of our training strategy and compare the performance of the two architectures. Our results show that the mirrored structure significantly enhances the robustness of the learned codebooks, and the training strategy balances the advantages between mirrored and non-mirrored structures, leading to improved high-fidelity speech reconstruction.


【24】 Discl-VC: Disentangled Discrete Tokens and In-Context Learning for  Controllable Zero-Shot Voice Conversion

标题: Discl-VC:分离离散令牌和上下文内学习以实现可控Zero-Shot语音转换
链接:https://arxiv.org/abs/2505.24291
作者: Kaidi Wang,  Wenhao Guan,  Ziyue Jiang,  Hukai Huang,  Peijie Chen,  Weijie Wu,  Qingyang Hong,  Lin Li 
摘要:目前,zero-shot语音转换系统能够合成看不见的说话者的语音。然而,大多数现有方法难以准确地复制源说话者的说话风格或模仿目标说话者的独特说话风格,从而限制了语音转换的可控性。在这项工作中,我们提出了Discl-VC,一种新的语音转换框架,从自我监督的语音表示中解开内容和韵律信息,并通过上下文学习与流匹配Transformer合成目标说话人的声音。为了能够精确控制生成的语音的韵律,我们引入了一个掩码生成Transformer,它基于提示以非自回归的方式预测离散的韵律标记。实验结果表明,该方法在zero-shot语音转换中具有良好的性能,在合成语音的韵律控制中具有较高的精度。
摘要:Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive speaking style of the target speaker, thereby limiting the controllability of voice conversion. In this work, we propose Discl-VC, a novel voice conversion framework that disentangles content and prosody information from self-supervised speech representations and synthesizes the target speaker's voice through in-context learning with a flow matching transformer. To enable precise control over the prosody of generated speech, we introduce a mask generative transformer that predicts discrete prosody tokens in a non-autoregressive manner based on prompts. Experimental results demonstrate the superior performance of Discl-VC in zero-shot voice conversion and its remarkable accuracy in prosody control for synthesized speech.


【25】 Dynamic Context-Aware Streaming Pretrained Language Model For Inverse  Text Normalization

标题: 用于反向文本规范化的动态上下文感知流预训练语言模型
链接:https://arxiv.org/abs/2505.24229
作者: Luong Ho,  Khanh Le,  Vinh Pham,  Bao Nguyen,  Tan Tran,  Duc Chau 
备注:Accepted to INTERSPEECH 2025
摘要:反向文本规范化(ITN)对于将口语自动语音识别(ASR)输出转换为格式良好的书面文本,增强可读性和可用性至关重要。尽管流媒体ITN很重要,但由于准确性、效率和适应性方面的挑战,特别是在资源匮乏和上下文有限的场景中,流媒体ITN在流媒体ASR中的集成在很大程度上尚未得到探索。在本文中,我们为ITN引入了一个流预训练语言模型,利用预训练语言表示来提高鲁棒性。为了解决流约束,我们在训练和推理过程中提出了动态上下文感知,从而实现自适应块大小调整和右上下文信息的集成。实验结果表明,我们的方法实现了与非流式ITN相当的准确性,并超过了越南数据集上现有的流式ITN模型,同时保持低延迟,确保无缝集成到ASR系统中。
摘要:Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integration of streaming ITN within streaming ASR remains largely unexplored due to challenges in accuracy, efficiency, and adaptability, particularly in low-resource and limited-context scenarios. In this paper, we introduce a streaming pretrained language model for ITN, leveraging pretrained linguistic representations for improved robustness. To address streaming constraints, we propose Dynamic Context-Aware during training and inference, enabling adaptive chunk size adjustments and the integration of right-context information. Experimental results demonstrate that our method achieves accuracy comparable to non-streaming ITN and surpasses existing streaming ITN models on a Vietnamese dataset, all while maintaining low latency, ensuring seamless integration into ASR systems.


【26】 Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with  Data Augmentation and LID-Aware CTC

标题: 改进ML-SURB 2.0上的多语言语音模型:使用数据增强和LID-Aware CTC进行微调
链接:https://arxiv.org/abs/2505.24200
作者: Qingzheng Wang,  Jiancheng Sun,  Yifan Peng,  Shinji Watanabe 
摘要:使用自监督或监督预训练语音基础模型(SFM)的多语言语音处理在语言识别(LID)和自动语音识别(ASR)等任务上取得了很好的性能。然而,这些模型在微调期间与有限的资源作斗争。本文通过探索多种适应SFM的策略,包括冻结上游训练,部分微调和低秩适应,增强了ML-SUPERB 2.0上的多语言LID和ASR。此外,我们采用数据增强来缓解Few-Shot设置中的性能差距,并引入LID连接时间分类(CTC)损失进行正则化。我们的方法在ML-SUPERB 2.0上实现了LID准确性14%的相对提高和ASR CER 30%的相对降低,在Interspeech 2025 ML-SUPERB 2.0挑战赛中获得第二名。
摘要:Multilingual speech processing with self-supervised or supervised pre-trained Speech Foundation Models (SFM) has achieved strong performance on tasks like Language Identification (LID) and Automatic Speech Recognition (ASR). However, these models struggle with limited resources during fine-tuning. This paper enhances multilingual LID and ASR on ML-SUPERB 2.0 by exploring multiple strategies for adapting SFMs, including frozen upstream training, partial fine-tuning, and low-rank adaptation. Furthermore, we employ data augmentation to mitigate performance gaps in few-shot settings and introduce LID Connectionist Temporal Classification (CTC) loss for regularization. Our approach achieves a 14% relative improvement in LID accuracy and a 30% relative reduction in ASR CER over the baseline on ML-SUPERB 2.0, securing second place in the Interspeech 2025 ML-SUPERB 2.0 Challenge.


【27】 FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing  System

标题: QUALureSense:在始终在线的音频传感系统中保护说话者属性
链接:https://arxiv.org/abs/2505.24115
作者: Bhawana Chhaglani,  Sarmistha Sarna Gomasta,  Yuvraj Agarwal,  Jeremy Gummeson,  Prashant Shenoy 
摘要:音频是一种丰富的感测模态,可用于各种人类活动识别任务。然而,智能手机和带有始终开启麦克风的智能扬声器的无处不在的性质导致了许多隐私问题,并且对部署这些基于音频的传感系统缺乏信任。本文解决了这个关键的挑战,保护用户隐私时,使用音频传感应用程序,同时保持实用程序。虽然以前的工作主要集中在保护可恢复的语音内容,我们表明,敏感的扬声器特定的属性,如年龄和性别仍然可以推断掩蔽语音后,并提出了一个全面的隐私评估框架,以评估这个扬声器属性泄漏。我们设计并实现了一个开源库,提供了一组可推广的隐私感知音频功能,可用于广泛的传感应用。我们提出了一个自适应的特定于任务的功能选择算法,优化的隐私效用成本权衡的基础上的应用程序的要求。通过我们广泛的评估,我们证明了在各种传感任务中的高实用性。我们的系统在保护用户特定隐私方面比现有的隐私技术高出60.6%。这项工作提供了一个基础框架,通过启用有效的隐私感知音频分类系统来确保对音频感测的信任。
摘要:Audio is a rich sensing modality that is useful for a variety of human activity recognition tasks. However, the ubiquitous nature of smartphones and smart speakers with always-on microphones has led to numerous privacy concerns and a lack of trust in deploying these audio-based sensing systems. This paper addresses this critical challenge of preserving user privacy when using audio for sensing applications while maintaining utility. While prior work focuses primarily on protecting recoverable speech content, we show that sensitive speaker-specific attributes such as age and gender can still be inferred after masking speech and propose a comprehensive privacy evaluation framework to assess this speaker attribute leakage. We design and implement FeatureSense, an open-source library that provides a set of generalizable privacy-aware audio features that can be used for wide range of sensing applications. We present an adaptive task-specific feature selection algorithm that optimizes the privacy-utility-cost trade-off based on the application requirements. Through our extensive evaluation, we demonstrate the high utility of FeatureSense across a diverse set of sensing tasks. Our system outperforms existing privacy techniques by 60.6% in preserving user-specific privacy. This work provides a foundational framework for ensuring trust in audio sensing by enabling effective privacy-aware audio classification systems.


【28】 Acoustic Classification of Maritime Vessels using Learnable Filterbanks

标题: 使用可学习滤片组对船舶进行声学分类
链接:https://arxiv.org/abs/2505.23964
作者: Jonas Elsborg,  Tejs Vegge,  Arghya Bhowmik 
备注:9 pages, 5 figures, 2 tables
摘要:基于声学特征可靠地监测和识别海上船只由于不同记录场景的可变性而变得复杂。一个强大的分类框架必须能够概括不同的声学环境和可变的源传感器距离。为此,我们提出了一个深度学习模型,该模型在不同的记录场景中具有强大的性能。使用可训练的频谱前端和时间特征编码器来学习Gabor滤波器组,该模型可以动态地强调不同的频率分量。在格鲁吉亚海峡的VTUAD水听器记录上进行训练后,我们的模型CATFISH在不同的源传感器距离上实现了最先进的96.63%的测试准确率,超过了之前的基准超过12个百分点。我们提出的模型,证明我们的架构选择,分析学习的Gabor滤波器,并进行消融研究传感器数据融合和基于注意力的池。
摘要:Reliably monitoring and recognizing maritime vessels based on acoustic signatures is complicated by the variability of different recording scenarios. A robust classification framework must be able to generalize across diverse acoustic environments and variable source-sensor distances. To this end, we present a deep learning model with robust performance across different recording scenarios. Using a trainable spectral front-end and temporal feature encoder to learn a Gabor filterbank, the model can dynamically emphasize different frequency components. Trained on the VTUAD hydrophone recordings from the Strait of Georgia, our model, CATFISH, achieves a state-of-the-art 96.63 % percent test accuracy across varying source-sensor distances, surpassing the previous benchmark by over 12 percentage points. We present the model, justify our architectural choices, analyze the learned Gabor filters, and perform ablation studies on sensor data fusion and attention-based pooling.


【29】 Patient-Aware Feature Alignment for Robust Lung Sound  Classification:Cohesion-Separation and Global Alignment Losses

标题: 用于鲁棒肺音分类的患者感知特征对齐:内聚分离和全局对齐损失
链接:https://arxiv.org/abs/2505.23834
作者: Seung Gyu Jeong,  Seong Eun Kim 
备注:Accepted INTERSPEECH 2025
摘要:肺音分级对呼吸系统疾病的早期诊断至关重要。然而,生物医学信号往往表现出患者间的差异,甚至在具有相同症状的患者之间,需要考虑个体差异的学习方法。我们提出了一个患者感知特征对齐(PAFA)框架,其中包含两个新的损失,患者凝聚分离损失(PCSL)和全局患者对齐损失(GPAL)。PCSL将同一患者的特征聚类,同时将这些特征与其他患者分离,以捕获患者的变异性,而GPAL将每个患者的质心绘制到全局中心,防止特征空间碎片化。在ICBHI数据集上,该方法取得了较好的效果,四类分类的分类率为64.84%,两类分类的分类率为72.08%.这些发现突出了PAFA捕捉个性化模式的能力,并在不同的患者群中展示了性能增益,为以患者为中心的医疗保健提供了更广泛的应用。
摘要:Lung sound classification is vital for early diagnosis of respiratory diseases. However, biomedical signals often exhibit inter-patient variability even among patients with the same symptoms, requiring a learning approach that considers individual differences. We propose a Patient-Aware Feature Alignment (PAFA) framework with two novel losses, Patient Cohesion-Separation Loss (PCSL) and Global Patient Alignment Loss (GPAL). PCSL clusters features of the same patient while separating those from other patients to capture patient variability, whereas GPAL draws each patient's centroid toward a global center, preventing feature space fragmentation. Our method achieves outstanding results on the ICBHI dataset with a score of 64.84\% for four-class and 72.08\% for two-class classification. These findings highlight PAFA's ability to capture individualized patterns and demonstrate performance gains in distinct patient clusters, offering broader applications for patient-centered healthcare.


【30】 SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks  via Watermarking

标题: SpeechVerification:强大的声学指纹,防止通过水印进行的篡改攻击
链接:https://arxiv.org/abs/2505.23821
作者: Lingfeng Yao (1),  Chenpei Huang (1),  Shengyao Wang (2),  Junpei Xue (2),  Hanqing Guo (3),  Jiang Liu (2),  Xun Chen (4),  Miao Pan (1) ((1) University of Houston (2) Waseda University (3) University of Hawaii at Mānoa (4) Independent Researcher) 
摘要:随着社交媒体的蓬勃发展,恶意篡改的公共言论,特别是来自有影响力的人物的言论,严重影响了社会稳定和公众信任。现有的语音篡改检测方法仍然不足:它们要么依赖于外部参考数据,要么对攻击不敏感,对压缩和重采样等良性操作不鲁棒。为了应对这些挑战,我们引入了SpeechVerifer,仅使用发布的语音本身来主动验证语音完整性,即,而不需要任何外部参考。受音频指纹和水印的启发,SpeechVerifier可以(i)有效地检测篡改攻击,(ii)对良性操作具有鲁棒性,(iii)仅基于发布的语音验证完整性。简言之,SpeechVerifier利用多尺度特征提取来捕获跨不同时间分辨率的语音特征。然后,它采用对比学习来生成指纹,可以检测不同粒度的修改。这些指纹被设计为对良性操作具有鲁棒性,但在发生恶意篡改时会显示出显著的变化。为了能够以自包含的方式进行语音验证,然后通过分段水印将生成的指纹嵌入到语音信号中。在没有外部参考的情况下,SpeechVerifier可以从发布的音频中检索指纹,并与嵌入的水印进行检查,以验证语音的完整性。大量的实验结果表明,提出的语音验证器是有效的检测篡改攻击和稳健的良性操作。
摘要:With the surge of social media, maliciously tampered public speeches, especially those from influential figures, have seriously affected social stability and public trust. Existing speech tampering detection methods remain insufficient: they either rely on external reference data or fail to be both sensitive to attacks and robust to benign operations, such as compression and resampling. To tackle these challenges, we introduce SpeechVerifer to proactively verify speech integrity using only the published speech itself, i.e., without requiring any external references. Inspired by audio fingerprinting and watermarking, SpeechVerifier can (i) effectively detect tampering attacks, (ii) be robust to benign operations and (iii) verify the integrity only based on published speeches. Briefly, SpeechVerifier utilizes multiscale feature extraction to capture speech features across different temporal resolutions. Then, it employs contrastive learning to generate fingerprints that can detect modifications at varying granularities. These fingerprints are designed to be robust to benign operations, but exhibit significant changes when malicious tampering occurs. To enable speech verification in a self-contained manner, the generated fingerprints are then embedded into the speech signal by segment-wise watermarking. Without external references, SpeechVerifier can retrieve the fingerprint from the published audio and check it with the embedded watermark to verify the integrity of the speech. Extensive experimental results demonstrate that the proposed SpeechVerifier is effective in detecting tampering attacks and robust to benign operations.


【31】 Learning Normal Patterns in Musical Loops

标题: 学习音乐循环中的正常模式
链接:https://arxiv.org/abs/2505.23784
作者: Shayan Dadman,  Bernt Arild Bremdal,  Børre Bang,  Rune Dalmo 
备注:27 pages, 10 figures
摘要:本文介绍了一个无监督的框架,通过异常检测技术检测音乐样本(循环)中的音频模式,解决音乐信息检索(MIR)的挑战。现有的方法往往受到手工制作的功能,特定领域的限制,或依赖于迭代的用户交互的依赖。我们通过将深度特征提取与无监督异常检测相结合的架构来解决这些限制。我们的方法利用预先训练的分层令牌语义音频Transformer(HTS-AT),与特征融合机制(FFM)配对,从可变长度的音频循环生成表示。这些嵌入使用单类深度支持向量数据描述(Deep SVDD)进行处理,该描述通过将标准音频模式映射到紧凑的潜在超球体来学习标准音频模式。对策划的低音和吉他数据集的评估将标准和残差自动编码器变体与隔离森林(IF)和主成分分析(PCA)方法等基线进行比较。结果表明,我们的Deep SVDD模型,特别是残差自动编码器变体,提供了改进的异常分离,特别是对于较大的变化。这项研究为处理不同的音频样本提供了一种灵活的、完全无监督的解决方案,克服了以前的结构和输入限制,同时通过基于距离的潜在空间评分实现了有效的模式识别。
摘要:This paper introduces an unsupervised framework for detecting audio patterns in musical samples (loops) through anomaly detection techniques, addressing challenges in music information retrieval (MIR). Existing methods are often constrained by reliance on handcrafted features, domain-specific limitations, or dependence on iterative user interaction. We address these limitations through an architecture combining deep feature extraction with unsupervised anomaly detection. Our approach leverages a pre-trained Hierarchical Token-semantic Audio Transformer (HTS-AT), paired with a Feature Fusion Mechanism (FFM), to generate representations from variable-length audio loops. These embeddings are processed using one-class Deep Support Vector Data Description (Deep SVDD), which learns normative audio patterns by mapping them to a compact latent hypersphere. Evaluations on curated bass and guitar datasets compare standard and residual autoencoder variants against baselines like Isolation Forest (IF) and and principle component analysis (PCA) methods. Results show our Deep SVDD models, especially the residual autoencoder variant, deliver improved anomaly separation, particularly for larger variations. This research contributes a flexible, fully unsupervised solution for processing diverse audio samples, overcoming previous structural and input limitations while enabling effective pattern identification through distance-based latent space scoring.


【32】 4,500 Seconds: Small Data Training Approaches for Deep UAV Audio  Classification

标题: 4,500秒:深度无人机音频分类的小数据训练方法
链接:https://arxiv.org/abs/2505.23782
作者: Andrew P. Berg,  Qian Zhang,  Mia Y. Wang 
备注:Accepted at the 14th International Conference on Data Science, Technology, and Applications (DATA), 2025
摘要:无人机的使用预计将在未来十年激增,这就需要加强安全措施,以防止侵犯领空和安全威胁。本研究探讨了无人机分类的深度学习方法,重点关注数据稀缺性这一关键问题。为了研究这一点,我们选择使用总共4,500秒的音频样本来训练模型,这些音频样本均匀分布在9类数据集中。我们利用参数有效微调(PEFT)和数据增强来缓解数据稀缺性。本文实现并比较了卷积神经网络(CNN)和基于注意力的Transformers的使用。我们的研究结果表明,CNN的准确率比Transformers高1- 2%,同时仍然具有更高的计算效率。然而,这些早期的发现指出了使用Transformers模型的潜力;这表明,通过更多的数据和进一步的优化,它们可以胜过CNN。未来的工作旨在扩大数据集,以更好地了解这些方法之间的权衡。
摘要:Unmanned aerial vehicle (UAV) usage is expected to surge in the coming decade, raising the need for heightened security measures to prevent airspace violations and security threats. This study investigates deep learning approaches to UAV classification focusing on the key issue of data scarcity. To investigate this we opted to train the models using a total of 4,500 seconds of audio samples, evenly distributed across a 9-class dataset. We leveraged parameter efficient fine-tuning (PEFT) and data augmentations to mitigate the data scarcity. This paper implements and compares the use of convolutional neural networks (CNNs) and attention-based transformers. Our results show that, CNNs outperform transformers by 1-2\% accuracy, while still being more computationally efficient. These early findings, however, point to potential in using transformers models; suggesting that with more data and further optimizations they could outperform CNNs. Future works aims to upscale the dataset to better understand the trade-offs between these approaches.


【33】 Unified AI for Accurate Audio Anomaly Detection

标题: 统一人工智能用于准确的音频异常检测
链接:https://arxiv.org/abs/2505.23781
作者: Hamideh Khaleghpour,  Brett McKinney 
备注:6 pages, 14 figures. Based on original research. Submitted to arXiv for public preprint
摘要:本文提出了一个统一的AI框架,通过集成先进的降噪,特征提取和机器学习建模技术,高精度的音频异常检测。该方法结合了频谱减法和自适应滤波来增强音频质量,然后使用传统方法(如MFCC)和深度嵌入预训练模型(如OpenL3)进行特征提取。建模管道结合了经典模型(SVM,随机森林),深度学习架构(CNN)和集成方法,以提高鲁棒性和准确性。在包括TORGO和LibriSpeech在内的基准数据集上进行评估,该框架在精确度,召回率和模糊语音与正常语音的分类方面表现出优异的性能。这项工作解决了噪声环境和实时应用中的挑战,并为基于音频的异常检测提供了一个可扩展的解决方案。
摘要:This paper presents a unified AI framework for high-accuracy audio anomaly detection by integrating advanced noise reduction, feature extraction, and machine learning modeling techniques. The approach combines spectral subtraction and adaptive filtering to enhance audio quality, followed by feature extraction using traditional methods like MFCCs and deep embeddings from pre-trained models such as OpenL3. The modeling pipeline incorporates classical models (SVM, Random Forest), deep learning architectures (CNNs), and ensemble methods to boost robustness and accuracy. Evaluated on benchmark datasets including TORGO and LibriSpeech, the proposed framework demonstrates superior performance in precision, recall, and classification of slurred vs. normal speech. This work addresses challenges in noisy environments and real-time applications and provides a scalable solution for audio-based anomaly detection.


【34】 More-than-Human Storytelling: Designing Longitudinal Narrative  Engagements with Generative AI

标题: 比人类更多的讲故事:用生成性人工智能设计纵向叙事参与
链接:https://arxiv.org/abs/2505.23780
作者: Émilie Fabre,  Katie Seaborn,  Shuta Koiwai,  Mizuki Watanabe,  Paul Riesch 
备注:CHI EA '25
摘要:与生成式AI(GenAI)讲故事代理的纵向参与是一个及时但不太明确的领域。我们使用“Dreamsmithy”探索了多代人的体验,这是一个日常的梦想制作应用程序,参与者(N = 28)每天与AI叙述者“Makoto”共同创作故事。通过为期两周的日记研究捕捉反思和互动。自反性主题分析揭示了诸如“摇摆不定的矛盾心理”和“社会时间关系”等主题,突出了随着时间的推移,个人和人工智能叙述者之间出现的复杂动态。研究结果表明,虽然人们欣赏个人笔记,反思的机会和人工智能的创造力,但叙事连贯性和控制的局限性偶尔会导致挫折感。这些结果强调了GenAI在纵向讲故事方面的潜力,但也提出了关于用户代理和道德的关键问题。我们为开发自适应的、超越人类的讲故事系统提供了初步的经验见解和设计考虑。
摘要:Longitudinal engagement with generative AI (GenAI) storytelling agents is a timely but less charted domain. We explored multi-generational experiences with "Dreamsmithy," a daily dream-crafting app, where participants (N = 28) co-created stories with AI narrator "Makoto" every day. Reflections and interactions were captured through a two-week diary study. Reflexive thematic analysis revealed themes likes "oscillating ambivalence" and "socio-chronological bonding," highlighting the complex dynamics that emerged between individuals and the AI narrator over time. Findings suggest that while people appreciated the personal notes, opportunities for reflection, and AI creativity, limitations in narrative coherence and control occasionally caused frustration. The results underscore the potential of GenAI for longitudinal storytelling, but also raise critical questions about user agency and ethics. We contribute initial empirical insights and design considerations for developing adaptive, more-than-human storytelling systems.


【35】 Automatic classification of stop realisation with wav2vec2.0

标题: 使用wav2vec2.0自动分类停止实现
链接:https://arxiv.org/abs/2505.23688
作者: James Tanner,  Morgan Sonderegger,  Jane Stuart-Smith,  Jeff Mielke,  Tyler Kendall 
备注:Accepted for Interspeech 2025. 5 pages, 3 figures
摘要:现代语音研究经常使用自动工具来注释语音数据,但用于注释许多可变语音现象的工具却很少。与此同时,预训练的自监督模型,如wav2vec2.0,已被证明在语音分类任务中表现良好,并潜在地编码细粒度的语音信息。我们证明了wav2vec2.0模型可以被训练来自动分类停止突发存在,在英语和日语中具有很高的准确性,在精心策划和未经准备的语音语料库中都是鲁棒的。停止实现中的可变性模式通过自动注释复制,并且紧密遵循手动注释的模式。这些结果表明,预先训练的语音模型作为语音语料库数据的自动注释和处理工具的潜力,使研究人员能够相对轻松地“扩大”语音研究的范围。
摘要:Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. At the same time, pre-trained self-supervised models, such as wav2vec2.0, have been shown to perform well at speech classification tasks and latently encode fine-grained phonetic information. We demonstrate that wav2vec2.0 models can be trained to automatically classify stop burst presence with high accuracy in both English and Japanese, robust across both finely-curated and unprepared speech corpora. Patterns of variability in stop realisation are replicated with the automatic annotations, and closely follow those of manual annotations. These results demonstrate the potential of pre-trained speech models as tools for the automatic annotation and processing of speech corpus data, enabling researchers to 'scale-up' the scope of phonetic research with relative ease.


机器翻译由腾讯交互翻译提供,仅供参考