今日论文合集:cs.SD语音12篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Voila: Voice-Language Foundation Models for Real-Time Autonomous  Interaction and Voice Role-Play
标题: Voila:实时自主交互和语音角色扮演的语音语言基础模型
链接:https://arxiv.org/abs/2505.02707
作者: Yemin Shi,  Yu Shu,  Siwei Dong,  Guangyi Liu,  Jaward Sesay,  Jingwen Li,  Zhiting Hu 
备注:18 pages, 7 figures, Website: this https URL
摘要:一个无缝融入日常生活的语音AI代理将以自主、实时和情感表达的方式与人类互动。它不仅会对命令做出反应,还会不断地倾听、推理和主动回应,从而促进流畅、动态和情感共鸣的互动。我们介绍Voila,这是一系列大型语音语言基础模型,朝着这一愿景迈出了一步。Voila超越了传统的管道系统,采用了新的端到端架构,实现了全双工、低延迟的对话,同时保留了丰富的声音细微差别,如音调、节奏和情感。它的响应延迟仅为195毫秒,超过了人类的平均响应时间。它的分层多尺度Transformer将大型语言模型(LLM)的推理功能与强大的声学建模集成在一起,从而实现自然的、具有人物感知的语音生成--用户只需编写文本指令即可定义说话者的身份、音调和其他特征。此外,Voila支持超过100万个预先构建的语音,并从短至10秒的简短音频样本中高效定制新的语音。除了口语对话之外,Voila还被设计为一个统一的模型,用于各种基于语音的应用程序,包括自动语音识别(ASR),文本到语音(TTS),以及最小的适应,多语言语音翻译。Voila是完全开源的,以支持开放式研究并加速下一代人机交互的进展。
摘要:A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, reason, and respond proactively, fostering fluid, dynamic, and emotionally resonant interactions. We introduce Voila, a family of large voice-language foundation models that make a step towards this vision. Voila moves beyond traditional pipeline systems by adopting a new end-to-end architecture that enables full-duplex, low-latency conversations while preserving rich vocal nuances such as tone, rhythm, and emotion. It achieves a response latency of just 195 milliseconds, surpassing the average human response time. Its hierarchical multi-scale Transformer integrates the reasoning capabilities of large language models (LLMs) with powerful acoustic modeling, enabling natural, persona-aware voice generation -- where users can simply write text instructions to define the speaker's identity, tone, and other characteristics. Moreover, Voila supports over one million pre-built voices and efficient customization of new ones from brief audio samples as short as 10 seconds. Beyond spoken dialogue, Voila is designed as a unified model for a wide range of voice-based applications, including automatic speech recognition (ASR), Text-to-Speech (TTS), and, with minimal adaptation, multilingual speech translation. Voila is fully open-sourced to support open research and accelerate progress toward next-generation human-machine interactions.


【2】 fastabx: A library for efficient computation of ABX discriminability

标题: fastabx:用于有效计算ABX辨别性的库
链接:https://arxiv.org/abs/2505.02692
作者: Maxime Poli,  Emmanuel Chemla,  Emmanuel Dupoux 
备注:8 pages, 6 figures
摘要:我们介绍fastabx,一个用于构建ABX判别任务的高性能Python库。ABX是对感兴趣的通用类别之间的分离的度量。它已被广泛用于评估语音辨别力的自我监督的语音表示。然而,由于缺乏适当的工具,其更广泛的采用受到限制。fastabx通过提供一个能够构建任何类型的ABX任务的框架来解决这一差距,同时在任务创建和计算表示之间的距离方面提供快速开发周期所需的效率。我们相信,fastabx将成为更广泛的表征学习社区的宝贵资源,使研究人员能够系统地研究在语音处理之外的几个领域中,可以直接从学习的表征中提取哪些信息。源代码可在https://github.com/bootphon/fastabx上获得。
摘要:We introduce fastabx, a high-performance Python library for building ABX discrimination tasks. ABX is a measure of the separation between generic categories of interest. It has been used extensively to evaluate phonetic discriminability in self-supervised speech representations. However, its broader adoption has been limited by the absence of adequate tools. fastabx addresses this gap by providing a framework capable of constructing any type of ABX task while delivering the efficiency necessary for rapid development cycles, both in task creation and in calculating distances between representations. We believe that fastabx will serve as a valuable resource for the broader representation learning community, enabling researchers to systematically investigate what information can be directly extracted from learned representations across several domains beyond speech processing. The source code is available at https://github.com/bootphon/fastabx.


【3】 LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive  Streaming Speech Synthesis

标题: LLaMA-Omni 2:基于LLM的实时语音聊天机器人,具有自回归流语音合成
链接:https://arxiv.org/abs/2505.02625
作者: Qingkai Fang,  Yan Zhou,  Shoutao Guo,  Shaolei Zhang,  Yang Feng 
备注:Preprint. Project: this https URL
摘要:实时、智能、自然的语音交互是下一代人机交互的重要组成部分。最近的进展展示了基于大型语言模型(LLM)构建智能语音聊天机器人的潜力。在本文中,我们介绍了LLaMA-Omni 2,一系列语音语言模型(SpeechLM),从0.5B到14 B参数,能够实现高质量的实时语音交互。LLaMA-Omni 2基于Qwen2.5系列模型构建,集成了语音编码器和自回归流语音解码器。尽管LLaMA-Omni 2只在20万个多轮语音对话样本上进行了训练,但它在几个口语问答和语音指令方面表现出了强大的性能,超过了之前最先进的SpeechLM,如GLM-4-Voice,后者是在数百万小时的语音数据上进行训练的。
摘要:Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs). In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving high-quality real-time speech interaction. LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder. Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-the-art SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data.


【4】 Automatic Proficiency Assessment in L2 English Learners

标题: 二语英语学习者的自动能力评估
链接:https://arxiv.org/abs/2505.02615
作者: Armita Mohammadi,  Alessandro Lameiras Koerich,  Laureano Moro-Velazquez,  Patrick Cardinal 
备注:6 pages
摘要:英语第二语言能力通常由英语教师或专家评估者进行感知评估,具有内在的内部和内部差异性。本文探讨了用于综合L2水平评估的深度学习技术,解决了语音信号及其相应的转录。我们使用不同的架构分析口语水平分类预测,包括2D CNN,基于频率的CNN,ResNet和预训练的wav2vec 2.0模型。此外,我们研究了基于文本的能力评估微调的BERT语言模型的资源限制。最后,我们解决了自发对话评估的复杂任务,通过wav2vec 2.0和BERT模型的单独应用程序管理长格式音频和扬声器交互。EFCamDat和ANGLISH数据集以及私人数据集的实验结果突出了深度学习的潜力,特别是预训练的wav2vec 2.0模型,用于强大的自动化L2水平评估。
摘要:Second language proficiency (L2) in English is usually perceptually evaluated by English teachers or expert evaluators, with the inherent intra- and inter-rater variability. This paper explores deep learning techniques for comprehensive L2 proficiency assessment, addressing both the speech signal and its correspondent transcription. We analyze spoken proficiency classification prediction using diverse architectures, including 2D CNN, frequency-based CNN, ResNet, and a pretrained wav2vec 2.0 model. Additionally, we examine text-based proficiency assessment by fine-tuning a BERT language model within resource constraints. Finally, we tackle the complex task of spontaneous dialogue assessment, managing long-form audio and speaker interactions through separate applications of wav2vec 2.0 and BERT models. Results from experiments on EFCamDat and ANGLISH datasets and a private dataset highlight the potential of deep learning, especially the pretrained wav2vec 2.0 model, for robust automated L2 proficiency evaluation.


【5】 Bemba Speech Translation: Exploring a Low-Resource African Language

标题: 本巴语音翻译:探索资源匮乏的非洲语言
链接:https://arxiv.org/abs/2505.02518
作者: Muhammad Hazim Al Farouq,  Aman Kassahun Wassie,  Yasmin Moslem 
备注:IWSLT 2025
摘要:本文介绍了我们的系统提交给国际会议口语翻译(IWITH 2025),低资源语言轨道,即本巴语到英语的语音翻译。我们建立了基于Whisper和NLLB-200的级联语音翻译系统,并采用了数据增强技术,如回译。我们研究了使用合成数据的效果并讨论了我们的实验设置。
摘要:This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2025), low-resource languages track, namely for Bemba-to-English speech translation. We built cascaded speech translation systems based on Whisper and NLLB-200, and employed data augmentation techniques, such as back-translation. We investigate the effect of using synthetic data and discuss our experimental setup.


【6】 Efficient Continual Learning in Keyword Spotting using Binary Neural  Networks

标题: 使用二进制神经网络进行关键词识别的高效连续学习
链接:https://arxiv.org/abs/2505.02469
作者: Quynh Nguyen-Phuong Vu,  Luciano Sebastian Martinez-Rau,  Yuxuan Zhang,  Nho-Duc Tran,  Bengt Oelmann,  Michele Magno,  Sebastian Bader 
备注:Accepted for publication on "2025 IEEE Sensors Applications Symposium"
摘要:关键字定位(KWS)是一项重要功能,可实现与无处不在的智能设备的交互。然而,在资源有限的设备中,KWS模型通常是静态的,因此不能适应新的场景,例如添加的关键字。为了克服这个问题,我们提出了一种基于二进制神经网络(BNN)的KWS连续学习(CL)方法。该框架利用BNN减少的计算和内存需求,同时结合了能够随着时间的推移无缝集成新关键字的技术。这项研究评估了16类用例的7种CL技术,报告了单个附加关键字的准确率超过95%,另外4个类的准确率高达86%。对CL阶段训练样本量的敏感性以及计算复杂性的差异正在进行评估。这些评估表明,基于批处理的算法对CL数据集的大小更敏感,并且计算复杂性之间的差异微不足道。这些研究结果突出了开发一种有效的和计算效率高的技术,不断整合新的关键字在KWS应用程序,是与资源受限的设备兼容的潜力。
摘要:Keyword spotting (KWS) is an essential function that enables interaction with ubiquitous smart devices. However, in resource-limited devices, KWS models are often static and can thus not adapt to new scenarios, such as added keywords. To overcome this problem, we propose a Continual Learning (CL) approach for KWS built on Binary Neural Networks (BNNs). The framework leverages the reduced computation and memory requirements of BNNs while incorporating techniques that enable the seamless integration of new keywords over time. This study evaluates seven CL techniques on a 16-class use case, reporting an accuracy exceeding 95% for a single additional keyword and up to 86% for four additional classes. Sensitivity to the amount of training samples in the CL phase, and differences in computational complexities are being evaluated. These evaluations demonstrate that batch-based algorithms are more sensitive to the CL dataset size, and that differences between the computational complexities are insignificant. These findings highlight the potential of developing an effective and computationally efficient technique for continuously integrating new keywords in KWS applications that is compatible with resource-constrained devices.


【7】 VAEmo: Efficient Representation Learning for Visual-Audio Emotion with  Knowledge Injection

标题: VAEMo:通过知识注入实现视觉音频情感的高效表示学习
链接:https://arxiv.org/abs/2505.02331
作者: Hao Cheng,  Zhiwei Zhao,  Yichao He,  Zhenzhen Hu,  Jia Li,  Meng Wang,  Richang Hong 
备注:Source code and pre-trained models will be available at this https URL
摘要:视听情绪识别(AVER)的目的是从非语言的视听(VA)线索推断人类的情绪,提供模态互补和语言不可知的优势。然而,由于情感表达的固有模糊性,跨模态表达差异以及可靠注释数据的稀缺性,AVER仍然具有挑战性。最近的自监督AVER方法引入了强大的多模态表示,但它们主要依赖于特定于模态的编码器和粗略的内容级对齐,限制了细粒度的情感语义建模。为了解决这些问题,我们提出了VAEmo,一个有效的两阶段框架,以情感为中心的联合VA表示学习与外部知识注入。在第一阶段,通过掩蔽重建和对比目标,在大规模以说话者为中心的VA语料库上预训练统一和轻量级的表示网络,减轻模态差距,并学习表达,补充表示,而无需情感标签。在第二阶段,多模态大型语言模型根据我们精心设计的思想链自动生成详细的情感描述,仅提示VA样本的一小部分;然后通过双路径对比学习将其相应的嵌入与VA表示对齐,从而注入这些丰富的文本语义,进一步弥合情感差距。对多个下游AVER基准测试的广泛实验表明,VAEmo以紧凑的设计实现了最先进的性能,突出了统一的跨模态编码和情感感知语义指导的优势,可用于高效,可推广的VA情感表示。
摘要:Audiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However, AVER remains challenging due to the inherent ambiguity of emotional expressions, cross-modal expressive disparities, and the scarcity of reliably annotated data. Recent self-supervised AVER approaches have introduced strong multimodal representations, yet they predominantly rely on modality-specific encoders and coarse content-level alignment, limiting fine-grained emotional semantic modeling. To address these issues, we propose VAEmo, an efficient two-stage framework for emotion-centric joint VA representation learning with external knowledge injection. In Stage 1, a unified and lightweight representation network is pre-trained on large-scale speaker-centric VA corpora via masked reconstruction and contrastive objectives, mitigating the modality gap and learning expressive, complementary representations without emotion labels. In Stage 2, multimodal large language models automatically generate detailed affective descriptions according to our well-designed chain-of-thought prompting for only a small subset of VA samples; these rich textual semantics are then injected by aligning their corresponding embeddings with VA representations through dual-path contrastive learning, further bridging the emotion gap. Extensive experiments on multiple downstream AVER benchmarks show that VAEmo achieves state-of-the-art performance with a compact design, highlighting the benefit of unified cross-modal encoding and emotion-aware semantic guidance for efficient, generalizable VA emotion representations.


【8】 MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface  Vibrations for Real-time Noise-Robust Speech Input

标题: MaskEditor:可拆卸的面罩表面振动夹式压电传感,用于实时噪音稳健的语音输入
链接:https://arxiv.org/abs/2505.02180
作者: Hirotaka Hiraki,  Jun Rekimoto 
备注:Augmented Humans 2025
摘要:口罩在医疗环境和传染病爆发期间是必不可少的,但会严重损害语言交流,特别是在有背景噪音的环境中。现有的解决方案通常需要大量的计算资源或损害卫生和舒适性。我们提出了一种新的传感方法,通过使用压电传感器检测面罩表面振动来仅捕获佩戴者的声音。我们开发的设备MaskClip采用不锈钢夹,并带有最佳定位的压电传感器,可选择性地捕获语音振动,同时从本质上过滤掉环境噪声。评估实验表明,与传统麦克风相比,在嘈杂的环境中具有6.1%的低字符错误率的优越性能。102名参与者的主观评价也显示出很高的满意度。这种方法显示出在穿戴防护设备时必须保持清晰语音通信的环境中的应用前景,例如医疗设施,洁净室和工业环境。
摘要:Masks are essential in medical settings and during infectious outbreaks but significantly impair speech communication, especially in environments with background noise. Existing solutions often require substantial computational resources or compromise hygiene and comfort. We propose a novel sensing approach that captures only the wearer's voice by detecting mask surface vibrations using a piezoelectric sensor. Our developed device, MaskClip, employs a stainless steel clip with an optimally positioned piezoelectric sensor to selectively capture speech vibrations while inherently filtering out ambient noise. Evaluation experiments demonstrated superior performance with a low Character Error Rate of 6.1\% in noisy environments compared to conventional microphones. Subjective evaluations by 102 participants also showed high satisfaction scores. This approach shows promise for applications in settings where clear voice communication must be maintained while wearing protective equipment, such as medical facilities, cleanrooms, and industrial environments.


【9】 Weakly-supervised Audio Temporal Forgery Localization via Progressive  Audio-language Co-learning Network

标题: 通过渐进式音频语言协同学习网络进行弱监督音频时态伪造本地化
链接:https://arxiv.org/abs/2505.01880
作者: Junyan Wu,  Wenbo Xu,  Wei Lu,  Xiangyang Luo,  Rui Yang,  Shize Guo 
备注:9pages, 5figures. This paper has been accepted for IJCAI2025
摘要:音频时间伪造定位(ATFL)的目的是找到被故意修改的部分欺骗音频的精确伪造区域。现有的ATFL方法依赖于使用细粒度的注释来训练高效的网络,这在现实世界的场景中是昂贵且具有挑战性的。为了应对这一挑战,本文提出了一种渐进式音频语言协同学习网络(LOCO),采用协同学习和自我监督的方式,以提高本地化性能弱监督的情况下。具体而言,音频语言的协同学习模块首先被设计为捕获伪造共识功能,从时间和全局的角度对齐语义。在该模块中,通过使用话语级注释与可学习提示一起构造伪造感知提示,其可以动态地将语义先验纳入时间内容特征。此外,伪造本地化模块被应用到产生伪造建议的基础上融合伪造类激活序列。最后,引入渐进细化策略来生成伪帧级标签,并利用监督语义对比学习来放大真实和虚假内容之间的语义区别,从而不断优化伪造感知特征。大量的实验表明,建议的LOCO实现SOTA性能的三个公共基准。
摘要:Audio temporal forgery localization (ATFL) aims to find the precise forgery regions of the partial spoof audio that is purposefully modified. Existing ATFL methods rely on training efficient networks using fine-grained annotations, which are obtained costly and challenging in real-world scenarios. To meet this challenge, in this paper, we propose a progressive audio-language co-learning network (LOCO) that adopts co-learning and self-supervision manners to prompt localization performance under weak supervision scenarios. Specifically, an audio-language co-learning module is first designed to capture forgery consensus features by aligning semantics from temporal and global perspectives. In this module, forgery-aware prompts are constructed by using utterance-level annotations together with learnable prompts, which can incorporate semantic priors into temporal content features dynamically. In addition, a forgery localization module is applied to produce forgery proposals based on fused forgery-class activation sequences. Finally, a progressive refinement strategy is introduced to generate pseudo frame-level labels and leverage supervised semantic contrastive learning to amplify the semantic distinction between real and fake content, thereby continuously optimizing forgery-aware features. Extensive experiments show that the proposed LOCO achieves SOTA performance on three public benchmarks.


【10】 FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech  Restoration

标题: FLOWER:用于通用语音恢复的基于流的估计高斯引导
链接:https://arxiv.org/abs/2505.01750
作者: Da-Hee Yang,  Jaeuk Lee,  Joon-Hyuk Chang 
摘要:我们介绍FLOWER,一种新颖的条件反射方法,设计用于语音恢复,将高斯指导集成到生成框架中。通过将干净语音变换成预定义的先验分布(例如,FLOWER使用归一化流网络(高斯分布)提取关键信息以指导生成模型。这种指导被纳入生成网络的每个块中,从而实现精确的恢复控制。实验结果表明,FLOWER在提高各种通用语音恢复任务的性能方面是有效的。
摘要:We introduce FLOWER, a novel conditioning method designed for speech restoration that integrates Gaussian guidance into generative frameworks. By transforming clean speech into a predefined prior distribution (e.g., Gaussian distribution) using a normalizing flow network, FLOWER extracts critical information to guide generative models. This guidance is incorporated into each block of the generative network, enabling precise restoration control. Experimental results demonstrate the effectiveness of FLOWER in improving performance across various general speech restoration tasks.


【11】 Low-Complexity Acoustic Scene Classification with Device Information in  the DCASE 2025 Challenge

标题: DUSE 2025挑战赛中使用设备信息进行低复杂度声学场景分类
链接:https://arxiv.org/abs/2505.01747
作者: Florian Schmid,  Paul Primus,  Toni Heittola,  Annamaria Mesaros,  Irene Martín-Morató,  Gerhard Widmer 
备注:Task Description Page: this https URL
摘要:本文介绍了DCASE 2025挑战赛的低复杂度声学场景分类与设备信息任务及其基线系统。今年的任务继续关注低复杂度模型、数据效率和前几个版本(2022- 2024)的设备不匹配,引入了一个关键变化:现在在推理时提供记录设备信息。这使得能够开发利用设备特性的特定于设备的模型--反映真实的部署场景,在这些场景中,模型是在了解底层硬件的情况下设计的。训练集与相应的DCASE 2024挑战中使用的25%子集相匹配,对外部数据的使用没有限制,突出了迁移学习作为中心主题。基线达到50.72%的准确率在这个十类问题与设备的一般模型,提高到51.89%时,使用可用的设备信息。
摘要:This paper presents the Low-Complexity Acoustic Scene Classification with Device Information Task of the DCASE 2025 Challenge and its baseline system. Continuing the focus on low-complexity models, data efficiency, and device mismatch from previous editions (2022--2024), this year's task introduces a key change: recording device information is now provided at inference time. This enables the development of device-specific models that leverage device characteristics -- reflecting real-world deployment scenarios in which a model is designed with awareness of the underlying hardware. The training set matches the 25% subset used in the corresponding DCASE 2024 challenge, with no restrictions on external data use, highlighting transfer learning as a central topic. The baseline achieves 50.72% accuracy on this ten-class problem with a device-general model, improving to 51.89% when using the available device information.


【12】 Transfer Learning-Based Deep Residual Learning for Speech Recognition in  Clean and Noisy Environments

标题: 基于迁移学习的深度剩余学习用于清洁和噪音环境中的语音识别
链接:https://arxiv.org/abs/2505.01632
作者: Noussaiba Djeffal,  Djamel Addou,  Hamza Kheddar,  Sid Ahmed Selouani 
备注:None
摘要:非平稳环境噪声对语音识别的影响一直是语音识别领域的研究热点。尽管取得了进展,但这一挑战仍然是一个主要关切。最近,数据驱动的监督方法,如深度神经网络,已经成为传统无监督方法的有前途的替代方案。通过广泛的培训,这些方法有可能克服各种现实生活中的声学环境所带来的挑战。在这种情况下,本文介绍了一种新的神经框架,将一个强大的前端到ASR系统在干净和嘈杂的环境。利用Aurora-2语音数据库,作者评估了Mel频率声学特征集的有效性,采用基于残差神经网络(ResNet)的迁移学习方法。实验结果表明,与卷积神经网络(CNN)和长短期记忆(LSTM)网络相比,识别准确率有显著提高。他们在干净模式下的准确率为98.94%,在嘈杂模式下的准确率为91.21%。
摘要:Addressing the detrimental impact of non-stationary environmental noise on automatic speech recognition (ASR) has been a persistent and significant research focus. Despite advancements, this challenge continues to be a major concern. Recently, data-driven supervised approaches, such as deep neural networks, have emerged as promising alternatives to traditional unsupervised methods. With extensive training, these approaches have the potential to overcome the challenges posed by diverse real-life acoustic environments. In this light, this paper introduces a novel neural framework that incorporates a robust frontend into ASR systems in both clean and noisy environments. Utilizing the Aurora-2 speech database, the authors evaluate the effectiveness of an acoustic feature set for Mel-frequency, employing the approach of transfer learning based on Residual neural network (ResNet). The experimental results demonstrate a significant improvement in recognition accuracy compared to convolutional neural networks (CNN) and long short-term memory (LSTM) networks. They achieved accuracies of 98.94% in clean and 91.21% in noisy mode.


eess.AS音频处理

【1】 FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech  Restoration
标题: FLOWER:用于通用语音恢复的基于流的估计高斯引导
链接:https://arxiv.org/abs/2505.01750
作者: Da-Hee Yang,  Jaeuk Lee,  Joon-Hyuk Chang 
摘要:我们介绍FLOWER,一种新颖的条件反射方法,设计用于语音恢复,将高斯指导集成到生成框架中。通过将干净语音变换成预定义的先验分布(例如,FLOWER使用归一化流网络(高斯分布)提取关键信息以指导生成模型。这种指导被纳入生成网络的每个块中,从而实现精确的恢复控制。实验结果表明,FLOWER在提高各种通用语音恢复任务的性能方面是有效的。
摘要:We introduce FLOWER, a novel conditioning method designed for speech restoration that integrates Gaussian guidance into generative frameworks. By transforming clean speech into a predefined prior distribution (e.g., Gaussian distribution) using a normalizing flow network, FLOWER extracts critical information to guide generative models. This guidance is incorporated into each block of the generative network, enabling precise restoration control. Experimental results demonstrate the effectiveness of FLOWER in improving performance across various general speech restoration tasks.


【2】 Low-Complexity Acoustic Scene Classification with Device Information in  the DCASE 2025 Challenge

标题: DUSE 2025挑战赛中使用设备信息进行低复杂度声学场景分类
链接:https://arxiv.org/abs/2505.01747
作者: Florian Schmid,  Paul Primus,  Toni Heittola,  Annamaria Mesaros,  Irene Martín-Morató,  Gerhard Widmer 
备注:Task Description Page: this https URL
摘要:本文介绍了DCASE 2025挑战赛的低复杂度声学场景分类与设备信息任务及其基线系统。今年的任务继续关注低复杂度模型、数据效率和前几个版本(2022- 2024)的设备不匹配,引入了一个关键变化:现在在推理时提供记录设备信息。这使得能够开发利用设备特性的特定于设备的模型--反映真实的部署场景,在这些场景中,模型是在了解底层硬件的情况下设计的。训练集与相应的DCASE 2024挑战中使用的25%子集相匹配,对外部数据的使用没有限制,突出了迁移学习作为中心主题。基线达到50.72%的准确率在这个十类问题与设备的一般模型,提高到51.89%时,使用可用的设备信息。
摘要:This paper presents the Low-Complexity Acoustic Scene Classification with Device Information Task of the DCASE 2025 Challenge and its baseline system. Continuing the focus on low-complexity models, data efficiency, and device mismatch from previous editions (2022--2024), this year's task introduces a key change: recording device information is now provided at inference time. This enables the development of device-specific models that leverage device characteristics -- reflecting real-world deployment scenarios in which a model is designed with awareness of the underlying hardware. The training set matches the 25% subset used in the corresponding DCASE 2024 challenge, with no restrictions on external data use, highlighting transfer learning as a central topic. The baseline achieves 50.72% accuracy on this ten-class problem with a device-general model, improving to 51.89% when using the available device information.


【3】 Transfer Learning-Based Deep Residual Learning for Speech Recognition in  Clean and Noisy Environments

标题: 基于迁移学习的深度剩余学习用于清洁和噪音环境中的语音识别
链接:https://arxiv.org/abs/2505.01632
作者: Noussaiba Djeffal,  Djamel Addou,  Hamza Kheddar,  Sid Ahmed Selouani 
备注:None
摘要:非平稳环境噪声对语音识别的影响一直是语音识别领域的研究热点。尽管取得了进展,但这一挑战仍然是一个主要关切。最近,数据驱动的监督方法,如深度神经网络,已经成为传统无监督方法的有前途的替代方案。通过广泛的培训,这些方法有可能克服各种现实生活中的声学环境所带来的挑战。在这种情况下,本文介绍了一种新的神经框架,将一个强大的前端到ASR系统在干净和嘈杂的环境。利用Aurora-2语音数据库,作者评估了Mel频率声学特征集的有效性,采用基于残差神经网络(ResNet)的迁移学习方法。实验结果表明,与卷积神经网络(CNN)和长短期记忆(LSTM)网络相比,识别准确率有显著提高。他们在干净模式下的准确率为98.94%,在嘈杂模式下的准确率为91.21%。
摘要:Addressing the detrimental impact of non-stationary environmental noise on automatic speech recognition (ASR) has been a persistent and significant research focus. Despite advancements, this challenge continues to be a major concern. Recently, data-driven supervised approaches, such as deep neural networks, have emerged as promising alternatives to traditional unsupervised methods. With extensive training, these approaches have the potential to overcome the challenges posed by diverse real-life acoustic environments. In this light, this paper introduces a novel neural framework that incorporates a robust frontend into ASR systems in both clean and noisy environments. Utilizing the Aurora-2 speech database, the authors evaluate the effectiveness of an acoustic feature set for Mel-frequency, employing the approach of transfer learning based on Residual neural network (ResNet). The experimental results demonstrate a significant improvement in recognition accuracy compared to convolutional neural networks (CNN) and long short-term memory (LSTM) networks. They achieved accuracies of 98.94% in clean and 91.21% in noisy mode.


【4】 fastabx: A library for efficient computation of ABX discriminability

标题: fastabx:用于有效计算ABX辨别性的库
链接:https://arxiv.org/abs/2505.02692
作者: Maxime Poli,  Emmanuel Chemla,  Emmanuel Dupoux 
备注:8 pages, 6 figures
摘要:我们介绍fastabx,一个用于构建ABX判别任务的高性能Python库。ABX是对感兴趣的通用类别之间的分离的度量。它已被广泛用于评估语音辨别力的自我监督的语音表示。然而,由于缺乏适当的工具,其更广泛的采用受到限制。fastabx通过提供一个能够构建任何类型的ABX任务的框架来解决这一差距,同时在任务创建和计算表示之间的距离方面提供快速开发周期所需的效率。我们相信,fastabx将成为更广泛的表征学习社区的宝贵资源,使研究人员能够系统地研究在语音处理之外的几个领域中,可以直接从学习的表征中提取哪些信息。源代码可在https://github.com/bootphon/fastabx上获得。
摘要:We introduce fastabx, a high-performance Python library for building ABX discrimination tasks. ABX is a measure of the separation between generic categories of interest. It has been used extensively to evaluate phonetic discriminability in self-supervised speech representations. However, its broader adoption has been limited by the absence of adequate tools. fastabx addresses this gap by providing a framework capable of constructing any type of ABX task while delivering the efficiency necessary for rapid development cycles, both in task creation and in calculating distances between representations. We believe that fastabx will serve as a valuable resource for the broader representation learning community, enabling researchers to systematically investigate what information can be directly extracted from learned representations across several domains beyond speech processing. The source code is available at https://github.com/bootphon/fastabx.


【5】 LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive  Streaming Speech Synthesis

标题: LLaMA-Omni 2:基于LLM的实时语音聊天机器人,具有自回归流语音合成
链接:https://arxiv.org/abs/2505.02625
作者: Qingkai Fang,  Yan Zhou,  Shoutao Guo,  Shaolei Zhang,  Yang Feng 
备注:Preprint. Project: this https URL
摘要:实时、智能、自然的语音交互是下一代人机交互的重要组成部分。最近的进展展示了基于大型语言模型(LLM)构建智能语音聊天机器人的潜力。在本文中,我们介绍了LLaMA-Omni 2,一系列语音语言模型(SpeechLM),从0.5B到14 B参数,能够实现高质量的实时语音交互。LLaMA-Omni 2基于Qwen2.5系列模型构建,集成了语音编码器和自回归流语音解码器。尽管LLaMA-Omni 2只在20万个多轮语音对话样本上进行了训练,但它在几个口语问答和语音指令方面表现出了强大的性能,超过了之前最先进的SpeechLM,如GLM-4-Voice,后者是在数百万小时的语音数据上进行训练的。
摘要:Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs). In this paper, we introduce LLaMA-Omni 2, a series of speech language models (SpeechLMs) ranging from 0.5B to 14B parameters, capable of achieving high-quality real-time speech interaction. LLaMA-Omni 2 is built upon the Qwen2.5 series models, integrating a speech encoder and an autoregressive streaming speech decoder. Despite being trained on only 200K multi-turn speech dialogue samples, LLaMA-Omni 2 demonstrates strong performance on several spoken question answering and speech instruction following benchmarks, surpassing previous state-of-the-art SpeechLMs like GLM-4-Voice, which was trained on millions of hours of speech data.


【6】 Automatic Proficiency Assessment in L2 English Learners

标题: 二语英语学习者的自动能力评估
链接:https://arxiv.org/abs/2505.02615
作者: Armita Mohammadi,  Alessandro Lameiras Koerich,  Laureano Moro-Velazquez,  Patrick Cardinal 
备注:6 pages
摘要:英语第二语言能力通常由英语教师或专家评估者进行感知评估,具有内在的内部和内部差异性。本文探讨了用于综合L2水平评估的深度学习技术,解决了语音信号及其相应的转录。我们使用不同的架构分析口语水平分类预测,包括2D CNN,基于频率的CNN,ResNet和预训练的wav2vec 2.0模型。此外,我们研究了基于文本的能力评估微调的BERT语言模型的资源限制。最后,我们解决了自发对话评估的复杂任务,通过wav2vec 2.0和BERT模型的单独应用程序管理长格式音频和扬声器交互。EFCamDat和ANGLISH数据集以及私人数据集的实验结果突出了深度学习的潜力,特别是预训练的wav2vec 2.0模型,用于强大的自动化L2水平评估。
摘要:Second language proficiency (L2) in English is usually perceptually evaluated by English teachers or expert evaluators, with the inherent intra- and inter-rater variability. This paper explores deep learning techniques for comprehensive L2 proficiency assessment, addressing both the speech signal and its correspondent transcription. We analyze spoken proficiency classification prediction using diverse architectures, including 2D CNN, frequency-based CNN, ResNet, and a pretrained wav2vec 2.0 model. Additionally, we examine text-based proficiency assessment by fine-tuning a BERT language model within resource constraints. Finally, we tackle the complex task of spontaneous dialogue assessment, managing long-form audio and speaker interactions through separate applications of wav2vec 2.0 and BERT models. Results from experiments on EFCamDat and ANGLISH datasets and a private dataset highlight the potential of deep learning, especially the pretrained wav2vec 2.0 model, for robust automated L2 proficiency evaluation.


【7】 Bemba Speech Translation: Exploring a Low-Resource African Language

标题: 本巴语音翻译:探索资源匮乏的非洲语言
链接:https://arxiv.org/abs/2505.02518
作者: Muhammad Hazim Al Farouq,  Aman Kassahun Wassie,  Yasmin Moslem 
备注:IWSLT 2025
摘要:本文介绍了我们的系统提交给国际会议口语翻译(IWITH 2025),低资源语言轨道,即本巴语到英语的语音翻译。我们建立了基于Whisper和NLLB-200的级联语音翻译系统,并采用了数据增强技术,如回译。我们研究了使用合成数据的效果并讨论了我们的实验设置。
摘要:This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2025), low-resource languages track, namely for Bemba-to-English speech translation. We built cascaded speech translation systems based on Whisper and NLLB-200, and employed data augmentation techniques, such as back-translation. We investigate the effect of using synthetic data and discuss our experimental setup.


【8】 MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface  Vibrations for Real-time Noise-Robust Speech Input

标题: MaskEditor:可拆卸的面罩表面振动夹式压电传感,用于实时噪音稳健的语音输入
链接:https://arxiv.org/abs/2505.02180
作者: Hirotaka Hiraki,  Jun Rekimoto 
备注:Augmented Humans 2025
摘要:口罩在医疗环境和传染病爆发期间是必不可少的,但会严重损害语言交流,特别是在有背景噪音的环境中。现有的解决方案通常需要大量的计算资源或损害卫生和舒适性。我们提出了一种新的传感方法,通过使用压电传感器检测面罩表面振动来仅捕获佩戴者的声音。我们开发的设备MaskClip采用不锈钢夹,并带有最佳定位的压电传感器,可选择性地捕获语音振动,同时从本质上过滤掉环境噪声。评估实验表明,与传统麦克风相比,在嘈杂的环境中具有6.1%的低字符错误率的优越性能。102名参与者的主观评价也显示出很高的满意度。这种方法显示出在穿戴防护设备时必须保持清晰语音通信的环境中的应用前景,例如医疗设施,洁净室和工业环境。
摘要:Masks are essential in medical settings and during infectious outbreaks but significantly impair speech communication, especially in environments with background noise. Existing solutions often require substantial computational resources or compromise hygiene and comfort. We propose a novel sensing approach that captures only the wearer's voice by detecting mask surface vibrations using a piezoelectric sensor. Our developed device, MaskClip, employs a stainless steel clip with an optimally positioned piezoelectric sensor to selectively capture speech vibrations while inherently filtering out ambient noise. Evaluation experiments demonstrated superior performance with a low Character Error Rate of 6.1\% in noisy environments compared to conventional microphones. Subjective evaluations by 102 participants also showed high satisfaction scores. This approach shows promise for applications in settings where clear voice communication must be maintained while wearing protective equipment, such as medical facilities, cleanrooms, and industrial environments.


【9】 Weakly-supervised Audio Temporal Forgery Localization via Progressive  Audio-language Co-learning Network

标题: 通过渐进式音频语言协同学习网络进行弱监督音频时态伪造本地化
链接:https://arxiv.org/abs/2505.01880
作者: Junyan Wu,  Wenbo Xu,  Wei Lu,  Xiangyang Luo,  Rui Yang,  Shize Guo 
备注:9pages, 5figures. This paper has been accepted for IJCAI2025
摘要:音频时间伪造定位(ATFL)的目的是找到被故意修改的部分欺骗音频的精确伪造区域。现有的ATFL方法依赖于使用细粒度的注释来训练高效的网络,这在现实世界的场景中是昂贵且具有挑战性的。为了应对这一挑战,本文提出了一种渐进式音频语言协同学习网络(LOCO),采用协同学习和自我监督的方式,以提高本地化性能弱监督的情况下。具体而言,音频语言的协同学习模块首先被设计为捕获伪造共识功能,从时间和全局的角度对齐语义。在该模块中,通过使用话语级注释与可学习提示一起构造伪造感知提示,其可以动态地将语义先验纳入时间内容特征。此外,伪造本地化模块被应用到产生伪造建议的基础上融合伪造类激活序列。最后,引入渐进细化策略来生成伪帧级标签,并利用监督语义对比学习来放大真实和虚假内容之间的语义区别,从而不断优化伪造感知特征。大量的实验表明,建议的LOCO实现SOTA性能的三个公共基准。
摘要:Audio temporal forgery localization (ATFL) aims to find the precise forgery regions of the partial spoof audio that is purposefully modified. Existing ATFL methods rely on training efficient networks using fine-grained annotations, which are obtained costly and challenging in real-world scenarios. To meet this challenge, in this paper, we propose a progressive audio-language co-learning network (LOCO) that adopts co-learning and self-supervision manners to prompt localization performance under weak supervision scenarios. Specifically, an audio-language co-learning module is first designed to capture forgery consensus features by aligning semantics from temporal and global perspectives. In this module, forgery-aware prompts are constructed by using utterance-level annotations together with learnable prompts, which can incorporate semantic priors into temporal content features dynamically. In addition, a forgery localization module is applied to produce forgery proposals based on fused forgery-class activation sequences. Finally, a progressive refinement strategy is introduced to generate pseudo frame-level labels and leverage supervised semantic contrastive learning to amplify the semantic distinction between real and fake content, thereby continuously optimizing forgery-aware features. Extensive experiments show that the proposed LOCO achieves SOTA performance on three public benchmarks.


机器翻译由腾讯交互翻译提供,仅供参考