今日论文合集:cs.SD语音10篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling
标题:通过增强说话人嵌入采样实现稳健的目标说话人数字化和分离
链接:https://arxiv.org/abs/2508.06393

作者:Md Asif Jalal, Luca Remaggi, Vasileios Moschopoulos, Thanasis Kotsiopoulos, Vandana Rajan, Karthikeyan Saravanan, Anastasis Drosou, Junho Heo, Hyuk Oh, Seokyeong Jeong
备注:Accepted to Interspeech 2025
摘要:传统的语音分离和说话人日志化方法依赖于目标说话人的先验知识或音频信号中预定数量的参与者。为了解决这些局限性,最近的进展集中在开发免注册的方法,能够识别目标,而无需显式的扬声器标记。本文介绍了一种新的方法来训练同时语音分离和diarization使用自动识别的目标说话人嵌入,在混合。我们提出的模型采用了一个双阶段的训练管道,旨在学习强大的扬声器表示功能,是弹性的背景噪声干扰。此外,我们提出了一个重叠的频谱损失函数,专门为提高日志精度重叠语音帧。实验结果表明,显着的性能增益相比,目前的SOTA基线,实现71%的相对改善DER和69%的cpWER。
摘要:Traditional speech separation and speaker diarization approaches rely on prior knowledge of target speakers or a predetermined number of participants in audio signals. To address these limitations, recent advances focus on developing enrollment-free methods capable of identifying targets without explicit speaker labeling. This work introduces a new approach to train simultaneous speech separation and diarization using automatic identification of target speaker embeddings, within mixtures. Our proposed model employs a dual-stage training pipeline designed to learn robust speaker representation features that are resilient to background noise interference. Furthermore, we present an overlapping spectral loss function specifically tailored for enhancing diarization accuracy during overlapped speech frames. Experimental results show significant performance gains compared to the current SOTA baseline, achieving 71% relative improvement in DER and 69% in cpWER.


【2】Improved Dysarthric Speech to Text Conversion via TTS Personalization
标题:通过TTC个性化改进结构障碍语音到文本的转换
链接:https://arxiv.org/abs/2508.06391

作者:Peter Mihajlik, Éva Székely, Piroska Barta, Máté Soma Kádár, Gergely Dobsinszki, László Tóth
摘要:我们提出了一个案例研究开发一个定制的语音到文本系统的匈牙利扬声器与严重的构音障碍。最先进的自动语音识别(ASR)模型与构音障碍语音的zero-shot转录斗争,产生高错误率。为了提高有限的真实构音障碍数据的性能,我们使用通过个性化文本到语音(TTS)系统生成的合成语音来微调ASR模型。我们介绍了一种方法,用于生成合成构音障碍的语音与控制的严重程度,通过利用给定的扬声器和扬声器嵌入插值的发病前录音,使ASR微调的连续的损伤。对真实的和合成的构音障碍语音进行微调,将字符错误率(CER)从36-51%(zero-shot)降低到7.3%。我们的单语FastConformer_Hu ASR模型在对相同数据进行微调时显著优于Whisper-turbo,并且包含合成语音有助于减少18%的相对CER。这些结果突出了个性化ASR系统在改善严重言语障碍患者无障碍环境方面的潜力。
摘要:We present a case study on developing a customized speech-to-text system for a Hungarian speaker with severe dysarthria. State-of-the-art automatic speech recognition (ASR) models struggle with zero-shot transcription of dysarthric speech, yielding high error rates. To improve performance with limited real dysarthric data, we fine-tune an ASR model using synthetic speech generated via a personalized text-to-speech (TTS) system. We introduce a method for generating synthetic dysarthric speech with controlled severity by leveraging premorbidity recordings of the given speaker and speaker embedding interpolation, enabling ASR fine-tuning on a continuum of impairments. Fine-tuning on both real and synthetic dysarthric speech reduces the character error rate (CER) from 36-51% (zero-shot) to 7.3%. Our monolingual FastConformer_Hu ASR model significantly outperforms Whisper-turbo when fine-tuned on the same data, and the inclusion of synthetic speech contributes to an 18% relative CER reduction. These results highlight the potential of personalized ASR systems for improving accessibility for individuals with severe speech impairments.


【3】SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
标题:SpeakerLM:利用多模式大型语言模型进行端到端多功能说话人规模化和识别
链接:https://arxiv.org/abs/2508.06372

作者:Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li
摘要:发言人日记和识别(SDR)任务旨在预测音频片段中的“谁在何时发言以及发言内容”,这在各种现实世界的多发言人场景中是一项至关重要的任务,例如会议转录和对话系统。现有的SDR系统通常采用级联框架,组合多个模块,例如说话人日记(SD)和自动语音识别(ASR)。级联系统受到几个限制,如错误传播,难以处理重叠的语音,以及缺乏联合优化探索SD和ASR任务之间的协同作用。为了解决这些局限性,我们引入了SpeakerLM,这是一种统一的SDR多模式大型语言模型,它以端到端的方式联合执行SD和ASR。此外,为了方便不同的现实世界的情况下,我们将一个灵活的扬声器注册机制到SpeakerLM,使SDR在不同的扬声器注册设置。SpeakerLM是在大规模真实数据上逐步开发的多阶段训练策略。大量的实验表明,SpeakerLM表现出强大的数据扩展能力和泛化能力,在域内和域外公共SDR基准测试中表现优于最先进的级联基线。此外,实验结果表明,所提出的说话人注册机制有效地确保了强大的SDR性能的SpeakerLM在不同的说话人注册条件和不同数量的注册说话人。
摘要:The Speaker Diarization and Recognition (SDR) task aims to predict "who spoke when and what" within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers.


【4】EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition
标题:CLARAugNet:用于语音情感识别的信号增强混合CNN-LSTM框架
链接:https://arxiv.org/abs/2508.06321

作者:Durjoy Chandra Paul, Gaurob Saha, Md Amjad Hossain
备注:To be published in ICCCNT 2025 (16th International Conference on Computing Communication and Networking Technologies)
摘要:语音中情感信号的识别对提高人机交互的有效性有着重要的影响。这项研究介绍了一种混合深度学习框架,它将长短期记忆(LSTM)层与一维卷积神经网络(1D-CNN)相结合,以实现可靠的语音情感识别(SER)。从语音信号中提取的特征的质量和多样性对SER系统的性能有很大的影响。一个全面的语音数据增强策略被用来结合传统的方法,如噪声添加,音调移动,和时间拉伸,与一个新的组合为基础的增强流水线,以提高泛化和减少过拟合。每个音频样本被转换成一个高维特征向量,使用均方根能量(RMSE),梅尔频率倒谱系数(MFCC),和零交叉率(ZCR)。我们的模型在IEMOCAP数据集上具有95.78\%的加权准确度和92.52\%的未加权准确度,并且具有ELU激活,具有96.75\%的加权准确度和91.28\%的未加权准确度。在RAVDESS数据集上,ReLU激活的加权准确率为94.53\%,未加权准确率为94.98\%,ELU激活的加权准确率为93.72\%,未加权准确率为94.64\%。这些结果突出了EAR-AugNet通过集成数据增强和混合建模在提高SER系统的鲁棒性和性能方面的有效性。
摘要:Recognizing emotional signals in speech has a significant impact on enhancing the effectiveness of human-computer interaction (HCI). This study introduces EmoAugNet, a hybrid deep learning framework, that incorporates Long Short-Term Memory (LSTM) layers with one-dimensional Convolutional Neural Networks (1D-CNN) to enable reliable Speech Emotion Recognition (SER). The quality and variety of the features that are taken from speech signals have a significant impact on how well SER systems perform. A comprehensive speech data augmentation strategy was used to combine both traditional methods, such as noise addition, pitch shifting, and time stretching, with a novel combination-based augmentation pipeline to enhance generalization and reduce overfitting. Each audio sample was transformed into a high-dimensional feature vector using root mean square energy (RMSE), Mel-frequency Cepstral Coefficient (MFCC), and zero-crossing rate (ZCR). Our model with ReLU activation has a weighted accuracy of 95.78\% and unweighted accuracy of 92.52\% on the IEMOCAP dataset and, with ELU activation, has a weighted accuracy of 96.75\% and unweighted accuracy of 91.28\%. On the RAVDESS dataset, we get a weighted accuracy of 94.53\% and 94.98\% unweighted accuracy for ReLU activation and 93.72\% weighted accuracy and 94.64\% unweighted accuracy for ELU activation. These results highlight EmoAugNet's effectiveness in improving the robustness and performance of SER systems through integated data augmentation and hybrid modeling.


【5】Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
标题:用于增强德语语音意图识别的大语言模型数据生成
链接:https://arxiv.org/abs/2508.06277

作者:Theresa Pekarek Rosin, Burak Can Kaplan, Stefan Wermter
备注:11 pages, 3 figures, accepted at KONVENS 2025
摘要:语音命令的意图识别(IR)对于人工智能(AI)辅助系统至关重要;然而,大多数现有方法仅限于短命令,并且主要针对英语开发。本文针对这些局限性,通过集中在IR从老年人的德语发言。我们提出了一种新的方法,该方法结合了一个适应的Whisper ASR模型,对老年人的德语语音(SVC-de)进行了微调,并在三个著名的大型语言模型(LLM)生成的合成文本数据集上训练了基于transformer的语言模型:LeoLM,Llama 3和ChatGPT。为了评估我们的方法的鲁棒性,我们用文本到语音模型生成合成语音,并进行广泛的跨数据集测试。我们的研究结果表明,合成LLM生成的数据显着提高了分类性能和鲁棒性,以不同的说话风格和看不见的词汇。值得注意的是,我们发现LeoLM,一个较小的,特定于领域的13 B LLM,在德国意图识别的数据集质量上超过了更大的ChatGPT(175 B)。我们的方法表明,生成式人工智能可以有效地弥合低资源领域的数据差距。我们提供数据生成和培训过程的详细文档,以确保透明度和可重复性。
摘要:Intent recognition (IR) for speech commands is essential for artificial intelligence (AI) assistant systems; however, most existing approaches are limited to short commands and are predominantly developed for English. This paper addresses these limitations by focusing on IR from speech by elderly German speakers. We propose a novel approach that combines an adapted Whisper ASR model, fine-tuned on elderly German speech (SVC-de), with Transformer-based language models trained on synthetic text datasets generated by three well-known large language models (LLMs): LeoLM, Llama3, and ChatGPT. To evaluate the robustness of our approach, we generate synthetic speech with a text-to-speech model and conduct extensive cross-dataset testing. Our results show that synthetic LLM-generated data significantly boosts classification performance and robustness to different speaking styles and unseen vocabulary. Notably, we find that LeoLM, a smaller, domain-specific 13B LLM, surpasses the much larger ChatGPT (175B) in dataset quality for German intent recognition. Our approach demonstrates that generative AI can effectively bridge data gaps in low-resource domains. We provide detailed documentation of our data generation and training process to ensure transparency and reproducibility.


【6】Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
标题:Llasa+:免费午餐,用于加速和流媒体基于大羊驼的语音合成
链接:https://arxiv.org/abs/2508.06262

作者:Wenjie Tian, Xinfa Zhu, Hanke Xie, Zhen Ye, Wei Xue, Lei Xie
摘要:文本到语音(TTS)的最新进展已经取得了令人印象深刻的自然性和灵活性,特别是随着大语言模型(LLM)为基础的方法的发展。然而,现有的自回归(AR)结构和大规模模型,如Llasa,仍然面临着推理延迟和流合成的重大挑战。为了克服这些局限性,我们引入了Llasa+,一个基于Llasa的加速流TTS模型。具体来说,为了加速生成过程,我们在冻结的主干之后引入了两个即插即用的多令牌预测(MTP)模块。这些模块允许模型在一个AR步骤中预测多个令牌。此外,为了减轻不准确的MTP造成的潜在错误传播,我们设计了一种新的验证算法,该算法利用冻结的主干来验证生成的令牌,从而使Llasa+在不牺牲生成质量的情况下实现加速。此外,我们设计了一个因果解码器,使流语音重建令牌。大量的实验表明,Llasa+在不牺牲生成质量的情况下实现了1.48倍的加速,尽管只在LibriTTS上训练。此外,MTP和验证框架可以应用于加速任何基于LLM的模型。所有代码和模型都可以在www.example.com上公开获得。
摘要:Recent progress in text-to-speech (TTS) has achieved impressive naturalness and flexibility, especially with the development of large language model (LLM)-based approaches. However, existing autoregressive (AR) structures and large-scale models, such as Llasa, still face significant challenges in inference latency and streaming synthesis. To deal with the limitations, we introduce Llasa+, an accelerated and streaming TTS model built on Llasa. Specifically, to accelerate the generation process, we introduce two plug-and-play Multi-Token Prediction (MTP) modules following the frozen backbone. These modules allow the model to predict multiple tokens in one AR step. Additionally, to mitigate potential error propagation caused by inaccurate MTP, we design a novel verification algorithm that leverages the frozen backbone to validate the generated tokens, thus allowing Llasa+ to achieve speedup without sacrificing generation quality. Furthermore, we design a causal decoder that enables streaming speech reconstruction from tokens. Extensive experiments show that Llasa+ achieves a 1.48X speedup without sacrificing generation quality, despite being trained only on LibriTTS. Moreover, the MTP-and-verification framework can be applied to accelerate any LLM-based model. All codes and models are publicly available at https://github.com/ASLP-lab/LLaSA_Plus.


【7】MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
标题:MeanAudio:通过Mean Flow快速可靠地生成文本到音频
链接:https://arxiv.org/abs/2508.06098

作者:Xiquan Li, Junxi Liu, Yuzhe Liang, Zhikang Niu, Wenxi Chen, Xie Chen
备注:9 pages, 3 figures
摘要:基于扩散和流的模型的最新发展显著地推进了文本到音频生成(TTA)。在实现高综合质量和可控性的同时,目前的TTA系统仍然存在推理速度慢的问题,这大大限制了它们的实用性。本文介绍了MeanAudio,这是一种新型的基于MeanFlow的模型,专为快速、忠实的文本到音频生成而设计。MeanAudio基于Flux风格的潜在Transformer构建,在训练期间回归平均速度场,通过直接从流轨迹的起点映射到终点来实现快速生成。通过将无分类器指导(CFG)纳入训练目标,MeanAudio在指导采样过程中不会产生额外的成本。为了进一步稳定训练,我们提出了一个带有流场混合的瞬时到平均值课程,鼓励模型首先学习基本的瞬时动态,然后逐渐适应平均流。该策略被证明是提高培训效率和生成质量的关键。实验结果表明,MeanAudio在单步音频生成方面达到了最先进的性能。具体而言,它在单台NVIDIA RTX 3090上实现了0.013的实时系数(RTF),比基于SOTA扩散的TTA系统加速100倍。此外,MeanAudio还在多步生成中表现出强大的性能,能够在连续的合成步骤中实现平滑和连贯的过渡。
摘要:Recent developments in diffusion- and flow- based models have significantly advanced Text-to-Audio Generation (TTA). While achieving great synthesis quality and controllability, current TTA systems still suffer from slow inference speed, which significantly limits their practical applicability. This paper presents MeanAudio, a novel MeanFlow-based model tailored for fast and faithful text-to-audio generation. Built on a Flux-style latent transformer, MeanAudio regresses the average velocity field during training, enabling fast generation by mapping directly from the start to the endpoint of the flow trajectory. By incorporating classifier-free guidance (CFG) into the training target, MeanAudio incurs no additional cost in the guided sampling process. To further stabilize training, we propose an instantaneous-to-mean curriculum with flow field mix-up, which encourages the model to first learn the foundational instantaneous dynamics, and then gradually adapt to mean flows. This strategy proves critical for enhancing training efficiency and generation quality. Experimental results demonstrate that MeanAudio achieves state-of-the-art performance in single-step audio generation. Specifically, it achieves a real time factor (RTF) of 0.013 on a single NVIDIA RTX 3090, yielding a 100x speedup over SOTA diffusion-based TTA systems. Moreover, MeanAudio also demonstrates strong performance in multi-step generation, enabling smooth and coherent transitions across successive synthesis steps.


【8】DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
标题:DAFMSVC:具有双重注意力机制和流量匹配的一次歌唱声音转换
链接:https://arxiv.org/abs/2508.05978

作者:Wei Chen, Binzhu Sha, Dan Luo, Jing Yang, Zhuo Wang, Fan Fan, Zhiyong Wu
备注:Accepted by INTERSPEECH 2025
摘要:歌唱声音转换(SVC)将源歌手的音色转换为目标,同时保留旋律和歌词。任何对任何SVC的关键挑战是在不降低质量的情况下使看不见的扬声器音色适应源音频。现有方法要么面临音色泄漏,要么无法在生成的音频中实现令人满意的音色相似性和质量。为了解决这些挑战,我们提出了DAFMSVC,其中来自源音频的自监督学习(SSL)特征被替换为来自目标音频的最相似的SSL特征,以防止音色泄漏。它还采用了双重交叉注意机制,用于自适应融合说话人嵌入,旋律和语言内容。此外,我们引入了一个流匹配模块,用于从融合的特征生成高质量的音频。实验结果表明,DAFMSVC显著提高了音色的相似性和自然度,在主观和客观评价方面都优于现有的方法。
摘要:Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing methods either face timbre leakage or fail to achieve satisfactory timbre similarity and quality in the generated audio. To address these challenges, we propose DAFMSVC, where the self-supervised learning (SSL) features from the source audio are replaced with the most similar SSL features from the target audio to prevent timbre leakage. It also incorporates a dual cross-attention mechanism for the adaptive fusion of speaker embeddings, melody, and linguistic content. Additionally, we introduce a flow matching module for high quality audio generation from the fused features. Experimental results show that DAFMSVC significantly enhances timbre similarity and naturalness, outperforming state-of-the-art methods in both subjective and objective evaluations.


【9】Training chord recognition models on artificially generated audio
标题:在人工生成的音频上训练和弦识别模型
链接:https://arxiv.org/abs/2508.05878

作者:Martyna Majchrzak, Jacek Mańdziuk
摘要:音乐信息检索中的一个挑战性问题是获得足够的非版权音频记录用于模型训练和评估。本研究比较了两种基于transformer的神经网络模型,用于音频记录中的和弦序列识别,并研究了使用人工生成的数据集进行此目的的有效性。这些模型在人工音频多轨(AAM),Schubert的Winterreise数据集和McGill Billboard数据集的各种组合上进行训练,并使用三个指标进行评估:Root,MajMin和Chord Content Metric(CCM)。实验证明,尽管人工生成的音乐和人类创作的音乐在复杂性和结构上确实存在差异,但前者在某些情况下是有用的。具体来说,AAM可以丰富由人类创作的音乐的较小训练数据集,或者甚至可以用作预测流行音乐中和弦序列的模型的独立训练集,如果没有其他数据可用的话。
摘要:One of the challenging problems in Music Information Retrieval is the acquisition of enough non-copyrighted audio recordings for model training and evaluation. This study compares two Transformer-based neural network models for chord sequence recognition in audio recordings and examines the effectiveness of using an artificially generated dataset for this purpose. The models are trained on various combinations of Artificial Audio Multitracks (AAM), Schubert's Winterreise Dataset, and the McGill Billboard Dataset and evaluated with three metrics: Root, MajMin and Chord Content Metric (CCM). The experiments prove that even though there are certainly differences in complexity and structure between artificially generated and human-composed music, the former can be useful in certain scenarios. Specifically, AAM can enrich a smaller training dataset of music composed by a human or can even be used as a standalone training set for a model that predicts chord sequences in pop music, if no other data is available.


【10】NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
标题:NanoCodec:迈向高质量超快速语音LLM推理
链接:https://arxiv.org/abs/2508.05835

作者:Edresson Casanova, Paarth Neekhara, Ryan Langman, Shehzeen Hussain, Subhankar Ghosh, Xuesong Yang, Ante Jukić, Jason Li, Boris Ginsburg
备注:Accepted to Interspeech 2025
摘要:大型语言模型(LLM)通过利用音频编解码器将音频离散化为令牌,从而使语言建模技术能够应用于语音数据,从而大大提高了音频处理。然而,现有的音频编解码器通常以高帧速率操作,导致缓慢的训练和推断,特别是对于自回归模型。为了解决这个问题,人们对低帧率音频编解码器越来越感兴趣,这减少了生成一秒音频所需的自回归步骤的数量。在本文中,我们进行消融研究,以检查帧速率,比特率和因果关系对编解码器重建质量的影响。基于我们的发现,我们推出了NanoCodec,这是一种最先进的音频编解码器,可实现仅12.5帧每秒(FPS)的高质量压缩。NanoCodec在各种比特率范围内都优于相关工作,为低延迟和高效的语音LLM训练和推理建立了新的基准。
摘要:Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling techniques to speech data. However, existing audio codecs often operate at high frame rates, leading to slow training and inference, particularly for autoregressive models. To address this, there is growing interest in low frame-rate audio codecs, which reduce the number of autoregressive steps required to generate one second of audio. In this paper, we conduct ablation studies to examine the impact of frame rate, bitrate, and causality on codec reconstruction quality. Based on our findings, we introduce NanoCodec, a state-of-the-art audio codec that achieves high-quality compression at just 12.5 frames per second (FPS). NanoCodec outperforms related works across various bitrate ranges, establishing a new benchmark for low-latency and efficient Speech LLM training and inference.


eess.AS音频处理


【1】Acoustic Non-Stationarity Objective Assessment with Hard Label Criteria for Supervised Learning Models
标题:具有监督学习模型硬标签标准的声学非平稳性客观评估
链接:https://arxiv.org/abs/2508.06405

作者:Guilherme Zucatelli, Ricardo Barioni, Gabriela Dantas
备注:Manuscript under review
摘要:客观的非平稳性措施是资源密集型的,并施加实时处理解决方案的关键限制。在本文中,提出了一种新的硬标签准则(HLC)算法生成一个全球非平稳标签的声学信号,使监督学习策略被训练为平稳估计。HLC首先评估国家的最先进的通用声学模型,证明这些模型编码平稳性信息。在此基础上,提出了基于HLC的声学非平稳性评估网络(NANSA)。NANSA模型优于竞争的方法,达到高达99%的分类准确率,同时解决了传统客观措施的计算不可行性。
摘要:Objective non-stationarity measures are resource intensive and impose critical limitations for real-time processing solutions. In this paper, a novel Hard Label Criteria (HLC) algorithm is proposed to generate a global non-stationarity label for acoustic signals, enabling supervised learning strategies to be trained as stationarity estimators. The HLC is first evaluated on state-of-the-art general-purpose acoustic models, demonstrating that these models encode stationarity information. Furthermore, the first-of-its-kind HLC-based Network for Acoustic Non-Stationarity Assessment (NANSA) is proposed. NANSA models outperform competing approaches, achieving up to 99\% classification accuracy, while solving the computational infeasibility of traditional objective measures.


【2】Use Cases for Voice Anonymization
标题:语音匿名化的用例
链接:https://arxiv.org/abs/2508.06356

作者:Sarina Meyer, Ngoc Thang Vu
备注:Accepted at SPSC 2025 - 5th Symposium on Security and Privacy in Speech Communication
摘要:语音匿名化系统的性能通常根据其隐藏说话者身份并保持数据对下游任务的效用的能力来衡量。这意味着匿名化应该满足的要求取决于使用它的上下文,并且在用例之间可能会有很大的不同。然而,这些用例很少在研究论文中详细说明。在本文中,我们研究了特定于用例的需求对语音匿名化方法设计的影响。我们进行了广泛的文献分析和用户研究,以收集可能的用例,并了解公众对这些工具的期望。基于这些研究,我们提出了语音匿名化用例的第一个分类,并推导出一组用于方法开发和评估的要求和设计标准。使用此方案,我们建议更多地关注面向用例的语音匿名系统的研究和开发。
摘要:The performance of a voice anonymization system is typically measured according to its ability to hide the speaker's identity and keep the data's utility for downstream tasks. This means that the requirements the anonymization should fulfill depend on the context in which it is used and may differ greatly between use cases. However, these use cases are rarely specified in research papers. In this paper, we study the implications of use case-specific requirements on the design of voice anonymization methods. We perform an extensive literature analysis and user study to collect possible use cases and to understand the expectations of the general public towards such tools. Based on these studies, we propose the first taxonomy of use cases for voice anonymization, and derive a set of requirements and design criteria for method development and evaluation. Using this scheme, we propose to focus more on use case-oriented research and development of voice anonymization systems.


【3】Egonoise Resilient Source Localization and Speech Enhancement for Drones Using a Hybrid Model and Learning-Based Approach
标题:使用混合模型和基于学习的方法的无人机弹性源定位和语音增强
链接:https://arxiv.org/abs/2508.06310

作者:Yihsuan Wu, Yukai Chiu, Michael Anthony, Mingsian R. Bai
摘要:无人机在搜索和救援任务,甚至军事行动中变得越来越重要。虽然大多数无人机都配备了摄像头视觉功能,但由于减轻转子产生的噪音的固有挑战,无人机试听领域仍然没有得到充分探索。在本文中,我们提出了一种新的技术来解决这个极低的信噪比(SNR)的麦克风嵌入式无人机遇到的问题。该技术是使用一种混合方法实现的,该方法结合了阵列信号处理(ASP)和深度神经网络(DNN),以增强安装在四轴飞行器上的六麦克风均匀圆形阵列捕获的语音信号。该系统通过波束控制结合通过广义旁瓣消除器-DeepFilterNet 2(GSC-DF 2)系统的语音增强来执行目标说话者的定位。为了验证该系统,DREGON数据集和实测数据。客观的评估表明,所提出的混合方法优于四个基线方法的SNR低至-30 dB的条件下的性能。
摘要:Drones are becoming increasingly important in search and rescue missions, and even military operations. While the majority of drones are equipped with camera vision capabilities, the realm of drone audition remains underexplored due to the inherent challenge of mitigating the egonoise generated by the rotors. In this paper, we present a novel technique to address this extremely low signal-to-noise ratio (SNR) problem encountered by the microphone-embedded drones. The technique is implemented using a hybrid approach that combines Array Signal Processing (ASP) and Deep Neural Networks (DNN) to enhance the speech signals captured by a six-microphone uniform circular array mounted on a quadcopter. The system performs localization of the target speaker through beamsteering in conjunction with speech enhancement through a Generalized Sidelobe Canceller-DeepFilterNet 2 (GSC-DF2) system. To validate the system, the DREGON dataset and measured data are employed. Objective evaluations of the proposed hybrid approach demonstrated its superior performance over four baseline methods in the SNR condition as low as -30 dB.


【4】Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
标题:利用LLM进行可扩展的非侵入性语音质量评估
链接:https://arxiv.org/abs/2508.06284

作者:Fredrik Cumlin, Xinyu Liang, Anubhab Ghosh, Saikat Chatterjee
备注:ECAI workshop paper
摘要:非侵入式语音质量评估(SQA)系统受到有限的训练数据和昂贵的人工注释的影响,阻碍了它们对实时会议呼叫的推广。在这项工作中,我们建议利用大型语言模型(LLM)作为语音质量的伪评分器来解决这些数据瓶颈。我们构建了LibriAugmented,这是一个由101,129个语音片段组成的数据集,这些语音片段具有由微调的听觉LLM(Vicuna-7 b-v1.5)标记的模拟退化。我们比较了三种训练策略:使用人类标记的数据,使用LLM标记的数据,以及使用DNSMOS Pro和DeePMOS的两阶段方法(对LLM标签进行预训练,然后对人类标签进行微调)。我们在几个跨语言和质量降级的数据集上进行测试。虽然与人类标记的训练相比,LLM标记的训练产生了混合的结果,但我们提供了经验证据,证明两阶段方法提高了泛化性能(例如,DNSMOS Pro在NISQA_TEST_LIVETALK上实现了0.63 vs. 0.55 PCC,在腾讯上实现了0.73 vs. 0.65 PCC(带混响)。我们的研究结果表明,使用LLM作为语音质量评估的可扩展伪评分器的潜力,提供了一个具有成本效益的解决方案的数据限制问题。
摘要:Non-intrusive speech quality assessment (SQA) systems suffer from limited training data and costly human annotations, hindering their generalization to real-time conferencing calls. In this work, we propose leveraging large language models (LLMs) as pseudo-raters for speech quality to address these data bottlenecks. We construct LibriAugmented, a dataset consisting of 101,129 speech clips with simulated degradations labeled by a fine-tuned auditory LLM (Vicuna-7b-v1.5). We compare three training strategies: using human-labeled data, using LLM-labeled data, and a two-stage approach (pretraining on LLM labels, then fine-tuning on human labels), using both DNSMOS Pro and DeePMOS. We test on several datasets across languages and quality degradations. While LLM-labeled training yields mixed results compared to human-labeled training, we provide empirical evidence that the two-stage approach improves the generalization performance (e.g., DNSMOS Pro achieves 0.63 vs. 0.55 PCC on NISQA_TEST_LIVETALK and 0.73 vs. 0.65 PCC on Tencent with reverb). Our findings demonstrate the potential of using LLMs as scalable pseudo-raters for speech quality assessment, offering a cost-effective solution to the data limitation problem.


【5】EchoFree: Towards Ultra Lightweight and Efficient Neural Acoustic Echo Cancellation
标题:EchoFree:迈向超轻型和高效的神经声学回声消除
链接:https://arxiv.org/abs/2508.06271

作者:Xingchen Li, Boyi Kang, Ziqian Wang, Zihan Zhang, Mingshuai Liu, Zhonghua Fu, Lei Xie
摘要:近年来,神经网络在声学回声消除中得到了广泛的应用。然而,现有的方法难以在保持性能的同时满足现实世界的低延迟和计算要求。为了应对这一挑战,我们提出了EchoFree,这是一个超轻量级的神经AEC框架,它将线性过滤与神经后过滤器相结合。具体来说,我们设计了一个神经后滤波器上的巴克尺度光谱功能。此外,我们引入了一个两阶段的优化策略,利用自监督学习(SSL)模型,以提高模型的性能。我们在ICASSP 2023 AEC挑战赛的盲测试集上评估了我们的方法。结果表明,我们的模型,只有278 K的参数和30 MMAC的计算复杂度,优于现有的低复杂度AEC模型,并实现了与最先进的轻量级模型DeepVQE-S的性能相当。音频示例可用。
摘要:In recent years, neural networks (NNs) have been widely applied in acoustic echo cancellation (AEC). However, existing approaches struggle to meet real-world low-latency and computational requirements while maintaining performance. To address this challenge, we propose EchoFree, an ultra lightweight neural AEC framework that combines linear filtering with a neural post filter. Specifically, we design a neural post-filter operating on Bark-scale spectral features. Furthermore, we introduce a two-stage optimization strategy utilizing self-supervised learning (SSL) models to improve model performance. We evaluate our method on the blind test set of the ICASSP 2023 AEC Challenge. The results demonstrate that our model, with only 278K parameters and 30 MMACs computational complexity, outperforms existing low-complexity AEC models and achieves performance comparable to that of state-of-the-art lightweight model DeepVQE-S. The audio examples are available.


【6】NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
标题:NanoCodec:迈向高质量超快速语音LLM推理
链接:https://arxiv.org/abs/2508.05835

作者:Edresson Casanova, Paarth Neekhara, Ryan Langman, Shehzeen Hussain, Subhankar Ghosh, Xuesong Yang, Ante Jukić, Jason Li, Boris Ginsburg
备注:Accepted to Interspeech 2025
摘要:大型语言模型(LLM)通过利用音频编解码器将音频离散化为令牌,从而使语言建模技术能够应用于语音数据,从而大大提高了音频处理。然而,现有的音频编解码器通常以高帧速率操作,导致缓慢的训练和推断,特别是对于自回归模型。为了解决这个问题,人们对低帧率音频编解码器越来越感兴趣,这减少了生成一秒音频所需的自回归步骤的数量。在本文中,我们进行消融研究,以检查帧速率,比特率和因果关系对编解码器重建质量的影响。基于我们的发现,我们推出了NanoCodec,这是一种最先进的音频编解码器,可实现仅12.5帧每秒(FPS)的高质量压缩。NanoCodec在各种比特率范围内都优于相关工作,为低延迟和高效的语音LLM训练和推理建立了新的基准。
摘要:Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling techniques to speech data. However, existing audio codecs often operate at high frame rates, leading to slow training and inference, particularly for autoregressive models. To address this, there is growing interest in low frame-rate audio codecs, which reduce the number of autoregressive steps required to generate one second of audio. In this paper, we conduct ablation studies to examine the impact of frame rate, bitrate, and causality on codec reconstruction quality. Based on our findings, we introduce NanoCodec, a state-of-the-art audio codec that achieves high-quality compression at just 12.5 frames per second (FPS). NanoCodec outperforms related works across various bitrate ranges, establishing a new benchmark for low-latency and efficient Speech LLM training and inference.


【7】Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
标题:Llasa+:免费午餐,用于加速和流媒体基于大羊驼的语音合成
链接:https://arxiv.org/abs/2508.06262

作者:Wenjie Tian, Xinfa Zhu, Hanke Xie, Zhen Ye, Wei Xue, Lei Xie
摘要:文本到语音(TTS)的最新进展已经取得了令人印象深刻的自然性和灵活性,特别是随着大语言模型(LLM)为基础的方法的发展。然而,现有的自回归(AR)结构和大规模模型,如Llasa,仍然面临着推理延迟和流合成的重大挑战。为了克服这些局限性,我们引入了Llasa+,一个基于Llasa的加速流TTS模型。具体来说,为了加速生成过程,我们在冻结的主干之后引入了两个即插即用的多令牌预测(MTP)模块。这些模块允许模型在一个AR步骤中预测多个令牌。此外,为了减轻不准确的MTP造成的潜在错误传播,我们设计了一种新的验证算法,该算法利用冻结的主干来验证生成的令牌,从而使Llasa+在不牺牲生成质量的情况下实现加速。此外,我们设计了一个因果解码器,使流语音重建令牌。大量的实验表明,Llasa+在不牺牲生成质量的情况下实现了1.48倍的加速,尽管只在LibriTTS上训练。此外,MTP和验证框架可以应用于加速任何基于LLM的模型。所有代码和模型都可以在https://github.com/ASLP-lab/LLaSA_Plus上公开获得。
摘要:Recent progress in text-to-speech (TTS) has achieved impressive naturalness and flexibility, especially with the development of large language model (LLM)-based approaches. However, existing autoregressive (AR) structures and large-scale models, such as Llasa, still face significant challenges in inference latency and streaming synthesis. To deal with the limitations, we introduce Llasa+, an accelerated and streaming TTS model built on Llasa. Specifically, to accelerate the generation process, we introduce two plug-and-play Multi-Token Prediction (MTP) modules following the frozen backbone. These modules allow the model to predict multiple tokens in one AR step. Additionally, to mitigate potential error propagation caused by inaccurate MTP, we design a novel verification algorithm that leverages the frozen backbone to validate the generated tokens, thus allowing Llasa+ to achieve speedup without sacrificing generation quality. Furthermore, we design a causal decoder that enables streaming speech reconstruction from tokens. Extensive experiments show that Llasa+ achieves a 1.48X speedup without sacrificing generation quality, despite being trained only on LibriTTS. Moreover, the MTP-and-verification framework can be applied to accelerate any LLM-based model. All codes and models are publicly available at https://github.com/ASLP-lab/LLaSA_Plus.


机器翻译由腾讯交互翻译提供,仅供参考