微信公众号:arXiv_Daily
cs.SD语音
标题:一种使用亚奈奎斯特采样和带宽扩展来节省可听音频功耗的实用方法
链接:https://arxiv.org/abs/2506.22321
摘要:Hearables是戴在耳朵上的可穿戴计算机。骨导麦克风(BCM)与气导麦克风(ACM)一起用于助听器中,作为噪声条件下多模态语音增强(SE)的支持模态。然而,现有的工作没有考虑以下实际方面的低功耗实现的听觉:(i)他们没有探讨如何降低采样频率和比特分辨率的模拟数字转换器(ADC)的听觉共同影响低功耗处理和多模态SE的语音质量和清晰度。(ii)他们没有讨论如何在不使用实际GAN鉴别器的情况下实现类似GAN的音频质量。以及(iii)它们不以亚奈奎斯特采样率处理来自ACM/BCM的信号,因为在它们的框架中,它们缺乏从它们的窄带部分的宽带重构方法。我们建议SUBARU(\textbf{Sub}-Nyquist \textbf{A}udio \textbf{R}esolution \textbf{U}psampling),实现了以下目标:SUBARU(i)有意在ADC中使用亚奈奎斯特采样和低位分辨率,实现了3.31倍的功耗降低;(ii)引入了新颖的多尺度和多周期虚拟鉴别器,在不使用GANs的对抗训练的情况下实现了类似GAN的音频质量;以及(iii)在移动平台和SE上在野外噪声条件下实现流操作,推理时间为1.74ms,内存占用小于13.77MB。
摘要:Hearables are wearable computers that are worn on the ear. Bone conduction microphones (BCMs) are used with air conduction microphones (ACMs) in hearables as a supporting modality for multimodal speech enhancement (SE) in noisy conditions. However, existing works don't consider the following practical aspects for low-power implementations on hearables: (i) They do not explore how lowering the sampling frequencies and bit resolutions in analog-to-digital converters (ADCs) of hearables jointly impact low-power processing and multimodal SE in terms of speech quality and intelligibility. (ii) They don't discuss how GAN-like audio quality can be achieved without using actual GAN discriminators. And (iii) They don't process signals from ACMs/BCMs at sub-Nyquist sampling rate because, in their frameworks, they lack a wideband reconstruction methodology from their narrowband parts. We propose SUBARU (\textbf{Sub}-Nyquist \textbf{A}udio \textbf{R}esolution \textbf{U}psampling), which achieves the following: SUBARU (i) intentionally uses sub-Nyquist sampling and low bit resolution in ADCs, achieving a 3.31x reduction in power consumption; (ii) introduces novel multi-scale and multi-period virtual discriminators, which achieve GAN-like audio quality without using GANs' adversarial training; and (iii) achieves streaming operations on mobile platforms and SE in in-the-wild noisy conditions with an inference time of 1.74ms and a memory footprint of less than 13.77MB.
【2】Reconstructing Intelligible Speech from the Pressure Sensor Data in HVACs
链接:https://arxiv.org/abs/2506.22311
摘要:压力传感器是现代供暖、通风和空调(HVAC)系统的集成组件。由于这些压力传感器在0-10 Pa范围内工作,支持0.5-2 kHz的高采样频率,并且通常放置在靠近人类的地方,因此它们可以用于窃听机密对话,因为人类语音具有类似的0-10 Pa可听范围和4 kHz带宽以获得可理解的质量。本文介绍了WaLi,它从低分辨率和噪声的压力传感器数据中重建可理解的语音,提供了以下技术贡献:(i)WaLi从压力传感器的最低0.5 kHz采样频率重建可理解的语音,而以前的工作只能检测热词/短语。WaLi使用复值构象和复杂全局注意力块(CGAB)来捕获存在于低分辨率压力传感器数据中的音素间和音素内依赖性。(ii)WaLi通过重建低频混叠分量的缺失频率的干净幅度和相位来处理从HVAC风扇和管道振动注入的瞬态噪声。对真实世界压力传感器的广泛测量研究表明,对于0.5 kHz至8 kHz的上采样,LSD为1.24,NISQA-MOS为1.78。我们认为,从隐私的角度来看,这种准确性水平构成了重大威胁,而这在压力传感器之前还没有得到解决。
摘要:Pressure sensors are an integrated component of modern Heating, Ventilation, and Air Conditioning (HVAC) systems. As these pressure sensors operate within the 0-10 Pa range, support high sampling frequencies of 0.5-2 kHz, and are often placed close to human proximity, they can be used to eavesdrop on confidential conversation, since human speech has a similar audible range of 0-10 Pa and a bandwidth of 4 kHz for intelligible quality. This paper presents WaLi, which reconstructs intelligible speech from the low-resolution and noisy pressure sensor data by providing the following technical contributions: (i) WaLi reconstructs intelligible speech from a minimum of 0.5 kHz sampling frequency of pressure sensors, whereas previous work can only detect hot words/phrases. WaLi uses complex-valued conformer and Complex Global Attention Block (CGAB) to capture inter-phoneme and intra-phoneme dependencies that exist in the low-resolution pressure sensor data. (ii) WaLi handles the transient noise injected from HVAC fans and duct vibrations, by reconstructing both the clean magnitude and phase of the missing frequencies of the low-frequency aliased components. Extensive measurement studies on real-world pressure sensors show an LSD of 1.24 and NISQA-MOS of 1.78 for 0.5 kHz to 8 kHz upsampling. We believe that such levels of accuracy pose a significant threat when viewed from a privacy perspective that has not been addressed before for pressure sensors.
【3】Fine-Tuning MIDI-to-Audio Alignment using a Neural Network on Piano Roll and CQT Representations
链接:https://arxiv.org/abs/2506.22237
备注:9 pages, 3 figures, 6 tables
摘要:在本文中,我们提出了一个神经网络的方法来同步人类钢琴演奏的音频记录与其相应的松散对齐的文件。该任务是使用卷积递归神经网络(CRNN)架构,有效地捕捉光谱和时间特征,通过处理未对齐的钢琴卷和声谱图作为输入,以估计对齐的钢琴卷。为了训练网络,我们创建了一个钢琴作品的数据集,其中包含模拟常见人类计时错误的增强型MIDI文件。该模型实现了高达20%的对准精度比行业标准的动态时间规整(DTW)方法在各种公差窗口。此外,将DTW与CRNN集成会产生额外的改进,提供增强的鲁棒性和一致性。这些发现证明了神经网络在推进最先进的MIDI到音频对齐方面的潜力。
摘要:In this paper, we present a neural network approach for synchronizing audio recordings of human piano performances with their corresponding loosely aligned MIDI files. The task is addressed using a Convolutional Recurrent Neural Network (CRNN) architecture, which effectively captures spectral and temporal features by processing an unaligned piano roll and a spectrogram as inputs to estimate the aligned piano roll. To train the network, we create a dataset of piano pieces with augmented MIDI files that simulate common human timing errors. The proposed model achieves up to 20% higher alignment accuracy than the industry-standard Dynamic Time Warping (DTW) method across various tolerance windows. Furthermore, integrating DTW with the CRNN yields additional improvements, offering enhanced robustness and consistency. These findings demonstrate the potential of neural networks in advancing state-of-the-art MIDI-to-audio alignment.
【4】SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition
链接:https://arxiv.org/abs/2506.22143
备注:Accepted for IEEE MLSP 2025
摘要:本文研究了不同语音SSL模型在阿拉伯方言(DA)和阿拉伯-英语码转换(CS)语音上的性能。为了解决数据稀缺的问题,引入了一种改进的音频拼接方法来生成人工CS语音数据。使用建议的拼接音频生成(SAGE)数据微调已经微调的SSL模型,在阿拉伯语和英语CS基准测试中,字错误率(WER)绝对提高了7.8%。此外,经验重放(ER)启发的方法,提出了提高DA和CS语音的概括,同时减轻灾难性的遗忘。集成域外3-gram语言模型将总体平均WER从31.7%降低到26.6%。针对代码切换基准的Few-Shot微调进一步将WER提高了4.9%。在阿拉伯语-英语CS基准测试中,WER为31.1%,超过了大规模多语言模型,包括USM和Whisper-large-v2(两者都大了10倍以上),分别为5.5%和8.4%。
摘要:This paper investigates the performance of various speech SSL models on dialectal Arabic (DA) and Arabic-English code-switched (CS) speech. To address data scarcity, a modified audio-splicing approach is introduced to generate artificial CS speech data. Fine-tuning an already fine-tuned SSL model with the proposed Spliced-Audio Generated (SAGE) data results in an absolute improvement on Word Error Rate (WER) of 7.8% on Arabic and English CS benchmarks. Additionally, an Experience Replay (ER) inspired approach is proposed to enhance generalisation across DA and CS speech while mitigating catastrophic forgetting. Integrating an out-of-domain 3-gram language model reduces the overall mean WER from 31.7% to 26.6%. Few-shot fine-tuning for code-switching benchmarks further improves WER by 4.9%. A WER of 31.1% on Arabic-English CS benchmarks surpasses large-scale multilingual models, including USM and Whisper-large-v2 (both over ten times larger) by an absolute margin of 5.5% and 8.4%, respectively.
【5】Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
链接:https://arxiv.org/abs/2506.22023
备注:17 pages, 8 figures, 5 tables
摘要:最近,自回归(AR)语言模型已经成为语音合成中的主导方法,提供表达生成和可扩展训练。然而,传统的AR语音合成模型依赖于下一个令牌预测范式往往遇到重大挑战时,处理长语音序列。这些模型通常难以构建稳定的帧到帧注意力,导致延迟增加和合成质量下降,从而限制了其实时应用的可行性。为了解决这些限制,我们引入了一种新的动态块自回归合成框架,称为DCAR,旨在提高AR语音生成的效率和可懂度鲁棒性。DCAR通过多标记预测训练引入了一种块到帧的注意机制,使用一个经过策略训练的轻量级模块在可变语音上下文中实现动态块预测。DCAR动态调整令牌预测跨度,显著降低序列长度依赖性,同时获得高合成质量。综合的实证评估表明,DCAR的性能大大优于传统的下一个令牌预测模型,在测试集上同时实现了高达72.27%的可懂度提升和2.61倍的推理加速。此外,我们进行全面的分析,以支持它作为下一代语音合成系统的通用基础。
摘要:Recently, autoregressive (AR) language models have emerged as a dominant approach in speech synthesis, offering expressive generation and scalable training. However, conventional AR speech synthesis models relying on the next-token prediction paradigm often encounter significant challenges when handling long speech sequences. These models often struggle to construct stable frame-to-frame attention, leading to increased latency and degraded synthesis quality, thereby limiting their feasibility for real-time applications. To address these limitations, we introduce a novel dynamic chunk-wise autoregressive synthesis framework, termed DCAR, designed to enhance both efficiency and intelligibility robustness in AR speech generation. DCAR introduces a chunk-to-frame attention mechanism through training with multi-token prediction, enabling dynamic chunk prediction in variable speech contexts using a lightweight module trained on-policy. DCAR dynamically adjusts the token prediction span, significantly reducing the sequence length dependency while obtaining high synthesis quality. Comprehensive empirical evaluations demonstrate that DCAR substantially outperforms traditional next-token prediction models, achieving up to 72.27% intelligibility improvement and 2.61x inference speedup simultaneously on the test set. Furthermore, we conduct comprehensive analysis to support it as a versatile foundation for next-generation speech synthesis systems.
【6】Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
链接:https://arxiv.org/abs/2506.21712
摘要:近年来,自监督语音Transformers的影响已经扩展到与说话者相关的应用。然而,很少有研究探讨这些模型如何编码说话人信息。在这项工作中,我们通过识别与扬声器信息相关的前馈层中的神经元来解决这一差距。具体来说,我们分析了与自监督特征和i向量的k均值聚类相关的神经元。我们的分析表明,这些集群对应于广泛的语音和性别类别,使它们适合识别代表说话者的神经元。通过在修剪过程中保护这些神经元,我们可以显着保留说话者相关任务的性能,证明它们在编码说话者信息中的关键作用。
摘要:In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speaker information. In this work, we address this gap by identifying neurons in the feed-forward layers that are correlated with speaker information. Specifically, we analyze neurons associated with k-means clusters of self-supervised features and i-vectors. Our analysis reveals that these clusters correspond to broad phonetic and gender classes, making them suitable for identifying neurons that represent speakers. By protecting these neurons during pruning, we can significantly preserve performance on speaker-related task, demonstrating their crucial role in encoding speaker information.
【7】Adapting Foundation Speech Recognition Models to Impaired Speech: A Semantic Re-chaining Approach for Personalization of German Speech
链接:https://arxiv.org/abs/2506.21622
摘要:由诸如脑瘫或遗传性疾病等疾病引起的语音障碍对自动语音识别(ASR)系统构成了重大挑战。尽管最近取得了进展,但由于训练数据有限以及收集和注释非规范语音样本的困难,像Whisper这样的ASR模型难以处理非规范语音。在这项工作中,我们提出了一个实用且轻量级的管道来个性化ASR模型,形式化单词的选择,并丰富了一个具有语义一致性的小型语音障碍数据集。应用到数据从一个孩子的结构性言语障碍,我们的方法显示出有希望的改善转录质量,表现出潜在的减少沟通障碍的个人与非典型的语音模式。
摘要:Speech impairments caused by conditions such as cerebral palsy or genetic disorders pose significant challenges for automatic speech recognition (ASR) systems. Despite recent advances, ASR models like Whisper struggle with non-normative speech due to limited training data and the difficulty of collecting and annotating non-normative speech samples. In this work, we propose a practical and lightweight pipeline to personalize ASR models, formalizing the selection of words and enriching a small, speech-impaired dataset with semantic coherence. Applied to data from a child with a structural speech impairment, our approach shows promising improvements in transcription quality, demonstrating the potential to reduce communication barriers for individuals with atypical speech patterns.
【8】IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
链接:https://arxiv.org/abs/2506.21619
摘要:大规模的文本到语音(TTS)模型通常分为自回归和非自回归系统。虽然自回归系统在语音自然度方面表现出一定的优势,但其逐个标记的生成机制使得难以精确地控制合成语音的持续时间。这是视频配音等需要严格视听同步的应用中的一个关键限制。本文介绍了IndexTTS 2,它提出了一种新的自回归模型友好的语音持续时间控制方法。该方法支持两种生成模式:一种允许显式指定生成的令牌的数量,以进行精确的持续时间控制;另一种不需要手动输入,让模型自由地生成语音,同时保留输入提示的韵律特征。此外,IndexTTS 2实现了情感表达和说话人身份之间的分离,使音色和情感能够独立控制。在zero-shot设置下,模型能够完美再现输入提示的情感特征。用户还可以提供单独的情感提示,甚至来自不同的说话者,允许模型在传达所需情感的同时重建目标音色。为了增强强烈情感表达的清晰度,我们将GPT潜在表征结合起来,以提高语音稳定性。同时,为了降低情感控制的障碍,我们通过对Qwen 3进行微调,设计了一种基于文本描述的软指令机制。这使得能够使用自然语言输入来有效地指导具有期望情感倾向的语音生成。实验结果表明,IndexTTS 2优于现有的国家的最先进的zero-shot TTS模型在单词错误率,说话人相似度和情感保真度。
摘要:Large-scale text-to-speech (TTS) models are typically categorized into autoregressive and non-autoregressive systems. Although autoregressive systems exhibit certain advantages in speech naturalness, their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This is a key limitation in applications such as video dubbing that require strict audio-visual synchronization. This paper introduces IndexTTS2, which proposes a novel and autoregressive-model-friendly method for speech duration control. The method supports two generation modes: one allows explicit specification of the number of generated tokens for precise duration control; the other does not require manual input and lets the model freely generate speech while preserving prosodic characteristics from the input prompt. Furthermore, IndexTTS2 achieves disentanglement between emotional expression and speaker identity, enabling independent control of timbre and emotion. In the zero-shot setting, the model can perfectly reproduce the emotional characteristics of the input prompt. Users may also provide a separate emotion prompt, even from a different speaker, allowing the model to reconstruct the target timbre while conveying the desired emotion. To enhance clarity during strong emotional expressions, we incorporate GPT latent representations to improve speech stability. Meanwhile, to lower the barrier for emotion control, we design a soft instruction mechanism based on textual descriptions by fine-tuning Qwen3. This enables effective guidance of speech generation with desired emotional tendencies using natural language input. Experimental results demonstrate that IndexTTS2 outperforms existing state-of-the-art zero-shot TTS models in word error rate, speaker similarity, and emotional fidelity.
【9】ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
链接:https://arxiv.org/abs/2506.21613
摘要:网上针对儿童的仇恨言论越来越普遍,突出表明迫切需要专门的数据集来解决这一关键问题。现有的仇恨言论数据集缺乏特定年龄的注释,无法捕捉细微差别的上下文,并且忽视了对儿童的独特情感影响。为了弥合这一差距,我们引入了ChildGuard1,这是一个来自现有语料库的策划数据集,并富含儿童特定的注释。ChildGuard收集了针对儿童的仇恨言论的各种背景,跨越了各个年龄组。我们对现有的最先进的仇恨言论检测方法进行了基准测试,包括大型语言模型(LLM),并评估了它们在检测和情境化针对儿童的仇恨言论方面的有效性。为了促进这一领域的进一步研究,我们公开发布了ChildGuard,为开发检测和减轻此类伤害的改进方法提供了坚实的基础。
摘要:The increasing prevalence of child-targeted hate speech online underscores the urgent need for specialized datasets to address this critical issue. Existing hate speech datasets lack agespecific annotations, fail to capture nuanced contexts, and overlook the unique emotional impact on children. To bridge this gap, we introduce ChildGuard1, a curated dataset derived from existing corpora and enriched with child-specific annotations. ChildGuard captures diverse contexts of child-targeted hate speech, spanning age groups. We benchmark existing state-of-the-art hate speech detection methods, including Large Language Models (LLMs), and assess their effectiveness in detecting and contextualizing child-targeted hate speech. To foster further research in this area, we publicly release ChildGuard, providing a robust foundation for developing improved methods to detect and mitigate such harm.
【10】Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASR
链接:https://arxiv.org/abs/2506.21577
备注:Accepted by Interspeech 2025
摘要:多语言自动语音识别(ASR)的最新进展是由像Whisper这样的大规模端到端模型驱动的。然而,语言干扰和扩展到看不见的语言(语言扩展)而不降低性能等挑战仍然存在。本文通过三个贡献来解决这些问题:1)整个软提示调整(整个SPT),其将软提示应用于编码器和解码器两者,从而增强特征提取和解码; 2)语音感知提示调整(LAPT),其利用跨语言相似性来使用轻量级提示矩阵对共享和语言特定的特征进行编码; 3)SPT-Whisper,一个将SPT集成到Whisper中并实现有效持续学习的工具包。FLEURS的三种语言的实验表明,Entire SPT和LAPT在语言扩展任务中分别比Decoder SPT高出5.0%和16.0%,为动态多语言ASR模型提供了一种有效的解决方案,并且计算开销最小。
摘要:Recent advancements in multilingual automatic speech recognition (ASR) have been driven by large-scale end-to-end models like Whisper. However, challenges such as language interference and expanding to unseen languages (language expansion) without degrading performance persist. This paper addresses these with three contributions: 1) Entire Soft Prompt Tuning (Entire SPT), which applies soft prompts to both the encoder and decoder, enhancing feature extraction and decoding; 2) Language-Aware Prompt Tuning (LAPT), which leverages cross-lingual similarities to encode shared and language-specific features using lightweight prompt matrices; 3) SPT-Whisper, a toolkit that integrates SPT into Whisper and enables efficient continual learning. Experiments across three languages from FLEURS demonstrate that Entire SPT and LAPT outperform Decoder SPT by 5.0% and 16.0% in language expansion tasks, respectively, providing an efficient solution for dynamic, multilingual ASR models with minimal computational overhead.
【11】Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
链接:https://arxiv.org/abs/2506.21576
备注:Accepted by Interspeech 2025
摘要:像Whisper这样的大规模多语言ASR模型在高资源环境中表现出色,但由于计算成本和灾难性遗忘,在低资源场景中面临挑战,例如稀有语言和代码切换(CS)。我们探索软提示调整(SPT),一个参数有效的方法,以提高CS ASR,同时保留先验知识。我们评估两种策略:(1)软提示和整个Whisper模型的完全微调(FFT),与传统方法相比,展示了改进的跨语言能力,以及(2)通过冻结模型参数和仅训练软提示来坚持SPT的原始设计。此外,我们还引入了SPT4ASR,这是一种不同SPT变体的组合。在SEAME和ASRU2019数据集上的实验表明,深度提示调优是最有效的SPT方法,我们的SPT4ASR方法在CS ASR中实现了进一步的错误减少,保持了与LoRA相似的参数效率,而不会降低现有语言的性能。
摘要:Large-scale multilingual ASR models like Whisper excel in high-resource settings but face challenges in low-resource scenarios, such as rare languages and code-switching (CS), due to computational costs and catastrophic forgetting. We explore Soft Prompt Tuning (SPT), a parameter-efficient method to enhance CS ASR while preserving prior knowledge. We evaluate two strategies: (1) full fine-tuning (FFT) of both soft prompts and the entire Whisper model, demonstrating improved cross-lingual capabilities compared to traditional methods, and (2) adhering to SPT's original design by freezing model parameters and only training soft prompts. Additionally, we introduce SPT4ASR, a combination of different SPT variants. Experiments on the SEAME and ASRU2019 datasets show that deep prompt tuning is the most effective SPT approach, and our SPT4ASR methods achieve further error reductions in CS ASR, maintaining parameter efficiency similar to LoRA, without degrading performance on existing languages.
【12】Efficient Multilingual ASR Finetuning via LoRA Language Experts
链接:https://arxiv.org/abs/2506.21555
备注:Accepted in Interspeech 2025
摘要:深度学习的最新进展显着增强了多语言自动语音识别(ASR),这是由于高级模型架构和可用的大规模多语言数据集的发展。尽管如此,多语言ASR仍然遭受着多语言性的诅咒,因为不同的语言往往会相互干扰,这使得ASR模型难以有效地识别多种语言,同时在它们之间共享模型容量。本文提出了一个高效的微调框架定制的多语言ASR通过准备LoRA语言专家的基础上Whisper。通过LoRA专家融合或知识蒸馏,我们的方法在目标语言上实现了比标准微调方法更好的识别性能。实验结果表明,该模型在语言感知和语言无关的情况下分别获得了约10%和15%的相对性能增益.
摘要:Recent advancements in deep learning have significantly enhanced multilingual automatic speech recognition (ASR) due to the development of advanced model architectures and available large-scale multilingual datasets. Despite that, multilingual ASR still suffers from the curse of multilinguality in that different languages tend to interfere with each other, making it difficult for the ASR model to identify multiple languages effectively while sharing model capacity across them. This paper proposes an efficient finetuning framework for customized multilingual ASR via prepared LoRA language experts based on Whisper. Through LoRA expert fusion or knowledge distillation, our approach achieves better recognition performance on target languages than standard fine-tuning methods. Experimental results demonstrate that the proposed models yield approximately 10\% and 15\% relative performance gains in language-aware and language-agnostic scenarios, respectively.
【13】WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation
链接:https://arxiv.org/abs/2506.22001
备注:Accepted by Interspeech2025
摘要:目前的多通道语音增强系统主要采用单输出结构,在多输入多输出(MIMO)处理过程中如何保持信号的时空完整性是一个很大的挑战。为了解决这个问题,我们提出了一种新的神经网络,称为WTFormer,MIMO语音增强,利用小波变换和多维协作注意力的多分辨率特性,有效地捕捉全球分布的空间特征,同时使用Conformer进行时频建模。在MUSIC算法的基础上,提出了一种多任务丢失策略进行优化训练,最大限度地保护了空间信息。LibriSpeech数据集上的实验结果表明,WTFormer可以实现与高级系统相当的去噪性能,同时仅用0.98M参数就可以保留更多的空间信息。
摘要:Current multi-channel speech enhancement systems mainly adopt single-output architecture, which face significant challenges in preserving spatio-temporal signal integrity during multiple-input multiple-output (MIMO) processing. To address this limitation, we propose a novel neural network, termed WTFormer, for MIMO speech enhancement that leverages the multi-resolution characteristics of wavelet transform and multi-dimensional collaborative attention to effectively capture globally distributed spatial features, while using Conformer for time-frequency modeling. A multi task loss strategy accompanying MUSIC algorithm is further proposed for optimization training to protect spatial information to the greatest extent. Experimental results on the LibriSpeech dataset show that WTFormer can achieve comparable denoising performance to advanced systems while preserving more spatial information with only 0.98M parameters.
【14】Explainable anomaly detection for sound spectrograms using pooling statistics with quantile differences
链接:https://arxiv.org/abs/2506.21921
摘要:异常检测是识别与数据集中几乎所有其他样本不同的罕见(即异常或异常)样本的任务。由于异常样本的模式通常是未知的先验,这项任务是非常具有挑战性的。因此,异常检测介于半监督学习和无监督学习之间。声音数据中的异常检测,通常称为“ASD”(异常声音检测),是一个子领域,涉及识别声学记录中新的和未知的影响。它对于工业4.0中的各种应用非常重要。这里,振动或声学数据通常从用于预测性维护的标准传感器信号获得。示例包括机器状态监控或质量保证,以跟踪组件或产品的状态。然而,智能算法的使用仍然是一个有争议的话题。管理层通常以降低成本和自动化为目标,而质量和维护专家则强调需要人类专业知识和全面的解决方案。在这项工作中,我们提出了一个异常检测方法,专门设计的频谱图。该方法是基于统计评估和理论动机。此外,它具有内在的可解释性,使其特别适用于工业环境中的应用。因此,该算法对于其中黑盒算法是不需要的或不适合的应用是相关的。
摘要:Anomaly detection is the task of identifying rarely occurring (i.e. anormal or anomalous) samples that differ from almost all other samples in a dataset. As the patterns of anormal samples are usually not known a priori, this task is highly challenging. Consequently, anomaly detection lies between semi- and unsupervised learning. The detection of anomalies in sound data, often called 'ASD' (Anomalous Sound Detection), is a sub-field that deals with the identification of new and yet unknown effects in acoustic recordings. It is of great importance for various applications in Industry 4.0. Here, vibrational or acoustic data are typically obtained from standard sensor signals used for predictive maintenance. Examples cover machine condition monitoring or quality assurance to track the state of components or products. However, the use of intelligent algorithms remains a controversial topic. Management generally aims for cost-reduction and automation, while quality and maintenance experts emphasize the need for human expertise and comprehensible solutions. In this work, we present an anomaly detection approach specifically designed for spectrograms. The approach is based on statistical evaluations and is theoretically motivated. In addition, it features intrinsic explainability, making it particularly suitable for applications in industrial settings. Thus, this algorithm is of relevance for applications in which black-box algorithms are unwanted or unsuitable.
标题:DistSoundStream:通过扩散解码实现高效语音令牌化
链接:https://arxiv.org/abs/2506.22362
摘要:基于令牌的语言建模是语音生成的一种重要方法,其中令牌是通过量化来自自监督学习(SSL)模型的特征并从神经语音编解码器中提取代码(通常称为语义令牌和声学令牌)来获得的。这些令牌通常是自回归建模的,推理速度受到令牌速率的约束。在这项工作中,我们提出了DiffSoundStream,这是一种通过两种技术提高非流场景中语音标记化效率的解决方案:(1)根据语义标记调节神经编解码器,以最大限度地减少语义和声学标记之间的冗余,以及(2)利用潜在扩散模型从语义和粗层次声学标记合成高质量波形。实验表明,在每秒50个令牌的情况下,DiffSoundStream实现了与以两倍令牌速率运行的标准SoundStream模型相当的语音质量。此外,我们只使用四个扩散采样步骤就可以实现分步蒸馏,只有很小的质量损失。
摘要:Token-based language modeling is a prominent approach for speech generation, where tokens are obtained by quantizing features from self-supervised learning (SSL) models and extracting codes from neural speech codecs, generally referred to as semantic tokens and acoustic tokens. These tokens are often modeled autoregressively, with the inference speed being constrained by the token rate. In this work, we propose DiffSoundStream, a solution that improves the efficiency of speech tokenization in non-streaming scenarios through two techniques: (1) conditioning the neural codec on semantic tokens to minimize redundancy between semantic and acoustic tokens, and (2) leveraging latent diffusion models to synthesize high-quality waveforms from semantic and coarse-level acoustic tokens. Experiments show that at 50 tokens per second, DiffSoundStream achieves speech quality on par with a standard SoundStream model operating at twice the token rate. Additionally, we achieve step-size distillation using just four diffusion sampling steps with only a minor quality loss.
【2】Cross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech Recognition
链接:https://arxiv.org/abs/2506.22194
备注:Accepted at INTERSPEECH 2025
摘要:本文提出了一种新的施主数据选择方法,以提高低资源的自动语音识别(ASR)。虽然ASR在高资源语言中表现良好,但由于训练数据有限,其准确性在低资源环境中下降。一个常见的解决方案是利用多语言自监督学习(SSL)模型与捐助者语言。然而,现有的方法依赖于语言级的相似性,忽略了剪辑级的变化。为了解决这一限制,我们提出了剪辑式声学标记分布相似性(CATDS),一种细粒度的选择方法,识别声学相关的捐助剪辑,以更好地与目标语言对齐。与现有的剪辑级选择方法不同,我们的方法与SSL模型的表示一致,并提供了更具挑战性但有价值的样本。实验结果表明,CATDS优于传统的选择方法,甚至可以利用捐助者的语言以前被认为是有害的。
摘要:This paper presents a novel donor data selection method to enhance low-resource automatic speech recognition (ASR). While ASR performs well in high-resource languages, its accuracy declines in low-resource settings due to limited training data. A common solution is to leverage multilingual self-supervised learning (SSL) models with donor languages. However, existing methods rely on language-level similarity, overlooking clip-level variations. To address this limitation, we propose clip-wise acoustic token distribution similarity (CATDS), a fine-grained selection method that identifies acoustically relevant donor clips for better alignment with the target language. Unlike existing clip-level selection methods, our method aligns with the representation of SSL models and offers more challenging yet valuable samples. Experimental results show that CATDS outperforms traditional selection methods and can even utilize donor languages previously considered detrimental.
【3】WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation
链接:https://arxiv.org/abs/2506.22001
备注:Accepted by Interspeech2025
摘要:目前的多通道语音增强系统主要采用单输出结构,在多输入多输出(MIMO)处理过程中如何保持信号的时空完整性是一个很大的挑战。为了解决这个问题,我们提出了一种新的神经网络,称为WTFormer,MIMO语音增强,利用小波变换和多维协作注意力的多分辨率特性,有效地捕捉全球分布的空间特征,同时使用Conformer进行时频建模。在MUSIC算法的基础上,提出了一种多任务丢失策略进行优化训练,最大限度地保护了空间信息。LibriSpeech数据集上的实验结果表明,WTFormer可以实现与高级系统相当的去噪性能,同时仅用0.98M参数就可以保留更多的空间信息。
摘要:Current multi-channel speech enhancement systems mainly adopt single-output architecture, which face significant challenges in preserving spatio-temporal signal integrity during multiple-input multiple-output (MIMO) processing. To address this limitation, we propose a novel neural network, termed WTFormer, for MIMO speech enhancement that leverages the multi-resolution characteristics of wavelet transform and multi-dimensional collaborative attention to effectively capture globally distributed spatial features, while using Conformer for time-frequency modeling. A multi task loss strategy accompanying MUSIC algorithm is further proposed for optimization training to protect spatial information to the greatest extent. Experimental results on the LibriSpeech dataset show that WTFormer can achieve comparable denoising performance to advanced systems while preserving more spatial information with only 0.98M parameters.
【4】HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment
链接:https://arxiv.org/abs/2506.21951
备注:Under Review, 3 pages + 1 References
摘要:现代语音质量预测模型是在重新采样到特定采样率的音频数据上训练的。当在测试时面对更高速率的音频时,这些模型可能会产生有偏见的分数。我们引入HighRateMOS,第一个非侵入性的平均意见评分(MOS)模型,明确考虑采样率。HighRateMOS集成了三种模型变体,利用以下信息:(i)语音采样率的可学习嵌入,(ii)Wav2vec 2.0自监督嵌入,(iii)多尺度CNN频谱特征,以及(iv)MFCC特征。在AudioMOS 2025 Track3中,HighRateMOS在八项指标中的五项中排名第一。我们的实验证实,建模的采样率直接导致更强大的和采样率不可知的语音质量预测。
摘要:Modern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When faced with higher-rate audio at test time, these models can produce biased scores. We introduce HighRateMOS, the first non-intrusive mean opinion score (MOS) model that explicitly considers sampling rate. HighRateMOS ensembles three model variants that exploit the following information: (i) a learnable embedding of speech sampling rate, (ii) Wav2vec 2.0 self-supervised embeddings, (iii) multi-scale CNN spectral features, and (iv) MFCC features. In AudioMOS 2025 Track3, HighRateMOS ranked first in five out of eight metrics. Our experiments confirm that modeling the sampling rate directly leads to more robust and sampling-rate-agnostic speech quality predictions.
【5】A Practical Approach to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension
链接:https://arxiv.org/abs/2506.22321
摘要:Hearables是戴在耳朵上的可穿戴计算机。骨导麦克风(BCM)与气导麦克风(ACM)一起用于助听器中,作为噪声条件下多模态语音增强(SE)的支持模态。然而,现有的工作没有考虑以下实际方面的低功耗实现的听觉:(i)他们没有探讨如何降低采样频率和比特分辨率的模拟数字转换器(ADC)的听觉共同影响低功耗处理和多模态SE的语音质量和清晰度。(ii)他们没有讨论如何在不使用实际GAN鉴别器的情况下实现类似GAN的音频质量。以及(iii)它们不以亚奈奎斯特采样率处理来自ACM/BCM的信号,因为在它们的框架中,它们缺乏从它们的窄带部分的宽带重构方法。我们建议SUBARU(\textbf{Sub}-Nyquist \textbf{A}udio \textbf{R}esolution \textbf{U}psampling),实现了以下目标:SUBARU(i)有意在ADC中使用亚奈奎斯特采样和低位分辨率,实现了3.31倍的功耗降低;(ii)引入了新颖的多尺度和多周期虚拟鉴别器,在不使用GANs的对抗训练的情况下实现了类似GAN的音频质量;以及(iii)在移动平台和SE上在野外噪声条件下实现流操作,推理时间为1.74ms,内存占用小于13.77MB。
摘要:Hearables are wearable computers that are worn on the ear. Bone conduction microphones (BCMs) are used with air conduction microphones (ACMs) in hearables as a supporting modality for multimodal speech enhancement (SE) in noisy conditions. However, existing works don't consider the following practical aspects for low-power implementations on hearables: (i) They do not explore how lowering the sampling frequencies and bit resolutions in analog-to-digital converters (ADCs) of hearables jointly impact low-power processing and multimodal SE in terms of speech quality and intelligibility. (ii) They don't discuss how GAN-like audio quality can be achieved without using actual GAN discriminators. And (iii) They don't process signals from ACMs/BCMs at sub-Nyquist sampling rate because, in their frameworks, they lack a wideband reconstruction methodology from their narrowband parts. We propose SUBARU (\textbf{Sub}-Nyquist \textbf{A}udio \textbf{R}esolution \textbf{U}psampling), which achieves the following: SUBARU (i) intentionally uses sub-Nyquist sampling and low bit resolution in ADCs, achieving a 3.31x reduction in power consumption; (ii) introduces novel multi-scale and multi-period virtual discriminators, which achieve GAN-like audio quality without using GANs' adversarial training; and (iii) achieves streaming operations on mobile platforms and SE in in-the-wild noisy conditions with an inference time of 1.74ms and a memory footprint of less than 13.77MB.
【6】Reconstructing Intelligible Speech from the Pressure Sensor Data in HVACs
链接:https://arxiv.org/abs/2506.22311
摘要:压力传感器是现代供暖、通风和空调(HVAC)系统的集成组件。由于这些压力传感器在0-10 Pa范围内工作,支持0.5-2 kHz的高采样频率,并且通常放置在靠近人类的地方,因此它们可以用于窃听机密对话,因为人类语音具有类似的0-10 Pa可听范围和4 kHz带宽以获得可理解的质量。本文介绍了WaLi,它从低分辨率和噪声的压力传感器数据中重建可理解的语音,提供了以下技术贡献:(i)WaLi从压力传感器的最低0.5 kHz采样频率重建可理解的语音,而以前的工作只能检测热词/短语。WaLi使用复值构象和复杂全局注意力块(CGAB)来捕获存在于低分辨率压力传感器数据中的音素间和音素内依赖性。(ii)WaLi通过重建低频混叠分量的缺失频率的干净幅度和相位来处理从HVAC风扇和管道振动注入的瞬态噪声。对真实世界压力传感器的广泛测量研究表明,对于0.5 kHz至8 kHz的上采样,LSD为1.24,NISQA-MOS为1.78。我们认为,从隐私的角度来看,这种准确性水平构成了重大威胁,而这在压力传感器之前还没有得到解决。
摘要:Pressure sensors are an integrated component of modern Heating, Ventilation, and Air Conditioning (HVAC) systems. As these pressure sensors operate within the 0-10 Pa range, support high sampling frequencies of 0.5-2 kHz, and are often placed close to human proximity, they can be used to eavesdrop on confidential conversation, since human speech has a similar audible range of 0-10 Pa and a bandwidth of 4 kHz for intelligible quality. This paper presents WaLi, which reconstructs intelligible speech from the low-resolution and noisy pressure sensor data by providing the following technical contributions: (i) WaLi reconstructs intelligible speech from a minimum of 0.5 kHz sampling frequency of pressure sensors, whereas previous work can only detect hot words/phrases. WaLi uses complex-valued conformer and Complex Global Attention Block (CGAB) to capture inter-phoneme and intra-phoneme dependencies that exist in the low-resolution pressure sensor data. (ii) WaLi handles the transient noise injected from HVAC fans and duct vibrations, by reconstructing both the clean magnitude and phase of the missing frequencies of the low-frequency aliased components. Extensive measurement studies on real-world pressure sensors show an LSD of 1.24 and NISQA-MOS of 1.78 for 0.5 kHz to 8 kHz upsampling. We believe that such levels of accuracy pose a significant threat when viewed from a privacy perspective that has not been addressed before for pressure sensors.
【7】Fine-Tuning MIDI-to-Audio Alignment using a Neural Network on Piano Roll and CQT Representations
链接:https://arxiv.org/abs/2506.22237
备注:9 pages, 3 figures, 6 tables
摘要:在本文中,我们提出了一个神经网络的方法来同步人类钢琴演奏的音频记录与其相应的松散对齐的文件。该任务是使用卷积递归神经网络(CRNN)架构,有效地捕捉光谱和时间特征,通过处理未对齐的钢琴卷和声谱图作为输入,以估计对齐的钢琴卷。为了训练网络,我们创建了一个钢琴作品的数据集,其中包含模拟常见人类计时错误的增强型MIDI文件。该模型实现了高达20%的对准精度比行业标准的动态时间规整(DTW)方法在各种公差窗口。此外,将DTW与CRNN集成会产生额外的改进,提供增强的鲁棒性和一致性。这些发现证明了神经网络在推进最先进的MIDI到音频对齐方面的潜力。
摘要:In this paper, we present a neural network approach for synchronizing audio recordings of human piano performances with their corresponding loosely aligned MIDI files. The task is addressed using a Convolutional Recurrent Neural Network (CRNN) architecture, which effectively captures spectral and temporal features by processing an unaligned piano roll and a spectrogram as inputs to estimate the aligned piano roll. To train the network, we create a dataset of piano pieces with augmented MIDI files that simulate common human timing errors. The proposed model achieves up to 20% higher alignment accuracy than the industry-standard Dynamic Time Warping (DTW) method across various tolerance windows. Furthermore, integrating DTW with the CRNN yields additional improvements, offering enhanced robustness and consistency. These findings demonstrate the potential of neural networks in advancing state-of-the-art MIDI-to-audio alignment.
【8】SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition
链接:https://arxiv.org/abs/2506.22143
备注:Accepted for IEEE MLSP 2025
摘要:本文研究了不同语音SSL模型在阿拉伯方言(DA)和阿拉伯-英语码转换(CS)语音上的性能。为了解决数据稀缺的问题,引入了一种改进的音频拼接方法来生成人工CS语音数据。使用建议的拼接音频生成(SAGE)数据微调已经微调的SSL模型,在阿拉伯语和英语CS基准测试中,字错误率(WER)绝对提高了7.8%。此外,经验重放(ER)启发的方法,提出了提高DA和CS语音的概括,同时减轻灾难性的遗忘。集成域外3-gram语言模型将总体平均WER从31.7%降低到26.6%。针对代码切换基准的Few-Shot微调进一步将WER提高了4.9%。在阿拉伯语-英语CS基准测试中,WER为31.1%,超过了大规模多语言模型,包括USM和Whisper-large-v2(两者都大了10倍以上),分别为5.5%和8.4%。
摘要:This paper investigates the performance of various speech SSL models on dialectal Arabic (DA) and Arabic-English code-switched (CS) speech. To address data scarcity, a modified audio-splicing approach is introduced to generate artificial CS speech data. Fine-tuning an already fine-tuned SSL model with the proposed Spliced-Audio Generated (SAGE) data results in an absolute improvement on Word Error Rate (WER) of 7.8% on Arabic and English CS benchmarks. Additionally, an Experience Replay (ER) inspired approach is proposed to enhance generalisation across DA and CS speech while mitigating catastrophic forgetting. Integrating an out-of-domain 3-gram language model reduces the overall mean WER from 31.7% to 26.6%. Few-shot fine-tuning for code-switching benchmarks further improves WER by 4.9%. A WER of 31.1% on Arabic-English CS benchmarks surpasses large-scale multilingual models, including USM and Whisper-large-v2 (both over ten times larger) by an absolute margin of 5.5% and 8.4%, respectively.
【9】Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
链接:https://arxiv.org/abs/2506.22023
备注:17 pages, 8 figures, 5 tables
摘要:最近,自回归(AR)语言模型已成为语音合成中的主导方法,提供表达生成和可扩展训练。然而,传统的AR语音合成模型依赖于下一个令牌预测范式往往遇到重大挑战时,处理长语音序列。这些模型通常难以构建稳定的帧到帧注意力,导致延迟增加和合成质量下降,从而限制了其实时应用的可行性。为了解决这些限制,我们引入了一种新的动态块自回归合成框架,称为DCAR,旨在提高AR语音生成的效率和可懂度鲁棒性。DCAR通过多标记预测训练引入了一种块到帧的注意机制,使用一个经过策略训练的轻量级模块在可变语音上下文中实现动态块预测。DCAR动态调整令牌预测跨度,显著降低序列长度依赖性,同时获得高合成质量。综合的实证评估表明,DCAR的性能大大优于传统的下一个令牌预测模型,在测试集上同时实现了高达72.27%的可懂度提升和2.61倍的推理加速。此外,我们进行全面的分析,以支持它作为下一代语音合成系统的通用基础。
摘要:Recently, autoregressive (AR) language models have emerged as a dominant approach in speech synthesis, offering expressive generation and scalable training. However, conventional AR speech synthesis models relying on the next-token prediction paradigm often encounter significant challenges when handling long speech sequences. These models often struggle to construct stable frame-to-frame attention, leading to increased latency and degraded synthesis quality, thereby limiting their feasibility for real-time applications. To address these limitations, we introduce a novel dynamic chunk-wise autoregressive synthesis framework, termed DCAR, designed to enhance both efficiency and intelligibility robustness in AR speech generation. DCAR introduces a chunk-to-frame attention mechanism through training with multi-token prediction, enabling dynamic chunk prediction in variable speech contexts using a lightweight module trained on-policy. DCAR dynamically adjusts the token prediction span, significantly reducing the sequence length dependency while obtaining high synthesis quality. Comprehensive empirical evaluations demonstrate that DCAR substantially outperforms traditional next-token prediction models, achieving up to 72.27% intelligibility improvement and 2.61x inference speedup simultaneously on the test set. Furthermore, we conduct comprehensive analysis to support it as a versatile foundation for next-generation speech synthesis systems.
【10】Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
链接:https://arxiv.org/abs/2506.21990
备注:Computer Vision and Pattern Recognition (CVPR) 2025 Workshops
摘要:Transformer编码器-解码器架构的发展导致了机器翻译、自动语音识别(ASR)和基于语音识别的聊天机等应用的重大突破。预先训练的模型在几个时期(大多数情况下少于5个)内使用大量通用数据进行训练,从而具有强大的泛化能力。然而,当应用于利基领域时,这些模型的性能确实会受到影响,例如在驾驶舱中转录飞行员的语音,这涉及到大量特定的词汇和多语言对话。本文研究和提高驾驶舱会话的耳语模型的转录准确性。我们收集了大约85分钟的驾驶舱模拟器录音和130分钟的飞行员访谈录音,并手动标记它们。演讲者是中年男子,会说德语和英语。为了提高转录的准确性,我们提出了多个规范化方案,以改善成绩单和提高单词错误率(WER)。然后,我们采用微调来提高ASR性能,利用低秩自适应(LoRA)的性能高效的微调。因此,WER从68.49%(未经标准化基线的预训练耳语大模型)下降到26.26%(采用拟议标准化方案的微调耳语大模型)。
摘要:The developments in transformer encoder-decoder architectures have led to significant breakthroughs in machine translation, Automatic Speech Recognition (ASR), and instruction-based chat machines, among other applications. The pre-trained models were trained on vast amounts of generic data over a few epochs (fewer than five in most cases), resulting in their strong generalization capabilities. Nevertheless, the performance of these models does suffer when applied to niche domains like transcribing pilot speech in the cockpit, which involves a lot of specific vocabulary and multilingual conversations. This paper investigates and improves the transcription accuracy of cockpit conversations with Whisper models. We have collected around 85 minutes of cockpit simulator recordings and 130 minutes of interview recordings with pilots and manually labeled them. The speakers are middle aged men speaking both German and English. To improve the accuracy of transcriptions, we propose multiple normalization schemes to refine the transcripts and improve Word Error Rate (WER). We then employ fine-tuning to enhance ASR performance, utilizing performance-efficient fine-tuning with Low-Rank Adaptation (LoRA). Hereby, WER decreased from 68.49 \% (pretrained whisper Large model without normalization baseline) to 26.26\% (finetuned whisper Large model with the proposed normalization scheme).
【11】Explainable anomaly detection for sound spectrograms using pooling statistics with quantile differences
链接:https://arxiv.org/abs/2506.21921
摘要:异常检测是识别与数据集中几乎所有其他样本不同的罕见(即异常或异常)样本的任务。由于异常样本的模式通常是未知的先验,这项任务是非常具有挑战性的。因此,异常检测介于半监督学习和无监督学习之间。声音数据中的异常检测,通常称为“ASD”(异常声音检测),是一个子领域,涉及识别声学记录中新的和未知的影响。它对于工业4.0中的各种应用非常重要。这里,振动或声学数据通常从用于预测性维护的标准传感器信号获得。示例包括机器状态监控或质量保证,以跟踪组件或产品的状态。然而,智能算法的使用仍然是一个有争议的话题。管理层通常以降低成本和自动化为目标,而质量和维护专家则强调需要人类专业知识和全面的解决方案。在这项工作中,我们提出了一个异常检测方法,专门设计的频谱图。该方法是基于统计评估和理论动机。此外,它具有内在的可解释性,使其特别适用于工业环境中的应用。因此,该算法对于其中黑盒算法是不需要的或不适合的应用是相关的。
摘要:Anomaly detection is the task of identifying rarely occurring (i.e. anormal or anomalous) samples that differ from almost all other samples in a dataset. As the patterns of anormal samples are usually not known a priori, this task is highly challenging. Consequently, anomaly detection lies between semi- and unsupervised learning. The detection of anomalies in sound data, often called 'ASD' (Anomalous Sound Detection), is a sub-field that deals with the identification of new and yet unknown effects in acoustic recordings. It is of great importance for various applications in Industry 4.0. Here, vibrational or acoustic data are typically obtained from standard sensor signals used for predictive maintenance. Examples cover machine condition monitoring or quality assurance to track the state of components or products. However, the use of intelligent algorithms remains a controversial topic. Management generally aims for cost-reduction and automation, while quality and maintenance experts emphasize the need for human expertise and comprehensible solutions. In this work, we present an anomaly detection approach specifically designed for spectrograms. The approach is based on statistical evaluations and is theoretically motivated. In addition, it features intrinsic explainability, making it particularly suitable for applications in industrial settings. Thus, this algorithm is of relevance for applications in which black-box algorithms are unwanted or unsuitable.
【12】Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
链接:https://arxiv.org/abs/2506.21712
摘要:近年来,自监督语音Transformers的影响已经扩展到与说话者相关的应用。然而,很少有研究探讨这些模型如何编码说话人信息。在这项工作中,我们通过识别与扬声器信息相关的前馈层中的神经元来解决这一差距。具体来说,我们分析了与自监督特征和i向量的k均值聚类相关的神经元。我们的分析表明,这些集群对应于广泛的语音和性别类别,使它们适合识别代表说话者的神经元。通过在修剪过程中保护这些神经元,我们可以显着保留说话者相关任务的性能,证明它们在编码说话者信息中的关键作用。
摘要:In recent years, the impact of self-supervised speech Transformers has extended to speaker-related applications. However, little research has explored how these models encode speaker information. In this work, we address this gap by identifying neurons in the feed-forward layers that are correlated with speaker information. Specifically, we analyze neurons associated with k-means clusters of self-supervised features and i-vectors. Our analysis reveals that these clusters correspond to broad phonetic and gender classes, making them suitable for identifying neurons that represent speakers. By protecting these neurons during pruning, we can significantly preserve performance on speaker-related task, demonstrating their crucial role in encoding speaker information.
【13】Adapting Foundation Speech Recognition Models to Impaired Speech: A Semantic Re-chaining Approach for Personalization of German Speech
链接:https://arxiv.org/abs/2506.21622
摘要:由诸如脑瘫或遗传性疾病等疾病引起的语音障碍对自动语音识别(ASR)系统构成了重大挑战。尽管最近取得了进展,但由于训练数据有限以及收集和注释非规范语音样本的困难,像Whisper这样的ASR模型难以处理非规范语音。在这项工作中,我们提出了一个实用且轻量级的管道来个性化ASR模型,形式化单词的选择,并丰富了一个具有语义一致性的小型语音障碍数据集。应用到数据从一个孩子的结构性言语障碍,我们的方法显示出有希望的改善转录质量,表现出潜在的减少沟通障碍的个人与非典型的语音模式。
摘要:Speech impairments caused by conditions such as cerebral palsy or genetic disorders pose significant challenges for automatic speech recognition (ASR) systems. Despite recent advances, ASR models like Whisper struggle with non-normative speech due to limited training data and the difficulty of collecting and annotating non-normative speech samples. In this work, we propose a practical and lightweight pipeline to personalize ASR models, formalizing the selection of words and enriching a small, speech-impaired dataset with semantic coherence. Applied to data from a child with a structural speech impairment, our approach shows promising improvements in transcription quality, demonstrating the potential to reduce communication barriers for individuals with atypical speech patterns.
【14】IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
链接:https://arxiv.org/abs/2506.21619
摘要:大规模的文本到语音(TTS)模型通常分为自回归和非自回归系统。虽然自回归系统在语音自然度方面表现出一定的优势,但其逐个标记的生成机制使得难以精确地控制合成语音的持续时间。这是视频配音等需要严格视听同步的应用中的一个关键限制。本文介绍了IndexTTS 2,它提出了一种新的自回归模型友好的语音持续时间控制方法。该方法支持两种生成模式:一种允许显式指定生成的令牌的数量,以进行精确的持续时间控制;另一种不需要手动输入,让模型自由地生成语音,同时保留输入提示的韵律特征。此外,IndexTTS 2实现了情感表达和说话人身份之间的分离,使音色和情感能够独立控制。在zero-shot设置下,模型能够完美再现输入提示的情感特征。用户还可以提供单独的情感提示,甚至来自不同的说话者,允许模型在传达所需情感的同时重建目标音色。为了增强强烈情感表达的清晰度,我们将GPT潜在表征结合起来,以提高语音稳定性。同时,为了降低情感控制的障碍,我们通过对Qwen 3进行微调,设计了一种基于文本描述的软指令机制。这使得能够使用自然语言输入来有效地指导具有期望情感倾向的语音生成。实验结果表明,IndexTTS 2优于现有的国家的最先进的zero-shot TTS模型在单词错误率,说话人相似度和情感保真度。
摘要:Large-scale text-to-speech (TTS) models are typically categorized into autoregressive and non-autoregressive systems. Although autoregressive systems exhibit certain advantages in speech naturalness, their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This is a key limitation in applications such as video dubbing that require strict audio-visual synchronization. This paper introduces IndexTTS2, which proposes a novel and autoregressive-model-friendly method for speech duration control. The method supports two generation modes: one allows explicit specification of the number of generated tokens for precise duration control; the other does not require manual input and lets the model freely generate speech while preserving prosodic characteristics from the input prompt. Furthermore, IndexTTS2 achieves disentanglement between emotional expression and speaker identity, enabling independent control of timbre and emotion. In the zero-shot setting, the model can perfectly reproduce the emotional characteristics of the input prompt. Users may also provide a separate emotion prompt, even from a different speaker, allowing the model to reconstruct the target timbre while conveying the desired emotion. To enhance clarity during strong emotional expressions, we incorporate GPT latent representations to improve speech stability. Meanwhile, to lower the barrier for emotion control, we design a soft instruction mechanism based on textual descriptions by fine-tuning Qwen3. This enables effective guidance of speech generation with desired emotional tendencies using natural language input. Experimental results demonstrate that IndexTTS2 outperforms existing state-of-the-art zero-shot TTS models in word error rate, speaker similarity, and emotional fidelity.
【15】ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
链接:https://arxiv.org/abs/2506.21613
摘要:网上针对儿童的仇恨言论越来越普遍,突出表明迫切需要专门的数据集来解决这一关键问题。现有的仇恨言论数据集缺乏特定年龄的注释,无法捕捉细微差别的上下文,并且忽视了对儿童的独特情感影响。为了弥合这一差距,我们引入了ChildGuard1,这是一个来自现有语料库的策划数据集,并富含儿童特定的注释。ChildGuard收集了针对儿童的仇恨言论的各种背景,跨越了各个年龄组。我们对现有的最先进的仇恨言论检测方法进行了基准测试,包括大型语言模型(LLM),并评估了它们在检测和情境化针对儿童的仇恨言论方面的有效性。为了促进这一领域的进一步研究,我们公开发布了ChildGuard,为开发检测和减轻此类伤害的改进方法提供了坚实的基础。
摘要:The increasing prevalence of child-targeted hate speech online underscores the urgent need for specialized datasets to address this critical issue. Existing hate speech datasets lack agespecific annotations, fail to capture nuanced contexts, and overlook the unique emotional impact on children. To bridge this gap, we introduce ChildGuard1, a curated dataset derived from existing corpora and enriched with child-specific annotations. ChildGuard captures diverse contexts of child-targeted hate speech, spanning age groups. We benchmark existing state-of-the-art hate speech detection methods, including Large Language Models (LLMs), and assess their effectiveness in detecting and contextualizing child-targeted hate speech. To foster further research in this area, we publicly release ChildGuard, providing a robust foundation for developing improved methods to detect and mitigate such harm.
【16】Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASR
链接:https://arxiv.org/abs/2506.21577
备注:Accepted by Interspeech 2025
摘要:多语言自动语音识别(ASR)的最新进展是由像Whisper这样的大规模端到端模型驱动的。然而,语言干扰和扩展到看不见的语言(语言扩展)而不降低性能等挑战仍然存在。本文主要从三个方面来解决这些问题:1)整个软提示调整(整个SPT),其将软提示应用于编码器和解码器两者,从而增强特征提取和解码; 2)语音感知提示调整(LAPT),其利用跨语言相似性来使用轻量级提示矩阵对共享和语言特定的特征进行编码; 3)SPT-Whisper,一个将SPT集成到Whisper中并实现有效持续学习的工具包。FLEURS的三种语言的实验表明,Entire SPT和LAPT在语言扩展任务中分别比Decoder SPT高出5.0%和16.0%,为动态多语言ASR模型提供了一种有效的解决方案,并且计算开销最小。
摘要:Recent advancements in multilingual automatic speech recognition (ASR) have been driven by large-scale end-to-end models like Whisper. However, challenges such as language interference and expanding to unseen languages (language expansion) without degrading performance persist. This paper addresses these with three contributions: 1) Entire Soft Prompt Tuning (Entire SPT), which applies soft prompts to both the encoder and decoder, enhancing feature extraction and decoding; 2) Language-Aware Prompt Tuning (LAPT), which leverages cross-lingual similarities to encode shared and language-specific features using lightweight prompt matrices; 3) SPT-Whisper, a toolkit that integrates SPT into Whisper and enables efficient continual learning. Experiments across three languages from FLEURS demonstrate that Entire SPT and LAPT outperform Decoder SPT by 5.0% and 16.0% in language expansion tasks, respectively, providing an efficient solution for dynamic, multilingual ASR models with minimal computational overhead.
【17】Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
链接:https://arxiv.org/abs/2506.21576
备注:Accepted by Interspeech 2025
摘要:像Whisper这样的大规模多语言ASR模型在高资源环境中表现出色,但由于计算成本和灾难性遗忘,在低资源场景中面临挑战,例如稀有语言和代码切换(CS)。我们探索软提示调整(SPT),一个参数有效的方法,以提高CS ASR,同时保留先验知识。我们评估两种策略:(1)软提示和整个Whisper模型的完全微调(FFT),与传统方法相比,展示了改进的跨语言能力,以及(2)通过冻结模型参数和仅训练软提示来坚持SPT的原始设计。此外,我们还引入了SPT4ASR,这是一种不同SPT变体的组合。在SEAME和ASRU2019数据集上的实验表明,深度提示调优是最有效的SPT方法,我们的SPT4ASR方法在CS ASR中实现了进一步的错误减少,保持了与LoRA相似的参数效率,而不会降低现有语言的性能。
摘要:Large-scale multilingual ASR models like Whisper excel in high-resource settings but face challenges in low-resource scenarios, such as rare languages and code-switching (CS), due to computational costs and catastrophic forgetting. We explore Soft Prompt Tuning (SPT), a parameter-efficient method to enhance CS ASR while preserving prior knowledge. We evaluate two strategies: (1) full fine-tuning (FFT) of both soft prompts and the entire Whisper model, demonstrating improved cross-lingual capabilities compared to traditional methods, and (2) adhering to SPT's original design by freezing model parameters and only training soft prompts. Additionally, we introduce SPT4ASR, a combination of different SPT variants. Experiments on the SEAME and ASRU2019 datasets show that deep prompt tuning is the most effective SPT approach, and our SPT4ASR methods achieve further error reductions in CS ASR, maintaining parameter efficiency similar to LoRA, without degrading performance on existing languages.
【18】Efficient Multilingual ASR Finetuning via LoRA Language Experts
链接:https://arxiv.org/abs/2506.21555
备注:Accepted in Interspeech 2025
摘要:深度学习的最新进展显着增强了多语言自动语音识别(ASR),这是由于高级模型架构和可用的大规模多语言数据集的发展。尽管如此,多语言ASR仍然遭受着多语言性的诅咒,因为不同的语言往往会相互干扰,这使得ASR模型难以有效地识别多种语言,同时在它们之间共享模型容量。本文提出了一个高效的微调框架定制的多语言ASR通过准备LoRA语言专家的基础上Whisper。通过LoRA专家融合或知识蒸馏,我们的方法在目标语言上实现了比标准微调方法更好的识别性能。实验结果表明,该模型在语言感知和语言无关的情况下分别获得了约10%和15%的相对性能增益.
摘要:Recent advancements in deep learning have significantly enhanced multilingual automatic speech recognition (ASR) due to the development of advanced model architectures and available large-scale multilingual datasets. Despite that, multilingual ASR still suffers from the curse of multilinguality in that different languages tend to interfere with each other, making it difficult for the ASR model to identify multiple languages effectively while sharing model capacity across them. This paper proposes an efficient finetuning framework for customized multilingual ASR via prepared LoRA language experts based on Whisper. Through LoRA expert fusion or knowledge distillation, our approach achieves better recognition performance on target languages than standard fine-tuning methods. Experimental results demonstrate that the proposed models yield approximately 10\% and 15\% relative performance gains in language-aware and language-agnostic scenarios, respectively.
机器翻译由腾讯交互翻译提供,仅供参考
