微信公众号:arXiv_Daily
cs.SD语音
标题:ASDA:用于自我监督表示学习的音频频谱图差异注意机制
链接:https://arxiv.org/abs/2507.02666
备注:Accepted at Interspeech2025
摘要:在音频自监督表示学习的最新进展中,标准Transformer架构已成为主要方法,但其注意力机制通常将一部分注意力权重分配给不相关的信息,从而可能损害模型的区分能力。为了解决这个问题,我们引入了一个差分注意力机制,它有效地减轻了无效的注意力分配,通过集成双softmax操作和适当调整的微分系数。实验结果表明,我们的ASDA模型在多个基准测试中达到了最先进的(SOTA)性能,包括音频分类(AS-2 M上的49.0% mAP,AS 20 K上的41.5% mAP),关键字定位(SPC-2上的98.3%准确率)和环境声音分类(ESC-50上的96.1%准确率)。这些结果突出了ASDA在音频任务中的有效性,为更广泛的应用铺平了道路。
摘要:In recent advancements in audio self-supervised representation learning, the standard Transformer architecture has emerged as the predominant approach, yet its attention mechanism often allocates a portion of attention weights to irrelevant information, potentially impairing the model's discriminative ability. To address this, we introduce a differential attention mechanism, which effectively mitigates ineffective attention allocation through the integration of dual-softmax operations and appropriately tuned differential coefficients. Experimental results demonstrate that our ASDA model achieves state-of-the-art (SOTA) performance across multiple benchmarks, including audio classification (49.0% mAP on AS-2M, 41.5% mAP on AS20K), keyword spotting (98.3% accuracy on SPC-2), and environmental sound classification (96.1% accuracy on ESC-50). These results highlight ASDA's effectiveness in audio tasks, paving the way for broader applications.
【2】De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks
链接:https://arxiv.org/abs/2507.02606
备注:Accepted by ICML 2025
摘要:语音生成模型的快速发展加剧了与语音克隆(VC)相关的隐私和安全问题。最近的研究调查了通过引入对抗性扰动来破坏未经授权的声音克隆。然而,有决心的攻击者可以减轻这些保护扰动并成功执行VC。在这项研究中,我们进行了第一次系统的评估,这些保护扰动对VC现实的威胁模型,包括扰动净化。我们的研究结果表明,虽然现有的净化方法可以中和相当一部分的保护扰动,他们仍然会导致VC模型的特征空间的扭曲,这降低了VC的性能。从这个角度来看,我们提出了一种新的两阶段的净化方法:(1)净化扰动语音;(2)使用音素指导来优化它,使其与干净的语音分布对齐。实验结果表明,我们的方法优于国家的最先进的净化方法,破坏VC防御。我们的研究揭示了基于对抗扰动的VC防御的局限性,并强调迫切需要更强大的解决方案来减轻VC带来的安全和隐私风险。代码和音频示例可在https://de-antifake.github.io上获得。
摘要:The rapid advancement of speech generation models has heightened privacy and security concerns related to voice cloning (VC). Recent studies have investigated disrupting unauthorized voice cloning by introducing adversarial perturbations. However, determined attackers can mitigate these protective perturbations and successfully execute VC. In this study, we conduct the first systematic evaluation of these protective perturbations against VC under realistic threat models that include perturbation purification. Our findings reveal that while existing purification methods can neutralize a considerable portion of the protective perturbations, they still lead to distortions in the feature space of VC models, which degrades the performance of VC. From this perspective, we propose a novel two-stage purification method: (1) Purify the perturbed speech; (2) Refine it using phoneme guidance to align it with the clean speech distribution. Experimental results demonstrate that our method outperforms state-of-the-art purification methods in disrupting VC defenses. Our study reveals the limitations of adversarial perturbation-based VC defenses and underscores the urgent need for more robust solutions to mitigate the security and privacy risks posed by VC. The code and audio samples are available at https://de-antifake.github.io.
【3】Padé Approximant Neural Networks for Enhanced Electric Motor Fault Diagnosis Using Vibration and Acoustic Data
链接:https://arxiv.org/abs/2507.02599
备注:Submitted to the Journal of Vibration Engineering & Technologies
摘要:目的:本研究的主要目的是通过利用Pad 'e近似神经元(PAON)模型来增强感应电机的故障诊断。虽然加速度计和麦克风是运动状态监测的标准,但具有非线性神经元架构的深度学习模型在诊断性能方面提供了有希望的改进。这项研究解决了这样一个问题:在使用振动和声学数据诊断电气和机械故障时,Pad\'e近似神经网络(Pad\' eNets)能否优于传统的卷积神经网络(CNN)和自组织操作神经网络(Self-ONNs)? 方法:我们评估并比较了三种深度学习架构的诊断能力:一维CNN,Self-ONNs和Pad\'eNets。这些模型在渥太华大学公开的恒速感应电机数据集上进行了测试,其中包括振动和声学传感器数据。Pad\'eNet模型旨在引入增强的非线性,并与无界激活函数(如Leaky ReLU)兼容。 结果与结论:Pad\'eNets始终优于基准模型,分别为加速度计1、2、3和声传感器实现了99.96%、98.26%、97.61%和98.33%的诊断准确度。Pad\'eNets增强的非线性,以及它们与无界激活函数的兼容性,显著提高了感应电机状态监测中的故障诊断性能。
摘要:Purpose: The primary aim of this study is to enhance fault diagnosis in induction machines by leveraging the Pad\'e Approximant Neuron (PAON) model. While accelerometers and microphones are standard in motor condition monitoring, deep learning models with nonlinear neuron architectures offer promising improvements in diagnostic performance. This research addresses the question: Can Pad\'e Approximant Neural Networks (Pad\'eNets) outperform conventional Convolutional Neural Networks (CNNs) and Self-Organized Operational Neural Networks (Self-ONNs) in diagnosing electrical and mechanical faults using vibration and acoustic data? Methods: We evaluate and compare the diagnostic capabilities of three deep learning architectures: one-dimensional CNNs, Self-ONNs, and Pad\'eNets. These models are tested on the University of Ottawa's publicly available constant-speed induction motor datasets, which include both vibration and acoustic sensor data. The Pad\'eNet model is designed to introduce enhanced nonlinearity and is compatible with unbounded activation functions such as Leaky ReLU. Results and Conclusion: Pad\'eNets consistently outperformed the baseline models, achieving diagnostic accuracies of 99.96%, 98.26%, 97.61%, and 98.33% for accelerometers 1, 2, 3, and the acoustic sensor, respectively. The enhanced nonlinearity of Pad\'eNets, together with their compatibility with unbounded activation functions, significantly improves fault diagnosis performance in induction motor condition monitoring.
【4】Benchmarking Akan ASR Models Across Domain-Specific Datasets: A Comparative Evaluation of Performance, Scalability, and Adaptability
链接:https://arxiv.org/abs/2507.02407
备注:This version has been reviewed and accepted for presentation at the Future Technologies Conference (FTC) 2025, to be held on 6 & 7 November 2025 in Munich, Germany. 17 pages, 4 figures, 1 table
摘要:大多数现有的自动语音识别(ASR)研究使用域内数据集来评估模型。然而,他们很少评估他们如何在不同的语音环境中推广。这项研究解决了这一差距,基准测试7阿肯ASR模型建立在Transformer架构,如耳语和Wav2Vec2,使用四个阿肯语音语料库,以确定其性能。这些数据集涵盖各个领域,包括文化相关的图像描述,非正式对话,圣经经文阅读和自发的金融对话。单词错误率和字符错误率的比较突出了域依赖性,模型仅在其训练域内表现最佳,而在不匹配的情况下显示出显着的准确性下降。该研究还确定了Whisper和Wav2Vec2架构之间的不同错误行为。虽然微调的Whisper Akan模型会导致更流畅但可能误导的转录错误,但Wav2Vec2在遇到不熟悉的输入时会产生更明显但更难解释的输出。在为低资源语言(LRL)应用程序选择架构时,应考虑ASR错误的可读性和透明度之间的权衡。这些发现强调了有针对性的域适应技术,自适应路由策略,和多语言的培训框架,阿肯和其他LRL的需要。
摘要:Most existing automatic speech recognition (ASR) research evaluate models using in-domain datasets. However, they seldom evaluate how they generalize across diverse speech contexts. This study addresses this gap by benchmarking seven Akan ASR models built on transformer architectures, such as Whisper and Wav2Vec2, using four Akan speech corpora to determine their performance. These datasets encompass various domains, including culturally relevant image descriptions, informal conversations, biblical scripture readings, and spontaneous financial dialogues. A comparison of the word error rate and character error rate highlighted domain dependency, with models performing optimally only within their training domains while showing marked accuracy degradation in mismatched scenarios. This study also identified distinct error behaviors between the Whisper and Wav2Vec2 architectures. Whereas fine-tuned Whisper Akan models led to more fluent but potentially misleading transcription errors, Wav2Vec2 produced more obvious yet less interpretable outputs when encountering unfamiliar inputs. This trade-off between readability and transparency in ASR errors should be considered when selecting architectures for low-resource language (LRL) applications. These findings highlight the need for targeted domain adaptation techniques, adaptive routing strategies, and multilingual training frameworks for Akan and other LRLs.
【5】Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
链接:https://arxiv.org/abs/2507.02391
备注:None
摘要:我们探索无监督语音增强使用扩散模型作为表达生成先验干净的语音。现有的方法通过一个近似的,噪声扰动的似然得分,结合无条件得分通过权衡超参数使用嘈杂的语音引导反向扩散过程。在这项工作中,我们提出了两种替代算法,直接模拟扩散状态的条件反向转移分布。第一种方法以原则性的方式将扩散先验与观察模型集成,从而消除了对超参数调整的需要。第二个定义了一个扩散过程中的嘈杂的语音本身,产生一个完全听话和准确的似然得分。在WSJ 0-QUT和VoiceBank-DEMAND数据集上的实验表明,与监督和无监督基线相比,增强指标得到了改进,对域偏移的鲁棒性更强。
摘要:We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise-perturbed likelihood score, combined with the unconditional score via a trade-off hyperparameter. In this work, we propose two alternative algorithms that directly model the conditional reverse transition distribution of diffusion states. The first method integrates the diffusion prior with the observation model in a principled way, removing the need for hyperparameter tuning. The second defines a diffusion process over the noisy speech itself, yielding a fully tractable and exact likelihood score. Experiments on the WSJ0-QUT and VoiceBank-DEMAND datasets demonstrate improved enhancement metrics and greater robustness to domain shifts compared to both supervised and unsupervised baselines.
【6】JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
链接:https://arxiv.org/abs/2507.02380
摘要:JoyTTS是一个端到端的语音聊天机器人,它将大型语言模型(LLM)与文本到语音(TTS)技术相结合,具有语音克隆功能。这个项目是建立在开源的MiniCPM-o和CosyVoice 2模型上的,并在2000小时的会话数据上进行了训练。我们还提供了完整的培训代码,以方便社区进一步开发和优化。在测试机器seed-tts-zh上,它实现了0.73的SS(说话人相似性)得分和5.09的WER(单词错误率)。代码和模型以及训练和推理脚本可在https://github.com/jdh-algo/JoyTTS.git上获得。
摘要:JoyTTS is an end-to-end spoken chatbot that combines large language models (LLM) with text-to-speech (TTS) technology, featuring voice cloning capabilities. This project is built upon the open-source MiniCPM-o and CosyVoice2 models and trained on 2000 hours of conversational data. We have also provided the complete training code to facilitate further development and optimization by the community. On the testing machine seed-tts-zh, it achieves a SS (speaker similarity) score of 0.73 and a WER (Word Error Rate) of 5.09. The code and models, along with training and inference scripts, are available at https://github.com/jdh-algo/JoyTTS.git.
【7】Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
链接:https://arxiv.org/abs/2507.02273
备注:ISMIR 2025
摘要:通用的音频表示已被证明是有效的,在不同的音乐信息检索应用程序,但他们的效用在智能音乐制作仍然有限的理解不足的音频效果(Fx)。虽然以前的方法强调了混合级别的音频效果分析,但这种关注对于需要乐器方面的音频效果理解的任务(例如自动混合)来说是不够的。在这项工作中,我们提出了Fx-Encoder++,一种新的模型,旨在从音乐混合物中提取乐器方面的音频效果表示。我们的方法利用了对比学习框架,并引入了一个“提取器”机制,当提供乐器查询(音频或文本)时,将混合级音频效果嵌入转换为乐器级音频效果嵌入。我们在检索和音频效果参数匹配任务中评估了我们的模型,测试了它在各种乐器上的性能。结果表明,Fx-Encoder++在混合级别上优于以前的方法,并显示出一种新颖的能力,可以从乐器中提取效果表示,解决智能音乐制作系统中的关键能力差距。
摘要:General-purpose audio representations have proven effective across diverse music information retrieval applications, yet their utility in intelligent music production remains limited by insufficient understanding of audio effects (Fx). Although previous approaches have emphasized audio effects analysis at the mixture level, this focus falls short for tasks demanding instrument-wise audio effects understanding, such as automatic mixing. In this work, we present Fx-Encoder++, a novel model designed to extract instrument-wise audio effects representations from music mixtures. Our approach leverages a contrastive learning framework and introduces an "extractor" mechanism that, when provided with instrument queries (audio or text), transforms mixture-level audio effects embeddings into instrument-wise audio effects embeddings. We evaluated our model across retrieval and audio effects parameter matching tasks, testing its performance across a diverse range of instruments. The results demonstrate that Fx-Encoder++ outperforms previous approaches at mixture level and show a novel ability to extract effects representation instrument-wise, addressing a critical capability gap in intelligent music production systems.
【8】Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
链接:https://arxiv.org/abs/2507.02176
备注:Accepted at SSW13 - Interspeech 2025 Speech Synthesis Workshop
摘要:语音身份建模是具有挑战性的,由于其多方面的性质。在生成语音系统中,身份通常使用自动说话人验证(ASV)嵌入来评估,其设计用于区分而不是表征身份。本文探讨了哪些方面的声音被捕获在这样的表示。我们发现,广泛使用的ASV嵌入主要集中在静态特征,如音色和音高范围,而忽略了动态元素,如节奏。我们还确定了混杂因素,损害扬声器相似性测量,并提出缓解策略。为了解决这些差距,我们提出了U3D,一个评估扬声器的动态节奏模式的指标。这项工作有助于不断改善的语音克隆系统的背景下,评估扬声器身份的一致性的持续挑战。我们公开发布代码。
摘要:Modeling voice identity is challenging due to its multifaceted nature. In generative speech systems, identity is often assessed using automatic speaker verification (ASV) embeddings, designed for discrimination rather than characterizing identity. This paper investigates which aspects of a voice are captured in such representations. We find that widely used ASV embeddings focus mainly on static features like timbre and pitch range, while neglecting dynamic elements such as rhythm. We also identify confounding factors that compromise speaker similarity measurements and suggest mitigation strategies. To address these gaps, we propose U3D, a metric that evaluates speakers' dynamic rhythm patterns. This work contributes to the ongoing challenge of assessing speaker identity consistency in the context of ever-better voice cloning systems. We publicly release our code.
【9】Parametric Neural Amp Modeling with Active Learning
链接:https://arxiv.org/abs/2507.02109
备注:Accepted at ISMIR 2025 as Late-Breaking Demo (LBD)
摘要:我们介绍PANAMA,一个主动学习框架,用于使用WaveNet架构训练端到端参数吉他放大器模型。使用\model,可以通过记录由主动学习策略确定的样本来创建虚拟放大器,以使用最少量的数据点(即,放大器旋钮设置)。我们表明,基于梯度的优化算法可以用来确定最佳的数据点进行采样,该方法有助于在有限数量的样本。
摘要:We introduce PANAMA, an active learning framework for the training of end-to-end parametric guitar amp models using a WaveNet-like architecture. With \model, one can create a virtual amp by recording samples that are determined by an active learning strategy to use a minimum amount of datapoints (i.e., amp knob settings). We show that gradient-based optimization algorithms can be used to determine the optimal datapoints to sample, and that the approach helps under a constrained number of samples.
【10】TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
链接:https://arxiv.org/abs/2507.02080
备注:9 pages, 2 figures, 2 tables
摘要:由于噪音以及音频和视觉模态之间的不对准,多模态情感识别通常会在效价唤醒估计中出现性能下降。为了解决这一挑战,我们引入了TAGF,一个时间感知的门控融合框架,用于多模态情感识别。TAGF基于时间动态自适应地调制递归注意输出的贡献。具体而言,TAGF采用了基于BiLSTM的时间门控机制来学习每个递归步骤的相对重要性,并有效地集成了多步跨模态特征。通过将时间感知嵌入到递归融合过程中,TAGF有效地捕获了情感表达的顺序演变和模态之间的复杂相互作用。在Aff-Wild 2数据集上的实验结果表明,与现有的基于递归注意力的模型相比,TAGF具有竞争力的性能。此外,TAGF表现出很强的鲁棒性跨模态错位和可靠的模型在现实世界的条件下的动态情感转变。
摘要:Multimodal emotion recognition often suffers from performance degradation in valence-arousal estimation due to noise and misalignment between audio and visual modalities. To address this challenge, we introduce TAGF, a Time-aware Gated Fusion framework for multimodal emotion recognition. The TAGF adaptively modulates the contribution of recursive attention outputs based on temporal dynamics. Specifically, the TAGF incorporates a BiLSTM-based temporal gating mechanism to learn the relative importance of each recursive step and effectively integrates multistep cross-modal features. By embedding temporal awareness into the recursive fusion process, the TAGF effectively captures the sequential evolution of emotional expressions and the complex interplay between modalities. Experimental results on the Aff-Wild2 dataset demonstrate that TAGF achieves competitive performance compared with existing recursive attention-based models. Furthermore, TAGF exhibits strong robustness to cross-modal misalignment and reliably models dynamic emotional transitions in real-world conditions.
【11】Acoustic evaluation of a neural network dedicated to the detection of animal vocalisations
链接:https://arxiv.org/abs/2507.01974
备注:None
摘要:长期记录器的可访问性,适应有时苛刻的现场条件,使部署广泛的动物种群监测活动,通过生态声学。自动信号检测方法的有效性越来越多地基于神经方法,通常仅通过机器学习指标进行评估,而性能的声学分析仍然很少。作为岩石雷鸟种群声学监测的一部分,我们在这里提出了一个简单的方法,声学分析的检测系统的性能。所提出的措施是基于相关的合成信号的信噪比,其检测概率。我们展示了这种措施如何提供有关系统的信息,并允许优化其培训。我们还展示了它如何使检测距离建模,从而提供了根据声音环境评估其动态和访问的空间密度的估计调用的可能性。
摘要:The accessibility of long-duration recorders, adapted to sometimes demanding field conditions, has enabled the deployment of extensive animal population monitoring campaigns through ecoacoustics. The effectiveness of automatic signal detection methods, increasingly based on neural approaches, is frequently evaluated solely through machine learning metrics, while acoustic analysis of performance remains rare. As part of the acoustic monitoring of Rock Ptarmigan populations, we propose here a simple method for acoustic analysis of the detection system's performance. The proposed measure is based on relating the signal-to-noise ratio of synthetic signals to their probability of detection. We show how this measure provides information about the system and allows optimisation of its training. We also show how it enables modelling of the detection distance, thus offering the possibility of evaluating its dynamics according to the sound environment and accessing an estimation of the spatial density of calls.
【12】Towards Perception-Informed Latent HRTF Representations
链接:https://arxiv.org/abs/2507.02815
备注:Accepted by IEEE WASPAA 2025, camera-ready version
摘要:个性化的头部相关传递函数(HRTF)对于确保耳机上的真实听觉体验至关重要,因为它们考虑到了影响听力的个体解剖差异。大多数HRTF个性化的机器学习方法都依赖于学习的低维潜在空间来为听众生成或选择自定义HRTF。然而,这些潜在表示通常以优化频谱重建但不优化感知兼容性的方式学习,这意味着它们可能不一定与感知距离对齐。在这项工作中,我们首先研究传统上学习的HRTF表示是否与使用基于属性的客观感知度量的感知关系很好地相关;然后,我们提出了一种方法,用于将HRTF显式嵌入到感知通知的潜在空间中,利用基于度量的损失函数和通过度量多维缩放(MMDS)的监督。最后,我们证明了这些学习表示的HRTF个性化的任务的适用性。我们认为,我们的方法有可能呈现个性化的空间音频,从而改善听觉体验。
摘要:Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning approaches to HRTF personalization rely on a learned low-dimensional latent space to generate or select custom HRTFs for a listener. However, these latent representations are typically learned in a manner that optimizes for spectral reconstruction but not for perceptual compatibility, meaning they may not necessarily align with perceptual distance. In this work, we first study whether traditionally learned HRTF representations are well correlated with perceptual relations using auditory-based objective perceptual metrics; we then propose a method for explicitly embedding HRTFs into a perception-informed latent space, leveraging a metric-based loss function and supervision via Metric Multidimensional Scaling (MMDS). Finally, we demonstrate the applicability of these learned representations to the task of HRTF personalization. We suggest that our method has the potential to render personalized spatial audio, leading to an improved listening experience.
【13】Self-Steering Deep Non-Linear Spatially Selective Filters for Efficient Extraction of Moving Speakers under Weak Guidance
链接:https://arxiv.org/abs/2507.02791
备注:Accepted at IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:深度非线性空间选择性滤波器的最新研究表明,对于已知方向的固定扬声器,具有计算量小的架构,具有出色的增强性能。然而,为了在动态场景中保持这种性能,资源密集型数据驱动的跟踪算法变得必要,以提供精确的空间指导的条件下的初始方向的目标扬声器。由于这种额外的计算开销阻碍了应用在资源受限的情况下,如实时语音增强,我们提出了一种新的策略,利用低复杂度的跟踪算法的粒子滤波器的形式代替。假设一个因果关系,顺序处理风格,我们引入时间反馈,利用增强的语音信号的空间选择性滤波器,以补偿有限的建模能力的粒子滤波器。合成数据集上的评估说明了两种算法之间的自回归相互作用如何大大提高跟踪精度,并导致强大的增强性能。一个真实世界录音的听力测试补充了这些发现,表明我们提出的自导向管道作为可比方法的首选的明显趋势。
摘要:Recent works on deep non-linear spatially selective filters demonstrate exceptional enhancement performance with computationally lightweight architectures for stationary speakers of known directions. However, to maintain this performance in dynamic scenarios, resource-intensive data-driven tracking algorithms become necessary to provide precise spatial guidance conditioned on the initial direction of a target speaker. As this additional computational overhead hinders application in resource-constrained scenarios such as real-time speech enhancement, we present a novel strategy utilizing a low-complexity tracking algorithm in the form of a particle filter instead. Assuming a causal, sequential processing style, we introduce temporal feedback to leverage the enhanced speech signal of the spatially selective filter to compensate for the limited modeling capabilities of the particle filter. Evaluation on a synthetic dataset illustrates how the autoregressive interplay between both algorithms drastically improves tracking accuracy and leads to strong enhancement performance. A listening test with real-world recordings complements these findings by indicating a clear trend towards our proposed self-steering pipeline as preferred choice over comparable methods.
【14】DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
链接:https://arxiv.org/abs/2507.02768
备注:Model and code available at: this https URL
摘要:我们介绍DeSTA2.5-Audio,这是一种通用的大型音频语言模型(LALM),旨在实现强大的听觉感知和听觉跟随,而不需要特定于任务的音频听觉调节。最近的LALM通常通过在大规模、手动策划或LLM合成的音频指令数据集上进行训练来增强具有听觉能力的大型语言模型(LLM)。然而,这些方法往往遭受灾难性的忘记LLM的原始语言能力。为了解决这个问题,我们重新审视了数据构建管道,并提出了DeSTA,这是一种自生成的跨模态对齐策略,其中骨干LLM生成自己的训练目标。这种方法保留了LLM的母语能力,同时建立有效的音频文本对齐,从而实现zero-shot泛化而无需特定于任务的调整。使用DeSTA,我们构建了DeSTA-AQA 5 M,这是一个大规模的任务不可知数据集,包含500万个训练样本,这些样本来自7,000小时的音频,涵盖50个不同的数据集,包括语音,环境声音和音乐。DeSTA2.5-Audio在各种音频语言基准测试中实现了最先进或具有竞争力的性能,包括Dynamic-SUPERB、MMAU、SAKURA、Speech-IFEval和VoiceBench。全面的比较研究表明,我们的自我生成的策略优于广泛采用的数据构建和训练策略的听觉感知和预防以下的能力。我们的研究结果强调了LALM开发中精心设计的数据构建的重要性,并为构建强大的通用LALM提供了实用的见解。
摘要:We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following, without requiring task-specific audio instruction-tuning. Recent LALMs typically augment Large Language Models (LLMs) with auditory capabilities by training on large-scale, manually curated or LLM-synthesized audio-instruction datasets. However, these approaches have often suffered from the catastrophic forgetting of the LLM's original language abilities. To address this, we revisit the data construction pipeline and propose DeSTA, a self-generated cross-modal alignment strategy in which the backbone LLM generates its own training targets. This approach preserves the LLM's native language proficiency while establishing effective audio-text alignment, thereby enabling zero-shot generalization without task-specific tuning. Using DeSTA, we construct DeSTA-AQA5M, a large-scale, task-agnostic dataset containing 5 million training samples derived from 7,000 hours of audio spanning 50 diverse datasets, including speech, environmental sounds, and music. DeSTA2.5-Audio achieves state-of-the-art or competitive performance across a wide range of audio-language benchmarks, including Dynamic-SUPERB, MMAU, SAKURA, Speech-IFEval, and VoiceBench. Comprehensive comparative studies demonstrate that our self-generated strategy outperforms widely adopted data construction and training strategies in both auditory perception and instruction-following capabilities. Our findings underscore the importance of carefully designed data construction in LALM development and offer practical insights for building robust, general-purpose LALMs.
【15】Multi-Utterance Speech Separation and Association Trained on Short Segments
链接:https://arxiv.org/abs/2507.02562
备注:5 pages, accepted by WASPAA 2025
摘要:当前基于深度神经网络(DNN)的语音分离面临着一个根本性的挑战-虽然由于计算限制,模型需要在短片段上进行训练,但现实世界的应用通常需要处理比训练期间更长的记录,每个说话者有多个话语。在本文中,我们研究了现有的方法如何在这种具有挑战性的情况下执行,并提出了一个频率-时间递归神经网络(FTRNN),有效地弥合这一差距。我们的FTRNN采用了一个全波段模块,在每个时间帧和一个子带模块,在每个频带的时间模式模型的频率依赖性。尽管在10秒的短固定长度段上进行了训练,但我们的模型在处理明显长于训练段(21-121秒)的信号时表现出了鲁棒的分离,并在训练期间超过了所看到的话语间隙中保留了说话者关联。与传统的分段分离缝合范式不同,我们的轻量级方法(0.9 M参数)在没有分段的情况下对长音频进行推断,消除了分段边界失真,同时简化了部署。实验结果表明,FTRNN的泛化能力的多话语语音分离和说话人关联。
摘要:Current deep neural network (DNN) based speech separation faces a fundamental challenge -- while the models need to be trained on short segments due to computational constraints, real-world applications typically require processing significantly longer recordings with multiple utterances per speaker than seen during training. In this paper, we investigate how existing approaches perform in this challenging scenario and propose a frequency-temporal recurrent neural network (FTRNN) that effectively bridges this gap. Our FTRNN employs a full-band module to model frequency dependencies within each time frame and a sub-band module that models temporal patterns in each frequency band. Despite being trained on short fixed-length segments of 10 s, our model demonstrates robust separation when processing signals significantly longer than training segments (21-121 s) and preserves speaker association across utterance gaps exceeding those seen during training. Unlike the conventional segment-separation-stitch paradigm, our lightweight approach (0.9 M parameters) performs inference on long audio without segmentation, eliminating segment boundary distortions while simplifying deployment. Experimental results demonstrate the generalization ability of FTRNN for multi-utterance speech separation and speaker association.
【1】Towards Perception-Informed Latent HRTF Representations
链接:https://arxiv.org/abs/2507.02815
备注:Accepted by IEEE WASPAA 2025, camera-ready version
摘要:个性化的头部相关传递函数(HRTF)对于确保耳机上的真实听觉体验至关重要,因为它们考虑到了影响听力的个体解剖差异。大多数HRTF个性化的机器学习方法都依赖于学习的低维潜在空间来为听众生成或选择自定义HRTF。然而,这些潜在表示通常以优化频谱重建但不优化感知兼容性的方式学习,这意味着它们可能不一定与感知距离对齐。在这项工作中,我们首先研究传统上学习的HRTF表示是否与使用基于属性的客观感知度量的感知关系很好地相关;然后,我们提出了一种方法,用于将HRTF显式嵌入到感知通知的潜在空间中,利用基于度量的损失函数和通过度量多维缩放(MMDS)的监督。最后,我们证明了这些学习表示的HRTF个性化的任务的适用性。我们认为,我们的方法有可能呈现个性化的空间音频,从而改善听觉体验。
摘要:Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning approaches to HRTF personalization rely on a learned low-dimensional latent space to generate or select custom HRTFs for a listener. However, these latent representations are typically learned in a manner that optimizes for spectral reconstruction but not for perceptual compatibility, meaning they may not necessarily align with perceptual distance. In this work, we first study whether traditionally learned HRTF representations are well correlated with perceptual relations using auditory-based objective perceptual metrics; we then propose a method for explicitly embedding HRTFs into a perception-informed latent space, leveraging a metric-based loss function and supervision via Metric Multidimensional Scaling (MMDS). Finally, we demonstrate the applicability of these learned representations to the task of HRTF personalization. We suggest that our method has the potential to render personalized spatial audio, leading to an improved listening experience.
【2】Self-Steering Deep Non-Linear Spatially Selective Filters for Efficient Extraction of Moving Speakers under Weak Guidance
链接:https://arxiv.org/abs/2507.02791
备注:Accepted at IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:深度非线性空间选择性滤波器的最新研究表明,对于已知方向的固定扬声器,具有计算量小的架构,具有出色的增强性能。然而,为了在动态场景中保持这种性能,资源密集型数据驱动的跟踪算法变得必要,以提供精确的空间指导的条件下的初始方向的目标扬声器。由于这种额外的计算开销阻碍了应用在资源受限的情况下,如实时语音增强,我们提出了一种新的策略,利用低复杂度的跟踪算法的粒子滤波器的形式代替。假设一个因果关系,顺序处理风格,我们引入时间反馈,利用增强的语音信号的空间选择性滤波器,以补偿有限的建模能力的粒子滤波器。合成数据集上的评估说明了两种算法之间的自回归相互作用如何大大提高跟踪精度,并导致强大的增强性能。一个真实世界录音的听力测试补充了这些发现,表明我们提出的自导向管道作为可比方法的首选的明显趋势。
摘要:Recent works on deep non-linear spatially selective filters demonstrate exceptional enhancement performance with computationally lightweight architectures for stationary speakers of known directions. However, to maintain this performance in dynamic scenarios, resource-intensive data-driven tracking algorithms become necessary to provide precise spatial guidance conditioned on the initial direction of a target speaker. As this additional computational overhead hinders application in resource-constrained scenarios such as real-time speech enhancement, we present a novel strategy utilizing a low-complexity tracking algorithm in the form of a particle filter instead. Assuming a causal, sequential processing style, we introduce temporal feedback to leverage the enhanced speech signal of the spatially selective filter to compensate for the limited modeling capabilities of the particle filter. Evaluation on a synthetic dataset illustrates how the autoregressive interplay between both algorithms drastically improves tracking accuracy and leads to strong enhancement performance. A listening test with real-world recordings complements these findings by indicating a clear trend towards our proposed self-steering pipeline as preferred choice over comparable methods.
【3】DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
链接:https://arxiv.org/abs/2507.02768
备注:Model and code available at: this https URL
摘要:我们介绍DeSTA2.5-Audio,这是一种通用的大型音频语言模型(LALM),旨在实现强大的听觉感知和听觉跟随,而不需要特定于任务的音频听觉调节。最近的LALM通常通过在大规模、手动策划或LLM合成的音频指令数据集上进行训练来增强具有听觉能力的大型语言模型(LLM)。然而,这些方法往往遭受灾难性的忘记LLM的原始语言能力。为了解决这个问题,我们重新审视了数据构建管道,并提出了DeSTA,这是一种自生成的跨模态对齐策略,其中骨干LLM生成自己的训练目标。这种方法保留了LLM的母语能力,同时建立有效的音频文本对齐,从而实现zero-shot泛化而无需特定于任务的调整。使用DeSTA,我们构建了DeSTA-AQA5M,这是一个大规模的任务不可知数据集,包含500万个训练样本,这些样本来自7,000小时的音频,涵盖50个不同的数据集,包括语音,环境声音和音乐。DeSTA2.5-Audio在各种音频语言基准测试中实现了最先进或具有竞争力的性能,包括Dynamic-SUPERB、MMAU、SAKURA、Speech-IFEval和VoiceBench。全面的比较研究表明,我们的自我生成的策略优于广泛采用的数据构建和训练策略的听觉感知和预防以下的能力。我们的研究结果强调了LALM开发中精心设计的数据构建的重要性,并为构建强大的通用LALM提供了实用的见解。
摘要:We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following, without requiring task-specific audio instruction-tuning. Recent LALMs typically augment Large Language Models (LLMs) with auditory capabilities by training on large-scale, manually curated or LLM-synthesized audio-instruction datasets. However, these approaches have often suffered from the catastrophic forgetting of the LLM's original language abilities. To address this, we revisit the data construction pipeline and propose DeSTA, a self-generated cross-modal alignment strategy in which the backbone LLM generates its own training targets. This approach preserves the LLM's native language proficiency while establishing effective audio-text alignment, thereby enabling zero-shot generalization without task-specific tuning. Using DeSTA, we construct DeSTA-AQA5M, a large-scale, task-agnostic dataset containing 5 million training samples derived from 7,000 hours of audio spanning 50 diverse datasets, including speech, environmental sounds, and music. DeSTA2.5-Audio achieves state-of-the-art or competitive performance across a wide range of audio-language benchmarks, including Dynamic-SUPERB, MMAU, SAKURA, Speech-IFEval, and VoiceBench. Comprehensive comparative studies demonstrate that our self-generated strategy outperforms widely adopted data construction and training strategies in both auditory perception and instruction-following capabilities. Our findings underscore the importance of carefully designed data construction in LALM development and offer practical insights for building robust, general-purpose LALMs.
【4】Multi-agent Auditory Scene Analysis
链接:https://arxiv.org/abs/2507.02755
备注:Submitted to Applied Intelligence
摘要:听觉场景分析(ASA)的目的是从声环境中检索信息,通过执行三个主要任务:声源定位,分离和分类。这些任务传统上使用线性数据流执行,其中首先定位声源;然后,使用它们的位置,将每个源分离成其自己的音频流;从每个音频流中提取与应用场景相关的信息(音频事件检测,说话人识别,情感分类等)。然而,运行这些任务线性地增加了整体响应时间,同时使最后一个任务(分离和分类)对第一个任务(位置)的错误高度敏感。在现有技术中已经采用了相当大量的努力和计算复杂性来开发尽可能最不容易出错的技术。然而,这样做会导致ASA系统在许多需要小计算占用空间和低响应时间的应用中不可行,例如生物声学、助听器设计、搜索和救援、人机交互等。为此,在这项工作中,提出了一种多代理方法来执行ASA,其中任务并行运行,在它们之间具有反馈回路以补偿局部误差,例如:使用分离输出的质量来校正定位误差;以及使用分类结果来降低定位对干扰的敏感度。其结果是一个多智能体的听觉场景分析(MASA)系统,对本地错误是强大的,没有相当大的复杂性增加,并具有较低的响应时间。完整的MASA系统是作为一个框架提供的,它使用开源工具进行声音采集和再现(JACK)和代理间通信(ROS 2),允许用户添加自己的代理。
摘要:Auditory scene analysis (ASA) aims to retrieve information from the acoustic environment, by carrying out three main tasks: sound source location, separation, and classification. These tasks are traditionally executed with a linear data flow, where the sound sources are first located; then, using their location, each source is separated into its own audio stream; from each of which, information is extracted that is relevant to the application scenario (audio event detection, speaker identification, emotion classification, etc.). However, running these tasks linearly increases the overall response time, while making the last tasks (separation and classification) highly sensitive to errors of the first task (location). A considerable amount of effort and computational complexity has been employed in the state-of-the-art to develop techniques that are the least error-prone possible. However, doing so gives rise to an ASA system that is non-viable in many applications that require a small computational footprint and a low response time, such as bioacoustics, hearing-aid design, search and rescue, human-robot interaction, etc. To this effect, in this work, a multi-agent approach is proposed to carry out ASA where the tasks are run in parallel, with feedback loops between them to compensate for local errors, such as: using the quality of the separation output to correct the location error; and using the classification result to reduce the localization's sensitivity towards interferences. The result is a multi-agent auditory scene analysis (MASA) system that is robust against local errors, without a considerable increase in complexity, and with a low response time. The complete proposed MASA system is provided as a framework that uses open-source tools for sound acquisition and reproduction (JACK) and inter-agent communication (ROS2), allowing users to add their own agents.
【5】Multi-Utterance Speech Separation and Association Trained on Short Segments
链接:https://arxiv.org/abs/2507.02562
备注:5 pages, accepted by WASPAA 2025
摘要:当前基于深度神经网络(DNN)的语音分离面临着一个根本性的挑战-虽然由于计算限制,模型需要在短片段上进行训练,但现实世界的应用通常需要处理比训练期间更长的记录,每个说话者有多个话语。在本文中,我们研究了现有的方法如何在这种具有挑战性的情况下执行,并提出了一个频率-时间递归神经网络(FTRNN),有效地弥合这一差距。我们的FTRNN采用了一个全波段模块,在每个时间帧和一个子带模块,在每个频带的时间模式模型的频率依赖性。尽管在10秒的短固定长度段上进行了训练,但我们的模型在处理明显长于训练段(21-121秒)的信号时表现出了鲁棒的分离,并在训练期间超过了所看到的话语间隙中保留了说话者关联。与传统的分段分离缝合范式不同,我们的轻量级方法(0.9 M参数)在没有分段的情况下对长音频进行推断,消除了分段边界失真,同时简化了部署。实验结果表明,FTRNN的泛化能力的多话语语音分离和说话人关联。
摘要:Current deep neural network (DNN) based speech separation faces a fundamental challenge -- while the models need to be trained on short segments due to computational constraints, real-world applications typically require processing significantly longer recordings with multiple utterances per speaker than seen during training. In this paper, we investigate how existing approaches perform in this challenging scenario and propose a frequency-temporal recurrent neural network (FTRNN) that effectively bridges this gap. Our FTRNN employs a full-band module to model frequency dependencies within each time frame and a sub-band module that models temporal patterns in each frequency band. Despite being trained on short fixed-length segments of 10 s, our model demonstrates robust separation when processing signals significantly longer than training segments (21-121 s) and preserves speaker association across utterance gaps exceeding those seen during training. Unlike the conventional segment-separation-stitch paradigm, our lightweight approach (0.9 M parameters) performs inference on long audio without segmentation, eliminating segment boundary distortions while simplifying deployment. Experimental results demonstrate the generalization ability of FTRNN for multi-utterance speech separation and speaker association.
【6】Open-Source System for Multilingual Translation and Cloned Speech Synthesis
链接:https://arxiv.org/abs/2507.02530
备注:Presented at Forum Acusticum Euronoise 2025
摘要:我们提出了一个专为多语言翻译和语音再生而设计的开源系统,解决了在不同语言环境中沟通和可访问性方面的挑战。该系统将Whisper语音识别与语音活动检测(VAD)集成在一起,以识别说话间隔,然后是大型语言模型(LLM)的管道。对于多语言应用程序,第一个LLM将语音分割成连贯的完整句子,然后第二个LLM进行翻译。对于语音再生,该系统使用具有语音克隆功能的文本到语音(TTS)模块来复制原始说话者的语音,保持自然性和说话者身份。 该系统的开源组件可以在本地或通过API运行,从而在各种用例中提供具有成本效益的部署。其中包括Zoom会话中的实时多语言翻译、公共广播的语音再生以及通过个人设备支持蓝牙的多语言播放。通过保留说话者的声音,该系统确保了无缝和身临其境的体验,无论是翻译还是再生语音。 这个开源项目与社区共享,以促进创新和可访问性。我们提供了详细的系统性能分析,包括延迟和单词准确性,展示了其在现实世界的多语言场景中实现包容性,适应性强的通信解决方案的潜力。
摘要:We present an open-source system designed for multilingual translation and speech regeneration, addressing challenges in communication and accessibility across diverse linguistic contexts. The system integrates Whisper for speech recognition with Voice Activity Detection (VAD) to identify speaking intervals, followed by a pipeline of Large Language Models (LLMs). For multilingual applications, the first LLM segments speech into coherent, complete sentences, which a second LLM then translates. For speech regeneration, the system uses a text-to-speech (TTS) module with voice cloning capabilities to replicate the original speaker's voice, maintaining naturalness and speaker identity. The system's open-source components can operate locally or via APIs, offering cost-effective deployment across various use cases. These include real-time multilingual translation in Zoom sessions, speech regeneration for public broadcasts, and Bluetooth-enabled multilingual playback through personal devices. By preserving the speaker's voice, the system ensures a seamless and immersive experience, whether translating or regenerating speech. This open-source project is shared with the community to foster innovation and accessibility. We provide a detailed system performance analysis, including latency and word accuracy, demonstrating its potential to enable inclusive, adaptable communication solutions in real-world multilingual scenarios.
【7】An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
链接:https://arxiv.org/abs/2507.02192
备注:5 pages
摘要:提出了一种新的迭代相位估计框架,称为多源Griffin-Lim算法(MSGLA),用于加性噪声条件下的语音增强(SE)。其核心思想是利用复值短时傅里叶变换(STFT)谱图的ad-hoc一致性约束来解决基于几何的相位估计中经常遇到的符号模糊性挑战。此外,我们引入了一个变形的几何约束框架的基础上的法律的正弦和余弦,制定了一个新的相位重建算法,使用噪声相位估计。我们首先通过一系列的Oracle实验验证了所提出的技术,在理想条件下证明了其有效性。然后,我们在VB-DMD和WSJ 0-CHiME 3数据集上评估了其性能,并表明所提出的MSGLA变体匹配良好或略优于现有算法,包括直接相位估计和基于DNN的符号预测,特别是在背景噪声抑制方面。
摘要:We propose a novel iterative phase estimation framework, termed multi-source Griffin-Lim algorithm (MSGLA), for speech enhancement (SE) under additive noise conditions. The core idea is to leverage the ad-hoc consistency constraint of complex-valued short-time Fourier transform (STFT) spectrograms to address the sign ambiguity challenge commonly encountered in geometry-based phase estimation. Furthermore, we introduce a variant of the geometric constraint framework based on the law of sines and cosines, formulating a new phase reconstruction algorithm using noise phase estimates. We first validate the proposed technique through a series of oracle experiments, demonstrating its effectiveness under ideal conditions. We then evaluate its performance on the VB-DMD and WSJ0-CHiME3 data sets, and show that the proposed MSGLA variants match well or slightly outperform existing algorithms, including direct phase estimation and DNN-based sign prediction, especially in terms of background noise suppression.
【8】Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
链接:https://arxiv.org/abs/2507.02115
备注:5 pages; 1 figure; Accepted to Speech Synthesis Workshop 2025 (SSW13)
摘要:合成第二语言(L2)语音是潜在的高度重视L2语言学习经验和反馈。然而,由于缺乏L2语音合成数据集,很难合成L2语音低资源的语言。在本文中,我们提供了一个实用的解决方案,编辑母语近似L2语音,并提出PPG 2Speech,一个基于扩散的多扬声器Phonetic-Posteriorgrams-to-Speech模型,能够编辑一个单一的音素没有文本对齐。我们使用Matcha-TTS的流匹配解码器作为骨干,将语音后验图(PPG)转换为基于外部扬声器嵌入和音高的梅尔频谱图。PPG 2Speech增强了Matcha-TTS的流匹配解码器,具有无分类器指导(CFG)和摇摆采样。我们还提出了一个新的特定于任务的客观评价指标,语音对齐一致性(PAC),编辑的PPG和从合成语音中提取的PPG之间的编辑效果。我们使用大约60小时的数据验证了我们的方法在芬兰语上的有效性,芬兰语是一种低资源的,几乎是语音语言。我们进行客观和主观的评价,我们的方法比较其自然度,说话人相似性,编辑效果与TTS为基础的编辑。我们的源代码发布在https://github.com/aalto-speech/PPG2Speech上。
摘要:Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback. However, due to the lack of L2 speech synthesis datasets, it is difficult to synthesize L2 speech for low-resourced languages. In this paper, we provide a practical solution for editing native speech to approximate L2 speech and present PPG2Speech, a diffusion-based multispeaker Phonetic-Posteriorgrams-to-Speech model that is capable of editing a single phoneme without text alignment. We use Matcha-TTS's flow-matching decoder as the backbone, transforming Phonetic Posteriorgrams (PPGs) to mel-spectrograms conditioned on external speaker embeddings and pitch. PPG2Speech strengthens the Matcha-TTS's flow-matching decoder with Classifier-free Guidance (CFG) and Sway Sampling. We also propose a new task-specific objective evaluation metric, the Phonetic Aligned Consistency (PAC), between the edited PPGs and the PPGs extracted from the synthetic speech for editing effects. We validate the effectiveness of our method on Finnish, a low-resourced, nearly phonetic language, using approximately 60 hours of data. We conduct objective and subjective evaluations of our approach to compare its naturalness, speaker similarity, and editing effectiveness with TTS-based editing. Our source code is published at https://github.com/aalto-speech/PPG2Speech.
【9】ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
链接:https://arxiv.org/abs/2507.02666
备注:Accepted at Interspeech2025
摘要:在音频自监督表示学习的最新进展中,标准Transformer架构已成为主要方法,但其注意力机制通常将一部分注意力权重分配给不相关的信息,从而可能损害模型的区分能力。为了解决这个问题,我们引入了一个差分注意力机制,它有效地减轻了无效的注意力分配,通过集成双softmax操作和适当调整的微分系数。实验结果表明,我们的ASDA模型在多个基准测试中达到了最先进的(SOTA)性能,包括音频分类(AS-2 M上的49.0% mAP,AS 20 K上的41.5% mAP),关键字定位(SPC-2上的98.3%准确率)和环境声音分类(ESC-50上的96.1%准确率)。这些结果突出了ASDA在音频任务中的有效性,为更广泛的应用铺平了道路。
摘要:In recent advancements in audio self-supervised representation learning, the standard Transformer architecture has emerged as the predominant approach, yet its attention mechanism often allocates a portion of attention weights to irrelevant information, potentially impairing the model's discriminative ability. To address this, we introduce a differential attention mechanism, which effectively mitigates ineffective attention allocation through the integration of dual-softmax operations and appropriately tuned differential coefficients. Experimental results demonstrate that our ASDA model achieves state-of-the-art (SOTA) performance across multiple benchmarks, including audio classification (49.0% mAP on AS-2M, 41.5% mAP on AS20K), keyword spotting (98.3% accuracy on SPC-2), and environmental sound classification (96.1% accuracy on ESC-50). These results highlight ASDA's effectiveness in audio tasks, paving the way for broader applications.
【10】De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks
链接:https://arxiv.org/abs/2507.02606
备注:Accepted by ICML 2025
摘要:语音生成模型的快速发展加剧了与语音克隆(VC)相关的隐私和安全问题。最近的研究调查了通过引入对抗性扰动来破坏未经授权的声音克隆。然而,有决心的攻击者可以减轻这些保护扰动并成功执行VC。在这项研究中,我们进行了第一次系统的评估,这些保护扰动对VC现实的威胁模型,包括扰动净化。我们的研究结果表明,虽然现有的净化方法可以中和相当一部分的保护扰动,他们仍然会导致VC模型的特征空间的扭曲,这降低了VC的性能。从这个角度来看,我们提出了一种新的两阶段的净化方法:(1)净化扰动语音;(2)使用音素指导来优化它,使其与干净的语音分布对齐。实验结果表明,我们的方法优于国家的最先进的净化方法,破坏VC防御。我们的研究揭示了基于对抗扰动的VC防御的局限性,并强调迫切需要更强大的解决方案来减轻VC带来的安全和隐私风险。代码和音频示例可在https://de-antifake.github.io上获得。
摘要:The rapid advancement of speech generation models has heightened privacy and security concerns related to voice cloning (VC). Recent studies have investigated disrupting unauthorized voice cloning by introducing adversarial perturbations. However, determined attackers can mitigate these protective perturbations and successfully execute VC. In this study, we conduct the first systematic evaluation of these protective perturbations against VC under realistic threat models that include perturbation purification. Our findings reveal that while existing purification methods can neutralize a considerable portion of the protective perturbations, they still lead to distortions in the feature space of VC models, which degrades the performance of VC. From this perspective, we propose a novel two-stage purification method: (1) Purify the perturbed speech; (2) Refine it using phoneme guidance to align it with the clean speech distribution. Experimental results demonstrate that our method outperforms state-of-the-art purification methods in disrupting VC defenses. Our study reveals the limitations of adversarial perturbation-based VC defenses and underscores the urgent need for more robust solutions to mitigate the security and privacy risks posed by VC. The code and audio samples are available at https://de-antifake.github.io.
【
11】Benchmarking Akan ASR Models Across Domain-Specific Datasets: A Comparative Evaluation of Performance, Scalability, and Adaptability
链接:https://arxiv.org/abs/2507.02407
备注:This version has been reviewed and accepted for presentation at the Future Technologies Conference (FTC) 2025, to be held on 6 & 7 November 2025 in Munich, Germany. 17 pages, 4 figures, 1 table
摘要:大多数现有的自动语音识别(ASR)研究使用域内数据集来评估模型。然而,他们很少评估他们如何在不同的语音环境中推广。这项研究解决了这一差距,基准测试7阿肯ASR模型建立在Transformer架构,如耳语和Wav2Vec2,使用四个阿肯语音语料库,以确定其性能。这些数据集涵盖各个领域,包括文化相关的图像描述,非正式对话,圣经经文阅读和自发的金融对话。单词错误率和字符错误率的比较突出了域依赖性,模型仅在其训练域内表现最佳,而在不匹配的情况下显示出显着的准确性下降。该研究还确定了Whisper和Wav2Vec2架构之间的不同错误行为。虽然微调的Whisper Akan模型会导致更流畅但可能误导的转录错误,但Wav2Vec2在遇到不熟悉的输入时会产生更明显但更难解释的输出。在为低资源语言(LRL)应用程序选择架构时,应考虑ASR错误的可读性和透明度之间的权衡。这些发现强调了有针对性的域适应技术,自适应路由策略,和多语言的培训框架,阿肯和其他LRL的需要。
摘要:Most existing automatic speech recognition (ASR) research evaluate models using in-domain datasets. However, they seldom evaluate how they generalize across diverse speech contexts. This study addresses this gap by benchmarking seven Akan ASR models built on transformer architectures, such as Whisper and Wav2Vec2, using four Akan speech corpora to determine their performance. These datasets encompass various domains, including culturally relevant image descriptions, informal conversations, biblical scripture readings, and spontaneous financial dialogues. A comparison of the word error rate and character error rate highlighted domain dependency, with models performing optimally only within their training domains while showing marked accuracy degradation in mismatched scenarios. This study also identified distinct error behaviors between the Whisper and Wav2Vec2 architectures. Whereas fine-tuned Whisper Akan models led to more fluent but potentially misleading transcription errors, Wav2Vec2 produced more obvious yet less interpretable outputs when encountering unfamiliar inputs. This trade-off between readability and transparency in ASR errors should be considered when selecting architectures for low-resource language (LRL) applications. These findings highlight the need for targeted domain adaptation techniques, adaptive routing strategies, and multilingual training frameworks for Akan and other LRLs.
【12】Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
链接:https://arxiv.org/abs/2507.02391
备注:None
摘要:我们探索无监督语音增强使用扩散模型作为表达生成先验干净的语音。现有的方法通过一个近似的,噪声扰动的似然得分,结合无条件得分通过权衡超参数使用嘈杂的语音引导反向扩散过程。在这项工作中,我们提出了两种替代算法,直接模拟扩散状态的条件反向转移分布。第一种方法以原则性的方式将扩散先验与观察模型集成,从而消除了对超参数调整的需要。第二个定义了一个扩散过程中的嘈杂的语音本身,产生一个完全听话和准确的似然得分。在WSJ 0-QUT和VoiceBank-DEMAND数据集上的实验表明,与监督和无监督基线相比,增强指标得到了改进,对域偏移的鲁棒性更强。
摘要:We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise-perturbed likelihood score, combined with the unconditional score via a trade-off hyperparameter. In this work, we propose two alternative algorithms that directly model the conditional reverse transition distribution of diffusion states. The first method integrates the diffusion prior with the observation model in a principled way, removing the need for hyperparameter tuning. The second defines a diffusion process over the noisy speech itself, yielding a fully tractable and exact likelihood score. Experiments on the WSJ0-QUT and VoiceBank-DEMAND datasets demonstrate improved enhancement metrics and greater robustness to domain shifts compared to both supervised and unsupervised baselines.
【13】JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
链接:https://arxiv.org/abs/2507.02380
摘要:JoyTTS是一个端到端的语音聊天机器人,它将大型语言模型(LLM)与文本到语音(TTS)技术相结合,具有语音克隆功能。这个项目是建立在开源的MiniCPM-o和CosyVoice 2模型上的,并在2000小时的会话数据上进行了训练。我们还提供了完整的培训代码,以方便社区进一步开发和优化。在测试机器seed-tts-zh上,它实现了0.73的SS(说话人相似性)得分和5.09的WER(单词错误率)。代码和模型以及训练和推理脚本可在https://github.com/jdh-algo/JoyTTS.git上获得。
摘要:JoyTTS is an end-to-end spoken chatbot that combines large language models (LLM) with text-to-speech (TTS) technology, featuring voice cloning capabilities. This project is built upon the open-source MiniCPM-o and CosyVoice2 models and trained on 2000 hours of conversational data. We have also provided the complete training code to facilitate further development and optimization by the community. On the testing machine seed-tts-zh, it achieves a SS (speaker similarity) score of 0.73 and a WER (Word Error Rate) of 5.09. The code and models, along with training and inference scripts, are available at https://github.com/jdh-algo/JoyTTS.git.
【14】Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
链接:https://arxiv.org/abs/2507.02273
备注:ISMIR 2025
摘要:通用的音频表示已被证明是有效的,在不同的音乐信息检索应用程序,但他们的效用在智能音乐制作仍然有限的理解不足的音频效果(Fx)。虽然以前的方法强调了混合级别的音频效果分析,但这种关注对于需要乐器方面的音频效果理解的任务(例如自动混合)来说是不够的。在这项工作中,我们提出了Fx-Encoder++,一种新的模型,旨在从音乐混合物中提取乐器方面的音频效果表示。我们的方法利用了对比学习框架,并引入了一个“提取器”机制,当提供乐器查询(音频或文本)时,将混合级音频效果嵌入转换为乐器级音频效果嵌入。我们在检索和音频效果参数匹配任务中评估了我们的模型,测试了它在各种乐器上的性能。结果表明,Fx-Encoder++在混合级别上优于以前的方法,并显示出一种新颖的能力,可以从乐器中提取效果表示,解决智能音乐制作系统中的关键能力差距。
摘要:General-purpose audio representations have proven effective across diverse music information retrieval applications, yet their utility in intelligent music production remains limited by insufficient understanding of audio effects (Fx). Although previous approaches have emphasized audio effects analysis at the mixture level, this focus falls short for tasks demanding instrument-wise audio effects understanding, such as automatic mixing. In this work, we present Fx-Encoder++, a novel model designed to extract instrument-wise audio effects representations from music mixtures. Our approach leverages a contrastive learning framework and introduces an "extractor" mechanism that, when provided with instrument queries (audio or text), transforms mixture-level audio effects embeddings into instrument-wise audio effects embeddings. We evaluated our model across retrieval and audio effects parameter matching tasks, testing its performance across a diverse range of instruments. The results demonstrate that Fx-Encoder++ outperforms previous approaches at mixture level and show a novel ability to extract effects representation instrument-wise, addressing a critical capability gap in intelligent music production systems.
【15】Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
链接:https://arxiv.org/abs/2507.02176
备注:Accepted at SSW13 - Interspeech 2025 Speech Synthesis Workshop
摘要:语音身份建模是具有挑战性的,由于其多方面的性质。在生成语音系统中,身份通常使用自动说话人验证(ASV)嵌入来评估,其设计用于区分而不是表征身份。本文探讨了哪些方面的声音被捕获在这样的表示。我们发现,广泛使用的ASV嵌入主要集中在静态特征,如音色和音高范围,而忽略了动态元素,如节奏。我们还确定了混杂因素,损害扬声器相似性测量,并提出缓解策略。为了解决这些差距,我们提出了U3D,一个评估扬声器的动态节奏模式的指标。这项工作有助于不断改善的语音克隆系统的背景下,评估扬声器身份的一致性的持续挑战。我们公开发布代码。
摘要:Modeling voice identity is challenging due to its multifaceted nature. In generative speech systems, identity is often assessed using automatic speaker verification (ASV) embeddings, designed for discrimination rather than characterizing identity. This paper investigates which aspects of a voice are captured in such representations. We find that widely used ASV embeddings focus mainly on static features like timbre and pitch range, while neglecting dynamic elements such as rhythm. We also identify confounding factors that compromise speaker similarity measurements and suggest mitigation strategies. To address these gaps, we propose U3D, a metric that evaluates speakers' dynamic rhythm patterns. This work contributes to the ongoing challenge of assessing speaker identity consistency in the context of ever-better voice cloning systems. We publicly release our code.
【16】Parametric Neural Amp Modeling with Active Learning
链接:https://arxiv.org/abs/2507.02109
备注:Accepted at ISMIR 2025 as Late-Breaking Demo (LBD)
摘要:我们介绍PANAMA,一个主动学习框架,用于使用WaveNet架构训练端到端参数吉他放大器模型。使用\model,可以通过记录由主动学习策略确定的样本来创建虚拟放大器,以使用最少量的数据点(即,放大器旋钮设置)。我们表明,基于梯度的优化算法可以用来确定最佳的数据点进行采样,该方法有助于在有限数量的样本。
摘要:We introduce PANAMA, an active learning framework for the training of end-to-end parametric guitar amp models using a WaveNet-like architecture. With \model, one can create a virtual amp by recording samples that are determined by an active learning strategy to use a minimum amount of datapoints (i.e., amp knob settings). We show that gradient-based optimization algorithms can be used to determine the optimal datapoints to sample, and that the approach helps under a constrained number of samples.
机器翻译由腾讯交互翻译提供,仅供参考
