微信公众号:arXiv_Daily
cs.SD语音
链接:http://arxiv.org/pdf/2507.08768v1
备注:Update with Acknowledgements of ICNSLP 2025 paper
摘要:在这项研究中,我们利用联合国教科文组织独特的20世纪中期无线电录音收集,以探讨现代现成的语言识别(LID)和说话人识别(SR)方法的鲁棒性,特别是在多语言说话者和跨年龄录音的影响方面。我们的研究结果表明,LID系统,如耳语,越来越善于处理第二语言和口音的讲话。然而,说话者嵌入仍然是语音处理管道中的一个脆弱组件,容易出现与通道、年龄和语言相关的偏见。如果档案馆的目标是采用SR方法为发言者索引,则需要克服的问题。
摘要:In this study, we leverage a unique UNESCO collection of mid-20th century radio recordings to probe the robustness of modern off-the-shelf language identification (LID) and speaker recognition (SR) methods, especially with respect to the impact of multilingual speakers and cross-age recordings. Our findings suggest that LID systems, such as Whisper, are increasingly adept at handling second-language and accented speech. However, speaker embeddings remain a fragile component of speech processing pipelines that is prone to biases related to the channel, age, and language. Issues which will need to be overcome should archives aim to employ SR methods for speaker indexing.
【2】Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
备注:Accepted at ICCV Workshop - Authenticity & Provenance in the age of Generative AI
摘要:生成式人工智能的最新进展使得语音深度伪造的创建变得广泛,这对数字信任构成了严重挑战。为了解决这个问题,已经提出了各种语音深度伪造检测策略,包括感兴趣的人(POI)方法,该方法专注于通过建模和分析特定个体独特的声音特征来识别其模仿。尽管现有方法性能出色,但其粒度有限且缺乏可解释性。在这项工作中,我们提出了一种基于POI的语音deepfake检测方法,该方法在音素级别上运行。我们的方法将参考音频分解为音素,以构建详细的扬声器配置文件。在推理中,将测试样本中的音素单独与此配置文件进行比较,从而实现对合成工件的细粒度检测。所提出的方法实现了可比的准确性,传统的方法,同时提供卓越的鲁棒性和可解释性,在多媒体取证的关键方面。通过专注于音素分析,这项工作探索了一个可解释的、以说话者为中心的deepfake检测的新方向。
摘要:Recent advances in generative AI have made the creation of speech deepfakes widely accessible, posing serious challenges to digital trust. To counter this, various speech deepfake detection strategies have been proposed, including Person-of-Interest (POI) approaches, which focus on identifying impersonations of specific individuals by modeling and analyzing their unique vocal traits. Despite their excellent performance, the existing methods offer limited granularity and lack interpretability. In this work, we propose a POI-based speech deepfake detection method that operates at the phoneme level. Our approach decomposes reference audio into phonemes to construct a detailed speaker profile. In inference, phonemes from a test sample are individually compared against this profile, enabling fine-grained detection of synthetic artifacts. The proposed method achieves comparable accuracy to traditional approaches while offering superior robustness and interpretability, key aspects in multimedia forensics. By focusing on phoneme analysis, this work explores a novel direction for explainable, speaker-centric deepfake detection.
备注:ACL 2025 Findings
摘要:端到端大型语音语言模型~( textbf{LSLMs})在响应延迟和语音理解能力方面表现出强大的潜力,展示了跨语音理解任务的一般智能。然而,由于缺乏数据集和严重偏差的训练任务,遵循语音指令的能力尚未完全实现。利用丰富的ASR数据集,以前的方法使用大语言模型~( textbf{LLM})来继续语音的语言信息以构建语音指令数据集。然而,由于LLM生成的结果和真实的人类反应之间的差距,延续方法进一步放大了这些缺点。鉴于人类收集和注释语音指令数据集的成本很高,使用语音合成来构建大规模语音指令数据集已成为一种平衡且鲁棒的替代方案。虽然现代的文本到语音(Text-To-Speech~( textbf{TTS}))模型已经达到了接近人类水平的合成质量,但是由于TTS模型中训练数据分布的限制,将分布外的文本指令适当地转换为语音是具有挑战性的。为了解决这个问题,我们提出了一个多LLM知识融合的查询重写框架,采用多个代理来注释和验证合成语音,从而可以在不依赖于人工注释的情况下构建高质量的语音指令数据集。实验表明,该方法通过zero-shot重写,将文本指令转换为更适合语音合成TTS模型的分布,使数据可用性从72%提高到93%.它还展示了重写任务,需要复杂的知识和上下文相关的能力的独特优势。
摘要:End-to-end Large Speech Language Models~( textbf{LSLMs}) demonstrate strong potential in response latency and speech comprehension capabilities, showcasing general intelligence across speech understanding tasks. However, the ability to follow speech instructions has not been fully realized due to the lack of datasets and heavily biased training tasks. Leveraging the rich ASR datasets, previous approaches have used Large Language Models~( textbf{LLMs}) to continue the linguistic information of speech to construct speech instruction datasets. Yet, due to the gap between LLM-generated results and real human responses, the continuation methods further amplify these shortcomings. Given the high costs of collecting and annotating speech instruction datasets by humans, using speech synthesis to construct large-scale speech instruction datasets has become a balanced and robust alternative. Although modern Text-To-Speech~( textbf{TTS}) models have achieved near-human-level synthesis quality, it is challenging to appropriately convert out-of-distribution text instruction to speech due to the limitations of the training data distribution in TTS models. To address this issue, we propose a query rewriting framework with multi-LLM knowledge fusion, employing multiple agents to annotate and validate the synthesized speech, making it possible to construct high-quality speech instruction datasets without relying on human annotation. Experiments show that this method can transform text instructions into distributions more suitable for TTS models for speech synthesis through zero-shot rewriting, increasing data usability from 72 % to 93 %. It also demonstrates unique advantages in rewriting tasks that require complex knowledge and context-related abilities.
备注:Accepted at ACM MM 2025
摘要:随着生成模型的最新进展,文本到音频(T2 A)生成已经取得了令人鼓舞的成果。然而,由于时间对准的音频-文本对的质量和数量有限,现有的T2 A方法难以处理包含精确定时控制的复杂文本提示,例如,“猫头鹰以2.4s-5.2s的速度鸣叫”。最近的工作已经探索了数据增强技术或引入定时条件作为模型输入,以实现定时调节的10秒T2 A生成,而它们的合成质量仍然有限。在这项工作中,我们提出了一种新的无训练定时控制的T2 A框架FreeAudio,首次尝试实现定时控制的长形式T2 A生成,例如,“猫头鹰在2.4s-5.2s鸣叫,蟋蟀在0 s-24 s鸣叫”。具体来说,我们首先采用LLM计划非重叠的时间窗口,并根据输入的文本和时间提示,用精炼的自然语言描述重新描述每个时间窗口。然后我们介绍:1)解耦和聚合注意力控制,用于精确的定时控制; 2)上下文潜在成分,用于局部平滑性和参考指导,用于全局一致性。大量的实验表明:1)FreeAudio在无需训练的方法中实现了最先进的定时条件T2 A合成质量,可与领先的基于训练的方法相媲美; 2)FreeAudio与基于训练的Stable Audio具有可比的长格式生成质量,为定时控制的长格式T2 A合成铺平了道路。演示示例可在以下网址获得:https: freeaudio.github.io FreeAudio摘要:Text-to-audio (T2A) generation has achieved promising results with the recent advances in generative models. However, because of the limited quality and quantity of temporally-aligned audio-text pairs, existing T2A methods struggle to handle the complex text prompts that contain precise timing control, e.g., "owl hooted at 2.4s-5.2s". Recent works have explored data augmentation techniques or introduced timing conditions as model inputs to enable timing-conditioned 10-second T2A generation, while their synthesis quality is still limited. In this work, we propose a novel training-free timing-controlled T2A framework, FreeAudio, making the first attempt to enable timing-controlled long-form T2A generation, e.g., "owl hooted at 2.4s-5.2s and crickets chirping at 0s-24s". Specifically, we first employ an LLM to plan non-overlapping time windows and recaption each with a refined natural language description, based on the input text and timing prompts. Then we introduce: 1) Decoupling and Aggregating Attention Control for precise timing control; 2) Contextual Latent Composition for local smoothness and Reference Guidance for global consistency. Extensive experiments show that: 1) FreeAudio achieves state-of-the-art timing-conditioned T2A synthesis quality among training-free methods and is comparable to leading training-based methods; 2) FreeAudio demonstrates comparable long-form generation quality with training-based Stable Audio and paves the way for timing-controlled long-form T2A synthesis. Demo samples are available at: https: freeaudio.github.io FreeAudio
备注:Accepted by ISMIR 2025
摘要:从乐谱生成富有表现力的音频表演需要模型来捕获乐器声学和人类解释。传统的音乐表演合成管道遵循两个阶段的方法,首先从乐谱中生成表现性的表演片段,然后将片段合成为音频。然而,合成模型往往难以概括不同的音乐来源,音乐风格和录音环境。为了解决这些挑战,我们提出了MIDI-VALLE,一个神经编解码器语言模型改编自VALLE框架,这是最初设计的zero-shot个性化的文本到语音(TTS)的合成。对于性能MIDI到音频的合成,我们改进了架构,以参考音频性能及其相应的参数为条件。与之前基于TTS的系统依赖于钢琴卷不同,MIDI-VALLE将声音和音频编码为离散标记,从而促进钢琴演奏的更一致和更强大的建模。此外,该模型的泛化能力通过在广泛而多样化的钢琴演奏数据集上进行训练而得到增强。评估结果表明,MIDI-VALLE的性能明显优于最先进的基线,在ATEPP和Maestro数据集上实现了超过75%的Frechet音频距离降低。在听力测试中,MIDI-VALLE获得了202票,而基线为58票,这表明合成质量得到了提高,并且在不同的性能测试输入中得到了推广。摘要:Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pipelines follow a two-stage approach, first generating expressive performance MIDI from a score, then synthesising the MIDI into audio. However, the synthesis models often struggle to generalise across diverse MIDI sources, musical styles, and recording environments. To address these challenges, we propose MIDI-VALLE, a neural codec language model adapted from the VALLE framework, which was originally designed for zero-shot personalised text-to-speech (TTS) synthesis. For performance MIDI-to-audio synthesis, we improve the architecture to condition on a reference audio performance and its corresponding MIDI. Unlike previous TTS-based systems that rely on piano rolls, MIDI-VALLE encodes both MIDI and audio as discrete tokens, facilitating a more consistent and robust modelling of piano performances. Furthermore, the model's generalisation ability is enhanced by training on an extensive and diverse piano performance dataset. Evaluation results show that MIDI-VALLE significantly outperforms a state-of-the-art baseline, achieving over 75% lower Frechet Audio Distance on the ATEPP and Maestro datasets. In the listening test, MIDI-VALLE received 202 votes compared to 58 for the baseline, demonstrating improved synthesis quality and generalisation across diverse performance MIDI inputs.
链接:http://arxiv.org/pdf/2507.08412v1
摘要:环境声音记录通常包含可理解的语音,这引起了对隐私的担忧,限制了数据的分析、共享和再利用。在本文中,我们介绍了一种方法,使语音难以理解,同时保持声学场景的完整性,和整体的音频质量。我们的方法包括反转波形段来扭曲语音内容。这一过程通过语音活动检测和语音分离管道得到增强,这允许更精确地定位语音。 为了证明所提出的方法的有效性,我们考虑了一个由三部分组成的评估协议,该协议评估:1)使用单词错误率(WER)的语音可懂度,2)使用来自广泛使用的预训练模型的声源分类准确度下降(SCAD)的声源可检测性,以及3)使用Fr 'echet音频距离(FAD)的音频质量,用我们的参考数据集计算,该数据集包含未改变的语音。实验结果表明,我们的方法实现了令人满意的语音清晰度降低(97.9%WER),声源可检测性的最小退化(2.7%SCAD),和高感知质量(FAD为1.40)。消融研究进一步突出了管道每个组件的贡献。我们还表明,将随机拼接到我们的语音内容隐私执法方法可以提高算法的鲁棒性,试图恢复干净的语音,在音频质量的轻微成本。
摘要:Environmental sound recordings often contain intelligible speech, raising privacy concerns that limit analysis, sharing and reuse of data. In this paper, we introduce a method that renders speech unintelligible while preserving both the integrity of the acoustic scene, and the overall audio quality. Our approach involves reversing waveform segments to distort speech content. This process is enhanced through a voice activity detection and speech separation pipeline, which allows for more precise targeting of speech. In order to demonstrate the effectivness of the proposed approach, we consider a three-part evaluation protocol that assesses: 1) speech intelligibility using Word Error Rate (WER), 2) sound sources detectability using Sound source Classification Accuracy-Drop (SCAD) from a widely used pre-trained model, and 3) audio quality using the Fr 'echet Audio Distance (FAD), computed with our reference dataset that contains unaltered speech. Experiments on this simulated evaluation dataset, which consists of linear mixtures of speech and environmental sound scenes, show that our method achieves satisfactory speech intelligibility reduction (97.9% WER), minimal degradation of the sound sources detectability (2.7% SCAD), and high perceptual quality (FAD of 1.40). An ablation study further highlights the contribution of each component of the pipeline. We also show that incorporating random splicing to our speech content privacy enforcement method can enhance the algorithm's robustness to attempt to recover the clean speech, at a slight cost of audio quality.
摘要:音频修复是指在损坏的音频记录中重建丢失的片段的任务。虽然先前的方法-包括基于波形和频谱的扩散模型-已经显示出对短间隙的有希望的结果,但是当间隙超过100毫秒(ms)时,它们的质量通常会降低。在这项工作中,我们介绍了一种新的基于离散扩散建模的修复方法,该方法对由预训练的音频标记器产生的标记化音频表示进行操作。我们的方法直接在离散潜在空间中对生成过程进行建模,从而实现丢失音频的稳定和语义一致的重建。我们评估的方法MusicNet数据集上使用客观和感知指标的间隙持续时间长达300毫秒。我们进一步评估了我们的MTG数据集上的方法,将间隙持续时间延长到500毫秒。实验结果表明,我们的方法实现了竞争力或优越的性能相比,现有的基线,特别是较长的差距,提供了一个强大的解决方案,恢复退化的音乐录音。我们提出的方法的音频示例可以在https: iftach21.github.io 上找到
摘要:Audio inpainting refers to the task of reconstructing missing segments in corrupted audio recordings. While prior approaches-including waveform and spectrogram-based diffusion models-have shown promising results for short gaps, they often degrade in quality when gaps exceed 100 milliseconds (ms). In this work, we introduce a novel inpainting method based on discrete diffusion modeling, which operates over tokenized audio representations produced by a pre-trained audio tokenizer. Our approach models the generative process directly in the discrete latent space, enabling stable and semantically coherent reconstruction of missing audio. We evaluate the method on the MusicNet dataset using both objective and perceptual metrics across gap durations up to 300 ms. We further evaluated our approach on the MTG dataset, extending the gap duration to 500 ms. Experimental results demonstrate that our method achieves competitive or superior performance compared to existing baselines, particularly for longer gaps, offering a robust solution for restoring degraded musical recordings. Audio examples of our proposed method can be found at https: iftach21.github.io摘要:高质量数据集的构建是现代文语转换(TTS)系统的基石。然而,可用数据规模的不断扩大带来了重大挑战,包括存储限制。针对这些问题,本文提出了一种基于主动学习的TTS语料库构建方法。与传统的前馈和模型无关的语料库构建方法不同,我们的方法在数据收集和模型训练之间反复交替,从而专注于获取对模型改进更有用的数据。这种方法能够构建数据高效的语料库。实验结果表明,使用我们的方法构建的语料库,使更高质量的语音合成比语料库的相同大小。
摘要:The construction of high-quality datasets is a cornerstone of modern text-to-speech (TTS) systems. However, the increasing scale of available data poses significant challenges, including storage constraints. To address these issues, we propose a TTS corpus construction method based on active learning. Unlike traditional feed-forward and model-agnostic corpus construction approaches, our method iteratively alternates between data collection and model training, thereby focusing on acquiring data that is more informative for model improvement. This approach enables the construction of a data-efficient corpus. Experimental results demonstrate that the corpus constructed using our method enables higher-quality speech synthesis than corpora of the same size.
备注:Working note submitted to CLEF 2025 under the LifeCLEF lab
摘要:BirdCLEF+ 2025挑战要求在90分钟的严格CPU推理期限内对206个物种进行分类,包括鸟类、哺乳动物、昆虫和两栖动物,这使得许多最先进的深度学习方法变得不切实际。为了解决这一限制,DS@GT BirdCLEF团队探索了两种策略。首先,我们通过优化来自生物声学模型动物园的预训练模型来建立有竞争力的基线。使用TFLite,我们为Perch模型实现了近10倍的推理加速,使其能够在大约16分钟内运行,并在赛后公共排行榜上获得0.729的最终ROC-AUC分数,在私人排行榜上获得0.711。动物园的最佳模型是BirdSetEfficientNetB 1,公开得分为0.810,私人得分为0.778。其次,我们介绍了一种新的,轻量级的管道命名为Spectrogram令牌Skip-Gram(STSG),将生物声学作为一个序列建模任务。该方法通过使用Faiss K-means对Mel谱图进行聚类,将音频转换为离散的“谱图令牌”,然后使用Word 2 Vec跳图模型以无监督的方式学习这些令牌的高质量上下文嵌入。对于分类,5秒窗口内的嵌入被平均并传递到线性模型。对于700分钟的测试集,预计推理时间为6分钟,STSG方法的最终ROC-AUC公共得分为0.559,私有得分为0.520,证明了使用静态嵌入进行生物声学分类的快速标记化方法的可行性。本文的支持代码可以在https: github.com dsgt-arc birdclef-2025上找到。
摘要:The BirdCLEF+ 2025 challenge requires classifying 206 species, including birds, mammals, insects, and amphibians, from soundscape recordings under a strict 90-minute CPU-only inference deadline, making many state-of-the-art deep learning approaches impractical. To address this constraint, the DS@GT BirdCLEF team explored two strategies. First, we establish competitive baselines by optimizing pre-trained models from the Bioacoustics Model Zoo for CPU inference. Using TFLite, we achieved a nearly 10x inference speedup for the Perch model, enabling it to run in approximately 16 minutes and achieve a final ROC-AUC score of 0.729 on the public leaderboard post-competition and 0.711 on the private leaderboard. The best model from the zoo was BirdSetEfficientNetB1, with a public score of 0.810 and a private score of 0.778. Second, we introduce a novel, lightweight pipeline named Spectrogram Token Skip-Gram (STSG) that treats bioacoustics as a sequence modeling task. This method converts audio into discrete "spectrogram tokens" by clustering Mel-spectrograms using Faiss K-means and then learns high-quality contextual embeddings for these tokens in an unsupervised manner with a Word2Vec skip-gram model. For classification, embeddings within a 5-second window are averaged and passed to a linear model. With a projected inference time of 6 minutes for a 700-minute test set, the STSG approach achieved a final ROC-AUC public score of 0.559 and a private score of 0.520, demonstrating the viability of fast tokenization approaches with static embeddings for bioacoustic classification. Supporting code for this paper can be found at https: github.com dsgt-arc birdclef-2025.
备注:Code, Datasets and Models: this https URL
摘要:Audio Flamingo 3(AF 3)是一个完全开放的最先进的(SOTA)大型音频语言模型,可以在语音,声音和音乐中推进推理和理解。AF 3介绍:(i)AF-Whisper,一个统一的音频编码器,使用一种新的策略进行训练,用于跨语音、声音和音乐的所有3种模式的联合表示学习;(ii)灵活的按需思维,允许模型在回答之前进行思维链类型的推理;(iii)多回合、多音频聊天;(iv)长音频理解和推理(包括演讲)最多10分钟;及(v)语音对语音互动。为了实现这些功能,我们提出了几个使用新策略策划的大规模训练数据集,包括AudioSkills-XL,LongAudio-XL,AF-Think和AF-Chat,并使用一种新的基于五阶段训练的训练策略训练AF 3。AF 3仅在开源音频数据上训练,在超过20个(长)音频理解和推理基准上实现了新的SOTA结果,超过了在更大数据集上训练的开放权重和闭源模型。
摘要:We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning across all 3 modalities of speech, sound, and music; (ii) flexible, on-demand thinking, allowing the model to do chain-of-thought-type reasoning before answering; (iii) multi-turn, multi-audio chat; (iv) long audio understanding and reasoning (including speech) up to 10 minutes; and (v) voice-to-voice interaction. To enable these capabilities, we propose several large-scale training datasets curated using novel strategies, including AudioSkills-XL, LongAudio-XL, AF-Think, and AF-Chat, and train AF3 with a novel five-stage curriculum-based training strategy. Trained on only open-source audio data, AF3 achieves new SOTA results on over 20+ (long) audio understanding and reasoning benchmarks, surpassing both open-weight and closed-source models trained on much larger datasets.
摘要:房间脉冲响应估计对于诸如语音去混响之类的任务是必不可少的,这提高了自动语音识别。大多数现有方法依赖于统计信号处理或旨在复制信号处理原理的深度神经网络。然而,结合统计和物理建模RIR估计仍然在很大程度上未被探索。本文提出了一种通过理论基础模型将这两个方面整合在一起的新方法。RIR被分解为可解释的参数:通过频率相关指数衰减(例如,建模壁吸收)和自回归滤波器(例如,建模麦克风响应)过滤的高斯白噪声。变分自由能成本函数使实际的参数估计。作为一个概念的证明,我们表明,干和混响的语音信号,所提出的方法优于经典的去卷积在嘈杂的环境中,通过客观的度量验证。
摘要:Room impulse response estimation is essential for tasks like speech dereverberation, which improves automatic speech recognition. Most existing methods rely on either statistical signal processing or deep neural networks designed to replicate signal processing principles. However, combining statistical and physical modeling for RIR estimation remains largely unexplored. This paper proposes a novel approach integrating both aspects through a theoretically grounded model. The RIR is decomposed into interpretable parameters: white Gaussian noise filtered by a frequency-dependent exponential decay (e.g. modeling wall absorption) and an autoregressive filter (e.g. modeling microphone response). A variational free-energy cost function enables practical parameter estimation. As a proof of concept, we show that given dry and reverberant speech signals, the proposed method outperforms classical deconvolution in noisy environments, as validated by objective metrics.
摘要:基于文本到语音的模型允许用户通过自然语言指令来控制语音的不同方面,例如说话速率和感知性别。虽然用户友好,这样的方法是一方面约束:控制仅限于在训练过程中暴露于模型的声学特征,另一方面过于灵活:相同的输入会产生无法控制的变化,这些变化反映在语料库统计中。 我们研究了一种新的微调制度,以解决这两个问题,在同一时间,利用模型的不可控的方差。通过对数千个合成样本的主成分分析,我们确定了占输出方差比例最高的潜在特征,并将它们作为新标签进行二次微调。我们评估所提出的方法在两个模型上训练的表达冰岛语语音语料库,一个与情感的披露和一个没有。在没有情绪披露的模型的情况下,该方法产生连续和离散的功能,提高整体的可控性的模型。
摘要:A Prompt-based Text-To-Speech model allows a user to control different aspects of speech, such as speaking rate and perceived gender, through natural language instruction. Although user-friendly, such approaches are on one hand constrained: control is limited to acoustic features exposed to the model during training, and too flexible on the other: the same inputs yields uncontrollable variation that are reflected in the corpus statistics. We investigate a novel fine-tuning regime to address both of these issues at the same time by exploiting the uncontrollable variance of the model. Through principal component analysis of thousands of synthesised samples, we determine latent features that account for the highest proportion of the output variance and incorporate them as new labels for secondary fine-tuning. We evaluate the proposed methods on two models trained on an expressive Icelandic speech corpus, one with emotional disclosure and one without. In the case of the model without emotional disclosure, the method yields both continuous and discrete features that improve overall controllability of the model.
备注:Submitted to APSIPA ASC 2025
摘要:自动说话人确认(ASV)系统经常受到欺骗攻击。最近的基于transformer的模型通过学习强特征表示来提高反欺骗性能。然而,这些模型通常需要很高的计算能力。为了解决这个问题,我们引入了RawTFNet,这是一个为音频信号设计的轻量级CNN模型。RawTFNet沿着时间和频率维度分离特征处理,这有助于捕获合成语音的细粒度细节。我们在ASVspoof 2021 LA和DF评估数据集上测试了RawTFNet。结果表明,RawTFNet达到了与最先进模型相当的性能,同时使用更少的计算资源。代码和模型将公开提供。
摘要:Automatic speaker verification (ASV) systems are often affected by spoofing attacks. Recent transformer-based models have improved anti-spoofing performance by learning strong feature representations. However, these models usually need high computing power. To address this, we introduce RawTFNet, a lightweight CNN model designed for audio signals. The RawTFNet separates feature processing along time and frequency dimensions, which helps to capture the fine-grained details of synthetic speech. We tested RawTFNet on the ASVspoof 2021 LA and DF evaluation datasets. The results show that RawTFNet reaches comparable performance to that of the state-of-the-art models, while also using fewer computing resources. The code and models will be made publicly available.
备注:14 pages, 9 figures, submitted to IEEEACM Transactions on Audio, Speech, and Language Processing
摘要:房间脉冲响应(RIR)准确地表征室内环境的声学特性,并在诸如增强现实(AR)和虚拟现实(VR)中的语音增强、语音识别和音频渲染等应用中发挥关键作用。现有的盲估计方法难以达到实际的精度。为了克服这一挑战,我们提出了动态音频室声学合成(DARAS)模型,这是一种新型的深度学习框架,专门用于单声道混响语音信号的盲RIR估计。首先,专用的深度音频编码器有效地提取相关的非线性潜在空间特征。第二,基于Mamba的自监督盲房间参数估计(MASS-BRPE)模块,利用高效的Mamba状态空间模型(SSM),准确地估计关键房间声学参数和特征。第三,该系统集成了混合路径交叉注意特征融合模块,增强了音频和房间声学特征之间的深度融合。最后,我们提出的动态声学调谐(DAT)解码器自适应分段早期反射和后期混响,以提高合成RIR的真实感。实验结果,包括基于MUSHRA的主观听力研究,表明,DARAS大大优于现有的基线模型,提供了一个强大的和有效的解决方案,在现实世界的声学环境中的实际盲RIR估计。
摘要:Room Impulse Responses (RIRs) accurately characterize acoustic properties of indoor environments and play a crucial role in applications such as speech enhancement, speech recognition, and audio rendering in augmented reality (AR) and virtual reality (VR). Existing blind estimation methods struggle to achieve practical accuracy. To overcome this challenge, we propose the dynamic audio-room acoustic synthesis (DARAS) model, a novel deep learning framework that is explicitly designed for blind RIR estimation from monaural reverberant speech signals. First, a dedicated deep audio encoder effectively extracts relevant nonlinear latent space features. Second, the Mamba-based self-supervised blind room parameter estimation (MASS-BRPE) module, utilizing the efficient Mamba state space model (SSM), accurately estimates key room acoustic parameters and features. Third, the system incorporates a hybrid-path cross-attention feature fusion module, enhancing deep integration between audio and room acoustic features. Finally, our proposed dynamic acoustic tuning (DAT) decoder adaptively segments early reflections and late reverberation to improve the realism of synthesized RIRs. Experimental results, including a MUSHRA-based subjective listening study, demonstrate that DARAS substantially outperforms existing baseline models, providing a robust and effective solution for practical blind RIR estimation in real-world acoustic environments.
备注:Submitted to APSIPA ASC 2025
摘要:自动说话人确认(ASV)系统经常受到欺骗攻击。最近的基于transformer的模型通过学习强特征表示来提高反欺骗性能。然而,这些模型通常需要很高的计算能力。为了解决这个问题,我们引入了RawTFNet,这是一个为音频信号设计的轻量级CNN模型。RawTFNet沿着时间和频率维度分离特征处理,这有助于捕获合成语音的细粒度细节。我们在ASVspoof 2021 LA和DF评估数据集上测试了RawTFNet。结果表明,RawTFNet达到了与最先进模型相当的性能,同时使用更少的计算资源。代码和模型将公开提供。
摘要:Automatic speaker verification (ASV) systems are often affected by spoofing attacks. Recent transformer-based models have improved anti-spoofing performance by learning strong feature representations. However, these models usually need high computing power. To address this, we introduce RawTFNet, a lightweight CNN model designed for audio signals. The RawTFNet separates feature processing along time and frequency dimensions, which helps to capture the fine-grained details of synthetic speech. We tested RawTFNet on the ASVspoof 2021 LA and DF evaluation datasets. The results show that RawTFNet reaches comparable performance to that of the state-of-the-art models, while also using fewer computing resources. The code and models will be made publicly available.
备注:14 pages, 9 figures, submitted to IEEEACM Transactions on Audio, Speech, and Language Processing
摘要:房间脉冲响应(RIR)准确地表征室内环境的声学特性,并在诸如增强现实(AR)和虚拟现实(VR)中的语音增强、语音识别和音频渲染等应用中发挥关键作用。现有的盲估计方法难以达到实际的精度。为了克服这一挑战,我们提出了动态音频室声学合成(DARAS)模型,这是一种新型的深度学习框架,专门用于单声道混响语音信号的盲RIR估计。首先,专用的深度音频编码器有效地提取相关的非线性潜在空间特征。第二,基于Mamba的自监督盲房间参数估计(MASS-BRPE)模块,利用高效的Mamba状态空间模型(SSM),准确地估计关键房间声学参数和特征。第三,该系统集成了混合路径交叉注意特征融合模块,增强了音频和房间声学特征之间的深度融合。最后,我们提出的动态声学调谐(DAT)解码器自适应分段早期反射和后期混响,以提高合成RIR的真实感。实验结果,包括基于MUSHRA的主观听力研究,表明,DARAS大大优于现有的基线模型,提供了一个强大的和有效的解决方案,在现实世界的声学环境中的实际盲RIR估计。
摘要:Room Impulse Responses (RIRs) accurately characterize acoustic properties of indoor environments and play a crucial role in applications such as speech enhancement, speech recognition, and audio rendering in augmented reality (AR) and virtual reality (VR). Existing blind estimation methods struggle to achieve practical accuracy. To overcome this challenge, we propose the dynamic audio-room acoustic synthesis (DARAS) model, a novel deep learning framework that is explicitly designed for blind RIR estimation from monaural reverberant speech signals. First, a dedicated deep audio encoder effectively extracts relevant nonlinear latent space features. Second, the Mamba-based self-supervised blind room parameter estimation (MASS-BRPE) module, utilizing the efficient Mamba state space model (SSM), accurately estimates key room acoustic parameters and features. Third, the system incorporates a hybrid-path cross-attention feature fusion module, enhancing deep integration between audio and room acoustic features. Finally, our proposed dynamic acoustic tuning (DAT) decoder adaptively segments early reflections and late reverberation to improve the realism of synthesized RIRs. Experimental results, including a MUSHRA-based subjective listening study, demonstrate that DARAS substantially outperforms existing baseline models, providing a robust and effective solution for practical blind RIR estimation in real-world acoustic environments.
链接:http://arxiv.org/pdf/2507.08768v1
备注:Update with Acknowledgements of ICNSLP 2025 paper
摘要:在这项研究中,我们利用联合国教科文组织独特的20世纪中期无线电录音收集,以探讨现代现成的语言识别(LID)和说话人识别(SR)方法的鲁棒性,特别是在多语言说话者和跨年龄录音的影响方面。我们的研究结果表明,LID系统,如耳语,越来越善于处理第二语言和口音的讲话。然而,说话者嵌入仍然是语音处理管道中的一个脆弱组件,容易出现与通道、年龄和语言相关的偏见。如果档案馆的目标是采用SR方法为发言者索引,则需要克服的问题。
摘要:In this study, we leverage a unique UNESCO collection of mid-20th century radio recordings to probe the robustness of modern off-the-shelf language identification (LID) and speaker recognition (SR) methods, especially with respect to the impact of multilingual speakers and cross-age recordings. Our findings suggest that LID systems, such as Whisper, are increasingly adept at handling second-language and accented speech. However, speaker embeddings remain a fragile component of speech processing pipelines that is prone to biases related to the channel, age, and language. Issues which will need to be overcome should archives aim to employ SR methods for speaker indexing.
【4】Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
备注:Accepted at ICCV Workshop - Authenticity & Provenance in the age of Generative AI
摘要:生成式人工智能的最新进展使得语音深度伪造的创建变得广泛,这对数字信任构成了严重挑战。为了解决这个问题,已经提出了各种语音深度伪造检测策略,包括感兴趣的人(POI)方法,该方法专注于通过建模和分析特定个体独特的声音特征来识别其模仿。尽管现有方法性能出色,但其粒度有限且缺乏可解释性。在这项工作中,我们提出了一种基于POI的语音deepfake检测方法,该方法在音素级别上运行。我们的方法将参考音频分解为音素,以构建详细的扬声器配置文件。在推理中,将测试样本中的音素单独与此配置文件进行比较,从而实现对合成工件的细粒度检测。所提出的方法实现了可比的准确性,传统的方法,同时提供卓越的鲁棒性和可解释性,在多媒体取证的关键方面。通过专注于音素分析,这项工作探索了一个可解释的、以说话者为中心的deepfake检测的新方向。
摘要:Recent advances in generative AI have made the creation of speech deepfakes widely accessible, posing serious challenges to digital trust. To counter this, various speech deepfake detection strategies have been proposed, including Person-of-Interest (POI) approaches, which focus on identifying impersonations of specific individuals by modeling and analyzing their unique vocal traits. Despite their excellent performance, the existing methods offer limited granularity and lack interpretability. In this work, we propose a POI-based speech deepfake detection method that operates at the phoneme level. Our approach decomposes reference audio into phonemes to construct a detailed speaker profile. In inference, phonemes from a test sample are individually compared against this profile, enabling fine-grained detection of synthetic artifacts. The proposed method achieves comparable accuracy to traditional approaches while offering superior robustness and interpretability, key aspects in multimedia forensics. By focusing on phoneme analysis, this work explores a novel direction for explainable, speaker-centric deepfake detection.
备注:ACL 2025 Findings
摘要:端到端大型语音语言模型~( textbf{LSLMs})在响应延迟和语音理解能力方面表现出强大的潜力,展示了跨语音理解任务的一般智能。然而,由于缺乏数据集和严重偏差的训练任务,遵循语音指令的能力尚未完全实现。利用丰富的ASR数据集,以前的方法使用大语言模型~( textbf{LLM})来继续语音的语言信息以构建语音指令数据集。然而,由于LLM生成的结果和真实的人类反应之间的差距,延续方法进一步放大了这些缺点。鉴于人类收集和注释语音指令数据集的成本很高,使用语音合成来构建大规模语音指令数据集已成为一种平衡且鲁棒的替代方案。虽然现代的文本到语音(Text-To-Speech~( textbf{TTS}))模型已经达到了接近人类水平的合成质量,但是由于TTS模型中训练数据分布的限制,将分布外的文本指令适当地转换为语音是具有挑战性的。为了解决这个问题,我们提出了一个多LLM知识融合的查询重写框架,采用多个代理来注释和验证合成语音,从而可以在不依赖于人工注释的情况下构建高质量的语音指令数据集。实验表明,该方法通过zero-shot重写,将文本指令转换为更适合语音合成TTS模型的分布,使数据可用性从72%提高到93%.它还展示了重写任务,需要复杂的知识和上下文相关的能力的独特优势。
摘要:End-to-end Large Speech Language Models~( textbf{LSLMs}) demonstrate strong potential in response latency and speech comprehension capabilities, showcasing general intelligence across speech understanding tasks. However, the ability to follow speech instructions has not been fully realized due to the lack of datasets and heavily biased training tasks. Leveraging the rich ASR datasets, previous approaches have used Large Language Models~( textbf{LLMs}) to continue the linguistic information of speech to construct speech instruction datasets. Yet, due to the gap between LLM-generated results and real human responses, the continuation methods further amplify these shortcomings. Given the high costs of collecting and annotating speech instruction datasets by humans, using speech synthesis to construct large-scale speech instruction datasets has become a balanced and robust alternative. Although modern Text-To-Speech~( textbf{TTS}) models have achieved near-human-level synthesis quality, it is challenging to appropriately convert out-of-distribution text instruction to speech due to the limitations of the training data distribution in TTS models. To address this issue, we propose a query rewriting framework with multi-LLM knowledge fusion, employing multiple agents to annotate and validate the synthesized speech, making it possible to construct high-quality speech instruction datasets without relying on human annotation. Experiments show that this method can transform text instructions into distributions more suitable for TTS models for speech synthesis through zero-shot rewriting, increasing data usability from 72 % to 93 %. It also demonstrates unique advantages in rewriting tasks that require complex knowledge and context-related abilities.
备注:Accepted at ACM MM 2025
摘要:随着生成模型的最新进展,文本到音频(T2 A)生成已经取得了令人鼓舞的成果。然而,由于时间对准的音频-文本对的质量和数量有限,现有的T2 A方法难以处理包含精确定时控制的复杂文本提示,例如,“猫头鹰以2.4s-5.2s的速度鸣叫”。最近的工作已经探索了数据增强技术或引入定时条件作为模型输入,以实现定时调节的10秒T2 A生成,而它们的合成质量仍然有限。在这项工作中,我们提出了一种新的无训练定时控制的T2 A框架FreeAudio,首次尝试实现定时控制的长形式T2 A生成,例如,“猫头鹰在2.4s-5.2s鸣叫,蟋蟀在0 s-24 s鸣叫”。具体来说,我们首先采用LLM计划非重叠的时间窗口,并根据输入的文本和时间提示,用精炼的自然语言描述重新描述每个时间窗口。然后我们介绍:1)解耦和聚合注意力控制,用于精确的定时控制; 2)上下文潜在成分,用于局部平滑性和参考指导,用于全局一致性。大量的实验表明:1)FreeAudio在无需训练的方法中实现了最先进的定时条件T2 A合成质量,可与领先的基于训练的方法相媲美; 2)FreeAudio与基于训练的Stable Audio具有可比的长格式生成质量,为定时控制的长格式T2 A合成铺平了道路。演示示例可在以下网址获得:https: freeaudio.github.io FreeAudio摘要:Text-to-audio (T2A) generation has achieved promising results with the recent advances in generative models. However, because of the limited quality and quantity of temporally-aligned audio-text pairs, existing T2A methods struggle to handle the complex text prompts that contain precise timing control, e.g., "owl hooted at 2.4s-5.2s". Recent works have explored data augmentation techniques or introduced timing conditions as model inputs to enable timing-conditioned 10-second T2A generation, while their synthesis quality is still limited. In this work, we propose a novel training-free timing-controlled T2A framework, FreeAudio, making the first attempt to enable timing-controlled long-form T2A generation, e.g., "owl hooted at 2.4s-5.2s and crickets chirping at 0s-24s". Specifically, we first employ an LLM to plan non-overlapping time windows and recaption each with a refined natural language description, based on the input text and timing prompts. Then we introduce: 1) Decoupling and Aggregating Attention Control for precise timing control; 2) Contextual Latent Composition for local smoothness and Reference Guidance for global consistency. Extensive experiments show that: 1) FreeAudio achieves state-of-the-art timing-conditioned T2A synthesis quality among training-free methods and is comparable to leading training-based methods; 2) FreeAudio demonstrates comparable long-form generation quality with training-based Stable Audio and paves the way for timing-controlled long-form T2A synthesis. Demo samples are available at: https: freeaudio.github.io FreeAudio
备注:Accepted by ISMIR 2025
摘要:从乐谱生成富有表现力的音频表演需要模型来捕获乐器声学和人类解释。传统的音乐表演合成管道遵循两个阶段的方法,首先从乐谱中生成表现性的表演片段,然后将片段合成为音频。然而,合成模型往往难以概括不同的音乐来源,音乐风格和录音环境。为了解决这些挑战,我们提出了MIDI-VALLE,一个神经编解码器语言模型改编自VALLE框架,这是最初设计的zero-shot个性化的文本到语音(TTS)的合成。对于性能MIDI到音频的合成,我们改进了架构,以参考音频性能及其相应的参数为条件。与之前基于TTS的系统依赖于钢琴卷不同,MIDI-VALLE将声音和音频编码为离散标记,从而促进钢琴演奏的更一致和更强大的建模。此外,该模型的泛化能力通过在广泛而多样化的钢琴演奏数据集上进行训练而得到增强。评估结果表明,MIDI-VALLE的性能明显优于最先进的基线,在ATEPP和Maestro数据集上实现了超过75%的Frechet音频距离降低。在听力测试中,MIDI-VALLE获得了202票,而基线为58票,这表明合成质量得到了提高,并且在不同的性能测试输入中得到了推广。摘要:Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pipelines follow a two-stage approach, first generating expressive performance MIDI from a score, then synthesising the MIDI into audio. However, the synthesis models often struggle to generalise across diverse MIDI sources, musical styles, and recording environments. To address these challenges, we propose MIDI-VALLE, a neural codec language model adapted from the VALLE framework, which was originally designed for zero-shot personalised text-to-speech (TTS) synthesis. For performance MIDI-to-audio synthesis, we improve the architecture to condition on a reference audio performance and its corresponding MIDI. Unlike previous TTS-based systems that rely on piano rolls, MIDI-VALLE encodes both MIDI and audio as discrete tokens, facilitating a more consistent and robust modelling of piano performances. Furthermore, the model's generalisation ability is enhanced by training on an extensive and diverse piano performance dataset. Evaluation results show that MIDI-VALLE significantly outperforms a state-of-the-art baseline, achieving over 75% lower Frechet Audio Distance on the ATEPP and Maestro datasets. In the listening test, MIDI-VALLE received 202 votes compared to 58 for the baseline, demonstrating improved synthesis quality and generalisation across diverse performance MIDI inputs.
链接:http://arxiv.org/pdf/2507.08412v1
摘要:环境声音记录通常包含可理解的语音,这引起了对隐私的担忧,限制了数据的分析、共享和再利用。在本文中,我们介绍了一种方法,使语音难以理解,同时保持声学场景的完整性,和整体的音频质量。我们的方法包括反转波形段来扭曲语音内容。这一过程通过语音活动检测和语音分离管道得到增强,这允许更精确地定位语音。 为了证明所提出的方法的有效性,我们考虑了一个由三部分组成的评估协议,该协议评估:1)使用单词错误率(WER)的语音可懂度,2)使用来自广泛使用的预训练模型的声源分类准确度下降(SCAD)的声源可检测性,以及3)使用Fr 'echet音频距离(FAD)的音频质量,用我们的参考数据集计算,该数据集包含未改变的语音。实验结果表明,我们的方法实现了令人满意的语音清晰度降低(97.9%WER),声源可检测性的最小退化(2.7%SCAD),和高感知质量(FAD为1.40)。消融研究进一步突出了管道每个组件的贡献。我们还表明,将随机拼接到我们的语音内容隐私执法方法可以提高算法的鲁棒性,试图恢复干净的语音,在音频质量的轻微成本。
摘要:Environmental sound recordings often contain intelligible speech, raising privacy concerns that limit analysis, sharing and reuse of data. In this paper, we introduce a method that renders speech unintelligible while preserving both the integrity of the acoustic scene, and the overall audio quality. Our approach involves reversing waveform segments to distort speech content. This process is enhanced through a voice activity detection and speech separation pipeline, which allows for more precise targeting of speech. In order to demonstrate the effectivness of the proposed approach, we consider a three-part evaluation protocol that assesses: 1) speech intelligibility using Word Error Rate (WER), 2) sound sources detectability using Sound source Classification Accuracy-Drop (SCAD) from a widely used pre-trained model, and 3) audio quality using the Fr 'echet Audio Distance (FAD), computed with our reference dataset that contains unaltered speech. Experiments on this simulated evaluation dataset, which consists of linear mixtures of speech and environmental sound scenes, show that our method achieves satisfactory speech intelligibility reduction (97.9% WER), minimal degradation of the sound sources detectability (2.7% SCAD), and high perceptual quality (FAD of 1.40). An ablation study further highlights the contribution of each component of the pipeline. We also show that incorporating random splicing to our speech content privacy enforcement method can enhance the algorithm's robustness to attempt to recover the clean speech, at a slight cost of audio quality.
摘要:音频修复是指在损坏的音频记录中重建丢失的片段的任务。虽然先前的方法-包括基于波形和频谱的扩散模型-已经显示出对短间隙的有希望的结果,但是当间隙超过100毫秒(ms)时,它们的质量通常会降低。在这项工作中,我们介绍了一种新的基于离散扩散建模的修复方法,该方法对由预训练的音频标记器产生的标记化音频表示进行操作。我们的方法直接在离散潜在空间中对生成过程进行建模,从而实现丢失音频的稳定和语义一致的重建。我们评估的方法MusicNet数据集上使用客观和感知指标的间隙持续时间长达300毫秒。我们进一步评估了我们的MTG数据集上的方法,将间隙持续时间延长到500毫秒。实验结果表明,我们的方法实现了竞争力或优越的性能相比,现有的基线,特别是较长的差距,提供了一个强大的解决方案,恢复退化的音乐录音。我们提出的方法的音频示例可以在https: iftach21.github.io 上找到
摘要:Audio inpainting refers to the task of reconstructing missing segments in corrupted audio recordings. While prior approaches-including waveform and spectrogram-based diffusion models-have shown promising results for short gaps, they often degrade in quality when gaps exceed 100 milliseconds (ms). In this work, we introduce a novel inpainting method based on discrete diffusion modeling, which operates over tokenized audio representations produced by a pre-trained audio tokenizer. Our approach models the generative process directly in the discrete latent space, enabling stable and semantically coherent reconstruction of missing audio. We evaluate the method on the MusicNet dataset using both objective and perceptual metrics across gap durations up to 300 ms. We further evaluated our approach on the MTG dataset, extending the gap duration to 500 ms. Experimental results demonstrate that our method achieves competitive or superior performance compared to existing baselines, particularly for longer gaps, offering a robust solution for restoring degraded musical recordings. Audio examples of our proposed method can be found at https: iftach21.github.io摘要:高质量数据集的构建是现代文语转换(TTS)系统的基石。然而,可用数据规模的不断扩大带来了重大挑战,包括存储限制。针对这些问题,本文提出了一种基于主动学习的TTS语料库构建方法。与传统的前馈和模型无关的语料库构建方法不同,我们的方法在数据收集和模型训练之间反复交替,从而专注于获取对模型改进更有用的数据。这种方法能够构建数据高效的语料库。实验结果表明,使用我们的方法构建的语料库,使更高质量的语音合成比语料库的相同大小。
摘要:The construction of high-quality datasets is a cornerstone of modern text-to-speech (TTS) systems. However, the increasing scale of available data poses significant challenges, including storage constraints. To address these issues, we propose a TTS corpus construction method based on active learning. Unlike traditional feed-forward and model-agnostic corpus construction approaches, our method iteratively alternates between data collection and model training, thereby focusing on acquiring data that is more informative for model improvement. This approach enables the construction of a data-efficient corpus. Experimental results demonstrate that the corpus constructed using our method enables higher-quality speech synthesis than corpora of the same size.
备注:Working note submitted to CLEF 2025 under the LifeCLEF lab
摘要:BirdCLEF+ 2025挑战要求在90分钟的严格CPU推理期限内对206个物种进行分类,包括鸟类、哺乳动物、昆虫和两栖动物,这使得许多最先进的深度学习方法变得不切实际。为了解决这一限制,DS@GT BirdCLEF团队探索了两种策略。首先,我们通过优化来自生物声学模型动物园的预训练模型来建立有竞争力的基线。使用TFLite,我们为Perch模型实现了近10倍的推理加速,使其能够在大约16分钟内运行,并在赛后公共排行榜上获得0.729的最终ROC-AUC分数,在私人排行榜上获得0.711。动物园的最佳模型是BirdSetEfficientNetB 1,公开得分为0.810,私人得分为0.778。其次,我们介绍了一种新的,轻量级的管道命名为Spectrogram令牌Skip-Gram(STSG),将生物声学作为一个序列建模任务。该方法通过使用Faiss K-means对Mel谱图进行聚类,将音频转换为离散的“谱图令牌”,然后使用Word 2 Vec跳图模型以无监督的方式学习这些令牌的高质量上下文嵌入。对于分类,5秒窗口内的嵌入被平均并传递到线性模型。对于700分钟的测试集,预计推理时间为6分钟,STSG方法的最终ROC-AUC公共得分为0.559,私有得分为0.520,证明了使用静态嵌入进行生物声学分类的快速标记化方法的可行性。本文的支持代码可以在https: github.com dsgt-arc birdclef-2025上找到。
摘要:The BirdCLEF+ 2025 challenge requires classifying 206 species, including birds, mammals, insects, and amphibians, from soundscape recordings under a strict 90-minute CPU-only inference deadline, making many state-of-the-art deep learning approaches impractical. To address this constraint, the DS@GT BirdCLEF team explored two strategies. First, we establish competitive baselines by optimizing pre-trained models from the Bioacoustics Model Zoo for CPU inference. Using TFLite, we achieved a nearly 10x inference speedup for the Perch model, enabling it to run in approximately 16 minutes and achieve a final ROC-AUC score of 0.729 on the public leaderboard post-competition and 0.711 on the private leaderboard. The best model from the zoo was BirdSetEfficientNetB1, with a public score of 0.810 and a private score of 0.778. Second, we introduce a novel, lightweight pipeline named Spectrogram Token Skip-Gram (STSG) that treats bioacoustics as a sequence modeling task. This method converts audio into discrete "spectrogram tokens" by clustering Mel-spectrograms using Faiss K-means and then learns high-quality contextual embeddings for these tokens in an unsupervised manner with a Word2Vec skip-gram model. For classification, embeddings within a 5-second window are averaged and passed to a linear model. With a projected inference time of 6 minutes for a 700-minute test set, the STSG approach achieved a final ROC-AUC public score of 0.559 and a private score of 0.520, demonstrating the viability of fast tokenization approaches with static embeddings for bioacoustic classification. Supporting code for this paper can be found at https: github.com dsgt-arc birdclef-2025.
备注:Code, Datasets and Models: this https URL
摘要:Audio Flamingo 3(AF 3)是一个完全开放的最先进的(SOTA)大型音频语言模型,可以在语音,声音和音乐中推进推理和理解。AF 3介绍:(i)AF-Whisper,一个统一的音频编码器,使用一种新的策略进行训练,用于跨语音、声音和音乐的所有3种模式的联合表示学习;(ii)灵活的按需思维,允许模型在回答之前进行思维链类型的推理;(iii)多回合、多音频聊天;(iv)长音频理解和推理(包括演讲)最多10分钟;及(v)语音对语音互动。为了实现这些功能,我们提出了几个使用新策略策划的大规模训练数据集,包括AudioSkills-XL,LongAudio-XL,AF-Think和AF-Chat,并使用一种新的基于五阶段训练的训练策略训练AF 3。AF 3仅在开源音频数据上训练,在超过20个(长)音频理解和推理基准上实现了新的SOTA结果,超过了在更大数据集上训练的开放权重和闭源模型。
摘要:We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning across all 3 modalities of speech, sound, and music; (ii) flexible, on-demand thinking, allowing the model to do chain-of-thought-type reasoning before answering; (iii) multi-turn, multi-audio chat; (iv) long audio understanding and reasoning (including speech) up to 10 minutes; and (v) voice-to-voice interaction. To enable these capabilities, we propose several large-scale training datasets curated using novel strategies, including AudioSkills-XL, LongAudio-XL, AF-Think, and AF-Chat, and train AF3 with a novel five-stage curriculum-based training strategy. Trained on only open-source audio data, AF3 achieves new SOTA results on over 20+ (long) audio understanding and reasoning benchmarks, surpassing both open-weight and closed-source models trained on much larger datasets.
摘要:房间脉冲响应估计对于诸如语音去混响之类的任务是必不可少的,这提高了自动语音识别。大多数现有方法依赖于统计信号处理或旨在复制信号处理原理的深度神经网络。然而,结合统计和物理建模RIR估计仍然在很大程度上未被探索。本文提出了一种通过理论基础模型将这两个方面整合在一起的新方法。RIR被分解为可解释的参数:通过频率相关指数衰减(例如,建模壁吸收)和自回归滤波器(例如,建模麦克风响应)过滤的高斯白噪声。变分自由能成本函数使实际的参数估计。作为一个概念的证明,我们表明,干和混响的语音信号,所提出的方法优于经典的去卷积在嘈杂的环境中,通过客观的度量验证。
摘要:Room impulse response estimation is essential for tasks like speech dereverberation, which improves automatic speech recognition. Most existing methods rely on either statistical signal processing or deep neural networks designed to replicate signal processing principles. However, combining statistical and physical modeling for RIR estimation remains largely unexplored. This paper proposes a novel approach integrating both aspects through a theoretically grounded model. The RIR is decomposed into interpretable parameters: white Gaussian noise filtered by a frequency-dependent exponential decay (e.g. modeling wall absorption) and an autoregressive filter (e.g. modeling microphone response). A variational free-energy cost function enables practical parameter estimation. As a proof of concept, we show that given dry and reverberant speech signals, the proposed method outperforms classical deconvolution in noisy environments, as validated by objective metrics.
摘要:基于文本到语音的模型允许用户通过自然语言指令来控制语音的不同方面,例如说话速率和感知性别。虽然用户友好,这样的方法是一方面约束:控制仅限于在训练过程中暴露于模型的声学特征,另一方面过于灵活:相同的输入会产生无法控制的变化,这些变化反映在语料库统计中。 我们研究了一种新的微调制度,以解决这两个问题,在同一时间,利用模型的不可控的方差。通过对数千个合成样本的主成分分析,我们确定了占输出方差比例最高的潜在特征,并将它们作为新标签进行二次微调。我们评估所提出的方法在两个模型上训练的表达冰岛语语音语料库,一个与情感的披露和一个没有。在没有情绪披露的模型的情况下,该方法产生连续和离散的功能,提高整体的可控性的模型。
摘要:A Prompt-based Text-To-Speech model allows a user to control different aspects of speech, such as speaking rate and perceived gender, through natural language instruction. Although user-friendly, such approaches are on one hand constrained: control is limited to acoustic features exposed to the model during training, and too flexible on the other: the same inputs yields uncontrollable variation that are reflected in the corpus statistics. We investigate a novel fine-tuning regime to address both of these issues at the same time by exploiting the uncontrollable variance of the model. Through principal component analysis of thousands of synthesised samples, we determine latent features that account for the highest proportion of the output variance and incorporate them as new labels for secondary fine-tuning. We evaluate the proposed methods on two models trained on an expressive Icelandic speech corpus, one with emotional disclosure and one without. In the case of the model without emotional disclosure, the method yields both continuous and discrete features that improve overall controllability of the model.
机器翻译由腾讯交互翻译提供,仅供参考
