微信公众号:arXiv_Daily
cs.SD语音
标题: ISAC:一个可逆且稳定的听觉过滤器组,具有可定制的ML集成内核
链接:https://arxiv.org/abs/2505.07709
备注:Accepted at the IEEE International Conference on Sampling Theory and Applications (SampTA) 2025
摘要:本文介绍了ISAC,这是一种可逆的、稳定的、感知驱动的滤波器组,专门设计用于集成到机器学习范式中。更准确地说,滤波器的中心频率和带宽被选择为遵循非线性听觉频率尺度,滤波器内核具有用户定义的最大时间支持并且可以用作可学习的卷积内核,并且存在对应的滤波器组,使得两者形成完美的重建对。ISAC提供了一个功能强大且用户友好的音频前端,适用于任何应用,包括分析合成方案。
摘要:This paper introduces ISAC, an invertible and stable, perceptually-motivated filter bank that is specifically designed to be integrated into machine learning paradigms. More precisely, the center frequencies and bandwidths of the filters are chosen to follow a non-linear, auditory frequency scale, the filter kernels have user-defined maximum temporal support and may serve as learnable convolutional kernels, and there exists a corresponding filter bank such that both form a perfect reconstruction pair. ISAC provides a powerful and user-friendly audio front-end suitable for any application, including analysis-synthesis schemes.
【2】 Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
标题: 轻量级端到端文本到语音合成,适用于低资源设备上应用链接:https://arxiv.org/abs/2505.07701
备注:Published as a conference paper at SSW 2023
摘要:最近的工作表明,以端到端(E2 E)方式直接从文本建模原始波形,比基于级联或两阶段方法的传统神经文本到语音(TTS)系统产生更自然的语音。然而,当前的E2 E最先进的模型计算复杂且消耗内存,使得它们不适合低资源场景中的实时离线设备上应用。为了解决这个问题,我们提出了一个轻量级E2 E-TTS(LE 2 E)模型,生成高质量的语音,需要最少的计算资源。我们在LJSpeech数据集上评估了所提出的模型,并表明它实现了最先进的性能,同时在模型参数方面最多可减少90\%$,在实时因素方面最多可提高10\times $。此外,我们证明了建议的E2 E培训范例实现更好的质量相比,在两个阶段的方法训练的等效架构。我们的研究结果表明,LE 2 E是一种很有前途的方法,为开发实时,高质量,低资源的TTS应用程序的设备上的应用程序。
摘要:Recent works have shown that modelling raw waveform directly from text in an end-to-end (E2E) fashion produces more natural-sounding speech than traditional neural text-to-speech (TTS) systems based on a cascade or two-stage approach. However, current E2E state-of-the-art models are computationally complex and memory-consuming, making them unsuitable for real-time offline on-device applications in low-resource scenarios. To address this issue, we propose a Lightweight E2E-TTS (LE2E) model that generates high-quality speech requiring minimal computational resources. We evaluate the proposed model on the LJSpeech dataset and show that it achieves state-of-the-art performance while being up to $90\%$ smaller in terms of model parameters and $10\times$ faster in real-time-factor. Furthermore, we demonstrate that the proposed E2E training paradigm achieves better quality compared to an equivalent architecture trained in a two-stage approach. Our results suggest that LE2E is a promising approach for developing real-time, high quality, low-resource TTS applications for on-device applications.
【3】 Multi-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge
标题: DUSE 2025挑战赛中多域音频问题回答声学内容推理链接:https://arxiv.org/abs/2505.07365
备注:Preprint. DCASE 2025 Audio QA Challenge: this https URL
摘要:我们提出了DCASE 2025挑战的任务5:一个跨越声音理解多个领域的音频问题分类(AQA)基准。该任务定义了三个QA子集(生物声学,时间音景和复杂QA),以测试音频语言模型在不同声学场景下的交互式问答。我们描述了数据集的组成(从海洋哺乳动物的声音和复杂的现实世界的剪辑),评估协议(前1名的准确性与答案洗牌的鲁棒性),和基线系统(Qwen 2-Audio-7 B,AudioFlamingo 2,Gemini-2-Flash)。对开发集的初步结果进行了比较,显示出模型和子集之间的强烈变化。这项挑战旨在将音频语言模型的音频理解和推理能力提升到人类水平的敏锐度,这对于使AI代理能够有效地感知和交互世界至关重要。
摘要:We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over diverse acoustic scenes. We describe the dataset composition (from marine mammal calls to soundscapes and complex real-world clips), the evaluation protocol (top-1 accuracy with answer-shuffling robustness), and baseline systems (Qwen2-Audio-7B, AudioFlamingo 2, Gemini-2-Flash). Preliminary results on the development set are compared, showing strong variation across models and subsets. This challenge aims to advance the audio understanding and reasoning capabilities of audio-language models toward human-level acuity, which are crucial for enabling AI agents to perceive and interact about the world effectively.
【4】 Predicting Music Track Popularity by Convolutional Neural Networks on Spotify Features and Spectrogram of Audio Waveform
标题: 通过基于Spotify特征和音频频谱图的卷积神经网络预测音乐曲目受欢迎程度链接:https://arxiv.org/abs/2505.07280
备注:12 pages, 6 figures, 4 tables
摘要:在数字流媒体领域,艺术家和行业专家预测音乐曲目的成功变得越来越具有挑战性。这项研究介绍了一种开创性的方法,该方法使用卷积神经网络(CNN)和Spotify数据分析来预测音乐曲目的受欢迎程度。我们的方法利用了Spotify的广泛功能,包括基于音频波形频谱图的声学属性、元数据和用户参与度指标,来捕捉影响曲目受欢迎程度的复杂模式和关系。使用涵盖各种流派和人口统计的大型数据集,我们基于CNN的模型在预测音乐曲目的流行度方面表现出令人印象深刻的有效性。此外,我们还进行了广泛的实验,以评估我们的模型在不同音乐风格和时间段的强度和适应性,结果令人鼓舞,获得了97%的F1分数。我们的研究不仅为数字音乐消费的动态格局提供了有价值的见解,还为音乐行业提供了先进的预测工具,用于评估和预测音乐曲目的成功。
摘要:In the digital streaming landscape, it's becoming increasingly challenging for artists and industry experts to predict the success of music tracks. This study introduces a pioneering methodology that uses Convolutional Neural Networks (CNNs) and Spotify data analysis to forecast the popularity of music tracks. Our approach takes advantage of Spotify's wide range of features, including acoustic attributes based on the spectrogram of audio waveform, metadata, and user engagement metrics, to capture the complex patterns and relationships that influence a track's popularity. Using a large dataset covering various genres and demographics, our CNN-based model shows impressive effectiveness in predicting the popularity of music tracks. Additionally, we've conducted extensive experiments to assess the strength and adaptability of our model across different musical styles and time periods, with promising results yielding a 97\% F1 score. Our study not only offers valuable insights into the dynamic landscape of digital music consumption but also provides the music industry with advanced predictive tools for assessing and predicting the success of music tracks.
【5】 Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
标题: 神经心理声学编码的多频段频率重建链接:https://arxiv.org/abs/2505.07235
摘要:实现高保真音频压缩,同时保持不同内容的感知质量仍然是神经音频编码(NAC)的关键挑战。我们介绍MUFFIN,一个完全卷积的神经心理声学编码(NPC)框架,利用心理声学引导的多频带频率重建。其核心是一个多频带频谱残差矢量量化(MBS-RVQ)模块,该模块基于感知显著性在频带上分配比特率。这种设计实现了有效的压缩,同时使用不同的码本将说话者身份从内容中分离出来。MUFFIN结合了transformer启发的卷积骨干和修改的snake激活,以提高细粒度光谱区域的分辨率。在多个基准上的实验结果表明,MUFFIN在重建质量上始终优于现有的方法。高压缩变体以最小的损耗实现了最先进的12.5 Hz速率。MUFFIN在下游生成任务中也被证明是有效的,突出了它作为与语言模型集成的令牌表示的承诺。音频样本和代码可用。
摘要:Achieving high-fidelity audio compression while preserving perceptual quality across diverse content remains a key challenge in Neural Audio Coding (NAC). We introduce MUFFIN, a fully convolutional Neural Psychoacoustic Coding (NPC) framework that leverages psychoacoustically guided multi-band frequency reconstruction. At its core is a Multi-Band Spectral Residual Vector Quantization (MBS-RVQ) module that allocates bitrate across frequency bands based on perceptual salience. This design enables efficient compression while disentangling speaker identity from content using distinct codebooks. MUFFIN incorporates a transformer-inspired convolutional backbone and a modified snake activation to enhance resolution in fine-grained spectral regions. Experimental results on multiple benchmarks demonstrate that MUFFIN consistently outperforms existing approaches in reconstruction quality. A high-compression variant achieves a state-of-the-art 12.5 Hz rate with minimal loss. MUFFIN also proves effective in downstream generative tasks, highlighting its promise as a token representation for integration with language models. Audio samples and code are available.
【6】 On the Cost and Benefits of Training Context with Utterance or Full Conversation Training: A Comparative Stud
标题: 带直言不讳或全面对话训练的训练环境的成本和收益:比较研究链接:https://arxiv.org/abs/2505.07202
摘要:为对话设计的现代TTS系统实现了高质量的话语,但通常仍然无法公开访问。是现有的开源架构不够,还是目前的培训技术不够?本文研究了会话语境中的主要模型及其潜在行为。在NVIDIA H100上使用20个GPU小时,我们实证研究了两种方法:基于上下文的话语级训练与完整的会话训练。结果表明,基于上下文的话语训练取得了优异的MOS分数(4.3/5.0 vs 3.7/5.0),并减少了37%的训练时间,而完整的对话方法遭受说话人相似性幻觉问题。这些研究结果提供了实用的指导方针会话TTS的发展,有利于话语水平的培训与上下文条件的资源效率和输出质量。
摘要:Modern TTS systems designed for conversations achieve high-quality utterances but often remain inaccessible publicly. Are existing open-source architectures inadequate, or are current training techniques insufficient? This paper investigates prominent models and their underlying behaviors regarding conversational context. Using 20 GPU-hours on an NVIDIA H100, we empirically examine two approaches: context-based utterance-level training versus full conversation training. Results demonstrate that context-based utterance training achieves superior MOS scores (4.3/5.0 vs 3.7/5.0) and reduces training time by 37%, while full conversation approaches suffer from speaker similarity hallucination issues. These findings provide practical guidelines for conversational TTS development, favoring utterance-level training with contextual conditioning for both resource efficiency and output quality.
【7】 Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
标题: 架起耳朵和眼睛:在可见声音识别中分析人类的音频和视觉大型语言模型并通过跨模式蒸馏减少他们的感觉差距链接:https://arxiv.org/abs/2505.06803
摘要:音频大语言模型(LLM)被认为是识别声音对象的专家,但它们相对于其他感官模式(如视觉或视听LLM)中的LLM以及使用耳朵、眼睛或两者的人类的性能仍然未被探索。为了研究这一点,我们系统地评估了音频,视觉和视听LLM,特别是Qwen 2-Audio,Qwen 2-VL和Qwen2.5-Omni,对人类识别来自仅音频,无声视频或有声视频输入的不同类别的声音对象。我们发现Qwen 2-Audio和Qwen 2-VL之间的性能差距与人耳和眼睛之间的感官差异相似。为了缩小这一差距,我们引入了一个跨模态蒸馏框架,其中一种模态的LLM作为教师,另一种作为学生,通过启发式模型预测声音类中的知识转移对学生更具挑战性。从Qwen 2-VL到Qwen 2-Audio的双向蒸馏,反之亦然,会带来显著的改进,特别是在具有挑战性的课程中。这项工作突出了感觉差距LLM从人类对齐的角度来看,并提出了一个原则性的方法,以提高模态特定的感知多模态LLM。
摘要:Audio large language models (LLMs) are considered experts at recognizing sound objects, yet their performance relative to LLMs in other sensory modalities, such as visual or audio-visual LLMs, and to humans using their ears, eyes, or both remains unexplored. To investigate this, we systematically evaluate audio, visual, and audio-visual LLMs, specifically Qwen2-Audio, Qwen2-VL, and Qwen2.5-Omni, against humans in recognizing sound objects of different classes from audio-only, silent video, or sounded video inputs. We uncover a performance gap between Qwen2-Audio and Qwen2-VL that parallels the sensory discrepancy between human ears and eyes. To reduce this gap, we introduce a cross-modal distillation framework, where an LLM in one modality serves as the teacher and another as the student, with knowledge transfer in sound classes predicted as more challenging to the student by a heuristic model. Distillation in both directions, from Qwen2-VL to Qwen2-Audio and vice versa, leads to notable improvements, particularly in challenging classes. This work highlights the sensory gap in LLMs from a human-aligned perspective and proposes a principled approach to enhancing modality-specific perception in multimodal LLMs.
【8】 Beyond Identity: A Generalizable Approach for Deepfake Audio Detection
标题: 超越身份:Deepfake音频检测的一种通用方法链接:https://arxiv.org/abs/2505.06766
备注:Submitted to IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM)
摘要:Deepfake音频对数字安全构成了越来越大的威胁,因为它有可能引发社会工程、欺诈和身份滥用。然而,现有的检测模型由于隐式身份泄漏而在数据集之间具有较差的泛化能力,其中模型无意中学习说话者特定的特征而不是操纵伪像。据我们所知,这是第一项明确分析和解决音频deepfake检测领域身份泄露的研究。这项工作提出了一个独立于身份的音频deepfake检测框架,通过鼓励模型专注于伪造特定的伪影,而不是过度拟合扬声器特征,来减轻身份泄露。我们的方法利用人工智能检测模块(ADM)在时域和频域中隔离合成工件,增强跨数据集的泛化。我们引入了新的动态伪影生成技术,包括频域交换,时域操作,和背景噪声增强,以加强学习的不变特征。在ASVspoof 2019、ADD 2022、FoR和In-The-Wild数据集上进行的大量实验表明,所提出的ADM增强模型的F1得分为0.230(ADD 2022)、0.604(FoR)和0.813(In-The-Wild),始终优于基线。动态频率交换被证明是在不同条件下最有效的策略。这些发现强调了基于伪影的学习在减轻隐式身份泄露以进行更普遍的音频deepfake检测方面的价值。
摘要:Deepfake audio presents a growing threat to digital security, due to its potential for social engineering, fraud, and identity misuse. However, existing detection models suffer from poor generalization across datasets, due to implicit identity leakage, where models inadvertently learn speaker-specific features instead of manipulation artifacts. To the best of our knowledge, this is the first study to explicitly analyze and address identity leakage in the audio deepfake detection domain. This work proposes an identity-independent audio deepfake detection framework that mitigates identity leakage by encouraging the model to focus on forgery-specific artifacts instead of overfitting to speaker traits. Our approach leverages Artifact Detection Modules (ADMs) to isolate synthetic artifacts in both time and frequency domains, enhancing cross-dataset generalization. We introduce novel dynamic artifact generation techniques, including frequency domain swaps, time domain manipulations, and background noise augmentation, to enforce learning of dataset-invariant features. Extensive experiments conducted on ASVspoof2019, ADD 2022, FoR, and In-The-Wild datasets demonstrate that the proposed ADM-enhanced models achieve F1 scores of 0.230 (ADD 2022), 0.604 (FoR), and 0.813 (In-The-Wild), consistently outperforming the baseline. Dynamic Frequency Swap proves to be the most effective strategy across diverse conditions. These findings emphasize the value of artifact-based learning in mitigating implicit identity leakage for more generalizable audio deepfake detection.
【9】 TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
标题: TS-SURB:语音自我监督学习模型的目标语音处理基准链接:https://arxiv.org/abs/2505.06660
备注:Accepted at ICASSP 2025
摘要:自监督学习(SSL)模型有显着先进的语音处理任务,已经提出了几个基准来验证其有效性。然而,以前的基准测试主要集中在单个说话者的情况下,在嘈杂的多说话者条件下对目标说话者任务的探索较少-这是一个更具挑战性但实用的案例。在本文中,我们介绍了目标扬声器语音处理通用性能基准(TS-SUPERB),其中包括四个广泛认可的目标扬声器处理任务,需要识别目标扬声器和提取信息的语音混合。在我们的基准测试中,从注册语音中提取的说话人嵌入被用作条件下游模型的线索。基准测试结果揭示了在目标说话人场景中评估SSL模型的重要性,表明性能不能轻易地从相关的单说话人任务中推断出来。此外,通过使用一个统一的基于SSL的目标语音编码器,由一个扬声器编码器和提取器模块,我们还研究跨TS任务的联合优化,以利用互信息,并证明其有效性。
摘要:Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in noisy, multi-talker conditions -- a more challenging yet practical case. In this paper, we introduce the Target-Speaker Speech Processing Universal Performance Benchmark (TS-SUPERB), which includes four widely recognized target-speaker processing tasks that require identifying the target speaker and extracting information from the speech mixture. In our benchmark, the speaker embedding extracted from enrollment speech is used as a clue to condition downstream models. The benchmark result reveals the importance of evaluating SSL models in target speaker scenarios, demonstrating that performance cannot be easily inferred from related single-speaker tasks. Moreover, by using a unified SSL-based target speech encoder, consisting of a speaker encoder and an extractor module, we also investigate joint optimization across TS tasks to leverage mutual information and demonstrate its effectiveness.
【10】 Interpretable SHAP-bounded Bayesian Optimization for Underwater Acoustic Metamaterial Coating Design
标题: 水下声学超材料涂层设计的可解释SHAP有界Bayesian优化链接:https://arxiv.org/abs/2505.06519
摘要:我们开发了一个可解释性通知贝叶斯优化框架,以优化基于嵌入式超材料功能的聚氨酯弹性体的水声涂层。采用数据驱动模型分析声学性能,特别是吸声性能与相应设计变量之间的关系。通过利用机器学习可解释性工具SHapley Additive exPlanations(SHAP),我们确定了影响目标函数的关键参数,并深入了解了这些参数如何影响吸声。从SHAP分析中获得的见解随后用于自动优化优化问题的边界,从而实现对设计空间的更有针对性和更有效的探索。 所提出的方法被应用到两个不同的硬度水平的聚氨酯材料,导致改进的最佳解决方案相比,没有形状知情的指导。值得注意的是,这些增强是在不增加模拟迭代次数的情况下实现的。我们的研究结果表明,SHAP的潜力,以简化优化过程,揭示隐藏的参数关系,并引导搜索到有前途的区域的设计空间。这项工作强调了可解释性技术与贝叶斯优化相结合的有效性,在严格的计算约束条件下,水下声学超材料的高效和成本效益的设计,可以推广到其他材料和工程优化问题。
摘要:We developed an interpretability informed Bayesian optimization framework to optimize underwater acoustic coatings based on polyurethane elastomers with embedded metamaterial features. A data driven model was employed to analyze the relationship between acoustic performance, specifically sound absorption and the corresponding design variables. By leveraging SHapley Additive exPlanations (SHAP), a machine learning interpretability tool, we identified the key parameters influencing the objective function and gained insights into how these parameters affect sound absorption. The insights derived from the SHAP analysis were subsequently used to automatically refine the bounds of the optimization problem automatically, enabling a more targeted and efficient exploration of the design space. The proposed approach was applied to two polyurethane materials with distinct hardness levels, resulting in improved optimal solutions compared to those obtained without SHAP-informed guidance. Notably, these enhancements were achieved without increasing the number of simulation iterations. Our findings demonstrate the potential of SHAP to streamline optimization processes by uncovering hidden parameter relationships and guiding the search toward promising regions of the design space. This work underscores the effectiveness of combining interpretability techniques with Bayesian optimization for the efficient and cost-effective design of underwater acoustic metamaterials under strict computational constraints and can be generalized towards other materials and engineering optimization problems.
【11】 Tri-MTL: A Triple Multitask Learning Approach for Respiratory Disease Diagnosis
标题: Tri-MTL:呼吸系统疾病诊断的三重多任务学习方法链接:https://arxiv.org/abs/2505.06271
备注:Accepted to EMBC 2025
摘要:听诊仍然是临床实践的基石,对于初始评估和持续监测都至关重要。临床医生听肺音,并结合病人的病史和测试结果作出诊断。考虑到这种强烈的关联,多任务学习(MTL)可以提供一个令人信服的框架来同时对这些关系进行建模,将呼吸声模式与疾病表现相结合。虽然MTL在医学应用中显示出相当大的前景,但在理解呼吸音、疾病表现和患者元数据属性之间的复杂相互作用方面仍然存在重大的研究差距。这项研究探讨了如何将MTL与先进的深度学习架构相结合,以增强呼吸音分类和疾病诊断。具体来说,我们扩展了最近的研究结果,元数据对呼吸声分类的有益影响,通过评估其有效性内的MTL框架。我们的综合实验揭示了显着改善肺音分类和诊断性能时,听诊器的信息被纳入MTL架构。
摘要:Auscultation remains a cornerstone of clinical practice, essential for both initial evaluation and continuous monitoring. Clinicians listen to the lung sounds and make a diagnosis by combining the patient's medical history and test results. Given this strong association, multitask learning (MTL) can offer a compelling framework to simultaneously model these relationships, integrating respiratory sound patterns with disease manifestations. While MTL has shown considerable promise in medical applications, a significant research gap remains in understanding the complex interplay between respiratory sounds, disease manifestations, and patient metadata attributes. This study investigates how integrating MTL with cutting-edge deep learning architectures can enhance both respiratory sound classification and disease diagnosis. Specifically, we extend recent findings regarding the beneficial impact of metadata on respiratory sound classification by evaluating its effectiveness within an MTL framework. Our comprehensive experiments reveal significant improvements in both lung sound classification and diagnostic performance when the stethoscope information is incorporated into the MTL architecture.
【12】 Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
标题: 扩散责任:分析生成性文本到音频扩散模型的能耗链接:https://arxiv.org/abs/2505.07615
摘要:文本到音频模型最近已经成为从文本描述生成声音的强大技术。然而,它们的高计算需求引起了对能源消耗和环境影响的担忧。在本文中,我们进行了7个国家的最先进的文本到音频的扩散为基础的生成模型的能量使用的分析,评估在多大程度上生成参数的变化影响在推理时的能量消耗。我们还旨在通过考虑所有选定模型的帕累托最优解决方案来确定音频质量和能耗之间的最佳平衡。我们的研究结果为性能和环境影响之间的权衡提供了见解,有助于开发更高效的生成音频模型。
摘要:Text-to-audio models have recently emerged as a powerful technology for generating sound from textual descriptions. However, their high computational demands raise concerns about energy consumption and environmental impact. In this paper, we conduct an analysis of the energy usage of 7 state-of-the-art text-to-audio diffusion-based generative models, evaluating to what extent variations in generation parameters affect energy consumption at inference time. We also aim to identify an optimal balance between audio quality and energy consumption by considering Pareto-optimal solutions across all selected models. Our findings provide insights into the trade-offs between performance and environmental impact, contributing to the development of more efficient generative audio models.
【13】 TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
标题: TACOS:用于语音音频预训练的时间对齐音频CaptiOnS链接:https://arxiv.org/abs/2505.07609
备注:submitted to the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025. Dataset (Zenodo): this https URL, Implementation (GitHub): this https URL
摘要:学习将音频与文本描述相关联对于一系列任务是有价值的,包括预训练、zero-shot分类、音频检索、音频字幕和文本条件音频生成。现有的对比语言音频预训练模型通常使用全局的剪辑级描述进行训练,仅提供弱的时间监督。我们假设,CLAP类语言音频模型-特别是,如果他们预计将产生帧级嵌入-可以受益于更强的时间监督。为了证实我们的假设,我们策划了一个由Freesound的大约12,000个音频记录组成的新数据集,每个音频记录都带有与音频记录中特定时间段相关的单句自由文本描述。我们使用大型语言模型来清除这些注释,删除对不可听事件、转录语音、错别字和注释者语言偏见的引用。我们进一步提出了一种逐帧对比训练策略,该策略学习将文本描述与音频记录中的时间区域对齐,并证明我们的模型在AudioSet Strong基准测试中评估时,与仅在全局字幕上训练的模型相比,具有更好的时间文本-音频对齐能力。数据集和我们的源代码分别在Zenodo和GitHub上提供。
摘要:Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text-conditioned audio generation. Existing contrastive language-audio pretrained models are typically trained using global, clip-level descriptions, which provide only weak temporal supervision. We hypothesize that CLAP-like language-audio models - particularly, if they are expected to produce frame-level embeddings - can benefit from a stronger temporal supervision. To confirm our hypothesis, we curate a novel dataset of approximately 12,000 audio recordings from Freesound, each annotated with single-sentence free-text descriptions linked to a specific temporal segment in an audio recording. We use large language models to clean these annotations by removing references to non-audible events, transcribed speech, typos, and annotator language bias. We further propose a frame-wise contrastive training strategy that learns to align text descriptions with temporal regions in an audio recording and demonstrate that our model has better temporal text-audio alignment abilities compared to models trained only on global captions when evaluated on the AudioSet Strong benchmark. The dataset and our source code are available on Zenodo and GitHub, respectively.
【14】 RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
标题: RADE:一种用于通过HF无线电频道传输语音的神经编解码器链接:https://arxiv.org/abs/2505.06671
备注:5 pages
摘要:语音压缩通常用于在诸如移动电话和双向一键通(PTT)无线电之类的应用中通过无线电信道发送语音。在经典系统中,语音编解码器与前向纠错、调制和无线电硬件相结合。在本文中,我们描述了一个自动编码器,取代了许多传统的信号处理元件与神经网络。编码器采用声码器特征集(短期频谱、音调、语音),并产生离散时间但连续值的正交幅度调制(QAM)符号。我们使用正交频域复用(OFDM)在高频(HF)无线电信道上发送和接收这些符号。解码器将接收到的QAM符号转换为适合于合成的声码器特征。自编码器已被训练为对加性高斯噪声和多径信道损伤具有鲁棒性,同时保持小于1 dB的峰均功率比(PAPR)。通过模拟和真实世界的高频无线电信道,我们已经实现了输出语音清晰度,明显超过了现有的模拟和数字无线电系统的信噪比范围。
摘要:Speech compression is commonly used to send voice over radio channels in applications such as mobile telephony and two-way push-to-talk (PTT) radio. In classical systems, the speech codec is combined with forward error correction, modulation and radio hardware. In this paper we describe an autoencoder that replaces many of the traditional signal processing elements with a neural network. The encoder takes a vocoder feature set (short term spectrum, pitch, voicing), and produces discrete time, but continuously valued quadrature amplitude modulation (QAM) symbols. We use orthogonal frequency domain multiplexing (OFDM) to send and receive these symbols over high frequency (HF) radio channels. The decoder converts received QAM symbols to vocoder features suitable for synthesis. The autoencoder has been trained to be robust to additive Gaussian noise and multipath channel impairments while simultaneously maintaining a Peak To Average Power Ratio (PAPR) of less than 1~dB. Over simulated and real world HF radio channels we have achieved output speech intelligibility that clearly surpasses existing analog and digital radio systems over a range of SNRs.
【15】 STRAUSS: Sonification Tools & Resources for Analysis Using Sound Synthesis
标题: 施特劳斯:使用声音合成进行分析的发声工具和资源链接:https://arxiv.org/abs/2504.01660
备注:4 pages, linking to documentation on ReadTheDocs (this https URL)
摘要:声音化,或使用非语言音频传达数据,是一种相对小众但不断发展的方法,用于在包括天文学,气候科学等多个专业领域呈现数据。STRAUSS Python包旨在提供这样一个工具,它建立在以前的方法之上,提供了一种强大的手段来探索表达数据的不同方式,并对输出音频及其格式进行精细控制。STRAUSS是一个免费的开源软件包,旨在将灵活有效的声音处理集成到数据工作流程中,类似于广泛使用的可视化软件包。STRAUSS的职责范围很广;它旨在能够在用于声音化非常特定的数据集的特设解决方案与未针对声音化进行优化或可能具有陡峭学习曲线的高技术作曲和声音设计工具之间架起桥梁。该规范提供了一系列的方法来发声的各种情况下(如科学教育,科学传播,技术数据分析等)。为此,STRAUSS包装了许多不同的声化方法的示例,并预设了支持“低屏障,高天花板”方法的配置。STRAUSS已被用于制作教育资源和分析工具。
摘要:Sonification, or conveying data using non-verbal audio, is a relatively niche but growing approach for presenting data across multiple specialist domains including astronomy, climate science, and beyond. The STRAUSS Python package aims to provide such a tool, which builds upon previous approaches to provide a powerful means to explore different ways of expressing data, with fine control over the output audio and its format. STRAUSS is a free, open source (FOSS) package, designed to allow flexible and effective sonification to be integrated into data workflows, in analogy to widely used visualisation packages. The remit of STRAUSS is broad; it is intended to be able to bridge between ad-hoc solutions for sonifying very particular datasets, and highly technical compositional and sound-design tools that are not optimised for sonification, or may have a steep learning curve. The code offers a range of approaches to sonification for a variety of contexts (e.g. science education, science communication, technical data analysis, etc). To this end, STRAUSS is packaged with a number of examples of different sonification approaches, and preset configurations to support "low-barrier, high-ceiling" approach. STRAUSS has been used to produce both educational resources and analysis tools.
【1】 Is MixIT Really Unsuitable for Correlated Sources? Exploring MixIT for Unsupervised Pre-training in Music Source Separation
标题: MixIT真的不适合相关来源吗?探索MixIT在音乐源分离中进行无监督预训练链接:https://arxiv.org/abs/2505.07631
备注:5 pages, 1 figure, 3 tables
摘要:在音乐源分离(MSS)中,获得孤立的源或茎是非常昂贵的,使得在未标记的数据上进行预训练成为一种有前途的方法。虽然在一般的声音分离中已经探索了与源无关的无监督学习,如混合不变训练(MixIT),但由于其隐含的源独立性假设,它们在MSS中很大程度上被忽视了。然而,我们假设将MixIT应用于MSS的困难来自MSS本身的不适定性质,其中茎定义是依赖于应用程序的,并且模型缺乏关于应该或不应该分离的明确知识,而不是来自高源间相关性。虽然MixIT没有假设任何源模型,并且难以处理这种模糊性,但我们的初步实验表明,它仍然可以在一定程度上分离乐器,这表明它具有无监督预训练的潜力。受这些见解的启发,本研究调查了MSS基于MixIT的预训练。我们首先使用MixIT对来自免费音乐档案的未标记数据进行预训练,然后在监督下在MUSDB 18上对其进行微调。使用最先进的MSS模型之一的带分裂TF-Locoformer,我们证明了基于MixIT的预训练从零开始就提高了训练的性能。
摘要:In music source separation (MSS), obtaining isolated sources or stems is highly costly, making pre-training on unlabeled data a promising approach. Although source-agnostic unsupervised learning like mixture-invariant training (MixIT) has been explored in general sound separation, they have been largely overlooked in MSS due to its implicit assumption of source independence. We hypothesize, however, that the difficulty of applying MixIT to MSS arises from the ill-posed nature of MSS itself, where stem definitions are application-dependent and models lack explicit knowledge of what should or should not be separated, rather than from high inter-source correlation. While MixIT does not assume any source model and struggles with such ambiguities, our preliminary experiments show that it can still separate instruments to some extent, suggesting its potential for unsupervised pre-training. Motivated by these insights, this study investigates MixIT-based pre-training for MSS. We first pre-train a model on in-the-wild, unlabeled data from the Free Music Archive using MixIT, and then fine-tune it on MUSDB18 with supervision. Using the band-split TF-Locoformer, one of the state-of-the-art MSS models, we demonstrate that MixIT-based pre-training improves the performance over training from scratch.
【2】 Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
标题: 扩散责任:分析生成性文本到音频扩散模型的能耗链接:https://arxiv.org/abs/2505.07615
摘要:文本到音频模型最近已经成为从文本描述生成声音的强大技术。然而,它们的高计算需求引起了对能源消耗和环境影响的担忧。在本文中,我们进行了7个国家的最先进的文本到音频的扩散为基础的生成模型的能量使用的分析,评估在多大程度上生成参数的变化影响在推理时的能量消耗。我们还旨在通过考虑所有选定模型的帕累托最优解决方案来确定音频质量和能耗之间的最佳平衡。我们的研究结果为性能和环境影响之间的权衡提供了见解,有助于开发更高效的生成音频模型。
摘要:Text-to-audio models have recently emerged as a powerful technology for generating sound from textual descriptions. However, their high computational demands raise concerns about energy consumption and environmental impact. In this paper, we conduct an analysis of the energy usage of 7 state-of-the-art text-to-audio diffusion-based generative models, evaluating to what extent variations in generation parameters affect energy consumption at inference time. We also aim to identify an optimal balance between audio quality and energy consumption by considering Pareto-optimal solutions across all selected models. Our findings provide insights into the trade-offs between performance and environmental impact, contributing to the development of more efficient generative audio models.
【3】 TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
标题: TACOS:用于语音音频预训练的时间对齐音频CaptiOnS链接:https://arxiv.org/abs/2505.07609
备注:submitted to the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025. Dataset (Zenodo): this https URL, Implementation (GitHub): this https URL
摘要:学习将音频与文本描述相关联对于一系列任务是有价值的,包括预训练、zero-shot分类、音频检索、音频字幕和文本条件音频生成。现有的对比语言音频预训练模型通常使用全局的剪辑级描述进行训练,仅提供弱的时间监督。我们假设,CLAP类语言音频模型-特别是,如果他们预计将产生帧级嵌入-可以受益于更强的时间监督。为了证实我们的假设,我们策划了一个由Freesound的大约12,000个音频记录组成的新数据集,每个音频记录都带有与音频记录中特定时间段相关的单句自由文本描述。我们使用大型语言模型来清除这些注释,删除对不可听事件、转录语音、错别字和注释者语言偏见的引用。我们进一步提出了一种逐帧对比训练策略,该策略学习将文本描述与音频记录中的时间区域对齐,并证明我们的模型在AudioSet Strong基准测试中评估时,与仅在全局字幕上训练的模型相比,具有更好的时间文本-音频对齐能力。数据集和我们的源代码分别在Zenodo和GitHub上提供。
摘要:Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text-conditioned audio generation. Existing contrastive language-audio pretrained models are typically trained using global, clip-level descriptions, which provide only weak temporal supervision. We hypothesize that CLAP-like language-audio models - particularly, if they are expected to produce frame-level embeddings - can benefit from a stronger temporal supervision. To confirm our hypothesis, we curate a novel dataset of approximately 12,000 audio recordings from Freesound, each annotated with single-sentence free-text descriptions linked to a specific temporal segment in an audio recording. We use large language models to clean these annotations by removing references to non-audible events, transcribed speech, typos, and annotator language bias. We further propose a frame-wise contrastive training strategy that learns to align text descriptions with temporal regions in an audio recording and demonstrate that our model has better temporal text-audio alignment abilities compared to models trained only on global captions when evaluated on the AudioSet Strong benchmark. The dataset and our source code are available on Zenodo and GitHub, respectively.
【4】 Collection: Datasets from AFAR Challenge
标题: 收集:AFAR挑战赛的数据集链接:https://arxiv.org/abs/2505.06823
备注:Submitted to IEEE Data Descriptions
摘要:本文介绍了一个全面的真实世界和数字双胞胎(DT)数据集,作为寻找漫游者(AFAR)挑战赛的一部分,由NSF高级无线空中实验和研究平台(AERPAW)测试平台组织,并在北卡罗来纳州罗利市的Lake Wheeler Field托管。AFAR挑战赛是一项涉及五个入围大学团队的竞赛,重点是促进无人机辅助射频(RF)源定位的创新。参与团队的任务是设计无人机飞行轨迹和定位算法,以检测隐藏的无人驾驶地面车辆(UGV)的位置,也称为漫游车,发射由GNU Radio生成的无线探测信号。竞赛的结构是首先在DT环境中评估解决方案,然后在AERPAW的室外无线测试平台中进行部署和测试。对于每个团队,UGV被放置在三个不同的位置,总共产生了30个数据集,其中15个在DT模拟环境中收集,15个在物理室外测试平台中收集。每个数据集包含接收信号强度(RSS)、接收信号质量(RSQ)、GPS坐标、无人机速度和无人机方位(滚转、俯仰和偏航)的时间同步测量。数据按团队、环境(DT和真实世界)和UGV位置组织到结构化文件夹中。该数据集支持无人机辅助射频源定位、空对地(A2G)无线传播建模、轨迹优化、信号预测、自主导航和DT验证等研究。该数据集从真实世界的实验中收集了大约30万个时间同步的样本,为训练和评估深度学习(DL)模型提供了坚实的基础。总体而言,AFAR数据集是推进无人机无线通信和传感系统中强大的现实解决方案的宝贵资源。
摘要:This paper presents a comprehensive real-world and Digital Twin (DT) dataset collected as part of the Find A Rover (AFAR) Challenge, organized by the NSF Aerial Experimentation and Research Platform for Advanced Wireless (AERPAW) testbed and hosted at the Lake Wheeler Field in Raleigh, North Carolina. The AFAR Challenge was a competition involving five finalist university teams, focused on promoting innovation in UAV-assisted radio frequency (RF) source localization. Participating teams were tasked with designing UAV flight trajectories and localization algorithms to detect the position of a hidden unmanned ground vehicle (UGV), also referred to as a rover, emitting wireless probe signals generated by GNU Radio. The competition was structured to evaluate solutions in a DT environment first, followed by deployment and testing in AERPAW's outdoor wireless testbed. For each team, the UGV was placed at three different positions, resulting in a total of 30 datasets, 15 collected in a DT simulation environment and 15 in a physical outdoor testbed. Each dataset contains time-synchronized measurements of received signal strength (RSS), received signal quality (RSQ), GPS coordinates, UAV velocity, and UAV orientation (roll, pitch, and yaw). Data is organized into structured folders by team, environment (DT and real-world), and UGV location. The dataset supports research in UAV-assisted RF source localization, air-to-ground (A2G) wireless propagation modeling, trajectory optimization, signal prediction, autonomous navigation, and DT validation. With approximately 300k time-synchronized samples collected from real-world experiments, the dataset provides a substantial foundation for training and evaluating deep learning (DL) models. Overall, the AFAR dataset serves as a valuable resource for advancing robust, real-world solutions in UAV-enabled wireless communications and sensing systems.
【5】 RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
标题: RADE:一种用于通过HF无线电频道传输语音的神经编解码器链接:https://arxiv.org/abs/2505.06671
备注:5 pages
摘要:语音压缩通常用于在诸如移动电话和双向一键通(PTT)无线电之类的应用中通过无线电信道发送语音。在经典系统中,语音编解码器与前向纠错、调制和无线电硬件相结合。在本文中,我们描述了一个自动编码器,取代了许多传统的信号处理元件与神经网络。编码器采用声码器特征集(短期频谱、音调、语音),并产生离散时间但连续值的正交幅度调制(QAM)符号。我们使用正交频域复用(OFDM)在高频(HF)无线电信道上发送和接收这些符号。解码器将接收到的QAM符号转换为适合于合成的声码器特征。自编码器已被训练为对加性高斯噪声和多径信道损伤具有鲁棒性,同时保持小于1 dB的峰均功率比(PAPR)。通过模拟和真实世界的高频无线电信道,我们已经实现了输出语音清晰度,明显超过了现有的模拟和数字无线电系统的信噪比范围。
摘要:Speech compression is commonly used to send voice over radio channels in applications such as mobile telephony and two-way push-to-talk (PTT) radio. In classical systems, the speech codec is combined with forward error correction, modulation and radio hardware. In this paper we describe an autoencoder that replaces many of the traditional signal processing elements with a neural network. The encoder takes a vocoder feature set (short term spectrum, pitch, voicing), and produces discrete time, but continuously valued quadrature amplitude modulation (QAM) symbols. We use orthogonal frequency domain multiplexing (OFDM) to send and receive these symbols over high frequency (HF) radio channels. The decoder converts received QAM symbols to vocoder features suitable for synthesis. The autoencoder has been trained to be robust to additive Gaussian noise and multipath channel impairments while simultaneously maintaining a Peak To Average Power Ratio (PAPR) of less than 1~dB. Over simulated and real world HF radio channels we have achieved output speech intelligibility that clearly surpasses existing analog and digital radio systems over a range of SNRs.
【6】 Spoken Language Understanding on Unseen Tasks With In-Context Learning
标题: 通过上下文学习对不可见任务的口语理解链接:https://arxiv.org/abs/2505.07731
摘要:口语理解(SLU)任务涉及各种技能,这些技能可以探测模型的信息提取,分类和/或生成能力。在这种情况下,特定于任务的训练数据可能并不总是可用的。而传统的特定于任务的SLU模型无法满足这样的要求,语音-文本大语言模型(LLM)提供了一个有前途的替代与新兴的能力。然而,开箱即用,我们的评估表明,零/Few-Shot性能突出的开源语音文本LLM SLU任务是不达标的。在本文中,我们介绍了一种新的方法,强大的任务无关的微调使用随机类标签。通过这种微调,我们说明了语音文本LLM在看不见的任务上的性能比标准方法有了显着提高。关键是,所提出的方法避免了在语音文本LLM中启用新任务的特定于任务的数据注释的要求。
摘要:Spoken language understanding (SLU) tasks involve diverse skills that probe the information extraction, classification and/or generation capabilities of models. In this setting, task-specific training data may not always be available. While traditional task-specific SLU models are unable to cater to such requirements, the speech-text large language models (LLMs) offer a promising alternative with emergent abilities. However, out of-the-box, our evaluations indicate that the zero/few-shot performance of prominent open-source speech-text LLMs on SLU tasks are not up to the mark. In this paper, we introduce a novel approach to robust task-agnostic fine-tuning using randomized class labels. With this proposed fine-tuning, we illustrate that the performance of the speech-text LLMs on an unseen task is significantly improved over standard approaches. Critically, the proposed approach avoids the requirement of task-specific data annotations for enabling new tasks in speech-text LLMs.
【7】 Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
标题: 轻量级端到端文本到语音合成,适用于低资源设备上应用链接:https://arxiv.org/abs/2505.07701
备注:Published as a conference paper at SSW 2023
摘要:最近的工作表明,以端到端(E2 E)方式直接从文本建模原始波形,比基于级联或两阶段方法的传统神经文本到语音(TTS)系统产生更自然的语音。然而,当前的E2 E最先进的模型计算复杂且消耗内存,使得它们不适合低资源场景中的实时离线设备上应用。为了解决这个问题,我们提出了一个轻量级E2 E-TTS(LE 2 E)模型,生成高质量的语音,需要最少的计算资源。我们在LJSpeech数据集上评估了所提出的模型,并表明它实现了最先进的性能,同时在模型参数方面最多可减少90\%$,在实时因素方面最多可提高10\times $。此外,我们表明,建议的E2 E培训范例实现更好的质量相比,在两个阶段的方法训练的等效架构。我们的研究结果表明,LE 2 E是一种很有前途的方法,为开发实时,高质量,低资源的TTS应用程序的设备上的应用程序。
摘要:Recent works have shown that modelling raw waveform directly from text in an end-to-end (E2E) fashion produces more natural-sounding speech than traditional neural text-to-speech (TTS) systems based on a cascade or two-stage approach. However, current E2E state-of-the-art models are computationally complex and memory-consuming, making them unsuitable for real-time offline on-device applications in low-resource scenarios. To address this issue, we propose a Lightweight E2E-TTS (LE2E) model that generates high-quality speech requiring minimal computational resources. We evaluate the proposed model on the LJSpeech dataset and show that it achieves state-of-the-art performance while being up to $90\%$ smaller in terms of model parameters and $10\times$ faster in real-time-factor. Furthermore, we demonstrate that the proposed E2E training paradigm achieves better quality compared to an equivalent architecture trained in a two-stage approach. Our results suggest that LE2E is a promising approach for developing real-time, high quality, low-resource TTS applications for on-device applications.
【8】 Multi-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge
标题: DUSE 2025挑战赛中多域音频问题回答声学内容推理链接:https://arxiv.org/abs/2505.07365
备注:Preprint. DCASE 2025 Audio QA Challenge: this https URL
摘要:我们提出了DCASE 2025挑战的任务5:一个跨越声音理解多个领域的音频问题分类(AQA)基准。该任务定义了三个QA子集(生物声学,时间音景和复杂QA),以测试音频语言模型在不同声学场景下的交互式问答。我们描述了数据集的组成(从海洋哺乳动物的声音和复杂的现实世界的剪辑),评估协议(前1名的准确性与答案洗牌的鲁棒性),和基线系统(Qwen 2-Audio-7 B,AudioFlamingo 2,Gemini-2-Flash)。对开发集的初步结果进行了比较,显示出模型和子集之间的强烈变化。这项挑战旨在将音频语言模型的音频理解和推理能力提升到人类水平的敏锐度,这对于使AI代理能够有效地感知和交互世界至关重要。
摘要:We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over diverse acoustic scenes. We describe the dataset composition (from marine mammal calls to soundscapes and complex real-world clips), the evaluation protocol (top-1 accuracy with answer-shuffling robustness), and baseline systems (Qwen2-Audio-7B, AudioFlamingo 2, Gemini-2-Flash). Preliminary results on the development set are compared, showing strong variation across models and subsets. This challenge aims to advance the audio understanding and reasoning capabilities of audio-language models toward human-level acuity, which are crucial for enabling AI agents to perceive and interact about the world effectively.
【9】 Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
标题: 神经心理声学编码的多频段频率重建链接:https://arxiv.org/abs/2505.07235
摘要:实现高保真音频压缩,同时保持不同内容的感知质量仍然是神经音频编码(NAC)的关键挑战。我们介绍MUFFIN,一个完全卷积的神经心理声学编码(NPC)框架,利用心理声学引导的多频带频率重建。其核心是一个多频带频谱残差矢量量化(MBS-RVQ)模块,该模块基于感知显著性在频带上分配比特率。这种设计实现了有效的压缩,同时使用不同的码本将说话者身份从内容中分离出来。MUFFIN结合了transformer启发的卷积骨干和修改的snake激活,以提高细粒度光谱区域的分辨率。在多个基准上的实验结果表明,MUFFIN在重建质量上始终优于现有的方法。高压缩变体以最小的损耗实现了最先进的12.5 Hz速率。MUFFIN在下游生成任务中也被证明是有效的,突出了它作为与语言模型集成的令牌表示的承诺。音频样本和代码可用。
摘要:Achieving high-fidelity audio compression while preserving perceptual quality across diverse content remains a key challenge in Neural Audio Coding (NAC). We introduce MUFFIN, a fully convolutional Neural Psychoacoustic Coding (NPC) framework that leverages psychoacoustically guided multi-band frequency reconstruction. At its core is a Multi-Band Spectral Residual Vector Quantization (MBS-RVQ) module that allocates bitrate across frequency bands based on perceptual salience. This design enables efficient compression while disentangling speaker identity from content using distinct codebooks. MUFFIN incorporates a transformer-inspired convolutional backbone and a modified snake activation to enhance resolution in fine-grained spectral regions. Experimental results on multiple benchmarks demonstrate that MUFFIN consistently outperforms existing approaches in reconstruction quality. A high-compression variant achieves a state-of-the-art 12.5 Hz rate with minimal loss. MUFFIN also proves effective in downstream generative tasks, highlighting its promise as a token representation for integration with language models. Audio samples and code are available.
【10】 On the Cost and Benefits of Training Context with Utterance or Full Conversation Training: A Comparative Stud
标题: 带直言不讳或全面对话训练的训练环境的成本和收益:比较研究链接:https://arxiv.org/abs/2505.07202
摘要:为对话设计的现代TTS系统实现了高质量的话语,但通常仍然无法公开访问。是现有的开源架构不够,还是目前的培训技术不够?本文研究了会话语境中的主要模型及其潜在行为。在NVIDIA H100上使用20个GPU小时,我们实证研究了两种方法:基于上下文的话语级训练与完整的会话训练。结果表明,基于上下文的话语训练取得了优异的MOS分数(4.3/5.0 vs 3.7/5.0),并减少了37%的训练时间,而完整的对话方法遭受说话人相似性幻觉问题。这些研究结果提供了实用的指导方针会话TTS的发展,有利于话语水平的培训与上下文条件的资源效率和输出质量。
摘要:Modern TTS systems designed for conversations achieve high-quality utterances but often remain inaccessible publicly. Are existing open-source architectures inadequate, or are current training techniques insufficient? This paper investigates prominent models and their underlying behaviors regarding conversational context. Using 20 GPU-hours on an NVIDIA H100, we empirically examine two approaches: context-based utterance-level training versus full conversation training. Results demonstrate that context-based utterance training achieves superior MOS scores (4.3/5.0 vs 3.7/5.0) and reduces training time by 37%, while full conversation approaches suffer from speaker similarity hallucination issues. These findings provide practical guidelines for conversational TTS development, favoring utterance-level training with contextual conditioning for both resource efficiency and output quality.
【11】 Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
标题: 架起耳朵和眼睛:在可见声音识别中分析人类的音频和视觉大型语言模型并通过跨模式蒸馏减少他们的感觉差距链接:https://arxiv.org/abs/2505.06803
摘要:音频大语言模型(LLM)被认为是识别声音对象的专家,但它们相对于其他感官模式(如视觉或视听LLM)中的LLM以及使用耳朵、眼睛或两者的人类的性能仍然未被探索。为了研究这一点,我们系统地评估了音频,视觉和视听LLM,特别是Qwen 2-Audio,Qwen 2-VL和Qwen2.5-Omni,对人类识别来自仅音频,无声视频或有声视频输入的不同类别的声音对象。我们发现Qwen 2-Audio和Qwen 2-VL之间的性能差距与人耳和眼睛之间的感官差异相似。为了缩小这一差距,我们引入了一个跨模态蒸馏框架,其中一种模态的LLM作为教师,另一种作为学生,通过启发式模型预测声音类中的知识转移对学生更具挑战性。从Qwen 2-VL到Qwen 2-Audio的双向蒸馏,反之亦然,会带来显著的改进,特别是在具有挑战性的课程中。这项工作突出了感觉差距LLM从人类对齐的角度来看,并提出了一个原则性的方法,以提高模态特定的感知多模态LLM。
摘要:Audio large language models (LLMs) are considered experts at recognizing sound objects, yet their performance relative to LLMs in other sensory modalities, such as visual or audio-visual LLMs, and to humans using their ears, eyes, or both remains unexplored. To investigate this, we systematically evaluate audio, visual, and audio-visual LLMs, specifically Qwen2-Audio, Qwen2-VL, and Qwen2.5-Omni, against humans in recognizing sound objects of different classes from audio-only, silent video, or sounded video inputs. We uncover a performance gap between Qwen2-Audio and Qwen2-VL that parallels the sensory discrepancy between human ears and eyes. To reduce this gap, we introduce a cross-modal distillation framework, where an LLM in one modality serves as the teacher and another as the student, with knowledge transfer in sound classes predicted as more challenging to the student by a heuristic model. Distillation in both directions, from Qwen2-VL to Qwen2-Audio and vice versa, leads to notable improvements, particularly in challenging classes. This work highlights the sensory gap in LLMs from a human-aligned perspective and proposes a principled approach to enhancing modality-specific perception in multimodal LLMs.
【12】 Beyond Identity: A Generalizable Approach for Deepfake Audio Detection
标题: 超越身份:Deepfake音频检测的一种通用方法链接:https://arxiv.org/abs/2505.06766
备注:Submitted to IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM)
摘要:Deepfake音频对数字安全构成了越来越大的威胁,因为它有可能引发社会工程、欺诈和身份滥用。然而,现有的检测模型由于隐式身份泄漏而在数据集之间具有较差的泛化能力,其中模型无意中学习说话者特定的特征而不是操纵伪像。据我们所知,这是第一项明确分析和解决音频deepfake检测领域身份泄露的研究。这项工作提出了一个独立于身份的音频deepfake检测框架,通过鼓励模型专注于伪造特定的伪影,而不是过度拟合扬声器特征,来减轻身份泄露。我们的方法利用人工智能检测模块(ADM)在时域和频域中隔离合成工件,增强跨数据集的泛化。我们引入了新的动态伪影生成技术,包括频域交换,时域操作,和背景噪声增强,以加强学习的不变特征。在ASVspoof 2019、ADD 2022、FoR和In-The-Wild数据集上进行的大量实验表明,所提出的ADM增强模型的F1得分为0.230(ADD 2022)、0.604(FoR)和0.813(In-The-Wild),始终优于基线。动态频率交换被证明是在不同条件下最有效的策略。这些发现强调了基于伪影的学习在减轻隐式身份泄露以进行更普遍的音频deepfake检测方面的价值。
摘要:Deepfake audio presents a growing threat to digital security, due to its potential for social engineering, fraud, and identity misuse. However, existing detection models suffer from poor generalization across datasets, due to implicit identity leakage, where models inadvertently learn speaker-specific features instead of manipulation artifacts. To the best of our knowledge, this is the first study to explicitly analyze and address identity leakage in the audio deepfake detection domain. This work proposes an identity-independent audio deepfake detection framework that mitigates identity leakage by encouraging the model to focus on forgery-specific artifacts instead of overfitting to speaker traits. Our approach leverages Artifact Detection Modules (ADMs) to isolate synthetic artifacts in both time and frequency domains, enhancing cross-dataset generalization. We introduce novel dynamic artifact generation techniques, including frequency domain swaps, time domain manipulations, and background noise augmentation, to enforce learning of dataset-invariant features. Extensive experiments conducted on ASVspoof2019, ADD 2022, FoR, and In-The-Wild datasets demonstrate that the proposed ADM-enhanced models achieve F1 scores of 0.230 (ADD 2022), 0.604 (FoR), and 0.813 (In-The-Wild), consistently outperforming the baseline. Dynamic Frequency Swap proves to be the most effective strategy across diverse conditions. These findings emphasize the value of artifact-based learning in mitigating implicit identity leakage for more generalizable audio deepfake detection.
【13】 TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
标题: TS-SURB:语音自我监督学习模型的目标语音处理基准链接:https://arxiv.org/abs/2505.06660
备注:Accepted at ICASSP 2025
摘要:自监督学习(SSL)模型有显着先进的语音处理任务,已经提出了几个基准来验证其有效性。然而,以前的基准测试主要集中在单个说话者的情况下,在嘈杂的多说话者条件下对目标说话者任务的探索较少-这是一个更具挑战性但实用的案例。在本文中,我们介绍了目标说话人语音处理通用性能基准(TS-SUPERB),其中包括四个广泛认可的目标说话人处理任务,需要识别目标说话人和提取信息的语音混合。在我们的基准测试中,从注册语音中提取的说话人嵌入被用作条件下游模型的线索。基准测试结果揭示了在目标说话人场景中评估SSL模型的重要性,表明性能不能轻易地从相关的单说话人任务中推断出来。此外,通过使用一个统一的基于SSL的目标语音编码器,由一个扬声器编码器和提取器模块,我们还研究跨TS任务的联合优化,以利用互信息,并证明其有效性。
摘要:Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in noisy, multi-talker conditions -- a more challenging yet practical case. In this paper, we introduce the Target-Speaker Speech Processing Universal Performance Benchmark (TS-SUPERB), which includes four widely recognized target-speaker processing tasks that require identifying the target speaker and extracting information from the speech mixture. In our benchmark, the speaker embedding extracted from enrollment speech is used as a clue to condition downstream models. The benchmark result reveals the importance of evaluating SSL models in target speaker scenarios, demonstrating that performance cannot be easily inferred from related single-speaker tasks. Moreover, by using a unified SSL-based target speech encoder, consisting of a speaker encoder and an extractor module, we also investigate joint optimization across TS tasks to leverage mutual information and demonstrate its effectiveness.
机器翻译由腾讯交互翻译提供,仅供参考
