微信公众号:arXiv_Daily
cs.SD语音
【1】CAK: Emergent Audio Effects from Minimal Deep Learning
标题:CAK:来自最小深度学习的紧急音频效果
链接:https://arxiv.org/abs/2508.02643
备注:8 pages, 3 figures, code and other resources at this https URL
摘要:我们证明了一个3x 3卷积核可以在来自个性化语料库的200个样本上训练时产生紧急音频效果。我们通过两个关键技术实现了这一点:(1)条件感知内核(CAK),其中输出=输入+(学习模式x控制),具有支持零控制身份保留的软门机制;(2)AuGAN(审计GAN),它从“这是真的吗?“to“是否应用了所请求的值?“我们的网络不是学习生成或检测fortunes,而是合作验证控制应用程序,发现独特的转换。学习的内核表现出对角结构,创建频率相关的时间偏移,能够根据输入特性产生音乐效果。我们的研究结果显示了对抗性训练从最少的数据中发现音频转换的潜力,从而为效果设计提供了新的方法。
摘要:We demonstrate that a single 3x3 convolutional kernel can produce emergent audio effects when trained on 200 samples from a personalized corpus. We achieve this through two key techniques: (1) Conditioning Aware Kernels (CAK), where output = input + (learned_pattern x control), with a soft-gate mechanism supporting identity preservation at zero control; and (2) AuGAN (Audit GAN), which reframes adversarial training from "is this real?" to "did you apply the requested value?" Rather than learning to generate or detect forgeries, our networks cooperate to verify control application, discovering unique transformations. The learned kernel exhibits a diagonal structure creating frequency-dependent temporal shifts that are capable of producing musical effects based on input characteristics. Our results show the potential of adversarial training to discover audio transformations from minimal data, enabling new approaches to effect design.
【2】Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
标题:可靠的音频Deepfake属性和模型识别:基于多级自动编码器的框架
链接:https://arxiv.org/abs/2508.02521
摘要:音频深度伪造的扩散对数字通信的信任构成了越来越大的威胁。虽然检测方法已经取得了进展,但将音频deepfake归因于其源模型仍然是一个尚未探索但至关重要的挑战。在本文中,我们介绍了LAVA(Layered Architecture for Voice Attribution),这是一种用于音频deepfake检测和模型识别的分层框架,它利用了仅在假音频上训练的卷积自动编码器提取的注意力增强的潜在表示。两个专门的分类器对这些特征进行操作:音频Deepfake属性(ADA),用于识别生成技术,以及音频Deepfake模型识别(ADMR),用于识别特定的生成模型实例。为了提高开集条件下的鲁棒性,我们采用了基于置信度的拒绝阈值。在ASVspoof2021、FakeOrReal和CodecFake上的实验显示了强大的性能:ADA分类器在所有数据集上的F1得分超过95%,ADMR模块在六个类上达到了96.31%的宏F1。对ASVpoof2019 LA的未知攻击和错误传播分析的额外测试证实了LAVA的鲁棒性和可靠性。该框架通过引入一种有监督的方法来在开放集条件下进行deepfake归因和模型识别,并在公共基准上进行验证,并伴随着公开发布的模型和代码,从而推动了该领域的发展。模型和代码可在https://www.github.com/adipiz99/lava-framework上获得。
摘要:The proliferation of audio deepfakes poses a growing threat to trust in digital communications. While detection methods have advanced, attributing audio deepfakes to their source models remains an underexplored yet crucial challenge. In this paper we introduce LAVA (Layered Architecture for Voice Attribution), a hierarchical framework for audio deepfake detection and model recognition that leverages attention-enhanced latent representations extracted by a convolutional autoencoder trained solely on fake audio. Two specialized classifiers operate on these features: Audio Deepfake Attribution (ADA), which identifies the generation technology, and Audio Deepfake Model Recognition (ADMR), which recognize the specific generative model instance. To improve robustness under open-set conditions, we incorporate confidence-based rejection thresholds. Experiments on ASVspoof2021, FakeOrReal, and CodecFake show strong performance: the ADA classifier achieves F1-scores over 95% across all datasets, and the ADMR module reaches 96.31% macro F1 across six classes. Additional tests on unseen attacks from ASVpoof2019 LA and error propagation analysis confirm LAVA's robustness and reliability. The framework advances the field by introducing a supervised approach to deepfake attribution and model recognition under open-set conditions, validated on public benchmarks and accompanied by publicly released models and code. Models and code are available at https://www.github.com/adipiz99/lava-framework.
【3】Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
标题:绘制语音情感识别深度学习15年的进展:一项复制研究
链接:https://arxiv.org/abs/2508.02448
备注:Code: this https URL Submitted for review
摘要:语音情感识别(SER)长期以来一直受益于深度学习方法的采用。更深的模型-具有更多的层次和更多的可训练参数-通常被SER社区认为是“更好的”。这就提出了一个问题--与早期的迭代相比,现代的深度神经网络好多少?除此之外,更重要的问题是如何向前迈进,这个问题仍然像以往一样尖锐。SER是远远没有解决的问题,因此,确定未来研究的最突出的途径是至关重要的。在目前的贡献,我们试图量化的进展,在15年的研究开始,具有里程碑意义的2009年INTERSPEECH情感挑战。我们进行了大规模的模型架构的调查,跨越基于音频的模型,依赖于语音输入和文本的模型,只依赖于transmittance。我们的研究结果指向收益递减和平台后,最近推出的Transformer架构。此外,我们展示了如何进步的看法是有条件的特定选择的模型进行比较。我们的研究结果对SER研究的最新水平和前进道路具有重要影响
摘要:Speech emotion recognition (SER) has long benefited from the adoption of deep learning methodologies. Deeper models -- with more layers and more trainable parameters -- are generally perceived as being `better' by the SER community. This raises the question -- \emph{how much better} are modern-era deep neural networks compared to their earlier iterations? Beyond that, the more important question of how to move forward remains as poignant as ever. SER is far from a solved problem; therefore, identifying the most prominent avenues of future research is of paramount importance. In the present contribution, we attempt a quantification of progress in the 15 years of research beginning with the introduction of the landmark 2009 INTERSPEECH Emotion Challenge. We conduct a large scale investigation of model architectures, spanning both audio-based models that rely on speech inputs and text-baed models that rely solely on transcriptions. Our results point towards diminishing returns and a plateau after the recent introduction of transformer architectures. Moreover, we demonstrate how perceptions of progress are conditioned on the particular selection of models that are compared. Our findings have important repercussions about the state-of-the-art in SER research and the paths forward
【4】Inference-time Scaling for Diffusion-based Audio Super-resolution
标题:基于扩散的音频超分辨率的推理时缩放
链接:https://arxiv.org/abs/2508.02391
摘要:扩散模型在生成任务中取得了显着的成功,包括音频超分辨率(SR)。在电影后期制作和专辑母版制作等许多应用中,大量的计算预算可用于实现卓越的音频质量。然而,虽然现有的扩散方法通常增加采样步骤以提高质量,但性能仍然受到采样过程的随机性的根本限制,导致高方差和质量受限的输出。在这里,而不是简单地增加采样步骤的数量,我们提出了一个不同的范例,通过推理时间缩放SR,探索多个解决方案的轨迹在采样过程中。设计了不同的任务验证器,介绍了随机搜索和零阶搜索两种搜索算法。通过验证器-算法组合积极引导高维解空间的探索,我们可以实现更强大和更高质量的输出。通过在不同的音频域(语音,音乐,音效)和频率范围内进行广泛的验证,我们展示了一致的性能增益,在美学上实现了高达9.70%的改进,在扬声器相似度上实现了5.88%的改进,在单词错误率上实现了15.20%的改进,在语音SR从4kHz到24 kHz的频谱距离上实现了46.98%的改进,展示了我们方法的有效性。音频样本可在https://racerk.github.io/tt-scale-audiosr/上获得。
摘要:Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are available for achieving superior audio quality. However, while existing diffusion approaches typically increase sampling steps to improve quality, the performance remains fundamentally limited by the stochastic nature of the sampling process, leading to high-variance and quality-limited outputs. Here, rather than simply increasing the number of sampling steps, we propose a different paradigm through inference-time scaling for SR, which explores multiple solution trajectories during the sampling process. Different task-specific verifiers are developed, and two search algorithms, including the random search and zero-order search for SR, are introduced. By actively guiding the exploration of the high-dimensional solution space through verifier-algorithm combinations, we enable more robust and higher-quality outputs. Through extensive validation across diverse audio domains (speech, music, sound effects) and frequency ranges, we demonstrate consistent performance gains, achieving improvements of up to 9.70% in aesthetics, 5.88% in speaker similarity, 15.20% in word error rate, and 46.98% in spectral distance for speech SR from 4kHz to 24kHz, showcasing the effectiveness of our approach. Audio samples are available at: https://racerk.github.io/tt-scale-audiosr/.
【5】Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
标题:通过语音分析检测COPD:丹麦语音和机器学习方法的数据集
链接:https://arxiv.org/abs/2508.02354
摘要:慢性阻塞性肺疾病(COPD)是一种严重的、使人衰弱的疾病,影响着全世界数百万人。使用非侵入性手段对其进行早期检测可以实现预防性干预,改善生活质量和患者预后,最近发现言语是一种有价值的生物标志物。然而,它在不同语言群体中的有效性仍有待观察。为此,音频数据收集96丹麦参与者进行三个语音任务(阅读,咳嗽,持续元音)。一半的参与者被诊断患有不同程度的COPD,另一半组成健康对照组。随后,我们使用openSMILE特征和学习的x向量嵌入研究了不同的基线模型。我们使用openSMILE特征和logistic回归获得了67%的最佳准确率。我们的研究结果支持基于语音的分析作为未来COPD医疗保健解决方案的一部分的非侵入性,远程和可扩展的筛查工具的潜力。
摘要:Chronic Obstructive Pulmonary Disease (COPD) is a serious and debilitating disease affecting millions around the world. Its early detection using non-invasive means could enable preventive interventions that improve quality of life and patient outcomes, with speech recently shown to be a valuable biomarker. Yet, its validity across different linguistic groups remains to be seen. To that end, audio data were collected from 96 Danish participants conducting three speech tasks (reading, coughing, sustained vowels). Half of the participants were diagnosed with different levels of COPD and the other half formed a healthy control group. Subsequently, we investigated different baseline models using openSMILE features and learnt x-vector embeddings. We obtained a best accuracy of 67% using openSMILE features and logistic regression. Our findings support the potential of speech-based analysis as a non-invasive, remote, and scalable screening tool as part of future COPD healthcare solutions.
【6】StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation
标题:StutterCut:不确定性引导的标准化切割,用于不流利分割
链接:https://arxiv.org/abs/2508.02255
备注:Accepted in Interspeech 2025
摘要:检测和分割不流利是有效的语音治疗和实时反馈的关键。然而,大多数方法仅在话语水平上对不流利进行分类。我们介绍StutterCut,一个半监督框架,制定不流畅分割作为一个图形划分问题,其中语音嵌入重叠窗口表示为图形节点。我们使用在弱(话语级)标签上训练的伪Oracle分类器来改进节点之间的连接,其影响力由Monte Carlo dropout的不确定性度量控制。此外,我们扩展了弱标记的FluencyBank数据集,将帧级不流畅边界的四种不流畅类型。与合成数据集相比,这提供了一个更现实的基准。在真实和合成数据集上的实验表明,StutterCut优于现有方法,实现了更高的F1分数和更精确的口吃发作检测。
摘要:Detecting and segmenting dysfluencies is crucial for effective speech therapy and real-time feedback. However, most methods only classify dysfluencies at the utterance level. We introduce StutterCut, a semi-supervised framework that formulates dysfluency segmentation as a graph partitioning problem, where speech embeddings from overlapping windows are represented as graph nodes. We refine the connections between nodes using a pseudo-oracle classifier trained on weak (utterance-level) labels, with its influence controlled by an uncertainty measure from Monte Carlo dropout. Additionally, we extend the weakly labelled FluencyBank dataset by incorporating frame-level dysfluency boundaries for four dysfluency types. This provides a more realistic benchmark compared to synthetic datasets. Experiments on real and synthetic datasets show that StutterCut outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection.
【7】WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
标题:WhiSQA:使用Whisper编码器功能的非侵入性语音质量预测
链接:https://arxiv.org/abs/2508.02210
备注:Accepted at SPECOM 2025
摘要:近年来,人们在开发基于神经网络的SQ预测器方面做出了重大的研究努力。虽然主要目标是开发非侵入性,即~无参考的度量来评估SE系统的性能,最近的工作还研究了下游语音任务的损失函数内的神经SQ预测器的直接推断。为了帮助SQ预测器的训练,已经创建了具有相应的人类质量标签的几个大型音频数据集。最近在这一领域的工作表明,来自大型无监督或半监督基础语音模型的语音表示是神经SQ预测的有用输入特征表示。在这项工作中,提出了一种新的和强大的SQ预测的基础上提取的ASR模型的特征表示,发现是一个强大的输入功能的SQ预测任务。该系统实现了更高的相关性与人类MOS评级比最近的方法在所有NISQA测试集,并显示出显着更好的域适应性相比,常用的DNSMOS度量。
摘要:There has been significant research effort developing neural-network-based predictors of SQ in recent years. While a primary objective has been to develop non-intrusive, i.e.~reference-free, metrics to assess the performance of SE systems, recent work has also investigated the direct inference of neural SQ predictors within the loss function of downstream speech tasks. To aid in the training of SQ predictors, several large datasets of audio with corresponding human labels of quality have been created. Recent work in this area has shown that speech representations derived from large unsupervised or semi-supervised foundational speech models are useful input feature representations for neural SQ prediction. In this work, a novel and robust SQ predictor is proposed based on feature representations extracted from an ASR model, found to be a powerful input feature for the SQ prediction task. The proposed system achieves higher correlation with human MOS ratings than recent approaches on all NISQA test sets and shows significantly better domain adaption compared to the commonly used DNSMOS metric.
【8】Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers
标题:隐藏在噪音中:通过潜在声学模式触发器揭开音频LLM对齐中的后门
链接:https://arxiv.org/abs/2508.02175
摘要:随着音频大语言模型(ALLM)成为语音处理的强大工具,其安全性问题迫切需要关注。虽然大量的研究已经探索了文本和视觉安全,但音频的独特特征带来了重大挑战。本文首先调查:ALLM容易受到利用声学触发器的后门攻击?针对这个问题,我们引入了隐藏在噪声中(HIN),这是一种新颖的后门攻击框架,旨在利用微妙的音频特定功能。HIN将声学修改应用于原始音频波形,例如改变时间动态和频谱定制噪声的策略注入。这些变化引入了ALLM的声学特征编码器捕获的一致模式,在音频流中嵌入了鲁棒的触发器。为了评估ALLM对基于音频特征的触发器的鲁棒性,我们开发了AudioSafe基准,评估了九种不同的风险类型。在AudioSafe和三个已建立的安全数据集上进行的广泛实验揭示了现有ALLM中的关键漏洞:(I)环境噪声和语音速率变化等音频特征的平均攻击成功率超过90%。(II)ALLM表现出显着的灵敏度差异,在声学功能,特别是显示最小的响应量作为触发器,和(III)中毒的样品夹杂物只会导致边际损失曲线波动,突出了攻击的隐身。
摘要:As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio's distinct characteristics present significant challenges. This paper first investigates: Is ALLM vulnerable to backdoor attacks exploiting acoustic triggers? In response to this issue, we introduce Hidden in the Noise (HIN), a novel backdoor attack framework designed to exploit subtle, audio-specific features. HIN applies acoustic modifications to raw audio waveforms, such as alterations to temporal dynamics and strategic injection of spectrally tailored noise. These changes introduce consistent patterns that an ALLM's acoustic feature encoder captures, embedding robust triggers within the audio stream. To evaluate ALLM robustness against audio-feature-based triggers, we develop the AudioSafe benchmark, assessing nine distinct risk types. Extensive experiments on AudioSafe and three established safety datasets reveal critical vulnerabilities in existing ALLMs: (I) audio features like environment noise and speech rate variations achieve over 90% average attack success rate. (II) ALLMs exhibit significant sensitivity differences across acoustic features, particularly showing minimal response to volume as a trigger, and (III) poisoned sample inclusion causes only marginal loss curve fluctuations, highlighting the attack's stealth.
【9】Unsupervised Multi-channel Speech Dereverberation via Diffusion
标题:无监督多通道扩散语音去回响
链接:https://arxiv.org/abs/2508.02071
摘要:本文研究了多通道单说话人盲去混响问题,即利用多通道混合信号恢复纯净的无回声语音。为了解决这个问题,我们提出了USD-DPS,{U}无监督的{S}peech {D}通过{D} ffffffective {P} posterior {S}混响。USD-DPS使用无条件干净语音扩散模型作为强先验,通过后验采样来解决问题。在每个扩散采样步骤中,我们估计所有麦克风通道的房间脉冲响应(RIR),这些响应进一步用于实施扩散指导的多通道混合一致性约束。对于多通道RIR估计,我们估计参考通道RIR通过优化子带RIR信号模型的RIR参数,与亚当优化器。我们使用前向卷积预测(FCP)来解析地估计非参考信道的RIR。我们发现,这种组合提供了采样效率和RIR先验建模之间的良好平衡,这表明无监督去混响方法的性能优越。https://usddps.github.io/USDDPS_demo/提供了音频演示页面。
摘要:We consider the problem of multi-channel single-speaker blind dereverberation, where multi-channel mixtures are used to recover the clean anechoic speech. To solve this problem, we propose USD-DPS, {U}nsupervised {S}peech {D}ereverberation via {D}iffusion {P}osterior {S}ampling. USD-DPS uses an unconditional clean speech diffusion model as a strong prior to solve the problem by posterior sampling. At each diffusion sampling step, we estimate all microphone channels' room impulse responses (RIRs), which are further used to enforce a multi-channel mixture consistency constraint for diffusion guidance. For multi-channel RIR estimation, we estimate reference-channel RIR by optimizing RIR parameters of a sub-band RIR signal model, with the Adam optimizer. We estimate non-reference channels' RIRs analytically using forward convolutive prediction (FCP). We found that this combination provides a good balance between sampling efficiency and RIR prior modeling, which shows superior performance among unsupervised dereverberation approaches. An audio demo page is provided in https://usddps.github.io/USDDPS_demo/.
【10】Marco-Voice Technical Report
标题:Marco-Voice技术报告
链接:https://arxiv.org/abs/2508.02038
备注:Technical Report
摘要:本文提出了一个多功能的语音合成系统,它集成了语音克隆和情感控制语音合成在一个统一的框架。这项工作的目标是解决长期存在的挑战,实现高度表达,可控和自然的语音生成,忠实地保留扬声器身份在不同的语言和情感背景。我们的方法引入了一个有效的说话人情感解纠缠机制与批量对比学习,使独立操纵的发言人身份和eemotional风格,以及旋转的情绪嵌入集成方法平滑的情绪控制。为了支持全面的训练和评估,我们构建了CSEMOTIONS,这是一个高质量的情感语音数据集,包含来自七个情感类别的六个专业演讲者的10小时普通话语音。大量实验表明,我们的系统Marco-Voice在客观和主观指标方面都取得了实质性改进。通过对MarcoVoice语音合成系统进行全面的评测和分析,结果表明,MarcoVoice在语音清晰度和情感丰富度方面具有竞争力,在表达性神经语音合成领域取得了实质性的进步。
摘要:This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. Our approach introduces an effective speaker-emotion disentanglement mechanism with in-batch contrastive learning, enabling independent manipulation of speaker identity and eemotional style, as well as rotational emotional embedding integration method for smooth emotion control. To support comprehensive training and evaluation, we construct CSEMOTIONS, a high-quality emotional speech dataset containing 10 hours of Mandarin speech from six professional speakers across seven emotional categories. Extensive experiments demonstrate that our system, Marco-Voice, achieves substantial improvements in both objective and subjective metrics. Comprehensive evaluations and analysis were conducted, results show that MarcoVoice delivers competitive performance in terms of speech clarity and emotional richness, representing a substantial advance in the field of expressive neural speech synthesis.
【11】Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
标题:通过分层边界建模本地化视听Deepfakes
链接:https://arxiv.org/abs/2508.02000
备注:Work in progress
摘要:基于内容驱动的部分操作的视听时间深度伪造定位仍然是一项极具挑战性的任务。在这种情况下,深度伪造区域通常只跨越几帧,其余大部分区域与原始区域保持相同。为了解决这个问题,我们提出了一个分层边界建模网络(HBMNet),它包括三个模块:一个视听特征编码器,提取有区别的帧级表示,一个粗建议生成器,预测候选边界区域,和一个细粒度概率生成器,使用双向边界内容概率细化这些建议。从模态的角度来看,我们通过专门的编码和融合来增强视听学习,并通过帧级监督来增强辨别力。从时间的角度来看,HBMNet集成了多尺度线索和双向边界内容关系。实验表明,编码和融合主要提高精度,而帧级监督提高召回率。每个模块(视听融合、时间尺度、双向性)都提供互补的优势,共同提高定位性能。HBMNet的性能优于BA-TFD和UMMA Former,并通过更多的训练数据显示出潜在的可扩展性。
摘要:Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest remaining identical to the original. To tackle this, we propose a Hierarchical Boundary Modeling Network (HBMNet), which includes three modules: an Audio-Visual Feature Encoder that extracts discriminative frame-level representations, a Coarse Proposal Generator that predicts candidate boundary regions, and a Fine-grained Probabilities Generator that refines these proposals using bidirectional boundary-content probabilities. From the modality perspective, we enhance audio-visual learning through dedicated encoding and fusion, reinforced by frame-level supervision to boost discriminability. From the temporal perspective, HBMNet integrates multi-scale cues and bidirectional boundary-content relationships. Experiments show that encoding and fusion primarily improve precision, while frame-level supervision boosts recall. Each module (audio-visual fusion, temporal scales, bi-directionality) contributes complementary benefits, collectively enhancing localization performance. HBMNet outperforms BA-TFD and UMMAFormer and shows improved potential scalability with more training data.
【12】Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life
标题:非言语发声及其挑战:情感、隐私、稀疏和现实生活
链接:https://arxiv.org/abs/2508.01960
摘要:非言语发声(NVV)是没有适当的语言(语义)意义,但传达内涵的简短“非词”话语-无论是这种情绪/影响或其他非语言信息。我们从一个历史性的概述开始:在过去的两个世纪里,它们是如何在心理学和语言学中得到解决的,它们后来是如何被忽视的,以及它们是如何随着情绪研究的出现而脱颖而出的。然后,我们给出了NVV的类型(形式方面)和NVV的功能的概述,以典型的NVV \textit{ah}为例。尽管NVV很有趣,但它们也带来了一系列挑战:隐私和一般道德考虑阻止它们在现实生活(私人)场景中被记录到足够的程度。孤立的,提示的(行动的)范例不一定在上下文中建模NVV;然而,这是迄今为止建模NVV时的首选策略,特别是在AI中。为了克服这些问题,我们主张采用基于语料库的方法。这保证了更真实的建模;然而,我们仍然面临着隐私和稀疏数据的问题。
摘要:Non-Verbal Vocalisations (NVVs) are short `non-word' utterances without proper linguistic (semantic) meaning but conveying connotations -- be this emotions/affects or other paralinguistic information. We start this contribution with a historic sketch: how they were addressed in psychology and linguistics in the last two centuries, how they were neglected later on, and how they came to the fore with the advent of emotion research. We then give an overview of types of NVVs (formal aspects) and functions of NVVs, exemplified with the typical NVV \textit{ah}. Interesting as they are, NVVs come, however, with a bunch of challenges that should be accounted for: Privacy and general ethical considerations prevent them of being recorded in real-life (private) scenarios to a sufficient extent. Isolated, prompted (acted) exemplars do not necessarily model NVVs in context; yet, this is the preferred strategy so far when modelling NVVs, especially in AI. To overcome these problems, we argue in favour of corpus-based approaches. This guarantees a more realistic modelling; however, we are still faced with privacy and sparse data problems.
【13】EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
标题:EgoTrigger:在全天节能智能眼镜中实现音频驱动图像捕获以增强人类记忆力
链接:https://arxiv.org/abs/2508.01915
备注:15 pages, 6 figres, 6 tables. Accepted to ISMAR 2025 as a TVCG journal paper
摘要:全天候智能眼镜可能会成为能够持续感知环境的平台,为我们的日常生活提供前所未有的帮助。然而,在执行连续感测的同时集成人类记忆增强所需的多模式AI代理,对全天使用提出了一个主要的能源效率挑战。实现这种平衡需要智能的环境感知传感器管理。我们的方法,EgoTrigger,利用麦克风的音频提示来选择性地激活功率密集型摄像头,从而实现高效的传感,同时保留对人类记忆增强的实质性效用。EgoTrigger使用轻量级音频模型(YAMNet)和自定义分类头来触发来自手-物交互(HOI)音频提示的图像捕获,例如抽屉打开或药瓶打开的声音。除了在QA-Ego 4D数据集上进行评估外,我们还介绍了人类记忆增强问答(HME-QA)数据集并进行了评估。我们的数据集包含来自全长Ego 4D视频的340个人类注释的第一人称QA对,这些视频经过精心策划,以确保它们包含音频,专注于对上下文理解和记忆至关重要的HOI时刻。我们的研究结果表明,EgoTrigger平均可以减少54%的帧,显著节省了两个耗电传感组件(例如,摄像机)和下游操作(例如,无线传输),同时在情景记忆任务的数据集上实现可比较的性能。我们相信这种情境感知触发策略代表了一个有希望的方向,可以实现节能,功能智能眼镜,能够全天使用-支持应用程序,如帮助用户回忆他们放置钥匙的位置或有关他们日常活动的信息(例如,服用药物)。
摘要:All-day smart glasses are likely to emerge as platforms capable of continuous contextual sensing, uniquely positioning them for unprecedented assistance in our daily lives. Integrating the multi-modal AI agents required for human memory enhancement while performing continuous sensing, however, presents a major energy efficiency challenge for all-day usage. Achieving this balance requires intelligent, context-aware sensor management. Our approach, EgoTrigger, leverages audio cues from the microphone to selectively activate power-intensive cameras, enabling efficient sensing while preserving substantial utility for human memory enhancement. EgoTrigger uses a lightweight audio model (YAMNet) and a custom classification head to trigger image capture from hand-object interaction (HOI) audio cues, such as the sound of a drawer opening or a medication bottle being opened. In addition to evaluating on the QA-Ego4D dataset, we introduce and evaluate on the Human Memory Enhancement Question-Answer (HME-QA) dataset. Our dataset contains 340 human-annotated first-person QA pairs from full-length Ego4D videos that were curated to ensure that they contained audio, focusing on HOI moments critical for contextual understanding and memory. Our results show EgoTrigger can use 54% fewer frames on average, significantly saving energy in both power-hungry sensing components (e.g., cameras) and downstream operations (e.g., wireless transmission), while achieving comparable performance on datasets for an episodic memory task. We believe this context-aware triggering strategy represents a promising direction for enabling energy-efficient, functional smart glasses capable of all-day use -- supporting applications like helping users recall where they placed their keys or information about their routine activities (e.g., taking medications).
【14】Generalizable Audio Deepfake Detection via Hierarchical Structure Learning and Feature Whitening in Poincaré sphere
标题:通过庞加莱球中的分层结构学习和特征白化进行可推广的音频深度伪造检测
链接:https://arxiv.org/abs/2508.01897
备注:Accepted for publication on Interspeech 2025
摘要:由于各种真实世界的欺骗攻击和域变化,音频深度伪造检测(ADD)面临着关键的泛化挑战。然而,现有的方法主要依赖于欧几里德距离,未能充分捕捉与攻击类别和域因素相关联的内在层次结构。为了解决这些问题,我们设计了一个新的框架点HierNet构造域不变的庞加莱球的层次表示。Poin-HierNet包括三个关键组件:1)Poincar\'e Prototype Learning(PPL),使用多个数据原型对齐样本特征并捕获人类标签之外的多级层次结构; 2)层次结构学习(HSL)利用顶部原型从数据原型建立树状层次结构; 3)Poincar特征白化(PFW)通过应用特征白化抑制领域敏感特征来增强领域不变性。我们在四个数据集上评估了我们的方法:ASVspoof 2019 LA,ASVspoof 2021 LA,ASVspoof 2021 DF和In-The-Wild。实验结果表明,点层次网络超过了国家的最先进的方法在等错误率。
摘要:Audio deepfake detection (ADD) faces critical generalization challenges due to diverse real-world spoofing attacks and domain variations. However, existing methods primarily rely on Euclidean distances, failing to adequately capture the intrinsic hierarchical structures associated with attack categories and domain factors. To address these issues, we design a novel framework Poin-HierNet to construct domain-invariant hierarchical representations in the Poincar\'e sphere. Poin-HierNet includes three key components: 1) Poincar\'e Prototype Learning (PPL) with several data prototypes aligning sample features and capturing multilevel hierarchies beyond human labels; 2) Hierarchical Structure Learning (HSL) leverages top prototypes to establish a tree-like hierarchical structure from data prototypes; and 3) Poincar\'e Feature Whitening (PFW) enhances domain invariance by applying feature whitening to suppress domain-sensitive features. We evaluate our approach on four datasets: ASVspoof 2019 LA, ASVspoof 2021 LA, ASVspoof 2021 DF, and In-The-Wild. Experimental results demonstrate that Poin-HierNet exceeds state-of-the-art methods in Equal Error Rate.
【15】Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
标题:通过声码器之前的显式带宽扩展增强歌唱声音合成中的频谱图现实性
链接:https://arxiv.org/abs/2508.01796
备注:7 pages, 8 figures
摘要:本文解决了增强声码器生成的歌声音频的现实主义的挑战,通过减轻合成和现实生活中的录音,特别是在高频频谱图组件之间的可区分的差距。我们提出的方法结合了两个创新:一个显式的线性谱图估计步骤,使用去噪扩散过程与基于DiT的神经网络架构优化的时间-频率数据,和一个重新设计的声码器的基础上Vocos专门处理大型线性谱图增加频率箱。这种集成方法可以产生具有高保真度频谱图的音频,这对于人类听众和机器分类器来说都是一个挑战,难以与真实的录音区分开来。客观和主观的评估表明,我们的精简方法保持高音频质量,同时实现这种现实主义。这项工作在克服当前声码技术的局限性方面取得了重大进展,特别是在对抗性攻击的背景下对假频谱图检测。
摘要:This paper addresses the challenge of enhancing the realism of vocoder-generated singing voice audio by mitigating the distinguishable disparities between synthetic and real-life recordings, particularly in high-frequency spectrogram components. Our proposed approach combines two innovations: an explicit linear spectrogram estimation step using denoising diffusion process with DiT-based neural network architecture optimized for time-frequency data, and a redesigned vocoder based on Vocos specialized in handling large linear spectrograms with increased frequency bins. This integrated method can produce audio with high-fidelity spectrograms that are challenging for both human listeners and machine classifiers to differentiate from authentic recordings. Objective and subjective evaluations demonstrate that our streamlined approach maintains high audio quality while achieving this realism. This work presents a substantial advancement in overcoming the limitations of current vocoding techniques, particularly in the context of adversarial attacks on fake spectrogram detection.
【16】Sonify Anything: Towards Context-Aware Sonic Interactions in AR
标题:Sonify Anything:在AR中实现上下文感知的声波交互
链接:https://arxiv.org/abs/2508.01789
摘要:在增强现实(AR)中,虚拟对象与真实对象交互。然而,虚拟对象的物理性的缺乏导致自然的声音交互的缺乏。当虚拟和真实物体碰撞时,要么不播放声音,要么播放普通声音。两者都导致了不一致的多感官体验,减少了互动和对象的真实性。与虚拟现实(VR)和游戏不同,在虚拟现实和游戏中,预定义的场景和交互允许播放预先录制的声音样本,AR需要实时声音合成,动态适应新的上下文和对象,以在交互期间提供视听一致性。为了增强AR中的真实-虚拟对象交互,我们提出了一个上下文感知声音的框架,使用计算机视觉的方法来识别和分割真实对象的材料。材料的物理特性和相互作用的冲击动力学用于使用物理建模合成实时生成基于材料的声音。在一项有24名参与者的用户研究中,我们将我们一致的基于材料的声音与通用的声音效果进行了比较,反映了AR应用程序中非上下文感知声音的当前标准。结果表明,基于材料的声音导致更真实的声音交互。基于材料的声音也使参与者能够以更高的准确性和信心区分视觉上相似的材料。这些发现表明,AR中基于材料的情境感知声波交互可以培养更强的现实感,并增强我们对现实世界环境的感知。
摘要:In Augmented Reality (AR), virtual objects interact with real objects. However, the lack of physicality of virtual objects leads to the absence of natural sonic interactions. When virtual and real objects collide, either no sound or a generic sound is played. Both lead to an incongruent multisensory experience, reducing interaction and object realism. Unlike in Virtual Reality (VR) and games, where predefined scenes and interactions allow for the playback of pre-recorded sound samples, AR requires real-time sound synthesis that dynamically adapts to novel contexts and objects to provide audiovisual congruence during interaction. To enhance real-virtual object interactions in AR, we propose a framework for context-aware sounds using methods from computer vision to recognize and segment the materials of real objects. The material's physical properties and the impact dynamics of the interaction are used to generate material-based sounds in real-time using physical modelling synthesis. In a user study with 24 participants, we compared our congruent material-based sounds to a generic sound effect, mirroring the current standard of non-context-aware sounds in AR applications. The results showed that material-based sounds led to significantly more realistic sonic interactions. Material-based sounds also enabled participants to distinguish visually similar materials with significantly greater accuracy and confidence. These findings show that context-aware, material-based sonic interactions in AR foster a stronger sense of realism and enhance our perception of real-world surroundings.
【17】Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
标题:Voxlect:全球方言和地区语言建模的语音基金会模型基准
链接:https://arxiv.org/abs/2508.01691
摘要:我们提出了Voxlect,一个新的基准建模方言和区域语言的全球语音基础模型。具体而言,我们报告了对英语、阿拉伯语、普通话和广东话、藏语、印度语、泰语、西班牙语、法语、德语、巴西葡萄牙语和意大利语的方言和区域语言变体的全面基准评估。我们的研究使用了来自30个提供方言信息的公开语音库的超过200万个训练话语。我们评估了几种广泛使用的语音基础模型在语音方言分类方面的性能。我们评估的鲁棒性的方言模型在嘈杂的条件下,并提出了一个错误分析,突出建模结果与地理连续性。除了对方言分类进行基准测试外,我们还展示了Voxlect支持的几个下游应用程序。具体来说,我们表明,Voxlect可以应用于增强现有的语音识别数据集与方言信息,使跨方言变化的ASR性能进行更详细的分析。Voxlect也被用作评估语音生成系统性能的工具。Voxlect在RAIL家族的许可下可在https://github.com/tiantiaf0627/voxlect上公开获得。
摘要:We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the license of the RAIL family at: https://github.com/tiantiaf0627/voxlect.
【18】From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
标题:从对比到共性:音频共性字幕增强多模式LLM中的音频文本跨模式理解
链接:https://arxiv.org/abs/2508.01659
摘要:在多模态大语言模型的预训练和微调过程中,音频字幕(AC)在增强音频-文本跨模态理解方面起着关键作用。为了进一步加强这种对齐,最近的工作提出了音频差异字幕(ADC),它采用多个音频输入,并鼓励模型描述它们的差异,从而促进细粒度的音频区分。然而,尽管它的有效性,使差异告诉和详细的歧视,ADC介绍了一个显着的语义差距之间的输入音频-往往丰富的不同的声音事件-和相对简短,差异为重点的输出字幕。这种对AC风格描述的偏离导致了与预训练目标的不匹配,导致了微调过程中的灾难性遗忘。为了缓解这个问题,我们提出了音频通用字幕(ACC),这是一种极具挑战性但更温和的替代方案,它鼓励模型捕获音频片段之间的共享语义,而不是强调它们的详细差异。实验结果表明,ACC不仅有效地提高了对主要字幕基准的音频文本理解,而且与ADC相比,ACC还更好地保留了各种语音和音乐相关下游任务的一般能力,例如人声分类(VSC),语音情感识别(SER),乐器分类(MIC)和音乐流派分类(MGC)。这些发现验证了ACC有助于更强大的跨模态理解,并在MLLM的背景下实现了泛化和特定任务性能之间的更好平衡。
摘要:Audio Captioning (AC) plays a pivotal role in enhancing audio-text cross-modal understanding during the pretraining and finetuning of multimodal large language models (MLLMs). To further strengthen this alignment, recent works have proposed Audio Difference Captioning (ADC), which takes multiple audio inputs and encourages the model to describe their differences, thereby promoting fine-grained audio discrimination. However, despite its effectiveness in enabling difference-telling and detailed discrimination, ADC introduces a notable semantic gap between the input audios-often rich in diverse sound events-and the relatively brief, difference-focused output captions. This deviation from AC-style descriptions leads to a mismatch with the pretraining objective, resulting in catastrophic forgetting during finetuning. To mitigate this issue, we propose Audio Commonality Captioning (ACC), a comparably challenging but gentler alternative that encourages the model to capture the shared semantics across audio clips rather than emphasizing their detailed differences. Experimental results demonstrate that ACC not only effectively enhances audio-text understanding on primary captioning benchmarks but also better preserves general capabilities across diverse speech and music-related downstream tasks, such as vocal sound classification (VSC), speech emotion recognition (SER), musical instrument classification (MIC), and music genre classification (MGC), compared to ADC. These findings validate that ACC contributes to more robust cross-modal understanding and achieves a better balance between generalization and task-specific performance in the context of MLLMs.
【19】DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
标题:DRKF:用于多模态情感识别的知识融合解耦表示
链接:https://arxiv.org/abs/2508.01644
备注:Published in ACM Multimedia 2025. 10 pages, 4 figures
摘要:多模态情感识别(MER)旨在通过整合和分析来自多个模态的信息来识别情感状态。然而,固有的模态异质性和不一致的情绪线索仍然是阻碍性能的关键挑战。为了解决这些问题,我们提出了一个解耦表示与知识融合(DRKF)方法MER。DRKF由两个主要模块组成:优化表示学习(ORL)模块和知识融合(KF)模块。ORL采用了一种对比互信息估计方法与渐进的模态增强解耦任务相关的共享表示和模态特定的功能,同时减轻模态异质性。KF包括一个轻量级的基于自我注意力的融合编码器(FE),该编码器识别主导模态并整合来自其他模态的情感信息以增强融合表示。为了处理潜在的错误,从不正确的主导模态选择情绪不一致的条件下,我们引入了一个情绪歧视子模块(ED),强制融合表示保留歧视性线索的情绪不一致。这确保了即使FE选择了不适当的主导模态,情感分类子模块(EC)仍然可以通过利用保留的不一致信息来进行准确的预测。实验结果表明,DRKF在IEMOCAP、MELD和M3ED上实现了最先进的(SOTA)性能。源代码可在https://github.com/PANPANKK/DRKF上公开获取。
摘要:Multimodal emotion recognition (MER) aims to identify emotional states by integrating and analyzing information from multiple modalities. However, inherent modality heterogeneity and inconsistencies in emotional cues remain key challenges that hinder performance. To address these issues, we propose a Decoupled Representations with Knowledge Fusion (DRKF) method for MER. DRKF consists of two main modules: an Optimized Representation Learning (ORL) Module and a Knowledge Fusion (KF) Module. ORL employs a contrastive mutual information estimation method with progressive modality augmentation to decouple task-relevant shared representations and modality-specific features while mitigating modality heterogeneity. KF includes a lightweight self-attention-based Fusion Encoder (FE) that identifies the dominant modality and integrates emotional information from other modalities to enhance the fused representation. To handle potential errors from incorrect dominant modality selection under emotionally inconsistent conditions, we introduce an Emotion Discrimination Submodule (ED), which enforces the fused representation to retain discriminative cues of emotional inconsistency. This ensures that even if the FE selects an inappropriate dominant modality, the Emotion Classification Submodule (EC) can still make accurate predictions by leveraging preserved inconsistency information. Experiments show that DRKF achieves state-of-the-art (SOTA) performance on IEMOCAP, MELD, and M3ED. The source code is publicly available at https://github.com/PANPANKK/DRKF.
【20】Automatic Melody Reduction via Shortest Path Finding
标题:通过最短路径寻找自动旋律简化
链接:https://arxiv.org/abs/2508.01571
备注:Accepted paper at ISMIR 2025. this https URL
摘要:旋律约简作为音乐作品的抽象表示,不仅是音乐分析的工具,也是结构化音乐生成的中间表示。先前的计算理论,如调性音乐的生成理论,提供了对音乐的深刻解释,但它们不是完全自动的,通常仅限于古典流派。在本文中,我们提出了一种新颖的和概念上简单的计算方法,用于旋律减少使用基于图形的表示,灵感来自计算音乐理论的原则,其中减少过程被制定为找到最短路径。我们评估我们的算法流行,民间和古典流派,实验结果表明,该算法产生的旋律缩减比其他常见的旋律缩减更忠实于原始旋律,并且在音乐上更连贯方法.作为下游任务,我们使用旋律缩减来生成象征性的音乐变奏。实验表明,我们的方法实现了更高的质量比国家的最先进的风格转移方法。
摘要:Melody reduction, as an abstract representation of musical compositions, serves not only as a tool for music analysis but also as an intermediate representation for structured music generation. Prior computational theories, such as the Generative Theory of Tonal Music, provide insightful interpretations of music, but they are not fully automatic and usually limited to the classical genre. In this paper, we propose a novel and conceptually simple computational method for melody reduction using a graph-based representation inspired by principles from computational music theories, where the reduction process is formulated as finding the shortest path. We evaluate our algorithm on pop, folk, and classical genres, and experimental results show that the algorithm produces melody reductions that are more faithful to the original melody and more musically coherent than other common melody downsampling methods. As a downstream task, we use melody reductions to generate symbolic music variations. Experiments show that our method achieves higher quality than state-of-the-art style transfer methods.
【21】ShrutiSense: Microtonal Modeling and Correction in Indian Classical Music
标题:ShrutiSense:印度古典音乐中的微音调建模和纠正
链接:https://arxiv.org/abs/2508.01498
摘要:印度古典音乐依赖于一个复杂的22 shrutis(音高间隔)的微色调系统,它提供了超过12音等律系统的表达细微差别。现有的符号音乐处理工具无法解释这些微音调的区别和文化特定的拉加语法,支配旋律运动。我们提出了ShrutiSense,一个全面的符号音高处理系统,专为印度古典音乐,解决两个关键任务:(1)纠正西化或损坏的音高序列,(2)完成缺失值的旋律序列。我们的方法采用互补模型不同的任务:一个Shruti感知有限状态转换器(FST),执行上下文校正内的22 shruti框架和语法约束的Shruti隐马尔可夫模型(GC-SHMM),采用拉加特定的转换规则的上下文完成。对五个ragas的模拟数据的综合评估表明,ShrutiSense(FST模型)在校正任务中实现了91.3%的shruti分类准确率,示例序列在0.2至0.4的腐败水平下显示出86.7-90.0%的准确率。该系统在高达+/-50美分的音高噪声下表现出强大的性能,在整个拉格斯保持一致的准确性(90.7-91.8%),从而保留了印度古典音乐表达的文化真实性。
摘要:Indian classical music relies on a sophisticated microtonal system of 22 shrutis (pitch intervals), which provides expressive nuance beyond the 12-tone equal temperament system. Existing symbolic music processing tools fail to account for these microtonal distinctions and culturally specific raga grammars that govern melodic movement. We present ShrutiSense, a comprehensive symbolic pitch processing system designed for Indian classical music, addressing two critical tasks: (1) correcting westernized or corrupted pitch sequences, and (2) completing melodic sequences with missing values. Our approach employs complementary models for different tasks: a Shruti-aware finite-state transducer (FST) that performs contextual corrections within the 22-shruti framework and a grammar-constrained Shruti hidden Markov model (GC-SHMM) that incorporates raga-specific transition rules for contextual completions. Comprehensive evaluation on simulated data across five ragas demonstrates that ShrutiSense (FST model) achieves 91.3% shruti classification accuracy for correction tasks, with example sequences showing 86.7-90.0% accuracy at corruption levels of 0.2 to 0.4. The system exhibits robust performance under pitch noise up to +/-50 cents, maintaining consistent accuracy across ragas (90.7-91.8%), thus preserving the cultural authenticity of Indian classical music expression.
【22】Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
标题:具有最佳传输的音调估计的翻译等变自监督学习
链接:https://arxiv.org/abs/2508.01493
备注:Extended Abstracts for the Late-Breaking Demo Session of the 26th International Society for Music Information Retrieval Conference
摘要:在本文中,我们提出了一个最佳运输目标学习一维的双等变系统,并证明其适用于单音调估计。我们的方法为训练最先进的自监督音高估计器提供了一种理论上接地,数值上更稳定,更简单的替代方案。
摘要:In this paper, we propose an Optimal Transport objective for learning one-dimensional translation-equivariant systems and demonstrate its applicability to single pitch estimation. Our method provides a theoretically grounded, more numerically stable, and simpler alternative for training state-of-the-art self-supervised pitch estimators.
【23】PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
标题:PESTO:具有自监督转置等变目标的实时音调估计
链接:https://arxiv.org/abs/2508.01488
备注:Accepted to the Transactions of the International Society for Music Information Retrieval
摘要:在本文中,我们介绍PESTO,一个自我监督的学习方法,使用暹罗架构的单音高估计。我们的模型处理一个变量$Q$变换(VQT)的各个帧,并预测音高分布。神经网络被设计为与翻译等变,特别是由于Toeplitz全连接层。此外,我们通过翻译和裁剪VQT帧来构建音高移位对,并使用一种新的基于类的转置等变目标训练我们的模型,从而消除了对注释数据的需求。由于这种架构和训练目标,我们的模型实现了卓越的性能,同时非常轻量级($130$k参数)。对音乐和语音数据集(MIR-1 K,MDB-stem-synth和PTDB)的评估表明,PESTO不仅优于自监督基线,而且与监督方法竞争,表现出优越的跨数据集泛化能力。最后,我们通过开发一个使用缓存卷积的流VQT实现来增强PESTO的实用性。结合我们的模型的低延迟(小于10 ms)和最小的参数计数,这使得PESTO特别适合实时应用。
摘要:In this paper, we introduce PESTO, a self-supervised learning approach for single-pitch estimation using a Siamese architecture. Our model processes individual frames of a Variable-$Q$ Transform (VQT) and predicts pitch distributions. The neural network is designed to be equivariant to translations, notably thanks to a Toeplitz fully-connected layer. In addition, we construct pitch-shifted pairs by translating and cropping the VQT frames and train our model with a novel class-based transposition-equivariant objective, eliminating the need for annotated data. Thanks to this architecture and training objective, our model achieves remarkable performances while being very lightweight ($130$k parameters). Evaluations on music and speech datasets (MIR-1K, MDB-stem-synth, and PTDB) demonstrate that PESTO not only outperforms self-supervised baselines but also competes with supervised methods, exhibiting superior cross-dataset generalization. Finally, we enhance PESTO's practical utility by developing a streamable VQT implementation using cached convolutions. Combined with our model's low latency (less than 10 ms) and minimal parameter count, this makes PESTO particularly suitable for real-time applications.
【24】Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation
标题:从乐谱到表演:高效的人性可控长歌生成,具有酒吧级符号符号
链接:https://arxiv.org/abs/2508.01394
摘要:歌曲生成被认为是音乐AIGC中最具挑战性的问题;然而,现有的方法还没有完全克服四个持续的限制:可控性,概括性,感知质量和持续时间。我们认为,这些缺点主要源于试图直接从原始音频中学习音乐理论的流行范式,这对当前模型来说仍然非常困难。为了解决这个问题,我们提出了Bar级AI Composing Helper(BACH),这是第一个明确设计用于通过人类可编辑的符号分数生成歌曲的模型。BACH引入了一个标记化策略和一个符号生成过程,为分层歌曲结构量身定制。因此,它在歌曲生成的效率、持续时间和感知质量方面都有了很大的提高。实验表明,BACH以较小的模型尺寸,在所有公开报道的歌曲生成系统中建立了新的SOTA,甚至超过了Suno等商业解决方案。人类评估进一步证实了其在多个主观指标上的优越性。
摘要:Song generation is regarded as the most challenging problem in music AIGC; nonetheless, existing approaches have yet to fully overcome four persistent limitations: controllability, generalizability, perceptual quality, and duration. We argue that these shortcomings stem primarily from the prevailing paradigm of attempting to learn music theory directly from raw audio, a task that remains prohibitively difficult for current models. To address this, we present Bar-level AI Composing Helper (BACH), the first model explicitly designed for song generation through human-editable symbolic scores. BACH introduces a tokenization strategy and a symbolic generative procedure tailored to hierarchical song structure. Consequently, it achieves substantial gains in the efficiency, duration, and perceptual quality of song generation. Experiments demonstrate that BACH, with a small model size, establishes a new SOTA among all publicly reported song generation systems, even surpassing commercial solutions such as Suno. Human evaluations further confirm its superiority across multiple subjective metrics.
【25】Foundation Models for Bioacoustics -- a Comparative Review
标题:生物声学基础模型--比较评论
链接:https://arxiv.org/abs/2508.01277
备注:Preprint
摘要:自动化生物声学分析对于生物多样性监测和保护至关重要,需要能够适应各种生物声学任务的高级深度学习模型。本文对大规模预训练的生物声学基础模型进行了全面的回顾,并系统地研究了它们在多个生物声学分类任务中的可移植性。我们概述了生物声学表示学习,包括主要的预训练数据源和基准。在此基础上,我们回顾生物声学基础模型,深入分析设计决策,如模型架构,预训练方案,和训练范式。此外,我们评估了BEANS和BirdSet基准分类任务的选定基础模型,比较了线性和专注探测策略下学习表示的通用性。我们全面的实验分析表明,BirdMAE,训练大规模的鸟鸣数据与自我监督的目标,达到最佳性能的BirdSet基准。在BEANS上,BEAT $_{NLM}$(NatureLM-audio大音频模型的提取编码器)稍好一些。这两种基于transformer的模型都需要仔细探测,以提取其表示的全部性能。ConvNext$_{BS}$和Perch模型在大规模鸟鸣数据的监督下进行训练,对于线性探测设置中BirdSet的被动声学监测分类任务仍然具有竞争力。训练一个新的线性分类器比在没有进一步训练的情况下评估这些模型有明显的优势。而在BEANS上,在AudioSet上用自我监督训练的基线模型BEAT在用专注的探测进行评估时表现优于鸟类特定模型。这些研究结果提供了有价值的指导,从业者选择适当的模型,以适应新的生物声学分类任务,通过探测。
摘要:Automated bioacoustic analysis is essential for biodiversity monitoring and conservation, requiring advanced deep learning models that can adapt to diverse bioacoustic tasks. This article presents a comprehensive review of large-scale pretrained bioacoustic foundation models and systematically investigates their transferability across multiple bioacoustic classification tasks. We overview bioacoustic representation learning including major pretraining data sources and benchmarks. On this basis, we review bioacoustic foundation models by thoroughly analysing design decisions such as model architecture, pretraining scheme, and training paradigm. Additionally, we evaluate selected foundation models on classification tasks from the BEANS and BirdSet benchmarks, comparing the generalisability of learned representations under both linear and attentive probing strategies. Our comprehensive experimental analysis reveals that BirdMAE, trained on large-scale bird song data with a self-supervised objective, achieves the best performance on the BirdSet benchmark. On BEANS, BEATs$_{NLM}$, the extracted encoder of the NatureLM-audio large audio model, is slightly better. Both transformer-based models require attentive probing to extract the full performance of their representations. ConvNext$_{BS}$ and Perch models trained with supervision on large-scale bird song data remain competitive for passive acoustic monitoring classification tasks of BirdSet in linear probing settings. Training a new linear classifier has clear advantages over evaluating these models without further training. While on BEANS, the baseline model BEATs trained with self-supervision on AudioSet outperforms bird-specific models when evaluated with attentive probing. These findings provide valuable guidance for practitioners selecting appropriate models to adapt them to new bioacoustic classification tasks via probing.
【26】Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
标题:多模式情感推理的基准和弥合情感冲突
链接:https://arxiv.org/abs/2508.01181
备注:ACM Multimedia 2025
摘要:尽管现有的多模态大语言模型在多模态情感推理方面表现出色,但它们往往忽略了涉及情感冲突的场景,其中来自不同模态的情感线索是不一致的。为了填补这一空白,我们首先介绍CA-MER,一个新的基准,旨在检查MLLM下现实的情感冲突。它由三个子集组成:视频对齐,音频对齐和一致性,其中只有一个或所有模态反映了真实的情感。然而,对我们的CA-MER的评估表明,目前最先进的情感MLLM系统地过度依赖于音频信号在情感冲突,忽视了视觉模态的关键线索。为了减轻这种偏见,我们提出了MoSEAR,一个参数有效的框架,促进平衡的模态集成。MoSEAR由两个模块组成:(1)MoSE,具有正则化门控机制的特定模态专家,可减少微调头部中的模态偏差;(2)AR,一种注意力重新分配机制,可在推理过程中重新平衡冻结骨干中的模态贡献。我们的框架提供了两个关键的优势:它减轻了情感冲突,并提高了一致的样本的性能,而不会导致音频和视觉模态之间的权衡。在多个基准测试(包括MER 2023、EMER、DFEW和我们的CA-MER)上的实验表明,MoSEAR实现了最先进的性能,特别是在模态冲突条件下。
摘要:Despite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where emotional cues from different modalities are inconsistent. To fill this gap, we first introduce CA-MER, a new benchmark designed to examine MLLMs under realistic emotion conflicts. It consists of three subsets: video-aligned, audio-aligned, and consistent, where only one or all modalities reflect the true emotion. However, evaluations on our CA-MER reveal that current state-of-the-art emotion MLLMs systematically over-rely on audio signal during emotion conflicts, neglecting critical cues from visual modality. To mitigate this bias, we propose MoSEAR, a parameter-efficient framework that promotes balanced modality integration. MoSEAR consists of two modules: (1)MoSE, modality-specific experts with a regularized gating mechanism that reduces modality bias in the fine-tuning heads; and (2)AR, an attention reallocation mechanism that rebalances modality contributions in frozen backbones during inference. Our framework offers two key advantages: it mitigates emotion conflicts and improves performance on consistent samples-without incurring a trade-off between audio and visual modalities. Experiments on multiple benchmarks-including MER2023, EMER, DFEW, and our CA-MER-demonstrate that MoSEAR achieves state-of-the-art performance, particularly under modality conflict conditions.
【27】Advancing the Foundation Model for Music Understanding
标题:推进音乐理解的基础模型
链接:https://arxiv.org/abs/2508.01178
摘要:音乐信息检索(MIR)领域是分散的,专门的模型擅长孤立的任务。在这项工作中,我们通过引入一个统一的基础模型MuFun来挑战这种范式,以实现整体的音乐理解。我们的模型采用了一种新颖的架构,可以联合处理器乐和抒情内容,并在大规模数据集上进行训练,涵盖体裁分类、音乐标签和问答等多种任务。为了促进强大的评估,我们还提出了一个新的基准多方面的音乐理解称为MuCUE(音乐综合理解评估)。实验表明,我们的模型在MuCUE任务中的性能明显优于现有的音频大语言模型,证明了其最先进的有效性和泛化能力。
摘要:The field of Music Information Retrieval (MIR) is fragmented, with specialized models excelling at isolated tasks. In this work, we challenge this paradigm by introducing a unified foundation model named MuFun for holistic music understanding. Our model features a novel architecture that jointly processes instrumental and lyrical content, and is trained on a large-scale dataset covering diverse tasks such as genre classification, music tagging, and question answering. To facilitate robust evaluation, we also propose a new benchmark for multi-faceted music understanding called MuCUE (Music Comprehensive Understanding Evaluation). Experiments show our model significantly outperforms existing audio large language models across the MuCUE tasks, demonstrating its state-of-the-art effectiveness and generalization ability.
【28】GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification
标题:GeHirNet:语音病理分类的性别意识分层模型
链接:https://arxiv.org/abs/2508.01172
摘要:基于人工智能的语音分析显示出疾病诊断的前景,但由于与性别相关的声学变化和罕见疾病数据的稀缺,现有的分类器往往无法准确识别特定的病理。我们提出了一个新的两阶段框架,首先使用ResNet-50在Mel谱图上识别性别特异性病理模式,然后进行性别条件性疾病分类。我们通过多尺度恢复和时间扭曲增强来解决类不平衡问题。在四个公共存储库的合并数据集上进行评估,我们的两阶段时间扭曲架构实现了最先进的性能(97.63%的准确率,95.25%的MCC),与单阶段基线相比,MCC提高了5%。这项工作推进了语音病理分类,同时通过对语音特征的分层建模来减少性别偏见。
摘要:AI-based voice analysis shows promise for disease diagnostics, but existing classifiers often fail to accurately identify specific pathologies because of gender-related acoustic variations and the scarcity of data for rare diseases. We propose a novel two-stage framework that first identifies gender-specific pathological patterns using ResNet-50 on Mel spectrograms, then performs gender-conditioned disease classification. We address class imbalance through multi-scale resampling and time warping augmentation. Evaluated on a merged dataset from four public repositories, our two-stage architecture with time warping achieves state-of-the-art performance (97.63\% accuracy, 95.25\% MCC), with a 5\% MCC improvement over single-stage baseline. This work advances voice pathology classification while reducing gender bias through hierarchical modeling of vocal characteristics.
【29】Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR
标题:用更少的钱听更多的:多模式检索和选择增强的对话基于LLM的ASB
链接:https://arxiv.org/abs/2508.01166
摘要:自动语音识别(ASR)的目的是将人类的语音内容转换为相应的文本。在会话场景中,有效地利用上下文可以提高其准确性。大型语言模型(LLM)卓越的长上下文理解和推理能力使基于LLM的ASR(LLM-ASR)能够利用历史上下文来识别具有高度上下文相关性的会话语音。然而,现有的会话LLM-ASR方法使用固定数量的先前话语或整个会话历史作为上下文,由于大量不相关和冗余信息而导致显著的ASR混淆和计算成本。本文提出了一种多模态检索和选择方法,名为MARS,增强会话LLM-ASR,使其能够检索和选择最相关的声学和文本的历史背景下,为当前的话语。具体地,多模态检索获得一组候选历史上下文,每个候选历史上下文与当前话语表现出高的声学或文本相似性。多模态选择计算每个检索到的候选历史上下文的声学和文本相似性,并通过采用我们提出的接近理想的排名方法来考虑这两个相似性,选择最佳的历史上下文。对Interspeech 2025多语言会话语音模型挑战数据集的评估表明,LLM-ASR在仅使用1.5K小时的数据进行训练并配备MARS时,其性能优于使用179 K小时数据进行训练的最先进的顶级系统。
摘要:Automatic Speech Recognition (ASR) aims to convert human speech content into corresponding text. In conversational scenarios, effectively utilizing context can enhance its accuracy. Large Language Models' (LLMs) exceptional long-context understanding and reasoning abilities enable LLM-based ASR (LLM-ASR) to leverage historical context for recognizing conversational speech, which has a high degree of contextual relevance. However, existing conversational LLM-ASR methods use a fixed number of preceding utterances or the entire conversation history as context, resulting in significant ASR confusion and computational costs due to massive irrelevant and redundant information. This paper proposes a multi-modal retrieval-and-selection method named MARS that augments conversational LLM-ASR by enabling it to retrieve and select the most relevant acoustic and textual historical context for the current utterance. Specifically, multi-modal retrieval obtains a set of candidate historical contexts, each exhibiting high acoustic or textual similarity to the current utterance. Multi-modal selection calculates the acoustic and textual similarities for each retrieved candidate historical context and, by employing our proposed near-ideal ranking method to consider both similarities, selects the best historical context. Evaluations on the Interspeech 2025 Multilingual Conversational Speech Language Model Challenge dataset show that the LLM-ASR, when trained on only 1.5K hours of data and equipped with the MARS, outperforms the state-of-the-art top-ranking system trained on 179K hours of data.
【30】Accessibility and Social Inclusivity: A Literature Review of Music Technology for Blind and Low Vision People
标题:可及性与社会包容性--盲与低视力人群音乐技术研究综述
链接:https://arxiv.org/abs/2508.00929
备注:Accepted by ASSETS'25 - The 27th International ACM SIGACCESS Conference on Computers and Accessibility
摘要:本文介绍了一个系统的文献综述音乐技术为盲人和低视力(BLV)的个人。音乐活动对BLV人特别有益。然而,尚未尝试采用一种系统的方法来组织关于为BLV人设计无障碍技术的知识。我们根据技术类型和BLV人参与研究的程度对现有研究进行分类。我们确定了BLV以人为本的音乐技术的六个主要类别,并强调了设计目标的四个关键趋势。基于这些类别,我们提出了四个一般性的见解,重点是(1)空间意识,(2)获取信息,(3)(非语言)交流,(4)记忆。确定的趋势表明,需要更多的实证研究,涉及BLV人在现实世界中的场景,以确保技术进步可以提高音乐体验和社会包容性。这项研究提出了合作的音乐技术和包容性的现实世界的测试与目标群体作为两个关键领域在目前的研究中缺失。它们是将重点从“无障碍技术”转移到“包容性技术”的基础步骤,适用于BLV个人在更广泛的无障碍研究领域。
摘要:This paper presents a systematic literature review of music technology tailored for blind and low vision (BLV) individuals. Music activities can be particularly beneficial for BLV people. However, a systematic approach to organizing knowledge on designing accessible technology for BLV people has yet to be attempted. We categorize the existing studies based on the type of technology and the extent of BLV people's involvement in the research. We identify six main categories of BLV people-oriented music technology and highlight four key trends in design goals. Based on these categories, we propose four general insights focusing on (1) spatial awareness, (2) access to information, (3) (non-verbal) communication, and (4) memory. The identified trends suggest that more empirical studies involving BLV people in real-world scenarios are needed to ensure that technological advancements can enhance musical experiences and social inclusion. This research proposes collaborative music technology and inclusive real-world testing with the target group as two key areas missing in current research. They serve as a foundational step in shifting the focus from ``accessible technology'' to ``inclusive technology'' for BLV individuals within the broader field of accessibility research.
【31】Reference-free Adversarial Sex Obfuscation in Speech
标题:无参考的言语中的对抗性性别混淆
链接:https://arxiv.org/abs/2508.02295
摘要:言语中的性别转换涉及数据收集的隐私风险,并且通常在输出中留下残留的性别特定线索,即使目标说话者参考不可用。我们介绍RASO用于无参考对抗性性别混淆。创新包括性别条件对抗学习框架,以将语言内容与性别相关的声学标记和明确的规则化相结合,以将基频分布和共振峰轨迹与从性别平衡的训练数据中学习到的性别中性特征相一致。RASO保留了语言内容,即使在半知情的攻击模型下进行评估,它也明显优于竞争性混淆方法。
摘要:Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation.
【32】Test-Time Training for Speech Enhancement
标题:语音增强的测试时间训练
链接:https://arxiv.org/abs/2508.01847
备注:Accepted to Interspeech 2025. 5 pages, 2 figures
摘要:本文介绍了一种新的应用测试时训练(TTT)语音增强,解决不可预测的噪声条件和域偏移所带来的挑战。该方法将主语音增强任务与自监督辅助任务相结合,采用Y型结构。该模型通过优化所提出的自监督任务(如噪声增强信号重建或掩蔽谱图预测),在推理时间内动态适应新的领域,从而绕过对标记数据的需求。我们进一步介绍了各种TTT策略,提供了适应和效率之间的权衡。对合成和真实世界数据集的评估显示,语音质量指标得到了一致的改善,优于基线模型。这项工作突出了TTT在语音增强中的有效性,为未来自适应和鲁棒语音处理的研究提供了见解。
摘要:This paper introduces a novel application of Test-Time Training (TTT) for Speech Enhancement, addressing the challenges posed by unpredictable noise conditions and domain shifts. This method combines a main speech enhancement task with a self-supervised auxiliary task in a Y-shaped architecture. The model dynamically adapts to new domains during inference time by optimizing the proposed self-supervised tasks like noise-augmented signal reconstruction or masked spectrogram prediction, bypassing the need for labeled data. We further introduce various TTT strategies offering a trade-off between adaptation and efficiency. Evaluations across synthetic and real-world datasets show consistent improvements across speech quality metrics, outperforming the baseline model. This work highlights the effectiveness of TTT in speech enhancement, providing insights for future research in adaptive and robust speech processing.
【1】Revisiting the Privacy of Low-Frequency Speech Signals: Exploring Resampling Methods, Evaluation Scenarios, and Speaker Characteristics
标题:重新审视低频语音信号的隐私:探索恢复方法、评估场景和说话者特征
链接:https://arxiv.org/abs/2508.02483
备注:Accepted at SPSC 2025 - 5th Symposium on Security and Privacy in Speech Communication
摘要:虽然现实生活中的音频记录提供了对社会动态和会话行为的洞察,但它们也引起了对个人敏感数据隐私的担忧。本文探讨了将录音限制为低频音频以保护口语内容的有效性。对于不同采样率的音频信号,我们比较了采用抗混叠滤波的效果。隐私增强通过自动语音识别模型的增加的单词错误率来衡量。对效用性能的影响是用语音活动检测模型来衡量的。我们的实验结果表明,对于干净的录音,以高达800 Hz的采样率训练的模型可以正确转录大多数单词。对于这两种模型,我们分析了说话者的性别和音高的影响,我们证明了缺少抗混叠滤波器会更严重地损害语音隐私。
摘要:While audio recordings in real life provide insights into social dynamics and conversational behavior, they also raise concerns about the privacy of personal, sensitive data. This article explores the effectiveness of restricting recordings to low-frequency audio to protect spoken content. For resampling the audio signals to different sampling rates, we compare the effect of employing anti-aliasing filtering. Privacy enhancement is measured by an increased word error rate of automatic speech recognition models. The impact on utility performance is measured with voice activity detection models. Our experimental results show that for clean recordings, models trained with a sampling rate of up to 800 Hz transcribe the majority of words correctly. For both models, we analyzed the impact of the speaker's sex and pitch, and we demonstrated that missing anti-aliasing filters more strongly compromise speech privacy.
【2】Reference-free Adversarial Sex Obfuscation in Speech
标题:无参考的言语中的对抗性性别混淆
链接:https://arxiv.org/abs/2508.02295
摘要:言语中的性别转换涉及数据收集的隐私风险,并且通常在输出中留下残留的性别特定线索,即使目标说话者参考不可用。我们介绍RASO用于无参考对抗性性别混淆。创新包括性别条件对抗学习框架,以将语言内容与性别相关的声学标记和明确的规则化相结合,以将基频分布和共振峰轨迹与从性别平衡的训练数据中学习到的性别中性特征相一致。RASO保留了语言内容,即使在半知情的攻击模型下进行评估,它也明显优于竞争性混淆方法。
摘要:Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation.
【3】Guiding an Automatic Speech Recognition Decoder Using Large Language Models
标题:使用大型语言模型指导自动语音识别解码器
链接:https://arxiv.org/abs/2508.02228
备注:11 pages, 2 figures. This work has been submitted to the IEEE for possible publication
摘要:自动语音识别(ASR)由声学模型(AM)和语言模型(LM)组成。AM基于一系列语言单元(通常是音素、字符或标记)来估计声学信号的概率,而LM则评估特定单词或标记序列的可能性。尽管大型语言模型(LLM)在各种任务中表现出了巨大的潜力,但将它们集成到ASR中仍然是一个开放的挑战。通过分解的最大后验概率(MAP)估计的话(或令牌)给定的声学信号,我们推导出一个迭代过程,有利于一个新的整合的AM和LLM,同时保持其可分性。这种方法使每个组件都能够使用自己的数据进行独立训练和改进,从而通过利用两个模型的优势来最大限度地提高系统的性能,而无需联合优化。我们将我们的方法与三种语言模型进行了比较:N-gram,GCNN和TransformerLM,这些语言模型跨越了各种语音风格的多个数据集,包括ALLSSTAR,WSJ 0和TED-LIUM 3。我们的实验涉及两个声学模型(wav 2 vec 2.0和HuBERT)和三个LLM(GPT-2,LLaMA 2和Falcon)。值得注意的是,我们的方法在解决复杂的语音句子,首字母缩略词和特定领域的词汇表现出特别的功效。
摘要:Automatic Speech Recognition (ASR) consists of an acoustic model (AM) and a language model (LM). The AM estimates the probability of an acoustic signal based on a sequence of linguistic units, typically phones, characters, or tokens, while the LM assesses the likelihood of a specific sequence of words or tokens. Although Large Language Models (LLMs) have demonstrated significant potential across various tasks, integrating them into ASR remains an open challenge. By decomposing the maximum a posteriori (MAP) estimator of words (or tokens) given the acoustic signal, we derive an iterative procedure that facilitates a novel integration of the AM and LLM, while maintaining their separability. This approach enables each component to be independently trained and improved using its own data, thereby maximizing the system's performance by leveraging the strengths of both models without requiring joint optimization. We illustrate the effectiveness of our method in comparison to three language models: N-gram, GCNN, and TransformerLM across multiple datasets spanning various speech styles, including ALLSSTAR, WSJ0, and TED-LIUM 3. Our experiments involved two acoustic models (wav2vec 2.0 and HuBERT) and three LLMs (GPT-2, LLaMA 2, and Falcon). Notably, our method demonstrates particular efficacy in addressing complex speech sentences, acronyms, and domain-specific vocabulary.
【4】Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
标题:长形式多说话者语音识别的字错误率定义和算法
链接:https://arxiv.org/abs/2508.02112
备注:Accepted for IEEE Transactions on Audio Speech and Language Processing (TASLP), vol. 33
摘要:作为评价语音识别器的主要指标,单词错误率(WER)已经以不同的方式进行了扩展,以处理长格式多说话者语音识别器产生的转录本。这些系统处理包含多个说话者和复杂说话模式的长转录,使得经典的WER不能被应用。存在对说话者混淆错误进行计数的说话者属性方法,诸如级联最小置换WER cpWER和时间约束cpWER(tcpWER),以及旨在忽略说话者混淆错误的说话者不可知方法,诸如最优参考组合WER(ORC-WER)和MIMO-WER。这些WER评估不同的方面和错误类型(例如,时间未对准)。没有进行详细的比较。因此,我们提出了一个统一的描述现有的WER,并强调何时使用哪个指标。为了进一步分析有多少错误是由说话人混淆引起的,我们提出了拨号不变的cpWER(DI-cpWER)。它忽略了说话人归因错误,与cpWER的差异反映了说话人混淆对WER的影响。由于错误类型不能可靠地自动分类,我们讨论的方法来可视化的参考和假设成绩单之间的序列比对,以方便发现错误的人判断。由于一些WER定义具有高计算复杂度,我们引入了一个贪婪算法来近似ORC-WER和DI-cpWER具有高精度($<0.1 $偏差在我们的实验)和多项式复杂度,而不是指数。为了提高度量的可扩展性,我们还将tcpWER的时间约束纳入ORC-WER和MIMO-WER中,也显着降低了计算复杂度。
摘要:The predominant metric for evaluating speech recognizers, the Word Error Rate (WER) has been extended in different ways to handle transcripts produced by long-form multi-talker speech recognizers. These systems process long transcripts containing multiple speakers and complex speaking patterns so that the classical WER cannot be applied. There are speaker-attributed approaches that count speaker confusion errors, such as the concatenated minimum-permutation WER cpWER and the time-constrained cpWER (tcpWER), and speaker-agnostic approaches, which aim to ignore speaker confusion errors, such as the Optimal Reference Combination WER (ORC-WER) and the MIMO-WER. These WERs evaluate different aspects and error types (e.g., temporal misalignment). A detailed comparison has not been made. We therefore present a unified description of the existing WERs and highlight when to use which metric. To further analyze how many errors are caused by speaker confusion, we propose the Diarization-invariant cpWER (DI-cpWER). It ignores speaker attribution errors and its difference to cpWER reflects the impact of speaker confusions on the WER. Since error types cannot reliably be classified automatically, we discuss ways to visualize sequence alignments between the reference and hypothesis transcripts to facilitate the spotting of errors by a human judge. Since some WER definitions have high computational complexity, we introduce a greedy algorithm to approximate the ORC-WER and DI-cpWER with high precision ($<0.1\%$ deviation in our experiments) and polynomial complexity instead of exponential. To improve the plausibility of the metrics, we also incorporate the time constraint from the tcpWER into ORC-WER and MIMO-WER, also significantly reducing the computational complexity.
【5】Test-Time Training for Speech Enhancement
标题:语音增强的测试时间训练
链接:https://arxiv.org/abs/2508.01847
备注:Accepted to Interspeech 2025. 5 pages, 2 figures
摘要:本文介绍了一种新的应用测试时训练(TTT)语音增强,解决不可预测的噪声条件和域偏移所带来的挑战。该方法将主语音增强任务与自监督辅助任务相结合,采用Y型结构。该模型通过优化所提出的自监督任务(如噪声增强信号重建或掩蔽谱图预测),在推理时间内动态适应新的领域,从而绕过对标记数据的需求。我们进一步介绍了各种TTT策略,提供了适应和效率之间的权衡。对合成和真实世界数据集的评估显示,语音质量指标得到了一致的改善,优于基线模型。这项工作突出了TTT在语音增强中的有效性,为未来自适应和鲁棒语音处理的研究提供了见解。
摘要:This paper introduces a novel application of Test-Time Training (TTT) for Speech Enhancement, addressing the challenges posed by unpredictable noise conditions and domain shifts. This method combines a main speech enhancement task with a self-supervised auxiliary task in a Y-shaped architecture. The model dynamically adapts to new domains during inference time by optimizing the proposed self-supervised tasks like noise-augmented signal reconstruction or masked spectrogram prediction, bypassing the need for labeled data. We further introduce various TTT strategies offering a trade-off between adaptation and efficiency. Evaluations across synthetic and real-world datasets show consistent improvements across speech quality metrics, outperforming the baseline model. This work highlights the effectiveness of TTT in speech enhancement, providing insights for future research in adaptive and robust speech processing.
【6】An Age-Agnostic System for Robust Speaker Verification
标题:用于鲁棒说话人验证的时间不可知系统
链接:https://arxiv.org/abs/2508.01637
备注:Accepted to the Interspeech 2025 Workshop on Child Computer Interaction
摘要:在说话人验证(SV)中,当成人训练的SV系统应用于儿童说话人验证(C-SV)时,儿童和成人语音之间的声学不匹配会导致性能不佳。虽然领域适应技术可以提高C-SV任务的性能,但它们往往以成年人SV(A-SV)任务的性能显著下降为代价。在这项研究中,我们提出了一个年龄不可知的说话人确认(AASV)系统,实现了强大的性能在C-SV和A-SV任务。我们的方法采用了一个域分类器,从语音中解开与年龄相关的属性,并随后使用提取的域信息扩展嵌入空间,形成一个统一的扬声器表示,是强大的和高度区分不同年龄组。OGI和VoxCeleb数据集上的实验证明了我们的方法在弥合SV性能差异方面的有效性,为包容性和年龄适应性SV系统奠定了基础。
摘要:In speaker verification (SV), the acoustic mismatch between children's and adults' speech leads to suboptimal performance when adult-trained SV systems are applied to children's speaker verification (C-SV). While domain adaptation techniques can enhance performance on C-SV tasks, they often do so at the expense of significant degradation in performance on adults' SV (A-SV) tasks. In this study, we propose an Age Agnostic Speaker Verification (AASV) system that achieves robust performance across both C-SV and A-SV tasks. Our approach employs a domain classifier to disentangle age-related attributes from speech and subsequently expands the embedding space using the extracted domain information, forming a unified speaker representation that is robust and highly discriminative across age groups. Experiments on the OGI and VoxCeleb datasets demonstrate the effectiveness of our approach in bridging SV performance disparities, laying the foundation for inclusive and age-adaptive SV systems.
【7】Lumename: Wearable Device for Hearing Impaired with Personalized ML-Based Auditory Detection and Haptic-Visual Alerts
标题:Lumpose:用于听力受损的可穿戴设备,具有个性化的基于ML的听觉检测和触觉视觉警报
链接:https://arxiv.org/abs/2508.01576
摘要:根据世界卫生组织的数据,全球有4.3亿人患有听力损失。对于他们来说,识别诸如名字之类的口头命令是困难的。为了解决这个问题,实时智能手表Lumpose利用设备上的机器学习来检测用户自定义的名称,然后生成触觉视觉警报。在训练过程中,为了克服对大型数据集的需求,Lumpose使用了新的音频调制技术来增强来自一个用户的样本,并生成额外的样本来代表不同的性别和年龄。约束随机迭代被用来找到模型架构内的最佳参数。这种方法产生了一个低资源和低功耗的TinyML模型,可以快速推断各种关键字样本,同时在基于Arduino Nano 33 BLE Sense的定制智能手表上保持91.67%的准确率。
摘要:According to the World Health Organization, 430 million people experience disabling hearing loss. For them, recognizing spoken commands such as one's name is difficult. To address this issue, Lumename, a real-time smartwatch, utilizes on-device machine learning to detect a user-customized name before generating a haptic-visual alert. During training, to overcome the need for large datasets, Lumename uses novel audio modulation techniques to augment samples from one user and generate additional samples to represent diverse genders and ages. Constrained random iterations were used to find optimal parameters within the model architecture. This approach resulted in a low-resource and low-power TinyML model that could quickly infer various keyword samples while remaining 91.67\% accurate on a custom-built smartwatch based on an Arduino Nano 33 BLE Sense.
【8】Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations
标题:多粒度自适应时频注意力框架在真实世界通信降级下的音频Deepfake检测
链接:https://arxiv.org/abs/2508.01467
摘要:高度可信的合成语音的兴起对音频通信构成了越来越大的威胁。虽然现有的音频深度伪造检测(ADD)方法在干净条件下表现出良好的性能,但在现实通信环境中的数据包丢失和语音编解码器压缩等劣化情况下,它们的有效性会显著下降。在这项工作中,我们提出了第一个统一的框架下,这种退化,这是为了有效地适应多种类型的时频(TF)表示强大的ADD。我们的框架的核心是一种新的多粒度自适应注意力(MGAA)架构,它采用了一组可定制的多尺度注意力头捕捉全球和当地的感受野在不同的TF粒度。一种新的自适应融合机制随后调整和融合这些注意力分支的TF区域的显著性的基础上,允许模型动态地重新分配其重点根据退化的特性。这使得能够有效地定位和放大细微的伪造痕迹。大量的实验表明,所提出的框架始终优于国家的最先进的基线在各种现实世界的通信退化的情况下,包括六个语音编解码器和五个级别的数据包丢失。此外,比较分析表明,MGAA增强的功能显着提高了真实和虚假音频类之间的可分性,并锐化决策边界。这些结果突出了我们的框架在现实世界的通信环境中的鲁棒性和实际部署潜力。
摘要:The rise of highly convincing synthetic speech poses a growing threat to audio communications. Although existing Audio Deepfake Detection (ADD) methods have demonstrated good performance under clean conditions, their effectiveness drops significantly under degradations such as packet losses and speech codec compression in real-world communication environments. In this work, we propose the first unified framework for robust ADD under such degradations, which is designed to effectively accommodate multiple types of Time-Frequency (TF) representations. The core of our framework is a novel Multi-Granularity Adaptive Attention (MGAA) architecture, which employs a set of customizable multi-scale attention heads to capture both global and local receptive fields across varying TF granularities. A novel adaptive fusion mechanism subsequently adjusts and fuses these attention branches based on the saliency of TF regions, allowing the model to dynamically reallocate its focus according to the characteristics of the degradation. This enables the effective localization and amplification of subtle forgery traces. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art baselines across various real-world communication degradation scenarios, including six speech codecs and five levels of packet losses. In addition, comparative analysis reveals that the MGAA-enhanced features significantly improve separability between real and fake audio classes and sharpen decision boundaries. These results highlight the robustness and practical deployment potential of our framework in real-world communication environments.
【9】Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
标题:调制谱图与具有多头注意力的SSL融合用于假语音检测
链接:https://arxiv.org/abs/2508.01034
摘要:虚假语音检测系统已经成为对抗语音deepfake的必要条件。由于缺乏多样的训练数据,目前的系统在域外语音样本上表现出较差的泛化能力。在本文中,我们试图解决域泛化问题,提出了一种新的语音表示使用自监督(SSL)语音嵌入和调制频谱图(MS)功能。一个融合策略是用来结合这两个语音表示引入一个新的前端的分类任务。建议的SSL+MS融合表示被传递到AASIST后端网络。在单语和多语种的假语音数据集上进行实验,以评估所提出的模型架构在跨数据集和多语种情况下的有效性。与基线相比,该模型在域内设置中分别在ASVspoof 2019和MLAAD数据集上实现了37%和20%的相对性能提升。在域外场景中,在ASVspoof 2019上训练的模型在MLAAD数据集上进行评估时显示出36%的相对改进。在所有评估的语言中,所提出的模型始终优于基线,表明增强的域泛化。
摘要:Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to address domain generalization issue by proposing a novel speech representation using self-supervised (SSL) speech embeddings and the Modulation Spectrogram (MS) feature. A fusion strategy is used to combine both speech representations to introduce a new front-end for the classification task. The proposed SSL+MS fusion representation is passed to the AASIST back-end network. Experiments are conducted on monolingual and multilingual fake speech datasets to evaluate the efficacy of the proposed model architecture in cross-dataset and multilingual cases. The proposed model achieves a relative performance improvement of 37% and 20% on the ASVspoof 2019 and MLAAD datasets, respectively, in in-domain settings compared to the baseline. In the out-of-domain scenario, the model trained on ASVspoof 2019 shows a 36% relative improvement when evaluated on the MLAAD dataset. Across all evaluated languages, the proposed model consistently outperforms the baseline, indicating enhanced domain generalization.
【10】CAK: Emergent Audio Effects from Minimal Deep Learning
标题:CAK:来自最小深度学习的紧急音频效果
链接:https://arxiv.org/abs/2508.02643
备注:8 pages, 3 figures, code and other resources at this https URL
摘要:我们证明了一个3x 3卷积核可以在来自个性化语料库的200个样本上训练时产生紧急音频效果。我们通过两个关键技术实现了这一点:(1)条件感知内核(CAK),其中输出=输入+(学习模式x控制),具有支持零控制身份保留的软门机制;(2)AuGAN(审计GAN),它从“这是真的吗?“to“是否应用了所请求的值?“我们的网络不是学习生成或检测fortunes,而是合作验证控制应用程序,发现独特的转换。学习的内核表现出对角结构,创建频率相关的时间偏移,能够根据输入特性产生音乐效果。我们的研究结果显示了对抗性训练从最少的数据中发现音频转换的潜力,从而为效果设计提供了新的方法。
摘要:We demonstrate that a single 3x3 convolutional kernel can produce emergent audio effects when trained on 200 samples from a personalized corpus. We achieve this through two key techniques: (1) Conditioning Aware Kernels (CAK), where output = input + (learned_pattern x control), with a soft-gate mechanism supporting identity preservation at zero control; and (2) AuGAN (Audit GAN), which reframes adversarial training from "is this real?" to "did you apply the requested value?" Rather than learning to generate or detect forgeries, our networks cooperate to verify control application, discovering unique transformations. The learned kernel exhibits a diagonal structure creating frequency-dependent temporal shifts that are capable of producing musical effects based on input characteristics. Our results show the potential of adversarial training to discover audio transformations from minimal data, enabling new approaches to effect design.
【11】Perception of dynamic multi-speaker auditory scenes under different modes of attention
标题:不同注意模式下动态多扬声器听觉场景的感知
链接:https://arxiv.org/abs/2508.02620
摘要:注意力不是单一的,而是以多种形式运作,以促进有效的认知处理。在听觉域中,注意力使听觉场景中的相关声音优先化,并且可以以自下而上的方式被场景中的元素吸引,或者以自上而下的方式指向特征、对象或整个场景。这些注意力模式是如何相互作用的,它们的神经基础是否不同,目前还不清楚。在这项工作中,我们调查的感知和神经相关的不同的注意力模式在一个受控的“鸡尾酒会”的范例,听众听相同的刺激,并出席一个空间位置(基于特征),扬声器(基于对象),或整个场景(全球或自由听),同时检测偏差在现场的声音的音高。我们的研究结果表明,基于对象的注意是更有效的感知比基于特征或全球的注意。此外,基于对象的注意和基于空间的注意涉及不同的神经机制,并受到自下而上的显著性的不同调制。值得注意的是,虽然自下而上的显着性艾滋病的听觉对象的初始隔离,它在对象跟踪中的作用降低,一旦注意力已经自愿分配。此外,从EEG数据解码的刺激包络揭示了一个源采样方案,在全球注意力模式,不存在于对象或空间模式。总的来说,研究表明,相同的声学场景的感知不同的听力任务,由自上而下和自下而上的过程之间的相互作用的指导。
摘要:Attention is not monolithic; rather, it operates in multiple forms to facilitate efficient cognitive processing. In the auditory domain, attention enables the prioritization of relevant sounds in an auditory scene and can be either attracted by elements in the scene in a bottom-up fashion or directed towards features, objects, or the entire scene in a top-down fashion. How these modes of attention interact and whether their neural underpinnings are distinct remains unclear. In this work, we investigate the perceptual and neural correlates of different attentional modes in a controlled "cocktail party" paradigm, where listeners listen to the same stimuli and attend to either a spatial location (feature-based), a speaker (object-based), or the entire scene (global or free-listening) while detecting deviations in pitch of a voice in the scene. Our findings indicate that object-based attention is more perceptually effective than feature-based or global attention. Furthermore, object-based and spatial-based attention engage distinct neural mechanisms and are differentially modulated by bottom-up salience. Notably, while bottom-up salience aids in the initial segregation of auditory objects, it plays a reduced role in object tracking once attention has been voluntarily allocated. In addition, decoding the stimulus envelope from the EEG data revealed a source-sampling scheme in the global attention mode that is not present in the object or spatial modes. Overall, the study shows that the perception of the same acoustic scene differs according to the listening task, guided by an interaction between top-down and bottom-up processes.
【12】StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation
标题:StutterCut:不确定性引导的标准化切割,用于不流利分割
链接:https://arxiv.org/abs/2508.02255
备注:Accepted in Interspeech 2025
摘要:检测和分割不流利是有效的语音治疗和实时反馈的关键。然而,大多数方法仅在话语水平上对不流利进行分类。我们引入了StutterCut,这是一个半监督框架,它将不流利性分割公式化为图分区问题,其中来自重叠窗口的语音嵌入表示为图节点。我们使用在弱(话语级)标签上训练的伪Oracle分类器来改进节点之间的连接,其影响力由Monte Carlo dropout的不确定性度量控制。此外,我们扩展了弱标记的FluencyBank数据集,将帧级不流畅边界的四种不流畅类型。与合成数据集相比,这提供了一个更现实的基准。在真实和合成数据集上的实验表明,StutterCut优于现有方法,实现了更高的F1分数和更精确的口吃发作检测。
摘要:Detecting and segmenting dysfluencies is crucial for effective speech therapy and real-time feedback. However, most methods only classify dysfluencies at the utterance level. We introduce StutterCut, a semi-supervised framework that formulates dysfluency segmentation as a graph partitioning problem, where speech embeddings from overlapping windows are represented as graph nodes. We refine the connections between nodes using a pseudo-oracle classifier trained on weak (utterance-level) labels, with its influence controlled by an uncertainty measure from Monte Carlo dropout. Additionally, we extend the weakly labelled FluencyBank dataset by incorporating frame-level dysfluency boundaries for four dysfluency types. This provides a more realistic benchmark compared to synthetic datasets. Experiments on real and synthetic datasets show that StutterCut outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection.
【13】WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
标题:WhiSQA:使用Whisper编码器功能的非侵入性语音质量预测
链接:https://arxiv.org/abs/2508.02210
备注:Accepted at SPECOM 2025
摘要:近年来,人们在开发基于神经网络的SQ预测器方面做出了重大的研究努力。虽然主要目标是开发非侵入性,即~无参考的度量来评估SE系统的性能,最近的工作还研究了下游语音任务的损失函数内的神经SQ预测器的直接推断。为了帮助SQ预测器的训练,已经创建了具有相应的人类质量标签的几个大型音频数据集。最近在这一领域的工作表明,来自大型无监督或半监督基础语音模型的语音表示是神经SQ预测的有用输入特征表示。在这项工作中,提出了一种新的和强大的SQ预测的基础上提取的ASR模型的特征表示,发现是一个强大的输入功能的SQ预测任务。该系统实现了更高的相关性与人类MOS评级比最近的方法在所有NISQA测试集,并显示出显着更好的域适应性相比,常用的DNSMOS度量。
摘要:There has been significant research effort developing neural-network-based predictors of SQ in recent years. While a primary objective has been to develop non-intrusive, i.e.~reference-free, metrics to assess the performance of SE systems, recent work has also investigated the direct inference of neural SQ predictors within the loss function of downstream speech tasks. To aid in the training of SQ predictors, several large datasets of audio with corresponding human labels of quality have been created. Recent work in this area has shown that speech representations derived from large unsupervised or semi-supervised foundational speech models are useful input feature representations for neural SQ prediction. In this work, a novel and robust SQ predictor is proposed based on feature representations extracted from an ASR model, found to be a powerful input feature for the SQ prediction task. The proposed system achieves higher correlation with human MOS ratings than recent approaches on all NISQA test sets and shows significantly better domain adaption compared to the commonly used DNSMOS metric.
【14】Unsupervised Multi-channel Speech Dereverberation via Diffusion
标题:无监督多通道扩散语音去回响
链接:https://arxiv.org/abs/2508.02071
摘要:本文研究了多通道单说话人盲去混响问题,即利用多通道混合信号恢复纯净的无回声语音。为了解决这个问题,我们提出了USD-DPS,通过{D}扩散{P}后验{S}采样进行{U}无监督{S}语音{D}混响。USD-DPS使用无条件干净语音扩散模型作为强先验,通过后验采样来解决问题。在每个扩散采样步骤,我们估计所有麦克风通道的房间脉冲响应(RIR),这是进一步用于执行多通道混合一致性约束的扩散指导。对于多通道RIR估计,我们估计参考通道RIR通过优化子带RIR信号模型的RIR参数,与亚当优化器。我们使用前向卷积预测(FCP)来解析地估计非参考信道的RIR。我们发现,这种组合提供了采样效率和RIR先验建模之间的良好平衡,这表明无监督去混响方法的性能优越。https://usddps.github.io/USDDPS_demo/提供了音频演示页面。
摘要:We consider the problem of multi-channel single-speaker blind dereverberation, where multi-channel mixtures are used to recover the clean anechoic speech. To solve this problem, we propose USD-DPS, {U}nsupervised {S}peech {D}ereverberation via {D}iffusion {P}osterior {S}ampling. USD-DPS uses an unconditional clean speech diffusion model as a strong prior to solve the problem by posterior sampling. At each diffusion sampling step, we estimate all microphone channels' room impulse responses (RIRs), which are further used to enforce a multi-channel mixture consistency constraint for diffusion guidance. For multi-channel RIR estimation, we estimate reference-channel RIR by optimizing RIR parameters of a sub-band RIR signal model, with the Adam optimizer. We estimate non-reference channels' RIRs analytically using forward convolutive prediction (FCP). We found that this combination provides a good balance between sampling efficiency and RIR prior modeling, which shows superior performance among unsupervised dereverberation approaches. An audio demo page is provided in https://usddps.github.io/USDDPS_demo/.
【15】Marco-Voice Technical Report
标题:Marco-Voice技术报告
链接:https://arxiv.org/abs/2508.02038
备注:Technical Report
摘要:本文提出了一个多功能的语音合成系统,它集成了语音克隆和情感控制语音合成在一个统一的框架。这项工作的目标是解决长期存在的挑战,实现高度表达,可控和自然的语音生成,忠实地保留扬声器身份在不同的语言和情感背景。我们的方法引入了一个有效的说话人情感解纠缠机制与批量对比学习,使独立操纵的发言人身份和eemotional风格,以及旋转的情绪嵌入集成方法平滑的情绪控制。为了支持全面的训练和评估,我们构建了CSEMOTIONS,这是一个高质量的情感语音数据集,包含来自七个情感类别的六个专业演讲者的10小时普通话语音。大量的实验表明,我们的系统,马可语音,实现了客观和主观指标的大幅改善。通过对MarcoVoice语音合成系统进行全面的评测和分析,结果表明,MarcoVoice在语音清晰度和情感丰富度方面具有竞争力,在表达性神经语音合成领域取得了实质性的进步。
摘要:This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. Our approach introduces an effective speaker-emotion disentanglement mechanism with in-batch contrastive learning, enabling independent manipulation of speaker identity and eemotional style, as well as rotational emotional embedding integration method for smooth emotion control. To support comprehensive training and evaluation, we construct CSEMOTIONS, a high-quality emotional speech dataset containing 10 hours of Mandarin speech from six professional speakers across seven emotional categories. Extensive experiments demonstrate that our system, Marco-Voice, achieves substantial improvements in both objective and subjective metrics. Comprehensive evaluations and analysis were conducted, results show that MarcoVoice delivers competitive performance in terms of speech clarity and emotional richness, representing a substantial advance in the field of expressive neural speech synthesis.
【16】Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
标题:通过分层边界建模本地化视听Deepfakes
链接:https://arxiv.org/abs/2508.02000
备注:Work in progress
摘要:基于内容驱动的部分操作的视听时间深度伪造定位仍然是一项极具挑战性的任务。在这种情况下,deepfake区域通常只跨越几帧,其余的大部分与原始区域相同。为了解决这个问题,我们提出了一个分层边界建模网络(HBMNet),它包括三个模块:一个视听特征编码器,提取有区别的帧级表示,一个粗建议生成器,预测候选边界区域,和一个细粒度概率生成器,使用双向边界内容概率细化这些建议。从模态的角度来看,我们通过专门的编码和融合来增强视听学习,并通过帧级监督来增强辨别力。从时间的角度来看,HBMNet集成了多尺度线索和双向边界内容关系。实验表明,编码和融合主要提高精度,而帧级监督提高召回率。每个模块(视听融合、时间尺度、双向性)都提供互补的优势,共同提高定位性能。HBMNet的性能优于BA-TFD和UMMA Former,并通过更多的训练数据显示出潜在的可扩展性。
摘要:Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest remaining identical to the original. To tackle this, we propose a Hierarchical Boundary Modeling Network (HBMNet), which includes three modules: an Audio-Visual Feature Encoder that extracts discriminative frame-level representations, a Coarse Proposal Generator that predicts candidate boundary regions, and a Fine-grained Probabilities Generator that refines these proposals using bidirectional boundary-content probabilities. From the modality perspective, we enhance audio-visual learning through dedicated encoding and fusion, reinforced by frame-level supervision to boost discriminability. From the temporal perspective, HBMNet integrates multi-scale cues and bidirectional boundary-content relationships. Experiments show that encoding and fusion primarily improve precision, while frame-level supervision boosts recall. Each module (audio-visual fusion, temporal scales, bi-directionality) contributes complementary benefits, collectively enhancing localization performance. HBMNet outperforms BA-TFD and UMMAFormer and shows improved potential scalability with more training data.
【17】Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
标题:通过声码器之前的显式带宽扩展增强歌唱声音合成中的频谱图现实性
链接:https://arxiv.org/abs/2508.01796
备注:7 pages, 8 figures
摘要:本文解决了增强声码器生成的歌声音频的现实主义的挑战,通过减轻合成和现实生活中的录音,特别是在高频频谱图组件之间的可区分的差距。我们提出的方法结合了两个创新:一个显式的线性谱图估计步骤,使用去噪扩散过程与基于DiT的神经网络架构优化的时间-频率数据,和一个重新设计的声码器的基础上Vocos专门处理大型线性谱图增加频率箱。这种集成方法可以产生具有高保真度频谱图的音频,这对于人类听众和机器分类器来说都是一个挑战,难以与真实的录音区分开来。客观和主观的评估表明,我们的精简方法保持高音频质量,同时实现这种现实主义。这项工作在克服当前声码技术的局限性方面取得了重大进展,特别是在对抗性攻击的背景下对假频谱图检测。
摘要:This paper addresses the challenge of enhancing the realism of vocoder-generated singing voice audio by mitigating the distinguishable disparities between synthetic and real-life recordings, particularly in high-frequency spectrogram components. Our proposed approach combines two innovations: an explicit linear spectrogram estimation step using denoising diffusion process with DiT-based neural network architecture optimized for time-frequency data, and a redesigned vocoder based on Vocos specialized in handling large linear spectrograms with increased frequency bins. This integrated method can produce audio with high-fidelity spectrograms that are challenging for both human listeners and machine classifiers to differentiate from authentic recordings. Objective and subjective evaluations demonstrate that our streamlined approach maintains high audio quality while achieving this realism. This work presents a substantial advancement in overcoming the limitations of current vocoding techniques, particularly in the context of adversarial attacks on fake spectrogram detection.
【18】Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
标题:具有最佳传输的音调估计的翻译等变自监督学习
链接:https://arxiv.org/abs/2508.01493
备注:Extended Abstracts for the Late-Breaking Demo Session of the 26th International Society for Music Information Retrieval Conference
摘要:在本文中,我们提出了一个最佳运输目标学习一维的双等变系统,并证明其适用于单音调估计。我们的方法为训练最先进的自监督音高估计器提供了一种理论上接地,数值上更稳定,更简单的替代方案。
摘要:In this paper, we propose an Optimal Transport objective for learning one-dimensional translation-equivariant systems and demonstrate its applicability to single pitch estimation. Our method provides a theoretically grounded, more numerically stable, and simpler alternative for training state-of-the-art self-supervised pitch estimators.
【19】Foundation Models for Bioacoustics -- a Comparative Review
标题:生物声学基础模型--比较评论
链接:https://arxiv.org/abs/2508.01277
备注:Preprint
摘要:自动化生物声学分析对于生物多样性监测和保护至关重要,需要能够适应各种生物声学任务的高级深度学习模型。本文对大规模预训练的生物声学基础模型进行了全面的回顾,并系统地研究了它们在多个生物声学分类任务中的可移植性。我们概述了生物声学表示学习,包括主要的预训练数据源和基准。在此基础上,我们回顾生物声学基础模型,深入分析设计决策,如模型架构,预训练方案,和训练范式。此外,我们评估了BEANS和BirdSet基准分类任务的选定基础模型,比较了线性和专注探测策略下学习表示的通用性。我们全面的实验分析表明,BirdMAE,训练大规模的鸟鸣数据与自我监督的目标,达到最佳性能的BirdSet基准。在BEANS上,BEAT $_{NLM}$(NatureLM-audio大音频模型的提取编码器)稍好一些。这两种基于transformer的模型都需要仔细探测,以提取其表示的全部性能。ConvNext$_{BS}$和Perch模型在大规模鸟鸣数据的监督下进行训练,对于线性探测设置中BirdSet的被动声学监测分类任务仍然具有竞争力。训练一个新的线性分类器比在没有进一步训练的情况下评估这些模型有明显的优势。而在BEANS上,在AudioSet上用自我监督训练的基线模型BEAT在用专注的探测进行评估时表现优于鸟类特定模型。这些研究结果提供了有价值的指导,从业者选择适当的模型,以适应新的生物声学分类任务,通过探测。
摘要:Automated bioacoustic analysis is essential for biodiversity monitoring and conservation, requiring advanced deep learning models that can adapt to diverse bioacoustic tasks. This article presents a comprehensive review of large-scale pretrained bioacoustic foundation models and systematically investigates their transferability across multiple bioacoustic classification tasks. We overview bioacoustic representation learning including major pretraining data sources and benchmarks. On this basis, we review bioacoustic foundation models by thoroughly analysing design decisions such as model architecture, pretraining scheme, and training paradigm. Additionally, we evaluate selected foundation models on classification tasks from the BEANS and BirdSet benchmarks, comparing the generalisability of learned representations under both linear and attentive probing strategies. Our comprehensive experimental analysis reveals that BirdMAE, trained on large-scale bird song data with a self-supervised objective, achieves the best performance on the BirdSet benchmark. On BEANS, BEATs$_{NLM}$, the extracted encoder of the NatureLM-audio large audio model, is slightly better. Both transformer-based models require attentive probing to extract the full performance of their representations. ConvNext$_{BS}$ and Perch models trained with supervision on large-scale bird song data remain competitive for passive acoustic monitoring classification tasks of BirdSet in linear probing settings. Training a new linear classifier has clear advantages over evaluating these models without further training. While on BEANS, the baseline model BEATs trained with self-supervision on AudioSet outperforms bird-specific models when evaluated with attentive probing. These findings provide valuable guidance for practitioners selecting appropriate models to adapt them to new bioacoustic classification tasks via probing.
【20】Accessibility and Social Inclusivity: A Literature Review of Music Technology for Blind and Low Vision People
标题:可及性与社会包容性--盲与低视力人群音乐技术研究综述
链接:https://arxiv.org/abs/2508.00929
备注:Accepted by ASSETS'25 - The 27th International ACM SIGACCESS Conference on Computers and Accessibility
摘要:本文介绍了一个系统的文献综述音乐技术为盲人和低视力(BLV)的个人。音乐活动对BLV人特别有益。然而,尚未尝试采用一种系统的方法来组织关于为BLV人设计无障碍技术的知识。我们根据技术类型和BLV人参与研究的程度对现有研究进行分类。我们确定了BLV以人为本的音乐技术的六个主要类别,并强调了设计目标的四个关键趋势。基于这些类别,我们提出了四个一般性的见解,重点是(1)空间意识,(2)获取信息,(3)(非语言)交流,(4)记忆。确定的趋势表明,需要更多的实证研究,涉及BLV人在现实世界中的场景,以确保技术进步可以提高音乐体验和社会包容性。这项研究提出了合作的音乐技术和包容性的现实世界的测试与目标群体作为两个关键领域在目前的研究中缺失。它们是将重点从“无障碍技术”转移到“包容性技术”的基础步骤,适用于BLV个人在更广泛的无障碍研究领域。
摘要:This paper presents a systematic literature review of music technology tailored for blind and low vision (BLV) individuals. Music activities can be particularly beneficial for BLV people. However, a systematic approach to organizing knowledge on designing accessible technology for BLV people has yet to be attempted. We categorize the existing studies based on the type of technology and the extent of BLV people's involvement in the research. We identify six main categories of BLV people-oriented music technology and highlight four key trends in design goals. Based on these categories, we propose four general insights focusing on (1) spatial awareness, (2) access to information, (3) (non-verbal) communication, and (4) memory. The identified trends suggest that more empirical studies involving BLV people in real-world scenarios are needed to ensure that technological advancements can enhance musical experiences and social inclusion. This research proposes collaborative music technology and inclusive real-world testing with the target group as two key areas missing in current research. They serve as a foundational step in shifting the focus from ``accessible technology'' to ``inclusive technology'' for BLV individuals within the broader field of accessibility research.
机器翻译由腾讯交互翻译提供,仅供参考
