微信公众号:arXiv_Daily
cs.SD语音
【1】FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
标题:FuID:用于生成性音乐推荐的情态融合语义ID
链接:https://arxiv.org/abs/2601.08764
摘要:生成式推荐系统通过利用语义ID来表示项目,取得了显著的进步。然而,独立地标记每个模态的现有方法面临两个关键限制:(1)降低效率的跨模态冗余,以及(2)无法捕获限制项目表示的模态间交互。我们引入FusID,一个模态融合的语义ID框架,通过三个关键组件解决这些限制:(i)多模态融合,其通过跨模态联合编码信息来学习统一的表示,(ii)表示学习,其使频繁共现的项目嵌入更接近,同时保持独特性并防止特征冗余,以及(iii)将融合的连续嵌入转换成多个离散令牌以减轻ID冲突的乘积量化。在多模式下一首歌曲推荐上进行评估(即,在播放列表延续(playlist continuation)基准测试中,FusID实现了零ID冲突,确保每个标记序列映射到恰好一首歌曲,减轻了码本利用不足,并且在MRR和Recall@k(k = 1,5,10,20)方面优于基线。
摘要:Generative recommendation systems have achieved significant advances by leveraging semantic IDs to represent items. However, existing approaches that tokenize each modality independently face two critical limitations: (1) redundancy across modalities that reduces efficiency, and (2) failure to capture inter-modal interactions that limits item representation. We introduce FusID, a modality-fused semantic ID framework that addresses these limitations through three key components: (i) multimodal fusion that learns unified representations by jointly encoding information across modalities, (ii) representation learning that brings frequently co-occurring item embeddings closer while maintaining distinctiveness and preventing feature redundancy, and (iii) product quantization that converts the fused continuous embeddings into multiple discrete tokens to mitigate ID conflict. Evaluated on a multimodal next-song recommendation (i.e., playlist continuation) benchmark, FusID achieves zero ID conflicts, ensuring that each token sequence maps to exactly one song, mitigates codebook underutilization, and outperforms baselines in terms of MRR and Recall@k (k = 1, 5, 10, 20).
【2】Robust CAPTCHA Using Audio Illusions in the Era of Large Language Models: from Evaluation to Advances
标题:在大型语言模型时代使用音频幻象的稳健验证码:从评估到进步
链接:https://arxiv.org/abs/2601.08516
摘要:网站广泛使用CAPTCHA来阻止机器人和垃圾邮件,这给人类带来了挑战,但自动化程序很难解决。为了提高可访问性,音频验证码被设计为补充视觉验证码。然而,音频CAPTCHA对高级大型音频语言模型(LALM)和自动语音识别(ASR)模型的鲁棒性仍然不清楚。 在本文中,我们介绍了AI-CAPTCHA,这是一个统一的框架,它提供了(i)一个评估框架ACEval,其中包括先进的基于LALM和ASR的求解器,以及(ii)一种新的音频CAPTCHA方法IllusionAudio,利用音频错觉。通过对七个广泛部署的音频CAPTCHA的广泛评估,我们发现大多数现有方法可以通过高级LALM和ASR模型以高成功率解决,暴露出关键的安全漏洞。 为了解决这些漏洞,我们设计了一种新的音频CAPTCHA方法,IllusionAudio,它利用了植根于人类听觉机制的感知错觉线索。大量的实验表明,我们的方法击败了所有测试的LALM和基于ASR的攻击,同时实现了100%的人类通过率,显着优于现有的音频CAPTCHA方法。
摘要:CAPTCHAs are widely used by websites to block bots and spam by presenting challenges that are easy for humans but difficult for automated programs to solve. To improve accessibility, audio CAPTCHAs are designed to complement visual ones. However, the robustness of audio CAPTCHAs against advanced Large Audio Language Models (LALMs) and Automatic Speech Recognition (ASR) models remains unclear. In this paper, we introduce AI-CAPTCHA, a unified framework that offers (i) an evaluation framework, ACEval, which includes advanced LALM- and ASR-based solvers, and (ii) a novel audio CAPTCHA approach, IllusionAudio, leveraging audio illusions. Through extensive evaluations of seven widely deployed audio CAPTCHAs, we show that most existing methods can be solved with high success rates by advanced LALMs and ASR models, exposing critical security weaknesses. To address these vulnerabilities, we design a new audio CAPTCHA approach, IllusionAudio, which exploits perceptual illusion cues rooted in human auditory mechanisms. Extensive experiments demonstrate that our method defeats all tested LALM- and ASR-based attacks while achieving a 100% human pass rate, significantly outperforming existing audio CAPTCHA methods.
【3】Decoding Order Matters in Autoregressive Speech Synthesis
标题:自回归语音合成中的解码顺序很重要
链接:https://arxiv.org/abs/2601.08450
摘要:自回归语音合成通常采用从左到右的顺序,但生成顺序是一种建模选择。我们通过掩码扩散框架研究解码顺序,该框架在训练和推理过程中逐步揭开位置并允许任意解码顺序。通过身份和随机排列之间的插值,我们表明,解码顺序的随机性影响语音质量。我们进一步比较固定策略,如\texttt{l2 r}和\texttt{r2 l}与自适应策略,如Top-$K$,发现固定顺序解码,包括占主导地位的从左到右的方法,是次优的,而自适应解码产生更好的性能。最后,由于掩蔽扩散需要离散的输入,我们量化的声学表示,发现即使1位量化可以支持合理的高质量的语音。
摘要:Autoregressive speech synthesis often adopts a left-to-right order, yet generation order is a modelling choice. We investigate decoding order through masked diffusion framework, which progressively unmasks positions and allows arbitrary decoding orders during training and inference. By interpolating between identity and random permutations, we show that randomness in decoding order affects speech quality. We further compare fixed strategies, such as \texttt{l2r} and \texttt{r2l} with adaptive ones, such as Top-$K$, finding that fixed-order decoding, including the dominating left-to-right approach, is suboptimal, while adaptive decoding yields better performance. Finally, since masked diffusion requires discrete inputs, we quantise acoustic representations and find that even 1-bit quantisation can support reasonably high-quality speech.
【4】Decodable but not structured: linear probing enables Underwater Acoustic Target Recognition with pretrained audio embeddings
标题:可解码但非结构化:线性探测通过预训练的音频嵌入实现水下声学目标识别
链接:https://arxiv.org/abs/2601.08358
摘要:船舶的人为噪音水平不断提高,大大加剧了水下声污染,对海洋生态系统构成威胁。这使得监测对于了解和量化船舶辐射噪声的影响至关重要。被动声学监测(PAM)系统被广泛部署用于此目的,在不同的音景中产生多年的水下录音。对如此大规模的数据进行手动分析是不切实际的,因此需要基于机器学习的自动化方法。近年来,水声目标自动识别(UATR)主要依赖于有监督学习,但受标注数据稀缺的限制。迁移学习(TL)提供了一个有希望的替代方案来减轻这种限制。在这项工作中,我们对UATR的迁移学习进行了第一次实证比较研究,评估了来自不同音频域的多个预训练音频模型。预训练的模型权重被冻结,并通过分类、聚类和基于相似性的评估来分析结果嵌入。分析表明,嵌入空间的几何结构在很大程度上是由记录特定的特征。然而,一个简单的线性探测器可以有效地抑制这种记录特定的信息,并隔离船舶类型的功能,从这些嵌入。因此,线性探测能够以较低的计算成本使用预训练的音频模型实现有效的自动UATR,从而显著减少对大量高质量标记的船舶录音的需求。
摘要:Increasing levels of anthropogenic noise from ships contribute significantly to underwater sound pollution, posing risks to marine ecosystems. This makes monitoring crucial to understand and quantify the impact of the ship radiated noise. Passive Acoustic Monitoring (PAM) systems are widely deployed for this purpose, generating years of underwater recordings across diverse soundscapes. Manual analysis of such large-scale data is impractical, motivating the need for automated approaches based on machine learning. Recent advances in automatic Underwater Acoustic Target Recognition (UATR) have largely relied on supervised learning, which is constrained by the scarcity of labeled data. Transfer Learning (TL) offers a promising alternative to mitigate this limitation. In this work, we conduct the first empirical comparative study of transfer learning for UATR, evaluating multiple pretrained audio models originating from diverse audio domains. The pretrained model weights are frozen, and the resulting embeddings are analyzed through classification, clustering, and similarity-based evaluations. The analysis shows that the geometrical structure of the embedding space is largely dominated by recording-specific characteristics. However, a simple linear probe can effectively suppress this recording-specific information and isolate ship-type features from these embeddings. As a result, linear probing enables effective automatic UATR using pretrained audio models at low computational cost, significantly reducing the need for a large amounts of high-quality labeled ship recordings.
【5】VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
标题:VoxCog:基于方言知识的端到端多语言认知障碍分类
链接:https://arxiv.org/abs/2601.07999
摘要:在这项工作中,我们提出了一个新的角度对认知障碍分类语音集成语音基础模型,明确识别语音方言。我们的动机是基于观察到阿尔茨海默病(AD)或轻度认知障碍(MCI)的个体通常会产生可测量的语音特征,例如较慢的清晰度和延长的声音,其方式类似于语音中看到的方言语音变化。基于这一想法,我们引入了VoxCog,这是一个端到端的框架,它使用预先训练的方言模型来检测AD或MCI,而不依赖于文本或图像等其他形式。通过在多个多语言数据集上进行AD和MCI检测的实验,我们证明了在语音基础模型之上使用方言分类器进行模型初始化可以持续提高AD或MCI的预测性能。与以前的方法相比,我们训练的模型产生了类似或更好的性能,这些方法使用不同的信号模态集成了几种计算方法。特别是,我们的端到端基于语音的模型在ADReSS 2020挑战和ADReSSo 2021挑战测试集上实现了87.5%和85.9%的准确率,优于使用基于多模态集成的计算或LLM的现有解决方案。
摘要:In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals with Alzheimer's Disease (AD) or mild cognitive impairment (MCI) often produce measurable speech characteristics, such as slower articulation rate and lengthened sounds, in a manner similar to dialectal phonetic variations seen in speech. Building on this idea, we introduce VoxCog, an end-to-end framework that uses pre-trained dialect models to detect AD or MCI without relying on additional modalities such as text or images. Through experiments on multiple multilingual datasets for AD and MCI detection, we demonstrate that model initialization with a dialect classifier on top of speech foundation models consistently improves the predictive performance of AD or MCI. Our trained models yield similar or often better performance compared to previous approaches that ensembled several computational methods using different signal modalities. Particularly, our end-to-end speech-based model achieves 87.5% and 85.9% accuracy on the ADReSS 2020 challenge and ADReSSo 2021 challenge test sets, outperforming existing solutions that use multimodal ensemble-based computation or LLMs.
【6】LJ-Spoof: A Generatively Varied Corpus for Audio Anti-Spoofing and Synthesis Source Tracing
标题:LJ-Spoof:一个用于音频反欺骗和合成源跟踪的生成性变化语料库
链接:https://arxiv.org/abs/2601.07958
摘要:特定于说话人的反欺骗和合成源跟踪是音频反欺骗中的核心挑战。由于缺乏系统地改变模型架构,合成管道和生成参数的数据集,进展受到阻碍。为了解决这一差距,我们介绍了LJ-Spoof,一个特定于说话者的,生成多样的语料库,系统地改变韵律,声码器,生成超参数,真正的提示源,训练制度和神经后处理。语料库涵盖一个扬声器,包括录音室质量的录音,30 TTS家庭,500 generatively变量子集,10个真正的神经处理变量,超过300万的话语。这种变化密集的设计实现了鲁棒的说话者条件反欺骗和细粒度的合成源跟踪。我们进一步将此数据集定位为实用的参考培训资源和反欺骗和源跟踪的基准评估套件。
摘要:Speaker-specific anti-spoofing and synthesis-source tracing are central challenges in audio anti-spoofing. Progress has been hampered by the lack of datasets that systematically vary model architectures, synthesis pipelines, and generative parameters. To address this gap, we introduce LJ-Spoof, a speaker-specific, generatively diverse corpus that systematically varies prosody, vocoders, generative hyperparameters, bona fide prompt sources, training regimes, and neural post-processing. The corpus spans one speakers-including studio-quality recordings-30 TTS families, 500 generatively variant subsets, 10 bona fide neural-processing variants, and more than 3 million utterances. This variation-dense design enables robust speaker-conditioned anti-spoofing and fine-grained synthesis-source tracing. We further position this dataset as both a practical reference training resource and a benchmark evaluation suite for anti-spoofing and source tracing.
【7】Elastic overtones: an equal temperament 12 tone music system with "perfect" fifths
链接:https://arxiv.org/abs/2601.08074
备注:14 pages, 4 figures, 6 audio files
摘要:八度音阶的转置12 π 1调音的不可能性源于数学事实,即,五次谐波的二次谐波不能精确地匹配基波的三次谐波。这反过来又源于西方音乐的整数谐波结构,以及随后的倍频程间隔作为频率2的倍数的基本特征,这是我们的音乐系统从具有振动元件的乐器的物理学中继承的一个属性,该属性很好地近似于一维。在当今的电子音乐时代,人们可以放松上述假设来构建一个类似的音乐系统,其中保留了标准音乐系统的所有结构属性,但谐波不是基频的整数倍,并且八度音不再是频率的2倍。这样就可以构建一个可转置的12和弦音乐系统,其中第五和弦的二次谐波与第三和弦的基波完全匹配。该系统的增强的谐波质量恢复到一个很好的近似的音乐品质的公正语调,同时保留建设的所有多功能性和调制能力的12 TET。
摘要:The impossibility of a transposable 12 semitone tuning of the octave arises from the mathematical fact that $2 \times 2^{7/12} \neq 3$ i.e., the second harmonic of the fifth can not exactly match the third harmonic of the fundamental. This in turn, stems from the whole number harmonic structure of western music, and the subsequent fundamental character of the octave interval as multiples of 2 in frequency, a property inherited by our music system from the physics of instruments with vibrating elements being to a good approximation one dimensional. In the current era of electronic music, one can relax the above assumptions to construct an analogous music system where all the structural properties of the standard music system are preserved, but where harmonics are not whole number multiples of the fundamental frequency, and the octave is no longer a factor of 2 in frequency. This now allows to construct a transposable 12 semitone music system where the second harmonic of the fifth exactly matches the third harmonic of the fundamental. The enhanced harmonic qualities of this system recover to a good approximation the musical qualities of Just Intonation, whilst retaining by construction all the versatility and modulating ability of 12TET.
【8】Tuberculosis Screening from Cough Audio: Baseline Models, Clinical Variables, and Uncertainty Quantification
标题:来自咳嗽音频的结核病筛查:基线模型、临床变量和不确定性量化
链接:https://arxiv.org/abs/2601.07969
摘要:在本文中,我们提出了一个标准化的框架,用于使用机器学习从咳嗽音频和常规收集的临床数据中自动检测结核病(TB)。虽然从音频进行结核病筛查已经引起了越来越多的关注,但进展难以衡量,因为现有的研究在数据集、队列定义、特征表示、模型族、验证协议和报告的指标方面存在很大差异。因此,报告的收益往往不能直接比较,目前还不清楚改进是源于建模的进步还是数据和评价的差异。我们通过使用咳嗽记录和来自几个国家的最近汇编的数据集的伴随临床元数据建立强有力的、有充分记录的结核病预测基线来解决这一差距。我们的管道是可重复的端到端,涵盖特征提取,多模态融合,咳嗽独立评估和不确定性量化,它报告了一套一致的临床相关指标,以实现公平的比较。我们进一步量化了仅咳嗽音频和融合(音频+临床元数据)模型的性能,并发布了完整的实验协议以促进基准测试。这一基线旨在作为一个共同的参考点,减少目前阻碍实地进展的方法差异。
摘要:In this paper, we propose a standardized framework for automatic tuberculosis (TB) detection from cough audio and routinely collected clinical data using machine learning. While TB screening from audio has attracted growing interest, progress is difficult to measure because existing studies vary substantially in datasets, cohort definitions, feature representations, model families, validation protocols, and reported metrics. Consequently, reported gains are often not directly comparable, and it remains unclear whether improvements stem from modeling advances or from differences in data and evaluation. We address this gap by establishing a strong, well-documented baseline for TB prediction using cough recordings and accompanying clinical metadata from a recently compiled dataset from several countries. Our pipeline is reproducible end-to-end, covering feature extraction, multimodal fusion, cougher-independent evaluation, and uncertainty quantification, and it reports a consistent suite of clinically relevant metrics to enable fair comparison. We further quantify performance for cough audio-only and fused (audio + clinical metadata) models, and release the full experimental protocol to facilitate benchmarking. This baseline is intended to serve as a common reference point and to reduce methodological variance that currently holds back progress in the field.
【1】Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
标题:通过TI-SDRM进行弱监督的Tabla Stroke转录:节奏感知格子重新评分框架
链接:https://arxiv.org/abs/2601.08537
摘要:Tabla Stroke Transcription(TST)是分析印度斯坦古典音乐节奏结构的核心,但由于复杂的节奏组织和缺乏强有力的注释数据,仍然具有挑战性。现有的方法在很大程度上依赖于具有起始级注释的完全监督学习,这在规模上是昂贵且不切实际的。这项工作解决了TST在弱监督设置,只使用符号笔划序列没有时间对齐。我们提出了一个框架,结合了基于CTC的声学模型与序列级节奏重新评分。声学模型产生一个解码网格,它是使用一个独立的静态-动态节奏模型(TI-SDRM)通过自适应插值机制将长期节奏结构与短期自适应动态相结合。我们策划了一个新的真实世界的Tabla独奏数据集和一个互补的合成数据集,建立了印度斯坦古典音乐中弱监督TST的第一个基准。实验表明,一致的和实质性的减少,在笔画错误率在声学解码,确认明确的节奏结构的重要性,准确的转录。
摘要:Tabla Stroke Transcription (TST) is central to the analysis of rhythmic structure in Hindustani classical music, yet remains challenging due to complex rhythmic organization and the scarcity of strongly annotated data. Existing approaches largely rely on fully supervised learning with onset-level annotations, which are costly and impractical at scale. This work addresses TST in a weakly supervised setting, using only symbolic stroke sequences without temporal alignment. We propose a framework that combines a CTC-based acoustic model with sequence-level rhythmic rescoring. The acoustic model produces a decoding lattice, which is refined using a \textbf{$T\bar{a}la$}-Independent Static--Dynamic Rhythmic Model (TI-SDRM) that integrates long-term rhythmic structure with short-term adaptive dynamics through an adaptive interpolation mechanism. We curate a new real-world tabla solo dataset and a complementary synthetic dataset, establishing the first benchmark for weakly supervised TST in Hindustani classical music. Experiments demonstrate consistent and substantial reductions in stroke error rate over acoustic-only decoding, confirming the importance of explicit rhythmic structure for accurate transcription.
【2】Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection
标题:异常声音检测代理任务的定量分析
链接:https://arxiv.org/abs/2601.08480
备注:13 pages, 5 figures, Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing
摘要:由于异常样本的稀缺性,异常声音检测(ASD)通常涉及自监督代理任务,以从正常声音数据中学习特征表示。在ASD研究中,诸如AutoEncoders之类的代理任务在明确的假设下运行,即在正常数据上训练的模型将增加与异常相关的重建错误。一个自然的延伸表明,提高代理任务的性能应提高ASD的能力,但是,这种关系很少受到系统的关注。这项研究通过定量分析代理任务指标与ASD性能之间的关系来解决这一研究空白,这五种配置分别是AutoEncoders、分类、源分离、对比学习和预训练模型。我们使用线性探测(线性可分性)和马氏距离(分布紧性)来评估学习的表示。我们的实验表明,强大的代理性能并不一定提高异常声音检测性能。具体来说,分类任务由于任务难度不足而经历性能饱和,而对比学习由于数据多样性有限而无法学习有意义的特征。值得注意的是,源分离是唯一的任务表现出强正相关性,这样改进的分离一致地提高异常检测。基于这些发现,我们强调了任务难度和目标对齐的至关重要性。最后,我们提出了一个三阶段的对齐验证协议,以指导ASD系统的高效代理任务的设计。
摘要:Anomalous sound detection (ASD) typically involves self-supervised proxy tasks to learn feature representations from normal sound data, owing to the scarcity of anomalous samples. In ASD research, proxy tasks such as AutoEncoders operate under the explicit assumption that models trained on normal data will increase the reconstruction errors related to anomalies. A natural extension suggests that improved proxy task performance should improve ASD capability; however, this relationship has received little systematic attention. This study addresses this research gap by quantitatively analyzing the relationship between proxy task metrics and ASD performance across five configurations, namely, AutoEncoders, classification, source separation, contrastive learning, and pre-trained models. We evaluate the learned representations using linear probe (linear separability) and Mahalanobis distance (distributional compactness). Our experiments reveal that strong proxy performance does not necessarily improve anomalous sound detection performance. Specifically, classification tasks experience performance saturation owing to insufficient task difficulty, whereas contrastive learning fails to learn meaningful features owing to limited data diversity. Notably, source separation is the only task demonstrating a strong positive correlation, such that improved separation consistently improves anomaly detection. Based on these findings, we highlight the critical importance of task difficulty and objective alignment. Finally, we propose a three-stage alignment verification protocol to guide the design of highly effective proxy tasks for ASD systems.
【3】Tuberculosis Screening from Cough Audio: Baseline Models, Clinical Variables, and Uncertainty Quantification
标题:来自咳嗽音频的结核病筛查:基线模型、临床变量和不确定性量化
链接:https://arxiv.org/abs/2601.07969
摘要:在本文中,我们提出了一个标准化的框架,用于使用机器学习从咳嗽音频和常规收集的临床数据中自动检测结核病(TB)。虽然从音频进行结核病筛查已经引起了越来越多的关注,但进展难以衡量,因为现有的研究在数据集、队列定义、特征表示、模型族、验证协议和报告的指标方面存在很大差异。因此,报告的收益往往不能直接比较,目前还不清楚改进是源于建模的进步还是数据和评价的差异。我们通过使用咳嗽记录和来自几个国家的最近汇编的数据集的伴随临床元数据建立强有力的、有充分记录的结核病预测基线来解决这一差距。我们的管道是可重复的端到端,涵盖特征提取,多模态融合,咳嗽独立评估和不确定性量化,它报告了一套一致的临床相关指标,以实现公平的比较。我们进一步量化了仅咳嗽音频和融合(音频+临床元数据)模型的性能,并发布了完整的实验协议以促进基准测试。这一基线旨在作为一个共同的参考点,减少目前阻碍实地进展的方法差异。
摘要:In this paper, we propose a standardized framework for automatic tuberculosis (TB) detection from cough audio and routinely collected clinical data using machine learning. While TB screening from audio has attracted growing interest, progress is difficult to measure because existing studies vary substantially in datasets, cohort definitions, feature representations, model families, validation protocols, and reported metrics. Consequently, reported gains are often not directly comparable, and it remains unclear whether improvements stem from modeling advances or from differences in data and evaluation. We address this gap by establishing a strong, well-documented baseline for TB prediction using cough recordings and accompanying clinical metadata from a recently compiled dataset from several countries. Our pipeline is reproducible end-to-end, covering feature extraction, multimodal fusion, cougher-independent evaluation, and uncertainty quantification, and it reports a consistent suite of clinically relevant metrics to enable fair comparison. We further quantify performance for cough audio-only and fused (audio + clinical metadata) models, and release the full experimental protocol to facilitate benchmarking. This baseline is intended to serve as a common reference point and to reduce methodological variance that currently holds back progress in the field.
【4】FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
标题:FuID:用于生成性音乐推荐的情态融合语义ID
链接:https://arxiv.org/abs/2601.08764
摘要:生成式推荐系统通过利用语义ID来表示项目,取得了显著的进步。然而,独立地标记每个模态的现有方法面临两个关键限制:(1)降低效率的跨模态冗余,以及(2)无法捕获限制项目表示的模态间交互。我们引入FusID,一个模态融合的语义ID框架,通过三个关键组件解决这些限制:(i)多模态融合,其通过跨模态联合编码信息来学习统一的表示,(ii)表示学习,其使频繁共现的项目嵌入更接近,同时保持独特性并防止特征冗余,以及(iii)将融合的连续嵌入转换成多个离散令牌以减轻ID冲突的乘积量化。在多模式下一首歌曲推荐上进行评估(即,在播放列表延续(playlist continuation)基准测试中,FusID实现了零ID冲突,确保每个标记序列映射到恰好一首歌曲,减轻了码本利用不足,并且在MRR和Recall@k(k = 1,5,10,20)方面优于基线。
摘要:Generative recommendation systems have achieved significant advances by leveraging semantic IDs to represent items. However, existing approaches that tokenize each modality independently face two critical limitations: (1) redundancy across modalities that reduces efficiency, and (2) failure to capture inter-modal interactions that limits item representation. We introduce FusID, a modality-fused semantic ID framework that addresses these limitations through three key components: (i) multimodal fusion that learns unified representations by jointly encoding information across modalities, (ii) representation learning that brings frequently co-occurring item embeddings closer while maintaining distinctiveness and preventing feature redundancy, and (iii) product quantization that converts the fused continuous embeddings into multiple discrete tokens to mitigate ID conflict. Evaluated on a multimodal next-song recommendation (i.e., playlist continuation) benchmark, FusID achieves zero ID conflicts, ensuring that each token sequence maps to exactly one song, mitigates codebook underutilization, and outperforms baselines in terms of MRR and Recall@k (k = 1, 5, 10, 20).
【5】Robust CAPTCHA Using Audio Illusions in the Era of Large Language Models: from Evaluation to Advances
标题:在大型语言模型时代使用音频幻象的稳健验证码:从评估到进步
链接:https://arxiv.org/abs/2601.08516
摘要:网站广泛使用CAPTCHA来阻止机器人和垃圾邮件,这给人类带来了挑战,但自动化程序很难解决。为了提高可访问性,音频验证码被设计为补充视觉验证码。然而,音频CAPTCHA对高级大型音频语言模型(LALM)和自动语音识别(ASR)模型的鲁棒性仍然不清楚。 在本文中,我们介绍了AI-CAPTCHA,这是一个统一的框架,它提供了(i)一个评估框架ACEval,其中包括先进的基于LALM和ASR的求解器,以及(ii)一种新的音频CAPTCHA方法IllusionAudio,利用音频错觉。通过对七个广泛部署的音频CAPTCHA的广泛评估,我们发现大多数现有方法可以通过高级LALM和ASR模型以高成功率解决,暴露出关键的安全漏洞。 为了解决这些漏洞,我们设计了一种新的音频CAPTCHA方法,IllusionAudio,它利用了植根于人类听觉机制的感知错觉线索。大量的实验表明,我们的方法击败了所有测试的LALM和基于ASR的攻击,同时实现了100%的人类通过率,显着优于现有的音频CAPTCHA方法。
摘要:CAPTCHAs are widely used by websites to block bots and spam by presenting challenges that are easy for humans but difficult for automated programs to solve. To improve accessibility, audio CAPTCHAs are designed to complement visual ones. However, the robustness of audio CAPTCHAs against advanced Large Audio Language Models (LALMs) and Automatic Speech Recognition (ASR) models remains unclear. In this paper, we introduce AI-CAPTCHA, a unified framework that offers (i) an evaluation framework, ACEval, which includes advanced LALM- and ASR-based solvers, and (ii) a novel audio CAPTCHA approach, IllusionAudio, leveraging audio illusions. Through extensive evaluations of seven widely deployed audio CAPTCHAs, we show that most existing methods can be solved with high success rates by advanced LALMs and ASR models, exposing critical security weaknesses. To address these vulnerabilities, we design a new audio CAPTCHA approach, IllusionAudio, which exploits perceptual illusion cues rooted in human auditory mechanisms. Extensive experiments demonstrate that our method defeats all tested LALM- and ASR-based attacks while achieving a 100% human pass rate, significantly outperforming existing audio CAPTCHA methods.
【6】Decoding Order Matters in Autoregressive Speech Synthesis
标题:自回归语音合成中的解码顺序很重要
链接:https://arxiv.org/abs/2601.08450
摘要:自回归语音合成通常采用从左到右的顺序,但生成顺序是一种建模选择。我们通过掩码扩散框架研究解码顺序,该框架在训练和推理过程中逐步揭开位置并允许任意解码顺序。通过身份和随机排列之间的插值,我们表明,解码顺序的随机性影响语音质量。我们进一步比较固定策略,如\texttt{l2 r}和\texttt{r2 l}与自适应策略,如Top-$K$,发现固定顺序解码,包括占主导地位的从左到右的方法,是次优的,而自适应解码产生更好的性能。最后,由于掩蔽扩散需要离散的输入,我们量化的声学表示,发现即使1位量化可以支持合理的高质量的语音。
摘要:Autoregressive speech synthesis often adopts a left-to-right order, yet generation order is a modelling choice. We investigate decoding order through masked diffusion framework, which progressively unmasks positions and allows arbitrary decoding orders during training and inference. By interpolating between identity and random permutations, we show that randomness in decoding order affects speech quality. We further compare fixed strategies, such as \texttt{l2r} and \texttt{r2l} with adaptive ones, such as Top-$K$, finding that fixed-order decoding, including the dominating left-to-right approach, is suboptimal, while adaptive decoding yields better performance. Finally, since masked diffusion requires discrete inputs, we quantise acoustic representations and find that even 1-bit quantisation can support reasonably high-quality speech.
【7】Decodable but not structured: linear probing enables Underwater Acoustic Target Recognition with pretrained audio embeddings
标题:可解码但非结构化:线性探测通过预训练的音频嵌入实现水下声学目标识别
链接:https://arxiv.org/abs/2601.08358
摘要:船舶的人为噪音水平不断提高,大大加剧了水下声污染,对海洋生态系统构成威胁。这使得监测对于了解和量化船舶辐射噪声的影响至关重要。被动声学监测(PAM)系统被广泛部署用于此目的,在不同的音景中产生多年的水下录音。对如此大规模的数据进行手动分析是不切实际的,因此需要基于机器学习的自动化方法。近年来,水声目标自动识别(UATR)主要依赖于有监督学习,但受标注数据稀缺的限制。迁移学习(TL)提供了一个有希望的替代方案来减轻这种限制。在这项工作中,我们对UATR的迁移学习进行了第一次实证比较研究,评估了来自不同音频域的多个预训练音频模型。预训练的模型权重被冻结,并通过分类、聚类和基于相似性的评估来分析结果嵌入。分析表明,嵌入空间的几何结构在很大程度上是由记录特定的特征。然而,一个简单的线性探测器可以有效地抑制这种记录特定的信息,并隔离船舶类型的功能,从这些嵌入。因此,线性探测能够以较低的计算成本使用预训练的音频模型实现有效的自动UATR,从而显著减少对大量高质量标记的船舶录音的需求。
摘要:Increasing levels of anthropogenic noise from ships contribute significantly to underwater sound pollution, posing risks to marine ecosystems. This makes monitoring crucial to understand and quantify the impact of the ship radiated noise. Passive Acoustic Monitoring (PAM) systems are widely deployed for this purpose, generating years of underwater recordings across diverse soundscapes. Manual analysis of such large-scale data is impractical, motivating the need for automated approaches based on machine learning. Recent advances in automatic Underwater Acoustic Target Recognition (UATR) have largely relied on supervised learning, which is constrained by the scarcity of labeled data. Transfer Learning (TL) offers a promising alternative to mitigate this limitation. In this work, we conduct the first empirical comparative study of transfer learning for UATR, evaluating multiple pretrained audio models originating from diverse audio domains. The pretrained model weights are frozen, and the resulting embeddings are analyzed through classification, clustering, and similarity-based evaluations. The analysis shows that the geometrical structure of the embedding space is largely dominated by recording-specific characteristics. However, a simple linear probe can effectively suppress this recording-specific information and isolate ship-type features from these embeddings. As a result, linear probing enables effective automatic UATR using pretrained audio models at low computational cost, significantly reducing the need for a large amounts of high-quality labeled ship recordings.
【8】Elastic overtones: an equal temperament 12 tone music system with "perfect" fifths
链接:https://arxiv.org/abs/2601.08074
备注:14 pages, 4 figures, 6 audio files
摘要:八度音阶的转置12 π 1调音的不可能性源于数学事实,即,五次谐波的二次谐波不能精确地匹配基波的三次谐波。这反过来又源于西方音乐的整数谐波结构,以及随后的倍频程间隔作为频率2的倍数的基本特征,这是我们的音乐系统从具有振动元件的乐器的物理学中继承的一个属性,该属性很好地近似于一维。在当今的电子音乐时代,人们可以放松上述假设来构建一个类似的音乐系统,其中保留了标准音乐系统的所有结构属性,但谐波不是基频的整数倍,并且倍频程不再是频率的2倍。这样就可以构建一个可转置的12和弦音乐系统,其中第五和弦的二次谐波与第三和弦的基波完全匹配。该系统的增强的谐波质量恢复到一个很好的近似的音乐品质的公正语调,同时保留建设的所有多功能性和调制能力的12 TET。
摘要:The impossibility of a transposable 12 semitone tuning of the octave arises from the mathematical fact that $2 \times 2^{7/12} \neq 3$ i.e., the second harmonic of the fifth can not exactly match the third harmonic of the fundamental. This in turn, stems from the whole number harmonic structure of western music, and the subsequent fundamental character of the octave interval as multiples of 2 in frequency, a property inherited by our music system from the physics of instruments with vibrating elements being to a good approximation one dimensional. In the current era of electronic music, one can relax the above assumptions to construct an analogous music system where all the structural properties of the standard music system are preserved, but where harmonics are not whole number multiples of the fundamental frequency, and the octave is no longer a factor of 2 in frequency. This now allows to construct a transposable 12 semitone music system where the second harmonic of the fifth exactly matches the third harmonic of the fundamental. The enhanced harmonic qualities of this system recover to a good approximation the musical qualities of Just Intonation, whilst retaining by construction all the versatility and modulating ability of 12TET.
【9】VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
标题:VoxCog:基于方言知识的端到端多语言认知障碍分类
链接:https://arxiv.org/abs/2601.07999
摘要:在这项工作中,我们提出了一个新的角度对认知障碍分类语音集成语音基础模型,明确识别语音方言。我们的动机是基于观察到阿尔茨海默病(AD)或轻度认知障碍(MCI)的个体通常会产生可测量的语音特征,例如较慢的清晰度和延长的声音,其方式类似于语音中看到的方言语音变化。基于这一想法,我们引入了VoxCog,这是一个端到端的框架,它使用预先训练的方言模型来检测AD或MCI,而不依赖于文本或图像等其他形式。通过在多个多语言数据集上进行AD和MCI检测的实验,我们证明了在语音基础模型之上使用方言分类器进行模型初始化可以持续提高AD或MCI的预测性能。与以前的方法相比,我们训练的模型产生了类似或更好的性能,这些方法使用不同的信号模态集成了几种计算方法。特别是,我们的端到端基于语音的模型在ADReSS 2020挑战和ADReSSo 2021挑战测试集上实现了87.5%和85.9%的准确率,优于使用基于多模态集成的计算或LLM的现有解决方案。
摘要:In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals with Alzheimer's Disease (AD) or mild cognitive impairment (MCI) often produce measurable speech characteristics, such as slower articulation rate and lengthened sounds, in a manner similar to dialectal phonetic variations seen in speech. Building on this idea, we introduce VoxCog, an end-to-end framework that uses pre-trained dialect models to detect AD or MCI without relying on additional modalities such as text or images. Through experiments on multiple multilingual datasets for AD and MCI detection, we demonstrate that model initialization with a dialect classifier on top of speech foundation models consistently improves the predictive performance of AD or MCI. Our trained models yield similar or often better performance compared to previous approaches that ensembled several computational methods using different signal modalities. Particularly, our end-to-end speech-based model achieves 87.5% and 85.9% accuracy on the ADReSS 2020 challenge and ADReSSo 2021 challenge test sets, outperforming existing solutions that use multimodal ensemble-based computation or LLMs.
【10】LJ-Spoof: A Generatively Varied Corpus for Audio Anti-Spoofing and Synthesis Source Tracing
标题:LJ-Spoof:一个用于音频反欺骗和合成源跟踪的生成性变化语料库
链接:https://arxiv.org/abs/2601.07958
摘要:特定于说话人的反欺骗和合成源跟踪是音频反欺骗中的核心挑战。由于缺乏系统地改变模型架构,合成管道和生成参数的数据集,进展受到阻碍。为了解决这一差距,我们介绍了LJ-Spoof,一个特定于说话者的,生成多样的语料库,系统地改变韵律,声码器,生成超参数,真正的提示源,训练制度和神经后处理。语料库涵盖一个扬声器,包括录音室质量的录音,30 TTS家庭,500 generatively变量子集,10个真正的神经处理变量,超过300万的话语。这种变化密集的设计实现了鲁棒的说话者条件反欺骗和细粒度的合成源跟踪。我们进一步将此数据集定位为实用的参考培训资源和反欺骗和源跟踪的基准评估套件。
摘要:Speaker-specific anti-spoofing and synthesis-source tracing are central challenges in audio anti-spoofing. Progress has been hampered by the lack of datasets that systematically vary model architectures, synthesis pipelines, and generative parameters. To address this gap, we introduce LJ-Spoof, a speaker-specific, generatively diverse corpus that systematically varies prosody, vocoders, generative hyperparameters, bona fide prompt sources, training regimes, and neural post-processing. The corpus spans one speakers-including studio-quality recordings-30 TTS families, 500 generatively variant subsets, 10 bona fide neural-processing variants, and more than 3 million utterances. This variation-dense design enables robust speaker-conditioned anti-spoofing and fine-grained synthesis-source tracing. We further position this dataset as both a practical reference training resource and a benchmark evaluation suite for anti-spoofing and source tracing.
机器翻译由腾讯交互翻译提供,仅供参考
