今日论文合集:cs.SD语音13篇,eess.AS音频处理15篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Machine Learning Approaches to Vocal Register Classification in  Contemporary Male Pop Music
标题: 当代男性流行音乐中人声语域分类的机器学习方法
链接:https://arxiv.org/abs/2505.11378
作者: Alexander Kim,  Charlotte Botha 
备注:8 pages, 8 figures
摘要:对于所有经验水平的歌手来说,学习技术曲目最艰巨的挑战之一是在passagio(胸部声音和头部声音区域之间的通道)及其周围导航位置和声音区域。特别是在流行音乐中,其中单个艺术家可以使用各种音色和纹理来实现期望的质量,可能难以识别歌手正在使用的音域内的声音音域。本文提出了两种方法,通过分析梅尔频谱图图像的纹理特征来分类男性流行音乐音频信号中的声区。此外,我们将讨论这些模型的声乐分析工具的实际集成,并介绍一个同时开发的软件称为AVRA,代表自动声乐语域分析。我们提出的方法通过支持向量机(SVM)和卷积神经网络(CNN)模型实现了一致的分类,这支持了在更多的声音类型和歌唱流派中更强大的分类可能性的承诺。
摘要:For singers of all experience levels, one of the most daunting challenges in learning technical repertoire is navigating placement and vocal register in and around the passagio (passage between chest voice and head voice registers). Particularly in pop music, where a single artist may use a variety of timbre's and textures to achieve a desired quality, it can be difficult to identify what vocal register within the vocal range a singer is using. This paper presents two methods for classifying vocal registers in an audio signal of male pop music through the analysis of textural features of mel-spectrogram images. Additionally, we will discuss the practical integration of these models for vocal analysis tools, and introduce a concurrently developed software called AVRA which stands for Automatic Vocal Register Analysis. Our proposed methods achieved consistent classification of vocal register through both Support Vector Machine (SVM) and Convolutional Neural Network (CNN) models, which supports the promise of more robust classification possibilities across more voice types and genres of singing.


【2】 LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors

标题: LegoLAM:使用CTCPosteriors将LLM与语音编码器连接起来
链接:https://arxiv.org/abs/2505.11352
作者: Rao Ma,  Tongzhou Chen,  Kartik Audhkhasi,  Bhuvana Ramabhadran 
摘要:最近,大规模预训练语音编码器和大型语言模型(LLM)已经发布,它们在包括自动语音识别(ASR)在内的一系列口语处理任务上表现出最先进的性能。为了有效地结合这两种模型以获得更好的性能,采用了连续语音提示和ASR纠错。然而,这些方法倾向于次优性能或不灵活。在本文中,我们提出了一个新的范例,LegoSLM,桥梁语音编码器和LLM使用ASR后验矩阵。语音编码器被训练以在LLM词汇表上生成连接主义时间分类(CTC)后验,其用于通过计算LLM输入嵌入的加权和来重建伪音频嵌入。这些嵌入与LLM输入空间中的文本嵌入连接在一起。使用性能良好的USM和Gemma模型作为例子,我们证明了我们提出的LegoSLM方法在ASR和语音翻译任务上都具有良好的性能。通过将USM与Gemma模型连接起来,我们可以在8个MLS测试集上获得超过USM-CTC基线的平均49%的WERR。经过训练的模型还在一系列设置中表现出模块性-在微调Gemma模型权重之后,语音编码器可以切换并以zero-shot方式与LLM组合。此外,我们建议使用softmax温度来控制USM和LLM的解码时间影响,这在域自适应中显示出有效性。
摘要:Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language processing tasks including Automatic Speech Recognition (ASR). To effectively combine both models for better performance, continuous speech prompts, and ASR error correction have been adopted. However, these methods are prone to suboptimal performance or are inflexible. In this paper, we propose a new paradigm, LegoSLM, that bridges speech encoders and LLMs using the ASR posterior matrices. The speech encoder is trained to generate Connectionist Temporal Classification (CTC) posteriors over the LLM vocabulary, which are used to reconstruct pseudo-audio embeddings by computing a weighted sum of the LLM input embeddings. These embeddings are concatenated with text embeddings in the LLM input space. Using the well-performing USM and Gemma models as an example, we demonstrate that our proposed LegoSLM method yields good performance on both ASR and speech translation tasks. By connecting USM with Gemma models, we can get an average of 49% WERR over the USM-CTC baseline on 8 MLS testsets. The trained model also exhibits modularity in a range of settings -- after fine-tuning the Gemma model weights, the speech encoder can be switched and combined with the LLM in a zero-shot fashion. Additionally, we propose to control the decode-time influence of the USM and LLM using a softmax temperature, which shows effectiveness in domain adaptation.


【3】 Improving Inference-Time Optimisation for Vocal Effects Style Transfer  with a Gaussian Prior

标题: 利用高斯先验改进声乐效果风格转移的推理时优化
链接:https://arxiv.org/abs/2505.11315
作者: Chin-Yun Yu,  Marco A. Martínez-Ramírez,  Junghyun Koo,  Wei-Hsiang Liao,  Yuki Mitsufuji,  György Fazekas 
备注:Submitted to WASPAA 2025
摘要:具有推理时间优化的风格转移(ST-ITO)是用于将参考音频的应用效果转移到原始音轨的最近方法。它优化效果参数,以最大限度地减少处理后的音频和参考之间的风格嵌入的距离。然而,这种方法平等地对待所有可能的配置,并且仅依赖于嵌入空间,这可能导致不切实际或有偏见的结果。我们通过在参数空间上引入从声乐预设数据集DiffVox导出的高斯先验来解决这个问题。由此产生的优化相当于最大后验估计。对MedleyDB数据集上的声音效果转移的评估显示,与基线相比,各项指标都有显著改善,包括盲音频效果估计器、最近邻方法和未校准的ST-ITO。建议的校准减少参数的均方误差高达33%,更好地匹配参考样式。16名参与者的主观评价证实了我们的方法的优越性,特别是在有限的数据制度。这项工作演示了如何在推理时间中结合先验知识来增强音频效果传输,为更有效和更逼真的音频处理系统铺平道路。
摘要:Style Transfer with Inference-Time Optimisation (ST-ITO) is a recent approach for transferring the applied effects of a reference audio to a raw audio track. It optimises the effect parameters to minimise the distance between the style embeddings of the processed audio and the reference. However, this method treats all possible configurations equally and relies solely on the embedding space, which can lead to unrealistic or biased results. We address this pitfall by introducing a Gaussian prior derived from a vocal preset dataset, DiffVox, over the parameter space. The resulting optimisation is equivalent to maximum-a-posteriori estimation. Evaluations on vocal effects transfer on the MedleyDB dataset show significant improvements across metrics compared to baselines, including a blind audio effects estimator, nearest-neighbour approaches, and uncalibrated ST-ITO. The proposed calibration reduces parameter mean squared error by up to 33% and matches the reference style better. Subjective evaluations with 16 participants confirm our method's superiority, especially in limited data regimes. This work demonstrates how incorporating prior knowledge in inference time enhances audio effects transfer, paving the way for more effective and realistic audio processing systems.


【4】 Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI  models in Sound Localization

标题: 看到声音,听到视觉:揭示声音定位中人工智能模型的情态偏差和冲突
链接:https://arxiv.org/abs/2505.11217
作者: Yanhao Jia,  Ji Xie,  S Jivaganesh,  Hao Li,  Xu Wu,  Mengmi Zhang 
备注:16 pages, 14 figures
摘要:想象一下,听到狗叫,转向声音,只看到一辆停着的汽车,而真正的,沉默的狗坐在其他地方。这种感官冲突测试感知,但人类通过优先考虑声音而不是误导性的视觉来可靠地解决它们。尽管多模态AI集成了视觉和音频,但人们对这些系统如何处理跨模态冲突或它们是否支持一种模态知之甚少。在这项研究中,我们系统地研究了AI声音定位中的模态偏差和冲突解决。我们评估领先的多模态模型,并将其与人类在心理物理学实验中的表现进行比较,包括六种视听条件,包括一致,冲突和缺乏线索。人类的表现一直优于人工智能,通过依赖听觉信息,表现出对冲突或缺失视觉的卓越弹性。相比之下,人工智能模型通常默认为视觉输入,将性能降低到接近偶然的水平。为了解决这个问题,我们使用通过3D模拟生成的立体声音频图像数据集来微调最先进的模型。即使训练数据有限,改进后的模型也超过了现有的基准。值得注意的是,它也反映了类似人类的水平定位偏差,有利于左右精度-可能是由于立体声音频结构反映了人耳的位置。这些发现强调了感官输入质量和系统架构如何塑造多模态表征的准确性。
摘要:Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere. Such sensory conflicts test perception, yet humans reliably resolve them by prioritizing sound over misleading visuals. Despite advances in multimodal AI integrating vision and audio, little is known about how these systems handle cross-modal conflicts or whether they favor one modality. In this study, we systematically examine modality bias and conflict resolution in AI sound localization. We assess leading multimodal models and benchmark them against human performance in psychophysics experiments across six audiovisual conditions, including congruent, conflicting, and absent cues. Humans consistently outperform AI, demonstrating superior resilience to conflicting or missing visuals by relying on auditory information. In contrast, AI models often default to visual input, degrading performance to near chance levels. To address this, we finetune a state-of-the-art model using a stereo audio-image dataset generated via 3D simulations. Even with limited training data, the refined model surpasses existing benchmarks. Notably, it also mirrors human-like horizontal localization bias favoring left-right precision-likely due to the stereo audio structure reflecting human ear placement. These findings underscore how sensory input quality and system architecture shape multimodal representation accuracy.


【5】 Audio Turing Test: Benchmarking the Human-likeness of Large Language  Model-based Text-to-Speech Systems in Chinese

标题: 音频图灵测试:基于大型语言模型的中文文本到语音系统的人形基准
链接:https://arxiv.org/abs/2505.11200
作者: Xihuai Wang,  Ziyi Zhao,  Siyu Ren,  Shao Zhang,  Song Li,  Xiaoyu Li,  Ziwen Wang,  Lin Qiu,  Guanglu Wan,  Xuezhi Cao,  Xunliang Cai,  Weinan Zhang 
备注:Under Review
摘要:大型语言模型(LLM)的最新进展显着改善了文本到语音(TTS)系统,增强了对语音风格,自然度和情感表达的控制,使TTS系统更接近人类水平的性能。虽然平均意见评分(MOS)仍然是TTS系统评估的标准,但它存在主观性,环境不一致性和有限的可解释性。现有的评测数据集也缺乏多维设计,往往忽略了说话风格、语境多样性和陷阱话语等因素,这在中文TTS评测中表现得尤为明显。为了解决这些挑战,我们引入了音频图灵测试(ATT),一个多维的中文语料库数据集ATT-Corpus配对一个简单的,图灵测试启发的评估协议。ATT不依赖复杂的MOS量表或直接的模型比较,而是要求评估者判断声音是否听起来像人类。这种简化减少了评级偏差,提高了评估的稳健性。为了进一步支持快速模型开发,我们还使用人类判断数据微调Qwen 2-Audio-Instruct作为Auto-ATT进行自动评估。实验结果表明,ATT有效地区分模型在特定的能力维度,使用其多维设计。Auto-ATT还展示了与人类评估的高度一致性,证实了其作为快速可靠的评估工具的价值。白盒ATT-Corpus和Auto-ATT可以在ATT拥抱脸合集(https://huggingface.co/collections/meituan/audio-turing-test-682446320368164faeaf38a4)中找到。
摘要:Recent advances in large language models (LLMs) have significantly improved text-to-speech (TTS) systems, enhancing control over speech style, naturalness, and emotional expression, which brings TTS Systems closer to human-level performance. Although the Mean Opinion Score (MOS) remains the standard for TTS System evaluation, it suffers from subjectivity, environmental inconsistencies, and limited interpretability. Existing evaluation datasets also lack a multi-dimensional design, often neglecting factors such as speaking styles, context diversity, and trap utterances, which is particularly evident in Chinese TTS evaluation. To address these challenges, we introduce the Audio Turing Test (ATT), a multi-dimensional Chinese corpus dataset ATT-Corpus paired with a simple, Turing-Test-inspired evaluation protocol. Instead of relying on complex MOS scales or direct model comparisons, ATT asks evaluators to judge whether a voice sounds human. This simplification reduces rating bias and improves evaluation robustness. To further support rapid model development, we also finetune Qwen2-Audio-Instruct with human judgment data as Auto-ATT for automatic evaluation. Experimental results show that ATT effectively differentiates models across specific capability dimensions using its multi-dimensional design. Auto-ATT also demonstrates strong alignment with human evaluations, confirming its value as a fast and reliable assessment tool. The white-box ATT-Corpus and Auto-ATT can be found in ATT Hugging Face Collection (https://huggingface.co/collections/meituan/audio-turing-test-682446320368164faeaf38a4).


【6】 $\mathcal{A}LLM4ADD$: Unlocking the Capabilities of Audio Large Language  Models for Audio Deepfake Detection

链接:https://arxiv.org/abs/2505.11079
作者: Hao Gu,  Jiangyan Yi,  Chenglong Wang,  Jianhua Tao,  Zheng Lian,  Jiayi He,  Yong Ren,  Yujie Chen,  Zhengqi Wen 
摘要:由于高保真音频生成模型的兴起及其滥用的可能性,音频深度伪造检测(ADD)变得越来越重要。鉴于音频大语言模型(ALLM)在各种音频处理任务中取得了重大进展,一个启发性的问题出现了:ALLM可以用来解决ADD吗?在本文中,我们首先进行全面的zero-shot评估的ALLM上添加,揭示他们在检测虚假音频的无效性。为了提高它们的性能,我们提出了一个全LM驱动的ADD框架$\mathcal{A} LLM 4ADD $。具体来说,我们将ADD任务重新定义为音频问题回答问题,提示模型问题:“这音频是假的还是真的?".然后,我们执行监督微调,使ALLM评估查询音频的真实性。进行了大量的实验来证明我们基于ALLM的方法可以在虚假音频检测中实现卓越的性能,特别是在数据稀缺的场景中。作为一项开创性的研究,我们预计这项工作将激励研究界利用ALLM开发更有效的ADD系统。
摘要:Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in various audio processing tasks, a heuristic question arises: Can ALLMs be leveraged to solve ADD?. In this paper, we first conduct a comprehensive zero-shot evaluation of ALLMs on ADD, revealing their ineffectiveness in detecting fake audio. To enhance their performance, we propose $\mathcal{A}LLM4ADD$, an ALLM-driven framework for ADD. Specifically, we reformulate ADD task as an audio question answering problem, prompting the model with the question: "Is this audio fake or real?". We then perform supervised fine-tuning to enable the ALLM to assess the authenticity of query audio. Extensive experiments are conducted to demonstrate that our ALLM-based method can achieve superior performance in fake audio detection, particularly in data-scarce scenarios. As a pioneering study, we anticipate that this work will inspire the research community to leverage ALLMs to develop more effective ADD systems.


【7】 CAMEO: Collection of Multilingual Emotional Speech Corpora

标题: CAMEO:多语言情感言语库集合
链接:https://arxiv.org/abs/2505.11051
作者: Iwona Christop,  Maciej Czajka 
备注:Under review at NeurIPS
摘要:本文介绍了CAMEO -一个精心策划的多语言情感语音数据集集合,旨在促进情感识别和其他语音相关任务的研究。主要目标是确保数据的容易访问,允许结果的可重复性,并提供一个标准化的基准,用于评估不同情绪状态和语言的语音情感识别(SER)系统。本文描述了数据集选择标准,策展和规范化过程,并提供了几个模型的性能结果。该集合,以及元数据和排行榜,通过拥抱脸平台公开提供。
摘要:This paper presents CAMEO -- a curated collection of multilingual emotional speech datasets designed to facilitate research in emotion recognition and other speech-related tasks. The main objectives were to ensure easy access to the data, to allow reproducibility of the results, and to provide a standardized benchmark for evaluating speech emotion recognition (SER) systems across different emotional states and languages. The paper describes the dataset selection criteria, the curation and normalization process, and provides performance results for several models. The collection, along with metadata, and a leaderboard, is publicly available via the Hugging Face platform.


【8】 Classifying Shelf Life Quality of Pineapples by Combining Audio and  Visual Features

标题: 结合视听特征对菠萝保质期质量进行分类
链接:https://arxiv.org/abs/2505.11020
作者: Yi-Lu Jiang,  Wen-Chang Chang,  Ching-Lin Wang,  Kung-Liang Hsu,  Chih-Yi Chiu 
摘要:使用非破坏性方法确定菠萝的保质期质量是减少浪费和增加收入的关键一步。本文建立了一个多通道多视角的菠萝分类模型,根据菠萝的听觉和视觉特征将其分为四个质量等级。出于研究目的,我们编译并发布了由500个菠萝组成的PQC 500数据集,其中包括两种模式:一种是通过多个麦克风敲击菠萝来记录声音,另一种是通过多个摄像头在不同位置拍摄照片,提供多模态和多视图视听功能。我们修改了对比视听掩蔽自编码器,通过丰富的音频和视觉对的组合来训练基于交叉模态的分类模型。此外,我们建议对训练数据进行紧凑的采样,以提高计算效率。在各种数据和模型配置下对实验进行了评估,结果表明,使用音频主要采样训练的跨模态模型可以产生84%的准确率,分别优于仅音频和仅视觉的单峰模型6%和18%。
摘要:Determining the shelf life quality of pineapples using non-destructive methods is a crucial step to reduce waste and increase income. In this paper, a multimodal and multiview classification model was constructed to classify pineapples into four quality levels based on audio and visual characteristics. For research purposes, we compiled and released the PQC500 dataset consisting of 500 pineapples with two modalities: one was tapping pineapples to record sounds by multiple microphones and the other was taking pictures by multiple cameras at different locations, providing multimodal and multi-view audiovisual features. We modified the contrastive audiovisual masked autoencoder to train the cross-modal-based classification model by abundant combinations of audio and visual pairs. In addition, we proposed to sample a compact size of training data for efficient computation. The experiments were evaluated under various data and model configurations, and the results demonstrated that the proposed cross-modal model trained using audio-major sampling can yield 84% accuracy, outperforming the unimodal models of only audio and only visual by 6% and 18%, respectively.


【9】 Survey of End-to-End Multi-Speaker Automatic Speech Recognition for  Monaural Audio

标题: 单耳音频端到端多说话人自动语音识别综述
链接:https://arxiv.org/abs/2505.10975
作者: Xinlu He,  Jacob Whitehill 
备注:13 pages. Submitted to IEEE/ACM Transaction on Audio Speech and Language Processing (TASLP)
摘要:单声道多说话人自动语音识别(ASR)仍然具有挑战性,由于数据稀缺和识别和归因于单个说话人的单词的固有困难,特别是在重叠语音中。最近的进展已经推动了从级联系统到端到端(E2 E)架构的转变,这减少了错误传播,并更好地利用语音内容和扬声器身份之间的协同作用。尽管E2 E多扬声器ASR取得了快速进展,但该领域缺乏对最新发展的全面回顾。该调查提供了多说话者ASR的E2 E神经方法的系统分类,突出了最新进展和比较分析。具体而言,我们分析:(1)体系结构范式(SIMO与~ SISO)的预分段音频,分析其独特的特点和权衡;(2)最近的架构和算法的改进,基于这两个范例;(3)扩展到长格式的语音,包括分段策略和说话人一致的假设拼接。此外,我们(4)评估和比较标准基准的方法。最后,我们讨论了开放的挑战和未来的研究方向,以建立强大的和可扩展的多扬声器ASR。
摘要:Monaural multi-speaker automatic speech recognition (ASR) remains challenging due to data scarcity and the intrinsic difficulty of recognizing and attributing words to individual speakers, particularly in overlapping speech. Recent advances have driven the shift from cascade systems to end-to-end (E2E) architectures, which reduce error propagation and better exploit the synergy between speech content and speaker identity. Despite rapid progress in E2E multi-speaker ASR, the field lacks a comprehensive review of recent developments. This survey provides a systematic taxonomy of E2E neural approaches for multi-speaker ASR, highlighting recent advances and comparative analysis. Specifically, we analyze: (1) architectural paradigms (SIMO vs.~SISO) for pre-segmented audio, analyzing their distinct characteristics and trade-offs; (2) recent architectural and algorithmic improvements based on these two paradigms; (3) extensions to long-form speech, including segmentation strategy and speaker-consistent hypothesis stitching. Further, we (4) evaluate and compare methods across standard benchmarks. We conclude with a discussion of open challenges and future research directions towards building robust and scalable multi-speaker ASR.


【10】 BanglaFake: Constructing and Evaluating a Specialized Bengali Deepfake  Audio Dataset

标题: BanglaFake:构建和评估专业的孟加拉Deepfake音频数据集
链接:https://arxiv.org/abs/2505.10885
作者: Istiaq Ahmed Fahad,  Kamruzzaman Asif,  Sifat Sikder 
备注:5 page
摘要:Deepfake音频检测对于像孟加拉语这样的低资源语言来说是具有挑战性的,因为数据集有限,声学特征也很微妙。为了解决这个问题,我们引入了BangalFake,这是一个孟加拉语Deepfake音频数据集,包含12,260个真实和13,260个deepfake话语。合成语音使用SOTA文本到语音(TTS)模型生成,确保高自然度和质量。我们通过定性和定量分析来评估数据集。来自30位母语人士的平均意见得分(MOS)显示Robust-MOS为3.40(自然度)和4.01(可理解性)。MFCC的t-SNE可视化突出了真实与虚假的差异化挑战。该数据集是推进孟加拉语deepfake检测的关键资源,解决了低资源语言研究的局限性。
摘要:Deepfake audio detection is challenging for low-resource languages like Bengali due to limited datasets and subtle acoustic features. To address this, we introduce BangalFake, a Bengali Deepfake Audio Dataset with 12,260 real and 13,260 deepfake utterances. Synthetic speech is generated using SOTA Text-to-Speech (TTS) models, ensuring high naturalness and quality. We evaluate the dataset through both qualitative and quantitative analyses. Mean Opinion Score (MOS) from 30 native speakers shows Robust-MOS of 3.40 (naturalness) and 4.01 (intelligibility). t-SNE visualization of MFCCs highlights real vs. fake differentiation challenges. This dataset serves as a crucial resource for advancing deepfake detection in Bengali, addressing the limitations of low-resource language research.


【11】 Multi-Stage Speaker Diarization for Noisy Classrooms

标题: 噪音教室的多阶段扬声器拨号
链接:https://arxiv.org/abs/2505.10879
作者: Ali Sartaz Khan,  Tolulope Ogunremi,  Ahmed Attia,  Dorottya Demszky 
摘要:发言人日记,识别过程中的“谁说话时”的录音,是理解课堂动态必不可少的。然而,教室环境带来了独特的挑战,包括录音质量差,背景噪音高,语音重叠,以及难以准确捕捉儿童的声音。本研究使用Nvidia的NeMo diarization pipeline研究了多阶段diarization模型的有效性。我们评估了去噪对日记准确性的影响,并比较了各种语音活动检测(VAD)模型,包括基于自监督变换的帧式VAD模型。我们还探索了一种混合VAD方法,该方法将自动语音识别(ASR)单词级时间戳与帧级VAD预测相结合。我们进行实验,使用两个数据集从英语口语教室分离教师与学生的讲话,并分离所有扬声器。我们的研究结果表明,去噪显着提高了拨号错误率(DER),降低了错过的语音率。此外,在去噪和噪声数据集上进行训练可以在噪声条件下获得可观的性能增益。混合VAD模型导致语音检测的进一步改进,在师生实验中达到低至17%的DER,在所有说话人实验中达到45%。然而,我们也确定了语音活动检测和扬声器混淆之间的权衡。总的来说,我们的研究突出了多阶段日记化模型的有效性,并整合了基于ASR的信息,以增强嘈杂的教室环境中的扬声器日记化。
摘要:Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording quality, high levels of background noise, overlapping speech, and the difficulty of accurately capturing children's voices. This study investigates the effectiveness of multi-stage diarization models using Nvidia's NeMo diarization pipeline. We assess the impact of denoising on diarization accuracy and compare various voice activity detection (VAD) models, including self-supervised transformer-based frame-wise VAD models. We also explore a hybrid VAD approach that integrates Automatic Speech Recognition (ASR) word-level timestamps with frame-level VAD predictions. We conduct experiments using two datasets from English speaking classrooms to separate teacher vs. student speech and to separate all speakers. Our results show that denoising significantly improves the Diarization Error Rate (DER) by reducing the rate of missed speech. Additionally, training on both denoised and noisy datasets leads to substantial performance gains in noisy conditions. The hybrid VAD model leads to further improvements in speech detection, achieving a DER as low as 17% in teacher-student experiments and 45% in all-speaker experiments. However, we also identified trade-offs between voice activity detection and speaker confusion. Overall, our study highlights the effectiveness of multi-stage diarization models and integrating ASR-based information for enhancing speaker diarization in noisy classroom environments.


【12】 NeoLightning: A Modern Reimagination of Gesture-Based Sound Design

标题: NeoLightning:基于手势的声音设计的现代再想象
链接:https://arxiv.org/abs/2505.10686
作者: Yonghyun Kim,  Sangheon Park,  Marcus Parker,  Donghoon Seu,  Alexandria Smith 
备注:Accepted to the 50th International Computer Music Conference (ICMC), 2025
摘要:本文介绍了新闪电,一个现代的重新解释的Buchla闪电。NeoLightning保留了Don Buchla的“Buchla Lightning”(于20世纪90年代推出)的创新精神,同时使其基于手势的交互对当代用户开放。虽然最初的Buchla Lightning和许多其他历史乐器在当时都是开创性的,但它们现在基本上不受支持,将用户交互限制为间接体验。为了解决这个问题,NeoLightning利用MediaPipe进行基于深度学习的手势识别,并采用Max/MSP和Processing进行实时多媒体处理。重新设计的系统提供精确、低延迟的手势识别和沉浸式3D交互。通过将原始Lightning的创造精神与现代进步相结合,NeoLightning重新定义了基于手势的音乐交互,扩展了表现力表现和交互式声音设计的可能性。
摘要:This paper introduces NeoLightning, a modern reinterpretation of the Buchla Lightning. NeoLightning preserves the innovative spirit of Don Buchla's "Buchla Lightning" (introduced in the 1990s) while making its gesture-based interaction accessible to contemporary users. While the original Buchla Lightning and many other historical instruments were groundbreaking in their time, they are now largely unsupported, limiting user interaction to indirect experiences. To address this, NeoLightning leverages MediaPipe for deep learning-based gesture recognition and employs Max/MSP and Processing for real-time multimedia processing. The redesigned system offers precise, low-latency gesture recognition and immersive 3D interaction. By merging the creative spirit of the original Lightning with modern advancements, NeoLightning redefines gesture-based musical interaction, expanding possibilities for expressive performance and interactive sound design.


【13】 LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models

标题: LipDistuser:使用条件扩散模型的唇转语音生成
链接:https://arxiv.org/abs/2505.11391
作者: Danilo de Oliveira,  Julius Richter,  Tal Peer,  Timo Germann 
摘要:我们提出了LipDiffuser,唇到语音生成合成自然和可理解的语音直接从无声的视频记录的条件扩散模型。我们的方法利用幅度保持消融扩散模型(MP-ADM)架构作为去噪模型。为了有效地调节模型,我们使用幅度保持特征线性调制(MP-FiLM)与扬声器嵌入结合视觉特征。然后,神经声码器从生成的梅尔频谱图重建语音波形。在LRS 3和TCD-TIMIT上的评估表明,LipDiffuser在感知语音质量和说话人相似性方面优于现有的唇到语音基线,同时在下游自动语音识别(ASR)中保持竞争力。这些发现也得到了正式听力实验的支持。广泛的消融研究和交叉数据集评估证实了我们的方法的有效性和泛化能力。
摘要:We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 and TCD-TIMIT demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition (ASR). These findings are also supported by a formal listening experiment. Extensive ablation studies and cross-dataset evaluation confirm the effectiveness and generalization capabilities of our approach.
eess.AS音频处理

【1】 LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models

标题: LipDistuser:使用条件扩散模型的唇转语音生成
链接:https://arxiv.org/abs/2505.11391
作者: Danilo de Oliveira,  Julius Richter,  Tal Peer,  Timo Germann 
摘要:我们提出了LipDiffuser,这是一个用于唇语生成的条件扩散模型,可以直接从无声视频记录中合成自然且可理解的语音。我们的方法利用幅度保持消融扩散模型(MP-ADM)架构作为去噪模型。为了有效地调节模型,我们使用幅度保持特征线性调制(MP-FiLM)与扬声器嵌入结合视觉特征。然后,神经声码器从生成的梅尔频谱图重建语音波形。在LRS 3和TCD-TIMIT上的评估表明,LipDiffuser在感知语音质量和说话人相似性方面优于现有的唇到语音基线,同时在下游自动语音识别(ASR)中保持竞争力。这些发现也得到了正式听力实验的支持。广泛的消融研究和交叉数据集评估证实了我们的方法的有效性和泛化能力。
摘要:We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach leverages the magnitude-preserving ablated diffusion model (MP-ADM) architecture as a denoiser model. To effectively condition the model, we incorporate visual features using magnitude-preserving feature-wise linear modulation (MP-FiLM) alongside speaker embeddings. A neural vocoder then reconstructs the speech waveform from the generated mel-spectrograms. Evaluations on LRS3 and TCD-TIMIT demonstrate that LipDiffuser outperforms existing lip-to-speech baselines in perceptual speech quality and speaker similarity, while remaining competitive in downstream automatic speech recognition (ASR). These findings are also supported by a formal listening experiment. Extensive ablation studies and cross-dataset evaluation confirm the effectiveness and generalization capabilities of our approach.


【2】 Anti-aliasing of neural distortion effects via model fine tuning

标题: 通过模型微调来抗锯齿神经失真效应
链接:https://arxiv.org/abs/2505.11375
作者: Alistair Carson,  Alec Wright,  Stefan Bilbao 
备注:Accepted for DAFx25
摘要:近年来,神经网络在吉他失真效果建模中变得无处不在。尽管它们能够产生感知上令人信服的模型,但当由高频和高增益输入驱动时,它们容易受到频率混叠的影响。当信号的带宽扩展到奈奎斯特频率之外时,非线性激活函数会产生所需的谐波失真和不需要的混叠失真。在这里,我们提出了一种通过教师-学生微调方法减少神经模型中混叠的方法,其中教师是一个预先训练的模型,其权重被冻结,而学生是一个具有可学习参数的副本。学生根据无混叠数据集进行微调,该数据集是通过将正弦曲线传递到原始模型并从输出光谱中去除非谐波分量而生成的。我们的研究结果表明,这种方法可以显著抑制长短期记忆网络(LSTM)和时间卷积网络(TCN)的混叠。在我们的大多数案例研究中,混叠的减少大于两倍过采样所实现的效果。所提出的方法的一个副作用是谐波失真分量也受到影响。发现这种不利影响与模型相关,LSTM模型在抗锯齿和保持与模拟参考设备的感知相似性之间提供了最佳平衡。
摘要:Neural networks have become ubiquitous with guitar distortion effects modelling in recent years. Despite their ability to yield perceptually convincing models, they are susceptible to frequency aliasing when driven by high frequency and high gain inputs. Nonlinear activation functions create both the desired harmonic distortion and unwanted aliasing distortion as the bandwidth of the signal is expanded beyond the Nyquist frequency. Here, we present a method for reducing aliasing in neural models via a teacher-student fine tuning approach, where the teacher is a pre-trained model with its weights frozen, and the student is a copy of this with learnable parameters. The student is fine-tuned against an aliasing-free dataset generated by passing sinusoids through the original model and removing non-harmonic components from the output spectra. Our results show that this method significantly suppresses aliasing for both long-short-term-memory networks (LSTM) and temporal convolutional networks (TCN). In the majority of our case studies, the reduction in aliasing was greater than that achieved by two times oversampling. One side-effect of the proposed method is that harmonic distortion components are also affected. This adverse effect was found to be model-dependent, with the LSTM models giving the best balance between anti-aliasing and preserving the perceived similarity to an analog reference device.


【3】 SongEval: A Benchmark Dataset for Song Aesthetics Evaluation

标题: SongEval:歌曲美学评估的基准数据集
链接:https://arxiv.org/abs/2505.10793
作者: Jixun Yao,  Guobin Ma,  Huixin Xue,  Huakang Chen,  Chunbo Hao,  Yuepeng Jiang,  Haohe Liu,  Ruibin Yuan,  Jin Xu,  Wei Xue,  Hao Liu,  Lei Xie 
摘要:美学在歌曲生成任务中是一个隐含的重要标准,它反映了人类超越客观度量的感知。然而,评估生成的歌曲的美学仍然是一个根本的挑战,因为音乐的欣赏是高度主观的。现有的评价指标,如嵌入式距离,是有限的,在反映主观和感知方面,定义音乐的吸引力。为了解决这个问题,我们引入了SongEval,这是第一个用于评估全长歌曲美学的开源大规模基准数据集。SongEval包括超过2,399首完整长度的歌曲,总计超过140小时,由16位具有音乐背景的专业注释者进行美学评级。每首歌曲都从五个关键方面进行评估:整体连贯性,记忆力,声音呼吸和措辞的自然性,歌曲结构的清晰度和整体音乐性。该数据集涵盖了英语和中文歌曲,涵盖了九种主流类型。此外,为了评估歌曲美学评价的有效性,我们进行了实验,使用SongEval预测美学分数,并表现出比现有的客观评价指标更好的性能,在预测人类感知的音乐质量。
摘要:Aesthetics serve as an implicit and important criterion in song generation tasks that reflect human perception beyond objective metrics. However, evaluating the aesthetics of generated songs remains a fundamental challenge, as the appreciation of music is highly subjective. Existing evaluation metrics, such as embedding-based distances, are limited in reflecting the subjective and perceptual aspects that define musical appeal. To address this issue, we introduce SongEval, the first open-source, large-scale benchmark dataset for evaluating the aesthetics of full-length songs. SongEval includes over 2,399 songs in full length, summing up to more than 140 hours, with aesthetic ratings from 16 professional annotators with musical backgrounds. Each song is evaluated across five key dimensions: overall coherence, memorability, naturalness of vocal breathing and phrasing, clarity of song structure, and overall musicality. The dataset covers both English and Chinese songs, spanning nine mainstream genres. Moreover, to assess the effectiveness of song aesthetic evaluation, we conduct experiments using SongEval to predict aesthetic scores and demonstrate better performance than existing objective evaluation metrics in predicting human-perceived musical quality.


【4】 Machine Learning Approaches to Vocal Register Classification in  Contemporary Male Pop Music

标题: 当代男性流行音乐中人声语域分类的机器学习方法
链接:https://arxiv.org/abs/2505.11378
作者: Alexander Kim,  Charlotte Botha 
备注:8 pages, 8 figures
摘要:对于所有经验水平的歌手来说,学习技术曲目最艰巨的挑战之一是在passagio(胸部声音和头部声音区域之间的通道)及其周围导航位置和声音区域。特别是在流行音乐中,其中单个艺术家可以使用各种音色和纹理来实现期望的质量,可能难以识别歌手正在使用的音域内的声音音域。本文提出了两种方法,通过分析梅尔频谱图图像的纹理特征来分类男性流行音乐音频信号中的声区。此外,我们将讨论这些模型的声乐分析工具的实际集成,并介绍一个同时开发的软件称为AVRA,代表自动声乐语域分析。我们提出的方法通过支持向量机(SVM)和卷积神经网络(CNN)模型实现了一致的分类,这支持了在更多的声音类型和歌唱流派中更强大的分类可能性的承诺。
摘要:For singers of all experience levels, one of the most daunting challenges in learning technical repertoire is navigating placement and vocal register in and around the passagio (passage between chest voice and head voice registers). Particularly in pop music, where a single artist may use a variety of timbre's and textures to achieve a desired quality, it can be difficult to identify what vocal register within the vocal range a singer is using. This paper presents two methods for classifying vocal registers in an audio signal of male pop music through the analysis of textural features of mel-spectrogram images. Additionally, we will discuss the practical integration of these models for vocal analysis tools, and introduce a concurrently developed software called AVRA which stands for Automatic Vocal Register Analysis. Our proposed methods achieved consistent classification of vocal register through both Support Vector Machine (SVM) and Convolutional Neural Network (CNN) models, which supports the promise of more robust classification possibilities across more voice types and genres of singing.


【5】 LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors

标题: LegoLAM:使用CTCPosteriors将LLM与语音编码器连接起来
链接:https://arxiv.org/abs/2505.11352
作者: Rao Ma,  Tongzhou Chen,  Kartik Audhkhasi,  Bhuvana Ramabhadran 
摘要:最近,大规模预训练语音编码器和大型语言模型(LLM)已经发布,它们在包括自动语音识别(ASR)在内的一系列口语处理任务上表现出最先进的性能。为了有效地结合这两种模型以获得更好的性能,采用了连续语音提示和ASR纠错。然而,这些方法倾向于次优性能或不灵活。在本文中,我们提出了一个新的范例,LegoSLM,桥梁语音编码器和LLM使用ASR后验矩阵。语音编码器被训练以在LLM词汇表上生成连接主义时间分类(CTC)后验,其用于通过计算LLM输入嵌入的加权和来重建伪音频嵌入。这些嵌入与LLM输入空间中的文本嵌入连接在一起。使用性能良好的USM和Gemma模型作为例子,我们证明了我们提出的LegoSLM方法在ASR和语音翻译任务上都具有良好的性能。通过将USM与Gemma模型连接起来,我们可以在8个MLS测试集上获得超过USM-CTC基线的平均49%的WERR。经过训练的模型还在一系列设置中表现出模块性-在微调Gemma模型权重之后,语音编码器可以切换并以zero-shot方式与LLM组合。此外,我们建议使用softmax温度来控制USM和LLM的解码时间影响,这在域自适应中显示出有效性。
摘要:Recently, large-scale pre-trained speech encoders and Large Language Models (LLMs) have been released, which show state-of-the-art performance on a range of spoken language processing tasks including Automatic Speech Recognition (ASR). To effectively combine both models for better performance, continuous speech prompts, and ASR error correction have been adopted. However, these methods are prone to suboptimal performance or are inflexible. In this paper, we propose a new paradigm, LegoSLM, that bridges speech encoders and LLMs using the ASR posterior matrices. The speech encoder is trained to generate Connectionist Temporal Classification (CTC) posteriors over the LLM vocabulary, which are used to reconstruct pseudo-audio embeddings by computing a weighted sum of the LLM input embeddings. These embeddings are concatenated with text embeddings in the LLM input space. Using the well-performing USM and Gemma models as an example, we demonstrate that our proposed LegoSLM method yields good performance on both ASR and speech translation tasks. By connecting USM with Gemma models, we can get an average of 49% WERR over the USM-CTC baseline on 8 MLS testsets. The trained model also exhibits modularity in a range of settings -- after fine-tuning the Gemma model weights, the speech encoder can be switched and combined with the LLM in a zero-shot fashion. Additionally, we propose to control the decode-time influence of the USM and LLM using a softmax temperature, which shows effectiveness in domain adaptation.


【6】 Improving Inference-Time Optimisation for Vocal Effects Style Transfer  with a Gaussian Prior

标题: 利用高斯先验改进声乐效果风格转移的推理时优化
链接:https://arxiv.org/abs/2505.11315
作者: Chin-Yun Yu,  Marco A. Martínez-Ramírez,  Junghyun Koo,  Wei-Hsiang Liao,  Yuki Mitsufuji,  György Fazekas 
备注:Submitted to WASPAA 2025
摘要:具有推理时间优化的风格转移(ST-ITO)是用于将参考音频的应用效果转移到原始音轨的最近方法。它优化效果参数,以最大限度地减少处理后的音频和参考之间的风格嵌入的距离。然而,这种方法平等地对待所有可能的配置,并且仅依赖于嵌入空间,这可能导致不切实际或有偏见的结果。我们通过在参数空间上引入从声乐预设数据集DiffVox导出的高斯先验来解决这个问题。由此产生的优化相当于最大后验估计。对MedleyDB数据集上的声音效果转移的评估显示,与基线相比,各项指标都有显著改善,包括盲音频效果估计器、最近邻方法和未校准的ST-ITO。建议的校准减少参数的均方误差高达33%,更好地匹配参考样式。16名参与者的主观评价证实了我们的方法的优越性,特别是在有限的数据制度。这项工作演示了如何在推理时间中结合先验知识来增强音频效果传输,为更有效和更逼真的音频处理系统铺平道路。
摘要:Style Transfer with Inference-Time Optimisation (ST-ITO) is a recent approach for transferring the applied effects of a reference audio to a raw audio track. It optimises the effect parameters to minimise the distance between the style embeddings of the processed audio and the reference. However, this method treats all possible configurations equally and relies solely on the embedding space, which can lead to unrealistic or biased results. We address this pitfall by introducing a Gaussian prior derived from a vocal preset dataset, DiffVox, over the parameter space. The resulting optimisation is equivalent to maximum-a-posteriori estimation. Evaluations on vocal effects transfer on the MedleyDB dataset show significant improvements across metrics compared to baselines, including a blind audio effects estimator, nearest-neighbour approaches, and uncalibrated ST-ITO. The proposed calibration reduces parameter mean squared error by up to 33% and matches the reference style better. Subjective evaluations with 16 participants confirm our method's superiority, especially in limited data regimes. This work demonstrates how incorporating prior knowledge in inference time enhances audio effects transfer, paving the way for more effective and realistic audio processing systems.


【7】 Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI  models in Sound Localization

标题: 看到声音,听到视觉:揭示声音定位中人工智能模型的情态偏差和冲突
链接:https://arxiv.org/abs/2505.11217
作者: Yanhao Jia,  Ji Xie,  S Jivaganesh,  Hao Li,  Xu Wu,  Mengmi Zhang 
备注:16 pages, 14 figures
摘要:None
摘要:Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere. Such sensory conflicts test perception, yet humans reliably resolve them by prioritizing sound over misleading visuals. Despite advances in multimodal AI integrating vision and audio, little is known about how these systems handle cross-modal conflicts or whether they favor one modality. In this study, we systematically examine modality bias and conflict resolution in AI sound localization. We assess leading multimodal models and benchmark them against human performance in psychophysics experiments across six audiovisual conditions, including congruent, conflicting, and absent cues. Humans consistently outperform AI, demonstrating superior resilience to conflicting or missing visuals by relying on auditory information. In contrast, AI models often default to visual input, degrading performance to near chance levels. To address this, we finetune a state-of-the-art model using a stereo audio-image dataset generated via 3D simulations. Even with limited training data, the refined model surpasses existing benchmarks. Notably, it also mirrors human-like horizontal localization bias favoring left-right precision-likely due to the stereo audio structure reflecting human ear placement. These findings underscore how sensory input quality and system architecture shape multimodal representation accuracy.


【8】 Audio Turing Test: Benchmarking the Human-likeness of Large Language  Model-based Text-to-Speech Systems in Chinese

标题: 音频图灵测试:基于大型语言模型的中文文本到语音系统的人形基准
链接:https://arxiv.org/abs/2505.11200
作者: Xihuai Wang,  Ziyi Zhao,  Siyu Ren,  Shao Zhang,  Song Li,  Xiaoyu Li,  Ziwen Wang,  Lin Qiu,  Guanglu Wan,  Xuezhi Cao,  Xunliang Cai,  Weinan Zhang 
备注:Under Review
摘要:大型语言模型(LLM)的最新进展显着改善了文本到语音(TTS)系统,增强了对语音风格,自然度和情感表达的控制,使TTS系统更接近人类水平的性能。虽然平均意见评分(MOS)仍然是TTS系统评估的标准,但它存在主观性,环境不一致性和有限的可解释性。现有的评测数据集也缺乏多维设计,往往忽略了说话风格、语境多样性和陷阱话语等因素,这在中文TTS评测中表现得尤为明显。为了解决这些挑战,我们引入了音频图灵测试(ATT),一个多维的中文语料库数据集ATT-Corpus配对一个简单的,图灵测试启发的评估协议。ATT不依赖复杂的MOS量表或直接的模型比较,而是要求评估者判断声音是否听起来像人。这种简化减少了评级偏差,提高了评估的稳健性。为了进一步支持快速模型开发,我们还使用人类判断数据微调Qwen 2-Audio-Instruct作为Auto-ATT进行自动评估。实验结果表明,ATT有效地区分模型在特定的能力维度,使用其多维设计。Auto-ATT还展示了与人类评估的高度一致性,证实了其作为快速可靠的评估工具的价值。白盒ATT-Corpus和Auto-ATT可以在ATT拥抱脸合集(https://huggingface.co/collections/meituan/audio-turing-test-682446320368164faeaf38a4)中找到。
摘要:Recent advances in large language models (LLMs) have significantly improved text-to-speech (TTS) systems, enhancing control over speech style, naturalness, and emotional expression, which brings TTS Systems closer to human-level performance. Although the Mean Opinion Score (MOS) remains the standard for TTS System evaluation, it suffers from subjectivity, environmental inconsistencies, and limited interpretability. Existing evaluation datasets also lack a multi-dimensional design, often neglecting factors such as speaking styles, context diversity, and trap utterances, which is particularly evident in Chinese TTS evaluation. To address these challenges, we introduce the Audio Turing Test (ATT), a multi-dimensional Chinese corpus dataset ATT-Corpus paired with a simple, Turing-Test-inspired evaluation protocol. Instead of relying on complex MOS scales or direct model comparisons, ATT asks evaluators to judge whether a voice sounds human. This simplification reduces rating bias and improves evaluation robustness. To further support rapid model development, we also finetune Qwen2-Audio-Instruct with human judgment data as Auto-ATT for automatic evaluation. Experimental results show that ATT effectively differentiates models across specific capability dimensions using its multi-dimensional design. Auto-ATT also demonstrates strong alignment with human evaluations, confirming its value as a fast and reliable assessment tool. The white-box ATT-Corpus and Auto-ATT can be found in ATT Hugging Face Collection (https://huggingface.co/collections/meituan/audio-turing-test-682446320368164faeaf38a4).


【9】 $\mathcal{A}LLM4ADD$: Unlocking the Capabilities of Audio Large Language  Models for Audio Deepfake Detection

链接:https://arxiv.org/abs/2505.11079
作者: Hao Gu,  Jiangyan Yi,  Chenglong Wang,  Jianhua Tao,  Zheng Lian,  Jiayi He,  Yong Ren,  Yujie Chen,  Zhengqi Wen 
摘要:由于高保真音频生成模型的兴起及其滥用的可能性,音频深度伪造检测(ADD)变得越来越重要。鉴于音频大语言模型(ALLM)在各种音频处理任务中取得了重大进展,一个启发性的问题出现了:ALLM可以用来解决ADD吗?在本文中,我们首先进行全面的zero-shot评估的ALLM上添加,揭示他们在检测虚假音频的无效性。为了提高它们的性能,我们提出了一个全LM驱动的ADD框架$\mathcal{A} LLM 4ADD $。具体来说,我们将ADD任务重新定义为音频问题回答问题,提示模型问题:“这音频是假的还是真的?".然后,我们执行监督微调,使ALLM评估查询音频的真实性。进行了大量的实验来证明我们基于ALLM的方法可以在虚假音频检测中实现卓越的性能,特别是在数据稀缺的场景中。作为一项开创性的研究,我们预计这项工作将激励研究界利用ALLM开发更有效的ADD系统。
摘要:Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in various audio processing tasks, a heuristic question arises: Can ALLMs be leveraged to solve ADD?. In this paper, we first conduct a comprehensive zero-shot evaluation of ALLMs on ADD, revealing their ineffectiveness in detecting fake audio. To enhance their performance, we propose $\mathcal{A}LLM4ADD$, an ALLM-driven framework for ADD. Specifically, we reformulate ADD task as an audio question answering problem, prompting the model with the question: "Is this audio fake or real?". We then perform supervised fine-tuning to enable the ALLM to assess the authenticity of query audio. Extensive experiments are conducted to demonstrate that our ALLM-based method can achieve superior performance in fake audio detection, particularly in data-scarce scenarios. As a pioneering study, we anticipate that this work will inspire the research community to leverage ALLMs to develop more effective ADD systems.


【10】 CAMEO: Collection of Multilingual Emotional Speech Corpora

标题: CAMEO:多语言情感言语库集合
链接:https://arxiv.org/abs/2505.11051
作者: Iwona Christop,  Maciej Czajka 
备注:Under review at NeurIPS
摘要:None
摘要:This paper presents CAMEO -- a curated collection of multilingual emotional speech datasets designed to facilitate research in emotion recognition and other speech-related tasks. The main objectives were to ensure easy access to the data, to allow reproducibility of the results, and to provide a standardized benchmark for evaluating speech emotion recognition (SER) systems across different emotional states and languages. The paper describes the dataset selection criteria, the curation and normalization process, and provides performance results for several models. The collection, along with metadata, and a leaderboard, is publicly available via the Hugging Face platform.


【11】 Classifying Shelf Life Quality of Pineapples by Combining Audio and  Visual Features

标题: 结合视听特征对菠萝保质期质量进行分类
链接:https://arxiv.org/abs/2505.11020
作者: Yi-Lu Jiang,  Wen-Chang Chang,  Ching-Lin Wang,  Kung-Liang Hsu,  Chih-Yi Chiu 
摘要:使用非破坏性方法确定菠萝的保质期质量是减少浪费和增加收入的关键一步。本文建立了一个多通道多视角的菠萝分类模型,根据菠萝的听觉和视觉特征将其分为四个质量等级。出于研究目的,我们编译并发布了由500个菠萝组成的PQC 500数据集,其中包括两种模式:一种是通过多个麦克风敲击菠萝来记录声音,另一种是通过多个摄像头在不同位置拍摄照片,提供多模态和多视图视听功能。我们修改了对比视听掩蔽自编码器,通过丰富的音频和视觉对的组合来训练基于交叉模态的分类模型。此外,我们建议对训练数据进行紧凑的采样,以提高计算效率。在各种数据和模型配置下对实验进行了评估,结果表明,使用音频主要采样训练的跨模态模型可以产生84%的准确率,分别优于仅音频和仅视觉的单峰模型6%和18%。
摘要:Determining the shelf life quality of pineapples using non-destructive methods is a crucial step to reduce waste and increase income. In this paper, a multimodal and multiview classification model was constructed to classify pineapples into four quality levels based on audio and visual characteristics. For research purposes, we compiled and released the PQC500 dataset consisting of 500 pineapples with two modalities: one was tapping pineapples to record sounds by multiple microphones and the other was taking pictures by multiple cameras at different locations, providing multimodal and multi-view audiovisual features. We modified the contrastive audiovisual masked autoencoder to train the cross-modal-based classification model by abundant combinations of audio and visual pairs. In addition, we proposed to sample a compact size of training data for efficient computation. The experiments were evaluated under various data and model configurations, and the results demonstrated that the proposed cross-modal model trained using audio-major sampling can yield 84% accuracy, outperforming the unimodal models of only audio and only visual by 6% and 18%, respectively.


【12】 Survey of End-to-End Multi-Speaker Automatic Speech Recognition for  Monaural Audio

标题: 单耳音频端到端多说话人自动语音识别综述
链接:https://arxiv.org/abs/2505.10975
作者: Xinlu He,  Jacob Whitehill 
备注:13 pages. Submitted to IEEE/ACM Transaction on Audio Speech and Language Processing (TASLP)
摘要:单声道多说话人自动语音识别(ASR)仍然具有挑战性,由于数据稀缺和识别和归因于单个说话人的单词的固有困难,特别是在重叠语音中。最近的进展已经推动了从级联系统到端到端(E2E)架构的转变,这减少了错误传播,并更好地利用语音内容和扬声器身份之间的协同作用。尽管E2E多扬声器ASR取得了快速进展,但该领域缺乏对最新发展的全面回顾。该调查提供了多说话者ASR的E2E神经方法的系统分类,突出了最新进展和比较分析。具体而言,我们分析:(1)体系结构范式(SIMO与~ SISO)的预分段音频,分析其独特的特点和权衡;(2)最近的架构和算法的改进,基于这两个范例;(3)扩展到长格式的语音,包括分段策略和说话人一致的假设拼接。此外,我们(4)评估和比较标准基准的方法。最后,我们讨论了开放的挑战和未来的研究方向,以建立强大的和可扩展的多扬声器ASR。
摘要:Monaural multi-speaker automatic speech recognition (ASR) remains challenging due to data scarcity and the intrinsic difficulty of recognizing and attributing words to individual speakers, particularly in overlapping speech. Recent advances have driven the shift from cascade systems to end-to-end (E2E) architectures, which reduce error propagation and better exploit the synergy between speech content and speaker identity. Despite rapid progress in E2E multi-speaker ASR, the field lacks a comprehensive review of recent developments. This survey provides a systematic taxonomy of E2E neural approaches for multi-speaker ASR, highlighting recent advances and comparative analysis. Specifically, we analyze: (1) architectural paradigms (SIMO vs.~SISO) for pre-segmented audio, analyzing their distinct characteristics and trade-offs; (2) recent architectural and algorithmic improvements based on these two paradigms; (3) extensions to long-form speech, including segmentation strategy and speaker-consistent hypothesis stitching. Further, we (4) evaluate and compare methods across standard benchmarks. We conclude with a discussion of open challenges and future research directions towards building robust and scalable multi-speaker ASR.


【13】 BanglaFake: Constructing and Evaluating a Specialized Bengali Deepfake  Audio Dataset

标题: BanglaFake:构建和评估专业的孟加拉Deepfake音频数据集
链接:https://arxiv.org/abs/2505.10885
作者: Istiaq Ahmed Fahad,  Kamruzzaman Asif,  Sifat Sikder 
备注:5 page
摘要:Deepfake音频检测对于像孟加拉语这样的低资源语言来说是具有挑战性的,因为数据集有限,声学特征也很微妙。为了解决这个问题,我们引入了BangalFake,这是一个孟加拉语Deepfake音频数据集,包含12,260个真实和13,260个deepfake话语。合成语音使用SOTA文本到语音(TTS)模型生成,确保高自然度和质量。我们通过定性和定量分析来评估数据集。来自30位母语人士的平均意见得分(MOS)显示Robust-MOS为3.40(自然度)和4.01(可理解性)。MFCC的t-SNE可视化突出了真实与虚假的差异化挑战。该数据集是推进孟加拉语deepfake检测的关键资源,解决了低资源语言研究的局限性。
摘要:Deepfake audio detection is challenging for low-resource languages like Bengali due to limited datasets and subtle acoustic features. To address this, we introduce BangalFake, a Bengali Deepfake Audio Dataset with 12,260 real and 13,260 deepfake utterances. Synthetic speech is generated using SOTA Text-to-Speech (TTS) models, ensuring high naturalness and quality. We evaluate the dataset through both qualitative and quantitative analyses. Mean Opinion Score (MOS) from 30 native speakers shows Robust-MOS of 3.40 (naturalness) and 4.01 (intelligibility). t-SNE visualization of MFCCs highlights real vs. fake differentiation challenges. This dataset serves as a crucial resource for advancing deepfake detection in Bengali, addressing the limitations of low-resource language research.


【14】 Multi-Stage Speaker Diarization for Noisy Classrooms

标题: 噪音教室的多阶段扬声器拨号
链接:https://arxiv.org/abs/2505.10879
作者: Ali Sartaz Khan,  Tolulope Ogunremi,  Ahmed Attia,  Dorottya Demszky 
摘要:发言人日记,识别过程中的“谁说话时”的录音,是理解课堂动态必不可少的。然而,教室环境带来了独特的挑战,包括录音质量差,背景噪音高,语音重叠,以及难以准确捕捉儿童的声音。本研究使用Nvidia的NeMo diarization pipeline研究了多阶段diarization模型的有效性。我们评估了去噪对日记准确性的影响,并比较了各种语音活动检测(VAD)模型,包括基于自监督变换的帧式VAD模型。我们还探索了一种混合VAD方法,该方法将自动语音识别(ASR)单词级时间戳与帧级VAD预测相结合。我们进行实验,使用两个数据集从英语口语教室分离教师与学生的讲话,并分离所有扬声器。我们的研究结果表明,去噪显着提高了拨号错误率(DER),降低了错过的语音率。此外,在去噪和噪声数据集上进行训练可以在噪声条件下获得可观的性能增益。混合VAD模型导致语音检测的进一步改进,在师生实验中达到低至17%的DER,在所有说话人实验中达到45%。然而,我们也确定了语音活动检测和扬声器混淆之间的权衡。总的来说,我们的研究突出了多阶段日记化模型的有效性,并整合了基于ASR的信息,以增强嘈杂的教室环境中的扬声器日记化。
摘要:Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording quality, high levels of background noise, overlapping speech, and the difficulty of accurately capturing children's voices. This study investigates the effectiveness of multi-stage diarization models using Nvidia's NeMo diarization pipeline. We assess the impact of denoising on diarization accuracy and compare various voice activity detection (VAD) models, including self-supervised transformer-based frame-wise VAD models. We also explore a hybrid VAD approach that integrates Automatic Speech Recognition (ASR) word-level timestamps with frame-level VAD predictions. We conduct experiments using two datasets from English speaking classrooms to separate teacher vs. student speech and to separate all speakers. Our results show that denoising significantly improves the Diarization Error Rate (DER) by reducing the rate of missed speech. Additionally, training on both denoised and noisy datasets leads to substantial performance gains in noisy conditions. The hybrid VAD model leads to further improvements in speech detection, achieving a DER as low as 17% in teacher-student experiments and 45% in all-speaker experiments. However, we also identified trade-offs between voice activity detection and speaker confusion. Overall, our study highlights the effectiveness of multi-stage diarization models and integrating ASR-based information for enhancing speaker diarization in noisy classroom environments.


【15】 NeoLightning: A Modern Reimagination of Gesture-Based Sound Design

标题: NeoLightning:基于手势的声音设计的现代再想象
链接:https://arxiv.org/abs/2505.10686
作者: Yonghyun Kim,  Sangheon Park,  Marcus Parker,  Donghoon Seu,  Alexandria Smith 
备注:Accepted to the 50th International Computer Music Conference (ICMC), 2025
摘要:本文介绍了新闪电,一个现代的重新解释的Buchla闪电。NeoLightning保留了Don Buchla的“Buchla Lightning”(于20世纪90年代推出)的创新精神,同时使其基于手势的交互对当代用户开放。虽然最初的Buchla Lightning和许多其他历史乐器在当时都是开创性的,但它们现在基本上不受支持,将用户交互限制为间接体验。为了解决这个问题,NeoLightning利用MediaPipe进行基于深度学习的手势识别,并采用Max/MSP和Processing进行实时多媒体处理。重新设计的系统提供精确、低延迟的手势识别和沉浸式3D交互。通过将原始Lightning的创造精神与现代进步相结合,NeoLightning重新定义了基于手势的音乐交互,扩展了表现力表现和交互式声音设计的可能性。
摘要:This paper introduces NeoLightning, a modern reinterpretation of the Buchla Lightning. NeoLightning preserves the innovative spirit of Don Buchla's "Buchla Lightning" (introduced in the 1990s) while making its gesture-based interaction accessible to contemporary users. While the original Buchla Lightning and many other historical instruments were groundbreaking in their time, they are now largely unsupported, limiting user interaction to indirect experiences. To address this, NeoLightning leverages MediaPipe for deep learning-based gesture recognition and employs Max/MSP and Processing for real-time multimedia processing. The redesigned system offers precise, low-latency gesture recognition and immersive 3D interaction. By merging the creative spirit of the original Lightning with modern advancements, NeoLightning redefines gesture-based musical interaction, expanding possibilities for expressive performance and interactive sound design.


机器翻译由腾讯交互翻译提供,仅供参考