今日论文合集:cs.SD语音16篇,eess.AS音频处理15篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Perch 2.0: The Bittern Lesson for Bioacoustics
标题:鲈鱼2.0:生物声学的苦涩课程
链接:https://arxiv.org/abs/2508.04665

作者:Bart van Merriënboer, Vincent Dumoulin, Jenny Hamer, Lauren Harrell, Andrea Burns, Tom Denton
摘要:Perch是一个高性能的生物声学预训练模型。它以监督的方式进行训练,为数千种发声物种提供现成的分类分数,并为迁移学习提供强大的嵌入。在这个新版本Perch 2.0中,我们从专门针对鸟类物种的训练扩展到大型多分类群数据集。该模型使用原型学习分类器以及新的源预测训练标准进行自蒸馏训练。Perch 2.0在BirdSet和BEANS基准测试中获得了最先进的性能。尽管几乎没有海洋训练数据,但它在海洋迁移学习任务上的表现也优于专门的海洋模型。我们提出的假设,为什么细粒度的物种分类是一个特别强大的预训练任务的生物声学。
摘要:Perch is a performant pre-trained model for bioacoustics. It was trained in supervised fashion, providing both off-the-shelf classification scores for thousands of vocalizing species as well as strong embeddings for transfer learning. In this new release, Perch 2.0, we expand from training exclusively on avian species to a large multi-taxa dataset. The model is trained with self-distillation using a prototype-learning classifier as well as a new source-prediction training criterion. Perch 2.0 obtains state-of-the-art performance on the BirdSet and BEANS benchmarks. It also outperforms specialized marine models on marine transfer learning tasks, despite having almost no marine training data. We present hypotheses as to why fine-grained species classification is a particularly robust pre-training task for bioacoustics.


【2】Live Music Models
标题:现场音乐模特
链接:https://arxiv.org/abs/2508.04651

作者:Lyria Team
摘要:我们介绍了一类新的生成模型的音乐称为现场音乐模型,产生一个连续的音乐流在实时同步用户控制。我们发布了Magenta RealTime,这是一个开放权重的现场音乐模型,可以使用文本或音频提示来控制声学风格。在音乐质量的自动度量方面,Magenta RealTime优于其他开放权重音乐生成模型,尽管使用的参数更少,并提供了首个实时生成功能。我们还发布了Lyria RealTime,这是一个基于API的模型,具有扩展控件,可以访问我们最强大的模型,并具有广泛的即时覆盖范围。这些模型展示了人工智能辅助音乐创作的新范式,强调现场音乐表演的人在回路中的互动。
摘要:We introduce a new class of generative models for music called live music models that produce a continuous stream of music in real-time with synchronized user control. We release Magenta RealTime, an open-weights live music model that can be steered using text or audio prompts to control acoustic style. On automatic metrics of music quality, Magenta RealTime outperforms other open-weights music generation models, despite using fewer parameters and offering first-of-its-kind live generation capabilities. We also release Lyria RealTime, an API-based model with extended controls, offering access to our most powerful model with wide prompt coverage. These models demonstrate a new paradigm for AI-assisted music creation that emphasizes human-in-the-loop interaction for live music performance.


【3】ESDD 2026: Environmental Sound Deepfake Detection Challenge Evaluation Plan
标题:ESDD 2026:环境声音Deepfake检测挑战评估计划
链接:https://arxiv.org/abs/2508.04529

作者:Han Yin, Yang Xiao, Rohan Kumar Das, Jisheng Bai, Ting Dang
摘要:音频生成系统的最新进展使得能够创建高度逼真和沉浸式的音景,这些音景越来越多地用于电影和虚拟现实。然而,这些音频生成器也引起了人们对潜在滥用的担忧,例如为假视频生成欺骗性音频内容和传播误导性信息。用于环境声音深度伪造检测(ESDD)的现有数据集在规模和音频类型方面受到限制。为了解决这一差距,我们提出了EnvSDD,这是第一个为ESDD设计的大规模策划数据集,包括45.25小时的真实声音和316.7小时的假声音。基于EnvSDD,我们正在发起环境声音Deepfake检测挑战赛。具体来说,我们提出了两个不同的轨道:ESDD在看不见的发电机和黑盒低资源ESDD,涵盖了在现实生活中遇到的各种挑战。该挑战将与2026年IEEE声学,语音和信号处理国际会议(ICASSP 2026)一起举行。
摘要:Recent advances in audio generation systems have enabled the creation of highly realistic and immersive soundscapes, which are increasingly used in film and virtual reality. However, these audio generators also raise concerns about potential misuse, such as generating deceptive audio content for fake videos and spreading misleading information. Existing datasets for environmental sound deepfake detection (ESDD) are limited in scale and audio types. To address this gap, we have proposed EnvSDD, the first large-scale curated dataset designed for ESDD, consisting of 45.25 hours of real and 316.7 hours of fake sound. Based on EnvSDD, we are launching the Environmental Sound Deepfake Detection Challenge. Specifically, we present two different tracks: ESDD in Unseen Generators and Black-Box Low-Resource ESDD, covering various challenges encountered in real-life scenarios. The challenge will be held in conjunction with the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026).


【4】NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
标题:NVSpeech:一个集成且可扩展的基于副语言发声的类人语音建模管道
链接:https://arxiv.org/abs/2508.04195

作者:Huan Liao, Qinke Ni, Yuancheng Wang, Yiheng Lu, Haoyue Zhan, Pengyuan Xie, Qiang Zhang, Zhizheng Wu
摘要:副语言发声包括非语言的声音,如笑声和呼吸,以及词汇化的感叹词,如“嗯”和“哦”,是自然口语交流的组成部分。尽管它们在传达情感、意图和干扰线索方面很重要,但在传统的自动语音识别(ASR)和文本到语音(TTS)系统中,这些线索在很大程度上被忽视。我们提出了NVSpeech,一个集成的和可扩展的管道,桥梁的识别和合成的非语言发声,包括数据集的建设,ASR建模,和可控的TTS。(1)我们引入了一个手动注释的数据集,包含48,430个人类口语话语,其中包含18个单词级别的非语言学类别。(2)我们开发了语言感知的ASR模型,该模型将语言线索视为内联可解码标记(例如,“你真有趣(笑声)”),使联合词汇和非语言转录。然后,该模型被用来自动标注一个大型语料库,第一个大规模的中文数据集的174,179话语(573小时)与词级对齐和反语言线索。(3)我们微调zero-shot TTS模型上的人类和自动标记的数据,使显式的控制,语言发声,允许上下文感知插入在任意标记的位置,类人的语音合成。通过统一识别和生成非语言发音,NVSpeech提供了第一个开放的,大规模的,单词级注释的管道,用于普通话的表达性语音建模,以可扩展和可控的方式集成识别和合成。数据集和音频演示可在https://nvspeech170k.github.io/上获得。
摘要:Paralinguistic vocalizations-including non-verbal sounds like laughter and breathing, as well as lexicalized interjections such as "uhm" and "oh"-are integral to natural spoken communication. Despite their importance in conveying affect, intent, and interactional cues, such cues remain largely overlooked in conventional automatic speech recognition (ASR) and text-to-speech (TTS) systems. We present NVSpeech, an integrated and scalable pipeline that bridges the recognition and synthesis of paralinguistic vocalizations, encompassing dataset construction, ASR modeling, and controllable TTS. (1) We introduce a manually annotated dataset of 48,430 human-spoken utterances with 18 word-level paralinguistic categories. (2) We develop the paralinguistic-aware ASR model, which treats paralinguistic cues as inline decodable tokens (e.g., "You're so funny [Laughter]"), enabling joint lexical and non-verbal transcription. This model is then used to automatically annotate a large corpus, the first large-scale Chinese dataset of 174,179 utterances (573 hours) with word-level alignment and paralingustic cues. (3) We finetune zero-shot TTS models on both human- and auto-labeled data to enable explicit control over paralinguistic vocalizations, allowing context-aware insertion at arbitrary token positions for human-like speech synthesis. By unifying the recognition and generation of paralinguistic vocalizations, NVSpeech offers the first open, large-scale, word-level annotated pipeline for expressive speech modeling in Mandarin, integrating recognition and synthesis in a scalable and controllable manner. Dataset and audio demos are available at https://nvspeech170k.github.io/.


【5】The State Of TTS: A Case Study with Human Fooling Rates
标题:TTS的状况:人类愚弄率的案例研究
链接:https://arxiv.org/abs/2508.04179

作者:Praveen Srinivasa Varadhan, Sherry Thomas, Sai Teja M. S., Suvrat Bhooshan, Mitesh M. Khapra
备注:Accepted at InterSpeech 2025
摘要:虽然近年来的主观评估表明TTS的快速发展,目前的TTS系统真的可以通过人类的欺骗测试在图灵类评估?我们引入了人类愚弄率(HFR),这是一个直接衡量机器生成的语音被误认为人类的频率的指标。我们对开源和商业TTS模型的大规模评估揭示了重要的见解:(i)基于CMOS的人类平等的主张往往在欺骗测试中失败,(ii)TTS进展应以人类语音达到高HFR的数据集为基准,因为对单调或表达较少的参考样本进行评估设置了低标准,(iii)商业模型在zero-shot设置中接近人类欺骗,而开放源码系统仍在与自然的对话语音作斗争; ㈣对高质量数据进行微调可以提高真实性,但不能完全弥合差距。我们的研究结果强调,除了现有的主观测试外,还需要更现实、以人为本的评估。
摘要:While subjective evaluations in recent years indicate rapid progress in TTS, can current TTS systems truly pass a human deception test in a Turing-like evaluation? We introduce Human Fooling Rate (HFR), a metric that directly measures how often machine-generated speech is mistaken for human. Our large-scale evaluation of open-source and commercial TTS models reveals critical insights: (i) CMOS-based claims of human parity often fail under deception testing, (ii) TTS progress should be benchmarked on datasets where human speech achieves high HFRs, as evaluating against monotonous or less expressive reference samples sets a low bar, (iii) Commercial models approach human deception in zero-shot settings, while open-source systems still struggle with natural conversational speech; (iv) Fine-tuning on high-quality data improves realism but does not fully bridge the gap. Our findings underscore the need for more realistic, human-centric evaluations alongside existing subjective tests.


【6】Efficient Scaling for LLM-based ASR
标题:基于LLM的ASB的高效扩展
链接:https://arxiv.org/abs/2508.04096

作者:Bingshen Mu, Yiwen Shao, Kun Wei, Dong Yu, Lei Xie
备注:Accepted by ASRU 2025
摘要:基于大语言模型(LLM)的自动语音识别(ASR)具有很强的性能,但通常会产生很高的计算成本。本文研究了如何有效地获得最佳LLM-ASR性能。通过全面和受控的实验,我们发现,在将语音编码器与LLM集成之前对其进行预训练,比LLM-ASR的联合后训练的标准实践具有更好的缩放效率。基于这种见解,我们提出了一种新的多阶段LLM-ASR训练策略EFIN:编码器优先集成。在所有评估的训练策略中,EFIN始终提供更好的性能(相对于21.1%CERR),同时显著降低计算预算(49.9%FLOP)。此外,我们推导出一个缩放律,近似ASR的错误率作为计算函数,LLM-ASR缩放提供了实际的指导。
摘要:Large language model (LLM)-based automatic speech recognition (ASR) achieves strong performance but often incurs high computational costs. This work investigates how to obtain the best LLM-ASR performance efficiently. Through comprehensive and controlled experiments, we find that pretraining the speech encoder before integrating it with the LLM leads to significantly better scaling efficiency than the standard practice of joint post-training of LLM-ASR. Based on this insight, we propose a new multi-stage LLM-ASR training strategy, EFIN: Encoder First Integration. Among all training strategies evaluated, EFIN consistently delivers better performance (relative to 21.1% CERR) with significantly lower computation budgets (49.9% FLOPs). Furthermore, we derive a scaling law that approximates ASR error rates as a computation function, providing practical guidance for LLM-ASR scaling.


【7】MiDashengLM: Efficient Audio Understanding with General Audio Captions
标题:MiDashengLM:通过通用音频说明高效的音频理解
链接:https://arxiv.org/abs/2508.03983

作者:Heinrich Dinkel, Gang Li, Jizhong Liu, Jian Luan, Yadong Niu, Xingwei Sun, Tianzi Wang, Qiyang Xiao, Junbo Zhang, Jiahao Zhou
摘要:目前的大型音频语言模型(LALM)方法通常依赖于封闭的数据源或专有模型,限制了它们的通用性和可访问性。本文介绍了MiDashengLM,这是一种新型的开放式音频语言模型,旨在通过使用我们新的ACAVCaps训练数据集使用通用音频字幕进行高效和全面的音频理解。MiDashengLM完全依赖于公开的预训练和监督微调(SFT)数据集,确保完全透明和可重复性。在其核心,MiDashengLM集成了Dasheng,一个开源的音频编码器,专门设计用于有效地处理各种听觉信息。与以前的作品主要集中在基于自动语音识别(ASR)的音频文本对齐,我们的策略集中在一般的音频字幕,融合语音,声音和音乐信息到一个文本表示,使复杂的音频场景的整体文本表示。最后,MiDashengLM在首次令牌时间(TTFT)方面提供了高达4倍的加速,并且吞吐量比同类模型高出20倍。检查点可在https://huggingface.co/mispeech/midashenglm-7b和https://github.com/xiaomi-research/dasheng-lm上查阅。
摘要:Current approaches for large audio language models (LALMs) often rely on closed data sources or proprietary models, limiting their generalization and accessibility. This paper introduces MiDashengLM, a novel open audio-language model designed for efficient and comprehensive audio understanding through the use of general audio captions using our novel ACAVCaps training dataset. MiDashengLM exclusively relies on publicly available pretraining and supervised fine-tuning (SFT) datasets, ensuring full transparency and reproducibility. At its core, MiDashengLM integrates Dasheng, an open-source audio encoder, specifically engineered to process diverse auditory information effectively. Unlike previous works primarily focused on Automatic Speech Recognition (ASR) based audio-text alignment, our strategy centers on general audio captions, fusing speech, sound and music information into one textual representation, enabling a holistic textual representation of complex audio scenes. Lastly, MiDashengLM provides an up to 4x speedup in terms of time-to-first-token (TTFT) and up to 20x higher throughput than comparable models. Checkpoints are available online at https://huggingface.co/mispeech/midashenglm-7b and https://github.com/xiaomi-research/dasheng-lm.


【8】Are Inherently Interpretable Models More Robust? A Study In Music Emotion Recognition
标题:固有可解释模型是否更稳健?音乐情感识别研究
链接:https://arxiv.org/abs/2508.03780

作者: Katharina HOEDT, Arthur Flexer, Gerhard Widmer
备注:8 pages, published in Proceedings of the 22nd Sound and Music Computing Conference 2025 (SMC-25)
摘要:深度学习模型所需的关键属性之一是能够概括未知样本。当提供与一个或多个训练样本(感知上)相似的新样本时,深度学习模型预计会产生相应的相似输出。成功预测类似输入的类似输出的模型通常被称为鲁棒模型。另一方面,深度学习模型已经被证明非常容易受到输入的微小(对抗性)扰动的影响,这会极大地改变模型的输出,同时暴露出它对虚假相关性的依赖。在这项工作中,我们研究了内在可解释的深度模型,即,与黑箱模型相比,被设计为更关注有意义和可解释的特征的深度模型对数据中不相关的扰动更鲁棒。我们通过比较可解释的和黑盒音乐情感识别(MER)模型在对抗性示例挑战时的鲁棒性来测试我们的假设。此外,我们还包括一个经过对抗训练的模型,该模型经过优化,在比较中更加强大。我们的研究结果表明,本质上更具可解释性的模型确实可以比黑箱模型更健壮,并以更低的计算成本达到与对抗训练模型相似的健壮性水平。
摘要:One of the desired key properties of deep learning models is the ability to generalise to unseen samples. When provided with new samples that are (perceptually) similar to one or more training samples, deep learning models are expected to produce correspondingly similar outputs. Models that succeed in predicting similar outputs for similar inputs are often called robust. Deep learning models, on the other hand, have been shown to be highly vulnerable to minor (adversarial) perturbations of the input, which manage to drastically change a model's output and simultaneously expose its reliance on spurious correlations. In this work, we investigate whether inherently interpretable deep models, i.e., deep models that were designed to focus more on meaningful and interpretable features, are more robust to irrelevant perturbations in the data, compared to their black-box counterparts. We test our hypothesis by comparing the robustness of an interpretable and a black-box music emotion recognition (MER) model when challenged with adversarial examples. Furthermore, we include an adversarially trained model, which is optimised to be more robust, in the comparison. Our results indicate that inherently more interpretable models can indeed be more robust than their black-box counterparts, and achieve similar levels of robustness as adversarially trained models, at lower computational cost.


【9】CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
标题:CoughViT:用于咳嗽音频表示学习的自我监督视觉Transformer
链接:https://arxiv.org/abs/2508.03764

作者:Justin Luong, Hao Xue, Flora D. Salim
备注:Accepted to ISWC
摘要:医生在诊断过程中常规评估呼吸音,以了解患者气道的状况。近年来,基于人工智能的呼吸音诊断系统在呼吸疾病检测方面取得了成功。这些系统代表了早期和可获得的诊断方面的重大进步,这对于及时治疗至关重要。然而,标签和数据稀缺仍然是关键挑战,特别是对于COVID-19以外的情况,限制了诊断性能和可靠的评估。在本文中,我们提出了CoughViT,这是一种用于学习通用咳嗽声音表示的新型预训练框架,以增强数据有限的任务中的诊断性能。为了解决标签稀缺性,我们采用屏蔽数据建模训练的特征编码器在自监督学习的方式。我们评估我们的方法对其他预训练策略的三个诊断重要的咳嗽分类任务。实验结果表明,我们的表示匹配或超过当前最先进的监督音频表示,以提高下游任务的性能。
摘要:Physicians routinely assess respiratory sounds during the diagnostic process, providing insight into the condition of a patient's airways. In recent years, AI-based diagnostic systems operating on respiratory sounds, have demonstrated success in respiratory disease detection. These systems represent a crucial advancement in early and accessible diagnosis which is essential for timely treatment. However, label and data scarcity remain key challenges, especially for conditions beyond COVID-19, limiting diagnostic performance and reliable evaluation. In this paper, we propose CoughViT, a novel pre-training framework for learning general-purpose cough sound representations, to enhance diagnostic performance in tasks with limited data. To address label scarcity, we employ masked data modelling to train a feature encoder in a self-supervised learning manner. We evaluate our approach against other pre-training strategies on three diagnostically important cough classification tasks. Experimental results show that our representations match or exceed current state-of-the-art supervised audio representations in enhancing performance on downstream tasks.


【10】Melodic and Metrical Elements of Expressiveness in Hindustani Vocal Music
标题:印度斯坦声乐表现力的旋律和节拍元素
链接:https://arxiv.org/abs/2508.04430

作者:Yash Bhake, Ankit Anand, Preeti Rao
备注:To appear in the proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR), Daejeon Korea, 2025
摘要:本文试图从艺术家在表演流行作品时所表现出的灵活性来研究北印度卡雅尔音乐的美学。我们研究表达的时间和音高变化的给定的抒情内容内和跨性能,并提出计算表示,可以区分不同的表现,同一首歌的表达。我们提出了必要的音频处理和注释程序,并讨论了我们的意见和见解,从分析数据集的两首歌曲在两个ragas每个呈现由10位著名艺术家。
摘要:This paper presents an attempt to study the aesthetics of North Indian Khayal music with reference to the flexibility exercised by artists in performing popular compositions. We study expressive timing and pitch variations of the given lyrical content within and across performances and propose computational representations that can discriminate between different performances of the same song in terms of expression. We present the necessary audio processing and annotation procedures, and discuss our observations and insights from the analysis of a dataset of two songs in two ragas each rendered by ten prominent artists.


【11】Text adaptation for speaker verification with speaker-text factorized embeddings
标题:使用说话人文本因子分解嵌入进行说话人验证的文本自适应
链接:https://arxiv.org/abs/2508.04425

作者:Yexin Yang, Shuai Wang, Xun Gong, Yanmin Qian, Kai Yu
备注:ICASSP 2020
摘要:预先收集的训练数据或注册数据与实际测试数据之间的文本不匹配会严重影响文本相关说话人确认(SV)系统的性能。虽然这个问题可以通过仔细地收集具有目标语音内容的数据来解决,但是这样的数据收集可能是昂贵且不灵活的。在本文中,我们提出了一个新的文本自适应框架,以解决文本不匹配的问题。本文提出了一种说话人-文本分解网络,将输入语音分解为说话人嵌入和文本嵌入,然后在后期将它们集成为单个表示。给定少量的与说话者无关的自适应话语,可以提取目标语音内容的文本嵌入,并将其用于使与文本无关的说话者嵌入适应于文本定制的说话者嵌入。RSR 2015上的实验表明,文本自适应可以显着改善文本不匹配条件的性能。
摘要:Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully collecting data with the target speech content, such data collection could be costly and inflexible. In this paper, we propose a novel text adaptation framework to address the text mismatch issue. Here, a speaker-text factorization network is proposed to factorize the input speech into speaker embeddings and text embeddings and then integrate them into a single representation in the later stage. Given a small amount of speaker-independent adaptation utterances, text embeddings of target speech content can be extracted and used to adapt the text-independent speaker embeddings to text-customized speaker embeddings. Experiments on RSR2015 show that text adaptation can significantly improve the performance of text mismatch conditions.


【12】Binaural Sound Event Localization and Detection Neural Network based on HRTF Localization Cues for Humanoid Robots
标题:基于HRTF定位线索的仿人机器人双耳声音事件定位和检测神经网络
链接:https://arxiv.org/abs/2508.04333

作者:Gyeong-Tae Lee
备注:200 pages
摘要:类人机器人需要同时进行声音事件类型和方向估计以用于态势感知,但是传统的双通道输入与高度估计和前后混淆作斗争。本文提出了一种双耳声音事件定位和检测(BiSELD)神经网络来解决这些挑战。BiSELDnet从双耳输入特征中学习时频模式和头部相关传递函数(HRTF)定位线索。介绍了一种新的八通道双耳时频特征(BTFF),包括左/右梅尔频谱图、V图、耳间时间差(ITD)图(低于1.5 kHz)、耳间电平差(ILD)图(高于5 kHz,具有前后不对称性)和频谱线索(SC)图(高于5 kHz,用于仰角)。在全向、水平和正中平面上证实了BTFF的有效性。BiSELDnets,特别是一个有效的三位一体模块的基础上,实现输出的方向向量的时间序列为每个声音事件类,使同时检测和定位。提出了向量激活图(VAM)可视化来分析网络学习,证实了BiSELDnet对海拔估计的N1陷波频率的关注。城市背景噪声条件下的比较评估表明,建议的BiSELD模型显着优于国家的最先进的(SOTA)SELD模型与双耳输入。
摘要:Humanoid robots require simultaneous sound event type and direction estimation for situational awareness, but conventional two-channel input struggles with elevation estimation and front-back confusion. This paper proposes a binaural sound event localization and detection (BiSELD) neural network to address these challenges. BiSELDnet learns time-frequency patterns and head-related transfer function (HRTF) localization cues from binaural input features. A novel eight-channel binaural time-frequency feature (BTFF) is introduced, comprising left/right mel-spectrograms, V-maps, an interaural time difference (ITD) map (below 1.5 kHz), an interaural level difference (ILD) map (above 5 kHz with front-back asymmetry), and spectral cue (SC) maps (above 5 kHz for elevation). The effectiveness of BTFF was confirmed across omnidirectional, horizontal, and median planes. BiSELDnets, particularly one based on the efficient Trinity module, were implemented to output time series of direction vectors for each sound event class, enabling simultaneous detection and localization. Vector activation map (VAM) visualization was proposed to analyze network learning, confirming BiSELDnet's focus on the N1 notch frequency for elevation estimation. Comparative evaluations under urban background noise conditions demonstrated that the proposed BiSELD model significantly outperforms state-of-the-art (SOTA) SELD models with binaural input.


【13】A Multi-stage Low-latency Enhancement System for Hearing Aids
标题:一种多阶段低延迟助听器增强系统
链接:https://arxiv.org/abs/2508.04283

作者:Chengwei Ouyang, Kexin Fei, Haoshuai Zhou, Congxi Lu, Linkai Li
备注:2 pages, 1 figure, 1 table. accepted to ICASSP 2023
摘要:本文为ICASSP 2023 Clarity Challenge提出了一个端到端系统。在这项工作中,我们介绍了四个主要的创新:(1)一个新的多级系统在幅度和复杂的域,以更好地利用相位信息;(2)一个非对称的窗口对,以实现更高的频率分辨率与5 ms的延迟约束;(3)集成的头部旋转信息和混合信号,以实现更好的增强;(4)后处理模块,其利用由基线系统提供的助听器放大级实现更高的助听器言语感知指数(HASPI)分数。
摘要:This paper proposes an end-to-end system for the ICASSP 2023 Clarity Challenge. In this work, we introduce four major novelties: (1) a novel multi-stage system in both the magnitude and complex domains to better utilize phase information; (2) an asymmetric window pair to achieve higher frequency resolution with the 5ms latency constraint; (3) the integration of head rotation information and the mixture signals to achieve better enhancement; (4) a post-processing module that achieves higher hearing aid speech perception index (HASPI) scores with the hearing aid amplification stage provided by the baseline system.


【14】Towards interpretable emotion recognition: Identifying key features with machine learning
标题:走向可解释的情感识别:用机器学习识别关键特征
链接:https://arxiv.org/abs/2508.04230

作者:Yacouba Kaloga, Ina Kodrasi
备注:None
摘要:无监督方法,如wav2vec2和HuBERT,在音频任务中实现了最先进的性能,导致从可解释特征的研究转移。然而,这些方法缺乏可解释性,限制了它们在医学等关键领域的适用性,在这些领域,理解特征相关性至关重要。为了更好地理解无监督模型的特征,识别与给定任务相关的可解释特征仍然至关重要。在这项工作中,我们专注于情感识别,并使用机器学习算法来识别和概括这项任务中最重要的可解释特征。虽然以前的研究探讨了情绪识别中的特征相关性,但它们往往受到狭窄背景的限制,并且呈现不一致的结果。我们的方法旨在克服这些限制,提供一个更广泛,更强大的框架,以确定最重要的可解释的功能。
摘要:Unsupervised methods, such as wav2vec2 and HuBERT, have achieved state-of-the-art performance in audio tasks, leading to a shift away from research on interpretable features. However, the lack of interpretability in these methods limits their applicability in critical domains like medicine, where understanding feature relevance is crucial. To better understand the features of unsupervised models, it remains critical to identify the interpretable features relevant to a given task. In this work, we focus on emotion recognition and use machine learning algorithms to identify and generalize the most important interpretable features for this task. While previous studies have explored feature relevance in emotion recognition, they are often constrained by narrow contexts and present inconsistent findings. Our approach aims to overcome these limitations, providing a broader and more robust framework for identifying the most important interpretable features.


【15】Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
标题:Deepfakes的多语言源跟踪:第一个基准
链接:https://arxiv.org/abs/2508.04143

作者:Xi Xuan, Yang Xiao, Rohan Kumar Das, Tomi Kinnunen
备注:Accepted at Interspeech SPSC 2025 - 5th Symposium on Security and Privacy in Speech Communication (Oral)
摘要:最近在生成人工智能方面取得的进展使得从几秒钟的音频中创建听起来自然的Deepfake语音变得越来越容易。虽然这些工具支持有用的应用程序,但它们也引起了严重的关注,因为它们可以用多种语言生成令人信服的假语音。目前的研究主要集中在检测虚假语音,但很少关注跟踪用于生成它的源模型。本文介绍了多语言语音deepfake源跟踪的第一个基准,涵盖单语言和跨语言场景。我们比较研究基于DSP和SSL的建模;研究SSL表示在不同语言上的微调如何影响跨语言泛化性能;并评估对看不见的语言和说话者的泛化。我们的研究结果提供了第一个全面的见解识别语音生成模型时,训练和推理语言不同的挑战。数据集、方案和代码可在https://github.com/xuanxixi/Multilingual-Source-Tracing上获得。
摘要:Recent progress in generative AI has made it increasingly easy to create natural-sounding deepfake speech from just a few seconds of audio. While these tools support helpful applications, they also raise serious concerns by making it possible to generate convincing fake speech in many languages. Current research has largely focused on detecting fake speech, but little attention has been given to tracing the source models used to generate it. This paper introduces the first benchmark for multilingual speech deepfake source tracing, covering both mono- and cross-lingual scenarios. We comparatively investigate DSP- and SSL-based modeling; examine how SSL representations fine-tuned on different languages impact cross-lingual generalization performance; and evaluate generalization to unseen languages and speakers. Our findings offer the first comprehensive insights into the challenges of identifying speech generation models when training and inference languages differ. The dataset, protocol and code are available at https://github.com/xuanxixi/Multilingual-Source-Tracing.


【16】Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
标题:并行GPT:协调Zero-Shot文本到语音的声学和语义信息的独立性和相互依赖性
链接:https://arxiv.org/abs/2508.04141

作者:Jingyuan Xing, Zhipeng Li, Jialong Mai, Xiaofen Xing, Xiangmin Xu
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:语音表示和大型语言模型的进步增强了zero-shot文本到语音(TTS)的性能。然而,现有的zero-shot TTS模型在捕获声学和语义特征之间的复杂相关性方面面临挑战,导致缺乏表达性和相似性。本文提出了一种结合自回归(AR)和非自回归(NAR)模块的TTS框架,以协调语音信息和语义信息的独立性和相互依赖性。AR模型利用所提出的并行标记器来同时合成顶级语义和声学标记。相比之下,考虑到相互依赖性,耦合NAR模型根据一般AR模型的输出预测详细代币。基于该结构设计的并行GPT,通过其并行结构来提高zero-shot文语合成。在英语和汉语数据集上的实验表明,该模型的合成质量和效率明显优于现有的zero-shot TTS模型。语音演示可在https://t1235-ch.github.io/pgpt/上获得。
摘要:Advances in speech representation and large language models have enhanced zero-shot text-to-speech (TTS) performance. However, existing zero-shot TTS models face challenges in capturing the complex correlations between acoustic and semantic features, resulting in a lack of expressiveness and similarity. The primary reason lies in the complex relationship between semantic and acoustic features, which manifests independent and interdependent aspects.This paper introduces a TTS framework that combines both autoregressive (AR) and non-autoregressive (NAR) modules to harmonize the independence and interdependence of acoustic and semantic information. The AR model leverages the proposed Parallel Tokenizer to synthesize the top semantic and acoustic tokens simultaneously. In contrast, considering the interdependence, the Coupled NAR model predicts detailed tokens based on the general AR model's output. Parallel GPT, built on this architecture, is designed to improve zero-shot text-to-speech synthesis through its parallel structure. Experiments on English and Chinese datasets demonstrate that the proposed model significantly outperforms the quality and efficiency of the synthesis of existing zero-shot TTS models. Speech demos are available at https://t1235-ch.github.io/pgpt/.


eess.AS音频处理


【1】UniTalker: Conversational Speech-Visual Synthesis
标题:UniTalker:对话演讲-视觉合成
链接:https://arxiv.org/abs/2508.04585

作者:Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, Haizhou Li
备注:15 pages, 8 figures
摘要:会话语音合成是用户-智能体交互领域的一项关键任务,旨在为用户生成更具表达力和同情心的语音。然而,众所周知,“倾听”和“目光接触”在真实世界的人际交往中对传达情感起着至关重要的作用。现有的CSS研究仅限于感知对话背景下的文本和语音,这限制了其有效性。此外,仅语音响应进一步限制了交互体验。为了解决这些限制,我们引入了一个会话语音视觉合成(CSVS)任务作为传统CSS的扩展。通过利用多模态对话上下文,它为用户提供连贯的视听响应。为此,我们开发了一个CSVS系统名为UniTalker,这是一个统一的模型,无缝集成多模态感知和多模态渲染能力。具体而言,它利用大规模的语言模型来全面理解对话上下文中的多模态线索,包括说话人,文本,语音和说话的面部动画。在此基础上,利用多任务序列预测技术,首先推断出目标话语的情感,然后生成移情语音和自然的说话人脸动画。为了确保生成的语音-视觉内容在情感、内容和持续时间方面保持一致,我们引入了三个关键优化:1)设计一个专门的神经地标编解码器来标记和重建面部表情序列。2)提出一种双模态语音-视觉硬对齐解码策略。3)在生成阶段应用情感引导渲染。综合的客观和主观实验表明,我们的模型合成更多的移情语音,并为用户提供更自然和情感一致的说话脸动画。
摘要:Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that "listening" and "eye contact" play crucial roles in conveying emotions during real-world interpersonal communication. Existing CSS research is limited to perceiving only text and speech within the dialogue context, which restricts its effectiveness. Moreover, speech-only responses further constrain the interactive experience. To address these limitations, we introduce a Conversational Speech-Visual Synthesis (CSVS) task as an extension of traditional CSS. By leveraging multimodal dialogue context, it provides users with coherent audiovisual responses. To this end, we develop a CSVS system named UniTalker, which is a unified model that seamlessly integrates multimodal perception and multimodal rendering capabilities. Specifically, it leverages a large-scale language model to comprehensively understand multimodal cues in the dialogue context, including speaker, text, speech, and the talking-face animations. After that, it employs multi-task sequence prediction to first infer the target utterance's emotion and then generate empathetic speech and natural talking-face animations. To ensure that the generated speech-visual content remains consistent in terms of emotion, content, and duration, we introduce three key optimizations: 1) Designing a specialized neural landmark codec to tokenize and reconstruct facial expression sequences. 2) Proposing a bimodal speech-visual hard alignment decoding strategy. 3) Applying emotion-guided rendering during the generation stage. Comprehensive objective and subjective experiments demonstrate that our model synthesizes more empathetic speech and provides users with more natural and emotionally consistent talking-face animations.


【2】Pitfalls and Limits in Automatic Dementia Assessment
标题:自动痴呆症评估的陷阱和限制
链接:https://arxiv.org/abs/2508.04512

作者: Franziska Braun, Christopher Witzl, Andreas Erzigkeit, Hartmut Lehfeld, Thomas Hillemacher, Tobias Bocklet, Korbinian Riedhammer
备注:Accepted at INTERSPEECH 2025
摘要:目前基于语音的痴呆症评估的工作集中在特征提取来预测评估量表,或现有测试程序的自动化。大多数研究毫无疑问地使用公共数据,很少进行详细的错误分析,主要集中在数值性能上。我们对自动化标准化痴呆评估,即Syndrom-Kurz-Test进行了深入分析。我们发现,虽然与人类注释者的整体相关性很高,但由于某些人为因素,我们观察到严重受损个体的相关性很高,而健康或轻度受损个体的相关性较低。言语产出随着认知能力的下降而下降,当测试评分依赖于单词命名时,会导致过度乐观的相关性。根据测试设计,回退处理引入了有利于某些群体的进一步偏见。这些缺陷仍然独立于数据集中的群体分布,需要对目标群体进行差异化分析。
摘要:Current work on speech-based dementia assessment focuses on either feature extraction to predict assessment scales, or on the automation of existing test procedures. Most research uses public data unquestioningly and rarely performs a detailed error analysis, focusing primarily on numerical performance. We perform an in-depth analysis of an automated standardized dementia assessment, the Syndrom-Kurz-Test. We find that while there is a high overall correlation with human annotators, due to certain artifacts, we observe high correlations for the severely impaired individuals, which is less true for the healthy or mildly impaired ones. Speech production decreases with cognitive decline, leading to overoptimistic correlations when test scoring relies on word naming. Depending on the test design, fallback handling introduces further biases that favor certain groups. These pitfalls remain independent of group distributions in datasets and require differentiated analysis of target groups.


【3】Melodic and Metrical Elements of Expressiveness in Hindustani Vocal Music
标题:印度斯坦声乐表现力的旋律和节拍元素
链接:https://arxiv.org/abs/2508.04430

作者:Yash Bhake, Ankit Anand, Preeti Rao
备注:To appear in the proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR), Daejeon Korea, 2025
摘要:本文试图从艺术家在表演流行作品时所表现出的灵活性来研究北印度卡雅尔音乐的美学。我们研究表达的时间和音高变化的给定的抒情内容内和跨性能,并提出计算表示,可以区分不同的表现,同一首歌的表达。我们提出了必要的音频处理和注释程序,并讨论了我们的意见和见解,从分析数据集的两首歌曲在两个ragas每个呈现由10位著名艺术家。
摘要:This paper presents an attempt to study the aesthetics of North Indian Khayal music with reference to the flexibility exercised by artists in performing popular compositions. We study expressive timing and pitch variations of the given lyrical content within and across performances and propose computational representations that can discriminate between different performances of the same song in terms of expression. We present the necessary audio processing and annotation procedures, and discuss our observations and insights from the analysis of a dataset of two songs in two ragas each rendered by ten prominent artists.


【4】Text adaptation for speaker verification with speaker-text factorized embeddings
标题:使用说话人文本因子分解嵌入进行说话人验证的文本自适应
链接:https://arxiv.org/abs/2508.04425

作者:Yexin Yang, Shuai Wang, Xun Gong, Yanmin Qian, Kai Yu
备注:ICASSP 2020
摘要:预先收集的训练数据或注册数据与实际测试数据之间的文本不匹配会严重影响文本相关说话人确认(SV)系统的性能。虽然这个问题可以通过仔细地收集具有目标语音内容的数据来解决,但是这样的数据收集可能是昂贵且不灵活的。在本文中,我们提出了一个新的文本自适应框架,以解决文本不匹配的问题。本文提出了一种说话人-文本分解网络,将输入语音分解为说话人嵌入和文本嵌入,然后在后期将它们集成为单个表示。给定少量的与说话者无关的自适应话语,可以提取目标语音内容的文本嵌入,并将其用于使与文本无关的说话者嵌入适应于文本定制的说话者嵌入。在RSR 2015上的实验表明,文本自适应可以显著提高文本失配条件下的性能。
摘要:Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully collecting data with the target speech content, such data collection could be costly and inflexible. In this paper, we propose a novel text adaptation framework to address the text mismatch issue. Here, a speaker-text factorization network is proposed to factorize the input speech into speaker embeddings and text embeddings and then integrate them into a single representation in the later stage. Given a small amount of speaker-independent adaptation utterances, text embeddings of target speech content can be extracted and used to adapt the text-independent speaker embeddings to text-customized speaker embeddings. Experiments on RSR2015 show that text adaptation can significantly improve the performance of text mismatch conditions.


【5】Binaural Sound Event Localization and Detection Neural Network based on HRTF Localization Cues for Humanoid Robots
标题:基于HRTF定位线索的仿人机器人双耳声音事件定位和检测神经网络
链接:https://arxiv.org/abs/2508.04333

作者:Gyeong-Tae Lee
备注:200 pages
摘要:类人机器人需要同时进行声音事件类型和方向估计以用于态势感知,但是传统的双通道输入与高度估计和前后混淆作斗争。本文提出了一种双耳声音事件定位和检测(BiSELD)神经网络来解决这些挑战。BiSELDnet从双耳输入特征中学习时频模式和头部相关传递函数(HRTF)定位线索。介绍了一种新的八通道双耳时频特征(BTFF),包括左/右梅尔频谱图、V图、耳间时间差(ITD)图(低于1.5 kHz)、耳间电平差(ILD)图(高于5 kHz,具有前后不对称性)和频谱线索(SC)图(高于5 kHz,用于仰角)。在全向、水平和正中平面上证实了BTFF的有效性。BiSELDnets,特别是一个有效的三位一体模块的基础上,实现输出的方向向量的时间序列为每个声音事件类,使同时检测和定位。矢量激活图(VAM)可视化被提出来分析网络学习,证实了BiSELDnet对海拔估计的N1陷波频率的关注。城市背景噪声条件下的比较评估表明,建议的BiSELD模型显着优于国家的最先进的(SOTA)SELD模型与双耳输入。
摘要:Humanoid robots require simultaneous sound event type and direction estimation for situational awareness, but conventional two-channel input struggles with elevation estimation and front-back confusion. This paper proposes a binaural sound event localization and detection (BiSELD) neural network to address these challenges. BiSELDnet learns time-frequency patterns and head-related transfer function (HRTF) localization cues from binaural input features. A novel eight-channel binaural time-frequency feature (BTFF) is introduced, comprising left/right mel-spectrograms, V-maps, an interaural time difference (ITD) map (below 1.5 kHz), an interaural level difference (ILD) map (above 5 kHz with front-back asymmetry), and spectral cue (SC) maps (above 5 kHz for elevation). The effectiveness of BTFF was confirmed across omnidirectional, horizontal, and median planes. BiSELDnets, particularly one based on the efficient Trinity module, were implemented to output time series of direction vectors for each sound event class, enabling simultaneous detection and localization. Vector activation map (VAM) visualization was proposed to analyze network learning, confirming BiSELDnet's focus on the N1 notch frequency for elevation estimation. Comparative evaluations under urban background noise conditions demonstrated that the proposed BiSELD model significantly outperforms state-of-the-art (SOTA) SELD models with binaural input.


【6】A Multi-stage Low-latency Enhancement System for Hearing Aids
标题:一种多阶段低延迟助听器增强系统
链接:https://arxiv.org/abs/2508.04283

作者:Chengwei Ouyang, Kexin Fei, Haoshuai Zhou, Congxi Lu, Linkai Li
备注:2 pages, 1 figure, 1 table. accepted to ICASSP 2023
摘要:本文为ICASSP 2023 Clarity Challenge提出了一个端到端系统。在这项工作中,我们介绍了四个主要的创新:(1)一个新的多级系统在幅度和复杂的域,以更好地利用相位信息;(2)一个非对称的窗口对,以实现更高的频率分辨率与5 ms的延迟约束;(3)集成的头部旋转信息和混合信号,以实现更好的增强;(4)后处理模块,其利用由基线系统提供的助听器放大级实现更高的助听器言语感知指数(HASPI)分数。
摘要:This paper proposes an end-to-end system for the ICASSP 2023 Clarity Challenge. In this work, we introduce four major novelties: (1) a novel multi-stage system in both the magnitude and complex domains to better utilize phase information; (2) an asymmetric window pair to achieve higher frequency resolution with the 5ms latency constraint; (3) the integration of head rotation information and the mixture signals to achieve better enhancement; (4) a post-processing module that achieves higher hearing aid speech perception index (HASPI) scores with the hearing aid amplification stage provided by the baseline system.


【7】Towards interpretable emotion recognition: Identifying key features with machine learning
标题:走向可解释的情感识别:用机器学习识别关键特征
链接:https://arxiv.org/abs/2508.04230

作者:Yacouba Kaloga, Ina Kodrasi
备注:None
摘要:无监督方法,如wav2vec2和HuBERT,在音频任务中实现了最先进的性能,导致从可解释特征的研究转移。然而,这些方法缺乏可解释性,限制了它们在医学等关键领域的适用性,在这些领域,理解特征相关性至关重要。为了更好地理解无监督模型的特征,识别与给定任务相关的可解释特征仍然至关重要。在这项工作中,我们专注于情感识别,并使用机器学习算法来识别和概括这项任务中最重要的可解释特征。虽然以前的研究探讨了情绪识别中的特征相关性,但它们往往受到狭窄背景的限制,并且呈现不一致的结果。我们的方法旨在克服这些限制,提供一个更广泛,更强大的框架,以确定最重要的可解释的功能。
摘要:Unsupervised methods, such as wav2vec2 and HuBERT, have achieved state-of-the-art performance in audio tasks, leading to a shift away from research on interpretable features. However, the lack of interpretability in these methods limits their applicability in critical domains like medicine, where understanding feature relevance is crucial. To better understand the features of unsupervised models, it remains critical to identify the interpretable features relevant to a given task. In this work, we focus on emotion recognition and use machine learning algorithms to identify and generalize the most important interpretable features for this task. While previous studies have explored feature relevance in emotion recognition, they are often constrained by narrow contexts and present inconsistent findings. Our approach aims to overcome these limitations, providing a broader and more robust framework for identifying the most important interpretable features.


【8】Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
标题:Deepfakes的多语言源跟踪:第一个基准
链接:https://arxiv.org/abs/2508.04143

作者:Xi Xuan, Yang Xiao, Rohan Kumar Das, Tomi Kinnunen
备注:Accepted at Interspeech SPSC 2025 - 5th Symposium on Security and Privacy in Speech Communication (Oral)
摘要:最近在生成人工智能方面取得的进展使得从几秒钟的音频中创建听起来自然的Deepfake语音变得越来越容易。虽然这些工具支持有用的应用程序,但它们也引起了严重的关注,因为它们可以用多种语言生成令人信服的假语音。目前的研究主要集中在检测虚假语音,但很少关注跟踪用于生成它的源模型。本文介绍了多语言语音deepfake源跟踪的第一个基准,涵盖单语言和跨语言场景。我们比较研究基于DSP和SSL的建模;研究SSL表示在不同语言上的微调如何影响跨语言泛化性能;并评估对看不见的语言和说话者的泛化。我们的研究结果提供了第一个全面的见解识别语音生成模型时,训练和推理语言不同的挑战。数据集、方案和代码可在https://github.com/xuanxixi/Multilingual-Source-Tracing上获得。
摘要:Recent progress in generative AI has made it increasingly easy to create natural-sounding deepfake speech from just a few seconds of audio. While these tools support helpful applications, they also raise serious concerns by making it possible to generate convincing fake speech in many languages. Current research has largely focused on detecting fake speech, but little attention has been given to tracing the source models used to generate it. This paper introduces the first benchmark for multilingual speech deepfake source tracing, covering both mono- and cross-lingual scenarios. We comparatively investigate DSP- and SSL-based modeling; examine how SSL representations fine-tuned on different languages impact cross-lingual generalization performance; and evaluate generalization to unseen languages and speakers. Our findings offer the first comprehensive insights into the challenges of identifying speech generation models when training and inference languages differ. The dataset, protocol and code are available at https://github.com/xuanxixi/Multilingual-Source-Tracing.


【9】Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
标题:并行GPT:协调Zero-Shot文本到语音的声学和语义信息的独立性和相互依赖性
链接:https://arxiv.org/abs/2508.04141

作者:Jingyuan Xing, Zhipeng Li, Jialong Mai, Xiaofen Xing, Xiangmin Xu
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:语音表示和大型语言模型的进步增强了zero-shot文本到语音(TTS)的性能。然而,现有的zero-shot TTS模型在捕获声学和语义特征之间的复杂相关性方面面临挑战,导致缺乏表达性和相似性。本文提出了一种结合自回归(AR)和非自回归(NAR)模块的TTS框架,以协调语音信息和语义信息的独立性和相互依赖性。AR模型利用所提出的并行标记器来同时合成顶级语义和声学标记。相比之下,考虑到相互依赖性,耦合NAR模型基于一般AR模型的输出来预测详细的令牌。基于该结构设计的并行GPT,通过其并行结构来提高zero-shot文语合成。在英语和汉语数据集上的实验表明,该模型的合成质量和效率明显优于现有的zero-shot TTS模型。语音演示可在https://t1235-ch.github.io/pgpt/上获得。
摘要:Advances in speech representation and large language models have enhanced zero-shot text-to-speech (TTS) performance. However, existing zero-shot TTS models face challenges in capturing the complex correlations between acoustic and semantic features, resulting in a lack of expressiveness and similarity. The primary reason lies in the complex relationship between semantic and acoustic features, which manifests independent and interdependent aspects.This paper introduces a TTS framework that combines both autoregressive (AR) and non-autoregressive (NAR) modules to harmonize the independence and interdependence of acoustic and semantic information. The AR model leverages the proposed Parallel Tokenizer to synthesize the top semantic and acoustic tokens simultaneously. In contrast, considering the interdependence, the Coupled NAR model predicts detailed tokens based on the general AR model's output. Parallel GPT, built on this architecture, is designed to improve zero-shot text-to-speech synthesis through its parallel structure. Experiments on English and Chinese datasets demonstrate that the proposed model significantly outperforms the quality and efficiency of the synthesis of existing zero-shot TTS models. Speech demos are available at https://t1235-ch.github.io/pgpt/.


【10】LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
标题:LCS-CSC:利用软对齐增强语音转录稳健性
链接:https://arxiv.org/abs/2508.03937

作者:Zongli Ye, Jiachen Lian, Akshaj Gupta, Xuanru Zhou, Krish Patel, Haodong Li, Hwi Joo Park, Chenxu Guo, Shuhe Li, Sam Wang, Cheol Jun Cho, Zoe Ezzes, Jet M.J. Vonk, Brittany T. Morin, Rian Bogley, Lisa Wauters, Zachary A. Miller, Maria Luisa Gorno-Tempini, Gopala Anumanchipalli
备注:2025 ASRU
摘要:语音转录对于细粒度语言分析和下游语音应用至关重要。虽然连接主义时间分类(CTC)是一种广泛使用的方法,由于其效率,这类任务,它往往在识别性能不足,特别是在不清楚和不流利的语音。在这项工作中,我们提出了LCS-CTC,一个两阶段的框架音素级语音识别,结合了一个相似性感知的局部对齐算法与受约束的CTC训练目标。通过预测细粒度的帧音素成本矩阵和应用修改后的最长公共子序列(LCS)算法,我们的方法确定高置信度的对齐区域,用于约束CTC解码路径空间,从而减少过拟合和提高泛化能力,这使得鲁棒的识别和文本自由的强制对齐。在LibriSpeech和PPA上的实验表明,LCS-CTC始终优于vanilla CTC基线,这表明它有潜力在流利和不流利的语音中统一音素建模。
摘要:Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used approach for such tasks due to its efficiency, it often falls short in recognition performance, especially under unclear and nonfluent speech. In this work, we propose LCS-CTC, a two-stage framework for phoneme-level speech recognition that combines a similarity-aware local alignment algorithm with a constrained CTC training objective. By predicting fine-grained frame-phoneme cost matrices and applying a modified Longest Common Subsequence (LCS) algorithm, our method identifies high-confidence alignment zones which are used to constrain the CTC decoding path space, thereby reducing overfitting and improving generalization ability, which enables both robust recognition and text-free forced alignment. Experiments on both LibriSpeech and PPA demonstrate that LCS-CTC consistently outperforms vanilla CTC baselines, suggesting its potential to unify phoneme modeling across fluent and non-fluent speech.


【11】Perch 2.0: The Bittern Lesson for Bioacoustics
标题:鲈鱼2.0:生物声学的苦涩课程
链接:https://arxiv.org/abs/2508.04665

作者:Bart van Merriënboer, Vincent Dumoulin, Jenny Hamer, Lauren Harrell, Andrea Burns, Tom Denton
摘要:Perch是一个高性能的生物声学预训练模型。它以监督的方式进行训练,为数千种发声物种提供现成的分类分数,并为迁移学习提供强大的嵌入。在这个新版本Perch 2.0中,我们从专门针对鸟类物种的训练扩展到大型多分类群数据集。该模型使用原型学习分类器以及新的源预测训练标准进行自蒸馏训练。Perch 2.0在BirdSet和BEANS基准测试中获得了最先进的性能。尽管几乎没有海洋训练数据,但它在海洋迁移学习任务上的表现也优于专门的海洋模型。我们提出的假设,为什么细粒度的物种分类是一个特别强大的预训练任务的生物声学。
摘要:Perch is a performant pre-trained model for bioacoustics. It was trained in supervised fashion, providing both off-the-shelf classification scores for thousands of vocalizing species as well as strong embeddings for transfer learning. In this new release, Perch 2.0, we expand from training exclusively on avian species to a large multi-taxa dataset. The model is trained with self-distillation using a prototype-learning classifier as well as a new source-prediction training criterion. Perch 2.0 obtains state-of-the-art performance on the BirdSet and BEANS benchmarks. It also outperforms specialized marine models on marine transfer learning tasks, despite having almost no marine training data. We present hypotheses as to why fine-grained species classification is a particularly robust pre-training task for bioacoustics.


【12】The State Of TTS: A Case Study with Human Fooling Rates
标题:TTS的状况:人类愚弄率的案例研究
链接:https://arxiv.org/abs/2508.04179

作者:Praveen Srinivasa Varadhan, Sherry Thomas, Sai Teja M. S., Suvrat Bhooshan, Mitesh M. Khapra
备注:Accepted at InterSpeech 2025
摘要:虽然近年来的主观评估表明TTS的快速发展,目前的TTS系统真的可以通过人类的欺骗测试在图灵类评估?我们引入了人类愚弄率(HFR),这是一个直接衡量机器生成的语音被误认为人类的频率的指标。我们对开源和商业TTS模型的大规模评估揭示了重要的见解:(i)基于CMOS的人类平等的主张往往在欺骗测试中失败,(ii)TTS进展应以人类语音达到高HFR的数据集为基准,因为对单调或表达较少的参考样本进行评估设置了低标准,(iii)商业模型在zero-shot设置中接近人类欺骗,而开放源码系统仍在与自然的对话语音作斗争; ㈣对高质量数据进行微调可以提高真实性,但不能完全弥合差距。我们的研究结果强调,除了现有的主观测试外,还需要更现实、以人为本的评估。
摘要:While subjective evaluations in recent years indicate rapid progress in TTS, can current TTS systems truly pass a human deception test in a Turing-like evaluation? We introduce Human Fooling Rate (HFR), a metric that directly measures how often machine-generated speech is mistaken for human. Our large-scale evaluation of open-source and commercial TTS models reveals critical insights: (i) CMOS-based claims of human parity often fail under deception testing, (ii) TTS progress should be benchmarked on datasets where human speech achieves high HFRs, as evaluating against monotonous or less expressive reference samples sets a low bar, (iii) Commercial models approach human deception in zero-shot settings, while open-source systems still struggle with natural conversational speech; (iv) Fine-tuning on high-quality data improves realism but does not fully bridge the gap. Our findings underscore the need for more realistic, human-centric evaluations alongside existing subjective tests.


【13】Efficient Scaling for LLM-based ASR
标题:基于LLM的ASB的高效扩展
链接:https://arxiv.org/abs/2508.04096

作者:Bingshen Mu, Yiwen Shao, Kun Wei, Dong Yu, Lei Xie
备注:Accepted by ASRU 2025
摘要:基于大语言模型(LLM)的自动语音识别(ASR)具有很强的性能,但通常会产生很高的计算成本。本文研究了如何有效地获得最佳LLM-ASR性能。通过全面和受控的实验,我们发现,在将语音编码器与LLM集成之前对其进行预训练,比LLM-ASR的联合后训练的标准实践具有更好的缩放效率。基于这种见解,我们提出了一种新的多阶段LLM-ASR训练策略EFIN:编码器优先集成。在所有评估的训练策略中,EFIN始终提供更好的性能(相对于21.1%CERR),同时显著降低计算预算(49.9%FLOP)。此外,我们推导出一个缩放律,近似ASR的错误率作为计算函数,LLM-ASR缩放提供了实际的指导。
摘要:Large language model (LLM)-based automatic speech recognition (ASR) achieves strong performance but often incurs high computational costs. This work investigates how to obtain the best LLM-ASR performance efficiently. Through comprehensive and controlled experiments, we find that pretraining the speech encoder before integrating it with the LLM leads to significantly better scaling efficiency than the standard practice of joint post-training of LLM-ASR. Based on this insight, we propose a new multi-stage LLM-ASR training strategy, EFIN: Encoder First Integration. Among all training strategies evaluated, EFIN consistently delivers better performance (relative to 21.1% CERR) with significantly lower computation budgets (49.9% FLOPs). Furthermore, we derive a scaling law that approximates ASR error rates as a computation function, providing practical guidance for LLM-ASR scaling.


【14】MiDashengLM: Efficient Audio Understanding with General Audio Captions
标题:MiDashengLM:通过通用音频说明高效的音频理解
链接:https://arxiv.org/abs/2508.03983

作者:Horizon Team
摘要:目前的大型音频语言模型(LALM)方法通常依赖于封闭的数据源或专有模型,限制了它们的通用性和可访问性。本文介绍了MiDashengLM,这是一种新型的开放式音频语言模型,旨在通过使用我们新的ACAVCaps训练数据集使用通用音频字幕进行高效和全面的音频理解。MiDashengLM完全依赖于公开的预训练和监督微调(SFT)数据集,确保完全透明和可重复性。在其核心,MiDashengLM集成了Dasheng,一个开源的音频编码器,专门设计用于有效地处理各种听觉信息。与以前的作品主要集中在基于自动语音识别(ASR)的音频文本对齐,我们的策略集中在一般的音频字幕,融合语音,声音和音乐信息到一个文本表示,使复杂的音频场景的整体文本表示。最后,MiDashengLM在首次令牌时间(TTFT)方面提供了高达4倍的加速,并且吞吐量比同类模型高出20倍。检查点可在https://huggingface.co/mispeech/midashenglm-7b和https://github.com/xiaomi-research/dasheng-lm上查阅。
摘要:Current approaches for large audio language models (LALMs) often rely on closed data sources or proprietary models, limiting their generalization and accessibility. This paper introduces MiDashengLM, a novel open audio-language model designed for efficient and comprehensive audio understanding through the use of general audio captions using our novel ACAVCaps training dataset. MiDashengLM exclusively relies on publicly available pretraining and supervised fine-tuning (SFT) datasets, ensuring full transparency and reproducibility. At its core, MiDashengLM integrates Dasheng, an open-source audio encoder, specifically engineered to process diverse auditory information effectively. Unlike previous works primarily focused on Automatic Speech Recognition (ASR) based audio-text alignment, our strategy centers on general audio captions, fusing speech, sound and music information into one textual representation, enabling a holistic textual representation of complex audio scenes. Lastly, MiDashengLM provides an up to 4x speedup in terms of time-to-first-token (TTFT) and up to 20x higher throughput than comparable models. Checkpoints are available online at https://huggingface.co/mispeech/midashenglm-7b and https://github.com/xiaomi-research/dasheng-lm.


【15】Are Inherently Interpretable Models More Robust? A Study In Music Emotion Recognition
标题:固有可解释模型是否更稳健?音乐情感识别研究
链接:https://arxiv.org/abs/2508.03780

作者: Katharina HOEDT, Arthur Flexer, Gerhard Widmer
备注:8 pages, published in Proceedings of the 22nd Sound and Music Computing Conference 2025 (SMC-25)
摘要:深度学习模型所需的关键属性之一是能够概括未知样本。当提供与一个或多个训练样本(感知上)相似的新样本时,深度学习模型预计会产生相应的相似输出。成功预测类似输入的类似输出的模型通常被称为鲁棒模型。另一方面,深度学习模型已经被证明非常容易受到输入的微小(对抗性)扰动的影响,这会极大地改变模型的输出,同时暴露出它对虚假相关性的依赖。在这项工作中,我们研究了内在可解释的深度模型,即,与黑箱模型相比,被设计为更关注有意义和可解释的特征的深度模型对数据中不相关的扰动更鲁棒。我们通过比较可解释的和黑盒音乐情感识别(MER)模型在对抗性示例挑战时的鲁棒性来测试我们的假设。此外,我们还包括一个经过对抗训练的模型,该模型经过优化,在比较中更加强大。我们的研究结果表明,本质上更具可解释性的模型确实可以比黑箱模型更健壮,并以更低的计算成本达到与对抗训练模型相似的健壮性水平。
摘要:One of the desired key properties of deep learning models is the ability to generalise to unseen samples. When provided with new samples that are (perceptually) similar to one or more training samples, deep learning models are expected to produce correspondingly similar outputs. Models that succeed in predicting similar outputs for similar inputs are often called robust. Deep learning models, on the other hand, have been shown to be highly vulnerable to minor (adversarial) perturbations of the input, which manage to drastically change a model's output and simultaneously expose its reliance on spurious correlations. In this work, we investigate whether inherently interpretable deep models, i.e., deep models that were designed to focus more on meaningful and interpretable features, are more robust to irrelevant perturbations in the data, compared to their black-box counterparts. We test our hypothesis by comparing the robustness of an interpretable and a black-box music emotion recognition (MER) model when challenged with adversarial examples. Furthermore, we include an adversarially trained model, which is optimised to be more robust, in the comparison. Our results indicate that inherently more interpretable models can indeed be more robust than their black-box counterparts, and achieve similar levels of robustness as adversarially trained models, at lower computational cost.


机器翻译由腾讯交互翻译提供,仅供参考