今日论文合集:cs.SD语音7篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Self-Supervised Models for Phoneme Recognition: Applications in  Children's Speech for Reading Learning
标题:音素识别的自我监督模型:在儿童阅读学习语音中的应用
链接:https://arxiv.org/abs/2503.04710
作者:Lucas Block Medin,  Thomas Pellegrini,  Lucile Gelin
备注:This paper was originally published in the Proceedings of Interspeech 2024. DOI: 10.21437/Interspeech.2024-1095
摘要:儿童语音识别仍然是一个欠发达的研究领域,由于缺乏数据(特别是非英语语言)和这项任务的具体困难。在以前的工作中探索了儿童语音识别的各种架构,在本文中,我们将讨论最近的自监督模型。我们首先比较了wav2vec 2.0,HuBERT和WavLM模型适用于法语儿童语音中的音素识别,并继续我们的实验,其中最好的,WavLM基地+。然后,我们通过在对子语音进行微调期间解冻其Transformer块来进一步调整它,这大大提高了它的性能,并使其显著优于我们的基本模型,即Transformer+CTC。最后,我们详细研究了这两个模型在实际应用条件下的行为,并表明WavLM base+对各种阅读任务和噪声水平更具鲁棒性。索引术语:语音识别,儿童语音,自我监督学习
摘要:Child speech recognition is still an underdeveloped area of research due tothe lack of data (especially on non-English languages) and the specificdifficulties of this task. Having explored various architectures for childspeech recognition in previous work, in this article we tackle recentself-supervised models. We first compare wav2vec 2.0, HuBERT and WavLM modelsadapted to phoneme recognition in French child speech, and continue ourexperiments with the best of them, WavLM base+. We then further adapt it byunfreezing its transformer blocks during fine-tuning on child speech, whichgreatly improves its performance and makes it significantly outperform our basemodel, a Transformer+CTC. Finally, we study in detail the behaviour of thesetwo models under the real conditions of our application, and show that WavLMbase+ is more robust to various reading tasks and noise levels. Index Terms:speech recognition, child speech, self-supervised learning

【2】 TAIL: Text-Audio Incremental Learning
标题:TAIL:文本音频增量学习
链接:https://arxiv.org/abs/2503.04258
作者:Yingfei Sun,  Xu Gu,  Wei Ji,  Hanbin Zhao,  Hao Fei,  Yifang Yin,  Roger Zimmermann
备注:4 figures, 5 tables
摘要:许多研究结合文本和音频来捕获多模态信息,但他们忽略了模型在新数据集上的泛化能力。引入新的数据集可能会影响原始数据集的特征空间,导致灾难性的遗忘。同时,大的模型参数会显著影响训练性能。为了解决这些局限性,我们引入了一个新的任务,称为文本音频增量学习(TAIL)任务的文本音频检索,并提出了一种新的方法,PTAT,提示调整音频文本增量学习。该方法利用快速调整优化模型参数,同时结合音频文本相似性和特征提取模块,以有效地减轻灾难性遗忘。我们在AudioCaps,Clotho,BBC Sound Effects和Audioset数据集上对我们的方法和以前的增量学习方法进行了基准测试,我们的方法明显优于以前的方法,特别是在旧数据集上表现出更强的抗遗忘能力。与全参数Finetune(Sequential)方法相比,该模型只需要2.42%的参数,性能提高了4.46%.
摘要:Many studies combine text and audio to capture multi-modal information butthey overlook the model's generalization ability on new datasets. Introducingnew datasets may affect the feature space of the original dataset, leading tocatastrophic forgetting. Meanwhile, large model parameters can significantlyimpact training performance. To address these limitations, we introduce a noveltask called Text-Audio Incremental Learning (TAIL) task for text-audioretrieval, and propose a new method, PTAT, Prompt Tuning for Audio-Textincremental learning. This method utilizes prompt tuning to optimize the modelparameters while incorporating an audio-text similarity and featuredistillation module to effectively mitigate catastrophic forgetting. Webenchmark our method and previous incremental learning methods on AudioCaps,Clotho, BBC Sound Effects and Audioset datasets, and our method outperformsprevious methods significantly, particularly demonstrating stronger resistanceto forgetting on older datasets. Compared to the full-parameters Finetune(Sequential) method, our model only requires 2.42\% of its parameters,achieving 4.46\% higher performance.

【3】 Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding  and Expert Reasoning Abilities
标题:音频火烈鸟2:具有长音频理解和专家推理能力的音频语言模型
链接:https://arxiv.org/abs/2503.03983
作者:Sreyan Ghosh,  Zhifeng Kong,  Sonal Kumar,  S Sakshi,  Jaehyeon Kim,  Wei Ping,  Rafael Valle,  Dinesh Manocha,  Bryan Catanzaro
摘要:对非语音声音和音乐的理解和推理对于人类和AI智能体与环境进行有效交互至关重要。在本文中,我们介绍了音频火烈鸟2(AF 2),一个音频语言模型(ALM)与先进的音频理解和推理能力。AF2利用(i)自定义CLAP模型,(ii)用于细粒度音频推理的合成音频QA数据,以及(iii)多阶段课程学习策略。AF2仅用3B参数小型语言模型就实现了最先进的性能,在20多个基准测试中超过了大型开源和专有模型。接下来,我们第一次将音频理解扩展到长音频片段(30秒到5分钟),并提出了LongAudio,这是一个大型且新颖的数据集,用于在长音频字幕和问答任务上训练ALM。在LongAudio上微调AF 2会在我们提出的LongAudioBench上产生出色的性能,LongAudioBench是一个专家注释的基准,用于评估长音频理解能力的ALM。我们进行了广泛的消融研究,以确认我们的方法的有效性。项目网址:https://research.nvidia.com/labs/adlr/AF2/。
摘要:Understanding and reasoning over non-speech sounds and music are crucial forboth humans and AI agents to interact effectively with their environments. Inthis paper, we introduce Audio Flamingo 2 (AF2), an Audio-Language Model (ALM)with advanced audio understanding and reasoning capabilities. AF2 leverages (i)a custom CLAP model, (ii) synthetic Audio QA data for fine-grained audioreasoning, and (iii) a multi-stage curriculum learning strategy. AF2 achievesstate-of-the-art performance with only a 3B parameter small language model,surpassing large open-source and proprietary models across over 20 benchmarks.Next, for the first time, we extend audio understanding to long audio segments(30 secs to 5 mins) and propose LongAudio, a large and novel dataset fortraining ALMs on long audio captioning and question-answering tasks.Fine-tuning AF2 on LongAudio leads to exceptional performance on our proposedLongAudioBench, an expert annotated benchmark for evaluating ALMs on long audiounderstanding capabilities. We conduct extensive ablation studies to confirmthe efficacy of our approach. Project Website:https://research.nvidia.com/labs/adlr/AF2/.

【4】 VoiceGRPO: Modern MoE Transformers with Group Relative Policy  Optimization GRPO for AI Voice Health Care Applications on Voice Pathology  Detection
标题:SecureGRPO:具有团体相对政策优化的现代MoE部TransformerGRPO,用于语音病理检测方面的人工智能语音医疗保健应用
链接:https://arxiv.org/abs/2503.03797
作者:Enkhtogtokh Togootogtokh,  Christian Klasen
摘要:本研究介绍一种新的人工智能技术,作为混合专家Transformers与组相对策略优化(GRPO)的语音健康护理应用程序的语音病理检测。通过架构创新,我们采用了受强化学习启发的高级训练范式,即近端策略优化(PPO)和分组正则化策略优化(GRPO),以提高模型的稳定性和性能。在综合生成的语音病理数据集上进行的实验表明,与传统方法相比,我们提出的模型显着提高了诊断准确性,F1评分和ROC-AUC。这些发现强调了将Transformer架构与新的训练策略相结合的潜力,以推进自动语音病理检测,并最终有助于更有效的医疗服务。我们用来训练和评估模型的代码可以在https://github.com/enkhtogtokh/voicegrpo上找到
摘要:This research introduces a novel AI techniques as Mixture-of-ExpertsTransformers with Group Relative Policy Optimization (GRPO) for voice healthcare applications on voice pathology detection. With the architecturalinnovations, we adopt advanced training paradigms inspired by reinforcementlearning, namely Proximal Policy Optimization (PPO) and Group-wise RegularizedPolicy Optimization (GRPO), to enhance model stability and performance.Experiments conducted on a synthetically generated voice pathology datasetdemonstrate that our proposed models significantly improve diagnostic accuracy,F1 score, and ROC-AUC compared to conventional approaches. These findingsunderscore the potential of integrating transformer architectures with noveltraining strategies to advance automated voice pathology detection andultimately contribute to more effective healthcare delivery. The code we usedto train and evaluate our models is available athttps://github.com/enkhtogtokh/voicegrpo

【5】 Efficient Finetuning for Dimensional Speech Emotion Recognition in the  Age of Transformers
标题:Transformer时代二维语音情感识别的高效微调
链接:https://arxiv.org/abs/2503.03756
作者:Aneesha Sampath,  James Tavernor,  Emily Mower Provost
备注:ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
摘要:准确的语音情感识别对于开发面向人类的系统至关重要。最近的进展包括微调大型的预训练Transformer模型,如Wav2Vec 2.0。然而,微调过程需要大量的计算资源,包括高内存GPU和大量的处理时间。随着对精确情感识别的需求持续增长,需要有效的微调方法来减少计算负担。我们的研究重点是维度情绪识别,预测属性,如激活(平静到兴奋)和效价(消极到积极)。我们提出了各种微调技术,包括全微调,部分微调的Transformer层,微调与混合精度,部分微调与缓存,低秩自适应(LoRA)的Wav2Vec 2.0基础模型。我们发现,混合精度的部分微调实现了与完全微调相当的性能,同时将训练速度提高了67%。缓存中间表示进一步提高了效率,产生了88%的加速和71%的可学习参数减少。我们建议以混合精度微调最后三个Transformer层,以平衡性能和训练效率,并添加中间表示缓存,以实现最佳速度和最小性能折衷。这些发现降低了微调语音情感识别系统的障碍,使更广泛的研究人员和从业者更容易获得准确的情感识别。
摘要:Accurate speech emotion recognition is essential for developing human-facingsystems. Recent advancements have included finetuning large, pretrainedtransformer models like Wav2Vec 2.0. However, the finetuning process requiressubstantial computational resources, including high-memory GPUs and significantprocessing time. As the demand for accurate emotion recognition continues togrow, efficient finetuning approaches are needed to reduce the computationalburden. Our study focuses on dimensional emotion recognition, predictingattributes such as activation (calm to excited) and valence (negative topositive). We present various finetuning techniques, including full finetuning,partial finetuning of transformer layers, finetuning with mixed precision,partial finetuning with caching, and low-rank adaptation (LoRA) on the Wav2Vec2.0 base model. We find that partial finetuning with mixed precision achievesperformance comparable to full finetuning while increasing training speed by67%. Caching intermediate representations further boosts efficiency, yieldingan 88% speedup and a 71% reduction in learnable parameters. We recommendfinetuning the final three transformer layers in mixed precision to balanceperformance and training efficiency, and adding intermediate representationcaching for optimal speed with minimal performance trade-offs. These findingslower the barriers to finetuning speech emotion recognition systems, makingaccurate emotion recognition more accessible to a broader range of researchersand practitioners.

【6】 Scaling Rich Style-Prompted Text-to-Speech Datasets
标题:扩展丰富风格的文本到语音数据集
链接:https://arxiv.org/abs/2503.04713
作者:Anuj Diwan,  Zhisheng Zheng,  David Harwath,  Eunsol Choi
摘要:我们介绍了副语言语音字幕(ParaSpeechCaps),一个大规模的数据集,注释语音话语丰富的风格字幕。虽然在小规模的人类注释数据集中已经探索了丰富的抽象标签(例如喉音,鼻音,疼痛),但现有的大规模数据集仅涵盖基本标签(例如低音,慢,大声)。我们结合了现成的文本和语音嵌入器,分类器和音频语言模型,首次自动扩展丰富的标签注释。ParaSpeechCaps涵盖了总共59个风格标签,包括说话者级别的内在标签和话语级别的情景标签。它包括342小时的人工标记数据(PSC-Base)和2427小时的自动注释数据(PSC-Scaled)。我们在ParaSpeechCaps上微调了Parler-TTS,一个开源的风格提示的TTS模型,并在结合现有丰富风格标签数据集的最佳基线上实现了改进的风格一致性(+7.9%的一致性MOS)和语音质量(+15.5%的自然度MOS)。我们消除了我们的几个数据集设计选择,为未来在这个领域的工作奠定基础。我们的数据集,模型和代码发布在https://github.com/ajd12342/paraspeechcaps。
摘要:We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scaledataset that annotates speech utterances with rich style captions. While richabstract tags (e.g. guttural, nasal, pained) have been explored in small-scalehuman-annotated datasets, existing large-scale datasets only cover basic tags(e.g. low-pitched, slow, loud). We combine off-the-shelf text and speechembedders, classifiers and an audio language model to automatically scale richtag annotations for the first time. ParaSpeechCaps covers a total of 59 styletags, including both speaker-level intrinsic tags and utterance-levelsituational tags. It consists of 342 hours of human-labelled data (PSC-Base)and 2427 hours of automatically annotated data (PSC-Scaled). We finetuneParler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, andachieve improved style consistency (+7.9% Consistency MOS) and speech quality(+15.5% Naturalness MOS) over the best performing baseline that combinesexisting rich style tag datasets. We ablate several of our dataset designchoices to lay the foundation for future work in this space. Our dataset,models and code are released at https://github.com/ajd12342/paraspeechcaps .

【7】 Frequency-Based Alignment of EEG and Audio Signals Using Contrastive  Learning and SincNet for Auditory Attention Detection
标题:使用对比学习和SincNet对脑电和音频信号进行基于频率的对齐以进行听觉注意力检测
链接:https://arxiv.org/abs/2503.04156
作者:Yuan Liao,  Yuhong Zhang,  Qiushi Han,  Yuhang Yang,  Weiwei Ding,  Yuzhe Gu,  Hengxin Yang,  Liya Huang
摘要:人类在复杂的声学环境中表现出非凡的听觉注意力集中能力,例如鸡尾酒会。听觉注意检测(AAD)的目的是通过分析大脑信号,如脑电图(EEG)数据来识别被关注的说话者。现有的AAD算法通常利用深度学习强大的非线性建模能力,很少考虑大脑中听觉处理的神经机制。本文提出了一种新的基于改进SincNet和对比学习的SincAlignNet网络,用于匹配听觉注意检测中的音频和EEG特征。SincNet组件模拟大脑在听觉注意过程中对音频的处理,而对比学习则指导模型学习EEG信号与受注意语音之间的关系。在推理过程中,我们计算了脑电和音频特征之间的余弦相似度,并探索了利用脑电数据直接推理说话人的方法。交叉试验评估结果表明,SincAlignNet在两个公开数据集KUL和DTU上的平均准确率分别为78.3%和92.2%,决策窗口为1秒,优于最新的AAD方法。该模型具有较强的可解释性,表明在男性和女性说话者的情景中,左右颞叶都更活跃。此外,我们发现,与使用64个电极相比,仅使用来自颞叶附近的6个电极的数据保持了类似或甚至更好的性能。这些发现表明,有效的低密度EEG在线解码是可以实现的,标志着在现实世界应用中神经引导助听器的实际实施迈出了重要的一步。代码可从以下网址获得:https://github.com/LiaoEuan/SincAlignNet。
摘要:Humans exhibit a remarkable ability to focus auditory attention in complexacoustic environments, such as cocktail parties. Auditory attention detection(AAD) aims to identify the attended speaker by analyzing brain signals, such aselectroencephalography (EEG) data. Existing AAD algorithms often leverage deeplearning's powerful nonlinear modeling capabilities, few consider the neuralmechanisms underlying auditory processing in the brain. In this paper, wepropose SincAlignNet, a novel network based on an improved SincNet andcontrastive learning, designed to align audio and EEG features for auditoryattention detection. The SincNet component simulates the brain's processing ofaudio during auditory attention, while contrastive learning guides the model tolearn the relationship between EEG signals and attended speech. Duringinference, we calculate the cosine similarity between EEG and audio featuresand also explore direct inference of the attended speaker using EEG data.Cross-trial evaluations results demonstrate that SincAlignNet outperformsstate-of-the-art AAD methods on two publicly available datasets, KUL and DTU,achieving average accuracies of 78.3% and 92.2%, respectively, with a 1-seconddecision window. The model exhibits strong interpretability, revealing that theleft and right temporal lobes are more active during both male and femalespeaker scenarios. Furthermore, we found that using data from only sixelectrodes near the temporal lobes maintains similar or even better performancecompared to using 64 electrodes. These findings indicate that efficientlow-density EEG online decoding is achievable, marking an important step towardthe practical implementation of neuro-guided hearing aids in real-worldapplications. Code is available at: https://github.com/LiaoEuan/SincAlignNet.

eess.AS音频处理

【1】 Scaling Rich Style-Prompted Text-to-Speech Datasets
标题:扩展丰富风格的文本到语音数据集
链接:https://arxiv.org/abs/2503.04713
作者:Anuj Diwan,  Zhisheng Zheng,  David Harwath,  Eunsol Choi
摘要:我们介绍了副语言语音字幕(ParaSpeechCaps),一个大规模的数据集,注释语音话语丰富的风格字幕。虽然在小规模的人类注释数据集中已经探索了丰富的抽象标签(例如喉音,鼻音,疼痛),但现有的大规模数据集仅涵盖基本标签(例如低音,慢,大声)。我们结合了现成的文本和语音嵌入器,分类器和音频语言模型,首次自动扩展丰富的标签注释。ParaSpeechCaps涵盖了总共59个风格标签,包括说话者级别的内在标签和话语级别的情景标签。它包括342小时的人工标记数据(PSC-Base)和2427小时的自动注释数据(PSC-Scaled)。我们在ParaSpeechCaps上微调了Parler-TTS,一个开源的风格提示的TTS模型,并在结合现有丰富风格标签数据集的最佳基线上实现了改进的风格一致性(+7.9%的一致性MOS)和语音质量(+15.5%的自然度MOS)。我们消除了我们的几个数据集设计选择,为未来在这个领域的工作奠定基础。我们的数据集,模型和代码发布在https://github.com/ajd12342/paraspeechcaps。
摘要:We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scaledataset that annotates speech utterances with rich style captions. While richabstract tags (e.g. guttural, nasal, pained) have been explored in small-scalehuman-annotated datasets, existing large-scale datasets only cover basic tags(e.g. low-pitched, slow, loud). We combine off-the-shelf text and speechembedders, classifiers and an audio language model to automatically scale richtag annotations for the first time. ParaSpeechCaps covers a total of 59 styletags, including both speaker-level intrinsic tags and utterance-levelsituational tags. It consists of 342 hours of human-labelled data (PSC-Base)and 2427 hours of automatically annotated data (PSC-Scaled). We finetuneParler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, andachieve improved style consistency (+7.9% Consistency MOS) and speech quality(+15.5% Naturalness MOS) over the best performing baseline that combinesexisting rich style tag datasets. We ablate several of our dataset designchoices to lay the foundation for future work in this space. Our dataset,models and code are released at https://github.com/ajd12342/paraspeechcaps .

【2】 Frequency-Based Alignment of EEG and Audio Signals Using Contrastive  Learning and SincNet for Auditory Attention Detection
标题:使用对比学习和SincNet对脑电和音频信号进行基于频率的对齐以进行听觉注意力检测
链接:https://arxiv.org/abs/2503.04156
作者:Yuan Liao,  Yuhong Zhang,  Qiushi Han,  Yuhang Yang,  Weiwei Ding,  Yuzhe Gu,  Hengxin Yang,  Liya Huang
摘要:人类在复杂的声学环境中表现出非凡的听觉注意力集中能力,例如鸡尾酒会。听觉注意检测(AAD)的目的是通过分析大脑信号,如脑电图(EEG)数据来识别被关注的说话者。现有的AAD算法通常利用深度学习强大的非线性建模能力,很少考虑大脑中听觉处理的神经机制。本文提出了一种新的基于改进SincNet和对比学习的SincAlignNet网络,用于匹配听觉注意检测中的音频和EEG特征。SincNet组件模拟大脑在听觉注意过程中对音频的处理,而对比学习则指导模型学习EEG信号与受注意语音之间的关系。在推理过程中,我们计算了脑电和音频特征之间的余弦相似度,并探索了利用脑电数据直接推理说话人的方法。交叉试验评估结果表明,SincAlignNet在两个公开数据集KUL和DTU上的平均准确率分别为78.3%和92.2%,决策窗口为1秒,优于最新的AAD方法。该模型具有较强的可解释性,表明在男性和女性说话者的情景中,左右颞叶都更活跃。此外,我们发现,与使用64个电极相比,仅使用来自颞叶附近的6个电极的数据保持了类似或甚至更好的性能。这些发现表明,有效的低密度EEG在线解码是可以实现的,标志着在现实世界中实际实施神经引导助听器的重要一步。代码可从以下网址获得:https://github.com/LiaoEuan/SincAlignNet。
摘要:Humans exhibit a remarkable ability to focus auditory attention in complexacoustic environments, such as cocktail parties. Auditory attention detection(AAD) aims to identify the attended speaker by analyzing brain signals, such aselectroencephalography (EEG) data. Existing AAD algorithms often leverage deeplearning's powerful nonlinear modeling capabilities, few consider the neuralmechanisms underlying auditory processing in the brain. In this paper, wepropose SincAlignNet, a novel network based on an improved SincNet andcontrastive learning, designed to align audio and EEG features for auditoryattention detection. The SincNet component simulates the brain's processing ofaudio during auditory attention, while contrastive learning guides the model tolearn the relationship between EEG signals and attended speech. Duringinference, we calculate the cosine similarity between EEG and audio featuresand also explore direct inference of the attended speaker using EEG data.Cross-trial evaluations results demonstrate that SincAlignNet outperformsstate-of-the-art AAD methods on two publicly available datasets, KUL and DTU,achieving average accuracies of 78.3% and 92.2%, respectively, with a 1-seconddecision window. The model exhibits strong interpretability, revealing that theleft and right temporal lobes are more active during both male and femalespeaker scenarios. Furthermore, we found that using data from only sixelectrodes near the temporal lobes maintains similar or even better performancecompared to using 64 electrodes. These findings indicate that efficientlow-density EEG online decoding is achievable, marking an important step towardthe practical implementation of neuro-guided hearing aids in real-worldapplications. Code is available at: https://github.com/LiaoEuan/SincAlignNet.

【3】 Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue  Models on Turn-taking Capabilities
标题:全速长凳:评估全速口语对话模型轮接能力的基准
链接:https://arxiv.org/abs/2503.04721
作者:Guan-Ting Lin,  Jiachen Lian,  Tingle Li,  Qirui Wang,  Gopala Anumanchipalli,  Alexander H. Liu,  Hung-yi Lee
摘要:口语对话建模引入了超越基于文本的语言建模的独特挑战,要求鲁棒的话轮转换,反向引导和实时交互。虽然大多数口语对话模型(SDM)依赖于半双工处理(一次处理一个回合的语音),但新兴的全双工SDM可以同时听和说,从而实现更自然、更吸引人的对话。然而,目前对这种模型的评估仍然有限,通常集中在基于话轮的度量或高级语料库分析(例如,转身间隙、停顿)。为了解决这一差距,我们提出了Full-Duplex-Bench,这是一个新的基准,可以系统地评估关键的对话行为:暂停处理、反向通道、话轮转换和中断管理。我们的框架使用自动指标的SDM的交互性能的一致性和可重复的评估。通过提供开放和标准化的评估基准,我们旨在推进口语对话建模,并鼓励开发更具交互性和自然性的对话系统。
摘要:Spoken dialogue modeling introduces unique challenges beyond text-basedlanguage modeling, demanding robust turn-taking, backchanneling, and real-timeinteraction. Although most Spoken Dialogue Models (SDMs) rely on half-duplexprocessing (handling speech one turn at a time), emerging full-duplex SDMs canlisten and speak simultaneously, enabling more natural and engagingconversations. However, current evaluations of such models remain limited,often focusing on turn-based metrics or high-level corpus analyses (e.g., turngaps, pauses). To address this gap, we present Full-Duplex-Bench, a newbenchmark that systematically evaluates key conversational behaviors: pausehandling, backchanneling, turn-taking, and interruption management. Ourframework uses automatic metrics for consistent and reproducible assessments ofSDMs' interactive performance. By offering an open and standardized evaluationbenchmark, we aim to advance spoken dialogue modeling and encourage thedevelopment of more interactive and natural dialogue systems.

【4】 Self-Supervised Models for Phoneme Recognition: Applications in  Children's Speech for Reading Learning
标题:音素识别的自我监督模型:在儿童阅读学习语音中的应用
链接:https://arxiv.org/abs/2503.04710
作者:Lucas Block Medin,  Thomas Pellegrini,  Lucile Gelin
备注:This paper was originally published in the Proceedings of Interspeech 2024. DOI: 10.21437/Interspeech.2024-1095
摘要:儿童语音识别仍然是一个欠发达的研究领域,由于缺乏数据(特别是非英语语言)和这项任务的具体困难。在以前的工作中探索了儿童语音识别的各种架构,在本文中,我们将讨论最近的自监督模型。我们首先比较了wav2vec 2.0,HuBERT和WavLM模型适用于法语儿童语音中的音素识别,并继续我们的实验,其中最好的,WavLM基地+。然后,我们通过在对子语音进行微调期间解冻其Transformer块来进一步调整它,这大大提高了它的性能,并使其显著优于我们的基本模型,即Transformer+CTC。最后,我们详细研究了这两个模型在实际应用条件下的行为,并表明WavLM base+对各种阅读任务和噪声水平更具鲁棒性。索引术语:语音识别,儿童语音,自我监督学习
摘要:Child speech recognition is still an underdeveloped area of research due tothe lack of data (especially on non-English languages) and the specificdifficulties of this task. Having explored various architectures for childspeech recognition in previous work, in this article we tackle recentself-supervised models. We first compare wav2vec 2.0, HuBERT and WavLM modelsadapted to phoneme recognition in French child speech, and continue ourexperiments with the best of them, WavLM base+. We then further adapt it byunfreezing its transformer blocks during fine-tuning on child speech, whichgreatly improves its performance and makes it significantly outperform our basemodel, a Transformer+CTC. Finally, we study in detail the behaviour of thesetwo models under the real conditions of our application, and show that WavLMbase+ is more robust to various reading tasks and noise levels. Index Terms:speech recognition, child speech, self-supervised learning

【5】 TAIL: Text-Audio Incremental Learning
标题:TAIL:文本音频增量学习
链接:https://arxiv.org/abs/2503.04258
作者:Yingfei Sun,  Xu Gu,  Wei Ji,  Hanbin Zhao,  Hao Fei,  Yifang Yin,  Roger Zimmermann
备注:4 figures, 5 tables
摘要:许多研究结合文本和音频来捕获多模态信息,但他们忽略了模型在新数据集上的泛化能力。引入新的数据集可能会影响原始数据集的特征空间,导致灾难性的遗忘。同时,大的模型参数会显著影响训练性能。为了解决这些局限性,我们引入了一个新的任务,称为文本音频增量学习(TAIL)任务的文本音频检索,并提出了一种新的方法,PTAT,提示调整音频文本增量学习。该方法利用快速调整优化模型参数,同时结合音频文本相似性和特征提取模块,以有效地减轻灾难性遗忘。我们在AudioCaps,Clotho,BBC Sound Effects和Audioset数据集上对我们的方法和以前的增量学习方法进行了基准测试,我们的方法明显优于以前的方法,特别是在旧数据集上表现出更强的抗遗忘能力。与全参数Finetune(Sequential)方法相比,该模型只需要2.42%的参数,性能提高了4.46%.
摘要:Many studies combine text and audio to capture multi-modal information butthey overlook the model's generalization ability on new datasets. Introducingnew datasets may affect the feature space of the original dataset, leading tocatastrophic forgetting. Meanwhile, large model parameters can significantlyimpact training performance. To address these limitations, we introduce a noveltask called Text-Audio Incremental Learning (TAIL) task for text-audioretrieval, and propose a new method, PTAT, Prompt Tuning for Audio-Textincremental learning. This method utilizes prompt tuning to optimize the modelparameters while incorporating an audio-text similarity and featuredistillation module to effectively mitigate catastrophic forgetting. Webenchmark our method and previous incremental learning methods on AudioCaps,Clotho, BBC Sound Effects and Audioset datasets, and our method outperformsprevious methods significantly, particularly demonstrating stronger resistanceto forgetting on older datasets. Compared to the full-parameters Finetune(Sequential) method, our model only requires 2.42\% of its parameters,achieving 4.46\% higher performance.

【6】 VoiceGRPO: Modern MoE Transformers with Group Relative Policy  Optimization GRPO for AI Voice Health Care Applications on Voice Pathology  Detection
标题:SecureGRPO:具有团体相对政策优化的现代MoE部TransformerGRPO,用于语音病理检测方面的人工智能语音医疗保健应用
链接:https://arxiv.org/abs/2503.03797
作者:Enkhtogtokh Togootogtokh,  Christian Klasen
摘要:本研究介绍一种新的人工智能技术,作为混合专家Transformers与组相对策略优化(GRPO)的语音健康护理应用程序的语音病理检测。通过架构创新,我们采用了受强化学习启发的高级训练范式,即近端策略优化(PPO)和分组正则化策略优化(GRPO),以提高模型的稳定性和性能。在综合生成的语音病理数据集上进行的实验表明,与传统方法相比,我们提出的模型显着提高了诊断准确性,F1评分和ROC-AUC。这些发现强调了将Transformer架构与新的训练策略相结合的潜力,以推进自动语音病理检测,并最终有助于更有效的医疗服务。我们用来训练和评估模型的代码可以在https://github.com/enkhtogtokh/voicegrpo上找到
摘要:This research introduces a novel AI techniques as Mixture-of-ExpertsTransformers with Group Relative Policy Optimization (GRPO) for voice healthcare applications on voice pathology detection. With the architecturalinnovations, we adopt advanced training paradigms inspired by reinforcementlearning, namely Proximal Policy Optimization (PPO) and Group-wise RegularizedPolicy Optimization (GRPO), to enhance model stability and performance.Experiments conducted on a synthetically generated voice pathology datasetdemonstrate that our proposed models significantly improve diagnostic accuracy,F1 score, and ROC-AUC compared to conventional approaches. These findingsunderscore the potential of integrating transformer architectures with noveltraining strategies to advance automated voice pathology detection andultimately contribute to more effective healthcare delivery. The code we usedto train and evaluate our models is available athttps://github.com/enkhtogtokh/voicegrpo

【7】 Efficient Finetuning for Dimensional Speech Emotion Recognition in the  Age of Transformers
标题:Transformer时代二维语音情感识别的高效微调
链接:https://arxiv.org/abs/2503.03756
作者:Aneesha Sampath,  James Tavernor,  Emily Mower Provost
备注:ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
摘要:准确的语音情感识别对于开发面向人类的系统至关重要。最近的进展包括微调大型的预训练Transformer模型,如Wav2Vec 2.0。然而,微调过程需要大量的计算资源,包括高内存GPU和大量的处理时间。随着对精确情感识别的需求持续增长,需要有效的微调方法来减少计算负担。我们的研究重点是维度情绪识别,预测属性,如激活(平静到兴奋)和效价(消极到积极)。我们提出了各种微调技术,包括全微调,部分微调的Transformer层,微调与混合精度,部分微调与缓存,低秩自适应(LoRA)的Wav2Vec 2.0基础模型。我们发现,混合精度的部分微调实现了与完全微调相当的性能,同时将训练速度提高了67%。缓存中间表示进一步提高了效率,产生了88%的加速和71%的可学习参数减少。我们建议以混合精度微调最后三个Transformer层,以平衡性能和训练效率,并添加中间表示缓存,以实现最佳速度和最小性能折衷。这些发现降低了微调语音情感识别系统的障碍,使更广泛的研究人员和从业者更容易获得准确的情感识别。
摘要:Accurate speech emotion recognition is essential for developing human-facingsystems. Recent advancements have included finetuning large, pretrainedtransformer models like Wav2Vec 2.0. However, the finetuning process requiressubstantial computational resources, including high-memory GPUs and significantprocessing time. As the demand for accurate emotion recognition continues togrow, efficient finetuning approaches are needed to reduce the computationalburden. Our study focuses on dimensional emotion recognition, predictingattributes such as activation (calm to excited) and valence (negative topositive). We present various finetuning techniques, including full finetuning,partial finetuning of transformer layers, finetuning with mixed precision,partial finetuning with caching, and low-rank adaptation (LoRA) on the Wav2Vec2.0 base model. We find that partial finetuning with mixed precision achievesperformance comparable to full finetuning while increasing training speed by67%. Caching intermediate representations further boosts efficiency, yieldingan 88% speedup and a 71% reduction in learnable parameters. We recommendfinetuning the final three transformer layers in mixed precision to balanceperformance and training efficiency, and adding intermediate representationcaching for optimal speed with minimal performance trade-offs. These findingslower the barriers to finetuning speech emotion recognition systems, makingaccurate emotion recognition more accessible to a broader range of researchersand practitioners.

机器翻译由腾讯交互翻译提供,仅供参考