今日论文合集:cs.SD语音8篇,eess.AS音频处理9篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Simultaneous Interpretation Corpus Construction by Large Language Models  in Distant Language Pair
标题:远程语言对中大型语言模型的同声传译数据库构建
链接:https://arxiv.org/abs/2404.12299
作者:Yusuke Sakai,Mana Makinae,Hidetaka Kamigaito,Taro Watanabe
备注:23 pages, 9 figures
摘要:在同声机器翻译(SiMT)系统中,使用同声传译(SI)语料库进行训练是实现高质量低延迟系统的有效方法。然而,由于注释者能力的限制,策划这样的语料库是非常具有挑战性的,因此,现有的SI语料库是有限的。因此,我们提出了一种方法,将现有的语音翻译语料库转换为解释风格的数据,保持原来的词序,并使用大型语言模型(LLM-SI-Corpus)保留整个源内容。我们证明了使用LLM-SI语料库在文本到文本和语音到文本设置中微调SiMT模型可以减少延迟,同时保持与离线数据集训练的模型相同的质量水平。LLM-SI-Corpus可以在\url{https://github.com/yusuke1997/LLM-SI-Corpus}上找到。
摘要:In Simultaneous Machine Translation (SiMT) systems, training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency systems. However, it is very challenging to curate such a corpus due to limitations in the abilities of annotators, and hence, existing SI corpora are limited. Therefore, we propose a method to convert existing speech translation corpora into interpretation-style data, maintaining the original word order and preserving the entire source content using Large Language Models (LLM-SI-Corpus). We demonstrate that fine-tuning SiMT models in text-to-text and speech-to-text settings with the LLM-SI-Corpus reduces latencies while maintaining the same level of quality as the models trained with offline datasets. The LLM-SI-Corpus is available at \url{https://github.com/yusuke1997/LLM-SI-Corpus}.

【2】 Dynamic Modality and View Selection for Multimodal Emotion Recognition  with Missing Modalities
标题:缺失模式的多模式情感识别的动态模式和视图选择
链接:https://arxiv.org/abs/2404.12251
作者:Luciana Trinkaus Menon,Luiz Carlos Ribeiro Neduziak,Jean Paul Barddal,Alessandro Lameiras Koerich,Alceu de Souza Britto Jr
备注:15 pages
摘要:人类情感的研究,传统上是心理学和神经科学等领域的基石,已经受到人工智能(AI)的深刻影响。多渠道,如语音(声音)和面部表情(图像),在理解人类情感方面至关重要。然而,人工智能在多模态情感识别(MER)方面的旅程面临着巨大的技术挑战。一个重要的障碍是人工智能模型如何管理特定模态的缺失-这在现实世界中经常发生。本研究的中心重点是评估两种策略的性能和弹性时,面对缺乏一种模态:一种新的多模态动态模态和视图选择和交叉注意机制。在RECOLA数据集上的结果表明,基于动态选择的方法是一种很有前途的MER方法。在缺少模态的情况下,所有基于动态选择的方法都优于基线。研究结论强调了音频和视频模态在情感预测中的复杂相互作用,展示了动态选择方法在处理缺失模态时的适应性。
摘要:The study of human emotions, traditionally a cornerstone in fields like psychology and neuroscience, has been profoundly impacted by the advent of artificial intelligence (AI). Multiple channels, such as speech (voice) and facial expressions (image), are crucial in understanding human emotions. However, AI's journey in multimodal emotion recognition (MER) is marked by substantial technical challenges. One significant hurdle is how AI models manage the absence of a particular modality - a frequent occurrence in real-world situations. This study's central focus is assessing the performance and resilience of two strategies when confronted with the lack of one modality: a novel multimodal dynamic modality and view selection and a cross-attention mechanism. Results on the RECOLA dataset show that dynamic selection-based methods are a promising approach for MER. In the missing modalities scenarios, all dynamic selection-based methods outperformed the baseline. The study concludes by emphasizing the intricate interplay between audio and video modalities in emotion prediction, showcasing the adaptability of dynamic selection methods in handling missing modalities.

【3】 Enhancing Suicide Risk Assessment: A Speech-Based Automated Approach in  Emergency Medicine
标题:加强自杀风险评估:急诊医学中基于语音的自动化方法
链接:https://arxiv.org/abs/2404.12132
作者:Shahin Amiriparian,Maurice Gerczuk,Justina Lutz,Wolfgang Strube,Irina Papazova,Alkomiet Hasan,Alexander Kathan,Björn W. Schuller
摘要:在急诊科,有自杀倾向风险的病人迟迟得不到专门的精神病评估和护理,这在及时干预方面造成了明显的差距,妨碍了在危急情况下提供充分的心理健康支持。为了解决这个问题,我们提出了一种非侵入性的,基于语音的自动自杀风险评估方法。在我们的研究中,我们从20美元的患者中收集了一个新的语音记录数据集,从中提取了三组特征,包括wav 2 vec,可解释的语音和声学特征以及基于深度学习的频谱表示。我们继续进行二进制分类,以评估自杀风险在一个留一个主题的方式。我们最有效的语音模型达到了66.2\,\%$的平衡准确率。此外,我们表明,将我们的语音模型与一系列患者的元数据(如自杀企图或枪支使用史)相结合,可以改善整体结果。元数据集成产生了平衡的准确性为94.4\,\%$,标志着28.2\,\%$的绝对改善,证明了我们提出的方法在急诊医学自动自杀风险评估的有效性。
摘要:The delayed access to specialized psychiatric assessments and care for patients at risk of suicidal tendencies in emergency departments creates a notable gap in timely intervention, hindering the provision of adequate mental health support during critical situations. To address this, we present a non-invasive, speech-based approach for automatic suicide risk assessment. For our study, we have collected a novel dataset of speech recordings from $20$ patients from which we extract three sets of features, including wav2vec, interpretable speech and acoustic features, and deep learning-based spectral representations. We proceed by conducting a binary classification to assess suicide risk in a leave-one-subject-out fashion. Our most effective speech model achieves a balanced accuracy of $66.2\,\%$. Moreover, we show that integrating our speech model with a series of patients' metadata, such as the history of suicide attempts or access to firearms, improves the overall result. The metadata integration yields a balanced accuracy of $94.4\,\%$, marking an absolute improvement of $28.2\,\%$, demonstrating the efficacy of our proposed approaches for automatic suicide risk assessment in emergency medicine.

【4】 TIMIT Speaker Profiling: A Comparison of Multi-task learning and  Single-task learning Approaches
标题:TIMIT演讲者剖析:多任务学习和单任务学习方法的比较
链接:https://arxiv.org/abs/2404.12077
作者:Rong Wang,Kun Sun
摘要:本研究采用深度学习技术在TIMIT数据集上探索四个说话人分析任务,即性别分类、口音分类、年龄估计和说话人识别,突出了多任务学习与单任务模型相比的潜力和挑战。这项研究的动机是双重的:首先,在说话人分析的背景下,经验性地评估多任务学习相对于单任务模型的优点和缺点;其次,强调熟练的特征工程对说话人识别任务的重要性。研究结果揭示了口音分类的挑战,多任务学习被发现对类似复杂性的任务有利。非顺序特征对于说话人识别是有利的,但是顺序特征可以作为复杂模型的起点。该研究强调了对深度学习模型进行细致实验和参数调整的必要性。
摘要:This study employs deep learning techniques to explore four speaker profiling tasks on the TIMIT dataset, namely gender classification, accent classification, age estimation, and speaker identification, highlighting the potential and challenges of multi-task learning versus single-task models. The motivation for this research is twofold: firstly, to empirically assess the advantages and drawbacks of multi-task learning over single-task models in the context of speaker profiling; secondly, to emphasize the undiminished significance of skillful feature engineering for speaker recognition tasks. The findings reveal challenges in accent classification, and multi-task learning is found advantageous for tasks of similar complexity. Non-sequential features are favored for speaker recognition, but sequential ones can serve as starting points for complex models. The study underscores the necessity of meticulous experimentation and parameter tuning for deep learning models.

【5】 MIDGET: Music Conditioned 3D Dance Generation
标题:MIDGET:音乐条件3D舞蹈一代链接:https://arxiv.org/abs/2404.12062
作者:Jinwu Wang,Wei Mao,Miaomiao Liu
备注:None
摘要:本文介绍了一种基于音乐的三维舞蹈生成模型MIDGET,该模型基于舞蹈运动矢量量化变分自动编码器(VQ—VAE)模型和运动生成预训练(GPT)模型,能够生成与音乐节奏相匹配的、充满活力的、高质量的舞蹈。为了解决该领域的挑战,我们引入了三个新的组件:1)基于Motion VQ—VAE模型的预训练记忆码本,用于存储不同的人体姿势代码,2)采用Motion GPT模型来生成具有音乐和运动编码器的姿势代码,3)用于音乐特征提取的简单框架。我们与现有的最先进的模型进行比较,并在AIST ++上进行消融实验,AIST ++是最大的公开可用的音乐舞蹈数据集。实验表明,我们提出的框架实现了最先进的性能的运动质量和它的对齐与音乐。
摘要:In this paper, we introduce a MusIc conditioned 3D Dance GEneraTion model, named MIDGET based on Dance motion Vector Quantised Variational AutoEncoder (VQ-VAE) model and Motion Generative Pre-Training (GPT) model to generate vibrant and highquality dances that match the music rhythm. To tackle challenges in the field, we introduce three new components: 1) a pre-trained memory codebook based on the Motion VQ-VAE model to store different human pose codes, 2) employing Motion GPT model to generate pose codes with music and motion Encoders, 3) a simple framework for music feature extraction. We compare with existing state-of-the-art models and perform ablation experiments on AIST++, the largest publicly available music-dance dataset. Experiments demonstrate that our proposed framework achieves state-of-the-art performance on motion quality and its alignment with the music.


【6】 Large Language Models: From Notes to Musical Form
标题:大型语言模型:从音符到音乐形式
链接:https://arxiv.org/abs/2404.11976
作者:Lilac Atassi
摘要:虽然基于学习的自动音乐生成方法的许多主题正在积极研究中,但音乐形式研究不足。特别是,最近基于深度学习模型的方法生成的音乐在最大的时间尺度上缺乏任何结构。在实践中,这种模型生成的超过一分钟的音乐要么是令人不快的重复,要么是没有方向的。本文采用了一种新的音乐生成模型,提出了一种新的生成具有形式的音乐的方法。实验结果表明,该方法可以产生2.5分钟长的音乐,被认为是作为音乐用于训练的模型一样愉快。本文首先回顾了一种基于语言模型的音乐生成方法(Transformer架构)。我们讨论了为什么学习音乐形式,这样的模型是不可行的。然后,我们讨论了我们提出的方法和实验。
摘要:While many topics of the learning-based approach to automated music generation are under active research, musical form is under-researched. In particular, recent methods based on deep learning models generate music that, at the largest time scale, lacks any structure. In practice, music longer than one minute generated by such models is either unpleasantly repetitive or directionless. Adapting a recent music generation model, this paper proposes a novel method to generate music with form. The experimental results show that the proposed method can generate 2.5-minute-long music that is considered as pleasant as the music used to train the model. The paper first reviews a recent music generation method based on language models (transformer architecture). We discuss why learning musical form by such models is infeasible. Then we discuss our proposed method and the experiments.

【7】 HyDiscGAN: A Hybrid Distributed cGAN for Audio-Visual Privacy  Preservation in Multimodal Sentiment Analysis
标题:HyDiscGAN:一种混合分布式cGAN,用于多模式情绪分析中的视听隐私保护
链接:https://arxiv.org/abs/2404.11938
作者:Zhuojia Wu,Qi Zhang,Duoqian Miao,Kun Yi,Wei Fan,Liang Hu备注:13 pages, IJCAI-2024
摘要:多模态情感分析(MSA)旨在识别多模态视频内容中说话者的情感倾向,引起了人们对多模态数据(如声纹和面部图像)相关隐私风险的严重关注。最近的分布式协作学习已被证实是一种有效的模式,在多模态任务的隐私保护。然而,他们往往忽视了不同模式之间的隐私差异,努力在性能和隐私保护之间取得平衡。因此,它提出了一个有趣的问题,最大限度地利用多式联运,以提高性能,同时保护必要的方式。本文形成了模态指定的第一次尝试(即,音频和视频)隐私保护。我们提出了一种新的混合分布式跨模态cGAN框架(HyDiscGAN),该框架学习多模态对齐以生成以可共享的去识别文本数据为条件的虚假音频和视觉特征。其目的是利用假特征来近似真实的音频和视频内容,以保证隐私保护,同时有效地提高性能。大量的实验表明,与最先进的MSA模型相比,HyDiscGAN可以在保护隐私的同时实现卓越或有竞争力的性能。
摘要:Multimodal Sentiment Analysis (MSA) aims to identify speakers' sentiment tendencies in multimodal video content, raising serious concerns about privacy risks associated with multimodal data, such as voiceprints and facial images. Recent distributed collaborative learning has been verified as an effective paradigm for privacy preservation in multimodal tasks. However, they often overlook the privacy distinctions among different modalities, struggling to strike a balance between performance and privacy preservation. Consequently, it poses an intriguing question of maximizing multimodal utilization to improve performance while simultaneously protecting necessary modalities. This paper forms the first attempt at modality-specified (i.e., audio and visual) privacy preservation in MSA tasks. We propose a novel Hybrid Distributed cross-modality cGAN framework (HyDiscGAN), which learns multimodality alignment to generate fake audio and visual features conditioned on shareable de-identified textual data. The objective is to leverage the fake features to approximate real audio and visual content to guarantee privacy preservation while effectively enhancing performance. Extensive experiments show that compared with the state-of-the-art MSA model, HyDiscGAN can achieve superior or competitive performance while preserving privacy.


【8】 Advancing Speech Translation: A Corpus of Mandarin-English  Conversational Telephone Speech标题:推进语音翻译:普通英语对话电话语音数据库
链接:https://arxiv.org/abs/2404.11619
作者:Shannon Wotherspoon,William Hartmann,Matthew Snover
备注:2 pages
摘要:本文介绍了一套英语翻译的一个123小时的子集的CallHome普通话中文数据和香港科技大学的普通话电话语音数据的语音翻译任务。配对的源语言语音和目标语言文本对于训练端到端语音翻译系统是必不可少的,并且相对于在更广泛可用的文本数据集上进行训练,还可以为级联系统提供实质性的性能改进。我们证明,微调通用翻译模型到我们的普通话-英语对话电话语音训练集,将目标域BLEU提高了8个点以上,突出了匹配训练数据的重要性。
摘要:This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language speech and target-language text is essential for training end-to-end speech translation systems and can provide substantial performance improvements for cascaded systems as well, relative to training on more widely available text data sets. We demonstrate that fine-tuning a general-purpose translation model to our Mandarin-English conversational telephone speech training set improves target-domain BLEU by more than 8 points, highlighting the importance of matched training data.

eess.AS音频处理

【1】Efficient High-Performance Bark-Scale Neural Network for Residual Echo  and Noise Suppression
标题:用于残余回声和噪音抑制的高效高性能树皮规模神经网络
链接:https://arxiv.org/abs/2404.11621
作者:Ernst Seidel,Pejman Mowlaee,Tim Fingscheidt
备注:accepted to ICASSP 2024; 5 pages, 3 figures
摘要:近年来,将神经网络引入语音增强领域带来了显著的进步。然而,许多提出的方法是相当苛刻的计算复杂性和内存占用方面。对于专用通信设备中的应用,如免提电话,免提汽车系统或智能手机,效率与性能一起发挥着重要作用。在这种情况下,我们提出了一个有效的,高性能的混合联合声学回声控制和噪声抑制系统,其中我们的主要贡献是后置滤波器NN,执行噪声和残留回声抑制。近端语音的保存被改进为用于NN后滤波器的Bark尺度听觉滤波器组。所提出的混合方法是基准与国家的最先进的方法和ICASSP 2023 AEC挑战盲测试集上证明了其有效性。我们证明,它提供了高质量的近端语音保存在双通话和近端语音条件。与此同时,它能够有效地消除回声泄漏,实现与端到端DeepVQE-S等已经很小的最先进模型相当的性能,同时只需要大约10%的计算复杂度。这使得它很容易在扬声器电话设备上实时实现。
摘要:In recent years, the introduction of neural networks (NNs) into the field of speech enhancement has brought significant improvements. However, many of the proposed methods are quite demanding in terms of computational complexity and memory footprint. For the application in dedicated communication devices, such as speakerphones, hands-free car systems, or smartphones, efficiency plays a major role along with performance. In this context, we present an efficient, high-performance hybrid joint acoustic echo control and noise suppression system, whereby our main contribution is the postfilter NN, performing both noise and residual echo suppression. The preservation of nearend speech is improved by a Bark-scale auditory filterbank for the NN postfilter. The proposed hybrid method is benchmarked with state-of-the-art methods and its effectiveness is demonstrated on the ICASSP 2023 AEC Challenge blind test set. We demonstrate that it offers high-quality nearend speech preservation during both double-talk and nearend speech conditions. At the same time, it is capable of efficient removal of echo leaks, achieving a comparable performance to already small state-of-the-art models such as the end-to-end DeepVQE-S, while requiring only around 10 % of its computational complexity. This makes it easily realtime implementable on a speakerphone device.

【2】 Advancing Speech Translation: A Corpus of Mandarin-English  Conversational Telephone Speech标题:推进语音翻译:普通英语对话电话语音数据库
链接:https://arxiv.org/abs/2404.11619
作者:Shannon Wotherspoon,William Hartmann,Matthew Snover
备注:2 pages
摘要:本文介绍了一套英语翻译的一个123小时的子集的CallHome普通话中文数据和香港科技大学的普通话电话语音数据的语音翻译任务。配对的源语言语音和目标语言文本对于训练端到端语音翻译系统是必不可少的,并且相对于在更广泛可用的文本数据集上进行训练,还可以为级联系统提供实质性的性能改进。我们证明,微调通用翻译模型到我们的普通话-英语对话电话语音训练集,将目标域BLEU提高了8个点以上,突出了匹配训练数据的重要性。
摘要:This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language speech and target-language text is essential for training end-to-end speech translation systems and can provide substantial performance improvements for cascaded systems as well, relative to training on more widely available text data sets. We demonstrate that fine-tuning a general-purpose translation model to our Mandarin-English conversational telephone speech training set improves target-domain BLEU by more than 8 points, highlighting the importance of matched training data.


【3】 Simultaneous Interpretation Corpus Construction by Large Language Models  in Distant Language Pair
标题:远程语言对中大型语言模型的同声传译数据库构建
链接:https://arxiv.org/abs/2404.12299
作者:Yusuke Sakai,Mana Makinae,Hidetaka Kamigaito,Taro Watanabe
备注:23 pages, 9 figures
摘要:在同声机器翻译(SiMT)系统中,使用同声传译(SI)语料库进行训练是实现高质量低延迟系统的有效方法。然而,由于注释者能力的限制,策划这样的语料库是非常具有挑战性的,因此,现有的SI语料库是有限的。因此,我们提出了一种方法,将现有的语音翻译语料库转换为解释风格的数据,保持原来的词序,并使用大型语言模型(LLM-SI-Corpus)保留整个源内容。我们证明了使用LLM-SI语料库在文本到文本和语音到文本设置中微调SiMT模型可以减少延迟,同时保持与离线数据集训练的模型相同的质量水平。LLM-SI-Corpus可以在\url{https://github.com/yusuke1997/LLM-SI-Corpus}上找到。
摘要:In Simultaneous Machine Translation (SiMT) systems, training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency systems. However, it is very challenging to curate such a corpus due to limitations in the abilities of annotators, and hence, existing SI corpora are limited. Therefore, we propose a method to convert existing speech translation corpora into interpretation-style data, maintaining the original word order and preserving the entire source content using Large Language Models (LLM-SI-Corpus). We demonstrate that fine-tuning SiMT models in text-to-text and speech-to-text settings with the LLM-SI-Corpus reduces latencies while maintaining the same level of quality as the models trained with offline datasets. The LLM-SI-Corpus is available at \url{https://github.com/yusuke1997/LLM-SI-Corpus}.


【4】 Dynamic Modality and View Selection for Multimodal Emotion Recognition  with Missing Modalities
标题:缺失模式的多模式情感识别的动态模式和视图选择
链接:https://arxiv.org/abs/2404.12251
作者:Luciana Trinkaus Menon,Luiz Carlos Ribeiro Neduziak,Jean Paul Barddal,Alessandro Lameiras Koerich,Alceu de Souza Britto Jr
备注:15 pages
摘要:人类情感的研究,传统上是心理学和神经科学等领域的基石,已经受到人工智能(AI)的深刻影响。多渠道,如语音(声音)和面部表情(图像),在理解人类情感方面至关重要。然而,人工智能在多模态情感识别(MER)方面的旅程面临着巨大的技术挑战。一个重要的障碍是人工智能模型如何管理特定模态的缺失-这在现实世界中经常发生。本研究的中心重点是评估两种策略的性能和弹性时,面对缺乏一种模态:一种新的多模态动态模态和视图选择和交叉注意机制。在RECOLA数据集上的结果表明,基于动态选择的方法是一种很有前途的MER方法。在缺少模态的情况下,所有基于动态选择的方法都优于基线。研究结论强调了音频和视频模态在情感预测中的复杂相互作用,展示了动态选择方法在处理缺失模态时的适应性。
摘要:The study of human emotions, traditionally a cornerstone in fields like psychology and neuroscience, has been profoundly impacted by the advent of artificial intelligence (AI). Multiple channels, such as speech (voice) and facial expressions (image), are crucial in understanding human emotions. However, AI's journey in multimodal emotion recognition (MER) is marked by substantial technical challenges. One significant hurdle is how AI models manage the absence of a particular modality - a frequent occurrence in real-world situations. This study's central focus is assessing the performance and resilience of two strategies when confronted with the lack of one modality: a novel multimodal dynamic modality and view selection and a cross-attention mechanism. Results on the RECOLA dataset show that dynamic selection-based methods are a promising approach for MER. In the missing modalities scenarios, all dynamic selection-based methods outperformed the baseline. The study concludes by emphasizing the intricate interplay between audio and video modalities in emotion prediction, showcasing the adaptability of dynamic selection methods in handling missing modalities.

【5】 Enhancing Suicide Risk Assessment: A Speech-Based Automated Approach in  Emergency Medicine
标题:加强自杀风险评估:急诊医学中基于语音的自动化方法
链接:https://arxiv.org/abs/2404.12132
作者:Shahin Amiriparian,Maurice Gerczuk,Justina Lutz,Wolfgang Strube,Irina Papazova,Alkomiet Hasan,Alexander Kathan,Björn W. Schuller
摘要:在急诊科,有自杀倾向风险的病人迟迟得不到专门的精神病评估和护理,这在及时干预方面造成了明显的差距,妨碍了在危急情况下提供充分的心理健康支持。为了解决这个问题,我们提出了一种非侵入性的,基于语音的自动自杀风险评估方法。在我们的研究中,我们从20美元的患者中收集了一个新的语音记录数据集,从中提取了三组特征,包括wav 2 vec,可解释的语音和声学特征以及基于深度学习的频谱表示。我们继续进行二进制分类,以评估自杀风险在一个留一个主题的方式。我们最有效的语音模型达到了66.2\,\%$的平衡准确率。此外,我们表明,将我们的语音模型与一系列患者的元数据(如自杀企图或枪支使用史)相结合,可以改善整体结果。元数据集成产生了平衡的准确性为94.4\,\%$,标志着28.2\,\%$的绝对改善,证明了我们提出的方法在急诊医学自动自杀风险评估的有效性。
摘要:The delayed access to specialized psychiatric assessments and care for patients at risk of suicidal tendencies in emergency departments creates a notable gap in timely intervention, hindering the provision of adequate mental health support during critical situations. To address this, we present a non-invasive, speech-based approach for automatic suicide risk assessment. For our study, we have collected a novel dataset of speech recordings from $20$ patients from which we extract three sets of features, including wav2vec, interpretable speech and acoustic features, and deep learning-based spectral representations. We proceed by conducting a binary classification to assess suicide risk in a leave-one-subject-out fashion. Our most effective speech model achieves a balanced accuracy of $66.2\,\%$. Moreover, we show that integrating our speech model with a series of patients' metadata, such as the history of suicide attempts or access to firearms, improves the overall result. The metadata integration yields a balanced accuracy of $94.4\,\%$, marking an absolute improvement of $28.2\,\%$, demonstrating the efficacy of our proposed approaches for automatic suicide risk assessment in emergency medicine.

【6】 TIMIT Speaker Profiling: A Comparison of Multi-task learning and  Single-task learning Approaches
标题:TIMIT演讲者剖析:多任务学习和单任务学习方法的比较
链接:https://arxiv.org/abs/2404.12077
作者:Rong Wang,Kun Sun
摘要:本研究采用深度学习技术在TIMIT数据集上探索四个说话人分析任务,即性别分类、口音分类、年龄估计和说话人识别,突出了多任务学习与单任务模型相比的潜力和挑战。这项研究的动机是双重的:首先,在说话人分析的背景下,经验性地评估多任务学习相对于单任务模型的优点和缺点;其次,强调熟练的特征工程对说话人识别任务的重要性。研究结果揭示了口音分类的挑战,多任务学习被发现对类似复杂性的任务有利。非顺序特征对于说话人识别是有利的,但是顺序特征可以作为复杂模型的起点。该研究强调了对深度学习模型进行细致实验和参数调整的必要性。
摘要:This study employs deep learning techniques to explore four speaker profiling tasks on the TIMIT dataset, namely gender classification, accent classification, age estimation, and speaker identification, highlighting the potential and challenges of multi-task learning versus single-task models. The motivation for this research is twofold: firstly, to empirically assess the advantages and drawbacks of multi-task learning over single-task models in the context of speaker profiling; secondly, to emphasize the undiminished significance of skillful feature engineering for speaker recognition tasks. The findings reveal challenges in accent classification, and multi-task learning is found advantageous for tasks of similar complexity. Non-sequential features are favored for speaker recognition, but sequential ones can serve as starting points for complex models. The study underscores the necessity of meticulous experimentation and parameter tuning for deep learning models.

【7】 MIDGET: Music Conditioned 3D Dance Generation
标题:MIDGET:音乐条件3D舞蹈一代
链接:https://arxiv.org/abs/2404.12062
作者:Jinwu Wang,Wei Mao,Miaomiao Liu
备注:None
摘要:本文介绍了一种基于音乐的三维舞蹈生成模型MIDGET,该模型基于舞蹈运动矢量量化变分自动编码器(VQ-VAE)模型和运动生成预训练(GPT)模型,能够生成与音乐节奏相匹配的、充满活力的、高质量的舞蹈。为了解决该领域的挑战,我们引入了三个新的组件:1)基于Motion VQ-VAE模型的预训练记忆码本,用于存储不同的人体姿势代码,2)采用Motion GPT模型来生成具有音乐和运动编码器的姿势代码,3)用于音乐特征提取的简单框架。我们与现有的最先进的模型进行比较,并在AIST++上进行消融实验,AIST++是最大的公开可用的音乐舞蹈数据集。实验表明,我们提出的框架实现了最先进的性能的运动质量和它的对齐与音乐。
摘要:In this paper, we introduce a MusIc conditioned 3D Dance GEneraTion model, named MIDGET based on Dance motion Vector Quantised Variational AutoEncoder (VQ-VAE) model and Motion Generative Pre-Training (GPT) model to generate vibrant and highquality dances that match the music rhythm. To tackle challenges in the field, we introduce three new components: 1) a pre-trained memory codebook based on the Motion VQ-VAE model to store different human pose codes, 2) employing Motion GPT model to generate pose codes with music and motion Encoders, 3) a simple framework for music feature extraction. We compare with existing state-of-the-art models and perform ablation experiments on AIST++, the largest publicly available music-dance dataset. Experiments demonstrate that our proposed framework achieves state-of-the-art performance on motion quality and its alignment with the music.


【8】 Large Language Models: From Notes to Musical Form
标题:大型语言模型:从音符到音乐形式
链接:https://arxiv.org/abs/2404.11976
作者:Lilac Atassi
摘要:虽然基于学习的自动音乐生成方法的许多主题正在积极研究中,但音乐形式研究不足。特别是,最近基于深度学习模型的方法生成的音乐在最大的时间尺度上缺乏任何结构。在实践中,这种模型生成的超过一分钟的音乐要么是令人不快的重复,要么是没有方向的。本文采用了一种新的音乐生成模型,提出了一种新的生成具有形式的音乐的方法。实验结果表明,该方法可以产生2.5分钟长的音乐,被认为是作为音乐用于训练的模型一样愉快。本文首先回顾了一种基于语言模型的音乐生成方法(Transformer架构)。我们讨论了为什么学习音乐形式,这样的模型是不可行的。然后,我们讨论了我们提出的方法和实验。
摘要:While many topics of the learning-based approach to automated music generation are under active research, musical form is under-researched. In particular, recent methods based on deep learning models generate music that, at the largest time scale, lacks any structure. In practice, music longer than one minute generated by such models is either unpleasantly repetitive or directionless. Adapting a recent music generation model, this paper proposes a novel method to generate music with form. The experimental results show that the proposed method can generate 2.5-minute-long music that is considered as pleasant as the music used to train the model. The paper first reviews a recent music generation method based on language models (transformer architecture). We discuss why learning musical form by such models is infeasible. Then we discuss our proposed method and the experiments.


【9】 HyDiscGAN: A Hybrid Distributed cGAN for Audio-Visual Privacy  Preservation in Multimodal Sentiment Analysis
标题:HyDiscGAN:一种混合分布式cGAN,用于多模式情绪分析中的视听隐私保护
链接:https://arxiv.org/abs/2404.11938
作者:Zhuojia Wu,Qi Zhang,Duoqian Miao,Kun Yi,Wei Fan,Liang Hu
备注:13 pages, IJCAI-2024
摘要:多模态情感分析(MSA)旨在识别多模态视频内容中说话者的情感倾向,引起了人们对多模态数据(如声纹和面部图像)相关隐私风险的严重关注。最近的分布式协作学习已被证实是一种有效的模式,在多模态任务的隐私保护。然而,他们往往忽视了不同模式之间的隐私差异,努力在性能和隐私保护之间取得平衡。因此,它提出了一个有趣的问题,最大限度地利用多式联运,以提高性能,同时保护必要的方式。本文形成了模态指定的第一次尝试(即,音频和视频)隐私保护。我们提出了一种新的混合分布式跨模态cGAN框架(HyDiscGAN),该框架学习多模态对齐以生成以可共享的去识别文本数据为条件的虚假音频和视觉特征。其目的是利用假特征来近似真实的音频和视频内容,以保证隐私保护,同时有效地提高性能。大量的实验表明,与最先进的MSA模型相比,HyDiscGAN可以在保护隐私的同时实现卓越或有竞争力的性能。
摘要:Multimodal Sentiment Analysis (MSA) aims to identify speakers' sentiment tendencies in multimodal video content, raising serious concerns about privacy risks associated with multimodal data, such as voiceprints and facial images. Recent distributed collaborative learning has been verified as an effective paradigm for privacy preservation in multimodal tasks. However, they often overlook the privacy distinctions among different modalities, struggling to strike a balance between performance and privacy preservation. Consequently, it poses an intriguing question of maximizing multimodal utilization to improve performance while simultaneously protecting necessary modalities. This paper forms the first attempt at modality-specified (i.e., audio and visual) privacy preservation in MSA tasks. We propose a novel Hybrid Distributed cross-modality cGAN framework (HyDiscGAN), which learns multimodality alignment to generate fake audio and visual features conditioned on shareable de-identified textual data. The objective is to leverage the fake features to approximate real audio and visual content to guarantee privacy preservation while effectively enhancing performance. Extensive experiments show that compared with the state-of-the-art MSA model, HyDiscGAN can achieve superior or competitive performance while preserving privacy.


机器翻译由腾讯交互翻译提供,仅供参考