本文经arXiv每日学术速递授权转载
【1】Source Separation & Automatic Transcription for Music
链接:https://arxiv.org/abs/2412.06703
摘要:源分离是在多个声音的听觉混合物中分离单个声音的过程[1],并且具有从语音增强和歌词转录[2]到音乐的数字音频制作的各种应用。此外,自动音乐转录(AMT)是将原始音乐音频转换为音乐家可以阅读的乐谱的过程[3]。从历史上看,这些任务面临着巨大的音频噪声,长时间的训练,以及由于版权限制而缺乏免费使用的数据等挑战。然而,深度学习的最新发展带来了新的有前途的方法来构建低失真的茎和从音频信号生成乐谱[4]。使用频谱图掩蔽,深度神经网络和MuseScore API,我们试图创建一个端到端的管道,允许初始音乐音频混合(例如... WAV文件),以被分离成乐器词干、转换成MIDI文件、以及转录成每个组成乐器的乐谱。
摘要:Source separation is the process of isolating individual sounds in anauditory mixture of multiple sounds [1], and has a variety of applicationsranging from speech enhancement and lyric transcription [2] to digital audioproduction for music. Furthermore, Automatic Music Transcription (AMT) is theprocess of converting raw music audio into sheet music that musicians can read[3]. Historically, these tasks have faced challenges such as significant audionoise, long training times, and lack of free-use data due to copyrightrestrictions. However, recent developments in deep learning have brought newpromising approaches to building low-distortion stems and generating sheetmusic from audio signals [4]. Using spectrogram masking, deep neural networks,and the MuseScore API, we attempt to create an end-to-end pipeline that allowsfor an initial music audio mixture (e.g...wav file) to be separated intoinstrument stems, converted into MIDI files, and transcribed into sheet musicfor each component instrument.
标题:MuMuMu-LLaMA:通过大型语言模型的多模式音乐理解和生成
链接:https://arxiv.org/abs/2412.06660
摘要:大型语言模型的研究在文本、语音、图像和视频方面取得了显著进展。然而,多模态音乐的理解和生成由于缺乏良好的注释数据集仍然未被充分探索。为了解决这个问题,我们引入了一个包含167.69小时多模态数据的数据集,包括文本、图像、视频和音乐注释。基于这个数据集,我们提出了MuMu-LLaMA,这是一个利用音乐,图像和视频的预训练编码器的模型。对于音乐生成,我们集成了AudioLDM 2和MusicGen。我们对四个任务的评估-音乐理解,文本到音乐生成,基于音乐编辑和多模态音乐生成-表明MuMu-LLaMA优于最先进的模型,显示了其多模态音乐应用的潜力。
摘要:Research on large language models has advanced significantly across text,speech, images, and videos. However, multi-modal music understanding andgeneration remain underexplored due to the lack of well-annotated datasets. Toaddress this, we introduce a dataset with 167.69 hours of multi-modal data,including text, images, videos, and music annotations. Based on this dataset,we propose MuMu-LLaMA, a model that leverages pre-trained encoders for music,images, and videos. For music generation, we integrate AudioLDM 2 and MusicGen.Our evaluation across four tasks--music understanding, text-to-musicgeneration, prompt-based music editing, and multi-modal musicgeneration--demonstrates that MuMu-LLaMA outperforms state-of-the-art models,showing its potential for multi-modal music applications.
链接:https://arxiv.org/abs/2412.06617
备注:Accepted for the NeurIPS 2024 Creative AI Track
摘要:“卧室制作人”的兴起使音乐创作民主化,同时也挑战制作人客观地评估他们的作品。为了解决这个问题,我们提出了AI TrackMate,这是一个基于LLM的音乐聊天机器人,旨在为音乐制作提供建设性的反馈。通过将LLM固有的音乐知识与直接的音轨分析相结合,AI TrackMate提供了特定于制作的见解,将其与纯文本方法区分开来。我们的框架集成了一个音乐分析模块,一个法学硕士可读的音乐报告和音乐制作导向的反馈指令,创建一个即插即用,与各种法学硕士兼容的免培训系统,并适应未来的发展。我们通过交互式Web界面展示了AI TrackMate的功能,并与音乐制作人进行了试点研究。通过将人工智能功能与独立制作人的需求联系起来,AI TrackMate提供了按需分析反馈,可能支持音乐制作的创作过程和技能发展。该系统解决了在不断发展的独立音乐制作领域对客观自我评估工具日益增长的需求。
摘要:The rise of "bedroom producers" has democratized music creation, whilechallenging producers to objectively evaluate their work. To address this, wepresent AI TrackMate, an LLM-based music chatbot designed to provideconstructive feedback on music productions. By combining LLMs' inherent musicalknowledge with direct audio track analysis, AI TrackMate offersproduction-specific insights, distinguishing it from text-only approaches. Ourframework integrates a Music Analysis Module, an LLM-Readable Music Report, andMusic Production-Oriented Feedback Instruction, creating a plug-and-play,training-free system compatible with various LLMs and adaptable to futureadvancements. We demonstrate AI TrackMate's capabilities through an interactiveweb interface and present findings from a pilot study with a music producer. Bybridging AI capabilities with the needs of independent producers, AI TrackMateoffers on-demand analytical feedback, potentially supporting the creativeprocess and skill development in music production. This system addresses thegrowing demand for objective self-assessment tools in the evolving landscape ofindependent music production.
标题:大型语言模型时代的可控语音合成:综述
链接:https://arxiv.org/abs/2412.06602
备注:A comprehensive survey on controllable TTS, 23 pages, 6 tables, 4 figures, 280 references
摘要:文本到语音(TTS),也称为语音合成,是一个重要的研究领域,旨在从文本生成听起来自然的人类语音。近年来,随着工业需求的不断增长,TTS技术已经从合成类人语音发展到实现可控语音生成。这包括对合成语音的各种属性(如情感、韵律、音色和持续时间)的细粒度控制。此外,深度学习的进步,如扩散和大型语言模型,在过去几年中显着增强了可控TTS。在本文中,我们进行了全面的调查可控TTS,涵盖的方法,从基本的控制技术,利用自然语言提示的方法,旨在提供一个清晰的了解目前的研究状况。我们研究了一般可控的TTS管道,挑战,模型架构和控制策略,提供了一个全面而清晰的分类现有的方法。此外,我们提供了数据集和评估指标的详细摘要,并阐明了可控TTS的应用和未来的发展方向。据我们所知,这份调查报告提供了第一个全面审查新兴的可控TTS方法,这可以作为一个有益的资源,为学术研究人员和行业从业者。
摘要:Text-to-speech (TTS), also known as speech synthesis, is a prominent researcharea that aims to generate natural-sounding human speech from text. Recently,with the increasing industrial demand, TTS technologies have evolved beyondsynthesizing human-like speech to enabling controllable speech generation. Thisincludes fine-grained control over various attributes of synthesized speechsuch as emotion, prosody, timbre, and duration. Besides, advancements in deeplearning, such as diffusion and large language models, have significantlyenhanced controllable TTS over the past several years. In this paper, weconduct a comprehensive survey of controllable TTS, covering approaches rangingfrom basic control techniques to methods utilizing natural language prompts,aiming to provide a clear understanding of the current state of research. Weexamine the general controllable TTS pipeline, challenges, model architectures,and control strategies, offering a comprehensive and clear taxonomy of existingmethods. Additionally, we provide a detailed summary of datasets and evaluationmetrics and shed some light on the applications and future directions ofcontrollable TTS. To the best of our knowledge, this survey paper provides thefirst comprehensive review of emerging controllable TTS methods, which canserve as a beneficial resource for both academic researchers and industrypractitioners.
标题:语音:情感丰富且上下文详细的语音注释库
链接:https://arxiv.org/abs/2412.06581
备注:4 pages, 1 figure. To appear in the Proceedings of the International Symposium on Chinese Spoken Language Processing, 7-10 November 2024, Beijing, China
摘要:文语转换(TTS)技术的进步显著提高了生成语音的质量,使其与目标说话人的音色和语调紧密匹配。然而,由于人类情感表达的内在复杂性,开发能够控制细微情感差异的TTS系统仍然是一个艰巨的挑战。现有的情感语音数据库往往遭受过于简单的标签计划,无法捕捉到广泛的情感状态,从而限制了情感合成的TTS应用程序的有效性。为此,最近的努力集中在建立使用自然语言注释来描述语音情感的数据库。然而,这些方法成本高昂,并且需要更多的情感深度来训练健壮的系统。在本文中,我们提出了一种新的过程,旨在通过系统地提取情感丰富的语音片段,并通过生成模型用详细的自然语言描述来注释它们,从而构建数据库。这种方法增强了数据库的情感粒度,并通过使用高级语言模型自动增强数据,显著降低了对昂贵的手动注释的依赖。由此产生的丰富的数据库提供了一个可扩展的和经济上可行的解决方案,为开发情绪控制的TTS系统开发一个更微妙和动态的基础。
摘要:Advances in text-to-speech (TTS) technology have significantly improved thequality of generated speech, closely matching the timbre and intonation of thetarget speaker. However, due to the inherent complexity of human emotionalexpression, the development of TTS systems capable of controlling subtleemotional differences remains a formidable challenge. Existing emotional speechdatabases often suffer from overly simplistic labelling schemes that fail tocapture a wide range of emotional states, thus limiting the effectiveness ofemotion synthesis in TTS applications. To this end, recent efforts havefocussed on building databases that use natural language annotations todescribe speech emotions. However, these approaches are costly and require moreemotional depth to train robust systems. In this paper, we propose a novelprocess aimed at building databases by systematically extracting emotion-richspeech segments and annotating them with detailed natural language descriptionsthrough a generative model. This approach enhances the emotional granularity ofthe database and significantly reduces the reliance on costly manualannotations by automatically augmenting the data with high-level languagemodels. The resulting rich database provides a scalable and economically viablesolution for developing a more nuanced and dynamic basis for developingemotionally controlled TTS systems.
标题:VidMusician:通过分层视觉特征实现语义节奏对齐的视频到音乐生成
链接:https://arxiv.org/abs/2412.06296
摘要:视频到音乐的生成在视频制作中呈现出巨大的潜力,要求生成的音乐在语义和节奏上都与视频一致。实现这种对齐需要先进的音乐生成功能,复杂的视频理解,以及学习两种模式之间对应关系的有效机制。在本文中,我们提出了VidMusician,一个参数高效的视频到音乐生成框架建立在文本到音乐模型。VidMusician利用分层视觉功能来确保视频和音乐之间的语义和节奏一致。具体来说,我们的方法利用全局视觉特征作为语义条件和局部视觉特征作为节奏线索。这些特征分别通过交叉注意和内注意机制集成到生成主干中。通过两个阶段的训练过程,我们逐步纳入语义和节奏特征,利用零初始化和身份初始化来保持骨干固有的音乐生成能力。此外,我们构建了一个多样化的视频音乐数据集DVMSet,包括各种场景,如宣传视频,广告和汇编。实验表明,VidMusician在多个评估指标上优于最先进的方法,并在AI生成的视频上表现出强大的性能。示例可在\url{https://youtu.be/EPOSXwtl1jw}获得。
摘要:Video-to-music generation presents significant potential in video production,requiring the generated music to be both semantically and rhythmically alignedwith the video. Achieving this alignment demands advanced music generationcapabilities, sophisticated video understanding, and an efficient mechanism tolearn the correspondence between the two modalities. In this paper, we proposeVidMusician, a parameter-efficient video-to-music generation framework builtupon text-to-music models. VidMusician leverages hierarchical visual featuresto ensure semantic and rhythmic alignment between video and music.Specifically, our approach utilizes global visual features as semanticconditions and local visual features as rhythmic cues. These features areintegrated into the generative backbone via cross-attention and in-attentionmechanisms, respectively. Through a two-stage training process, weincrementally incorporate semantic and rhythmic features, utilizing zeroinitialization and identity initialization to maintain the inherentmusic-generative capabilities of the backbone. Additionally, we construct adiverse video-music dataset, DVMSet, encompassing various scenarios, such aspromo videos, commercials, and compilations. Experiments demonstrate thatVidMusician outperforms state-of-the-art methods across multiple evaluationmetrics and exhibits robust performance on AI-generated videos. Samples areavailable at \url{https://youtu.be/EPOSXwtl1jw}.
标题:Sound 2Vision:通过跨模式潜在对齐从音频生成多样化的视觉效果
链接:https://arxiv.org/abs/2412.06209
备注:Under-review
摘要:音频如何描述我们周围的世界?在这项工作中,我们提出了一种方法,从不同的在野外的声音产生的视觉场景的图像。由于听觉和视觉信号之间存在显着的信息差距,这种跨模态生成任务具有挑战性。我们通过设计一个模型来应对这一挑战,该模型通过丰富音频功能与视觉信息并将其转换为视觉潜在空间来调整视听模式。然后将这些特征馈送到预训练的图像生成器中以生成图像。为了提高图像质量,我们使用声源定位来选择具有强交叉模态相关性的视听对。与以前的工作相比,我们的方法在VEGAS和VGGSound数据集上取得了更好的结果,并通过对输入波形或潜在空间的简单操作来控制生成过程。此外,我们分析了学习的嵌入空间的几何特性,并证明了我们的学习方法有效地对齐视听信号的跨模态生成。基于这种分析,我们表明,我们的方法是不可知的具体设计选择,通过整合各种模型架构和不同类型的视听数据显示其普适性。
摘要:How does audio describe the world around us? In this work, we propose amethod for generating images of visual scenes from diverse in-the-wild sounds.This cross-modal generation task is challenging due to the significantinformation gap between auditory and visual signals. We address this challengeby designing a model that aligns audio-visual modalities by enriching audiofeatures with visual information and translating them into the visual latentspace. These features are then fed into the pre-trained image generator toproduce images. To enhance image quality, we use sound source localization toselect audio-visual pairs with strong cross-modal correlations. Our methodachieves substantially better results on the VEGAS and VGGSound datasetscompared to previous work and demonstrates control over the generation processthrough simple manipulations to the input waveform or latent space.Furthermore, we analyze the geometric properties of the learned embedding spaceand demonstrate that our learning approach effectively aligns audio-visualsignals for cross-modal generation. Based on this analysis, we show that ourmethod is agnostic to specific design choices, showing its generalizability byintegrating various model architectures and different types of audio-visualdata.
标题:用于视听事件本地化的试点引导多模式语义通信
链接:https://arxiv.org/abs/2412.06208
摘要:多模式语义通信集成了文本、图像和音频等各种数据形式,显着提高了通信效率和可靠性。在人工智能、自动驾驶、智能家居等领域具有广阔的应用前景。然而,目前的研究主要依赖于模拟信道,并假设恒定的信道状态(完美CSI),这是不足以解决动态物理信道和噪声在现实世界的场景。现有的方法往往集中在单一的模态任务,并未能处理多模态流数据,如视频和音频,以及它们相应的任务。此外,当前的语义编码和解码模块主要传输单一模态特征,忽视了多模态语义增强和识别任务的需求。 为了解决这些挑战,本文提出了一个试点指导的框架,专门为视听事件本地化任务量身定制的多模态语义通信。该框架利用数字导频码和信道模块来指导实际场景中模拟信道的状态,并基于动态信道状态设计了考虑时频特性的基于欧拉的多模态语义编解码。该方法有效地处理多模态流源数据,特别是对于视听事件定位任务。大量的数值实验表明,所提出的框架在信道变化和支持各种通信场景的鲁棒性。实验结果表明,该框架优于现有的基准方法的信噪比(SNR),突出了其在语义通信质量的优势。
摘要:Multimodal semantic communication, which integrates various data modalitiessuch as text, images, and audio, significantly enhances communicationefficiency and reliability. It has broad application prospects in fields suchas artificial intelligence, autonomous driving, and smart homes. However,current research primarily relies on analog channels and assumes constantchannel states (perfect CSI), which is inadequate for addressing dynamicphysical channels and noise in real-world scenarios. Existing methods oftenfocus on single modality tasks and fail to handle multimodal stream data, suchas video and audio, and their corresponding tasks. Furthermore, currentsemantic encoding and decoding modules mainly transmit single modalityfeatures, neglecting the need for multimodal semantic enhancement andrecognition tasks. To address these challenges, this paper proposes a pilot-guided framework formultimodal semantic communication specifically tailored for audio-visual eventlocalization tasks. This framework utilizes digital pilot codes and channelmodules to guide the state of analog channels in real-wold scenarios anddesigns Euler-based multimodal semantic encoding and decoding that considertime-frequency characteristics based on dynamic channel state. This approacheffectively handles multimodal stream source data, especially for audio-visualevent localization tasks. Extensive numerical experiments demonstrate therobustness of the proposed framework in channel changes and its support forvarious communication scenarios. The experimental results show that theframework outperforms existing benchmark methods in terms of Signal-to-NoiseRatio (SNR), highlighting its advantage in semantic communication quality.
标题:M6:多生成器、多领域、多语言和文化、多流派、多乐器机器生成音乐检测数据库
链接:https://arxiv.org/abs/2412.06001
摘要:机器生成音乐(MGM)已经成为一种强大的工具,用于音乐治疗,个性化编辑和音乐社区的创作灵感。然而,它的不受管制的使用威胁着娱乐,教育和艺术部门,降低了高质量的人类作品的价值。因此,检测机器生成音乐(MGMD)对于保护这些领域至关重要,但该领域缺乏全面的数据集来支持有意义的进展。为了解决这一差距,我们引入了\textbf{M6},这是一个为MGMD研究量身定制的大规模基准数据集。M6以其多样性而闻名,包括多种生成器,领域,语言,文化背景,流派和工具。我们概述了我们的数据选择和收集方法,并附有详细的数据分析,提供所有WAV形式的音乐。此外,我们使用基本的二进制分类模型提供基线性能分数,说明了MGMD的复杂性和显著的改进空间。通过提供一个强大的和多方面的资源,我们的目标是使未来的研究开发更有效的检测方法MGM。我们相信,M6将成为应对这一社会挑战的关键一步。数据集和代码将免费提供,以支持该领域的开放式合作和创新。
摘要:Machine-generated music (MGM) has emerged as a powerful tool withapplications in music therapy, personalised editing, and creative inspirationfor the music community. However, its unregulated use threatens theentertainment, education, and arts sectors by diminishing the value ofhigh-quality human compositions. Detecting machine-generated music (MGMD) is,therefore, critical to safeguarding these domains, yet the field lackscomprehensive datasets to support meaningful progress. To address this gap, weintroduce \textbf{M6}, a large-scale benchmark dataset tailored for MGMDresearch. M6 is distinguished by its diversity, encompassing multiplegenerators, domains, languages, cultural contexts, genres, and instruments. Weoutline our methodology for data selection and collection, accompanied bydetailed data analysis, providing all WAV form of music. Additionally, weprovide baseline performance scores using foundational binary classificationmodels, illustrating the complexity of MGMD and the significant room forimprovement. By offering a robust and multifaceted resource, we aim to empowerfuture research to develop more effective detection methods for MGM. We believeM6 will serve as a critical step toward addressing this societal challenge. Thedataset and code will be freely available to support open collaboration andinnovation in this field.
标题:当视觉模型满足参数高效的后备适配器时,无需大规模音频预训练
链接:https://arxiv.org/abs/2412.05951
备注:5 pages, 3 figures
摘要:最近的研究表明,预训练的视觉模型可以提高音频下游任务的性能。为了进一步提高性能,通常需要使用大规模音频数据的额外预训练阶段,以将音频特定知识注入视觉模型中。然而,这种方法需要大量的音频数据和精心设计的目标函数。在这项工作中,我们建议绕过预训练阶段,直接用我们的Look Aside Adapter(LoAA)微调视觉模型,以实现高效的音频理解。音频频谱数据表示在两个异构的维度时间和频率,我们完善适配器,以促进跨这些维度的令牌之间的交互。我们的实验表明,我们的适配器允许视觉模型在各种音频和语音任务中达到或超过预训练音频模型的性能,为在音频应用中利用视觉模型提供了一种资源高效和有效的解决方案。
摘要:Recent studies show that pretrained vision models can boost performance inaudio downstream tasks. To enhance the performance further, an additionalpretraining stage with large scale audio data is typically required to infuseaudio specific knowledge into the vision model. However, such approachesrequire extensive audio data and a carefully designed objective function. Inthis work, we propose bypassing the pretraining stage by directly fine-tuningthe vision model with our Look Aside Adapter (LoAA) designed for efficientaudio understanding. Audio spectrum data is represented across twoheterogeneous dimensions time and frequency and we refine adapters tofacilitate interactions between tokens across these dimensions. Our experimentsdemonstrate that our adapters allow vision models to reach or surpass theperformance of pretrained audio models in various audio and speech tasks,offering a resource efficient and effective solution for leveraging visionmodels in audio applications.
标题:用于可控视频到音乐检索的半监督对比学习
链接:https://arxiv.org/abs/2412.05831
备注:4 pages + 1 reference page, 2 figures, 2 tables. Under review
摘要:内容创作者经常使用音乐来增强他们的视频,从电影中的配乐到视频博客和社交媒体内容中的背景音乐。然而,识别视频的最佳音乐可能是一项困难且耗时的任务。为了解决这一挑战,我们提出了一个新的框架,自动检索匹配的音乐剪辑为给定的视频,反之亦然。我们的方法利用了注释的音乐标签,以及视觉和音乐元素之间固有的艺术对应。不同于以往的跨模态音乐检索作品,我们的方法结合了自我监督和监督的训练目标。我们使用自监督和标签监督对比学习训练音乐和视频之间的联合嵌入空间。我们通过对监督训练组件使用音乐流派标签来展示我们方法的有效性,并且我们的框架可以推广到其他音乐注释(例如,情感、工具等)。此外,我们的方法可以细粒度地控制检索过程在推理时关注自监督信息与标签信息的程度。我们通过各种视频到音乐和音乐到视频检索任务来评估学习到的嵌入。实验表明,该方法成功地结合了自监督和监督目标,是有效的可控音乐视频检索。
摘要:Content creators often use music to enhance their videos, from soundtracks inmovies to background music in video blogs and social media content. However,identifying the best music for a video can be a difficult and time-consumingtask. To address this challenge, we propose a novel framework for automaticallyretrieving a matching music clip for a given video, and vice versa. Ourapproach leverages annotated music labels, as well as the inherent artisticcorrespondence between visual and music elements. Distinct from previouscross-modal music retrieval works, our method combines both self-supervised andsupervised training objectives. We use self-supervised and label-supervisedcontrastive learning to train a joint embedding space between music and video.We show the effectiveness of our approach by using music genre labels for thesupervised training component, and our framework can be generalized to othermusic annotations (e.g., emotion, instrument, etc.). Furthermore, our methodenables fine-grained control over how much the retrieval process focuses onself-supervised vs. label information at inference time. We evaluate thelearned embeddings through a variety of video-to-music and music-to-videoretrieval tasks. Our experiments show that the proposed approach successfullycombines self-supervised and supervised objectives and is effective forcontrollable music-video retrieval.
标题:将流派分类和和声打击特征与扩散模型相结合用于音乐视频生成
链接:https://arxiv.org/abs/2412.05694
摘要:本研究提出了一种新的方法,用于生成音乐可视化使用扩散模型,结合音频输入与用户选择的艺术品。该过程包括两个主要阶段:图像生成和视频创建。首先,音乐字幕和体裁分类,其次是检索的艺术风格的描述。扩散模型然后基于用户的输入图像和所导出的艺术风格描述来生成图像。视频生成阶段利用相同的扩散模型来内插帧,由从和声和和弦的关键音乐特征导出的音频能量矢量控制。该方法在各种流派中表现出良好的效果,并引入了一个新的度量标准,视听同步(AVS),以定量评估视觉和音频元素之间的同步。对比分析表明,与线性插值相比,使用所提出的方法与音频能量向量生成的视频的AVS值显着更高。这种方法在各个领域都有潜在的应用,包括独立音乐视频创作,电影制作,现场音乐活动以及增强公共空间的视听体验。
摘要:This study presents a novel method for generating music visualisers usingdiffusion models, combining audio input with user-selected artwork. The processinvolves two main stages: image generation and video creation. First, musiccaptioning and genre classification are performed, followed by the retrieval ofartistic style descriptions. A diffusion model then generates images based onthe user's input image and the derived artistic style descriptions. The videogeneration stage utilises the same diffusion model to interpolate frames,controlled by audio energy vectors derived from key musical features ofharmonics and percussives. The method demonstrates promising results acrossvarious genres, and a new metric, Audio-Visual Synchrony (AVS), is introducedto quantitatively evaluate the synchronisation between visual and audioelements. Comparative analysis shows significantly higher AVS values for videosgenerated using the proposed method with audio energy vectors, compared tolinear interpolation. This approach has potential applications in diversefields, including independent music video creation, film production, live musicevents, and enhancing audio-visual experiences in public spaces.
标题:WavFusion:迈向wav2vec 2.0多模式语音情感识别
链接:https://arxiv.org/abs/2412.05558
备注:Accepted by 31st International Conference on MultiMedia Modeling (MMM2025)
摘要:由于人类情感固有的复杂性和多样性,语音情感识别(SER)仍然是一个具有挑战性的关键任务。为了解决这个问题,研究人员试图通过多模态学习融合来自其他模态的信息。然而,现有的多模态融合技术往往忽略了跨模态相互作用的复杂性,从而导致次优的特征表示。在本文中,我们提出了WavFusion,这是一个多模态语音情感识别框架,它解决了有效多模态融合、模态之间的异质性和区分性表示学习中的关键研究问题。通过利用门控跨模态注意力机制和多模态同质特征差异学习,WavFusion在基准数据集上展示了比现有最先进方法更好的性能。我们的工作强调了捕捉细微差别的跨模态交互和学习区分表示以获得准确的多模态SER的重要性。两个基准数据集(IEMOCAP和MELD)的实验结果表明,WavFusion在情感识别方面成功地超过了最先进的策略。
摘要:Speech emotion recognition (SER) remains a challenging yet crucial task dueto the inherent complexity and diversity of human emotions. To address thisproblem, researchers attempt to fuse information from other modalities viamultimodal learning. However, existing multimodal fusion techniques oftenoverlook the intricacies of cross-modal interactions, resulting in suboptimalfeature representations. In this paper, we propose WavFusion, a multimodalspeech emotion recognition framework that addresses critical research problemsin effective multimodal fusion, heterogeneity among modalities, anddiscriminative representation learning. By leveraging a gated cross-modalattention mechanism and multimodal homogeneous feature discrepancy learning,WavFusion demonstrates improved performance over existing state-of-the-artmethods on benchmark datasets. Our work highlights the importance of capturingnuanced cross-modal interactions and learning discriminative representationsfor accurate multimodal SER. Experimental results on two benchmark datasets(IEMOCAP and MELD) demonstrate that WavFusion succeeds over thestate-of-the-art strategies on emotion recognition.
标题:pyAMPACT:用于性能数据估计和多模式处理的分数音频对齐工具包
链接:https://arxiv.org/abs/2412.05436
备注:International Society for Music Information Retrieval, Late Breaking Demo
摘要:pyAMPACT(基于Python的自动音乐表现分析和比较工具包)链接符号和音频音乐表示,以促进对音频中的表现数据的分数信息估计,以及符号和音频音乐表示与各种注释的一般链接。pyAMPACT可以读取一系列符号格式,并可以将音符链接的音频描述符/性能数据输出到MEI格式的文件中。音频分析使用分数对齐来计算符号表示中的每个音符的重要性的时间-频率区域,从该时间-频率区域来估计参数的范围。这些包括调音、动态和音色相关的性能描述符,以及来自乐谱对齐的时序相关信息。除了性能数据估计之外,pyAMPACT还通过其将符号表示和注释链接到音频的基础设施来促进多模态调查。
摘要:pyAMPACT (Python-based Automatic Music Performance Analysis and ComparisonToolkit) links symbolic and audio music representations to facilitatescore-informed estimation of performance data in audio as well as generallinking of symbolic and audio music representations with a variety ofannotations. pyAMPACT can read a range of symbolic formats and can outputnote-linked audio descriptors/performance data into MEI-formatted files. Theaudio analysis uses score alignment to calculate time-frequency regions ofimportance for each note in the symbolic representation from which to estimatea range of parameters. These include tuning-, dynamics-, and timbre-relatedperformance descriptors, with timing-related information available from thescore alignment. Beyond performance data estimation, pyAMPACT also facilitatesmulti-modal investigations through its infrastructure for linking symbolicrepresentations and annotations to audio.
标题:重温你的记忆:通过脑电引导的视听生成重建情感背景化记忆
链接:https://arxiv.org/abs/2412.05296
备注:Codes and the dataset will be released upon acceptance
摘要:在本文中,我们介绍了RecallAffectiveMemory,一个新的任务,旨在重建自传体记忆,通过视听生成引导的影响提取的脑电图(EEG)信号。为了支持这一开创性的任务,我们提出了EEG-AffectiveMemory数据集,其中包括文本描述,视觉,音乐和EEG记录,这些记录是在9名参与者的记忆回忆过程中收集的。此外,我们提出了RYM(回忆你的记忆),一个三阶段的框架,用于生成同步的视听内容,同时保持动态的个人记忆影响轨迹。实验结果表明,我们的方法可以忠实地重建所有受试者的情感情境化视听记忆,定性和定量,与参与者报告强烈的情感一致性之间的回忆和生成的内容。我们的方法通过基于神经的情感理解推进了情感解码研究及其在个性化媒体创作中的实际应用。
摘要:In this paper, we introduce RecallAffectiveMemory, a novel task designed toreconstruct autobiographical memories through audio-visual generation guided byaffect extracted from electroencephalogram (EEG) signals. To support thispioneering task, we present the EEG-AffectiveMemory dataset, which encompassestextual descriptions, visuals, music, and EEG recordings collected duringmemory recall from nine participants. Furthermore, we propose RYM (Recall YourMemory), a three-stage framework for generating synchronized audio-visualcontents while maintaining dynamic personal memory affect trajectories.Experimental results indicate that our method can faithfully reconstructaffect-contextualized audio-visual memory across all subjects, bothqualitatively and quantitatively, with participants reporting strong affectiveconcordance between their recalled memories and the generated content. Ourapproaches advance affect decoding research and its practical applications inpersonalized media creation via neural-based affect comprehension.
标题:利用即时学习和预设编码进行阿尔茨海默病检测
链接:https://arxiv.org/abs/2412.06259
备注:Accepted by ISCSLP 2024
摘要:与其他临床筛查技术相比,基于语音和语言的阿尔茨海默病(AD)自动检测方法具有非侵入性,成本效益和方便的特点。以前的研究已经证明了微调预训练语言模型(PLM)用于AD检测的有效性。然而,这种传统的微调方法的目标,它涉及只输入成绩单,是不一致的,在PLM的预训练阶段使用的掩蔽语言建模(MLM)任务。在本文中,我们研究了基于文本的PLM微调,通过将提示模板插入到成绩单输入中,将分类任务转换为MLM任务。我们还探讨了将暂停信息强制对齐到手动成绩单的影响。此外,我们比较了各种自动语音识别(ASR)模型的性能,并选择Whisper模型生成基于ASR的成绩单与手动成绩单进行比较。此外,多数表决和集成技术应用于不同的PLM(BERT和RoBERTa)使用不同的随机种子。最终,我们使用手动转录本获得了95.8%的最大检测准确率(平均值为87.9%,标准值为3.3%),仅使用ADReSS测试集上的转录本即可实现AD检测的最佳性能。
摘要:Compared to other clinical screening techniques, speech-and-language-basedautomated Alzheimer's disease (AD) detection methods are characterized by theirnon-invasiveness, cost-effectiveness, and convenience. Previous studies havedemonstrated the efficacy of fine-tuning pre-trained language models (PLMs) forAD detection. However, the objective of this traditional fine-tuning method,which involves inputting only transcripts, is inconsistent with the maskedlanguage modeling (MLM) task used during the pre-training phase of PLMs. Inthis paper, we investigate prompt-based fine-tuning of PLMs, converting theclassification task into a MLM task by inserting prompt templates into thetranscript inputs. We also explore the impact of incorporating pauseinformation from forced alignment into manual transcripts. Additionally, wecompare the performance of various automatic speech recognition (ASR) modelsand select the Whisper model to generate ASR-based transcripts for comparisonwith manual transcripts. Furthermore, majority voting and ensemble techniquesare applied across different PLMs (BERT and RoBERTa) using different randomseeds. Ultimately, we obtain maximum detection accuracy of 95.8% (with mean87.9%, std 3.3%) using manual transcripts, achieving state-of-the-artperformance for AD detection using only transcripts on the ADReSS test set.
标题:SQ-Whisper:基于扬声器查询的目标扬声器ASB Whisper模型
链接:https://arxiv.org/abs/2412.05589
备注:Accepted by IEEE/ACM TASLP
摘要:得益于海量和多样化的数据源,语音基础模型表现出强大的泛化能力和知识转移能力,可用于广泛的下游任务。然而,一个局限性,从他们的独家处理单说话人的语音输入,使他们在识别多说话人重叠的语音,在现实世界中常见的情况下无效。在这项研究中,我们深入到语音基础模型的适应,以消除干扰扬声器重叠的语音和执行目标说话人自动语音识别(TS-ASR)。首先,我们利用Whisper模型作为适应的基础,并将其与现有的目标说话人自适应技术的集成进行了彻底的比较。然后,我们提出了一个创新的模型,称为扬声器查询耳语(SQ-Whisper),它采用了一组数量的可训练的查询捕获扬声器提示重叠语音的基础上目标扬声器登记。这些提示用于引导模型提取特定于说话人的特征,并准确地识别目标说话人的语音。实验结果表明,我们的方法有效地适应预先训练的语音基础模型TS-ASR。与鲁棒的TS-HuBERT模型相比,所提出的SQ-Whisper显著提高了性能,在Libri 2 Mix和WSJ 0 - 2 Mix数据集上,单词错误率(WER)分别降低了15%和10%。随着数据的增加,我们在Libri 2 Mix测试集上建立了14.6%的新的最先进的WER,在WSJ 0 - 2 Mix测试集上建立了4.4%的WER。此外,我们评估了我们的模型在现实世界的AMI会议数据集,这表明了一致的改进,比其他适应方法。
摘要:Benefiting from massive and diverse data sources, speech foundation modelsexhibit strong generalization and knowledge transfer capabilities to a widerange of downstream tasks. However, a limitation arises from their exclusivehandling of single-speaker speech input, making them ineffective in recognizingmulti-speaker overlapped speech, a common occurrence in real-world scenarios.In this study, we delve into the adaptation of speech foundation models toeliminate interfering speakers from overlapping speech and performtarget-speaker automatic speech recognition (TS-ASR). Initially, we utilize theWhisper model as the foundation for adaptation and conduct a thoroughcomparison of its integration with existing target-speaker adaptationtechniques. We then propose an innovative model termed Speaker-Querying Whisper(SQ-Whisper), which employs a set number of trainable queries to capturespeaker prompts from overlapping speech based on target-speaker enrollment.These prompts serve to steer the model in extracting speaker-specific featuresand accurately recognizing target-speaker transcriptions. Experimental resultsdemonstrate that our approach effectively adapts the pre-trained speechfoundation model to TS-ASR. Compared with the robust TS-HuBERT model, theproposed SQ-Whisper significantly improves performance, yielding up to 15% and10% relative reductions in word error rates (WERs) on the Libri2Mix andWSJ0-2Mix datasets, respectively. With data augmentation, we establish newstate-of-the-art WERs of 14.6% on the Libri2Mix Test set and 4.4% on theWSJ0-2Mix Test set. Furthermore, we evaluate our model on the real-world AMImeeting dataset, which shows consistent improvement over other adaptationmethods.
标题:利用即时学习和预设编码进行阿尔茨海默病检测
链接:https://arxiv.org/abs/2412.06259
备注:Accepted by ISCSLP 2024
摘要:与其他临床筛查技术相比,基于语音和语言的阿尔茨海默病(AD)自动检测方法具有非侵入性,成本效益和方便的特点。以前的研究已经证明了微调预训练语言模型(PLM)用于AD检测的有效性。然而,这种传统的微调方法的目标,它涉及只输入成绩单,是不一致的,在PLM的预训练阶段使用的掩蔽语言建模(MLM)任务。在本文中,我们研究了基于文本的PLM微调,通过将提示模板插入到成绩单输入中,将分类任务转换为MLM任务。我们还探讨了将暂停信息强制对齐到手动成绩单的影响。此外,我们比较了各种自动语音识别(ASR)模型的性能,并选择Whisper模型生成基于ASR的成绩单与手动成绩单进行比较。此外,多数表决和集成技术应用于不同的PLM(BERT和RoBERTa)使用不同的随机种子。最终,我们使用手动转录本获得了95.8%的最大检测准确率(平均值为87.9%,标准值为3.3%),仅使用ADReSS测试集上的转录本即可实现AD检测的最佳性能。
摘要:Compared to other clinical screening techniques, speech-and-language-basedautomated Alzheimer's disease (AD) detection methods are characterized by theirnon-invasiveness, cost-effectiveness, and convenience. Previous studies havedemonstrated the efficacy of fine-tuning pre-trained language models (PLMs) forAD detection. However, the objective of this traditional fine-tuning method,which involves inputting only transcripts, is inconsistent with the maskedlanguage modeling (MLM) task used during the pre-training phase of PLMs. Inthis paper, we investigate prompt-based fine-tuning of PLMs, converting theclassification task into a MLM task by inserting prompt templates into thetranscript inputs. We also explore the impact of incorporating pauseinformation from forced alignment into manual transcripts. Additionally, wecompare the performance of various automatic speech recognition (ASR) modelsand select the Whisper model to generate ASR-based transcripts for comparisonwith manual transcripts. Furthermore, majority voting and ensemble techniquesare applied across different PLMs (BERT and RoBERTa) using different randomseeds. Ultimately, we obtain maximum detection accuracy of 95.8% (with mean87.9%, std 3.3%) using manual transcripts, achieving state-of-the-artperformance for AD detection using only transcripts on the ADReSS test set.
标题:SQ-Whisper:基于扬声器查询的目标扬声器ASB Whisper模型
链接:https://arxiv.org/abs/2412.05589
备注:Accepted by IEEE/ACM TASLP
摘要:得益于海量和多样化的数据源,语音基础模型表现出强大的泛化能力和知识转移能力,可用于广泛的下游任务。然而,一个局限性,从他们的独家处理单说话人的语音输入,使他们在识别多说话人重叠的语音,在现实世界中常见的情况下无效。在这项研究中,我们深入到语音基础模型的适应,以消除干扰扬声器重叠的语音和执行目标说话人自动语音识别(TS-ASR)。首先,我们利用Whisper模型作为适应的基础,并将其与现有的目标说话人自适应技术的集成进行了彻底的比较。然后,我们提出了一个创新的模型,称为扬声器查询耳语(SQ-Whisper),它采用了一组数量的可训练的查询捕获扬声器提示重叠语音的基础上目标扬声器登记。这些提示用于引导模型提取特定于说话人的特征,并准确地识别目标说话人的语音。实验结果表明,我们的方法有效地适应预先训练的语音基础模型TS-ASR。与鲁棒的TS-HuBERT模型相比,所提出的SQ-Whisper显著提高了性能,在Libri 2 Mix和WSJ 0 - 2 Mix数据集上,单词错误率(WER)分别降低了15%和10%。随着数据的增加,我们在Libri 2 Mix测试集上建立了14.6%的新的最先进的WER,在WSJ 0 - 2 Mix测试集上建立了4.4%的WER。此外,我们评估了我们的模型在现实世界的AMI会议数据集,这表明了一致的改进,比其他适应方法。
摘要:Benefiting from massive and diverse data sources, speech foundation modelsexhibit strong generalization and knowledge transfer capabilities to a widerange of downstream tasks. However, a limitation arises from their exclusivehandling of single-speaker speech input, making them ineffective in recognizingmulti-speaker overlapped speech, a common occurrence in real-world scenarios.In this study, we delve into the adaptation of speech foundation models toeliminate interfering speakers from overlapping speech and performtarget-speaker automatic speech recognition (TS-ASR). Initially, we utilize theWhisper model as the foundation for adaptation and conduct a thoroughcomparison of its integration with existing target-speaker adaptationtechniques. We then propose an innovative model termed Speaker-Querying Whisper(SQ-Whisper), which employs a set number of trainable queries to capturespeaker prompts from overlapping speech based on target-speaker enrollment.These prompts serve to steer the model in extracting speaker-specific featuresand accurately recognizing target-speaker transcriptions. Experimental resultsdemonstrate that our approach effectively adapts the pre-trained speechfoundation model to TS-ASR. Compared with the robust TS-HuBERT model, theproposed SQ-Whisper significantly improves performance, yielding up to 15% and10% relative reductions in word error rates (WERs) on the Libri2Mix andWSJ0-2Mix datasets, respectively. With data augmentation, we establish newstate-of-the-art WERs of 14.6% on the Libri2Mix Test set and 4.4% on theWSJ0-2Mix Test set. Furthermore, we evaluate our model on the real-world AMImeeting dataset, which shows consistent improvement over other adaptationmethods.
标题:音乐的来源分离和自动转录
链接:https://arxiv.org/abs/2412.06703
摘要:源分离是在多个声音的听觉混合物中分离单个声音的过程[1],并且具有从语音增强和歌词转录[2]到音乐的数字音频制作的各种应用。此外,自动音乐转录(AMT)是将原始音乐音频转换为音乐家可以阅读的乐谱的过程[3]。从历史上看,这些任务面临着巨大的音频噪声,长时间的训练,以及由于版权限制而缺乏免费使用的数据等挑战。然而,深度学习的最新发展带来了新的有前途的方法来构建低失真的茎和从音频信号生成乐谱[4]。使用频谱图掩蔽,深度神经网络和MuseScore API,我们试图创建一个端到端的管道,允许初始音乐音频混合(例如... WAV文件),以被分离成乐器词干、转换成MIDI文件、以及转录成每个组成乐器的乐谱。
摘要:Source separation is the process of isolating individual sounds in anauditory mixture of multiple sounds [1], and has a variety of applicationsranging from speech enhancement and lyric transcription [2] to digital audioproduction for music. Furthermore, Automatic Music Transcription (AMT) is theprocess of converting raw music audio into sheet music that musicians can read[3]. Historically, these tasks have faced challenges such as significant audionoise, long training times, and lack of free-use data due to copyrightrestrictions. However, recent developments in deep learning have brought newpromising approaches to building low-distortion stems and generating sheetmusic from audio signals [4]. Using spectrogram masking, deep neural networks,and the MuseScore API, we attempt to create an end-to-end pipeline that allowsfor an initial music audio mixture (e.g...wav file) to be separated intoinstrument stems, converted into MIDI files, and transcribed into sheet musicfor each component instrument.
标题:MuMuMu-LLaMA:通过大型语言模型的多模式音乐理解和生成
链接:https://arxiv.org/abs/2412.06660
摘要:大型语言模型的研究在文本、语音、图像和视频方面取得了显著进展。然而,多模态音乐的理解和生成由于缺乏良好的注释数据集仍然未被充分探索。为了解决这个问题,我们引入了一个包含167.69小时多模态数据的数据集,包括文本、图像、视频和音乐注释。基于这个数据集,我们提出了MuMu-LLaMA,这是一个利用音乐,图像和视频的预训练编码器的模型。对于音乐生成,我们集成了AudioLDM 2和MusicGen。我们对四个任务的评估-音乐理解,文本到音乐生成,基于音乐编辑和多模态音乐生成-表明MuMu-LLaMA优于最先进的模型,显示了其多模态音乐应用的潜力。
摘要:Research on large language models has advanced significantly across text,speech, images, and videos. However, multi-modal music understanding andgeneration remain underexplored due to the lack of well-annotated datasets. Toaddress this, we introduce a dataset with 167.69 hours of multi-modal data,including text, images, videos, and music annotations. Based on this dataset,we propose MuMu-LLaMA, a model that leverages pre-trained encoders for music,images, and videos. For music generation, we integrate AudioLDM 2 and MusicGen.Our evaluation across four tasks--music understanding, text-to-musicgeneration, prompt-based music editing, and multi-modal musicgeneration--demonstrates that MuMu-LLaMA outperforms state-of-the-art models,showing its potential for multi-modal music applications.
链接:https://arxiv.org/abs/2412.06617
备注:Accepted for the NeurIPS 2024 Creative AI Track
摘要:“卧室制作人”的兴起使音乐创作民主化,同时也挑战制作人客观地评估他们的作品。为了解决这个问题,我们提出了AI TrackMate,这是一个基于LLM的音乐聊天机器人,旨在为音乐制作提供建设性的反馈。通过将LLM固有的音乐知识与直接的音轨分析相结合,AI TrackMate提供了特定于制作的见解,将其与纯文本方法区分开来。我们的框架集成了一个音乐分析模块,一个法学硕士可读的音乐报告和音乐制作导向的反馈指令,创建一个即插即用,与各种法学硕士兼容的免培训系统,并适应未来的发展。我们通过交互式Web界面展示了AI TrackMate的功能,并与音乐制作人进行了试点研究。通过将人工智能功能与独立制作人的需求联系起来,AI TrackMate提供了按需分析反馈,可能支持音乐制作的创作过程和技能发展。该系统解决了在不断发展的独立音乐制作领域对客观自我评估工具日益增长的需求。
摘要:The rise of "bedroom producers" has democratized music creation, whilechallenging producers to objectively evaluate their work. To address this, wepresent AI TrackMate, an LLM-based music chatbot designed to provideconstructive feedback on music productions. By combining LLMs' inherent musicalknowledge with direct audio track analysis, AI TrackMate offersproduction-specific insights, distinguishing it from text-only approaches. Ourframework integrates a Music Analysis Module, an LLM-Readable Music Report, andMusic Production-Oriented Feedback Instruction, creating a plug-and-play,training-free system compatible with various LLMs and adaptable to futureadvancements. We demonstrate AI TrackMate's capabilities through an interactiveweb interface and present findings from a pilot study with a music producer. Bybridging AI capabilities with the needs of independent producers, AI TrackMateoffers on-demand analytical feedback, potentially supporting the creativeprocess and skill development in music production. This system addresses thegrowing demand for objective self-assessment tools in the evolving landscape ofindependent music production.
标题:大型语言模型时代的可控语音合成:综述
链接:https://arxiv.org/abs/2412.06602
备注:A comprehensive survey on controllable TTS, 23 pages, 6 tables, 4 figures, 280 references
摘要:文本到语音(TTS),也称为语音合成,是一个突出的研究领域,旨在从文本生成自然的人类语音。近年来,随着工业需求的不断增长,TTS技术已经从合成类人语音发展到实现可控语音生成。这包括对合成语音的各种属性(如情感、韵律、音色和持续时间)的细粒度控制。此外,深度学习的进步,如扩散和大型语言模型,在过去几年中显着增强了可控TTS。在本文中,我们进行了全面的调查可控TTS,涵盖的方法,从基本的控制技术,利用自然语言提示的方法,旨在提供一个清晰的了解目前的研究状况。我们研究了一般可控的TTS管道,挑战,模型架构和控制策略,提供了一个全面而清晰的分类现有的方法。此外,我们提供了数据集和评估指标的详细摘要,并阐明了可控TTS的应用和未来的发展方向。据我们所知,这份调查报告提供了第一个全面审查新兴的可控TTS方法,这可以作为一个有益的资源,为学术研究人员和行业从业者。
摘要:Text-to-speech (TTS), also known as speech synthesis, is a prominent researcharea that aims to generate natural-sounding human speech from text. Recently,with the increasing industrial demand, TTS technologies have evolved beyondsynthesizing human-like speech to enabling controllable speech generation. Thisincludes fine-grained control over various attributes of synthesized speechsuch as emotion, prosody, timbre, and duration. Besides, advancements in deeplearning, such as diffusion and large language models, have significantlyenhanced controllable TTS over the past several years. In this paper, weconduct a comprehensive survey of controllable TTS, covering approaches rangingfrom basic control techniques to methods utilizing natural language prompts,aiming to provide a clear understanding of the current state of research. Weexamine the general controllable TTS pipeline, challenges, model architectures,and control strategies, offering a comprehensive and clear taxonomy of existingmethods. Additionally, we provide a detailed summary of datasets and evaluationmetrics and shed some light on the applications and future directions ofcontrollable TTS. To the best of our knowledge, this survey paper provides thefirst comprehensive review of emerging controllable TTS methods, which canserve as a beneficial resource for both academic researchers and industrypractitioners.
标题:语音:情感丰富且上下文详细的语音注释库
链接:https://arxiv.org/abs/2412.06581
备注:4 pages, 1 figure. To appear in the Proceedings of the International Symposium on Chinese Spoken Language Processing, 7-10 November 2024, Beijing, China
摘要:文语转换(TTS)技术的进步显著提高了生成语音的质量,使其与目标说话人的音色和语调紧密匹配。然而,由于人类情感表达的内在复杂性,开发能够控制细微情感差异的TTS系统仍然是一个艰巨的挑战。现有的情感语音数据库往往遭受过于简单的标签计划,无法捕捉到广泛的情感状态,从而限制了情感合成的TTS应用程序的有效性。为此,最近的努力集中在建立使用自然语言注释来描述语音情感的数据库。然而,这些方法成本高昂,并且需要更多的情感深度来训练健壮的系统。在本文中,我们提出了一种新的过程,旨在通过系统地提取情感丰富的语音片段,并通过生成模型用详细的自然语言描述来注释它们,从而构建数据库。这种方法增强了数据库的情感粒度,并通过使用高级语言模型自动增强数据,显著降低了对昂贵的手动注释的依赖。由此产生的丰富的数据库提供了一个可扩展的和经济上可行的解决方案,为开发情绪控制的TTS系统开发一个更微妙和动态的基础。
摘要:Advances in text-to-speech (TTS) technology have significantly improved thequality of generated speech, closely matching the timbre and intonation of thetarget speaker. However, due to the inherent complexity of human emotionalexpression, the development of TTS systems capable of controlling subtleemotional differences remains a formidable challenge. Existing emotional speechdatabases often suffer from overly simplistic labelling schemes that fail tocapture a wide range of emotional states, thus limiting the effectiveness ofemotion synthesis in TTS applications. To this end, recent efforts havefocussed on building databases that use natural language annotations todescribe speech emotions. However, these approaches are costly and require moreemotional depth to train robust systems. In this paper, we propose a novelprocess aimed at building databases by systematically extracting emotion-richspeech segments and annotating them with detailed natural language descriptionsthrough a generative model. This approach enhances the emotional granularity ofthe database and significantly reduces the reliance on costly manualannotations by automatically augmenting the data with high-level languagemodels. The resulting rich database provides a scalable and economically viablesolution for developing a more nuanced and dynamic basis for developingemotionally controlled TTS systems.
标题:VidMusician:通过分层视觉特征实现语义节奏对齐的视频到音乐生成
链接:https://arxiv.org/abs/2412.06296
摘要:视频到音乐的生成在视频制作中呈现出巨大的潜力,要求生成的音乐在语义和节奏上都与视频一致。实现这种对齐需要先进的音乐生成功能,复杂的视频理解,以及学习两种模式之间对应关系的有效机制。在本文中,我们提出了VidMusician,一个参数高效的视频到音乐生成框架建立在文本到音乐模型。VidMusician利用分层视觉功能来确保视频和音乐之间的语义和节奏一致。具体来说,我们的方法利用全局视觉特征作为语义条件和局部视觉特征作为节奏线索。这些特征分别通过交叉注意和内注意机制集成到生成主干中。通过两个阶段的训练过程,我们逐步纳入语义和节奏特征,利用零初始化和身份初始化来保持骨干固有的音乐生成能力。此外,我们构建了一个多样化的视频音乐数据集DVMSet,包括各种场景,如宣传视频,广告和汇编。实验表明,VidMusician在多个评估指标上优于最先进的方法,并在AI生成的视频上表现出强大的性能。示例可在\url{https://youtu.be/EPOSXwtl1jw}获得。
摘要:Video-to-music generation presents significant potential in video production,requiring the generated music to be both semantically and rhythmically alignedwith the video. Achieving this alignment demands advanced music generationcapabilities, sophisticated video understanding, and an efficient mechanism tolearn the correspondence between the two modalities. In this paper, we proposeVidMusician, a parameter-efficient video-to-music generation framework builtupon text-to-music models. VidMusician leverages hierarchical visual featuresto ensure semantic and rhythmic alignment between video and music.Specifically, our approach utilizes global visual features as semanticconditions and local visual features as rhythmic cues. These features areintegrated into the generative backbone via cross-attention and in-attentionmechanisms, respectively. Through a two-stage training process, weincrementally incorporate semantic and rhythmic features, utilizing zeroinitialization and identity initialization to maintain the inherentmusic-generative capabilities of the backbone. Additionally, we construct adiverse video-music dataset, DVMSet, encompassing various scenarios, such aspromo videos, commercials, and compilations. Experiments demonstrate thatVidMusician outperforms state-of-the-art methods across multiple evaluationmetrics and exhibits robust performance on AI-generated videos. Samples areavailable at \url{https://youtu.be/EPOSXwtl1jw}.
标题:Sound 2Vision:通过跨模式潜在对齐从音频生成多样化的视觉效果
链接:https://arxiv.org/abs/2412.06209
备注:Under-review
摘要:音频如何描述我们周围的世界?在这项工作中,我们提出了一种方法,从不同的在野外的声音产生的视觉场景的图像。由于听觉和视觉信号之间存在显著的信息差,这种跨模态生成任务具有挑战性。我们通过设计一个模型来应对这一挑战,该模型通过丰富音频功能与视觉信息并将其转换为视觉潜在空间来调整视听模式。然后将这些特征馈送到预训练的图像生成器中以生成图像。为了提高图像质量,我们使用声源定位来选择具有强交叉模态相关性的视听对。与以前的工作相比,我们的方法在VEGAS和VGGSound数据集上取得了更好的结果,并通过对输入波形或潜在空间的简单操作来控制生成过程。此外,我们分析了学习的嵌入空间的几何特性,并证明了我们的学习方法有效地对齐视听信号的跨模态生成。基于这一分析,我们表明我们的方法对特定的设计选择是不可知的,通过集成各种模型架构和不同类型的视听数据来展示其通用性。
摘要:How does audio describe the world around us? In this work, we propose amethod for generating images of visual scenes from diverse in-the-wild sounds.This cross-modal generation task is challenging due to the significantinformation gap between auditory and visual signals. We address this challengeby designing a model that aligns audio-visual modalities by enriching audiofeatures with visual information and translating them into the visual latentspace. These features are then fed into the pre-trained image generator toproduce images. To enhance image quality, we use sound source localization toselect audio-visual pairs with strong cross-modal correlations. Our methodachieves substantially better results on the VEGAS and VGGSound datasetscompared to previous work and demonstrates control over the generation processthrough simple manipulations to the input waveform or latent space.Furthermore, we analyze the geometric properties of the learned embedding spaceand demonstrate that our learning approach effectively aligns audio-visualsignals for cross-modal generation. Based on this analysis, we show that ourmethod is agnostic to specific design choices, showing its generalizability byintegrating various model architectures and different types of audio-visualdata.
标题:用于视听事件本地化的试点引导多模式语义通信
链接:https://arxiv.org/abs/2412.06208
摘要:多模态语义通信集成了文本、图像和音频等各种数据模态,显著提高了通信效率和可靠性。在人工智能、自动驾驶、智能家居等领域具有广阔的应用前景。然而,目前的研究主要依赖于模拟信道,并假设恒定的信道状态(完美CSI),这是不足以解决动态物理信道和噪声在现实世界的场景。现有的方法往往集中在单一的模态任务,并未能处理多模态流数据,如视频和音频,以及它们相应的任务。此外,目前的语义编码和解码模块主要传输单一的模态特征,忽略了多模态语义增强和识别任务的需要。 为了解决这些挑战,本文提出了一个试点指导的框架,专门为视听事件本地化任务量身定制的多模态语义通信。该框架利用数字导频码和信道模型来指导真实场景中模拟信道的状态,并基于动态信道状态设计了考虑时频特性的基于欧拉的多模态语义编解码。该方法有效地处理多模态流源数据,特别是对于视听事件定位任务。大量的数值实验表明,所提出的框架在信道变化和支持各种通信场景的鲁棒性。实验结果表明,该框架优于现有的基准方法的信噪比(SNR),突出了其在语义通信质量的优势。
摘要:Multimodal semantic communication, which integrates various data modalitiessuch as text, images, and audio, significantly enhances communicationefficiency and reliability. It has broad application prospects in fields suchas artificial intelligence, autonomous driving, and smart homes. However,current research primarily relies on analog channels and assumes constantchannel states (perfect CSI), which is inadequate for addressing dynamicphysical channels and noise in real-world scenarios. Existing methods oftenfocus on single modality tasks and fail to handle multimodal stream data, suchas video and audio, and their corresponding tasks. Furthermore, currentsemantic encoding and decoding modules mainly transmit single modalityfeatures, neglecting the need for multimodal semantic enhancement andrecognition tasks. To address these challenges, this paper proposes a pilot-guided framework formultimodal semantic communication specifically tailored for audio-visual eventlocalization tasks. This framework utilizes digital pilot codes and channelmodules to guide the state of analog channels in real-wold scenarios anddesigns Euler-based multimodal semantic encoding and decoding that considertime-frequency characteristics based on dynamic channel state. This approacheffectively handles multimodal stream source data, especially for audio-visualevent localization tasks. Extensive numerical experiments demonstrate therobustness of the proposed framework in channel changes and its support forvarious communication scenarios. The experimental results show that theframework outperforms existing benchmark methods in terms of Signal-to-NoiseRatio (SNR), highlighting its advantage in semantic communication quality.
标题:M6:多生成器、多领域、多语言和文化、多流派、多乐器机器生成音乐检测数据库
链接:https://arxiv.org/abs/2412.06001
摘要:机器生成音乐(MGM)已经成为一种强大的工具,用于音乐治疗,个性化编辑和音乐社区的创作灵感。然而,它的不受管制的使用威胁着娱乐,教育和艺术部门,降低了高质量的人类作品的价值。因此,检测机器生成音乐(MGMD)对于保护这些领域至关重要,但该领域缺乏全面的数据集来支持有意义的进展。为了解决这一差距,我们引入了\textbf{M6},这是一个为MGMD研究量身定制的大规模基准数据集。M6以其多样性而闻名,包括多种生成器,领域,语言,文化背景,流派和工具。我们概述了我们的数据选择和收集方法,并附有详细的数据分析,提供所有WAV形式的音乐。此外,我们使用基本的二进制分类模型提供基线性能分数,说明了MGMD的复杂性和显著的改进空间。通过提供一个强大的和多方面的资源,我们的目标是使未来的研究开发更有效的检测方法MGM。我们相信,M6将成为应对这一社会挑战的关键一步。数据集和代码将免费提供,以支持该领域的开放式合作和创新。
摘要:Machine-generated music (MGM) has emerged as a powerful tool withapplications in music therapy, personalised editing, and creative inspirationfor the music community. However, its unregulated use threatens theentertainment, education, and arts sectors by diminishing the value ofhigh-quality human compositions. Detecting machine-generated music (MGMD) is,therefore, critical to safeguarding these domains, yet the field lackscomprehensive datasets to support meaningful progress. To address this gap, weintroduce \textbf{M6}, a large-scale benchmark dataset tailored for MGMDresearch. M6 is distinguished by its diversity, encompassing multiplegenerators, domains, languages, cultural contexts, genres, and instruments. Weoutline our methodology for data selection and collection, accompanied bydetailed data analysis, providing all WAV form of music. Additionally, weprovide baseline performance scores using foundational binary classificationmodels, illustrating the complexity of MGMD and the significant room forimprovement. By offering a robust and multifaceted resource, we aim to empowerfuture research to develop more effective detection methods for MGM. We believeM6 will serve as a critical step toward addressing this societal challenge. Thedataset and code will be freely available to support open collaboration andinnovation in this field.
标题:当视觉模型满足参数高效的后备适配器时,无需大规模音频预训练
链接:https://arxiv.org/abs/2412.05951
备注:5 pages, 3 figures
摘要:最近的研究表明,预训练的视觉模型可以提高音频下游任务的性能。为了进一步增强性能,通常需要使用大规模音频数据的额外预训练阶段,以将音频特定知识注入视觉模型中。然而,这种方法需要大量的音频数据和精心设计的目标函数。在这项工作中,我们建议绕过预训练阶段,直接使用我们的Look Aside Adapter(LoAA)对视觉模型进行微调,以实现高效的音频理解。音频频谱数据表示在两个异构的维度时间和频率,我们完善适配器,以促进跨这些维度的令牌之间的交互。我们的实验表明,我们的适配器允许视觉模型在各种音频和语音任务中达到或超过预训练音频模型的性能,为在音频应用中利用视觉模型提供了一种资源高效和有效的解决方案。
摘要:Recent studies show that pretrained vision models can boost performance inaudio downstream tasks. To enhance the performance further, an additionalpretraining stage with large scale audio data is typically required to infuseaudio specific knowledge into the vision model. However, such approachesrequire extensive audio data and a carefully designed objective function. Inthis work, we propose bypassing the pretraining stage by directly fine-tuningthe vision model with our Look Aside Adapter (LoAA) designed for efficientaudio understanding. Audio spectrum data is represented across twoheterogeneous dimensions time and frequency and we refine adapters tofacilitate interactions between tokens across these dimensions. Our experimentsdemonstrate that our adapters allow vision models to reach or surpass theperformance of pretrained audio models in various audio and speech tasks,offering a resource efficient and effective solution for leveraging visionmodels in audio applications.
标题:用于可控视频到音乐检索的半监督对比学习
链接:https://arxiv.org/abs/2412.05831
备注:4 pages + 1 reference page, 2 figures, 2 tables. Under review
摘要:内容创作者经常使用音乐来增强他们的视频,从电影中的配乐到视频博客和社交媒体内容中的背景音乐。然而,识别视频的最佳音乐可能是一项困难且耗时的任务。为了解决这一挑战,我们提出了一个新的框架,自动检索匹配的音乐剪辑为给定的视频,反之亦然。我们的方法利用了注释的音乐标签,以及视觉和音乐元素之间固有的艺术对应。不同于以往的跨模态音乐检索作品,我们的方法结合了自我监督和监督的训练目标。我们使用自监督和标签监督对比学习训练音乐和视频之间的联合嵌入空间。我们通过对监督训练组件使用音乐流派标签来展示我们方法的有效性,并且我们的框架可以推广到其他音乐注释(例如,情感、工具等)。此外,我们的方法可以细粒度地控制检索过程在推理时关注自监督信息与标签信息的程度。我们通过各种视频到音乐和音乐到视频检索任务来评估学习到的嵌入。实验表明,该方法成功地结合了自监督和监督目标,是有效的可控音乐视频检索。
摘要:Content creators often use music to enhance their videos, from soundtracks inmovies to background music in video blogs and social media content. However,identifying the best music for a video can be a difficult and time-consumingtask. To address this challenge, we propose a novel framework for automaticallyretrieving a matching music clip for a given video, and vice versa. Ourapproach leverages annotated music labels, as well as the inherent artisticcorrespondence between visual and music elements. Distinct from previouscross-modal music retrieval works, our method combines both self-supervised andsupervised training objectives. We use self-supervised and label-supervisedcontrastive learning to train a joint embedding space between music and video.We show the effectiveness of our approach by using music genre labels for thesupervised training component, and our framework can be generalized to othermusic annotations (e.g., emotion, instrument, etc.). Furthermore, our methodenables fine-grained control over how much the retrieval process focuses onself-supervised vs. label information at inference time. We evaluate thelearned embeddings through a variety of video-to-music and music-to-videoretrieval tasks. Our experiments show that the proposed approach successfullycombines self-supervised and supervised objectives and is effective forcontrollable music-video retrieval.
标题:将流派分类和和声打击特征与扩散模型相结合用于音乐视频生成
链接:https://arxiv.org/abs/2412.05694
摘要:本研究提出了一种新的方法,用于生成音乐可视化使用扩散模型,结合音频输入与用户选择的艺术品。该过程包括两个主要阶段:图像生成和视频创建。首先,音乐字幕和体裁分类,其次是检索的艺术风格的描述。扩散模型然后基于用户的输入图像和所导出的艺术风格描述来生成图像。视频生成阶段利用相同的扩散模型来内插帧,由从和声和和弦的关键音乐特征导出的音频能量矢量控制。该方法在各种流派中表现出良好的效果,并引入了一个新的度量标准,视听同步(AVS),以定量评估视觉和音频元素之间的同步。对比分析表明,与线性插值相比,使用所提出的方法与音频能量向量生成的视频的AVS值显着更高。这种方法在各个领域都有潜在的应用,包括独立音乐视频创作,电影制作,现场音乐活动以及增强公共空间的视听体验。
摘要:This study presents a novel method for generating music visualisers usingdiffusion models, combining audio input with user-selected artwork. The processinvolves two main stages: image generation and video creation. First, musiccaptioning and genre classification are performed, followed by the retrieval ofartistic style descriptions. A diffusion model then generates images based onthe user's input image and the derived artistic style descriptions. The videogeneration stage utilises the same diffusion model to interpolate frames,controlled by audio energy vectors derived from key musical features ofharmonics and percussives. The method demonstrates promising results acrossvarious genres, and a new metric, Audio-Visual Synchrony (AVS), is introducedto quantitatively evaluate the synchronisation between visual and audioelements. Comparative analysis shows significantly higher AVS values for videosgenerated using the proposed method with audio energy vectors, compared tolinear interpolation. This approach has potential applications in diversefields, including independent music video creation, film production, live musicevents, and enhancing audio-visual experiences in public spaces.
标题:WavFusion:迈向wav2vec 2.0多模式语音情感识别
链接:https://arxiv.org/abs/2412.05558
备注:Accepted by 31st International Conference on MultiMedia Modeling (MMM2025)
摘要:由于人类情感固有的复杂性和多样性,语音情感识别(SER)仍然是一个具有挑战性的关键任务。为了解决这个问题,研究人员试图通过多模态学习融合来自其他模态的信息。然而,现有的多模态融合技术往往忽略了跨模态相互作用的复杂性,从而导致次优的特征表示。在本文中,我们提出了WavFusion,多模态语音情感识别框架,解决关键的研究问题,有效的多模态融合,模态之间的异质性,和歧视性表示学习。通过利用门控跨模态注意力机制和多模态同质特征差异学习,WavFusion在基准数据集上展示了比现有最先进方法更好的性能。我们的工作强调了捕捉细微差别的跨模态交互和学习区分表示以获得准确的多模态SER的重要性。两个基准数据集(IEMOCAP和MELD)的实验结果表明,WavFusion在情感识别方面成功地超过了最先进的策略。
摘要:Speech emotion recognition (SER) remains a challenging yet crucial task dueto the inherent complexity and diversity of human emotions. To address thisproblem, researchers attempt to fuse information from other modalities viamultimodal learning. However, existing multimodal fusion techniques oftenoverlook the intricacies of cross-modal interactions, resulting in suboptimalfeature representations. In this paper, we propose WavFusion, a multimodalspeech emotion recognition framework that addresses critical research problemsin effective multimodal fusion, heterogeneity among modalities, anddiscriminative representation learning. By leveraging a gated cross-modalattention mechanism and multimodal homogeneous feature discrepancy learning,WavFusion demonstrates improved performance over existing state-of-the-artmethods on benchmark datasets. Our work highlights the importance of capturingnuanced cross-modal interactions and learning discriminative representationsfor accurate multimodal SER. Experimental results on two benchmark datasets(IEMOCAP and MELD) demonstrate that WavFusion succeeds over thestate-of-the-art strategies on emotion recognition.
标题:pyAMPACT:用于性能数据估计和多模式处理的分数音频对齐工具包
链接:https://arxiv.org/abs/2412.05436
备注:International Society for Music Information Retrieval, Late Breaking Demo
摘要:pyAMPACT(基于Python的自动音乐表现分析和比较工具包)链接符号和音频音乐表示,以促进对音频中的表现数据的分数信息估计,以及符号和音频音乐表示与各种注释的一般链接。pyAMPACT可以读取一系列符号格式,并可以将音符链接的音频描述符/性能数据输出到MEI格式的文件中。音频分析使用分数对齐来计算符号表示中的每个音符的重要性的时间-频率区域,从该时间-频率区域来估计参数的范围。这些包括调音、动态和音色相关的性能描述符,以及来自乐谱对齐的时序相关信息。除了性能数据估计之外,pyAMPACT还通过其将符号表示和注释链接到音频的基础设施来促进多模态调查。
摘要:pyAMPACT (Python-based Automatic Music Performance Analysis and ComparisonToolkit) links symbolic and audio music representations to facilitatescore-informed estimation of performance data in audio as well as generallinking of symbolic and audio music representations with a variety ofannotations. pyAMPACT can read a range of symbolic formats and can outputnote-linked audio descriptors/performance data into MEI-formatted files. Theaudio analysis uses score alignment to calculate time-frequency regions ofimportance for each note in the symbolic representation from which to estimatea range of parameters. These include tuning-, dynamics-, and timbre-relatedperformance descriptors, with timing-related information available from thescore alignment. Beyond performance data estimation, pyAMPACT also facilitatesmulti-modal investigations through its infrastructure for linking symbolicrepresentations and annotations to audio.
标题:重温你的记忆:通过脑电引导的视听生成重建情感背景化记忆
链接:https://arxiv.org/abs/2412.05296
备注:Codes and the dataset will be released upon acceptance
摘要:在本文中,我们介绍了RecallAffectiveMemory,一个新的任务,旨在重建自传体记忆,通过视听生成引导的影响提取的脑电图(EEG)信号。为了支持这一开创性的任务,我们提出了EEG-AffectiveMemory数据集,其中包括文本描述,视觉,音乐和EEG记录,这些记录是在9名参与者的记忆回忆过程中收集的。此外,我们提出了RYM(回忆你的记忆),一个三阶段的框架,用于生成同步的视听内容,同时保持动态的个人记忆影响轨迹。实验结果表明,我们的方法可以在定性和定量方面忠实地重建所有受试者的情感情境化视听记忆,参与者报告他们回忆的记忆和生成的内容之间存在很强的情感一致性。我们的方法通过基于神经的情感理解推进了情感解码研究及其在个性化媒体创作中的实际应用。
摘要:In this paper, we introduce RecallAffectiveMemory, a novel task designed toreconstruct autobiographical memories through audio-visual generation guided byaffect extracted from electroencephalogram (EEG) signals. To support thispioneering task, we present the EEG-AffectiveMemory dataset, which encompassestextual descriptions, visuals, music, and EEG recordings collected duringmemory recall from nine participants. Furthermore, we propose RYM (Recall YourMemory), a three-stage framework for generating synchronized audio-visualcontents while maintaining dynamic personal memory affect trajectories.Experimental results indicate that our method can faithfully reconstructaffect-contextualized audio-visual memory across all subjects, bothqualitatively and quantitatively, with participants reporting strong affectiveconcordance between their recalled memories and the generated content. Ourapproaches advance affect decoding research and its practical applications inpersonalized media creation via neural-based affect comprehension.
