今日论文合集:cs.SD语音5篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Mellow: a small audio language model for reasoning
标题:MASYS:用于推理的小型音频语言模型
链接:https://arxiv.org/abs/2503.08540
作者:Soham Deshmukh,  Satvik Dixit,  Rita Singh,  Bhiksha Raj
备注:Checkpoint and dataset available at: this https URL
摘要:多模态音频语言模型(ALM)可以理解和推理音频和文本。通常情况下,推理性能与模型大小相关,超过80亿个参数的模型可以获得最佳结果。然而,尽管边缘设备具有潜在的应用,但之前没有研究过使小型音频语言模型能够执行推理任务。为了解决这一差距,我们引入了一个专门为推理设计的小型音频语言模型。MPEG4在现有的小型音频语言模型中实现了最先进的性能,并在推理能力方面超过了几个较大的模型。例如,MMAU的MPEG4得分为52.11,与SoTA Qwen 2 Audio(得分为52.5)相当,但使用的参数少了50倍,训练的数据少了60倍(音频小时)。为了训练Mogram,我们引入了ReasonAQA,这是一个旨在增强模型中基于音频的推理的数据集。它由现有数据集(30%的数据)和合成生成的数据(70%)组成。合成数据集来自音频字幕数据集,其中大型语言模型(LLM)生成详细的多项选择题,重点关注音频事件,对象,声学场景,信号属性,语义和听众情绪。为了评估Mackay的推理能力,我们在一组不同的任务上对其进行基准测试,评估分布内和分布外的数据,包括音频理解,演绎推理和比较推理。最后,我们进行了广泛的消融研究,以探索投影层选择,合成数据生成方法和语言模型预训练对推理性能的影响。我们的训练数据集、发现和基线为开发能够推理的小型ALM铺平了道路。
摘要:Multimodal Audio-Language Models (ALMs) can understand and reason over bothaudio and text. Typically, reasoning performance correlates with model size,with the best results achieved by models exceeding 8 billion parameters.However, no prior work has explored enabling small audio-language models toperform reasoning tasks, despite the potential applications for edge devices.To address this gap, we introduce Mellow, a small Audio-Language Modelspecifically designed for reasoning. Mellow achieves state-of-the-artperformance among existing small audio-language models and surpasses severallarger models in reasoning capabilities. For instance, Mellow scores 52.11 onMMAU, comparable to SoTA Qwen2 Audio (which scores 52.5) while using 50 timesfewer parameters and being trained on 60 times less data (audio hrs). To trainMellow, we introduce ReasonAQA, a dataset designed to enhance audio-groundedreasoning in models. It consists of a mixture of existing datasets (30% of thedata) and synthetically generated data (70%). The synthetic dataset is derivedfrom audio captioning datasets, where Large Language Models (LLMs) generatedetailed and multiple-choice questions focusing on audio events, objects,acoustic scenes, signal properties, semantics, and listener emotions. Toevaluate Mellow's reasoning ability, we benchmark it on a diverse set of tasks,assessing on both in-distribution and out-of-distribution data, including audiounderstanding, deductive reasoning, and comparative reasoning. Finally, weconduct extensive ablation studies to explore the impact of projection layerchoices, synthetic data generation methods, and language model pretraining onreasoning performance. Our training dataset, findings, and baseline pave theway for developing small ALMs capable of reasoning.

【2】 ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
标题:ESPnet-SDP:口语对话系统的统一工具包和演示
链接:https://arxiv.org/abs/2503.08533
作者:Siddhant Arora,  Yifan Peng,  Jiatong Shi,  Jinchuan Tian,  William Chen,  Shikhar Bharadwaj,  Hayato Futami,  Yosuke Kashiwagi,  Emiru Tsunoo,  Shuichiro Shimizu,  Vaibhav Srivastav,  Shinji Watanabe
备注:Accepted at NAACL 2025 Demo Track
摘要:None
摘要:Advancements in audio foundation models (FMs) have fueled interest inend-to-end (E2E) spoken dialogue systems, but different web interfaces for eachsystem makes it challenging to compare and contrast them effectively. Motivatedby this, we introduce an open-source, user-friendly toolkit designed to buildunified web interfaces for various cascaded and E2E spoken dialogue systems.Our demo further provides users with the option to get on-the-fly automatedevaluation metrics such as (1) latency, (2) ability to understand user input,(3) coherence, diversity, and relevance of system response, and (4)intelligibility and audio quality of system output. Using the evaluationmetrics, we compare various cascaded and E2E spoken dialogue systems with ahuman-human conversation dataset as a proxy. Our analysis demonstrates that thetoolkit allows researchers to effortlessly compare and contrast differenttechnologies, providing valuable insights such as current E2E systems havingpoorer audio quality and less diverse responses. An example demo produced usingour toolkit is publicly available here:https://huggingface.co/spaces/Siddhant/Voice_Assistant_Demo.

【3】 FilmComposer: LLM-Driven Music Production for Silent Film Clips
标题:FilmComposer:LLM驱动的无声电影剪辑音乐制作
链接:https://arxiv.org/abs/2503.08147
作者:Zhifeng Xie,  Qile He,  Youjia Zhu,  Qiwei He,  Mengtian Li
备注:Project page: this https URL
摘要:在这项工作中,我们实现了无声电影剪辑使用LLM驱动的方法的音乐制作。鉴于电影音乐制作的强烈专业需求,我们提出了FilmComposer,模拟专业音乐家的实际工作流程。FilmComposer是第一个将大型生成模型与多智能体方法相结合的软件,利用了波形音乐和符号音乐生成的优势。此外,FilmComposer是第一个专注于电影音乐制作的三个核心要素-音频质量,音乐性和音乐发展-并引入了各种控制,如节奏,语义和视觉效果,以增强这些关键方面。具体来说,FilmComposer包括视觉处理模块,节奏可控的MusicGen,以及多代理评估,安排和混合。此外,我们的框架可以无缝集成到实际的音乐制作管道中,并允许用户在每一步进行干预,提供强大的交互性和高度的创作自由。此外,考虑到缺乏专业和高质量的电影音乐数据集,我们提出了MusicPro-7 k,其中包括7,418个电影片段,音乐,描述,节奏点和主旋律。最后,我们提出的标准指标和新的专业指标都表明,我们的模型生成的音乐在质量、与视频的一致性、多样性、音乐性和音乐发展方面达到了最先进的性能。项目页面:https://apple-jun.github.io/FilmComposer.github.io/
摘要:In this work, we implement music production for silent film clips usingLLM-driven method. Given the strong professional demands of film musicproduction, we propose the FilmComposer, simulating the actual workflows ofprofessional musicians. FilmComposer is the first to combine large generativemodels with a multi-agent approach, leveraging the advantages of both waveformmusic and symbolic music generation. Additionally, FilmComposer is the first tofocus on the three core elements of music production for film-audio quality,musicality, and musical development-and introduces various controls, such asrhythm, semantics, and visuals, to enhance these key aspects. Specifically,FilmComposer consists of the visual processing module, rhythm-controllableMusicGen, and multi-agent assessment, arrangement and mix. In addition, ourframework can seamlessly integrate into the actual music production pipelineand allows user intervention in every step, providing strong interactivity anda high degree of creative freedom. Furthermore, we propose MusicPro-7k whichincludes 7,418 film clips, music, description, rhythm spots and main melody,considering the lack of a professional and high-quality film music dataset.Finally, both the standard metrics and the new specialized metrics we proposedemonstrate that the music generated by our model achieves state-of-the-artperformance in terms of quality, consistency with video, diversity, musicality,and musical development. Project page:https://apple-jun.github.io/FilmComposer.github.io/

【4】 Boundary Regression for Leitmotif Detection in Music Audio
标题:边界回归用于音乐音频中的主旋律检测
链接:https://arxiv.org/abs/2503.07977
作者:Sihun Lee,  Dasaem Jeong
备注:2 pages, 1 figure; presented at the 2024 ISMIR conference Late-Breaking Demo
摘要:主旋律是以各种形式在整首乐曲中重复出现的乐句。由于不同的变化和仪器,检测从音频记录的leitmotifs的发生是一个极具挑战性的任务。主基序检测可以被处理为音频事件检测的子类别,其中主基序活动在帧级被预测。然而,由于主旋律体现了独特的、连贯的音乐结构,因此在视觉对象检测中,类似于边界框回归的更全面的方法可能会有所帮助。这种方法捕捉了整个主题,而不是将其分割成单个帧,从而保持了其音乐的完整性并产生了更有用的预测。我们提出了我们的实验结果,解决主旋律检测作为一个边界回归任务。
摘要:Leitmotifs are musical phrases that are reprised in various forms throughouta piece. Due to diverse variations and instrumentation, detecting theoccurrence of leitmotifs from audio recordings is a highly challenging task.Leitmotif detection may be handled as a subcategory of audio event detection,where leitmotif activity is predicted at the frame level. However, asleitmotifs embody distinct, coherent musical structures, a more holisticapproach akin to bounding box regression in visual object detection can behelpful. This method captures the entirety of a motif rather than fragmentingit into individual frames, thereby preserving its musical integrity andproducing more useful predictions. We present our experimental results ontackling leitmotif detection as a boundary regression task.

【5】 YuE: Scaling Open Foundation Models for Long-Form Music Generation
标题:YuE:扩展开放基金会模型以促进长篇音乐一代
链接:https://arxiv.org/abs/2503.08638

作者:Ruibin Yuan,  Hanfeng Lin,  Shuyue Guo,  Ge Zhang,  Jiahao Pan,  Yongyi Zang,  Haohe Liu,  Yiming Liang,  Wenye Ma,  Xingjian Du,  Xinrun Du,  Zhen Ye,  Tianyu Zheng,  Yinghao Ma,  Minghao Liu,  Zeyue Tian,  Ziya Zhou,  Liumeng Xue,  Xingwei Qu,  Yizhi Li,  Shangda Wu,  Tianhao Shen,  Ziyang Ma,  Jun Zhan,  Chunhui Wang,  Yatian Wang,  Xiaowei Chi,  Xinyue Zhang,  Zhenzhu Yang,  Xiangzhou Wang,  Shansong Liu,  Lingrui Mei,  Peng Li,  Junjie Wang,  Jianwei Yu,  Guojian Pang,  Xu Li,  Zihao Wang,  Xiaohuan Zhou,  Lijun Yu,  Emmanouil Benetos,  Yong Chen,  Chenghua Lin,  Xie Chen,  Gus Xia,  Zhaoxiang Zhang,  Chao Zhang,  Wenhu Chen,  Xinyu Zhou,  Xipeng Qiu,  Roger Dannenberg,  Jiaheng Liu,  Jian Yang,  Wenhao Huang,  Wei Xue,  Xu Tan,  Yike Guo
备注:this https URL
摘要:我们解决了长格式音乐生成的任务-特别是具有挑战性的\textbf{歌词到歌曲}的问题-通过引入YuE,一个开放的基础模型的LLaMA 2架构的基础上的家庭。具体来说,YuE可扩展到数万亿个令牌,并生成长达五分钟的音乐,同时保持抒情对齐,连贯的音乐结构,并在适当的伴奏下进行声乐旋律。它通过以下方式实现这一点:(1)轨道解耦的下一个令牌预测,以克服密集的混合信号;(2)结构渐进式条件反射,用于长上下文歌词对齐;以及(3)多任务,多阶段预训练配方,以收敛和概括。此外,我们重新设计了用于音乐生成的上下文学习技术,实现了多功能的风格迁移(例如,将日本城市流行音乐转换为英语说唱,同时保留原始伴奏)和双向生成。通过广泛的评估,我们证明了YuE在音乐性和声乐敏捷性方面与某些专有系统相匹配甚至超越。此外,对YuE进行微调可以实现额外的控件并增强对尾部语言的支持。此外,除了生成,我们表明YuE的学习表示可以很好地执行音乐理解任务,YuE的结果匹配或超过最先进的方法在MARBLE基准。关键词:lyrics 2song,歌曲生成,长格式,基础模型,音乐生成
摘要:We tackle the task of long-form music generation--particularly thechallenging \textbf{lyrics-to-song} problem--by introducing YuE, a family ofopen foundation models based on the LLaMA2 architecture. Specifically, YuEscales to trillions of tokens and generates up to five minutes of music whilemaintaining lyrical alignment, coherent musical structure, and engaging vocalmelodies with appropriate accompaniment. It achieves this through (1)track-decoupled next-token prediction to overcome dense mixture signals, (2)structural progressive conditioning for long-context lyrical alignment, and (3)a multitask, multiphase pre-training recipe to converge and generalize. Inaddition, we redesign the in-context learning technique for music generation,enabling versatile style transfer (e.g., converting Japanese city pop into anEnglish rap while preserving the original accompaniment) and bidirectionalgeneration. Through extensive evaluation, we demonstrate that YuE matches oreven surpasses some of the proprietary systems in musicality and vocal agility.In addition, fine-tuning YuE enables additional controls and enhanced supportfor tail languages. Furthermore, beyond generation, we show that YuE's learnedrepresentations can perform well on music understanding tasks, where theresults of YuE match or exceed state-of-the-art methods on the MARBLEbenchmark. Keywords: lyrics2song, song generation, long-form, foundation model,music generation


eess.AS音频处理

【1】 YuE: Scaling Open Foundation Models for Long-Form Music Generation
标题:YuE:扩展开放基金会模型以促进长篇音乐一代
链接:https://arxiv.org/abs/2503.08638

作者:Ruibin Yuan,  Hanfeng Lin,  Shuyue Guo,  Ge Zhang,  Jiahao Pan,  Yongyi Zang,  Haohe Liu,  Yiming Liang,  Wenye Ma,  Xingjian Du,  Xinrun Du,  Zhen Ye,  Tianyu Zheng,  Yinghao Ma,  Minghao Liu,  Zeyue Tian,  Ziya Zhou,  Liumeng Xue,  Xingwei Qu,  Yizhi Li,  Shangda Wu,  Tianhao Shen,  Ziyang Ma,  Jun Zhan,  Chunhui Wang,  Yatian Wang,  Xiaowei Chi,  Xinyue Zhang,  Zhenzhu Yang,  Xiangzhou Wang,  Shansong Liu,  Lingrui Mei,  Peng Li,  Junjie Wang,  Jianwei Yu,  Guojian Pang,  Xu Li,  Zihao Wang,  Xiaohuan Zhou,  Lijun Yu,  Emmanouil Benetos,  Yong Chen,  Chenghua Lin,  Xie Chen,  Gus Xia,  Zhaoxiang Zhang,  Chao Zhang,  Wenhu Chen,  Xinyu Zhou,  Xipeng Qiu,  Roger Dannenberg,  Jiaheng Liu,  Jian Yang,  Wenhao Huang,  Wei Xue,  Xu Tan,  Yike Guo
备注:this https URL
摘要:我们解决了长格式音乐生成的任务-特别是具有挑战性的\textbf{歌词到歌曲}的问题-通过引入YuE,一个开放的基础模型的LLaMA 2架构的基础上的家庭。具体来说,YuE可扩展到数万亿个令牌,并生成长达五分钟的音乐,同时保持抒情对齐,连贯的音乐结构,并在适当的伴奏下进行声乐旋律。它通过以下方式实现这一点:(1)轨道解耦的下一个令牌预测,以克服密集的混合信号;(2)结构渐进式条件反射,用于长上下文歌词对齐;以及(3)多任务,多阶段预训练配方,以收敛和概括。此外,我们重新设计了用于音乐生成的上下文学习技术,实现了多功能的风格迁移(例如,将日本城市流行音乐转换为英语说唱,同时保留原始伴奏)和双向生成。通过广泛的评估,我们证明了YuE在音乐性和声乐敏捷性方面与某些专有系统相匹配甚至超越。此外,对YuE进行微调可以实现额外的控件并增强对尾部语言的支持。此外,除了生成,我们表明YuE的学习表示可以很好地执行音乐理解任务,YuE的结果匹配或超过最先进的方法在MARBLE基准。关键词:lyrics 2song,歌曲生成,长格式,基础模型,音乐生成
摘要:We tackle the task of long-form music generation--particularly thechallenging \textbf{lyrics-to-song} problem--by introducing YuE, a family ofopen foundation models based on the LLaMA2 architecture. Specifically, YuEscales to trillions of tokens and generates up to five minutes of music whilemaintaining lyrical alignment, coherent musical structure, and engaging vocalmelodies with appropriate accompaniment. It achieves this through (1)track-decoupled next-token prediction to overcome dense mixture signals, (2)structural progressive conditioning for long-context lyrical alignment, and (3)a multitask, multiphase pre-training recipe to converge and generalize. Inaddition, we redesign the in-context learning technique for music generation,enabling versatile style transfer (e.g., converting Japanese city pop into anEnglish rap while preserving the original accompaniment) and bidirectionalgeneration. Through extensive evaluation, we demonstrate that YuE matches oreven surpasses some of the proprietary systems in musicality and vocal agility.In addition, fine-tuning YuE enables additional controls and enhanced supportfor tail languages. Furthermore, beyond generation, we show that YuE's learnedrepresentations can perform well on music understanding tasks, where theresults of YuE match or exceed state-of-the-art methods on the MARBLEbenchmark. Keywords: lyrics2song, song generation, long-form, foundation model,music generation


【2】 Mellow: a small audio language model for reasoning
标题:MASYS:用于推理的小型音频语言模型
链接:https://arxiv.org/abs/2503.08540
作者:Soham Deshmukh,  Satvik Dixit,  Rita Singh,  Bhiksha Raj
备注:Checkpoint and dataset available at: this https URL
摘要:多模态音频语言模型(ALM)可以理解和推理音频和文本。通常情况下,推理性能与模型大小相关,超过80亿个参数的模型可以获得最佳结果。然而,尽管边缘设备具有潜在的应用,但之前没有研究过使小型音频语言模型能够执行推理任务。为了解决这一差距,我们引入了一个专门为推理设计的小型音频语言模型。MPEG4在现有的小型音频语言模型中实现了最先进的性能,并在推理能力方面超过了几个较大的模型。例如,MMAU的MPEG4得分为52.11,与SoTA Qwen 2 Audio(得分为52.5)相当,但使用的参数少了50倍,训练的数据少了60倍(音频小时)。为了训练Mogram,我们引入了ReasonAQA,这是一个旨在增强模型中基于音频的推理的数据集。它由现有数据集(30%的数据)和合成生成的数据(70%)组成。合成数据集来自音频字幕数据集,其中大型语言模型(LLM)生成详细的多项选择题,重点关注音频事件,对象,声学场景,信号属性,语义和听众情绪。为了评估Mackay的推理能力,我们在一组不同的任务上对其进行基准测试,评估分布内和分布外的数据,包括音频理解,演绎推理和比较推理。最后,我们进行了广泛的消融研究,以探索投影层选择,合成数据生成方法和语言模型预训练对推理性能的影响。我们的训练数据集、发现和基线为开发能够推理的小型ALM铺平了道路。
摘要:Multimodal Audio-Language Models (ALMs) can understand and reason over bothaudio and text. Typically, reasoning performance correlates with model size,with the best results achieved by models exceeding 8 billion parameters.However, no prior work has explored enabling small audio-language models toperform reasoning tasks, despite the potential applications for edge devices.To address this gap, we introduce Mellow, a small Audio-Language Modelspecifically designed for reasoning. Mellow achieves state-of-the-artperformance among existing small audio-language models and surpasses severallarger models in reasoning capabilities. For instance, Mellow scores 52.11 onMMAU, comparable to SoTA Qwen2 Audio (which scores 52.5) while using 50 timesfewer parameters and being trained on 60 times less data (audio hrs). To trainMellow, we introduce ReasonAQA, a dataset designed to enhance audio-groundedreasoning in models. It consists of a mixture of existing datasets (30% of thedata) and synthetically generated data (70%). The synthetic dataset is derivedfrom audio captioning datasets, where Large Language Models (LLMs) generatedetailed and multiple-choice questions focusing on audio events, objects,acoustic scenes, signal properties, semantics, and listener emotions. Toevaluate Mellow's reasoning ability, we benchmark it on a diverse set of tasks,assessing on both in-distribution and out-of-distribution data, including audiounderstanding, deductive reasoning, and comparative reasoning. Finally, weconduct extensive ablation studies to explore the impact of projection layerchoices, synthetic data generation methods, and language model pretraining onreasoning performance. Our training dataset, findings, and baseline pave theway for developing small ALMs capable of reasoning.

【3】 ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
标题:ESPnet-SDP:口语对话系统的统一工具包和演示
链接:https://arxiv.org/abs/2503.08533
作者:Siddhant Arora,  Yifan Peng,  Jiatong Shi,  Jinchuan Tian,  William Chen,  Shikhar Bharadwaj,  Hayato Futami,  Yosuke Kashiwagi,  Emiru Tsunoo,  Shuichiro Shimizu,  Vaibhav Srivastav,  Shinji Watanabe
备注:Accepted at NAACL 2025 Demo Track
摘要:音频基础模型(FM)的进步激发了人们对端到端(E2 E)口语对话系统的兴趣,但每个系统的不同Web界面使得有效地比较和对比它们变得具有挑战性。出于这一动机,我们介绍了一个开源的,用户友好的工具包,旨在建立各种级联和E2 E口语对话系统的统一的Web界面。我们的演示还为用户提供了获得实时自动评估指标的选项,例如(1)延迟,(2)理解用户输入的能力,(3)系统响应的一致性,多样性和相关性,以及(4)系统输出的可理解性和音频质量。使用评估指标,我们比较了各种级联和E2 E口语对话系统与人-人对话数据集作为代理。我们的分析表明,该工具包使研究人员能够毫不费力地比较和对比不同的技术,提供有价值的见解,如当前的E2 E系统具有较差的音频质量和较少的多样性的反应。使用我们的工具包生成的示例演示可在此处公开获得:https://huggingface.co/spaces/Siddhant/Voice_Assistant_Demo。
摘要:Advancements in audio foundation models (FMs) have fueled interest inend-to-end (E2E) spoken dialogue systems, but different web interfaces for eachsystem makes it challenging to compare and contrast them effectively. Motivatedby this, we introduce an open-source, user-friendly toolkit designed to buildunified web interfaces for various cascaded and E2E spoken dialogue systems.Our demo further provides users with the option to get on-the-fly automatedevaluation metrics such as (1) latency, (2) ability to understand user input,(3) coherence, diversity, and relevance of system response, and (4)intelligibility and audio quality of system output. Using the evaluationmetrics, we compare various cascaded and E2E spoken dialogue systems with ahuman-human conversation dataset as a proxy. Our analysis demonstrates that thetoolkit allows researchers to effortlessly compare and contrast differenttechnologies, providing valuable insights such as current E2E systems havingpoorer audio quality and less diverse responses. An example demo produced usingour toolkit is publicly available here:https://huggingface.co/spaces/Siddhant/Voice_Assistant_Demo.

【4】 FilmComposer: LLM-Driven Music Production for Silent Film Clips
标题:FilmComposer:LLM驱动的无声电影剪辑音乐制作
链接:https://arxiv.org/abs/2503.08147
作者:Zhifeng Xie,  Qile He,  Youjia Zhu,  Qiwei He,  Mengtian Li
备注:Project page: this https URL
摘要:在这项工作中,我们实现了无声电影剪辑使用LLM驱动的方法的音乐制作。鉴于电影音乐制作的强烈专业需求,我们提出了FilmComposer,模拟专业音乐家的实际工作流程。FilmComposer是第一个将大型生成模型与多智能体方法相结合的软件,利用了波形音乐和符号音乐生成的优势。此外,FilmComposer是第一个专注于电影音乐制作的三个核心要素-音频质量,音乐性和音乐发展-并引入了各种控制,如节奏,语义和视觉效果,以增强这些关键方面。具体来说,FilmComposer包括视觉处理模块,节奏可控的MusicGen,以及多代理评估,安排和混合。此外,我们的框架可以无缝集成到实际的音乐制作管道中,并允许用户在每一步进行干预,提供强大的交互性和高度的创作自由。此外,考虑到缺乏专业和高质量的电影音乐数据集,我们提出了MusicPro-7 k,其中包括7,418个电影片段,音乐,描述,节奏点和主旋律。最后,我们提出的标准指标和新的专业指标都表明,我们的模型生成的音乐在质量、与视频的一致性、多样性、音乐性和音乐发展方面达到了最先进的性能。项目页面:https://apple-jun.github.io/FilmComposer.github.io/
摘要:In this work, we implement music production for silent film clips usingLLM-driven method. Given the strong professional demands of film musicproduction, we propose the FilmComposer, simulating the actual workflows ofprofessional musicians. FilmComposer is the first to combine large generativemodels with a multi-agent approach, leveraging the advantages of both waveformmusic and symbolic music generation. Additionally, FilmComposer is the first tofocus on the three core elements of music production for film-audio quality,musicality, and musical development-and introduces various controls, such asrhythm, semantics, and visuals, to enhance these key aspects. Specifically,FilmComposer consists of the visual processing module, rhythm-controllableMusicGen, and multi-agent assessment, arrangement and mix. In addition, ourframework can seamlessly integrate into the actual music production pipelineand allows user intervention in every step, providing strong interactivity anda high degree of creative freedom. Furthermore, we propose MusicPro-7k whichincludes 7,418 film clips, music, description, rhythm spots and main melody,considering the lack of a professional and high-quality film music dataset.Finally, both the standard metrics and the new specialized metrics we proposedemonstrate that the music generated by our model achieves state-of-the-artperformance in terms of quality, consistency with video, diversity, musicality,and musical development. Project page:https://apple-jun.github.io/FilmComposer.github.io/

【5】 Boundary Regression for Leitmotif Detection in Music Audio
标题:边界回归用于音乐音频中的主旋律检测
链接:https://arxiv.org/abs/2503.07977
作者:Sihun Lee,  Dasaem Jeong
备注:2 pages, 1 figure; presented at the 2024 ISMIR conference Late-Breaking Demo
摘要:主旋律是以各种形式在整首乐曲中重复出现的乐句。由于不同的变化和仪器,检测从音频记录的leitmotifs的发生是一个极具挑战性的任务。主基序检测可以被处理为音频事件检测的子类别,其中主基序活动在帧级被预测。然而,由于主旋律体现了独特的、连贯的音乐结构,因此在视觉对象检测中,类似于边界框回归的更全面的方法可能会有所帮助。这种方法捕捉了整个主题,而不是将其分割成单个帧,从而保持了其音乐的完整性并产生了更有用的预测。我们提出了我们的实验结果,解决主旋律检测作为一个边界回归任务。
摘要:Leitmotifs are musical phrases that are reprised in various forms throughouta piece. Due to diverse variations and instrumentation, detecting theoccurrence of leitmotifs from audio recordings is a highly challenging task.Leitmotif detection may be handled as a subcategory of audio event detection,where leitmotif activity is predicted at the frame level. However, asleitmotifs embody distinct, coherent musical structures, a more holisticapproach akin to bounding box regression in visual object detection can behelpful. This method captures the entirety of a motif rather than fragmentingit into individual frames, thereby preserving its musical integrity andproducing more useful predictions. We present our experimental results ontackling leitmotif detection as a boundary regression task.

机器翻译由腾讯交互翻译提供,仅供参考