今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
标题: MuChoMusic:在多模式音频语言模型中评估音乐理解
作者:Benno Weck,Ilaria Manco,Emmanouil Benetos,Elio Quinton,George Fazekas,Dmitry Bogdanov
备注:Accepted at ISMIR 2024. Data: this https URL Code: this https URL Supplementary material: this https URL
链接:点击下载PDF文件
摘要:联合处理音频和语言的多模态模型在音频理解中具有很大的前景,并且越来越多地被应用于音乐领域。通过允许用户通过文本进行查询并获得有关给定音频输入的信息,这些模型有可能通过基于语言的界面实现各种音乐理解任务。然而,他们的评估带来了相当大的挑战,目前还不清楚如何有效地评估他们的能力,正确地解释与音乐相关的输入与当前的方法。出于这一动机,我们介绍了MuChoMusic,这是一个用于评估多模态语言模型中音乐理解的基准,专注于音频。MuChoMusic包含1,187个多项选择题,全部由人工注释员验证,来自两个公开的音乐数据集的644首音乐曲目,涵盖各种类型。基准中的问题是精心设计的,以评估知识和推理能力,涵盖基本的音乐概念及其与文化和功能背景的关系。通过基准提供的整体分析,我们评估了五个开源模型,并确定了几个陷阱,包括过度依赖语言模态,指出需要更好的多模态集成。数据和代码都是开源的。摘要:Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about a given audio input, these models have the potential to enable a variety of music understanding tasks via language-based interfaces. However, their evaluation poses considerable challenges, and it remains unclear how to effectively assess their ability to correctly interpret music-related inputs with current methods. Motivated by this, we introduce MuChoMusic, a benchmark for evaluating music understanding in multimodal language models focused on audio. MuChoMusic comprises 1,187 multiple-choice questions, all validated by human annotators, on 644 music tracks sourced from two publicly available music datasets, and covering a wide variety of genres. Questions in the benchmark are crafted to assess knowledge and reasoning abilities across several dimensions that cover fundamental musical concepts and their relation to cultural and functional contexts. Through the holistic analysis afforded by the benchmark, we evaluate five open-source models and identify several pitfalls, including an over-reliance on the language modality, pointing to a need for better multimodal integration. Data and code are open-sourced.

【2】 Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
标题: 视听广义Zero-Shot学习的分布外检测:一个通用框架
作者:Liuyuan Wen
链接:点击下载PDF文件
摘要:广义零次学习(Generalized Zero-Shot Learning,简称GSHL)是一项具有挑战性的任务,需要对可见和不可见的类进行准确分类。在这一领域,视听语言作为一个非常令人兴奋但困难的任务出现,考虑到包括视觉和声学特征作为多模态输入。在这一领域的现有努力主要利用嵌入为基础的或生成为基础的方法。然而,生成式训练困难且不稳定,而基于嵌入的方法经常遇到域偏移问题。因此,我们发现将这两种方法集成到一个统一的框架中,以利用它们的优点,同时减轻各自的缺点是有希望的。我们的研究介绍了一个一般框架,采用了分布(OOD)检测,旨在利用这两种方法的优势。我们首先使用生成对抗网络来合成看不见的特征,从而能够训练OOD检测器以及可见和不可见类的分类器。该检测器确定测试特征是否属于可见或不可见的类,然后对每个特征类型使用单独的分类器进行分类。我们在三个流行的视听数据集上测试了我们的框架,并观察到与现有的最先进的作品相比有显着的改进。代码可以在https: github.com liuyuan-wen AV-OOD-GZSL中找到。摘要:Generalized Zero-Shot Learning (GZSL) is a challenging task requiring accurate classification of both seen and unseen classes. Within this domain, Audio-visual GZSL emerges as an extremely exciting yet difficult task, given the inclusion of both visual and acoustic features as multi-modal inputs. Existing efforts in this field mostly utilize either embedding-based or generative-based methods. However, generative training is difficult and unstable, while embedding-based methods often encounter domain shift problem. Thus, we find it promising to integrate both methods into a unified framework to leverage their advantages while mitigating their respective disadvantages. Our study introduces a general framework employing out-of-distribution (OOD) detection, aiming to harness the strengths of both approaches. We first employ generative adversarial networks to synthesize unseen features, enabling the training of an OOD detector alongside classifiers for seen and unseen classes. This detector determines whether a test feature belongs to seen or unseen classes, followed by classification utilizing separate classifiers for each feature type. We test our framework on three popular audio-visual datasets and observe a significant improvement comparing to existing state-of-the-art works. Codes can be found in https: github.com liuyuan-wen AV-OOD-GZSL.

【3】 Nested Music Transformer: Sequentially Decoding Compound Tokens in Symbolic Music and Audio Generation
标题: 嵌套音乐Transformer:在符号音乐和音频生成中顺序解码复合代币
作者:Jiwoo Ryu,Hao-Wen Dong,Jongmin Jung,Dasaem Jeong
备注:Accepted at 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:用复合记号表示符号音乐,其中每个记号由表示不同音乐特征或属性的若干不同子记号组成,提供了减少序列长度的优点。虽然之前的研究已经验证了复合标记在音乐序列建模中的有效性,但同时预测所有子标记可能会导致次优结果,因为它可能无法完全捕获它们之间的相互依赖性。我们介绍嵌套的音乐Transformer(NMT),为解码复合令牌自回归,类似于处理扁平令牌,但具有低内存使用量的架构。NMT由两个Transformers组成:对复合令牌序列进行建模的主解码器和对每个复合令牌的子令牌进行建模的子解码器。实验结果表明,NMT应用于复合令牌可以提高性能,在处理各种符号音乐数据集和离散音频令牌从MAESTRO数据集更好的困惑。摘要:Representing symbolic music with compound tokens, where each token consists of several different sub-tokens representing a distinct musical feature or attribute, offers the advantage of reducing sequence length. While previous research has validated the efficacy of compound tokens in music sequence modeling, predicting all sub-tokens simultaneously can lead to suboptimal results as it may not fully capture the interdependencies between them. We introduce the Nested Music Transformer (NMT), an architecture tailored for decoding compound tokens autoregressively, similar to processing flattened tokens, but with low memory usage. The NMT consists of two transformers: the main decoder that models a sequence of compound tokens and the sub-decoder for modeling sub-tokens of each compound token. The experiment results showed that applying the NMT to compound tokens can enhance the performance in terms of better perplexity in processing various symbolic music datasets and discrete audio tokens from the MAESTRO dataset.

【4】 Six Dragons Fly Again: Reviving 15th-Century Korean Court Music with Transformers and Novel Encoding
标题: 六龙再次飞翔:用Transformer和新颖编码复兴15世纪韩国宫廷音乐
作者:Danbinaerin Han,Mark Gotham,Dongmin Kim,Hannah Park,Sihun Lee,Dasaem Jeong
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:我们介绍一个项目,复兴了15世纪韩国宫廷音乐,Chihwapyeong和Chwipungyeong,根据诗歌《龙飞上天之歌》而创作。作为韩国音乐记谱法系统Jeongganbo的最早例子之一,剩余的版本只包括一个基本的旋律。我们的研究团队受韩国传统音乐中心的委托,旨在将这首古老的旋律转化为一个六声部合奏的可表演的安排。使用通过定制的光学音乐识别获得的Jeongganbo数据,我们训练了一个类似BERT的掩蔽语言模型和一个编码器-解码器Transformer模型。我们还提出了一个编码方案,严格遵循Jeongganbo的结构,并表示注意持续时间的位置。由此产生的机器转换版本的芝瓦平和Chwipungyeong由专家评估,并由国立Gugak中心的宫廷音乐管弦乐队演奏。我们的工作表明,如果结合精心的设计,生成模型可以成功地应用于训练数据有限的传统音乐。摘要:We introduce a project that revives a piece of 15th-century Korean court music, Chihwapyeong and Chwipunghyeong, composed upon the poem Songs of the Dragon Flying to Heaven. One of the earliest examples of Jeongganbo, a Korean musical notation system, the remaining version only consists of a rudimentary melody. Our research team, commissioned by the National Gugak (Korean Traditional Music) Center, aimed to transform this old melody into a performable arrangement for a six-part ensemble. Using Jeongganbo data acquired through bespoke optical music recognition, we trained a BERT-like masked language model and an encoder-decoder transformer model. We also propose an encoding scheme that strictly follows the structure of Jeongganbo and denotes note durations as positions. The resulting machine-transformed version of Chihwapyeong and Chwipunghyeong were evaluated by experts and performed by the Court Music Orchestra of National Gugak Center. Our work demonstrates that generative models can successfully be applied to traditional music with limited training data if combined with careful design.

【5】 Expressive MIDI-format Piano Performance Generation
标题: 表现力丰富的MIDI格式钢琴演奏生成
作者:Jingwei Liu
备注:4 pages, 2 figures
链接:点击下载PDF文件
摘要:这项工作提出了一个生成神经网络,能够生成具有表现力的钢琴演奏。音乐的表现力体现在生动的微计时,丰富的复调纹理,多样的动态,和持续踏板效果。该模型从数据处理到神经网络设计等多个方面都具有创新性。我们声称,这种象征性的音乐生成模型克服了象征性音乐的共同批评,并能够生成表现力的音乐流一样好,如果不是更好的几代人与原始音频。一个缺点是,由于提交的时间有限,模型没有经过微调和充分训练,因此在某些点上,生成可能听起来不连贯和随机。尽管如此,该模型显示出其强大的生成能力,生成富有表现力的钢琴作品。摘要:This work presents a generative neural network that's able to generate expressive piano performance in MIDI format. The musical expressivity is reflected by vivid micro-timing, rich polyphonic texture, varied dynamics, and the sustain pedal effects. This model is innovative from many aspects of data processing to neural network design. We claim that this symbolic music generation model overcame the common critics of symbolic music and is able to generate expressive music flows as good as, if not better than generations with raw audio. One drawback is that, due to the limited time for submission, the model is not fine-tuned and sufficiently trained, thus the generation may sound incoherent and random at certain points. Despite that, this model shows its powerful generative ability to generate expressive piano pieces.

【6】 Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
标题: 通过多阶段训练改进用于声音事件检测的音频频谱图变换器
作者:Florian Schmid,Paul Primus,Tobias Morocutti,Jonathan Greif,Gerhard Widmer
备注:Technical Report describing our system for DCASE2024 Challenge Task 4: this https URL Code: this https URL arXiv admin note: text overlap with arXiv:2407.12997
链接:点击下载PDF文件
摘要:本技术报告描述了CP-JKU团队提交的任务4声音事件检测与异构训练数据集和DCASE 24挑战的潜在缺失标签。我们在两阶段训练过程中对联合DESED和MAESTRO数据集上的三个大型音频频谱图Transformers,PaSST,BEAT和ATST进行微调。第一阶段紧密匹配基线系统设置并训练CRNN模型,同时保持大型预训练的Transformer模型冻结。在第二阶段,CRNN和Transformer都使用加权自监督损耗进行微调。在第二阶段之后,我们使用所有三个微调Transformers的集合来计算训练集中所有音频片段的强伪标签。然后,在第二次迭代中,我们重复两阶段的训练过程,并包含基于伪标签的蒸馏损失,从而大幅提高单模型性能。此外,我们在带有强时间标签的AudioSet子集上预训练PaSST和ATST,然后在Task 4数据集上对其进行微调。摘要:This technical report describes the CP-JKU team's submission for Task 4 Sound Event Detection with Heterogeneous Training Datasets and Potentially Missing Labels of the DCASE 24 Challenge. We fine-tune three large Audio Spectrogram Transformers, PaSST, BEATs, and ATST, on the joint DESED and MAESTRO datasets in a two-stage training procedure. The first stage closely matches the baseline system setup and trains a CRNN model while keeping the large pre-trained transformer model frozen. In the second stage, both CRNN and transformer are fine-tuned using heavily weighted self-supervised losses. After the second stage, we compute strong pseudo-labels for all audio clips in the training set using an ensemble of all three fine-tuned transformers. Then, in a second iteration, we repeat the two-stage training process and include a distillation loss based on the pseudo-labels, boosting single-model performance substantially. Additionally, we pre-train PaSST and ATST on the subset of AudioSet that comes with strong temporal labels, before fine-tuning them on the Task 4 datasets.


eess.AS音频处理
【1】 Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
标题: 通过多阶段训练改进用于声音事件检测的音频频谱图变换器
作者:Florian Schmid,Paul Primus,Tobias Morocutti,Jonathan Greif,Gerhard Widmer
备注:Technical Report describing our system for DCASE2024 Challenge Task 4: this https URL Code: this https URL arXiv admin note: text overlap with arXiv:2407.12997
链接:点击下载PDF文件
摘要:本技术报告描述了CP-JKU团队提交的任务4声音事件检测与异构训练数据集和DCASE 24挑战的潜在缺失标签。我们在两阶段训练过程中对联合DESED和MAESTRO数据集上的三个大型音频频谱图Transformers,PaSST,BEAT和ATST进行微调。第一阶段紧密匹配基线系统设置并训练CRNN模型,同时保持大型预训练的Transformer模型冻结。在第二阶段,CRNN和Transformer都使用加权自监督损耗进行微调。在第二阶段之后,我们使用所有三个微调Transformers的集合来计算训练集中所有音频片段的强伪标签。然后,在第二次迭代中,我们重复两阶段的训练过程,并包含基于伪标签的蒸馏损失,从而大幅提高单模型性能。此外,我们在带有强时间标签的AudioSet子集上预训练PaSST和ATST,然后在Task 4数据集上对其进行微调。摘要:This technical report describes the CP-JKU team's submission for Task 4 Sound Event Detection with Heterogeneous Training Datasets and Potentially Missing Labels of the DCASE 24 Challenge. We fine-tune three large Audio Spectrogram Transformers, PaSST, BEATs, and ATST, on the joint DESED and MAESTRO datasets in a two-stage training procedure. The first stage closely matches the baseline system setup and trains a CRNN model while keeping the large pre-trained transformer model frozen. In the second stage, both CRNN and transformer are fine-tuned using heavily weighted self-supervised losses. After the second stage, we compute strong pseudo-labels for all audio clips in the training set using an ensemble of all three fine-tuned transformers. Then, in a second iteration, we repeat the two-stage training process and include a distillation loss based on the pseudo-labels, boosting single-model performance substantially. Additionally, we pre-train PaSST and ATST on the subset of AudioSet that comes with strong temporal labels, before fine-tuning them on the Task 4 datasets.

【2】 MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
标题: MuChoMusic:在多模式音频语言模型中评估音乐理解
作者:Benno Weck,Ilaria Manco,Emmanouil Benetos,Elio Quinton,George Fazekas,Dmitry Bogdanov
备注:Accepted at ISMIR 2024. Data: this https URL Code: this https URL Supplementary material: this https URL
链接:点击下载PDF文件
摘要:联合处理音频和语言的多模态模型在音频理解中具有很大的前景,并且越来越多地被应用于音乐领域。通过允许用户通过文本进行查询并获得有关给定音频输入的信息,这些模型有可能通过基于语言的界面实现各种音乐理解任务。然而,他们的评估带来了相当大的挑战,目前还不清楚如何有效地评估他们的能力,正确地解释与音乐相关的输入与当前的方法。出于这一动机,我们介绍了MuChoMusic,这是一个用于评估多模态语言模型中音乐理解的基准,专注于音频。MuChoMusic包含1,187个多项选择题,全部由人工注释员验证,来自两个公开的音乐数据集的644首音乐曲目,涵盖各种类型。基准中的问题是精心设计的,以评估知识和推理能力,涵盖基本的音乐概念及其与文化和功能背景的关系。通过基准提供的整体分析,我们评估了五个开源模型,并确定了几个陷阱,包括过度依赖语言模态,指出需要更好的多模态集成。数据和代码都是开源的。摘要:Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about a given audio input, these models have the potential to enable a variety of music understanding tasks via language-based interfaces. However, their evaluation poses considerable challenges, and it remains unclear how to effectively assess their ability to correctly interpret music-related inputs with current methods. Motivated by this, we introduce MuChoMusic, a benchmark for evaluating music understanding in multimodal language models focused on audio. MuChoMusic comprises 1,187 multiple-choice questions, all validated by human annotators, on 644 music tracks sourced from two publicly available music datasets, and covering a wide variety of genres. Questions in the benchmark are crafted to assess knowledge and reasoning abilities across several dimensions that cover fundamental musical concepts and their relation to cultural and functional contexts. Through the holistic analysis afforded by the benchmark, we evaluate five open-source models and identify several pitfalls, including an over-reliance on the language modality, pointing to a need for better multimodal integration. Data and code are open-sourced.

【3】 Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
标题: 视听广义Zero-Shot学习的分布外检测:一个通用框架
作者:Liuyuan Wen
链接:点击下载PDF文件
摘要:广义零次学习(Generalized Zero-Shot Learning,简称GSHL)是一项具有挑战性的任务,需要对可见和不可见的类进行准确分类。在这一领域,视听语言作为一个非常令人兴奋但困难的任务出现,考虑到包括视觉和声学特征作为多模态输入。在这一领域的现有努力主要利用嵌入为基础的或生成为基础的方法。然而,生成式训练困难且不稳定,而基于嵌入的方法经常遇到域偏移问题。因此,我们发现将这两种方法集成到一个统一的框架中,以利用它们的优点,同时减轻各自的缺点是有希望的。我们的研究介绍了一个一般框架,采用了分布(OOD)检测,旨在利用这两种方法的优势。我们首先使用生成对抗网络来合成看不见的特征,从而能够训练OOD检测器以及可见和不可见类的分类器。该检测器确定测试特征是否属于可见或不可见的类,然后对每个特征类型使用单独的分类器进行分类。我们在三个流行的视听数据集上测试了我们的框架,并观察到与现有的最先进的作品相比有显着的改进。代码可以在https: github.com liuyuan-wen AV-OOD-GZSL中找到。摘要:Generalized Zero-Shot Learning (GZSL) is a challenging task requiring accurate classification of both seen and unseen classes. Within this domain, Audio-visual GZSL emerges as an extremely exciting yet difficult task, given the inclusion of both visual and acoustic features as multi-modal inputs. Existing efforts in this field mostly utilize either embedding-based or generative-based methods. However, generative training is difficult and unstable, while embedding-based methods often encounter domain shift problem. Thus, we find it promising to integrate both methods into a unified framework to leverage their advantages while mitigating their respective disadvantages. Our study introduces a general framework employing out-of-distribution (OOD) detection, aiming to harness the strengths of both approaches. We first employ generative adversarial networks to synthesize unseen features, enabling the training of an OOD detector alongside classifiers for seen and unseen classes. This detector determines whether a test feature belongs to seen or unseen classes, followed by classification utilizing separate classifiers for each feature type. We test our framework on three popular audio-visual datasets and observe a significant improvement comparing to existing state-of-the-art works. Codes can be found in https: github.com liuyuan-wen AV-OOD-GZSL.

【4】 Nested Music Transformer: Sequentially Decoding Compound Tokens in Symbolic Music and Audio Generation
标题: 嵌套音乐Transformer:在符号音乐和音频生成中顺序解码复合代币
作者:Jiwoo Ryu,Hao-Wen Dong,Jongmin Jung,Dasaem Jeong
备注:Accepted at 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:用复合记号表示符号音乐,其中每个记号由表示不同音乐特征或属性的若干不同子记号组成,提供了减少序列长度的优点。虽然之前的研究已经验证了复合标记在音乐序列建模中的有效性,但同时预测所有子标记可能会导致次优结果,因为它可能无法完全捕获它们之间的相互依赖性。我们介绍嵌套的音乐Transformer(NMT),为解码复合令牌自回归,类似于处理扁平令牌,但具有低内存使用量的架构。NMT由两个Transformers组成:对复合令牌序列进行建模的主解码器和对每个复合令牌的子令牌进行建模的子解码器。实验结果表明,NMT应用于复合令牌可以提高性能,在处理各种符号音乐数据集和离散音频令牌从MAESTRO数据集更好的困惑。摘要:Representing symbolic music with compound tokens, where each token consists of several different sub-tokens representing a distinct musical feature or attribute, offers the advantage of reducing sequence length. While previous research has validated the efficacy of compound tokens in music sequence modeling, predicting all sub-tokens simultaneously can lead to suboptimal results as it may not fully capture the interdependencies between them. We introduce the Nested Music Transformer (NMT), an architecture tailored for decoding compound tokens autoregressively, similar to processing flattened tokens, but with low memory usage. The NMT consists of two transformers: the main decoder that models a sequence of compound tokens and the sub-decoder for modeling sub-tokens of each compound token. The experiment results showed that applying the NMT to compound tokens can enhance the performance in terms of better perplexity in processing various symbolic music datasets and discrete audio tokens from the MAESTRO dataset.

【5】 Six Dragons Fly Again: Reviving 15th-Century Korean Court Music with Transformers and Novel Encoding
标题: 六龙再次飞翔:用Transformer和新颖编码复兴15世纪韩国宫廷音乐
作者:Danbinaerin Han,Mark Gotham,Dongmin Kim,Hannah Park,Sihun Lee,Dasaem Jeong
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:我们介绍一个项目,复兴了15世纪韩国宫廷音乐,Chihwapyeong和Chwipungyeong,根据诗歌《龙飞上天之歌》而创作。作为韩国音乐记谱法系统Jeongganbo的最早例子之一,剩余的版本只包括一个基本的旋律。我们的研究团队受韩国传统音乐中心的委托,旨在将这首古老的旋律转化为一个六声部合奏的可表演的安排。我们使用通过定制光学音乐识别获取的Jeongganbo数据,训练了类似BERT的掩蔽语言模型和编码器-解码器Transformer模型。我们还提出了一个编码方案,严格遵循Jeongganbo的结构,并表示注意持续时间的位置。由此产生的机器转换版本的芝瓦平和Chwipungyeong由专家评估,并由国立Gugak中心的宫廷音乐管弦乐队演奏。我们的工作表明,如果结合精心的设计,生成模型可以成功地应用于训练数据有限的传统音乐。摘要:We introduce a project that revives a piece of 15th-century Korean court music, Chihwapyeong and Chwipunghyeong, composed upon the poem Songs of the Dragon Flying to Heaven. One of the earliest examples of Jeongganbo, a Korean musical notation system, the remaining version only consists of a rudimentary melody. Our research team, commissioned by the National Gugak (Korean Traditional Music) Center, aimed to transform this old melody into a performable arrangement for a six-part ensemble. Using Jeongganbo data acquired through bespoke optical music recognition, we trained a BERT-like masked language model and an encoder-decoder transformer model. We also propose an encoding scheme that strictly follows the structure of Jeongganbo and denotes note durations as positions. The resulting machine-transformed version of Chihwapyeong and Chwipunghyeong were evaluated by experts and performed by the Court Music Orchestra of National Gugak Center. Our work demonstrates that generative models can successfully be applied to traditional music with limited training data if combined with careful design.

【6】 Expressive MIDI-format Piano Performance Generation
标题: 表现力丰富的MIDI格式钢琴演奏生成
作者:Jingwei Liu
备注:4 pages, 2 figures
链接:点击下载PDF文件
摘要:这项工作提出了一个生成神经网络,能够生成具有表现力的钢琴演奏。音乐的表现力体现在生动的微计时,丰富的复调纹理,多样的动态,和持续踏板效果。该模型从数据处理到神经网络设计等多个方面都具有创新性。我们声称,这种象征性的音乐生成模型克服了象征性音乐的共同批评,并能够生成表现力的音乐流一样好,如果不是更好的几代人与原始音频。一个缺点是,由于提交的时间有限,模型没有经过微调和充分训练,因此在某些点上,生成可能听起来不连贯和随机。尽管如此,该模型显示出其强大的生成能力,生成富有表现力的钢琴作品。摘要:This work presents a generative neural network that's able to generate expressive piano performance in MIDI format. The musical expressivity is reflected by vivid micro-timing, rich polyphonic texture, varied dynamics, and the sustain pedal effects. This model is innovative from many aspects of data processing to neural network design. We claim that this symbolic music generation model overcame the common critics of symbolic music and is able to generate expressive music flows as good as, if not better than generations with raw audio. One drawback is that, due to the limited time for submission, the model is not fine-tuned and sufficiently trained, thus the generation may sound incoherent and random at certain points. Despite that, this model shows its powerful generative ability to generate expressive piano pieces.


机器翻译,仅供参考