今天跟大家分享一篇语音相关的论文合集:cs.SD语音8篇,eess.AS音频处理9篇。本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
【1】 Sketching the Expression: Flexible Rendering of Expressive Piano Performance with Self-Supervised Learning
标题:表达的素描:通过自我监督学习灵活呈现富有表现力的钢琴演奏
链接:https://arxiv.org/abs/2208.14867
作者:Seungyeon Rhyu,Sarah Kim,Kyogu Lee机构:Music and Audio Research Group (MARG), Seoul National University, South Korea, Krust Universe, South Korea, Graduate School of AI, AI Institute, Seoul National University, South Korea备注:8 pages, 4 figures, the 23rd International Society for Music Information Retrieval Conference, Bengaluru, India, 2022摘要:我们提出了一种用于渲染具有灵活的音乐表达的象征性钢琴演奏的系统。为了创造出传达各种情感或细微差别的新的音乐表演,有必要主动地控制音乐表达。然而,先前的方法仅限于遵循作曲家的音乐表达指导方针或仅处理音乐属性的一部分。我们的目标是使用一个条件VAE框架来解开钢琴演奏的整体音乐表现和结构属性。它从潜在表征和给定的音符结构中随机产生表达参数。此外,我们还采用了自监督的方法,强制潜变量代表目标属性。最后,我们利用两步编码器和解码器来学习层次依赖性,以增强输出的自然度。实验结果表明,该系统能够稳定地生成与给定乐谱相关的演奏参数,学习解纠缠的表示,并独立地控制音乐属性.摘要:We propose a system for rendering a symbolic piano performance with flexible musical expression. It is necessary to actively control musical expression for creating a new music performance that conveys various emotions or nuances. However, previous approaches were limited to following the composer's guidelines of musical expression or dealing with only a part of the musical attributes. We aim to disentangle the entire musical expression and structural attribute of piano performance using a conditional VAE framework. It stochastically generates expressive parameters from latent representations and given note structures. In addition, we employ self-supervised approaches that force the latent variables to represent target attributes. Finally, we leverage a two-step encoder and decoder that learn hierarchical dependency to enhance the naturalness of the output. Experimental results show that our system can stably generate performance parameters relevant to the given musical scores, learn disentangled representations, and control musical attributes independently of each other.
【2】 Cadence Detection in Symbolic Classical Music using Graph Neural Networks
标题:基于图神经网络的符号化古典音乐韵律检测
链接:https://arxiv.org/abs/2208.14819
作者:Emmanouil Karystinaios,Gerhard Widmer机构:Institute of Computational Perception, Johannes Kepler University Linz, Austria, LIT AI Lab, Linz Institute of Technology, Austria备注:In proceedings of the International Society for Music Information Retrieval Conference 2022 (ISMIR)摘要:从对位复调开始到今天,抑扬顿挫是一种复杂的结构,它一直在推动着音乐的发展。检测这样的结构对于诸如音乐分析、基调检测或音乐分段的许多MIR任务是至关重要的。然而,自动节奏检测仍然具有挑战性,主要是因为它涉及到高级音乐元素的组合,如和声、声部引导和节奏。在这项工作中,我们提出了一个符号分数的图形表示作为解决节奏检测任务的中间手段。我们使用图卷积网络将节奏检测作为不平衡节点分类问题来处理。我们得到的结果与现有技术大致相当,我们提出了一个能够在多个粒度级别(从单个音符到节拍)进行预测的模型,这要归功于细粒度、逐个音符的表示。此外,我们的实验表明,图形卷积可以学习非局部特征,帮助节奏检测,使我们不必设计专门的特征来编码非局部上下文。我们认为,这种建模乐谱和分类任务的通用方法具有许多潜在的优势,超出了这里介绍的特定识别任务。摘要:Cadences are complex structures that have been driving music from the beginning of contrapuntal polyphony until today. Detecting such structures is vital for numerous MIR tasks such as musicological analysis, key detection, or music segmentation. However, automatic cadence detection remains challenging mainly because it involves a combination of high-level musical elements like harmony, voice leading, and rhythm. In this work, we present a graph representation of symbolic scores as an intermediate means to solve the cadence detection task. We approach cadence detection as an imbalanced node classification problem using a Graph Convolutional Network. We obtain results that are roughly on par with the state of the art, and we present a model capable of making predictions at multiple levels of granularity, from individual notes to beats, thanks to the fine-grained, note-by-note representation. Moreover, our experiments suggest that graph convolution can learn non-local features that assist in cadence detection, freeing us from the need of having to devise specialized features that encode non-local context. We argue that this general approach to modeling musical scores and classification tasks has a number of potential advantages, beyond the specific recognition task presented here.
【3】 Domain Shift-oriented Machine Anomalous Sound Detection Model Based on Self-Supervised Learning
标题:基于自监督学习的面向域移位的机器异声检测模型
链接:https://arxiv.org/abs/2208.14812
作者:Jing-ke Yan,Xin Wang,Qin Wang,Qin Qin,Huang-he Li,Peng-fei Ye,Yue-ping He,Jing Zeng摘要:得益于深度学习的发展,基于自监督学习的机器异常声音检测研究取得了显著成果。然而,在同一机器的不同操作条件下,测试集和训练集的声学特性存在差异(域移位)。现有的域偏移检测方法难以在低计算开销的情况下稳定地学习域偏移特征。针对这些问题,提出了一种基于自监督学习的面向域转移的机器异常声音检测模型(TranSelf-DyGCN)。首先,设计了时频域特征建模网络,捕获全局和局部的空间和时域特征,提高了机器异常声音检测在域偏移下的稳定性。然后,采用动态图卷积网络(Dynamic Graph Convolutional Network,DyGCN)对领域转移特征之间的相互依赖关系进行建模,使模型能够有效地感知领域转移特征.最后,利用域自适应网络(DAN)补偿域偏移引起的性能下降,使模型更好地适应自监督环境下的异常声音。在DCASE 2020任务2和DCASE 2022任务2上验证了所建议模型的性能。摘要:Thanks to the development of deep learning, research on machine anomalous sound detection based on self-supervised learning has made remarkable achievements. However, there are differences in the acoustic characteristics of the test set and the training set under different operating conditions of the same machine (domain shifts). It is challenging for the existing detection methods to learn the domain shifts features stably with low computation overhead. To address these problems, we propose a domain shift-oriented machine anomalous sound detection model based on self-supervised learning (TranSelf-DyGCN) in this paper. Firstly, we design a time-frequency domain feature modeling network to capture global and local spatial and time-domain features, thus improving the stability of machine anomalous sound detection stability under domain shifts. Then, we adopt a Dynamic Graph Convolutional Network (DyGCN) to model the inter-dependence relationship between domain shifts features, enabling the model to perceive domain shifts features efficiently. Finally, we use a Domain Adaptive Network (DAN) to compensate for the performance decrease caused by domain shifts, making the model adapt to anomalous sound better in the self-supervised environment. The performance of the suggested model is validated on DCASE 2020 task 2 and DCASE 2022 task 2.
【4】 Harmonization and Evaluation; Tweaking the Parameters on Human Listeners
标题:协调与评估;调整人类听者的参数
链接:https://arxiv.org/abs/2208.14750
作者:Filippo Carnovalini,Alessandro Pelizzo,Antonio Rodà,Sergio Canazza机构:a Centro di Sonologia Computazionale (CSC), Dept. of Information Engineering, University of备注:Accepted for publication in 9th International Conference on Kansei Engineering and Emotion Research 2022摘要:感性模型被用来研究音乐的内涵意义。在多媒体和混合现实中,越来越多地使用自动生成的旋律。重要的是要考虑这种音乐是否传达了什么情感以及传达了什么情感。对计算机生成的旋律进行评估并不是一项微不足道的任务。考虑到难以定义生成的音乐作品的质量的有用的量化度量,研究人员经常求助于人的评估。在这些评估中,通常要求评委评估一组生成的作品以及一些基准作品。后者通常由人组成。虽然这种评估相对常见,但众所周知,在设计实验时应谨慎,因为人类可能受到各种因素的影响。在本文中,我们研究了评委必须评估的音频文件中和声的影响,以了解伴奏是否会改变对生成旋律的评估。为了做到这一点,我们用两种不同的算法生成旋律,并用我们为这个实验设计的自动工具协调它们,然后让60多名参与者对旋律进行评估。通过使用统计分析,我们表明,通过强调判断之间的差异,协调确实影响了评估过程。摘要:Kansei models were used to study the connotative meaning of music. In multimedia and mixed reality, automatically generated melodies are increasingly being used. It is important to consider whether and what feelings are communicated by this music. Evaluation of computer-generated melodies is not a trivial task. Considered the difficulty of defining useful quantitative metrics of the quality of a generated musical piece, researchers often resort to human evaluation. In these evaluations, often the judges are required to evaluate a set of generated pieces along with some benchmark pieces. The latter are often composed by humans. While this kind of evaluation is relatively common, it is known that care should be taken when designing the experiment, as humans can be influenced by a variety of factors. In this paper, we examine the impact of the presence of harmony in audio files that judges must evaluate, to see whether having an accompaniment can change the evaluation of generated melodies. To do so, we generate melodies with two different algorithms and harmonize them with an automatic tool that we designed for this experiment, and ask more than sixty participants to evaluate the melodies. By using statistical analyses, we show harmonization does impact the evaluation process, by emphasizing the differences among judgements.
【5】 A New Corpus for Computational Music Research and A Novel Method for Musical Structure Analysis
标题:计算音乐研究的新语料库和音乐结构分析的新方法
链接:https://arxiv.org/abs/2208.14747
作者:Filippo Carnovalini,Antonio Rodà,Nicholas Harley,Steven T. Homer,Geraint A. Wiggins机构:Università degli Studi di Padova, Padova, Italy, Vrije Universiteit Brussel, &, Queen Mary University of London, Pleinlaan , Brussel, Belgium, Mile End Road, London E,NS, UK摘要:音乐的计算模型,虽然提供了旋律发展的良好描述,但仍然不能完全掌握由旋律材料的重复、移调和重用组成的一般结构。我们提出了一个强结构化的巴洛克风格的Allemandes语料库,并描述了一种自顶向下的方法来抽象他们的音乐内容的共享结构,使用树表示法从Schenkerian启发的每一个作品的分析之间的成对差异产生,从而提供了一个丰富的语料库的层次描述。摘要:Computational models of music, while providing good descriptions of melodic development, still cannot fully grasp the general structure comprised of repetitions, transpositions, and reuse of melodic material. We present a corpus of strongly structured baroque allemandes, and describe a top-down approach to abstract the shared structure of their musical content using tree representations produced from pairwise differences between the Schenkerian-inspired analyses of each piece, thereby providing a rich hierarchical description of the corpus.
【6】 Open Challenges in Musical Metacreation
标题:音乐元创作中的开放挑战
链接:https://arxiv.org/abs/2208.14734
机构:University of Padua, Department of Information Engineering, Padua, Italy摘要:音乐元创造试图从计算机算法中获得创作行为。在本文中,我简要分析了该领域是如何从算法合成发展到专注于寻找创造力的,并指出了在追求这一目标过程中的一些问题。最后,我认为算法的混合可以是一个有用的研究方向。摘要:Musical Metacreation tries to obtain creative behaviors from computers algorithms composing music. In this paper I briefly analyze how this field evolved from algorithmic composition to be focused on the search for creativity, and I point out some issues in pursuing this goal. Finally, I argue that hybridization of algorithms can be a useful direction for research.
【7】 A Real-Time Tempo and Meter Tracking System for Rhythmic Improvis
标题:一种节奏性即兴演奏的实时节奏和节拍跟踪系统
链接:https://arxiv.org/abs/2208.14717
作者:Filippo Carnovalini,Antonio Rodà机构:University of Padova (IT)摘要:音乐是一种表达形式,通常需要玩家之间的互动。如果一个人希望以这样一种音乐方式与计算机进行交互,那么机器就必须能够解释人类给出的输入,以发现其音乐意义。在这项工作中,我们提出了一个能够检测基本节奏特征的系统,该系统可以允许应用程序将其输出与用户给出的节奏同步,而无需对可能的输入有任何预先协议或要求。详细描述了该系统,并通过仿真使用定量指标给出了评估。评估表明,该系统在一定的设置下可以一致地检测速度和节拍,并且可以为进一步开发导致系统对节奏变化的输入鲁棒的系统奠定坚实的基础。摘要:Music is a form of expression that often requires interaction between players. If one wishes to interact in such a musical way with a computer, it is necessary for the machine to be able to interpret the input given by the human to find its musical meaning. In this work, we propose a system capable of detecting basic rhythmic features that can allow an application to synchronize its output with the rhythm given by the user, without having any prior agreement or requirement on the possible input. The system is described in detail and an evaluation is given through simulation using quantitative metrics. The evaluation shows that the system can detect tempo and meter consistently under certain settings, and could be a solid base for further developments leading to a system robust to rhythmically changing inputs.
【8】 Audiogram Digitization Tool for Audiological Reports
标题:用于听力学报告的听力图数字化工具
链接:https://arxiv.org/abs/2208.14621
作者:François Charih,James R. Green机构:Systems and Computer Engineering, Carleton University, Ottawa, ON, Canada摘要:许多私营和公共保险公司对那些听力损失直接归因于工作场所过度暴露于噪音的工人进行赔偿。索赔评估过程通常是漫长的,并且需要人类裁决者的大量努力,他们必须解释通常通过传真或等同物发送的手工记录的听力图。在这项工作中,我们提出了一个与安大略省工作场所安全保险委员会合作开发的解决方案,以简化裁决过程。特别是,我们提出了第一个听力图数字化算法,能够从扫描或传真的听力学报告中自动提取听阈,作为概念验证。该算法提取的大多数阈值在5dB精度内,允许以半监督方式大大减少将听力图转换为数字格式所需的时间,并且是朝向自动化判定过程的第一步。数字化算法的源代码和我们NIHL注释门户的基于桌面的实现可在GitHub上公开获得(https://github.com/GreenCUBIC/AudiogramDigitization)。摘要:A number of private and public insurers compensate workers whose hearing loss can be directly attributed to excessive exposure to noise in the workplace. The claim assessment process is typically lengthy and requires significant effort from human adjudicators who must interpret hand-recorded audiograms, often sent via fax or equivalent. In this work, we present a solution developed in partnership with the Workplace Safety Insurance Board of Ontario to streamline the adjudication process. In particular, we present the first audiogram digitization algorithm capable of automatically extracting the hearing thresholds from a scanned or faxed audiology report as a proof-of-concept. The algorithm extracts most thresholds within 5 dB accuracy, allowing to substantially lessen the time required to convert an audiogram into digital format in a semi-supervised fashion, and is a first step towards the automation of the adjudication process. The source code for the digitization algorithm and a desktop-based implementation of our NIHL annotation portal is publicly available on GitHub (https://github.com/GreenCUBIC/AudiogramDigitization).
【1】 Singing Beat Tracking With Self-supervised Front-end and Linear Transformers
标题:基于自监控前端和线性Transformer的演唱节拍跟踪
链接:https://arxiv.org/abs/2208.14578
作者:Mojtaba Heydari,Zhiyao Duan机构:Department of Electrical and Computer Engineering, University of Rochester, Wilson Blvd, Rochester, NY , USA备注:23rd International Society for Music Information Retrieval Conference (ISMIR 2022)摘要:在没有音乐伴奏的情况下跟踪歌声的节拍可以在音乐制作、自动歌曲编排和社交媒体交互中找到许多应用。它的主要挑战是缺乏强节奏和和声模式,这些模式对音乐节奏分析通常很重要。即使对于人类听众来说,这也是一项具有挑战性的任务。结果,现有的音乐节拍跟踪系统不能对歌唱声音提供令人满意的性能。本文提出了一种新的歌唱节拍跟踪算法,并提出了解决该算法的第一种方法。该方法通过使用预先训练的自监督WavLM和DistilHuBERT语音表示作为前端来利用歌唱声音的语义信息,并使用自注意编码器层来预测节拍。为了训练和测试该系统,我们对完整的歌曲使用源分离和节拍跟踪来获得分离的演唱声音及其节拍注释,然后进行手动校正。在GTZAN数据集上的741个分离音轨上的实验表明,该系统在节拍跟踪精度方面远远优于几种最先进的音乐节拍跟踪方法.消融研究还证实了预训练的自监督语音表示相对于通用频谱特征的优点。摘要:Tracking beats of singing voices without the presence of musical accompaniment can find many applications in music production, automatic song arrangement, and social media interaction. Its main challenge is the lack of strong rhythmic and harmonic patterns that are important for music rhythmic analysis in general. Even for human listeners, this can be a challenging task. As a result, existing music beat tracking systems fail to deliver satisfactory performance on singing voices. In this paper, we propose singing beat tracking as a novel task, and propose the first approach to solving this task. Our approach leverages semantic information of singing voices by employing pre-trained self-supervised WavLM and DistilHuBERT speech representations as the front-end and uses a self-attention encoder layer to predict beats. To train and test the system, we obtain separated singing voices and their beat annotations using source separation and beat tracking on complete songs, followed by manual corrections. Experiments on the 741 separated vocal tracks of the GTZAN dataset show that the proposed system outperforms several state-of-the-art music beat tracking methods by a large margin in terms of beat tracking accuracy. Ablation studies also confirm the advantages of pre-trained self-supervised speech representations over generic spectral features.
【2】 Sketching the Expression: Flexible Rendering of Expressive Piano Performance with Self-Supervised Learning
标题:表达的素描:通过自我监督学习灵活呈现富有表现力的钢琴演奏
链接:https://arxiv.org/abs/2208.14867
* 与cs.SD语音【1】为同一篇
作者:Seungyeon Rhyu,Sarah Kim,Kyogu Lee机构:Music and Audio Research Group (MARG), Seoul National University, South Korea, Krust Universe, South Korea, Graduate School of AI, AI Institute, Seoul National University, South Korea备注:8 pages, 4 figures, the 23rd International Society for Music Information Retrieval Conference, Bengaluru, India, 2022摘要:我们提出了一种用于渲染具有灵活的音乐表达的象征性钢琴演奏的系统。为了创造出传达各种情感或细微差别的新的音乐表演,有必要主动地控制音乐表达。然而,先前的方法仅限于遵循作曲家的音乐表达指导方针或仅处理音乐属性的一部分。我们的目标是使用一个条件VAE框架来解开钢琴演奏的整体音乐表现和结构属性。它从潜在表征和给定的音符结构中随机产生表达参数。此外,我们还采用了自监督的方法,强制潜变量代表目标属性。最后,我们利用两步编码器和解码器来学习层次依赖性,以增强输出的自然度。实验结果表明,该系统能够稳定地生成与给定乐谱相关的演奏参数,学习解纠缠的表示,并独立地控制音乐属性.摘要:We propose a system for rendering a symbolic piano performance with flexible musical expression. It is necessary to actively control musical expression for creating a new music performance that conveys various emotions or nuances. However, previous approaches were limited to following the composer's guidelines of musical expression or dealing with only a part of the musical attributes. We aim to disentangle the entire musical expression and structural attribute of piano performance using a conditional VAE framework. It stochastically generates expressive parameters from latent representations and given note structures. In addition, we employ self-supervised approaches that force the latent variables to represent target attributes. Finally, we leverage a two-step encoder and decoder that learn hierarchical dependency to enhance the naturalness of the output. Experimental results show that our system can stably generate performance parameters relevant to the given musical scores, learn disentangled representations, and control musical attributes independently of each other.
【3】 Cadence Detection in Symbolic Classical Music using Graph Neural Networks
标题:基于图神经网络的符号化古典音乐韵律检测
链接:https://arxiv.org/abs/2208.14819
* 与cs.SD语音【2】为同一篇
作者:Emmanouil Karystinaios,Gerhard Widmer机构:Institute of Computational Perception, Johannes Kepler University Linz, Austria, LIT AI Lab, Linz Institute of Technology, Austria备注:In proceedings of the International Society for Music Information Retrieval Conference 2022 (ISMIR)摘要:从对位复调开始到今天,抑扬顿挫是一种复杂的结构,它一直在推动着音乐的发展。检测这样的结构对于诸如音乐分析、基调检测或音乐分段的许多MIR任务是至关重要的。然而,自动节奏检测仍然具有挑战性,主要是因为它涉及到高级音乐元素的组合,如和声、声部引导和节奏。在这项工作中,我们提出了一个符号分数的图形表示作为解决节奏检测任务的中间手段。我们使用图卷积网络将节奏检测作为不平衡节点分类问题来处理。我们得到的结果与现有技术大致相当,我们提出了一个能够在多个粒度级别(从单个音符到节拍)进行预测的模型,这要归功于细粒度、逐个音符的表示。此外,我们的实验表明,图形卷积可以学习非局部特征,帮助节奏检测,使我们不必设计专门的特征来编码非局部上下文。我们认为,这种建模乐谱和分类任务的通用方法具有许多潜在的优势,超出了这里介绍的特定识别任务。摘要:Cadences are complex structures that have been driving music from the beginning of contrapuntal polyphony until today. Detecting such structures is vital for numerous MIR tasks such as musicological analysis, key detection, or music segmentation. However, automatic cadence detection remains challenging mainly because it involves a combination of high-level musical elements like harmony, voice leading, and rhythm. In this work, we present a graph representation of symbolic scores as an intermediate means to solve the cadence detection task. We approach cadence detection as an imbalanced node classification problem using a Graph Convolutional Network. We obtain results that are roughly on par with the state of the art, and we present a model capable of making predictions at multiple levels of granularity, from individual notes to beats, thanks to the fine-grained, note-by-note representation. Moreover, our experiments suggest that graph convolution can learn non-local features that assist in cadence detection, freeing us from the need of having to devise specialized features that encode non-local context. We argue that this general approach to modeling musical scores and classification tasks has a number of potential advantages, beyond the specific recognition task presented here.
【4】 Domain Shift-oriented Machine Anomalous Sound Detection Model Based on Self-Supervised Learning
标题:基于自监督学习的面向域移位的机器异声检测模型
链接:https://arxiv.org/abs/2208.14812
* 与cs.SD语音【3】为同一篇
作者:Jing-ke Yan,Xin Wang,Qin Wang,Qin Qin,Huang-he Li,Peng-fei Ye,Yue-ping He,Jing Zeng摘要:得益于深度学习的发展,基于自监督学习的机器异常声音检测研究取得了显著成果。然而,在同一机器的不同操作条件下,测试集和训练集的声学特性存在差异(域移位)。现有的域偏移检测方法难以在低计算开销的情况下稳定地学习域偏移特征。针对这些问题,提出了一种基于自监督学习的面向域转移的机器异常声音检测模型(TranSelf-DyGCN)。首先,设计了时频域特征建模网络,捕获全局和局部的空间和时域特征,提高了机器异常声音检测在域偏移下的稳定性。然后,采用动态图卷积网络(Dynamic Graph Convolutional Network,DyGCN)对领域转移特征之间的相互依赖关系进行建模,使模型能够有效地感知领域转移特征.最后,利用域自适应网络(DAN)补偿域偏移引起的性能下降,使模型更好地适应自监督环境下的异常声音。在DCASE 2020任务2和DCASE 2022任务2上验证了所建议模型的性能。摘要:Thanks to the development of deep learning, research on machine anomalous sound detection based on self-supervised learning has made remarkable achievements. However, there are differences in the acoustic characteristics of the test set and the training set under different operating conditions of the same machine (domain shifts). It is challenging for the existing detection methods to learn the domain shifts features stably with low computation overhead. To address these problems, we propose a domain shift-oriented machine anomalous sound detection model based on self-supervised learning (TranSelf-DyGCN) in this paper. Firstly, we design a time-frequency domain feature modeling network to capture global and local spatial and time-domain features, thus improving the stability of machine anomalous sound detection stability under domain shifts. Then, we adopt a Dynamic Graph Convolutional Network (DyGCN) to model the inter-dependence relationship between domain shifts features, enabling the model to perceive domain shifts features efficiently. Finally, we use a Domain Adaptive Network (DAN) to compensate for the performance decrease caused by domain shifts, making the model adapt to anomalous sound better in the self-supervised environment. The performance of the suggested model is validated on DCASE 2020 task 2 and DCASE 2022 task 2.
【5】 Harmonization and Evaluation; Tweaking the Parameters on Human Listeners
标题:协调与评估;调整人类听者的参数
链接:https://arxiv.org/abs/2208.14750
* 与cs.SD语音【4】为同一篇
作者:Filippo Carnovalini,Alessandro Pelizzo,Antonio Rodà,Sergio Canazza机构:a Centro di Sonologia Computazionale (CSC), Dept. of Information Engineering, University of备注:Accepted for publication in 9th International Conference on Kansei Engineering and Emotion Research 2022摘要:感性模型被用来研究音乐的内涵意义。在多媒体和混合现实中,越来越多地使用自动生成的旋律。重要的是要考虑这种音乐是否传达了什么情感以及传达了什么情感。对计算机生成的旋律进行评估并不是一项微不足道的任务。考虑到难以定义生成的音乐作品的质量的有用的量化度量,研究人员经常求助于人的评估。在这些评估中,通常要求评委评估一组生成的作品以及一些基准作品。后者通常由人组成。虽然这种评估相对常见,但众所周知,在设计实验时应谨慎,因为人类可能受到各种因素的影响。在本文中,我们研究了评委必须评估的音频文件中和声的影响,以了解伴奏是否会改变对生成旋律的评估。为了做到这一点,我们用两种不同的算法生成旋律,并用我们为这个实验设计的自动工具协调它们,然后让60多名参与者对旋律进行评估。通过使用统计分析,我们表明,通过强调判断之间的差异,协调确实影响了评估过程。摘要:Kansei models were used to study the connotative meaning of music. In multimedia and mixed reality, automatically generated melodies are increasingly being used. It is important to consider whether and what feelings are communicated by this music. Evaluation of computer-generated melodies is not a trivial task. Considered the difficulty of defining useful quantitative metrics of the quality of a generated musical piece, researchers often resort to human evaluation. In these evaluations, often the judges are required to evaluate a set of generated pieces along with some benchmark pieces. The latter are often composed by humans. While this kind of evaluation is relatively common, it is known that care should be taken when designing the experiment, as humans can be influenced by a variety of factors. In this paper, we examine the impact of the presence of harmony in audio files that judges must evaluate, to see whether having an accompaniment can change the evaluation of generated melodies. To do so, we generate melodies with two different algorithms and harmonize them with an automatic tool that we designed for this experiment, and ask more than sixty participants to evaluate the melodies. By using statistical analyses, we show harmonization does impact the evaluation process, by emphasizing the differences among judgements.
【6】 A New Corpus for Computational Music Research and A Novel Method for Musical Structure Analysis
标题:计算音乐研究的新语料库和音乐结构分析的新方法
链接:https://arxiv.org/abs/2208.14747
* 与cs.SD语音【5】为同一篇
作者:Filippo Carnovalini,Antonio Rodà,Nicholas Harley,Steven T. Homer,Geraint A. Wiggins机构:Università degli Studi di Padova, Padova, Italy, Vrije Universiteit Brussel, &, Queen Mary University of London, Pleinlaan , Brussel, Belgium, Mile End Road, London E,NS, UK备注:None摘要:音乐的计算模型,虽然提供了旋律发展的良好描述,但仍然不能完全掌握由旋律材料的重复、移调和重用组成的一般结构。我们提出了一个强结构化的巴洛克风格的Allemandes语料库,并描述了一种自顶向下的方法来抽象他们的音乐内容的共享结构,使用树表示法从Schenkerian启发的每一个作品的分析之间的成对差异产生,从而提供了一个丰富的语料库的层次描述。摘要:Computational models of music, while providing good descriptions of melodic development, still cannot fully grasp the general structure comprised of repetitions, transpositions, and reuse of melodic material. We present a corpus of strongly structured baroque allemandes, and describe a top-down approach to abstract the shared structure of their musical content using tree representations produced from pairwise differences between the Schenkerian-inspired analyses of each piece, thereby providing a rich hierarchical description of the corpus.
【7】 Open Challenges in Musical Metacreation
标题:音乐元创作中的开放挑战
链接:https://arxiv.org/abs/2208.14734
* 与cs.SD语音【6】为同一篇
机构:University of Padua, Department of Information Engineering, Padua, Italy摘要:音乐元创造试图从计算机算法中获得创作行为。在本文中,我简要分析了该领域是如何从算法合成发展到专注于寻找创造力的,并指出了在追求这一目标过程中的一些问题。最后,我认为算法的混合可以是一个有用的研究方向。摘要:Musical Metacreation tries to obtain creative behaviors from computers algorithms composing music. In this paper I briefly analyze how this field evolved from algorithmic composition to be focused on the search for creativity, and I point out some issues in pursuing this goal. Finally, I argue that hybridization of algorithms can be a useful direction for research.
【8】 A Real-Time Tempo and Meter Tracking System for Rhythmic Improvis
标题:一种节奏性即兴演奏的实时节奏和节拍跟踪系统
链接:https://arxiv.org/abs/2208.14717
* 与cs.SD语音【7】为同一篇
作者:Filippo Carnovalini,Antonio Rodà机构:University of Padova (IT)摘要:音乐是一种表达形式,通常需要玩家之间的互动。如果一个人希望以这样一种音乐方式与计算机进行交互,那么机器就必须能够解释人类给出的输入,以发现其音乐意义。在这项工作中,我们提出了一个能够检测基本节奏特征的系统,该系统可以允许应用程序将其输出与用户给出的节奏同步,而无需对可能的输入有任何预先协议或要求。详细描述了该系统,并通过仿真使用定量指标给出了评估。评估表明,该系统在一定的设置下可以一致地检测速度和节拍,并且可以为进一步开发导致系统对节奏变化的输入鲁棒的系统奠定坚实的基础。摘要:Music is a form of expression that often requires interaction between players. If one wishes to interact in such a musical way with a computer, it is necessary for the machine to be able to interpret the input given by the human to find its musical meaning. In this work, we propose a system capable of detecting basic rhythmic features that can allow an application to synchronize its output with the rhythm given by the user, without having any prior agreement or requirement on the possible input. The system is described in detail and an evaluation is given through simulation using quantitative metrics. The evaluation shows that the system can detect tempo and meter consistently under certain settings, and could be a solid base for further developments leading to a system robust to rhythmically changing inputs.
【9】 Time-Frequency Localization Using Deep Convolutional Maxout Neural Network in Persian Speech Recognition
标题:深度卷积最大输出神经网络在波斯语语音识别中的时频定位
链接:https://arxiv.org/abs/2108.03818
作者:Arash Dehghani,Seyyed Ali Seyyedsalehi机构:Amirkabir University of Technology (Tehran Polytechnic), Hafez, Ave., Tehran, Iran摘要:针对波斯语语音识别问题,提出了一种基于神经网络的信息时频局部化结构。研究表明,哺乳动物初级听觉皮层和中脑某些神经元的感受野的时频可塑性使定位机制提高了识别能力。在过去的几年中,已经做了很多工作来定位ASR系统中的时间—频率信息,使用诸如HMM、TDNN、CNNs和LSTM-RNN的方法的空间或时间不变性特性。然而,这些模型中的大多数具有大的参数体积,并且训练起来具有挑战性。为此,提出了一种时频卷积最大输出神经网络(TFCMNN)结构,将时域和频域并行的一维卷积最大输出神经网络同时独立地应用于频谱图,然后将它们的输出级联并联合应用于全连接的最大输出网络进行分类。为了提高该结构的性能,我们使用了新开发的方法和模型,如Dropout、maxout和权重归一化。在FARSDAT数据集上设计并实现了两组实验,以评估该模型与传统1D-CMNN模型相比的性能。实验结果表明,TFCMNN模型的平均识别率比传统的一维CMNN模型的平均识别率高1.6%左右。此外,TFCMNN模型的平均训练时间比传统模型的平均训练时间少约17小时。因此,如在其他来源中证明的,ASR系统中的时间—频率定位增加了系统准确性并且加速了训练过程。摘要:In this paper, a CNN-based structure for the time-frequency localization of information is proposed for Persian speech recognition. Research has shown that the receptive fields' spectrotemporal plasticity of some neurons in mammals' primary auditory cortex and midbrain makes localization facilities improve recognition performance. Over the past few years, much work has been done to localize time-frequency information in ASR systems, using the spatial or temporal immutability properties of methods such as HMMs, TDNNs, CNNs, and LSTM-RNNs. However, most of these models have large parameter volumes and are challenging to train. For this purpose, we have presented a structure called Time-Frequency Convolutional Maxout Neural Network (TFCMNN) in which parallel time-domain and frequency-domain 1D-CMNNs are applied simultaneously and independently to the spectrogram, and then their outputs are concatenated and applied jointly to a fully connected Maxout network for classification. To improve the performance of this structure, we have used newly developed methods and models such as Dropout, maxout, and weight normalization. Two sets of experiments were designed and implemented on the FARSDAT dataset to evaluate the performance of this model compared to conventional 1D-CMNN models. According to the experimental results, the average recognition score of TFCMNN models is about 1.6% higher than the average of conventional 1D-CMNN models. In addition, the average training time of the TFCMNN models is about 17 hours lower than the average training time of traditional models. Therefore, as proven in other sources, time-frequency localization in ASR systems increases system accuracy and speeds up the training process.
机器翻译,仅供参考