今天跟大家分享一篇语音相关的论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily
cs.SD语音

【1】 The match file format: Encoding Alignments between Scores and  Performances

标题:匹配文件格式:对分数和表演之间的对齐进行编码

链接:https://arxiv.org/abs/2206.01104

作者:Francesco Foscarin,Emmanouil Karystinaios,Silvan David Peter,Carlos Cancino-Chacón,Maarten Grachten,Gerhard Widmer
机构:Johannes Kepler University, Carlos Cancino-Chac´on∗, carlos eduardo., Independent Researcher
摘要:本文介绍了match的规范:一种文件格式,它将MIDI人体表演与音符、节拍和重拍级别对齐扩展到相应的乐谱。这可以对与各种任务相关的性能进行高级分析,例如表现性性能建模、乐谱跟踪、音乐转录和演奏者分类。匹配文件包括一组分数相关描述符,使其也可用作基本分数表示。对于需要使用结构分数元素的应用程序(例如,语音、部件、梁、音调),可以轻松地将匹配文件与符号分数组合。为了支持我们工作的实际应用,我们发布了Vienna4x22分数和成绩数据集的修正和升级版本,该数据集与比赛文件保持一致。
摘要:This paper presents the specifications of match: a file format that extends a MIDI human performance with note-, beat-, and downbeat-level alignments to a corresponding musical score. This enables advanced analyses of the performance that are relevant for various tasks, such as expressive performance modeling, score following, music transcription, and performer classification. The match file includes a set of score-related descriptors that makes it usable also as a bare-bones score representation. For applications that require the use of structural score elements (e.g., voices, parts, beams, slurs), the match file can be easily combined with the symbolic score. To support the practical application of our work, we release a corrected and upgraded version of the Vienna4x22 dataset of scores and performances aligned with match files.


【2】 Partitura: A Python Package for Symbolic Music Processing

标题:Partitura:一种用于符号音乐处理的Python包

链接:https://arxiv.org/abs/2206.01071

作者:Carlos Cancino-Chacón,Silvan David Peter,Emmanouil Karystinaios,Francesco Foscarin,Maarten Grachten,Gerhard Widmer
机构:Johannes Kepler University, Independent Researcher
摘要:Partitura是一个轻量级Python包,用于处理符号音乐信息。它可以方便地访问音乐信息检索任务中常用的功能,如音符数组(定时音调事件列表)和2D钢琴滚动矩阵,以及其他乐谱元素,如时间和关键签名、演奏指令和重复结构。Partitura可以加载乐谱(MEI、MusicXML、Kern和MIDI格式)、MIDI表演以及乐谱与表演的对齐。该软件包包括一些音乐分析工具,如自动音高拼写、关键签名识别和语音分离。Partitura是一个开源项目,可在https://github.com/CPJKU/partitura/.
摘要:Partitura is a lightweight Python package for handling symbolic musical information. It provides easy access to features commonly used in music information retrieval tasks, like note arrays (lists of timed pitched events) and 2D piano roll matrices, as well as other score elements such as time and key signatures, performance directives, and repeat structures. Partitura can load musical scores (in MEI, MusicXML, Kern, and MIDI formats), MIDI performances, and score-to-performance alignments. The package includes some tools for music analysis, such as automatic pitch spelling, key signature identification, and voice separation. Partitura is an open-source project and is available at https://github.com/CPJKU/partitura/.


【3】 Musical Instrument Recognition by XGBoost Combining Feature Fusion

标题:基于XGBoost结合特征融合的乐器识别

链接:https://arxiv.org/abs/2206.00901

作者:Yijie Liu,Yanfang Yin,Qigang Zhu,Wenzhuo Cui
机构:Shandong Univ Sci & Technol, Dept Elect Engn & Informat Technol, Jinan , Peoples R China;,
摘要:乐器分类是音乐信息检索的研究热点之一。为了解决现有乐器分类模型性能较差的问题,提出了一种基于多通道特征融合和XGBoost的乐器分类算法。在数据集音频特征提取和融合的基础上,将特征输入到XGBoost模型中进行训练;其次,通过比较不同的特征组合和朴素贝叶斯等经典机器学习模型,验证了该算法在乐器分类任务中的优越性能。该算法在Medley solos DB数据集上的准确率达到97.65%,优于现有模型。实验结果为乐器分类特征工程中的特征选择提供了参考。
摘要:Musical instrument classification is one of the focuses of Music Information Retrieval (MIR). In order to solve the problem of poor performance of current musical instrument classification models, we propose a musical instrument classification algorithm based on multi-channel feature fusion and XGBoost. Based on audio feature extraction and fusion of the dataset, the features are input into the XGBoost model for training; secondly, we verified the superior performance of the algorithm in the musical instrument classification task by com-paring different feature combinations and several classical machine learning models such as Naive Bayes. The algorithm achieves an accuracy of 97.65% on the Medley-solos-DB dataset, outperforming existing models. The experiments provide a reference for feature selection in feature engineering for musical instrument classification.


【4】 Self-supervised Learning of Audio Representations from Audio-Visual Data  using Spatial Alignment

标题:基于空间对齐的视听数据音频表示的自监督学习

链接:https://arxiv.org/abs/2206.00970

作者:Shanshan Wang,Archontis Politis,Annamaria Mesaros,Tuomas Virtanen
机构:Tampere University
摘要:从视听数据中学习为表达视听内容之间的对应关系提供了许多可能性,类似于人类对听觉和视觉信息的感知。在这项工作中,我们提出了一种基于视听空间对齐(AVSA)的自监督表示学习方法,AVSA是一种比视听对应(AVC)更复杂的对齐任务。除了通信之外,AVSA还从声音和视觉内容的空间位置学习。基于360$^\text{o}$视频和环境音音频,我们提出了使用对象检测选择视觉对象,并将音频信号向检测到的对象进行波束形成,试图学习对象之间的空间对齐及其产生的声音。我们研究使用空间音频特征来表示音频输入,以及不同的音频格式:环境音、单声道和立体声。实验结果表明,与对数mel谱图特征相比,一阶环境声强度向量(FOA-IV)的AVSA提高了10$\%$;面向对象crops的加入也显著提高了人类行为识别下游任务的性能。设计了许多纯音频下游任务,用于测试学习的音频特征表示的有效性,获得与最新的基于环境音和双耳音频的声场景分类方法相当的性能。
摘要:Learning from audio-visual data offers many possibilities to express correspondence between the audio and visual content, similar to the human perception that relates aural and visual information. In this work, we present a method for self-supervised representation learning based on audio-visual spatial alignment (AVSA), a more sophisticated alignment task than the audio-visual correspondence (AVC). In addition to the correspondence, AVSA also learns from the spatial location of acoustic and visual content. Based on 360$^\text{o}$ video and Ambisonics audio, we propose selection of visual objects using object detection, and beamforming of the audio signal towards the detected objects, attempting to learn the spatial alignment between objects and the sound they produce. We investigate the use of spatial audio features to represent the audio input, and different audio formats: Ambisonics, mono, and stereo. Experimental results show a 10 $\%$ improvement on AVSA for the first order ambisonics intensity vector (FOA-IV) in comparison with log-mel spectrogram features; the addition of object-oriented crops also brings significant performance increases for the human action recognition downstream task. A number of audio-only downstream tasks are devised for testing the effectiveness of the learnt audio feature representation, obtaining performance comparable to state-of-the-art methods on acoustic scene classification from ambisonic and binaural audio.


【5】 Pronunciation Dictionary-Free Multilingual Speech Synthesis by Combining  Unsupervised and Supervised Phonetic Representations

标题:无监督和有监督语音表示相结合的无发音词典多语种语音合成

链接:https://arxiv.org/abs/2206.00951

作者:Chang Liu,Zhen-Hua Ling,Ling-Hui Chen
机构:NERC-SLIP, University of Science and Technology of China, Hefei, China, FLYTEK Research, Hefei, China
备注:Submitted to Interspeech 2022
摘要:本文提出了一种结合无监督语音表示(UPR)和有监督语音表示(SPR)的多语言语音合成方法,以避免对目标语言语音词典的依赖。该方法采用预训练的wav2vec 2.0模型提取UPR,并利用连接主义时间分类(CTC)损失建立与语言无关的自动语音识别(LI-ASR)模型,从目标语言的音频数据中提取段级SPR。然后,设计了一个声学模型,该模型首先从文本中分别预测UPRs和SPR,然后将预测的UPRs和SPR组合生成mel谱图。我们在六种语言上的实验结果表明,该方法优于直接从字符或音素序列预测mel谱图的方法和仅使用UPR或SPR的烧蚀模型。
摘要:This paper proposes a multilingual speech synthesis method which combines unsupervised phonetic representations (UPR) and supervised phonetic representations (SPR) to avoid reliance on the pronunciation dictionaries of target languages. In this method, a pretrained wav2vec 2.0 model is adopted to extract UPRs and a language-independent automatic speech recognition (LI-ASR) model is built with a connectionist temporal classification (CTC) loss to extract segment-level SPRs from the audio data of target languages. Then, an acoustic model is designed, which first predicts UPRs and SPRs from texts separately and then combines the predicted UPRs and SPRs to generate mel-spectrograms. The results of our experiments on six languages show that the proposed method outperformed the methods that directly predicted mel-spectrograms from character or phoneme sequences and the ablated models that utilized only UPRs or SPRs.


【6】 Squeezeformer: An Efficient Transformer for Automatic Speech Recognition

标题:Squeezeformer:一种高效的自动语音识别转换器

链接:https://arxiv.org/abs/2206.00888

作者:Sehoon Kim,Amir Gholami,Albert Shaw,Nicholas Lee,Karttikeya Mangalam,Jitendra Malik,Michael W. Mahoney,Kurt Keutzer
机构:University of California, Berkeley,ICSI
摘要:最近提出的构象模型基于其捕获局部和全局特征的混合注意卷积结构,已成为各种下游语音任务的事实骨干模型。然而,通过一系列系统的研究,我们发现一致性架构的设计选择并不是最优的。在重新检查Conformer宏观和微观架构的设计选择后,我们提出了挤压成型器模型,在相同的训练方案下,该模型始终优于最先进的ASR模型。特别是,对于宏体系结构,挤压成型器结合了(i)时间U网结构,这降低了长序列上多头注意模块的成本,以及(ii)前馈模块的简单块结构,随后是多头注意或卷积模块,而不是Conformer中提出的Macaron结构。此外,对于微架构,挤压成型器(i)简化卷积块中的激活,(ii)移除冗余层归一化操作,以及(iii)合并有效的深度向下采样层以有效地对输入信号进行子采样。在没有外部语言模型的情况下,Squeezeformer在Librispeech测试other上取得了7.5%、6.5%和6.0%的最新结果。在失败次数相同的情况下,这比一致性CTC好3.1%、1.4%和0.6%。我们的代码是开源的,可以在线获取。
摘要:The recently proposed Conformer model has become the de facto backbone model for various downstream speech tasks based on its hybrid attention-convolution architecture that captures both local and global features. However, through a series of systematic studies, we find that the Conformer architecture's design choices are not optimal. After reexamining the design choices for both the macro and micro-architecture of Conformer, we propose the Squeezeformer model, which consistently outperforms the state-of-the-art ASR models under the same training schemes. In particular, for the macro-architecture, Squeezeformer incorporates (i) the Temporal U-Net structure, which reduces the cost of the multi-head attention modules on long sequences, and (ii) a simpler block structure of feed-forward module, followed up by multi-head attention or convolution modules, instead of the Macaron structure proposed in Conformer. Furthermore, for the micro-architecture, Squeezeformer (i) simplifies the activations in the convolutional block, (ii) removes redundant Layer Normalization operations, and (iii) incorporates an efficient depth-wise downsampling layer to efficiently sub-sample the input signal. Squeezeformer achieves state-of-the-art results of 7.5%, 6.5%, and 6.0% word-error-rate on Librispeech test-other without external language models. This is 3.1%, 1.4%, and 0.6% better than Conformer-CTC with the same number of FLOPs. Our code is open-sourced and available online.


eess.AS音频处理

【1】 Self-supervised Learning of Audio Representations from Audio-Visual Data  using Spatial Alignment

标题:基于空间对齐的视听数据音频表示的自监督学习

链接:https://arxiv.org/abs/2206.00970

作者:Shanshan Wang,Archontis Politis,Annamaria Mesaros,Tuomas Virtanen
机构:Tampere University
摘要:从视听数据中学习为表达视听内容之间的对应关系提供了许多可能性,类似于人类对听觉和视觉信息的感知。在这项工作中,我们提出了一种基于视听空间对齐(AVSA)的自监督表示学习方法,AVSA是一种比视听对应(AVC)更复杂的对齐任务。除了通信之外,AVSA还从声音和视觉内容的空间位置学习。基于360$^\text{o}$视频和环境音音频,我们提出了使用对象检测选择视觉对象,并将音频信号向检测到的对象进行波束形成,试图学习对象之间的空间对齐及其产生的声音。我们研究使用空间音频特征来表示音频输入,以及不同的音频格式:环境音、单声道和立体声。实验结果表明,与对数mel谱图特征相比,一阶环境声强度向量(FOA-IV)的AVSA提高了10$\%$;面向对象crops的加入也显著提高了人类行为识别下游任务的性能。设计了许多纯音频下游任务,用于测试学习的音频特征表示的有效性,获得与最新的基于环境音和双耳音频的声场景分类方法相当的性能。
摘要:Learning from audio-visual data offers many possibilities to express correspondence between the audio and visual content, similar to the human perception that relates aural and visual information. In this work, we present a method for self-supervised representation learning based on audio-visual spatial alignment (AVSA), a more sophisticated alignment task than the audio-visual correspondence (AVC). In addition to the correspondence, AVSA also learns from the spatial location of acoustic and visual content. Based on 360$^\text{o}$ video and Ambisonics audio, we propose selection of visual objects using object detection, and beamforming of the audio signal towards the detected objects, attempting to learn the spatial alignment between objects and the sound they produce. We investigate the use of spatial audio features to represent the audio input, and different audio formats: Ambisonics, mono, and stereo. Experimental results show a 10 $\%$ improvement on AVSA for the first order ambisonics intensity vector (FOA-IV) in comparison with log-mel spectrogram features; the addition of object-oriented crops also brings significant performance increases for the human action recognition downstream task. A number of audio-only downstream tasks are devised for testing the effectiveness of the learnt audio feature representation, obtaining performance comparable to state-of-the-art methods on acoustic scene classification from ambisonic and binaural audio.


【2】 Pronunciation Dictionary-Free Multilingual Speech Synthesis by Combining  Unsupervised and Supervised Phonetic Representations

标题:无监督和有监督语音表示相结合的无发音词典多语种语音合成

链接:https://arxiv.org/abs/2206.00951

作者:Chang Liu,Zhen-Hua Ling,Ling-Hui Chen
机构:NERC-SLIP, University of Science and Technology of China, Hefei, China, FLYTEK Research, Hefei, China
备注:Submitted to Interspeech 2022
摘要:本文提出了一种结合无监督语音表示(UPR)和有监督语音表示(SPR)的多语言语音合成方法,以避免对目标语言语音词典的依赖。该方法采用预训练的wav2vec 2.0模型提取UPR,并利用连接主义时间分类(CTC)损失建立与语言无关的自动语音识别(LI-ASR)模型,从目标语言的音频数据中提取段级SPR。然后,设计了一个声学模型,该模型首先从文本中分别预测UPRs和SPR,然后将预测的UPRs和SPR组合生成mel谱图。我们在六种语言上的实验结果表明,该方法优于直接从字符或音素序列预测mel谱图的方法和仅使用UPR或SPR的烧蚀模型。
摘要:This paper proposes a multilingual speech synthesis method which combines unsupervised phonetic representations (UPR) and supervised phonetic representations (SPR) to avoid reliance on the pronunciation dictionaries of target languages. In this method, a pretrained wav2vec 2.0 model is adopted to extract UPRs and a language-independent automatic speech recognition (LI-ASR) model is built with a connectionist temporal classification (CTC) loss to extract segment-level SPRs from the audio data of target languages. Then, an acoustic model is designed, which first predicts UPRs and SPRs from texts separately and then combines the predicted UPRs and SPRs to generate mel-spectrograms. The results of our experiments on six languages show that the proposed method outperformed the methods that directly predicted mel-spectrograms from character or phoneme sequences and the ablated models that utilized only UPRs or SPRs.


【3】 Squeezeformer: An Efficient Transformer for Automatic Speech Recognition

标题:Squeezeformer:一种高效的自动语音识别转换器

链接:https://arxiv.org/abs/2206.00888

作者:Sehoon Kim,Amir Gholami,Albert Shaw,Nicholas Lee,Karttikeya Mangalam,Jitendra Malik,Michael W. Mahoney,Kurt Keutzer
机构:University of California, Berkeley,ICSI
摘要:最近提出的构象模型基于其捕获局部和全局特征的混合注意卷积结构,已成为各种下游语音任务的事实骨干模型。然而,通过一系列系统的研究,我们发现一致性架构的设计选择并不是最优的。在重新检查Conformer宏观和微观架构的设计选择后,我们提出了挤压成型器模型,在相同的训练方案下,该模型始终优于最先进的ASR模型。特别是,对于宏体系结构,挤压成型器结合了(i)时间U网结构,这降低了长序列上多头注意模块的成本,以及(ii)前馈模块的简单块结构,随后是多头注意或卷积模块,而不是Conformer中提出的Macaron结构。此外,对于微架构,挤压成型器(i)简化卷积块中的激活,(ii)移除冗余层归一化操作,以及(iii)合并有效的深度向下采样层以有效地对输入信号进行子采样。在没有外部语言模型的情况下,Squeezeformer在Librispeech测试other上取得了7.5%、6.5%和6.0%的最新结果。在失败次数相同的情况下,这比一致性CTC好3.1%、1.4%和0.6%。我们的代码是开源的,可以在线获取。
摘要:The recently proposed Conformer model has become the de facto backbone model for various downstream speech tasks based on its hybrid attention-convolution architecture that captures both local and global features. However, through a series of systematic studies, we find that the Conformer architecture's design choices are not optimal. After reexamining the design choices for both the macro and micro-architecture of Conformer, we propose the Squeezeformer model, which consistently outperforms the state-of-the-art ASR models under the same training schemes. In particular, for the macro-architecture, Squeezeformer incorporates (i) the Temporal U-Net structure, which reduces the cost of the multi-head attention modules on long sequences, and (ii) a simpler block structure of feed-forward module, followed up by multi-head attention or convolution modules, instead of the Macaron structure proposed in Conformer. Furthermore, for the micro-architecture, Squeezeformer (i) simplifies the activations in the convolutional block, (ii) removes redundant Layer Normalization operations, and (iii) incorporates an efficient depth-wise downsampling layer to efficiently sub-sample the input signal. Squeezeformer achieves state-of-the-art results of 7.5%, 6.5%, and 6.0% word-error-rate on Librispeech test-other without external language models. This is 3.1%, 1.4%, and 0.6% better than Conformer-CTC with the same number of FLOPs. Our code is open-sourced and available online.


【4】 The match file format: Encoding Alignments between Scores and  Performances

标题:匹配文件格式:对分数和表演之间的对齐进行编码

链接:https://arxiv.org/abs/2206.01104

作者:Francesco Foscarin,Emmanouil Karystinaios,Silvan David Peter,Carlos Cancino-Chacón,Maarten Grachten,Gerhard Widmer
机构:Johannes Kepler University, Carlos Cancino-Chac´on∗, carlos eduardo., Independent Researcher
摘要:本文介绍了match的规范:一种文件格式,它将MIDI人体表演与音符、节拍和重拍级别对齐扩展到相应的乐谱。这可以对与各种任务相关的性能进行高级分析,例如表现性性能建模、乐谱跟踪、音乐转录和演奏者分类。匹配文件包括一组分数相关描述符,使其也可用作基本分数表示。对于需要使用结构分数元素的应用程序(例如,语音、部件、梁、音调),可以轻松地将匹配文件与符号分数组合。为了支持我们工作的实际应用,我们发布了Vienna4x22分数和成绩数据集的修正和升级版本,该数据集与比赛文件保持一致。
摘要:This paper presents the specifications of match: a file format that extends a MIDI human performance with note-, beat-, and downbeat-level alignments to a corresponding musical score. This enables advanced analyses of the performance that are relevant for various tasks, such as expressive performance modeling, score following, music transcription, and performer classification. The match file includes a set of score-related descriptors that makes it usable also as a bare-bones score representation. For applications that require the use of structural score elements (e.g., voices, parts, beams, slurs), the match file can be easily combined with the symbolic score. To support the practical application of our work, we release a corrected and upgraded version of the Vienna4x22 dataset of scores and performances aligned with match files.


【5】 Partitura: A Python Package for Symbolic Music Processing

标题:Partitura:一种用于符号音乐处理的Python包

链接:https://arxiv.org/abs/2206.01071

作者:Carlos Cancino-Chacón,Silvan David Peter,Emmanouil Karystinaios,Francesco Foscarin,Maarten Grachten,Gerhard Widmer
机构:Johannes Kepler University, Independent Researcher
摘要:Partitura是一个轻量级Python包,用于处理符号音乐信息。它可以方便地访问音乐信息检索任务中常用的功能,如音符数组(定时音调事件列表)和2D钢琴滚动矩阵,以及其他乐谱元素,如时间和关键签名、演奏指令和重复结构。Partitura可以加载乐谱(MEI、MusicXML、Kern和MIDI格式)、MIDI表演以及乐谱与表演的对齐。该软件包包括一些音乐分析工具,如自动音高拼写、关键签名识别和语音分离。Partitura是一个开源项目,可在https://github.com/CPJKU/partitura/.
摘要:Partitura is a lightweight Python package for handling symbolic musical information. It provides easy access to features commonly used in music information retrieval tasks, like note arrays (lists of timed pitched events) and 2D piano roll matrices, as well as other score elements such as time and key signatures, performance directives, and repeat structures. Partitura can load musical scores (in MEI, MusicXML, Kern, and MIDI formats), MIDI performances, and score-to-performance alignments. The package includes some tools for music analysis, such as automatic pitch spelling, key signature identification, and voice separation. Partitura is an open-source project and is available at https://github.com/CPJKU/partitura/.


【6】 Musical Instrument Recognition by XGBoost Combining Feature Fusion

标题:基于XGBoost结合特征融合的乐器识别

链接:https://arxiv.org/abs/2206.00901

作者:Yijie Liu,Yanfang Yin,Qigang Zhu,Wenzhuo Cui
机构:Shandong Univ Sci & Technol, Dept Elect Engn & Informat Technol, Jinan , Peoples R China;,
摘要:乐器分类是音乐信息检索的研究热点之一。为了解决现有乐器分类模型性能较差的问题,提出了一种基于多通道特征融合和XGBoost的乐器分类算法。在数据集音频特征提取和融合的基础上,将特征输入到XGBoost模型中进行训练;其次,通过比较不同的特征组合和朴素贝叶斯等经典机器学习模型,验证了该算法在乐器分类任务中的优越性能。该算法在Medley solos DB数据集上的准确率达到97.65%,优于现有模型。实验结果为乐器分类特征工程中的特征选择提供了参考。
摘要:Musical instrument classification is one of the focuses of Music Information Retrieval (MIR). In order to solve the problem of poor performance of current musical instrument classification models, we propose a musical instrument classification algorithm based on multi-channel feature fusion and XGBoost. Based on audio feature extraction and fusion of the dataset, the features are input into the XGBoost model for training; secondly, we verified the superior performance of the algorithm in the musical instrument classification task by com-paring different feature combinations and several classical machine learning models such as Naive Bayes. The algorithm achieves an accuracy of 97.65% on the Medley-solos-DB dataset, outperforming existing models. The experiments provide a reference for feature selection in feature engineering for musical instrument classification.


机器翻译,仅供参考