今天跟大家分享一篇语音相关的论文合集:cs.SD语音8篇,eess.AS音频处理8篇。本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
【1】 Towards robust music source separation on loud commercial music
标题:在吵闹的商业音乐上实现稳健的音乐源分离
链接:https://arxiv.org/abs/2208.14355
作者:Chang-Bin Jeon,Kyogu Lee备注:Accepted to ISMIR 2022摘要:与过去相比,如今的商业音乐具有极高的响度和高度压缩的动态范围。然而,在音乐源分离中,这些特性没有得到充分的考虑,导致实验室与现实世界的领域不匹配。在本文中,我们证实了这种域失配对音乐源分离网络的性能有负面影响。为此,我们首先通过模仿音乐母带制作过程创建了域外评估数据集musdb-L和XL。然后,我们定量地验证了最先进的算法在我们的数据集中的性能显著恶化。最后,提出了LimitAug数据增强方法,该方法在训练数据采样过程中使用在线限幅器,以减少域失配。我们证实,它不仅缓解了域外数据集的性能下降,而且还提高了域内数据的性能。摘要:Nowadays, commercial music has extreme loudness and heavily compressed dynamic range compared to the past. Yet, in music source separation, these characteristics have not been thoroughly considered, resulting in the domain mismatch between the laboratory and the real world. In this paper, we confirmed that this domain mismatch negatively affect the performance of the music source separation networks. To this end, we first created the out-of-domain evaluation datasets, musdb-L and XL, by mimicking the music mastering process. Then, we quantitatively verify that the performance of the state-of-the-art algorithms significantly deteriorated in our datasets. Lastly, we proposed LimitAug data augmentation method to reduce the domain mismatch, which utilizes an online limiter during the training data sampling process. We confirmed that it not only alleviates the performance degradation on our out-of-domain datasets, but also results in higher performance on in-domain data.
【2】 MeloForm: Generating Melody with Musical Form based on Expert Systems and Neural Networks
标题:MeloForm:基于专家系统和神经网络的乐曲形式生成
链接:https://arxiv.org/abs/2208.14345
作者:Peiling Lu,Xu Tan,Botao Yu,Tao Qin,Sheng Zhao,Tie-Yan Liu摘要:人类在创作音乐时,通常是按照音乐的形式来组织音乐元素,以表达音乐的思想。然而,对于基于神经网络的音乐生成,由于缺乏关于音乐形式的标记数据,这是很难做到的。本文利用专家系统和神经网络技术开发了一个具有曲式的旋律生成系统MeloForm。具体来说,1)我们设计了一个专家系统,通过根据预先给定的音乐形式将音乐元素从母题发展到乐句,然后发展到具有重复和变化的部分来生成旋律; 2)针对生成的旋律缺乏音乐丰富性的问题,设计了Transformer基于Transformer的细化模型,在不改变旋律音乐形式的前提下,对旋律进行细化。MeloForm的优势在于通过专家系统实现精确的音乐形式控制,以及通过神经模型实现丰富的音乐学习。主观和客观的实验评价均表明,MeloForm生成的旋律具有精确的曲式控制,准确率为97.79%,在结构、主题、丰富性和整体质量方面的主观评价得分比基线系统分别高出0.75、0.50、0.86和0.89分,且不需要任何曲式数据。此外,MeloForm还可以支持多种曲式,如韵文和合唱曲式、回旋曲式、变奏曲式、奏鸣曲曲式等。摘要:Human usually composes music by organizing elements according to the musical form to express music ideas. However, for neural network-based music generation, it is difficult to do so due to the lack of labelled data on musical form. In this paper, we develop MeloForm, a system that generates melody with musical form using expert systems and neural networks. Specifically, 1) we design an expert system to generate a melody by developing musical elements from motifs to phrases then to sections with repetitions and variations according to pre-given musical form; 2) considering the generated melody is lack of musical richness, we design a Transformer based refinement model to improve the melody without changing its musical form. MeloForm enjoys the advantages of precise musical form control by expert systems and musical richness learning via neural models. Both subjective and objective experimental evaluations demonstrate that MeloForm generates melodies with precise musical form control with 97.79% accuracy, and outperforms baseline systems in terms of subjective evaluation score by 0.75, 0.50, 0.86 and 0.89 in structure, thematic, richness and overall quality, without any labelled musical form data. Besides, MeloForm can support various kinds of forms, such as verse and chorus form, rondo form, variational form, sonata form, etc.
【3】 HPPNet: Modeling the Harmonic Structure and Pitch Invariance in Piano Transcription
标题:HPPNet:钢琴改编中的调和结构和音调不变性建模
链接:https://arxiv.org/abs/2208.14339
作者:Weixing Wei,Peilin Li,Yi Yu,Wei Li摘要:虽然神经网络模型在钢琴改编方面取得了重大进展,但由于需要更大的模型尺寸和更高的计算能力,它们变得更加消耗资源。本文尝试应用更多的钢琴先验知识来减少模型的规模,提高转录性能。钢琴音符的声音包含各种泛音,并且键的音高不会随时间而改变。为了充分利用这些潜在的信息,提出了HPPNet网络,该网络使用谐波膨胀卷积来捕获谐波结构,并使用频率分组递归神经网络来建模基音随时间的不变性。在MAESTRO数据集上的实验结果表明,我们的钢琴转录系统在帧和音符评分方面都达到了最高水平(帧F1为93.15%,音符F1为97.18%)。此外,该模型的规模远小于之前最先进的深度学习模型。摘要:While neural network models are making significant progress in piano transcription, they are becoming more resource-consuming due to requiring larger model size and more computing power. In this paper, we attempt to apply more prior about piano to reduce model size and improve the transcription performance. The sound of a piano note contains various overtones, and the pitch of a key does not change over time. To make full use of such latent information, we propose HPPNet that using the Harmonic Dilated Convolution to capture the harmonic structures and the Frequency Grouped Recurrent Neural Network to model the pitch-invariance over time. Experimental results on the MAESTRO dataset show that our piano transcription system achieves state-of-the-art performance both in frame and note scores (frame F1 93.15%, note F1 97.18%). Moreover, the model size is much smaller than the previous state-of-the-art deep learning models.
【4】 Gridless 3D Recovery of Image Sources from Room Impulse Responses
标题:室内脉冲响应图像源的无网格三维恢复
链接:https://arxiv.org/abs/2208.14017
作者:Tom Sprunck,Yannick Privat,Cédric Foy,Antoine Deleforge摘要:给定由脉冲图像源的稀疏分布产生的声场,这些源的连续3D位置和幅度是否可以从有限位置集合处的场的离散的、带宽受限的测量中恢复,多通道房间脉冲响应?借鉴超分辨率成像的最新进展,证明了这个非线性、非凸反问题可以有效地松弛为R3中Radon测度空间上的凸线性反问题.这里引入的线性算子源于自由场非齐次波动方程的基本解,并结合接收器的响应。本文提出了一种改进的滑动Frank-Wolfe算法来数值求解网格外问题,在连续的三维空间中。仿真实验表明,该方法可以在任意矩形房间内使用任意放置的32通道紧凑型球形麦克风阵列,实现数百个图像源的近精确恢复。文中还分析了噪声、采样率和阵列直径对这些结果的影响。摘要:Given a sound field generated by a sparse distribution of impulse image sources, can the continuous 3D positions and amplitudes of these sources be recovered from discrete, bandlimited measurements of the field at a finite set of locations, e.g., a multichannel room impulse response? Borrowing from recent advances in super-resolution imaging, it is shown that this nonlinear, non-convex inverse problem can be efficiently relaxed into a convex linear inverse problem over the space of Radon measures in R3. The linear operator introduced here stems from the fundamental solution of the free-field inhomogenous wave equation combined with the receivers' responses. An adaptation of the Sliding Frank-Wolfe algorithm is proposed to numerically solve the problem off-the-grid, i.e., in continuous 3D space. Simulated experiments show that the approach achieves near-exact recovery of hundreds of image sources using an arbitrarily placed compact 32-channel spherical microphone array in random rectangular rooms. The impact of noise, sampling rate and array diameter on these results is also examined.
【5】 Video-based Cross-modal Auxiliary Network for Multimodal Sentiment Analysis
标题:基于视频的跨通道辅助网络多通道情感分析
链接:https://arxiv.org/abs/2208.13954
作者:Rongfei Chen,Wenju Zhou,Yang Li,Huiyu Zhou摘要:多模态情感分析因其在多模态交互中的信息互补性而具有广泛的应用前景。以往的研究多集中于研究有效的联合表示,而很少考虑多模态融合的单峰特征提取不足和数据冗余问题。本文提出了一种基于视频的跨模态辅助网络(VCAN),该网络由音频特征映射模块和跨模态选择模块组成。第一模块被设计为实质上增加音频特征提取中的特征多样性,旨在通过提供更全面的声学表示来提高分类准确度。为了使模型能够处理冗余视觉特征,第二模块被寻址为在集成视听数据期间有效地过滤冗余视觉帧。此外,引入了由多个图像分类网络组成的分类器组来预测情感极性和情感类别。在RAVDESS、CMU-MOSI和CMU-MOSEI基准测试上的大量实验结果表明,VCAN在提高多模态情感分析的分类准确率方面明显优于现有的方法.摘要:Multimodal sentiment analysis has a wide range of applications due to its information complementarity in multimodal interactions. Previous works focus more on investigating efficient joint representations, but they rarely consider the insufficient unimodal features extraction and data redundancy of multimodal fusion. In this paper, a Video-based Cross-modal Auxiliary Network (VCAN) is proposed, which is comprised of an audio features map module and a cross-modal selection module. The first module is designed to substantially increase feature diversity in audio feature extraction, aiming to improve classification accuracy by providing more comprehensive acoustic representations. To empower the model to handle redundant visual features, the second module is addressed to efficiently filter the redundant visual frames during integrating audiovisual data. Moreover, a classifier group consisting of several image classification networks is introduced to predict sentiment polarities and emotion categories. Extensive experimental results on RAVDESS, CMU-MOSI, and CMU-MOSEI benchmarks indicate that VCAN is significantly superior to the state-of-the-art methods for improving the classification accuracy of multimodal sentiment analysis.
【6】 A Study on the relationship between the geometrical shapes and the biometrical acoustic characteristics of human ear canal
标题:人体耳道几何形状与生物声学特性关系的研究
链接:https://arxiv.org/abs/2208.14182
作者:Riki Kimura,Shunsuke Tanaka,Naoki Wakui,Naoki Kodama,Shohei Yano摘要:耳声认证是一种新的生物特征识别方法,它利用了用户耳道声学特性的差异。然而,关于引起声学特性差异的因素的报道很少。我们研究了耳道形状与用户间相似性方面的声学特征之间的关系。我们使用磁共振成像(MRI)来测量耳道几何形状。结果,形状相似度和声学特征相似度之间的相关系数高于0.7,并且确定系数高于0.5。这表明耳道形状的差异是重要因素之一。摘要:Ear acoustic authentication is a new biometrics method and it utilizes the differences in acoustic characteristics of the ear canal between users. However, there have been few reports on the factors that cause differences in the acoustic characteristics. We investigate the relationship between ear canal shapes and acoustic characteristics in terms of user-to-user similarity. We used magnetic resonance imaging (MRI) to measure ear canal geometry. As a result, the correlation coefficient between shape similarity and acoustic characteristic similarity is higher than 0.7 and the coefficient of determination is higher than 0.5. This suggests that the difference in the shape of the ear canal is one of the important factors.
【7】 Classify Respiratory Abnormality in Lung Sounds Using STFT and a Fine-Tuned ResNet18 Network
标题:利用短时傅立叶变换和改进的ResNet18网络对肺音呼吸异常进行分类
链接:https://arxiv.org/abs/2208.13943
作者:Zizhao Chen,Hongliang Wang,Chia-Hui Yeh,Xilin Liu摘要:识别肺音中的模式对于检测和监测呼吸系统疾病至关重要。用于分析呼吸声音的当前技术需要领域专家并且需要进行解释。因此,需要一种精确且自动的呼吸声分类系统。在这项工作中,我们采取了一种数据驱动的方法来分类异常肺音。我们比较了使用三种不同的特征提取技术(短时傅立叶变换(STFT)、Mel谱图和Wav2vec)以及三种不同的分类Transformer(预训练的ResNet18、LightCNN和Audio Spectrogram Transformer)的性能。我们的主要贡献包括对不同音频特征提取器和基于神经网络的分类器进行基准测试,以及使用STFT和微调的ResNet18网络实现完整的流水线。在IEEE BioCAS 2022呼吸音分类大挑战中,对于任务1-1、1-2、2-1和2-2,所提出的方法在测试集上分别实现了0.89、0.80、0.71、0.36的谐波得分。摘要:Recognizing patterns in lung sounds is crucial to detecting and monitoring respiratory diseases. Current techniques for analyzing respiratory sounds demand domain experts and are subject to interpretation. Hence an accurate and automatic respiratory sound classification system is desired. In this work, we took a data-driven approach to classify abnormal lung sounds. We compared the performance using three different feature extraction techniques, which are short-time Fourier transformation (STFT), Mel spectrograms, and Wav2vec, as well as three different classifiers, including pre-trained ResNet18, LightCNN, and Audio Spectrogram Transformer. Our key contributions include the bench-marking of different audio feature extractors and neural network based classifiers, and the implementation of a complete pipeline using STFT and a fine-tuned ResNet18 network. The proposed method achieved Harmonic Scores of 0.89, 0.80, 0.71, 0.36 for tasks 1-1, 1-2, 2-1 and 2-2, respectively on the testing sets in the IEEE BioCAS 2022 Grand Challenge on Respiratory Sound Classification.
【8】 A Language Agnostic Multilingual Streaming On-Device ASR System
标题:一种与语言无关的设备上多语种流媒体ASR系统
链接:https://arxiv.org/abs/2208.13916
作者:Bo Li,Tara N. Sainath,Ruoming Pang,Shuo-yiin Chang,Qiumin Xu,Trevor Strohman,Vince Chen,Qiao Liang,Heguang Liu,Yanzhang He,Parisa Haghani,Sameer Bidichandani备注:Accepted in Interspeech 2022摘要:在英语语音搜索任务上,设备上端到端(E2E)模型在质量和延迟方面都比传统模型有所改进。E2E模型对于多语言自动语音识别(ASR)也显示出有希望的结果。在这篇文章中,我们将我们先前的容量解决方案扩展到流应用,并提出了一个流多语言E2E ASR系统,该系统完全运行在设备上,具有与单个单语言模型相当的质量和延迟。为了实现这一点,我们提出了一个编码器端点指针模型和一个语音结束(EOU)联合层,以获得更好的质量和延迟折衷。我们的系统是以一种语言不可知的方式构建的,允许它本机支持实时的句间代码切换。为了解决大型模型的可行性问题,我们进行了设备上分析,并将耗时的LSTM解码器替换为最近开发的嵌入式解码器。通过这些更改,我们成功地在移动设备上以低于实时的速度运行了这样一个系统。摘要:On-device end-to-end (E2E) models have shown improvements over a conventional model on English Voice Search tasks in both quality and latency. E2E models have also shown promising results for multilingual automatic speech recognition (ASR). In this paper, we extend our previous capacity solution to streaming applications and present a streaming multilingual E2E ASR system that runs fully on device with comparable quality and latency to individual monolingual models. To achieve that, we propose an Encoder Endpointer model and an End-of-Utterance (EOU) Joint Layer for a better quality and latency trade-off. Our system is built in a language agnostic manner allowing it to natively support intersentential code switching in real time. To address the feasibility concerns on large models, we conducted on-device profiling and replaced the time consuming LSTM decoder with the recently developed Embedding decoder. With these changes, we managed to run such a system on a mobile device in less than real time.
【1】 A Study on the relationship between the geometrical shapes and the biometrical acoustic characteristics of human ear canal
标题:人体耳道几何形状与生物声学特性关系的研究
链接:https://arxiv.org/abs/2208.14182
* 与cs.SD语音【6】为同一篇
作者:Riki Kimura,Shunsuke Tanaka,Naoki Wakui,Naoki Kodama,Shohei Yano摘要:耳声认证是一种新的生物特征识别方法,它利用了用户耳道声学特性的差异。然而,关于引起声学特性差异的因素的报道很少。我们研究了耳道形状与用户间相似性方面的声学特征之间的关系。我们使用磁共振成像(MRI)来测量耳道几何形状。结果,形状相似度和声学特征相似度之间的相关系数高于0.7,并且确定系数高于0.5。这表明耳道形状的差异是重要因素之一。摘要:Ear acoustic authentication is a new biometrics method and it utilizes the differences in acoustic characteristics of the ear canal between users. However, there have been few reports on the factors that cause differences in the acoustic characteristics. We investigate the relationship between ear canal shapes and acoustic characteristics in terms of user-to-user similarity. We used magnetic resonance imaging (MRI) to measure ear canal geometry. As a result, the correlation coefficient between shape similarity and acoustic characteristic similarity is higher than 0.7 and the coefficient of determination is higher than 0.5. This suggests that the difference in the shape of the ear canal is one of the important factors.
【2】 Classify Respiratory Abnormality in Lung Sounds Using STFT and a Fine-Tuned ResNet18 Network
标题:利用短时傅立叶变换和改进的ResNet18网络对肺音呼吸异常进行分类
链接:https://arxiv.org/abs/2208.13943
* 与cs.SD语音【7】为同一篇
作者:Zizhao Chen,Hongliang Wang,Chia-Hui Yeh,Xilin Liu摘要:识别肺音中的模式对于检测和监测呼吸系统疾病至关重要。用于分析呼吸声音的当前技术需要领域专家并且需要进行解释。因此,需要一种精确且自动的呼吸声分类系统。在这项工作中,我们采取了一种数据驱动的方法来分类异常肺音。我们比较了使用三种不同的特征提取技术(短时傅立叶变换(STFT)、Mel谱图和Wav2vec)以及三种不同的分类Transformer(预训练的ResNet18、LightCNN和Audio Spectrogram Transformer)的性能。我们的主要贡献包括对不同音频特征提取器和基于神经网络的分类器进行基准测试,以及使用STFT和微调的ResNet18网络实现完整的流水线。在IEEE BioCAS 2022呼吸音分类大挑战中,对于任务1-1、1-2、2-1和2-2,所提出的方法在测试集上分别实现了0.89、0.80、0.71、0.36的谐波得分。摘要:Recognizing patterns in lung sounds is crucial to detecting and monitoring respiratory diseases. Current techniques for analyzing respiratory sounds demand domain experts and are subject to interpretation. Hence an accurate and automatic respiratory sound classification system is desired. In this work, we took a data-driven approach to classify abnormal lung sounds. We compared the performance using three different feature extraction techniques, which are short-time Fourier transformation (STFT), Mel spectrograms, and Wav2vec, as well as three different classifiers, including pre-trained ResNet18, LightCNN, and Audio Spectrogram Transformer. Our key contributions include the bench-marking of different audio feature extractors and neural network based classifiers, and the implementation of a complete pipeline using STFT and a fine-tuned ResNet18 network. The proposed method achieved Harmonic Scores of 0.89, 0.80, 0.71, 0.36 for tasks 1-1, 1-2, 2-1 and 2-2, respectively on the testing sets in the IEEE BioCAS 2022 Grand Challenge on Respiratory Sound Classification.
【3】 A Language Agnostic Multilingual Streaming On-Device ASR System
标题:一种与语言无关的设备上多语种流媒体ASR系统
链接:https://arxiv.org/abs/2208.13916
* 与cs.SD语音【8】为同一篇
作者:Bo Li,Tara N. Sainath,Ruoming Pang,Shuo-yiin Chang,Qiumin Xu,Trevor Strohman,Vince Chen,Qiao Liang,Heguang Liu,Yanzhang He,Parisa Haghani,Sameer Bidichandani备注:Accepted in Interspeech 2022摘要:在英语语音搜索任务上,设备上端到端(E2E)模型在质量和延迟方面都比传统模型有所改进。E2E模型对于多语言自动语音识别(ASR)也显示出有希望的结果。在这篇文章中,我们将我们先前的容量解决方案扩展到流应用,并提出了一个流多语言E2E ASR系统,该系统完全运行在设备上,具有与单个单语言模型相当的质量和延迟。为了实现这一点,我们提出了一个编码器端点指针模型和一个语音结束(EOU)联合层,以获得更好的质量和延迟折衷。我们的系统是以一种语言不可知的方式构建的,允许它本机支持实时的句间代码切换。为了解决大型模型的可行性问题,我们进行了设备上分析,并将耗时的LSTM解码器替换为最近开发的嵌入式解码器。通过这些更改,我们成功地在移动设备上以低于实时的速度运行了这样一个系统。摘要:On-device end-to-end (E2E) models have shown improvements over a conventional model on English Voice Search tasks in both quality and latency. E2E models have also shown promising results for multilingual automatic speech recognition (ASR). In this paper, we extend our previous capacity solution to streaming applications and present a streaming multilingual E2E ASR system that runs fully on device with comparable quality and latency to individual monolingual models. To achieve that, we propose an Encoder Endpointer model and an End-of-Utterance (EOU) Joint Layer for a better quality and latency trade-off. Our system is built in a language agnostic manner allowing it to natively support intersentential code switching in real time. To address the feasibility concerns on large models, we conducted on-device profiling and replaced the time consuming LSTM decoder with the recently developed Embedding decoder. With these changes, we managed to run such a system on a mobile device in less than real time.
【4】 Towards robust music source separation on loud commercial music
标题:在吵闹的商业音乐上实现稳健的音乐源分离
链接:https://arxiv.org/abs/2208.14355
* 与cs.SD语音【1】为同一篇
作者:Chang-Bin Jeon,Kyogu Lee备注:Accepted to ISMIR 2022摘要:与过去相比,如今的商业音乐具有极高的响度和高度压缩的动态范围。然而,在音乐源分离中,这些特性没有得到充分的考虑,导致实验室与现实世界的领域不匹配。在本文中,我们证实了这种域失配对音乐源分离网络的性能有负面影响。为此,我们首先通过模仿音乐母带制作过程创建了域外评估数据集musdb-L和XL。然后,我们定量地验证了最先进的算法在我们的数据集中的性能显著恶化。最后,提出了LimitAug数据增强方法,该方法在训练数据采样过程中使用在线限幅器,以减少域失配。我们证实,它不仅缓解了域外数据集的性能下降,而且还提高了域内数据的性能。摘要:Nowadays, commercial music has extreme loudness and heavily compressed dynamic range compared to the past. Yet, in music source separation, these characteristics have not been thoroughly considered, resulting in the domain mismatch between the laboratory and the real world. In this paper, we confirmed that this domain mismatch negatively affect the performance of the music source separation networks. To this end, we first created the out-of-domain evaluation datasets, musdb-L and XL, by mimicking the music mastering process. Then, we quantitatively verify that the performance of the state-of-the-art algorithms significantly deteriorated in our datasets. Lastly, we proposed LimitAug data augmentation method to reduce the domain mismatch, which utilizes an online limiter during the training data sampling process. We confirmed that it not only alleviates the performance degradation on our out-of-domain datasets, but also results in higher performance on in-domain data.
【5】 MeloForm: Generating Melody with Musical Form based on Expert Systems and Neural Networks
标题:MeloForm:基于专家系统和神经网络的乐曲形式生成
链接:https://arxiv.org/abs/2208.14345
* 与cs.SD语音【2】为同一篇
作者:Peiling Lu,Xu Tan,Botao Yu,Tao Qin,Sheng Zhao,Tie-Yan Liu摘要:人类在创作音乐时,通常是按照音乐的形式来组织音乐元素,以表达音乐的思想。然而,对于基于神经网络的音乐生成,由于缺乏关于音乐形式的标记数据,这是很难做到的。本文利用专家系统和神经网络技术开发了一个具有曲式的旋律生成系统MeloForm。具体来说,1)我们设计了一个专家系统,通过根据预先给定的音乐形式将音乐元素从母题发展到乐句,然后发展到具有重复和变化的部分来生成旋律; 2)针对生成的旋律缺乏音乐丰富性的问题,设计了Transformer基于Transformer的细化模型,在不改变旋律音乐形式的前提下,对旋律进行细化。MeloForm的优势在于通过专家系统实现精确的音乐形式控制,以及通过神经模型实现丰富的音乐学习。主观和客观的实验评价均表明,MeloForm生成的旋律具有精确的曲式控制,准确率为97.79%,在结构、主题、丰富性和整体质量方面的主观评价得分比基线系统分别高出0.75、0.50、0.86和0.89分,且不需要任何曲式数据。此外,MeloForm还可以支持多种曲式,如韵文和合唱曲式、回旋曲式、变奏曲式、奏鸣曲曲式等。摘要:Human usually composes music by organizing elements according to the musical form to express music ideas. However, for neural network-based music generation, it is difficult to do so due to the lack of labelled data on musical form. In this paper, we develop MeloForm, a system that generates melody with musical form using expert systems and neural networks. Specifically, 1) we design an expert system to generate a melody by developing musical elements from motifs to phrases then to sections with repetitions and variations according to pre-given musical form; 2) considering the generated melody is lack of musical richness, we design a Transformer based refinement model to improve the melody without changing its musical form. MeloForm enjoys the advantages of precise musical form control by expert systems and musical richness learning via neural models. Both subjective and objective experimental evaluations demonstrate that MeloForm generates melodies with precise musical form control with 97.79% accuracy, and outperforms baseline systems in terms of subjective evaluation score by 0.75, 0.50, 0.86 and 0.89 in structure, thematic, richness and overall quality, without any labelled musical form data. Besides, MeloForm can support various kinds of forms, such as verse and chorus form, rondo form, variational form, sonata form, etc.
【6】 HPPNet: Modeling the Harmonic Structure and Pitch Invariance in Piano Transcription
标题:HPPNet:钢琴改编中的调和结构和音调不变性建模
链接:https://arxiv.org/abs/2208.14339
* 与cs.SD语音【3】为同一篇
作者:Weixing Wei,Peilin Li,Yi Yu,Wei Li摘要:虽然神经网络模型在钢琴改编方面取得了重大进展,但由于需要更大的模型尺寸和更高的计算能力,它们变得更加消耗资源。本文尝试应用更多的钢琴先验知识来减少模型的规模,提高转录性能。钢琴音符的声音包含各种泛音,并且键的音高不会随时间而改变。为了充分利用这些潜在的信息,提出了HPPNet网络,该网络使用谐波膨胀卷积来捕获谐波结构,并使用频率分组递归神经网络来建模基音随时间的不变性。在MAESTRO数据集上的实验结果表明,我们的钢琴转录系统在帧和音符评分方面都达到了最高水平(帧F1为93.15%,音符F1为97.18%)。此外,该模型的规模远小于之前最先进的深度学习模型。摘要:While neural network models are making significant progress in piano transcription, they are becoming more resource-consuming due to requiring larger model size and more computing power. In this paper, we attempt to apply more prior about piano to reduce model size and improve the transcription performance. The sound of a piano note contains various overtones, and the pitch of a key does not change over time. To make full use of such latent information, we propose HPPNet that using the Harmonic Dilated Convolution to capture the harmonic structures and the Frequency Grouped Recurrent Neural Network to model the pitch-invariance over time. Experimental results on the MAESTRO dataset show that our piano transcription system achieves state-of-the-art performance both in frame and note scores (frame F1 93.15%, note F1 97.18%). Moreover, the model size is much smaller than the previous state-of-the-art deep learning models.
【7】 Gridless 3D Recovery of Image Sources from Room Impulse Responses
标题:室内脉冲响应图像源的无网格三维恢复
链接:https://arxiv.org/abs/2208.14017
* 与cs.SD语音【4】为同一篇
作者:Tom Sprunck,Yannick Privat,Cédric Foy,Antoine Deleforge摘要:给定由脉冲图像源的稀疏分布产生的声场,这些源的连续3D位置和幅度是否可以从有限位置集合处的场的离散的、带宽受限的测量中恢复,多通道房间脉冲响应?借鉴超分辨率成像的最新进展,证明了这个非线性、非凸反问题可以有效地松弛为R3中Radon测度空间上的凸线性反问题.这里引入的线性算子源于自由场非齐次波动方程的基本解,并结合接收器的响应。本文提出了一种改进的滑动Frank-Wolfe算法来数值求解网格外问题,在连续的三维空间中。仿真实验表明,该方法可以在任意矩形房间内使用任意放置的32通道紧凑型球形麦克风阵列,实现数百个图像源的近精确恢复。文中还分析了噪声、采样率和阵列直径对这些结果的影响。摘要:Given a sound field generated by a sparse distribution of impulse image sources, can the continuous 3D positions and amplitudes of these sources be recovered from discrete, bandlimited measurements of the field at a finite set of locations, e.g., a multichannel room impulse response? Borrowing from recent advances in super-resolution imaging, it is shown that this nonlinear, non-convex inverse problem can be efficiently relaxed into a convex linear inverse problem over the space of Radon measures in R3. The linear operator introduced here stems from the fundamental solution of the free-field inhomogenous wave equation combined with the receivers' responses. An adaptation of the Sliding Frank-Wolfe algorithm is proposed to numerically solve the problem off-the-grid, i.e., in continuous 3D space. Simulated experiments show that the approach achieves near-exact recovery of hundreds of image sources using an arbitrarily placed compact 32-channel spherical microphone array in random rectangular rooms. The impact of noise, sampling rate and array diameter on these results is also examined.
【8】 Video-based Cross-modal Auxiliary Network for Multimodal Sentiment Analysis
标题:基于视频的跨通道辅助网络多通道情感分析
链接:https://arxiv.org/abs/2208.13954
* 与cs.SD语音【5】为同一篇
作者:Rongfei Chen,Wenju Zhou,Yang Li,Huiyu Zhou摘要:多模态情感分析因其在多模态交互中的信息互补性而具有广泛的应用前景。以往的研究多集中于研究有效的联合表示,而很少考虑多模态融合的单峰特征提取不足和数据冗余问题。本文提出了一种基于视频的跨模态辅助网络(VCAN),该网络由音频特征映射模块和跨模态选择模块组成。第一模块被设计为实质上增加音频特征提取中的特征多样性,旨在通过提供更全面的声学表示来提高分类准确度。为了使模型能够处理冗余视觉特征,第二模块被寻址为在集成视听数据期间有效地过滤冗余视觉帧。此外,引入了由多个图像分类网络组成的分类器组来预测情感极性和情感类别。在RAVDESS、CMU-MOSI和CMU-MOSEI基准测试上的大量实验结果表明,VCAN在提高多模态情感分析的分类准确率方面明显优于现有的方法.摘要:Multimodal sentiment analysis has a wide range of applications due to its information complementarity in multimodal interactions. Previous works focus more on investigating efficient joint representations, but they rarely consider the insufficient unimodal features extraction and data redundancy of multimodal fusion. In this paper, a Video-based Cross-modal Auxiliary Network (VCAN) is proposed, which is comprised of an audio features map module and a cross-modal selection module. The first module is designed to substantially increase feature diversity in audio feature extraction, aiming to improve classification accuracy by providing more comprehensive acoustic representations. To empower the model to handle redundant visual features, the second module is addressed to efficiently filter the redundant visual frames during integrating audiovisual data. Moreover, a classifier group consisting of several image classification networks is introduced to predict sentiment polarities and emotion categories. Extensive experimental results on RAVDESS, CMU-MOSI, and CMU-MOSEI benchmarks indicate that VCAN is significantly superior to the state-of-the-art methods for improving the classification accuracy of multimodal sentiment analysis.
机器翻译,仅供参考