今天跟大家分享一篇语音相关的论文合集:cs.SD语音10篇,eess.AS音频处理13篇。
cs.SD语音
【1】 AnimeTAB: A new guitar tablature dataset of anime and game music
标题:AnimeTAB:一个新的动漫和游戏音乐吉他表数据集
链接:https://arxiv.org/abs/2210.03027
作者:Yuecheng Zhou,Yaolong Ju,Lingyun Xie机构:Communication University of China, Mcgill University摘要:虽然吉他谱已经成为MIR研究中的一个热门话题,但还没有这样一个专注于动画和视频游戏配乐的吉他谱数据集,这些音乐在年轻人中有着令人惊讶的广泛和不断增长的观众。本文提出了一个MusicXML格式的指法吉他谱数据集AnimeTAB,为研究者和吉他演奏者提供了更高质量的吉他谱。AnimeTAB包含412个完整音轨和547个片段,后者用音乐结构(前奏、独唱、合唱和过渡)注释。附带的分析工具包TABprocessor可进一步方便其使用。这包括旋律和低音线提取、调检测和和弦标记等功能,这些功能都是使用基于规则的算法实现的。我们根据手动注释的真实数据评估了这些功能中的每一个。最后,作为一个实例,我们利用TABprocessor对AnimeTAB进行了音乐和技术分析。我们的数据和代码已公开提供给作曲家、表演者和音乐信息检索(MIR)研究人员等。摘要:While guitar tablature has become a popular topic in MIR research, there exists no such a guitar tablature dataset that focuses on the soundtracks of anime and video games, which have a surprisingly broad and growing audience among the youths. In this paper, we present AnimeTAB, a fingerstyle guitar tablature dataset in MusicXML format, which provides more high-quality guitar tablature for both researchers and guitar players. AnimeTAB contains 412 full tracks and 547 clips, the latter are annotated with musical structures (intro, verse, chorus, and bridge). An accompanying analysis toolkit, TABprocessor, is included to further facilitate its use. This includes functions for melody and bassline extraction, key detection, and chord labeling, which are implemented using rule-based algorithms. We evaluated each of these functions against a manually annotated ground truth. Finally, as an example, we performed a music and technique analysis of AnimeTAB using TABprocessor. Our data and code have been made publicly available for composers, performers, and music information retrieval (MIR) researchers alike.
【2】 WakeUpNet: A Mobile-Transformer based Framework for End-to-End Streaming Voice Trigger
标题:WakeUpNet:一种基于移动Transformer的端到端流语音触发框架
链接:https://arxiv.org/abs/2210.02904
作者:Zixing Zhang,Thorin Farnsworth,Senling Lin,Salah Karout机构:Huawei Technologies Research & Development (UK) Ltd摘要:端到端模型逐渐成为语音触发的主流技术,其目标是以较小的占用空间实现最大的预测准确度。本文提出了一种基于Transformer编码器的端到端语音触发框架WakeupNet。该框架的目的是探索Transformer的上下文捕获能力,因为顺序信息对于唤醒字检测至关重要。然而,传统的Transformer编码器太大,不适合我们的任务。为了解决这个问题,我们引入了不同的模型压缩方法,将普通模型压缩成一个微小的模型,称为mobile-Transformer。为了评估mobile-Transformer的性能,我们在一个大型公共数据集HiMia上进行了大量的实验.实验结果表明,无论是在干净环境下还是在有噪声环境下,所提出的mobile-Transformer模型都明显优于其他常用的语音触发模型。摘要:End-to-end models have gradually become the main technical stream for voice trigger, aiming to achieve an utmost prediction accuracy but with a small footprint. In present paper, we propose an end-to-end voice trigger framework, namely WakeupNet, which is basically structured on a Transformer encoder. The purpose of this framework is to explore the context-capturing capability of Transformer, as sequential information is vital for wakeup-word detection. However, the conventional Transformer encoder is too large to fit our task. To address this issue, we introduce different model compression approaches to shrink the vanilla one into a tiny one, called mobile-Transformer. To evaluate the performance of mobile-Transformer, we conduct extensive experiments on a large public-available dataset HiMia. The obtained results indicate that introduced mobile-Transformer significantly outperforms other frequently used models for voice trigger in both clean and noisy scenarios.
【3】 Matching Text and Audio Embeddings: Exploring Transfer-learning Strategies for Language-based Audio Retrieval
标题:匹配文本和音频嵌入:探索基于语言的音频检索的迁移学习策略
链接:https://arxiv.org/abs/2210.02833
作者:Benno Weck,Miguel Pérez Fernández,Holger Kirchhoff,Xavier Serra机构:Huawei Technologies, Munich Research Center, Germany, Universitat Pompeu Fabra, Music Technology Group, Spain备注:5 pages, 2 figures. Accepted at Detection and Classification of Acoustic Scenes and Events 2022 (DCASE2022)摘要:我们对用于跨模态(文本到音频)检索的大规模预训练深度学习模型进行了分析。我们在度量学习框架中使用由这些模型提取的嵌入来连接音频和文本的匹配对。浅层神经网络将嵌入映射到公共维度。我们的系统是DCASE Challenge 2022的基于语言的音频检索任务的一个扩展,采用了RoBERTa基础模型作为文本嵌入提取器。预训练的PANN模型提取音频嵌入。为了提高模型的泛化能力,我们研究了如何使用从在线平台Freesound收集的音频和相关噪声文本进行预训练来提高我们方法的性能。此外,我们的消融研究表明,损失函数的正确选择和预训练模型的微调在训练竞争性检索系统中是必不可少的。摘要:We present an analysis of large-scale pretrained deep learning models used for cross-modal (text-to-audio) retrieval. We use embeddings extracted by these models in a metric learning framework to connect matching pairs of audio and text. Shallow neural networks map the embeddings to a common dimensionality. Our system, which is an extension of our submission to the Language-based Audio Retrieval Task of the DCASE Challenge 2022, employs the RoBERTa foundation model as the text embedding extractor. A pretrained PANNs model extracts the audio embeddings. To improve the generalisation of our model, we investigate how pretraining with audio and associated noisy text collected from the online platform Freesound improves the performance of our method. Furthermore, our ablation study reveals that the proper choice of the loss function and fine-tuning the pretrained models are essential in training a competitive retrieval system.
【4】 Melody Infilling with User-Provided Structural Context
标题:用用户提供的结构上下文填充旋律
链接:https://arxiv.org/abs/2210.02829
作者:Chih-Pin Tan,Alvin W. Y. Su,Yi-Hsuan Yang机构:Alvin W.Y. Su, National Cheng Kung University, Academia Sinica, Taiwan AI Labs摘要:本文提出了一种基于Transformer的乐谱填充模型,用于生成一个能够填充过去和未来背景之间空白的音乐片段。虽然现有的填充方法可以生成与给定上下文局部平滑连接的段落,但是它们没有考虑音乐的音乐形式或结构,并且因此可能生成过度平滑的结果。针对这一问题,本文提出了一种结构感知条件化方法,该方法采用一种新颖的注意力选择模块,将用户提供的结构相关信息提供给Transformer进行填充。主观和客观评价结果表明,与现有的两种结构不可知填充模型相比,该模型能够有效地利用结构信息,生成质量更高的流行风格旋律。摘要:This paper proposes a novel Transformer-based model for music score infilling, to generate a music passage that fills in the gap between given past and future contexts. While existing infilling approaches can generate a passage that connects smoothly locally with the given contexts, they do not take into account the musical form or structure of the music and may therefore generate overly smooth results. To address this issue, we propose a structure-aware conditioning approach that employs a novel attention-selecting module to supply user-provided structure-related information to the Transformer for infilling. With both objective and subjective evaluations, we show that the proposed model can harness the structural information effectively and generate melodies in the style of pop of higher quality than the two existing structure-agnostic infilling models.
【5】 The Sound of Silence: Efficiency of First Digit Features in Synthetic Audio Detection
标题:沉默之声:合成音频检测中第一个数字特征的效率
链接:https://arxiv.org/abs/2210.02746
作者:Daniele Mari,Federica Latora,Simone Milani机构:Department of Information Engineering, University of Padova, Padova, Italy摘要:近年来,生成神经策略和音频处理技术的结合促进了合成语音合成或变换算法的广泛应用。事实证明,这种能力在许多法律和信息流程(新闻、生物特征认证、法庭上的音频证据等)中是有害的。因此,由于伪造技术的异质性,开发有效的检测算法既至关重要又具有挑战性。 本文研究了合成语音检测中静音部分的区分作用,并展示了从MFCC系数中提取的第一位统计量如何有效地实现鲁棒检测。该方法不依赖于大规模的神经网络检测结构,计算量小,在多种算法上都是有效的,并且在ASVSpoof数据集的大多数类中获得了90%以上的准确率.摘要:The recent integration of generative neural strategies and audio processing techniques have fostered the widespread of synthetic speech synthesis or transformation algorithms. This capability proves to be harmful in many legal and informative processes (news, biometric authentication, audio evidence in courts, etc.). Thus, the development of efficient detection algorithms is both crucial and challenging due to the heterogeneity of forgery techniques. This work investigates the discriminative role of silenced parts in synthetic speech detection and shows how first digit statistics extracted from MFCC coefficients can efficiently enable a robust detection. The proposed procedure is computationally-lightweight and effective on many different algorithms since it does not rely on large neural detection architecture and obtains an accuracy above 90\% in most of the classes of the ASVSpoof dataset.
【6】 PSVRF: Learning to restore Pitch-Shifted Voice without reference
标题:PSVRF:学习在没有参考的情况下恢复音调变化的声音
链接:https://arxiv.org/abs/2210.02731
作者:Yangfu Li,Xiaodan Lin,Jiaxin Yang机构:School of Information Science and Engineering, Huaqiao University, Xiamen, China摘要:基音缩放算法对自动说话人确认(ASV)系统的安全性有着重要影响。虽然已经提出了许多反电子欺骗算法来识别基音偏移语音,甚至将其恢复为原始版本,但这些算法要么性能不佳,要么需要原始语音作为参考,限制了应用前景。本文提出了一种无参考的高质量基音偏移语音恢复方法PSVRF$^1$。在AISHELL-1和AISHELL-3上的实验结果表明,PSVRF能够恢复出各种基音缩放技术伪装的语音,显著提高了ASV系统对基音缩放攻击的鲁棒性。此外,PSVRF的性能甚至超过了最先进的基于参考的方法。摘要:Pitch scaling algorithms have a significant impact on the security of Automatic Speaker Verification (ASV) systems. Although numerous anti-spoofing algorithms have been proposed to identify the pitch-shifted voice and even restore it to the original version, they either have poor performance or require the original voice as a reference, limiting the prospects of applications. In this paper, we propose a no-reference approach termed PSVRF$^1$ for high-quality restoration of pitch-shifted voice. Experiments on AISHELL-1 and AISHELL-3 demonstrate that PSVRF can restore the voice disguised by various pitch-scaling techniques, which obviously enhances the robustness of ASV systems to pitch-scaling attacks. Furthermore, the performance of PSVRF even surpasses that of the state-of-the-art reference-based approach.
【7】 Feasibility on Detecting Door Slamming towards Monitoring Early Signs of Domestic Violence
标题:检测摔门行为监测家庭暴力早期征兆的可行性
链接:https://arxiv.org/abs/2210.02642
作者:Osian Morgan,Hakan Kayan,Charith Perera机构:School of Computer Science and, Informatics, Cardiff University, Cardiff, UK备注:In Proceedings of the 2022 IEEE/ACM Seventh International Conference on Internet-of-Things Design and Implementation (IoTDI) 2022摘要:通过使用低成本的微控制器和TinyML,我们研究了检测家庭暴力和其他反社会行为的潜在早期预警信号的可行性。我们创建了一个机器学习模型,通过分析音频数据并将其输入卷积神经网络来对样本进行分类,以确定门是否被激进地关闭。在试验条件下,无背景噪声时,准确率为88.89%,当混合相对体积为样品体积的0.5倍的各种背景噪声时,准确率降至87.50%。然后将该模型部署在连接到门的Arduino Nano BLE 33 Sense上,并且仅在检测到大于预定义阈值加速度的加速度时才开始采样。然后,模型所做的预测可以通过BLE发送到另一个设备,例如Raspberry Pi的智能手机。摘要:By using low-cost microcontrollers and TinyML, we investigate the feasibility of detecting potential early warning signs of domestic violence and other anti-social behaviors within the home. We created a machine learning model to determine if a door was closed aggressively by analyzing audio data and feeding this into a convolutional neural network to classify the sample. Under test conditions, with no background noise, accuracy of 88.89\% was achieved, declining to 87.50\% when assorted background noises were mixed in at a relative volume of 0.5 times that of the sample. The model is then deployed on an Arduino Nano BLE 33 Sense attached to the door, and only begins sampling once an acceleration greater than a predefined threshold acceleration is detected. The predictions made by the model can then be sent via BLE to another device, such as a smartphone of Raspberry Pi.
【8】 JoeyS2T: Minimalistic Speech-to-Text Modeling with JoeyNMT
标题:JoeyS2T:基于JoeyNMT的简约语音到文本建模
链接:https://arxiv.org/abs/2210.02545
作者:Mayumi Ohta,Julia Kreutzer,Stefan Riezler机构:Heidelberg University, Germany, Google Research, Computational Linguistics & IWR摘要:JoeyS2T是JoeyNMT的一个扩展,用于语音到文本任务,如自动语音识别和端到端语音翻译。它继承了JoeyNMT的核心理念,JoeyNMT是一个基于PyTorch的极简主义NMT工具包,追求简单性和可访问性。JoeyS2T的工作流程是独立的,从数据预处理开始,经过模型训练和预测到评估,并无缝集成到JoeyNMT的紧凑和简单的代码库中。在JoeyNMT最先进的基于Transformer的编解码器架构之上,JoeyS2T提供了面向语音的组件,如卷积层、SpecAugment、CTC丢失和WER评估。尽管JoeyS2T与以前的实现相比很简单,但它在英语语音识别和英语到德语语音翻译基准测试中的表现很有竞争力。该实施附带了一个演练教程,可在www.example.com上获得https://github.com/may-/joeys2t。摘要:JoeyS2T is a JoeyNMT extension for speech-to-text tasks such as automatic speech recognition and end-to-end speech translation. It inherits the core philosophy of JoeyNMT, a minimalist NMT toolkit built on PyTorch, seeking simplicity and accessibility. JoeyS2T's workflow is self-contained, starting from data pre-processing, over model training and prediction to evaluation, and is seamlessly integrated into JoeyNMT's compact and simple code base. On top of JoeyNMT's state-of-the-art Transformer-based encoder-decoder architecture, JoeyS2T provides speech-oriented components such as convolutional layers, SpecAugment, CTC-loss, and WER evaluation. Despite its simplicity compared to prior implementations, JoeyS2T performs competitively on English speech recognition and English-to-German speech translation benchmarks. The implementation is accompanied by a walk-through tutorial and available on https://github.com/may-/joeys2t.
【9】 ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
标题:ASVspoof 2021:走向野外的欺骗和深度假冒语音检测
链接:https://arxiv.org/abs/2210.02437
作者:Xuechen Liu,Xin Wang,Md Sahidullah,Jose Patino,Héctor Delgado,Tomi Kinnunen,Massimiliano Todisco,Junichi Yamagishi,Nicholas Evans,Andreas Nautsch,Kong Aik Lee备注:Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing摘要:基准测试倡议支持对语音和语言处理中突出问题的竞争解决方案进行有意义的比较。连续的基准评估通常反映了从理想的实验室条件到野外条件的逐步演变。ASVspoof,欺骗和deepfake检测倡议和挑战系列,也遵循了同样的趋势。本文提供了ASVspoof 2021挑战赛的总结和37支参赛队伍的结果。对于逻辑访问任务,结果表明对抗方案对新引入的编码和传输效应是鲁棒的。物理访问任务的结果表明,与模拟的物理空间相反,在真实环境中检测重放攻击的潜力,但是对模拟和真实声学环境之间的变化缺乏鲁棒性。DF任务是2021年版的新任务,旨在解决在线发布的被操纵、压缩的语音数据的检测问题。虽然检测解决方案提供了对压缩效应的一些弹性,但它们缺乏跨不同源数据集的通用性。除了总结每项任务的最佳执行系统、对影响数据因素的新分析和隐藏数据子集的结果之外,本文还包括对挑战后结果的回顾、主要挑战限制的概述和ASVspoof未来的路线图。链接到ASVspoof挑战和相关资源:https://www.asvspoof.org/index2021.html摘要:Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 37 participating teams. For the logical access task, results indicate that countermeasures solutions are robust to newly introduced encoding and transmission effects. Results for the physical access task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The DF task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof. Link to the ASVspoof challenge and related resources: https://www.asvspoof.org/index2021.html
【10】 TC-SKNet with GridMask for Low-complexity Classification of Acoustic scene
标题:基于网格掩码的TC-SKNet低复杂度声场景分类
链接:https://arxiv.org/abs/2210.02287
作者:Luyuan Xie,Yan Zhong,Lin Yang,Zhaoyu Yan,Zhonghai Wu,Junjie Wang机构:School of Software and Microelectronics, Peking University, Beijing, China备注:Accepted to APSIPA ASC 2022摘要:卷积神经网络(CNNs)在低复杂度分类任务(如声学场景分类(ASCs))中具有良好的性能。然而,关于目标语音长度与卷积核大小之间关系的研究却很少。本文将选择性核网络与时域卷积相结合(TC-SKNet),通过调节卷积核的感受野,在保持低复杂度的同时,解决了目标语音长度可变的问题。GridMask是一种通过屏蔽部分原始数据或特征区域的数据扩充策略。由于辍学的作用,提高了模型的泛化能力。在我们的实验中,GridMask带来的性能增益比ASC中的频谱扩展更强。最后,采用AutoML搜索TC-SKNet的最佳结构和GridMask的超参数,以提高分类性能。结果,59.87% TC-SKNet的峰值准确度与SOTA的峰值准确度相当,但参数仅使用20.9K。摘要:Convolution neural networks (CNNs) have good performance in low-complexity classification tasks such as acoustic scene classifications (ASCs). However, there are few studies on the relationship between the length of target speech and the size of the convolution kernels. In this paper, we combine Selective Kernel Network with Temporal-Convolution (TC-SKNet) to adjust the receptive field of convolution kernels to solve the problem of variable length of target voice while keeping low-complexity. GridMask is a data augmentation strategy by masking part of the raw data or feature area. It can enhance the generalization of the model as the role of dropout. In our experiments, the performance gain brought by GridMask is stronger than spectrum augmentation in ASCs. Finally, we adopt AutoML to search best structure of TC-SKNet and hyperparameters of GridMask for improving the classification performance. As a result, a peak accuracy of 59.87% TC-SKNet is equivalent to that of SOTA, but the parameters only use 20.9 K.
【1】 Fully Unsupervised Training of Few-shot Keyword Spotting
标题:Few-Shot关键词识别的全无监督训练
链接:https://arxiv.org/abs/2210.02732
作者:Dongjune Lee,Minchan Kim,Sung Hwan Mun,Min Hyun Han,Nam Soo Kim机构:Department of Electrical and Computer Engineering and INMC, Seoul National University, Seoul, South Korea备注:Accepted by IEEE SLT 2022摘要:为了训练一个少镜头关键词识别(FS-KWS)模型,需要一个包含大量目标关键词的大规模标记数据集,以便在仅有少量注册样本的情况下推广到任意目标关键词.为了减少标记所带来的数据收集开销,提出了一种新的FS-KWS系统,该系统仅在合成数据上训练.所提出的系统基于度量学习,使得能够使用距离度量来检测目标关键字。利用语音合成模型,用伪音素代替文本生成语音,我们很容易获得大量具有相同语义的多视图样本。考虑到度量学习本质上不需要标记数据,这些样本足以用于训练。我们框架中的所有组件都不需要任何监督,这使得我们的方法不需要监督。在真实数据集上的实验结果表明,该方法在没有真实数据集和标记数据集的情况下仍具有较好的性能.摘要:For training a few-shot keyword spotting (FS-KWS) model, a large labeled dataset containing massive target keywords has known to be essential to generalize to arbitrary target keywords with only a few enrollment samples. To alleviate the expensive data collection with labeling, in this paper, we propose a novel FS-KWS system trained only on synthetic data. The proposed system is based on metric learning enabling target keywords to be detected using distance metrics. Exploiting the speech synthesis model that generates speech with pseudo phonemes instead of texts, we easily obtain a large collection of multi-view samples with the same semantics. These samples are sufficient for training, considering metric learning does not intrinsically necessitate labeled data. All of the components in our framework do not require any supervision, making our method unsupervised. Experimental results on real datasets show our proposed method is competitive even without any labeled and real datasets.
【2】 Exploration of A Self-Supervised Speech Model: A Study on Emotional Corpora
标题:基于情感语料库的自监督语音模型研究
链接:https://arxiv.org/abs/2210.02595
作者:Yuanchao Li,Yumnah Mohamied,Peter Bell,Catherine Lai机构:Centre for Speech Technology Research, University of Edinburgh摘要:在过去的几年中,自监督语音模型已经快速发展,并且已经证明在各种下游任务中使用是可行的。最近的一些工作已经开始研究这些模型的特征,但许多问题还没有得到充分解决。本文以情感语料库为研究对象,探讨了一种流行的自我监督模型--wav 2 vec 2.0。通过一系列的定量分析,我们主要论证了:1)wav2vec2.0似乎丢弃了对于单词识别目的不太有用的副语言信息; 2)对于情感识别,单独来自中间层的表示的性能与通过层平均得到的表示的性能一样好,而最终层在某些情况下导致最差的性能; 3)当前的自监督模型可能不是利用非词汇特征的下游任务的最佳解决方案。我们的工作提供了新的发现,将有助于未来的研究在这一领域和理论基础的使用现有的模型。摘要:Self-supervised speech models have grown fast during the past few years and have proven feasible for use in various downstream tasks. Some recent work has started to look at the characteristics of these models, yet many concerns have not been fully addressed. In this work, we conduct a study on emotional corpora to explore a popular self-supervised model -- wav2vec 2.0. Via a set of quantitative analysis, we mainly demonstrate that: 1) wav2vec 2.0 appears to discard paralinguistic information that is less useful for word recognition purposes; 2) for emotion recognition, representations from the middle layer alone perform as well as those derived from layer averaging, while the final layer results in the worst performance in some cases; 3) current self-supervised models may not be the optimal solution for downstream tasks that make use of non-lexical features. Our work provides novel findings that will aid future research in this area and theoretical basis for the use of existing models.
【3】 AnimeTAB: A new guitar tablature dataset of anime and game music
标题:AnimeTAB:一个新的动漫和游戏音乐吉他表数据集
链接:https://arxiv.org/abs/2210.03027
* 与cs.SD语音【1】为同一篇
作者:Yuecheng Zhou,Yaolong Ju,Lingyun Xie机构:Communication University of China, Mcgill University摘要:虽然吉他谱已经成为MIR研究中的一个热门话题,但还没有这样一个专注于动画和视频游戏配乐的吉他谱数据集,这些音乐在年轻人中有着令人惊讶的广泛和不断增长的观众。本文提出了一个MusicXML格式的指法吉他谱数据集AnimeTAB,为研究者和吉他演奏者提供了更高质量的吉他谱。AnimeTAB包含412个完整音轨和547个片段,后者用音乐结构(前奏、独唱、合唱和过渡)注释。附带的分析工具包TABprocessor可进一步方便其使用。这包括旋律和低音线提取、调检测和和弦标记等功能,这些功能都是使用基于规则的算法实现的。我们根据手动注释的真实数据评估了这些功能中的每一个。最后,作为一个实例,我们利用TABprocessor对AnimeTAB进行了音乐和技术分析。我们的数据和代码已公开提供给作曲家、表演者和音乐信息检索(MIR)研究人员等。摘要:While guitar tablature has become a popular topic in MIR research, there exists no such a guitar tablature dataset that focuses on the soundtracks of anime and video games, which have a surprisingly broad and growing audience among the youths. In this paper, we present AnimeTAB, a fingerstyle guitar tablature dataset in MusicXML format, which provides more high-quality guitar tablature for both researchers and guitar players. AnimeTAB contains 412 full tracks and 547 clips, the latter are annotated with musical structures (intro, verse, chorus, and bridge). An accompanying analysis toolkit, TABprocessor, is included to further facilitate its use. This includes functions for melody and bassline extraction, key detection, and chord labeling, which are implemented using rule-based algorithms. We evaluated each of these functions against a manually annotated ground truth. Finally, as an example, we performed a music and technique analysis of AnimeTAB using TABprocessor. Our data and code have been made publicly available for composers, performers, and music information retrieval (MIR) researchers alike.
【4】 WakeUpNet: A Mobile-Transformer based Framework for End-to-End Streaming Voice Trigger
标题:WakeUpNet:一种基于移动Transformer的端到端流语音触发框架
链接:https://arxiv.org/abs/2210.02904
* 与cs.SD语音【2】为同一篇
作者:Zixing Zhang,Thorin Farnsworth,Senling Lin,Salah Karout机构:Huawei Technologies Research & Development (UK) Ltd摘要:端到端模型逐渐成为语音触发的主流技术,其目标是以较小的占用空间实现最大的预测准确度。本文提出了一种基于Transformer编码器的端到端语音触发框架WakeupNet。该框架的目的是探索Transformer的上下文捕获能力,因为顺序信息对于唤醒字检测至关重要。然而,传统的Transformer编码器太大,不适合我们的任务。为了解决这个问题,我们引入了不同的模型压缩方法,将普通模型压缩成一个微小的模型,称为mobile-Transformer。为了评估mobile-Transformer的性能,我们在一个大型公共数据集HiMia上进行了大量的实验.实验结果表明,无论是在干净环境下还是在有噪声环境下,所提出的mobile-Transformer模型都明显优于其他常用的语音触发模型。摘要:End-to-end models have gradually become the main technical stream for voice trigger, aiming to achieve an utmost prediction accuracy but with a small footprint. In present paper, we propose an end-to-end voice trigger framework, namely WakeupNet, which is basically structured on a Transformer encoder. The purpose of this framework is to explore the context-capturing capability of Transformer, as sequential information is vital for wakeup-word detection. However, the conventional Transformer encoder is too large to fit our task. To address this issue, we introduce different model compression approaches to shrink the vanilla one into a tiny one, called mobile-Transformer. To evaluate the performance of mobile-Transformer, we conduct extensive experiments on a large public-available dataset HiMia. The obtained results indicate that introduced mobile-Transformer significantly outperforms other frequently used models for voice trigger in both clean and noisy scenarios.
【5】 Matching Text and Audio Embeddings: Exploring Transfer-learning Strategies for Language-based Audio Retrieval
标题:匹配文本和音频嵌入:探索基于语言的音频检索的迁移学习策略
链接:https://arxiv.org/abs/2210.02833
* 与cs.SD语音【3】为同一篇
作者:Benno Weck,Miguel Pérez Fernández,Holger Kirchhoff,Xavier Serra机构:Huawei Technologies, Munich Research Center, Germany, Universitat Pompeu Fabra, Music Technology Group, Spain备注:5 pages, 2 figures. Accepted at Detection and Classification of Acoustic Scenes and Events 2022 (DCASE2022)摘要:我们对用于跨模态(文本到音频)检索的大规模预训练深度学习模型进行了分析。我们在度量学习框架中使用由这些模型提取的嵌入来连接音频和文本的匹配对。浅层神经网络将嵌入映射到公共维度。我们的系统是DCASE Challenge 2022的基于语言的音频检索任务的一个扩展,采用了RoBERTa基础模型作为文本嵌入提取器。预训练的PANN模型提取音频嵌入。为了提高模型的泛化能力,我们研究了如何使用从在线平台Freesound收集的音频和相关噪声文本进行预训练来提高我们方法的性能。此外,我们的消融研究表明,损失函数的正确选择和预训练模型的微调在训练竞争性检索系统中是必不可少的。摘要:We present an analysis of large-scale pretrained deep learning models used for cross-modal (text-to-audio) retrieval. We use embeddings extracted by these models in a metric learning framework to connect matching pairs of audio and text. Shallow neural networks map the embeddings to a common dimensionality. Our system, which is an extension of our submission to the Language-based Audio Retrieval Task of the DCASE Challenge 2022, employs the RoBERTa foundation model as the text embedding extractor. A pretrained PANNs model extracts the audio embeddings. To improve the generalisation of our model, we investigate how pretraining with audio and associated noisy text collected from the online platform Freesound improves the performance of our method. Furthermore, our ablation study reveals that the proper choice of the loss function and fine-tuning the pretrained models are essential in training a competitive retrieval system.
【6】 Melody Infilling with User-Provided Structural Context
标题:用用户提供的结构上下文填充旋律
链接:https://arxiv.org/abs/2210.02829
* 与cs.SD语音【4】为同一篇
作者:Chih-Pin Tan,Alvin W. Y. Su,Yi-Hsuan Yang机构:Alvin W.Y. Su, National Cheng Kung University, Academia Sinica, Taiwan AI Labs摘要:本文提出了一种基于Transformer的乐谱填充模型,用于生成一个能够填充过去和未来背景之间空白的音乐片段。虽然现有的填充方法可以生成与给定上下文局部平滑连接的段落,但是它们没有考虑音乐的音乐形式或结构,并且因此可能生成过度平滑的结果。针对这一问题,本文提出了一种结构感知条件化方法,该方法采用一种新颖的注意力选择模块,将用户提供的结构相关信息提供给Transformer进行填充。主观和客观评价结果表明,与现有的两种结构不可知填充模型相比,该模型能够有效地利用结构信息,生成质量更高的流行风格旋律。摘要:This paper proposes a novel Transformer-based model for music score infilling, to generate a music passage that fills in the gap between given past and future contexts. While existing infilling approaches can generate a passage that connects smoothly locally with the given contexts, they do not take into account the musical form or structure of the music and may therefore generate overly smooth results. To address this issue, we propose a structure-aware conditioning approach that employs a novel attention-selecting module to supply user-provided structure-related information to the Transformer for infilling. With both objective and subjective evaluations, we show that the proposed model can harness the structural information effectively and generate melodies in the style of pop of higher quality than the two existing structure-agnostic infilling models.
【7】 The Sound of Silence: Efficiency of First Digit Features in Synthetic Audio Detection
标题:沉默之声:合成音频检测中第一个数字特征的效率
链接:https://arxiv.org/abs/2210.02746
* 与cs.SD语音【5】为同一篇
作者:Daniele Mari,Federica Latora,Simone Milani机构:Department of Information Engineering, University of Padova, Padova, Italy摘要:近年来,生成神经策略和音频处理技术的结合促进了合成语音合成或变换算法的广泛应用。事实证明,这种能力在许多法律和信息流程(新闻、生物特征认证、法庭上的音频证据等)中是有害的。因此,由于伪造技术的异质性,开发有效的检测算法既至关重要又具有挑战性。 本文研究了合成语音检测中静音部分的区分作用,并展示了从MFCC系数中提取的第一位统计量如何有效地实现鲁棒检测。该方法不依赖于大规模的神经网络检测结构,计算量小,在多种算法上都是有效的,并且在ASVSpoof数据集的大多数类中获得了90%以上的准确率.摘要:The recent integration of generative neural strategies and audio processing techniques have fostered the widespread of synthetic speech synthesis or transformation algorithms. This capability proves to be harmful in many legal and informative processes (news, biometric authentication, audio evidence in courts, etc.). Thus, the development of efficient detection algorithms is both crucial and challenging due to the heterogeneity of forgery techniques. This work investigates the discriminative role of silenced parts in synthetic speech detection and shows how first digit statistics extracted from MFCC coefficients can efficiently enable a robust detection. The proposed procedure is computationally-lightweight and effective on many different algorithms since it does not rely on large neural detection architecture and obtains an accuracy above 90\% in most of the classes of the ASVSpoof dataset.
【8】 PSVRF: Learning to restore Pitch-Shifted Voice without reference
标题:PSVRF:学习在没有参考的情况下恢复音调变化的声音
链接:https://arxiv.org/abs/2210.02731
* 与cs.SD语音【6】为同一篇
作者:Yangfu Li,Xiaodan Lin,Jiaxin Yang机构:School of Information Science and Engineering, Huaqiao University, Xiamen, China摘要:基音缩放算法对自动说话人确认(ASV)系统的安全性有着重要影响。虽然已经提出了许多反电子欺骗算法来识别基音偏移语音,甚至将其恢复为原始版本,但这些算法要么性能不佳,要么需要原始语音作为参考,限制了应用前景。本文提出了一种无参考的高质量基音偏移语音恢复方法PSVRF$^1$。在AISHELL-1和AISHELL-3上的实验结果表明,PSVRF能够恢复出各种基音缩放技术伪装的语音,显著提高了ASV系统对基音缩放攻击的鲁棒性。此外,PSVRF的性能甚至超过了最先进的基于参考的方法。摘要:Pitch scaling algorithms have a significant impact on the security of Automatic Speaker Verification (ASV) systems. Although numerous anti-spoofing algorithms have been proposed to identify the pitch-shifted voice and even restore it to the original version, they either have poor performance or require the original voice as a reference, limiting the prospects of applications. In this paper, we propose a no-reference approach termed PSVRF$^1$ for high-quality restoration of pitch-shifted voice. Experiments on AISHELL-1 and AISHELL-3 demonstrate that PSVRF can restore the voice disguised by various pitch-scaling techniques, which obviously enhances the robustness of ASV systems to pitch-scaling attacks. Furthermore, the performance of PSVRF even surpasses that of the state-of-the-art reference-based approach.
【9】 Feasibility on Detecting Door Slamming towards Monitoring Early Signs of Domestic Violence
标题:检测摔门行为监测家庭暴力早期征兆的可行性
链接:https://arxiv.org/abs/2210.02642
* 与cs.SD语音【7】为同一篇
作者:Osian Morgan,Hakan Kayan,Charith Perera机构:School of Computer Science and, Informatics, Cardiff University, Cardiff, UK备注:In Proceedings of the 2022 IEEE/ACM Seventh International Conference on Internet-of-Things Design and Implementation (IoTDI) 2022摘要:通过使用低成本的微控制器和TinyML,我们研究了检测家庭暴力和其他反社会行为的潜在早期预警信号的可行性。我们创建了一个机器学习模型,通过分析音频数据并将其输入卷积神经网络来对样本进行分类,以确定门是否被激进地关闭。在试验条件下,无背景噪声时,准确率为88.89%,当混合相对体积为样品体积的0.5倍的各种背景噪声时,准确率降至87.50%。然后将该模型部署在连接到门的Arduino Nano BLE 33 Sense上,并且仅在检测到大于预定义阈值加速度的加速度时才开始采样。然后,模型所做的预测可以通过BLE发送到另一个设备,例如Raspberry Pi的智能手机。摘要:By using low-cost microcontrollers and TinyML, we investigate the feasibility of detecting potential early warning signs of domestic violence and other anti-social behaviors within the home. We created a machine learning model to determine if a door was closed aggressively by analyzing audio data and feeding this into a convolutional neural network to classify the sample. Under test conditions, with no background noise, accuracy of 88.89\% was achieved, declining to 87.50\% when assorted background noises were mixed in at a relative volume of 0.5 times that of the sample. The model is then deployed on an Arduino Nano BLE 33 Sense attached to the door, and only begins sampling once an acceleration greater than a predefined threshold acceleration is detected. The predictions made by the model can then be sent via BLE to another device, such as a smartphone of Raspberry Pi.
【10】 JoeyS2T: Minimalistic Speech-to-Text Modeling with JoeyNMT
标题:JoeyS2T:基于JoeyNMT的简约语音到文本建模
链接:https://arxiv.org/abs/2210.02545
* 与cs.SD语音【8】为同一篇
作者:Mayumi Ohta,Julia Kreutzer,Stefan Riezler机构:Heidelberg University, Germany, Google Research, Computational Linguistics & IWR摘要:JoeyS2T是JoeyNMT的一个扩展,用于语音到文本任务,如自动语音识别和端到端语音翻译。它继承了JoeyNMT的核心理念,JoeyNMT是一个基于PyTorch的极简主义NMT工具包,追求简单性和可访问性。JoeyS2T的工作流程是独立的,从数据预处理开始,经过模型训练和预测到评估,并无缝集成到JoeyNMT的紧凑和简单的代码库中。在JoeyNMT最先进的基于Transformer的编解码器架构之上,JoeyS2T提供了面向语音的组件,如卷积层、SpecAugment、CTC丢失和WER评估。尽管JoeyS2T与以前的实现相比很简单,但它在英语语音识别和英语到德语语音翻译基准测试中的表现很有竞争力。该实施附带了一个演练教程,可在www.example.com上获得https://github.com/may-/joeys2t。摘要:JoeyS2T is a JoeyNMT extension for speech-to-text tasks such as automatic speech recognition and end-to-end speech translation. It inherits the core philosophy of JoeyNMT, a minimalist NMT toolkit built on PyTorch, seeking simplicity and accessibility. JoeyS2T's workflow is self-contained, starting from data pre-processing, over model training and prediction to evaluation, and is seamlessly integrated into JoeyNMT's compact and simple code base. On top of JoeyNMT's state-of-the-art Transformer-based encoder-decoder architecture, JoeyS2T provides speech-oriented components such as convolutional layers, SpecAugment, CTC-loss, and WER evaluation. Despite its simplicity compared to prior implementations, JoeyS2T performs competitively on English speech recognition and English-to-German speech translation benchmarks. The implementation is accompanied by a walk-through tutorial and available on https://github.com/may-/joeys2t.
【11】 Evaluation of Automatic Single Cough Segmentation
标题:单次咳嗽自动分割算法的评价
链接:https://arxiv.org/abs/2210.02057
作者:Bagus Tris Atmaja,Zanjabila,Suyanto,Akira Sasou机构:National Institute of Advanced Industrial Science and Technology, Japan, Institut Teknologi Sepuluh Nopember, Indonesia备注:3 figure,s 3 tables, submitted to IJST摘要:目前,基于语音信号诊断疾病的研究正在迅速增加,包括咳嗽相关疾病。当将咳嗽声信号训练成深度学习模型时,需要通过将若干咳嗽信号分割成单独的咳嗽信号来具有标准输入。先前的研究已经发展为将咳嗽信号与非咳嗽信号分割。本研究评估了将单个音频文件中的多个咳嗽信号分割成多个单咳嗽信号的方法。我们评估了三种不同的方法,采用手动分割作为基线和自动分割。两种自动分割方法的分割精度分别为73%和70%,而手工分割的分割精度仅为49%。计算正确的单次咳嗽分割数的听力测试的一致性显示自动分割方法具有中等相关性,并且与手动分割相当。摘要:Research on diagnosing diseases based on voice signals currently are rapidly increasing, including cough-related diseases. When training the cough sound signals into deep learning models, it is necessary to have a standard input by segmenting several cough signals into individual cough signals. Previous research has been developed to segment cough signals from non-cough signals. This research evaluates the segmentation methods of several cough signals from a single audio file into several single-cough signals. We evaluate three different methods employing manual segmentation as a baseline and automatic segmentation. The results by two automatic segmentation methods obtained precisions of 73% and 70% compared to 49% by manual segmentation. The agreements of listening tests to count the number of correct single-cough segmentations show fair and moderate correlations for automatic segmentation methods and are comparable with manual segmentation.
【12】 ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
标题:ASVspoof 2021:走向野外的欺骗和深度假冒语音检测
链接:https://arxiv.org/abs/2210.02437
* 与cs.SD语音【9】为同一篇
作者:Xuechen Liu,Xin Wang,Md Sahidullah,Jose Patino,Héctor Delgado,Tomi Kinnunen,Massimiliano Todisco,Junichi Yamagishi,Nicholas Evans,Andreas Nautsch,Kong Aik Lee备注:Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing摘要:基准测试倡议支持对语音和语言处理中突出问题的竞争解决方案进行有意义的比较。连续的基准评估通常反映了从理想的实验室条件到野外条件的逐步演变。ASVspoof,欺骗和deepfake检测倡议和挑战系列,也遵循了同样的趋势。本文提供了ASVspoof 2021挑战赛的总结和37支参赛队伍的结果。对于逻辑访问任务,结果表明对抗方案对新引入的编码和传输效应是鲁棒的。物理访问任务的结果表明,与模拟的物理空间相反,在真实环境中检测重放攻击的潜力,但是对模拟和真实声学环境之间的变化缺乏鲁棒性。DF任务是2021年版的新任务,旨在解决在线发布的被操纵、压缩的语音数据的检测问题。虽然检测解决方案提供了对压缩效应的一些弹性,但它们缺乏跨不同源数据集的通用性。除了总结每项任务的最佳执行系统、对影响数据因素的新分析和隐藏数据子集的结果之外,本文还包括对挑战后结果的回顾、主要挑战限制的概述和ASVspoof未来的路线图。链接到ASVspoof挑战和相关资源:https://www.asvspoof.org/index2021.html摘要:Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 37 participating teams. For the logical access task, results indicate that countermeasures solutions are robust to newly introduced encoding and transmission effects. Results for the physical access task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The DF task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof. Link to the ASVspoof challenge and related resources: https://www.asvspoof.org/index2021.html
【13】 TC-SKNet with GridMask for Low-complexity Classification of Acoustic scene
标题:基于网格掩码的TC-SKNet低复杂度声场景分类
链接:https://arxiv.org/abs/2210.02287
* 与cs.SD语音【10】为同一篇
作者:Luyuan Xie,Yan Zhong,Lin Yang,Zhaoyu Yan,Zhonghai Wu,Junjie Wang机构:School of Software and Microelectronics, Peking University, Beijing, China备注:Accepted to APSIPA ASC 2022摘要:卷积神经网络(CNNs)在低复杂度分类任务(如声学场景分类(ASCs))中具有良好的性能。然而,关于目标语音长度与卷积核大小之间关系的研究却很少。本文将选择性核网络与时域卷积相结合(TC-SKNet),通过调节卷积核的感受野,在保持低复杂度的同时,解决了目标语音长度可变的问题。GridMask是一种通过屏蔽部分原始数据或特征区域的数据扩充策略。由于辍学的作用,提高了模型的泛化能力。在我们的实验中,GridMask带来的性能增益比ASC中的频谱扩展更强。最后,采用AutoML搜索TC-SKNet的最佳结构和GridMask的超参数,以提高分类性能。结果,59.87% TC-SKNet的峰值准确度与SOTA的峰值准确度相当,但参数仅使用20.9K。摘要:Convolution neural networks (CNNs) have good performance in low-complexity classification tasks such as acoustic scene classifications (ASCs). However, there are few studies on the relationship between the length of target speech and the size of the convolution kernels. In this paper, we combine Selective Kernel Network with Temporal-Convolution (TC-SKNet) to adjust the receptive field of convolution kernels to solve the problem of variable length of target voice while keeping low-complexity. GridMask is a data augmentation strategy by masking part of the raw data or feature area. It can enhance the generalization of the model as the role of dropout. In our experiments, the performance gain brought by GridMask is stronger than spectrum augmentation in ASCs. Finally, we adopt AutoML to search best structure of TC-SKNet and hyperparameters of GridMask for improving the classification performance. As a result, a peak accuracy of 59.87% TC-SKNet is equivalent to that of SOTA, but the parameters only use 20.9 K.
机器翻译,仅供参考