今天跟大家分享一篇语音相关的论文合集:cs.SD语音9篇,eess.AS音频处理10篇。

cs.SD语音

【1】 Audio Deepfake Detection Based on a Combination of F0 Information and  Real Plus Imaginary Spectrogram Features

标题:基于F0信息和实虚谱图特征的音频深度伪造检测

链接:https://arxiv.org/abs/2208.01214

作者:Jun Xue,Cunhang Fan,Zhao Lv,Jianhua Tao,Jiangyan Yi,Chengshi Zheng,Zhengqi Wen,Minmin Yuan,Shegang Shao
机构:School of Computer Science and, Technology, Anhui University, Hefei, China, NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China, Institute of Acoustics, Chinese, Qiyuan Laboratory, National Environmental Protection, Engineering and Technology Center
摘要:近年来,研究者们提出了大量的声学特征(对数功率谱图、线性频率倒谱系数、常数Q倒谱系数等)用于音频deepfake检测,获得了良好的性能,并表明不同子带对音频deepfake检测的贡献不同,但缺乏对子带中具体信息的解释。并且这些特征也丢失了相位等信息,受合成语音机理的启发,(F0)信息被用来改善合成语音的质量,然而合成语音的F0仍然太平均,这与真实语音有很大的差别,F_ 0可以作为鉴别真假语音的重要信息,但由于F_ 0的不规则分布,该方法选取包含F0大部分的频带作为输入特征,同时,为了充分利用相位和全频带信息,提出了将实、虚谱图特征作为补充特征,并对不相交的子带分别建模,最后,对F0,在ASVspoof2019LA数据集上的实验结果表明,该系统对音频deepfake检测任务是非常有效的,实现了0.43%的等效误码率(EER),这超过了几乎所有的系统。
摘要:Recently, pioneer research works have proposed a large number of acoustic features (log power spectrogram, linear frequency cepstral coefficients, constant Q cepstral coefficients, etc.) for audio deepfake detection, obtaining good performance, and showing that different subbands have different contributions to audio deepfake detection. However, this lacks an explanation of the specific information in the subband, and these features also lose information such as phase. Inspired by the mechanism of synthetic speech, the fundamental frequency (F0) information is used to improve the quality of synthetic speech, while the F0 of synthetic speech is still too average, which differs significantly from that of real speech. It is expected that F0 can be used as important information to discriminate between bonafide and fake speech, while this information cannot be used directly due to the irregular distribution of F0. Insteadly, the frequency band containing most of F0 is selected as the input feature. Meanwhile, to make full use of the phase and full-band information, we also propose to use real and imaginary spectrogram features as complementary input features and model the disjoint subbands separately. Finally, the results of F0, real and imaginary spectrogram features are fused. Experimental results on the ASVspoof 2019 LA dataset show that our proposed system is very effective for the audio deepfake detection task, achieving an equivalent error rate (EER) of 0.43%, which surpasses almost all systems.


【2】 Analog Gated Recurrent Neural Network for Detecting Chewing Events

标题:用于咀嚼事件检测的模拟门控递归神经网络

链接:https://arxiv.org/abs/2208.01201

作者:Kofi Odame,Maria Nyamukuru,Mohsen Shahghasemi,Shengjie Bi,David Kotz
备注:11 pages, 16 figures
摘要:提出了一种新颖的门控递归神经网络,用于检测人在咀嚼食物的时间.我们用0.18 μ m CMOS工艺将神经网络实现为定制的模拟集成电路.神经网络用从安装在志愿者乳突骨上的接触式麦克风收集的6.4小时的数据训练.当用1.6小时的先前未见过的数据测试时,该神经网络以24秒的时间分辨率识别咀嚼事件。它实现了91%的召回率和94%的F1得分,同时消耗了1.1微瓦的功率。一个基于新型模拟神经网络的检测整个进食事件——如正餐和小吃——的系统估计消耗了18微瓦的功率。8uW的功率。
摘要:We present a novel gated recurrent neural network to detect when a person is chewing on food. We implemented the neural network as a custom analog integrated circuit in a 0.18 um CMOS technology. The neural network was trained on 6.4 hours of data collected from a contact microphone that was mounted on volunteers' mastoid bones. When tested on 1.6 hours of previously-unseen data, the neural network identified chewing events at a 24-second time resolution. It achieved a recall of 91% and an F1-score of 94% while consuming 1.1 uW of power. A system for detecting whole eating episodes -- like meals and snacks -- that is based on the novel analog neural network consumes an estimated 18.8uW of power.


【3】 SampleMatch: Drum Sample Retrieval by Musical Context

标题:样本匹配:基于音乐上下文的鼓样本检索

链接:https://arxiv.org/abs/2208.01141

作者:Stefan Lattner
机构:Sony Computer Science Laboratories (CSL), Paris, France
备注:8 pages, 3 figures, 1 table; Accepted at the ISMIR conference, Bengaluru, India, 2022
摘要:现代数字音乐制作通常涉及组合大量声学元素来编辑一首乐曲。此类元素的重要类型是鼓样本,鼓样本确定乐曲的打击乐组成部分的特征。艺术家必须使用他们的审美判断来评估给定鼓样本是否适合当前音乐背景。然而,摘要从一个庞大的鼓库中选取鼓样本是一件繁琐的工作,而且可能会中断创作流程.在这项工作中,我们探索了基于从数据中学习到的美学原则的鼓样本自动检索.结果表明,艺术家可以通过在制作过程的不同阶段适合于某些音乐上下文来对他们的库中的样本进行排序(即,通过适合于不完整的歌曲混合)。为此,我们使用对比学习来最大化源自与混合相同的歌曲的鼓样本的分数。我们进行了一个听力测试来确定人工评分是否与自动评分函数相匹配,我们还进行了客观的定量分析来评估我们方法的有效性。
摘要:Modern digital music production typically involves combining numerous acoustic elements to compile a piece of music. Important types of such elements are drum samples, which determine the characteristics of the percussive components of the piece. Artists must use their aesthetic judgement to assess whether a given drum sample fits the current musical context. However, selecting drum samples from a potentially large library is tedious and may interrupt the creative flow. In this work, we explore the automatic drum sample retrieval based on aesthetic principles learned from data. As a result, artists can rank the samples in their library by fit to some musical context at different stages of the production process (i.e., by fit to incomplete song mixtures). To this end, we use contrastive learning to maximize the score of drum samples originating from the same song as the mixture. We conduct a listening test to determine whether the human ratings match the automatic scoring function. We also perform objective quantitative analyses to evaluate the efficacy of our approach.


【4】 Low-complexity CNNs for Acoustic Scene Classification

标题:基于低复杂度神经网络的声学场景分类

链接:https://arxiv.org/abs/2208.01555

作者:Arshdeep Singh,James A King,Xubo Liu,Wenwu Wang,Mark D. Plumbley
机构:Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey, UK
备注:Technical Report DCASE 2022 TASK 1. arXiv admin note: substantial text overlap with arXiv:2207.11529
摘要:本技术报告描述了SurreyAudioTeam22s提交的DCASE 2022 ASC任务1,低复杂度声学场景分类(ASC)。该任务有两个规则,(a)ASC框架应具有最大128K的参数,以及(b)最多应该有3000万次乘加运算在本报告中,我们提出了一个低复杂度的ASC系统,它遵循该任务的预定规则。
摘要:This technical report describes the SurreyAudioTeam22s submission for DCASE 2022 ASC Task 1, Low-Complexity Acoustic Scene Classification (ASC). The task has two rules, (a) the ASC framework should have maximum 128K parameters, and (b) there should be a maximum of 30 millions multiply-accumulate operations (MACs) per inference. In this report, we present low-complexity systems for ASC that follow the rules intended for the task.


【5】 Voice Analysis for Stress Detection and Application in Virtual Reality  to Improve Public Speaking in Real-time: A Review

标题:用于重音检测的语音分析和在虚拟现实中的应用,以实时改善公众演讲:回顾

链接:https://arxiv.org/abs/2208.01041

作者:Arushi,Roberto Dillon,Ai Ni Teoh,Denise Dillon
机构:A Review   Arushi James Cook University, au Roberto Dillon James Cook University, au Ai Ni Teoh James Cook University, au Denise Dillon James Cook University
备注:41 pages, 7 figures, 4 tables
摘要:在公开演讲期间的压力是常见的并且不利地影响表现和自信。已经进行了广泛的研究以开发各种模型来识别情绪状态。然而,已经进行了很少的研究来使用语音分析实时检测公开演讲期间的压力。在这种情况下,当前的评审表明,算法的应用没有得到适当的探索,这有助于确定创建合适的测试环境的主要障碍,同时说明本文提出了一种可以集成到虚拟现实系统中的应力检测计算模型(VR)应用程序来创建智能虚拟听众,以提高公众演讲技能。所开发的模型在与VR集成时,将能够通过分析与指示压力的生理参数相关的语音特征来实时检测过度的压力,并帮助用户逐渐控制过度的压力和提高公共演讲能力性能指标
摘要:Stress during public speaking is common and adversely affects performance and self-confidence. Extensive research has been carried out to develop various models to recognize emotional states. However, minimal research has been conducted to detect stress during public speaking in real time using voice analysis. In this context, the current review showed that the application of algorithms was not properly explored and helped identify the main obstacles in creating a suitable testing environment while accounting for current complexities and limitations. In this paper, we present our main idea and propose a stress detection computational algorithmic model that could be integrated into a Virtual Reality (VR) application to create an intelligent virtual audience for improving public speaking skills. The developed model, when integrated with VR, will be able to detect excessive stress in real time by analysing voice features correlated to physiological parameters indicative of stress and help users gradually control excessive stress and improve public speaking performance

【6】 Jazz Contrafact Detection

标题:Jazz冲突检测

链接:https://arxiv.org/abs/2208.00792

作者:C. Bunks,T. Weyde
机构:Licensed under a Creative, Commons Attribution ,., International License (CC BY ,.,)., bedding for chords based on a co-occurrence matrix [,], computed from a corpus of symbolic jazz chord progres-, sions., As many chords in our corpus occur only rarely, the
备注:8 pages, 6 figures, 4 tables
摘要:在爵士乐中,矛盾音是在一个已经存在但经常被重新和声的和弦进行上创作的一个新旋律。由于重新和声会引入各种各样的变化,因此检测矛盾音是一项具有挑战性的任务。本文提出了一种新颖的表示和弦进行的向量空间模型,并将其用于矛盾音的检测。该过程应用音乐理论的原理来降低和弦空间的维数,该方法首先确定一个公共密钥签名表示,然后计算一个弦共现矩阵,矩阵的行构成向量空间的基,弦级数在向量空间中表示为分段线性函数,通过计算膜面积这一新的距离度量来评估和声相似性.为了说明该方法的有效性,我们将其应用于包含2,612个和弦进行的Impro-Visor语料库,并给出示例来证明其解释重调和和发现矛盾的能力。
摘要:In jazz, a contrafact is a new melody composed over an existing, but often reharmonized chord progression. Because reharmonization can introduce a wide range of variations, detecting contrafacts is a challenging task. This paper develops a novel vector-space model to represent chord progressions, and uses it for contrafact detection. The process applies principles from music theory to reduce the dimensionality of chord space, determine a common key signature representation, and compute a chordal co-occurrence matrix. The rows of the matrix form a basis for the vector space in which chord progressions are represented as piecewise linear functions, and harmonic similarity is evaluated by computing the membrane area, a novel distance metric. To illustrate our method's effectiveness, we apply it to the Impro-Visor corpus of 2,612 chord progressions, and present examples demonstrating its ability to account for reharmonizations and find contrafacts.


【7】 A 23 $μ$W Keyword Spotting IC with Ring-Oscillator-Based Time-Domain  Feature Extraction

标题: 基于环形振荡器时域特征提取的23 μ W关键字识别芯片

链接:https://arxiv.org/abs/2208.00693

作者:Kwantae Kim,Chang Gao,Rui Graça,Ilya Kiselev,Hoi-Jun Yoo,Tobi Delbruck,Shih-Chii Liu
机构:University of Z¨urichand ETH Z¨urich
备注:14 pages, 21 figures, 2 tables
摘要:这篇文章介绍了第一个关键字识别(KWS)IC,其使用基于环形振荡器的时域处理技术用于其模拟特征提取器其对时间编码方案的广泛使用允许以除模拟前端的电压到时间转换级之外的全时域方式处理模拟音频信号。受益于基于数字逻辑门的基本构建块,与传统的电压域设计相比,它提供了更好的技术可扩展性。采用65纳米CMOS工艺制造,KWS集成电路样机包括模拟场效应管和数字神经网络分类器在内,芯片面积为2.03mm,功耗为23 μ W。测量结果验证了所提出的IC在Google语音命令数据集(GSCD)上执行12类KWS任务时具有〉86%的准确度和12.4ms的延迟。
摘要:This article presents the first keyword spotting (KWS) IC which uses a ring-oscillator-based time-domain processing technique for its analog feature extractor (FEx). Its extensive usage of time-encoding schemes allows the analog audio signal to be processed in a fully time-domain manner except for the voltage-to-time conversion stage of the analog front-end. Benefiting from fundamental building blocks based on digital logic gates, it offers a better technology scalability compared to conventional voltage-domain designs. Fabricated in a 65 nm CMOS process, the prototyped KWS IC occupies 2.03mm$^{2}$ and dissipates 23 $\mu$W power consumption including analog FEx and digital neural network classifier. The 16-channel time-domain FEx achieves 54.89 dB dynamic range for 16 ms frame shift size while consuming 9.3 $\mu$W. The measurement result verifies that the proposed IC performs a 12-class KWS task on the Google Speech Command Dataset (GSCD) with >86% accuracy and 12.4 ms latency.


【8】 UAVM: A Unified Model for Audio-Visual Learning

标题:UAVM:一个统一的视听教学模型

链接:https://arxiv.org/abs/2208.00061

作者:Yuan Gong,Alexander H. Liu,Andrew Rouditchenko,James Glass
机构:Massachusetts Instituteof Technology
摘要:传统的视听模型都有独立的音频和视频分支,我们设计了一个统一的视听模型UAVM(Unified Audio-Visual Model).本文描述了UAVM,报道了它在VGGSound上最新的视听事件分类准确率为65.8%,并描述了该模型的有趣特性.
摘要:Conventional audio-visual models have independent audio and video branches. We design a unified model for audio and video processing called Unified Audio-Visual Model (UAVM). In this paper, we describe UAVM, report its new state-of-the-art audio-visual event classification accuracy of 65.8% on VGGSound, and describe the intriguing properties of the model.


【9】 DENT-DDSP: Data-efficient noisy speech generator using differentiable  digital signal processors for explicit distortion modelling and noise-robust  speech recognition

标题:DENT-DDSP:利用可微分数字信号处理器进行显式失真建模和噪声鲁棒语音识别的数据高效噪声语音发生器

链接:https://arxiv.org/abs/2208.00987

作者:Z. Guo,C. Chen,E. S. Chng
机构:Chng Eng Siong 1 1Nanyang Technological University
摘要:自动语音识别的性能(ASR)系统在噪声条件下性能急剧下降。(EDM)作为一种特征补偿步骤,通过模拟干净语音的带噪语音,能够在这种条件下增强ASR系统,但现有的失真模型要么不可训练,要么无法解释,而且往往缺乏可控性和泛化能力.本文提出了一种基于EDM的ASR系统,我们提出了一个完全可解释和可控制模型:DENT-DDSP实现电火花加工。DENT-DDSP采用新颖的可微分数字信号处理技术实验结果表明,DENT-DDSP模拟的噪声数据在多尺度光谱损失方面达到了最高的模拟逼真度此外,为了验证DENT-DDSP模拟的数据是否能够替代噪声鲁棒ASR任务中稀缺的域内噪声数据,使用仿真数据和实际数据训练具有相同体系结构的几个下游ASR模型。实验结果表明,用DENT DDSP模拟的噪声数据训练的模型与基准测试系统相比,在误字率(WER)方面相差2 .7%,取得了相似的性能,模型代码已在线发布.
摘要:The performances of automatic speech recognition (ASR) systems degrade drastically under noisy conditions. Explicit distortion modelling (EDM), as a feature compensation step, is able to enhance ASR systems under such conditions by simulating the in-domain noisy speeches from the clean counterparts. Yet, existing distortion models are either non-trainable or unexplainable and often lack controllability and generalization ability. In this paper, we propose a fully explainable and controllable model: DENT-DDSP to achieve EDM. DENT-DDSP utilizes novel differentiable digital signal processing (DDSP) components and requires only 10 seconds of training data to achieve high fidelity. The experiment shows that the simulated noisy data from DENT-DDSP achieves the highest simulation fidelity compared to other baseline models in terms of multi-scale spectral loss (MSSL). Moreover, to validate whether the data simulated by DENT-DDSP are able to replace the scarce in-domain noisy data in the noise-robust ASR tasks, several downstream ASR models with the same architecture are trained using the simulated data and the real data. The experiment shows that the model trained with the simulated noisy data from DENT-DDSP achieves similar performances to the benchmark with a 2.7\% difference in terms of word error rate (WER). The code of the model is released online.


eess.AS音频处理

【1】 Low-complexity CNNs for Acoustic Scene Classification

标题:基于低复杂度神经网络的声学场景分类

链接:https://arxiv.org/abs/2208.01555

* 与cs.SD语音【4】为同一篇

作者:Arshdeep Singh,James A King,Xubo Liu,Wenwu Wang,Mark D. Plumbley
机构:Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey, UK
备注:Technical Report DCASE 2022 TASK 1. arXiv admin note: substantial text overlap with arXiv:2207.11529
摘要:本技术报告描述了SurreyAudioTeam22s提交的DCASE 2022 ASC任务1,低复杂度声学场景分类(ASC)。该任务有两个规则,(a)ASC框架应具有最大128K的参数,以及(b)最多应该有3000万次乘加运算在本报告中,我们提出了一个低复杂度的ASC系统,它遵循该任务的预定规则。
摘要:This technical report describes the SurreyAudioTeam22s submission for DCASE 2022 ASC Task 1, Low-Complexity Acoustic Scene Classification (ASC). The task has two rules, (a) the ASC framework should have maximum 128K parameters, and (b) there should be a maximum of 30 millions multiply-accumulate operations (MACs) per inference. In this report, we present low-complexity systems for ASC that follow the rules intended for the task.


【2】 Voice Analysis for Stress Detection and Application in Virtual Reality  to Improve Public Speaking in Real-time: A Review

标题:用于重音检测的语音分析和在虚拟现实中的应用,以实时改善公众演讲:回顾

链接:https://arxiv.org/abs/2208.01041

* 与cs.SD语音【5】为同一篇

作者:Arushi,Roberto Dillon,Ai Ni Teoh,Denise Dillon
机构:A Review   Arushi James Cook University, au Roberto Dillon James Cook University, au Ai Ni Teoh James Cook University, au Denise Dillon James Cook University
备注:41 pages, 7 figures, 4 tables
摘要:在公开演讲期间的压力是常见的并且不利地影响表现和自信。已经进行了广泛的研究以开发各种模型来识别情绪状态。然而,已经进行了很少的研究来使用语音分析实时检测公开演讲期间的压力。在这种情况下,当前的评审表明,算法的应用没有得到适当的探索,这有助于确定创建合适的测试环境的主要障碍,同时说明本文提出了一种可以集成到虚拟现实系统中的应力检测计算模型(VR)应用程序来创建智能虚拟听众,以提高公众演讲技能。所开发的模型在与VR集成时,将能够通过分析与指示压力的生理参数相关的语音特征来实时检测过度的压力,并帮助用户逐渐控制过度的压力和提高公共演讲能力性能指标
摘要:Stress during public speaking is common and adversely affects performance and self-confidence. Extensive research has been carried out to develop various models to recognize emotional states. However, minimal research has been conducted to detect stress during public speaking in real time using voice analysis. In this context, the current review showed that the application of algorithms was not properly explored and helped identify the main obstacles in creating a suitable testing environment while accounting for current complexities and limitations. In this paper, we present our main idea and propose a stress detection computational algorithmic model that could be integrated into a Virtual Reality (VR) application to create an intelligent virtual audience for improving public speaking skills. The developed model, when integrated with VR, will be able to detect excessive stress in real time by analysing voice features correlated to physiological parameters indicative of stress and help users gradually control excessive stress and improve public speaking performance


【3】 Audio Deepfake Detection Based on a Combination of F0 Information and  Real Plus Imaginary Spectrogram Features

标题:基于F0信息和实虚谱图特征的音频深度伪造检测

链接:https://arxiv.org/abs/2208.01214

* 与cs.SD语音【1】为同一篇

作者:Jun Xue,Cunhang Fan,Zhao Lv,Jianhua Tao,Jiangyan Yi,Chengshi Zheng,Zhengqi Wen,Minmin Yuan,Shegang Shao
机构:School of Computer Science and, Technology, Anhui University, Hefei, China, NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China, Institute of Acoustics, Chinese, Qiyuan Laboratory, National Environmental Protection, Engineering and Technology Center
摘要:近年来,研究者们提出了大量的声学特征(对数功率谱图、线性频率倒谱系数、常数Q倒谱系数等)用于音频deepfake检测,获得了良好的性能,并表明不同子带对音频deepfake检测的贡献不同,但缺乏对子带中具体信息的解释。并且这些特征也丢失了相位等信息,受合成语音机理的启发,(F0)信息被用来改善合成语音的质量,然而合成语音的F0仍然太平均,这与真实语音有很大的差别,F_ 0可以作为鉴别真假语音的重要信息,但由于F_ 0的不规则分布,该方法选取包含F0大部分的频带作为输入特征,同时,为了充分利用相位和全频带信息,提出了将实、虚谱图特征作为补充特征,并对不相交的子带分别建模,最后,对F0,在ASVspoof2019LA数据集上的实验结果表明,该系统对音频deepfake检测任务是非常有效的,实现了0.43%的等效误码率(EER),这超过了几乎所有的系统。
摘要:Recently, pioneer research works have proposed a large number of acoustic features (log power spectrogram, linear frequency cepstral coefficients, constant Q cepstral coefficients, etc.) for audio deepfake detection, obtaining good performance, and showing that different subbands have different contributions to audio deepfake detection. However, this lacks an explanation of the specific information in the subband, and these features also lose information such as phase. Inspired by the mechanism of synthetic speech, the fundamental frequency (F0) information is used to improve the quality of synthetic speech, while the F0 of synthetic speech is still too average, which differs significantly from that of real speech. It is expected that F0 can be used as important information to discriminate between bonafide and fake speech, while this information cannot be used directly due to the irregular distribution of F0. Insteadly, the frequency band containing most of F0 is selected as the input feature. Meanwhile, to make full use of the phase and full-band information, we also propose to use real and imaginary spectrogram features as complementary input features and model the disjoint subbands separately. Finally, the results of F0, real and imaginary spectrogram features are fused. Experimental results on the ASVspoof 2019 LA dataset show that our proposed system is very effective for the audio deepfake detection task, achieving an equivalent error rate (EER) of 0.43%, which surpasses almost all systems.


【4】 Analog Gated Recurrent Neural Network for Detecting Chewing Events

标题:用于咀嚼事件检测的模拟门控递归神经网络

链接:https://arxiv.org/abs/2208.01201

* 与cs.SD语音【2】为同一篇

作者:Kofi Odame,Maria Nyamukuru,Mohsen Shahghasemi,Shengjie Bi,David Kotz
备注:11 pages, 16 figures
摘要:提出了一种新颖的门控递归神经网络,用于检测人在咀嚼食物的时间.我们用0.18 μ m CMOS工艺将神经网络实现为定制的模拟集成电路.神经网络用从安装在志愿者乳突骨上的接触式麦克风收集的6.4小时的数据训练.当用1.6小时的先前未见过的数据测试时,该神经网络以24秒的时间分辨率识别咀嚼事件。它实现了91%的召回率和94%的F1得分,同时消耗了1.1微瓦的功率。一个基于新型模拟神经网络的检测整个进食事件——如正餐和小吃——的系统估计消耗了18微瓦的功率。8uW的功率。
摘要:We present a novel gated recurrent neural network to detect when a person is chewing on food. We implemented the neural network as a custom analog integrated circuit in a 0.18 um CMOS technology. The neural network was trained on 6.4 hours of data collected from a contact microphone that was mounted on volunteers' mastoid bones. When tested on 1.6 hours of previously-unseen data, the neural network identified chewing events at a 24-second time resolution. It achieved a recall of 91% and an F1-score of 94% while consuming 1.1 uW of power. A system for detecting whole eating episodes -- like meals and snacks -- that is based on the novel analog neural network consumes an estimated 18.8uW of power.


【5】 SampleMatch: Drum Sample Retrieval by Musical Context

标题:样本匹配:基于音乐上下文的鼓样本检索

链接:https://arxiv.org/abs/2208.01141

* 与cs.SD语音【3】为同一篇

作者:Stefan Lattner
机构:Sony Computer Science Laboratories (CSL), Paris, France
备注:8 pages, 3 figures, 1 table; Accepted at the ISMIR conference, Bengaluru, India, 2022
摘要:现代数字音乐制作通常涉及组合大量声学元素来编辑一首乐曲。此类元素的重要类型是鼓样本,鼓样本确定乐曲的打击乐组成部分的特征。艺术家必须使用他们的审美判断来评估给定鼓样本是否适合当前音乐背景。然而,摘要从一个庞大的鼓库中选取鼓样本是一件繁琐的工作,而且可能会中断创作流程.在这项工作中,我们探索了基于从数据中学习到的美学原则的鼓样本自动检索.结果表明,艺术家可以通过在制作过程的不同阶段适合于某些音乐上下文来对他们的库中的样本进行排序(即,通过适合于不完整的歌曲混合)。为此,我们使用对比学习来最大化源自与混合相同的歌曲的鼓样本的分数。我们进行了一个听力测试来确定人工评分是否与自动评分函数相匹配,我们还进行了客观的定量分析来评估我们方法的有效性。
摘要:Modern digital music production typically involves combining numerous acoustic elements to compile a piece of music. Important types of such elements are drum samples, which determine the characteristics of the percussive components of the piece. Artists must use their aesthetic judgement to assess whether a given drum sample fits the current musical context. However, selecting drum samples from a potentially large library is tedious and may interrupt the creative flow. In this work, we explore the automatic drum sample retrieval based on aesthetic principles learned from data. As a result, artists can rank the samples in their library by fit to some musical context at different stages of the production process (i.e., by fit to incomplete song mixtures). To this end, we use contrastive learning to maximize the score of drum samples originating from the same song as the mixture. We conduct a listening test to determine whether the human ratings match the automatic scoring function. We also perform objective quantitative analyses to evaluate the efficacy of our approach.


【6】 Amino Acid Classification in 2D NMR Spectra via Acoustic Signal  Embeddings

标题:基于声信号嵌入的二维核磁共振氨基酸分类

链接:https://arxiv.org/abs/2208.00935

作者:Jia Qi Yip,Dianwen Ng,Bin Ma,Konstantin Pervushin,Eng Siong Chng
机构:§ Alibaba Group, † Nanyang Technological University, Singapore
摘要:核磁共振(核磁共振)用于结构生物学,用来实验测定蛋白质的结构,这在生物学的许多领域都有应用,是药物开发的重要组成部分,不幸的是,{1} NMR {2}数据的采集成本可能高达数千美元,而且专家可能需要数周时间才能将观察到的共振分配给特定的化学基团。因此,NMR社区对使用深度学习自动化NMR数据注释的兴趣越来越大。由于NMR和音频数据之间的相似性,我们建议声学信号处理中使用的方法也可以应用于NMR。使用模拟氨基酸数据集,我们证明了通过用可训练卷积编码器替换滤波器组,来自说话人确认模型的声学信号嵌入可以通过将每个氨基酸作为唯一的说话人来处理而用于2D NMR谱中的氨基酸分类。在大小与46小时的音频相当的NMR数据集上,我们在20类问题上获得了97.7%的分类性能,与现有的基于NMR的模型相比,我们还通过使用声学嵌入模型获得了23%的相对改进。
摘要:Nuclear Magnetic Resonance (NMR) is used in structural biology to experimentally determine the structure of proteins, which is used in many areas of biology and is an important part of drug development. Unfortunately, NMR data can cost thousands of dollars per sample to collect and it can take a specialist weeks to assign the observed resonances to specific chemical groups. There has thus been growing interest in the NMR community to use deep learning to automate NMR data annotation. Due to similarities between NMR and audio data, we propose that methods used in acoustic signal processing can be applied to NMR as well. Using a simulated amino acid dataset, we show that by swapping out filter banks with a trainable convolutional encoder, acoustic signal embeddings from speaker verification models can be used for amino acid classification in 2D NMR spectra by treating each amino acid as a unique speaker. On an NMR dataset comparable in size with of 46 hours of audio, we achieve a classification performance of 97.7% on a 20-class problem. We also achieve a 23% relative improvement by using an acoustic embedding model compared to an existing NMR-based model.


【7】 DENT-DDSP: Data-efficient noisy speech generator using differentiable  digital signal processors for explicit distortion modelling and noise-robust  speech recognition

标题:DENT-DDSP:利用可微分数字信号处理器进行显式失真建模和噪声鲁棒语音识别的数据高效噪声语音发生器

链接:https://arxiv.org/abs/2208.00987

* 与cs.SD语音【9】为同一篇

作者:Z. Guo,C. Chen,E. S. Chng
机构:Chng Eng Siong 1 1Nanyang Technological University
摘要:自动语音识别的性能(ASR)系统在噪声条件下性能急剧下降。(EDM)作为一种特征补偿步骤,通过模拟干净语音的带噪语音,能够在这种条件下增强ASR系统,但现有的失真模型要么不可训练,要么无法解释,而且往往缺乏可控性和泛化能力.本文提出了一种基于EDM的ASR系统,我们提出了一个完全可解释和可控制模型:DENT-DDSP实现电火花加工。DENT-DDSP采用新颖的可微分数字信号处理技术实验结果表明,DENT-DDSP模拟的噪声数据在多尺度光谱损失方面达到了最高的模拟逼真度此外,为了验证DENT-DDSP模拟的数据是否能够替代噪声鲁棒ASR任务中稀缺的域内噪声数据,使用仿真数据和实际数据训练具有相同体系结构的几个下游ASR模型。实验结果表明,用DENT DDSP模拟的噪声数据训练的模型与基准测试系统相比,在误字率(WER)方面相差2 .7%,取得了相似的性能,模型代码已在线发布.
摘要:The performances of automatic speech recognition (ASR) systems degrade drastically under noisy conditions. Explicit distortion modelling (EDM), as a feature compensation step, is able to enhance ASR systems under such conditions by simulating the in-domain noisy speeches from the clean counterparts. Yet, existing distortion models are either non-trainable or unexplainable and often lack controllability and generalization ability. In this paper, we propose a fully explainable and controllable model: DENT-DDSP to achieve EDM. DENT-DDSP utilizes novel differentiable digital signal processing (DDSP) components and requires only 10 seconds of training data to achieve high fidelity. The experiment shows that the simulated noisy data from DENT-DDSP achieves the highest simulation fidelity compared to other baseline models in terms of multi-scale spectral loss (MSSL). Moreover, to validate whether the data simulated by DENT-DDSP are able to replace the scarce in-domain noisy data in the noise-robust ASR tasks, several downstream ASR models with the same architecture are trained using the simulated data and the real data. The experiment shows that the model trained with the simulated noisy data from DENT-DDSP achieves similar performances to the benchmark with a 2.7\% difference in terms of word error rate (WER). The code of the model is released online.


【8】 Jazz Contrafact Detection

标题:Jazz冲突检测

链接:https://arxiv.org/abs/2208.00792

* 与cs.SD语音【6】为同一篇

作者:C. Bunks,T. Weyde
机构:Licensed under a Creative, Commons Attribution ,., International License (CC BY ,.,)., bedding for chords based on a co-occurrence matrix [,], computed from a corpus of symbolic jazz chord progres-, sions., As many chords in our corpus occur only rarely, the
备注:8 pages, 6 figures, 4 tables
摘要:在爵士乐中,矛盾音是在一个已经存在但经常被重新和声的和弦进行上创作的一个新旋律。由于重新和声会引入各种各样的变化,因此检测矛盾音是一项具有挑战性的任务。本文提出了一种新颖的表示和弦进行的向量空间模型,并将其用于矛盾音的检测。该过程应用音乐理论的原理来降低和弦空间的维数,该方法首先确定一个公共密钥签名表示,然后计算一个弦共现矩阵,矩阵的行构成向量空间的基,弦级数在向量空间中表示为分段线性函数,通过计算膜面积这一新的距离度量来评估和声相似性.为了说明该方法的有效性,我们将其应用于包含2,612个和弦进行的Impro-Visor语料库,并给出示例来证明其解释重调和和发现矛盾的能力。
摘要:In jazz, a contrafact is a new melody composed over an existing, but often reharmonized chord progression. Because reharmonization can introduce a wide range of variations, detecting contrafacts is a challenging task. This paper develops a novel vector-space model to represent chord progressions, and uses it for contrafact detection. The process applies principles from music theory to reduce the dimensionality of chord space, determine a common key signature representation, and compute a chordal co-occurrence matrix. The rows of the matrix form a basis for the vector space in which chord progressions are represented as piecewise linear functions, and harmonic similarity is evaluated by computing the membrane area, a novel distance metric. To illustrate our method's effectiveness, we apply it to the Impro-Visor corpus of 2,612 chord progressions, and present examples demonstrating its ability to account for reharmonizations and find contrafacts.


【9】 A 23 $μ$W Keyword Spotting IC with Ring-Oscillator-Based Time-Domain  Feature Extraction

标题:基于环形振荡器时域特征提取的23 μ W关键字识别芯片

链接:https://arxiv.org/abs/2208.00693

* 与cs.SD语音【7】为同一篇

作者:Kwantae Kim,Chang Gao,Rui Graça,Ilya Kiselev,Hoi-Jun Yoo,Tobi Delbruck,Shih-Chii Liu
机构:University of Z¨urichand ETH Z¨urich
备注:14 pages, 21 figures, 2 tables
摘要:这篇文章介绍了第一个关键字识别(KWS)IC,其使用基于环形振荡器的时域处理技术用于其模拟特征提取器其对时间编码方案的广泛使用允许以除模拟前端的电压到时间转换级之外的全时域方式处理模拟音频信号。受益于基于数字逻辑门的基本构建块,与传统的电压域设计相比,它提供了更好的技术可扩展性。采用65纳米CMOS工艺制造,KWS集成电路样机包括模拟场效应管和数字神经网络分类器在内,芯片面积为2.03mm,功耗为23 μ W。测量结果验证了所提出的IC在Google语音命令数据集(GSCD)上执行12类KWS任务时具有〉86%的准确度和12.4ms的延迟。
摘要:This article presents the first keyword spotting (KWS) IC which uses a ring-oscillator-based time-domain processing technique for its analog feature extractor (FEx). Its extensive usage of time-encoding schemes allows the analog audio signal to be processed in a fully time-domain manner except for the voltage-to-time conversion stage of the analog front-end. Benefiting from fundamental building blocks based on digital logic gates, it offers a better technology scalability compared to conventional voltage-domain designs. Fabricated in a 65 nm CMOS process, the prototyped KWS IC occupies 2.03mm$^{2}$ and dissipates 23 $\mu$W power consumption including analog FEx and digital neural network classifier. The 16-channel time-domain FEx achieves 54.89 dB dynamic range for 16 ms frame shift size while consuming 9.3 $\mu$W. The measurement result verifies that the proposed IC performs a 12-class KWS task on the Google Speech Command Dataset (GSCD) with >86% accuracy and 12.4 ms latency.


【10】 UAVM: A Unified Model for Audio-Visual Learning

标题:UAVM:一个统一的视听教学模型

链接:https://arxiv.org/abs/2208.00061

* 与cs.SD语音【8】为同一篇

作者:Yuan Gong,Alexander H. Liu,Andrew Rouditchenko,James Glass
机构:Massachusetts Instituteof Technology
摘要:传统的视听模型都有独立的音频和视频分支,我们设计了一个统一的视听模型UAVM(Unified Audio-Visual Model).本文描述了UAVM,报道了它在VGGSound上最新的视听事件分类准确率为65.8%,并描述了该模型的有趣特性.
摘要:Conventional audio-visual models have independent audio and video branches. We design a unified model for audio and video processing called Unified Audio-Visual Model (UAVM). In this paper, we describe UAVM, report its new state-of-the-art audio-visual event classification accuracy of 65.8% on VGGSound, and describe the intriguing properties of the model.


机器翻译,仅供参考