【1】 Learning Unsupervised Hierarchies of Audio Concepts
标题:学习音频概念的无监督层次
链接:https://arxiv.org/abs/2207.11231
作者:Darius Afchar,Romain Hennequin,Vincent Guigue机构:Deezer Research, France, MLIA, Sorbonne Université, France摘要:音乐信号很难从它们的低级特征中解读出来,也许比图像更难:例如,突出显示频谱图或图像的一部分通常不足以传达与人类真正相关的高级思想。在计算机视觉中,提出了概念学习,以将解释调整到正确的抽象级别(例如,从射线照片中检测临床概念)。这些方法尚未用于MIR。在本文中,我们将概念学习应用于音乐领域,因为音乐领域有其特殊性,例如,音乐概念是典型的非独立性和混合性(例如,流派、乐器、情绪),不同于先前的假设解开概念的工作,我们提出一种从音频中学习大量音乐概念,然后自动将它们分层以揭示它们之间的相互关系的方法。我们在音乐流媒体服务的播放列表数据集上进行了实验,作为不同概念的一些注释示例。评估表明,挖掘的层次结构与概念的地面真实层次结构(如果可用)和一般情况下的概念相似性代理源都是一致的。摘要:Music signals are difficult to interpret from their low-level features, perhaps even more than images: e.g. highlighting part of a spectrogram or an image is often insufficient to convey high-level ideas that are genuinely relevant to humans. In computer vision, concept learning was therein proposed to adjust explanations to the right abstraction level (e.g. detect clinical concepts from radiographs). These methods have yet to be used for MIR. In this paper, we adapt concept learning to the realm of music, with its particularities. For instance, music concepts are typically non-independent and of mixed nature (e.g. genre, instruments, mood), unlike previous work that assumed disentangled concepts. We propose a method to learn numerous music concepts from audio and then automatically hierarchise them to expose their mutual relationships. We conduct experiments on datasets of playlists from a music streaming service, serving as a few annotated examples for diverse concepts. Evaluations show that the mined hierarchies are aligned with both ground-truth hierarchies of concepts -- when available -- and with proxy sources of concept similarity in the general case.
【2】 Inference skipping for more efficient real-time speech enhancement with parallel RNNs
标题:基于并行RNN的跳过推理实时语音增强
链接:https://arxiv.org/abs/2207.11108
作者:Xiaohuai Le,Tong Lei,Kai Chen,Jing Lu机构:Institute of Acoustics, Nanjing University备注:11 pages, 8 figures, accepted by IEEE/ACM TASLP摘要:深度神经网络基于DNN的语音增强模型由于其良好的性能而受到广泛的关注.然而,由于DNN的计算量大,难以在实时应用中部署一个功能强大的DNN.典型的压缩方法如剪枝和量化等都没有很好地利用数据的特性.本文提出了一种基于DNN的语音增强模型,将Skip-RNN策略引入到并行RNN语音增强模型中,RNN的状态在不中断输出掩码更新的情况下间歇性地更新,这导致计算负荷的显著减少而没有明显的音频伪像。为了更好地平衡语音和噪声之间的差异,我们进一步利用语音活动检测来调整跳过策略,(VAD)引导,节省了更多的计算量。实验采用了一种高性能的语音增强模型——双路卷积递归网络(DPCRN)的实验结果表明,与网络剪枝和直接训练一个较小的模型等策略相比,本文提出的算法具有更好的性能,同时也验证了本文提出的算法在其他两个竞争性语音增强模型上的推广性。摘要:Deep neural network (DNN) based speech enhancement models have attracted extensive attention due to their promising performance. However, it is difficult to deploy a powerful DNN in real-time applications because of its high computational cost. Typical compression methods such as pruning and quantization do not make good use of the data characteristics. In this paper, we introduce the Skip-RNN strategy into speech enhancement models with parallel RNNs. The states of the RNNs update intermittently without interrupting the update of the output mask, which leads to significant reduction of computational load without evident audio artifacts. To better leverage the difference between the voice and the noise, we further regularize the skipping strategy with voice activity detection (VAD) guidance, saving more computational load. Experiments on a high-performance speech enhancement model, dual-path convolutional recurrent network (DPCRN), show the superiority of our strategy over strategies like network pruning or directly training a smaller model. We also validate the generalization of the proposed strategy on two other competitive speech enhancement models.
【3】 Head-Related Transfer Function Interpolation from Spatially Sparse Measurements Using Autoencoder with Source Position Conditioning
标题:使用带源位置调节的自动编码器从空间稀疏测量值进行头部相关传递函数插值
链接:https://arxiv.org/abs/2207.10967
作者:Yuki Ito,Tomohiko Nakamura,Shoichi Koyama,Hiroshi Saruwatari机构:Graduate School of Information Science and Technology, The University of Tokyo, -,-, Hongo, Bunkyo-ku, Tokyo ,-, Japan备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2022摘要:本文提出了一种头相关传递函数的方法本文提出的方法是基于正则化线性回归的HRTF插值方法与基于自编码器的HRTF插值方法的类比(RLR)和自动编码器。通过这种类比,我们发现基于RLR的方法的主要特点是将HRTF分解为与源位置相关的和与源位置无关的两个因子,我们设计编码器和解码器使得它们的权值和偏置是从源位置产生的,我们引入了一个聚集模块,减少了隐变量对源位置的依赖性,从而获得了与源位置无关的每个对象的表示.数值实验表明,所提出的方法可以很好地适用于不可见对象,并且只需要一个源位置就可以获得插值性能.第八测量值与基于RLR的方法的测量值相当。摘要:We propose a method of head-related transfer function (HRTF) interpolation from sparsely measured HRTFs using an autoencoder with source position conditioning. The proposed method is drawn from an analogy between an HRTF interpolation method based on regularized linear regression (RLR) and an autoencoder. Through this analogy, we found the key feature of the RLR-based method that HRTFs are decomposed into source-position-dependent and source-position-independent factors. On the basis of this finding, we design the encoder and decoder so that their weights and biases are generated from source positions. Furthermore, we introduce an aggregation module that reduces the dependence of latent variables on source position for obtaining a source-position-independent representation of each subject. Numerical experiments show that the proposed method can work well for unseen subjects and achieve an interpolation performance with only one-eighth measurements comparable to that of the RLR-based method.
【4】 Physics-informed convolutional neural network with bicubic spline interpolation for sound field estimation
标题:基于双三次样条插值卷积神经网络的声场估计
链接:https://arxiv.org/abs/2207.10937
作者:Kazuhide Shigemi,Shoichi Koyama,Tomohiko Nakamura,Hiroshi Saruwatari机构:Graduate School of Information Science and Technology, The University of Tokyo, -,-, Hongo, Bunkyo-ku, Tokyo ,-, Japan备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2022摘要:提出了一种基于样条插值的物理信息卷积神经网络(PICNN)的声场估计方法.现有的声场估计方法大多基于波函数展开,使估计函数满足Helmholtz方程,但这些方法仅依赖于物理性质,而基于样条插值的PICNN的声场估计方法存在着收敛速度慢、收敛速度慢等缺点.文中给出了一种基于样条插值的PICNN的声场估计方法.最近的基于神经网络的学习方法在训练数据可用的情况下从稀疏测量进行估计方面具有优势,但是,由于没有考虑物理特性,估计的函数可以是物理上不可行的解。本文提出了一种利用损失函数补偿Helmholtz方程偏差的PICNN算法来估计声场,在PICNN的框架下引入双三次样条插值,实验结果表明,该方法可以从稀疏的观测数据中获得精确的物理上可行的估计.摘要:A sound field estimation method based on a physics-informed convolutional neural network (PICNN) using spline interpolation is proposed. Most of the sound field estimation methods are based on wavefunction expansion, making the estimated function satisfy the Helmholtz equation. However, these methods rely only on physical properties; thus, they suffer from a significant deterioration of accuracy when the number of measurements is small. Recent learning-based methods based on neural networks have advantages in estimating from sparse measurements when training data are available. However, since physical properties are not taken into consideration, the estimated function can be a physically infeasible solution. We propose the application of PICNN to the sound field estimation problem by using a loss function that penalizes deviation from the Helmholtz equation. Since the output of CNN is a spatially discretized pressure distribution, it is difficult to directly evaluate the Helmholtz-equation loss function. Therefore, we incorporate bicubic spline interpolation in the PICNN framework. Experimental results indicated that accurate and physically feasible estimation from sparse measurements can be achieved with the proposed method.
【5】 ASR Error Detection via Audio-Transcript entailment
标题:通过音频转录本限定的ASR错误检测
链接:https://arxiv.org/abs/2207.10849
作者:Nimshi Venkat Meripo,Sandeep Konam备注:Accepted to Interspeech 2022摘要:尽管最新的自动语音识别性能有所提高(ASR)系统中,转录错误仍然是不可避免的。这些错误在医疗保健等关键领域中用于帮助临床文档时,会产生相当大的影响。因此,检测ASR错误是防止错误进一步传播到下游应用程序的关键第一步。为此,我们提出了一种新的基于音频转录本蕴涵的端到端ASR错误检测方法,我们是第一个将该问题构造为音频段和其相应的抄本段之间的端到端蕴涵任务的人。我们的直觉是,当没有识别错误时,音频和文本之间应该存在双向蕴涵,反之亦然,所提出的模型利用声学编码器和语言编码器分别对语音和文本建模。由于实验中使用的是医生和病人之间的对话,因此我们的实验重点放在了医学术语上,我们提出的模型达到了分类错误率(CER)中,26.2%针对所有转录错误,23%针对医疗错误,分别比强基线改善了12%和15.4%。摘要:Despite improved performances of the latest Automatic Speech Recognition (ASR) systems, transcription errors are still unavoidable. These errors can have a considerable impact in critical domains such as healthcare, when used to help with clinical documentation. Therefore, detecting ASR errors is a critical first step in preventing further error propagation to downstream applications. To this end, we propose a novel end-to-end approach for ASR error detection using audio-transcript entailment. To the best of our knowledge, we are the first to frame this problem as an end-to-end entailment task between the audio segment and its corresponding transcript segment. Our intuition is that there should be a bidirectional entailment between audio and transcript when there is no recognition error and vice versa. The proposed model utilizes an acoustic encoder and a linguistic encoder to model the speech and transcript respectively. The encoded representations of both modalities are fused to predict the entailment. Since doctor-patient conversations are used in our experiments, a particular emphasis is placed on medical terms. Our proposed model achieves classification error rates (CER) of 26.2% on all transcription errors and 23% on medical errors specifically, leading to improvements upon a strong baseline by 12% and 15.4%, respectively.
【6】 End-to-End and Self-Supervised Learning for ComParE 2022 Stuttering Sub-Challenge
标题:ComParE 2022口吃子挑战赛端到端和自我监督学习
链接:https://arxiv.org/abs/2207.10817
作者:Shakeel Ahmad Sheikh,Md Sahidullah,Fabrice Hirsch,Slim Ouni机构:Universite de Lorraine, CNRS, Inria, LORIA, F-, Nancy, France, Universite Paul-Valery Montpellier, CNRS, Praxiling, Montpellier, France备注:Accepted in ACM MM 2022 Conference : Grand Challenges, "\c{opyright} {Owner/Author | ACM} {2022}. This is the author's version of the work. It is posted here for your personal use. Not for redistribution摘要:在这篇文章中,我们提出了以自监督方式训练的基于端到端和语音嵌入的系统,以参加ACM多媒体2022 ComParE挑战赛,特别是口吃子挑战赛,特别是,我们利用预训练的Wav2Vec2.0模型的嵌入来检测口吃(SD)在KSoF数据集上。在嵌入提取之后,我们对几种SD方法进行了测试。我们提出的基于自监督的SD系统达到了36.9%和41.0%的UAR。在验证集和测试集上分别为31.32%。(验证集)和1.49%(测试集)高于最佳(DeepSpectrum)挑战基线此外,我们证明了具有Mel-频率倒谱系数的级联层嵌入(MFCC)特征进一步提高了验证集和测试集上的UAR,我们证明了Wav2Vec2.0的所有层的信息总和超过CBL的相对幅度为45.91%和5.69%分别在验证集和测试集上。大挑战:计算副语言学挑战摘要:In this paper, we present end-to-end and speech embedding based systems trained in a self-supervised fashion to participate in the ACM Multimedia 2022 ComParE Challenge, specifically the stuttering sub-challenge. In particular, we exploit the embeddings from the pre-trained Wav2Vec2.0 model for stuttering detection (SD) on the KSoF dataset. After embedding extraction, we benchmark with several methods for SD. Our proposed self-supervised based SD system achieves a UAR of 36.9% and 41.0% on validation and test sets respectively, which is 31.32% (validation set) and 1.49% (test set) higher than the best (DeepSpectrum) challenge baseline (CBL). Moreover, we show that concatenating layer embeddings with Mel-frequency cepstral coefficients (MFCCs) features further improves the UAR of 33.81% and 5.45% on validation and test sets respectively over the CBL. Finally, we demonstrate that the summing information across all the layers of Wav2Vec2.0 surpasses the CBL by a relative margin of 45.91% and 5.69% on validation and test sets respectively. Grand-challenge: Computational Paralinguistics ChallengE
【7】 Smart speaker design and implementation with biometric authentication and advanced voice interaction capability
标题:具有生物认证和高级语音交互功能的智能扬声器设计与实现
链接:https://arxiv.org/abs/2207.10811
作者:Bharath Sudharsan,Peter Corcoran,Muhammad Intizar Ali机构:Data Science Institute, National University of Ireland Galway.摘要:半导体技术的进步已经在提高芯片组的性能和容量的同时减小了尺寸和成本,人工智能框架和库的进步为在资源受限的消费类物联网设备边缘容纳更多人工智能提供了可能性。传感器如今已成为我们环境中不可或缺的一部分,可提供持续的数据流以构建智能应用。一个示例是具有多个互连设备的智能家庭场景。在这样的智能环境中,为了方便快速地访问基于网络的服务和个人信息,例如日历、笔记、电子邮件、提醒、银行等,用户将第三方技能或来自亚马逊商店的技能链接到他们的智能扬声器。此外,在当前的智能家庭场景中,可以使用多种智能家庭产品,例如智能安全摄像机、视频门铃、智能插头、智能一氧化碳监视器和智能门锁,因为智能扬声器通过智能扬声器用户的账户链接到这样的服务和设备。任何人都可以通过语音命令使用智能扬声器。如果这样做,用户的数据隐私,家庭安全和其他方面会受到损害。最近推出的,Tensor Cam的AI相机,东芝的Symbio,Facebook的Portal是具有人工智能功能的带摄像头的智能音箱。虽然它们是带摄像头的,然而它们除了呼叫唤醒字之外没有认证方案。本文概述了智能音箱用户由于缺乏认证方案而面临的网络安全风险,并讨论了为解决这些风险而开发的最先进的支持摄像头、基于麦克风阵列的现代Alexa智能音箱原型。摘要:Advancements in semiconductor technology have reduced dimensions and cost while improving the performance and capacity of chipsets. In addition, advancement in the AI frameworks and libraries brings possibilities to accommodate more AI at the resource-constrained edge of consumer IoT devices. Sensors are nowadays an integral part of our environment which provide continuous data streams to build intelligent applications. An example could be a smart home scenario with multiple interconnected devices. In such smart environments, for convenience and quick access to web-based service and personal information such as calendars, notes, emails, reminders, banking, etc, users link third-party skills or skills from the Amazon store to their smart speakers. Also, in current smart home scenarios, several smart home products such as smart security cameras, video doorbells, smart plugs, smart carbon monoxide monitors, and smart door locks, etc. are interlinked to a modern smart speaker via means of custom skill addition. Since smart speakers are linked to such services and devices via the smart speaker user's account. They can be used by anyone with physical access to the smart speaker via voice commands. If done so, the data privacy, home security and other aspects of the user get compromised. Recently launched, Tensor Cam's AI Camera, Toshiba's Symbio, Facebook's Portal are camera-enabled smart speakers with AI functionalities. Although they are camera-enabled, yet they do not have an authentication scheme in addition to calling out the wake-word. This paper provides an overview of cybersecurity risks faced by smart speaker users due to lack of authentication scheme and discusses the development of a state-of-the-art camera-enabled, microphone array-based modern Alexa smart speaker prototype to address these risks.
【8】 A Proposal for Foley Sound Synthesis Challenge
标题:Foley声音合成挑战赛提案
链接:https://arxiv.org/abs/2207.10760
作者:Keunwoo Choi,Sangshin Oh,Minsung Kang,Brian McFee机构:Gaudio Lab, Inc., Seoul, South Korea, New York University, New York, USA摘要:“Foley”指的是在后期制作期间添加到多媒体以增强其感知的声学特性的声音效果,例如通过模拟脚步声、周围环境声音或屏幕上的可见物体。虽然Foley传统上由Foley艺术家制作,对建立在声音合成和生成模型的最新进展上的自动或机器辅助技术的兴趣日益增加。为了促进更多人参与这一不断发展的研究领域,我们提出了一个自动Foley合成的挑战。通过对音频和机器学习领域先前成功挑战的案例研究,我们设定了所提出挑战的目标:严格、统一、高效地评估不同的Foley合成系统,总体目标是吸引研究团体的积极参与。我们概述了Foley声音合成挑战的细节和设计考虑,包括任务定义、数据集要求和评估标准。摘要:"Foley" refers to sound effects that are added to multimedia during post-production to enhance its perceived acoustic properties, e.g., by simulating the sounds of footsteps, ambient environmental sounds, or visible objects on the screen. While foley is traditionally produced by foley artists, there is increasing interest in automatic or machine-assisted techniques building upon recent advances in sound synthesis and generative models. To foster more participation in this growing research area, we propose a challenge for automatic foley synthesis. Through case studies on successful previous challenges in audio and machine learning, we set the goals of the proposed challenge: rigorous, unified, and efficient evaluation of different foley synthesis systems, with an overarching goal of drawing active participation from the research community. We outline the details and design considerations of a foley sound synthesis challenge, including task definition, dataset requirements, and evaluation criteria.
【9】 DNN-Free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online FastMNMF
标题:基于块在线FastMNMF的帧在线波束形成的无DNN低时延自适应语音增强
链接:https://arxiv.org/abs/2207.10934
作者:Aditya Arie Nugraha,Kouhei Sekiguchi,Mathieu Fontaine,Yoshiaki Bando,Kazuyoshi Yoshii机构:Center for Advanced Intelligence Project (AIP), RIKEN, Japan, Graduate School of Informatics, Kyoto University, Japan, LTCI, T´el´ecom Paris, Institut Polytechnique de Paris, France, National Institute of Advanced Industrial Science and Technology (AIST), Japan摘要:本文介绍了一种实用的双处理语音增强系统,该系统采用环境敏感帧—在线波束形成技术(前端)借助无环境数据块在线源分离(后端)。若要使用最小方差无失真响应(MVDR)波束形成,可以训练深度神经网络,(DNN),它估计用于计算源协方差矩阵的时频掩码提出了一种基于反向传播的DNN运行时自适应算法来处理训练—测试条件不匹配的情况,可以尝试利用称为快速多信道非负矩阵分解的现有技术盲源分离方法来直接估计源协方差矩阵然而,在实践中,DNN和FastMNMF都不能以帧在线的方式更新,这是由于其计算量很大的迭代性质。我们的无DNN系统利用块在线FastMNMF给出的最新源谱图的后验估计来推导当前源协方差矩阵,用于帧在线波束形成。评估表明,我们的帧在线波束形成系统具有更好的性能。在线系统可以快速响应由干扰说话人移动引起的场景变化,并且在字错误率方面比现有的具有基于DNN的波束形成的块在线系统性能好5.0个点。摘要:This paper describes a practical dual-process speech enhancement system that adapts environment-sensitive frame-online beamforming (front-end) with help from environment-free block-online source separation (back-end). To use minimum variance distortionless response (MVDR) beamforming, one may train a deep neural network (DNN) that estimates time-frequency masks used for computing the covariance matrices of sources (speech and noise). Backpropagation-based run-time adaptation of the DNN was proposed for dealing with the mismatched training-test conditions. Instead, one may try to directly estimate the source covariance matrices with a state-of-the-art blind source separation method called fast multichannel non-negative matrix factorization (FastMNMF). In practice, however, neither the DNN nor the FastMNMF can be updated in a frame-online manner due to its computationally-expensive iterative nature. Our DNN-free system leverages the posteriors of the latest source spectrograms given by block-online FastMNMF to derive the current source covariance matrices for frame-online beamforming. The evaluation shows that our frame-online system can quickly respond to scene changes caused by interfering speaker movements and outperformed an existing block-online system with DNN-based beamforming by 5.0 points in terms of the word error rate.
【1】 DNN-Free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online FastMNMF
标题:基于块在线FastMNMF的帧在线波束形成的无DNN低时延自适应语音增强
链接:https://arxiv.org/abs/2207.10934
* 与cs.SD语音【9】为同一篇
作者:Aditya Arie Nugraha,Kouhei Sekiguchi,Mathieu Fontaine,Yoshiaki Bando,Kazuyoshi Yoshii机构:Center for Advanced Intelligence Project (AIP), RIKEN, Japan, Graduate School of Informatics, Kyoto University, Japan, LTCI, T´el´ecom Paris, Institut Polytechnique de Paris, France, National Institute of Advanced Industrial Science and Technology (AIST), Japan摘要:本文介绍了一种实用的双处理语音增强系统,该系统采用环境敏感帧—在线波束形成技术(前端)借助无环境数据块在线源分离(后端)。若要使用最小方差无失真响应(MVDR)波束形成,可以训练深度神经网络,(DNN),它估计用于计算源协方差矩阵的时频掩码提出了一种基于反向传播的DNN运行时自适应算法来处理训练—测试条件不匹配的情况,可以尝试利用称为快速多信道非负矩阵分解的现有技术盲源分离方法来直接估计源协方差矩阵然而,在实践中,DNN和FastMNMF都不能以帧在线的方式更新,这是由于其计算量很大的迭代性质。我们的无DNN系统利用块在线FastMNMF给出的最新源谱图的后验估计来推导当前源协方差矩阵,用于帧在线波束形成。评估表明,我们的帧在线波束形成系统具有更好的性能。在线系统可以快速响应由干扰说话人移动引起的场景变化,并且在字错误率方面比现有的具有基于DNN的波束形成的块在线系统性能好5.0个点。摘要:This paper describes a practical dual-process speech enhancement system that adapts environment-sensitive frame-online beamforming (front-end) with help from environment-free block-online source separation (back-end). To use minimum variance distortionless response (MVDR) beamforming, one may train a deep neural network (DNN) that estimates time-frequency masks used for computing the covariance matrices of sources (speech and noise). Backpropagation-based run-time adaptation of the DNN was proposed for dealing with the mismatched training-test conditions. Instead, one may try to directly estimate the source covariance matrices with a state-of-the-art blind source separation method called fast multichannel non-negative matrix factorization (FastMNMF). In practice, however, neither the DNN nor the FastMNMF can be updated in a frame-online manner due to its computationally-expensive iterative nature. Our DNN-free system leverages the posteriors of the latest source spectrograms given by block-online FastMNMF to derive the current source covariance matrices for frame-online beamforming. The evaluation shows that our frame-online system can quickly respond to scene changes caused by interfering speaker movements and outperformed an existing block-online system with DNN-based beamforming by 5.0 points in terms of the word error rate.
【2】 Learning Unsupervised Hierarchies of Audio Concepts
标题:学习音频概念的无监督层次
链接:https://arxiv.org/abs/2207.11231
* 与cs.SD语音【1】为同一篇
作者:Darius Afchar,Romain Hennequin,Vincent Guigue机构:Deezer Research, France, MLIA, Sorbonne Université, France摘要:音乐信号很难从它们的低级特征中解读出来,也许比图像更难:例如,突出显示频谱图或图像的一部分通常不足以传达与人类真正相关的高级思想。在计算机视觉中,提出了概念学习,以将解释调整到正确的抽象级别(例如,从射线照片中检测临床概念)。这些方法尚未用于MIR。在本文中,我们将概念学习应用于音乐领域,因为音乐领域有其特殊性,例如,音乐概念是典型的非独立性和混合性(例如,流派、乐器、情绪),不同于先前的假设解开概念的工作,我们提出一种从音频中学习大量音乐概念,然后自动将它们分层以揭示它们之间的相互关系的方法。我们在音乐流媒体服务的播放列表数据集上进行了实验,作为不同概念的一些注释示例。评估表明,挖掘的层次结构与概念的地面真实层次结构(如果可用)和一般情况下的概念相似性代理源都是一致的。摘要:Music signals are difficult to interpret from their low-level features, perhaps even more than images: e.g. highlighting part of a spectrogram or an image is often insufficient to convey high-level ideas that are genuinely relevant to humans. In computer vision, concept learning was therein proposed to adjust explanations to the right abstraction level (e.g. detect clinical concepts from radiographs). These methods have yet to be used for MIR. In this paper, we adapt concept learning to the realm of music, with its particularities. For instance, music concepts are typically non-independent and of mixed nature (e.g. genre, instruments, mood), unlike previous work that assumed disentangled concepts. We propose a method to learn numerous music concepts from audio and then automatically hierarchise them to expose their mutual relationships. We conduct experiments on datasets of playlists from a music streaming service, serving as a few annotated examples for diverse concepts. Evaluations show that the mined hierarchies are aligned with both ground-truth hierarchies of concepts -- when available -- and with proxy sources of concept similarity in the general case.
【3】 Inference skipping for more efficient real-time speech enhancement with parallel RNNs
标题:基于并行RNN的跳过推理实时语音增强
链接:https://arxiv.org/abs/2207.11108
* 与cs.SD语音【2】为同一篇
作者:Xiaohuai Le,Tong Lei,Kai Chen,Jing Lu机构:Institute of Acoustics, Nanjing University备注:11 pages, 8 figures, accepted by IEEE/ACM TASLP摘要:深度神经网络基于DNN的语音增强模型由于其良好的性能而受到广泛的关注.然而,由于DNN的计算量大,难以在实时应用中部署一个功能强大的DNN.典型的压缩方法如剪枝和量化等都没有很好地利用数据的特性.本文提出了一种基于DNN的语音增强模型,将Skip-RNN策略引入到并行RNN语音增强模型中,RNN的状态在不中断输出掩码更新的情况下间歇性地更新,这导致计算负荷的显著减少而没有明显的音频伪像。为了更好地平衡语音和噪声之间的差异,我们进一步利用语音活动检测来调整跳过策略,(VAD)引导,节省了更多的计算量。实验采用了一种高性能的语音增强模型——双路卷积递归网络(DPCRN)的实验结果表明,与网络剪枝和直接训练一个较小的模型等策略相比,本文提出的算法具有更好的性能,同时也验证了本文提出的算法在其他两个竞争性语音增强模型上的推广性。摘要:Deep neural network (DNN) based speech enhancement models have attracted extensive attention due to their promising performance. However, it is difficult to deploy a powerful DNN in real-time applications because of its high computational cost. Typical compression methods such as pruning and quantization do not make good use of the data characteristics. In this paper, we introduce the Skip-RNN strategy into speech enhancement models with parallel RNNs. The states of the RNNs update intermittently without interrupting the update of the output mask, which leads to significant reduction of computational load without evident audio artifacts. To better leverage the difference between the voice and the noise, we further regularize the skipping strategy with voice activity detection (VAD) guidance, saving more computational load. Experiments on a high-performance speech enhancement model, dual-path convolutional recurrent network (DPCRN), show the superiority of our strategy over strategies like network pruning or directly training a smaller model. We also validate the generalization of the proposed strategy on two other competitive speech enhancement models.
【4】 Head-Related Transfer Function Interpolation from Spatially Sparse Measurements Using Autoencoder with Source Position Conditioning
标题:使用带源位置调节的自动编码器从空间稀疏测量值进行头部相关传递函数插值
链接:https://arxiv.org/abs/2207.10967
* 与cs.SD语音【3】为同一篇
作者:Yuki Ito,Tomohiko Nakamura,Shoichi Koyama,Hiroshi Saruwatari机构:Graduate School of Information Science and Technology, The University of Tokyo, -,-, Hongo, Bunkyo-ku, Tokyo ,-, Japan备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2022摘要:本文提出了一种头相关传递函数的方法本文提出的方法是基于正则化线性回归的HRTF插值方法与基于自编码器的HRTF插值方法的类比(RLR)和自动编码器。通过这种类比,我们发现基于RLR的方法的主要特点是将HRTF分解为与源位置相关的和与源位置无关的两个因子,我们设计编码器和解码器使得它们的权值和偏置是从源位置产生的,我们引入了一个聚集模块,减少了隐变量对源位置的依赖性,从而获得了与源位置无关的每个对象的表示.数值实验表明,所提出的方法可以很好地适用于不可见对象,并且只需要一个源位置就可以获得插值性能.第八测量值与基于RLR的方法的测量值相当。摘要:We propose a method of head-related transfer function (HRTF) interpolation from sparsely measured HRTFs using an autoencoder with source position conditioning. The proposed method is drawn from an analogy between an HRTF interpolation method based on regularized linear regression (RLR) and an autoencoder. Through this analogy, we found the key feature of the RLR-based method that HRTFs are decomposed into source-position-dependent and source-position-independent factors. On the basis of this finding, we design the encoder and decoder so that their weights and biases are generated from source positions. Furthermore, we introduce an aggregation module that reduces the dependence of latent variables on source position for obtaining a source-position-independent representation of each subject. Numerical experiments show that the proposed method can work well for unseen subjects and achieve an interpolation performance with only one-eighth measurements comparable to that of the RLR-based method.
【5】 Physics-informed convolutional neural network with bicubic spline interpolation for sound field estimation
标题:基于双三次样条插值卷积神经网络的声场估计
链接:https://arxiv.org/abs/2207.10937
* 与cs.SD语音【4】为同一篇
作者:Kazuhide Shigemi,Shoichi Koyama,Tomohiko Nakamura,Hiroshi Saruwatari机构:Graduate School of Information Science and Technology, The University of Tokyo, -,-, Hongo, Bunkyo-ku, Tokyo ,-, Japan备注:Accepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2022摘要:提出了一种基于样条插值的物理信息卷积神经网络(PICNN)的声场估计方法.现有的声场估计方法大多基于波函数展开,使估计函数满足Helmholtz方程,但这些方法仅依赖于物理性质,而基于样条插值的PICNN的声场估计方法存在着收敛速度慢、收敛速度慢等缺点.文中给出了一种基于样条插值的PICNN的声场估计方法.最近的基于神经网络的学习方法在训练数据可用的情况下从稀疏测量进行估计方面具有优势,但是,由于没有考虑物理特性,估计的函数可以是物理上不可行的解。本文提出了一种利用损失函数补偿Helmholtz方程偏差的PICNN算法来估计声场,在PICNN的框架下引入双三次样条插值,实验结果表明,该方法可以从稀疏的观测数据中获得精确的物理上可行的估计.摘要:A sound field estimation method based on a physics-informed convolutional neural network (PICNN) using spline interpolation is proposed. Most of the sound field estimation methods are based on wavefunction expansion, making the estimated function satisfy the Helmholtz equation. However, these methods rely only on physical properties; thus, they suffer from a significant deterioration of accuracy when the number of measurements is small. Recent learning-based methods based on neural networks have advantages in estimating from sparse measurements when training data are available. However, since physical properties are not taken into consideration, the estimated function can be a physically infeasible solution. We propose the application of PICNN to the sound field estimation problem by using a loss function that penalizes deviation from the Helmholtz equation. Since the output of CNN is a spatially discretized pressure distribution, it is difficult to directly evaluate the Helmholtz-equation loss function. Therefore, we incorporate bicubic spline interpolation in the PICNN framework. Experimental results indicated that accurate and physically feasible estimation from sparse measurements can be achieved with the proposed method.
【6】 ASR Error Detection via Audio-Transcript entailment
标题:通过音频转录本限定的ASR错误检测
链接:https://arxiv.org/abs/2207.10849
* 与cs.SD语音【5】为同一篇
作者:Nimshi Venkat Meripo,Sandeep Konam备注:Accepted to Interspeech 2022摘要:尽管最新的自动语音识别性能有所提高(ASR)系统中,转录错误仍然是不可避免的。这些错误在医疗保健等关键领域中用于帮助临床文档时,会产生相当大的影响。因此,检测ASR错误是防止错误进一步传播到下游应用程序的关键第一步。为此,我们提出了一种新的基于音频转录本蕴涵的端到端ASR错误检测方法,我们是第一个将该问题构造为音频段和其相应的抄本段之间的端到端蕴涵任务的人。我们的直觉是,当没有识别错误时,音频和文本之间应该存在双向蕴涵,反之亦然,所提出的模型利用声学编码器和语言编码器分别对语音和文本建模。由于实验中使用的是医生和病人之间的对话,因此我们的实验重点放在了医学术语上,我们提出的模型达到了分类错误率(CER)中,26.2%针对所有转录错误,23%针对医疗错误,分别比强基线改善了12%和15.4%。摘要:Despite improved performances of the latest Automatic Speech Recognition (ASR) systems, transcription errors are still unavoidable. These errors can have a considerable impact in critical domains such as healthcare, when used to help with clinical documentation. Therefore, detecting ASR errors is a critical first step in preventing further error propagation to downstream applications. To this end, we propose a novel end-to-end approach for ASR error detection using audio-transcript entailment. To the best of our knowledge, we are the first to frame this problem as an end-to-end entailment task between the audio segment and its corresponding transcript segment. Our intuition is that there should be a bidirectional entailment between audio and transcript when there is no recognition error and vice versa. The proposed model utilizes an acoustic encoder and a linguistic encoder to model the speech and transcript respectively. The encoded representations of both modalities are fused to predict the entailment. Since doctor-patient conversations are used in our experiments, a particular emphasis is placed on medical terms. Our proposed model achieves classification error rates (CER) of 26.2% on all transcription errors and 23% on medical errors specifically, leading to improvements upon a strong baseline by 12% and 15.4%, respectively.
【7】 End-to-End and Self-Supervised Learning for ComParE 2022 Stuttering Sub-Challenge
标题:ComParE 2022口吃子挑战赛端到端和自我监督学习
链接:https://arxiv.org/abs/2207.10817
* 与cs.SD语音【6】为同一篇
作者:Shakeel Ahmad Sheikh,Md Sahidullah,Fabrice Hirsch,Slim Ouni机构:Universite de Lorraine, CNRS, Inria, LORIA, F-, Nancy, France, Universite Paul-Valery Montpellier, CNRS, Praxiling, Montpellier, France备注:Accepted in ACM MM 2022 Conference : Grand Challenges, "\c{opyright} {Owner/Author | ACM} {2022}. This is the author's version of the work. It is posted here for your personal use. Not for redistribution摘要:在这篇文章中,我们提出了以自监督方式训练的基于端到端和语音嵌入的系统,以参加ACM多媒体2022 ComParE挑战赛,特别是口吃子挑战赛,特别是,我们利用预训练的Wav2Vec2.0模型的嵌入来检测口吃(SD)在KSoF数据集上。在嵌入提取之后,我们对几种SD方法进行了测试。我们提出的基于自监督的SD系统达到了36.9%和41.0%的UAR。在验证集和测试集上分别为31.32%。(验证集)和1.49%(测试集)高于最佳(DeepSpectrum)挑战基线此外,我们证明了具有Mel-频率倒谱系数的级联层嵌入(MFCC)特征进一步提高了验证集和测试集上的UAR,我们证明了Wav2Vec2.0的所有层的信息总和超过CBL的相对幅度为45.91%和5.69%分别在验证集和测试集上。大挑战:计算副语言学挑战摘要:In this paper, we present end-to-end and speech embedding based systems trained in a self-supervised fashion to participate in the ACM Multimedia 2022 ComParE Challenge, specifically the stuttering sub-challenge. In particular, we exploit the embeddings from the pre-trained Wav2Vec2.0 model for stuttering detection (SD) on the KSoF dataset. After embedding extraction, we benchmark with several methods for SD. Our proposed self-supervised based SD system achieves a UAR of 36.9% and 41.0% on validation and test sets respectively, which is 31.32% (validation set) and 1.49% (test set) higher than the best (DeepSpectrum) challenge baseline (CBL). Moreover, we show that concatenating layer embeddings with Mel-frequency cepstral coefficients (MFCCs) features further improves the UAR of 33.81% and 5.45% on validation and test sets respectively over the CBL. Finally, we demonstrate that the summing information across all the layers of Wav2Vec2.0 surpasses the CBL by a relative margin of 45.91% and 5.69% on validation and test sets respectively. Grand-challenge: Computational Paralinguistics ChallengE
【8】 Smart speaker design and implementation with biometric authentication and advanced voice interaction capability
标题:具有生物认证和高级语音交互功能的智能扬声器设计与实现
链接:https://arxiv.org/abs/2207.10811
* 与cs.SD语音【7】为同一篇
作者:Bharath Sudharsan,Peter Corcoran,Muhammad Intizar Ali机构:Data Science Institute, National University of Ireland Galway.摘要:半导体技术的进步已经在提高芯片组的性能和容量的同时减小了尺寸和成本,人工智能框架和库的进步为在资源受限的消费类物联网设备边缘容纳更多人工智能提供了可能性。传感器如今已成为我们环境中不可或缺的一部分,可提供持续的数据流以构建智能应用。一个示例是具有多个互连设备的智能家庭场景。在这样的智能环境中,为了方便快速地访问基于网络的服务和个人信息,例如日历、笔记、电子邮件、提醒、银行等,用户将第三方技能或来自亚马逊商店的技能链接到他们的智能扬声器。此外,在当前的智能家庭场景中,可以使用多种智能家庭产品,例如智能安全摄像机、视频门铃、智能插头、智能一氧化碳监视器和智能门锁,因为智能扬声器通过智能扬声器用户的账户链接到这样的服务和设备。任何人都可以通过语音命令使用智能扬声器。如果这样做,用户的数据隐私,家庭安全和其他方面会受到损害。最近推出的,Tensor Cam的AI相机,东芝的Symbio,Facebook的Portal是具有人工智能功能的带摄像头的智能音箱。虽然它们是带摄像头的,然而它们除了呼叫唤醒字之外没有认证方案。本文概述了智能音箱用户由于缺乏认证方案而面临的网络安全风险,并讨论了为解决这些风险而开发的最先进的支持摄像头、基于麦克风阵列的现代Alexa智能音箱原型。摘要:Advancements in semiconductor technology have reduced dimensions and cost while improving the performance and capacity of chipsets. In addition, advancement in the AI frameworks and libraries brings possibilities to accommodate more AI at the resource-constrained edge of consumer IoT devices. Sensors are nowadays an integral part of our environment which provide continuous data streams to build intelligent applications. An example could be a smart home scenario with multiple interconnected devices. In such smart environments, for convenience and quick access to web-based service and personal information such as calendars, notes, emails, reminders, banking, etc, users link third-party skills or skills from the Amazon store to their smart speakers. Also, in current smart home scenarios, several smart home products such as smart security cameras, video doorbells, smart plugs, smart carbon monoxide monitors, and smart door locks, etc. are interlinked to a modern smart speaker via means of custom skill addition. Since smart speakers are linked to such services and devices via the smart speaker user's account. They can be used by anyone with physical access to the smart speaker via voice commands. If done so, the data privacy, home security and other aspects of the user get compromised. Recently launched, Tensor Cam's AI Camera, Toshiba's Symbio, Facebook's Portal are camera-enabled smart speakers with AI functionalities. Although they are camera-enabled, yet they do not have an authentication scheme in addition to calling out the wake-word. This paper provides an overview of cybersecurity risks faced by smart speaker users due to lack of authentication scheme and discusses the development of a state-of-the-art camera-enabled, microphone array-based modern Alexa smart speaker prototype to address these risks.
【9】 A Proposal for Foley Sound Synthesis Challenge
标题:Foley声音合成挑战赛提案
链接:https://arxiv.org/abs/2207.10760
* 与cs.SD语音【8】为同一篇
作者:Keunwoo Choi,Sangshin Oh,Minsung Kang,Brian McFee机构:Gaudio Lab, Inc., Seoul, South Korea, New York University, New York, USA摘要:“Foley”指的是在后期制作期间添加到多媒体以增强其感知的声学特性的声音效果,例如通过模拟脚步声、周围环境声音或屏幕上的可见物体。虽然Foley传统上由Foley艺术家制作,对建立在声音合成和生成模型的最新进展上的自动或机器辅助技术的兴趣日益增加。为了促进更多人参与这一不断发展的研究领域,我们提出了一个自动Foley合成的挑战。通过对音频和机器学习领域先前成功挑战的案例研究,我们设定了所提出挑战的目标:严格、统一、高效地评估不同的Foley合成系统,总体目标是吸引研究团体的积极参与。我们概述了Foley声音合成挑战的细节和设计考虑,包括任务定义、数据集要求和评估标准。摘要:"Foley" refers to sound effects that are added to multimedia during post-production to enhance its perceived acoustic properties, e.g., by simulating the sounds of footsteps, ambient environmental sounds, or visible objects on the screen. While foley is traditionally produced by foley artists, there is increasing interest in automatic or machine-assisted techniques building upon recent advances in sound synthesis and generative models. To foster more participation in this growing research area, we propose a challenge for automatic foley synthesis. Through case studies on successful previous challenges in audio and machine learning, we set the goals of the proposed challenge: rigorous, unified, and efficient evaluation of different foley synthesis systems, with an overarching goal of drawing active participation from the research community. We outline the details and design considerations of a foley sound synthesis challenge, including task definition, dataset requirements, and evaluation criteria.
机器翻译,仅供参考