今天跟大家分享一篇语音相关的论文合集:cs.SD语音7篇,eess.AS音频处理7篇。

cs.SD语音

【1】 Keyword Spotting System and Evaluation of Pruning and Quantization  Methods on Low-power Edge Microcontrollers

标题:低功耗边缘微控制器关键词识别系统及剪枝和量化方法评价

链接:https://arxiv.org/abs/2208.02765

作者:Jingyi Wang,Shengchen Li
备注:Submitted to DCASE2022 Workshop. Code available at: this https URL
摘要:关键字定位(KWS)有利于基于语音的用户与边缘的低功耗设备交互。边缘设备通常是始终在线的,因此边缘计算可节省带宽并保护隐私。这些设备通常具有有限的存储空间、计算性能、功耗和成本,例如,基于Cortex-M的微控制器。这些设备面临的挑战是满足深度学习的高计算和低延迟要求。本文首先介绍了我们的小型KWS系统运行在STM32F7微控制器上,采用Cortex-M7内核@216MHz和512KB静态RAM。我们选择的卷积神经网络(CNN)架构简化了KWS的运算量,满足了边缘设备的限制,我们的基线系统每37ms生成一个分类结果,包括实时音频特征提取部分。本文进一步评估了在微控制器上不同的剪枝和量化方法,包括不同的稀疏度粒度、跳过零权重、权重优先的循环顺序和SIMD指令的实际性能,结果表明,对于微控制器,加速非结构化剪枝模型是相当大的挑战,并且结构化剪枝比非结构化剪枝更友好,实验结果也验证了该算法对量化和SIMD指令的性能改善。
摘要:Keyword spotting (KWS) is beneficial for voice-based user interactions with low-power devices at the edge. The edge devices are usually always-on, so edge computing brings bandwidth savings and privacy protection. The devices typically have limited memory spaces, computational performances, power and costs, for example, Cortex-M based microcontrollers. The challenge is to meet the high computation and low-latency requirements of deep learning on these devices. This paper firstly shows our small-footprint KWS system running on STM32F7 microcontroller with Cortex-M7 core @216MHz and 512KB static RAM. Our selected convolutional neural network (CNN) architecture has simplified number of operations for KWS to meet the constraint of edge devices. Our baseline system generates classification results for each 37ms including real-time audio feature extraction part. This paper further evaluates the actual performance for different pruning and quantization methods on microcontroller, including different granularity of sparsity, skipping zero weights, weight-prioritized loop order, and SIMD instruction. The result shows that for microcontrollers, there are considerable challenges for accelerate unstructured pruned models, and the structured pruning is more friendly than unstructured pruning. The result also verified that the performance improvement for quantization and SIMD instruction.


【2】 Impact Makes a Sound and Sound Makes an Impact: Sound Guides  Representations and Explorations

标题:影响产生声音,声音产生影响:声音引导表象与探索

链接:https://arxiv.org/abs/2208.02680

作者:Xufeng Zhao,Cornelius Weber,Muhammad Burhan Hafez,Stefan Wermter
备注:Accepted at IROS 2022
摘要:声音是现实世界中信息量最大、最丰富的形式之一,同时也是移动设备上小型廉价传感器无需接触即可感知的鲁棒形式。尽管深度学习能够从多个感官输入中提取信息,但在机器人动作的控制和学习中,声音的使用却很少。对于无监督强化学习,我们通过基于物理学的声音模拟来构建真实的机器人操作场景,并提出了内在声音好奇模块(Intrinsic Sound Curiosity Module ISCM为强化学习者提供反馈以学习鲁棒的表征并奖励更有效的探索行为.我们进行了在预训练期间启用声音而在适应期间禁用声音的实验,并证明了ISCM学习的表示优于仅视觉基线学习的表示,预先训练的策略应用于下游任务时可以加速学习过程。
摘要:Sound is one of the most informative and abundant modalities in the real world while being robust to sense without contacts by small and cheap sensors that can be placed on mobile devices. Although deep learning is capable of extracting information from multiple sensory inputs, there has been little use of sound for the control and learning of robotic actions. For unsupervised reinforcement learning, an agent is expected to actively collect experiences and jointly learn representations and policies in a self-supervised way. We build realistic robotic manipulation scenarios with physics-based sound simulation and propose the Intrinsic Sound Curiosity Module (ISCM). The ISCM provides feedback to a reinforcement learner to learn robust representations and to reward a more efficient exploration behavior. We perform experiments with sound enabled during pre-training and disabled during adaptation, and show that representations learned by ISCM outperform the ones by vision-only baselines and pre-trained policies can accelerate the learning process when applied to downstream tasks.


【3】 Tokyo Kion-On: Query-Based Generative Sonification of Atmospheric Data

标题:东京京安:基于查询的大气数据生成式发音

链接:https://arxiv.org/abs/2208.02494

作者:Stefano Kalonaris
备注:To appear in: Proceedings of the 27th International Conference on Auditory Display (ICAD 2022)摘要:在日益关注的环境问题中,数据的交互式显示成为探索和理解气候变化对地球生态系统完整性影响的重要工具。本文介绍了东京建恩,东京1876年至2021年气温的基于查询的发音模型。该系统使用被称为LSTM的递归神经网络结构,其注意力在日本旋律的小数据集上训练并以所述大气数据为条件。在描述了模型的实现之后,给出了音乐结果的简要比较说明,并讨论了暴露的超参数如何促进数据的主动和非线性探索。
摘要:Amid growing environmental concerns, interactive displays of data constitute an important tool for exploring and understanding the impact of climate change on the planet's ecosystemic integrity. This paper presents Tokyo kion-on, a query-based sonification model of Tokyo's air temperature from 1876 to 2021. The system uses a recurrent neural network architecture known as LSTM with attention trained on a small dataset of Japanese melodies and conditioned upon said atmospheric data. After describing the model's implementation, a brief comparative illustration of the musical results is presented, along with a discussion on how the exposed hyper-parameters can promote active and non-linear exploration of the data.


【4】 Estimating Visual Information From Audio Through Manifold Learning

标题:通过流形学习从音频估计视觉信息

链接:https://arxiv.org/abs/2208.02337

作者:Fabrizio Pedersoli,Dryden Wiebe,Amin Banitalebi,Yong Zhang,Kwang Moo Yi
摘要:提出了一种新的基于音频信号的视觉信息提取方法,它克服了基于视觉的视觉信息提取方法的一些局限性,即不需要“视线,”对遮挡和光照变化具有鲁棒性,并且可以在视觉/激光雷达传感器失效时作为后备,因此,基于音频的方法甚至对于其中仅对视觉信息感兴趣的应用也是有用的我们的框架基于流形学习并且包括两个步骤:首先,我们训练矢量量化变分自动编码器以学习我们感兴趣的特定视觉模态的数据流形。其次,我们训练一个音频变换网络来将多通道音频信号映射到对应的视觉样本的潜在表示.我们证明了我们的方法能够使用公开可用的音频/视觉数据集从音频产生有意义的图像.特别地,我们考虑从音频预测以下视觉模态:深度和语义分割。我们希望我们的研究结果能够促进从音频中提取视觉信息的进一步研究。代码可在以下网址获得:https://github.com/ubc-vision/audio_manifold.
摘要:We propose a new framework for extracting visual information about a scene only using audio signals. Audio-based methods can overcome some of the limitations of vision-based methods i.e., they do not require "line-of-sight", are robust to occlusions and changes in illumination, and can function as a backup in case vision/lidar sensors fail. Therefore, audio-based methods can be useful even for applications in which only visual information is of interest Our framework is based on Manifold Learning and consists of two steps. First, we train a Vector-Quantized Variational Auto-Encoder to learn the data manifold of the particular visual modality we are interested in. Second, we train an Audio Transformation network to map multi-channel audio signals to the latent representation of the corresponding visual sample. We show that our method is able to produce meaningful images from audio using a publicly available audio/visual dataset. In particular, we consider the prediction of the following visual modalities from audio: depth and semantic segmentation. We hope the findings of our work can facilitate further research in visual information extraction from audio. Code is available at: https://github.com/ubc-vision/audio_manifold.


【5】 Adversarial Attacks on ASR Systems: An Overview

标题:对ASR系统的敌对攻击:概述

链接:https://arxiv.org/abs/2208.02250

作者:Xiao Zhang,Hao Tan,Xuan Huang,Denghui Zhang,Keke Tang,Zhaoquan Gu
摘要:随着硬件和算法的发展,ASR(自动语音识别)系统进化了很多,随着模型变得越来越简单,开发和部署的难度变得越来越容易,ASR系统也越来越接近我们的生活,一方面我们经常使用ASR的APP或API来生成字幕和录制会议,另一方面,在过去的几年里,有许多针对ASR系统的对抗性攻击实例,通过在波形中加入一个小的扰动,本文首先介绍了ASR系统的发展、攻击的不同假设以及如何评估这些攻击,然后从两个攻击假设:与其他研究不同的是,本文重点研究了它们对ASR系统中哪一层波形的干扰,它们之间的相互关系,以及它们的实现方法,重点研究了它们的工作效果。
摘要:With the development of hardware and algorithms, ASR(Automatic Speech Recognition) systems evolve a lot. As The models get simpler, the difficulty of development and deployment become easier, ASR systems are getting closer to our life. On the one hand, we often use APPs or APIs of ASR to generate subtitles and record meetings. On the other hand, smart speaker and self-driving car rely on ASR systems to control AIoT devices. In past few years, there are a lot of works on adversarial examples attacks against ASR systems. By adding a small perturbation to the waveforms, the recognition results make a big difference. In this paper, we describe the development of ASR system, different assumptions of attacks, and how to evaluate these attacks. Next, we introduce the current works on adversarial examples attacks from two attack assumptions: white-box attack and black-box attack. Different from other surveys, we pay more attention to which layer they perturb waveforms in ASR system, the relationship between these attacks, and their implementation methods. We focus on the effect of their works.


【6】 Data-driven Attention and Data-independent DCT based Global Context  Modeling for Text-independent Speaker Recognition

标题:基于数据驱动注意力和数据无关DCT的全局上下文建模与文本无关说话人识别

链接:https://arxiv.org/abs/2208.02778

作者:Wei Xia,John H. L. Hansen
摘要:语音信号是高维、长、变长的序列,具有复杂的层次结构,信号在不同的时频上可能包含不同的信息,因此,学习有效的说话人表征对于说话人确认任务的可靠性至关重要(TF)位置。例如,可能更有益的是集中于音素类的高能量部分,例如摩擦音。标准卷积层对相邻局部区域的操作不能捕获复杂的TF全局上下文信息,提出了一种通用的全局时频上下文建模框架来利用上下文信息进行说话人表示建模,提出了一种基于数据驱动的注意力上下文模型来描述不同时频位置之间的长程和非局部关系,提出了一种基于2D-DCT的数据无关上下文模型,以提高模型的可解释性.提出了DCT注意机制以提高DCT基形式的建模能力,全局上下文信息被用于通过计算全局上下文和局部特征之间的相似性来重新校准显著的时间—频率位置。与标准的ResNet模型和Squeeze\&Excitation模块相比,本文提出的轻量化模块可以很容易地集成到说话人模型中,增加的计算量很小,有效地提高了说话人确认的性能。实验结果表明,本文提出得全局上下文建模框架能够有效地改善说话人表征得学习效果,实现通道和时间上得一致性.频率特征重新校准。
摘要:Learning an effective speaker representation is crucial for achieving reliable performance in speaker verification tasks. Speech signals are high-dimensional, long, and variable-length sequences that entail a complex hierarchical structure. Signals may contain diverse information at each time-frequency (TF) location. For example, it may be more beneficial to focus on high-energy parts for phoneme classes such as fricatives. The standard convolutional layer that operates on neighboring local regions cannot capture the complex TF global context information. In this study, a general global time-frequency context modeling framework is proposed to leverage the context information specifically for speaker representation modeling. First, a data-driven attention-based context model is introduced to capture the long-range and non-local relationship across different time-frequency locations. Second, a data-independent 2D-DCT based context model is proposed to improve model interpretability. A multi-DCT attention mechanism is presented to improve modeling power with alternate DCT base forms. Finally, the global context information is used to recalibrate salient time-frequency locations by computing the similarity between the global context and local features. The proposed lightweight blocks can be easily incorporated into a speaker model with little additional computational costs and effectively improves the speaker verification performance compared to the standard ResNet model and Squeeze\&Excitation block by a large margin. Detailed ablation studies are also performed to analyze various factors that may impact performance of the proposed individual modules. Results from experiments show that the proposed global context modeling framework can efficiently improve the learned speaker representations by achieving channel-wise and time-frequency feature recalibration.


【7】 Domestic Activity Clustering from Audio via Depthwise Separable  Convolutional Autoencoder Network

标题:基于深度可分离卷积自动编码网络的音频家庭活动聚类

链接:https://arxiv.org/abs/2208.02406

作者:Yanxiong Li,Wenchang Cao,Konstantinos Drossos,Tuomas Virtanen
备注:6 pages, 5 figures, 4 tables. Accepted by IEEE MMSP 2022
摘要:从音频中自动估计家庭活动可以用来解决许多问题,例如减少护理老年人的人力成本。本研究重点解决从音频中进行家庭活动聚类的问题。家庭活动聚类的目标是以无监督的方式将属于相同家庭活动类别的音频片段聚类到一个聚类中。本文,提出了一种基于深度可分卷积自编码器网络的家庭活动聚类方法,该方法通过深度可分卷积自编码器学习初始嵌入,并设计了一种面向聚类的损失算法,以联合优化嵌入精化和聚类分配(SINS数据集的衍生物),用于声场景和事件的探测和分类挑战我们的方法得到了归一化互信息(NMI)得分为54.46%,聚类准确率(CA)得分为63.64%,在NMI和CA方面优于现有方法。此外,该方法的计算复杂度和内存需求均低于已有的基于深度模型的方法.代码:https://github.com/vinceasvp/domestic-activity-clustering-from-audio
摘要:Automatic estimation of domestic activities from audio can be used to solve many problems, such as reducing the labor cost for nursing the elderly people. This study focuses on solving the problem of domestic activity clustering from audio. The target of domestic activity clustering is to cluster audio clips which belong to the same category of domestic activity into one cluster in an unsupervised way. In this paper, we propose a method of domestic activity clustering using a depthwise separable convolutional autoencoder network. In the proposed method, initial embeddings are learned by the depthwise separable convolutional autoencoder, and a clustering-oriented loss is designed to jointly optimize embedding refinement and cluster assignment. Different methods are evaluated on a public dataset (a derivative of the SINS dataset) used in the challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) in 2018. Our method obtains the normalized mutual information (NMI) score of 54.46%, and the clustering accuracy (CA) score of 63.64%, and outperforms state-of-the-art methods in terms of NMI and CA. In addition, both computational complexity and memory requirement of our method is lower than that of previous deep-model-based methods. Codes: https://github.com/vinceasvp/domestic-activity-clustering-from-audio


eess.AS音频处理

【1】 Data-driven Attention and Data-independent DCT based Global Context  Modeling for Text-independent Speaker Recognition

标题:基于数据驱动注意力和数据无关DCT的全局上下文建模与文本无关说话人识别

链接:https://arxiv.org/abs/2208.02778

* 与cs.SD语音【6】为同一篇

作者:Wei Xia,John H. L. Hansen
摘要:语音信号是高维、长、变长的序列,具有复杂的层次结构,信号在不同的时频上可能包含不同的信息,因此,学习有效的说话人表征对于说话人确认任务的可靠性至关重要(TF)位置。例如,可能更有益的是集中于音素类的高能量部分,例如摩擦音。标准卷积层对相邻局部区域的操作不能捕获复杂的TF全局上下文信息,提出了一种通用的全局时频上下文建模框架来利用上下文信息进行说话人表示建模,提出了一种基于数据驱动的注意力上下文模型来描述不同时频位置之间的长程和非局部关系,提出了一种基于2D-DCT的数据无关上下文模型,以提高模型的可解释性.提出了DCT注意机制以提高DCT基形式的建模能力,全局上下文信息被用于通过计算全局上下文和局部特征之间的相似性来重新校准显著的时间—频率位置。与标准的ResNet模型和Squeeze\&Excitation模块相比,本文提出的轻量化模块可以很容易地集成到说话人模型中,增加的计算量很小,有效地提高了说话人确认的性能。实验结果表明,本文提出得全局上下文建模框架能够有效地改善说话人表征得学习效果,实现通道和时间上得一致性.频率特征重新校准。
摘要:Learning an effective speaker representation is crucial for achieving reliable performance in speaker verification tasks. Speech signals are high-dimensional, long, and variable-length sequences that entail a complex hierarchical structure. Signals may contain diverse information at each time-frequency (TF) location. For example, it may be more beneficial to focus on high-energy parts for phoneme classes such as fricatives. The standard convolutional layer that operates on neighboring local regions cannot capture the complex TF global context information. In this study, a general global time-frequency context modeling framework is proposed to leverage the context information specifically for speaker representation modeling. First, a data-driven attention-based context model is introduced to capture the long-range and non-local relationship across different time-frequency locations. Second, a data-independent 2D-DCT based context model is proposed to improve model interpretability. A multi-DCT attention mechanism is presented to improve modeling power with alternate DCT base forms. Finally, the global context information is used to recalibrate salient time-frequency locations by computing the similarity between the global context and local features. The proposed lightweight blocks can be easily incorporated into a speaker model with little additional computational costs and effectively improves the speaker verification performance compared to the standard ResNet model and Squeeze\&Excitation block by a large margin. Detailed ablation studies are also performed to analyze various factors that may impact performance of the proposed individual modules. Results from experiments show that the proposed global context modeling framework can efficiently improve the learned speaker representations by achieving channel-wise and time-frequency feature recalibration.


【2】 Domestic Activity Clustering from Audio via Depthwise Separable  Convolutional Autoencoder Network

标题:基于深度可分离卷积自动编码网络的音频家庭活动聚类

链接:https://arxiv.org/abs/2208.02406

* 与cs.SD语音【7】为同一篇

作者:Yanxiong Li,Wenchang Cao,Konstantinos Drossos,Tuomas Virtanen
备注:6 pages, 5 figures, 4 tables. Accepted by IEEE MMSP 2022
摘要:从音频中自动估计家庭活动可以用来解决许多问题,例如减少护理老年人的人力成本。本研究重点解决从音频中进行家庭活动聚类的问题。家庭活动聚类的目标是以无监督的方式将属于相同家庭活动类别的音频片段聚类到一个聚类中。本文,提出了一种基于深度可分卷积自编码器网络的家庭活动聚类方法,该方法通过深度可分卷积自编码器学习初始嵌入,并设计了一种面向聚类的损失算法,以联合优化嵌入精化和聚类分配(SINS数据集的衍生物),用于声场景和事件的探测和分类挑战我们的方法得到了归一化互信息(NMI)得分为54.46%,聚类准确率(CA)得分为63.64%,在NMI和CA方面优于现有方法。此外,该方法的计算复杂度和内存需求均低于已有的基于深度模型的方法.代码:https://github.com/vinceasvp/domestic-activity-clustering-from-audio
摘要:Automatic estimation of domestic activities from audio can be used to solve many problems, such as reducing the labor cost for nursing the elderly people. This study focuses on solving the problem of domestic activity clustering from audio. The target of domestic activity clustering is to cluster audio clips which belong to the same category of domestic activity into one cluster in an unsupervised way. In this paper, we propose a method of domestic activity clustering using a depthwise separable convolutional autoencoder network. In the proposed method, initial embeddings are learned by the depthwise separable convolutional autoencoder, and a clustering-oriented loss is designed to jointly optimize embedding refinement and cluster assignment. Different methods are evaluated on a public dataset (a derivative of the SINS dataset) used in the challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) in 2018. Our method obtains the normalized mutual information (NMI) score of 54.46%, and the clustering accuracy (CA) score of 63.64%, and outperforms state-of-the-art methods in terms of NMI and CA. In addition, both computational complexity and memory requirement of our method is lower than that of previous deep-model-based methods. Codes: https://github.com/vinceasvp/domestic-activity-clustering-from-audio


【3】 Keyword Spotting System and Evaluation of Pruning and Quantization  Methods on Low-power Edge Microcontrollers

标题:低功耗边缘微控制器关键词识别系统及剪枝和量化方法评价

链接:https://arxiv.org/abs/2208.02765

* 与cs.SD语音【1】为同一篇

作者:Jingyi Wang,Shengchen Li
备注:Submitted to DCASE2022 Workshop. Code available at: this https URL
摘要:关键字定位(KWS)有利于基于语音的用户与边缘的低功耗设备交互。边缘设备通常是始终在线的,因此边缘计算可节省带宽并保护隐私。这些设备通常具有有限的存储空间、计算性能、功耗和成本,例如,基于Cortex-M的微控制器。这些设备面临的挑战是满足深度学习的高计算和低延迟要求。本文首先介绍了我们的小型KWS系统运行在STM32F7微控制器上,采用Cortex-M7内核@216MHz和512KB静态RAM。我们选择的卷积神经网络(CNN)架构简化了KWS的运算量,满足了边缘设备的限制,我们的基线系统每37ms生成一个分类结果,包括实时音频特征提取部分。本文进一步评估了在微控制器上不同的剪枝和量化方法,包括不同的稀疏度粒度、跳过零权重、权重优先的循环顺序和SIMD指令的实际性能,结果表明,对于微控制器,加速非结构化剪枝模型是相当大的挑战,并且结构化剪枝比非结构化剪枝更友好,实验结果也验证了该算法对量化和SIMD指令的性能改善。
摘要:Keyword spotting (KWS) is beneficial for voice-based user interactions with low-power devices at the edge. The edge devices are usually always-on, so edge computing brings bandwidth savings and privacy protection. The devices typically have limited memory spaces, computational performances, power and costs, for example, Cortex-M based microcontrollers. The challenge is to meet the high computation and low-latency requirements of deep learning on these devices. This paper firstly shows our small-footprint KWS system running on STM32F7 microcontroller with Cortex-M7 core @216MHz and 512KB static RAM. Our selected convolutional neural network (CNN) architecture has simplified number of operations for KWS to meet the constraint of edge devices. Our baseline system generates classification results for each 37ms including real-time audio feature extraction part. This paper further evaluates the actual performance for different pruning and quantization methods on microcontroller, including different granularity of sparsity, skipping zero weights, weight-prioritized loop order, and SIMD instruction. The result shows that for microcontrollers, there are considerable challenges for accelerate unstructured pruned models, and the structured pruning is more friendly than unstructured pruning. The result also verified that the performance improvement for quantization and SIMD instruction.


【4】 Impact Makes a Sound and Sound Makes an Impact: Sound Guides  Representations and Explorations

标题:影响产生声音,声音产生影响:声音引导表象与探索

链接:https://arxiv.org/abs/2208.02680

* 与cs.SD语音【2】为同一篇

作者:Xufeng Zhao,Cornelius Weber,Muhammad Burhan Hafez,Stefan Wermter
备注:Accepted at IROS 2022
摘要:声音是现实世界中信息量最大、最丰富的形式之一,同时也是移动设备上小型廉价传感器无需接触即可感知的鲁棒形式。尽管深度学习能够从多个感官输入中提取信息,但在机器人动作的控制和学习中,声音的使用却很少。对于无监督强化学习,我们通过基于物理学的声音模拟来构建真实的机器人操作场景,并提出了内在声音好奇模块(Intrinsic Sound Curiosity Module ISCM为强化学习者提供反馈以学习鲁棒的表征并奖励更有效的探索行为.我们进行了在预训练期间启用声音而在适应期间禁用声音的实验,并证明了ISCM学习的表示优于仅视觉基线学习的表示,预先训练的策略应用于下游任务时可以加速学习过程。
摘要:Sound is one of the most informative and abundant modalities in the real world while being robust to sense without contacts by small and cheap sensors that can be placed on mobile devices. Although deep learning is capable of extracting information from multiple sensory inputs, there has been little use of sound for the control and learning of robotic actions. For unsupervised reinforcement learning, an agent is expected to actively collect experiences and jointly learn representations and policies in a self-supervised way. We build realistic robotic manipulation scenarios with physics-based sound simulation and propose the Intrinsic Sound Curiosity Module (ISCM). The ISCM provides feedback to a reinforcement learner to learn robust representations and to reward a more efficient exploration behavior. We perform experiments with sound enabled during pre-training and disabled during adaptation, and show that representations learned by ISCM outperform the ones by vision-only baselines and pre-trained policies can accelerate the learning process when applied to downstream tasks.


【5】 Tokyo Kion-On: Query-Based Generative Sonification of Atmospheric Data

标题:东京京安:基于查询的大气数据生成式发音

链接:https://arxiv.org/abs/2208.02494

* 与cs.SD语音【3】为同一篇

作者:Stefano Kalonaris
备注:To appear in: Proceedings of the 27th International Conference on Auditory Display (ICAD 2022)摘要:在日益关注的环境问题中,数据的交互式显示成为探索和理解气候变化对地球生态系统完整性影响的重要工具。本文介绍了东京建恩,东京1876年至2021年气温的基于查询的发音模型。该系统使用被称为LSTM的递归神经网络结构,其注意力在日本旋律的小数据集上训练并以所述大气数据为条件。在描述了模型的实现之后,给出了音乐结果的简要比较说明,并讨论了暴露的超参数如何促进数据的主动和非线性探索。
摘要:Amid growing environmental concerns, interactive displays of data constitute an important tool for exploring and understanding the impact of climate change on the planet's ecosystemic integrity. This paper presents Tokyo kion-on, a query-based sonification model of Tokyo's air temperature from 1876 to 2021. The system uses a recurrent neural network architecture known as LSTM with attention trained on a small dataset of Japanese melodies and conditioned upon said atmospheric data. After describing the model's implementation, a brief comparative illustration of the musical results is presented, along with a discussion on how the exposed hyper-parameters can promote active and non-linear exploration of the data.


【6】 Estimating Visual Information From Audio Through Manifold Learning

标题:通过流形学习从音频估计视觉信息

链接:https://arxiv.org/abs/2208.02337

* 与cs.SD语音【4】为同一篇

作者:Fabrizio Pedersoli,Dryden Wiebe,Amin Banitalebi,Yong Zhang,Kwang Moo Yi
摘要:提出了一种新的基于音频信号的视觉信息提取方法,它克服了基于视觉的视觉信息提取方法的一些局限性,即不需要“视线,”对遮挡和光照变化具有鲁棒性,并且可以在视觉/激光雷达传感器失效时作为后备,因此,基于音频的方法甚至对于其中仅对视觉信息感兴趣的应用也是有用的我们的框架基于流形学习并且包括两个步骤:首先,我们训练矢量量化变分自动编码器以学习我们感兴趣的特定视觉模态的数据流形。其次,我们训练一个音频变换网络来将多通道音频信号映射到对应的视觉样本的潜在表示.我们证明了我们的方法能够使用公开可用的音频/视觉数据集从音频产生有意义的图像.特别地,我们考虑从音频预测以下视觉模态:深度和语义分割。我们希望我们的研究结果能够促进从音频中提取视觉信息的进一步研究。代码可在以下网址获得:https://github.com/ubc-vision/audio_manifold.
摘要:We propose a new framework for extracting visual information about a scene only using audio signals. Audio-based methods can overcome some of the limitations of vision-based methods i.e., they do not require "line-of-sight", are robust to occlusions and changes in illumination, and can function as a backup in case vision/lidar sensors fail. Therefore, audio-based methods can be useful even for applications in which only visual information is of interest Our framework is based on Manifold Learning and consists of two steps. First, we train a Vector-Quantized Variational Auto-Encoder to learn the data manifold of the particular visual modality we are interested in. Second, we train an Audio Transformation network to map multi-channel audio signals to the latent representation of the corresponding visual sample. We show that our method is able to produce meaningful images from audio using a publicly available audio/visual dataset. In particular, we consider the prediction of the following visual modalities from audio: depth and semantic segmentation. We hope the findings of our work can facilitate further research in visual information extraction from audio. Code is available at: https://github.com/ubc-vision/audio_manifold.


【7】 Adversarial Attacks on ASR Systems: An Overview

标题:对ASR系统的敌对攻击:概述

链接:https://arxiv.org/abs/2208.02250

* 与cs.SD语音【5】为同一篇

作者:Xiao Zhang,Hao Tan,Xuan Huang,Denghui Zhang,Keke Tang,Zhaoquan Gu
摘要:随着硬件和算法的发展,ASR(自动语音识别)系统进化了很多,随着模型变得越来越简单,开发和部署的难度变得越来越容易,ASR系统也越来越接近我们的生活,一方面我们经常使用ASR的APP或API来生成字幕和录制会议,另一方面,在过去的几年里,有许多针对ASR系统的对抗性攻击实例,通过在波形中加入一个小的扰动,本文首先介绍了ASR系统的发展、攻击的不同假设以及如何评估这些攻击,然后从两个攻击假设:与其他研究不同的是,本文重点研究了它们对ASR系统中哪一层波形的干扰,它们之间的相互关系,以及它们的实现方法,重点研究了它们的工作效果。
摘要:With the development of hardware and algorithms, ASR(Automatic Speech Recognition) systems evolve a lot. As The models get simpler, the difficulty of development and deployment become easier, ASR systems are getting closer to our life. On the one hand, we often use APPs or APIs of ASR to generate subtitles and record meetings. On the other hand, smart speaker and self-driving car rely on ASR systems to control AIoT devices. In past few years, there are a lot of works on adversarial examples attacks against ASR systems. By adding a small perturbation to the waveforms, the recognition results make a big difference. In this paper, we describe the development of ASR system, different assumptions of attacks, and how to evaluate these attacks. Next, we introduce the current works on adversarial examples attacks from two attack assumptions: white-box attack and black-box attack. Different from other surveys, we pay more attention to which layer they perturb waveforms in ASR system, the relationship between these attacks, and their implementation methods. We focus on the effect of their works.


机器翻译,仅供参考