今天跟大家分享一篇语音相关的论文合集:cs.SD语音1篇,eess.AS音频处理3篇。本文经arXiv每日学术速递授权转载
【1】 Few-Shot Audio-Visual Learning of Environment Acoustics
标题:环境声学的Few-Shot视听学习
链接:https://arxiv.org/abs/2206.04006
作者:Sagnik Majumder,Changan Chen,Ziad Al-Halah,Kristen Grauman摘要:房间脉冲响应(RIR)功能捕捉周围物理环境如何改变听者听到的声音,并对AR、VR和机器人技术中的各种应用产生影响。传统的估计RIR的方法假设在整个环境中有密集的几何和/或声音测量,而我们探索如何基于在空间中观察到的稀疏图像集和回波推断RIR。为了实现这一目标,我们引入了一种基于转换器的方法,该方法使用自我注意构建丰富的声学上下文,然后通过交叉注意预测任意查询源-接收器位置的RIR。此外,我们还设计了一个新的训练目标,以改进RIR预测和目标之间的声学特征匹配。在使用最先进的3D环境视听模拟器进行的实验中,我们证明了我们的方法成功地生成任意RIR,优于最先进的方法,并且与传统方法有很大的不同,我们还以少量镜头的方式将其推广到新环境。项目:http://vision.cs.utexas.edu/projects/fs_rir.摘要:Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughout the environment, we explore how to infer RIRs based on a sparse set of images and echoes observed in the space. Towards that goal, we introduce a transformer-based method that uses self-attention to build a rich acoustic context, then predicts RIRs of arbitrary query source-receiver locations through cross-attention. Additionally, we design a novel training objective that improves the match in the acoustic signature between the RIR predictions and the targets. In experiments using a state-of-the-art audio-visual simulator for 3D environments, we demonstrate that our method successfully generates arbitrary RIRs, outperforming state-of-the-art methods and--in a major departure from traditional methods--generalizing to novel environments in a few-shot manner. Project: http://vision.cs.utexas.edu/projects/fs_rir.
【1】 On the Integration of Acoustics and LiDAR: a Multi-Modal Approach to Acoustic Reflector Estimation
标题:声学与激光雷达的融合:一种多模声反射面估计方法
链接:https://arxiv.org/abs/2206.03885
作者:Ellen Riemens,Pablo Martínez-Nuevo,Jorge Martinez,Martin Møller,Richard C. Hendriks备注:5 pages, 9 figures, to be published in EUSIPCO 2022摘要:了解房间声学特性,例如声学反射器的位置,可以更好地再现预期的声场。当前使用麦克风测量进行房间边界检测的最先进方法通常侧重于二维设置,在实际场景中使用时会导致模型不匹配。三维任意反射器的检测遇到了实际限制,例如,需要球形阵列和增加的计算复杂性。此外,扬声器可能没有文献中通常假设的全向方向性模式,这使得在某些方向上检测声反射器更具挑战性。在该方法中,将激光雷达传感器添加到扬声器中,以提高墙壁检测精度和鲁棒性。这有两种方法。首先,水平反射器引入的模型失配可以通过使用激光雷达传感器检测反射器来解决,以消除预处理中2D问题对反射器的不利影响。其次,提出了一种基于激光雷达的方法来补偿定向扬声器发射的能量很小的挑战方向。我们通过仿真表明,这种多模态方法,即结合麦克风和激光雷达传感器,提高了墙壁检测的鲁棒性和准确性。摘要:Having knowledge on the room acoustic properties, e.g., the location of acoustic reflectors, allows to better reproduce the sound field as intended. Current state-of-the-art methods for room boundary detection using microphone measurements typically focus on a two-dimensional setting, causing a model mismatch when employed in real-life scenarios. Detection of arbitrary reflectors in three dimensions encounters practical limitations, e.g., the need for a spherical array and the increased computational complexity. Moreover, loudspeakers may not have an omnidirectional directivity pattern, as usually assumed in the literature, making the detection of acoustic reflectors in some directions more challenging. In the proposed method, a LiDAR sensor is added to a loudspeaker to improve wall detection accuracy and robustness. This is done in two ways. First, the model mismatch introduced by horizontal reflectors can be resolved by detecting reflectors with the LiDAR sensor to enable elimination of their detrimental influence from the 2D problem in pre-processing. Second, a LiDAR-based method is proposed to compensate for the challenging directions where the directive loudspeaker emits little energy. We show via simulations that this multi-modal approach, i.e., combining microphone and LiDAR sensors, improves the robustness and accuracy of wall detection.
【2】 Low-complexity acoustic scene classification in DCASE 2022 Challenge
标题:DCASE 2022挑战赛中低复杂度的声学场景分类
链接:https://arxiv.org/abs/2206.03835
作者:Irene Martín-Morató,Francesco Paissan,Alberto Ancilotto,Toni Heittola,Annamaria Mesaros,Elisabetta Farella,Alessio Brutti,Tuomas Virtanen摘要:本文分析了DCASE 2022挑战中低复杂度声学场景分类任务的结果。这项任务是前几年的延续。在此版本中,对低复杂性解决方案的要求进行了修改,包括:对参数数量(包括零值参数)的限制为128 K,强制采用INT8数字格式,以及在推理时限制3000万次乘法累加操作。所提供的基线系统是一个卷积神经网络,它采用参数的训练后量化,产生46512个参数,以及2923万个乘法和累加运算,分别在128K和3000万的设定限制下。基线系统对由9个不同设备的音频组成的开发数据的准确率为42.9%,对数损失为1.575。将在质询截止日期后对提交的系统进行分析。摘要:This paper analyzes the outcome of the Low-Complexity Acoustic Scene Classification task in DCASE 2022 Challenge. The task is a continuation from the previous years. In this edition, the requirement for low-complexity solutions were modified including: a limit of 128 K on the number of parameters, including the zero-valued ones, imposed INT8 numerical format, and a limit of 30 million multiply-accumulate operations at inference time. The provided baseline system is a convolutional neural network which employs post-training quantization of parameters, resulting in 46512 parameters, and 29.23 million multiply-and-accumulate operations, well under the set limits of 128K and 30 million, respectively. The baseline system has a 42.9% accuracy and a log-loss of 1.575 on the development data consisting of audio from 9 different devices. An analysis of the submitted systems will be provided after the challenge deadline.
【3】 Few-Shot Audio-Visual Learning of Environment Acoustics
标题:环境声学的Few-Shot视听学习
链接:https://arxiv.org/abs/2206.04006
作者:Sagnik Majumder,Changan Chen,Ziad Al-Halah,Kristen Grauman摘要:房间脉冲响应(RIR)功能捕捉周围物理环境如何改变听者听到的声音,并对AR、VR和机器人技术中的各种应用产生影响。传统的估计RIR的方法假设在整个环境中有密集的几何和/或声音测量,而我们探索如何基于在空间中观察到的稀疏图像集和回波推断RIR。为了实现这一目标,我们引入了一种基于转换器的方法,该方法使用自我注意构建丰富的声学上下文,然后通过交叉注意预测任意查询源-接收器位置的RIR。此外,我们还设计了一个新的训练目标,以改进RIR预测和目标之间的声学特征匹配。在使用最先进的3D环境视听模拟器进行的实验中,我们证明了我们的方法成功地生成任意RIR,优于最先进的方法,并且与传统方法有很大的不同,我们还以少量镜头的方式将其推广到新环境。项目:http://vision.cs.utexas.edu/projects/fs_rir.摘要:Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughout the environment, we explore how to infer RIRs based on a sparse set of images and echoes observed in the space. Towards that goal, we introduce a transformer-based method that uses self-attention to build a rich acoustic context, then predicts RIRs of arbitrary query source-receiver locations through cross-attention. Additionally, we design a novel training objective that improves the match in the acoustic signature between the RIR predictions and the targets. In experiments using a state-of-the-art audio-visual simulator for 3D environments, we demonstrate that our method successfully generates arbitrary RIRs, outperforming state-of-the-art methods and--in a major departure from traditional methods--generalizing to novel environments in a few-shot manner. Project: http://vision.cs.utexas.edu/projects/fs_rir.
机器翻译,仅供参考