今天跟大家分享一篇语音相关的论文合集:cs.SD语音7篇,eess.AS音频处理8篇。
cs.SD语音

【1】 STOP: A dataset for Spoken Task Oriented Semantic Parsing

标题:STOP:面向口语任务的语义分析数据集

链接:https://arxiv.org/abs/2207.10643

作者:Paden Tomasello,Po-Chun Hsu,Akshat Shrivastava,Daniel Lazar,Duc Le,Adithya Sagar,Ali Elkahky,Jade Copet,Wei-Ning Hsu,Yossef Mordechay,Robin Algayres,Tu Ahn Nguyen,Emmanuel Dupoux,Luke Zettlemoyer,Abdelrahman Mohamed
机构:Meta AI
摘要:端到端的口语理解(SLU)使用单个模型直接从音频预测意图,它通过利用中间文本表示中丢失的声学信息和防止自动语音识别中的级联错误,有望提高辅助系统的性能此外,当在设备上部署辅助系统时,具有一个统一的模型具有效率优势。本文提出了面向口语任务的语义分析算法(SpokenTask-OrientedSemantic Parsing(STOP)数据集,这是公开提供的最大、最复杂的SLU数据集。此外,我们定义了低资源分割来建立在有限的标记数据可用时提高SLU的基准。此外,除了人类录制的音频之外,我们正在发布一个TTS生成的版本,以基准测试端到端SLU系统的低资源域适应性能。初步实验表明,端到端SLU模型的性能略差于其级联对应模型,我们希望这能鼓励未来在此方向上的工作。
摘要:End-to-end spoken language understanding (SLU) predicts intent directly from audio using a single model. It promises to improve the performance of assistant systems by leveraging acoustic information lost in the intermediate textual representation and preventing cascading errors from Automatic Speech Recognition (ASR). Further, having one unified model has efficiency advantages when deploying assistant systems on-device. However, the limited number of public audio datasets with semantic parse labels hinders the research progress in this area. In this paper, we release the Spoken Task-Oriented semantic Parsing (STOP) dataset, the largest and most complex SLU dataset to be publicly available. Additionally, we define low-resource splits to establish a benchmark for improving SLU when limited labeled data is available. Furthermore, in addition to the human-recorded audio, we are releasing a TTS-generated version to benchmark the performance for low-resource domain adaptation of end-to-end SLU systems. Initial experimentation show end-to-end SLU models performing slightly worse than their cascaded counterparts, which we hope encourages future work in this direction.


【2】 Knowledge Transfer and Distillation from Autoregressive to  Non-Autoregressive Speech Recognition

标题:自回归到非自回归语音识别的知识转移与提取

链接:https://arxiv.org/abs/2207.10600

作者:Xun Gong,Zhikai Zhou,Yanmin Qian
机构:MoE Key Lab of Artificial Intelligence, AI Institute, X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai, China
备注:Accepted to Interspeech 2022
摘要:现代非自回归语音识别系统的目标是提高推理速度;提出了一种新的知识转移和提取结构,利用AR模型中的知识来提高NAR的性能,同时减小模型的规模.为了进一步提高NAR的性能,我们设计了帧级和序列级的转移学习目标,提出了一种基于Mask-CTC的波束搜索方法以扩大推理阶段的搜索空间.实验表明,与AR教师具有相同大小的NAR学生在AISHELL-1开发/测试集上获得8/16%的相对CER减少,并且在LibriSpeech测试—清洁/其他集上获得超过25%的相对WER减少。此外,在AISHELL-1和LibriSpeech基准测试中,约9倍的NAR模型通过所建议的知识转移和提炼实现了约25%的相对CER/WER降低。
摘要:Modern non-autoregressive~(NAR) speech recognition systems aim to accelerate the inference speed; however, they suffer from performance degradation compared with autoregressive~(AR) models as well as the huge model size issue. We propose a novel knowledge transfer and distillation architecture that leverages knowledge from AR models to improve the NAR performance while reducing the model's size. Frame- and sequence-level objectives are well-designed for transfer learning. To further boost the performance of NAR, a beam search method on Mask-CTC is developed to enlarge the search space during the inference stage. Experiments show that the proposed NAR beam search relatively reduces CER by over 5% on AISHELL-1 benchmark with a tolerable real-time-factor~(RTF) increment. By knowledge transfer, the NAR student who has the same size as the AR teacher obtains relative CER reductions of 8/16% on AISHELL-1 dev/test sets, and over 25% relative WER reductions on LibriSpeech test-clean/other sets. Moreover, the ~9x smaller NAR models achieve ~25% relative CER/WER reductions on both AISHELL-1 and LibriSpeech benchmarks with the proposed knowledge transfer and distillation.


【3】 Surrey System for DCASE 2022 Task 5: Few-shot Bioacoustic Event  Detection with Segment-level Metric Learning

标题:DCASE 2022任务5的Surrey系统:基于分段尺度学习的Few-Shot生物声学事件检测

链接:https://arxiv.org/abs/2207.10547

作者:Haohe Liu,Xubo Liu,Xinhao Mei,Qiuqiang Kong,Wenwu Wang,Mark D. Plumbley
机构:Centre for Vision, Speech, and Signal Processing (CVSSP), University of Surrey, UK,  Speech, Audio, and Music Intelligence (SAMI) Group, ByteDance, China
备注:Technical Report of the system that ranks 2nd in the DCASE Challenge Task 5. arXiv admin note: text overlap with arXiv:2207.07773
摘要:针对DCASE 2022中的少镜头few-Shot声学事件检测问题,提出了一种基于片段级度量学习的少镜头生物声学事件检测系统(任务5)我们更好地利用每个声音类内的负数据来构建损失函数,并利用直推推理在评价集上获得较好的自适应性,我们发现与Δ mel频率倒谱系数级联的每通道能量归一化是最有效的组合。我们的最终系统在DCASE任务5验证集上达到了68.74的f-测量值,大大超过了基线性能29.5。//github.com/haoheliu/DCASE_2022_Task_5.
摘要:Few-shot audio event detection is a task that detects the occurrence time of a novel sound class given a few examples. In this work, we propose a system based on segment-level metric learning for the DCASE 2022 challenge of few-shot bioacoustic event detection (task 5). We make better utilization of the negative data within each sound class to build the loss function, and use transductive inference to gain better adaptation on the evaluation set. For the input feature, we find the per-channel energy normalization concatenated with delta mel-frequency cepstral coefficients to be the most effective combination. We also introduce new data augmentation and post-processing procedures for this task. Our final system achieves an f-measure of 68.74 on the DCASE task 5 validation set, outperforming the baseline performance of 29.5 by a large margin. Our system is fully open-sourced at https://github.com/haoheliu/DCASE_2022_Task_5.


【4】 Room geometry blind inference based on the localization of real sound  source and first order reflections

标题:基于真实声源定位和一阶反射的房间几何盲推断

链接:https://arxiv.org/abs/2207.10478

作者:Shan Gao,Xihong Wu,Tianshu Qu
机构:Peking University
摘要:传统的基于声学信号的房间几何形状盲推断技术是基于环境的先验知识,如房间脉冲响应来进行的(RIR)或声源位置,这将限制其在已知场景下的应用。提出了一种利用直达波与一阶反射波之间的几何关系进行房间几何重建的方法。该方法除了利用紧凑型麦克风阵列本身的信息外,不需要对环境参数进行任何预知,设计了基于学习的DNN模型,提高了直达波和一阶反射波定位结果的准确性和完整性,(DOA)和到达时间差首先利用所提出的DCNN和TD-CNN模型估计直达波和反射波的时差信息,该方法具有比传统方法更高的灵敏度和精度,仿真和实测结果表明,在不同的混响环境下,与传统方法相比,该方法具有较好的精度和有效性。
摘要:The conventional room geometry blind inference techniques with acoustic signals are conducted based on the prior knowledge of the environment, such as the room impulse response (RIR) or the sound source position, which will limit its application under known scenarios. To solve this problem, we have proposed a room geometry reconstruction method in this paper by using the geometric relation between the direct signal and first-order reflections. In addition to the information of the compact microphone array itself, this method does not need any precognition of the environmental parameters. Besides, the learning-based DNN models are designed and used to improve the accuracy and integrity of the localization results of the direct source and first-order reflections. The direction of arrival (DOA) and time difference of arrival (TDOA) information of the direct and reflected signals are firstly estimated using the proposed DCNN and TD-CNN models, which have higher sensitivity and accuracy than the conventional methods. Then the position of the sound source is inferred by integrating the DOA, TDOA and array height using the proposed DNN model. After that, the positions of image sources and corresponding boundaries are derived based on the geometric relation. Experimental results of both simulations and real measurements verify the effectiveness and accuracy of the proposed techniques compared with the conventional methods under different reverberant environments.


【5】 Deep Audio Waveform Prior

标题:前一个深度音频波形

链接:https://arxiv.org/abs/2207.10441

作者:Arnon Turetzky,Tzvi Michelson,Yossi Adi,Shmuel Peleg
机构:School of Computer Science and Engineering, The Hebrew University of Jerusalem
备注:Interspeech 2022
摘要:卷积神经网络包含用于生成看起来自然的图像的强先验[1]。这些先验使得能够以无监督的方式进行图像去噪、超分辨率和修复。先前尝试在音频中演示类似的想法,即深度音频先验,(i)使用手工挑选的结构,例如谐波卷积,(i i)仅与谱图输入一起工作,和(iii)主要用于消除高斯噪声[2]。在这项工作中,我们表明,现有的SOTA架构的音频源分离包含深先验,即使工作与原始波形。当输入白噪声时,通过训练神经网络生成单个受损信号,可以发现深度先验。具有相关深度先验的网络可能在收敛到受损信号之前生成更清晰的信号版本。我们用几种受损情况演示了这种恢复效果:背景噪声、混响和信号中的间隙(音频修补)。
摘要:Convolutional neural networks contain strong priors for generating natural looking images [1]. These priors enable image denoising, super resolution, and inpainting in an unsupervised manner. Previous attempts to demonstrate similar ideas in audio, namely deep audio priors, (i) use hand picked architectures such as harmonic convolutions, (ii) only work with spectrogram input, and (iii) have been used mostly for eliminating Gaussian noise [2]. In this work we show that existing SOTA architectures for audio source separation contain deep priors even when working with the raw waveform. Deep priors can be discovered by training a neural network to generate a single corrupted signal when given white noise as input. A network with relevant deep priors is likely to generate a cleaner version of the signal before converging on the corrupted signal. We demonstrate this restoration effect with several corruptions: background noise, reverberations, and a gap in the signal (audio inpainting).


【6】 Spatial Aware Multi-Task Learning Based Speech Separation

标题:基于空间感知多任务学习的语音分离

链接:https://arxiv.org/abs/2207.10229

作者:Wei Sun,Mei Wang,Lili Qiu
机构:The University of Texas at Austin, USA
摘要:在新冠肺炎疫情期间,在线会议已经成为我们生活中不可或缺的一部分。由于其便利性和广泛的覆盖范围,这种趋势可能会继续下去。然而,来自其他家庭成员、室友、办公室同事的背景噪音不仅降低了语音质量,还引发了严重的隐私问题。本文开发了一种新颖的系统,称为基于空间感知的多任务学习分离(SAMS),以便在电话会议期间提取目标用户的音频信号。(i)从用户的语音和听不见的跟踪声音中生成细粒度的位置嵌入,其中包含用户的位置和丰富的多径信息;(ii)使用多任务学习开发源分离神经网络,以联合优化源分离和定位;(iii)显著加快推理速度,提供实时性保证.我们的测试台实验证明了我们方法的有效性
摘要:During the Covid, online meetings have become an indispensable part of our lives. This trend is likely to continue due to their convenience and broad reach. However, background noise from other family members, roommates, office-mates not only degrades the voice quality but also raises serious privacy issues. In this paper, we develop a novel system, called Spatial Aware Multi-task learning-based Separation (SAMS), to extract audio signals from the target user during teleconferencing. Our solution consists of three novel components: (i) generating fine-grained location embeddings from the user's voice and inaudible tracking sound, which contains the user's position and rich multipath information, (ii) developing a source separation neural network using multi-task learning to jointly optimize source separation and location, and (iii) significantly speeding up inference to provide a real-time guarantee. Our testbed experiments demonstrate the effectiveness of our approach


【7】 AudioScopeV2: Audio-Visual Attention Architectures for Calibrated  Open-Domain On-Screen Sound Separation

标题:音频示波器V2:用于校准的开放域屏幕上声音分离的视听注意力体系结构

链接:https://arxiv.org/abs/2207.10141

作者:Efthymios Tzinis,Scott Wisdom,Tal Remez,John R. Hershey
机构:Google Research,  University of Illinois Urbana-Champaign
备注:ECCV 2022
摘要:本文介绍了一个通用的视听屏幕声音分离系统AudioScopeV2,它能够通过观看野外视频学习分离声音并将其与屏幕对象相关联,指出了以往视听屏幕声音分离工作的几个局限性,包括时空注意力分辨率粗糙、音频分离模型收敛性差我们提出的跨模态和自注意网络体系结构随着时间的推移以更精细的分辨率捕获视听依赖性,并且我们还提出了能够缩放到更长视频而不牺牲太多性能的有效的可分离变体.我们还发现仅在音频上预训练分离模型极大地改进了结果.对于训练和评估,我们从一个大型的野外视频数据库中收集了新的人类对屏幕声音的注释(YFCC100M)。这个新的数据集更加多样化和具有挑战性。最后,我们提出了一个校准过程,允许精确调谐屏幕上重建与屏幕外抑制,这大大简化了不同工作点模型之间的性能比较.总体而言,我们的实验结果表明,在更一般的条件下,我们的屏幕分离性能比以前的方法有显著的改善,而附加的计算复杂度最小.
摘要:We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify several limitations of previous work on audio-visual on-screen sound separation, including the coarse resolution of spatio-temporal attention, poor convergence of the audio separation model, limited variety in training and evaluation data, and failure to account for the trade off between preservation of on-screen sounds and suppression of off-screen sounds. We provide solutions to all of these issues. Our proposed cross-modal and self-attention network architectures capture audio-visual dependencies at a finer resolution over time, and we also propose efficient separable variants that are capable of scaling to longer videos without sacrificing much performance. We also find that pre-training the separation model only on audio greatly improves results. For training and evaluation, we collected new human annotations of onscreen sounds from a large database of in-the-wild videos (YFCC100M). This new dataset is more diverse and challenging. Finally, we propose a calibration procedure that allows exact tuning of on-screen reconstruction versus off-screen suppression, which greatly simplifies comparing performance between models with different operating points. Overall, our experimental results show marked improvements in on-screen separation performance under much more general conditions than previous methods with minimal additional computational complexity.


eess.AS音频处理

【1】 Jointly Predicting Emotion, Age, and Country Using Pre-Trained Acoustic  Embedding

标题:使用预先训练的声学嵌入来联合预测情绪、年龄和国家

链接:https://arxiv.org/abs/2207.10333

作者:Bagus Tris Atmaja,Zanjabila,Akira Sasou
机构:∗ National Institute of Advanced Industrial Science and Technology, Japan, † Sepuluh Nopember Institute of Technology, Indonesia
摘要:在本文中,我们展示了使用预训练模型提取声学嵌入以联合预测(多任务学习)三个任务的益处:情感、年龄和国籍.在语音情感语料库上用wav2vec2.0大型鲁棒模型训练预训练模型.情感和年龄任务是回归问题,国籍预测是分类任务.用三个度量的单一调和平均值来评价多任务学习的性能.分类器是一个具有两个独立层和共享层的线性网络,包括从情感语音数据集上训练的模型中提取的声学嵌入)、种子数、批量大小和归一化处理,以从语音中预测副语言信息。
摘要:In this paper, we demonstrated the benefit of using pre-trained model to extract acoustic embedding to jointly predict (multitask learning) three tasks: emotion, age, and native country. The pre-trained model was trained with wav2vec 2.0 large robust model on the speech emotion corpus. The emotion and age tasks were regression problems, while country prediction was a classification task. A single harmonic mean from three metrics was used to evaluate the performance of multitask learning. The classifier was a linear network with two independent layers and shared layers, including the output layers. This study explores multitask learning on different acoustic features (including the acoustic embedding extracted from a model trained on an affective speech dataset), seed numbers, batch sizes, and normalizations for predicting paralinguistic information from speech.


【2】 STOP: A dataset for Spoken Task Oriented Semantic Parsing

标题:STOP:面向口语任务的语义分析数据集

链接:https://arxiv.org/abs/2207.10643

* 与cs.SD语音【1】为同一篇

作者:Paden Tomasello,Po-Chun Hsu,Akshat Shrivastava,Daniel Lazar,Duc Le,Adithya Sagar,Ali Elkahky,Jade Copet,Wei-Ning Hsu,Yossef Mordechay,Robin Algayres,Tu Ahn Nguyen,Emmanuel Dupoux,Luke Zettlemoyer,Abdelrahman Mohamed
机构:Meta AI
摘要:端到端的口语理解(SLU)使用单个模型直接从音频预测意图,它通过利用中间文本表示中丢失的声学信息和防止自动语音识别中的级联错误,有望提高辅助系统的性能此外,当在设备上部署辅助系统时,具有一个统一的模型具有效率优势。本文提出了面向口语任务的语义分析算法(SpokenTask-OrientedSemantic Parsing(STOP)数据集,这是公开提供的最大、最复杂的SLU数据集。此外,我们定义了低资源分割来建立在有限的标记数据可用时提高SLU的基准。此外,除了人类录制的音频之外,我们正在发布一个TTS生成的版本,以基准测试端到端SLU系统的低资源域适应性能。初步实验表明,端到端SLU模型的性能略差于其级联对应模型,我们希望这能鼓励未来在此方向上的工作。
摘要:End-to-end spoken language understanding (SLU) predicts intent directly from audio using a single model. It promises to improve the performance of assistant systems by leveraging acoustic information lost in the intermediate textual representation and preventing cascading errors from Automatic Speech Recognition (ASR). Further, having one unified model has efficiency advantages when deploying assistant systems on-device. However, the limited number of public audio datasets with semantic parse labels hinders the research progress in this area. In this paper, we release the Spoken Task-Oriented semantic Parsing (STOP) dataset, the largest and most complex SLU dataset to be publicly available. Additionally, we define low-resource splits to establish a benchmark for improving SLU when limited labeled data is available. Furthermore, in addition to the human-recorded audio, we are releasing a TTS-generated version to benchmark the performance for low-resource domain adaptation of end-to-end SLU systems. Initial experimentation show end-to-end SLU models performing slightly worse than their cascaded counterparts, which we hope encourages future work in this direction.


【3】 Knowledge Transfer and Distillation from Autoregressive to  Non-Autoregressive Speech Recognition

标题:自回归到非自回归语音识别的知识转移与提取

链接:https://arxiv.org/abs/2207.10600

* 与cs.SD语音【2】为同一篇

作者:Xun Gong,Zhikai Zhou,Yanmin Qian
机构:MoE Key Lab of Artificial Intelligence, AI Institute, X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai, China
备注:Accepted to Interspeech 2022
摘要:现代非自回归语音识别系统的目标是提高推理速度;提出了一种新的知识转移和提取结构,利用AR模型中的知识来提高NAR的性能,同时减小模型的规模.为了进一步提高NAR的性能,我们设计了帧级和序列级的转移学习目标,提出了一种基于Mask-CTC的波束搜索方法以扩大推理阶段的搜索空间.实验表明,与AR教师具有相同大小的NAR学生在AISHELL-1开发/测试集上获得8/16%的相对CER减少,并且在LibriSpeech测试—清洁/其他集上获得超过25%的相对WER减少。此外,在AISHELL-1和LibriSpeech基准测试中,约9倍的NAR模型通过所建议的知识转移和提炼实现了约25%的相对CER/WER降低。
摘要:Modern non-autoregressive~(NAR) speech recognition systems aim to accelerate the inference speed; however, they suffer from performance degradation compared with autoregressive~(AR) models as well as the huge model size issue. We propose a novel knowledge transfer and distillation architecture that leverages knowledge from AR models to improve the NAR performance while reducing the model's size. Frame- and sequence-level objectives are well-designed for transfer learning. To further boost the performance of NAR, a beam search method on Mask-CTC is developed to enlarge the search space during the inference stage. Experiments show that the proposed NAR beam search relatively reduces CER by over 5% on AISHELL-1 benchmark with a tolerable real-time-factor~(RTF) increment. By knowledge transfer, the NAR student who has the same size as the AR teacher obtains relative CER reductions of 8/16% on AISHELL-1 dev/test sets, and over 25% relative WER reductions on LibriSpeech test-clean/other sets. Moreover, the ~9x smaller NAR models achieve ~25% relative CER/WER reductions on both AISHELL-1 and LibriSpeech benchmarks with the proposed knowledge transfer and distillation.


【4】 Surrey System for DCASE 2022 Task 5: Few-shot Bioacoustic Event  Detection with Segment-level Metric Learning

标题:DCASE 2022任务5的Surrey系统:基于分段尺度学习的Few-Shot生物声学事件检测

链接:https://arxiv.org/abs/2207.10547

* 与cs.SD语音【3】为同一篇

作者:Haohe Liu,Xubo Liu,Xinhao Mei,Qiuqiang Kong,Wenwu Wang,Mark D. Plumbley
机构:Centre for Vision, Speech, and Signal Processing (CVSSP), University of Surrey, UK,  Speech, Audio, and Music Intelligence (SAMI) Group, ByteDance, China
备注:Technical Report of the system that ranks 2nd in the DCASE Challenge Task 5. arXiv admin note: text overlap with arXiv:2207.07773
摘要:针对DCASE 2022中的少镜头few-Shot声学事件检测问题,提出了一种基于片段级度量学习的少镜头生物声学事件检测系统(任务5)我们更好地利用每个声音类内的负数据来构建损失函数,并利用直推推理在评价集上获得较好的自适应性,我们发现与Δ mel频率倒谱系数级联的每通道能量归一化是最有效的组合。我们的最终系统在DCASE任务5验证集上达到了68.74的f-测量值,大大超过了基线性能29.5。//github.com/haoheliu/DCASE_2022_Task_5.
摘要:Few-shot audio event detection is a task that detects the occurrence time of a novel sound class given a few examples. In this work, we propose a system based on segment-level metric learning for the DCASE 2022 challenge of few-shot bioacoustic event detection (task 5). We make better utilization of the negative data within each sound class to build the loss function, and use transductive inference to gain better adaptation on the evaluation set. For the input feature, we find the per-channel energy normalization concatenated with delta mel-frequency cepstral coefficients to be the most effective combination. We also introduce new data augmentation and post-processing procedures for this task. Our final system achieves an f-measure of 68.74 on the DCASE task 5 validation set, outperforming the baseline performance of 29.5 by a large margin. Our system is fully open-sourced at https://github.com/haoheliu/DCASE_2022_Task_5.


【5】 Room geometry blind inference based on the localization of real sound  source and first order reflections

标题:基于真实声源定位和一阶反射的房间几何盲推断

链接:https://arxiv.org/abs/2207.10478

* 与cs.SD语音【4】为同一篇

作者:Shan Gao,Xihong Wu,Tianshu Qu
机构:Peking University
摘要:传统的基于声学信号的房间几何形状盲推断技术是基于环境的先验知识,如房间脉冲响应来进行的(RIR)或声源位置,这将限制其在已知场景下的应用。提出了一种利用直达波与一阶反射波之间的几何关系进行房间几何重建的方法。该方法除了利用紧凑型麦克风阵列本身的信息外,不需要对环境参数进行任何预知,设计了基于学习的DNN模型,提高了直达波和一阶反射波定位结果的准确性和完整性,(DOA)和到达时间差首先利用所提出的DCNN和TD-CNN模型估计直达波和反射波的时差信息,该方法具有比传统方法更高的灵敏度和精度,仿真和实测结果表明,在不同的混响环境下,与传统方法相比,该方法具有较好的精度和有效性。
摘要:The conventional room geometry blind inference techniques with acoustic signals are conducted based on the prior knowledge of the environment, such as the room impulse response (RIR) or the sound source position, which will limit its application under known scenarios. To solve this problem, we have proposed a room geometry reconstruction method in this paper by using the geometric relation between the direct signal and first-order reflections. In addition to the information of the compact microphone array itself, this method does not need any precognition of the environmental parameters. Besides, the learning-based DNN models are designed and used to improve the accuracy and integrity of the localization results of the direct source and first-order reflections. The direction of arrival (DOA) and time difference of arrival (TDOA) information of the direct and reflected signals are firstly estimated using the proposed DCNN and TD-CNN models, which have higher sensitivity and accuracy than the conventional methods. Then the position of the sound source is inferred by integrating the DOA, TDOA and array height using the proposed DNN model. After that, the positions of image sources and corresponding boundaries are derived based on the geometric relation. Experimental results of both simulations and real measurements verify the effectiveness and accuracy of the proposed techniques compared with the conventional methods under different reverberant environments.


【6】 Deep Audio Waveform Prior

标题:前一个深度音频波形

链接:https://arxiv.org/abs/2207.10441

* 与cs.SD语音【5】为同一篇

作者:Arnon Turetzky,Tzvi Michelson,Yossi Adi,Shmuel Peleg
机构:School of Computer Science and Engineering, The Hebrew University of Jerusalem
备注:Interspeech 2022
摘要:卷积神经网络包含用于生成看起来自然的图像的强先验[1]。这些先验使得能够以无监督的方式进行图像去噪、超分辨率和修复。先前尝试在音频中演示类似的想法,即深度音频先验,(i)使用手工挑选的结构,例如谐波卷积,(i i)仅与谱图输入一起工作,和(iii)主要用于消除高斯噪声[2]。在这项工作中,我们表明,现有的SOTA架构的音频源分离包含深先验,即使工作与原始波形。当输入白噪声时,通过训练神经网络生成单个受损信号,可以发现深度先验。具有相关深度先验的网络可能在收敛到受损信号之前生成更清晰的信号版本。我们用几种受损情况演示了这种恢复效果:背景噪声、混响和信号中的间隙(音频修补)。
摘要:Convolutional neural networks contain strong priors for generating natural looking images [1]. These priors enable image denoising, super resolution, and inpainting in an unsupervised manner. Previous attempts to demonstrate similar ideas in audio, namely deep audio priors, (i) use hand picked architectures such as harmonic convolutions, (ii) only work with spectrogram input, and (iii) have been used mostly for eliminating Gaussian noise [2]. In this work we show that existing SOTA architectures for audio source separation contain deep priors even when working with the raw waveform. Deep priors can be discovered by training a neural network to generate a single corrupted signal when given white noise as input. A network with relevant deep priors is likely to generate a cleaner version of the signal before converging on the corrupted signal. We demonstrate this restoration effect with several corruptions: background noise, reverberations, and a gap in the signal (audio inpainting).


【7】 Spatial Aware Multi-Task Learning Based Speech Separation

标题:基于空间感知多任务学习的语音分离

链接:https://arxiv.org/abs/2207.10229

* 与cs.SD语音【6】为同一篇

作者:Wei Sun,Mei Wang,Lili Qiu
机构:The University of Texas at Austin, USA
摘要:在新冠肺炎疫情期间,在线会议已经成为我们生活中不可或缺的一部分。由于其便利性和广泛的覆盖范围,这种趋势可能会继续下去。然而,来自其他家庭成员、室友、办公室同事的背景噪音不仅降低了语音质量,还引发了严重的隐私问题。本文开发了一种新颖的系统,称为基于空间感知的多任务学习分离(SAMS),以便在电话会议期间提取目标用户的音频信号。(i)从用户的语音和听不见的跟踪声音中生成细粒度的位置嵌入,其中包含用户的位置和丰富的多径信息;(ii)使用多任务学习开发源分离神经网络,以联合优化源分离和定位;(iii)显著加快推理速度,提供实时性保证.我们的测试台实验证明了我们方法的有效性
摘要:During the Covid, online meetings have become an indispensable part of our lives. This trend is likely to continue due to their convenience and broad reach. However, background noise from other family members, roommates, office-mates not only degrades the voice quality but also raises serious privacy issues. In this paper, we develop a novel system, called Spatial Aware Multi-task learning-based Separation (SAMS), to extract audio signals from the target user during teleconferencing. Our solution consists of three novel components: (i) generating fine-grained location embeddings from the user's voice and inaudible tracking sound, which contains the user's position and rich multipath information, (ii) developing a source separation neural network using multi-task learning to jointly optimize source separation and location, and (iii) significantly speeding up inference to provide a real-time guarantee. Our testbed experiments demonstrate the effectiveness of our approach


【8】 AudioScopeV2: Audio-Visual Attention Architectures for Calibrated  Open-Domain On-Screen Sound Separation

标题:音频示波器V2:用于校准的开放域屏幕上声音分离的视听注意力体系结构

链接:https://arxiv.org/abs/2207.10141

* 与cs.SD语音【7】为同一篇

作者:Efthymios Tzinis,Scott Wisdom,Tal Remez,John R. Hershey
机构:Google Research,  University of Illinois Urbana-Champaign
备注:ECCV 2022
摘要:本文介绍了一个通用的视听屏幕声音分离系统AudioScopeV2,它能够通过观看野外视频学习分离声音并将其与屏幕对象相关联,指出了以往视听屏幕声音分离工作的几个局限性,包括时空注意力分辨率粗糙、音频分离模型收敛性差我们提出的跨模态和自注意网络体系结构随着时间的推移以更精细的分辨率捕获视听依赖性,并且我们还提出了能够缩放到更长视频而不牺牲太多性能的有效的可分离变体.我们还发现仅在音频上预训练分离模型极大地改进了结果.对于训练和评估,我们从一个大型的野外视频数据库中收集了新的人类对屏幕声音的注释(YFCC100M)。这个新的数据集更加多样化和具有挑战性。最后,我们提出了一个校准过程,允许精确调谐屏幕上重建与屏幕外抑制,这大大简化了不同工作点模型之间的性能比较.总体而言,我们的实验结果表明,在更一般的条件下,我们的屏幕分离性能比以前的方法有显著的改善,而附加的计算复杂度最小.
摘要:We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify several limitations of previous work on audio-visual on-screen sound separation, including the coarse resolution of spatio-temporal attention, poor convergence of the audio separation model, limited variety in training and evaluation data, and failure to account for the trade off between preservation of on-screen sounds and suppression of off-screen sounds. We provide solutions to all of these issues. Our proposed cross-modal and self-attention network architectures capture audio-visual dependencies at a finer resolution over time, and we also propose efficient separable variants that are capable of scaling to longer videos without sacrificing much performance. We also find that pre-training the separation model only on audio greatly improves results. For training and evaluation, we collected new human annotations of onscreen sounds from a large database of in-the-wild videos (YFCC100M). This new dataset is more diverse and challenging. Finally, we propose a calibration procedure that allows exact tuning of on-screen reconstruction versus off-screen suppression, which greatly simplifies comparing performance between models with different operating points. Overall, our experimental results show marked improvements in on-screen separation performance under much more general conditions than previous methods with minimal additional computational complexity.


机器翻译,仅供参考