今日论文合集:cs.SD语音8篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Exploration of Adapter for Noise Robust Automatic Speech Recognition

标题:用于抗噪自动语音识别的适配器探讨
链接:https://arxiv.org/abs/2402.18275
作者:Hao Shi,Tatsuya Kawahara
摘要:采用鲁棒的自动语音识别(ASR)系统来处理不可见的噪声场景至关重要。将适配器集成到神经网络中已经成为迁移学习的一种有效技术。本文深入研究了基于自适应的噪声鲁棒ASR自适应。我们使用CHiME-4数据集进行了实验。结果表明,在浅层中插入适配器会产生更好的效果,并且仅在浅层内进行适配和跨所有层进行适配之间没有显著差异。此外,仿真数据有助于提高系统在实际噪声条件下的性能。然而,当数据量相同时,真实数据比模拟数据更有效。多条件训练对于适配器训练仍然有效。此外,将适配器集成到基于语音增强的ASR系统中会产生实质性的改进。
摘要:Adapting a robust automatic speech recognition (ASR) system to tackle unseen noise scenarios is crucial. Integrating adapters into neural networks has emerged as a potent technique for transfer learning. This paper thoroughly investigates adapter-based noise-robust ASR adaptation. We conducted the experiments using the CHiME--4 dataset. The results show that inserting the adapter in the shallow layer yields superior effectiveness, and there is no significant difference between adapting solely within the shallow layer and adapting across all layers. Besides, the simulated data helps the system to improve its performance under real noise conditions. Nonetheless, when the amount of data is the same, the real data is more effective than the simulated data. Multi-condition training remains valid for adapter training. Furthermore, integrating adapters into speech enhancement-based ASR systems yields substantial improvements.


【2】ConvDTW-ACS: Audio Segmentation for Track Type Detection During Car  Manufacturing
标题:ConvDTW-ACS:用于汽车制造过程中轨道类型检测的音频分割
链接:https://arxiv.org/abs/2402.18204
作者:Álvaro López-Chilet,Zhaoyi Liu,Jon Ander Gómez,Carlos Alvarez,Marivi Alonso Ortiz,Andres Orejuela Mesa,David Newton,Friedrich Wolf-Monheim,Sam Michiels,Danny Hughes
备注:12 pages, 2 figures
摘要:本文提出了一种方法,声学约束分割(ACS)在音频记录的车辆驾驶通过生产测试轨道,划定边界的表面类型的轨道。ACS是经典声学分段的一种变体,其中标签的序列是已知的、连续的和不变的,这在这项工作中特别有用,因为测试轨道具有标准的表面类型配置。所提出的ConvDTW-ACS方法利用卷积神经网络对从完整音频频谱图中提取的重叠图像块进行分类。然后,我们的自定义动态时间规整算法将预测概率序列与轨道中的表面类型序列对齐,从中可以提取表面类型边界的时间戳。该方法在从西班牙瓦伦西亚的福特制造厂收集的真实世界数据集上进行了评估,在音频内界定轨道中表面的边界时,平均误差为166毫秒。结果表明,所提出的方法在准确分割不同表面类型方面是有效的,这可以使更专业的AI系统的开发,以改善质量检测过程。
摘要:This paper proposes a method for Acoustic Constrained Segmentation (ACS) in audio recordings of vehicles driven through a production test track, delimiting the boundaries of surface types in the track. ACS is a variant of classical acoustic segmentation where the sequence of labels is known, contiguous and invariable, which is especially useful in this work as the test track has a standard configuration of surface types. The proposed ConvDTW-ACS method utilizes a Convolutional Neural Network for classifying overlapping image chunks extracted from the full audio spectrogram. Then, our custom Dynamic Time Warping algorithm aligns the sequence of predicted probabilities to the sequence of surface types in the track, from which timestamps of the surface type boundaries can be extracted. The method was evaluated on a real-world dataset collected from the Ford Manufacturing Plant in Valencia (Spain), achieving a mean error of 166 milliseconds when delimiting, within the audio, the boundaries of the surfaces in the track. The results demonstrate the effectiveness of the proposed method in accurately segmenting different surface types, which could enable the development of more specialized AI systems to improve the quality inspection process.


【3】AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response
标题:基于挑战-响应的人工智能深伪音频呼叫标记
链接:https://arxiv.org/abs/2402.18085
作者:Govind Mittal,Arthur Jakobsson,Kelly O. Marshall,Chinmay Hegde,Nasir Memon
备注:Dataset will be made public by end of March 2024
摘要:诈骗者正在积极利用人工智能语音克隆技术进行社会工程攻击,这种情况因音频实时Deepfakes(RTF)的出现而严重恶化。RTDF可以通过电话实时克隆目标的声音,使这些交互高度互动,从而更具说服力。我们的研究充满信心地解决了现有文献中关于deepfake检测的空白,这在很大程度上对RTDF威胁无效。我们引入了一种强大的基于挑战-响应的方法来检测deepfake音频呼叫,开创了音频挑战的全面分类。我们的评估针对领先的语音克隆系统提出了20项潜在挑战。我们编制了一个新的开源挑战数据集,来自100名智能手机和台式机用户的贡献,产生了18,600个原始样本和160万个deepfake样本。通过对该数据集进行严格的机器和人工评估,我们分别实现了86%的deepfake检测率和80%的AUC得分。值得注意的是,利用一组11个挑战大大提高了检测能力。我们的研究结果表明,将人类直觉与机器精度相结合可以提供互补优势。因此,我们开发了一种创新的人类-人工智能协作系统,将人类识别与算法准确性相结合,将最终的联合准确率提高到82.9%。该系统突出了人工智能辅助预筛选在呼叫验证过程中的显著优势。样品可以在https://mittalgovind.github.io/autch-samples/上听到
摘要:Scammers are aggressively leveraging AI voice-cloning technology for social engineering attacks, a situation significantly worsened by the advent of audio Real-time Deepfakes (RTDFs). RTDFs can clone a target's voice in real-time over phone calls, making these interactions highly interactive and thus far more convincing. Our research confidently addresses the gap in the existing literature on deepfake detection, which has largely been ineffective against RTDF threats. We introduce a robust challenge-response-based method to detect deepfake audio calls, pioneering a comprehensive taxonomy of audio challenges. Our evaluation pitches 20 prospective challenges against a leading voice-cloning system. We have compiled a novel open-source challenge dataset with contributions from 100 smartphone and desktop users, yielding 18,600 original and 1.6 million deepfake samples. Through rigorous machine and human evaluations of this dataset, we achieved a deepfake detection rate of 86% and an 80% AUC score, respectively. Notably, utilizing a set of 11 challenges significantly enhances detection capabilities. Our findings reveal that combining human intuition with machine precision offers complementary advantages. Consequently, we have developed an innovative human-AI collaborative system that melds human discernment with algorithmic accuracy, boosting final joint accuracy to 82.9%. This system highlights the significant advantage of AI-assisted pre-screening in call verification processes. Samples can be heard at https://mittalgovind.github.io/autch-samples/

【4】Mixer is more than just a model
标题:搅拌机不仅仅是一种模式
链接:https://arxiv.org/abs/2402.18007
作者:Qingfeng Ji,Yuxin Wang,Letong Sun
摘要:最近,MLP结构重新流行起来,MLP-Mixer就是一个突出的例子。在计算机视觉领域,MLP Mixer以其从通道和令牌角度提取数据信息的能力而闻名,有效地充当通道和令牌信息的融合。事实上,Mixer代表了一种融合通道和令牌信息的信息提取范例。Mixer的本质在于它能够混合来自不同视角的信息,集中体现了神经网络架构领域中“混合”的真正概念。除了通道和令牌考虑之外,还可以从各种角度创建更定制的混合器,以更好地满足特定的任务要求。本研究的重点是音频识别领域,介绍了一种新的模型命名为音频频谱图混合器与滚动时间和Hermit FFT(ASM-RH),结合了从时域和频域的见解。实验结果表明,ASM-RH特别适合音频数据,并在多个分类任务中产生有希望的结果。
摘要:Recently, MLP structures have regained popularity, with MLP-Mixer standing out as a prominent example. In the field of computer vision, MLP-Mixer is noted for its ability to extract data information from both channel and token perspectives, effectively acting as a fusion of channel and token information. Indeed, Mixer represents a paradigm for information extraction that amalgamates channel and token information. The essence of Mixer lies in its ability to blend information from diverse perspectives, epitomizing the true concept of "mixing" in the realm of neural network architectures. Beyond channel and token considerations, it is possible to create more tailored mixers from various perspectives to better suit specific task requirements. This study focuses on the domain of audio recognition, introducing a novel model named Audio Spectrogram Mixer with Roll-Time and Hermit FFT (ASM-RH) that incorporates insights from both time and frequency domains. Experimental results demonstrate that ASM-RH is particularly well-suited for audio data and yields promising outcomes across multiple classification tasks.


【5】ByteComposer: a Human-like Melody Composition Method based on Language  Model Agent
标题:ByteComposer:一种基于语言模型代理的仿人旋律合成方法
链接:https://arxiv.org/abs/2402.17785
作者:Xia Liang,Jiaju Lin,Xinjian Du
摘要:大型语言模型(LLM)在多模态理解和生成任务方面取得了令人鼓舞的进展。然而,如何设计一个符合人类和可解释的旋律创作系统仍然是探索不足。为了解决这个问题,我们提出了ByteComposer,这是一个代理框架,它通过四个独立的步骤来模拟人类的创造性管道:“概念分析-草稿组成-自我评估和修改-美学选择”。该框架无缝地融合了LLM的交互和知识理解功能与现有的符号音乐生成模型,从而实现了与人类创作者相媲美的旋律合成代理。我们进行了广泛的实验GPT 4和几个开源的大型语言模型,这证实了我们的框架的有效性。此外,专业音乐作曲家进行了多维度的评估,最终的结果表明,在音乐创作的各个方面,ByteComposer代理达到了新手旋律作曲家的水平。
摘要:Large Language Models (LLM) have shown encouraging progress in multimodal understanding and generation tasks. However, how to design a human-aligned and interpretable melody composition system is still under-explored. To solve this problem, we propose ByteComposer, an agent framework emulating a human's creative pipeline in four separate steps : "Conception Analysis - Draft Composition - Self-Evaluation and Modification - Aesthetic Selection". This framework seamlessly blends the interactive and knowledge-understanding features of LLMs with existing symbolic music generation models, thereby achieving a melody composition agent comparable to human creators. We conduct extensive experiments on GPT4 and several open-source large language models, which substantiate our framework's effectiveness. Furthermore, professional music composers were engaged in multi-dimensional evaluations, the final results demonstrated that across various facets of music composition, ByteComposer agent attains the level of a novice melody composer.


【6】Improvement Of Audiovisual Quality Estimation Using A Nonlinear  Autoregressive Exogenous Neural Network And Bitstream Parameters
标题:利用非线性自回归外源神经网络和码流参数改进视听质量估计
链接:https://arxiv.org/abs/2402.18056
作者:Koffi Kossi,Stephane Coulombe,Christian Desrosiers,Ghyslain Gagnon
摘要:随着对视听服务需求的不断增长,电信服务提供商和应用程序开发人员不得不确保其服务提供最佳的用户体验。特别是,视频会议等服务对网络状况非常敏感。因此,它们的性能应该被实时监控,以便根据任何网络扰动调整参数。在本文中,我们开发了一个参数模型,用于估计在视频会议服务的视听质量。我们的模型是开发与非线性自回归外源(NARX)递归神经网络和估计的感知质量的平均意见得分(MOS)。我们使用公开的INRS比特流视听质量数据集来验证我们的模型。该数据集包含比特流参数,如每帧丢失,比特率和视频持续时间。我们将所提出的模型与基于机器学习的最先进方法进行比较,并显示我们的模型在均方误差(MSE=0.150)和Pearson相关系数(R=0.931)方面优于这些方法。
摘要:With the increasing demand for audiovisual services, telecom service providers and application developers are compelled to ensure that their services provide the best possible user experience. Particularly, services such as videoconferencing are very sensitive to network conditions. Therefore, their performance should be monitored in real time in order to adjust parameters to any network perturbation. In this paper, we developed a parametric model for estimating the perceived audiovisual quality in videoconference services. Our model is developed with the nonlinear autoregressive exogenous (NARX) recurrent neural network and estimates the perceived quality in terms of mean opinion score (MOS). We validate our model using the publicly available INRS bitstream audiovisual quality dataset. This dataset contains bitstream parameters such as loss per frame, bit rate and video duration. We compare the proposed model against state-of-the-art methods based on machine learning and show our model to outperform these methods in terms of mean square error (MSE=0.150) and Pearson correlation coefficient (R=0.931)


【7】NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
标题:NIIRF:用于HRTF上采样和个性化的神经IIR滤波域
链接:https://arxiv.org/abs/2402.17907
作者:Yoshiki Masuyama,Gordon Wichern,François G. Germain,Zexu Pan,Sameer Khurana,Chiori Hori,Jonathan Le Roux
备注:Accepted to ICASSP 2024
摘要:头部相关传递函数(HRTF)对于沉浸式音频是重要的,并且已经研究了它们的空间插值以对有限测量进行上采样。最近,从声源方向映射到HRTF的神经场(NFs)得到了关注。现有的基于NF的方法集中于从给定的声源方向估计HRTF的幅度,并且幅度被转换为有限脉冲响应(FIR)滤波器。我们提出了神经无限脉冲响应滤波器场(NIIRF)的方法,而不是估计级联IIR滤波器的系数。IIR滤波器模仿HRTF的模态性质,因此与FIR滤波器相比,需要更少的系数来很好地近似它们。我们发现,我们的方法可以在多个数据集上匹配现有的基于NF的方法的性能,甚至在测量稀疏时优于它们。我们还探讨了个性化的NF到一个主题的方法和实验发现低级别的适应是有效的。
摘要:Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimating the magnitude of the HRTF from a given sound source direction, and the magnitude is converted to a finite impulse response (FIR) filter. We propose the neural infinite impulse response filter field (NIIRF) method that instead estimates the coefficients of cascaded IIR filters. IIR filters mimic the modal nature of HRTFs, thus needing fewer coefficients to approximate them well compared to FIR filters. We find that our method can match the performance of existing NF-based methods on multiple datasets, even outperforming them when measurements are sparse. We also explore approaches to personalize the NF to a subject and experimentally find low-rank adaptation to be effective.


【8】Wavelet Scattering Transform for Bioacustics: Application to Watkins  Marine Mammal Sound Database
标题:小波散射变换在生物声学中的应用:Watkins海洋哺乳动物声音数据库
链接:https://arxiv.org/abs/2402.17775
作者:Davide Carbone,Alessandro Licciardi
摘要:海洋哺乳动物的交流是一个复杂的领域,受到声音多样性和环境因素的阻碍。沃特金斯海洋哺乳动物声音数据库(WMMD)是机器学习应用程序中使用的广泛标记数据集。然而,在文献中发现的数据准备,预处理和分类的方法是相当不同的。本研究首先简要回顾了数据集上的最新基准,重点是澄清数据准备和预处理方法。随后,我们提出了应用小波散射变换(WST)的标准方法的基础上的短时傅立叶变换(STFT)。该研究还使用具有剩余层的ad-hoc深度架构来处理分类任务。我们优于现有的分类体系结构的准确性6 $\%$使用WST和8 $\%$使用梅尔频谱预处理,有效地减少了一半的错误分类的样本数量,并达到最高的准确性96 $\%$。
摘要:Marine mammal communication is a complex field, hindered by the diversity of vocalizations and environmental factors. The Watkins Marine Mammal Sound Database (WMMD) is an extensive labeled dataset used in machine learning applications. However, the methods for data preparation, preprocessing, and classification found in the literature are quite disparate. This study first focuses on a brief review of the state-of-the-art benchmarks on the dataset, with an emphasis on clarifying data preparation and preprocessing methods. Subsequently, we propose the application of the Wavelet Scattering Transform (WST) in place of standard methods based on the Short-Time Fourier Transform (STFT). The study also tackles a classification task using an ad-hoc deep architecture with residual layers. We outperform the existing classification architecture by $6\%$ in accuracy using WST and $8\%$ using Mel spectrogram preprocessing, effectively reducing by half the number of misclassified samples, and reaching a top accuracy of $96\%$.

eess.AS音频处理
【1】Why does music source separation benefit from cacophony?
标题:为什么音乐来源分离会受益于刺耳的声音?
链接:https://arxiv.org/abs/2402.18407
作者:Chang-Bin Jeon,Gordon Wichern,François G. Germain,Jonathan Le Roux
备注:ICASSP 2024 Workshop on Explainable AI for Speech and Audio
摘要:在音乐源分离中,标准的训练数据增强过程是通过随机组合来自不同歌曲的乐器茎来创建新的训练样本。这些随机混合与真实音乐相比具有不匹配的特性,例如,不同的词干没有一致的节拍或音调,导致不和谐的声音。在这项工作中,我们研究了为什么随机混合在训练最先进的音乐源分离模型时是有效的,尽管它会产生明显的分布偏移。此外,我们研究了为什么性能水平,尽管潜在的无限组合,并检查音乐源分离性能的灵敏度,在混合物中的乐器源的节拍和音调的差异。
摘要:In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These random mixes have mismatched characteristics compared to real music, e.g., the different stems do not have consistent beat or tonality, resulting in a cacophony. In this work, we investigate why random mixing is effective when training a state-of-the-art music source separation model in spite of the apparent distribution shift it creates. Additionally, we examine why performance levels off despite potentially limitless combinations, and examine the sensitivity of music source separation performance to differences in beat and tonality of the instrumental sources in a mixture.

【2】Improvement Of Audiovisual Quality Estimation Using A Nonlinear  Autoregressive Exogenous Neural Network And Bitstream Parameters
标题:利用非线性自回归外源神经网络和码流参数改进视听质量估计
链接:https://arxiv.org/abs/2402.18056
作者:Koffi Kossi,Stephane Coulombe,Christian Desrosiers,Ghyslain Gagnon
摘要:随着对视听服务需求的不断增长,电信服务提供商和应用程序开发人员不得不确保其服务提供最佳的用户体验。特别是,视频会议等服务对网络状况非常敏感。因此,它们的性能应该被实时监控,以便根据任何网络扰动调整参数。在本文中,我们开发了一个参数模型,用于估计在视频会议服务的视听质量。我们的模型是开发与非线性自回归外源(NARX)递归神经网络和估计的感知质量的平均意见得分(MOS)。我们使用公开的INRS比特流视听质量数据集来验证我们的模型。该数据集包含比特流参数,如每帧丢失,比特率和视频持续时间。我们将所提出的模型与基于机器学习的最先进方法进行比较,并显示我们的模型在均方误差(MSE=0.150)和Pearson相关系数(R=0.931)方面优于这些方法。
摘要:With the increasing demand for audiovisual services, telecom service providers and application developers are compelled to ensure that their services provide the best possible user experience. Particularly, services such as videoconferencing are very sensitive to network conditions. Therefore, their performance should be monitored in real time in order to adjust parameters to any network perturbation. In this paper, we developed a parametric model for estimating the perceived audiovisual quality in videoconference services. Our model is developed with the nonlinear autoregressive exogenous (NARX) recurrent neural network and estimates the perceived quality in terms of mean opinion score (MOS). We validate our model using the publicly available INRS bitstream audiovisual quality dataset. This dataset contains bitstream parameters such as loss per frame, bit rate and video duration. We compare the proposed model against state-of-the-art methods based on machine learning and show our model to outperform these methods in terms of mean square error (MSE=0.150) and Pearson correlation coefficient (R=0.931)


【3】NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
标题:NIIRF:用于HRTF上采样和个性化的神经IIR滤波域
链接:https://arxiv.org/abs/2402.17907
作者:Yoshiki Masuyama,Gordon Wichern,François G. Germain,Zexu Pan,Sameer Khurana,Chiori Hori,Jonathan Le Roux
备注:Accepted to ICASSP 2024
摘要:头部相关传递函数(HRTF)对于沉浸式音频是重要的,并且已经研究了它们的空间插值以对有限测量进行上采样。最近,从声源方向映射到HRTF的神经场(NFs)得到了关注。现有的基于NF的方法集中于从给定的声源方向估计HRTF的幅度,并且幅度被转换为有限脉冲响应(FIR)滤波器。我们提出了神经无限脉冲响应滤波器场(NIIRF)的方法,而不是估计级联IIR滤波器的系数。IIR滤波器模仿HRTF的模态性质,因此与FIR滤波器相比,需要更少的系数来很好地近似它们。我们发现,我们的方法可以在多个数据集上匹配现有的基于NF的方法的性能,甚至在测量稀疏时优于它们。我们还探讨了个性化的NF到一个主题的方法和实验发现低级别的适应是有效的。
摘要:Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimating the magnitude of the HRTF from a given sound source direction, and the magnitude is converted to a finite impulse response (FIR) filter. We propose the neural infinite impulse response filter field (NIIRF) method that instead estimates the coefficients of cascaded IIR filters. IIR filters mimic the modal nature of HRTFs, thus needing fewer coefficients to approximate them well compared to FIR filters. We find that our method can match the performance of existing NF-based methods on multiple datasets, even outperforming them when measurements are sparse. We also explore approaches to personalize the NF to a subject and experimentally find low-rank adaptation to be effective.

【4】Wavelet Scattering Transform for Bioacustics: Application to Watkins  Marine Mammal Sound Database
标题:生物声学的小波散射变换:在Watkins海洋哺乳动物声音数据库中的应用
链接:https://arxiv.org/abs/2402.17775
作者:Davide Carbone,Alessandro Licciardi
摘要:海洋哺乳动物的交流是一个复杂的领域,受到声音多样性和环境因素的阻碍。沃特金斯海洋哺乳动物声音数据库(WMMD)是机器学习应用程序中使用的广泛标记数据集。然而,在文献中发现的数据准备,预处理和分类的方法是相当不同的。本研究首先简要回顾了数据集上的最新基准,重点是澄清数据准备和预处理方法。随后,我们提出了应用小波散射变换(WST)的标准方法的基础上的短时傅立叶变换(STFT)。该研究还使用具有剩余层的ad-hoc深度架构来处理分类任务。我们优于现有的分类体系结构的准确性6 $\%$使用WST和8 $\%$使用梅尔频谱预处理,有效地减少了一半的错误分类的样本数量,并达到最高的准确性96 $\%$。
摘要:Marine mammal communication is a complex field, hindered by the diversity of vocalizations and environmental factors. The Watkins Marine Mammal Sound Database (WMMD) is an extensive labeled dataset used in machine learning applications. However, the methods for data preparation, preprocessing, and classification found in the literature are quite disparate. This study first focuses on a brief review of the state-of-the-art benchmarks on the dataset, with an emphasis on clarifying data preparation and preprocessing methods. Subsequently, we propose the application of the Wavelet Scattering Transform (WST) in place of standard methods based on the Short-Time Fourier Transform (STFT). The study also tackles a classification task using an ad-hoc deep architecture with residual layers. We outperform the existing classification architecture by $6\%$ in accuracy using WST and $8\%$ using Mel spectrogram preprocessing, effectively reducing by half the number of misclassified samples, and reaching a top accuracy of $96\%$.


【5】EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous  Driving
标题:EchoTrack:用于自动驾驶的听觉参考多目标跟踪
链接:https://arxiv.org/abs/2402.18302
作者:Jiacheng Lin,Jiajun Chen,Kunyu Peng,Xuan He,Zhiyong Li,Rainer Stiefelhagen,Kailun Yang
备注:The source code and datasets will be made publicly available at this https URL
摘要:本文介绍了听觉参考多目标跟踪(AR-MOT)的任务,它动态跟踪特定的对象在视频序列中的音频表达的基础上,并出现作为一个具有挑战性的问题,在自动驾驶。由于缺乏对音频和视频的语义建模能力,现有的工作主要集中在基于文本的多目标跟踪上,这往往以跟踪质量、交互效率甚至辅助系统的安全性为代价,限制了此类方法在自动驾驶中的应用。本文从音视频融合和音视频跟踪的角度对AR-MOT问题进行了深入研究。我们提出了EchoTrack,这是一个具有双流Vision Transformers的端到端AR-MOT框架。双流与我们的双向频域交叉注意力融合模块(Bi-FCFM)交织在一起,该模块双向融合了来自频域和时空域的音频和视频特征。此外,我们提出了视听对比跟踪学习(ACTL)制度,通过学习不同的音频和视频对象之间的同质特征有效地提取表达式和视觉对象之间的同质语义特征。除了架构设计,我们还建立了第一套大规模的AR-MOT基准测试,包括Echo-KITTI,Echo-KITTI+和Echo-BDD。在已建立的基准上进行的大量实验表明了所提出的EchoTrack模型及其组件的有效性。源代码和数据集将在https://github.com/lab206/EchoTrack上公开。
摘要:This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autonomous driving. Due to the lack of semantic modeling capacity in audio and video, existing works have mainly focused on text-based multi-object tracking, which often comes at the cost of tracking quality, interaction efficiency, and even the safety of assistance systems, limiting the application of such methods in autonomous driving. In this paper, we delve into the problem of AR-MOT from the perspective of audio-video fusion and audio-video tracking. We put forward EchoTrack, an end-to-end AR-MOT framework with dual-stream vision transformers. The dual streams are intertwined with our Bidirectional Frequency-domain Cross-attention Fusion Module (Bi-FCFM), which bidirectionally fuses audio and video features from both frequency- and spatiotemporal domains. Moreover, we propose the Audio-visual Contrastive Tracking Learning (ACTL) regime to extract homogeneous semantic features between expressions and visual objects by learning homogeneous features between different audio and video objects effectively. Aside from the architectural design, we establish the first set of large-scale AR-MOT benchmarks, including Echo-KITTI, Echo-KITTI+, and Echo-BDD. Extensive experiments on the established benchmarks demonstrate the effectiveness of the proposed EchoTrack model and its components. The source code and datasets will be made publicly available at https://github.com/lab206/EchoTrack.


【6】Exploration of Adapter for Noise Robust Automatic Speech Recognition
标题:用于抗噪自动语音识别的适配器探讨
链接:https://arxiv.org/abs/2402.18275
作者:Hao Shi,Tatsuya Kawahara
摘要:采用鲁棒的自动语音识别(ASR)系统来处理不可见的噪声场景至关重要。将适配器集成到神经网络中已经成为迁移学习的一种有效技术。本文深入研究了基于自适应的噪声鲁棒ASR自适应。我们使用CHiME-4数据集进行了实验。结果表明,在浅层中插入适配器会产生更好的效果,并且仅在浅层内进行适配和跨所有层进行适配之间没有显著差异。此外,仿真数据有助于提高系统在实际噪声条件下的性能。然而,当数据量相同时,真实数据比模拟数据更有效。多条件训练对于适配器训练仍然有效。此外,将适配器集成到基于语音增强的ASR系统中会产生实质性的改进。
摘要:Adapting a robust automatic speech recognition (ASR) system to tackle unseen noise scenarios is crucial. Integrating adapters into neural networks has emerged as a potent technique for transfer learning. This paper thoroughly investigates adapter-based noise-robust ASR adaptation. We conducted the experiments using the CHiME--4 dataset. The results show that inserting the adapter in the shallow layer yields superior effectiveness, and there is no significant difference between adapting solely within the shallow layer and adapting across all layers. Besides, the simulated data helps the system to improve its performance under real noise conditions. Nonetheless, when the amount of data is the same, the real data is more effective than the simulated data. Multi-condition training remains valid for adapter training. Furthermore, integrating adapters into speech enhancement-based ASR systems yields substantial improvements.

【7】ConvDTW-ACS: Audio Segmentation for Track Type Detection During Car  Manufacturing
标题:ConvDTW-ACS:汽车制造过程中用于轨道类型检测的音频分割
链接:https://arxiv.org/abs/2402.18204
作者:Álvaro López-Chilet,Zhaoyi Liu,Jon Ander Gómez,Carlos Alvarez,Marivi Alonso Ortiz,Andres Orejuela Mesa,David Newton,Friedrich Wolf-Monheim,Sam Michiels,Danny Hughes
备注:12 pages, 2 figures
摘要:本文提出了一种方法,声学约束分割(ACS)在音频记录的车辆驾驶通过生产测试轨道,划定边界的表面类型的轨道。ACS是经典声学分段的一种变体,其中标签的序列是已知的、连续的和不变的,这在这项工作中特别有用,因为测试轨道具有标准的表面类型配置。所提出的ConvDTW-ACS方法利用卷积神经网络对从完整音频频谱图中提取的重叠图像块进行分类。然后,我们的自定义动态时间规整算法将预测概率序列与轨道中的表面类型序列对齐,从中可以提取表面类型边界的时间戳。该方法在从西班牙瓦伦西亚的福特制造厂收集的真实世界数据集上进行了评估,在音频内界定轨道中表面的边界时,平均误差为166毫秒。结果表明,所提出的方法在准确分割不同表面类型方面是有效的,这可以使更专业的AI系统的开发,以改善质量检测过程。
摘要:This paper proposes a method for Acoustic Constrained Segmentation (ACS) in audio recordings of vehicles driven through a production test track, delimiting the boundaries of surface types in the track. ACS is a variant of classical acoustic segmentation where the sequence of labels is known, contiguous and invariable, which is especially useful in this work as the test track has a standard configuration of surface types. The proposed ConvDTW-ACS method utilizes a Convolutional Neural Network for classifying overlapping image chunks extracted from the full audio spectrogram. Then, our custom Dynamic Time Warping algorithm aligns the sequence of predicted probabilities to the sequence of surface types in the track, from which timestamps of the surface type boundaries can be extracted. The method was evaluated on a real-world dataset collected from the Ford Manufacturing Plant in Valencia (Spain), achieving a mean error of 166 milliseconds when delimiting, within the audio, the boundaries of the surfaces in the track. The results demonstrate the effectiveness of the proposed method in accurately segmenting different surface types, which could enable the development of more specialized AI systems to improve the quality inspection process.

【8】AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response
标题:基于挑战-响应的人工智能深伪音频呼叫标记
链接:https://arxiv.org/abs/2402.18085
作者:Govind Mittal,Arthur Jakobsson,Kelly O. Marshall,Chinmay Hegde,Nasir Memon
备注:Dataset will be made public by end of March 2024
摘要:诈骗者正在积极利用人工智能语音克隆技术进行社会工程攻击,这种情况因音频实时Deepfakes(RTF)的出现而严重恶化。RTDF可以通过电话实时克隆目标的声音,使这些交互高度互动,从而更具说服力。我们的研究充满信心地解决了现有文献中关于deepfake检测的空白,这在很大程度上对RTDF威胁无效。我们引入了一种强大的基于挑战-响应的方法来检测deepfake音频呼叫,开创了音频挑战的全面分类。我们的评估针对领先的语音克隆系统提出了20项潜在挑战。我们编制了一个新的开源挑战数据集,来自100名智能手机和台式机用户的贡献,产生了18,600个原始样本和160万个deepfake样本。通过对该数据集进行严格的机器和人工评估,我们分别实现了86%的deepfake检测率和80%的AUC得分。值得注意的是,利用一组11个挑战大大提高了检测能力。我们的研究结果表明,将人类直觉与机器精度相结合可以提供互补优势。因此,我们开发了一种创新的人类-人工智能协作系统,将人类识别与算法准确性相结合,将最终的联合准确率提高到82.9%。该系统突出了人工智能辅助预筛选在呼叫验证过程中的显著优势。样品可以在https://mittalgovind.github.io/autch-samples/上听到
摘要:Scammers are aggressively leveraging AI voice-cloning technology for social engineering attacks, a situation significantly worsened by the advent of audio Real-time Deepfakes (RTDFs). RTDFs can clone a target's voice in real-time over phone calls, making these interactions highly interactive and thus far more convincing. Our research confidently addresses the gap in the existing literature on deepfake detection, which has largely been ineffective against RTDF threats. We introduce a robust challenge-response-based method to detect deepfake audio calls, pioneering a comprehensive taxonomy of audio challenges. Our evaluation pitches 20 prospective challenges against a leading voice-cloning system. We have compiled a novel open-source challenge dataset with contributions from 100 smartphone and desktop users, yielding 18,600 original and 1.6 million deepfake samples. Through rigorous machine and human evaluations of this dataset, we achieved a deepfake detection rate of 86% and an 80% AUC score, respectively. Notably, utilizing a set of 11 challenges significantly enhances detection capabilities. Our findings reveal that combining human intuition with machine precision offers complementary advantages. Consequently, we have developed an innovative human-AI collaborative system that melds human discernment with algorithmic accuracy, boosting final joint accuracy to 82.9%. This system highlights the significant advantage of AI-assisted pre-screening in call verification processes. Samples can be heard at https://mittalgovind.github.io/autch-samples/

【9】Mixer is more than just a model
标题:搅拌机不仅仅是一种模式
链接:https://arxiv.org/abs/2402.18007
作者:Qingfeng Ji,Yuxin Wang,Letong Sun
摘要:最近,MLP结构重新流行起来,MLP-Mixer就是一个突出的例子。在计算机视觉领域,MLP Mixer以其从通道和令牌角度提取数据信息的能力而闻名,有效地充当通道和令牌信息的融合。事实上,Mixer代表了一种融合通道和令牌信息的信息提取范例。Mixer的本质在于它能够混合来自不同视角的信息,集中体现了神经网络架构领域中“混合”的真正概念。除了通道和令牌考虑之外,还可以从各种角度创建更定制的混合器,以更好地满足特定的任务要求。本研究的重点是音频识别领域,介绍了一种新的模型命名为音频频谱图混合器与滚动时间和Hermit FFT(ASM-RH),结合了从时域和频域的见解。实验结果表明,ASM-RH特别适合音频数据,并在多个分类任务中产生有希望的结果。
摘要:Recently, MLP structures have regained popularity, with MLP-Mixer standing out as a prominent example. In the field of computer vision, MLP-Mixer is noted for its ability to extract data information from both channel and token perspectives, effectively acting as a fusion of channel and token information. Indeed, Mixer represents a paradigm for information extraction that amalgamates channel and token information. The essence of Mixer lies in its ability to blend information from diverse perspectives, epitomizing the true concept of "mixing" in the realm of neural network architectures. Beyond channel and token considerations, it is possible to create more tailored mixers from various perspectives to better suit specific task requirements. This study focuses on the domain of audio recognition, introducing a novel model named Audio Spectrogram Mixer with Roll-Time and Hermit FFT (ASM-RH) that incorporates insights from both time and frequency domains. Experimental results demonstrate that ASM-RH is particularly well-suited for audio data and yields promising outcomes across multiple classification tasks.

【10】ByteComposer: a Human-like Melody Composition Method based on Language  Model Agent
标题:ByteComposer:一种基于语言模型代理的仿人旋律合成方法
链接:https://arxiv.org/abs/2402.17785
作者:Xia Liang,Jiaju Lin,Xinjian Du
摘要:大型语言模型(LLM)在多模态理解和生成任务方面取得了令人鼓舞的进展。然而,如何设计一个符合人类和可解释的旋律创作系统仍然是探索不足。为了解决这个问题,我们提出了ByteComposer,这是一个代理框架,它通过四个独立的步骤来模拟人类的创造性管道:“概念分析-草稿组成-自我评估和修改-美学选择”。该框架无缝地融合了LLM的交互和知识理解功能与现有的符号音乐生成模型,从而实现了与人类创作者相媲美的旋律合成代理。我们进行了广泛的实验GPT 4和几个开源的大型语言模型,这证实了我们的框架的有效性。此外,专业音乐作曲家进行了多维度的评估,最终的结果表明,在音乐创作的各个方面,ByteComposer代理达到了新手旋律作曲家的水平。
摘要:Large Language Models (LLM) have shown encouraging progress in multimodal understanding and generation tasks. However, how to design a human-aligned and interpretable melody composition system is still under-explored. To solve this problem, we propose ByteComposer, an agent framework emulating a human's creative pipeline in four separate steps : "Conception Analysis - Draft Composition - Self-Evaluation and Modification - Aesthetic Selection". This framework seamlessly blends the interactive and knowledge-understanding features of LLMs with existing symbolic music generation models, thereby achieving a melody composition agent comparable to human creators. We conduct extensive experiments on GPT4 and several open-source large language models, which substantiate our framework's effectiveness. Furthermore, professional music composers were engaged in multi-dimensional evaluations, the final results demonstrated that across various facets of music composition, ByteComposer agent attains the level of a novice melody composer.


【11】Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion  Latent Aligners
标题:视听:基于扩散潜伏期对准的开放视听生成
链接:https://arxiv.org/abs/2402.17723
作者:Yazhou Xing,Yingqing He,Zeyue Tian,Xintao Wang,Qifeng Chen
备注:Accepted to CVPR 2024. Project website: this https URL
摘要:视频和音频内容创作是电影行业和专业用户的核心技术。现有的基于扩散的视频和音频生成方法是分开处理的,这阻碍了技术从学术界向工业界的转移。在这项工作中,我们的目标是填补空白,精心设计的优化为基础的框架,跨视听和联合视听生成。我们观察到强大的生成能力的现成的视频或音频生成模型。因此,我们建议用共享的潜在表示空间来桥接现有的强模型,而不是从头开始训练巨型模型。具体来说,我们提出了一个多模态潜在对齐器与预训练的ImageBind模型。我们的潜在对齐器与分类器指导共享类似的核心,该分类器指导在推理时间期间指导扩散去噪过程。通过精心设计的优化策略和损失函数,我们展示了我们的方法在联合视频-音频生成,视觉导向音频生成和音频导向视觉生成任务上的优越性能。该项目的网站可以在https://yzxing87.github.io/Seeing-and-Hearing/上找到
摘要:Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from academia to industry. In this work, we aim at filling the gap, with a carefully designed optimization-based framework for cross-visual-audio and joint-visual-audio generation. We observe the powerful generation ability of off-the-shelf video or audio generation models. Thus, instead of training the giant models from scratch, we propose to bridge the existing strong models with a shared latent representation space. Specifically, we propose a multimodality latent aligner with the pre-trained ImageBind model. Our latent aligner shares a similar core as the classifier guidance that guides the diffusion denoising process during inference time. Through carefully designed optimization strategy and loss functions, we show the superior performance of our method on joint video-audio generation, visual-steered audio generation, and audio-steered visual generation tasks. The project website can be found at https://yzxing87.github.io/Seeing-and-Hearing/


机器翻译由腾讯交互翻译提供,仅供参考