今日论文合集:cs.SD语音11篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Deep functional multiple index models with an application to SER

标题:深度函数多指标模型及其在SER中的应用

链接:https://arxiv.org/abs/2403.17562

作者:Matthieu Saumard,Abir El Haj,Thibault Napoleon

备注:5 pages, 1 figure

摘要:语音情感识别(SER)在提高人机交互和语音处理能力方面起着至关重要的作用。我们介绍了一种专门为函数数据模型设计的新型深度学习架构,称为多索引函数模型。我们的主要创新在于将自适应基础层和自动数据转换搜索集成到深度学习框架中。仿真结果表明,该模型具有良好的性能。这使我们能够提取适合块级SER的功能,基于梅尔频率倒谱系数(MFCC)。我们证明了我们的方法在基准IEMOCAP数据库上的有效性,与现有的方法相比,取得了良好的性能。

摘要:Speech Emotion Recognition (SER) plays a crucial role in advancing human-computer interaction and speech processing capabilities. We introduce a novel deep-learning architecture designed specifically for the functional data model known as the multiple-index functional model. Our key innovation lies in integrating adaptive basis layers and an automated data transformation search within the deep learning framework. Simulations for this new model show good performances. This allows us to extract features tailored for chunk-level SER, based on Mel Frequency Cepstral Coefficients (MFCCs). We demonstrate the effectiveness of our approach on the benchmark IEMOCAP database, achieving good performance compared to existing methods.


【2】 Detection of Deepfake Environmental Audio
标题:检测Deepfake环境音频
链接:https://arxiv.org/abs/2403.17529
作者:Hafsa Ouajdi,Oussama Hadder,Modan Tailleur,Mathieu Lagrange,Laurie M. Heller
摘要:随着深度生成模型的质量不断提高,能够辨别手头的音频数据是记录的还是合成的变得越来越重要。虽然已经广泛地研究了伪语音信号的检测,但是对于伪环境音频的检测并非如此。  我们提出了一个简单而有效的管道,用于检测基于CLAP音频嵌入的假环境声音。我们使用来自2023年DCASE挑战任务的音频数据对该检测器进行评估。  我们的实验表明,由44个最先进的合成器生成的假声音平均可以检测到98%的准确率。我们表明,使用在环境音频上学习的音频嵌入比标准的VGGish音频嵌入更有益,因为它提供了10%的检测性能提高。非正式聆听不正确的否定示例展示了检测器错过的假声音的可听特征,例如失真和难以置信的背景噪声。
摘要:With the ever-rising quality of deep generative models, it is increasingly important to be able to discern whether the audio data at hand have been recorded or synthesized. Although the detection of fake speech signals has been studied extensively, this is not the case for the detection of fake environmental audio.  We propose a simple and efficient pipeline for detecting fake environmental sounds based on the CLAP audio embedding. We evaluate this detector using audio data from the 2023 DCASE challenge task on Foley sound synthesis.  Our experiments show that fake sounds generated by 44 state-of-the-art synthesizers can be detected on average with 98% accuracy. We show that using an audio embedding learned on environmental audio is beneficial over a standard VGGish one as it provides a 10% increase in detection performance. Informal listening to Incorrect Negative examples demonstrates audible features of fake sounds missed by the detector such as distortion and implausible background noise.

【3】 Correlation of Fréchet Audio Distance With Human Perception of  Environmental Audio Is Embedding Dependant
标题:Fréchet音频距离与人对环境音频感知的相关性是嵌入依赖关系
链接:https://arxiv.org/abs/2403.17508
作者:Modan Tailleur,Junwon Lee,Mathieu Lagrange,Keunwoo Choi,Laurie M. Heller,Keisuke Imoto,Yuki Okamoto
摘要:本文探讨了是否考虑替代特定领域的嵌入来计算Fr\'echet音频距离(FAD)度量可以帮助FAD更好地与环境声音的感知评级相关。我们使用了VGGish、PANN、MS-CLAP、L-CLAP和MERT的嵌入,这些嵌入是为音乐或环境声音评估量身定制的。FAD分数是针对来自DCASE 2023任务7数据集的声音计算的。使用来自相同任务的感知数据,我们发现PANNs-WGM-LogMel产生FAD分数与音频质量和感知拟合的感知评级之间的最佳相关性,Spearman相关性高于0.5。我们还发现,音乐特定的嵌入导致显着较低的结果。有趣的是,VGGish,用于原始Fr\'echet计算的嵌入,产生了低于0.1的相关性。这些结果强调了嵌入的FAD度量设计的选择至关重要。
摘要:This paper explores whether considering alternative domain-specific embeddings to calculate the Fr\'echet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings from VGGish, PANNs, MS-CLAP, L-CLAP, and MERT, which are tailored for either music or environmental sound evaluation. The FAD scores were calculated for sounds from the DCASE 2023 Task 7 dataset. Using perceptual data from the same task, we find that PANNs-WGM-LogMel produces the best correlation between FAD scores and perceptual ratings of both audio quality and perceived fit with a Spearman correlation higher than 0.5. We also find that music-specific embeddings resulted in significantly lower results. Interestingly, VGGish, the embedding used for the original Fr\'echet calculation, yielded a correlation below 0.1. These results underscore the critical importance of the choice of embedding for the FAD metric design.


【4】 Learning to Visually Localize Sound Sources from Mixtures without Prior  Source Knowledge
标题:学习在没有先验声源知识的情况下从混合声中视觉定位声源
链接:https://arxiv.org/abs/2403.17420
作者:Dongjin Kim,Sung Jin Um,Sangmin Lee,Jung Uk Kim
备注:Accepted at CVPR 2024
摘要:多声源定位任务的目标是从混合中单独定位声源。虽然最近的多声源定位方法已经显示出改进的性能,但由于它们依赖于关于要分离的对象的数量的先验信息,因此它们面临挑战。在本文中,为了克服这一限制,我们提出了一种新的多声源定位方法,可以进行定位,而无需先验知识的声源的数量。为了实现这一目标,我们提出了一个迭代对象识别(IOI)模块,它可以识别发声对象的迭代方式。在找到发声对象的区域后,我们设计了对象相似性感知聚类(OSC)损失,以指导IOI模块有效地组合同一对象的区域,但也区分不同的对象和背景。它使我们的方法能够在没有任何先验知识的情况下对发声对象进行准确的定位。在MUSIC和VGGSound基准测试上的大量实验结果表明,与现有的单源和多源方法相比,所提出的方法具有显著的性能改进。我们的代码可在:https://github.com/VisualAIKHU/NoPrior_MultiSSL
摘要:The goal of the multi-sound source localization task is to localize sound sources from the mixture individually. While recent multi-sound source localization methods have shown improved performance, they face challenges due to their reliance on prior information about the number of objects to be separated. In this paper, to overcome this limitation, we present a novel multi-sound source localization method that can perform localization without prior knowledge of the number of sound sources. To achieve this goal, we propose an iterative object identification (IOI) module, which can recognize sound-making objects in an iterative manner. After finding the regions of sound-making objects, we devise object similarity-aware clustering (OSC) loss to guide the IOI module to effectively combine regions of the same object but also distinguish between different objects and backgrounds. It enables our method to perform accurate localization of sound-making objects without any prior knowledge. Extensive experimental results on the MUSIC and VGGSound benchmarks show the significant performance improvements of the proposed method over the existing methods for both single and multi-source. Our code is available at: https://github.com/VisualAIKHU/NoPrior_MultiSSL


【5】 Exploring and Applying Audio-Based Sentiment Analysis in Music
标题:基于音频的情感分析在音乐中的探索与应用
链接:https://arxiv.org/abs/2403.17379
作者:Etash Jhanji
备注:5 pages, 7 figures, 2 tables. For source code, see this https URL
摘要:情感分析是文本处理的一个不断探索的领域,涉及文本的意见,情感和主观性的计算分析。然而,这一思想并不局限于文本和语音,事实上,它可以应用于其他模式。事实上,人类在文本中表达自己的方式并不像在音乐中那样深刻。计算模型解释音乐情感的能力在很大程度上是未开发的,可能在治疗和音乐排队中有影响和用途。在本文中,两个单独的任务得到解决。本研究旨在(1)预测音乐片段随时间的情感,(2)确定音乐之后的下一个情感值,以确保无缝过渡。利用来自音乐数据库中的情感数据,其中包含从免费音乐档案中选择的歌曲片段,并根据多名志愿者在Russel的circumplex情感模型中报告的效价和唤醒水平进行注释,模型被训练用于这两项任务。总的来说,这些模型的性能反映了它们能够有效和准确地执行设计任务。
摘要:Sentiment analysis is a continuously explored area of text processing that deals with the computational analysis of opinions, sentiments, and subjectivity of text. However, this idea is not limited to text and speech, in fact, it could be applied to other modalities. In reality, humans do not express themselves in text as deeply as they do in music. The ability of a computational model to interpret musical emotions is largely unexplored and could have implications and uses in therapy and musical queuing. In this paper, two individual tasks are addressed. This study seeks to (1) predict the emotion of a musical clip over time and (2) determine the next emotion value after the music in a time series to ensure seamless transitions. Utilizing data from the Emotions in Music Database, which contains clips of songs selected from the Free Music Archive annotated with levels of valence and arousal as reported on Russel's circumplex model of affect by multiple volunteers, models are trained for both tasks. Overall, the performance of these models reflected that they were able to perform the tasks they were designed for effectively and accurately.


【6】 Low-Latency Neural Speech Phase Prediction based on Parallel Estimation  Architecture and Anti-Wrapping Losses for Speech Generation Tasks
标题:基于并行估计结构和抗包裹损耗的语音生成任务低延迟神经语音相位预测
链接:https://arxiv.org/abs/2403.17378
作者:Yang Ai,Zhen-Hua Ling
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing. arXiv admin note: substantial text overlap with arXiv:2211.15974
摘要:本文提出了一种新的神经语音相位预测模型,它直接从幅度谱预测包裹相位谱。所提出的模型是一个级联的残差卷积网络和并行估计架构。并行估计结构是直接包裹相位预测的核心模块。该架构由两个并行的线性卷积层和一个相位计算公式组成,模仿了从复谱的实部和虚部计算相位谱的过程,并将预测的相位值严格限制在主值区间。为了避免相位缠绕引起的误差扩展问题,我们设计了反缠绕训练损耗,其定义在预测的缠绕相位谱和自然相位谱之间,通过使用反缠绕函数激活瞬时相位误差、群延迟误差和瞬时角频率误差。我们从数学上证明了反包裹函数应具有三个性质,即奇偶性、周期性和单调性。我们还通过结合因果卷积和知识蒸馏训练策略来实现低延迟的可流相位预测。对于分析合成和特定语音生成任务,实验结果表明,我们提出的神经语音相位预测模型优于迭代相位估计算法和基于神经网络的相位预测方法的相位预测精度,效率和鲁棒性。与基于HiFi-GAN的波形重构方法相比,该模型在保证合成语音质量的同时,也显示出了显著的效率优势。据我们所知,我们是第一个直接预测语音相位谱的幅度谱,只有通过神经网络。
摘要:This paper presents a novel neural speech phase prediction model which predicts wrapped phase spectra directly from amplitude spectra. The proposed model is a cascade of a residual convolutional network and a parallel estimation architecture. The parallel estimation architecture is a core module for direct wrapped phase prediction. This architecture consists of two parallel linear convolutional layers and a phase calculation formula, imitating the process of calculating the phase spectra from the real and imaginary parts of complex spectra and strictly restricting the predicted phase values to the principal value interval. To avoid the error expansion issue caused by phase wrapping, we design anti-wrapping training losses defined between the predicted wrapped phase spectra and natural ones by activating the instantaneous phase error, group delay error and instantaneous angular frequency error using an anti-wrapping function. We mathematically demonstrate that the anti-wrapping function should possess three properties, namely parity, periodicity and monotonicity. We also achieve low-latency streamable phase prediction by combining causal convolutions and knowledge distillation training strategies. For both analysis-synthesis and specific speech generation tasks, experimental results show that our proposed neural speech phase prediction model outperforms the iterative phase estimation algorithms and neural network-based phase prediction methods in terms of phase prediction precision, efficiency and robustness. Compared with HiFi-GAN-based waveform reconstruction method, our proposed model also shows outstanding efficiency advantages while ensuring the quality of synthesized speech. To the best of our knowledge, we are the first to directly predict speech phase spectra from amplitude spectra only via neural networks.


【7】 Theoretical Analysis of Quality of Conventional Beamforming for Phased  Microphone Arrays
标题:相控阵传声器阵列常规波束形成质量的理论分析
链接:https://arxiv.org/abs/2403.17376
作者:Dheepak Khatri,Kenneth Granlund
摘要:理论研究进行分析不同类型的麦克风阵列设计的方向性响应。一维(线性)和二维(平面)麦克风阵列类型被认为是,和延迟和总和波束形成和传统的波束形成技术被用来定位声源。无量纲参数G的特征在于简化和标准化作为阵列几何形状和声源参数的函数的1-D和2-D麦克风阵列的抑制性能。该参数G然后用于确定用于远场声音定位的2-D麦克风阵列的改进设计。介绍并详细分析了一种称为等面积阵列的设计。该设计示出具有有利的抑制性能相比,其他传统使用的2-D平面麦克风阵列。
摘要:A theoretical study is performed to analyze the directional response of different types of microphone array designs. 1-D (linear) and 2-D (planar) microphone array types are considered, and the delay and sum beamforming and conventional beamforming techniques are employed to localize the sound source. A non-dimensional parameter, G, is characterized to simplify and standardize the rejection performance of both 1-D and 2-D microphone arrays as a function of array geometry and sound source parameters. This parameter G is then used to determine an improved design of a 2-D microphone array for far-field sound localization. One such design, termed the Equi-area array is introduced and analyzed in detail. The design is shown to have an advantageous rejection performance compared to other conventionally used 2-D planar microphone arrays.

【8】 Accuracy enhancement method for speech emotion recognition from  spectrogram using temporal frequency correlation and positional information  learning through knowledge transfer
标题:基于时频相关和知识转移位置信息学习的语谱图语音情感识别准确率提高方法
链接:https://arxiv.org/abs/2403.17327
作者:Jeong-Yoon Kim,Seung-Ho Lee
摘要:本文提出了一种利用Vision Transformer(ViT)处理语音谱图中频率(y轴)和时间(x轴)相关性,并通过知识传递在ViT之间传递位置信息,从而提高语音情感识别准确率的方法。所提出的方法具有以下独创性:i)我们使用垂直分割的对数梅尔频谱图的补丁来分析频率随时间的相关性。这种类型的补丁允许我们将特定情绪的最相关频率与它们发出的时间相关联。ii)我们提出使用图像坐标编码,这是一种适用于ViT的绝对位置编码。通过将图像的x,y坐标归一化为-1至1并将它们连接到图像,我们可以有效地为ViT提供有效的绝对位置信息。iii)通过特征图匹配,教师网络的地点和位置信息被有效地传输到学生网络。教师网络是通过图像坐标编码包含卷积干的局部性和绝对位置信息的ViT,而学生网络是在基本ViT结构中缺乏位置编码的结构。在特征图匹配阶段,我们通过平均绝对误差(L1损失)进行训练,以最小化两个网络的特征图之间的差异。为了验证所提出的方法,三个情感数据集(SAVEE,CARDB,和CREMA-D)组成的语音转换成对数梅尔频谱进行比较实验。实验结果表明,所提出的方法显着优于国家的最先进的方法在加权精度方面,同时需要显着减少浮点运算(FLOPs)。总的来说,所提出的方法提供了一个有前途的解决方案,通过提供更高的效率和性能的SER。
摘要:In this paper, we propose a method to improve the accuracy of speech emotion recognition (SER) by using vision transformer (ViT) to attend to the correlation of frequency (y-axis) with time (x-axis) in spectrogram and transferring positional information between ViT through knowledge transfer. The proposed method has the following originality i) We use vertically segmented patches of log-Mel spectrogram to analyze the correlation of frequencies over time. This type of patch allows us to correlate the most relevant frequencies for a particular emotion with the time they were uttered. ii) We propose the use of image coordinate encoding, an absolute positional encoding suitable for ViT. By normalizing the x, y coordinates of the image to -1 to 1 and concatenating them to the image, we can effectively provide valid absolute positional information for ViT. iii) Through feature map matching, the locality and location information of the teacher network is effectively transmitted to the student network. Teacher network is a ViT that contains locality of convolutional stem and absolute position information through image coordinate encoding, and student network is a structure that lacks positional encoding in the basic ViT structure. In feature map matching stage, we train through the mean absolute error (L1 loss) to minimize the difference between the feature maps of the two networks. To validate the proposed method, three emotion datasets (SAVEE, EmoDB, and CREMA-D) consisting of speech were converted into log-Mel spectrograms for comparison experiments. The experimental results show that the proposed method significantly outperforms the state-of-the-art methods in terms of weighted accuracy while requiring significantly fewer floating point operations (FLOPs). Overall, the proposed method offers an promising solution for SER by providing improved efficiency and performance.

【9】 Synthesizing Soundscapes: Leveraging Text-to-Audio Models for  Environmental Sound Classification
标题:合成声景:利用文本到音频模型进行环境声分类
链接:https://arxiv.org/abs/2403.17864
作者:Francesca Ronchini,Luca Comanducci,Fabio Antonacci
备注:Submitted to EUSIPCO 2024
摘要:在过去的几年里,文本到音频模型已经成为自动音频生成的一个重大进步。虽然它们代表了令人印象深刻的技术进步,但它们在音频应用开发中的使用效果仍然不确定。本文旨在研究这些方面,特别是集中在环境声音的分类任务。本研究分析了两种不同的环境分类系统的性能时,从文本到音频模型生成的数据用于训练。考虑两种情况:a)当训练数据集由来自两个不同的文本到音频模型的数据增强时;以及b)当训练数据集仅由生成的合成音频组成时。在这两种情况下,分类任务的性能在真实数据上进行测试。结果表明,文本到音频模型是有效的数据集增强,而模型的性能下降时,只依赖于生成的音频。
摘要:In the past few years, text-to-audio models have emerged as a significant advancement in automatic audio gener- ation. Although they represent impressive technological progress, the effectiveness of their use in the development of audio applications remains uncertain. This paper aims to investigate these aspects, specifically focusing on the task of classification of environmental sounds. This study analyzes the performance of two different environmental classification systems when data generated from text-to-audio models is used for training. Two cases are considered: a) when the training dataset is augmented by data coming from two different text-to-audio models; and b) when the training dataset consists solely of synthetic audio generated. In both cases, the performance of the classification task is tested on real data. Results indicate that text-to-audio models are effective for dataset augmentation, whereas the performance of the models drops when relying on only generated audio.


【10】 Speaker Distance Estimation in Enclosures from Single-Channel Audio
标题:基于单声道音频的音箱说话人距离估计
链接:https://arxiv.org/abs/2403.17514
作者:Michael Neri,Archontis Politis,Daniel Krause,Marco Carli,Tuomas Virtanen
备注:Accepted for publication in IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:音频距离估计在声学场景分析、声源定位和房间建模等各种应用中起着至关重要的作用。大多数研究主要集中在采用分类方法,其中距离被离散化为不同的类别,从而实现更平滑的模型训练并实现更高的准确性,但对所获得的声源位置的精度施加限制。朝着这个方向,在本文中,我们提出了一种新的方法,使用卷积递归神经网络与注意力模块的音频信号的连续距离估计。注意力机制使模型能够专注于相关的时间和光谱特征,增强其捕获细粒度距离相关信息的能力。为了评估我们所提出的方法的有效性,我们在四个数据集(我们的合成数据集,QMULTIMIT,VoiceHome-2和STARSS 23)上使用三个真实度(合成房间脉冲响应,卷积语音测量响应和真实录音)的受控环境中的音频记录进行了广泛的实验。实验结果表明,该模型实现了0.11米的绝对误差在无噪声的合成场景。此外,结果显示,在混合场景中,绝对误差约为1.30米。该算法在真实场景中的性能,其中不可预测的环境因素和噪声是普遍的,产生约0.50米的绝对误差。出于可重复研究的目的,我们在https://github.com/michaelneri/audio-distance-estimation上提供模型,代码和合成数据集。
摘要:Distance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where distances are discretized into distinct categories, enabling smoother model training and achieving higher accuracy but imposing restrictions on the precision of the obtained sound source position. Towards this direction, in this paper we propose a novel approach for continuous distance estimation from audio signals using a convolutional recurrent neural network with an attention module. The attention mechanism enables the model to focus on relevant temporal and spectral features, enhancing its ability to capture fine-grained distance-related information. To evaluate the effectiveness of our proposed method, we conduct extensive experiments using audio recordings in controlled environments with three levels of realism (synthetic room impulse response, measured response with convolved speech, and real recordings) on four datasets (our synthetic dataset, QMULTIMIT, VoiceHome-2, and STARSS23). Experimental results show that the model achieves an absolute error of 0.11 meters in a noiseless synthetic scenario. Moreover, the results showed an absolute error of about 1.30 meters in the hybrid scenario. The algorithm's performance in the real scenario, where unpredictable environmental factors and noise are prevalent, yields an absolute error of approximately 0.50 meters. For reproducible research purposes we make model, code, and synthetic datasets available at https://github.com/michaelneri/audio-distance-estimation.

【11】 Infrastructure-less Localization from Indoor Environmental Sounds Based  on Spectral Decomposition and Spatial Likelihood Model
标题:基于频谱分解和空间似然模型的室内环境声无基础设施定位
链接:https://arxiv.org/abs/2403.17402
作者:Satoki Ogiso,Yoshiaki Bando,Takeshi Kurata,Takashi Okuma
备注:6 pages, 6 figures, accepted to IEEE/SICE SII 2023
摘要:使用所附接的传感器单元的人和/或资产跟踪有助于了解他们的活动。用于人体跟踪技术的最常见的室内定位方法需要昂贵的基础设施、部署和维护。为了克服这个问题,环境声音已经被用于无基础设施的定位。虽然它们实现了房间级分类,但它们存在两个问题:低信噪比(SNR)条件和覆盖区域内声音的非唯一性。针对这些问题,提出了一种基于监督谱分解和空间似然的麦克风定位方法。所提出的方法进行了评估与实际记录在一个实验室的大小为12 × 30米。实验结果表明,与简单的特征(梅尔倒谱系数:MFCC)相比,在低信噪比条件下,该方法具有较好的鲁棒性。此外,所提出的方法可以很容易地集成到先验分布,这是从其他贝叶斯定位。该方法可用于评估环境声音的空间似然性。
摘要:Human and/or asset tracking using an attached sensor units helps understand their activities. Most common indoor localization methods for human tracking technologies require expensive infrastructures, deployment and maintenance. To overcome this problem, environmental sounds have been used for infrastructure-free localization. While they achieve room-level classification, they suffer from two problems: low signal-to-noise-ratio (SNR) condition and non-uniqueness of sound over the coverage area. A microphone localization method was proposed using supervised spectral decomposition and spatial likelihood to solve these problems. The proposed method was evaluated with actual recordings in an experimental room with a size of 12 x 30 m. The results showed that the proposed method with supervised NMF was robust under low-SNR condition compared to a simple feature (mel frequency cepstrum coefficient: MFCC). Additionally, the proposed method could be easily integrated with prior distribution, which is available from other Bayesian localizations. The proposed method can be used to evaluate the spatial likelihood from environmental sounds.

eess.AS音频处理
【1】 Synthesizing Soundscapes: Leveraging Text-to-Audio Models for  Environmental Sound Classification
标题:合成声景:利用文本到音频模型进行环境声音分类
链接:https://arxiv.org/abs/2403.17864
作者:Francesca Ronchini,Luca Comanducci,Fabio Antonacci
备注:Submitted to EUSIPCO 2024
摘要:在过去的几年里,文本到音频模型已经成为自动音频生成的一个重大进步。虽然它们代表了令人印象深刻的技术进步,但它们在音频应用开发中的使用效果仍然不确定。本文旨在研究这些方面,特别是集中在环境声音的分类任务。本研究分析了两种不同的环境分类系统的性能时,从文本到音频模型生成的数据用于训练。考虑两种情况:a)当训练数据集由来自两个不同的文本到音频模型的数据增强时;以及b)当训练数据集仅由生成的合成音频组成时。在这两种情况下,分类任务的性能在真实数据上进行测试。结果表明,文本到音频模型是有效的数据集增强,而模型的性能下降时,只依赖于生成的音频。
摘要:In the past few years, text-to-audio models have emerged as a significant advancement in automatic audio gener- ation. Although they represent impressive technological progress, the effectiveness of their use in the development of audio applications remains uncertain. This paper aims to investigate these aspects, specifically focusing on the task of classification of environmental sounds. This study analyzes the performance of two different environmental classification systems when data generated from text-to-audio models is used for training. Two cases are considered: a) when the training dataset is augmented by data coming from two different text-to-audio models; and b) when the training dataset consists solely of synthetic audio generated. In both cases, the performance of the classification task is tested on real data. Results indicate that text-to-audio models are effective for dataset augmentation, whereas the performance of the models drops when relying on only generated audio.

【2】 Speaker Distance Estimation in Enclosures from Single-Channel Audio
标题:基于单声道音频的音箱说话人距离估计
链接:https://arxiv.org/abs/2403.17514
作者:Michael Neri,Archontis Politis,Daniel Krause,Marco Carli,Tuomas Virtanen
备注:Accepted for publication in IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:音频距离估计在声学场景分析、声源定位和房间建模等各种应用中起着至关重要的作用。大多数研究主要集中在采用分类方法,其中距离被离散化为不同的类别,从而实现更平滑的模型训练并实现更高的准确性,但对所获得的声源位置的精度施加限制。朝着这个方向,在本文中,我们提出了一种新的方法,使用卷积递归神经网络与注意力模块的音频信号的连续距离估计。注意力机制使模型能够专注于相关的时间和光谱特征,增强其捕获细粒度距离相关信息的能力。为了评估我们所提出的方法的有效性,我们在四个数据集(我们的合成数据集,QMULTIMIT,VoiceHome-2和STARSS 23)上使用三个真实度(合成房间脉冲响应,卷积语音测量响应和真实录音)的受控环境中的音频记录进行了广泛的实验。实验结果表明,该模型实现了0.11米的绝对误差在无噪声的合成场景。此外,结果显示,在混合场景中,绝对误差约为1.30米。该算法在真实场景中的性能,其中不可预测的环境因素和噪声是普遍的,产生约0.50米的绝对误差。出于可重复研究的目的,我们在https://github.com/michaelneri/audio-distance-estimation上提供模型,代码和合成数据集。
摘要:Distance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where distances are discretized into distinct categories, enabling smoother model training and achieving higher accuracy but imposing restrictions on the precision of the obtained sound source position. Towards this direction, in this paper we propose a novel approach for continuous distance estimation from audio signals using a convolutional recurrent neural network with an attention module. The attention mechanism enables the model to focus on relevant temporal and spectral features, enhancing its ability to capture fine-grained distance-related information. To evaluate the effectiveness of our proposed method, we conduct extensive experiments using audio recordings in controlled environments with three levels of realism (synthetic room impulse response, measured response with convolved speech, and real recordings) on four datasets (our synthetic dataset, QMULTIMIT, VoiceHome-2, and STARSS23). Experimental results show that the model achieves an absolute error of 0.11 meters in a noiseless synthetic scenario. Moreover, the results showed an absolute error of about 1.30 meters in the hybrid scenario. The algorithm's performance in the real scenario, where unpredictable environmental factors and noise are prevalent, yields an absolute error of approximately 0.50 meters. For reproducible research purposes we make model, code, and synthetic datasets available at https://github.com/michaelneri/audio-distance-estimation.


【3】 Infrastructure-less Localization from Indoor Environmental Sounds Based  on Spectral Decomposition and Spatial Likelihood Model
标题:基于谱分解和空间似然模型的室内环境声无结构定位
链接:https://arxiv.org/abs/2403.17402
作者:Satoki Ogiso,Yoshiaki Bando,Takeshi Kurata,Takashi Okuma
备注:6 pages, 6 figures, accepted to IEEE/SICE SII 2023
摘要:使用所附接的传感器单元的人和/或资产跟踪有助于了解他们的活动。用于人体跟踪技术的最常见的室内定位方法需要昂贵的基础设施、部署和维护。为了克服这个问题,环境声音已经被用于无基础设施的定位。虽然它们实现了房间级分类,但它们存在两个问题:低信噪比(SNR)条件和覆盖区域内声音的非唯一性。针对这些问题,提出了一种基于监督谱分解和空间似然的麦克风定位方法。所提出的方法进行了评估与实际记录在一个实验室的大小为12 × 30米。实验结果表明,与简单的特征(梅尔倒谱系数:MFCC)相比,在低信噪比条件下,该方法具有较好的鲁棒性。此外,所提出的方法可以很容易地集成到先验分布,这是从其他贝叶斯定位。该方法可用于评估环境声音的空间似然性。
摘要:Human and/or asset tracking using an attached sensor units helps understand their activities. Most common indoor localization methods for human tracking technologies require expensive infrastructures, deployment and maintenance. To overcome this problem, environmental sounds have been used for infrastructure-free localization. While they achieve room-level classification, they suffer from two problems: low signal-to-noise-ratio (SNR) condition and non-uniqueness of sound over the coverage area. A microphone localization method was proposed using supervised spectral decomposition and spatial likelihood to solve these problems. The proposed method was evaluated with actual recordings in an experimental room with a size of 12 x 30 m. The results showed that the proposed method with supervised NMF was robust under low-SNR condition compared to a simple feature (mel frequency cepstrum coefficient: MFCC). Additionally, the proposed method could be easily integrated with prior distribution, which is available from other Bayesian localizations. The proposed method can be used to evaluate the spatial likelihood from environmental sounds.


【4】 Deep functional multiple index models with an application to SER
标题:深度泛函多指标模型及其在SER中的应用
链接:https://arxiv.org/abs/2403.17562
作者:Matthieu Saumard,Abir El Haj,Thibault Napoleon
备注:5 pages, 1 figure
摘要:语音情感识别(SER)在提高人机交互和语音处理能力方面起着至关重要的作用。我们介绍了一种专门为函数数据模型设计的新型深度学习架构,称为多索引函数模型。我们的主要创新在于将自适应基础层和自动数据转换搜索集成到深度学习框架中。仿真结果表明,该模型具有良好的性能。这使我们能够提取适合块级SER的功能,基于梅尔频率倒谱系数(MFCC)。我们证明了我们的方法在基准IEMOCAP数据库上的有效性,与现有的方法相比,取得了良好的性能。
摘要:Speech Emotion Recognition (SER) plays a crucial role in advancing human-computer interaction and speech processing capabilities. We introduce a novel deep-learning architecture designed specifically for the functional data model known as the multiple-index functional model. Our key innovation lies in integrating adaptive basis layers and an automated data transformation search within the deep learning framework. Simulations for this new model show good performances. This allows us to extract features tailored for chunk-level SER, based on Mel Frequency Cepstral Coefficients (MFCCs). We demonstrate the effectiveness of our approach on the benchmark IEMOCAP database, achieving good performance compared to existing methods.


【5】 Detection of Deepfake Environmental Audio
标题:检测Deepfake环境音频
链接:https://arxiv.org/abs/2403.17529
作者:Hafsa Ouajdi,Oussama Hadder,Modan Tailleur,Mathieu Lagrange,Laurie M. Heller
摘要:随着深度生成模型的质量不断提高,能够辨别手头的音频数据是记录的还是合成的变得越来越重要。虽然已经广泛地研究了伪语音信号的检测,但是对于伪环境音频的检测并非如此。  我们提出了一个简单而有效的管道,用于检测基于CLAP音频嵌入的假环境声音。我们使用来自2023年DCASE挑战任务的音频数据对该检测器进行评估。  我们的实验表明,由44个最先进的合成器生成的假声音平均可以检测到98%的准确率。我们表明,使用在环境音频上学习的音频嵌入比标准的VGGish音频嵌入更有益,因为它提供了10%的检测性能提高。非正式聆听不正确的否定示例展示了检测器错过的假声音的可听特征,例如失真和难以置信的背景噪声。
摘要:With the ever-rising quality of deep generative models, it is increasingly important to be able to discern whether the audio data at hand have been recorded or synthesized. Although the detection of fake speech signals has been studied extensively, this is not the case for the detection of fake environmental audio.  We propose a simple and efficient pipeline for detecting fake environmental sounds based on the CLAP audio embedding. We evaluate this detector using audio data from the 2023 DCASE challenge task on Foley sound synthesis.  Our experiments show that fake sounds generated by 44 state-of-the-art synthesizers can be detected on average with 98% accuracy. We show that using an audio embedding learned on environmental audio is beneficial over a standard VGGish one as it provides a 10% increase in detection performance. Informal listening to Incorrect Negative examples demonstrates audible features of fake sounds missed by the detector such as distortion and implausible background noise.

【6】 Correlation of Fréchet Audio Distance With Human Perception of  Environmental Audio Is Embedding Dependant
标题:Fréchet音频距离与人对环境音频感知的相关性是嵌入依赖关系
链接:https://arxiv.org/abs/2403.17508
作者:Modan Tailleur,Junwon Lee,Mathieu Lagrange,Keunwoo Choi,Laurie M. Heller,Keisuke Imoto,Yuki Okamoto
摘要:本文探讨了是否考虑替代特定领域的嵌入来计算Fr\'echet音频距离(FAD)度量可以帮助FAD更好地与环境声音的感知评级相关。我们使用了VGGish、PANN、MS-CLAP、L-CLAP和MERT的嵌入,这些嵌入是为音乐或环境声音评估量身定制的。FAD分数是针对来自DCASE 2023任务7数据集的声音计算的。使用来自相同任务的感知数据,我们发现PANNs-WGM-LogMel产生FAD分数与音频质量和感知拟合的感知评级之间的最佳相关性,Spearman相关性高于0.5。我们还发现,音乐特定的嵌入导致显着较低的结果。有趣的是,VGGish,用于原始Fr\'echet计算的嵌入,产生了低于0.1的相关性。这些结果强调了嵌入的FAD度量设计的选择至关重要。
摘要:This paper explores whether considering alternative domain-specific embeddings to calculate the Fr\'echet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings from VGGish, PANNs, MS-CLAP, L-CLAP, and MERT, which are tailored for either music or environmental sound evaluation. The FAD scores were calculated for sounds from the DCASE 2023 Task 7 dataset. Using perceptual data from the same task, we find that PANNs-WGM-LogMel produces the best correlation between FAD scores and perceptual ratings of both audio quality and perceived fit with a Spearman correlation higher than 0.5. We also find that music-specific embeddings resulted in significantly lower results. Interestingly, VGGish, the embedding used for the original Fr\'echet calculation, yielded a correlation below 0.1. These results underscore the critical importance of the choice of embedding for the FAD metric design.

【7】 Learning to Visually Localize Sound Sources from Mixtures without Prior  Source Knowledge
标题:学习在没有先验声源知识的情况下从混合声中视觉定位声源
链接:https://arxiv.org/abs/2403.17420
作者:Dongjin Kim,Sung Jin Um,Sangmin Lee,Jung Uk Kim
备注:Accepted at CVPR 2024
摘要:多声源定位任务的目标是从混合中单独定位声源。虽然最近的多声源定位方法已经显示出改进的性能,但由于它们依赖于关于要分离的对象的数量的先验信息,因此它们面临挑战。在本文中,为了克服这一限制,我们提出了一种新的多声源定位方法,可以进行定位,而无需先验知识的声源的数量。为了实现这一目标,我们提出了一个迭代对象识别(IOI)模块,它可以识别发声对象的迭代方式。在找到发声对象的区域后,我们设计了对象相似性感知聚类(OSC)损失,以指导IOI模块有效地组合同一对象的区域,但也区分不同的对象和背景。它使我们的方法能够在没有任何先验知识的情况下对发声对象进行准确的定位。在MUSIC和VGGSound基准测试上的大量实验结果表明,与现有的单源和多源方法相比,所提出的方法具有显著的性能改进。我们的代码可在:https://github.com/VisualAIKHU/NoPrior_MultiSSL
摘要:The goal of the multi-sound source localization task is to localize sound sources from the mixture individually. While recent multi-sound source localization methods have shown improved performance, they face challenges due to their reliance on prior information about the number of objects to be separated. In this paper, to overcome this limitation, we present a novel multi-sound source localization method that can perform localization without prior knowledge of the number of sound sources. To achieve this goal, we propose an iterative object identification (IOI) module, which can recognize sound-making objects in an iterative manner. After finding the regions of sound-making objects, we devise object similarity-aware clustering (OSC) loss to guide the IOI module to effectively combine regions of the same object but also distinguish between different objects and backgrounds. It enables our method to perform accurate localization of sound-making objects without any prior knowledge. Extensive experimental results on the MUSIC and VGGSound benchmarks show the significant performance improvements of the proposed method over the existing methods for both single and multi-source. Our code is available at: https://github.com/VisualAIKHU/NoPrior_MultiSSL

【8】 Exploring and Applying Audio-Based Sentiment Analysis in Music
标题:基于音频的情感分析在音乐中的探索与应用
链接:https://arxiv.org/abs/2403.17379
作者:Etash Jhanji
备注:5 pages, 7 figures, 2 tables. For source code, see this https URL
摘要:情感分析是文本处理的一个不断探索的领域,涉及文本的意见,情感和主观性的计算分析。然而,这一思想并不局限于文本和语音,事实上,它可以应用于其他模式。事实上,人类在文本中表达自己的方式并不像在音乐中那样深刻。计算模型解释音乐情感的能力在很大程度上是未开发的,可能在治疗和音乐排队中有影响和用途。在本文中,两个单独的任务得到解决。本研究旨在(1)预测音乐片段随时间的情感,(2)确定音乐之后的下一个情感值,以确保无缝过渡。利用来自音乐数据库中的情感数据,其中包含从免费音乐档案中选择的歌曲片段,并根据多名志愿者在Russel的circumplex情感模型中报告的效价和唤醒水平进行注释,模型被训练用于这两项任务。总的来说,这些模型的性能反映了它们能够有效和准确地执行设计任务。
摘要:Sentiment analysis is a continuously explored area of text processing that deals with the computational analysis of opinions, sentiments, and subjectivity of text. However, this idea is not limited to text and speech, in fact, it could be applied to other modalities. In reality, humans do not express themselves in text as deeply as they do in music. The ability of a computational model to interpret musical emotions is largely unexplored and could have implications and uses in therapy and musical queuing. In this paper, two individual tasks are addressed. This study seeks to (1) predict the emotion of a musical clip over time and (2) determine the next emotion value after the music in a time series to ensure seamless transitions. Utilizing data from the Emotions in Music Database, which contains clips of songs selected from the Free Music Archive annotated with levels of valence and arousal as reported on Russel's circumplex model of affect by multiple volunteers, models are trained for both tasks. Overall, the performance of these models reflected that they were able to perform the tasks they were designed for effectively and accurately.

【9】 Low-Latency Neural Speech Phase Prediction based on Parallel Estimation  Architecture and Anti-Wrapping Losses for Speech Generation Tasks
标题:基于并行估计结构和抗包裹损耗的语音生成任务低延迟神经语音相位预测
链接:https://arxiv.org/abs/2403.17378
作者:Yang Ai,Zhen-Hua Ling
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing. arXiv admin note: substantial text overlap with arXiv:2211.15974
摘要:本文提出了一种新的神经语音相位预测模型,它直接从幅度谱预测包裹相位谱。所提出的模型是一个级联的残差卷积网络和并行估计架构。并行估计结构是直接包裹相位预测的核心模块。该架构由两个并行的线性卷积层和一个相位计算公式组成,模仿了从复谱的实部和虚部计算相位谱的过程,并将预测的相位值严格限制在主值区间。为了避免相位缠绕引起的误差扩展问题,我们设计了反缠绕训练损耗,其定义在预测的缠绕相位谱和自然相位谱之间,通过使用反缠绕函数激活瞬时相位误差、群延迟误差和瞬时角频率误差。我们从数学上证明了反包裹函数应具有三个性质,即奇偶性、周期性和单调性。我们还通过结合因果卷积和知识蒸馏训练策略来实现低延迟的可流相位预测。对于分析合成和特定语音生成任务,实验结果表明,我们提出的神经语音相位预测模型优于迭代相位估计算法和基于神经网络的相位预测方法的相位预测精度,效率和鲁棒性。与基于HiFi-GAN的波形重构方法相比,该模型在保证合成语音质量的同时,也显示出了显著的效率优势。据我们所知,我们是第一个直接预测语音相位谱的幅度谱,只有通过神经网络。
摘要:This paper presents a novel neural speech phase prediction model which predicts wrapped phase spectra directly from amplitude spectra. The proposed model is a cascade of a residual convolutional network and a parallel estimation architecture. The parallel estimation architecture is a core module for direct wrapped phase prediction. This architecture consists of two parallel linear convolutional layers and a phase calculation formula, imitating the process of calculating the phase spectra from the real and imaginary parts of complex spectra and strictly restricting the predicted phase values to the principal value interval. To avoid the error expansion issue caused by phase wrapping, we design anti-wrapping training losses defined between the predicted wrapped phase spectra and natural ones by activating the instantaneous phase error, group delay error and instantaneous angular frequency error using an anti-wrapping function. We mathematically demonstrate that the anti-wrapping function should possess three properties, namely parity, periodicity and monotonicity. We also achieve low-latency streamable phase prediction by combining causal convolutions and knowledge distillation training strategies. For both analysis-synthesis and specific speech generation tasks, experimental results show that our proposed neural speech phase prediction model outperforms the iterative phase estimation algorithms and neural network-based phase prediction methods in terms of phase prediction precision, efficiency and robustness. Compared with HiFi-GAN-based waveform reconstruction method, our proposed model also shows outstanding efficiency advantages while ensuring the quality of synthesized speech. To the best of our knowledge, we are the first to directly predict speech phase spectra from amplitude spectra only via neural networks.


【10】 Theoretical Analysis of Quality of Conventional Beamforming for Phased  Microphone Arrays
标题:相控阵传声器常规波束形成质量的理论分析
链接:https://arxiv.org/abs/2403.17376
作者:Dheepak Khatri,Kenneth Granlund
摘要:理论研究进行分析不同类型的麦克风阵列设计的方向性响应。一维(线性)和二维(平面)麦克风阵列类型被认为是,和延迟和总和波束形成和传统的波束形成技术被用来定位声源。无量纲参数G的特征在于简化和标准化作为阵列几何形状和声源参数的函数的1-D和2-D麦克风阵列的抑制性能。该参数G然后用于确定用于远场声音定位的2-D麦克风阵列的改进设计。介绍并详细分析了一种称为等面积阵列的设计。该设计示出具有有利的抑制性能相比,其他传统使用的2-D平面麦克风阵列。
摘要:A theoretical study is performed to analyze the directional response of different types of microphone array designs. 1-D (linear) and 2-D (planar) microphone array types are considered, and the delay and sum beamforming and conventional beamforming techniques are employed to localize the sound source. A non-dimensional parameter, G, is characterized to simplify and standardize the rejection performance of both 1-D and 2-D microphone arrays as a function of array geometry and sound source parameters. This parameter G is then used to determine an improved design of a 2-D microphone array for far-field sound localization. One such design, termed the Equi-area array is introduced and analyzed in detail. The design is shown to have an advantageous rejection performance compared to other conventionally used 2-D planar microphone arrays.

【11】 Accuracy enhancement method for speech emotion recognition from  spectrogram using temporal frequency correlation and positional information  learning through knowledge transfer
标题:基于时频相关和知识转移位置信息学习的语谱图语音情感识别准确率提高方法
链接:https://arxiv.org/abs/2403.17327
作者:Jeong-Yoon Kim,Seung-Ho Lee
摘要:本文提出了一种利用Vision Transformer(ViT)处理语音谱图中频率(y轴)和时间(x轴)相关性,并通过知识传递在ViT之间传递位置信息,从而提高语音情感识别准确率的方法。所提出的方法具有以下独创性:i)我们使用垂直分割的对数梅尔频谱图的补丁来分析频率随时间的相关性。这种类型的补丁允许我们将特定情绪的最相关频率与它们发出的时间相关联。ii)我们提出使用图像坐标编码,这是一种适用于ViT的绝对位置编码。通过将图像的x,y坐标归一化为-1至1并将它们连接到图像,我们可以有效地为ViT提供有效的绝对位置信息。iii)通过特征图匹配,教师网络的地点和位置信息被有效地传输到学生网络。教师网络是通过图像坐标编码包含卷积干的局部性和绝对位置信息的ViT,而学生网络是在基本ViT结构中缺乏位置编码的结构。在特征图匹配阶段,我们通过平均绝对误差(L1损失)进行训练,以最小化两个网络的特征图之间的差异。为了验证所提出的方法,三个情感数据集(SAVEE,CARDB,和CREMA-D)组成的语音转换成对数梅尔频谱进行比较实验。实验结果表明,所提出的方法显着优于国家的最先进的方法在加权精度方面,同时需要显着减少浮点运算(FLOPs)。总的来说,所提出的方法提供了一个有前途的解决方案,通过提供更高的效率和性能的SER。
摘要:In this paper, we propose a method to improve the accuracy of speech emotion recognition (SER) by using vision transformer (ViT) to attend to the correlation of frequency (y-axis) with time (x-axis) in spectrogram and transferring positional information between ViT through knowledge transfer. The proposed method has the following originality i) We use vertically segmented patches of log-Mel spectrogram to analyze the correlation of frequencies over time. This type of patch allows us to correlate the most relevant frequencies for a particular emotion with the time they were uttered. ii) We propose the use of image coordinate encoding, an absolute positional encoding suitable for ViT. By normalizing the x, y coordinates of the image to -1 to 1 and concatenating them to the image, we can effectively provide valid absolute positional information for ViT. iii) Through feature map matching, the locality and location information of the teacher network is effectively transmitted to the student network. Teacher network is a ViT that contains locality of convolutional stem and absolute position information through image coordinate encoding, and student network is a structure that lacks positional encoding in the basic ViT structure. In feature map matching stage, we train through the mean absolute error (L1 loss) to minimize the difference between the feature maps of the two networks. To validate the proposed method, three emotion datasets (SAVEE, EmoDB, and CREMA-D) consisting of speech were converted into log-Mel spectrograms for comparison experiments. The experimental results show that the proposed method significantly outperforms the state-of-the-art methods in terms of weighted accuracy while requiring significantly fewer floating point operations (FLOPs). Overall, the proposed method offers an promising solution for SER by providing improved efficiency and performance.


机器翻译由腾讯交互翻译提供,仅供参考