今日论文合集:cs.SD语音14篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】Adapting WavLM for Speech Emotion Recognition
标题:将WavLM应用于语音情感识别
链接:https://arxiv.org/abs/2405.04485
作者:Daria Diatlova,Anton Udalov,Vitalii Shutov,Egor Spirin
摘要:最近,语音自监督模型(SSL)的下游任务的使用已经引起了很多关注。虽然大型预训练模型通常优于从头开始训练的较小模型,但有关最佳微调策略的问题仍然普遍存在。在本文中,我们探讨了微调策略的WavLM大模型的语音情感识别任务的MSP播客语料库。更具体地说,我们进行了一系列的实验,重点是使用性别和语义信息的话语。然后,我们总结了我们的研究结果,并描述了我们用于提交2024年语音情感识别挑战赛的最终模型。
摘要:Recently, the usage of speech self-supervised models (SSL) for downstream tasks has been drawing a lot of attention. While large pre-trained models commonly outperform smaller models trained from scratch, questions regarding the optimal fine-tuning strategies remain prevalent. In this paper, we explore the fine-tuning strategies of the WavLM Large model for the speech emotion recognition task on the MSP Podcast Corpus. More specifically, we perform a series of experiments focusing on using gender and semantic information from utterances. We then sum up our findings and describe the final model we used for submission to Speech Emotion Recognition Challenge 2024.

【2】 Universal Spatial Audio Transcoder
标题:通用空间音频代码转换器
链接:https://arxiv.org/abs/2405.04471
作者:Amaia Sagasti,Davide Scaini,Daniel Arteaga
备注:12 pages, 8 figures. Accepted for presentation at the AES 156th Convention, Madrid, Spain (June 2024)
摘要:本文解决了与不同空间音频格式之间的转换和空间音频格式到特定扬声器布局的解码相关联的挑战。现有的方法通常依赖于布局重映射工具,这可能无法保证从心理声学角度来看的最佳转换。为了克服这些挑战,我们提出了通用空间音频转码器(USAT)的方法和相应的开源实现。USAT为任何输入空间音频格式生成最佳解码器或转码器,使其适应任何输出格式或2D/3D扬声器配置。该算法利用基于心理声学原理的优化技术,最大限度地保留了空间信息。我们提出了几种音频格式的解码和转码的例子,并表明USAT方法是有利的,在该领域中最常见的方法相比。
摘要:This paper addresses the challenges associated with both the conversion between different spatial audio formats and the decoding of a spatial audio format to a specific loudspeaker layout. Existing approaches often rely on layout remapping tools, which may not guarantee optimal conversion from a psychoacoustic perspective. To overcome these challenges, we present the Universal Spatial Audio Transcoder(USAT) method and its corresponding open source implementation. USAT generates an optimal decoder or transcoder for any input spatial audio format, adapting it to any output format or 2D/3D loudspeaker configuration. Drawing upon optimization techniques based on psychoacoustic principles, the algorithm maximizes the preservation of spatial information. We present examples of the decoding and transcoding of several audio formats, and show that USAT approach is advantageous compared to the most common methods in the field.


【3】 Detecting music deepfakes is easy but actually hard
标题:检测音乐Deepfakes很容易,但实际上很难
链接:https://arxiv.org/abs/2405.04181
作者:Darius Afchar,Gabriel Meseguer Brocal,Romain Hennequin
备注:Under review
摘要:面对生成模型的新时代,人工生成内容的检测已成为至关重要的问题。在用户友好的平台上在几秒钟内创建可信的分钟长的音乐deepfake的能力对流媒体服务和人类艺术家的不公平竞争构成了真正的欺诈威胁。本文展示了在包含真实音频和假重建的数据集上训练分类器的可能性(以及令人惊讶的容易性),达到了令人信服的99.8%的准确率。据我们所知,这标志着音乐deepfake检测器的首次发布,这是一种有助于监管音乐伪造的工具。尽管如此,根据其他领域数十年来关于伪造检测的文献,我们强调,好的测试分数并不是故事的结局。我们从简单的机器学习框架中退一步,揭示了这种部署的检测器可能存在的许多问题:校准、对音频操作的鲁棒性、对看不见的模型的泛化、可解释性和追索的可能性。这第二部分作为该领域未来研究步骤的定位,并对繁荣的虚假内容检查器市场提出警告。
摘要:In the face of a new era of generative models, the detection of artificially generated content has become a matter of utmost importance. The ability to create credible minute-long music deepfakes in a few seconds on user-friendly platforms poses a real threat of fraud on streaming services and unfair competition to human artists. This paper demonstrates the possibility (and surprising ease) of training classifiers on datasets comprising real audio and fake reconstructions, achieving a convincing accuracy of 99.8%. To our knowledge, this marks the first publication of a music deepfake detector, a tool that will help in the regulation of music forgery. Nevertheless, informed by decades of literature on forgery detection in other fields, we stress that a good test score is not the end of the story. We step back from the straightforward ML framework and expose many facets that could be problematic with such a deployed detector: calibration, robustness to audio manipulation, generalisation to unseen models, interpretability and possibility for recourse. This second part acts as a position for future research steps in the field and a caveat to a flourishing market of fake content checkers.


【4】 Fine-grained Speech Sentiment Analysis in Chinese Psychological Support  Hotlines Based on Large-scale Pre-trained Model
标题:基于大规模预训练模型的中国心理支持热线细粒度言语情绪分析
链接:https://arxiv.org/abs/2405.04128
作者:Zhonglong Chen,Changwei Song,Yining Chen,Jianqiang Li,Guanghui Fu,Yongsheng Tong,Qing Zhao
摘要:自杀和自杀行为仍然是公共政策和医疗保健的重大挑战。为此,世界各地设立了心理支助热线,为处于精神危机中的个人提供即时帮助。这些热线的有效性在很大程度上取决于准确识别来电者的情绪状态,特别是表明自杀风险增加的潜在负面情绪。然而,对心理干预的高需求往往导致专业操作人员的短缺,突出了对有效的语音情感识别模型的需求。该模型将自动检测和分析呼叫者的情绪,便于整合到热线服务中。此外,它将使心理支持热线互动的大规模数据分析,以探索跨人群的心理现象和行为。我们的研究利用了中国最大的自杀热线北京心理支持热线的数据。我们分析了来自105个呼叫者的语音数据,包含20,630个片段,并将其分为11种类型的负面情绪。我们使用大规模的预训练模型开发了一个负面情绪识别模型和一个细粒度的多标签分类模型。我们的实验表明,负面情绪识别模型达到了最大的F1—得分为76.96%。然而,它在细粒度多标签分类任务中表现出有限的功效,最好的模型仅获得41.74%的加权F1分数。我们对这项任务进行了错误分析,讨论了未来可能的改进,并考虑了我们研究的临床应用可能性。所有代码都是公开的。
摘要:Suicide and suicidal behaviors remain significant challenges for public policy and healthcare. In response, psychological support hotlines have been established worldwide to provide immediate help to individuals in mental crises. The effectiveness of these hotlines largely depends on accurately identifying callers' emotional states, particularly underlying negative emotions indicative of increased suicide risk. However, the high demand for psychological interventions often results in a shortage of professional operators, highlighting the need for an effective speech emotion recognition model. This model would automatically detect and analyze callers' emotions, facilitating integration into hotline services. Additionally, it would enable large-scale data analysis of psychological support hotline interactions to explore psychological phenomena and behaviors across populations. Our study utilizes data from the Beijing psychological support hotline, the largest suicide hotline in China. We analyzed speech data from 105 callers containing 20,630 segments and categorized them into 11 types of negative emotions. We developed a negative emotion recognition model and a fine-grained multi-label classification model using a large-scale pre-trained model. Our experiments indicate that the negative emotion recognition model achieves a maximum F1-score of 76.96%. However, it shows limited efficacy in the fine-grained multi-label classification task, with the best model achieving only a 41.74% weighted F1-score. We conducted an error analysis for this task, discussed potential future improvements, and considered the clinical application possibilities of our study. All the codes are public available.

【5】 Comparative Study of Recurrent Neural Networks for Virtual Analog Audio  Effects Modeling
标题:虚拟模拟音效建模中的回归神经网络比较研究
链接:https://arxiv.org/abs/2405.04124
作者:Riccardo Simionato,Stefano Fasciani
备注:arXiv admin note: text overlap with arXiv:1810.06603 by other authors
摘要:模拟电子电路是音乐设备的一个重要类别的核心。其电子元件的非线性特性使模拟音乐设备具有独特的音色和音质,使其非常受欢迎。人工神经网络在模拟音频效果电路的仿真中,特别是在递归网络中,已经迅速得到普及。虽然神经方法在精确建模失真电路方面取得了成功,但它们需要考虑参数调节和低延迟响应的架构改进。在这篇文章中,我们将探讨最新机器学习技术在虚拟模拟建模中的应用。我们比较了状态空间模型和线性递归单元对更常见的长短期记忆网络。这些在序列到序列建模任务中显示出有前途的能力,在信号历史编码中显示出显着的改进。我们的比较研究使用这些黑盒神经建模技术与各种音频效果。我们使用多个指标来评估性能和局限性,旨在评估模型准确复制音频信号中的能量包络、频率内容和瞬态的能力。为了将控制参数,我们采用的特征明智的线性调制方法。长短期记忆网络在模拟失真和均衡器方面表现出更好的准确性,而状态空间模型,其次是长短期记忆网络,当集成在编码器解码器结构中时,在模拟饱和和压缩方面优于其他模型。当考虑长时变特性时,状态空间模型表现出最大的准确性。长短期记忆,特别是线性递归单元网络更倾向于引入音频伪影。
摘要:Analog electronic circuits are at the core of an important category of musical devices. The nonlinear features of their electronic components give analog musical devices a distinctive timbre and sound quality, making them highly desirable. Artificial neural networks have rapidly gained popularity for the emulation of analog audio effects circuits, particularly recurrent networks. While neural approaches have been successful in accurately modeling distortion circuits, they require architectural improvements that account for parameter conditioning and low latency response. In this article, we explore the application of recent machine learning advancements for virtual analog modeling. We compare State Space models and Linear Recurrent Units against the more common Long Short Term Memory networks. These have shown promising ability in sequence to sequence modeling tasks, showing a notable improvement in signal history encoding. Our comparative study uses these black box neural modeling techniques with a variety of audio effects. We evaluate the performance and limitations using multiple metrics aiming to assess the models' ability to accurately replicate energy envelopes, frequency contents, and transients in the audio signal. To incorporate control parameters we employ the Feature wise Linear Modulation method. Long Short Term Memory networks exhibit better accuracy in emulating distortions and equalizers, while the State Space model, followed by Long Short Term Memory networks when integrated in an encoder decoder structure, outperforms others in emulating saturation and compression. When considering long time variant characteristics, the State Space model demonstrates the greatest accuracy. The Long Short Term Memory and, in particular, Linear Recurrent Unit networks present more tendency to introduce audio artifacts.


【6】 Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
标题:基于动态图的自适应语音情感表示学习
链接:https://arxiv.org/abs/2405.03956
作者:Yingxue Gao,Huan Zhao,Zixing Zhang
备注:None
摘要:图表示学习因其在提取代表节点嵌入方面的强大非线性拟合能力而成为研究热点。然而,对于语音信号这样的序列数据,大多数传统的方法仅仅关注序列内创建的静态图,而在很大程度上忽略了这些数据的内在演变模式。这可能会降低顺序数据的图形表示学习的效率。为此,我们提出了一种自适应的图表示学习方法的基础上动态进化的图,这是连续构建的一系列连续性分割的滑动窗口。在这样做的时候,最好是在一个长序列中捕获局部和全局上下文信息。此外,我们引入了一种加权的方法来更新节点表示,而不是传统的平均值,其中的权重是通过一种新的矩阵计算的基础上的程度相邻节点。最后,我们构造了一个可学习的图卷积层,它结合了图结构损失和分类损失来优化图结构。为了验证所提出的方法的有效性,我们进行了实验的语音情感识别IEMOCAP和RAVDESS数据集。实验结果表明,该方法优于最新的(非)基于图的模型。
摘要:Graph representation learning has become a hot research topic due to its powerful nonlinear fitting capability in extracting representative node embeddings. However, for sequential data such as speech signals, most traditional methods merely focus on the static graph created within a sequence, and largely overlook the intrinsic evolving patterns of these data. This may reduce the efficiency of graph representation learning for sequential data. For this reason, we propose an adaptive graph representation learning method based on dynamically evolved graphs, which are consecutively constructed on a series of subsequences segmented by a sliding window. In doing this, it is better to capture local and global context information within a long sequence. Moreover, we introduce a weighted approach to update the node representation rather than the conventional average one, where the weights are calculated by a novel matrix computation based on the degree of neighboring nodes. Finally, we construct a learnable graph convolutional layer that combines the graph structure loss and classification loss to optimize the graph structure. To verify the effectiveness of the proposed method, we conducted experiments for speech emotion recognition on the IEMOCAP and RAVDESS datasets. Experimental results show that the proposed method outperforms the latest (non-)graph-based models.

【7】 Intelligent Cardiac Auscultation for Murmur Detection via  Parallel-Attentive Models with Uncertainty Estimation
标题:通过具有不确定性估计的注意力模型进行智能心脏听诊以检测低语
链接:https://arxiv.org/abs/2405.03953
作者:Zixing Zhang,Tao Pang,Jing Han,Björn W. Schuller
备注:None
摘要:心脏杂音是心血管疾病的常见表现,可为早期心脏异常提供重要线索。虽然目前大多数研究方法主要关注模型的准确性,但它们往往忽略了其他重要方面,例如机器学习算法的可解释性和预测的不确定性。本文介绍了一种基于并行注意模型的心脏杂音检测方法,该方法由两个分支组成:一个是基于自注意模块的,另一个是基于卷积网络的。与传统方法不同,这种结构更适合处理序列数据中的长期依赖关系,从而有效地捕获心脏杂音的局部和全局特征。此外,我们认识到理解医学领域模型预测的不确定性对临床决策的重要性。因此,我们已经纳入了一个有效的不确定性估计方法的基础上,蒙特卡罗辍学到我们的模型。此外,我们还采用了温度标度来校准概率模型的预测,提高了其可靠性。在CirCor Digiscope数据集上进行的心脏杂音检测实验中,我们提出的方法实现了79.8%的加权准确率和65.1%的F1,代表了最先进的结果。
摘要:Heart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focus on the accuracy of models, they often overlook other important aspects such as the interpretability of machine learning algorithms and the uncertainty of predictions. This paper introduces a heart murmur detection method based on a parallel-attentive model, which consists of two branches: One is based on a self-attention module and the other one is based on a convolutional network. Unlike traditional approaches, this structure is better equipped to handle long-term dependencies in sequential data, and thus effectively captures the local and global features of heart murmurs. Additionally, we acknowledge the significance of understanding the uncertainty of model predictions in the medical field for clinical decision-making. Therefore, we have incorporated an effective uncertainty estimation method based on Monte Carlo Dropout into our model. Furthermore, we have employed temperature scaling to calibrate the predictions of our probabilistic model, enhancing its reliability. In experiments conducted on the CirCor Digiscope dataset for heart murmur detection, our proposed method achieves a weighted accuracy of 79.8% and an F1 of 65.1%, representing state-of-the-art results.


【8】 HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's  Disease Detection From Spontaneous Speech
标题:HAFFormer:从自发言语检测阿尔茨海默病的分层免注意框架
链接:https://arxiv.org/abs/2405.03952
作者:Zhongren Dong,Zixing Zhang,Weixiang Xu,Jing Han,Jianjun Ou,Björn W. Schuller
备注:None
摘要:从自发语音中自动检测阿尔茨海默病(Alzheimer 'sDisease,AD)对AD的早期诊断有重要意义。最近的方法高度依赖于Transformer架构,因为它在建模长距离上下文依赖关系方面的效率很高。然而,当在边缘设备上部署此类模型时,与自我注意力和音频长度相关的计算复杂性的二次增加构成了挑战。在这种情况下,我们构建了一个新的框架,即分层无注意力Transformer(HAFFormer),以更好地处理长语音AD检测。具体来说,我们采用多尺度相关卷积的无注意模块来代替自注意,从而避免了昂贵的计算,并采用基于GELU的门控线性单元来代替前馈层,旨在自动过滤掉冗余信息。此外,我们设计了一个层次结构,迫使它学习各种信息颗粒,从框架级到对话级。通过在ADReSS—M数据集上进行广泛的实验,引入的HAFFormer可以实现与其他近期工作竞争的结果(82.6%的准确率),但与标准Transformer相比,计算复杂性和模型大小显著降低。这显示了HAFFormer在处理用于AD检测的长音频方面的效率。
摘要:Automatically detecting Alzheimer's Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectures due to its efficiency in modelling long-range context dependencies. However, the quadratic increase in computational complexity associated with self-attention and the length of audio poses a challenge when deploying such models on edge devices. In this context, we construct a novel framework, namely Hierarchical Attention-Free Transformer (HAFFormer), to better deal with long speech for AD detection. Specifically, we employ an attention-free module of Multi-Scale Depthwise Convolution to replace the self-attention and thus avoid the expensive computation, and a GELU-based Gated Linear Unit to replace the feedforward layer, aiming to automatically filter out the redundant information. Moreover, we design a hierarchical structure to force it to learn a variety of information grains, from the frame level to the dialogue level. By conducting extensive experiments on the ADReSS-M dataset, the introduced HAFFormer can achieve competitive results (82.6% accuracy) with other recent work, but with significant computational complexity and model size reduction compared to the standard Transformer. This shows the efficiency of HAFFormer in dealing with long audio for AD detection.

【9】 A 65nm 36nJ/Decision Bio-inspired Temporal-Sparsity-Aware Digital  Keyword Spotting IC with 0.6V Near-Threshold SRAM
标题:65纳米36 nJ/Decision受生物启发的时间稀疏感知数字关键字发现IC,配备0.6V近阈值静态存储器链接:https://arxiv.org/abs/2405.03905
作者:Qinyu Chen,Kwantae Kim,Chang Gao,Sheng Zhou,Taekwang Jang,Tobi Delbruck,Shih-Chii Liu摘要:本文介绍了,据作者所知,第一个细粒度的时间稀疏意识的关键字定位(KWS)IC利用从输入帧和网络隐藏状态提取的相邻特征向量之间的时间相似性,消除不必要的操作和内存访问。这款KWS IC采用生物启发的Delta门控递归神经网络({\Delta} RNN)分类器,实现了90.5%的11类Google语音命令数据集(GSCD)KWS准确率和36nJ/决策的能耗。在87%的时间稀疏度下,计算延迟和每次推理的能量分别减少了2.4$\times $/3.4$\times $。65纳米设计占用0.78mm$^2 $,并具有两个额外的模块,一个紧凑的0.084mm$^2 $基于数字无限脉冲响应(IIR)的带通滤波器(BPF)音频特征提取器(FEx)和一个24 kB 0.6V近Vth重量SRAM,与标准SRAM相比,读取功率低6.6$\times $。
摘要:This paper introduces, to the best of the authors' knowledge, the first fine-grained temporal sparsity-aware keyword spotting (KWS) IC leveraging temporal similarities between neighboring feature vectors extracted from input frames and network hidden states, eliminating unnecessary operations and memory accesses. This KWS IC, featuring a bio-inspired delta-gated recurrent neural network ({\Delta}RNN) classifier, achieves an 11-class Google Speech Command Dataset (GSCD) KWS accuracy of 90.5% and energy consumption of 36nJ/decision. At 87% temporal sparsity, computing latency and energy per inference are reduced by 2.4$\times$/3.4$\times$, respectively. The 65nm design occupies 0.78mm$^2$ and features two additional blocks, a compact 0.084mm$^2$ digital infinite-impulse-response (IIR)-based band-pass filter (BPF) audio feature extractor (FEx) and a 24kB 0.6V near-Vth weight SRAM with 6.6$\times$ lower read power compared to the standard SRAM.

【10】 BERP: A Blind Estimator of Room Acoustic and Physical Parameters for  Single-Channel Noisy Speech Signals
标题:BRP:单通道含噪语音信号房间声学和物理参数的盲估计
链接:https://arxiv.org/abs/2405.04476
作者:Lijun Wang,Yixian Lu,Ziyan Gao,Kai Li,Jianqiang Huang,Yuntao Kong,Shogo Okada备注:Submitted to IEEE/ACM Transaction on Audio Speech and Language Processing (TASLP)
摘要:房间声学参数(RAP)和房间物理参数(RPP)是用于参数化围绕收听者的局部环境的声场的房间声学特性(RAC)的基本度量,为各种应用提供全面的指示。当前的RAP和RPP估计方法要么在真实背景噪声的上下文中不能覆盖广泛的真实世界声学环境,要么缺乏用于从有噪声的单通道语音信号(特别是声源距离、声源的到达方向(DOA)和占用水平)盲估计RAP和RPP的通用框架。另一方面,本文通过引入一种新的随机房间脉冲响应(RIR)模型,即稀疏随机脉冲响应(SSIR)模型,并赋予BERP一个统一的编码器和多个独立的预测器来并行估计RPP和SSIR参数,提出了一种新的通用盲估计框架--房间声物理参数盲估计器(BERP).该估计框架使得能够通过单独使用有噪声的单通道语音信号来计算高效且通用地估计房间参数。最后,所有的RAP可以同时从合成的RIR从SSIR模型与估计的参数。为了评估所提出的BERP和SSIR模型的有效性,我们从几个公开的数据集中编译了一个特定于任务的数据集。结果表明,BERP达到了最先进的(SOTA)性能。此外,与SSIR RIR模型相关的评估结果也证明了其有效性。代码可以在GitHub上找到。
摘要:Room acoustic parameters (RAPs) and room physical parameters ( RPPs) are essential metrics for parameterizing the room acoustical characteristics (RAC) of a sound field around a listener's local environment, offering comprehensive indications for various applications. The current RAPs and RPPs estimation methods either fall short of covering broad real-world acoustic environments in the context of real background noise or lack universal frameworks for blindly estimating RAPs and RPPs from noisy single-channel speech signals, particularly sound source distances, direction-of-arrival (DOA) of sound sources, and occupancy levels. On the other hand, in this paper, we propose a novel universal blind estimation framework called the blind estimator of room acoustical and physical parameters (BERP), by introducing a new stochastic room impulse response (RIR) model, namely, the sparse stochastic impulse response (SSIR) model, and endowing the BERP with a unified encoder and multiple separate predictors to estimate RPPs and SSIR parameters in parallel. This estimation framework enables the computationally efficient and universal estimation of room parameters by solely using noisy single-channel speech signals. Finally, all the RAPs can be simultaneously derived from the RIRs synthesized from SSIR model with the estimated parameters. To evaluate the effectiveness of the proposed BERP and SSIR models, we compile a task-specific dataset from several publicly available datasets. The results reveal that the BERP achieves state-of-the-art (SOTA) performance. Moreover, the evaluation results pertaining to the SSIR RIR model also demonstrated its efficacy. The code is available on GitHub.


【11】 BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion  Models
标题:BUDDy:使用扩散模型的单通道盲无监督去回响
链接:https://arxiv.org/abs/2405.04272
作者:Eloi Moliner,Jean-Marie Lemercier,Simon Welker,Timo Gerkmann,Vesa Välimäki
备注:Submitted to IWAENC 2024
摘要:本文提出了一种基于扩散模型的后验采样的无监督单通道联合盲去混响和房间脉冲响应估计方法。我们参数化的混响算子使用滤波器与指数衰减的每个频率子带,并迭代地估计相应的参数的语音话语得到细化沿反向扩散轨迹。测量一致性准则强制所生成的语音与混响测量的保真度,而无条件扩散模型实现用于干净语音生成的强先验。在没有任何房间脉冲响应的知识,也没有任何耦合的混响消声数据,我们可以成功地在各种声学场景中执行去混响。我们的方法显着优于以前的盲无监督基线,我们证明了其增加的鲁棒性看不见的声学条件相比,盲监督方法。音频样本和代码可在线获得。
摘要:In this paper, we present an unsupervised single-channel method for joint blind dereverberation and room impulse response estimation, based on posterior sampling with diffusion models. We parameterize the reverberation operator using a filter with exponential decay for each frequency subband, and iteratively estimate the corresponding parameters as the speech utterance gets refined along the reverse diffusion trajectory. A measurement consistency criterion enforces the fidelity of the generated speech with the reverberant measurement, while an unconditional diffusion model implements a strong prior for clean speech generation. Without any knowledge of the room impulse response nor any coupled reverberant-anechoic data, we can successfully perform dereverberation in various acoustic scenarios. Our method significantly outperforms previous blind unsupervised baselines, and we demonstrate its increased robustness to unseen acoustic conditions in comparison to blind supervised methods. Audio samples and code are available online.


【12】 Speaker Characterization by means of Attention Pooling
标题:通过注意力池进行说话者定性
链接:https://arxiv.org/abs/2405.04096
作者:Federico Costa,Miquel India,Javier Hernando
备注:None
摘要:用于说话人验证的最先进的深度学习系统通常基于说话人嵌入提取器。这些架构通常由特征提取器前端和池化层组成,以将可变长度的话语编码成固定长度的说话者向量。作者最近提出了使用双多头自注意池进行说话人识别,放置在基于CNN的前端和一组完全连接的层之间。这已被证明是一种很好的方法,可以有效地选择前端从语音信号中捕获的最相关特征。在本文中,我们通过将该架构适用于其他不同的说话人特征化任务,如情感识别,性别分类和COVID-19检测,展示了出色的实验结果。
摘要:State-of-the-art Deep Learning systems for speaker verification are commonly based on speaker embedding extractors. These architectures are usually composed of a feature extractor front-end together with a pooling layer to encode variable-length utterances into fixed-length speaker vectors. The authors have recently proposed the use of a Double Multi-Head Self-Attention pooling for speaker recognition, placed between a CNN-based front-end and a set of fully connected layers. This has shown to be an excellent approach to efficiently select the most relevant features captured by the front-end from the speech signal. In this paper we show excellent experimental results by adapting this architecture to other different speaker characterization tasks, such as emotion recognition, sex classification and COVID-19 detection.


【13】 MMGER: Multi-modal and Multi-granularity Generative Error Correction  with LLM for Joint Accent and Speech Recognition
标题:MMGER:利用LLM进行多模式和多粒度生成式错误纠正,用于联合口音和语音识别
链接:https://arxiv.org/abs/2405.03152
作者:Bingshen Mu,Yangze Li,Qijie Shao,Kun Wei,Xucheng Wan,Naijun Zheng,Huan Zhou,Lei Xie
摘要:尽管自动语音识别(ASR)取得了显着的进步,但当面临不利条件时,性能往往会下降。生成纠错(GER)利用大型语言模型(LLM)的卓越文本理解能力,在ASR纠错中提供令人印象深刻的性能,其中N-best假设为转录预测提供了有价值的信息。然而,GER遇到的挑战,如固定的N-best假设,声学信息的利用不足,和有限的特异性多口音的情况下。在本文中,我们探讨了GER在多口音场景中的应用。口音代表了标准发音规范的偏差,同时ASR和口音识别(AR)的多任务学习框架有效地解决了多口音场景,使其成为一个突出的解决方案。在这项工作中,我们提出了一个统一的ASR-AR GER模型,命名为MMGER,利用多模态校正和多粒度校正。采用多任务ASR-AR学习来提供动态1-最佳假设和口音嵌入。多模态校正通过将语音的声学特征与对应的字符级1最佳假设序列强制对齐来实现细粒度的帧级校正。多粒度校正通过在细粒度多模态校正之上结合常规1-最佳假设来补充全局语言信息,以实现粗粒度话语级校正。MMGER有效地缓解了GER的局限性,并为多口音场景定制了基于LLM的ASR纠错。在多口音普通话KeSpeech数据集上进行的实验证明了MMGER的有效性,与完善的标准基线相比,AR准确率相对提高了26.72%,ASR字符错误率相对降低了27.55%。
摘要:Despite notable advancements in automatic speech recognition (ASR), performance tends to degrade when faced with adverse conditions. Generative error correction (GER) leverages the exceptional text comprehension capabilities of large language models (LLM), delivering impressive performance in ASR error correction, where N-best hypotheses provide valuable information for transcription prediction. However, GER encounters challenges such as fixed N-best hypotheses, insufficient utilization of acoustic information, and limited specificity to multi-accent scenarios. In this paper, we explore the application of GER in multi-accent scenarios. Accents represent deviations from standard pronunciation norms, and the multi-task learning framework for simultaneous ASR and accent recognition (AR) has effectively addressed the multi-accent scenarios, making it a prominent solution. In this work, we propose a unified ASR-AR GER model, named MMGER, leveraging multi-modal correction, and multi-granularity correction. Multi-task ASR-AR learning is employed to provide dynamic 1-best hypotheses and accent embeddings. Multi-modal correction accomplishes fine-grained frame-level correction by force-aligning the acoustic features of speech with the corresponding character-level 1-best hypothesis sequence. Multi-granularity correction supplements the global linguistic information by incorporating regular 1-best hypotheses atop fine-grained multi-modal correction to achieve coarse-grained utterance-level correction. MMGER effectively mitigates the limitations of GER and tailors LLM-based ASR error correction for the multi-accent scenarios. Experiments conducted on the multi-accent Mandarin KeSpeech dataset demonstrate the efficacy of MMGER, achieving a 26.72% relative improvement in AR accuracy and a 27.55% relative reduction in ASR character error rate, compared to a well-established standard baseline.


【14】 Analysis about Theoretical Foundations for Method to Enhancing ASR  Performance using OCR Word Frequency Differences
标题:利用OCR词频差提高ASB性能方法的理论基础分析
链接:https://arxiv.org/abs/2405.02995
作者:Kyudan Jung,Nam-Joon Kim,Hyun Gon Ryu,Hyuk-Jae Lee
摘要:随着人们对大型语言模型(LLM)的兴趣日益增长,自动语音识别(ASR)中准确性的重要性变得更加突出。对于包含专业术语的讲座尤其如此,传统ASR模型的成功率往往很低,这带来了一个具有挑战性的问题。提出了一种基于词频差异的专业术语自动识别方法。通过实验和数据分析,我们调查了这一建议是否有效地解决了这个问题。此外,我们还介绍了幂律作为相对频率的理论基础
摘要:As interest in large language models (LLMs) grows, the importance of accuracy in automatic speech recognition (ASR) has become more pronounced. This is particularly true for lectures that include specialized terminology, where the success rate of traditional ASR models tends to be low, posing a challenging problem. A method to improve ASR performance for specialized terminology using the word frequency difference approach has been proposed. Through experiments and data analysis, we investigate whether this proposal effectively addresses the issue. Additionally, we introduce the power law as the theoretical foundation for the relative frequency


eess.AS音频处理
【1】 BERP: A Blind Estimator of Room Acoustic and Physical Parameters for  Single-Channel Noisy Speech Signals
标题:BRP:单通道含噪语音信号房间声学和物理参数的盲估计
链接:https://arxiv.org/abs/2405.04476
作者:Lijun Wang,Yixian Lu,Ziyan Gao,Kai Li,Jianqiang Huang,Yuntao Kong,Shogo Okada
备注:Submitted to IEEE/ACM Transaction on Audio Speech and Language Processing (TASLP)
摘要:房间声学参数(RAP)和房间物理参数(RPP)是用于参数化围绕收听者的局部环境的声场的房间声学特性(RAC)的基本度量,为各种应用提供全面的指示。当前的RAP和RPP估计方法要么在真实背景噪声的上下文中不能覆盖广泛的真实世界声学环境,要么缺乏用于从有噪声的单通道语音信号(特别是声源距离、声源的到达方向(DOA)和占用水平)盲估计RAP和RPP的通用框架。另一方面,本文通过引入一种新的随机房间脉冲响应(RIR)模型,即稀疏随机脉冲响应(SSIR)模型,并赋予BERP一个统一的编码器和多个独立的预测器来并行估计RPP和SSIR参数,提出了一种新的通用盲估计框架--房间声物理参数盲估计器(BERP).该估计框架使得能够通过单独使用有噪声的单通道语音信号来计算高效且通用地估计房间参数。最后,所有的RAP可以同时从合成的RIR从SSIR模型与估计的参数。为了评估所提出的BERP和SSIR模型的有效性,我们从几个公开的数据集中编译了一个特定于任务的数据集。结果表明,BERP达到了最先进的(SOTA)性能。此外,与SSIR RIR模型相关的评估结果也证明了其有效性。代码可以在GitHub上找到。
摘要:Room acoustic parameters (RAPs) and room physical parameters ( RPPs) are essential metrics for parameterizing the room acoustical characteristics (RAC) of a sound field around a listener's local environment, offering comprehensive indications for various applications. The current RAPs and RPPs estimation methods either fall short of covering broad real-world acoustic environments in the context of real background noise or lack universal frameworks for blindly estimating RAPs and RPPs from noisy single-channel speech signals, particularly sound source distances, direction-of-arrival (DOA) of sound sources, and occupancy levels. On the other hand, in this paper, we propose a novel universal blind estimation framework called the blind estimator of room acoustical and physical parameters (BERP), by introducing a new stochastic room impulse response (RIR) model, namely, the sparse stochastic impulse response (SSIR) model, and endowing the BERP with a unified encoder and multiple separate predictors to estimate RPPs and SSIR parameters in parallel. This estimation framework enables the computationally efficient and universal estimation of room parameters by solely using noisy single-channel speech signals. Finally, all the RAPs can be simultaneously derived from the RIRs synthesized from SSIR model with the estimated parameters. To evaluate the effectiveness of the proposed BERP and SSIR models, we compile a task-specific dataset from several publicly available datasets. The results reveal that the BERP achieves state-of-the-art (SOTA) performance. Moreover, the evaluation results pertaining to the SSIR RIR model also demonstrated its efficacy. The code is available on GitHub.


【2】 BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion  Models
标题:BUDDy:使用扩散模型的单通道盲无监督去回响
链接:https://arxiv.org/abs/2405.04272
作者:Eloi Moliner,Jean-Marie Lemercier,Simon Welker,Timo Gerkmann,Vesa Välimäki
备注:Submitted to IWAENC 2024
摘要:本文提出了一种基于扩散模型的后验采样的无监督单通道联合盲去混响和房间脉冲响应估计方法。我们参数化的混响算子使用滤波器与指数衰减的每个频率子带,并迭代地估计相应的参数的语音话语得到细化沿反向扩散轨迹。测量一致性准则强制所生成的语音与混响测量的保真度,而无条件扩散模型实现用于干净语音生成的强先验。在没有任何房间脉冲响应的知识,也没有任何耦合的混响消声数据,我们可以成功地在各种声学场景中执行去混响。我们的方法显着优于以前的盲无监督基线,我们证明了其增加的鲁棒性看不见的声学条件相比,盲监督方法。音频样本和代码可在线获得。
摘要:In this paper, we present an unsupervised single-channel method for joint blind dereverberation and room impulse response estimation, based on posterior sampling with diffusion models. We parameterize the reverberation operator using a filter with exponential decay for each frequency subband, and iteratively estimate the corresponding parameters as the speech utterance gets refined along the reverse diffusion trajectory. A measurement consistency criterion enforces the fidelity of the generated speech with the reverberant measurement, while an unconditional diffusion model implements a strong prior for clean speech generation. Without any knowledge of the room impulse response nor any coupled reverberant-anechoic data, we can successfully perform dereverberation in various acoustic scenarios. Our method significantly outperforms previous blind unsupervised baselines, and we demonstrate its increased robustness to unseen acoustic conditions in comparison to blind supervised methods. Audio samples and code are available online.

【3】 Speaker Characterization by means of Attention Pooling
标题:通过注意力池进行说话者定性
链接:https://arxiv.org/abs/2405.04096
作者:Federico Costa,Miquel India,Javier Hernando
备注:None
摘要:用于说话人验证的最先进的深度学习系统通常基于说话人嵌入提取器。这些架构通常由特征提取器前端和池化层组成,以将可变长度的话语编码成固定长度的说话者向量。作者最近提出了使用双多头自注意池进行说话人识别,放置在基于CNN的前端和一组完全连接的层之间。这已被证明是一种很好的方法,可以有效地选择前端从语音信号中捕获的最相关特征。在本文中,我们通过将该架构适用于其他不同的说话人特征化任务,如情感识别,性别分类和COVID-19检测,展示了出色的实验结果。
摘要:State-of-the-art Deep Learning systems for speaker verification are commonly based on speaker embedding extractors. These architectures are usually composed of a feature extractor front-end together with a pooling layer to encode variable-length utterances into fixed-length speaker vectors. The authors have recently proposed the use of a Double Multi-Head Self-Attention pooling for speaker recognition, placed between a CNN-based front-end and a set of fully connected layers. This has shown to be an excellent approach to efficiently select the most relevant features captured by the front-end from the speech signal. In this paper we show excellent experimental results by adapting this architecture to other different speaker characterization tasks, such as emotion recognition, sex classification and COVID-19 detection.


【4】 Adapting WavLM for Speech Emotion Recognition
标题:将WavLM应用于语音情感识别
链接:https://arxiv.org/abs/2405.04485
作者:Daria Diatlova,Anton Udalov,Vitalii Shutov,Egor Spirin
摘要:最近,语音自监督模型(SSL)的下游任务的使用已经引起了很多关注。虽然大型预训练模型通常优于从头开始训练的较小模型,但有关最佳微调策略的问题仍然普遍存在。在本文中,我们探讨了微调策略的WavLM大模型的语音情感识别任务的MSP播客语料库。更具体地说,我们进行了一系列的实验,重点是使用性别和语义信息的话语。然后,我们总结了我们的研究结果,并描述了我们用于提交2024年语音情感识别挑战赛的最终模型。
摘要:Recently, the usage of speech self-supervised models (SSL) for downstream tasks has been drawing a lot of attention. While large pre-trained models commonly outperform smaller models trained from scratch, questions regarding the optimal fine-tuning strategies remain prevalent. In this paper, we explore the fine-tuning strategies of the WavLM Large model for the speech emotion recognition task on the MSP Podcast Corpus. More specifically, we perform a series of experiments focusing on using gender and semantic information from utterances. We then sum up our findings and describe the final model we used for submission to Speech Emotion Recognition Challenge 2024.

【5】 Universal Spatial Audio Transcoder
标题:通用空间音频代码转换器
链接:https://arxiv.org/abs/2405.04471
作者:Amaia Sagasti,Davide Scaini,Daniel Arteaga
备注:12 pages, 8 figures. Accepted for presentation at the AES 156th Convention, Madrid, Spain (June 2024)
摘要:本文解决了与不同空间音频格式之间的转换和空间音频格式到特定扬声器布局的解码相关联的挑战。现有的方法通常依赖于布局重映射工具,这可能无法保证从心理声学角度来看的最佳转换。为了克服这些挑战,我们提出了通用空间音频转码器(USAT)的方法和相应的开源实现。USAT为任何输入空间音频格式生成最佳解码器或转码器,使其适应任何输出格式或2D/3D扬声器配置。该算法利用基于心理声学原理的优化技术,最大限度地保留了空间信息。我们提出了几种音频格式的解码和转码的例子,并表明USAT方法是有利的,在该领域中最常见的方法相比。
摘要:This paper addresses the challenges associated with both the conversion between different spatial audio formats and the decoding of a spatial audio format to a specific loudspeaker layout. Existing approaches often rely on layout remapping tools, which may not guarantee optimal conversion from a psychoacoustic perspective. To overcome these challenges, we present the Universal Spatial Audio Transcoder(USAT) method and its corresponding open source implementation. USAT generates an optimal decoder or transcoder for any input spatial audio format, adapting it to any output format or 2D/3D loudspeaker configuration. Drawing upon optimization techniques based on psychoacoustic principles, the algorithm maximizes the preservation of spatial information. We present examples of the decoding and transcoding of several audio formats, and show that USAT approach is advantageous compared to the most common methods in the field.


【6】 Detecting music deepfakes is easy but actually hard
标题:检测音乐Deepfakes很容易,但实际上很难
链接:https://arxiv.org/abs/2405.04181
作者:Darius Afchar,Gabriel Meseguer Brocal,Romain Hennequin
备注:Under review
摘要:面对生成模型的新时代,人工生成内容的检测已成为至关重要的问题。在用户友好的平台上在几秒钟内创建可信的分钟长的音乐deepfake的能力对流媒体服务和人类艺术家的不公平竞争构成了真正的欺诈威胁。本文展示了在包含真实音频和假重建的数据集上训练分类器的可能性(以及令人惊讶的容易性),达到了令人信服的99.8%的准确率。据我们所知,这标志着音乐deepfake检测器的首次发布,这是一种有助于监管音乐伪造的工具。尽管如此,根据其他领域数十年来关于伪造检测的文献,我们强调,好的测试分数并不是故事的结局。我们从简单的机器学习框架中退一步,揭示了这种部署的检测器可能存在的许多问题:校准、对音频操作的鲁棒性、对看不见的模型的泛化、可解释性和追索的可能性。这第二部分作为该领域未来研究步骤的定位,并对繁荣的虚假内容检查器市场提出警告。
摘要:In the face of a new era of generative models, the detection of artificially generated content has become a matter of utmost importance. The ability to create credible minute-long music deepfakes in a few seconds on user-friendly platforms poses a real threat of fraud on streaming services and unfair competition to human artists. This paper demonstrates the possibility (and surprising ease) of training classifiers on datasets comprising real audio and fake reconstructions, achieving a convincing accuracy of 99.8%. To our knowledge, this marks the first publication of a music deepfake detector, a tool that will help in the regulation of music forgery. Nevertheless, informed by decades of literature on forgery detection in other fields, we stress that a good test score is not the end of the story. We step back from the straightforward ML framework and expose many facets that could be problematic with such a deployed detector: calibration, robustness to audio manipulation, generalisation to unseen models, interpretability and possibility for recourse. This second part acts as a position for future research steps in the field and a caveat to a flourishing market of fake content checkers.


【7】 Fine-grained Speech Sentiment Analysis in Chinese Psychological Support  Hotlines Based on Large-scale Pre-trained Model
标题:基于大规模预训练模型的中国心理支持热线细粒度言语情绪分析
链接:https://arxiv.org/abs/2405.04128
作者:Zhonglong Chen,Changwei Song,Yining Chen,Jianqiang Li,Guanghui Fu,Yongsheng Tong,Qing Zhao
摘要:自杀和自杀行为仍然是公共政策和医疗保健的重大挑战。为此,世界各地设立了心理支助热线,为处于精神危机中的个人提供即时帮助。这些热线的有效性在很大程度上取决于准确识别来电者的情绪状态,特别是表明自杀风险增加的潜在负面情绪。然而,对心理干预的高需求往往导致专业操作人员的短缺,突出了对有效的语音情感识别模型的需求。该模型将自动检测和分析呼叫者的情绪,便于整合到热线服务中。此外,它将使心理支持热线互动的大规模数据分析,以探索跨人群的心理现象和行为。我们的研究利用了中国最大的自杀热线北京心理支持热线的数据。我们分析了来自105个呼叫者的语音数据,包含20,630个片段,并将其分为11种类型的负面情绪。我们使用大规模的预训练模型开发了一个负面情绪识别模型和一个细粒度的多标签分类模型。我们的实验表明,负面情绪识别模型达到了最大的F1-得分为76.96%。然而,它在细粒度多标签分类任务中表现出有限的功效,最好的模型仅获得41.74%的加权F1分数。我们对这项任务进行了错误分析,讨论了未来可能的改进,并考虑了我们研究的临床应用可能性。所有代码都是公开的。
摘要:Suicide and suicidal behaviors remain significant challenges for public policy and healthcare. In response, psychological support hotlines have been established worldwide to provide immediate help to individuals in mental crises. The effectiveness of these hotlines largely depends on accurately identifying callers' emotional states, particularly underlying negative emotions indicative of increased suicide risk. However, the high demand for psychological interventions often results in a shortage of professional operators, highlighting the need for an effective speech emotion recognition model. This model would automatically detect and analyze callers' emotions, facilitating integration into hotline services. Additionally, it would enable large-scale data analysis of psychological support hotline interactions to explore psychological phenomena and behaviors across populations. Our study utilizes data from the Beijing psychological support hotline, the largest suicide hotline in China. We analyzed speech data from 105 callers containing 20,630 segments and categorized them into 11 types of negative emotions. We developed a negative emotion recognition model and a fine-grained multi-label classification model using a large-scale pre-trained model. Our experiments indicate that the negative emotion recognition model achieves a maximum F1-score of 76.96%. However, it shows limited efficacy in the fine-grained multi-label classification task, with the best model achieving only a 41.74% weighted F1-score. We conducted an error analysis for this task, discussed potential future improvements, and considered the clinical application possibilities of our study. All the codes are public available.


【8】 Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
标题:基于动态图的自适应语音情感表示学习
链接:https://arxiv.org/abs/2405.03956
作者:Yingxue Gao,Huan Zhao,Zixing Zhang
备注:None
摘要:图表示学习因其在提取代表节点嵌入方面的强大非线性拟合能力而成为研究热点。然而,对于语音信号这样的序列数据,大多数传统的方法仅仅关注序列内创建的静态图,而在很大程度上忽略了这些数据的内在演变模式。这可能会降低顺序数据的图形表示学习的效率。为此,我们提出了一种自适应的图表示学习方法的基础上动态进化的图,这是连续构建的一系列连续性分割的滑动窗口。在这样做的时候,最好是在一个长序列中捕获局部和全局上下文信息。此外,我们引入了一种加权的方法来更新节点表示,而不是传统的平均值,其中的权重是通过一种新的矩阵计算的基础上的程度相邻节点。最后,我们构造了一个可学习的图卷积层,它结合了图结构损失和分类损失来优化图结构。为了验证所提出的方法的有效性,我们进行了实验的语音情感识别IEMOCAP和RAVDESS数据集。实验结果表明,该方法优于最新的(非)基于图的模型。
摘要:Graph representation learning has become a hot research topic due to its powerful nonlinear fitting capability in extracting representative node embeddings. However, for sequential data such as speech signals, most traditional methods merely focus on the static graph created within a sequence, and largely overlook the intrinsic evolving patterns of these data. This may reduce the efficiency of graph representation learning for sequential data. For this reason, we propose an adaptive graph representation learning method based on dynamically evolved graphs, which are consecutively constructed on a series of subsequences segmented by a sliding window. In doing this, it is better to capture local and global context information within a long sequence. Moreover, we introduce a weighted approach to update the node representation rather than the conventional average one, where the weights are calculated by a novel matrix computation based on the degree of neighboring nodes. Finally, we construct a learnable graph convolutional layer that combines the graph structure loss and classification loss to optimize the graph structure. To verify the effectiveness of the proposed method, we conducted experiments for speech emotion recognition on the IEMOCAP and RAVDESS datasets. Experimental results show that the proposed method outperforms the latest (non-)graph-based models.

【9】 Intelligent Cardiac Auscultation for Murmur Detection via  Parallel-Attentive Models with Uncertainty Estimation
标题:通过具有不确定性估计的注意力模型进行智能心脏听诊以检测低语
链接:https://arxiv.org/abs/2405.03953
作者:Zixing Zhang,Tao Pang,Jing Han,Björn W. Schuller
备注:None
摘要:心脏杂音是心血管疾病的常见表现,可为早期心脏异常提供重要线索。虽然目前大多数研究方法主要关注模型的准确性,但它们往往忽略了其他重要方面,例如机器学习算法的可解释性和预测的不确定性。本文介绍了一种基于并行注意模型的心脏杂音检测方法,该方法由两个分支组成:一个是基于自注意模块的,另一个是基于卷积网络的。与传统方法不同,这种结构更适合处理序列数据中的长期依赖关系,从而有效地捕获心脏杂音的局部和全局特征。此外,我们认识到理解医学领域模型预测的不确定性对临床决策的重要性。因此,我们已经纳入了一个有效的不确定性估计方法的基础上,蒙特卡罗辍学到我们的模型。此外,我们还采用了温度标度来校准概率模型的预测,提高了其可靠性。在CirCor Digiscope数据集上进行的心脏杂音检测实验中,我们提出的方法实现了79.8%的加权准确率和65.1%的F1,代表了最先进的结果。
摘要:Heart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focus on the accuracy of models, they often overlook other important aspects such as the interpretability of machine learning algorithms and the uncertainty of predictions. This paper introduces a heart murmur detection method based on a parallel-attentive model, which consists of two branches: One is based on a self-attention module and the other one is based on a convolutional network. Unlike traditional approaches, this structure is better equipped to handle long-term dependencies in sequential data, and thus effectively captures the local and global features of heart murmurs. Additionally, we acknowledge the significance of understanding the uncertainty of model predictions in the medical field for clinical decision-making. Therefore, we have incorporated an effective uncertainty estimation method based on Monte Carlo Dropout into our model. Furthermore, we have employed temperature scaling to calibrate the predictions of our probabilistic model, enhancing its reliability. In experiments conducted on the CirCor Digiscope dataset for heart murmur detection, our proposed method achieves a weighted accuracy of 79.8% and an F1 of 65.1%, representing state-of-the-art results.


【10】 HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's  Disease Detection From Spontaneous Speech
标题:HAFFormer:从自发言语检测阿尔茨海默病的分层免注意框架
链接:https://arxiv.org/abs/2405.03952
作者:Zhongren Dong,Zixing Zhang,Weixiang Xu,Jing Han,Jianjun Ou,Björn W. Schuller
备注:None
摘要:从自发言语中自动检测阿尔茨海默病(Alzheimer 'sDisease,AD)对AD的早期诊断有重要意义。最近的方法高度依赖于Transformer架构,因为它在建模长距离上下文依赖关系方面的效率很高。然而,当在边缘设备上部署此类模型时,与自我注意力和音频长度相关的计算复杂性的二次增加构成了挑战。在这种情况下,我们构建了一个新的框架,即分层无注意力Transformer(HAFFormer),以更好地处理长语音AD检测。具体来说,我们采用多尺度相关卷积的无注意模块来代替自注意,从而避免了昂贵的计算,并采用基于GELU的门控线性单元来代替前馈层,旨在自动过滤掉冗余信息。此外,我们设计了一个层次结构,迫使它学习各种信息颗粒,从框架级到对话级。通过在ADReSS-M数据集上进行广泛的实验,引入的HAFFormer可以实现与其他近期工作竞争的结果(82.6%的准确率),但与标准Transformer相比,计算复杂性和模型大小显著降低。这显示了HAFFormer在处理用于AD检测的长音频方面的效率。
摘要:Automatically detecting Alzheimer's Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectures due to its efficiency in modelling long-range context dependencies. However, the quadratic increase in computational complexity associated with self-attention and the length of audio poses a challenge when deploying such models on edge devices. In this context, we construct a novel framework, namely Hierarchical Attention-Free Transformer (HAFFormer), to better deal with long speech for AD detection. Specifically, we employ an attention-free module of Multi-Scale Depthwise Convolution to replace the self-attention and thus avoid the expensive computation, and a GELU-based Gated Linear Unit to replace the feedforward layer, aiming to automatically filter out the redundant information. Moreover, we design a hierarchical structure to force it to learn a variety of information grains, from the frame level to the dialogue level. By conducting extensive experiments on the ADReSS-M dataset, the introduced HAFFormer can achieve competitive results (82.6% accuracy) with other recent work, but with significant computational complexity and model size reduction compared to the standard Transformer. This shows the efficiency of HAFFormer in dealing with long audio for AD detection.

【11】 A 65nm 36nJ/Decision Bio-inspired Temporal-Sparsity-Aware Digital  Keyword Spotting IC with 0.6V Near-Threshold SRAM
标题:65纳米36 nJ/Decision受生物启发的时间稀疏感知数字关键字发现IC,配备0.6V近阈值静态存储器
链接:https://arxiv.org/abs/2405.03905
作者:Qinyu Chen,Kwantae Kim,Chang Gao,Sheng Zhou,Taekwang Jang,Tobi Delbruck,Shih-Chii Liu
摘要:本文介绍了,据作者所知,第一个细粒度的时间稀疏意识的关键字定位(KWS)IC利用从输入帧和网络隐藏状态提取的相邻特征向量之间的时间相似性,消除不必要的操作和内存访问。这款KWS IC采用生物启发的Delta门控递归神经网络({\Delta} RNN)分类器,实现了90.5%的11类Google语音命令数据集(GSCD)KWS准确率和36nJ/决策的能耗。在87%的时间稀疏度下,计算延迟和每次推理的能量分别减少了2.4$\times $/3.4$\times $。65纳米设计占用0.78mm$^2 $,并具有两个额外的模块,一个紧凑的0.084mm$^2 $基于数字无限脉冲响应(IIR)的带通滤波器(BPF)音频特征提取器(FEx)和一个24 kB 0.6V近Vth重量SRAM,与标准SRAM相比,读取功率低6.6$\times $。
摘要:This paper introduces, to the best of the authors' knowledge, the first fine-grained temporal sparsity-aware keyword spotting (KWS) IC leveraging temporal similarities between neighboring feature vectors extracted from input frames and network hidden states, eliminating unnecessary operations and memory accesses. This KWS IC, featuring a bio-inspired delta-gated recurrent neural network ({\Delta}RNN) classifier, achieves an 11-class Google Speech Command Dataset (GSCD) KWS accuracy of 90.5% and energy consumption of 36nJ/decision. At 87% temporal sparsity, computing latency and energy per inference are reduced by 2.4$\times$/3.4$\times$, respectively. The 65nm design occupies 0.78mm$^2$ and features two additional blocks, a compact 0.084mm$^2$ digital infinite-impulse-response (IIR)-based band-pass filter (BPF) audio feature extractor (FEx) and a 24kB 0.6V near-Vth weight SRAM with 6.6$\times$ lower read power compared to the standard SRAM.

机器翻译由腾讯交互翻译提供,仅供参考