今日论文合集:cs.SD语音1篇,eess.AS音频处理3篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】A lightweight dual-stage framework for personalized speech enhancement  based on DeepFilterNet2
标题:基于DeepLayer Net 2的个性化语音增强轻量级双级框架
链接:https://arxiv.org/abs/2404.08022
作者:Thomas Serre,Mathieu Fontaine,Éric Benhaim,Geoffroy Dutour,Slim Essid
备注:None
摘要:在嘈杂的声学环境中,从多个扬声器中分离出所需扬声器的声音是一项具有挑战性的任务。个性化语音增强(PSE)通过利用说话人声音的先验知识来努力克服这一点。最近的研究工作已经产生了有前途的PSE模型,尽管通常伴随着计算密集型架构,不适合资源受限的嵌入式设备。在本文中,我们介绍了一种新的方法来personalize一个轻量级的双阶段语音增强(SE)模型,并实现它在DeepFilterNet2,SE模型以其最先进的性能而闻名。我们寻求一个最佳的集成模型内的说话人信息,探索不同的位置,为一体化的嵌入式双阶段增强架构。我们还研究了一个定制的训练策略,当DeepFilterNet2适应PSE任务时。我们表明,我们的个性化方法大大提高了DeepFilterNet2的性能,同时保持最小的计算开销。
摘要:Isolating the desired speaker's voice amidst multiplespeakers in a noisy acoustic context is a challenging task. Per-sonalized speech enhancement (PSE) endeavours to achievethis by leveraging prior knowledge of the speaker's voice.Recent research efforts have yielded promising PSE mod-els, albeit often accompanied by computationally intensivearchitectures, unsuitable for resource-constrained embeddeddevices. In this paper, we introduce a novel method to per-sonalize a lightweight dual-stage Speech Enhancement (SE)model and implement it within DeepFilterNet2, a SE modelrenowned for its state-of-the-art performance. We seek anoptimal integration of speaker information within the model,exploring different positions for the integration of the speakerembeddings within the dual-stage enhancement architec-ture. We also investigate a tailored training strategy whenadapting DeepFilterNet2 to a PSE task. We show that ourpersonalization method greatly improves the performancesof DeepFilterNet2 while preserving minimal computationaloverhead.

eess.AS音频处理
【1】 The Impact of Speech Anonymization on Pathology and Its Limits
标题:言语神经化对病理学的影响及其局限性
链接:https://arxiv.org/abs/2404.08064
作者:Soroosh Tayebi Arasteh,Tomas Arias-Vergara,Paula Andrea Perez-Toro,Tobias Weise,Kai Packhaeuser,Maria Schuster,Elmar Noeth,Andreas Maier,Seung Hee Yang
摘要:将语音整合到医疗保健中,由于其作为包含个人生物特征信息的非侵入性生物标志物的潜力,加剧了隐私问题。作为回应,说话人匿名化的目的是隐藏个人身份信息,同时保留关键的语言内容。然而,匿名化技术的应用病理性言论,隐私是特别重要的一个关键领域,还没有得到广泛的研究。这项研究调查了匿名化对来自多个德国机构的2,700多名发言者的病理性言论的影响,重点是隐私,病理效用和人口公平性。我们探讨了基于训练和基于信号处理的匿名化方法,并记录了各种疾病的实质性隐私改善,证明了相等的错误率增加了1933%,对效用的总体影响最小。特定的疾病,如构音障碍,发音困难,唇腭裂经历了最小的效用变化,而语言障碍表现出轻微的改善。我们的研究结果强调,匿名化的影响在不同的疾病中差异很大。这就需要特定于疾病的匿名化策略,以最佳地平衡隐私与诊断效用。此外,我们的公平性分析显示,在大多数人口统计数据中,匿名化效果是一致的。这项研究证明了匿名化在病理性语音中增强隐私的有效性,同时也强调了定制方法来应对反转攻击的重要性。
摘要:Integration of speech into healthcare has intensified privacy concerns due to its potential as a non-invasive biomarker containing individual biometric information. In response, speaker anonymization aims to conceal personally identifiable information while retaining crucial linguistic content. However, the application of anonymization techniques to pathological speech, a critical area where privacy is especially vital, has not been extensively examined. This study investigates anonymization's impact on pathological speech across over 2,700 speakers from multiple German institutions, focusing on privacy, pathological utility, and demographic fairness. We explore both training-based and signal processing-based anonymization methods, and document substantial privacy improvements across disorders-evidenced by equal error rate increases up to 1933%, with minimal overall impact on utility. Specific disorders such as Dysarthria, Dysphonia, and Cleft Lip and Palate experienced minimal utility changes, while Dysglossia showed slight improvements. Our findings underscore that the impact of anonymization varies substantially across different disorders. This necessitates disorder-specific anonymization strategies to optimally balance privacy with diagnostic utility. Additionally, our fairness analysis revealed consistent anonymization effects across most of the demographics. This study demonstrates the effectiveness of anonymization in pathological speech for enhancing privacy, while also highlighting the importance of customized approaches to account for inversion attacks.


【2】 Guided Masked Self-Distillation Modeling for Distributed Multimedia  Sensor Event Analysis
标题:分布式多媒体传感器事件分析的引导掩蔽自蒸馏建模
链接:https://arxiv.org/abs/2404.08264
作者:Masahiro Yasuda,Noboru Harada,Yasunori Ohishi,Shoichiro Saito,Akira Nakayama,Nobutaka Ono备注:13page, 7figure, under review
摘要:分布式传感器的观测对于分析复杂和广泛的现实世界环境中的一系列人类和机器活动(本文中称为“事件”)至关重要。这是因为在这样的环境中,从单个传感器获得的信息往往丢失或分散;来自多个位置和模式的观测应该被整合,以全面分析事件。然而,尚未建立一种学习方法来提取有效地结合这种分布式观测的联合表示。因此,我们提出了引导掩蔽自适应蒸馏建模(引导MELD)的传感器间的关系建模。Guided-MELD的基本思想是学习用来自检测事件所需的其他传感器的信息来补充来自屏蔽传感器的信息。引导MELD预期使系统能够有效地提取由传感器获得的碎片或冗余的目标事件信息,而不过度依赖于任何特定的传感器。为了验证所提出的方法在分布式多媒体传感器事件分析的新任务中的有效性,我们记录了两个新的数据集,适合的问题设置:MM-Store和MM-Office。这些数据集包括便利店和办公室中的人类活动,使用分布式摄像机和麦克风记录。在这些数据集上的实验结果表明,所提出的Guided-MELD提高了事件标记和检测性能,优于传统的传感器间关系建模方法。此外,所提出的方法进行鲁棒性,即使传感器减少。
摘要:Observations with distributed sensors are essential in analyzing a series of human and machine activities (referred to as 'events' in this paper) in complex and extensive real-world environments. This is because the information obtained from a single sensor is often missing or fragmented in such an environment; observations from multiple locations and modalities should be integrated to analyze events comprehensively. However, a learning method has yet to be established to extract joint representations that effectively combine such distributed observations. Therefore, we propose Guided Masked sELf-Distillation modeling (Guided-MELD) for inter-sensor relationship modeling. The basic idea of Guided-MELD is to learn to supplement the information from the masked sensor with information from other sensors needed to detect the event. Guided-MELD is expected to enable the system to effectively distill the fragmented or redundant target event information obtained by the sensors without being overly dependent on any specific sensors. To validate the effectiveness of the proposed method in novel tasks of distributed multimedia sensor event analysis, we recorded two new datasets that fit the problem setting: MM-Store and MM-Office. These datasets consist of human activities in a convenience store and an office, recorded using distributed cameras and microphones. Experimental results on these datasets show that the proposed Guided-MELD improves event tagging and detection performance and outperforms conventional inter-sensor relationship modeling methods. Furthermore, the proposed method performed robustly even when sensors were reduced.

【3】 A lightweight dual-stage framework for personalized speech enhancement  based on DeepFilterNet2
标题:基于DeepLayer Net 2的个性化语音增强轻量级双级框架
链接:https://arxiv.org/abs/2404.08022
作者:Thomas Serre,Mathieu Fontaine,Éric Benhaim,Geoffroy Dutour,Slim Essid
备注:None
摘要:在嘈杂的声学环境中,从多个扬声器中分离出所需扬声器的声音是一项具有挑战性的任务。个性化语音增强(PSE)通过利用说话人声音的先验知识来努力克服这一点。最近的研究工作已经产生了有前途的PSE模型,尽管通常伴随着计算密集型架构,不适合资源受限的嵌入式设备。在本文中,我们介绍了一种新的方法来personalize一个轻量级的双阶段语音增强(SE)模型,并实现它在DeepFilterNet2,SE模型以其最先进的性能而闻名。我们寻求一个最佳的集成模型内的说话人信息,探索不同的位置,为一体化的嵌入式双阶段增强架构。我们还研究了一个定制的训练策略,当DeepFilterNet2适应PSE任务时。我们表明,我们的个性化方法大大提高了DeepFilterNet2的性能,同时保持最小的计算开销。
摘要:Isolating the desired speaker's voice amidst multiplespeakers in a noisy acoustic context is a challenging task. Per-sonalized speech enhancement (PSE) endeavours to achievethis by leveraging prior knowledge of the speaker's voice.Recent research efforts have yielded promising PSE mod-els, albeit often accompanied by computationally intensivearchitectures, unsuitable for resource-constrained embeddeddevices. In this paper, we introduce a novel method to per-sonalize a lightweight dual-stage Speech Enhancement (SE)model and implement it within DeepFilterNet2, a SE modelrenowned for its state-of-the-art performance. We seek anoptimal integration of speaker information within the model,exploring different positions for the integration of the speakerembeddings within the dual-stage enhancement architec-ture. We also investigate a tailored training strategy whenadapting DeepFilterNet2 to a PSE task. We show that ourpersonalization method greatly improves the performancesof DeepFilterNet2 while preserving minimal computationaloverhead.


机器翻译由腾讯交互翻译提供,仅供参考