今天跟大家分享一篇语音相关的论文合集:cs.SD语音7篇,eess.AS音频处理7篇。
【1】 TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial Training
标题:TGAVC:利用文本引导和对抗训练改进自动编码器语音转换
链接:https://arxiv.org/abs/2208.04035
作者:Huaizhen Tang,Xulong Zhang,Jianzong Wang,Ning Cheng,Zhen Zeng,Edward Xiao,Jing Xiao机构:Ping An Technology (Shenzhen) Co., Ltd., University of Science and Technology of China., Aquinas International Academy, CA, USA摘要:非并行多对多语音转换是语音处理领域中一个有趣而又富有挑战性的课题.近年来,基于条件自动编码器的AutoVC方法利用信息约束瓶颈,将说话人身份和语音内容分离,取得了很好的转换效果.但是,由于采用的是单纯的自动编码器训练方法,很难对说话人身份和语音内容的分离效果进行评价.该文提出了一种基于条件自动编码器的多对多语音转换方法,一种新颖的语音转换框架,名为(TGAVC)提出了一种基于文本转录产生的期望内容嵌入来指导语音内容的提取,以更有效地从语音中分离出内容和音质.此外,通过对抗性训练消除从语音中提取的估计内容嵌入中的说话人身份信息,在AIShell-3数据集上的实验表明,该模型在转换后语音的自然度和相似度方面优于AutoVC.摘要:Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excellent conversion results by disentangling the speaker identity and the speech content using information-constraining bottlenecks. However, due to the pure autoencoder training method, it is difficult to evaluate the separation effect of content and speaker identity. In this paper, a novel voice conversion framework, named $\boldsymbol T$ext $\boldsymbol G$uided $\boldsymbol A$utoVC(TGAVC), is proposed to more effectively separate content and timbre from speech, where an expected content embedding produced based on the text transcriptions is designed to guide the extraction of voice content. In addition, the adversarial training is applied to eliminate the speaker identity information in the estimated content embedding extracted from speech. Under the guidance of the expected content embedding and the adversarial training, the content encoder is trained to extract speaker-independent content embedding from speech. Experiments on AIShell-3 dataset show that the proposed model outperforms AutoVC in terms of naturalness and similarity of converted speech.
【2】 Debiased Cross-modal Matching for Content-based Micro-video Background Music Recommendation
标题:基于内容的微视频背景音乐推荐的去偏跨模态匹配
链接:https://arxiv.org/abs/2208.03633
作者:Jinng Yi,Zhenzhong Chen摘要:微视频背景音乐推荐是一项复杂的任务,其中视频与上传者选择的背景音乐之间的匹配度是一个主要问题,然而,用户生成内容的选择针对目前UGC(UGC)中存在的知识局限性和上传者对音乐的历史偏好等问题,我们提出了一种去偏的跨模态(DebCM)匹配模型来减轻这种选择偏差的影响,具体来说,我们设计了一个教师—学生网络来利用音乐视频片段的匹配,这是专业制作的内容(PGC),以专业的音乐匹配技术,以更好地缓解用户知识不足所造成的偏差。PGC数据由教师网络捕获,通过基于知识库的知识转移来指导学生网络的上传者选择的UGC数据的匹配。此外,上传者对音乐流派的个人偏好被识别为虚假地将音乐嵌入和背景音乐选择相关联的混杂因素,从而导致学习的推荐器系统过度推荐来自多数组的音乐。为了解决学生网络的UGC数据中的这种混杂因素,我们利用后门调整来消除音乐嵌入和预测得分之间的虚假相关性(MC)估计量,以批水平平均值作为近似值,以避免整合调整计算的整个混杂空间。体裁数据集证明了所提出的方法对选择偏差的有效性。代码可在以下网站上公开获得:\url{https://github.com/jing-1/DebCM}.摘要:Micro-video background music recommendation is a complicated task where the matching degree between videos and uploader-selected background music is a major issue. However, the selection of the user-generated content (UGC) is biased caused by knowledge limitations and historical preferences among music of each uploader. In this paper, we propose a Debiased Cross-Modal (DebCM) matching model to alleviate the influence of such selection bias. Specifically, we design a teacher-student network to utilize the matching of segments of music videos, which is professional-generated content (PGC) with specialized music-matching techniques, to better alleviate the bias caused by insufficient knowledge of users. The PGC data is captured by a teacher network to guide the matching of uploader-selected UGC data of the student network by KL-based knowledge transfer. In addition, uploaders' personal preferences of music genres are identified as confounders that spuriously correlate music embeddings and background music selections, resulting in the learned recommender system to over-recommend music from the majority groups. To resolve such confounders in the UGC data of the student network, backdoor adjustment is utilized to deconfound the spurious correlation between music embeddings and prediction scores. We further utilize Monte Carlo (MC) estimator with batch-level average as the approximations to avoid integrating the entire confounder space calculated by the adjustment. Extensive experiments on the TT-150k-genre dataset demonstrate the effectiveness of the proposed method towards the selection bias. The code is publicly available on: \url{https://github.com/jing-1/DebCM}.
【3】 Chronological Self-Training for Real-Time Speaker Diarization
标题:基于时序自训练的实时说话人识别
链接:https://arxiv.org/abs/2208.03393
作者:Dirk Padfield,Daniel J. Liebling摘要:Diarization根据说话者的声音将音频流划分为多个段。包括注册步骤的实时Diarization系统应限制注册训练样本以减少用户交互时间。尽管在少量样本上训练会产生较差的性能,我们表明使用按时间顺序的自训练方法可以显著地提高准确性。我们研究了训练时间和分类性能之间的权衡,发现1秒足以达到95%以上的准确率。我们评估了来自6种不同语言的700个音频对话文件,每个文件大约10分钟,并证明了平均日记错误率为低至10%。摘要:Diarization partitions an audio stream into segments based on the voices of the speakers. Real-time diarization systems that include an enrollment step should limit enrollment training samples to reduce user interaction time. Although training on a small number of samples yields poor performance, we show that the accuracy can be improved dramatically using a chronological self-training approach. We studied the tradeoff between training time and classification performance and found that 1 second is sufficient to reach over 95% accuracy. We evaluated on 700 audio conversation files of about 10 minutes each from 6 different languages and demonstrated average diarization error rates as low as 10%.
【4】 Variational Autoencoders for Anomaly Detection in Respiratory Sounds
标题:用于呼吸音异常检测的变分自动编码器
链接:https://arxiv.org/abs/2208.03326
作者:Michele Cozzatti,Federico Simonetta,Stavros Ntalampiras机构:Ntalampiras[,−,−,−,], LIM – Music Informatics Laboratory, Department of Computer Science, University of Milano备注:Published at ICANN 2022摘要:本文提出了一种基于弱监督机器学习的方法,旨在提供一种工具来提醒患者可能的呼吸系统疾病.各种类型的病理可能影响呼吸系统,潜在地导致严重的疾病,在某些情况下,甚至死亡.一般来说,有效的预防措施被认为是改善病人健康状况的主要因素。该方法利用变分自动编码器结构,允许使用有限复杂度的训练流水线和相对小尺寸的数据集,重要的是,该方法提供了57%的准确率,这与现有的强监督方法是一致的。摘要:This paper proposes a weakly-supervised machine learning-based approach aiming at a tool to alert patients about possible respiratory diseases. Various types of pathologies may affect the respiratory system, potentially leading to severe diseases and, in certain cases, death. In general, effective prevention practices are considered as major actors towards the improvement of the patient's health condition. The proposed method strives to realize an easily accessible tool for the automatic diagnosis of respiratory diseases. Specifically, the method leverages Variational Autoencoder architectures permitting the usage of training pipelines of limited complexity and relatively small-sized datasets. Importantly, it offers an accuracy of 57 %, which is in line with the existing strongly-supervised approaches.
【5】 FRA-RIR: Fast Random Approximation of the Image-source Method
标题:FRA-RIR:像源法的快速随机逼近
链接:https://arxiv.org/abs/2208.04101
摘要:现代语音处理系统的训练往往需要大量的模拟房间脉冲响应然而,模拟真实的RIR数据通常需要精确的物理建模,并且这种模拟过程的加速通常需要诸如图形处理单元之类的某些计算平台本文提出了一种快速随机逼近的图像源法—-FRA-RIR FRA-RIR用一系列随机近似代替了标准ISM中的物理模拟,实验表明,FRA-RIR算法不仅比现有的ISM算法速度快,而且可以应用于实时数据生成流水线中.基于标准计算平台上的RIR仿真工具,而且当使用仿真RIR进行训练时,还提高了在真实世界RIR上评估的语音去噪系统的性能。FRA-RIR的Python实现可在线获得\footnote{\url{https://github.com/yluo42/FRA-RIR}}.摘要:The training of modern speech processing systems often requires a large amount of simulated room impulse response (RIR) data in order to allow the systems to generalize well in real-world, reverberant environments. However, simulating realistic RIR data typically requires accurate physical modeling, and the acceleration of such simulation process typically requires certain computational platforms such as a graphics processing unit (GPU). In this paper, we propose FRA-RIR, a fast random approximation method of the widely-used image-source method (ISM), to efficiently generate realistic RIR data without specific computational devices. FRA-RIR replaces the physical simulation in the standard ISM by a series of random approximations, which significantly speeds up the simulation process and enables its application in on-the-fly data generation pipelines. Experiments show that FRA-RIR can not only be significantly faster than other existing ISM-based RIR simulation tools on standard computational platforms, but also improves the performance of speech denoising systems evaluated on real-world RIR when trained with simulated RIR. A Python implementation of FRA-RIR is available online\footnote{\url{https://github.com/yluo42/FRA-RIR}}.
【6】 SSDPT: Self-Supervised Dual-Path Transformer for Anomalous Sound Detection in Machine Condition Monitoring
标题:SSDPT:用于机器状态监测中异常声音检测的自监督双路Transformer
链接:https://arxiv.org/abs/2208.03421
作者:Jisheng Bai,Jianfeng Chen,Mou Wang,Muhammad Saad Ayub,Qingli Yan摘要:机器状态监测中的异常声音检测在工业4.0的发展中有着巨大的潜力。然而,机器的异常声音在正常情况下通常是不存在的。因此,所使用的模型必须学习正常声音的声学表征以进行训练,并在测试时检测异常声音。本文在分析了机器异常声音检测的基础上,提出了一种新的机器异常声音检测方法。我们提出了一种自监督双路径Transformer(SSDPT)网络对机器监测中的异常声音进行检测,SSDPT网络将声音特征分割成多个段,并使用多个DPT块进行时间和频率建模。DPT模块使用注意力模块交替地对分割的声学特征的频率和时间分量的交互信息进行建模,为了解决缺少异常声音的问题,我们采用了一种自监督学习的方法,用正常声音来训练网络,具体来说,该方法随机地掩蔽和重构声学特征,并对机器身份信息进行联合分类以提高异常声音检测的性能.实验结果表明,与现有的异常声音检测方法相比,SSDPT网络实现了谐波平均AUC分数的显著提高。摘要:Anomalous sound detection for machine condition monitoring has great potential in the development of Industry 4.0. However, these anomalous sounds of machines are usually unavailable in normal conditions. Therefore, the models employed have to learn acoustic representations with normal sounds for training, and detect anomalous sounds while testing. In this article, we propose a self-supervised dual-path Transformer (SSDPT) network to detect anomalous sounds in machine monitoring. The SSDPT network splits the acoustic features into segments and employs several DPT blocks for time and frequency modeling. DPT blocks use attention modules to alternately model the interactive information about the frequency and temporal components of the segmented acoustic features. To address the problem of lack of anomalous sound, we adopt a self-supervised learning approach to train the network with normal sound. Specifically, this approach randomly masks and reconstructs the acoustic features, and jointly classifies machine identity information to improve the performance of anomalous sound detection. We evaluated our method on the DCASE2021 task2 dataset. The experimental results show that the SSDPT network achieves a significant increase in the harmonic mean AUC score, in comparison to present state-of-the-art methods of anomalous sound detection.
【7】 Virtual Analog Modeling of Distortion Circuits Using Neural Ordinary Differential Equations
标题:基于神经常微分方程的失真电路虚拟模拟建模
链接:https://arxiv.org/abs/2205.01897
作者:Jan Wilczek,Alec Wright,Vesa Välimäki,Emanuël Habets机构:WolfSound, Katowice, Poland, Acoustics Lab, Dept. Signal Processing and Acoustics, Aalto University, Espoo, Finland, Emanuël A. P. Habets, International Audio Laboratories Erlangen ‡, Erlangen, Germany, emanuel.habets备注:8 pages, 10 figures, accepted for DAFx 2022 conference, for associated audio examples, see this https URL摘要:深度学习的最新研究表明神经网络可以学习控制动力系统的微分方程,我们将此概念应用于虚拟模拟(VA)建模学习常微分方程(ODE)控制一阶和二阶二极管削波器。所提出的模型实现了与最先进的递归神经网络相当的性能我们证明了这种方法不需要过采样,并且允许在训练完成之后增加采样率,这导致增加的精度。使用复杂的数值解算器允许以较慢的处理为代价来增加精度。通过这种方式学习的常微分方程不需要闭合形式,但仍然可以进行物理解释。摘要:Recent research in deep learning has shown that neural networks can learn differential equations governing dynamical systems. In this paper, we adapt this concept to Virtual Analog (VA) modeling to learn the ordinary differential equations (ODEs) governing the first-order and the second-order diode clipper. The proposed models achieve performance comparable to state-of-the-art recurrent neural networks (RNNs) albeit using fewer parameters. We show that this approach does not require oversampling and allows to increase the sampling rate after the training has completed, which results in increased accuracy. Using a sophisticated numerical solver allows to increase the accuracy at the cost of slower processing. ODEs learned this way do not require closed forms but are still physically interpretable.
【1】 FRA-RIR: Fast Random Approximation of the Image-source Method
标题:FRA-RIR:像源法的快速随机逼近
链接:https://arxiv.org/abs/2208.04101
* 与cs.SD语音【5】为同一篇
摘要:现代语音处理系统的训练往往需要大量的模拟房间脉冲响应然而,模拟真实的RIR数据通常需要精确的物理建模,并且这种模拟过程的加速通常需要诸如图形处理单元之类的某些计算平台本文提出了一种快速随机逼近的图像源法—-FRA-RIR FRA-RIR用一系列随机近似代替了标准ISM中的物理模拟,实验表明,FRA-RIR算法不仅比现有的ISM算法速度快,而且可以应用于实时数据生成流水线中.基于标准计算平台上的RIR仿真工具,而且当使用仿真RIR进行训练时,还提高了在真实世界RIR上评估的语音去噪系统的性能。FRA-RIR的Python实现可在线获得\footnote{\url{https://github.com/yluo42/FRA-RIR}}.摘要:The training of modern speech processing systems often requires a large amount of simulated room impulse response (RIR) data in order to allow the systems to generalize well in real-world, reverberant environments. However, simulating realistic RIR data typically requires accurate physical modeling, and the acceleration of such simulation process typically requires certain computational platforms such as a graphics processing unit (GPU). In this paper, we propose FRA-RIR, a fast random approximation method of the widely-used image-source method (ISM), to efficiently generate realistic RIR data without specific computational devices. FRA-RIR replaces the physical simulation in the standard ISM by a series of random approximations, which significantly speeds up the simulation process and enables its application in on-the-fly data generation pipelines. Experiments show that FRA-RIR can not only be significantly faster than other existing ISM-based RIR simulation tools on standard computational platforms, but also improves the performance of speech denoising systems evaluated on real-world RIR when trained with simulated RIR. A Python implementation of FRA-RIR is available online\footnote{\url{https://github.com/yluo42/FRA-RIR}}.
【2】 SSDPT: Self-Supervised Dual-Path Transformer for Anomalous Sound Detection in Machine Condition Monitoring
标题:SSDPT:用于机器状态监测中异常声音检测的自监督双路Transformer
链接:https://arxiv.org/abs/2208.03421
* 与cs.SD语音【6】为同一篇
作者:Jisheng Bai,Jianfeng Chen,Mou Wang,Muhammad Saad Ayub,Qingli Yan摘要:机器状态监测中的异常声音检测在工业4.0的发展中有着巨大的潜力。然而,机器的异常声音在正常情况下通常是不存在的。因此,所使用的模型必须学习正常声音的声学表征以进行训练,并在测试时检测异常声音。本文在分析了机器异常声音检测的基础上,提出了一种新的机器异常声音检测方法。我们提出了一种自监督双路径Transformer(SSDPT)网络对机器监测中的异常声音进行检测,SSDPT网络将声音特征分割成多个段,并使用多个DPT块进行时间和频率建模。DPT模块使用注意力模块交替地对分割的声学特征的频率和时间分量的交互信息进行建模,为了解决缺少异常声音的问题,我们采用了一种自监督学习的方法,用正常声音来训练网络,具体来说,该方法随机地掩蔽和重构声学特征,并对机器身份信息进行联合分类以提高异常声音检测的性能.实验结果表明,与现有的异常声音检测方法相比,SSDPT网络实现了谐波平均AUC分数的显著提高。摘要:Anomalous sound detection for machine condition monitoring has great potential in the development of Industry 4.0. However, these anomalous sounds of machines are usually unavailable in normal conditions. Therefore, the models employed have to learn acoustic representations with normal sounds for training, and detect anomalous sounds while testing. In this article, we propose a self-supervised dual-path Transformer (SSDPT) network to detect anomalous sounds in machine monitoring. The SSDPT network splits the acoustic features into segments and employs several DPT blocks for time and frequency modeling. DPT blocks use attention modules to alternately model the interactive information about the frequency and temporal components of the segmented acoustic features. To address the problem of lack of anomalous sound, we adopt a self-supervised learning approach to train the network with normal sound. Specifically, this approach randomly masks and reconstructs the acoustic features, and jointly classifies machine identity information to improve the performance of anomalous sound detection. We evaluated our method on the DCASE2021 task2 dataset. The experimental results show that the SSDPT network achieves a significant increase in the harmonic mean AUC score, in comparison to present state-of-the-art methods of anomalous sound detection.
【3】 TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial Training
标题:TGAVC:利用文本引导和对抗训练改进自动编码器语音转换
链接:https://arxiv.org/abs/2208.04035
* 与cs.SD语音【1】为同一篇
作者:Huaizhen Tang,Xulong Zhang,Jianzong Wang,Ning Cheng,Zhen Zeng,Edward Xiao,Jing Xiao机构:Ping An Technology (Shenzhen) Co., Ltd., University of Science and Technology of China., Aquinas International Academy, CA, USA摘要:非并行多对多语音转换是语音处理领域中一个有趣而又富有挑战性的课题.近年来,基于条件自动编码器的AutoVC方法利用信息约束瓶颈,将说话人身份和语音内容分离,取得了很好的转换效果.但是,由于采用的是单纯的自动编码器训练方法,很难对说话人身份和语音内容的分离效果进行评价.该文提出了一种基于条件自动编码器的多对多语音转换方法,一种新颖的语音转换框架,名为(TGAVC)提出了一种基于文本转录产生的期望内容嵌入来指导语音内容的提取,以更有效地从语音中分离出内容和音质.此外,通过对抗性训练消除从语音中提取的估计内容嵌入中的说话人身份信息,在AIShell-3数据集上的实验表明,该模型在转换后语音的自然度和相似度方面优于AutoVC.摘要:Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excellent conversion results by disentangling the speaker identity and the speech content using information-constraining bottlenecks. However, due to the pure autoencoder training method, it is difficult to evaluate the separation effect of content and speaker identity. In this paper, a novel voice conversion framework, named $\boldsymbol T$ext $\boldsymbol G$uided $\boldsymbol A$utoVC(TGAVC), is proposed to more effectively separate content and timbre from speech, where an expected content embedding produced based on the text transcriptions is designed to guide the extraction of voice content. In addition, the adversarial training is applied to eliminate the speaker identity information in the estimated content embedding extracted from speech. Under the guidance of the expected content embedding and the adversarial training, the content encoder is trained to extract speaker-independent content embedding from speech. Experiments on AIShell-3 dataset show that the proposed model outperforms AutoVC in terms of naturalness and similarity of converted speech.
【4】 When can I Speak? Predicting initiation points for spoken dialogue agents
标题:我什么时候可以说话?预测口语对话主体的起始点
链接:https://arxiv.org/abs/2208.03812
作者:Siyan Li,Ashwin Paranjape,Christopher D. Manning摘要:当前的口语对话系统在长时间的沉默之后开始其转换(700-1000ms),这导致很少的实时反馈、缓慢的响应和整体呆板的对话流。人类通常在200ms内做出响应,并且成功地提前预测起始点将允许口语对话代理做同样的事情。在本工作中,我们使用来自预先训练的语音表示模型的韵律特征来预测到启动的前置时间(wav2vec 1.0)对来自预先训练的语言模型的用户音频和单词特征进行操作(GPT-2)。为了评估错误,我们提出了两个度量标准:预测提前期和真实提前期。我们在Switchboard语料库上对模型进行训练和评估,发现我们的方法在两个指标上都优于先前工作中的特征,并且大大优于等待的普通方法700ms的静默。摘要:Current spoken dialogue systems initiate their turns after a long period of silence (700-1000ms), which leads to little real-time feedback, sluggish responses, and an overall stilted conversational flow. Humans typically respond within 200ms and successfully predicting initiation points in advance would allow spoken dialogue agents to do the same. In this work, we predict the lead-time to initiation using prosodic features from a pre-trained speech representation model (wav2vec 1.0) operating on user audio and word features from a pre-trained language model (GPT-2) operating on incremental transcriptions. To evaluate errors, we propose two metrics w.r.t. predicted and true lead times. We train and evaluate the models on the Switchboard Corpus and find that our method outperforms features from prior work on both metrics and vastly outperforms the common approach of waiting for 700ms of silence.
【5】 Debiased Cross-modal Matching for Content-based Micro-video Background Music Recommendation
标题:基于内容的微视频背景音乐推荐的去偏跨模态匹配
链接:https://arxiv.org/abs/2208.03633
* 与cs.SD语音【2】为同一篇
作者:Jinng Yi,Zhenzhong Chen摘要:微视频背景音乐推荐是一项复杂的任务,其中视频与上传者选择的背景音乐之间的匹配度是一个主要问题,然而,用户生成内容的选择针对目前UGC(UGC)中存在的知识局限性和上传者对音乐的历史偏好等问题,我们提出了一种去偏的跨模态(DebCM)匹配模型来减轻这种选择偏差的影响,具体来说,我们设计了一个教师—学生网络来利用音乐视频片段的匹配,这是专业制作的内容(PGC),以专业的音乐匹配技术,以更好地缓解用户知识不足所造成的偏差。PGC数据由教师网络捕获,通过基于知识库的知识转移来指导学生网络的上传者选择的UGC数据的匹配。此外,上传者对音乐流派的个人偏好被识别为虚假地将音乐嵌入和背景音乐选择相关联的混杂因素,从而导致学习的推荐器系统过度推荐来自多数组的音乐。为了解决学生网络的UGC数据中的这种混杂因素,我们利用后门调整来消除音乐嵌入和预测得分之间的虚假相关性(MC)估计量,以批水平平均值作为近似值,以避免整合调整计算的整个混杂空间。体裁数据集证明了所提出的方法对选择偏差的有效性。代码可在以下网站上公开获得:\url{https://github.com/jing-1/DebCM}.摘要:Micro-video background music recommendation is a complicated task where the matching degree between videos and uploader-selected background music is a major issue. However, the selection of the user-generated content (UGC) is biased caused by knowledge limitations and historical preferences among music of each uploader. In this paper, we propose a Debiased Cross-Modal (DebCM) matching model to alleviate the influence of such selection bias. Specifically, we design a teacher-student network to utilize the matching of segments of music videos, which is professional-generated content (PGC) with specialized music-matching techniques, to better alleviate the bias caused by insufficient knowledge of users. The PGC data is captured by a teacher network to guide the matching of uploader-selected UGC data of the student network by KL-based knowledge transfer. In addition, uploaders' personal preferences of music genres are identified as confounders that spuriously correlate music embeddings and background music selections, resulting in the learned recommender system to over-recommend music from the majority groups. To resolve such confounders in the UGC data of the student network, backdoor adjustment is utilized to deconfound the spurious correlation between music embeddings and prediction scores. We further utilize Monte Carlo (MC) estimator with batch-level average as the approximations to avoid integrating the entire confounder space calculated by the adjustment. Extensive experiments on the TT-150k-genre dataset demonstrate the effectiveness of the proposed method towards the selection bias. The code is publicly available on: \url{https://github.com/jing-1/DebCM}.
【6】 Chronological Self-Training for Real-Time Speaker Diarization
标题:基于时序自训练的实时说话人识别
链接:https://arxiv.org/abs/2208.03393
* 与cs.SD语音【3】为同一篇
作者:Dirk Padfield,Daniel J. Liebling摘要:Diarization根据说话者的声音将音频流划分为多个段。包括注册步骤的实时Diarization系统应限制注册训练样本以减少用户交互时间。尽管在少量样本上训练会产生较差的性能,我们表明使用按时间顺序的自训练方法可以显著地提高准确性。我们研究了训练时间和分类性能之间的权衡,发现1秒足以达到95%以上的准确率。我们评估了来自6种不同语言的700个音频对话文件,每个文件大约10分钟,并证明了平均日记错误率为低至10%。摘要:Diarization partitions an audio stream into segments based on the voices of the speakers. Real-time diarization systems that include an enrollment step should limit enrollment training samples to reduce user interaction time. Although training on a small number of samples yields poor performance, we show that the accuracy can be improved dramatically using a chronological self-training approach. We studied the tradeoff between training time and classification performance and found that 1 second is sufficient to reach over 95% accuracy. We evaluated on 700 audio conversation files of about 10 minutes each from 6 different languages and demonstrated average diarization error rates as low as 10%.
【7】 Variational Autoencoders for Anomaly Detection in Respiratory Sounds
标题:用于呼吸音异常检测的变分自动编码器
链接:https://arxiv.org/abs/2208.03326
* 与cs.SD语音【4】为同一篇
作者:Michele Cozzatti,Federico Simonetta,Stavros Ntalampiras机构:Ntalampiras[,−,−,−,], LIM – Music Informatics Laboratory, Department of Computer Science, University of Milano备注:Published at ICANN 2022摘要:本文提出了一种基于弱监督机器学习的方法,旨在提供一种工具来提醒患者可能的呼吸系统疾病.各种类型的病理可能影响呼吸系统,潜在地导致严重的疾病,在某些情况下,甚至死亡.一般来说,有效的预防措施被认为是改善病人健康状况的主要因素。该方法利用变分自动编码器结构,允许使用有限复杂度的训练流水线和相对小尺寸的数据集,重要的是,该方法提供了57%的准确率,这与现有的强监督方法是一致的。摘要:This paper proposes a weakly-supervised machine learning-based approach aiming at a tool to alert patients about possible respiratory diseases. Various types of pathologies may affect the respiratory system, potentially leading to severe diseases and, in certain cases, death. In general, effective prevention practices are considered as major actors towards the improvement of the patient's health condition. The proposed method strives to realize an easily accessible tool for the automatic diagnosis of respiratory diseases. Specifically, the method leverages Variational Autoencoder architectures permitting the usage of training pipelines of limited complexity and relatively small-sized datasets. Importantly, it offers an accuracy of 57 %, which is in line with the existing strongly-supervised approaches.
机器翻译,仅供参考