今天跟大家分享一篇语音相关的论文合集:cs.SD语音3篇,eess.AS音频处理4篇。

cs.SD语音
【1】 Avoid Overfitting User Specific Information in Federated Keyword  Spotting
标题:避免联合关键词定位中用户特定信息的过度匹配
链接:https://arxiv.org/abs/2206.08864
作者:Xin-Chun Li,Jin-Lin Tang,Shaoming Song,Bingshuai Li,Yinchuan Li,Yunfeng Shao,Le Gan,De-Chuan Zhan
备注:Accepted by Interspeech 2022
摘要:关键词识别(KWS)旨在为不同的用户准确有效地从其他信号中区分特定的唤醒词。最近的工作利用各种深度网络来训练KWS模型,所有用户的语音数据都集中在一起,而不考虑数据隐私。联邦KWS(Federated KWS)可以作为一种解决方案,而无需直接共享用户数据。然而,少量的数据、不同的用户习惯和不同的口音可能会导致致命的问题,例如过度拟合或体重差异。因此,我们提出了几种策略,以鼓励该模型不要过度拟合FedKWS中的特定于用户的信息。具体而言,我们首先提出了一种对抗性学习策略,该策略针对过度拟合的局部模型更新下载的全局模型,并明确鼓励全局模型捕获用户不变的信息。此外,我们还提出了一种自适应的局部训练策略,让具有更多训练数据和更均匀类分布的客户执行更多的局部更新步骤。同样,这一战略可以削弱数据质量较差的用户的负面影响。我们提出的FedKWS UI可以显式和隐式地学习FedKWS中的用户不变信息。在联邦Google语音命令上的大量实验结果验证了FedKWS UI的有效性。

摘要:Keyword spotting (KWS) aims to discriminate a specific wake-up word from other signals precisely and efficiently for different users. Recent works utilize various deep networks to train KWS models with all users' speech data centralized without considering data privacy. Federated KWS (FedKWS) could serve as a solution without directly sharing users' data. However, the small amount of data, different user habits, and various accents could lead to fatal problems, e.g., overfitting or weight divergence. Hence, we propose several strategies to encourage the model not to overfit user-specific information in FedKWS. Specifically, we first propose an adversarial learning strategy, which updates the downloaded global model against an overfitted local model and explicitly encourages the global model to capture user-invariant information. Furthermore, we propose an adaptive local training strategy, letting clients with more training data and more uniform class distributions undertake more local update steps. Equivalently, this strategy could weaken the negative impacts of those users whose data is less qualified. Our proposed FedKWS-UI could explicitly and implicitly learn user-invariant information in FedKWS. Abundant experimental results on federated Google Speech Commands verify the effectiveness of FedKWS-UI.


【2】 What can Speech and Language Tell us About the Working Alliance in  Psychotherapy

标题:关于心理治疗中的工作联盟,言语和语言能告诉我们什么

链接:https://arxiv.org/abs/2206.08835

作者:Sebastian P. Bayerl,Gabriel Roccabruna,Shammur Absar Chowdhury,Tommaso Ciulli,Morena Danieli,Korbinian Riedhammer,Giuseppe Riccardi

备注:Accepted at Interspeech 2022

摘要:我们对会话分析问题及其在健康领域的应用感兴趣。认知行为疗法是一种结构化的心理治疗方法,允许治疗师帮助患者识别和修改恶意的想法、行为或行为。这种合作努力可以使用工作联盟量表(观察者评分缩短)进行评估,这是一个涵盖任务、目标和关系的12项量表,对治疗结果有相关影响。在这项工作中,我们调查了这种联盟清单与患者和心理治疗师之间的口语对话(会话)之间的关系。我们提供了为期八周的电子治疗,收集了他们的音频和视频通话,并手动转录。专业治疗师对口头对话进行了注释和WAI评分评估。我们调查了语音和语言特征及其与WAI项目的关系。特征类型包括转向动力学、词汇夹带和从语音和语言信号中提取的会话描述符。我们的发现提供了强有力的证据,证明这些特征中的一个子集是工作联盟的有力指标。据我们所知,这是第一个利用言语和语言来描述工作联盟的新颖研究。

摘要:We are interested in the problem of conversational analysis and its application to the health domain. Cognitive Behavioral Therapy is a structured approach in psychotherapy, allowing the therapist to help the patient to identify and modify the malicious thoughts, behavior, or actions. This cooperative effort can be evaluated using the Working Alliance Inventory Observer-rated Shortened - a 12 items inventory covering task, goal, and relationship - which has a relevant influence on therapeutic outcomes. In this work, we investigate the relation between this alliance inventory and the spoken conversations (sessions) between the patient and the psychotherapist. We have delivered eight weeks of e-therapy, collected their audio and video call sessions, and manually transcribed them. The spoken conversations have been annotated and evaluated with WAI ratings by professional therapists. We have investigated speech and language features and their association with WAI items. The feature types include turn dynamics, lexical entrainment, and conversational descriptors extracted from the speech and language signals. Our findings provide strong evidence that a subset of these features are strong indicators of working alliance. To the best of our knowledge, this is the first and a novel study to exploit speech and language for characterising working alliance.


【3】 Simultaneous Speech Extraction for Multiple Target Speakers under the  Meeting Scenarios(V1)

标题:会议场景下多目标说话人的同时语音提取(V1)

链接:https://arxiv.org/abs/2206.08525

作者:Bang Zeng,Weiqing Wang,Yuanyuan Bao,Ming Li
备注:4pages, 3 figures
摘要:近年来,会议场景下的目标语音分离或提取技术已成为研究的热点。我们提出了一种基于说话人二值化的多目标语音分离系统(SD-MTSS),它可以同时从混合语音中提取每个说话人的声音,而不需要像以前的解决方案那样需要一系列独立的过程。SD-MTSS由说话人日记(SD)模块和多目标语音分离(MTSS)模块组成。前者可以推断出混合语音的目标说话人语音活动检测(TSVAD)状态,并得到不同说话人的单说话人音频片段作为参考语音。后者使用混合音频和参考语音作为输入,然后生成估计掩码。通过利用TSVAD决策和估计的掩码,我们的SD-MTSS模型可以在转换记录中同时提取每个说话人的语音,而无需预先添加注册音频。实验结果表明,与WSJ0-2mix-extr数据集上最先进的SpEx+基线相比,我们的MTSS模型在很大程度上优于我们的基线,分别实现了1.38dB SDR、1.34dB SI-SNR和0.13 PESQ的改善。SD-MTSS系统也比Alimeeting数据集的基线有了显著的改进。
摘要:Recently, the target speech separation or extraction techniques under the meeting scenario have become a hot research trend. We propose a speaker diarization aware multiple target speech separation system (SD-MTSS) to simultaneously extract the voice of each speaker from the mixed speech, rather than requiring a succession of independent processes as presented in previous solutions. SD-MTSS consists of a speaker diarization (SD) module and a multiple target speech separation (MTSS) module. The former one infers the target speaker voice activity detection (TSVAD) states of the mixture, as well as gets different speakers' single-talker audio segments as the reference speech. The latter one employs both the mixed audio and reference speech as inputs, and then it generates an estimated mask. By exploiting the TSVAD decision and the estimated mask, our SD-MTSS model can extract the speech of each speaker concurrently in a conversion recording without additional enrollment audio in advance.Experimental results show that our MTSS model outperforms our baselines with a large margin, achieving 1.38dB SDR, 1.34dB SI-SNR, and 0.13 PESQ improvements over the state-of-the-art SpEx+ baseline on the WSJ0-2mix-extr dataset, respectively. The SD-MTSS system makes a significant improvement than the baseline on the Alimeeting dataset as well.

eess.AS音频处理

【1】 NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling  Rates

标题:NU-Wave 2:一种适用于不同采样率的通用神经音频上采样模型

链接:https://arxiv.org/abs/2206.08545

作者:Seungu Han,Junhyeok Lee
备注:Accepted to Interspeech 2022
摘要:传统上,音频超分辨率模型固定初始和目标采样率,这就需要针对每对采样率对模型进行训练。我们介绍了NU Wave 2,这是一种用于神经音频上采样的扩散模型,它可以通过单个模型从不同采样率的输入中生成48 kHz音频信号。基于NU-Wave的体系结构,NU-Wave 2使用短时傅立叶卷积(STFC)产生谐波以解决NU-Wave的主要故障模式,并结合带宽谱特征变换(BSFT)在频域中调节输入的带宽。我们通过实验证明,与其他模型相比,NU Wave 2可以在不考虑输入采样率的情况下产生高分辨率音频,同时需要的参数更少。官方代码和音频样本可在https://mindslab-ai.github.io/nuwave2.
摘要:Conventionally, audio super-resolution models fixed the initial and the target sampling rates, which necessitate the model to be trained for each pair of sampling rates. We introduce NU-Wave 2, a diffusion model for neural audio upsampling that enables the generation of 48 kHz audio signals from inputs of various sampling rates with a single model. Based on the architecture of NU-Wave, NU-Wave 2 uses short-time Fourier convolution (STFC) to generate harmonics to resolve the main failure modes of NU-Wave, and incorporates bandwidth spectral feature transform (BSFT) to condition the bandwidths of inputs in the frequency domain. We experimentally demonstrate that NU-Wave 2 produces high-resolution audio regardless of the sampling rate of input while requiring fewer parameters than other models. The official code and the audio samples are available at https://mindslab-ai.github.io/nuwave2.


【2】 Simultaneous Speech Extraction for Multiple Target Speakers under the  Meeting Scenarios(V1)

标题:会议场景下多目标说话人的同时语音提取(V1)

链接:https://arxiv.org/abs/2206.08525

* 与cs.SD语音【3】为同一篇

作者:Bang Zeng,Weiqing Wang,Yuanyuan Bao,Ming Li
备注:4pages, 3 figures
摘要:近年来,会议场景下的目标语音分离或提取技术已成为研究的热点。我们提出了一种基于说话人二值化的多目标语音分离系统(SD-MTSS),它可以同时从混合语音中提取每个说话人的声音,而不需要像以前的解决方案那样需要一系列独立的过程。SD-MTSS由说话人日记(SD)模块和多目标语音分离(MTSS)模块组成。前者可以推断出混合语音的目标说话人语音活动检测(TSVAD)状态,并得到不同说话人的单说话人音频片段作为参考语音。后者使用混合音频和参考语音作为输入,然后生成估计掩码。通过利用TSVAD决策和估计的掩码,我们的SD-MTSS模型可以在转换记录中同时提取每个说话人的语音,而无需预先添加注册音频。实验结果表明,与WSJ0-2mix-extr数据集上最先进的SpEx+基线相比,我们的MTSS模型在很大程度上优于我们的基线,分别实现了1.38dB SDR、1.34dB SI-SNR和0.13 PESQ的改善。SD-MTSS系统也比Alimeeting数据集的基线有了显著的改进。
摘要:Recently, the target speech separation or extraction techniques under the meeting scenario have become a hot research trend. We propose a speaker diarization aware multiple target speech separation system (SD-MTSS) to simultaneously extract the voice of each speaker from the mixed speech, rather than requiring a succession of independent processes as presented in previous solutions. SD-MTSS consists of a speaker diarization (SD) module and a multiple target speech separation (MTSS) module. The former one infers the target speaker voice activity detection (TSVAD) states of the mixture, as well as gets different speakers' single-talker audio segments as the reference speech. The latter one employs both the mixed audio and reference speech as inputs, and then it generates an estimated mask. By exploiting the TSVAD decision and the estimated mask, our SD-MTSS model can extract the speech of each speaker concurrently in a conversion recording without additional enrollment audio in advance.Experimental results show that our MTSS model outperforms our baselines with a large margin, achieving 1.38dB SDR, 1.34dB SI-SNR, and 0.13 PESQ improvements over the state-of-the-art SpEx+ baseline on the WSJ0-2mix-extr dataset, respectively. The SD-MTSS system makes a significant improvement than the baseline on the Alimeeting dataset as well.


【3】 Avoid Overfitting User Specific Information in Federated Keyword  Spotting

标题:避免联合关键词定位中用户特定信息的过度匹配链接:https://arxiv.org/abs/2206.08864

* 与cs.SD语音【1】为同一篇

作者:Xin-Chun Li,Jin-Lin Tang,Shaoming Song,Bingshuai Li,Yinchuan Li,Yunfeng Shao,Le Gan,De-Chuan Zhan

备注:Accepted by Interspeech 2022

摘要:关键词识别(KWS)旨在为不同的用户准确有效地从其他信号中区分特定的唤醒词。最近的工作利用各种深度网络来训练KWS模型,所有用户的语音数据都集中在一起,而不考虑数据隐私。联邦KWS(Federated KWS)可以作为一种解决方案,而无需直接共享用户数据。然而,少量的数据、不同的用户习惯和不同的口音可能会导致致命的问题,例如过度拟合或体重差异。因此,我们提出了几种策略,以鼓励该模型不要过度拟合FedKWS中的特定于用户的信息。具体而言,我们首先提出了一种对抗性学习策略,该策略针对过度拟合的局部模型更新下载的全局模型,并明确鼓励全局模型捕获用户不变的信息。此外,我们还提出了一种自适应的局部训练策略,让具有更多训练数据和更均匀类分布的客户执行更多的局部更新步骤。同样,这一战略可以削弱数据质量较差的用户的负面影响。我们提出的FedKWS UI可以显式和隐式地学习FedKWS中的用户不变信息。在联邦Google语音命令上的大量实验结果验证了FedKWS UI的有效性。

摘要:Keyword spotting (KWS) aims to discriminate a specific wake-up word from other signals precisely and efficiently for different users. Recent works utilize various deep networks to train KWS models with all users' speech data centralized without considering data privacy. Federated KWS (FedKWS) could serve as a solution without directly sharing users' data. However, the small amount of data, different user habits, and various accents could lead to fatal problems, e.g., overfitting or weight divergence. Hence, we propose several strategies to encourage the model not to overfit user-specific information in FedKWS. Specifically, we first propose an adversarial learning strategy, which updates the downloaded global model against an overfitted local model and explicitly encourages the global model to capture user-invariant information. Furthermore, we propose an adaptive local training strategy, letting clients with more training data and more uniform class distributions undertake more local update steps. Equivalently, this strategy could weaken the negative impacts of those users whose data is less qualified. Our proposed FedKWS-UI could explicitly and implicitly learn user-invariant information in FedKWS. Abundant experimental results on federated Google Speech Commands verify the effectiveness of FedKWS-UI.


【4】 What can Speech and Language Tell us About the Working Alliance in  Psychotherapy

标题:关于心理治疗中的工作联盟,言语和语言能告诉我们什么
链接:https://arxiv.org/abs/2206.08835
* 与cs.SD语音【2】为同一篇

作者:Sebastian P. Bayerl,Gabriel Roccabruna,Shammur Absar Chowdhury,Tommaso Ciulli,Morena Danieli,Korbinian Riedhammer,Giuseppe Riccardi

备注:Accepted at Interspeech 2022

摘要:我们对会话分析问题及其在健康领域的应用感兴趣。认知行为疗法是一种结构化的心理治疗方法,允许治疗师帮助患者识别和修改恶意的想法、行为或行为。这种合作努力可以使用工作联盟量表(观察者评分缩短)进行评估,这是一个涵盖任务、目标和关系的12项量表,对治疗结果有相关影响。在这项工作中,我们调查了这种联盟清单与患者和心理治疗师之间的口语对话(会话)之间的关系。我们提供了为期八周的电子治疗,收集了他们的音频和视频通话,并手动转录。专业治疗师对口头对话进行了注释和WAI评分评估。我们调查了语音和语言特征及其与WAI项目的关系。特征类型包括转向动力学、词汇夹带和从语音和语言信号中提取的会话描述符。我们的发现提供了强有力的证据,证明这些特征中的一个子集是工作联盟的有力指标。据我们所知,这是第一个利用言语和语言来描述工作联盟的新颖研究。

摘要:We are interested in the problem of conversational analysis and its application to the health domain. Cognitive Behavioral Therapy is a structured approach in psychotherapy, allowing the therapist to help the patient to identify and modify the malicious thoughts, behavior, or actions. This cooperative effort can be evaluated using the Working Alliance Inventory Observer-rated Shortened - a 12 items inventory covering task, goal, and relationship - which has a relevant influence on therapeutic outcomes. In this work, we investigate the relation between this alliance inventory and the spoken conversations (sessions) between the patient and the psychotherapist. We have delivered eight weeks of e-therapy, collected their audio and video call sessions, and manually transcribed them. The spoken conversations have been annotated and evaluated with WAI ratings by professional therapists. We have investigated speech and language features and their association with WAI items. The feature types include turn dynamics, lexical entrainment, and conversational descriptors extracted from the speech and language signals. Our findings provide strong evidence that a subset of these features are strong indicators of working alliance. To the best of our knowledge, this is the first and a novel study to exploit speech and language for characterising working alliance.