今日论文合集:cs.SD语音4篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Springboard, Roadblock or "Crutch"?: How Transgender Users Leverage  Voice Changers for Gender Presentation in Social Virtual Reality

链接:https://arxiv.org/abs/2402.08217

作者:Kassie Povinelli,Yuhang Zhao

备注:None

摘要:社交虚拟现实(VR)是跨性别者通过化身探索自己身份的重要平台,并在在线社区中培养个人联系。然而,它提出了一个挑战:化身体现和语音表示之间的脱节,往往导致性别歧视和骚扰。之前的研究承认了这个问题,但忽略了变声器的潜在解决方案。我们采访了13名跨性别和性别差异的社交VR平台用户,重点关注他们使用和不使用变声器的体验。我们发现,使用变声器不仅减少了与声音有关的骚扰,而且还可以让他们通过听到自己修改后的声音和其他人对他们修改后的声音的反应来体验性别欣快感,激励他们进行声音训练和药物治疗以实现所需的声音。此外,我们还确定了当前变声器技术的技术障碍,以及缓解跨性别和性别歧视用户面临的问题的潜在改进。

摘要:Social virtual reality (VR) serves as a vital platform for transgender individuals to explore their identities through avatars and foster personal connections within online communities. However, it presents a challenge: the disconnect between avatar embodiment and voice representation, often leading to misgendering and harassment. Prior research acknowledges this issue but overlooks the potential solution of voice changers. We interviewed 13 transgender and gender-nonconforming users of social VR platforms, focusing on their experiences with and without voice changers. We found that using a voice changer not only reduces voice-related harassment, but also allows them to experience gender euphoria through both hearing their modified voice and the reactions of others to their modified voice, motivating them to pursue voice training and medication to achieve desired voices. Furthermore, we identified the technical barriers to current voice changer technology and potential improvements to alleviate the problems that transgender and gender-nonconforming users face.


【2】 Benchmarking multi-component signal processing methods in the  time-frequency plane
标题:时频平面上多分量信号处理方法的基准测试
链接:https://arxiv.org/abs/2402.08521
作者:Juan M. Miramont,Rémi Bardenet,Pierre Chainais,Francois Auger
摘要:时频平面上的信号处理有着悠久的历史,并且仍然是方法创新的领域。例如,自2015年以来,已经提出了基于频谱图的零值的检测和去噪,与专注于频谱图的较大值的长期历史形成对比。然而,与优化和机器学习等邻近领域不同,时频信号处理缺乏广泛采用的基准测试工具。在这项工作中,我们贡献了一个开源的,基于Python的工具箱称为MCSM-Benchs基准多分量信号分析方法,我们展示了我们的工具箱上的三个时间频率基准。首先,我们比较了不同的方法,信号检测的基础上的频谱图的零点,包括未开发的变化,以前提出的检测测试。其次,我们比较了基于零的去噪方法的经典和新的方法的基础上的大值和脊的频谱图。最后,我们比较了这些方法对典型的频谱图阈值策略的去噪性能,在后处理文物通常被称为音乐噪声。在低水平上,所获得的结果提供了新的见解的评估方法,特别是研究方向,以进一步发展零基方法。在更高的层面上,我们的基准测试体现了使用公共、协作、通用框架进行基准测试的好处。
摘要:Signal processing in the time-frequency plane has a long history and remains a field of methodological innovation. For instance, detection and denoising based on the zeros of the spectrogram have been proposed since 2015, contrasting with a long history of focusing on larger values of the spectrogram. Yet, unlike neighboring fields like optimization and machine learning, time-frequency signal processing lacks widely-adopted benchmarking tools. In this work, we contribute an open-source, Python-based toolbox termed MCSM-Benchs for benchmarking multi-component signal analysis methods, and we demonstrate our toolbox on three time-frequency benchmarks. First, we compare different methods for signal detection based on the zeros of the spectrogram, including unexplored variations of previously proposed detection tests. Second, we compare zero-based denoising methods to both classical and novel methods based on large values and ridges of the spectrogram. Finally, we compare the denoising performance of these methods against typical spectrogram thresholding strategies, in terms of post-processing artifacts commonly referred to as musical noise. At a low level, the obtained results provide new insight on the assessed approaches, and in particular research directions to further develop zero-based methods. At a higher level, our benchmarks exemplify the benefits of using a public, collaborative, common framework for benchmarking.

【3】 Channel-Combination Algorithms for Robust Distant Voice Activity and  Overlapped Speech Detection
标题:稳健远端语音活动和重叠语音检测的信道组合算法
链接:https://arxiv.org/abs/2402.08312
作者:Théo Mariotte,Anthony Larcher,Silvio Montrésor,Jean-Hugh Thomas
备注:14 pages, 5 figures, accepted at IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)
摘要:语音激活检测(VAD)和重叠语音检测(OSD)是说话人日志化的关键预处理任务。在会议上下文中,使用远程设备捕获语音通常更容易。然而,这种考虑导致严重的性能下降。研究了一种统一的有监督学习框架来解决远距离多麦克风联合VAD和OSD(VAD+OSD)问题。本文研究了各种多通道VAD+OSD前端,加权和合并传入通道。我们提出了三种算法的基础上的自注意力通道组合器(SACC),以前提出的文献。在AMI会议语料库上进行的实验表明,通道组合方法在远距离语音场景中带来了显着的VAD+OSD改善。具体来说,我们探讨了使用学习复杂的组合权重,并证明了这种方法在可解释性方面的好处。通道组合为基础的VAD+OSD系统进行评估的最后后端任务,即扬声器日记,并显示出显着的改善。最后,由于多通道系统是在给定固定阵列配置的情况下训练的,因此它们可能无法推广到其他阵列设置,例如麦克风数量不匹配。一个通道数不变的损失,提出了学习一个独特的特征表示,无论可用的麦克风的数量。对不匹配的阵列配置进行的评估突出了这种训练策略的鲁棒性。
摘要:Voice Activity Detection (VAD) and Overlapped Speech Detection (OSD) are key pre-processing tasks for speaker diarization. In the meeting context, it is often easier to capture speech with a distant device. This consideration however leads to severe performance degradation. We study a unified supervised learning framework to solve distant multi-microphone joint VAD and OSD (VAD+OSD). This paper investigates various multi-channel VAD+OSD front-ends that weight and combine incoming channels. We propose three algorithms based on the Self-Attention Channel Combinator (SACC), previously proposed in the literature. Experiments conducted on the AMI meeting corpus exhibit that channel combination approaches bring significant VAD+OSD improvements in the distant speech scenario. Specifically, we explore the use of learned complex combination weights and demonstrate the benefits of such an approach in terms of explainability. Channel combination-based VAD+OSD systems are evaluated on the final back-end task, i.e. speaker diarization, and show significant improvements. Finally, since multi-channel systems are trained given a fixed array configuration, they may fail in generalizing to other array set-ups, e.g. mismatched number of microphones. A channel-number invariant loss is proposed to learn a unique feature representation regardless of the number of available microphones. The evaluation conducted on mismatched array configurations highlights the robustness of this training strategy.


【4】 Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement  with Conformer-based Metric GAN
标题:基于一致性度量GaN的无限制全局相位偏差感知单声道语音增强
链接:https://arxiv.org/abs/2402.08252
作者:Shiqi Zhang,Zheng Qiu,Daiki Takeuchi,Noboru Harada,Shoji Makino
备注:Accepted by ICASSP 2024
摘要:近年来,随着神经网络的迅速发展,各种网络在单通道语音增强领域对含噪语音幅度谱的增强能力变得异常突出。然而,使用神经网络增强相位谱往往是无效的,这仍然是一个具有挑战性的问题。在本文中,我们发现,人耳不能敏感地感知精确的相位谱和偏置相位(BP)谱之间的差异。因此,我们提出了一种相位重建的优化方法,允许自由的整体相位偏差,而不是重建精确的相位谱。我们将其应用于基于一致性的度量生成对抗网络(CMGAN)基线模型,该模型放松了现有精确相位的限制,并为神经网络提供了更广阔的学习空间。结果表明,该方法实现了一个新的国家的最先进的性能,而不会产生额外的计算开销。
摘要:With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often ineffective, which remains a challenging problem. In this paper, we found that the human ear cannot sensitively perceive the difference between a precise phase spectrum and a biased phase (BP) spectrum. Therefore, we propose an optimization method of phase reconstruction, allowing freedom on the global-phase bias instead of reconstructing the precise phase spectrum. We applied it to a Conformer-based Metric Generative Adversarial Networks (CMGAN) baseline model, which relaxes the existing constraints of precise phase and gives the neural network a broader learning space. Results show that this method achieves a new state-of-the-art performance without incurring additional computational overhead.

eess.AS音频处理
【1】 Benchmarking multi-component signal processing methods in the  time-frequency plane
标题:时频平面基准多分量信号处理方法
链接:https://arxiv.org/abs/2402.08521
作者:Juan M. Miramont,Rémi Bardenet,Pierre Chainais,Francois Auger
摘要:时频平面上的信号处理有着悠久的历史,并且仍然是方法创新的领域。例如,自2015年以来,已经提出了基于频谱图的零值的检测和去噪,与专注于频谱图的较大值的长期历史形成对比。然而,与优化和机器学习等邻近领域不同,时频信号处理缺乏广泛采用的基准测试工具。在这项工作中,我们贡献了一个开源的,基于Python的工具箱称为MCSM-Benchs基准多分量信号分析方法,我们展示了我们的工具箱上的三个时间频率基准。首先,我们比较了不同的方法,信号检测的基础上的频谱图的零点,包括未开发的变化,以前提出的检测测试。其次,我们比较了基于零的去噪方法的经典和新的方法的基础上的大值和脊的频谱图。最后,我们比较了这些方法对典型的频谱图阈值策略的去噪性能,在后处理文物通常被称为音乐噪声。在低水平上,所获得的结果提供了新的见解的评估方法,特别是研究方向,以进一步发展零基方法。在更高的层面上,我们的基准测试体现了使用公共、协作、通用框架进行基准测试的好处。
摘要:Signal processing in the time-frequency plane has a long history and remains a field of methodological innovation. For instance, detection and denoising based on the zeros of the spectrogram have been proposed since 2015, contrasting with a long history of focusing on larger values of the spectrogram. Yet, unlike neighboring fields like optimization and machine learning, time-frequency signal processing lacks widely-adopted benchmarking tools. In this work, we contribute an open-source, Python-based toolbox termed MCSM-Benchs for benchmarking multi-component signal analysis methods, and we demonstrate our toolbox on three time-frequency benchmarks. First, we compare different methods for signal detection based on the zeros of the spectrogram, including unexplored variations of previously proposed detection tests. Second, we compare zero-based denoising methods to both classical and novel methods based on large values and ridges of the spectrogram. Finally, we compare the denoising performance of these methods against typical spectrogram thresholding strategies, in terms of post-processing artifacts commonly referred to as musical noise. At a low level, the obtained results provide new insight on the assessed approaches, and in particular research directions to further develop zero-based methods. At a higher level, our benchmarks exemplify the benefits of using a public, collaborative, common framework for benchmarking.


【2】 Channel-Combination Algorithms for Robust Distant Voice Activity and  Overlapped Speech Detection
标题:稳健远端语音活动和重叠语音检测的信道组合算法
链接:https://arxiv.org/abs/2402.08312
作者:Théo Mariotte,Anthony Larcher,Silvio Montrésor,Jean-Hugh Thomas
备注:14 pages, 5 figures, accepted at IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)
摘要:语音激活检测(VAD)和重叠语音检测(OSD)是说话人日志化的关键预处理任务。在会议上下文中,使用远程设备捕获语音通常更容易。然而,这种考虑导致严重的性能下降。研究了一种统一的有监督学习框架来解决远距离多麦克风联合VAD和OSD(VAD+OSD)问题。本文研究了各种多通道VAD+OSD前端,加权和合并传入通道。我们提出了三种算法的基础上的自注意力通道组合器(SACC),以前提出的文献。在AMI会议语料库上进行的实验表明,通道组合方法在远距离语音场景中带来了显着的VAD+OSD改善。具体来说,我们探讨了使用学习复杂的组合权重,并证明了这种方法在可解释性方面的好处。通道组合为基础的VAD+OSD系统进行评估的最后后端任务,即扬声器日记,并显示出显着的改善。最后,由于多通道系统是在给定固定阵列配置的情况下训练的,因此它们可能无法推广到其他阵列设置,例如麦克风数量不匹配。一个通道数不变的损失,提出了学习一个独特的特征表示,无论可用的麦克风的数量。对不匹配的阵列配置进行的评估突出了这种训练策略的鲁棒性。
摘要:Voice Activity Detection (VAD) and Overlapped Speech Detection (OSD) are key pre-processing tasks for speaker diarization. In the meeting context, it is often easier to capture speech with a distant device. This consideration however leads to severe performance degradation. We study a unified supervised learning framework to solve distant multi-microphone joint VAD and OSD (VAD+OSD). This paper investigates various multi-channel VAD+OSD front-ends that weight and combine incoming channels. We propose three algorithms based on the Self-Attention Channel Combinator (SACC), previously proposed in the literature. Experiments conducted on the AMI meeting corpus exhibit that channel combination approaches bring significant VAD+OSD improvements in the distant speech scenario. Specifically, we explore the use of learned complex combination weights and demonstrate the benefits of such an approach in terms of explainability. Channel combination-based VAD+OSD systems are evaluated on the final back-end task, i.e. speaker diarization, and show significant improvements. Finally, since multi-channel systems are trained given a fixed array configuration, they may fail in generalizing to other array set-ups, e.g. mismatched number of microphones. A channel-number invariant loss is proposed to learn a unique feature representation regardless of the number of available microphones. The evaluation conducted on mismatched array configurations highlights the robustness of this training strategy.

【3】 Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement  with Conformer-based Metric GAN
标题:基于一致性度量GaN的无限制全局相位偏差感知单声道语音增强
链接:https://arxiv.org/abs/2402.08252
作者:Shiqi Zhang,Zheng Qiu,Daiki Takeuchi,Noboru Harada,Shoji Makino
备注:Accepted by ICASSP 2024
摘要:近年来,随着神经网络的迅速发展,各种网络在单通道语音增强领域对含噪语音幅度谱的增强能力变得异常突出。然而,使用神经网络增强相位谱往往是无效的,这仍然是一个具有挑战性的问题。在本文中,我们发现,人耳不能敏感地感知精确的相位谱和偏置相位(BP)谱之间的差异。因此,我们提出了一种相位重建的优化方法,允许自由的整体相位偏差,而不是重建精确的相位谱。我们将其应用于基于一致性的度量生成对抗网络(CMGAN)基线模型,该模型放松了现有精确相位的限制,并为神经网络提供了更广阔的学习空间。结果表明,该方法实现了一个新的国家的最先进的性能,而不会产生额外的计算开销。
摘要:With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often ineffective, which remains a challenging problem. In this paper, we found that the human ear cannot sensitively perceive the difference between a precise phase spectrum and a biased phase (BP) spectrum. Therefore, we propose an optimization method of phase reconstruction, allowing freedom on the global-phase bias instead of reconstructing the precise phase spectrum. We applied it to a Conformer-based Metric Generative Adversarial Networks (CMGAN) baseline model, which relaxes the existing constraints of precise phase and gives the neural network a broader learning space. Results show that this method achieves a new state-of-the-art performance without incurring additional computational overhead.


【4】 Springboard, Roadblock or "Crutch"?: How Transgender Users Leverage  Voice Changers for Gender Presentation in Social Virtual Reality
链接:https://arxiv.org/abs/2402.08217
作者:Kassie Povinelli,Yuhang Zhao
备注:None
摘要:社交虚拟现实(VR)是跨性别者通过化身探索自己身份的重要平台,并在在线社区中培养个人联系。然而,它提出了一个挑战:化身体现和语音表示之间的脱节,往往导致性别歧视和骚扰。之前的研究承认了这个问题,但忽略了变声器的潜在解决方案。我们采访了13名跨性别和性别差异的社交VR平台用户,重点关注他们使用和不使用变声器的体验。我们发现,使用变声器不仅减少了与声音有关的骚扰,而且还可以让他们通过听到自己修改后的声音和其他人对他们修改后的声音的反应来体验性别欣快感,激励他们进行声音训练和药物治疗以实现所需的声音。此外,我们还确定了当前变声器技术的技术障碍,以及缓解跨性别和性别歧视用户面临的问题的潜在改进。
摘要:Social virtual reality (VR) serves as a vital platform for transgender individuals to explore their identities through avatars and foster personal connections within online communities. However, it presents a challenge: the disconnect between avatar embodiment and voice representation, often leading to misgendering and harassment. Prior research acknowledges this issue but overlooks the potential solution of voice changers. We interviewed 13 transgender and gender-nonconforming users of social VR platforms, focusing on their experiences with and without voice changers. We found that using a voice changer not only reduces voice-related harassment, but also allows them to experience gender euphoria through both hearing their modified voice and the reactions of others to their modified voice, motivating them to pursue voice training and medication to achieve desired voices. Furthermore, we identified the technical barriers to current voice changer technology and potential improvements to alleviate the problems that transgender and gender-nonconforming users face.


【5】 BASE TTS: Lessons from building a billion-parameter Text-to-Speech model  on 100K hours of data
标题:基本TTS:基于100K小时数据构建10亿参数文本到语音模型的经验教训
链接:https://arxiv.org/abs/2402.08093
作者:Mateusz Łajszczak,Guillermo Cámbara,Yang Li,Fatih Beyhan,Arent van Korlaar,Fan Yang,Arnaud Joly,Álvaro Martín-Cortinas,Ammar Abbas,Adam Michalski,Alexis Moinet,Sri Karlapati,Ewa Muszyńska,Haohan Guo,Bartosz Putrycz,Soledad López Gambino,Kayeon Yoo,Elena Sokolova,Thomas Drugman
备注:v1
摘要:我们引入了一个文本到语音(TTS)模型,称为BASE TTS,它代表$\textbf{B}$ig $\textbf{A}$自适应$\textbf{S}$可处理的TTS与$\textbf{E}$mergent能力。BASE TTS是迄今为止最大的TTS模型,基于10万小时的公共领域语音数据进行训练,在语音自然度方面达到了新的水平。它部署了一个10亿参数的自回归Transformer,将原始文本转换为离散代码(“语音代码”),然后是一个基于卷积的解码器,将这些语音代码以增量,流式方式转换为波形。此外,我们的语音代码是使用一种新的语音标记化技术,其特点是扬声器ID解开和压缩字节对编码。与广泛报道的大型语言模型在接受越来越多的数据训练时的“涌现能力”相呼应,我们发现,使用10 K+小时和500 M+参数构建的BASE TTS变体开始在文本复杂的句子上表现出自然的韵律。我们设计并分享了一个专门的数据集来衡量这些文本到语音的新兴能力。我们通过对包括公开的大规模文本到语音系统(YourTTS,Bark和TortoiseTTS)在内的基线进行评估,展示了BASE TTS最先进的自然性。该模型生成的音频样本可以在https://amazon-ltts-paper.com/上听到。
摘要:We introduce a text-to-speech (TTS) model called BASE TTS, which stands for $\textbf{B}$ig $\textbf{A}$daptive $\textbf{S}$treamable TTS with $\textbf{E}$mergent abilities. BASE TTS is the largest TTS model to-date, trained on 100K hours of public domain speech data, achieving a new state-of-the-art in speech naturalness. It deploys a 1-billion-parameter autoregressive Transformer that converts raw texts into discrete codes ("speechcodes") followed by a convolution-based decoder which converts these speechcodes into waveforms in an incremental, streamable manner. Further, our speechcodes are built using a novel speech tokenization technique that features speaker ID disentanglement and compression with byte-pair encoding. Echoing the widely-reported "emergent abilities" of large language models when trained on increasing volume of data, we show that BASE TTS variants built with 10K+ hours and 500M+ parameters begin to demonstrate natural prosody on textually complex sentences. We design and share a specialized dataset to measure these emergent abilities for text-to-speech. We showcase state-of-the-art naturalness of BASE TTS by evaluating against baselines that include publicly available large-scale text-to-speech systems: YourTTS, Bark and TortoiseTTS. Audio samples generated by the model can be heard at https://amazon-ltts-paper.com/.

机器翻译由腾讯交互翻译提供,仅供参考