今天跟大家分享一篇语音相关的论文合集:cs.SD语音4篇,eess.AS音频处理3篇。本文经arXiv每日学术速递授权转载
【1】 Speech Artifact Removal from EEG Recordings of Spoken Word Production with Tensor Decomposition
标题:基于张量分解的语音产生脑电记录中的语音伪迹去除作者:Holy Lovenia,Hiroki Tanaka,Sakriani Sakti,Ayu Purwarianti,Satoshi Nakamura机构:Nara Institute of Science and Technology, Japan, RIKEN, Center for Advanced Intelligence Project AIP, Japan, Department of Informatics, Bandung Institute of Technology, Indonesia摘要:由于语音伪影的未被发现的特征,涉及到口语词产生的大脑活动的研究相当不发达,语音伪影污染了脑电图(EEG)信号,阻碍了对潜在认知过程的检查。为了进一步利用语音产生进行脑电研究,提出了一种利用三模式张量分解(时间x空间x频率)进行语音伪影去除的方法。张量分解能够同时检测多个模式,这符合EEG数据的多向性。在一个图片命名任务中,我们通过在嘴巴附近放置两个电极来记录嘴唇肌电图,从而收集带有语音伪影的原始数据。根据我们的评估,计算了总平均语音伪影和唇肌电之间的相关值,张量分解在检测语音伪影(0.985)和生成干净数据(0.101)方面优于以前基于独立分量分析(ICA)和盲源分离(BSS)的方法。我们提出的方法正确地保留了与语音无关的成分,并通过计算无EOG的总平均原始数据与语音开始前清理的数据之间的相关值(0.92-0.94)进行了验证。摘要:Research about brain activities involving spoken word production is considerably underdeveloped because of the undiscovered characteristics of speech artifacts, which contaminate electroencephalogram (EEG) signals and prevent the inspection of the underlying cognitive processes. To fuel further EEG research with speech production, a method using three-mode tensor decomposition (time x space x frequency) is proposed to perform speech artifact removal. Tensor decomposition enables simultaneous inspection of multiple modes, which suits the multi-way nature of EEG data. In a picture-naming task, we collected raw data with speech artifacts by placing two electrodes near the mouth to record lip EMG. Based on our evaluation, which calculated the correlation values between grand-averaged speech artifacts and the lip EMG, tensor decomposition outperformed the former methods that were based on independent component analysis (ICA) and blind source separation (BSS), both in detecting speech artifact (0.985) and producing clean data (0.101). Our proposed method correctly preserved the components unrelated to speech, which was validated by computing the correlation value between the grand-averaged raw data without EOG and cleaned data before the speech onset (0.92-0.94).
【2】 Towards Context-Aware Neural Performance-Score Synchronisation
标题:走向情境感知的神经表现--分数同步
链接:https://arxiv.org/abs/2206.00454
机构:Supervisors:, Professor Simon Dixon, Dr Daniel Wolff, Independent Assessor:, Dr Emmanouil Benetos备注:PhD Thesis, Queen Mary University of London (190 pages)摘要:音乐可以以多种形式表示,例如以音频形式表示为表演记录,以符号形式表示为计算机可读乐谱,或以图像形式表示为乐谱扫描。音乐同步通过在多个音乐表示之间生成精确的映射,提供了一种以统一方式在多个音乐表示之间导航的方法,使其适用于音乐教育、性能分析、自动伴奏和音乐编辑等众多领域。传统的同步方法使用知识驱动和随机方法计算对齐,通常使用手工制作的功能。这些方法通常无法很好地推广到不同的乐器、声环境和记录条件,通常假设性能和分数之间存在完全的结构一致性。该博士通过在三个方面提出数据驱动、上下文感知的对齐方法,进一步推动了性能分数同步研究的发展:首先,我通过采用基于度量学习的方法来替代手工制作的功能,该方法适用于不同的声学设置,并在数据匮乏的情况下表现良好。其次,我讨论了绩效和分数之间结构性差异的处理,这是标准对齐方法的一个常见限制。最后,我避免了对特征工程和动态规划的依赖,提出了一种完全数据驱动的同步方法,该方法使用神经框架计算对齐,同时对性能和分数之间的结构差异也具有鲁棒性。摘要:Music can be represented in multiple forms, such as in the audio form as a recording of a performance, in the symbolic form as a computer readable score, or in the image form as a scan of the sheet music. Music synchronisation provides a way to navigate among multiple representations of music in a unified manner by generating an accurate mapping between them, lending itself applicable to a myriad of domains like music education, performance analysis, automatic accompaniment and music editing. Traditional synchronisation methods compute alignment using knowledge-driven and stochastic approaches, typically employing handcrafted features. These methods are often unable to generalise well to different instruments, acoustic environments and recording conditions, and normally assume complete structural agreement between the performances and the scores. This PhD furthers the development of performance-score synchronisation research by proposing data-driven, context-aware alignment approaches, on three fronts: Firstly, I replace the handcrafted features by employing a metric learning based approach that is adaptable to different acoustic settings and performs well in data-scarce conditions. Secondly, I address the handling of structural differences between the performances and scores, which is a common limitation of standard alignment methods. Finally, I eschew the reliance on both feature engineering and dynamic programming, and propose a completely data-driven synchronisation method that computes alignments using a neural framework, whilst also being robust to structural differences between the performances and scores.
【3】 Towards Generalisable Audio Representations for Audio-Visual Navigation
标题:面向视听导航的通用音频表示法
链接:https://arxiv.org/abs/2206.00393
作者:Shunqi Mao,Chaoyi Zhang,Heng Wang,Weidong Cai机构:School of Computer Science, University of Sydney, Australia备注:CVPR 2022 Embodied AI Workshop摘要:在视听导航(AVN)中,智能代理需要根据其视听感知,在复杂的3D环境中导航到不断发声的对象。虽然现有的方法试图通过精确设计的路径规划或复杂的任务设置来提高导航性能,但没有一种方法能够在任务设置不变的情况下改进对闻所未闻声音的模型泛化。因此,我们提出了一种基于对比学习的方法,通过规范音频编码器来应对这一挑战,其中可以从不同类别的各种音频信号中学习声音不可知的目标驱动的潜在表征。此外,我们考虑了两种数据增强策略来丰富训练声音。我们证明,我们的设计可以轻松地装配到现有AVN框架中,以立即获得性能增益(副本上的SPL为13.4%$\uparrow$,MP3D上的SPL为12.2%$\uparrow$)。我们的项目位于https://AV-GeN.github.io/.摘要:In audio-visual navigation (AVN), an intelligent agent needs to navigate to a constantly sound-making object in complex 3D environments based on its audio and visual perceptions. While existing methods attempt to improve the navigation performance with preciously designed path planning or intricate task settings, none has improved the model generalisation on unheard sounds with task settings unchanged. We thus propose a contrastive learning-based method to tackle this challenge by regularising the audio encoder, where the sound-agnostic goal-driven latent representations can be learnt from various audio signals of different classes. In addition, we consider two data augmentation strategies to enrich the training sounds. We demonstrate that our designs can be easily equipped to existing AVN frameworks to obtain an immediate performance gain (13.4%$\uparrow$ in SPL on Replica and 12.2%$\uparrow$ in SPL on MP3D). Our project is available at https://AV-GeN.github.io/.
【4】 AdaVITS: Tiny VITS for Low Computing Resource Speaker Adaptation
标题:AdaVITS:用于低计算资源说话人适配的微型VITS
链接:https://arxiv.org/abs/2206.00208
作者:Kun Song,Heyang Xue,Xinsheng Wang,Jian Cong,Yongmao Zhang,Lei Xie,Bing Yang,Xiong Zhang,Dan Su机构:Audio, Speech and Language Processing Group,School of Computer Science,School of Software, Northwestern Polytechnical University, Xi’an, China, Cloud and Smart Industries Group, Tencent Technology Co., Ltd., China摘要:文本到语音合成(TTS)中的说话人自适应是对预先训练好的TTS模型进行微调,以适应数据有限的新目标说话人。虽然已经对此项任务进行了大量的工作,但由于轻量级模型的要求和较低的计算复杂性所带来的挑战,很少针对低计算资源场景进行工作。本文提出了一种基于微型VITS的TTS模型AdaVITS,用于低计算资源的说话人自适应。为了有效降低VITS的参数和计算复杂度,提出了一种基于iSTFT的波形构造解码器,以取代原VITS中资源消耗较大的基于上采样的解码器。此外,还引入了NanoFlow来共享流块之间的密度估计,以减少先前编码器的参数。此外,为了降低文本编码器的计算复杂度,将缩放点注意替换为线性注意。为了解决简化模型带来的不稳定性,不再使用原始的文本编码器,而是通过文本到PPG模块将语音后验概率(PPG)用作语言特征,然后将其用作编码器的输入。实验表明,AdaVITS模型参数为8.97M,计算复杂度为0.72GFlops,能够在说话人自适应中生成稳定、自然的语音。摘要:Speaker adaptation in text-to-speech synthesis (TTS) is to finetune a pre-trained TTS model to adapt to new target speakers with limited data. While much effort has been conducted towards this task, seldom work has been performed for low computational resource scenarios due to the challenges raised by the requirement of the lightweight model and less computational complexity. In this paper, a tiny VITS-based TTS model, named AdaVITS, for low computing resource speaker adaptation is proposed. To effectively reduce parameters and computational complexity of VITS, an iSTFT-based wave construction decoder is proposed to replace the upsampling-based decoder which is resource-consuming in the original VITS. Besides, NanoFlow is introduced to share the density estimate across flow blocks to reduce the parameters of the prior encoder. Furthermore, to reduce the computational complexity of the textual encoder, scaled-dot attention is replaced with linear attention. To deal with the instability caused by the simplified model, instead of using the original text encoder, phonetic posteriorgram (PPG) is utilized as linguistic feature via a text-to-PPG module, which is then used as input for the encoder. Experiment shows that AdaVITS can generate stable and natural speech in speaker adaptation with 8.97M model parameters and 0.72GFlops computational complexity.
【1】 Speech Artifact Removal from EEG Recordings of Spoken Word Production with Tensor Decomposition标题:基于张量分解的语音产生脑电记录中的语音伪迹去除作者:Holy Lovenia,Hiroki Tanaka,Sakriani Sakti,Ayu Purwarianti,Satoshi Nakamura机构:Nara Institute of Science and Technology, Japan, RIKEN, Center for Advanced Intelligence Project AIP, Japan, Department of Informatics, Bandung Institute of Technology, Indonesia摘要:由于语音伪影的未被发现的特征,涉及到口语词产生的大脑活动的研究相当不发达,语音伪影污染了脑电图(EEG)信号,阻碍了对潜在认知过程的检查。为了进一步利用语音产生进行脑电研究,提出了一种利用三模式张量分解(时间x空间x频率)进行语音伪影去除的方法。张量分解能够同时检测多个模式,这符合EEG数据的多向性。在一个图片命名任务中,我们通过在嘴巴附近放置两个电极来记录嘴唇肌电图,从而收集带有语音伪影的原始数据。根据我们的评估,计算了总平均语音伪影和唇肌电之间的相关值,张量分解在检测语音伪影(0.985)和生成干净数据(0.101)方面优于以前基于独立分量分析(ICA)和盲源分离(BSS)的方法。我们提出的方法正确地保留了与语音无关的成分,并通过计算无EOG的总平均原始数据与语音开始前清理的数据之间的相关值(0.92-0.94)进行了验证。摘要:Research about brain activities involving spoken word production is considerably underdeveloped because of the undiscovered characteristics of speech artifacts, which contaminate electroencephalogram (EEG) signals and prevent the inspection of the underlying cognitive processes. To fuel further EEG research with speech production, a method using three-mode tensor decomposition (time x space x frequency) is proposed to perform speech artifact removal. Tensor decomposition enables simultaneous inspection of multiple modes, which suits the multi-way nature of EEG data. In a picture-naming task, we collected raw data with speech artifacts by placing two electrodes near the mouth to record lip EMG. Based on our evaluation, which calculated the correlation values between grand-averaged speech artifacts and the lip EMG, tensor decomposition outperformed the former methods that were based on independent component analysis (ICA) and blind source separation (BSS), both in detecting speech artifact (0.985) and producing clean data (0.101). Our proposed method correctly preserved the components unrelated to speech, which was validated by computing the correlation value between the grand-averaged raw data without EOG and cleaned data before the speech onset (0.92-0.94).
【2】 Towards Generalisable Audio Representations for Audio-Visual Navigation
标题:面向视听导航的通用音频表示法
链接:https://arxiv.org/abs/2206.00393
作者:Shunqi Mao,Chaoyi Zhang,Heng Wang,Weidong Cai机构:School of Computer Science, University of Sydney, Australia备注:CVPR 2022 Embodied AI Workshop摘要:在视听导航(AVN)中,智能代理需要根据其视听感知,在复杂的3D环境中导航到不断发声的对象。虽然现有的方法试图通过精确设计的路径规划或复杂的任务设置来提高导航性能,但没有一种方法能够在任务设置不变的情况下改进对闻所未闻声音的模型泛化。因此,我们提出了一种基于对比学习的方法,通过规范音频编码器来应对这一挑战,其中可以从不同类别的各种音频信号中学习声音不可知的目标驱动的潜在表征。此外,我们考虑了两种数据增强策略来丰富训练声音。我们证明,我们的设计可以轻松地装配到现有AVN框架中,以立即获得性能增益(副本上的SPL为13.4%$\uparrow$,MP3D上的SPL为12.2%$\uparrow$)。我们的项目位于https://AV-GeN.github.io/.摘要:In audio-visual navigation (AVN), an intelligent agent needs to navigate to a constantly sound-making object in complex 3D environments based on its audio and visual perceptions. While existing methods attempt to improve the navigation performance with preciously designed path planning or intricate task settings, none has improved the model generalisation on unheard sounds with task settings unchanged. We thus propose a contrastive learning-based method to tackle this challenge by regularising the audio encoder, where the sound-agnostic goal-driven latent representations can be learnt from various audio signals of different classes. In addition, we consider two data augmentation strategies to enrich the training sounds. We demonstrate that our designs can be easily equipped to existing AVN frameworks to obtain an immediate performance gain (13.4%$\uparrow$ in SPL on Replica and 12.2%$\uparrow$ in SPL on MP3D). Our project is available at https://AV-GeN.github.io/.
【3】 AdaVITS: Tiny VITS for Low Computing Resource Speaker Adaptation
标题:AdaVITS:用于低计算资源说话人适配的微型VITS
链接:https://arxiv.org/abs/2206.00208
作者:Kun Song,Heyang Xue,Xinsheng Wang,Jian Cong,Yongmao Zhang,Lei Xie,Bing Yang,Xiong Zhang,Dan Su机构:Audio, Speech and Language Processing Group,School of Computer Science,School of Software, Northwestern Polytechnical University, Xi’an, China, Cloud and Smart Industries Group, Tencent Technology Co., Ltd., China摘要:文本到语音合成(TTS)中的说话人自适应是对预先训练好的TTS模型进行微调,以适应数据有限的新目标说话人。虽然已经对此项任务进行了大量的工作,但由于轻量级模型的要求和较低的计算复杂性所带来的挑战,很少针对低计算资源场景进行工作。本文提出了一种基于微型VITS的TTS模型AdaVITS,用于低计算资源的说话人自适应。为了有效降低VITS的参数和计算复杂度,提出了一种基于iSTFT的波形构造解码器,以取代原VITS中资源消耗较大的基于上采样的解码器。此外,还引入了NanoFlow来共享流块之间的密度估计,以减少先前编码器的参数。此外,为了降低文本编码器的计算复杂度,将缩放点注意替换为线性注意。为了解决简化模型带来的不稳定性,不再使用原始的文本编码器,而是通过文本到PPG模块将语音后验概率(PPG)用作语言特征,然后将其用作编码器的输入。实验表明,AdaVITS模型参数为8.97M,计算复杂度为0.72GFlops,能够在说话人自适应中生成稳定、自然的语音。摘要:Speaker adaptation in text-to-speech synthesis (TTS) is to finetune a pre-trained TTS model to adapt to new target speakers with limited data. While much effort has been conducted towards this task, seldom work has been performed for low computational resource scenarios due to the challenges raised by the requirement of the lightweight model and less computational complexity. In this paper, a tiny VITS-based TTS model, named AdaVITS, for low computing resource speaker adaptation is proposed. To effectively reduce parameters and computational complexity of VITS, an iSTFT-based wave construction decoder is proposed to replace the upsampling-based decoder which is resource-consuming in the original VITS. Besides, NanoFlow is introduced to share the density estimate across flow blocks to reduce the parameters of the prior encoder. Furthermore, to reduce the computational complexity of the textual encoder, scaled-dot attention is replaced with linear attention. To deal with the instability caused by the simplified model, instead of using the original text encoder, phonetic posteriorgram (PPG) is utilized as linguistic feature via a text-to-PPG module, which is then used as input for the encoder. Experiment shows that AdaVITS can generate stable and natural speech in speaker adaptation with 8.97M model parameters and 0.72GFlops computational complexity.
机器翻译,仅供参考