今日论文合集:cs.SD语音8篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Speech foundation models on intelligibility prediction for  hearing-impaired listeners
标题:用于听障人士清晰度预测的语音基础模型
链接:https://arxiv.org/abs/2401.14289
作者:Santiago Cuervo,Ricard Marxer
备注:To be presented in ICASSP 2024
摘要:语音基础模型(SFM)已经在许多语音处理任务上进行了基准测试,通常以最小的自适应实现最先进的性能。然而,SFM范式已显着较少探索感兴趣的应用程序的语音感知社区。在本文中,我们提出了一个这样的应用程序:语音清晰度预测的10个SFM系统的评价。我们专注于清晰度预测挑战2(CPC 2)的非侵入性设置,其中的任务是预测听力受损的听众从噪声录音中正确感知的单词的百分比。我们提出了一个简单的方法,学习一个轻量级的专业预测头上冻结的SFM来解决这个问题。我们的研究结果揭示了统计上显着的差异,在性能跨SFM。我们的方法在CPC 2中获胜,证明了它对语音感知应用的承诺。
摘要:Speech foundation models (SFMs) have been benchmarked on many speech processing tasks, often achieving state-of-the-art performance with minimal adaptation. However, the SFM paradigm has been significantly less explored for applications of interest to the speech perception community. In this paper we present a systematic evaluation of 10 SFMs on one such application: Speech intelligibility prediction. We focus on the non-intrusive setup of the Clarity Prediction Challenge 2 (CPC2), where the task is to predict the percentage of words correctly perceived by hearing-impaired listeners from speech-in-noise recordings. We propose a simple method that learns a lightweight specialized prediction head on top of frozen SFMs to approach the problem. Our results reveal statistically significant differences in performance across SFMs. Our method resulted in the winning submission in the CPC2, demonstrating its promise for speech perception applications.


【2】 TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down  Fusion
标题:TDFNet:一种高效的自顶向下融合音视频语音分离模型
链接:https://arxiv.org/abs/2401.14185
作者:Samuel Pegg,Kai Li,Xiaolin Hu
备注:None
摘要:视听语音分离由于其在语音识别、日记化、场景分析和辅助技术等领域的潜在应用,近年来获得了巨大的关注。设计一个轻量级的视听语音分离网络对于低延迟应用非常重要,但是现有的方法通常需要更高的计算成本和更多的参数来实现更好的分离性能。在本文中,我们提出了一个视听语音分离模型,称为自顶向下融合网络(TDFNet),一个国家的最先进的(SOTA)模型视听语音分离,它建立在TDANet,一个音频语音分离方法的架构。TDANet是TDFNet中听觉和视觉网络的架构基础,提供了一个参数较少的高效模型。在LRS 2 - 2 Mix数据集上,与之前的SOTA方法CTCNet相比,TDFNet在所有性能指标上实现了高达10%的性能提升。值得注意的是,这些结果是使用更少的参数和CTCNet的乘法累加操作(MAC)的28%来实现的。从本质上讲,我们的方法为视听领域内的语音分离挑战提供了一种高效的解决方案,在最佳利用视觉信息方面取得了重大进展。
摘要:Audio-visual speech separation has gained significant traction in recent years due to its potential applications in various fields such as speech recognition, diarization, scene analysis and assistive technologies. Designing a lightweight audio-visual speech separation network is important for low-latency applications, but existing methods often require higher computational costs and more parameters to achieve better separation performance. In this paper, we present an audio-visual speech separation model called Top-Down-Fusion Net (TDFNet), a state-of-the-art (SOTA) model for audio-visual speech separation, which builds upon the architecture of TDANet, an audio-only speech separation method. TDANet serves as the architectural foundation for the auditory and visual networks within TDFNet, offering an efficient model with fewer parameters. On the LRS2-2Mix dataset, TDFNet achieves a performance increase of up to 10\% across all performance metrics compared with the previous SOTA method CTCNet. Remarkably, these results are achieved using fewer parameters and only 28\% of the multiply-accumulate operations (MACs) of CTCNet. In essence, our method presents a highly effective and efficient solution to the challenges of speech separation within the audio-visual domain, making significant strides in harnessing visual information optimally.

【3】 Scaling NVIDIA's multi-speaker multi-lingual TTS systems with voice  cloning to Indic Languages
标题:通过将语音克隆到印度语来扩展NVIDIA的多扬声器多语言TTS系统
链接:https://arxiv.org/abs/2401.13851
作者:Akshit Arora,Rohan Badlani,Sungwon Kim,Rafael Valle,Bryan Catanzaro
备注:Presentation accepted at ICASSP 2024
摘要:在本文中,我们描述了NVIDIA为MMITS-VC(带语音克隆的多扬声器、多语言印度文文语转换系统)2024挑战赛开发的文语转换系统模型。在音轨1和2中,我们利用RAD-MMM通过额外训练5分钟的目标说话人数据来执行Few-Shot TTS。在Track 3中,我们利用P-Flow通过在挑战数据集以及外部数据集上进行训练来执行zero-shot TTS。我们使用HiFi-GAN声码器进行所有提交。RAD-MMM在轨道1和2上表现得很有竞争力,而P-Flow在轨道3上排名第一,平均意见得分(MOS)为4.4,说话人相似性得分(SMOS)为3.62。
摘要:In this paper, we describe the TTS models developed by NVIDIA for the MMITS-VC (Multi-speaker, Multi-lingual Indic TTS with Voice Cloning) 2024 Challenge. In Tracks 1 and 2, we utilize RAD-MMM to perform few-shot TTS by training additionally on 5 minutes of target speaker data. In Track 3, we utilize P-Flow to perform zero-shot TTS by training on the challenge dataset as well as external datasets. We use HiFi-GAN vocoders for all submissions. RAD-MMM performs competitively on Tracks 1 and 2, while P-Flow ranks first on Track 3, with mean opinion score (MOS) 4.4 and speaker similarity score (SMOS) of 3.62.

【4】 VALL-T: Decoder-Only Generative Transducer for Robust and  Decoding-Controllable Text-to-Speech
标题:VALL-T:用于鲁棒和解码可控的文本到语音转换的仅解码器生成转换器
链接:https://arxiv.org/abs/2401.14321
作者:Chenpeng Du,Yiwei Guo,Hankun Wang,Yifan Yang,Zhikang Niu,Shuai Wang,Hui Zhang,Xie Chen,Kai Yu
摘要:最近的TTS模型,如SPEAR-TTS和VALL-E,具有仅解码器Transformer架构,实现了令人印象深刻的自然度,并展示了在语音提示下进行zero-shot自适应的能力。然而,这种仅解码器的TTS模型缺乏单调对齐约束,有时会导致幻觉问题,例如发音错误,单词跳过和难以停止。为了解决这个问题,我们提出了VALL-T,一个生成的转换器模型,它引入了输入音素序列的相对位置嵌入,明确表示单调的生成过程,同时保持解码器的架构,只有Transformer。因此,VALL-T保留了基于自适应的zero-shot自适应能力,并表现出更好的鲁棒性,对幻觉的字错误率相对降低了28.3%。此外,在解码期间VALL-T中对齐的可控性促进了未转录语音提示的使用,即使是未知语言。它还能够通过利用对齐的上下文窗口合成冗长的语音。
摘要:Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hallucination issues such as mispronunciation, word skipping and difficulty in stopping. To address this limitation, we propose VALL-T, a generative Transducer model that introduces shifting relative position embeddings for input phoneme sequence, explicitly indicating the monotonic generation process while maintaining the architecture of decoder-only Transformer. Consequently, VALL-T retains the capability of prompt-based zero-shot adaptation and demonstrates better robustness against hallucinations with a relative reduction of 28.3\% in the word error rate. Furthermore, the controllability of alignment in VALL-T during decoding facilitates the use of untranscribed speech prompts, even in unknown languages. It also enables the synthesis of lengthy speech by utilizing an aligned context window.


【5】 Improving Design of Input Condition Invariant Speech Enhancement
标题:输入条件不变语音增强的改进设计
链接:https://arxiv.org/abs/2401.14271
作者:Wangyou Zhang,Jee-weon Jung,Shinji Watanabe,Yanmin Qian
备注:Accepted by ICASSP 2024, 5 pages, 2 figures, 3 tables
摘要:建立一个通用的语音增强(SE)系统,可以处理任意输入是一个需求,但未充分探索的研究课题。朝着这个最终目标,一个方向是建立一个单一的模型,处理不同的音频持续时间,采样频率,麦克风的变化,在嘈杂和混响的情况下,我们在这里定义为“输入条件不变的SE”。最近提出的这种模型表现出良好的性能,然而,其多通道性能在实际条件下严重下降。在本文中,我们提出了新的架构,以改善输入条件不变的SE模型,使模拟条件下的性能仍然具有竞争力,而真正的条件退化大大减轻。为此,我们重新设计了组成这样一个系统的关键组件。首先,我们确定,通道建模模块的泛化看不见的情况下,可以是次优的,并重新设计这个模块。我们进一步引入两阶段培训策略,以提高培训效率。其次,我们提出了两种新的双路径时频块,表现出更好的性能与更少的参数和计算成本相比,现有的方法。结合所有建议,在各种公共数据集上的实验验证了所提出的模型的有效性,在真实条件下的性能显着提高。包含完整模型详细信息的配方发布于https://github.com/espnet/espnet。
摘要:Building a single universal speech enhancement (SE) system that can handle arbitrary input is a demanded but underexplored research topic. Towards this ultimate goal, one direction is to build a single model that handles diverse audio duration, sampling frequencies, and microphone variations in noisy and reverberant scenarios, which we define here as "input condition invariant SE". Such a model was recently proposed showing promising performance; however, its multi-channel performance degraded severely in real conditions. In this paper we propose novel architectures to improve the input condition invariant SE model so that performance in simulated conditions remains competitive while real condition degradation is much mitigated. For this purpose, we redesign the key components that comprise such a system. First, we identify that the channel-modeling module's generalization to unseen scenarios can be sub-optimal and redesign this module. We further introduce a two-stage training strategy to enhance training efficiency. Second, we propose two novel dual-path time-frequency blocks, demonstrating superior performance with fewer parameters and computational costs compared to the existing method. All proposals combined, experiments on various public datasets validate the efficacy of the proposed model, with significantly improved performance on real conditions. Recipe with full model details is released at https://github.com/espnet/espnet.

【6】 Combined Generative and Predictive Modeling for Speech Super-resolution
标题:语音超分辨率的产生式和预测式联合建模
链接:https://arxiv.org/abs/2401.14269
作者:Heming Wang,Eric W. Healy,DeLiang Wang
摘要:语音超分辨率(SR)是从低分辨率输入中恢复高分辨率语音的任务。现有的模型采用模拟数据和受约束的实验设置,这限制了对真实世界SR的泛化。已知预测模型在固定的实验设置中表现良好,但在不利条件下可能会引入伪影。另一方面,生成模型学习目标数据的分布,并且在看不见的条件下具有更好的性能。在这项研究中,我们提出了一种新的两阶段方法,结合了预测和生成模型的优势。具体来说,我们采用基于扩散的模型,其条件是预测模型的输出。我们的实验表明,该模型显着优于单阶段同行和现有的基准SR数据集上的强基线。此外,我们在扩散过程的推理过程中引入了一种重绘技术,使所提出的模型即使在不匹配的条件下也能重新生成高频分量。另一个贡献是收集和评估真实的SR录音,使用相同的麦克风在不同的本地采样率。我们让这个数据集可以免费访问,以加速实现真实世界的语音超分辨率。
摘要:Speech super-resolution (SR) is the task that restores high-resolution speech from low-resolution input. Existing models employ simulated data and constrained experimental settings, which limit generalization to real-world SR. Predictive models are known to perform well in fixed experimental settings, but can introduce artifacts in adverse conditions. On the other hand, generative models learn the distribution of target data and have a better capacity to perform well on unseen conditions. In this study, we propose a novel two-stage approach that combines the strengths of predictive and generative models. Specifically, we employ a diffusion-based model that is conditioned on the output of a predictive model. Our experiments demonstrate that the model significantly outperforms single-stage counterparts and existing strong baselines on benchmark SR datasets. Furthermore, we introduce a repainting technique during the inference of the diffusion process, enabling the proposed model to regenerate high-frequency components even in mismatched conditions. An additional contribution is the collection of and evaluation on real SR recordings, using the same microphone at different native sampling rates. We make this dataset freely accessible, to accelerate progress towards real-world speech super-resolution.

【7】 Intelli-Z: Toward Intelligible Zero-Shot TTS
标题:Inteli-Z:迈向可理解的Zero-ShotTTS
链接:https://arxiv.org/abs/2401.13921
作者:Sunghee Jung,Won Jang,Jaesam Yoon,Bongwan Kim
摘要:尽管最近的许多研究已经提出了使用大规模真实世界数据的zero-shot TTS的新框架,但关注zero-shot TTS可懂度的研究相对较少。Zero-shot TTS需要额外的努力来确保清晰的发音和语音质量,这是由于其内在要求在推理阶段用新的核心参数(扬声器嵌入或声学提示)替换核心参数。在这项研究中,我们提出了一个专注于可懂度的zero-shot TTS模型,我们称之为Intelli-Z。Intelli-Z通过使用多说话者TTS作为其老师来学习说话者嵌入,并使用循环一致性损失进行训练,以包括用于训练的不匹配的文本-语音对。此外,它有选择地聚集沿时间维度的说话人嵌入,以尽量减少在推理阶段的参考语音的文本内容的干扰。我们证实了所提出的方法与消融研究的有效性。当应用前两种方法时,对于未被看见的说话者,平均意见得分(MOS)增加了9%,并且当应用选择性时间聚合时,它进一步提高了16%。
摘要:Although numerous recent studies have suggested new frame- works for zero-shot TTS using large-scale, real-world data, studies that focus on the intelligibility of zero-shot TTS are relatively scarce. Zero-shot TTS demands additional efforts to ensure clear pronunciation and speech quality due to its inherent requirement of replacing a core parameter (speaker embedding or acoustic prompt) with a new one at the in- ference stage. In this study, we propose a zero-shot TTS model focused on intelligibility, which we refer to as Intelli- Z. Intelli-Z learns speaker embeddings by using multi-speaker TTS as its teacher and is trained with a cycle-consistency loss to include mismatched text-speech pairs for training. Addi- tionally, it selectively aggregates speaker embeddings along the temporal dimension to minimize the interference of the text content of reference speech at the inference stage. We substantiate the effectiveness of the proposed methods with an ablation study. The Mean Opinion Score (MOS) increases by 9% for unseen speakers when the first two methods are ap- plied, and it further improves by 16% when selective temporal aggregation is applied.

【8】 Bayesian adaptive learning to latent variables via Variational Bayes and  Maximum a Posteriori
标题:基于变分贝叶斯和最大后验概率的潜在变量贝叶斯自适应学习
链接:https://arxiv.org/abs/2401.13766
作者:Hu Hu,Sabato Marco Siniscalchi,Chin-Hui Lee
备注:ASRU2023 Bayesian Symposium. arXiv admin note: text overlap with arXiv:2110.08598
摘要:在这项工作中,我们的目标是通过专注于估计深度神经网络(DNN)模型中的潜在变量来建立贝叶斯自适应学习框架。潜变量实际上编码了可传递的分布信息和结构关系。因此,源潜变量的分布(先验)可以与从目标数据(可能性)学习的知识相结合,以产生目标潜变量的分布(后验),其目标是解决训练和测试条件之间的声学失配。先验知识转移是通过变分贝叶斯(VB)。此外,我们还研究了基于最大后验概率(MAP)的贝叶斯自适应。声学场景分类中的设备自适应实验结果表明,我们提出的方法可以获得很好的改善目标设备,并始终优于其他切割边缘算法。
摘要:In this work, we aim to establish a Bayesian adaptive learning framework by focusing on estimating latent variables in deep neural network (DNN) models. Latent variables indeed encode both transferable distributional information and structural relationships. Thus the distributions of the source latent variables (prior) can be combined with the knowledge learned from the target data (likelihood) to yield the distributions of the target latent variables (posterior) with the goal of addressing acoustic mismatches between training and testing conditions. The prior knowledge transfer is accomplished through Variational Bayes (VB). In addition, we also investigate Maximum a Posteriori (MAP) based Bayesian adaptation. Experimental results on device adaptation in acoustic scene classification show that our proposed approaches can obtain good improvements on target devices, and consistently outperforms other cut-edging algorithms.


eess.AS音频处理
【1】 VALL-T: Decoder-Only Generative Transducer for Robust and  Decoding-Controllable Text-to-Speech
标题:VALL-T:用于稳健和解码可控的文本到语音转换的仅解码器生成式转换器
链接:https://arxiv.org/abs/2401.14321
作者:Chenpeng Du,Yiwei Guo,Hankun Wang,Yifan Yang,Zhikang Niu,Shuai Wang,Hui Zhang,Xie Chen,Kai Yu
摘要:最近的TTS模型,如SPEAR-TTS和VALL-E,具有仅解码器Transformer架构,实现了令人印象深刻的自然度,并展示了在语音提示下进行zero-shot自适应的能力。然而,这种仅解码器的TTS模型缺乏单调对齐约束,有时会导致幻觉问题,例如发音错误,单词跳过和难以停止。为了解决这个问题,我们提出了VALL-T,一个生成的转换器模型,它引入了输入音素序列的相对位置嵌入,明确表示单调的生成过程,同时保持解码器的架构,只有Transformer。因此,VALL-T保留了基于自适应的zero-shot自适应能力,并表现出更好的鲁棒性,对幻觉的字错误率相对降低了28.3%。此外,在解码期间VALL-T中对齐的可控性促进了未转录语音提示的使用,即使是未知语言。它还能够通过利用对齐的上下文窗口合成冗长的语音。
摘要:Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hallucination issues such as mispronunciation, word skipping and difficulty in stopping. To address this limitation, we propose VALL-T, a generative Transducer model that introduces shifting relative position embeddings for input phoneme sequence, explicitly indicating the monotonic generation process while maintaining the architecture of decoder-only Transformer. Consequently, VALL-T retains the capability of prompt-based zero-shot adaptation and demonstrates better robustness against hallucinations with a relative reduction of 28.3\% in the word error rate. Furthermore, the controllability of alignment in VALL-T during decoding facilitates the use of untranscribed speech prompts, even in unknown languages. It also enables the synthesis of lengthy speech by utilizing an aligned context window.


【2】 Improving Design of Input Condition Invariant Speech Enhancement
标题:输入条件不变语音增强的改进设计
链接:https://arxiv.org/abs/2401.14271
作者:Wangyou Zhang,Jee-weon Jung,Shinji Watanabe,Yanmin Qian
备注:Accepted by ICASSP 2024, 5 pages, 2 figures, 3 tables
摘要:建立一个通用的语音增强(SE)系统,可以处理任意输入是一个需求,但未充分探索的研究课题。朝着这个最终目标,一个方向是建立一个单一的模型,处理不同的音频持续时间,采样频率,麦克风的变化,在嘈杂和混响的情况下,我们在这里定义为“输入条件不变的SE”。最近提出的这种模型表现出良好的性能,然而,其多通道性能在实际条件下严重下降。在本文中,我们提出了新的架构,以改善输入条件不变的SE模型,使模拟条件下的性能仍然具有竞争力,而真正的条件退化大大减轻。为此,我们重新设计了组成这样一个系统的关键组件。首先,我们确定,通道建模模块的泛化看不见的情况下,可以是次优的,并重新设计这个模块。我们进一步引入两阶段培训策略,以提高培训效率。其次,我们提出了两种新的双路径时频块,表现出更好的性能与更少的参数和计算成本相比,现有的方法。结合所有建议,在各种公共数据集上的实验验证了所提出的模型的有效性,在真实条件下的性能显着提高。包含完整模型详细信息的配方发布于https://github.com/espnet/espnet。
摘要:Building a single universal speech enhancement (SE) system that can handle arbitrary input is a demanded but underexplored research topic. Towards this ultimate goal, one direction is to build a single model that handles diverse audio duration, sampling frequencies, and microphone variations in noisy and reverberant scenarios, which we define here as "input condition invariant SE". Such a model was recently proposed showing promising performance; however, its multi-channel performance degraded severely in real conditions. In this paper we propose novel architectures to improve the input condition invariant SE model so that performance in simulated conditions remains competitive while real condition degradation is much mitigated. For this purpose, we redesign the key components that comprise such a system. First, we identify that the channel-modeling module's generalization to unseen scenarios can be sub-optimal and redesign this module. We further introduce a two-stage training strategy to enhance training efficiency. Second, we propose two novel dual-path time-frequency blocks, demonstrating superior performance with fewer parameters and computational costs compared to the existing method. All proposals combined, experiments on various public datasets validate the efficacy of the proposed model, with significantly improved performance on real conditions. Recipe with full model details is released at https://github.com/espnet/espnet.

【3】 Combined Generative and Predictive Modeling for Speech Super-resolution
标题:语音超分辨率的产生式和预测式联合建模
链接:https://arxiv.org/abs/2401.14269
作者:Heming Wang,Eric W. Healy,DeLiang Wang
摘要:语音超分辨率(SR)是从低分辨率输入中恢复高分辨率语音的任务。现有的模型采用模拟数据和受约束的实验设置,这限制了对真实世界SR的泛化。已知预测模型在固定的实验设置中表现良好,但在不利条件下可能会引入伪影。另一方面,生成模型学习目标数据的分布,并且在看不见的条件下具有更好的性能。在这项研究中,我们提出了一种新的两阶段方法,结合了预测和生成模型的优势。具体来说,我们采用基于扩散的模型,其条件是预测模型的输出。我们的实验表明,该模型显着优于单阶段同行和现有的基准SR数据集上的强基线。此外,我们在扩散过程的推理过程中引入了一种重绘技术,使所提出的模型即使在不匹配的条件下也能重新生成高频分量。另一个贡献是收集和评估真实的SR录音,使用相同的麦克风在不同的本地采样率。我们让这个数据集可以免费访问,以加速实现真实世界的语音超分辨率。
摘要:Speech super-resolution (SR) is the task that restores high-resolution speech from low-resolution input. Existing models employ simulated data and constrained experimental settings, which limit generalization to real-world SR. Predictive models are known to perform well in fixed experimental settings, but can introduce artifacts in adverse conditions. On the other hand, generative models learn the distribution of target data and have a better capacity to perform well on unseen conditions. In this study, we propose a novel two-stage approach that combines the strengths of predictive and generative models. Specifically, we employ a diffusion-based model that is conditioned on the output of a predictive model. Our experiments demonstrate that the model significantly outperforms single-stage counterparts and existing strong baselines on benchmark SR datasets. Furthermore, we introduce a repainting technique during the inference of the diffusion process, enabling the proposed model to regenerate high-frequency components even in mismatched conditions. An additional contribution is the collection of and evaluation on real SR recordings, using the same microphone at different native sampling rates. We make this dataset freely accessible, to accelerate progress towards real-world speech super-resolution.


【4】 Intelli-Z: Toward Intelligible Zero-Shot TTS
标题:Inteli-Z:迈向可理解的Zero-ShotTTS
链接:https://arxiv.org/abs/2401.13921
作者:Sunghee Jung,Won Jang,Jaesam Yoon,Bongwan Kim
摘要:尽管最近的许多研究已经提出了使用大规模真实世界数据的zero-shot TTS的新框架,但关注zero-shot TTS可懂度的研究相对较少。Zero-shot TTS需要额外的努力来确保清晰的发音和语音质量,这是由于其内在要求在推理阶段用新的核心参数(扬声器嵌入或声学提示)替换核心参数。在这项研究中,我们提出了一个专注于可懂度的zero-shot TTS模型,我们称之为Intelli-Z。Intelli-Z通过使用多说话者TTS作为其老师来学习说话者嵌入,并使用循环一致性损失进行训练,以包括用于训练的不匹配的文本-语音对。此外,它有选择地聚集沿时间维度的说话人嵌入,以尽量减少在推理阶段的参考语音的文本内容的干扰。我们证实了所提出的方法与消融研究的有效性。当应用前两种方法时,对于未被看见的说话者,平均意见得分(MOS)增加了9%,并且当应用选择性时间聚合时,它进一步提高了16%。
摘要:Although numerous recent studies have suggested new frame- works for zero-shot TTS using large-scale, real-world data, studies that focus on the intelligibility of zero-shot TTS are relatively scarce. Zero-shot TTS demands additional efforts to ensure clear pronunciation and speech quality due to its inherent requirement of replacing a core parameter (speaker embedding or acoustic prompt) with a new one at the in- ference stage. In this study, we propose a zero-shot TTS model focused on intelligibility, which we refer to as Intelli- Z. Intelli-Z learns speaker embeddings by using multi-speaker TTS as its teacher and is trained with a cycle-consistency loss to include mismatched text-speech pairs for training. Addi- tionally, it selectively aggregates speaker embeddings along the temporal dimension to minimize the interference of the text content of reference speech at the inference stage. We substantiate the effectiveness of the proposed methods with an ablation study. The Mean Opinion Score (MOS) increases by 9% for unseen speakers when the first two methods are ap- plied, and it further improves by 16% when selective temporal aggregation is applied.


【5】 Bayesian adaptive learning to latent variables via Variational Bayes and  Maximum a Posteriori
标题:基于变分贝叶斯和最大后验概率的潜在变量贝叶斯自适应学习
链接:https://arxiv.org/abs/2401.13766
作者:Hu Hu,Sabato Marco Siniscalchi,Chin-Hui Lee
备注:ASRU2023 Bayesian Symposium. arXiv admin note: text overlap with arXiv:2110.08598
摘要:在这项工作中,我们的目标是通过专注于估计深度神经网络(DNN)模型中的潜在变量来建立贝叶斯自适应学习框架。潜变量实际上编码了可传递的分布信息和结构关系。因此,源潜变量的分布(先验)可以与从目标数据(可能性)学习的知识相结合,以产生目标潜变量的分布(后验),其目标是解决训练和测试条件之间的声学失配。先验知识转移是通过变分贝叶斯(VB)。此外,我们还研究了基于最大后验概率(MAP)的贝叶斯自适应。声学场景分类中的设备自适应实验结果表明,我们提出的方法可以获得很好的改善目标设备,并始终优于其他切割边缘算法。
摘要:In this work, we aim to establish a Bayesian adaptive learning framework by focusing on estimating latent variables in deep neural network (DNN) models. Latent variables indeed encode both transferable distributional information and structural relationships. Thus the distributions of the source latent variables (prior) can be combined with the knowledge learned from the target data (likelihood) to yield the distributions of the target latent variables (posterior) with the goal of addressing acoustic mismatches between training and testing conditions. The prior knowledge transfer is accomplished through Variational Bayes (VB). In addition, we also investigate Maximum a Posteriori (MAP) based Bayesian adaptation. Experimental results on device adaptation in acoustic scene classification show that our proposed approaches can obtain good improvements on target devices, and consistently outperforms other cut-edging algorithms.


【6】 Speech foundation models on intelligibility prediction for  hearing-impaired listeners
标题:用于听障者可懂度预测的语音基础模型
链接:https://arxiv.org/abs/2401.14289
作者:Santiago Cuervo,Ricard Marxer
备注:To be presented in ICASSP 2024
摘要:语音基础模型(SFM)已经在许多语音处理任务上进行了基准测试,通常以最小的自适应实现最先进的性能。然而,SFM范式已显着较少探索感兴趣的应用程序的语音感知社区。在本文中,我们提出了一个这样的应用程序:语音清晰度预测的10个SFM系统的评价。我们专注于清晰度预测挑战2(CPC 2)的非侵入性设置,其中的任务是预测听力受损的听众从噪声录音中正确感知的单词的百分比。我们提出了一个简单的方法,学习一个轻量级的专业预测头上冻结的SFM来解决这个问题。我们的研究结果揭示了统计上显着的差异,在性能跨SFM。我们的方法在CPC 2中获胜,证明了它对语音感知应用的承诺。
摘要:Speech foundation models (SFMs) have been benchmarked on many speech processing tasks, often achieving state-of-the-art performance with minimal adaptation. However, the SFM paradigm has been significantly less explored for applications of interest to the speech perception community. In this paper we present a systematic evaluation of 10 SFMs on one such application: Speech intelligibility prediction. We focus on the non-intrusive setup of the Clarity Prediction Challenge 2 (CPC2), where the task is to predict the percentage of words correctly perceived by hearing-impaired listeners from speech-in-noise recordings. We propose a simple method that learns a lightweight specialized prediction head on top of frozen SFMs to approach the problem. Our results reveal statistically significant differences in performance across SFMs. Our method resulted in the winning submission in the CPC2, demonstrating its promise for speech perception applications.


【7】 TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down  Fusion
标题:TDFNet:一种自顶向下融合的高效视听语音分离模型
链接:https://arxiv.org/abs/2401.14185
作者:Samuel Pegg,Kai Li,Xiaolin Hu
备注:None
摘要:视听语音分离由于其在语音识别、日记化、场景分析和辅助技术等领域的潜在应用,近年来获得了巨大的关注。设计一个轻量级的视听语音分离网络对于低延迟应用非常重要,但是现有的方法通常需要更高的计算成本和更多的参数来实现更好的分离性能。在本文中,我们提出了一个视听语音分离模型,称为自顶向下融合网络(TDFNet),一个国家的最先进的(SOTA)模型视听语音分离,它建立在TDANet,一个音频语音分离方法的架构。TDANet是TDFNet中听觉和视觉网络的架构基础,提供了一个参数较少的高效模型。在LRS 2 - 2 Mix数据集上,与之前的SOTA方法CTCNet相比,TDFNet在所有性能指标上实现了高达10%的性能提升。值得注意的是,这些结果是使用更少的参数和CTCNet的乘法累加操作(MAC)的28%来实现的。从本质上讲,我们的方法为视听领域内的语音分离挑战提供了一种高效的解决方案,在最佳利用视觉信息方面取得了重大进展。
摘要:Audio-visual speech separation has gained significant traction in recent years due to its potential applications in various fields such as speech recognition, diarization, scene analysis and assistive technologies. Designing a lightweight audio-visual speech separation network is important for low-latency applications, but existing methods often require higher computational costs and more parameters to achieve better separation performance. In this paper, we present an audio-visual speech separation model called Top-Down-Fusion Net (TDFNet), a state-of-the-art (SOTA) model for audio-visual speech separation, which builds upon the architecture of TDANet, an audio-only speech separation method. TDANet serves as the architectural foundation for the auditory and visual networks within TDFNet, offering an efficient model with fewer parameters. On the LRS2-2Mix dataset, TDFNet achieves a performance increase of up to 10\% across all performance metrics compared with the previous SOTA method CTCNet. Remarkably, these results are achieved using fewer parameters and only 28\% of the multiply-accumulate operations (MACs) of CTCNet. In essence, our method presents a highly effective and efficient solution to the challenges of speech separation within the audio-visual domain, making significant strides in harnessing visual information optimally.

【8】 Scaling NVIDIA's multi-speaker multi-lingual TTS systems with voice  cloning to Indic Languages
标题:将NVIDIA的多扬声器多语言TTS系统与印度语语音克隆相结合
链接:https://arxiv.org/abs/2401.13851
作者:Akshit Arora,Rohan Badlani,Sungwon Kim,Rafael Valle,Bryan Catanzaro
备注:Presentation accepted at ICASSP 2024
摘要:在本文中,我们描述了NVIDIA为MMITS-VC(带语音克隆的多扬声器、多语言印度文文语转换系统)2024挑战赛开发的文语转换系统模型。在音轨1和2中,我们利用RAD-MMM通过额外训练5分钟的目标说话人数据来执行Few-Shot TTS。在Track 3中,我们利用P-Flow通过在挑战数据集以及外部数据集上进行训练来执行zero-shot TTS。我们使用HiFi-GAN声码器进行所有提交。RAD-MMM在轨道1和2上表现得很有竞争力,而P-Flow在轨道3上排名第一,平均意见得分(MOS)为4.4,说话人相似性得分(SMOS)为3.62。
摘要:In this paper, we describe the TTS models developed by NVIDIA for the MMITS-VC (Multi-speaker, Multi-lingual Indic TTS with Voice Cloning) 2024 Challenge. In Tracks 1 and 2, we utilize RAD-MMM to perform few-shot TTS by training additionally on 5 minutes of target speaker data. In Track 3, we utilize P-Flow to perform zero-shot TTS by training on the challenge dataset as well as external datasets. We use HiFi-GAN vocoders for all submissions. RAD-MMM performs competitively on Tracks 1 and 2, while P-Flow ranks first on Track 3, with mean opinion score (MOS) 4.4 and speaker similarity score (SMOS) of 3.62.
机器翻译由腾讯交互翻译提供,仅供参考