← 返回资源分享
DiffWave: A Versatile Diffusion Model for Audio Synthesis
论文
论文
发布时间2020-09-21
发表ICLR 2021 1 · arXiv:2009.09761
作者:Wei Ping,Jiaji Huang,Kexin Zhao,Bryan Catanzaro,Zhifeng Kong
详细介绍
In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts the white noise signal into structured waveform through a Markov chain with a constant number of steps at synthesis. It is efficiently trained by optimizing a variant of variational bound on the data likelihood. DiffWave produces high-fidelity audios in different waveform generation tasks, including neural vocoding conditioned on mel spectrogram, class-conditional generation, and unconditional generation. We demonstrate that DiffWave matches a strong WaveNet vocoder in terms of speech quality (MOS: 4.44 versus 4.43), while synthesizing orders of magnitude faster. In particular, it significantly outperforms autoregressive and GAN-based waveform models in the challenging unconditional generation task in terms of audio quality and sample diversity from various automatic and human evaluations.
代码仓库 (11)
neillu23/cdiffuse官方PyTorch
keonlee9420/DiffSingerPyTorch
philsyn/diffwave-unconditionalPyTorch
lmnt-com/diffwavePyTorch
revsic/tf-diffwaveTensorFlow
revsic/jax-variational-diffwaveJAX
neillu23/DiffuSEPyTorch
revsic/torch-diffusion-waveganPyTorch
philsyn/diffwave-vocoderPyTorch
rf5/diffwave-unconditionalPyTorch
