← 返回资源分享
Adversarial Audio Synthesis
论文
论文
发布时间2018-02-12
发表ICLR 2019 5 · arXiv:1802.04208
作者:Chris Donahue,Julian McAuley,Miller Puckette
详细介绍
Audio signals are sampled at high temporal resolutions, and learning to
synthesize audio requires capturing structure across a range of timescales.
Generative adversarial networks (GANs) have seen wide success at generating
images that are both locally and globally coherent, but they have seen little
application to audio generation. In this paper we introduce WaveGAN, a first
attempt at applying GANs to unsupervised synthesis of raw-waveform audio.
WaveGAN is capable of synthesizing one second slices of audio waveforms with
global coherence, suitable for sound effect generation. Our experiments
demonstrate that, without labels, WaveGAN learns to produce intelligible words
when trained on a small-vocabulary speech dataset, and can also synthesize
audio from other domains such as drums, bird vocalizations, and piano. We
compare WaveGAN to a method which applies GANs designed for image generation on
image-like audio feature representations, finding both approaches to be
promising.
代码仓库 (22)
chrisdonahue/wavegan官方TensorFlow
acheketa/cwavegan官方TensorFlow
adrienchaton/BERGANPyTorch
MaxHolmberg96/WaveGANTensorFlow
SilverEngineered/WaveGanTensorFlow
cristiprg/wavegan-forkTensorFlow
MurreyCode/waveganTensorFlow
IBM/MAX-Audio-Sample-GeneratorTensorFlow
mahotani/ADVERSARIAL-AUDIO-SYNTHESIS
paechi/wavegan-asrPyTorch
