← 返回资源分享
GANSynth: Adversarial Neural Audio Synthesis
论文
论文
发布时间2019-02-23
发表ICLR 2019 5 · arXiv:1902.08710
作者:Chris Donahue,Jesse Engel,Kumar Krishna Agrawal,Adam Roberts,Ishaan Gulrajani,Shuo Chen
详细介绍
Efficient audio synthesis is an inherently difficult machine learning task,
as human perception is sensitive to both global structure and fine-scale
waveform coherence. Autoregressive models, such as WaveNet, model local
structure at the expense of global latent structure and slow iterative
sampling, while Generative Adversarial Networks (GANs), have global latent
conditioning and efficient parallel sampling, but struggle to generate
locally-coherent audio waveforms. Herein, we demonstrate that GANs can in fact
generate high-fidelity and locally-coherent audio by modeling log magnitudes
and instantaneous frequencies with sufficient frequency resolution in the
spectral domain. Through extensive empirical investigations on the NSynth
dataset, we demonstrate that GANs are able to outperform strong WaveNet
baselines on automated and human evaluation metrics, and efficiently generate
audio several orders of magnitude faster than their autoregressive
counterparts.
代码仓库 (6)
tensorflow/magentaTensorFlow
Ipsedo/MusicGANPyTorch
lonce/sonyGanForkPyTorch
elsalmi/qiskit
Ipsedo/MusicDiffusionPyTorch
Ipsedo/MusicDiffusionModelPyTorch
