← 返回资源分享
Tacotron: Towards End-to-End Speech Synthesis
论文
论文
发布时间2017-03-29
发表arXiv:1703.10135
作者:Ron J. Weiss,Zhifeng Chen,Yonghui Wu,RJ Skerry-Ryan,Ying Xiao,Yuxuan Wang,Daisy Stanton,Rob Clark,Rif A. Saurous,Samy Bengio,Navdeep Jaitly,Zongheng Yang,Yannis Agiomyrgiannakis,Quoc Le
详细介绍
A text-to-speech synthesis system typically consists of multiple stages, such
as a text analysis frontend, an acoustic model and an audio synthesis module.
Building these components often requires extensive domain expertise and may
contain brittle design choices. In this paper, we present Tacotron, an
end-to-end generative text-to-speech model that synthesizes speech directly
from characters. Given pairs, the model can be trained completely
from scratch with random initialization. We present several key techniques to
make the sequence-to-sequence framework perform well for this challenging task.
Tacotron achieves a 3.82 subjective 5-scale mean opinion score on US English,
outperforming a production parametric system in terms of naturalness. In
addition, since Tacotron generates speech at the frame level, it's
substantially faster than sample-level autoregressive methods.
代码仓库 (31)
PaddlePaddle/PaddleSpeech官方PaddlePaddle
coqui-ai/TTSPyTorch
CorentinJ/Real-Time-Voice-CloningTensorFlow
tigthor/Voice-Cloning-AIPyTorch
anandaswarup/TTSPyTorch
choiHkk/TacotronPyTorch
dipjyoti92/SC-WaveRNNPyTorch
izzajalandoni/tts_modelsPyTorch
cchinchristopherj/Concert-of-Whales
r9y9/tacotron_pytorchPyTorch
