← 返回资源分享
Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
论文
论文
发布时间2018-03-23
发表ICML 2018 7 · arXiv:1803.09017
作者:Ye Jia,Yu Zhang,Fei Ren,RJ Skerry-Ryan,Eric Battenberg,Ying Xiao,Yuxuan Wang,Daisy Stanton,Joel Shor,Rif A. Saurous
详细介绍
In this work, we propose "global style tokens" (GSTs), a bank of embeddings
that are jointly trained within Tacotron, a state-of-the-art end-to-end speech
synthesis system. The embeddings are trained with no explicit labels, yet learn
to model a large range of acoustic expressiveness. GSTs lead to a rich set of
significant results. The soft interpretable "labels" they generate can be used
to control synthesis in novel ways, such as varying speed and speaking style -
independently of the text content. They can also be used for style transfer,
replicating the speaking style of a single audio clip across an entire
long-form text corpus. When trained on noisy, unlabeled found data, GSTs learn
to factorize noise and speaker identity, providing a path towards highly
scalable but robust speech synthesis.
代码仓库 (11)
PaddlePaddle/PaddleSpeech官方PaddlePaddle
cnlinxi/style-token_tacotron2TensorFlow
jinhan/tacotron2-gstPyTorch
acetylSv/GST-tacotronTensorFlow
hash2430/pitchtronPyTorch
keonlee9420/Cross-Speaker-Emotion-TransferPyTorch
syang1993/gst-tacotronTensorFlow
foamliu/GST-Tacotron-v2PyTorch
KinglittleQ/GST-TacotronPyTorch
CODEJIN/GST_TacotronTensorFlow
