← 返回资源分享
Attention Is All You Need
论文
论文
发布时间2017-06-12
发表NeurIPS 2017 12 · arXiv:1706.03762
作者:Noam Shazeer,Jakob Uszkoreit,Llion Jones,Niki Parmar,Ashish Vaswani,Lukasz Kaiser,Aidan N. Gomez,Illia Polosukhin
详细介绍
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
代码仓库 (573)
tensorflow/tensor2tensor官方TensorFlow
studio-ousia/luke官方PyTorch
PaddlePaddle/PaddleSpeech官方PaddlePaddle
jbdel/OMG_UMONS_submission官方TensorFlow
joongbo/tta官方TensorFlow
ming024/FastSpeech2官方PyTorch
bzhangGo/zero官方TensorFlow
rupakdas18/SemEval-2017-Task-4-A-B-C-using-BERT官方PyTorch
brightmart/bert_customizedTensorFlow
lovedavidsilva/bert_old_versionTensorFlow
