← 返回资源分享
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
论文
论文
发布时间2016-11-04
发表arXiv:1611.01462
作者:Richard Socher,Hakan Inan,Khashayar Khosravi
详细介绍
Recurrent neural networks have been very successful at predicting sequences
of words in tasks such as language modeling. However, all such models are based
on the conventional classification framework, where the model is trained
against one-hot targets, and each word is represented both as an input and as
an output in isolation. This causes inefficiencies in learning both in terms of
utilizing all of the information and in terms of the number of parameters
needed to train. We introduce a novel theoretical framework that facilitates
better learning in language modeling, and show that our framework leads to
tying together the input embedding and the output projection matrices, greatly
reducing the number of trainable variables. Our framework leads to state of the
art performance on the Penn Treebank with a variety of network models.
代码仓库 (5)
rdspring1/PyTorch_GBW_LMPyTorch
JianGoForIt/YellowFin_PytorchPyTorch
InnerPeace-Wu/im2p-tensorflowTensorFlow
floydhub/word-language-modelPyTorch
Ravoxsg/Word-level-language-modelingPyTorch
