← 返回资源分享
Exploring the Limits of Language Modeling
论文
论文
发布时间2016-02-07
发表arXiv:1602.02410
作者:Yonghui Wu,Oriol Vinyals,Mike Schuster,Noam Shazeer,Rafal Jozefowicz
详细介绍
In this work we explore recent advances in Recurrent Neural Networks for
large scale Language Modeling, a task central to language understanding. We
extend current models to deal with two key challenges present in this task:
corpora and vocabulary sizes, and complex, long term structure of language. We
perform an exhaustive study on techniques such as character Convolutional
Neural Networks or Long-Short Term Memory, on the One Billion Word Benchmark.
Our best single model significantly improves state-of-the-art perplexity from
51.3 down to 30.0 (whilst reducing the number of parameters by a factor of 20),
while an ensemble of models sets a new record by improving perplexity from 41.0
down to 23.7. We also release these models for the NLP and ML community to
study and improve upon.
代码仓库 (10)
okuchaiev/f-lm官方TensorFlow
tensorflow/modelsTensorFlow
rdspring1/PyTorch_GBW_LMPyTorch
UnofficialJuliaMirror/DeepMark-deepmarkPyTorch
tensorflow/models/tree/master/research/lm_1bTensorFlow
jmichaelov/does-surprisal-explain-n400PyTorch
UnofficialJuliaMirrorSnapshots/DeepMark-deepmarkPyTorch
rafaljozefowicz/lmTensorFlow
DeepMark/deepmarkPyTorch
dmlc/gluon-nlpMXNet
