← 返回资源分享
Regularizing and Optimizing LSTM Language Models
论文
论文
发布时间2017-08-07
发表ICLR 2018 1 · arXiv:1708.02182
作者:Richard Socher,Stephen Merity,Nitish Shirish Keskar
详细介绍
Recurrent neural networks (RNNs), such as long short-term memory networks
(LSTMs), serve as a fundamental building block for many sequence learning
tasks, including machine translation, language modeling, and question
answering. In this paper, we consider the specific problem of word-level
language modeling and investigate strategies for regularizing and optimizing
LSTM-based models. We propose the weight-dropped LSTM which uses DropConnect on
hidden-to-hidden weights as a form of recurrent regularization. Further, we
introduce NT-ASGD, a variant of the averaged stochastic gradient method,
wherein the averaging trigger is determined using a non-monotonic condition as
opposed to being tuned by the user. Using these and other regularization
strategies, we achieve state-of-the-art word level perplexities on two data
sets: 57.3 on Penn Treebank and 65.8 on WikiText-2. In exploring the
effectiveness of a neural cache in conjunction with our proposed model, we
achieve an even lower state-of-the-art perplexity of 52.8 on Penn Treebank and
52.0 on WikiText-2.
代码仓库 (49)
Noahs-ARK/groc官方PyTorch
nkcr/overlap-ml官方PyTorch
mnhng/hier-char-emb官方PyTorch
salesforce/awd-lstm-lm官方PyTorch
muellerzr/CodeFest_2019
alexandra-chron/ntua-slp-wassa-iest2018PyTorch
google-research/google-research/tree/master/enas_lmTensorFlow
SachinIchake/KALMPyTorch
vganesh46/awd-lstm-pytorch-implementationPyTorch
chris-tng/semi-supervised-nlpPyTorch
