← 返回资源分享
Compressive Transformers for Long-Range Sequence Modelling
论文
论文
发布时间2019-11-13
发表ICLR 2020 1 · arXiv:1911.05507
作者:Jack W. Rae,Timothy P. Lillicrap,Siddhant M. Jayakumar,Anna Potapenko
详细介绍
We present the Compressive Transformer, an attentive sequence model which compresses past memories for long-range sequence learning. We find the Compressive Transformer obtains state-of-the-art language modelling results in the WikiText-103 and Enwik8 benchmarks, achieving 17.1 ppl and 0.97 bpc respectively. We also find it can model high-frequency speech effectively and can be used as a memory mechanism for RL, demonstrated on an object matching task. To promote the domain of long-range sequence learning, we propose a new open-vocabulary language modelling benchmark derived from books, PG-19.
代码仓库 (6)
ViktorStagge/CompressiveTransformer
labmlai/annotated_deep_learning_paper_implementationsPyTorch
deepmind/pg19
lucidrains/compressive-transformer-pytorchPyTorch
lucidrains/block-recurrent-transformer-pytorchPyTorch
google-deepmind/pg19
