← 返回资源分享
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
论文
论文
发布时间2020-12-31
发表arXiv:2101.00027
作者:Jason Phang,Leo Gao,Stella Biderman,Sid Black,Laurence Golding,Travis Hoppe,Charles Foster,Horace He,Anish Thite,Noa Nabeshima,Shawn Presser,Connor Leahy
详细介绍
Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.
代码仓库 (23)
THUDM/GLM官方PyTorch
EleutherAI/The-Pile官方
suu990901/LLaMA-InfoEntropy-Loss官方JAX
jackbandy/bookcorpus-datasheet
codedotal/gpt-code-clippy
eleutherai/gpt-neoxPyTorch
EleutherAI/gpt-neoTensorFlow
Wikidepia/indonesia_dataset
alrope123/prompt-waywardnessPyTorch
kevinlai219/neo-gpt3TensorFlow
