Training Data

Training data shapes what a model learns. We build and document open corpora so researchers can inspect that foundation and study how changing it changes a model.

The Pile brought together diverse sources of text in an 800GB language-modeling dataset. The Common Pile extends our work on open training data with an 8TB collection of public domain and openly licensed text.

Releasing a corpus also means documenting where its contents came from and how they were selected. This connects dataset construction to research on filtering, licensing, and the effects of data composition.

Selected Papers

Newest first. Showing 12 of 38.

All 38 papers in this area