Bergson: An Open Source Library for Data Attribution
Quirke, Jaburi, Johnston, Li, Paulo, Martres, Gupta, Biderman, and Belrose. EMNLP System Demonstrations, 2026
Training data shapes what a model learns. We build and document open corpora so researchers can inspect that foundation and study how changing it changes a model.
The Pile brought together diverse sources of text in an 800GB language-modeling dataset. The Common Pile extends our work on open training data with an 8TB collection of public domain and openly licensed text.
Releasing a corpus also means documenting where its contents came from and how they were selected. This connects dataset construction to research on filtering, licensing, and the effects of data composition.
Newest first. Showing 12 of 38.
Quirke, Jaburi, Johnston, Li, Paulo, Martres, Gupta, Biderman, and Belrose. EMNLP System Demonstrations, 2026
Matlin and 10 others. COLM, 2026
Schaeffer, Kazdan, Abbasi, Liu, Miranda, Ahmed, Puri, Mireshghallah, and Koyejo. Foundations of Deep Generative Models @ ICML, 2026
Ortiz Suarez, Burchell, Arnett, and 94 others. ACL, 2026
Son, Kim, Arnett, Ko, Lee, Kang, Jiang, Yun, Lee, Lee, Kim, Park, Hong, Lee, Yi, Shin, Bok, Shin, Ji, Kim, Jung, Asai, Neubig, Welleck, and 51 others. NeurIPS, 2026
Kowal, Paulo, Jaburi, Tseng, McKinney, Heimersheim, Tucker, Gleave, and Pelrine. arXiv, 2026
Jaburi, Paulo, Shabalin, Quirke, and Belrose. Mech Interp Workshop @ NeurIPS, 2025
Kandpal, Lester, Raffel, Majstorović, Biderman, and 22 others. NeurIPS Datasets and Benchmarks, 2025
Sai Prashanth, Deng, O'Brien, Jyothir S V, Khan, and 7 others. ICLR, 2025
Longpre, Singh, Cherep, Tiwary, Materzynska, and 38 others. ICLR, 2025
Bradshaw, and Colton. ICLR, 2025
Roger, Humane, Kaplan, Gupta, Sun, Adamopoulos, Lim, Anthony, Fennell, and Rish. arXiv, 2025