Privacy and Security

What models remember about their training data, how that can be measured and predicted, and what it means for the people whose data was used.

Memorization sits at the intersection of privacy, copyright, and training dynamics. Using Pythia, we showed that memorization is emergent and partly predictable from smaller models and earlier checkpoints, and later work characterizes it as several distinct phenomena rather than one. This connects directly to our research on data attribution and open-weight safety.

Selected Papers

Newest first.

All 6 papers in this area