Interpretability
Peeking inside the black box of machine learning algorithms to build robust understandings of what they do and why.
Current Projects
Featured
As models get smarter, humans won't always be able to independently check if a model's claims are true or false. We aim to circumvent this issue by directly eliciting latent knowledge (ELK) inside the model’s activations.
Releases
Featured
A suite of models designed to enable controlled scientific research on transparently trained LLMs
A library implementing the Tuned Lens, along with other tools for extracting, manipulating, and studying the learned representations of transformers across layers.
Publications
Featured

