Open-Weight Safety
Pretraining data filtering
Pretraining data filtering is a promising safety intervention that deserves further investment. It changes what a model learns in the first place, rather than relying only on restrictions added after training.
In Deep Ignorance, we trained 6.9-billion-parameter models on data filtered to exclude targeted biological knowledge used as a proxy in safety evaluations. The filtered models were substantially more resistant to the adversarial fine-tuning we tested than models protected by the post-training safeguards in our comparison. We observed no degradation on unrelated capability evaluations.
This is a reason to invest in pretraining as a site for safety interventions. Further work can test how filtering performs at larger scales, which kinds of knowledge it can reliably target, and how it can complement other methods.