Behavioral Safety

A model may know the answer to a question but refuse to give it, deny that it knows, agree with a user who is wrong, change its answer under a superficial reframing, or behave differently when it infers something about the user or the evaluation setting. We study these failures where capability, training incentives, and deployment context meet.

Many current safety claims are made at the level of behavior, while post-training is increasingly designed to shape what a model says, when it says it, and how it presents its reasoning. We focus on behavioral failures that are empirically tractable, deployment-relevant, and likely to reveal gaps in current alignment training: deception, sycophancy, evaluation awareness, sandbagging, and unfaithful chain-of-thought.

The central questions are practical as well as scientific. When does an unreliable answer reflect missing knowledge, and when does it reflect a learned policy? Which failures are stable properties of the model, and which are induced by prompt framing, user modeling, or the details of the post-training procedure? How much should we trust a behavioral evaluation when the model may be sensitive to whether it appears to be in an evaluation, deployment, or training-like setting?

To study failures that are hard to observe in deployed systems, we build model organisms and evaluation suites in open-weight models with reproducible training procedures. The aim is not to create arbitrary pathological models but to test whether plausible post-training pipelines produce the failures we care about, how they scale with capability and training pressure, and which mitigations survive contact with white-box and black-box evaluation. We also stress-test alignment methods that are increasingly used in practice, such as character training, deliberative alignment, inoculation prompting, safety-focused RLHF and DPO, and on-policy distillation, asking what each changes, what makes it work, and when its apparent benefits disappear.

Current Directions

Distinguishing ignorance from suppression
Methods that isolate model components with a causal role in deception-like behavior, building on the Liars' Bench benchmark, to tell missing knowledge apart from learned misreporting.
Model organisms
Realistic organisms of sycophancy, evaluation awareness, reward-seeking, and user-conditioned behavior, produced by plausible training pipelines rather than by construction.
Stress-testing alignment training
For each widely used method and a near-frontier open model: what behavior changes, what new failure modes appear, and where the benefits stop.

Selected Papers

Newest first. Showing 12 of 29.

All 29 papers in this area