PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Okamoto and Erol. arXiv, 2026
A model may know the answer to a question but refuse to give it, deny that it knows, agree with a user who is wrong, change its answer under a superficial reframing, or behave differently when it infers something about the user or the evaluation setting. We study these failures where capability, training incentives, and deployment context meet.
Many current safety claims are made at the level of behavior, while post-training is increasingly designed to shape what a model says, when it says it, and how it presents its reasoning. We focus on behavioral failures that are empirically tractable, deployment-relevant, and likely to reveal gaps in current alignment training: deception, sycophancy, evaluation awareness, sandbagging, and unfaithful chain-of-thought.
The central questions are practical as well as scientific. When does an unreliable answer reflect missing knowledge, and when does it reflect a learned policy? Which failures are stable properties of the model, and which are induced by prompt framing, user modeling, or the details of the post-training procedure? How much should we trust a behavioral evaluation when the model may be sensitive to whether it appears to be in an evaluation, deployment, or training-like setting?
To study failures that are hard to observe in deployed systems, we build model organisms and evaluation suites in open-weight models with reproducible training procedures. The aim is not to create arbitrary pathological models but to test whether plausible post-training pipelines produce the failures we care about, how they scale with capability and training pressure, and which mitigations survive contact with white-box and black-box evaluation. We also stress-test alignment methods that are increasingly used in practice, such as character training, deliberative alignment, inoculation prompting, safety-focused RLHF and DPO, and on-policy distillation, asking what each changes, what makes it work, and when its apparent benefits disappear.
Newest first. Showing 12 of 29.
Okamoto and Erol. arXiv, 2026
Weckbecker, Müller, Hagag, and Mulet. Agents in the Wild @ ICML, 2026
Reuel, Ghosh, Chim, and 32 others. ICML, 2026
Zur, Ying, Loftus, Şahin, Yu, Quirke, Rott Shaham, Shapira, Orgad, and Bau. Mech Interp Workshop @ NeurIPS, 2025
Jaburi, Paulo, Shabalin, Quirke, and Belrose. Mech Interp Workshop @ NeurIPS, 2025
Mickel, De-Arteaga, Liu, and Tian. NeurIPS, 2025
Imran and Chatterjee. Workshop on LLM Persona Modeling @ NeurIPS, 2025
Kamachee, Casper, Ding, Yew, Reuel, Biderman, and Hadfield-Menell. arXiv, 2025
Laurito, Belrose, Mallen, Kozaronek, Roger, and 12 others. Journal of Open Source Software, 2025
O'Brien, Majercak, Fernandes, Edgar, Chen, Nori, Carignan, Horvitz, and Poursabzi-Sangdeh. Actionable Interpretability Workshop @ ICML, 2025
François, Péran, Bdeir, Dziri, and Hawkins, and 15 others. Columbia Convening on AI Openness and Safety, 2025
Kolbeinsson, O'Brien, Huang, Gao, Liu, and 6 others. ICLR, 2025