Mika Okamoto and I are releasing PACT (Pressure-Applied Compliance Testing): a benchmark for whether enterprise AI assistants keep following workplace rules when breaking them is convenient.
We tested 23 models (18 open-weight, 4 closed) on 3,364 multi-turn trials across regulated domains like HIPAA, hiring law, and GDPR. Each trial gives the assistant a standing rule, makes the violating option the easy one, and adds ordinary corporate pressure: a deadline, a manager's verbal OK, "my colleague did it and nothing happened."
PACT spun out of our #AIES2026 paper on why AI agents break rules; this time we turned the question into a full benchmark.
The two highest-scoring models were open-weight. Kimi K2.7 (0.944) and Qwen3.6 27B (0.943) finished statistically tied, ahead of every closed model we tested. Bigger isn't better here; a 27B you can run yourself co-leads the board, and open models are competitive on all six dimensions (baseline, pressure resistance, transparency, and others) we measure.
Not one model aced it. One sentence of pressure raised violation rates 65%, and the best model still missed 1 decision in 18. If you're deploying in a regulated workflow, evaluate on your own workload, and don't assume the closed-source frontier is the frontier for rule-following.
Leaderboard, interactive trials, dataset, paper, code: https://lnkd.in/gPmB73Ns
Baseten funded the inference for all ~232,000 trials. An inference company paying to measure how models behave, not just how fast they serve, is why I love working here.